AI Engineer

Computer-Use 2.0: Agents Just Got Multi-Cursor — Francesco Bonacci, Cua

1849 summary words 8 min summary Watch video

Start with the signal

8 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Computer-use agents are moving from foreground screenshot-and-click loops toward background, cross-platform execution supported by verifiable GUI benchmarks and infrastructure that keeps expensive training GPUs continuously utilized.
  • Why it matters: The talk provides concrete architecture patterns and performance data for three hard parts of agent operations: reliable desktop control, trustworthy evaluation, and economical reinforcement-learning infrastructure.
  • Best use: Use it as a reference architecture for evaluating background computer-use capabilities and for designing desktop-agent test harnesses, world-model probes, and warm sandbox pools.

Executive Summary

Cua frames “computer-use 2.0” as an agent operating applications without taking over the user's visible screen. Its driver observes both accessibility-tree state and screenshots, attempts background execution through accessibility interfaces first, and falls back to pixel-level background clicks when necessary. The intended abstraction is a cross-platform desktop-control layer spanning macOS, Windows, and Linux despite their differing behavior.

The evaluation layer, KuaBench, packages each task as a setup function, a GUI-action oracle, and an evaluator that inspects the resulting environment. Cua reports more than 130 verifiable tasks across 42 environments and five platforms. Its KiCAD benchmark shows how weak current agents remain on professional workflows: the leading agent fully passed only six of 25 tasks, all involving edits to existing schematics, while blank-schematic tasks had a 0% success rate.

The strongest performance result is that replacing an agent's built-in computer tool with KuaDriver increased pass rate on a 4K basic benchmark from roughly 62% to 80% while reducing token usage by 34%. Cua attributes this primarily to focusing the model on the relevant application window instead of the entire desktop. The team also tests benchmark environments for reward hacking and can fork recorded trajectories to measure whether a model correctly predicts future reward, hidden state, or observations.

For reinforcement-learning workloads, Cua recommends demand-scaled pools of prewarmed sandboxes. This moves environment startup latency away from expensive GPU workers, whose utilization would otherwise fall while large desktop images reset or launch. The talk is highly relevant, although the title's “multi-cursor” framing is not explicitly demonstrated or benchmarked as concurrent multi-agent control.

Key Takeaways

  • Claim: Background computer use requires a layered control strategy rather than relying exclusively on screenshots or accessibility APIs. | Evidence: The driver retrieves both an accessibility-tree representation and a screenshot through a Windows-state call, tries background execution through the accessibility tree, and falls back to a pixel-level background click when that fails. The team says it supports macOS, Windows, and Linux. | Implication: Ken should treat accessibility, vision, and pixel actuation as complementary execution modes behind one adapter, with platform-specific fallbacks rather than a universal GUI-control assumption. | Caveat: Background behavior differs by operating system, and the macOS implementation relies in part on undocumented APIs in Apple's frameworks, creating compatibility and maintenance risk.
  • Claim: A useful GUI-agent benchmark must make environment setup, a valid action trajectory, and end-state verification reproducible. | Evidence: KuaBench defines tasks with three components: a setup function, an oracle expressed as GUI actions, and an evaluator that probes the environment. Its catalog reportedly contains more than 130 verifiable tasks across 42 environments and five platforms, reproducible through a CLI. | Implication: Agent evaluations should verify state changes inside the environment rather than score only visual trajectories or model self-reports. | Caveat: The talk describes the benchmark's structure and scale but does not provide external validation of its coverage or compare it systematically with other desktop benchmarks.
  • Claim: Current computer-use agents perform substantially better when modifying an existing artifact than when constructing a professional artifact from scratch. | Evidence: On KuaBench KiCAD, developed with Snorkel AI, the top tested agent fully passed six of 25 electrical-engineering tasks. Every successful task involved editing an existing schematic; success fell to 0% when starting from a blank schematic, and no tested model exceeded 30% reward. | Implication: Near-term production workflows should provide templates, partially completed artifacts, or constrained edit tasks instead of expecting agents to originate complex professional work from an empty state. | Caveat: The transcript does not name the tested models or describe the 25 tasks in enough detail to generalize the exact percentages beyond this benchmark.
  • Claim: Reducing the agent's visual and interaction scope can improve both task success and token efficiency. | Evidence: On the KuaBench basic dataset at 4K resolution, switching from an agent's built-in computer tool to KuaDriver reportedly raised pass rate from about 62% to 80% while using 34% fewer tokens. The stated reason is that KuaDriver focuses on one window instead of the full desktop. | Implication: Before changing models, Ken should test whether tighter observation boundaries and application-level targeting yield a cheaper, larger reliability gain. | Caveat: Only one benchmark configuration is cited, and the talk does not isolate how much of the gain comes from window focus versus other implementation differences.
  • Claim: Benchmark trustworthiness requires adversarial testing of the evaluator itself, not merely repeated agent runs. | Evidence: Before admitting a task to the dataset, Cua runs a matrix of agents against it to attempt reward hacking and environment breakage, then compiles the findings into a CodeRabbit-style review. Only tasks surviving that pipeline are accepted. | Implication: Ken should treat eval construction as a security problem: test for shortcuts, unintended state mutations, and evaluator exploits before using scores for routing or deployment decisions. | Caveat: The presentation does not quantify how many candidate tasks fail this process or whether human reviewers independently audit accepted tasks.
  • Claim: Recorded desktop trajectories can be converted into tests of an agent's world model, not just its ability to click correctly. | Evidence: Cua says each recorded run can be forked at any point to restore the computer state. A model can then predict reward, internal state, or another observation, and that prediction can be compared with the actual forked outcome. | Implication: Trajectory forking could help distinguish agents that understand application state from agents that merely imitate action patterns, which is useful for model routing and failure diagnosis. | Caveat: No experimental results are provided showing that these prediction scores correlate with safer behavior or better completion rates.
  • Claim: Demand-scaled pools of warm desktop sandboxes can reduce reinforcement-learning cost by preventing GPUs from waiting for environments to start or reset. | Evidence: Cua describes an autoscaler that estimates current sandbox demand from GPU workers and grows a warm pool accordingly. The speaker notes that sandbox infrastructure may be two to four times cheaper than GPUs and that some desktop environments can be roughly 40 GB, making startup optimization alone insufficient. | Implication: For computer-use RL, capacity planning should optimize total accelerator utilization rather than minimize sandbox count; modestly redundant CPU or VM capacity may be economically preferable to idle GPUs. | Caveat: The talk provides no measured utilization gain, dollar savings, pool hit rate, or end-to-end training benchmark.

Detailed Brief

Cross-platform reliability and release engineering

  • Claims: Cua treats operating-system variance as infrastructure the driver should absorb so the calling agent does not need separate interaction logic for every platform.; The project is designed to let external agents connect to the local operating system after the driver is installed.
  • Evidence: The team maintains approximately eight application harnesses to check that new releases do not break existing behavior.; Named early adopters or contributors include Clicky, Hermes, Quencoder, H Company, and Droid Factory.
  • Caveats: The transcript does not describe permission boundaries, credential isolation, user confirmation, rollback, or defenses against destructive background actions.; The product naming is inconsistent in the transcript—Quadriver, CodeDriver, and KuaDriver appear to refer to the driver project—so the exact branding should be verified against the repository.
  • Implications: The ongoing maintenance burden is likely to sit in application and operating-system compatibility rather than in the agent-facing API.; A serious deployment evaluation should include upgrade regression tests and permission controls, not just task-completion benchmarks.

Mobile background-control boundary

  • Claims: The speakers consider Android more permissive than iOS for background workloads, particularly when containers or platform activity frameworks can be used.; Their Android direction appears closer to background tool use than unrestricted hidden manipulation of arbitrary GUI interfaces.
  • Evidence: The team is discussing Android harness work with Hermes.; The Q&A mentions running an Ubuntu or GUI Docker container within Android and working through the activity framework.; The infrastructure offering currently lists instant sandboxes for Windows, Linux, and Android, with macOS described as forthcoming.
  • Caveats: The mobile answer is exploratory and does not establish production support, performance, or a path for equivalent iOS background control.
  • Implications: Android and iOS should be treated as different execution products rather than assumed to share the desktop driver's abstraction cleanly.

Notable Concepts & Terms

  • Computer-use 1.0: The conventional loop in which an agent observes a screenshot, reasons, and issues foreground clicks, typing, or scrolling while occupying the user's screen.
  • KuaDriver / Quadriver / CodeDriver: The transcript's varying names for Cua's local, cross-platform driver that exposes desktop state and supports background GUI actions.
  • Accessibility-tree-first execution: A control hierarchy that first uses structured operating-system UI metadata, then falls back to pixel-level interaction when structured execution fails.
  • KuaBench: Cua's reproducible GUI-agent benchmark framework built around setup, oracle, and evaluator functions.
  • KuaBench KiCAD: A Snorkel AI collaboration that evaluates agents on electrical-engineering workflows and verifies outcomes by simulating the resulting circuits.
  • Reward hacking: An agent exploiting weaknesses in an environment or evaluator to receive credit without completing the intended task; Cua adversarially tests tasks for it.
  • Trajectory forking: Restoring the computer to an intermediate point in a recorded run so predicted state or reward can be compared with the actual subsequent outcome.
  • Demand-based warm sandbox pool: A dynamically sized reserve of ready environments intended to eliminate costly GPU waiting during RL rollout generation.

Operator Notes / Why Ken Should Care

  • Run a controlled pilot comparing full-desktop screenshots, window-scoped observations, and accessibility-first execution on the same internal tasks; track completion rate, tokens, latency, and recovery frequency.
  • Add blank-state creation tasks and template-based editing tasks as separate benchmark categories so aggregate scores do not hide the capability gap shown in KiCAD.
  • Require every GUI eval to pass adversarial shortcut and reward-hacking review before its score influences model routing, release gates, or vendor selection.
  • Instrument GPU idle time attributable specifically to sandbox launch and reset; use the result to decide whether a warm-pool architecture is economically justified.
  • Perform a security review before adopting background control, covering least-privilege accessibility permissions, credential exposure, destructive-action confirmation, audit logs, and rollback.
  • Verify the project's current repository, product naming, platform support, and licensing because the transcript uses multiple driver names and describes macOS sandbox support as forthcoming.
  • Ask Cua for benchmark methodology, tested model names, per-task results, and raw cost/utilization data before treating the reported improvements as procurement-grade evidence.

Source/Metadata

  • Title: Computer-Use 2.0: Agents Just Got Multi-Cursor — Francesco Bonacci, Cua
  • Transcript words: 2516
  • Duration seconds: 1001
  • Timestamp note: No timestamps or chapter markers were present in the provided transcript.
Full transcript 2286 words · 11 min read
0:00

[SPEAKER_01] Thank you for taking the time for coming over here.

0:12

SPEAKER_01

I'm Francesco, I'm the CEO of the company. Alongside me, a couple of other folks, my CTO, Dillon, and my chief of infra, Rob, they're going to work on the stage in a while. But before we do that, who's excited for some computer using agent talk happening now? Are you guys excited? Lovely. If I were to ask, what was a computer using agent one year ago, probably half the crowd would say I don't have any idea what computer use means. So today I'm going to take you on a journey from our vision where we come from so far on computer user, this new shape of agents that are talking and up to modern intelligence.

0:34

SPEAKER_01

So we're going to start with the vision of the Quadriver, where we're coming from. And how many of you guys have been working with computer users for one year? How about three years? Lovely. Okay.

1:08

SPEAKER_01

So our team has plenty of experience. We go all the way back to our time on Microsoft.

1:14

SPEAKER_01

We were working on this type of GUI agents. We were calling them back in the days. And there is an example of old-fashioned human agent loop. We basically refer to this as a human loop where you will have an agent loop. You will take a screenshot that the agents will have to reason and plan through. And then you will work with an action space in terms of clicking, typing, scrolling around. So this is what we refer to as the old-fashioned computer use 1.0 just to set the tone for this space. And we've come a long way since this type of computer using agents.

1:38

SPEAKER_01

So this is the old-fashioned way of representing this agent loop as a human would do. Over two months ago, we released a project in the open source. It's called Quadriver. We made it working in the background. That means your computer user will not take over your screen as the computer use 1.0 kind of agent loop was doing back in the days. And it all started from Codex releasing their computer user model two months ago. So we took the challenge because we were really working with this type of background computer user. So over one weekend, we had something together. And the trick here is really not having your agents take over your screen.

2:39

SPEAKER_01

So there is a lot of dark magic happening behind the scenes just to give you some context. There are some undocumented API living in the Apple framework. And it basically ships with your laptop. And as you can see here in the demo, you have an AI agent that is not taking over control of your laptop. We made it working not only for Mac OS but also spanning across Windows and Linux. This is the very first driver that is living on your laptop. And it lets any AI agents connect to the underlying operating system, either using accessibility trees or a screenshot-level approach. We take both. This is what really the agents see. You will have to install Quadriver.

3:12

SPEAKER_01

The agents will take a snapshot of the Windows state. And you will have to observe. And we really take one different action path to really make ground computer use happening. So you really have to observe the space. So in this case, just by calling get Windows state, you get an accessibility tree representation plus a screenshot. And then you will try a background execution using accessibility tree. And if that doesn't work, we go all the way and make the heavy lifting for you and just try a pixel background click. This is best practice for background at this stage. It's not behaving the same way on Mac OS, Windows, and Linux.

4:03

SPEAKER_01

So we do some of the heavy lifting for you so that your AI agent can run on this tour on your MacBook. How do we manage to not break anything between release cycles? We have a lot of investment happening behind the scenes when we test new releases.

4:26

SPEAKER_01

We have about eight different application harnesses that we use for making sure that we don't break anything among different releases. Among our early adopters, you can see Clicky, Hermes, Quencoder, H Company, and Droid Factory. Huge thanks to them for using CodeDriver and basically releasing a lot of upstream contribution in our framework. Without further ado, I'm going to move to the next part of the presentation, which is going to be intelligence. And I'm going to have our CTO, Dillon, cover that. [SPEAKER_00] Hello. So thank you, Francesco. [SPEAKER_00] With CodeDriver, we gave an agent hands.

5:02

SPEAKER_01

[SPEAKER_00] But then the question becomes, how can you trust the agent to use those hands correctly and not leave anything broken behind? [SPEAKER_00] And to answer that, we had to build KuaBench. [SPEAKER_00] So for a show of hands, who here has heard of Terminal Bench or Harbor? [SPEAKER_00] Yeah, so a few of you have heard of it. [SPEAKER_00] And if you've ever authored a task for Terminal Bench, then this might look familiar. [SPEAKER_00] But in KuaBench, a task is made of three pieces. [SPEAKER_00] The setup function, which sets up the machine into initial state. [SPEAKER_00] The oracle function, which provides a golden trajectory for the task.

6:05

SPEAKER_01

[SPEAKER_00] And the evaluator, which probes the environment to check if the agent successfully completed the task. [SPEAKER_00] Unlike Terminal Bench, the oracle here is GUI actions. [SPEAKER_00] So it looks kind of like a pile of GUI when you write that. [SPEAKER_00] And writing environments takes scale and expertise.

6:34

SPEAKER_01

[SPEAKER_00] On desktop, there's more than five platforms that we target.

6:37

SPEAKER_00

And we try to collapse that into a single Python file. So using the KuaBench SDK, you can write a GUI that works across every desktop platform in a single Python file. And use the same SDK to probe that GUI to get usable agent data. Anyone or any agent can author one of these tasks. And when you put that to work, you get a real catalog. We have currently over 130 verifiable tasks, 42 environments, and across five platforms. So it looks like a pile of GUI when you write that. And writing environments takes scale and expertise. On desktop, there's more than five platforms that we target. And we try to collapse that into a single Python file.

7:21

SPEAKER_00

So using the KuaBench SDK, you can write a GUI that works across every desktop platform in a single Python file. And use the same SDK to probe that GUI to get usable agent data. Anyone or any agent can author one of these tasks. And when you put that to work, you get a real catalog. We have currently over 130 verifiable tasks, 42 environments, and across five platforms. And each of these are easily reproducible using our CLI. And the latest addition to our data sets is one that we're proud of. With collaboration with Snorkel AI, we built KuaBench KiCAD, which tests computer use agents on electrical engineering tasks,

7:59

SPEAKER_00

using software by real professionals, and evaluator functions that actually simulate the circuits. But the results are humbling. The top agent that we tested only got a full pass on six out of 25 of these tasks. Of those six, 100% of them involved editing an existing schematic. And when we start the task from a blank schematic, the success rate drops to 0%. And across all the models that we tested, the leaderboard is flat. No model has achieved more than 30% reward. But once you can score something, you can improve it. If we take a look at the KuaBench basic data set, scaled up to 4K resolution, testing an agent, they typically get around 62% pass rate.

8:53

SPEAKER_00

But when you switch the agent computer tool from the built-in one to KuaDriver, the pass rate jumps from 62% to 80% using 34% less tokens.

9:04

SPEAKER_00

And this is primarily because KuaDriver focuses on a window rather than the entire desktop. But our evals might say you can trust model XYZ at task whatever. But how can you know that the task, how can you know that the eval can be trusted? So before we test a task against any agent, we first try to break the environment ourselves. We have a matrix of agents attempt to do reward hacking and attempting to break the environment. And we take all that data and we compile it into a nice CodeRabbit-style code review. And only tasks that survive our pipeline can enter the data set. And if you ask us how we trust that agent, the answer is that it's just evals all the way down.

9:43

SPEAKER_00

But to measure the intelligence of an agent, you can't just measure its ability to successfully perform actions. You also have to measure its ability to understand the world that it's operating in. Every run that we record can be forked through any moment in its trajectory to give you the state of the computer at that moment. From there, we can probe a model asking to predict the reward, the internal state, or any other observation of the computer and compare it against the fork. And that prediction is the world model of the agent made measurable. And with that, I'll let Robert take the stage. Thank you, Dylan. [SPEAKER_02] Hello, everybody.

10:36

SPEAKER_00

[SPEAKER_02] I am the Chief Infra Officer at Kua. [SPEAKER_02] And I'm here to talk to you about how you're probably leaving a lot of money on the table with idle GPUs [SPEAKER_02] if you do RL training for computer use agents. [SPEAKER_02] So I want to introduce this diagram to you all. [SPEAKER_02] Can I get, is there general familiarity with this diagram?

10:54

SPEAKER_02

Or is this something that most of us haven't seen before? Anyone? I don't believe anyone. Awesome. Very niche. Almost everything on this is not really important for what we're talking about, but the blue portions are. And what those represent are GPUs generating tokens for RL training. And if you zoom in on this a little bit, you can see how this typically looks with a sandbox environment. You're going to be generating some tokens, and then you finish your task on a sandbox, and then you're waiting for either a new sandbox to spin up or for your existing one to reset. The problem here is that this is just pure cost. Your GPU really isn't doing anything useful here.

11:29

SPEAKER_02

And, I don't know if you've heard, but GPU time is pretty expensive right now. So, as you're scaling, this cost really compounds, and you really want to focus on minimizing this if possible. So, one thing that you might try to do is minimize the startup time of your sandbox. And, I mean, you should do that. That's a great thing to do. But, especially for computer-use style environments, sometimes this can be a little impractical. Your researchers might give you a 40-gigabyte environment, and that might just be necessary, and it takes a long time to pull that down and start it up. So, how do you design your training infrastructure so that you can minimize the GPU startup,

12:24

SPEAKER_02

or minimize the startup time of the sandbox, even when the sandbox is not well designed to be startup quickly? So, the way we do this is a pool. And this is supposed to be animated, but it's not animating. So, I guess I'll just explain to you orally. What that is, we have a set of GPUs here, which all want to use sandbox. And what we will do is use a demand-based autoscaler to detect how many GPUs currently need a sandbox. And we can grow the pool to be that size on demand. And what that means is that if you have a warm pool that you want to allocate to your GPU cluster, you don't actually need to know upfront what that warm pool size is.

13:12

SPEAKER_02

We can figure out what that warm pool size should be for you on demand. And that might even change over the course of your multi-day training run. You might start needing a lot of sandboxes, but then as your generations get longer, you might need less. So, these also could be two to four times cheaper than your GPUs. So, having a little redundancy here, you still wind up saving money because you're maximizing the use of your GPU time. Come see me after if you want to see the animation because it's cool. So, now when you have this redundancy in your pool, you're paying the cost of that startup time on the infrastructure side, not on the GPU side.

13:55

SPEAKER_02

So, your GPU workers have full utilization. And because we use this, we can give you instant sandboxes for your GPUs for Windows, Linux, Android, and macOS is coming up. And I'm going to hand it back to Francesco to close it out for us. Lovely. Thank you, Dillon. So, having a little bit of redundancy here, you still wind up saving money because you're maximizing the use of your GPU time. Yeah, come see me after if you want to see the animation because it's cool. So, now when you have this redundancy in your pool, you're paying the cost of that startup time on the infrastructure side, not on the GPU side. So your GPU workers have full utilization.

15:02

SPEAKER_02

And then because we use this, we can give you instant sandboxes for your GPUs for Windows, Linux, Android, and macOS is coming up. And I'm going to hand it back to Francesco to close it out for us.

15:04

SPEAKER_01

[SPEAKER_02] Lovely. [SPEAKER_02] Thank you, Dillon. Thank you, Rob, for taking this over. We do have plenty of time for Q&A.

15:30

SPEAKER_01

So, if you guys have any questions, I'm happy to take them. Either for Quadriver, QuaBench, basically what Dillon presented, or QuaFleet, which is what Robert covered. Any questions? Otherwise, we can wrap this up. Oh, I see.

15:41

[SPEAKER_01] What if it's possible to operate peer-reviewed agents in the background? [SPEAKER_01] How do you do this for Mac? [SPEAKER_02] Mm-hmm. Like, a mobile? So, the story for mobile, Android, there is very far you can go. We are talking with the Hermes team, because they do have a harness that runs on Android. I guess, if you're talking about background, there is some level of background that can happen if you containerize a workload. And, basically, on Android, you can even run your own container or Ubuntu or a GUI Docker container within Android.

16:27

SPEAKER_01

But, yeah, the Android ecosystem, especially compared to iOS, is more inclined to that form of background computer use.

16:38

SPEAKER_01

But it's more towards tool use than really controlling GUI interface. We work with the activity framework and do tool use in the background.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note