[SPEAKER_01] Thank you for taking the time for coming over here.
SPEAKER_01
I'm Francesco, I'm the CEO of the company. Alongside me, a couple of other folks, my CTO, Dillon, and my chief of infra, Rob, they're going to work on the stage in a while. But before we do that, who's excited for some computer using agent talk happening now? Are you guys excited? Lovely. If I were to ask, what was a computer using agent one year ago, probably half the crowd would say I don't have any idea what computer use means. So today I'm going to take you on a journey from our vision where we come from so far on computer user, this new shape of agents that are talking and up to modern intelligence.
SPEAKER_01
So we're going to start with the vision of the Quadriver, where we're coming from. And how many of you guys have been working with computer users for one year? How about three years? Lovely. Okay.
SPEAKER_01
So our team has plenty of experience. We go all the way back to our time on Microsoft.
SPEAKER_01
We were working on this type of GUI agents. We were calling them back in the days. And there is an example of old-fashioned human agent loop. We basically refer to this as a human loop where you will have an agent loop. You will take a screenshot that the agents will have to reason and plan through. And then you will work with an action space in terms of clicking, typing, scrolling around. So this is what we refer to as the old-fashioned computer use 1.0 just to set the tone for this space. And we've come a long way since this type of computer using agents.
SPEAKER_01
So this is the old-fashioned way of representing this agent loop as a human would do. Over two months ago, we released a project in the open source. It's called Quadriver. We made it working in the background. That means your computer user will not take over your screen as the computer use 1.0 kind of agent loop was doing back in the days. And it all started from Codex releasing their computer user model two months ago. So we took the challenge because we were really working with this type of background computer user. So over one weekend, we had something together. And the trick here is really not having your agents take over your screen.
SPEAKER_01
So there is a lot of dark magic happening behind the scenes just to give you some context. There are some undocumented API living in the Apple framework. And it basically ships with your laptop. And as you can see here in the demo, you have an AI agent that is not taking over control of your laptop. We made it working not only for Mac OS but also spanning across Windows and Linux. This is the very first driver that is living on your laptop. And it lets any AI agents connect to the underlying operating system, either using accessibility trees or a screenshot-level approach. We take both. This is what really the agents see. You will have to install Quadriver.
SPEAKER_01
The agents will take a snapshot of the Windows state. And you will have to observe. And we really take one different action path to really make ground computer use happening. So you really have to observe the space. So in this case, just by calling get Windows state, you get an accessibility tree representation plus a screenshot. And then you will try a background execution using accessibility tree. And if that doesn't work, we go all the way and make the heavy lifting for you and just try a pixel background click. This is best practice for background at this stage. It's not behaving the same way on Mac OS, Windows, and Linux.
SPEAKER_01
So we do some of the heavy lifting for you so that your AI agent can run on this tour on your MacBook. How do we manage to not break anything between release cycles? We have a lot of investment happening behind the scenes when we test new releases.
SPEAKER_01
We have about eight different application harnesses that we use for making sure that we don't break anything among different releases. Among our early adopters, you can see Clicky, Hermes, Quencoder, H Company, and Droid Factory. Huge thanks to them for using CodeDriver and basically releasing a lot of upstream contribution in our framework. Without further ado, I'm going to move to the next part of the presentation, which is going to be intelligence. And I'm going to have our CTO, Dillon, cover that. [SPEAKER_00] Hello. So thank you, Francesco. [SPEAKER_00] With CodeDriver, we gave an agent hands.
SPEAKER_01
[SPEAKER_00] But then the question becomes, how can you trust the agent to use those hands correctly and not leave anything broken behind? [SPEAKER_00] And to answer that, we had to build KuaBench. [SPEAKER_00] So for a show of hands, who here has heard of Terminal Bench or Harbor? [SPEAKER_00] Yeah, so a few of you have heard of it. [SPEAKER_00] And if you've ever authored a task for Terminal Bench, then this might look familiar. [SPEAKER_00] But in KuaBench, a task is made of three pieces. [SPEAKER_00] The setup function, which sets up the machine into initial state. [SPEAKER_00] The oracle function, which provides a golden trajectory for the task.
SPEAKER_01
[SPEAKER_00] And the evaluator, which probes the environment to check if the agent successfully completed the task. [SPEAKER_00] Unlike Terminal Bench, the oracle here is GUI actions. [SPEAKER_00] So it looks kind of like a pile of GUI when you write that. [SPEAKER_00] And writing environments takes scale and expertise.
SPEAKER_01
[SPEAKER_00] On desktop, there's more than five platforms that we target.
SPEAKER_00
And we try to collapse that into a single Python file. So using the KuaBench SDK, you can write a GUI that works across every desktop platform in a single Python file. And use the same SDK to probe that GUI to get usable agent data. Anyone or any agent can author one of these tasks. And when you put that to work, you get a real catalog. We have currently over 130 verifiable tasks, 42 environments, and across five platforms. So it looks like a pile of GUI when you write that. And writing environments takes scale and expertise. On desktop, there's more than five platforms that we target. And we try to collapse that into a single Python file.
SPEAKER_00
So using the KuaBench SDK, you can write a GUI that works across every desktop platform in a single Python file. And use the same SDK to probe that GUI to get usable agent data. Anyone or any agent can author one of these tasks. And when you put that to work, you get a real catalog. We have currently over 130 verifiable tasks, 42 environments, and across five platforms. And each of these are easily reproducible using our CLI. And the latest addition to our data sets is one that we're proud of. With collaboration with Snorkel AI, we built KuaBench KiCAD, which tests computer use agents on electrical engineering tasks,
SPEAKER_00
using software by real professionals, and evaluator functions that actually simulate the circuits. But the results are humbling. The top agent that we tested only got a full pass on six out of 25 of these tasks. Of those six, 100% of them involved editing an existing schematic. And when we start the task from a blank schematic, the success rate drops to 0%. And across all the models that we tested, the leaderboard is flat. No model has achieved more than 30% reward. But once you can score something, you can improve it. If we take a look at the KuaBench basic data set, scaled up to 4K resolution, testing an agent, they typically get around 62% pass rate.
SPEAKER_00
But when you switch the agent computer tool from the built-in one to KuaDriver, the pass rate jumps from 62% to 80% using 34% less tokens.
SPEAKER_00
And this is primarily because KuaDriver focuses on a window rather than the entire desktop. But our evals might say you can trust model XYZ at task whatever. But how can you know that the task, how can you know that the eval can be trusted? So before we test a task against any agent, we first try to break the environment ourselves. We have a matrix of agents attempt to do reward hacking and attempting to break the environment. And we take all that data and we compile it into a nice CodeRabbit-style code review. And only tasks that survive our pipeline can enter the data set. And if you ask us how we trust that agent, the answer is that it's just evals all the way down.
SPEAKER_00
But to measure the intelligence of an agent, you can't just measure its ability to successfully perform actions. You also have to measure its ability to understand the world that it's operating in. Every run that we record can be forked through any moment in its trajectory to give you the state of the computer at that moment. From there, we can probe a model asking to predict the reward, the internal state, or any other observation of the computer and compare it against the fork. And that prediction is the world model of the agent made measurable. And with that, I'll let Robert take the stage. Thank you, Dylan. [SPEAKER_02] Hello, everybody.
SPEAKER_00
[SPEAKER_02] I am the Chief Infra Officer at Kua. [SPEAKER_02] And I'm here to talk to you about how you're probably leaving a lot of money on the table with idle GPUs [SPEAKER_02] if you do RL training for computer use agents. [SPEAKER_02] So I want to introduce this diagram to you all. [SPEAKER_02] Can I get, is there general familiarity with this diagram?
SPEAKER_02
Or is this something that most of us haven't seen before? Anyone? I don't believe anyone. Awesome. Very niche. Almost everything on this is not really important for what we're talking about, but the blue portions are. And what those represent are GPUs generating tokens for RL training. And if you zoom in on this a little bit, you can see how this typically looks with a sandbox environment. You're going to be generating some tokens, and then you finish your task on a sandbox, and then you're waiting for either a new sandbox to spin up or for your existing one to reset. The problem here is that this is just pure cost. Your GPU really isn't doing anything useful here.
SPEAKER_02
And, I don't know if you've heard, but GPU time is pretty expensive right now. So, as you're scaling, this cost really compounds, and you really want to focus on minimizing this if possible. So, one thing that you might try to do is minimize the startup time of your sandbox. And, I mean, you should do that. That's a great thing to do. But, especially for computer-use style environments, sometimes this can be a little impractical. Your researchers might give you a 40-gigabyte environment, and that might just be necessary, and it takes a long time to pull that down and start it up. So, how do you design your training infrastructure so that you can minimize the GPU startup,
SPEAKER_02
or minimize the startup time of the sandbox, even when the sandbox is not well designed to be startup quickly? So, the way we do this is a pool. And this is supposed to be animated, but it's not animating. So, I guess I'll just explain to you orally. What that is, we have a set of GPUs here, which all want to use sandbox. And what we will do is use a demand-based autoscaler to detect how many GPUs currently need a sandbox. And we can grow the pool to be that size on demand. And what that means is that if you have a warm pool that you want to allocate to your GPU cluster, you don't actually need to know upfront what that warm pool size is.
SPEAKER_02
We can figure out what that warm pool size should be for you on demand. And that might even change over the course of your multi-day training run. You might start needing a lot of sandboxes, but then as your generations get longer, you might need less. So, these also could be two to four times cheaper than your GPUs. So, having a little redundancy here, you still wind up saving money because you're maximizing the use of your GPU time. Come see me after if you want to see the animation because it's cool. So, now when you have this redundancy in your pool, you're paying the cost of that startup time on the infrastructure side, not on the GPU side.
SPEAKER_02
So, your GPU workers have full utilization. And because we use this, we can give you instant sandboxes for your GPUs for Windows, Linux, Android, and macOS is coming up. And I'm going to hand it back to Francesco to close it out for us. Lovely. Thank you, Dillon. So, having a little bit of redundancy here, you still wind up saving money because you're maximizing the use of your GPU time. Yeah, come see me after if you want to see the animation because it's cool. So, now when you have this redundancy in your pool, you're paying the cost of that startup time on the infrastructure side, not on the GPU side. So your GPU workers have full utilization.
SPEAKER_02
And then because we use this, we can give you instant sandboxes for your GPUs for Windows, Linux, Android, and macOS is coming up. And I'm going to hand it back to Francesco to close it out for us.
SPEAKER_01
[SPEAKER_02] Lovely. [SPEAKER_02] Thank you, Dillon. Thank you, Rob, for taking this over. We do have plenty of time for Q&A.
SPEAKER_01
So, if you guys have any questions, I'm happy to take them. Either for Quadriver, QuaBench, basically what Dillon presented, or QuaFleet, which is what Robert covered. Any questions? Otherwise, we can wrap this up. Oh, I see.
[SPEAKER_01] What if it's possible to operate peer-reviewed agents in the background? [SPEAKER_01] How do you do this for Mac? [SPEAKER_02] Mm-hmm. Like, a mobile? So, the story for mobile, Android, there is very far you can go. We are talking with the Hermes team, because they do have a harness that runs on Android. I guess, if you're talking about background, there is some level of background that can happen if you containerize a workload. And, basically, on Android, you can even run your own container or Ubuntu or a GUI Docker container within Android.
SPEAKER_01
But, yeah, the Android ecosystem, especially compared to iOS, is more inclined to that form of background computer use.
SPEAKER_01
But it's more towards tool use than really controlling GUI interface. We work with the activity framework and do tool use in the background.