Dhrou Bhattra Reviewer Reviewer Welcome. Let's get started. So my name is Dhrou Bhattra. I want to talk to you about an argument. There's an argument online that says, roughly speaking, AI agents will be the drivers of actions on the web, not humans, that there will be more agents clicking buttons, reading things, buying things for us than human eyeballs on the web. You drive one step deeper into the argument and you ask how this will happen because today the web is extremely hostile to automated traffic, and the answer usually is the web will be agentified. You drive one step deeper into what that means and how that will happen,
and the answer usually is with APIs that agents will use, and this will call via some set of protocols, and the set of protocols grows over time. It was initially supposed to be MCP servers, then web MCP, and for payments there are 20 different competing protocols, and every single company wants to introduce their own protocol along the way. But okay, there will be APIs is the argument. My claim today and argument, and what I hope to convince you today, is that this last bit is wrong. I think the first two I generally agree with, this last bit that suddenly the web will provide you APIs for accessing things is just delusional.
And my claim is computer use agents and computer use models will agentify the web, not APIs, and more specifically the long tail of the web. The head of the distribution, the most popular website, perhaps will give you the API, but the long tail will not. And the story starts usually with these sorts of use cases. I agree that they also frustrate me. When people show a use case of, find me business class flights from Miami to Palma de Mallorca for August, and somehow they expect that behind the scenes it will be a computer use agent clicking buttons on flights.google.com. I find this example, this is my own example,
I'm not dunking on anybody else, I find this example bizarre. Why would you possibly do it this way? Aren't you aware that there are already aggregators? You send them a variable, they will send you a JSON. Your LLMs can do tool calling. Why would you click buttons? The purpose of those clicking buttons is to generate this behind the scenes. So we're on the same page there. The next step usually is, okay, let's roll this out to some other case. I want to know, does my favorite restaurant or does this restaurant have gluten-free items on the restaurant's menu? And so I assume that there's going to be a similar endpoint somewhere
for that restaurant website or via an aggregator, oops, that accepts a query that I can filter through, where I can ask for, give me your menu items that are gluten-free. I want to, for those of you who can already see, I want to rid you of the delusion that such an endpoint exists. And I want to show you, just for the sake of being on the same page, what restaurant pages look like. There are three that are on the screen. This is the first one. This is what we imagine a prototypical, this is easy mode. You go to a web page, there's no API, but at least it's text. It's in German, sure, a different part of the world. Models can speak languages.
I see text and I see pricings. Maybe this is easily scrapable. Okay. This is easy mode. The medium mode is a page like this, where of course the thing that you're looking for is slightly hidden away. You press the menu button, it actually takes you to a PDF. Okay, fine. If your agent has to do this, it needs a PDF reader mode, fine. I can highlight these. Okay, so this is released text, no problem. I will dump this into ChatGPT or whatever and it'll do it. Here's the hard mode. This is what a hard mode website for a restaurant looks like, where it's just pictures of the people who created the restaurant, of where they're located, of menu items.
And you're like, where is the menu? Does anybody speak Spanish? Nuestra carta? See? Okay, let's check. That's our menu. You click there. Oh my God. What am I staring at? Okay, you click on it. Oh, it's a gallery of individual, what is this? This is pixelated, there's no text here. I am looking at JPEGs embedded in a gallery which contain the PDF items. You download this and you put it into ChatGPT and it's struggling doing OCR on this thing. Okay, so this is what the web looks like. You're telling me this will give you an API endpoint that you can pass gluten-free items and it'll tell you what that is. That's example number one. Let's take a look at another example.
This is much more enterprise-focused. I am a business. I want to sell to a school district. There's about 15,000 to 20,000 school districts in the US. And I want to ask the question, is this school district where I give you mydistrict.gov, is this school district procuring a laptop right now? Right? It's a simple question. And I want a similar endpoint. I want to say RFP status open and topic laptop. And you think this will exist. Let me tell you what these websites look like. Again, this is what easy mode looks like. This is Ithaca Public Schools. Okay, there's some portal. Maybe enrollment is the wrong place to go. This is what slightly harder mode looks like.
This is a different public school. You look at the menu under finance. There's purchasing. Under purchasing, there are certain solicitations. There's a lot of questions. Scan PDF again. Okay, very nice. So, no text in here. But this is how they tell you about what they have purchased. And for ultimate boss level, I want to show you a different school district where, in order to find information, you have to file a request for access of information under Freedom of Information Act. And then what they do is they scan your email that you sent to them. And they will scan that email, put it on a Google Drive, and then attach PDFs associated with your request.
These are the people you're telling me will give you an MCP server? The amount of delusion here is off the chart. If there was ever a time to say go touch grass, I think this is it. Okay. So, this is what... My claim here is that the web we forget is massive. It is extremely big.
The number of websites out there, the active websites are somewhere around 200 million. The total number of websites is sitting at a billion. Infrastructure changes very slowly. You can imagine, as an engineer, getting unfettered access to software systems and letting your favorite coding agent rip and generate an API endpoint. Even if that technological problem is solved, which I agree seems on the horizon, it should eventually be solved, you're not going to get unfettered access to these institutions. And these institutions change very slowly. There are still places that are faxing each other. You're not going to be able to suddenly change this overnight. Okay.
So, at this point in time, you might be thinking, fine. I'm stuck with the infra as it is. So there's not going to be APIs available. But I have coding agents. Why don't I just throw them at the HTML? There is, after all, if the browser is doing it, it's a piece of code. My coding agent should be able to do it as well. Fair point. I also thought this way two years ago. Let me tell you what the web actually looks like. If I ask you the question, what was the final score of this game between Minnesota Timberwolves and the Brooklyn Nets? Here's what the web page looks like. This is a modern website. It's NBA.com. You and I can see the score, 125 versus 109. Okay.
Behind the scenes, when you actually load the page and read it, this is what initially gets loaded. There's an empty placeholder initially. And you wait a few hundred milliseconds to a few seconds, depending on your network connection. And your browser makes an asynchronous call later to fetch the information. So the browser makes a call to an endpoint that it's extracting that information from. That endpoint responds with the JSON that contains the answer. If you just read the HTML when you load it, the answer is not in the HTML. And so your chatbot doesn't have access to that either. So okay, you go, fine, these are just details.
I just have to add some waits and sleeps and sure. Okay, fair enough. I will take you to another example. I ask you the question on a product web page, is the 25mm Osmium Cube in stock or out of stock? You're on this product website. You scroll down. There is a drop-down. That drop-down is telling you and me as humans, three things are sold out. And you wait a few hundred milliseconds to a few seconds, depending on your network connection. And your browser makes an asynchronous call later to fetch the information. So, the browser makes a call to an endpoint that it's extracting that information from. That endpoint responds with the JSON that contains the answer.
If you just read the HTML when you load it, the answer is not in the HTML. And so, your chatbot doesn't have access to that either. So, okay, you go, fine, these are just details. I just have to add some weights and sleeps, and sure. Okay, fair enough. I will take you to another example. I ask you the question on a product web page, is the 25mm Osmium Cube in stock or out of stock? You're on this product website. You scroll down. There is a drop-down. That drop-down is telling you and I as humans three things are sold out. One thing is in stock, even though it doesn't actually say in stock. There's no text there that says in stock.
But you understand that grayed out means sold out. And sometimes the sold out won't actually be there as text. Sometimes it will just be unclickable, grayed out. Okay, surely this information must be in the code somewhere. You go and read the HTML, and it turns out there is an option selector. It actually doesn't say any of the things that I'm seeing on screen. It doesn't say sold out. It doesn't say available. So, what's going on? It turns out behind the scenes, the browser makes a call, gets a JSON object, which is the variable product count. That contains a variable called quantity. That quantity is 10 sometimes, zero sometimes.
That's just how many things can the backend support right now. Some of those quantities are zero. And there's a different rendering script that, any time there's zero, grays it out and makes it unclickable. Fundamentally, what is happening here is this information that you are seeing on screen is not written somewhere as pure text. It is calculated. It is rendered. And for people who work in the browser industry, they understand this. But often, people who are coming from an AI background like me, we didn't always understand this. The browser is a rendering engine. You are seeing pixels on screen. Think of it as a game engine.
And you're asking, can I not read the source code of the game and predict exactly what the pixels are going to be? Well, yes, eventually. But right now you're asking for an exact inversion of that process. Fundamentally, the web was built for human eyes. Pixels are the source of the truth because the consumers of the websites are humans. That is what it was built for. And so there is an implication here that we have to grapple with, which is machines will need vision to operate those things because the web was built for human consumption.
In a way, this is the bitter lesson for web agents, that the more you end up writing scaffolds around existing websites, it doesn't actually generalize to the long tail of the web. The thing that generalizes is the thing that it was designed for, which is the most general solution, just pixels in. This is what we and some others in the area have been working on. We have a model called Navigator. The first version of the model went out in November last year. The first version acted purely like a human. Screenshot in, button clicks, and scrolls out. I will tell you in a second that is not what you should settle on, but it is a general solution.
It lets you do things like this. I give you an e-commerce website, and I tell you there is this discount code. Please go.
The discount code is applicable under certain constraints. Maybe it's only on a product, it's only on a certain set of dates. Maybe it's only with this minimum cart threshold. I describe that in natural language, and I tell you, tell me if this discount code is valid or not. There is no API for this. The way to do it is just like a human can. You go to that website, you find the product that is described, you add it to cart, you apply the discount code, and you check whether the claim of 22% off was met or not. And that is screenshot in, button click out. This trajectory took 20, 30, 40 steps, depending on the sophistication of the task.
If you can do it on your browser, this model can accomplish it in principle. In practice, of course, there are accuracy gaps and so on. But in principle, this task is solvable. Whereas in a lot of earlier cases, even in principle, that task may not be solvable. So my claim is the web was built for human eyes, machines will need vision. But of course, they do not need to be limited to human ways. Just because for the long tail, you need to have a capability does not mean that is the only way you should do it. Here is an example showing that. The next version of the model that we trained can also write JavaScript on demand. So this is a Chrome extension.
On the right, you see an action that says execute JS, value default, text select. On the left, you saw that the model filled out multiple form fields simultaneously.
The reason why it could do that is because it wrote a little bit of a function. It can read the code when necessary. It can write code when necessary because the browser, after all, is an engine that can execute code. But it has a formal verification system built in. It is seeing the screenshot. That is the source of the truth. So it knows whether it succeeded or not. So click buttons when you have to, write code when you have to, and look at the result through pixels because that is the source of truth. And of course, because these things are machines, you can string them into multi-agent systems.
So you can have an orchestrator that is launching multiple navigators in parallel, each with a cloud sandbox instance. They are clicking buttons on multiple websites. So you can accomplish things that would be superhuman because no human would be able to parallelize over that many instances. Around here, usually, in this flow of an argument, is when people start asking, are computer use models actually good enough for these tasks? There is actually a perception online that it's not clear whether progress on computer use has been fast, and there are questions about why progress has been slow.
That's not the reality I am seeing, and that's not the reality that the numbers back up. This is a popular benchmark, online mind to web. No benchmark is perfect. The point isn't that this is the right solution. But on the x-axis are release times of different models. On the y-axis is performance, which is human eval on this. So a human went in, looked at the trajectory, and decided whether it was correct or not. And basically, this particular version of the benchmark is saturated. The last model that we just released, Navigator N1.5, is sitting at 97% human eval. Eight trajectories out of 300 are incorrect.
At this point in time, you should just retire the benchmark, build something harder. There's about 30 to 50 steps of interaction that are happening. The next step is to go for something larger. So at least in numbers, what I'm seeing, we're seeing steady progress in computer use agents being able to do more and more things. And this is the model that I showed: pixels in, button clicks, and code out. Around this time, usually, is when people start asking questions. Okay, so they're getting better. But aren't these things slow? Because after all, you're looking at a screen, you're clicking a button, there are lots of buttons to click.
And these things are expensive if you're running them for hundreds of things. That claim, I think, is largely true. There is some truth to it. But I think people forget how much you can optimize these things out. So these are our results compared to the frontier models. Opus 4.7, GPT 5.5, on a couple of browser use benchmarks. We're slightly better, but I think that's within statistical noise in terms of accuracy. That improvement, I wouldn't beat the drum on. What I would emphasize is latency per step and cost per task. In terms of latency, this is a smaller footprint model. That's why it's a lot faster than some of the trillion-parameter-plus models.
And there are corresponding cost savings. So if you have something like, on these data sets, 20, 30 steps of interaction, you're looking at 80 cents per task versus $2.30. And that makes a big difference. And so the models are getting cheaper in that sense, that you can launch them at scale. So hopefully, at a high level, I've convinced you that there is something off with this argument. This is what my goal was. This is where I started. AI agents are going to be the primary drivers of action. That seems an uncontestable statement because the underlying intelligence of the models is becoming larger and larger.
There is a certain gain of productivity and efficiency and just ease of life that you get. In terms of latency, this is a smaller footprint model. That's why it's a lot faster than some of the trillion-parameter-plus models. And there are corresponding cost savings. So if you have something like, on these data sets, 20, 30 steps of interaction, you're looking at 80 cents per task versus $2.30. And that makes a big difference. And so the models are getting cheaper in the sense that you can launch them at scale. So hopefully, at a high level, I've convinced you that there is something off with this argument. This is what my goal was. This is where I started.
AI agents are going to be the primary drivers of action. That seems an uncontestable statement because the underlying intelligence of the models is becoming larger and larger. There is a certain gain of productivity and efficiency and just ease of life that you get. So it makes sense. So I think this hypothesis that suddenly, overnight, 30 years of infrastructure that was built layer upon layer for human consumption will, in what, two, five, ten years, be reinvented is, I think, a fantasy. But this is ultimately what you want, right? Ultimately, you want an endpoint that some higher-level entity can go to, and I say, I want you to do X. There is some task.
Maybe I give you that task description in natural language. Maybe I have some programmatic description with parameters.