Computer-use models will agentify the web, not APIs — Dhruv Batra, Yutori
Description
To learn what a US school district is buying, you file a Freedom of Information Act request. Someone scans the email you sent, puts the scan on Google Drive, and attaches the relevant PDFs. Dhruv Batra's question is whether anyone seriously expects that office to publish an MCP server. He grants the popular claim that agents will drive most of the action on the web, then rejects its usual next step, that the web will meet them with APIs. The head of the distribution might. The long tail, some 200 million active sites where infrastructure changes over decades, will not. Reading the HTML instead does not save you, because much of what you see was never written down anywhere. A basketball score is missing from the page that first loads and arrives later as JSON. A product page contains no text reading sold out, only a quantity of zero and a script that grays the option out. State is calculated and rendered rather than stored, which makes the browser closer to a game engine than a document and makes pixels the source of truth. He calls this the bitter lesson for web agents: scaffolding built per site does not generalize, and the general solution is the one the web was actually built for. Their Navigator model runs screenshot in and clicks out, now writes JavaScript when that is quicker, and checks the result on screen. It misses 8 of 300 trajectories on a benchmark he thinks should be retired. Speaker info: - https://x.com/DhruvBatra_ - https://www.linkedin.com/in/dhruv-batra-dbatra/ - https://dhruvbatra.com Timestamps: 0:00 - The argument, and the part of it that is wrong 3:23 - Restaurant menus on easy, medium, and hard mode 5:31 - School district procurement, up to a FOIA request 7:38 - 200 million active sites that change slowly 8:42 - Why reading the HTML does not rescue it 10:14 - Sold out is not text, it is a rendered zero 11:33 - The browser is a rendering engine, pixels are the truth 12:53 - Navigator, and writing JavaScript when that is faster 15:51 - Are c
Summary
Generated by gpt-5.6-terraAt-a-Glance
- Verdict: Watch fully
- Core thesis: The web's long tail will be made agent-accessible primarily by vision-based computer-use models that operate existing interfaces, not by a universal rollout of APIs, MCP servers, or new agent protocols.
- Why it matters: This is a direct architecture argument for agent systems: robust web automation needs pixels as the ground-truth verification layer, with code and API access used opportunistically rather than assumed.
- Best use: Use it to pressure-test an API-first agent roadmap and extract the hybrid browser-agent pattern: visual navigation for universality, JavaScript for acceleration, and parallel sandboxed workers for scale.
Executive Summary
Dhruv Batra argues that the common progression from “agents will act on the web” to “the web will expose APIs for them” is mistaken. Large, sophisticated services may offer APIs, but the roughly 200 million active websites and institutional long tail will not modernize quickly enough to provide clean, structured, agent-ready interfaces. Their information is often buried in menus, PDFs, scanned documents, images, outdated portals, or manual processes.
The core technical point is that a webpage is not equivalent to its initial HTML. Modern sites render information asynchronously and conditionally: an NBA score may arrive after page load via JSON, while inventory state may be calculated from backend quantity data and represented only by a grayed-out, unclickable visual control. Since web experiences were built for human visual consumption, Batra argues that screenshots/pixels must be the final source of truth for general web agents.
His proposed operating model is hybrid rather than dogmatically human-like. Yutori's Navigator begins with screenshots in and browser actions out, but a newer version can generate JavaScript to fill fields or inspect/manipulate a page when that is faster. The agent then visually verifies the result. Multiple navigators can run in isolated cloud sandboxes under an orchestrator, allowing parallel execution beyond human capacity.
Batra claims progress is already substantial: Navigator N1.5 reportedly achieved 97% human-evaluated success on the Online Mind2Web benchmark, with 8 of 300 trajectories wrong. He also positions smaller specialized computer-use models as economically viable, citing roughly $0.80 per 20–30-step task versus $2.30 for frontier models on selected browser benchmarks. The presentation is strategically useful but product-promotional, and its benchmark and cost claims are presented without enough methodological detail to treat as independent proof.
Key Takeaways
- Claim: API-first agentification will cover the web's head, but computer-use agents are necessary for the heterogeneous long tail. | Evidence: Batra contrasts straightforward flight aggregation APIs with restaurant sites whose menus appear as German text, linked PDFs, or pixelated JPEG galleries, and school-district procurement information spread across nested portals, scanned PDFs, or even FOIA request workflows. He estimates about 200 million active websites and about 1 billion total websites. | Implication: Design web-agent systems with graceful API/tool use for structured sources, but retain browser automation as the universal fallback rather than treating it as an edge case. | Caveat: He does not argue against APIs where they already exist; his claim is that expecting widespread, rapid API retrofitting across legacy institutions is unrealistic.
- Claim: Reading HTML or relying on coding agents to parse page source is insufficient for reliable general web automation because meaningful state is often rendered after load or inferred visually. | Evidence: On NBA.com, the displayed 125–109 score is absent from initial HTML and arrives through a later asynchronous request. On an e-commerce product page, stock status is conveyed by grayed-out or disabled options; the underlying state is a quantity value in JSON plus a separate rendering script, not explicit “in stock” or “sold out” text. | Implication: Agent success criteria should validate rendered UI state, not merely DOM extraction, API-response parsing, or completion of a code path. | Caveat: Source inspection can still be valuable when the page exposes useful state, but it cannot be assumed to map cleanly to what a user sees.
- Claim: Pixels should be the source of truth for broad web agents because the web was designed as a rendering system for human eyes. | Evidence: Batra analogizes the browser to a game engine: deriving what pixels will appear from source code requires inverting a complex rendering process. Navigator's original interaction contract is screenshot input with clicks and scrolls as outputs. | Implication: For workflows involving visual affordances, dynamic content, localization, or poorly structured sites, build observation and verification around screenshots and UI effects. | Caveat: Visual grounding creates real accuracy, latency, and cost burdens, which Batra acknowledges rather than claiming browser agents are universally frictionless.
- Claim: The optimal computer-use system is hybrid: act visually when necessary, generate code when useful, then visually verify the outcome. | Evidence: Yutori's newer Navigator can write JavaScript on demand; the demonstration fills multiple form fields simultaneously through an executed function. Batra describes the screenshot as a built-in form of outcome verification after code execution. | Implication: Separate an agent's universal control plane from its accelerators: permit DOM/JS shortcuts, but treat them as optimizations beneath a visual validation layer. | Caveat: The transcript supplies a product demo, not a detailed account of safeguards for generated JavaScript, session handling, or irreversible actions.
- Claim: Computer-use agents can solve workflows that have no practical structured integration path, such as validating a conditional e-commerce discount. | Evidence: The cited task requires finding a specified product, adding it to cart, applying a discount code subject to product/date/minimum-cart constraints, and checking whether the promised 22% discount is actually realized; Batra estimates such trajectories may require 20–40 steps. | Implication: Prioritize browser agents for exception-heavy, multi-step workflows where the business value comes from executing and validating the real user journey rather than retrieving a data field. | Caveat: He explicitly says these tasks are solvable in principle but still have practical accuracy gaps.
- Claim: Specialized computer-use models may now be good and cheap enough for scaled deployment, not merely demos. | Evidence: Batra reports Navigator N1.5 at 97% human-evaluated performance on Online Mind2Web, with 8 incorrect trajectories out of 300, and cites approximately $0.80 per 20–30-step browser task versus $2.30 for Opus 4.7 and GPT 5.5 on selected benchmarks. | Implication: Re-evaluate prior assumptions that browser agents are categorically too slow or costly, but run task-level production evaluations with measured success, recovery rate, latency, and all-in cost. | Caveat: The benchmark is acknowledged as imperfect and reportedly saturated; the comparisons are vendor-provided, task-specific, and omit key operational details such as failure recovery, compute configuration, and production reliability.
Detailed Brief
Architecture pattern: orchestrated, sandboxed browser workers
- Claims: Computer-use workers need not remain constrained to a single serial human-like session.; A higher-level orchestrator can dispatch multiple navigator agents simultaneously across separate web targets.
- Evidence: Batra describes an orchestrator launching multiple navigators in parallel, each with its own cloud sandbox instance.; He frames this as superhuman parallelism: many websites can be navigated concurrently in ways a single human cannot.
- Caveats: The talk does not address coordination conflicts, credential isolation, anti-bot defenses, rate limits, or the controls needed for financial, contractual, or other consequential actions.
- Implications: The strategic value of computer use is not only access to inaccessible interfaces; it can become a distributed execution layer for broad web research and operations.; Scaling this pattern requires operational governance as much as model quality.
What the talk does—and does not—claim
- Claims: Batra accepts the broader prediction that agents will increasingly drive web activity and agrees that structured endpoints are ultimately desirable.; His disagreement is with the timing and universality of infrastructure migration, not with APIs as a technical ideal.
- Evidence: He states that an endpoint accepting natural-language or programmatic task specifications is ultimately what higher-level systems want.; He argues that decades of layered, human-oriented web infrastructure will not be reinvented overnight, citing institutions that still use fax-like and manual processes.
- Caveats: The transcript ends while he is beginning to describe the desired higher-level endpoint model, so the proposed end-state architecture is incomplete.
- Implications: Treat browser control as an enduring compatibility layer, while selectively replacing it with direct integrations where economics, reliability, and permissions justify doing so.
Notable Concepts & Terms
- Long tail of the web: The vast population of smaller, legacy, and institutionally slow websites that lack reliable structured integrations and drive the case for computer use.
- Computer-use model: A model that observes screenshots and performs browser actions such as clicking, scrolling, and typing to operate existing web interfaces.
- Pixels as source of truth: The principle that rendered UI state, rather than initial HTML or inferred backend structure, is the authoritative signal for whether a browser task succeeded.
- Navigator / Navigator N1.5: Yutori's computer-use model, presented as progressing from screenshot-to-action control to a hybrid system that can also execute generated JavaScript.
- Online Mind2Web: The browser-agent benchmark used to support the claim that current computer-use performance has reached 97% human-evaluated success, though Batra says it is now saturated.
- Hybrid visual-and-code execution: The proposed agent strategy: use visual interaction for generality, JavaScript/DOM-level operations for speed, and screenshots for verification.
- MCP servers / Web MCP: Examples of proposed agent-access protocols that Batra considers unlikely to become a universal solution for the web's long tail.
Operator Notes / Why Ken Should Care
- Adopt a routing policy for web tasks: direct API first when available and authorized; browser/computer use when no stable integration exists; visual postcondition checks for both dynamic UI and high-value workflows.
- Prototype a hybrid browser-worker harness that gives agents controlled JavaScript execution but requires screenshot-based confirmation before reporting completion.
- Benchmark target workflows on end-to-end success, retry/recovery behavior, latency, and all-in cost—not HTML extraction accuracy or browser benchmark scores alone.
- Set explicit approval, credential-isolation, rate-limit, and audit controls before parallelizing browser agents across purchasing, submission, login, or external communication workflows.
- Monitor whether Yutori's reported 97% benchmark performance transfers to the specific long-tail sites and adversarial edge cases relevant to Ken; do not extrapolate directly from a saturated benchmark.
Source/Metadata
- Title: Computer-use models will agentify the web, not APIs — Dhruv Batra, Yutori
- Transcript words: 3422
- Duration seconds: 1259
- Timestamp note: No timestamps or chapters were present in the supplied transcript; the transcript also contains duplicated passages and ends mid-thought.
Transcript
Dhrou Bhattra Reviewer Reviewer Welcome. Let's get started. So my name is Dhrou Bhattra. I want to talk to you about an argument. There's an argument online that says, roughly speaking, AI agents will be the drivers of actions on the web, not humans, that there will be more agents clicking buttons, reading things, buying things for us than human eyeballs on the web. You drive one step deeper into the argument and you ask how this will happen because today the web is extremely hostile to automated traffic, and the answer usually is the web will be agentified. You drive one step deeper into what that means and how that will happen, and the answer usually is with APIs that agents will use, and this will call via some set of protocols, and the set of protocols grows over time. It was initially supposed to be MCP servers, then web MCP, and for payments there are 20 different competing protocols, and every single company wants to introduce their own protocol along the way. But okay, there will be APIs is the argument. My claim today and argument, and what I hope to convince you today, is that this last bit is wrong. I think the first two I generally agree with, this last bit that suddenly the web will provide you APIs for accessing things is just delusional. And my claim is computer use agents and computer use models will agentify the web, not APIs, and more specifically the long tail of the web. The head of the distribution, the most popular website, perhaps will give you the API, but the long tail will not. And the story starts usually with these sorts of use cases. I agree that they also frustrate me. When people show a use case of, find me business class flights from Miami to Palma de Mallorca for August, and somehow they expect that behind the scenes it will be a computer use agent clicking buttons on flights.google.com. I find this example, this is my own example, I'm not dunking on anybody else, I find this example bizarre. Why would you possibly do it this way? Aren't you aware that there are already aggregators? You send them a variable, they will send you a JSON. Your LLMs can do tool calling. Why would you click buttons? The purpose of those clicking buttons is to generate this behind the scenes. So we're on the same page there. The next step usually is, okay, let's roll this out to some other case. I want to know, does my favorite restaurant or does this restaurant have gluten-free items on the restaurant's menu? And so I assume that there's going to be a similar endpoint somewhere for that restaurant website or via an aggregator, oops, that accepts a query that I can filter through, where I can ask for, give me your menu items that are gluten-free. I want to, for those of you who can already see, I want to rid you of the delusion that such an endpoint exists. And I want to show you, just for the sake of being on the same page, what restaurant pages look like. There are three that are on the screen. This is the first one. This is what we imagine a prototypical, this is easy mode. You go to a web page, there's no API, but at least it's text. It's in German, sure, a different part of the world. Models can speak languages. I see text and I see pricings. Maybe this is easily scrapable. Okay. This is easy mode. The medium mode is a page like this, where of course the thing that you're looking for is slightly hidden away. You press the menu button, it actually takes you to a PDF. Okay, fine. If your agent has to do this, it needs a PDF reader mode, fine. I can highlight these. Okay, so this is released text, no problem. I will dump this into ChatGPT or whatever and it'll do it. Here's the hard mode. This is what a hard mode website for a restaurant looks like, where it's just pictures of the people who created the restaurant, of where they're located, of menu items. And you're like, where is the menu? Does anybody speak Spanish? Nuestra carta? See? Okay, let's check. That's our menu. You click there. Oh my God. What am I staring at? Okay, you click on it. Oh, it's a gallery of individual, what is this? This is pixelated, there's no text here. I am looking at JPEGs embedded in a gallery which contain the PDF items. You download this and you put it into ChatGPT and it's struggling doing OCR on this thing. Okay, so this is what the web looks like. You're telling me this will give you an API endpoint that you can pass gluten-free items and it'll tell you what that is. That's example number one. Let's take a look at another example. This is much more enterprise-focused. I am a business. I want to sell to a school district. There's about 15,000 to 20,000 school districts in the US. And I want to ask the question, is this school district where I give you mydistrict.gov, is this school district procuring a laptop right now? Right? It's a simple question. And I want a similar endpoint. I want to say RFP status open and topic laptop. And you think this will exist. Let me tell you what these websites look like. Again, this is what easy mode looks like. This is Ithaca Public Schools. Okay, there's some portal. Maybe enrollment is the wrong place to go. This is what slightly harder mode looks like. This is a different public school. You look at the menu under finance. There's purchasing. Under purchasing, there are certain solicitations. There's a lot of questions. Scan PDF again. Okay, very nice. So, no text in here. But this is how they tell you about what they have purchased. And for ultimate boss level, I want to show you a different school district where, in order to find information, you have to file a request for access of information under Freedom of Information Act. And then what they do is they scan your email that you sent to them. And they will scan that email, put it on a Google Drive, and then attach PDFs associated with your request. These are the people you're telling me will give you an MCP server? The amount of delusion here is off the chart. If there was ever a time to say go touch grass, I think this is it. Okay. So, this is what... My claim here is that the web we forget is massive. It is extremely big. The number of websites out there, the active websites are somewhere around 200 million. The total number of websites is sitting at a billion. Infrastructure changes very slowly. You can imagine, as an engineer, getting unfettered access to software systems and letting your favorite coding agent rip and generate an API endpoint. Even if that technological problem is solved, which I agree seems on the horizon, it should eventually be solved, you're not going to get unfettered access to these institutions. And these institutions change very slowly. There are still places that are faxing each other. You're not going to be able to suddenly change this overnight. Okay. So, at this point in time, you might be thinking, fine. I'm stuck with the infra as it is. So there's not going to be APIs available. But I have coding agents. Why don't I just throw them at the HTML? There is, after all, if the browser is doing it, it's a piece of code. My coding agent should be able to do it as well. Fair point. I also thought this way two years ago. Let me tell you what the web actually looks like. If I ask you the question, what was the final score of this game between Minnesota Timberwolves and the Brooklyn Nets? Here's what the web page looks like. This is a modern website. It's NBA.com. You and I can see the score, 125 versus 109. Okay. Behind the scenes, when you actually load the page and read it, this is what initially gets loaded. There's an empty placeholder initially. And you wait a few hundred milliseconds to a few seconds, depending on your network connection. And your browser makes an asynchronous call later to fetch the information. So the browser makes a call to an endpoint that it's extracting that information from. That endpoint responds with the JSON that contains the answer. If you just read the HTML when you load it, the answer is not in the HTML. And so your chatbot doesn't have access to that either. So okay, you go, fine, these are just details. I just have to add some waits and sleeps and sure. Okay, fair enough. I will take you to another example. I ask you the question on a product web page, is the 25mm Osmium Cube in stock or out of stock? You're on this product website. You scroll down. There is a drop-down. That drop-down is telling you and me as humans, three things are sold out. And you wait a few hundred milliseconds to a few seconds, depending on your network connection. And your browser makes an asynchronous call later to fetch the information. So, the browser makes a call to an endpoint that it's extracting that information from. That endpoint responds with the JSON that contains the answer. If you just read the HTML when you load it, the answer is not in the HTML. And so, your chatbot doesn't have access to that either. So, okay, you go, fine, these are just details. I just have to add some weights and sleeps, and sure. Okay, fair enough. I will take you to another example. I ask you the question on a product web page, is the 25mm Osmium Cube in stock or out of stock? You're on this product website. You scroll down. There is a drop-down. That drop-down is telling you and I as humans three things are sold out. One thing is in stock, even though it doesn't actually say in stock. There's no text there that says in stock. But you understand that grayed out means sold out. And sometimes the sold out won't actually be there as text. Sometimes it will just be unclickable, grayed out. Okay, surely this information must be in the code somewhere. You go and read the HTML, and it turns out there is an option selector. It actually doesn't say any of the things that I'm seeing on screen. It doesn't say sold out. It doesn't say available. So, what's going on? It turns out behind the scenes, the browser makes a call, gets a JSON object, which is the variable product count. That contains a variable called quantity. That quantity is 10 sometimes, zero sometimes. That's just how many things can the backend support right now. Some of those quantities are zero. And there's a different rendering script that, any time there's zero, grays it out and makes it unclickable. Fundamentally, what is happening here is this information that you are seeing on screen is not written somewhere as pure text. It is calculated. It is rendered. And for people who work in the browser industry, they understand this. But often, people who are coming from an AI background like me, we didn't always understand this. The browser is a rendering engine. You are seeing pixels on screen. Think of it as a game engine. And you're asking, can I not read the source code of the game and predict exactly what the pixels are going to be? Well, yes, eventually. But right now you're asking for an exact inversion of that process. Fundamentally, the web was built for human eyes. Pixels are the source of the truth because the consumers of the websites are humans. That is what it was built for. And so there is an implication here that we have to grapple with, which is machines will need vision to operate those things because the web was built for human consumption. In a way, this is the bitter lesson for web agents, that the more you end up writing scaffolds around existing websites, it doesn't actually generalize to the long tail of the web. The thing that generalizes is the thing that it was designed for, which is the most general solution, just pixels in. This is what we and some others in the area have been working on. We have a model called Navigator. The first version of the model went out in November last year. The first version acted purely like a human. Screenshot in, button clicks, and scrolls out. I will tell you in a second that is not what you should settle on, but it is a general solution. It lets you do things like this. I give you an e-commerce website, and I tell you there is this discount code. Please go. The discount code is applicable under certain constraints. Maybe it's only on a product, it's only on a certain set of dates. Maybe it's only with this minimum cart threshold. I describe that in natural language, and I tell you, tell me if this discount code is valid or not. There is no API for this. The way to do it is just like a human can. You go to that website, you find the product that is described, you add it to cart, you apply the discount code, and you check whether the claim of 22% off was met or not. And that is screenshot in, button click out. This trajectory took 20, 30, 40 steps, depending on the sophistication of the task. If you can do it on your browser, this model can accomplish it in principle. In practice, of course, there are accuracy gaps and so on. But in principle, this task is solvable. Whereas in a lot of earlier cases, even in principle, that task may not be solvable. So my claim is the web was built for human eyes, machines will need vision. But of course, they do not need to be limited to human ways. Just because for the long tail, you need to have a capability does not mean that is the only way you should do it. Here is an example showing that. The next version of the model that we trained can also write JavaScript on demand. So this is a Chrome extension. On the right, you see an action that says execute JS, value default, text select. On the left, you saw that the model filled out multiple form fields simultaneously. The reason why it could do that is because it wrote a little bit of a function. It can read the code when necessary. It can write code when necessary because the browser, after all, is an engine that can execute code. But it has a formal verification system built in. It is seeing the screenshot. That is the source of the truth. So it knows whether it succeeded or not. So click buttons when you have to, write code when you have to, and look at the result through pixels because that is the source of truth. And of course, because these things are machines, you can string them into multi-agent systems. So you can have an orchestrator that is launching multiple navigators in parallel, each with a cloud sandbox instance. They are clicking buttons on multiple websites. So you can accomplish things that would be superhuman because no human would be able to parallelize over that many instances. Around here, usually, in this flow of an argument, is when people start asking, are computer use models actually good enough for these tasks? There is actually a perception online that it's not clear whether progress on computer use has been fast, and there are questions about why progress has been slow. That's not the reality I am seeing, and that's not the reality that the numbers back up. This is a popular benchmark, online mind to web. No benchmark is perfect. The point isn't that this is the right solution. But on the x-axis are release times of different models. On the y-axis is performance, which is human eval on this. So a human went in, looked at the trajectory, and decided whether it was correct or not. And basically, this particular version of the benchmark is saturated. The last model that we just released, Navigator N1.5, is sitting at 97% human eval. Eight trajectories out of 300 are incorrect. At this point in time, you should just retire the benchmark, build something harder. There's about 30 to 50 steps of interaction that are happening. The next step is to go for something larger. So at least in numbers, what I'm seeing, we're seeing steady progress in computer use agents being able to do more and more things. And this is the model that I showed: pixels in, button clicks, and code out. Around this time, usually, is when people start asking questions. Okay, so they're getting better. But aren't these things slow? Because after all, you're looking at a screen, you're clicking a button, there are lots of buttons to click. And these things are expensive if you're running them for hundreds of things. That claim, I think, is largely true. There is some truth to it. But I think people forget how much you can optimize these things out. So these are our results compared to the frontier models. Opus 4.7, GPT 5.5, on a couple of browser use benchmarks. We're slightly better, but I think that's within statistical noise in terms of accuracy. That improvement, I wouldn't beat the drum on. What I would emphasize is latency per step and cost per task. In terms of latency, this is a smaller footprint model. That's why it's a lot faster than some of the trillion-parameter-plus models. And there are corresponding cost savings. So if you have something like, on these data sets, 20, 30 steps of interaction, you're looking at 80 cents per task versus $2.30. And that makes a big difference. And so the models are getting cheaper in that sense, that you can launch them at scale. So hopefully, at a high level, I've convinced you that there is something off with this argument. This is what my goal was. This is where I started. AI agents are going to be the primary drivers of action. That seems an uncontestable statement because the underlying intelligence of the models is becoming larger and larger. There is a certain gain of productivity and efficiency and just ease of life that you get. In terms of latency, this is a smaller footprint model. That's why it's a lot faster than some of the trillion-parameter-plus models. And there are corresponding cost savings. So if you have something like, on these data sets, 20, 30 steps of interaction, you're looking at 80 cents per task versus $2.30. And that makes a big difference. And so the models are getting cheaper in the sense that you can launch them at scale. So hopefully, at a high level, I've convinced you that there is something off with this argument. This is what my goal was. This is where I started. AI agents are going to be the primary drivers of action. That seems an uncontestable statement because the underlying intelligence of the models is becoming larger and larger. There is a certain gain of productivity and efficiency and just ease of life that you get. So it makes sense. So I think this hypothesis that suddenly, overnight, 30 years of infrastructure that was built layer upon layer for human consumption will, in what, two, five, ten years, be reinvented is, I think, a fantasy. But this is ultimately what you want, right? Ultimately, you want an endpoint that some higher-level entity can go to, and I say, I want you to do X. There is some task. Maybe I give you that task description in natural language. Maybe I have some programmatic description with parameters.