AI Engineer

Bringing agents onto the world wide web — Paul Klein IV, Browserbase

2032 summary words 9 min summary Watch video

Start with the signal

9 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Reliable web agents are now constrained less by model capability than by the engineering harness around them: multimodal execution, reusable skills and memory, deterministic browser infrastructure, agent-native web interfaces, authentication, trust, and observability.
  • Why it matters: This is a concrete architecture and market thesis for turning browser automation from fragile demos into production agent infrastructure, directly relevant to AI operations, control planes, OpenClaw-style deployment, security, and workflow design.
  • Best use: Use the talk to pressure-test the browser layer of Ken's agent systems and identify which capabilities should be built internally versus sourced from a specialized platform.

Executive Summary

Paul Klein argues that browser and computer-use agents have reached an inflection point: models have improved enough at long-horizon tasks and interface use that waiting for a better base model is no longer the main strategy. The remaining gap is a large “capabilities overhang” caused by weak agent harnesses and unreliable operating environments. His claim is that domain-specific scaffolding can extract materially better outcomes from the same underlying model.

For web automation, Klein's proposed stack has three layers. First, agents should be multimodal and use the cheapest reliable execution method rather than blindly clicking through a browser; that can mean visual interaction, DOM/accessibility signals, network interception, or a coding agent that writes and replays an API-like request script. Second, the harness should accumulate skills, memory, and compressed task-specific context so the agent does not repeatedly rediscover a site and flood the model with irrelevant page tokens. Third, production infrastructure must provide consistent rendering, scale, and session behavior.

He broadens the problem beyond an individual agent stack: the web itself needs agent-readable interfaces, secure agent authentication, and a trust framework that distinguishes authorized agents from malicious bots. He highlights accessibility metadata, WebMCP, website-published agent instructions, service-account or delegated-access patterns, and emerging agent identity approaches as pieces of the solution, while acknowledging that authentication and trust are unresolved industry problems.

The presentation is also a Browserbase product thesis. Klein positions Browserbase as the managed platform that supplies browser runtime, sandboxing, models, observability, scaling, and optimization loops, allowing builders to treat web automation as a sub-agent rather than rebuild it. The strongest non-promotional takeaway is architectural: web agents should be operated as observable, learning systems with a controlled execution environment—not as one-off prompts driving an arbitrary browser.

Key Takeaways

  • Claim: The principal blocker to production web agents has shifted from model intelligence to harness engineering and execution reliability. | Evidence: Klein says computer-use models have improved through major recent investment in reinforcement-learning environments and human trajectories, and contrasts this with coding, where custom harnesses such as Cursor and Factory can outperform a baseline model using the same underlying model. | Implication: Ken should treat browser-agent quality as a systems-engineering problem—tool design, state, evaluation, recovery, and runtime control—not defer the roadmap pending another model release. | Caveat: He does not claim custom harnesses will permanently outperform models trained with durable RL; his narrower claim is that a harness improves present-day performance and should be measured against the model-only baseline.
  • Claim: Reliable web automation requires hybrid execution, not screen-clicking alone. | Evidence: Klein says effective production agents combine browser use with coding: they may inspect or intercept network requests, have a coding agent generate a script, and replay those requests when that is more context-efficient and repeatable than visual UI interaction. | Implication: Design browser agents with an execution router: prefer stable APIs or replayable network actions when available, use DOM/accessibility controls next, and reserve visual clicking for cases that require it. | Caveat: This approach depends on the task and site; browser interaction remains necessary where no stable request-level path exists or where the UI is the actual control surface.
  • Claim: Skills, memory, and context compression are necessary to make recurring web tasks reliable and economical. | Evidence: Browserbase's browse.sh is presented as a way to publish site-level skills before an agent visits a website; Klein argues that agents should reuse what they have learned and receive an optimized subset of relevant tokens rather than the entire page. | Implication: For recurring workflows, persist site/task playbooks, selectors or approved actions, error patterns, and successful traces; avoid presenting raw full-page HTML or screenshots by default. | Caveat: The transcript provides no independent performance numbers for browse.sh or a specific token-reduction benchmark.
  • Claim: Browser-agent infrastructure must be deterministic and enterprise-operable, not improvised from consumer machines. | Evidence: Klein uses the example of people running OpenClaw via home Mac minis to obtain Mac OS access and home-IP CAPTCHA handling, arguing that this does not translate into a SOC 2-compliant deployment supporting thousands of customer agents. He emphasizes fixed layout, viewport, rendering, and input/output consistency across runs. | Implication: Separate local prototyping from production architecture: standardize browser images, viewport and locale settings, network routes, identity/session handling, retries, and auditability before scaling workflows. | Caveat: A controlled runtime reduces variance but cannot eliminate website changes, bot defenses, third-party outages, or dynamic personalized content.
  • Claim: An agent-first web needs explicit machine-readable affordances, not just better perception models. | Evidence: Klein points to accessibility trees and ARIA tags as more useful than raw DOM/HTML alone; he also cites Chrome's WebMCP, under which sites can expose MCP servers in-page so an agent can invoke approved actions such as submitting a registration form. He mentions LLMs.txt, skills.md, and agents.md as related publishing patterns. | Implication: When Ken controls a web product, expose stable agent actions and semantic metadata rather than forcing agents to infer workflows from pixels; when consuming third-party sites, prefer officially exposed interfaces over UI reverse engineering. | Caveat: Adoption of these emerging conventions is uneven, and the transcript does not establish a settled standard or broad production availability.
  • Claim: Authentication and trusted agent identity—not task execution—may become the largest enterprise gating factors for web agents. | Evidence: Klein identifies insecure password sharing and constantly managed service-account permissions as inadequate patterns, calls for human approval of sensitive actions, cites WorkOS AuthMD as an emerging agent sign-up/account-discovery approach, and argues for a “VeriSign moment” that can distinguish legitimate agents from bad bots. | Implication: Do not deploy browser agents with broad standing credentials as the default. Build delegated, scoped, revocable access; require approval for consequential actions; and track the emerging identity and attestation layer as a strategic dependency. | Caveat: He explicitly says the industry has not solved trusted agent identity or web-wide good-agent verification; CAPTCHA remains an imperfect control rather than a durable trust mechanism.
  • Claim: Observability must feed a continuous-improvement loop because browser agents need to learn from actual execution traces. | Evidence: Klein calls for screen recordings, logs, and network activity for every agent run, then feeding those artifacts back into the agent to improve later attempts; he describes Browserbase's auto-browse as a multi-loop optimization feature. | Implication: Instrument browser workflows as production systems with replayable traces, outcome labels, failure taxonomy, and privacy-aware retention; use those records to improve skills and routing rather than relying on anecdotal debugging. | Caveat: Self-improvement based on traces requires governance: captured sessions can contain credentials, personal data, and unsafe actions, and optimization should not autonomously broaden permissions or change high-impact behavior.

Detailed Brief

Market thesis: computer use as AI diffusion into operational long-tail businesses

  • Claims: Klein views non-coding computer use as a larger opportunity than coding because much of the real economy still depends on employees operating legacy web applications and forms.; The target market is not limited to AI-native companies; browser agents can address operational work at organizations that lack APIs and run on legacy, bespoke, or lightly digitized systems.; He argues that builders should focus on customer workflows rather than reconstructing a browser-agent platform from scratch.
  • Evidence: He gives examples of a logistics company in Singapore, a bank in South Africa, and a lumber factory in Mexico, describing businesses built around PHP websites, forms, and humans clicking buttons.; Browserbase claims to see millions of browser sessions per month and positions that scale as a source of accumulated edge-case knowledge.
  • Caveats: The examples are illustrative rather than quantified case studies; the talk provides no deployment ROI, success-rate, labor-savings, or failure-rate data.; Legacy-web opportunities frequently carry the highest compliance, authorization, data-quality, and exception-handling burden.
  • Implications: Prioritize bounded, repetitive, high-volume workflows where a legacy browser interface is the only practical integration path.; Evaluate opportunity size by the operational friction removed and exception rate, not merely whether an agent can complete a happy-path browser demo.

What a platform-grade browser-agent control plane should provide

  • Claims: Klein's buying/building checklist is a scalable platform, model-agnostic execution, agent identity brokerage, and observability.; Model portability matters because browser agents should be able to move to improved or better-priced models without rebuilding the surrounding system.; A browser agent is best treated as a specialized sub-agent inside a wider agentic system, rather than necessarily the top-level orchestrator.
  • Evidence: He describes Browserbase Agents as packaging a runtime, harness, sandbox, code execution, fetch and search tools, and models behind a prompt-driven interface.; He states that a platform should support both one agent and thousands, despite substantially different operational requirements at those scales.
  • Caveats: This section is closely tied to Browserbase's commercial positioning, so its product-completeness claims should be validated through hands-on testing against Ken's workflows.; Model agnosticism alone does not ensure portability if prompts, skills, tool schemas, or evaluation baselines are tightly coupled to one model's behavior.
  • Implications: Require provider-neutral abstractions and a model-switching evaluation suite before committing browser workflows to a stack.; Benchmark any managed platform on deterministic execution, credential isolation, trace quality, error recovery, scale behavior, and access-control integration—not only task completion in a demo.

Notable Concepts & Terms

  • Capabilities overhang: Klein's label for the gap between what current models can likely accomplish and what deployed systems extract because surrounding engineering is inadequate.
  • Agent harness: The scaffolding around a model—tools, execution environment, context selection, memory, sub-agents, and feedback loops—that enables useful real-world action.
  • Hybrid browser/code execution: Using browser interaction alongside code generation, network inspection, and request replay to choose the most reliable and efficient automation path.
  • WebMCP: A Chrome-related concept Klein cites in which a website can expose MCP-based approved actions within a page, letting agents use explicit tools rather than infer UI operations.
  • Accessibility tree / ARIA tags: Semantic page metadata that can give agents more structured, action-relevant signals than raw HTML, screenshots, or broad DOM dumps.
  • Agent identity / WebBot auth: Emerging approaches to authenticate an agent's provenance and distinguish authorized automation from malicious bot traffic.
  • AuthMD: A WorkOS initiative cited as a mechanism for an agent to discover how to sign up for a website and obtain its own account.
  • Observability-driven self-improvement: Capturing recordings, logs, and network traces from agent runs and using them to improve later routing, skills, and execution behavior.

Operator Notes / Why Ken Should Care

  • Create a browser-agent reference architecture that explicitly supports three execution modes: approved/API or network-request actions, semantic DOM/accessibility actions, and visual computer use; define routing criteria and fallback rules.
  • Run a production-readiness review for any OpenClaw or browser-agent deployment: standardized runtime image, fixed viewport/locale, session isolation, credential vaulting, scoped permissions, approval gates, audit logs, and trace-retention controls.
  • Build a small benchmark suite from real recurring workflows, including page changes, login expiration, CAPTCHA/anti-bot responses, partial failures, and high-impact confirmation steps; compare model-only behavior against a skill- and memory-enabled harness.
  • For products Ken controls, add agent-facing affordances: strong accessibility semantics, stable action endpoints or MCP-style tools, documented permitted actions, and an agent-compatible authorization flow.
  • Monitor agent identity and delegated-auth standards rather than treating CAPTCHA bypass or password injection as a scalable long-term access strategy.
  • Evaluate Browserbase or alternatives as a build-versus-buy decision using controlled pilot workflows and independently measured success rate, recovery rate, cost per completed task, security controls, and model portability.

Source/Metadata

  • Title: Bringing agents onto the world wide web — Paul Klein IV, Browserbase
  • Transcript words: 4050
  • Duration seconds: 1105
  • Timestamp note: No timestamps or chapter markers were present in the supplied transcript.
Full transcript 3701 words · 18 min read
0:00

Paul Klein Reviewer Hello. Very sleepy crowd in the computer use room. Have we all given up at this point? What's going on? Thank you for coming in to my talk. My name is Paul Klein. I'm the founder of BrowserBase, and I'm going to talk about bringing agents onto the World Wide Web. If you're in this audience, in this track, you've done computer use, you tried operator when it came out, and you're probably, why isn't this happening yet? It seems obvious. Well, we'll address some of the high-level needs of computer use to really serve what I think is the largest category of AI agents, the agents that actually go out and do work on your behalf in the real world.

0:49

We'll talk through some of the technical stuff, but really trying to focus on there's a huge model capabilities overhang in this category specifically, and you all here can hopefully solve it. So thank you for coming. Well, of course, as we all know, the web wasn't built for agents. It was built for people, and that becomes very challenging as we're building systems to try and interact with it and automate it. So when we're thinking about building agents that interact with people for systems, we have to really wonder why it wasn't built for us, and why it struggles. When you've done any sort of automation in the past, you've run into so many roadblocks. The pages change.

1:27

The web was built in a very context-inefficient way. It's a lot of text, a lot of tokens. And when you're running any sort of browser agent or web agent right now, you get a broken browser that doesn't spin up. You have pages that don't work. You have blockers or other problems that really limit you. And I actually started my career doing web automation, maintaining these scripts every single day. It was very painful. So in my world, agents have made a huge advancement and allowed me to write durable web automation scripts, but we still haven't gotten to agents yet. And the question I want to ask is why. We're sitting here in this room, we're thinking,

2:00

computer use we saw a year ago. Has progress stalled? Why are web agents and browser agents not as big as they could be? And I think it really comes down to a few things. Until recently, the bottleneck was the models. The models one year ago really weren't good at long context horizon tasks. But that's clearly been solved in a major way. Models can do more and more complex tasks than ever. And, of course, in AI, you always have to update your priors. It's clear to me that anything I believed six months ago, I have to revisit every single week because these models are progressing at an insanely fast pace. So I don't think it's the models.

2:39

And especially, models are now much better at using interfaces. We've seen this kind of capabilities improvement in computer use models. In the last year, a lot of investment was made in RL environments for coding. And in the last six months, just as much investment has been made in RL environments for computer use. And computer use models are getting better. And you can see this in the evals. When you train things on human trajectories in our own environments that model our real world, the real web, you can make better models. So the models are getting there, I promise. But, okay, so the models are good. Why do agents still struggle to use the web?

3:15

What are the problems here? If it's a model problem, Dwarakish says if the models were good enough, diffusion would just happen. There's still a lot of work to be done here. And to me, it's no longer just the models. I'd argue that agents are missing the right harness and tools. And if you aren't familiar with an agent harness, if you haven't been on Twitter in the last few weeks, it's the scaffolding and systems around your model that enable it to actually interact with the world. A lot of talks in talking about harness engineering we're going to spend a lot of time on today. But I really think you can invest a lot in a harness and get a lot more out of the models

3:47

and extract that overhang out of the models. Karpathy actually tweeted this back in November 2023, and I thought it was just so forward-looking that this, what he described, these systems around an LLM, it's the harness. It's the tools that it can access. And if you fast-forward three years later, a lot of what we're doing every single day is building towards this: a code interpreter for the LLM, audio and video input like screenshots, a browser, and other LLMs as sub-agents. All of these principles have held true. So if you're ever wondering what should I build next, just go look at Karpathy's old talks. He's a pretty good predictor of the future.

4:19

And applying this to coding, we know that harnesses work really, really well with coding. On the graphic on the right, you can see Factory, when it compared to Cloud Code, using the same model but using their custom harness. And it turns out when you build a harness optimized for the domain that your agent is operating in, it can actually achieve above-model results in that domain. Harness engineering is a real thing. And I'd say Cursor actually started this. Cursor was the first one that was doing model engineering, or harness engineering, on top of the original LLMs. And a lot of what we've done at BrowserBase with browser models has been harness engineering.

5:00

And I think building a good harness is an engineering problem. You don't have to be a lab to build a good harness. And a lot of us in the room maybe aren't working at labs. Your company can make a great harness for your domain and actually improve model results. You don't just have to wait for the models to catch up. And once again, I believe the models are quite capable now. And if you look at this, you can see that it's not just Cursor. It's not just Factory. Many, many different types of companies are building coding harnesses on top of models and overperforming on the model's capabilities. Now, it's not clear yet if custom harnesses are going to beat out durable

5:32

RL models. But we're not going to debate that today. We know that adding a harness on a model improves results. Whether or not Cloud Code will be the best harness ever or not, I think that's a different conversation. But you should still have some sort of harness on your model and measure the performance versus the baseline model. And what I really want to get back to is that there is a massive capabilities overhang in computer use. The models are good enough, but we haven't done the engineering work to solve it. And I love this Greg Brockman tweet where he says, whenever I don't use codecs for a task, I ask myself why.

6:02

And it feels like the task is outside the capabilities of the model. The overhang is there. The actual work we can do is missing. And when you look at the amount of task completion you can get with coding, it's so much higher than CUA because we haven't actually really pushed the models far enough and given it the right tools. So to me, not only is this important because I think that non-coding is a much bigger opportunity than coding. If you look at this in recent slide, there are so many use cases that are in the non-coding domain that can benefit from computer use. It's a problem worth investing in.

6:39

And the wrong answer is to sit around and just wait for the models to get better. You can actually solve this today. Solving overhang is an engineering problem. And this is the work that we can do within our companies and within our agents, especially within the computer use domain, to build reliable web agents. And when I think about browser agents that work, it really comes down to three different types of things. They're multimodal, they're harness engineered, and they have reliable infrastructure. And I'll go through each of these three. First, they're multimodal. You no longer have to use a single model to actually interact with the task.

7:11

And we see this with coding agents all the time. Sometimes you'll use a smarter model for a more complex page. Sometimes a dumber model for a simpler page. And maybe you're using a combination of coding and computer use to actually power your agent. This is a really important insight. It turns out automating the web isn't always just clicking the button on the screen. It might be intercepting the network requests and writing a coding agent, or having a coding agent write a script to actually replay those network requests. The most reliable browser agents that we see in production right now are often writing code alongside using the browser to actually automate a task.

7:41

If you've done any sort of personal automation work in your life, you might see cloud code output a script more often than using cloud in Chrome. Because that's a very context-efficient way to automate a repeatable task. There's also harness engineering. It turns out, sure, we can write scripts and use models. But doing these things repeatedly, you want to benefit from things like memory and skills. Sometimes a dumber model for a simpler page. And maybe you're using a combination of coding and computer use to actually power your agent. This is a really important insight. It turns out automating the web isn't always just clicking the button on the screen.

8:22

It might be intercepting the network requests and writing a coding agent, or having a coding agent write a script to actually replay those network requests. The most reliable browser agents that we see in production right now are often writing code alongside using the browser to actually automate a task. If you've done any personal automation work in your life, you might see cloud code output a script more often than using cloud in Chrome. Because that's a very context-efficient way to automate a repeatable task. There's also harness engineering. It turns out, sure, we can write scripts and use models.

8:47

But doing these things repeatedly, you want to benefit from things like memory and skills. We launched something called browse.sh, which actually publishes skills for websites. So before your agent even goes to the website, it can observe what types of tasks it can do. WebMCP is very useful for this. It's a part of pulling in existing knowledge to optimize a website. Your agent doesn't have to discover something in the first place if it's done it before. It can use its memory and its skills to actually make it better. And you should think about trying to build skills into your agents.

9:20

If your agent is using CLIs to control websites, like the Playwright CLI, you can actually give it skills and context to be more effective there. And this results in much more optimized token usage. If you're throwing everything on the page to a model, you're going to get subpar results. And it's going to cost you a lot. The right harness should not only present the right tools, but present an optimized amount of tokens that are compressed to get exactly the right repeatable result every single time. And finally, the infrastructure here is extremely important.

9:46

Because when you're running browser agents in production, you want an environment that's going to work everywhere, every time. And I think a lot of work still needs to be done here. This is a lot of what our company does. Because computer use environments are pretty complex to scale up. It's funny. When OpenClaw came out, everyone started buying Mac minis, which to me feels like an infrastructure problem, right? You're running OpenClaw on a Mac mini in your house because that's the best way to run Mac OS that you can SSH into and then end up solving the CAPTCHAs because of your home IP address.

10:21

That is not something you can do when you're building thousands of agents for customers in production. I've yet to see a SOC 2-compliant Mac mini setup at scale. But please tell me afterwards if you found one. I'm very curious about it. The infrastructure problem that needs to be solved here is also an engineering problem. And most importantly, consistency in this infrastructure is important. When your agent is running across a website multiple times, you want it to see the same inputs and outputs, the same page layout, the same size.

10:55

If your infrastructure renders a page in a mobile layout one time and then in a desktop layout the second time, it's going to have inconsistent results. Consistency in the infrastructure is the nice base layer on top of your harness and on top of your models to actually get good results with this. I also think we have to improve the web itself. So there's a whole other side of this problem that's very interesting, which is how are we going to make it so the web works well with agents? And I think this is arguably the harder challenge because we're not just engineering on our own systems anymore.

11:25

We have to be evangelists to the web and to the broader world that, hey, you want agents to come to your website. So accessibility is the first thing I want to talk about. There's been a lot of really cool stuff here. Now, when you look at what best-in-class browser agents are doing, they're not just consuming the raw DOM and HTML of the page anymore. They're looking at subsections of that, like the accessibility tree, the ARIA tags. These are labeled components of a page that can help show your agent where it needs to click and why. Chrome just added WebMCP, which I think is really, really cool.

12:03

Websites can now publish MCP servers within their page that your agent can take advantage of without pre-installing the actual MCP. It can now issue tool calls to a website, like submit the registration form, in a way that's not only context-efficient, but is website-approved and blessed. More and more work can go into accessibility. And we've seen things like LLMs.txt, skills.md, agents.md, all being published alongside our websites. We need to see more of that to build the agent-first web. I think authentication is actually an even bigger problem here, too, because once your agent can actually go to a website, how can it log in on your behalf?

12:35

There's been a lot of different paradigms here. Maybe you're just giving your agent your password, but doing that securely can be very challenging. Maybe you're creating a service account for your agent where it has some limited access and you constantly have to give it new permissions. Authentication for agents is the next thing to be solved once you solve the harnessing capability problems. And doing that securely, where you can have a human loop approve certain actions on a website, is going to be a major challenge for unlocking computers for the enterprise. The biggest gate to building agents that actually work in prod is going to be the systems it has access to.

12:51

And authentication is something that needs to be solved in our industry to make this possible. I've seen a lot of really cool stuff come out. Work OS just launched AuthMD, which is a new way for your agent that goes to a website to find how to sign up on that website and get its own accounts. And if you're building software now, you should think about what does my agent-first sign-up and login flow look like? Because agents are going to be using your software whether you like it or not. It's best to let them use it securely. Finally, I want to talk about trust. The web was built to stop bad bots. But now there are good agents and bad bots.

13:30

How do we delineate between the two? And the CAPTCHA has been the tool in our tool chest for a very long time. But as we all know, CAPTCHAs are not as effective as we think against agents. And trying to identify these good agents is very important. There's been a lot of cool frameworks and work done on things like WebBot auth and more authenticated ways to say, this is my agent, it's coming from me, and you can follow me along on the web. But I still don't think we've solved the issue yet. And a big unlock to agents' access on the web, along with authentication, is actually how can we trust these agents?

13:55

And I think there needs to be almost a VeriSign moment for web agents where who can be the certificate issuer and say, my agent is trusted and this agent vendor is trusted. Nobody's come out and done that yet. I think those are really, really big opportunities. So building reliable browser agents is not a model problem. It's an engineering problem that all of us can solve. But doing that engineering is a full-time job. And if you are working in this space, I'd love to meet you. But if you aren't and you just want to build something that works, I might have a few ideas. You really don't have to reinvent the wheel here. There's been a lot of stuff happening.

14:30

And it's a consortium of companies that are continuing to push the world forward on what's possible when you want to automate the web. And I think there are a few things in here that are really important for a great solution, right? It has to be a scalable platform that serves your infrastructure needs. You can want to run one agent, but also thousands of agents. And the challenge is that those different levels of scale are very, very important. You want browser agents that are model-agnostic. As a developer, I don't want to be locked into a single model provider. As models continually change and get better, I want to be able to move my agent around.

14:57

That's why you need model-agnostic infrastructure. You need somebody to solve agent identity. Somebody who's going to go out and negotiate with the anti-bot providers of the world and say, we are the platform for trusted agents. And we are the ones that can help broker the access for your agents as you use the web. And finally, you need observability. When you're building these agents that go to any website in the world, you need to see where they're going and why, and how you can make sure that it's improving every iteration. Every agent you run should get better every single time. You need screen recordings, logs, network activity.

15:30

And you need to feed that back into your agent so it can self-improve. We published something called auto-browse earlier this year. That's a really interesting way to see how is my agent able to improve itself over multiple loops. And the feed-in of data to that from observability is extremely important to make your agents get better over time. And that's what we're building here at BrowserBase. we are the platform for trusted agents. And we are the ones that can help broker the access for your agents as you use the web. And finally, you need observability.

16:13

When you're building these agents that go to any website in the world, you need to see where they're going and why, and how you can make sure that it's improving every iteration. Every agent you run should get better every single time. You need screen recordings, logs, network activity. And you need to feed that back into your agent so it can self-improve. We published something called auto-browse earlier this year. That's a really interesting way to see how is my agent able to improve itself over multiple loops. And the feed-in of data to that from observability is extremely important to make your agents get better over time.

16:36

And that's what we're building here at BrowserBase. We power browser agents, web data extraction, and all these use cases across the entire web to make your agents work well. And what I've been extremely surprised by in building this company is the plethora of use cases. Of course, there are the large AI-native companies that use companies like BrowserBase to power their browser agents, but there's also all these little companies across the world that can benefit from automation. And my core belief with this company is that solving computer use accelerates the diffusion of AI to the real economy.

16:56

And as much as I love our bubble here in San Francisco, the real economy is companies like the logistics company in Singapore, the bank in South Africa, or the lumber factory in Mexico. These people are built on PHP websites with forums and human beings clicking buttons every single day. That's a huge opportunity for you to go solve, to build browser agents for them, and hopefully you can use the right infrastructure to power those things. And that's why we built BrowserBase agents, by the way. This is our new product we launched yesterday, because we want to give everyone a battery-included agent and harness for everything they need to automate the web.

17:16

The goal here is that you shouldn't reinvent the wheel. You shouldn't have to figure all this out and optimize it. You should benefit from the platform of scale that we've seen millions and millions of sessions every single month and understand how we've solved the edge cases for you so you don't have to solve them on your own. I have a quick little demo here. The way it works is instead of having to pull our tools together, you can actually put in a prompt, and we will stand up the harness, the runtime, the sandbox, the code execution, the fetch, the search tools, and the models to actually accomplish a task for you.

17:31

And what's beautiful is, as this agent is running, it's looking at its steps, and it's remembering what it can do, and learning from it so it can do them again in the future. The future for you is not having to reinvent the wheel every single time. It's actually being able to use an agent that's purpose-built for browsing the web and pull it in as a sub-agent of your larger agentic system. This is not the main thing you should be focusing your time on. You should be focusing your time on actually solving customer problems, not trying to rebuild the best-in-class browser agents. The optimization feature is quite cool.

17:51

It's going to look back and actually understand, hey, how can I do this better after looking back? This is this data feedback loop that I've talked about before, and I think it's what makes agents really, really special. I want to end with this last note. Based on the attendance in the room, I do think a lot of people have stepped back from computers because they've had so much challenges over the past year making browser agents work in production, but I can tell you firsthand from our customers we see it working.

18:01

And actually, I think one year from now, this room is going to be overfilled with people because the models are getting better, the techniques are getting better, the tools are getting better. It's just on us to build better things. Thank you all for having me today. I really appreciate it. Thank you.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note