AI Engineer

The agent-ready web: Simplify user actions with WebMCP — Tara Agyemang, Google

3686 summary words 16 min summary Watch video

Start with the signal

16 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: WebMCP is a proposed web standard from Google Chrome that lets websites define structured tools/capabilities (like a menu) for AI agents instead of forcing agents to parse DOM/accessibility trees and take screenshots, dramatically reducing token usage and improving reliability for agent-driven web interactions.
  • Why it matters: This is Google Chrome shipping a standard that could fundamentally change how agents interact with websites—moving from brittle screen-scraping (thousands of tokens per action) to structured, declarative APIs (like turning every site into an agent-friendly API).
  • Best use: Ken should understand WebMCP as a potential foundational shift in agent/web architecture; practical for any product building agent features that interact with websites, and critical for understanding Google's direction on agent tooling.

Executive Summary

Tara Agyemang from Google Chrome DevRel presents WebMCP (Web Model Context Protocol), an experimental web standard that enables websites to expose structured tool definitions for AI agents. The core problem: agents currently burn thousands of tokens parsing DOM trees, accessibility trees, and screenshots just to click a button or fill a form. WebMCP provides a declarative or imperative API so sites can register 'tools' (actions with JSON schemas) that agents can call directly, similar to how MCP works server-side but for client-side browser interactions.

The talk includes live demos: a maze game navigable only by agent tool calls, and a concert ticket site where an agent chains three tools (search concert, open page, purchase ticket) to complete a multi-step purchase from a single prompt. Tara emphasizes that WebMCP is complementary to existing web foundations—sites should first optimize semantic HTML, accessibility, and performance before layering in WebMCP tools. The standard is in early preview (Chrome 146+, Canary recommended), API is changing weekly, and Google is actively soliciting feedback.

WebMCP offers two implementation modes: declarative (add HTML attributes to forms, browser auto-generates tool schema) and imperative (register custom tools via JavaScript with manual schema, execute block, and return values). The imperative API is more common for complex multi-step flows. Tara stresses keeping UI in sync with agent tool calls and returning informative results so agents know what happened. The vision is user-agent hybrid workflows—users browse normally, hand off to agent for complex tasks, then resume manual control.

Key Takeaways

  • Claim: Current agent web interactions are token-heavy and brittle because agents parse entire DOM, accessibility tree, and screenshots just to perform simple actions like clicking a button. | Evidence: Tara's demo shows an agent attempting to buy concert tickets: it must parse HTML, analyze accessibility tree, take screenshots, measure click coordinates—'I don't even want to guess at how many tokens you probably just use trying to do this. It's probably a lot.' Then an ad loading can shift page layout and break the click. | Caveat: Tara does not provide exact token counts or failure rate data for traditional screen-scraping approaches, so the efficiency gain is qualitative. | Implication: For Ken: any agent product that scrapes websites today is burning tokens and facing reliability issues. WebMCP could cut costs and errors if sites adopt it, but adoption is the key risk. | Timestamp: 01:30
  • Claim: WebMCP allows websites to define structured tools (name, description, schema, execute function) that agents can discover and call directly, functioning as 'the USB-C of AI agent interactions'—a standard menu instead of guessing. | Evidence: The maze game demo shows tools appearing in the Model Context Tool Inspector extension: 'Start Maze Game' tool on the landing page, then 'move', 'look', 'pick up', 'drop', 'use' tools on the game page. Agent calls 'move' tool with direction parameters parsed from natural language prompts like 'R' for right. | Caveat: API is experimental and changing weekly ('the code that I've shown might be different next week'). Chrome 146+ only, requires experimental flags unless using Canary. No mention of Firefox/Safari support timeline. | Implication: Ken should treat this as early-stage but high-signal: Google is investing in a standard that could become the default agent/web interaction layer. Early adopters could gain agent-driven traffic; Ken's products could integrate WebMCP to reduce token costs and improve UX. | Timestamp: 05:00
  • Claim: WebMCP is client-side only (browser window must be open) and complements server-side MCP; it's 'inspired by MCP like JavaScript is inspired by Java'—implements the tools portion of MCP for in-browser agents. | Evidence: Tara clarifies: 'MCP enables AI agents to connect to applications on the server side... WebMCP is different... it's the implementation of the tools part of the MCP... allows engineers to provide tools to in-browser AI agents... all of the tools live in the browser.' | Caveat: No discussion of how WebMCP and server-side MCP interoperate, or whether a unified protocol is planned. Also no mention of security/privacy implications of agents executing arbitrary tools in the browser. | Implication: For Ken: WebMCP is narrow in scope (client-side only) but powerful for user-facing agent interactions. Products that combine server-side MCP (data access) with WebMCP (UI actions) could enable end-to-end agent workflows. | Timestamp: 10:45
  • Claim: The declarative API auto-generates tool schemas from standard HTML forms by adding attributes like 'tool-name' and 'tool-description'; the imperative API requires manual schema creation and JavaScript execute blocks for complex multi-step flows. | Evidence: Declarative example: add 'tool-name' and 'tool-description' attributes to a form, browser generates JSON schema using form fields as parameters. Imperative example: call 'registerTool()' with name, description, manual schema, and execute function that runs JavaScript to manipulate DOM and return results to the agent. | Caveat: Tara notes 'the imperative API is probably the one that's most used because people have more complex UI flows,' implying the declarative API is limited to simple forms. No guidance on when to use one vs. the other beyond 'standard form element' vs. 'something more complicated.' | Implication: Ken should assume most real-world implementations will require imperative API and custom tooling, not just form annotations. This means engineering effort to define and test tools, but also flexibility to expose any site capability as a structured tool. | Timestamp: 14:20
  • Claim: WebMCP enables hybrid user-agent workflows: users browse normally, hand off complex tasks to agents (e.g., 'complete the maze,' 'buy two VIP tickets to Summer Vibes Festival'), then resume manual control, with UI staying in sync throughout. | Evidence: Maze demo: user issues 'complete the maze' prompt, agent chains tool calls (move, look, pick up, use) until task is done. Concert demo: single prompt 'buy two VIP tickets to Summer Vibes Festival' triggers three chained tool calls (search concerts, open concert page, purchase ticket with quantity=2 and section=VIP), UI updates at each step (VIP selected, quantity=2), then suggests user manually complete checkout. | Caveat: Tara admits the 'complete the maze' prompt is inefficient ('sometimes you'll see it go backwards all the way to the start, and then go forwards again') and requires prompt refinement ('if you just say, the exit is in the bottom right corner, it'll be more efficient'). Also, the concert demo did not show actual payment, suggesting manual human approval is expected for sensitive actions. | Implication: For Ken: WebMCP is not full autonomy; it's human-in-the-loop automation for tedious multi-step tasks. Products should design for agent handoff points and keep UI transparent so users trust what the agent is doing. This could be a UX pattern for Ken's products: let agents handle filtering, data entry, navigation, but require human confirmation for high-stakes actions. | Timestamp: 06:30, 17:00
  • Claim: Before implementing WebMCP, sites should first optimize web foundations: semantic HTML, robust accessibility standards, page performance (Core Web Vitals), and good UX flows—'you're already halfway to getting an agent-ready website' by doing this. | Evidence: Tara states: 'making your site accessible for everyone makes it accessible to AI agents by default. So if you improve your semantic HTML, if you focus on robust accessibility standards, and if you improve your page performance... you're already halfway there.' | Caveat: No specific metrics or case studies showing how much accessibility/performance improvements alone help agents vs. WebMCP. This is presented as conventional wisdom rather than measured evidence. | Implication: Ken's products should prioritize accessibility and semantic HTML as a baseline—it benefits human users and agents alike. WebMCP is additive, not a replacement for good web engineering. This also means well-built sites may not need WebMCP immediately. | Timestamp: 03:00
  • Claim: WebMCP is in early preview (Chrome 146+), API is changing weekly, and Google is actively seeking feedback and bug reports to iterate toward a stable standard. | Evidence: Tara lists setup steps: Chrome 146+, recommend Canary to avoid enabling experimental flags in main browser, enable 'WebMCP testing flag' via chrome://flags, install Model Context Tool Inspector extension. She emphasizes: 'This API is very experimental. It will change. It has been changing over the past few weeks... we want people to try it out. We want feedback.' | Caveat: No timeline given for when WebMCP will exit preview or become stable. No mention of cross-browser support (Firefox, Safari). No discussion of governance (is this a W3C process, WHATWG, Chrome-only?). | Implication: For Ken: this is a bet on Google's vision for agent/web interaction. Early adopters get influence on the standard but risk API churn. Ken should monitor adoption signals (other browsers, developer uptake) before heavy investment, but prototyping now could yield early mover advantage if WebMCP becomes standard. | Timestamp: 19:00

Detailed Brief

Problem: Agents waste tokens and fail unpredictably on today's web

  • Claims: Agents currently parse entire DOM, accessibility tree, and take screenshots to understand web pages and perform actions; Token usage is extremely high for simple actions (Tara avoids estimating but implies thousands of tokens per interaction); Interactions are brittle: dynamic content (ads loading) can shift layout and break agent actions
  • Evidence: Demo scenario: agent trying to buy two concert tickets must parse HTML, analyze accessibility tree, screenshot page, measure click coordinates—'I don't even want to guess at how many tokens you probably just use'; Example failure mode: 'maybe your ad has loaded at the top of the page, pushed all your content down, and your AI agent couldn't even click the right place in the end'; Maze game is intentionally unbrowsable by humans, forcing agent-only interaction to demonstrate tool calling
  • Caveats: No quantitative token usage data or error rates provided; No comparison to other approaches (e.g., browser automation frameworks like Playwright with AI)
  • Implications: Ken's agent products that scrape websites are likely burning budget and facing reliability issues; Any agent interaction layer that reduces token usage by orders of magnitude could be a competitive moat; Sites that adopt WebMCP early could attract agent-driven traffic (analogous to mobile-friendly sites in early smartphone era)

Solution: WebMCP—structured tool definitions for websites

  • Claims: WebMCP is a proposed web standard that lets sites define capabilities as structured tools (JSON schemas with execute functions); Agents discover tools via browser APIs and call them directly instead of guessing from page structure; This is analogous to 'the USB-C of AI agent interactions'—a standard interface vs. proprietary adapters; WebMCP is client-side only (browser must be open), inspired by server-side MCP but implements only the 'tools' portion for in-browser agents
  • Evidence: Model Context Tool Inspector extension shows discovered tools in a side panel (e.g., 'Start Maze Game' tool with description and schema); Agent can call tools directly: 'registerTool()' function with name, description, schema, execute block; Live demos show multi-tool chaining: maze game (move, look, pick up, drop, use) and concert site (search concerts, open concert page, purchase ticket)
  • Caveats: API is experimental, changing weekly, Chrome 146+ only, requires Canary or experimental flags; No mention of Firefox/Safari support or W3C/WHATWG standardization process; Security/privacy implications not discussed (what prevents malicious tools? how are user credentials handled?); No discussion of tool versioning, deprecation, or compatibility across agent implementations
  • Implications: Ken should view this as Google Chrome's bet on the future of agent/web interaction—high signal but early stage; If WebMCP becomes standard, sites without it may be invisible to agents (like non-indexed sites in search); Ken's products could implement WebMCP to reduce costs and improve agent reliability, but should plan for API churn; Opportunity to build tooling/frameworks that simplify WebMCP implementation for common use cases (e.g., e-commerce, forms, dashboards)

Implementation: Declarative and Imperative APIs

  • Claims: Declarative API: add HTML attributes (tool-name, tool-description, agent-invoked) to standard forms; browser auto-generates JSON schema from form fields; Imperative API: call registerTool() with manual schema, execute function (JavaScript), and return values for complex multi-step flows; Imperative API is 'probably the one that's most used because people have more complex UI flows'; Tools should have descriptive descriptions so agents know when to call them; execute blocks should validate input, manipulate DOM, and return success/failure info to agent
  • Evidence: Declarative example: <form tool-name='exampleTool' tool-description='...'>, browser generates schema with form fields as parameters; Imperative example: registerTool({name, description, schema, execute: function(params) { / validate, create DOM nodes, return result / }}); Concert demo uses imperative API: three tools (searchConcerts, openConcertPage, purchaseTicket) registered separately, agent chains them based on single user prompt; Maze demo also uses imperative API: multiple tools for navigation and item management
  • Caveats: No guidance on when to use declarative vs. imperative beyond 'standard form' vs. 'more complicated'; No discussion of error handling, retries, or what happens when tools fail or return unexpected results; No mention of tool discoverability beyond Model Context Tool Inspector extension—how do agents in production find tools?; Best practices for schema design, tool naming, and description writing not covered in depth
  • Implications: Ken should assume most real-world implementations require imperative API and engineering effort; Tool design is critical: poor descriptions or schemas will cause agents to misuse tools or ignore them; Ken's products could offer WebMCP-as-a-service: auto-generate tools from site structure, provide testing/debugging, optimize descriptions for agent use; There's a UX design opportunity: define interaction patterns for agent handoff (e.g., 'complete this form for me' vs. 'filter products matching X')

User Experience: Hybrid human-agent workflows

  • Claims: WebMCP enables users to browse normally, hand off tasks to agents, and resume manual control at any point; UI should stay in sync with agent actions so users see what's happening; Agents can chain multiple tool calls to complete complex tasks from a single prompt (e.g., 'complete the maze,' 'buy two VIP tickets to Summer Vibes Festival'); For sensitive actions (e.g., checkout/payment), sites should require manual human approval
  • Evidence: Maze demo: user issues 'complete the maze' prompt, agent chains tool calls autonomously until task done (though inefficiently); Concert demo: single prompt triggers three tool calls (search, open page, purchase), UI updates at each step (VIP selected, quantity=2), then Tara suggests manual checkout: 'in real life... you'll probably want your user to manually do that step so they know that they're spending real money'; Model Context Tool Inspector shows tool calls and results in real-time so users can follow agent's reasoning
  • Caveats: Agent efficiency depends on prompt quality: 'complete the maze' is inefficient, 'the exit is in the bottom right corner' is better—implying users need to learn prompt engineering; No discussion of how to handle agent errors or undo actions (e.g., if agent buys wrong tickets); No mention of rate limiting, abuse prevention, or what happens if agent makes thousands of tool calls; Tara's demos are simple (maze, concert tickets); no examples of truly complex workflows (e.g., multi-page medical forms, tax filing)
  • Implications: Ken's products should design for hybrid workflows: identify high-friction tasks agents can automate, but keep humans in the loop for high-stakes decisions; Transparency is critical: show users what the agent is doing (like Model Context Tool Inspector does) to build trust; Prompt design matters: products may need to guide users toward effective prompts or auto-suggest better prompts based on context; This UX pattern (human→agent→human handoff) could become standard for complex web apps; Ken should explore how it fits his products

Adoption and Ecosystem

  • Claims: WebMCP is in early preview, API changing weekly, Google seeking feedback and bug reports; Setup requires Chrome 146+, Canary recommended, experimental flag enabled, Model Context Tool Inspector extension installed; Resources: Chrome DevRel blog post (sign up for docs, best practices, implementation details), GitHub repo with demos and evals CLI tool; Google wants developers to test WebMCP on their sites today and provide feedback to iterate toward stable standard
  • Evidence: Tara lists setup steps: 'WebMCP is enabled in Chrome version 146 upwards. I recommend using Chrome Canary... enable the WebMCP testing flag... install the model context tool inspector extension'; GitHub repo includes 6-7 demos (maze, concert tickets, etc.) and evals CLI tool for testing; Early preview program offers 'access to all of our initial documentation... information on best practices, all the extra implementation details'; No mention of production readiness timeline, cross-browser support, or W3C/WHATWG process
  • Caveats: No timeline for stable release or broader adoption; No discussion of browser compatibility (Firefox, Safari) or standards governance; No data on current adoption (how many sites have implemented WebMCP? which agents support it?); Security, privacy, and abuse prevention not addressed; No mention of how WebMCP interacts with existing web standards (ARIA, web components, etc.)
  • Implications: Ken should monitor adoption signals: if major sites (e-commerce, SaaS, gov) implement WebMCP, it's a strong signal to invest; Early adopters can influence the standard and gain first-mover advantage (agent-driven traffic, reduced support costs for agent users); Risk: if WebMCP remains Chrome-only or fails to gain traction, investment could be wasted; Opportunity: build tools/services that help sites adopt WebMCP (schema generators, testing frameworks, analytics on agent usage); Ken's products could experiment with WebMCP now to understand trade-offs, then scale if/when standard matures

Notable Concepts & Terms

  • WebMCP (Web Model Context Protocol): Proposed web standard from Google Chrome that lets websites define structured tools (JSON schemas + execute functions) for AI agents to discover and call, reducing token usage and improving reliability vs. screen-scraping. Client-side only (browser must be open), inspired by server-side MCP.
  • Declarative API vs. Imperative API: Two WebMCP implementation modes. Declarative: add HTML attributes to forms, browser auto-generates tool schema. Imperative: manually register tools via JavaScript with schema and execute block. Imperative is more common for complex multi-step flows.
  • Model Context Tool Inspector: Chrome extension built by Google Chrome DevRel that lists discovered WebMCP tools on a page, shows tool calls/results in real-time, and allows users to interact with agents via prompts or direct tool calls. Used for debugging WebMCP implementations.
  • registerTool() function: JavaScript function in WebMCP imperative API that registers a custom tool. Takes object with name, description, schema (JSON), and execute function (manipulates DOM, returns result to agent). Core mechanism for exposing site capabilities to agents.
  • USB-C of AI agent interactions: Metaphor used to describe WebMCP: a standard interface for agents to interact with websites, vs. proprietary/brittle screen-scraping. Implies interoperability, reduced friction, and broad adoption potential (like USB-C for hardware).
  • Agent-invoked attribute: Boolean HTML attribute in WebMCP declarative API that indicates whether a form was filled in by an agent or a human. Allows sites to track/audit agent usage and potentially apply different validation or rate limits.
  • Hybrid human-agent workflows: UX pattern where users browse normally, hand off complex tasks to agents (via prompts), and resume manual control at any point. WebMCP enables this by keeping UI in sync with agent tool calls and allowing seamless handoff.
  • Evals CLI tool: Command-line tool in the WebMCP GitHub repo for testing WebMCP implementations. Likely automates agent interactions to measure tool call success rates, latency, token usage, etc. No details provided in talk.

Operator Notes / Why Ken Should Care

  • WebMCP is Google's bet on structured agent/web interaction—if it becomes standard, it's foundational infrastructure for Ken's agent systems. Early prototyping now could yield competitive advantage.
  • Token cost reduction is the key economic argument: screen-scraping burns thousands of tokens per action; WebMCP likely cuts this by 10-100x. Ken should model cost savings for his products.
  • Security/privacy implications are entirely unaddressed in this talk. Ken should investigate: how are credentials handled? can malicious sites define dangerous tools? how is user consent managed?
  • Cross-browser support is a major question mark. If WebMCP remains Chrome-only, it's a niche feature. If Firefox/Safari adopt, it's a web standard. Ken should track this.
  • Hybrid workflows (human→agent→human) could be a key UX pattern for Ken's products. Consider: what tasks are high-friction enough for agent handoff? where do users need to stay in control?
  • WebMCP-as-a-service opportunity: auto-generate tools from site structure, provide testing/debugging, optimize descriptions for agent use, analytics on agent behavior. Potential product line for Ken.
  • Accessibility-first approach is emphasized: semantic HTML + accessibility standards benefit agents by default. Ken's products should prioritize this even without WebMCP.
  • Imperative API is the real implementation for complex flows; declarative API is limited to simple forms. Ken should plan for engineering effort to define/test custom tools.
  • API churn risk: Tara admits code may change weekly. Ken should prototype but avoid heavy production investment until API stabilizes.
  • Agent-driven traffic could be the next SEO: sites without WebMCP may be invisible to agents. Ken should consider how this affects GTM for products targeting agent users.

Watch Map

  • 00:00: Introduction and overview of WebMCP concept
  • 01:30: Problem statement: agents waste tokens parsing DOM/accessibility trees/screenshots
  • 03:00: Web foundations first: accessibility, semantic HTML, performance
  • 05:00: Live demo: Maze game navigable only by agent tool calls
  • 10:45: WebMCP vs. MCP: client-side vs. server-side, complementary
  • 12:30: Use cases: flight booking, product filtering, medical/financial forms
  • 14:20: APIs: declarative (HTML attributes) vs. imperative (registerTool)
  • 17:00: Live demo: concert ticket purchase via chained tool calls
  • 19:00: Setup instructions: Chrome 146+, Canary, experimental flags, extension
  • 20:30: Resources: blog post, GitHub repo with demos and evals CLI
  • 21:00: Wrap-up: agents are already using the web, WebMCP turns sites into high-performance APIs

Source/Metadata

  • Title: The agent-ready web: Simplify user actions with WebMCP — Tara Agyemang, Google
  • Transcript words: 5597
  • Duration seconds: 1293
  • Timestamp note: Timestamps estimated from video duration (1293s/21m33s) and transcript segment boundaries; some timestamps inferred from content flow rather than explicit chapter markers.
Full transcript 3190 words · 25 min read
0:14

SPEAKER_00

Hello, hello. Hello, can you hear me okay? Let's get started. So we are going to be talking a little bit about WebMCP. Has anybody, just out of curiosity, has anybody already played around with WebMCP? Only a few people. Okay, great. Those few people, you have a bit of a head start. But for everyone else, we'll be going into a bit more of the background, how it works, what it does. So my name is Tara. I am part of the Google Chrome team. I'm a developer relations engineer, and I'm here with a few of my colleagues from Google Chrome alongside the DeepMind team too. So we'd be really interested in talking to you afterwards around the DeepMind booth if you have thoughts around Web and AI and the intersection between the two. That is where my focus is these days. So let's get into it. The past few decades, we have been building the Web for human actions and human eyes, and we've been trying to optimize for that. But these days, it's not just humans that are using the Web. We have agents using the Web on human behalf too. And we are seeing an increasing number of agents using the Web. But the problem is the agents are having to do so much work to do simple actions on the sites that we've built. And just to give you a bit of an example of this, this is a website that I've biocoded. And it's a concert website for selling tickets for concerts. And we have Gemini and Chrome panel on the side here. And let's say you've come along to this website and you've typed this prompt, you want to buy two tickets to the Afrobeats festival, you've given it the details. The AI agent has to do so much work to make this happen. So it'll probably look at the HTML, because usually the agents will pass the entire DOM just to understand what's happening on your page. Then it will look into the accessibility tree just to understand the structure of your HTML page. Then maybe it'll take a screenshot of the page, analyze all the different elements that it couldn't see in the HTML and the accessibility tree. And then maybe it will measure how far down it needs to click, how far across, where the exact element that it needs to click. And then it'll click that element. And as you can see, this process is quite long. It can be brittle. And I don't even want to guess at how many tokens you probably just use trying to do this. It's probably a lot. And then after all that, maybe your ad has loaded at the top of the page, pushed all your content down, and your AI agent couldn't even click the right place in the end. So there's so much to think about. But before we go into this proposed web standard, it's worth mentioning that you can do so much by improving web foundations first. So making your site accessible for everyone makes it accessible to AI agents by default. So if you improve your semantic HTML, if you focus on robust accessibility standards, and if you improve your page performance, make it load really quickly, think about those core web vitals, and then improve really good user experience flows through your site, you're already halfway to getting an agent-ready website. And it's only once you have those in place, then it makes sense to start thinking about WebMCP. So if you're not already aware, the Web Model Context Protocol is a proposed web standard. And that gives you the ability to define your site's capabilities as structured tools for your AI agents to use. And so you might have heard references to this as the USB-C of AI agent interactions. And that's because instead of any agent guessing what your website does, you're kind of giving the AI agent a menu of tools that it can use and actions that it can take. And so because of this, we're seeing that WebMCP significantly improves the performance and the reliability of agents navigating your website. So let's see it in action. Hopefully, Gemini treats me well today. So this is the Maze Escape game built by our team in Chrome DevRel. And just on the side here, we have a Chrome extension. I'll show you a link to that afterwards. But this is the Model Context Tool Inspector. And so we're using this. This is a standard Chrome extension that lives in your side panel. And it lists out all the tools that it finds on your website. So at the moment, it can only see one tool. And that's the Start Maze Game Tool. And then at the bottom down here, it gives you two options to interact with the page. So you can interact via a prompt, like a user would prompt normally via the AI agent. Or you can call tools directly at the bottom, but we won't be looking at that one today. So this specific Maze Game is actually more unique in that you actually can't browse it by clicking around the UI. You can only use this app with the AI tooling. So let's start a new Maze Game here. You can also choose your model on the side. So let's stick with the Gemini 2.5. So you'll see that at the bottom, when you send a prompt, it gives you all the information. So the new prompt to start a new Maze Game. And the AI agent, Gemini in our case, has called that tool Start Game. The tool itself has returned this information. And then the AI has read that and given me this response. And so now we have our Maze. And you'll notice that on this page, we have a bunch of new tools in the scope of this page, whereas the previous page, it only had that one tool. This page, we've got a bunch of tools to help us navigate the Maze. So in this Maze, you can move around with the north, south, east, west directions. You can look to see where you are in the Maze and which directions are open. And then you can pick up items, drop items, use items as you navigate this maze. And if I pop in some prompts, I can see that I can move down. So maybe after that, then right. So I'm going to move down and right. The AI agent should use my prompt, match it to the specific tools. So in this case, the move tool, it's taken my direction of down and right, match that to the north, south, east direction, and send that off to the tool that we have registered on this page. And then it's moved it down and right. And so you can do, and because it's an AI agent, it can understand a whole bunch of different things. So I could just say, right, up, maybe right again. Let's try that. And so the AI agent has seen that R stands for right, mapped that to the direction, and then called the move tool with those information. And because it's an AI agent, it can just keep repeating the same tool calls until it thinks that it's done what needs to be done. So I could even say, complete the maze. And then the AI agent should use all the tools available to just keep moving around the maze, to pick up items, to use the items when it needs to, because it has all the information in the tools available. This specific prompt was not the most efficient. So sometimes you'll see it'll go backwards all the way to the start, and then go forwards again. But the more that you refine the prompt,

0:25

SPEAKER_00

And so the AI agent has seen that R stands for right, mapped that to the direction, and then called the move tool with that information. And because it's an AI agent, it can just keep repeating the same tool calls until it thinks that it's done what needs to be done. So I could say, complete the maze. And then the AI agent should use all the tools available to just keep moving around the maze, to pick up items, to use the items when it needs to, because it has all the information in the tools available. This specific prompt was not the most efficient. So sometimes you'll see it go backwards all the way to the start, and then go forwards again. But the more that you refine the prompt, the better the agent knows how to complete the maze in the most efficient way. For example, if you just say, the exit is in the bottom right corner, it'll be more efficient in its instructions to get to that direction. So I won't continue this because it can take quite a while to complete this maze. But if we go back to the slides here...

0:30

SPEAKER_00

So this is the model context tool inspector that I mentioned. So this is the web extension that our team in Chrome DevRel built. The QR code there is if you want to see where that is in the Chrome web store, but anyone can use that and grab it from the web store.

0:36

SPEAKER_00

So essentially, WebMCP unlocks this new approach to using the web, where your users don't have to spend a lot of time trying to figure out how to use more complicated sites. And they can figure out their own workflow. So they can choose to browse your website the normal way for a bit, then they can hand over control to their AI agent, and the AI agent takes steps on their behalf. And then your user can come in at any time to take control again and browse your site the way they normally would. And so that ability to simplify user journeys and make those user journeys for people easier has been a large part of the reason we've seen interest and excitement in this new standard.

0:41

SPEAKER_00

So I want to pause for a minute to address the question that some people have, which is what is the difference between WebMCP and MCP. But you can see them as being complementary to each other. Whereas MCP enables AI agents to connect to applications on the server side, and you'd need to set up your own server for the agent to access, and then the agent can access the information anywhere at any time. WebMCP is different in that it's inspired by MCP. I like to think of it as how JavaScript is inspired by Java. And in short, WebMCP is the implementation of the tools part of the MCP. And so WebMCP allows engineers to provide tools to in-browser AI agents. And it's very specific for the client side features. So you have to have your browser window open for WebMCP to work. And then you can use it to help your agent interact with the browser. So all of the tools live in the browser. But you can imagine this for quite a few different types of use cases. So imagine those websites that are really complicated, that have a lot of steps that a user needs to take, maybe like booking a flight, or filtering products on a normal shopping website, or filling in complicated medical forms or financial forms, or to trigger fixes that need to be hidden on a page. Or if you're like me, you're just on a normal shopping site. And you're trying to find the right black faux leather clutch bag that can fit your mobile phone in. And instead of going through all the little filters, you just want to ask your AI agent to do it for you. So these are examples where any user can ask whatever AI agent they are using to complete these things on their behalf. So the user doesn't have to manually do this. And they don't have to fill in each input, they don't have to select each checkbox. And using WebMCP in these cases can mean that you can make those actions much easier for users. So let's look at the APIs. WebMCP proposes two approaches for implementation. So you've got the declarative API and the imperative API. Let's start with the declarative API.

0:48

SPEAKER_00

So if you have a normal HTML form, you can just add a few attributes to the HTML to get this to work. So we've got the tool name and tool description here. And then your browser will automatically generate a JSON schema that the agent can use to read using the form fields as parameters for the tool. So here's an example of what the JSON schema would look like for this form HTML. And there are a whole bunch of other attributes that can be used. So there's an agent invoked Boolean attribute. So you can tell whether your form was filled in by an agent or if it was filled in by a human. And there's a lot of more specific attributes that can be used for things like that too. But essentially, you want to use the declarative API when you have a standard form element. But when you have something more complicated, that's when you want to go back to the imperative API. So this is where you can register and define your own custom tools for when you have more complex, maybe multi-step UI flows. So here is an example. So at the bottom, we have this register tool function. And when you call register tool with an object like this, you need to manually create your own schema similar to the one that we had in the declarative API that was generated. You name your tool and give it the description. And you want to make sure you have really descriptive descriptions that enable the AI agent to know when it should be calling this tool. And then you have the execute block, which is essentially where you call normal JavaScript. Maybe you already have functions that you're using that you can call in here, maybe do a light wrapper. In this add to the item example, you can validate and trim text input, for example. And then you create the DOM elements or DOM nodes and add them to your page. And then you want to return some information to the AI agent so it knows what happened if everything happened successfully. So it can use that information for its next steps. So those are the two APIs. The imperative API is probably the one that's most used because people have more complex UI flows that they want the agent to complete.

0:54

SPEAKER_00

So if we go back to my Vibe Coded demo, I have added a few tools here. So we have a few featured events in the demo and then all of the events available down here. And then you can go in and purchase tickets on an individual concert page. So I have noticed that this works much better with Gemini 3.1. So I'm going to try that one.

1:18

SPEAKER_00

If we wanted to buy tickets to one of these festivals, let's buy tickets to the Summer Vibes Festival. Two VIP tickets because VIP only for me. Let's send that prompt. So the AI saw the tool search concerts, which it has called to find the specific concert via the concert name. And the tool returned the information about the concert, including the ID for that concert. Then it has called the second tool, open concert page with the concert ID. And that has opened this Summer Vibes Festival page. And then this new page has separate tools. This one here called purchase ticket. And it's called that in the third tool call here.

1:24

SPEAKER_00

If we wanted to buy tickets to one of these festivals, let's buy tickets to the Summer Vibes Festival. Let's see, two VIP tickets because VIP only for me. Let's send that prompt. So the AI saw the tool search concerts, which it has called to find the specific concert via the concert name. And the tool returned the information about the concert, including the ID for that concert. Then it has called the second tool, open concert page with the concert ID. And that has opened this Summer Vibes Festival page.

1:53

SPEAKER_00

And then this new page has separate tools. This one here called purchase ticket. And it's called that in the third tool call here, with a quantity two and the section name. And then I've got a little notification to say, oh, you bought your tickets. You spent £356. Great. I'll put that on the Google's credit card. But you can see as well, in each step, it's updated the UI to make sure the user can also see what's happening. So you always want to make sure that your UI is in sync with the tool calls that are happening. So we've got the VIP selected, we've got the quantity selected, and then in real life, it will go through to some checkout page.

2:23

SPEAKER_00

You'll probably want your user to manually do that step so they know that they're spending real money. Let's head back. Let's head back. So if you're interested in trying this out, it's probably worth just understanding the status of where we're at with WebMCP. So we're still in early preview stage. This API is very experimental. It will change. It has been changing over the past few weeks. And so the code that I've shown might be different next week. But that's because we want people to try it out. We want feedback. We want to know the best way to use this API. And if you're interested in doing that, these are a few steps to get set up.

3:00

SPEAKER_00

So WebMCP is enabled in Chrome version 146 upwards. I recommend using Chrome Canary just so you can keep things separate. Otherwise, in the normal Chrome, you have to enable experimental flags. And you might not want to do that on your normal browser. Once you have Chrome Canary, you'll need to enable the WebMCP testing flag by putting this flag in your URL. And then install the model context tool inspector extension from the Chrome Web Store that I mentioned earlier. Just so you can play around and debug and see what your tools are doing. Then, these are the two resources that I recommend taking a look at.

3:47

SPEAKER_00

So this is our main blog post that gives you information on the early preview program for WebMCP. So if you sign up there, you get access to all of our initial documentation. And you get extra information about the program, information on best practices, all the extra implementation details that you might want to use while you're testing it out. So we're going to have a link to the URL and all of the API information. So that is the first one. And the second one is the GitHub repository of all the tools. So we've got the inspector tool here. We've got all the demos.

4:58

SPEAKER_00

So you can see the maze demo code is live there for you to play around with. There's about six, seven different demos you can try out. And there's an evals CLI tool you can use to help you start testing your own sites in the WebMCP tools on your own sites today. So I mentioned we're still in early preview. That's because we're looking for feedback. So try it out. Let us know what you think. If you have any friction points. If you find any bugs, we'd love to know that so we can keep iterating on this API and eventually move on to the next stage and start getting WebMCP in front of more users.

6:09

SPEAKER_00

But to wrap up, AI agents are already using the web. We don't have to settle for these token heavy, brittle, screen scraping processes that we have today. Instead, we can use WebMCP tools to turn every website into a high performance API for agents and at the same time build incredible user experiences for the users of our sites. So now that you have the tools and the context, please give it a go and try making your website's agent ready today. Thank you very much. Thank you. which directions are open. And then you can pick up items, drop items, use items as you navigate this

7:02

SPEAKER_00

maze. And if I pop in some prompts, I can see that I can move down. So maybe after that, then write. So I'm going to move down and write. The AI agent should use my prompt, match it to the specific tools. So in this case, the move tool, it's taken my direction of down and right, match that to the north, south, east direction, and send that off to the tool that we have registered on this page. And then it's moved it down down and right. And so you can do, and because it's an AI agent, it can understand a whole bunch of different things. So I could just say, right, up, maybe right again. Let's try that.

7:52

SPEAKER_00

And so the AI agent has seen that R stands for right, mapped that to the direction, and then called the move tool with those information. And because it's an AI agent, it can just keep repeating the same tool calls until it thinks that it's done what needs to be done. So I could even say, complete the maze. And then the AI agent should use all the tools available to just keep moving around the maze, to pick up items, to use the items when it needs to, because it has all the information in the tools available. This specific prompt was not the most efficient. So sometimes you'll see it'll go backwards

8:32

SPEAKER_00

all the way to the start, and then go forwards again. But the more that you refine the prompt, the better the agent knows how to complete the maze in the most efficient way. For example, if you just say, the exit is in the bottom right corner, it'll be more efficient in its instructions to get to that direction. So I won't continue this, because it can take quite a while to complete this maze. But if we go back to the slides here...

9:08

SPEAKER_00

So this is the model context tool inspector that I mentioned. So this is the web extension that our team in Chrome DevRel built. The QR code there is if you want to see where that is in the Chrome web store, but anyone can use that and grab it from the web store. So essentially, WebMCP kind of unlocks this new approach to using the web, where your users don't have to spend a lot of time trying to figure out how to use more complicated sites. And they can figure out their own workflow. So they can choose to browse your website the normal way for a bit, then they can hand over control to their AI agent, and the AI agent takes steps on their behalf. And then

9:50

SPEAKER_00

your user can come in at any time to take control again and browse your site again the way they normally would. And so that ability to simplify user journeys and make those user journeys for people easier has been a large part of the reason we've seen interest and excitement in this new standard. So I want to pause for a minute just to address the question that some people have, that's what is the difference between WebMCP and MCP. But you can kind of see them as being complementary to each other. So whereas MCP enables AI agents to connect to applications on the server side,

10:33

SPEAKER_00

and you'd need to set up your own server for the agent to access, and then the agent can access the information anywhere at any time. WebMCP is different in that it's kind of inspired by MCP. I like to think of it as how JavaScript is inspired by Java. And that's in short, WebMCP is the implementation of the tools part of the MCP. And so WebMCP allows engineers to provide tools to in-browser AI agents. And it's very specific for the client side features. So you have to have your browser window open for WebMCP to work. And then you can use it to help your agent interact with the

11:19

SPEAKER_00

browser. So all of the tools live in the browser. But you can imagine this for quite a few different types of use cases. So imagine those websites that are really complicated, that have a lot of steps that a user needs to take, maybe like booking a flight, or filtering products on a normal shopping website, or filling in complicated medical forms or financial forms, or to trigger fixes that need to be hidden on a page, that are hidden on a page. Or if you're like me, you're just on a normal shopping site. And you're trying to find the right black faux leather clutch bag that can fit your mobile phone in. And instead of going through all the

12:06

SPEAKER_00

little filters, you just want to ask your AI agent to do it for you. So these are a bunch of examples where any user can ask whatever AI agent they are using to complete these things on their behalf. So the user doesn't have to manually do this. And they don't have to fill in each input, they don't have to select each checkbox. And using WebMCP in these cases can mean that you can make those actions much easier for users. So let's look at the APIs. WebMCP proposes two approaches for implementation. So you've got the declarative API and the imperative API. Let's start with the declarative API.

12:48

SPEAKER_00

So if you have a normal HTML form, you can just add a few attributes to the HTML to get this to work. So we've got the tool name and tool description here. And then your browser will automatically generate a JSON schema that the agent can use to read using the form fields as parameters for the tool. So here's an example of what the JSON schema would look like for this form HTML. And there are a whole bunch of other attributes that can be used. So there's like an agent invoked Boolean attribute. So you can tell whether your form was filled in by an agent or if it was filled in by a human. And there's lots of like more specific attributes that can be used for

13:33

SPEAKER_00

things like that too. But essentially, you want to use the declarative API when you have a standard form element. But when you have something more complicated, that's when you want to go back to the imperative API. So this is where you can register and define your own custom tools for when you have more complex, maybe multi-step UI flows. So here is an example. So at the bottom, we have this register tool function. And when you call register tool with an object like this, you need to manually create your own schema similar to the one that we had in the declarative API that was generated. You name your tool and give it the description. And you

14:18

SPEAKER_00

want to make sure you have really descriptive descriptions that enable the AI agent to know when it should be calling this tool. And then you have the execute block, which is essentially where you call normal JavaScript. Maybe you already have functions that you're using that you can call in here, maybe do a light wrapper. In this add to the item example, you can validate and trim text input, for example. And then you create the DOM elements or DOM nodes and add them to your page. And then you want to return some information to the AI agent so it knows what happened if everything happened successfully. So it

14:59

SPEAKER_00

can use that information for its next steps. So those are the two APIs. The imperative API is probably the one that's most used because people have more complex UI flows that it wants the agent to complete. So if we go back to my Vibe Coded demo, I have added a few tools here. So we have a few featured events in the demo and then all of the events available down here. And then you can go in and purchase tickets on an individual concert page.

15:48

SPEAKER_00

So I have noticed that this works much better with Gemini 3.1. So I'm going to try that one. If we wanted to buy tickets to one of these festivals, let's buy tickets to the Summer Vibes Festival.

16:11

SPEAKER_00

Let's see, two VIP tickets because VIP only for me.

16:21

SPEAKER_00

Let's send that prompt. So the AI saw the tool search concerts, which it has called to find the specific concert via the concert name. And the tool returned the information about the concert, including the ID for that concert. Then it has called the second tool, open concert page with the concert ID. And that has opened this Summer Vibes Festival page. And then this new page has separate tools. This one here called purchase ticket. And it's called that in the third tool call here, with a quantity two and the section name. And then I've got a little notification to say, oh, you bought your tickets. You spent £356. Great. I'll put that on the Google's credit card.

17:16

SPEAKER_00

But you can see as well, like in each step, it's updated the UI to make sure the user can also see what's happening. So you always want to make sure that your UI is in sync with the tool calls that are happening. So we've got the VIP selected, we've got the quantity selected, and then in real life, it will go through to some checkout page. You'll probably want your user to manually do that step so they know that they're spending real money.

17:44

SPEAKER_00

Let's head back. Let's head back. So if you're interested in trying this out, it's probably worth just understanding the status of where we're at with WebMCP. So we're still in early preview stage. This API is very experimental. It will change. It has been changing over the past few weeks. And so the code that I've shown might be different next week. But that's because we want people to try it out. We want feedback. We want to know the best way to use this API. And if you're interested in doing that, these are a few steps to get set up. So WebMCP is enabled in Chrome version 146 upwards. I recommend using Chrome Canary just so you can keep things separate.

18:32

SPEAKER_00

Otherwise, in the normal Chrome, you have to enable experimental flags. And you might not want to do that on your normal browser. Once you have Chrome Canary, you'll need to enable the WebMCP testing flag by putting this flag in your URL. And then install the model context tool inspector extension from the Chrome Web Store that I mentioned earlier. Just so you can play around and debug and see what your tools are doing.

19:04

SPEAKER_00

Then, these are the two resources that I recommend taking a look at. So this is our main blog post that gives you information on the early preview program for WebMCP. So if you sign up there, you get access to all of our initial documentation. And you get extra information about the program, information on best practices, all the extra implementation details that you might want to use while you're testing it out. So we're going to have a link to the URL and all of the API information. So that is the first one. And the second one is the GitHub repository of all the tools. So we've got the inspector tool here. We've got all the demos.

19:45

SPEAKER_00

So you can see the maze demo code is live there for you to play around with. There's about six, seven different demos you can try out. And there's an evals CLI tool you can use to help you start testing your own sites in the WebMCP tools on your own sites today.

20:05

SPEAKER_00

So I mentioned we're still in early preview. That's because we're looking for feedback. So try it out. Let us know what you think. If you have any friction points. If you find any bugs, we'd love to know that so we can keep iterating on this API and eventually move on to the next stage and start getting WebMCP in front of more users. But to wrap up, AI agents are already using the web. We don't have to settle for these token heavy, brittle, screen scraping processes that we have today. Instead, we can use WebMCP tools to turn every website into a high performance API for agents and at the same time build incredible user experiences for the users of our sites.

21:01

SPEAKER_00

So now that you have the tools and the context, please give it a go and try making your website's agent ready today. Thank you very much. Thank you.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note