SPEAKER_00
Hello, hello. Hello, can you hear me okay? Let's get started. So we are going to be talking a little bit about WebMCP. Has anybody, just out of curiosity, has anybody already played around with WebMCP? Only a few people. Okay, great. Those few people, you have a bit of a head start. But for everyone else, we'll be going into a bit more of the background, how it works, what it does. So my name is Tara. I am part of the Google Chrome team. I'm a developer relations engineer, and I'm here with a few of my colleagues from Google Chrome alongside the DeepMind team too. So we'd be really interested in talking to you afterwards around the DeepMind booth if you have thoughts around Web and AI and the intersection between the two. That is where my focus is these days. So let's get into it. The past few decades, we have been building the Web for human actions and human eyes, and we've been trying to optimize for that. But these days, it's not just humans that are using the Web. We have agents using the Web on human behalf too. And we are seeing an increasing number of agents using the Web. But the problem is the agents are having to do so much work to do simple actions on the sites that we've built. And just to give you a bit of an example of this, this is a website that I've biocoded. And it's a concert website for selling tickets for concerts. And we have Gemini and Chrome panel on the side here. And let's say you've come along to this website and you've typed this prompt, you want to buy two tickets to the Afrobeats festival, you've given it the details. The AI agent has to do so much work to make this happen. So it'll probably look at the HTML, because usually the agents will pass the entire DOM just to understand what's happening on your page. Then it will look into the accessibility tree just to understand the structure of your HTML page. Then maybe it'll take a screenshot of the page, analyze all the different elements that it couldn't see in the HTML and the accessibility tree. And then maybe it will measure how far down it needs to click, how far across, where the exact element that it needs to click. And then it'll click that element. And as you can see, this process is quite long. It can be brittle. And I don't even want to guess at how many tokens you probably just use trying to do this. It's probably a lot. And then after all that, maybe your ad has loaded at the top of the page, pushed all your content down, and your AI agent couldn't even click the right place in the end. So there's so much to think about. But before we go into this proposed web standard, it's worth mentioning that you can do so much by improving web foundations first. So making your site accessible for everyone makes it accessible to AI agents by default. So if you improve your semantic HTML, if you focus on robust accessibility standards, and if you improve your page performance, make it load really quickly, think about those core web vitals, and then improve really good user experience flows through your site, you're already halfway to getting an agent-ready website. And it's only once you have those in place, then it makes sense to start thinking about WebMCP. So if you're not already aware, the Web Model Context Protocol is a proposed web standard. And that gives you the ability to define your site's capabilities as structured tools for your AI agents to use. And so you might have heard references to this as the USB-C of AI agent interactions. And that's because instead of any agent guessing what your website does, you're kind of giving the AI agent a menu of tools that it can use and actions that it can take. And so because of this, we're seeing that WebMCP significantly improves the performance and the reliability of agents navigating your website. So let's see it in action. Hopefully, Gemini treats me well today. So this is the Maze Escape game built by our team in Chrome DevRel. And just on the side here, we have a Chrome extension. I'll show you a link to that afterwards. But this is the Model Context Tool Inspector. And so we're using this. This is a standard Chrome extension that lives in your side panel. And it lists out all the tools that it finds on your website. So at the moment, it can only see one tool. And that's the Start Maze Game Tool. And then at the bottom down here, it gives you two options to interact with the page. So you can interact via a prompt, like a user would prompt normally via the AI agent. Or you can call tools directly at the bottom, but we won't be looking at that one today. So this specific Maze Game is actually more unique in that you actually can't browse it by clicking around the UI. You can only use this app with the AI tooling. So let's start a new Maze Game here. You can also choose your model on the side. So let's stick with the Gemini 2.5. So you'll see that at the bottom, when you send a prompt, it gives you all the information. So the new prompt to start a new Maze Game. And the AI agent, Gemini in our case, has called that tool Start Game. The tool itself has returned this information. And then the AI has read that and given me this response. And so now we have our Maze. And you'll notice that on this page, we have a bunch of new tools in the scope of this page, whereas the previous page, it only had that one tool. This page, we've got a bunch of tools to help us navigate the Maze. So in this Maze, you can move around with the north, south, east, west directions. You can look to see where you are in the Maze and which directions are open. And then you can pick up items, drop items, use items as you navigate this maze. And if I pop in some prompts, I can see that I can move down. So maybe after that, then right. So I'm going to move down and right. The AI agent should use my prompt, match it to the specific tools. So in this case, the move tool, it's taken my direction of down and right, match that to the north, south, east direction, and send that off to the tool that we have registered on this page. And then it's moved it down and right. And so you can do, and because it's an AI agent, it can understand a whole bunch of different things. So I could just say, right, up, maybe right again. Let's try that. And so the AI agent has seen that R stands for right, mapped that to the direction, and then called the move tool with those information. And because it's an AI agent, it can just keep repeating the same tool calls until it thinks that it's done what needs to be done. So I could even say, complete the maze. And then the AI agent should use all the tools available to just keep moving around the maze, to pick up items, to use the items when it needs to, because it has all the information in the tools available. This specific prompt was not the most efficient. So sometimes you'll see it'll go backwards all the way to the start, and then go forwards again. But the more that you refine the prompt,
SPEAKER_00
And so the AI agent has seen that R stands for right, mapped that to the direction, and then called the move tool with that information. And because it's an AI agent, it can just keep repeating the same tool calls until it thinks that it's done what needs to be done. So I could say, complete the maze. And then the AI agent should use all the tools available to just keep moving around the maze, to pick up items, to use the items when it needs to, because it has all the information in the tools available. This specific prompt was not the most efficient. So sometimes you'll see it go backwards all the way to the start, and then go forwards again. But the more that you refine the prompt, the better the agent knows how to complete the maze in the most efficient way. For example, if you just say, the exit is in the bottom right corner, it'll be more efficient in its instructions to get to that direction. So I won't continue this because it can take quite a while to complete this maze. But if we go back to the slides here...
SPEAKER_00
So this is the model context tool inspector that I mentioned. So this is the web extension that our team in Chrome DevRel built. The QR code there is if you want to see where that is in the Chrome web store, but anyone can use that and grab it from the web store.
SPEAKER_00
So essentially, WebMCP unlocks this new approach to using the web, where your users don't have to spend a lot of time trying to figure out how to use more complicated sites. And they can figure out their own workflow. So they can choose to browse your website the normal way for a bit, then they can hand over control to their AI agent, and the AI agent takes steps on their behalf. And then your user can come in at any time to take control again and browse your site the way they normally would. And so that ability to simplify user journeys and make those user journeys for people easier has been a large part of the reason we've seen interest and excitement in this new standard.
SPEAKER_00
So I want to pause for a minute to address the question that some people have, which is what is the difference between WebMCP and MCP. But you can see them as being complementary to each other. Whereas MCP enables AI agents to connect to applications on the server side, and you'd need to set up your own server for the agent to access, and then the agent can access the information anywhere at any time. WebMCP is different in that it's inspired by MCP. I like to think of it as how JavaScript is inspired by Java. And in short, WebMCP is the implementation of the tools part of the MCP. And so WebMCP allows engineers to provide tools to in-browser AI agents. And it's very specific for the client side features. So you have to have your browser window open for WebMCP to work. And then you can use it to help your agent interact with the browser. So all of the tools live in the browser. But you can imagine this for quite a few different types of use cases. So imagine those websites that are really complicated, that have a lot of steps that a user needs to take, maybe like booking a flight, or filtering products on a normal shopping website, or filling in complicated medical forms or financial forms, or to trigger fixes that need to be hidden on a page. Or if you're like me, you're just on a normal shopping site. And you're trying to find the right black faux leather clutch bag that can fit your mobile phone in. And instead of going through all the little filters, you just want to ask your AI agent to do it for you. So these are examples where any user can ask whatever AI agent they are using to complete these things on their behalf. So the user doesn't have to manually do this. And they don't have to fill in each input, they don't have to select each checkbox. And using WebMCP in these cases can mean that you can make those actions much easier for users. So let's look at the APIs. WebMCP proposes two approaches for implementation. So you've got the declarative API and the imperative API. Let's start with the declarative API.
SPEAKER_00
So if you have a normal HTML form, you can just add a few attributes to the HTML to get this to work. So we've got the tool name and tool description here. And then your browser will automatically generate a JSON schema that the agent can use to read using the form fields as parameters for the tool. So here's an example of what the JSON schema would look like for this form HTML. And there are a whole bunch of other attributes that can be used. So there's an agent invoked Boolean attribute. So you can tell whether your form was filled in by an agent or if it was filled in by a human. And there's a lot of more specific attributes that can be used for things like that too. But essentially, you want to use the declarative API when you have a standard form element. But when you have something more complicated, that's when you want to go back to the imperative API. So this is where you can register and define your own custom tools for when you have more complex, maybe multi-step UI flows. So here is an example. So at the bottom, we have this register tool function. And when you call register tool with an object like this, you need to manually create your own schema similar to the one that we had in the declarative API that was generated. You name your tool and give it the description. And you want to make sure you have really descriptive descriptions that enable the AI agent to know when it should be calling this tool. And then you have the execute block, which is essentially where you call normal JavaScript. Maybe you already have functions that you're using that you can call in here, maybe do a light wrapper. In this add to the item example, you can validate and trim text input, for example. And then you create the DOM elements or DOM nodes and add them to your page. And then you want to return some information to the AI agent so it knows what happened if everything happened successfully. So it can use that information for its next steps. So those are the two APIs. The imperative API is probably the one that's most used because people have more complex UI flows that they want the agent to complete.
SPEAKER_00
So if we go back to my Vibe Coded demo, I have added a few tools here. So we have a few featured events in the demo and then all of the events available down here. And then you can go in and purchase tickets on an individual concert page. So I have noticed that this works much better with Gemini 3.1. So I'm going to try that one.
SPEAKER_00
If we wanted to buy tickets to one of these festivals, let's buy tickets to the Summer Vibes Festival. Two VIP tickets because VIP only for me. Let's send that prompt. So the AI saw the tool search concerts, which it has called to find the specific concert via the concert name. And the tool returned the information about the concert, including the ID for that concert. Then it has called the second tool, open concert page with the concert ID. And that has opened this Summer Vibes Festival page. And then this new page has separate tools. This one here called purchase ticket. And it's called that in the third tool call here.
SPEAKER_00
If we wanted to buy tickets to one of these festivals, let's buy tickets to the Summer Vibes Festival. Let's see, two VIP tickets because VIP only for me. Let's send that prompt. So the AI saw the tool search concerts, which it has called to find the specific concert via the concert name. And the tool returned the information about the concert, including the ID for that concert. Then it has called the second tool, open concert page with the concert ID. And that has opened this Summer Vibes Festival page.
SPEAKER_00
And then this new page has separate tools. This one here called purchase ticket. And it's called that in the third tool call here, with a quantity two and the section name. And then I've got a little notification to say, oh, you bought your tickets. You spent £356. Great. I'll put that on the Google's credit card. But you can see as well, in each step, it's updated the UI to make sure the user can also see what's happening. So you always want to make sure that your UI is in sync with the tool calls that are happening. So we've got the VIP selected, we've got the quantity selected, and then in real life, it will go through to some checkout page.
SPEAKER_00
You'll probably want your user to manually do that step so they know that they're spending real money. Let's head back. Let's head back. So if you're interested in trying this out, it's probably worth just understanding the status of where we're at with WebMCP. So we're still in early preview stage. This API is very experimental. It will change. It has been changing over the past few weeks. And so the code that I've shown might be different next week. But that's because we want people to try it out. We want feedback. We want to know the best way to use this API. And if you're interested in doing that, these are a few steps to get set up.
SPEAKER_00
So WebMCP is enabled in Chrome version 146 upwards. I recommend using Chrome Canary just so you can keep things separate. Otherwise, in the normal Chrome, you have to enable experimental flags. And you might not want to do that on your normal browser. Once you have Chrome Canary, you'll need to enable the WebMCP testing flag by putting this flag in your URL. And then install the model context tool inspector extension from the Chrome Web Store that I mentioned earlier. Just so you can play around and debug and see what your tools are doing. Then, these are the two resources that I recommend taking a look at.
SPEAKER_00
So this is our main blog post that gives you information on the early preview program for WebMCP. So if you sign up there, you get access to all of our initial documentation. And you get extra information about the program, information on best practices, all the extra implementation details that you might want to use while you're testing it out. So we're going to have a link to the URL and all of the API information. So that is the first one. And the second one is the GitHub repository of all the tools. So we've got the inspector tool here. We've got all the demos.
SPEAKER_00
So you can see the maze demo code is live there for you to play around with. There's about six, seven different demos you can try out. And there's an evals CLI tool you can use to help you start testing your own sites in the WebMCP tools on your own sites today. So I mentioned we're still in early preview. That's because we're looking for feedback. So try it out. Let us know what you think. If you have any friction points. If you find any bugs, we'd love to know that so we can keep iterating on this API and eventually move on to the next stage and start getting WebMCP in front of more users.
SPEAKER_00
But to wrap up, AI agents are already using the web. We don't have to settle for these token heavy, brittle, screen scraping processes that we have today. Instead, we can use WebMCP tools to turn every website into a high performance API for agents and at the same time build incredible user experiences for the users of our sites. So now that you have the tools and the context, please give it a go and try making your website's agent ready today. Thank you very much. Thank you. which directions are open. And then you can pick up items, drop items, use items as you navigate this
SPEAKER_00
maze. And if I pop in some prompts, I can see that I can move down. So maybe after that, then write. So I'm going to move down and write. The AI agent should use my prompt, match it to the specific tools. So in this case, the move tool, it's taken my direction of down and right, match that to the north, south, east direction, and send that off to the tool that we have registered on this page. And then it's moved it down down and right. And so you can do, and because it's an AI agent, it can understand a whole bunch of different things. So I could just say, right, up, maybe right again. Let's try that.
SPEAKER_00
And so the AI agent has seen that R stands for right, mapped that to the direction, and then called the move tool with those information. And because it's an AI agent, it can just keep repeating the same tool calls until it thinks that it's done what needs to be done. So I could even say, complete the maze. And then the AI agent should use all the tools available to just keep moving around the maze, to pick up items, to use the items when it needs to, because it has all the information in the tools available. This specific prompt was not the most efficient. So sometimes you'll see it'll go backwards
SPEAKER_00
all the way to the start, and then go forwards again. But the more that you refine the prompt, the better the agent knows how to complete the maze in the most efficient way. For example, if you just say, the exit is in the bottom right corner, it'll be more efficient in its instructions to get to that direction. So I won't continue this, because it can take quite a while to complete this maze. But if we go back to the slides here...
SPEAKER_00
So this is the model context tool inspector that I mentioned. So this is the web extension that our team in Chrome DevRel built. The QR code there is if you want to see where that is in the Chrome web store, but anyone can use that and grab it from the web store. So essentially, WebMCP kind of unlocks this new approach to using the web, where your users don't have to spend a lot of time trying to figure out how to use more complicated sites. And they can figure out their own workflow. So they can choose to browse your website the normal way for a bit, then they can hand over control to their AI agent, and the AI agent takes steps on their behalf. And then
SPEAKER_00
your user can come in at any time to take control again and browse your site again the way they normally would. And so that ability to simplify user journeys and make those user journeys for people easier has been a large part of the reason we've seen interest and excitement in this new standard. So I want to pause for a minute just to address the question that some people have, that's what is the difference between WebMCP and MCP. But you can kind of see them as being complementary to each other. So whereas MCP enables AI agents to connect to applications on the server side,
SPEAKER_00
and you'd need to set up your own server for the agent to access, and then the agent can access the information anywhere at any time. WebMCP is different in that it's kind of inspired by MCP. I like to think of it as how JavaScript is inspired by Java. And that's in short, WebMCP is the implementation of the tools part of the MCP. And so WebMCP allows engineers to provide tools to in-browser AI agents. And it's very specific for the client side features. So you have to have your browser window open for WebMCP to work. And then you can use it to help your agent interact with the
SPEAKER_00
browser. So all of the tools live in the browser. But you can imagine this for quite a few different types of use cases. So imagine those websites that are really complicated, that have a lot of steps that a user needs to take, maybe like booking a flight, or filtering products on a normal shopping website, or filling in complicated medical forms or financial forms, or to trigger fixes that need to be hidden on a page, that are hidden on a page. Or if you're like me, you're just on a normal shopping site. And you're trying to find the right black faux leather clutch bag that can fit your mobile phone in. And instead of going through all the
SPEAKER_00
little filters, you just want to ask your AI agent to do it for you. So these are a bunch of examples where any user can ask whatever AI agent they are using to complete these things on their behalf. So the user doesn't have to manually do this. And they don't have to fill in each input, they don't have to select each checkbox. And using WebMCP in these cases can mean that you can make those actions much easier for users. So let's look at the APIs. WebMCP proposes two approaches for implementation. So you've got the declarative API and the imperative API. Let's start with the declarative API.
SPEAKER_00
So if you have a normal HTML form, you can just add a few attributes to the HTML to get this to work. So we've got the tool name and tool description here. And then your browser will automatically generate a JSON schema that the agent can use to read using the form fields as parameters for the tool. So here's an example of what the JSON schema would look like for this form HTML. And there are a whole bunch of other attributes that can be used. So there's like an agent invoked Boolean attribute. So you can tell whether your form was filled in by an agent or if it was filled in by a human. And there's lots of like more specific attributes that can be used for
SPEAKER_00
things like that too. But essentially, you want to use the declarative API when you have a standard form element. But when you have something more complicated, that's when you want to go back to the imperative API. So this is where you can register and define your own custom tools for when you have more complex, maybe multi-step UI flows. So here is an example. So at the bottom, we have this register tool function. And when you call register tool with an object like this, you need to manually create your own schema similar to the one that we had in the declarative API that was generated. You name your tool and give it the description. And you
SPEAKER_00
want to make sure you have really descriptive descriptions that enable the AI agent to know when it should be calling this tool. And then you have the execute block, which is essentially where you call normal JavaScript. Maybe you already have functions that you're using that you can call in here, maybe do a light wrapper. In this add to the item example, you can validate and trim text input, for example. And then you create the DOM elements or DOM nodes and add them to your page. And then you want to return some information to the AI agent so it knows what happened if everything happened successfully. So it
SPEAKER_00
can use that information for its next steps. So those are the two APIs. The imperative API is probably the one that's most used because people have more complex UI flows that it wants the agent to complete. So if we go back to my Vibe Coded demo, I have added a few tools here. So we have a few featured events in the demo and then all of the events available down here. And then you can go in and purchase tickets on an individual concert page.
SPEAKER_00
So I have noticed that this works much better with Gemini 3.1. So I'm going to try that one. If we wanted to buy tickets to one of these festivals, let's buy tickets to the Summer Vibes Festival.
SPEAKER_00
Let's see, two VIP tickets because VIP only for me.
SPEAKER_00
Let's send that prompt. So the AI saw the tool search concerts, which it has called to find the specific concert via the concert name. And the tool returned the information about the concert, including the ID for that concert. Then it has called the second tool, open concert page with the concert ID. And that has opened this Summer Vibes Festival page. And then this new page has separate tools. This one here called purchase ticket. And it's called that in the third tool call here, with a quantity two and the section name. And then I've got a little notification to say, oh, you bought your tickets. You spent £356. Great. I'll put that on the Google's credit card.
SPEAKER_00
But you can see as well, like in each step, it's updated the UI to make sure the user can also see what's happening. So you always want to make sure that your UI is in sync with the tool calls that are happening. So we've got the VIP selected, we've got the quantity selected, and then in real life, it will go through to some checkout page. You'll probably want your user to manually do that step so they know that they're spending real money.
SPEAKER_00
Let's head back. Let's head back. So if you're interested in trying this out, it's probably worth just understanding the status of where we're at with WebMCP. So we're still in early preview stage. This API is very experimental. It will change. It has been changing over the past few weeks. And so the code that I've shown might be different next week. But that's because we want people to try it out. We want feedback. We want to know the best way to use this API. And if you're interested in doing that, these are a few steps to get set up. So WebMCP is enabled in Chrome version 146 upwards. I recommend using Chrome Canary just so you can keep things separate.
SPEAKER_00
Otherwise, in the normal Chrome, you have to enable experimental flags. And you might not want to do that on your normal browser. Once you have Chrome Canary, you'll need to enable the WebMCP testing flag by putting this flag in your URL. And then install the model context tool inspector extension from the Chrome Web Store that I mentioned earlier. Just so you can play around and debug and see what your tools are doing.
SPEAKER_00
Then, these are the two resources that I recommend taking a look at. So this is our main blog post that gives you information on the early preview program for WebMCP. So if you sign up there, you get access to all of our initial documentation. And you get extra information about the program, information on best practices, all the extra implementation details that you might want to use while you're testing it out. So we're going to have a link to the URL and all of the API information. So that is the first one. And the second one is the GitHub repository of all the tools. So we've got the inspector tool here. We've got all the demos.
SPEAKER_00
So you can see the maze demo code is live there for you to play around with. There's about six, seven different demos you can try out. And there's an evals CLI tool you can use to help you start testing your own sites in the WebMCP tools on your own sites today.
SPEAKER_00
So I mentioned we're still in early preview. That's because we're looking for feedback. So try it out. Let us know what you think. If you have any friction points. If you find any bugs, we'd love to know that so we can keep iterating on this API and eventually move on to the next stage and start getting WebMCP in front of more users. But to wrap up, AI agents are already using the web. We don't have to settle for these token heavy, brittle, screen scraping processes that we have today. Instead, we can use WebMCP tools to turn every website into a high performance API for agents and at the same time build incredible user experiences for the users of our sites.
SPEAKER_00
So now that you have the tools and the context, please give it a go and try making your website's agent ready today. Thank you very much. Thank you.