AI Engineer

Rebuilding the web for agents — Liad Yosef, MCP Apps

2994 summary words 13 min summary Watch video

Start with the signal

13 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Personal AI assistants will replace browsers and websites as the primary web interface, forcing companies to redesign their entire web presence around agent accessibility rather than human UI—and the infrastructure standards, discovery mechanisms, and readiness benchmarks for this transition are being defined right now.
  • Why it matters: Ken is building OpenClaw and agent systems; this talk contains specific architecture choices, failed patterns to avoid, emerging standards battles, and operational data on how agents actually navigate websites versus how engineers think they do—including the surprising finding that human-defined agent best practices like llms.txt are mostly ignored by actual agents.
  • Best use: Watch for the concrete discovery/registry architecture decisions, the empirical agent journey data contradicting conventional wisdom, and the nearly-headless web pattern that solves the last-mile interaction problem; pause at the examples of how Claude chose PostHog over Mixpanel purely on API quality—this is the future selection mechanism for all B2B tools.

Executive Summary

Liad Yosef, co-creator of MCP apps (the spec underlying ChatGPT apps, Claude apps, Copilot, and GitHub chat-based apps) and co-founder of Aura (an agentic web research lab), argues that personal AI assistants are replacing browsers and websites as the primary interface to the web. He presents MCP apps as the critical standard that enables services to send UI chunks into chat contexts, preserving brand identity while solving the last-mile interaction problem—allowing users to see Booking.com or venue seat selectors embedded directly in Claude or ChatGPT. This solves a three-way value exchange: apps retain UI/brand, users gain trust and familiarity, and chat hosts gain access to real-world capabilities without negotiating individual deals or building support infrastructure.

The core insight is that websites must now optimize for agent experience, not just human UX. Major companies (Salesforce, Cloudflare, Sentry) are going headless because agent traffic has already surpassed human traffic on the web, and anecdotal evidence (his nine-year-old using ChatGPT instead of Google, a 70-year-old tourist with only WhatsApp/camera/ChatGPT on her phone) shows the generational shift is real. Yosef rejects both browser-use agents (calling them a 'faster horse solution' that wastes time parsing human UI) and Google's WebMCP co-browsing approach (which still forces agents to interact with dashboards nobody wants). Instead, he advocates the 'nearly headless web': agents interact headlessly with APIs/MCPs but return UI atoms for last-mile human decisions like choosing hotels or seats.

Aura's research revealed a critical failure mode: ~50% of tested websites published llms.txt (the supposed de facto agent-readiness standard), but almost no agents actually used it—they went straight to docs pages and homepages. The only agents that read llms.txt did so because the docs explicitly referenced it. This led to the realization that human-defined best practices become obsolete as models improve (OpenAI stopped publishing tool best practices for this reason), so Aura built tools to let agents define their own needs: Aura Journey (traces agent paths across websites with any harness), Aura Directory (an ARD-compliant discovery layer for agentic resources), and readiness benchmarks that surface agent feedback on websites. The talk includes live examples showing Claude, Vercel's harness, and ChatGPT taking completely different paths through the same site with the same intent.

The discovery problem is the next frontier: traditional web search relies on human SEO; per-agent registries are closed silos; central MCP registries face curation/governance challenges. Aura built a directory based on emerging standards (AI Catalog.json from Anthropic/Google/MCP/A2A for how websites expose themselves; ARD for how directories expose themselves to agents) that scans domains, generates catalog files, and lets agents query for agentic resources. The operational lesson for companies: agent accessibility overlaps heavily with human accessibility (because LLMs are like users with vision disabilities), but the selection criteria have changed—Claude recommended PostHog over Mixpanel purely because PostHog had better MCP/API integration, ignoring a decade of Mixpanel UX polish and the speaker's stated preference.

Key Takeaways

  • Claim: Personal AI assistants are replacing browsers and websites as the primary web interface, shifting user behavior from tab-browsing across multiple UIs to a single assistant that composes UI atoms from multiple services in one conversation. | Evidence: Yosef's nine-year-old uses ChatGPT instead of Google for search; a 70-year-old Georgian tourist had only three apps (WhatsApp, camera, ChatGPT); Yosef's colleagues who connected Jira MCP to their IDEs no longer visit Jira's website; Cloudflare CEO reported agent traffic surpassing human traffic. | Implication: Ken's OpenClaw and agent orchestration systems must assume users will route all web interactions through their personal assistant, not visit websites directly, which changes how services should expose capabilities and measure engagement.
  • Claim: MCP apps solve the last-mile interaction problem by letting services send UI chunks (maps, seat selectors, visualizations) into chat contexts, creating a three-way value exchange where apps keep brand/UI, users gain trust/familiarity, and chat hosts access real-world capabilities without building everything. | Evidence: Claude example showing Booking.com map embedded in chat for hotel search; ChatGPT embedding visualizations from MCP servers; ChatGPT/Claude/Copilot/GitHub all adopted MCP apps spec; services benefit because the chat host provides user context (calendar, preferences) so Booking doesn't need to integrate with Google Calendar or Amazon separately. | Implication: OpenClaw should support MCP apps for last-mile UI when full autonomy isn't appropriate (e.g., honeymoon hotel booking vs. email organization); this is the standard for preserving service differentiation in an agent-mediated world.
  • Claim: Browser-use agents and Google's WebMCP co-browsing are fundamentally flawed approaches because they force agents to interact with UI designed for human perception limitations (filters, pagination, sorting) rather than providing direct API/MCP access. | Evidence: Yosef calls browser agents 'a faster horse solution' and asks why anyone would push a Salesforce dashboard to Gemini just to have it click buttons in a UI no one wants to use; Google's WebMCP exposes JavaScript tools for agents to interact with websites while co-browsing, but still requires navigating human-facing dashboards. | Implication: Ken should architect OpenClaw to prefer headless API/MCP integrations over browser automation whenever possible; browser-use should be a fallback for legacy systems, not the primary pattern.
  • Claim: Almost 50% of tested websites publish llms.txt (the supposed de facto agent-readiness standard), but nearly zero agents actually use it—agents go straight to docs pages and homepages, and the minority that read llms.txt only did so because docs explicitly referenced the file. | Evidence: Aura's empirical testing of tens of thousands of agent journeys across websites; Yosef ran multiple harnesses (Claude, Vercel's harness, ChatGPT) on the same sites and observed completely different navigation paths; the 40% of agents that used llms.txt did so only because docs pages pointed to it. | Implication: Ken should not assume standardized agent-readiness files like llms.txt will guide actual agent behavior; instead, optimize docs pages and homepage structure, and use empirical agent journey data (like Aura Journey) to understand how agents actually navigate your services.
  • Claim: Human-defined agent best practices become obsolete as models improve, so agents must define their own needs through feedback loops rather than static standards like three-paragraph tool descriptions. | Evidence: Six months ago MCP server best practice was three-paragraph descriptions; now three lines suffice; OpenAI stopped publishing tool best practices because they become outdated with each model release; Aura's benchmarks include agent-generated feedback on websites. | Implication: OpenClaw's tool/MCP design should be model-adaptive and rely on empirical feedback loops (agent journey data, success/failure metrics) rather than static documentation patterns; invest in instrumentation to understand how different models interact with your systems.
  • Claim: Claude chose PostHog over Mixpanel for Aura's analytics purely because PostHog had better MCP and API integration, overriding the team's decade of Mixpanel experience and stated preference—demonstrating that agent selection criteria prioritize integration quality over brand loyalty or human UX. | Evidence: Aura team explicitly told Claude they preferred Mixpanel; Claude insisted on PostHog, citing better MCP/API; team followed Claude's recommendation despite contrary preference; Hermes (another agent) remembers when it had to spin up a browser to use a product and avoids it in the future. | Implication: Ken's GTM strategy and investment thesis must account for agent-driven vendor selection where API/MCP quality trumps traditional differentiation (UX polish, brand, sales relationships); companies with inferior agent integration will lose even when humans prefer them.
  • Claim: The discovery layer for agentic resources (MCPs, APIs, auth methods) is the next unsolved problem, and traditional approaches (web search, per-agent registries, central MCP registries) all fail for structural reasons. | Evidence: Web search is based on human SEO/PageRank and not flexible enough; per-agent/per-chat registries are closed and require manual app submission; central MCP registries face curation/governance challenges (who decides what gets listed); Yosef cites the historical pattern of discovery layers enabling revolutions (search for web, app stores for mobile, feeds for social). | Implication: Ken should monitor and potentially contribute to emerging discovery standards (AI Catalog.json for how websites expose themselves, ARD for how directories expose to agents) and consider how OpenClaw will discover and rank agentic resources—this is a strategic choke point.
  • Claim: Aura built a standards-compliant directory (Aura Directory) that generates AI Catalog.json files for scanned domains and lets agents query for agentic resources (MCP servers, APIs, auth methods) in an ARD-compliant way, solving the discovery problem through agent-queryable infrastructure. | Evidence: Example shown for monday.com where AI Catalog.json exposes MCP server and API server; directory is ARD (Agentic Resource Discovery) compliant, a standard co-developed by multiple companies; any agent can query the directory with any query; directory is at aura.directory. | Implication: Ken should evaluate Aura Directory as a discovery layer for OpenClaw and understand the AI Catalog.json and ARD standards—if these become the de facto standards, services integrated with OpenClaw will need to publish catalog files and Ken may want to contribute to or adopt these standards early.
  • Claim: Agent accessibility and human accessibility are nearly identical because LLMs are like users with vision disabilities—they don't see the website and rely on structural signals, so improvements in one improve the other. | Evidence: Aura's research finding that agent-ready practices overlap with accessibility best practices; LLMs navigate websites without visual parsing, relying on semantic HTML, clear structure, and metadata similar to screen readers. | Implication: Ken should audit OpenClaw's own website and any user-facing services for accessibility compliance as a proxy for agent readiness, and bake accessibility into all new UI development as a dual-purpose investment.

Detailed Brief

The Nearly-Headless Web Architecture Pattern

  • Claims: User experience is splitting into agent experience (90% of interaction) plus how users experience the agent's interactions; websites must be agent-ready for their own agent, customers' personal assistants (ChatGPT, OpenClaw), and residual human customers who still browse directly.; The 'nearly headless' pattern means agents interact with websites headlessly via API/MCP but return UI resources (atoms) for last-mile human decisions when full autonomy isn't appropriate, unlike fully autonomous tasks (email organization) that return only a 'done' confirmation.
  • Evidence: Salesforce, Cloudflare, and Sentry all went headless recently; Sentry's David Cramer published 'Designing for Agents' days ago stating interaction will no longer be exclusively through their web app and they must design API-first.; Spectrum example: autonomous email organization needs no UI; honeymoon hotel booking needs last-mile interaction for seat/view selection; MCP apps provide the UI atoms for the latter case.
  • Implications: OpenClaw's architecture should distinguish between fully autonomous workflows and last-mile interaction workflows, using MCP apps or equivalent for the latter; services that don't expose agent-ready APIs/MCPs will be excluded from agent-mediated transactions even if their human UI is superior.

Aura's Research Tools and Empirical Methodology

  • Claims: Aura built three free tools: (1) readiness benchmarks at aura.ai that score websites on agent-readiness and surface agent feedback, (2) Aura Journey at journey.aura.ai that traces agent paths across websites with any harness and intent, (3) Aura Directory at aura.directory that provides an ARD-compliant queryable directory of agentic resources.; Running multiple harnesses (Claude, Vercel's harness, ChatGPT) on the same website with the same intent reveals completely different navigation patterns, proving that agent behavior is model-specific and cannot be optimized with static best practices.
  • Evidence: Aura Journey ran tens of thousands of agent journeys to map the web; leaderboard shows company rankings on agent readiness; live demo showed Claude/Haiku/Eve/ChatGPT taking different paths through Airbnb.; Aura Directory generates AI Catalog.json files for every scanned domain (example: monday.com) exposing MCP servers, APIs, auth methods; directory itself is queryable by agents.
  • Implications: Ken should test OpenClaw's own services with Aura Journey to understand how different models interact with them, use the benchmarks to identify readiness gaps, and potentially integrate Aura Directory for agentic resource discovery; the multi-harness methodology is reusable for Ken's own testing.

Emerging Standards and Governance Landscape

  • Claims: AI Catalog.json is a standard co-developed by Anthropic, Google, MCP, and A2A that defines how websites expose themselves to agents (MCP servers, APIs, auth, pricing).; ARD (Agentic Resource Discovery) is a standard co-developed by multiple companies that defines how discovery layers/directories expose themselves to agents in a queryable way.; These standards are emerging now and will likely determine which directories and discovery mechanisms become the default for agent ecosystems.
  • Evidence: Yosef showed AI Catalog.json file structure for monday.com; referenced ARD compliance for Aura Directory; mentioned co-development by major labs/companies.; Historical pattern: every platform revolution came with a discovery layer (web/search, mobile/app stores, social/feeds); agentic web is at the point of defining its discovery layer.
  • Caveats: Standards are still emerging and may fragment or consolidate; governance models for central registries remain unresolved.
  • Implications: Ken should track AI Catalog.json and ARD standards development closely, consider publishing catalog files for OpenClaw-integrated services early, and evaluate whether to contribute to standards bodies or adopt early to influence direction; missing these standards may mean exclusion from agent discovery ecosystems.

Notable Concepts & Terms

  • MCP apps: The spec (co-created by Yosef) underlying ChatGPT apps, Claude apps, Copilot, and GitHub chat-based apps; enables services to send UI chunks into chat contexts for last-mile interaction while preserving brand/identity.
  • Nearly headless web: Yosef's term for the architecture pattern where agents interact with websites headlessly via API/MCP but return UI atoms (maps, selectors, visualizations) for last-mile human decisions when full autonomy isn't appropriate.
  • llms.txt: A supposed de facto standard file for agent readiness (similar to robots.txt for crawlers), but Aura's research found almost no agents actually use it—they go to docs and homepage instead, making it largely ineffective in practice.
  • AI Catalog.json: Emerging standard (Anthropic/Google/MCP/A2A) defining how websites expose agentic resources (MCP servers, APIs, auth, pricing) to agents; Aura Directory generates these for scanned domains.
  • ARD (Agentic Resource Discovery): Emerging standard defining how discovery layers/directories expose themselves to agents in a queryable way; Aura Directory is ARD-compliant, meaning any agent can query it for agentic resources.
  • WebMCP: Google's co-browsing standard where websites expose JavaScript tools for agents to interact with while the user browses; Yosef rejects this as still forcing agents to navigate human-facing dashboards.
  • Aura Journey: Free tool (journey.aura.ai) that traces agent navigation paths across websites with any harness/intent, revealing empirical agent behavior; Aura ran tens of thousands of journeys to map the web.
  • Agent experience (AX): The new design target for websites; how agents interact with your service's APIs/MCPs/auth, which will determine 90% of interaction, separate from how humans experience the agent's interactions.
  • Last-mile interaction: The residual human decision-making in agent workflows where full autonomy isn't appropriate (choosing hotels, seats, viewing 3D models); MCP apps solve this by embedding service UI in chat contexts.
  • Aura Directory: ARD-compliant directory (aura.directory) built by Yosef's team that scans domains, generates AI Catalog.json files, and lets agents query for agentic resources (MCPs, APIs, auth); addresses the discovery problem for agentic resources.

Operator Notes / Why Ken Should Care

  • Test OpenClaw's website and integrated services with Aura Journey (journey.aura.ai) using multiple harnesses (Claude, ChatGPT, custom) to understand how different models navigate your systems—do not assume llms.txt or static best practices will guide actual agent behavior.
  • Publish AI Catalog.json files for OpenClaw and any services you integrate, following the Anthropic/Google/MCP/A2A standard, to ensure discoverability in agent ecosystems; generate these via Aura Directory or manually.
  • Audit OpenClaw's website and user-facing services for accessibility compliance (semantic HTML, clear structure, screen reader compatibility) as a dual-purpose investment in agent readiness—Aura's research proves the two overlap almost completely.
  • Architect OpenClaw to distinguish between fully autonomous workflows (return 'done' only) and last-mile interaction workflows (return UI atoms via MCP apps or equivalent) rather than defaulting to browser automation for everything.
  • Evaluate Aura Directory (aura.directory) as a discovery layer for agentic resources and consider contributing to or adopting ARD and AI Catalog.json standards early—these may become the de facto discovery standards, and missing them means exclusion from agent-mediated transactions.
  • Monitor which companies in your portfolio or investment pipeline have agent-ready APIs/MCPs versus only human-facing UIs; factor agent integration quality into GTM strategy and due diligence, as Claude's PostHog-over-Mixpanel decision shows this will override brand loyalty and human preference.
  • Instrument OpenClaw to capture agent journey data (which endpoints agents hit, in what order, with what success/failure rates) as a feedback loop for model-adaptive design rather than relying on static documentation patterns.
  • Avoid defaulting to browser automation (computer-use, WebMCP) for OpenClaw integrations; treat it as a fallback for legacy systems without APIs rather than the primary pattern, as it forces agents to parse UI designed for human perception limitations.
  • Track David Cramer (Sentry), Salesforce, Cloudflare as leading indicators of the headless shift; if major SaaS platforms are deprioritizing web UI in favor of API-first, this is a structural signal for where the market is heading.
  • Consider running Aura's readiness benchmarks (aura.ai) on competitor services or portfolio companies to identify agent-readiness gaps before agents (or your customers' agents) encounter them in production.

Source/Metadata

  • Title: Rebuilding the web for agents — Liad Yosef, MCP Apps
  • Transcript words: 6354
  • Duration seconds: 1240
  • Timestamp note: Timestamps were not present in the transcript; some approximations may apply to long-duration videos.
Full transcript 3356 words · 28 min read
0:00

Hi, everyone. Hello, hello.

0:12

So we're going to talk about the agentic web, and more specifically, what it means and how do we make the web ready for agents. One disclaimer: I did this talk yesterday, so it might be outdated because things are moving really fast in this space.

0:23

I need to introduce myself just to get some context. I'm the co-creator and maintainer of a spec called MCP apps. MCP apps is the underlying spec behind ChatGPT apps, Claude apps, Copilot, GitHub. Every chat-based app that you see is based on the spec MCP apps in the MCP committee. I'm also the co-founder of a company called Aura, where we research agentic-human interactions. And I built and led the agentic storefronts in Shopify. So the topic of agentic web is really close to my heart.

0:36

A little bit of primer about MCP apps, if you're not familiar. MCP apps are the spec that paved the way to agentic web. MCP apps were released as a spec and as a standard a few months ago, with the support of Claude first as the first client, and then all clients followed. If you play with any chat-based app and you pulled any visualization and interaction UI layer from an MCP server, you probably used MCP apps.

0:40

The good thing about MCP apps is that everybody wins. If the servers or the providers can send UI chunks into the chats, then the apps gain their brand and identity. They get to keep their UI and not be reduced to a database or just text-based information. Users gain trust and familiarity because if you ask ChatGPT, "Hey, book some hotels for me," and you see Booking.com, then you know it's Booking.com. You know who you're interacting with. And the hosts of the chats gain access to the world of capabilities. A few months ago people asked, "Why won't OpenAI just build everything from scratch?" But OpenAI will not negotiate deals with hotels or support users that want to change seats in a venue. We need those services. And MCP apps actually solve for this last mile of interaction.

0:47

So MCP apps solved our last mile of interaction, right? What does the last mile mean? The last mile is as agents become better and models become better, they can do things more autonomously. But we are still at the end of the chain as humans. We still need to be able to choose the hotel or to view a 3D model or to choose a seat at the venue. All those last mile interactions can change.

0:55

This is an example from Claude, PR on MCP apps. This is how it looks. When you're embedding apps inside chats, you get this unified experience of how UI feels inside the same context of a chat. In the same conversation, you get data from Booking, you get data from Altrails, you can get data from everything you want to interact with. This is already happening. MCP apps is already widely supported across every chat agent except Gemini, but that's coming. The interesting part is that it brings us to what we call the agentic web.

1:09

What is the agentic web? The agentic web is not a web of agents. It's not the web as we know it today. It's a shift. It's a change in the way that we see websites and the way that we see browsers.

1:12

Because up until now, we spent 20 years, two decades, perfecting this experience. If I want to do a project or fulfill a task or plan an anniversary, I have to open tabs in my browser, and I have to go through these tabs, and I have to convey my intent differently to my browser. And I have to do each of those services. If I want to plan an anniversary, then I need to learn Google's UI, Amazon's UI, Booking's UI, and the other Booking UI, and the other Amazon UI. All that just conveys the same intent to those interfaces. Every company perfected this user flow. But we don't need to do it anymore.

1:18

We can take these interfaces and break them into atoms. I don't need 99% of Booking's dashboard. I don't need 99% of Airbnb's dashboard. I definitely don't need Jira's dashboard. But I do want to convey my intent to those services. So why can't I let my personal assistant use that?

1:28

If I have a personal assistant, I can take these atoms and it composes these atoms and says, "I see that you have an anniversary coming. This is the view from Google that says you have an anniversary coming. So I can book, I can buy things for you, I can book your hotel. Now, Claude knows me, so he knows that I prefer to book hotels in nature. So it knows to pull the map from Booking. I don't need to think about it."

1:39

What Booking and Amazon and Google benefit from is that they have this integration layer that they don't need to develop. Because Claude has the context for me. Booking doesn't need to develop integration with my calendar. Amazon doesn't need to develop integration with Booking. This is a win-win-win.

1:44

This is the view that we're aiming for. And this is very different from what's happening now, which is browsing into different services just to stare at different text boxes that ask what I want to do there. But no one is going to do that. In a year, no one is going to browse to Amazon's agents and Etsy's agents and Expedia's agents. I don't want to do that. I have my own personal assistant. I don't want to use your agent. I want to use my agent.

1:55

Now, if we think about it, websites as the source of truth are something that's going to go away. Why would they need the website? Why would they need to open the tab? Why would they need the information that the website is working really hard to convey?

2:02

The immediate response to that is that we have browser agents. We have computer use agents. We have very smart agents or assistants that can browse the web for us. But that doesn't make sense. That's a faster horse solution. Why would they let an agent that doesn't know anything about human perception limitations work with filters and paginations and sorting and all those UIs that we spent decades perfecting just to fulfill its task?

2:11

Google is very bullish on the other end of the spectrum, which is co-browsing and websites. They have a standard called WebMCP. WebMCP says, instead of the agent taking screenshots of the websites and trying to figure out what's happening, let's have the website expose some tools in the website—JavaScript tools—and the agent will work with it. You can see here an example of Gemini in Chrome that buys stuff for me or finds deals for me as I browse the web, which is very nice, but it's very simple. What happens if we push a Salesforce dashboard to Gemini on Chrome? Why would I want that? Why would I want to browse to a Salesforce dashboard with Gemini to click buttons for me in a dashboard that I don't even want to use? This is definitely not what we want.

2:13

The shift is that assistants are becoming our entry point to the web. We see it everywhere. Every major lab wants its app to be the everything app. The shift is happening. Friends of mine that connect Jira's MCP to their IDEs don't go to Jira's website anymore. My mother uses ChatGPT for everything. If she could book a doctor's appointment using ChatGPT, she won't go to that clinic's websites anymore. You have a long tail of websites that no one is going to see because it's much easier for me to connect them using my personal assistant. So websites will become obsolete, browsers will become obsolete, and personal assistants will become our only gate to the web.

2:24

I know what it sounds. It sounds old. They're not going to go completely obsolete, but we see it with the younger generation today. My nine-year-old goes to ChatGPT. He doesn't go to Google when he wants to search for stuff. I was traveling to Georgia, the country in Eastern Europe, and there was an older lady there that asked me to take a picture of her, a 70-year-old. On her phone, she had three apps: WhatsApp, camera, and ChatGPT. That's it. The shift is happening.

2:32

Very smart people, like David Cramer from Sentry, said a year ago, "I will take this bet against anyone who thinks web stuff is going to be obsolete in 25 years." That was one year ago. A few days ago, David Cramer published "Designing for Agents," which says, "We have to recognize that interaction with products at Sentry will no longer be exclusively through our web application."

2:34

Yeah, they're not going to go completely obsolete, but we see it with the younger generation today. My nine-year-old goes to ChatGPT. He doesn't go to Google when he wants to search for stuff. I was traveling to Georgia, the country in Eastern Europe, and there was an older lady there that asked me to take a picture of her, a 70-year-old. And on her phone, she had three apps: WhatsApp, camera, and ChatGPT. That's it. The shift is happening.

2:41

And very smart people, like David Kremer from Sentry, said a year ago, I will take this bet against anyone who thinks web stuff is going to be obsolete in 25 years. That was one year ago. A few days ago, David Kremer published "Designing for Agents," which says we have to recognize that interaction with products at Sentry will no longer be exclusively through our web application. And we have to design for API first as a surface. So everything is going headless. Salesforce recently went headless, which is very important because Salesforce's main differentiator from its competitors is its UX. It's the way that it conveys everything to the user. But it went headless. Cloudflare went headless. Sentry went headless.

2:42

This is a tweet by a Cloudflare CEO saying that agents' traffic surpassed human traffic on the web. And we have this spectrum of interaction because if my agent is autonomous, I don't need this last mile of UI. If I have an open claw, I just say, yeah, organize my emails. It comes back. It says, done. Good. But if I want to book a hotel for my honeymoon, I probably need this last mile of interaction. So we call it the nearly headless web. It's not completely headless. The agents will interact with your website headlessly, but then it will bring back these UI resources for you.

2:47

So user experience as we know it is being split. Now it's the agent experience of your website plus how a user experiences the agent that experiences your website. But 90% is going to rely on agent experiences, which means your website needs to be agent ready for everything. It needs to be agent ready for your agent. If I'm booking, I need to make bookings that are ready for bookings agent. I need to make it ready for my customers' ChatGPT or my customers' open claw. And I need to be ready for my human customer who doesn't have an agent but still wants to browse to booking.com. So I need to be agent ready.

2:51

And what we found out is that agent ready means a lot of things. We heard a lot about AEO, SEO, GEO, how to get discovered. But discovery is only the first step. Because once an agent knows about you, it still needs to know what you are and how to interact with you and how to authenticate to you and how to headlessly pay you.

2:53

And actually, when we built analytics for our product, we asked Claude what's the best analytics service. And it recommended PostHog, which is an analytics service. And we said, no, we prefer Mixpanel. Because we know Mixpanel, we worked with Mixpanel for a decade, we know how to work with it. And Claude insisted on PostHog. Because it said PostHog has better MCP and API, and I can integrate to it better. So we don't have brand loyalty. We went with PostHog if that's what Claude recommended. But it made us think about Mixpanel. Mixpanel spent a decade perfecting their UX and developer experience. We just left it just because Claude prefers PostHog. And it will always happen. Hermes, for example, if it uses your product and it needs to spin up a browser to use your product, it will remember that. And next time, it won't go to your product anymore.

2:59

So at Aura, which is an agentic web research lab, we started researching. We raised some capital, and we started researching what does it mean for the web to be ready for agents. And we have these readiness benchmarks, which is interesting. You can run it. It's free. You can go to aura.ai. You can run any website. You get this score and benchmark according to a lot of protocols and best practices, and you have agentic feedback. So the agent actually returns the feedback about your website. And we have this leaderboard of how companies and products rank according to this benchmark.

3:02

And we started mapping the web, and everything was nice. But we hit one insight, or one unexpected result. We found out anyone here heard about llms.txt? llms.txt is the de facto standard to be agent-ready. You say, yeah, if your website published an llms.txt, agents know how to interact with you. And you have auth.md and pricing.md and x402 and a lot of standards. And we found out that almost 50% of the websites that we tested, that we ran, published llms.txt. But none of the agents that we ran on this website actually used llms.txt. Actually, almost all of the agents went straight to the docs page and then the homepage. And the 40% that did use llms.txt used it only because the docs pointed out that there's a file called llms.txt that they need to use.

3:11

So then it hit us. We said, it doesn't make a lot of sense for us as humans to define to agents what they need. No one is doing it anymore. Even OpenAI, they don't publish best practices for tools anymore because they say that every time they publish best practices, the models become better and these practices become obsolete. Six months ago, the best practice for an MCP server was, yeah, have three paragraph descriptions so agents will know how to interact with you. And now three lines are enough. So best practices become obsolete. We need the agents to define what the agents need. We need agents' feedback on this website.

3:14

So we built Aura's Journey. That's a really cool. That's also free. You can go to journey.aura.ai and you can run on any website, any intent, with any agent and see the path of the agent as it tries to interact with the website. So, for example, airbnb.com, choose Claude, and we run it on the website. And you can see in real time how the agent goes and what it tries to look for in the website.

3:19

And we did it tens of thousands of times just to understand what agents really look for when they try to interact with websites. And the cool thing is not running just one harness. It's running multiple harnesses. As you see here, Claude and Eve, which is Vercel's harness, and ChatGPT on the same website with the same intent. Okay? So this is Claude. This is Haiku. This is Eve. See how different the agent journey looks. And this is ChatGPT. So ChatGPT could find results better. You can just see the journey across the website. And we need to understand why. We need to understand why websites publish auth.md files but agents don't look for auth.md files. And what do they look for?

3:19

So in Aura, you can actually go to any business question. So for any domain, you have business goals or questions. And you can see the paths that the agents are taking. And the interesting part of it is that now we know how agents interact with websites. What's the next big milestone for the agentic web? And for those of you who heard the previous talk, it's discovery. But it's not the discovery that we think of. It's not a CEO. It's not a GEO. It's not an AEO. Because we have to remember that every revolution came with its discovery layer. The web revolution came with search. Mobile came with app stores. Social came with feeds. What is the discovery layer for agentic resources? What is the discovery layer for MCPs? What is the discovery layer for openapi.json? What do we even look for? Maybe Airbnb has a better MCP than booking.com. But booking.com has better API than Airbnb. So how do we do it?

3:23

Web search. Classic web search. It's not enough. Because it's based on human SEO and human page rank. And it's not flexible enough. Custom registries, like per agent or per chats, they're not enough. Because they're closed. And they require every app to submit itself to these registries. And a central registry, like an MCP registry, that's not enough. Because who will do the curation? What's the governance model? How do we decide which resource comes to that registry?

3:24

There are emerging standards around it. There's an AI Catalog.json, which is a standard by Anthropic and Google, MCP, and A2A, which standardizes how a website exposes itself to agents. And there's the agentic resource discovery standard, which is by all these companies and more, which basically standardizes how a discovery layer or a directory exposes itself to agents. So we built Oracle Directory. Because we're a research lab for the agentic web, we built this. And in Oracle Directory, we take all the domains that we scanned or scanned ourselves.

3:40

Custom registries, like per agent or per chats, they're not enough because they're closed and they require every app to submit itself to these registries. And a central registry, like an MCP registry, that's not enough because who will do the curation? What's the governance model? How do we decide which resource comes to that registry?

3:44

There are emerging standards around it. There's an AI Catalog.json, which is a standard by Anthropic, Google, MCP, and A2A, which standardizes how a website exposes itself to agents. And there's the agentic resource discovery standard, which is by all these companies and more, which basically standardizes how a discovery layer or a directory exposes itself to agents.

3:47

So we built Aura Directory because we're a research lab for the agentic web. We take all the domains that we scanned ourselves and we put it in a directory that agents can actually query. We expose the AI Catalog.json files. So you can see here, for example, for monday.com, you can see that we generate this JSON file that tells agents, this is the MCP server for Monday, this is the API server for monday.com, and the agent can just query that. So for every domain, you can go to AI Catalog.json. And we expose the registry, the directory itself. So for every entry on the directory, we can tell the agent, for example, Vercel, these are the agentic resources, and this is how you access them.

3:53

This directory is fully ARD or agentic resource discovery compliant, so any agent can just query that directory with any query. This is aura.directory if you want to take a look. All of it is part of the Aura multiverse—journey, directory, and the ranker. We found some very interesting insights. For example, being agent-ready and being human-accessible is very similar because LLMs, when they come to your website, they're like users or people with vision disabilities. They don't see your website. They need other signals to understand how to work with your website. Making your website human-accessible helps agent accessibility and vice versa.

4:02

To wrap up, these are MCP apps. This is what I've been working on in the past few months, and this was the last piece for the agentic web. This actually brought the agentic web. Now agents start to roam the web. Agents need different things. They need different roads. They need different infrastructure. But we don't want to rebuild the web for agents. We want to make the web agent accessible. So let's just make sure the web is prepared for when agents are coming. Thank you very much. The agentic web is not a web of agents. It's not the web as we know it today. It's a shift. It's a change. It's a change in the way that we see websites and the way that we see browsers.

4:17

Because up until now, we spent 20 years, two decades, perfecting this experience. Meaning that if I want to somehow do a project or fulfill a task or plan an anniversary, I have to open tabs in my browser, and I have to go through these tabs, and I have to convey my intent differently to my browser. And I have to do each of those services, right? Meaning that if I just want to plan an anniversary, then I need to learn Google's UI, and Amazon's UI, and Booking UI, and the other Booking UI, and the other Amazon UI. All that just convey the same intent to those interfaces. And every company perfected this user flow. But we don't need to do it anymore, right?

5:02

We can just take these interfaces and break them into atoms. Because I don't need 99% of Booking's dashboard. I don't need 99% of Airbnb's dashboard. I definitely don't need Jira's dashboard. But I do want to convey my intent to those services. So why can't I let my personal assistant use that? So if I have a personal assistant, I can just take these atoms, and it composes these atoms, and it says, yeah, I see that you have an anniversary coming. This is the view from Google that says that you have an anniversary coming. So I can book, I can buy things for you, I can book your hotel. Now, Claude knows me, so he knows that I prefer to book hotels in the nature, right?

5:46

So it knows to pull the map from Booking. I don't need to think about it. And what Booking and Amazon and Google benefits from this is that they have this integration layer that they don't need to develop. Because Claude has the context for me, right? So Booking doesn't need to develop integration with my calendar. Amazon doesn't need to develop integration with Booking. So it's a win-win-win. This is the view that we're aiming for. And this is very different from what the things are going, which is this, which is basically browsing into different services just to stare at different text boxes that ask what I want to do there. But no one is going to do that.

6:27

In a year, no one is going to browse to Amazon's agents and Etsy's agents and Expedia's agents. I don't want to do that. I have my own personal assistant. I don't want to use your agent. I want to use my agent. Okay? Now, if we think about it, it means that websites as the source of truth are something that's going to go away. Because why would they need the website? Why would they need to open the tab? Why would they need the information that the website is working really hard to convey to me? And the immediate response to that is that, yeah, we have browser agents. We have computer use agents. We have very smart agents or assistants that can browse the web for us.

7:08

But that doesn't make sense. Because that's a faster horse solution. Why would they let an agent that doesn't know anything about the human perception limitations work with filters and paginations and sorting and all of those UIs that we spent decades perfecting for us just for it to fulfill its task? So Google is very bullish on the other end of the spectrum, which is co-browsing and websites. And they have a standard called WebMCP. WebMCP says, okay, instead of the agent taking screenshots of the websites and trying to figure out what's happening, let's have the website expose some tools in the website, JavaScript tools, and the agent will work with it.

7:51

And you can see here an example of Gemini in Chrome that buys stuff for me or finds deals for me as I browse the web, which is very nice, but it's very simple. What happens if we push Salesforce dashboard to Gemini on Chrome? Why would I want that? Why would I want to browse to Salesforce dashboard with Gemini to click buttons for me in a dashboard that I don't even want to use? I mean, this is definitely not what we want. And the shift is that assistants are becoming our entry point to the web, right? We see it everywhere. Every major lab wants its app to be the everything app. And the shift is happening.

8:34

And friends of mine that connect Jira's MCP to their IDs, they don't go to Jira's website anymore, right? And my mother uses chat GPT for everything. If she could book a doctor's appointment using chat GPT, she won't go to that clinic's websites anymore. And you have a long tail of websites that no one is going to see because it's much easier for me to connect them using my personal assistant. So websites will become obsolete, browsers will become obsolete, and personal assistant will become our only gate to the web. And I know what it sounds like. It sounds like old men yelling at a literal cloud, right?

9:07

I mean, yeah, they're not going to go completely obsolete, but we see it with the younger generation today. My nine-year-old goes to chat GPT. He doesn't go to Google when he wants to search for stuff. I was traveling to Georgia, the country in Eastern Europe, and there was an older lady there that asked me to take a picture of her, like a 70-year-old. And on her phone, she had three apps, WhatsApp, camera, and chat GPT. That's it. The shift is happening. And very smart people, like David Kremer from Sentry, said a year ago, I will take this bet against anyone who thinks Web stuff is going to be obsolete in 25 years. That was one year ago.

9:48

A few days ago, David Kremer published Designing for Agents, which says, we have to recognize that interaction with products at Sentry will no longer be exclusively through our web application. And we have to design for API first as a surface. So everything is going headless. Salesforce recently went headless, which is very important because Salesforce's main differentiator from its competitors is its UX. It's the way that it conveys everything to the user. But it went headless. Cloudflow went headless. Sentry went headless. This is a tweet by a Cloudflow CEO saying that agents' traffic surpassed human traffic in the web.

10:24

And we have this spectrum of interaction because if my agent is autonomous, I don't need this last mile of UI. If I have an open claw, I just say, yeah, organize my emails. It comes back. It says, done. Good. But if I want to book a hotel for my honeymoon, I probably need this last mile of interaction. So we call it the nearly headless web. It's not completely headless. The agents will interact with your website headlessly, but then it will bring back these UI resources for you. So user experience as we know it is being split. Now it's the agent experience of your website plus how a user experiences the agent that experiences your website.

11:05

But 90% is going to rely on agent experiences, which means your website needs to be agent ready for everything. It needs to be agent ready for your agent. If I'm booking, I need to make bookings that are ready for bookings agent. I need to make it ready for my customers' chat GPT or my customers' open claw. And I need to be ready for my human customer who doesn't have an agent but still wants to browse to booking.com. So I need to be agent ready. And what we found out is that agent ready means a lot of things. So we heard a lot about AEO, SEO, GEO, how to get discovered. But discovery is only the first step.

11:45

Because once an agent knows about you, it still needs to know what you are and how to interact with you and how to authenticate to you and how to headlessly pay you. And actually, when we built, it's an anecdote, but when we built analytics for our product, we asked CloudCode what's the best analytics service. And it recommended PostHog, which is an analytics service. And we said, no, we prefer Mixpanel. Because we know Mixpanel, we worked with Mixpanel for a decade, we know how to work with it. And CloudCode insisted on PostHog. Because it said, PostHog has better MCP and API, and I can integrate to it better. So we don't have brand loyalty, right?

12:18

We went with PostHog, if that's what CloudCode recommended. But it made us think about Mixpanel. Mixpanel spent a decade perfecting their UX and developer experience. We just left it just because CloudCode prefers PostHog. And it will always happen. Hermes, for example, if it uses your product and it needs to spin up a browser to use your product, it will remember that. And next time, it won't go to your product anymore, right? So at Aura, which is an agentic web research lab, we started researching. We raised some capital, and we started researching what does it mean for the web to be ready for agents. And we have these readiness benchmarks, which is interesting.

12:56

So you can run. It's free. You can go to Aura.ai. You can run any website. You get this score and benchmark according to a lot of protocols and best practices, and you have agentic feedback. So the agent actually returns the feedback about your website. And we have this leaderboard of how companies and products rank according to this benchmark. And we started mapping the web, and everything was nice. But we hit one insight, or one unexpected result. We found out... Anyone here heard about LMS.txt? LMS.txt, that's like the de facto standard to be agent-ready. You say, yeah, if your website published an LMS.txt, agents know how to interact with you.

13:44

And you have auth.md and pricing.md and X402 and a lot of standards. And we found out that almost 50% of the website that we tested, that we ran, published LMS.txt. But none of the agents that we ran on this website actually used LMS.txt. Actually, almost all of the agents went straight to the docs page and then the homepage. And the 40% that did use LMS.txt used it only because the docs pointed out that there's a file called LMS.txt that they need to use. So then it hit us. We said, it doesn't make a lot of sense for us as humans to define to agents what they need. No one is doing it anymore.

14:26

Even OpenAI, they don't publish best practices for tools anymore because they say that every time they publish best practices, the models become better and these practices become obsolete. Six months ago, the best practice for an MCP server was, yeah, have three paragraph descriptions so agents will know how to interact with you. And now three lines are enough. So best practices become obsolete. We need the agents to define what the agents need, right? We need agents' feedback on this website. So we built Aura's Journey. And Aura's Journey, that's a really cool. That's also free.

14:57

You can go to journey.aura.ai and you can run on any website, any intent, with any agent and see the path of the agent as it tries to interact with the website. So, for example, atia.com, choose Cloud Code, and we run it on the website. And you can see in real time how the agent goes and what it tries to look for in the website. And we did it tens of thousands of times just to understand what agents really look for when they look, when they try to interact with websites. And the cool thing is not running just one harness. It's running multiple harnesses. As you see here, Cloud Code and Eve, which is Vercel's harness, and ChagPT on the same website with the same intent.

15:45

Okay? So this is Cloud Code. This is Haiku. This is Eve. See how different the agent journey looks like. And this is ChagPT. So ChagPT could find results better. See, you can just see the journey across the website. And we need to understand why. We need to understand why does it happen. Why websites publish auth.md files but agents don't look for auth.md files. And what do they look for? So in Aura, you can actually go to any business question. So for any domain, you have business goals or questions. And you can see the paths that the agents are taking. And the interesting part of it is that, okay, now we know how agents interact with websites.

16:32

What's the next big milestone for the agentic web? And for those of you who heard the previous talk, it's discovery. Right? But it's not the discovery that we think of. It's not a CEO. It's not a GEO. It's not a AEO. Because we have to remember that every revolution came with its discovery layer. Right? The web revolution came with search. Mobile came with app stores. Social came with feeds. What is the discovery layer for agentic resources? What is the discovery layer for MCPs? What is the discovery layer for openapi.json? What do we even look for? I mean, maybe Airbnb has a better MCP than booking.com. But booking.com has better API than Airbnb. So how do we do it?

17:16

Web search. Classic web search. It's not enough. Because it's based on human SEO and human page rank. And it's not flexible enough. Custom registries, like per agent or per chats, they're not enough. Because they're closed. And they require every app to submit itself to these registries. And a central registry, like an MCP registry, that's not enough. Because who will do the curation? What's the governance model? How do we decide which resource comes to that registry? There are emerging standards around it. There's an AI Catalog.json, which is a standard by Anthropocop and AI, Google, MCP, and A2A, which standardizes how a website exposes itself to agents.

18:00

And there's the agentic resource discovery standard, which is by all these companies and more, which is basically standardized how a discovery layer or a directory exposes itself to agents. So we built Oracle Directory, right? Because we're a research lab for the agentic web. So we built this. And in Oracle Directory, we take all the domains that we scanned or scanned ourselves, and we put it in a directory that agents can actually query, right? We expose the AI Catalog.json files. So you can see here, for example, for monday.com, you can see that we generate this JSON file that basically tells agents, yeah, this is the MCP server for Monday.

18:40

This is the API server for monday.com. And you can just, the agent can just query that, right? So for every domain, you can just go to AI Catalog.json. And we expose the registry, the directory itself. So for every entry on the directory, we can tell the agent, yeah, for example, Vercel, these are the agentic resources, and this is how you access them. And obviously, this directory is fully ARD or agentic resource discovery compliant, so any agent can just query that directory with any query, right? So this is Aura.directory, if you want to take a look. All of it is part of the Aura multiverse, journey, directory, and the ranker.

19:22

And we found out some very other interesting insights. For example, being agent-ready and being human-accessible is very similar. Because LLMs, when they come to your website, they're like users or people with vision disabilities. They don't see your website. They need other signals to understand how to work with your website. And making your website human-accessible helps agent accessibility and vice versa. So to wrap up, these are MCP apps. This is what I've been working on in the past few months. And this was the last piece for the agentic web. This actually brought the agentic web. And now agents start to roam the web. And agents need different things.

20:03

They need different roads. They need different infras. But we don't want to rebuild the web for agents. We want to make the web agent accessible. So let's just make sure the web is prepared for when agents are coming.

20:18

Thank you very much. Thank you.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note