AI Engineer

The Missing Layer in Agentic AI — Giedrius Šteimantas, Oxylabs

1572 summary words 7 min summary Watch video

Start with the signal

7 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Web-capable agents need a dedicated access infrastructure layer—search, validated extraction, stealth browsing, proxies, and geolocation—rather than treating browser automation as the universal web interface.
  • Why it matters: The talk gives a reusable cost-and-reliability architecture for agents that must discover, verify, and act on open-web information without wasting LLM tokens on CAPTCHAs, invalid pages, or raw site payloads.
  • Best use: Use it as a design pattern for separating low-cost web discovery and verification from browser-only transactional execution, while evaluating whether Oxylabs or an equivalent provider should supply that infrastructure.

Executive Summary

Šteimantas argues that the missing layer in many agentic systems is not another model or orchestration framework but web-access infrastructure. His example is a shopping agent that can discuss a user's style, find products, validate them, and buy them. The original implementation used browser automation for every stage and consequently became slow, expensive, unreliable, and vulnerable to CAPTCHAs.

His proposed architecture assigns different tools to different workflow stages. For discovery, the agent searches the indexed web through a compact search API instead of driving retailer search pages in a browser. For product verification and selection, it retrieves valid, localized, LLM-ready page content through a scraping API. Only at purchase time—where pages are dynamic and the agent must interact with inputs—does it use a browser.

The central operational rule is to validate web content before sending it to an LLM. A 200 response and a large payload do not establish that a product page was actually retrieved; the payload may be a CAPTCHA or block page. Filtering failed retrievals before model inference prevents the agent from spending tokens determining that bad content is bad, and improves the candidate set available for downstream decisions.

The presentation is also a product pitch for Oxylabs' Fast Search API, Web Scraper API, and Headless Browser. Still, the underlying lessons are portable: use browsers selectively, return lightweight structured or Markdown content rather than raw HTML, make retrieval failures explicit, maintain location consistency across stages, and treat web-access cost and success rates as first-class agent metrics.

Key Takeaways

  • Claim: Browser automation should be the exception, not the default interface for an agent operating on the web. | Evidence: The speaker's friend's shopping agent used a browser for discovery, verification, and checkout; CAPTCHAs, retries, JavaScript-heavy sites, and running many browser instances made it slow, unreliable, and difficult to price per transaction. The rebuilt workflow reserves Playwright MCP for the final purchase interaction. | Implication: Design agent tool routing around task requirements: search/extract through APIs first, and escalate to browser control only for interaction that cannot be represented as retrieval. | Caveat: A browser remains necessary when the task requires dynamic, stateful interaction with inputs, such as selecting a size, adding an item to a cart, and completing checkout.
  • Claim: Content validity must be established before LLM processing; HTTP status and payload size are insufficient health checks. | Evidence: The speaker describes teams checking only for HTTP 200 and content size, then passing raw HTML to an LLM. In his example, if 10 sites are opened but only 3 return real product content, sending all 10 payloads to the model wastes roughly 70% of tokens on CAPTCHA or blocked pages. | Implication: Put a retrieval-validation gate ahead of model calls, with explicit block/CAPTCHA/error classification and exclusion of invalid responses from the decision context.
  • Claim: Discovery should leverage web search rather than a deterministic list of retailer pages. | Evidence: The original system queried search pages on a predefined retailer list through a browser, limiting inventory choice to those sites. The replacement uses agent-generated fan-out queries against indexed search engines, returning selected URLs in compact JSON responses reportedly under 2,000 tokens and averaging under 700 ms. | Implication: For open-web research or commerce agents, make discovery an adaptable query-and-ranking step rather than hard-coding a small supplier universe unless coverage control is more important than breadth. | Caveat: The performance and pricing figures are vendor claims for Oxylabs' Fast Search API, not independently benchmarked in the transcript.
  • Claim: The verification layer should deliver successful, LLM-ready content and surface access failures loudly. | Evidence: The proposed Web Scraper API returns an explicit error on CAPTCHAs or other blocks, supports Markdown rather than raw HTML, can process hundreds of requests in parallel, and can render dynamic sites through a browser behind the API when needed. The speaker says customers pay only for successful retrievals. | Implication: Separate extraction from agent reasoning, require structured failure states, and compare providers on usable-content yield rather than nominal request success or HTTP status alone. | Caveat: “Only valid content,” high success rates, and pay-for-success are provider-specific claims; an operator should independently test coverage, failure definitions, latency, and total cost for target domains.
  • Claim: Geolocation is a correctness requirement, not merely a scraping optimization, for commerce workflows. | Evidence: The speaker notes that e-commerce sites display location-dependent stock, sizes, and availability; the original agent could find an item during discovery but discover it was unavailable at checkout. The rebuilt verification and browser stages both apply localization. | Implication: Carry a consistent user-location context across discovery, verification, and execution; otherwise, an agent can make decisions on inventory that cannot actually be purchased.
  • Claim: For unavoidable browser execution, hardened browser infrastructure can be a drop-in reliability layer. | Evidence: Both implementations use Playwright MCP and an LLM for checkout, but the revised version replaces the ordinary browser with Oxylabs Headless Browser, described as Playwright MCP-compatible and equipped with browser-level stealth, residential proxies, and geolocation. | Implication: If deploying browser agents against third-party sites, assess authorization and policy constraints alongside technical stealth, and use managed browser infrastructure only where the use case is permitted. | Caveat: Automating retail checkouts and bypassing access controls can raise site-terms, consent, fraud, and compliance concerns that the presentation does not address.

Detailed Brief

Reference workflow: discover, decide, confirm, execute

  • Claims: The shopping-agent workflow has four functional stages: discovery of purchasable product pages, agent decision based on price/stock/description, user confirmation, and transactional execution.; The user remains the final approver before a purchase is made, rather than allowing the agent to execute autonomously after product selection.; The speaker frames the infrastructure layer as a way to let product builders concentrate on agent behavior rather than repeatedly solving web-access mechanics.
  • Evidence: The agent first converses about the user's style, translates preferences into shopping prompts, evaluates candidate items against those prompts, and asks the user to accept or reject the proposed purchase.; In the proposed implementation, discovery uses compact search results, verification retrieves localized page information, and checkout performs size selection, cart addition, and purchase in a browser.
  • Caveats: The transcript demonstrates an architecture and a vendor solution but does not provide measured end-to-end results such as checkout completion rate, absolute transaction cost, false-valid-content rate, or comparison against alternative providers.; The transcript's closing material is duplicated, so it adds no additional evidence beyond the main presentation.
  • Implications: A transactional agent should preserve a clear approval boundary after recommendation and before irreversible external action.; Evaluate agent workflows by end-to-end usable outcomes—validated candidate coverage, decision quality, completion rate, latency, and cost per successful transaction—not by whether individual page fetches return HTTP 200.

Notable Concepts & Terms

  • Missing infrastructure layer: The speaker's name for the web-access capability beneath an agent: search, extraction, anti-blocking measures, proxying, rendering, and localization.
  • Use a browser when you absolutely have to: The governing design rule: prefer lightweight retrieval and structured APIs; reserve browser automation for genuinely interactive, dynamic workflows.
  • Content validation: Detecting CAPTCHAs, block pages, and invalid retrievals before sending content to an LLM, reducing wasted context and misleading downstream reasoning.
  • Fan-out queries: An agent-generated set of search queries used to discover candidate URLs across the indexed web rather than within a fixed retailer list.
  • Playwright MCP: The browser-control interface used in the execution stage, where the agent must manipulate dynamic webpage inputs and transaction state.
  • Geolocation consistency: Using the same location context in verification and checkout so displayed stock, sizes, pricing, and final availability align.
  • Fail loudly: Return explicit retrieval errors for blocks or CAPTCHAs instead of silently returning misleading page content that the LLM must interpret.
  • LLM-ready Markdown: A lighter alternative to raw HTML intended to lower token consumption and simplify page understanding during agent decision-making.

Operator Notes / Why Ken Should Care

  • Add a pre-LLM web-retrieval gate that labels each response as usable content, CAPTCHA/block, authentication wall, empty page, or extraction failure; log usable-content yield by domain.
  • Implement tool routing that defaults to search and extraction APIs, with browser escalation triggered only by required interaction, client-side rendering, or authenticated state.
  • Carry location, locale, currency, and session assumptions as explicit workflow state from candidate discovery through final execution.
  • Benchmark any web-access provider on target domains using cost per validated page and cost per completed workflow, not advertised request price or generic success rate.
  • Establish policy controls before deploying stealth/proxy-backed browser automation: permitted target domains, user consent, checkout confirmation, credential handling, and transaction limits.

Source/Metadata

  • Title: The Missing Layer in Agentic AI — Giedrius Šteimantas, Oxylabs
  • Transcript words: 2594
  • Duration seconds: 904
  • Timestamp note: No timestamps or chapters were provided. The latter portion of the transcript repeats the verification, browser-execution, and closing sections.
Full transcript 1988 words · 12 min read
0:13

What a beautiful voice. Thank you for coming. Today I'm going to talk a lot about Missing Layer by Gentic AI and explain a little bit about how web scraping infrastructure can actually help you. But first, let me talk a little bit about my friend's idea. So my friend had this idea. He built this AI chatbot that chatted with people about their style, and it was supposed to help them pick out new items as some sort of personal shopper. And once those items were picked out, this chatbot would produce prompts that a shopping agent would then take and attempt to find them online and purchase them for the customers.

0:32

This idea is not new, and it could be applicable to many scenarios. But my friend was good at building agents, and he ran into different problems and asked me for advice. And when he ran it, he would usually, instead of product pages or whatever, get things like that. He would get CAPTCHAs. And, of course, he was doing it very, very quickly. So he webcoded the whole thing without having a thought about infrastructure and underlying layers and how it should work at all. He was using a browser automation framework for everything. And it was slow, expensive, and unreliable. So, in the end, he made a product that does not work and is expensive to run. So he asked me for help.

1:08

And I was a little bit reluctant at first because I don't like giving out professional advice for free. But I took a look at it, and I got a little curious, I have to be honest. I noticed that he was missing something. He was missing a layer, an infrastructure layer, that would allow this agent to operate freely on the open web. My name is Gedros. I work for Oxylapse, where in the past 10 years we've helped companies that trained large language models get their data. And now we use this infrastructure to help AI agents access web at scale and at low cost.

1:30

And before we go into this agent and see how we can build it, I wanted to talk a little bit about the scraping industry and how we operate. And the principles that we operate on can be summed up by one sentence. Cost matters. And the first principle is: use a browser when you absolutely have to. Validate content. HTTP response 200 does not mean that we are good to go. Lighter content is preferred. Websites are full of JavaScript, CSS, and HTML, and there are a lot of bytes that do not deliver any value whatsoever. And today, I will demonstrate how these principles are also applicable when building agents that interact with the web.

2:03

So, coming back to my friend's agent, let's take a look and see how we could do a better job in making this agent more lively. So here's how my friend set it all up. Four different stages. Discovery. The agent was supposed to find product pages on websites where these items can be bought. And then a decision stage, where an agent can decide what products to buy based on the content of these pages. So the agent has to visit them, verify that the stock is there, the price is right, and the description fits the prompt. And once that decision is made, the user is given a choice whether to go ahead with the purchase or reject it altogether.

2:26

Then, of course, we go to execution right away. Then execution was making the purchase. But the problem was that sometimes it worked and sometimes it did not. That was a little problematic. So let's dissect it step by step and see how we could build this differently while improving performance and reducing the cost dramatically by using the same principles from the scraping industry. So the first stage: discovery. My friend chose to go with a predefined list of websites, major retailers, and query their search pages in order to find his products. He used the browser automation tool for that. And it kind of worked, but it did have challenges.

3:02

So the browser automation tool lacked what we call stealth. So they would get CAPTCHAs and sometimes fail to access the sites altogether. This would break down the flow. So a retry mechanism would have to be put in place, making the whole process very long, costly, and sometimes the sites would not be accessed at all. And also, as a result, it became very difficult to predict the final cost per transaction. The list of websites that my friend was checking was also deterministic. So selection of items would only be limited to the few choices he put in. Websites themselves were heavy on JavaScript, making the whole process very slow and costly.

3:34

And finally, even if it worked, items ended up being unavailable at checkout because in the discovery phase he was not able to use geolocation capabilities. And a lot of e-commerce websites take your user's location into account when displaying stock options, sizes, and so on. Now, we solve these problems at Oxalabs every day. So when scraping, you always want the results to appear on the first try and not use a browser unless absolutely necessary. However, for this specific discovery phase, you also want to allow your agent to search the web. Doing so with a browser is very cumbersome.

3:58

That is why I chose to use a product that we built especially for agents, Fast Search API. It returns a compact JSON, which is less than 2,000 tokens per response, has fast response times, less than 700 milliseconds on average, and it has a high success rate at a predictable low price. And most importantly, it gives your agent access to many popular search engines that all of these websites have been indexed in already a long time ago. So in the discovery phase, instead of a predefined list and the browser, we give the agent a tool to search the web, Fast Search API. The agent formulates fan-out queries and selects the relevant URLs from search results.

4:10

Since the responses are quite small and there's no need for complicated models, we can have the agent run quite quickly in this stage. Now the agent has searched the web and selected some relevant URLs. It is time for the agent to visit those pages to see what they're all about in order to confirm price, stock level, description, and product details, and so on. With this, we can go to the decision phase. This is where the agent selects the items we will purchase. For this, my friend also used the browser. He ran many browsers on payroll so the whole process could happen faster, and that is not a bad thing. He managed to get some results.

4:46

However, many of the results would end up like this. As a result, the agent would be left with very few choices, with the majority of popular retailers being left out. It's a good thing he did well with observability, so he actually noticed when it happened. But what we see when working with these types of customers is that they often fail to detect the failure. They end up checking only the content size and HTTP response code and then feeding this large HTML to an LLM. Now, a large language model, of course, can distinguish between valid e-shop content and the CAPTCHA, but we need to spend tokens in order to do that.

5:09

And when we attempt to open 10 websites but only 3 return valid content, and feed all 10 to the model, it is a problem. It means that we waste 70% of the tokens. And that is a little crazy, in my opinion. So I noticed this problem as well. My initial hunch was compression. It was to compress the output. But then I thought, wait, the problem is not the compression. The problem is that the content is not valid. We need to make sure that the content is valid before even attempting any compression. This will lead to more options for the agent to choose from and fewer wasted tokens. And then I remembered rule number one of scraping. Use the browser when you absolutely need it.

6:05

Otherwise, look for other solutions. So I tried to rebuild the stage without a browser and only by using Oxlabs Web Scraper API. And this gave me many benefits. But firstly, only valid content was returned. In case of CAPTCHAs or other blocks, the request would fail with an explicit error message, so I know not to include it when sending it to a large language model. But the success rates are quite high.

6:44

And even for protected websites, that wasn't that much of a problem. So no browser was needed, and everything is a lightweight REST API. I can run hundreds of requests in parallel and receive content at the same time. Also, the API supports Markdown. So no need to submit raw HTML to LLMs. If a website is dynamic, it runs a full browser under the hood to render the content. And finally, it supports geolocation options. So I can localize my results and get relevant content. The best part? Customers only pay for successful results. Customers only pay for successful results. So actually, yeah, that's the best thing about it. No cure, no pay.

7:50

If the scraper fails, there's no cost. And it fails loudly. So now we have all of the information to make a decision.

8:06

We present the decision to the user, and the user makes the final call. Once it's affirmative, we move to the last stage of the workflow, the purchase. So I remember what I said a couple of times about browsers. But this time is different. This time, you absolutely need to use a browser. We need to process inputs, and the content is highly dynamic. Now this time, my implementation and my friend's implementation do not differ much. We both used Playwright MCP with a browser and a large language model. The main problem my friend faced, however, just like in the previous stages using a browser, was access.

8:43

Just like in the beginning, as he was using the browser, he was getting CAPTCHAs into oblivion, making it impossible to automate the flow.

8:50

Well, the fix was quite easy. I just connected the Oxlabs Headless Browser. Since it supports Playwright MCP, it's just a drop-in replacement. With this replacement, I hardened this agent with years of scraping experience and got proper stealth done at the browser source code level, a residential proxy attached to it out of the box, and, most importantly in this case, a geolocation capability. So my results are localized the same way as in the verification stage. So if we run it, we actually have a browser that accesses the content and can actually automate the flow by selecting the right size from the prompt, adding it to cart, and completing the purchase. And boom.

9:30

We have an agent that commands a powerful infrastructure, hardened by years of web scraping experience. Not only does it open up the web, but it also saves time on implementation and token cost. And if I can leave you with a few lessons we learned today, it is that when building agents, use the same principles from the scraping industry. Use the browser when you absolutely need to.

9:53

You have to validate content before feeding it to large language models. And most importantly, fill the missing layer with the proper infrastructure so you can focus on building stuff. But remember, cost matters. Thank you very much. The problem is that the content is not valid. We need to make sure that the content is valid before even attempting any compression. This will lead to more options for the agent to choose from and fewer wasted tokens. And then I remember rule number one of scraping. Use the browser when you absolutely need it. Otherwise, look for other solutions. So I tried to rebuild the stage without a browser, and only by using Oxlabs Web Scraper API.

10:36

And this gave me many benefits. But firstly, only valid content was returned. In case of CAPTCHAs or other blocks, the request would fail with an explicit error message, so I know not to include it when sending to a large language model. But the success rates are quite high. And even for protected websites, that wasn't that much of a problem. So no browser was needed. And everything is a lightweight REST API. I can run hundreds of requests in parallel and receive content at the same time. Also, the API supports Markdown. So no need to submit raw HTML to LLMs. If a website is dynamic, it runs a full browser under the hood to render the content.

11:24

And finally, it supports geolocation options. So I can localize my results and get relevant content. The best part? Customers only pay for successful results. Customers only pay for successful results. So actually, yeah. That's what's the best thing about it. No cure, no pay. If the scraper fails, there's no cost. And it fails loudly.

11:54

So now we have all of the information to make a decision. We present the decision to the user, and the user makes the final call. Once it's affirmative, we move to the last stage of the workflow, the purchase. So I remember what I said a couple of times about browsers. This time, but this time is different. This time, you absolutely need to use a browser. We need to process inputs, and the content is highly dynamic. Now this time, my implementation, my friend's implementation does not differ much. We both used Playwright MCP with a browser and a large language model.

12:37

The main problem my friend faced, however, just like in the previous stages using browser, was access. Just like in the beginning, as he was using the browser, he was getting captured into oblivion, making it impossible to automate the flow. Well, the fix was quite easy. I just connected the Oxlabs Headless Browser, since it supports Playwright MCP, it's just a drop-in replacement. With this replacement, I hardened this agent with years of scraping experience, and got proper stealth done at the browser source code level, a residential proxy attached to it out of the box, and most importantly in this case, a geolocation capability.

13:23

So my results are localized the same way as in the verification stage. So if we run it, we actually have a browser that access the content and can actually automate the flow by, you know, selecting the right size from the prompt, add it to cart, and complete the purchase. And boom! We have an agent that commands a powerful infrastructure, hardened by years of web scraping experience. Not only does it open up the web, but also saves the time on implementation and token cost. And if I can leave you a few lessons we learned today, was that, you know, when building agents use the same principles from the scraping industry. Use the browser when you absolutely need to.

14:24

You have to validate content before feeding it to large language models. And most importantly, fill the missing layer with the proper infrastructure, so you can focus on building stuff. But remember, cost matters. Thank you very much.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note