From MCP to Scale: Pipelines That Build Themselves — Rafael Levi, Bright Data
Description
Scraping is not the hard part anymore. Maintaining scrapers is. This session shows what it looks like when an agent uses MCP to inspect a site, understand its structure, generate a production scraper, and keep that pipeline working when the site changes. Using Bright Data's MCP, APIs, and browser infrastructure, the flow moves from one-off extraction to something much more useful: agents that build parsers, save tokens by switching from page parsing to reusable scripts, and repair broken collection jobs without a human getting dragged in at 2am. If you're thinking about web data, automation, or agents that operate beyond a single prompt, this is a practical look at pipeline building at scale. Speaker info: - https://il.linkedin.com/in/rafael-levi - https://github.com/ScrapeAlchemist
Summary
Generated by claude-sonnet-4-5At-a-Glance
- Verdict: Watch fully
- Core thesis: LLM agents can now build, run, and self-heal web scrapers at scale using Bright Data's MCP (Model Context Protocol), saving ~62%+ tokens vs. parsing raw HTML and eliminating manual scraper maintenance entirely.
- Why it matters: This solves the recurring cost and maintenance nightmares of web scraping at scale—your agents can write, execute, and fix scrapers autonomously, bypassing anti-bot systems (Cloudflare, Akamai, etc.) with 150M+ residential IPs and human-mimicking browser automation.
- Best use: If Ken is building agent workflows that depend on public web data (market research, price monitoring, lead gen, real-time alerts), this is essential for understanding how to delegate scraper creation/maintenance to LLMs while slashing token costs and avoiding blocks. Also watch for the live demo showing Claude Code building a Walmart scraper in minutes vs. the old multi-day manual approach.
Executive Summary
Rafael Levi (Bright Data) demonstrates how LLM agents can autonomously build production-grade web scrapers using Bright Data's MCP toolkit and GitHub skill sets, bypassing traditional manual coding and ongoing maintenance. The core workflow: the agent reads Bright Data's scraper best-practice skills, uses the MCP to fetch HTML/markdown from protected sites (Walmart, Amazon, real estate, etc.), identifies selectors, writes the scraper code, executes it, and monitors/self-heals when data validation fails. This replaces weeks of initial development and constant firefighting when selectors break.
The token economics are the headline: instead of feeding full HTML to an LLM for parsing (expensive, slow), the agent writes a lightweight scraper once and runs it repeatedly. In the demo, scraping three pages of products saved ~62% tokens vs. naive HTML parsing; for larger jobs (10K products), this compounds to ~1M token savings. The MCP provides 66 tools including curl-based fetching with automatic CAPTCHA solving, scrape-as-markdown for text extraction, 500 pre-built domain APIs (Amazon, etc.), and remote browser infrastructure with 150M+ IPs that mimic human behavior (pre-recorded mouse movements, realistic typing cadence).
Beyond enterprise scale, Rafael shows personal use cases: a scheduled apartment-hunting listener that auto-notifies when a listing under a price threshold appears; a restaurant booking bot that snags tables at packed spots. The system runs validation loops every 30 minutes, and if a data field is missing, the agent fixes the scraper in ~5 minutes without human intervention. Legal caveat: Bright Data only handles public data (no login-gated content); they've won lawsuits (Meta, Elon Musk/Twitter) establishing that public data scraping is legal, but users must still respect site ToS to avoid liability.
Live demo friction: Rafael uses Claude Code to build a Walmart scraper (later switched to very.com after cached data), showing the full flow from prompt to execution. The key takeaway: what used to take 1–1.5 days of manual selector work now takes ~3 minutes, and the scraper self-maintains. For Ken's agent ops: this is a template for offloading data collection to agents, especially for protected/geo-restricted sites, with massive token cost savings and zero manual maintenance if structured correctly.
Key Takeaways
- Claim: LLM agents can build and maintain scrapers autonomously using Bright Data's MCP and GitHub skill sets, eliminating manual coding and ongoing maintenance. | Evidence: Rafael demos Claude Code reading Bright Data's scraper best-practice repo, extracting HTML/selectors via MCP, writing a Walmart scraper in ~3 minutes vs. the old 1–1.5 day manual process; the agent monitors data every 30 minutes and self-heals within ~5 minutes if validation fails. | Caveat: The MCP is 'mostly useful in about 20% of domains'—those with heavy anti-bot systems (Akamai, Cloudflare, DataDome); simpler sites may not need this infrastructure. Also, live demo had slowdowns, and the speaker acknowledges network/conference WiFi congestion. | Implication: For Ken: if your agents need real-time or scheduled data from protected sites (real estate, e-commerce, marketplaces), Bright Data's MCP removes the scaling bottleneck and maintenance burden. You can deploy autonomous data pipelines that run indefinitely without human babysitting. | Timestamp: 01:30 / 03:45
- Claim: Using a scraper-builder approach saves ~62%+ tokens vs. feeding raw HTML to an LLM for parsing; for larger jobs (10K products), this compounds to ~1M+ token savings. | Evidence: Demo on very.com showed 62% token reduction; Rafael states 'for three pages I've done this before, it saves about a million tokens' when comparing scraper execution vs. LLM parsing every HTML page. Executing the scraper costs ~60 tokens; parsing full HTML per page costs thousands. | Caveat: The 62% figure is on the low end; Rafael notes the very.com HTML was 'maybe structured,' implying less complex sites yield lower savings. Also, the token count depends on how much data is being scraped—smaller jobs see less dramatic absolute savings. | Implication: For Ken: if you're running high-volume data ops (product catalogs, price monitoring, lead databases), switching from per-page LLM parsing to agent-built scrapers can slash your OpenAI/Claude bill by 50–90%. Build once, execute cheaply forever. | Timestamp: 15:20 / 18:45
- Claim: Bright Data's MCP provides 66 tools including curl-based fetching with auto-CAPTCHA solving, scrape-as-markdown (text-only extraction), 500 pre-built domain APIs, and remote browser infrastructure with 150M+ residential IPs that mimic human behavior. | Evidence: Rafael lists MCP tools: curl requests that solve CAPTCHAs and return HTML/markdown; pre-built APIs for Amazon, etc., so agents can skip scraper-building entirely; remote browsers with pre-recorded mouse movements and realistic typing to evade bot detection. Example: Walmart scrape without MCP gets blocked ('robot or human' screen); with MCP, it returns product data instantly. | Caveat: The MCP is limited to public data—no login-gated content. Users must still check site ToS; even though Bright Data won lawsuits (Meta, Musk), individual users can be sued if they violate explicit anti-scraping terms. Also, the 5,000 free requests require a Bright Data account (free tier). | Implication: For Ken: this MCP is a 'give your agent web superpowers' toolkit. If you're building agents that need to navigate Cloudflare-protected sites, solve CAPTCHAs, or appear human, this is the missing infra layer. The pre-built APIs mean you can skip scraper-building for ~500 major domains entirely. | Timestamp: 10:15 / 12:40
- Claim: Agents can run scheduled validation loops (e.g., every 30 minutes) to monitor scrapers and auto-fix when data validation fails, replacing manual on-call firefighting. | Evidence: Rafael: 'I have every 30 minutes, I have an LLM client's pool… checks what the data is collected, makes sure everything is fine… If a data point is missing something, your agent fixes it. Five minutes. You don't have to wake up in the middle of the night.' He also ran a personal apartment-hunting listener that notified him when a listing appeared and 'now I live there.' | Caveat: No mention of how the validation schema is defined or what happens if the agent can't fix the scraper (e.g., site structure changes too drastically). Also, the 'five-minute fix' assumes the issue is simple (missing field, broken selector); more fundamental changes might require human review. | Implication: For Ken: if you're running data pipelines for clients or internal ops, you can now build self-healing systems that eliminate weekend/night pages. Set validation rules, let the agent monitor, and only get alerts if it truly can't resolve an issue. This is ops automation 2.0. | Timestamp: 02:10 / 04:30
- Claim: Bright Data's remote browsers mimic real human behavior (pre-recorded mouse movements, realistic typing cadence, errors) to evade sophisticated bot detection systems. | Evidence: Rafael: 'When your agent clicks, it's not a teleportation. There's a mouse pre-recorded, like a real human being moving it. When it types, it will type a little slower, speed up, maybe even mistake… if the website has a tracker that's constantly sending to the server what the user is doing, it will look like a real human being.' Walmart's 'click and hold for 30 seconds' CAPTCHA is solved automatically. | Caveat: Rafael doesn't detail how the pre-recorded behaviors are generated or whether they're updated to stay ahead of evolving detection systems. Also, he notes that even 'low models' (e.g., Claude Haiku) work fine for browser automation, so you don't need GPT-4 for this. | Implication: For Ken: if your agents need to interact with sites beyond just scraping (form fills, search queries, booking flows), Bright Data's browser automation means you can skip Puppeteer/Playwright headaches and anti-bot cat-and-mouse games. The human mimicry is baked in. | Timestamp: 21:50 / 23:30
- Claim: Public data scraping is legal per court rulings (Bright Data beat Meta and Musk lawsuits), but users must respect site ToS to avoid individual liability. | Evidence: Rafael: 'We were sued by Meta. We were sued by Elon Musk a month after he took over. The judge said it's very simple. Public data is public data. It doesn't matter how you collect it… It's like walking on the street. You write down the prices on the counter and then you sell it to somebody.' But he warns: 'If they say, don't scrape, don't use robots. And you do, then that company can actually sue you.' | Caveat: The legal landscape varies by jurisdiction; these rulings apply to Bright Data's cases but aren't blanket protection for all users. Also, 'public data' is narrowly defined—anything behind login, accepted ToS checkpoints, or explicit anti-scraping clauses is off-limits. | Implication: For Ken: you can scrape public-facing data (product listings, prices, reviews) without fear, but always audit site ToS before scaling. If a client asks you to scrape login-gated data, that's a no-go. Use Bright Data's infra as a compliance moat. | Timestamp: 17:00 / 18:20
Detailed Brief
MCP-Powered Scraper Pipeline: How It Works
- Claims: Bright Data's MCP + GitHub skill sets enable LLM agents to autonomously build, run, and maintain scrapers without manual coding.; The agent reads best-practice skills from Bright Data's GitHub repo, extracts HTML/markdown via MCP, identifies selectors, writes scraper code, executes it, and self-heals when validation fails.; Demo: Claude Code builds a Walmart scraper in ~3 minutes vs. 1–1.5 days manually; later switched to very.com for a 'clean test' with no cached data.
- Evidence: Rafael: 'Build me a scraper. Two inputs. Keyword search. Max pages for walmart.com… It's going to go to our GitHub. It's going to get all the skill sets… extract the HTML… find all the selectors… build you a scraper.'; Live demo showed Claude Code reading the skills repo, calling scrape-as-markdown MCP tool, parsing selectors, writing a Python scraper with Web Unlocker API, and executing it to return JSON product data.; Validation loop: 'Every 30 minutes, I have an LLM client's pool… If a data point is missing something, your agent fixes it. Five minutes.'; Personal use case: Rafael set up an apartment-hunting listener that ran every 30 minutes, notified him of a listing, and 'now I live there.'
- Caveats: MCP is 'mostly useful in about 20% of domains'—those with Cloudflare, Akamai, DataDome; simpler sites don't need this.; Live demo had slowdowns due to conference WiFi congestion; Rafael noted 'always when you do a live, it slows down.'; No details on how validation schemas are defined or what happens if the agent can't fix a scraper (e.g., site redesign).
- Implications: For Ken: you can now delegate scraper creation/maintenance to agents, eliminating the 'wake up in the middle of the night' firefighting when scrapers break.; This is a template for self-healing data pipelines: set validation rules, let the agent monitor, only intervene if it truly fails.; If you're building agent ops for clients (e.g., e-commerce price monitoring, real estate alerts), this is your infra foundation.
Token Economics: Scraper-Builder vs. LLM Parsing
- Claims: Building a scraper once and executing it repeatedly saves ~62%+ tokens vs. feeding full HTML to an LLM for parsing every page.; For larger jobs (10K products), this compounds to ~1M+ token savings; executing a scraper costs ~60 tokens vs. thousands for LLM parsing.; Scrape-as-markdown (text-only extraction) further reduces token usage by stripping HTML tags.
- Evidence: Demo on very.com: 'It's about a 62% save of tokens. This is not a high number. This website, I guess, it has maybe a structured HTML.'; Rafael: 'For the three pages I've done this before, it saves about a million tokens just from building the scraper and using a script to parse the HTML.'; Token comparison for Walmart demo: 'To execute the script, maybe like 60 tokens… We're talking about literally a thousand tokens where if it needs to go through the JSON, it's like 10,000 tokens.'; Scrape-as-markdown: 'It's only pulling the text. It's not pulling the HTML, just the text itself… to save tokens.'
- Caveats: The 62% figure is on the low end; Rafael implies more complex/unstructured HTML sites would yield higher savings.; Token savings depend on job scale—small jobs (e.g., 10 products) see less dramatic absolute savings.; No breakdown of cost per token or total $ savings; the 1M token claim is anecdotal (Rafael: 'I've done this before').
- Implications: For Ken: if you're running high-volume data ops (product catalogs, lead gen, monitoring), switching to agent-built scrapers can cut your LLM bill 50–90%.; The ROI scales with job size: one-off scrapes see modest savings, but recurring pipelines (daily/hourly) see exponential ROI.; Use scrape-as-markdown for tasks where you only need text (e.g., summarization, search), not full DOM structure.
MCP Toolkit & Infrastructure Deep Dive
- Claims: Bright Data's MCP provides 66 tools: curl-based fetching with auto-CAPTCHA solving, scrape-as-markdown, 500 pre-built domain APIs, and remote browser infrastructure.; 150M+ residential IPs ensure scrapers/browsers appear as real users; pre-recorded mouse movements and typing cadence evade bot detection.; Remote browsers can be geo-restricted (e.g., US IP) and handle complex interactions (form fills, search queries, booking flows).
- Evidence: Rafael: 'The MCP gives you agent 66 tools… it can send a curl to any URL, and our system will literally get the HTML back. It will solve a capture if needed… It knows exactly what headers and cookies the website needs.'; Pre-built APIs: 'We have about 500 different APIs pre-built for different domains. For Amazon, we have pre-built API… it doesn't even need to build a scraper.'; Browser automation: 'I can literally right now open a thousand browsers on this laptop that are running on our servers… When your agent clicks, it's not a teleportation. There's a mouse pre-recorded, like a real human being moving it.'; Human mimicry: 'When it types, it will type a little slower, speed up, maybe even mistake… pre-recorded typing, pre-recorded mouse movements.'; Example: Walmart demo without MCP got 'robot or human' verification screen; with MCP, returned product data instantly.
- Caveats: MCP is limited to public data—no login-gated content. Rafael: 'We only deal with public data. Nothing behind login.'; 5,000 free MCP requests require a Bright Data account (free tier); no pricing details for paid tiers.; No mention of how pre-recorded behaviors are updated to stay ahead of evolving bot detection systems.; Rafael notes the MCP is 'mostly useful' for 20% of domains (Cloudflare, Akamai, etc.); simpler sites don't need it.
- Implications: For Ken: this MCP is the 'give your agent web superpowers' toolkit—bypasses anti-bot systems, solves CAPTCHAs, appears human.; If you're building agents that need to navigate protected sites (real estate, marketplaces, SaaS tools), this is your infra layer.; The pre-built APIs mean you can skip scraper-building for ~500 major domains (Amazon, etc.)—just call the API and get JSON.; For browser automation (form fills, booking flows), you can use low-cost models (Claude Haiku) since the human mimicry is baked into the infra.
Legal & Compliance: Public Data Boundaries
- Claims: Public data scraping is legal per court rulings (Bright Data beat Meta and Musk lawsuits), but users must respect site ToS to avoid individual liability.; Bright Data only handles public data—no login-gated content, no accepting ToS checkpoints.; If a site's ToS explicitly bans scraping/robots, scraping violates the ToS and opens users to lawsuits.
- Evidence: Rafael: 'We were sued by Meta. We were sued by Elon Musk a month after he took over. The judge said it's very simple. Public data is public data. It doesn't matter how you collect it. It's like walking on the street. You write down the prices on the counter and then you sell it to somebody.'; Compliance: 'We only deal with public data. Nothing behind login. We don't accept terms and conditions.'; Warning: 'If they say, don't scrape, don't use robots. And you do, then that company can actually sue you. This is happening, you probably heard in the news all the time. LinkedIn is suing them. Everybody's suing each other.'; Example: 'Elon Musk took over Twitter, locked it down. It used to be so open and now it's all locked in.'
- Caveats: The legal landscape varies by jurisdiction; Bright Data's case law applies to their specific lawsuits, not blanket protection for all users.; 'Public data' is narrowly defined—anything behind login, ToS checkpoints, or explicit anti-scraping clauses is off-limits.; No discussion of GDPR/privacy implications for scraping public data (e.g., personal info on public profiles).
- Implications: For Ken: you can scrape public-facing data (product listings, prices, reviews) without legal fear, but always audit site ToS before scaling.; If a client asks you to scrape login-gated data (e.g., private LinkedIn profiles), that's a hard no—violates Bright Data's policy and opens you to liability.; Use Bright Data's infra as a compliance moat: they've already litigated the 'public data is public' argument, so you're piggybacking on their legal wins.
Notable Concepts & Terms
- MCP (Model Context Protocol): Bright Data's toolkit that gives LLM agents 66 tools for web access: curl-based fetching with auto-CAPTCHA solving, scrape-as-markdown, 500 pre-built domain APIs, and remote browser infrastructure. The key unlock for agents to build/run scrapers autonomously.
- Scrape-as-Markdown: A Bright Data MCP tool that extracts only the text from a webpage (no HTML tags), drastically reducing token usage for LLM parsing. Used when you need content, not DOM structure.
- Self-Healing Pipeline: Rafael's term for scrapers that autonomously monitor data quality (e.g., every 30 minutes), detect validation failures (missing fields), and fix themselves without human intervention. Eliminates on-call firefighting.
- Web Unlocker API: The Bright Data API that the agent-built scraper uses to send requests and receive HTML/JSON, bypassing anti-bot systems. Same underlying tech as the MCP curl tool.
- Bright Data Skills (GitHub repo): A GitHub repository containing scraper best practices, selector identification techniques, and other skills that LLM agents read to learn how to build scrapers. The 'training data' for the agent's scraper-building workflow.
- Human Mimicry (Browser Automation): Bright Data's remote browsers use pre-recorded mouse movements, realistic typing cadence, and even typos to appear human to bot detection systems. Key for evading sophisticated trackers that monitor user behavior.
- Public Data Legal Precedent: Bright Data's court wins (Meta, Musk lawsuits) establish that scraping public-facing data is legal, regardless of collection method. However, users must still respect site ToS to avoid individual liability.
Operator Notes / Why Ken Should Care
- For agent systems: this is a reference architecture for autonomous data pipelines—agents read skills, build scrapers, execute, monitor, and self-heal. The MCP is the missing infra layer for web access at scale.
- For AI ops: the token economics are the headline—scraper-builder approach saves 62%+ tokens vs. LLM parsing, compounding to 1M+ token savings at scale. This is how you slash your OpenAI/Claude bill.
- For content/business: if you're doing market research, price monitoring, lead gen, or real-time alerts, this enables autonomous data collection from protected sites (Cloudflare, Akamai, etc.) without manual maintenance.
- For investing: Bright Data is positioned as the 'web access infra layer' for AI agents. If agents become the dominant interface for web tasks, this is a critical dependency. Their legal wins (Meta, Musk) also de-risk the compliance angle.
- For GTM: the personal use cases (apartment hunting, restaurant booking) show how low-friction this is for non-technical users. If Bright Data can package this as a no-code agent builder, it's a wedge into prosumer/SMB markets.
- For workflow: the 'build once, execute cheaply' pattern is the key—agents should write code, not parse raw data. This is the template for scaling any repetitive web task (scraping, monitoring, booking, etc.).
Watch Map
- 00:00: Intro: LLMs can build scrapers, not just parse HTML; agenda overview.
- 01:30: Bright Data Skills repo walkthrough; how agents learn to build scrapers.
- 03:45: Live demo starts: Claude Code builds Walmart scraper (later switches to very.com).
- 10:15: MCP deep dive: 66 tools, curl fetching, CAPTCHA solving, scrape-as-markdown, pre-built APIs.
- 12:40: Walmart demo without MCP (gets blocked) vs. with MCP (returns data instantly).
- 15:20: Token economics breakdown: 62% savings for very.com, ~1M tokens saved for larger jobs.
- 17:00: Legal discussion: public data is legal (Meta/Musk lawsuits), but respect site ToS.
- 21:50: Browser automation: pre-recorded mouse movements, realistic typing, human mimicry.
- 24:30: Q&A: form fills, geo-restrictions, login limitations, personal use cases (apartment hunting, restaurant booking).
Source/Metadata
- Title: From MCP to Scale: Pipelines That Build Themselves — Rafael Levi, Bright Data
- Transcript words: 6949
- Duration seconds: 1525
- Timestamp note: Timestamps are estimated based on speaker cues and demo flow; exact chapter markers were not present in the transcript.
Transcript
All righty, let's dive into it. So in the previous session, I don't know how many of you were here, but I talked about how MCP gives access to LLMs to websites that are behind capture bot detection systems, and so on. And in this session, I want to talk about how do you actually collect data at scales with LLM? Because a lot of times I'm seeing on Reddit and other social media, it's like, oh, I need to scan 10,000 products, but that's so many tokens if I need to parse everything with LLM. So obviously you don't do that, right? So the whole session is how do you build pipelines, right, with LLM? Instead of telling, hey, LLM, can you go and parse this for me, build a scrape that's going to parse it for me, and I will demonstrate how easy it is with our skills, right? So Bright Data has skill sets that actually teaches your LLM on how to build a pipeline. With our MCP, it can actually extract the HTML so that it knows what are the selectors that it needs to parse and so on. So the scrape tags, right? Write the scraper. I don't know, how many of you actually wrote scrapers before? Okay, so you know the headache when all of a sudden things are missing, data is missing, you got to wake up. I don't know if you guys did it for clients, and clients are like, oh my God, there's no... So that's what it used to be. You write a scraper and you maintain it. As a matter of fact, sometimes you maintain it more than it takes you to write it, right? So everything, especially if the website is constantly changing its selectors or if it's a React website, it gets really complicated, right? So an agent, it solves all that headache, right? It can explore with our MCP. It understands what the data is needed. It writes the scraper and actually runs it and executes it and maintains it. As a matter of fact, I do collections on a daily basis, and I have every 30 minutes, I have an LLM client's pool. It pulls up. It checks what the data is collected, makes sure that everything is fine. Everything is fine. It shuts down. If, for example, there's a always set a validation for data, right? Let's say a data point is missing something, your agent fixes it. Five minutes. You don't have to wake up in the middle of the night. So let me demonstrate. So for this demo, I want to use Claude Code. I don't know how many of you use Claude, if you're using Codex or whatever it is. I prefer Claude Code. I like how it's... I like it. It does a great job. So the way that you build scrapers now, and it's ridiculous because back in the day, it took weeks to set it up. We created a Bright Data, the GitHub page, Bright Data Skills. Here, your agent has everything that it might need. Scrape builds its best practices, how to... Everything. Everything that it needs in order to build a scraper. And what I want to show is, first of all, let's... I'm going to start simple. Go to... Build me a scraper. And I wanted to record it because I thought it was going to take a long time, but now it takes a little bit, three minutes. Build me a scraper. Two inputs. Keyword search. Max pages for walmart.com. Everybody's familiar with Walmart. Walmart has a very aggressive anti-bot systems. Anybody tried to scrape it. It just doesn't work. But through Bright Data... The website name is wrong. It might be a remote. It will figure it out, but you're right. Let's just correct that. Walmart.com. Run a search for headphones. Collect three pages. So what it's going to do right now, it's going to go to our GitHub. It's going to get all the skill sets that it needs, how to build scrapers. Then it's going to go with the MCP that's already connected to the cloud. It's going to extract the HTML from the page. It's going to find all the selectors that it needs. It's going to build you a scraper. Right now, you can see a scraper as a markdown is basically extracting the text. For the skills, we don't need the HTML. We just need the text. Scraper as a markdown is a part of the MCP of Bright Data MCP. The MCP has 5,000 requests for free. If anybody wants to try it, you guys can try it for free. You will need to just open an account with Bright Data, which doesn't cost you anything. And let it do its thing. So why is this better than having an LLM parse every single page, every HTML? I'm literally afterwards, I'm going to ask it to tell you how much tokens it saves. So for the three pages I've done this before, it saves about a million tokens just from building the scraper and using a script to parse the HTML. And even if you are not doing any scraping for production, but let's say that you want to find the best headphones for the money, right? It used to be you go to Google, you go to some CNET where they compare everything. But here you can use the marketplace reviews. So I can tell it, hey, scan the five pages, find me the best reviewed headphones. So it could be even useful for you in a personal way, right? So instead of it getting blocked, it can actually find you things that you need on websites that are behind Cloudflare. Oh, I forgot to delete the old one. No, no, stop, stop, stop. Hold on a second. Delete the old scraper. I totally, I was testing it and I totally forgot to clean up. Yeah, let's do a different website. Pick a website. What's a popular marketplace that's aggressive and blocking in UK? Amazon. Amazon, it's not that aggressive. You'll be surprised. I can scrape it with data center IPs. What's a popular website that everybody uses here? Very. Very. V-E-R-Y.com. V-E-R-Y.com. V-E-R-Y.com. Like that? Y. Y.com. Yeah. Okay. Oh, that's the other K? Okay, perfect. So this way, I don't want to... Stop it. New session. Let's do that. Let's do that. Where's the new session? Oh, there it is. Really? What is wrong with my paste? Okay. It's clothing star? Okay, so let's do headphones again. The same thing, right? So I don't know. I've never used this website. I've... Clean test. No cheating. So what else does this give you? So again, market research. For example, I was looking for a new apartment, right? I wanted to move a house. Let's do that. Let's do that. Where's the new session? Oh, there it is. Really? What is wrong with my paste? Okay. It's clothing star? Okay, so let's do headphones again. The same thing, right? So I don't know. I've never used this website. I've... Clean test. No cheating. So what else does this give you? So again, market research. For example, I was looking for a new apartment, right? I wanted to move a house. So with a Cloud Code, with Bright Data, I set up a listener. Literally just telling, Hey, listen, build me a scraper that will run every half an hour. When in this area, there's a house, private house under this price, notify me. That's all I did. In a few days, I got a notification and now I live there. So these things are not only useful on a scale where I need to scrape millions of records, but even for your personal use, it's so easy these days. Cloud Code does an amazing job, right? I know Codex does a great job as well, but I prefer Cloud Code because it just gives me less headache. Codex sometimes takes me on a wild goose chase. How many of you have Vype code? It's amazing. Why not? Right? Anybody can build anything now. So did you build it? I've got the structure, the search patterns, I don't know, with the version, let me build... Oh, okay. It's building a scraper. This used to take... I don't know, if you're having a good day, maybe a full day, maybe a day and a half, right? You need to go check out the selectors, figure it out. It used to be interesting and fun, but you don't need to do that anymore. So what I'm trying to show you guys is that with our MCP, with our infrastructure, we have over 150 million IPs, without unlocking technology, where we, even if it needs to run a remote browser, so if you want to, it can actually write you a browser automation, the browsers are running on our system. So I can literally right now open a thousand browsers on this laptop that are running on our servers, and they can go do things whatever you need it to. So it already got 90 products. Okay, it just needs to fix the Unicode because it's in pounds. Almost done. Can you also go into the ingredients of all the MCP? Sure, I did it in a previous session, but while it's loading, let me tell you. So the MCP gives you agent 66 tools. Some of them is we have a system where it can send a curl to any URL, and our system will literally get the HTML back. It will solve a capture if needed and send it with the token. It will, it knows exactly what headers and cookies the website needs. So basically, it will make sure that the server thinks it's a browser and serve the HTML back. So your agent can literally send curls, pull data without any questions. It can pull a full HTML. It can pull just the scrape as a markdown. Markdown, it's to save tokens, right? You just want the text of the page. You don't care about the HTML tags. It doesn't like the pound. Come on, you can do it. So that's number one. Second, we have about 500 different APIs pre-built for different domains. So instead of actually getting a markdown, you can actually get a JSON of the product. For example, for Amazon, we have pre-built API. You can listen, when you add your agent, it can be like, okay, go and check on Amazon something. It doesn't even need to build a scraper. It can literally just send the request and get the data back. On top of that, remote browser infrastructure and anything that needs anything that your agent needs to access the web, it has, right? So you don't get blocked. While this is running, I want to show you, for example, tell what to Walmart, do a search for headphones without MCP. So I'm telling it, go to Walmart, do a search for headphones and tell me what is the first result without the bright data MCP. It's going to do a fetch, which is going to get blocked. Product security verification screen, robot or human, obviously, right? That's the first thing. Now do the same with Brett. Oh my God. So now it's using scrape as a markdown. It's doing a search. It's only pulling the text. It's not pulling the HTML, just the text itself. Always when you do a live, it slows down. The time is like, oh, actually it could be. It could be opening a browser and actually holding the button. You know, Walmart has like one of those click and hold and it needs to hold it for 30 seconds. So it is we have built a capture solving solution. We literally have in house, we have an AI solving capture. Moving, clicking things and, wow, it's perfect. So it already got first result blue headphone. Basically there you go. It already has the results of the headphones without robot or human with full product listing name. So we here already have output. Tell me if you were to do this manually. How much more tokens have you used? How many did you save? Let's just give it a breakdown. And this is probably the biggest issue, right? Tokens is expensive. I'm going through millions and millions of tokens a day. I used to go more, but I optimized it. So all of you are looking, how do we save money on the LLM? How do we waste less tokens of web access? And [SPEAKER_01] Bright Data has the solution. Instead of parsing the full HTML, create a scraper. Instead of using, okay, it pulls the HTML and it extracts the data. It builds the parser. But it's working very slow right now. I feel like everybody's vibe coding. So basically here's the breakdown, right? What do we have here? The big swings output, parsing any products, input tokens, output tokens, total save. So it's about a 62% save of tokens. This is not a high number. This website, I guess, it has maybe a structured HTML. I'm not sure. But this is what I wanted to show you. And the best thing about it is that it will maintain it. If it breaks it, it will fix it. I can set up a loop, right? So for example, in close, I can set a schedule. Every 30 minutes, you do and run and check something. And this is what we are giving you guys. Okay? With our MCP, your agent has access to all the web, without getting blocked. No capture will stop it. No robots, So it's about a 62% save of tokens. This is not a high number. This website, I guess, it has maybe a structured HTML. I'm not sure. But this is what I wanted to show you. And the best thing about it is that it will maintain it. If it breaks it, it will fix it. I can set up a loop, right? So for example, in close, I can set a schedule. Every 30 minutes, you do and run and check something. And this is what we are giving you guys. Okay? With our MCP, your agent has access to all the web without getting blocked. No capture will stop it. No robots, anti-bot systems will stop it. And you have no headaches of that. I don't know how many of you have, you guys probably dealt with blocking, right? When you build a scraper, it gets really complicated. So this kind of removes all the headache for it. Any questions? Where is the code? Excuse me? The code, the code to generate the scraper, right? What is it? Oh, I can, if you want to see the code? It's here. We have the scraper. Let's check it out. Let's check it out. So it's using API Web Unlocker. This is the same thing as a scraper's markdown with an MCP. It sends a simple request and gets the data back. This is the parser that it built. It looks pretty good. This is the schema that it built for the output. It's doing a keyword search. It has the max page, right? So I also, for two inputs, the keyword search and how many pages to scrape. And hold on. Let's see. Let's see. We can run it through here as well. Where's the variable? API key, zone key. Let's run it. Do another run on laptops. What are the two pages results? And now you kind of control it. You don't even need to run it, right? It will trigger it for you. It will maintain it for you. You can create a script that will automatically do all of this. I prefer the visual interface for the demonstration. Sometimes looking at our code, it's a little more complicated. But right now, I can ask it any questions. It will actually run the script. So it's not wasting tokens. To execute the script, maybe like 60 tokens, right? Right now. So it's going to literally get all this data for a hundred tokens. And then you can do whatever. You can ask it questions about it. But now we're working with a JSON. JSON is much more token efficient than any markdown or any other file, right? It's a little bit than usual. Can you do this on all, with authorization as well? No, we only talk, we only deal with public data. So behind login, it's private data. And just a heads up. If you create an account, you accept terms and conditions of the website. Always check terms and conditions of the website. If they say, don't scrape, don't use robots. And you do, then that company can actually sue you. And this is happening, you probably heard in the news all the time. LinkedIn is suing them. Everybody's suing each other because the data right now is like the new gold, the oil, right? Everybody's trying to say, oh, this is my data. Elon Musk took over Twitter, locked it down. That's it. Literally, it used to be so open and now it's all locked in. A few accounts that you can actually scrape. So we're seeing this on every single level. Very slow today. And so yes, also same thing. If you accept, there's like a terms and if you do a search and there's a checkpoint, I accept terms and conditions. Always be careful with that. So we only deal with public data. Nothing behind login. We don't accept terms and conditions. And we actually won a few lawsuits. We were sued by Meta. We were sued by Elon Musk a month after he took over. And the judge said it's very simple. Public data is public data. It doesn't matter how you collect it. Doesn't matter what you do with it. It's public. You know, it's like walking on the street. You write down the prices on the counter and then you sell it to somebody. It's public data. It's available. You can do whatever you want with it. Okay. So as usual, it gets stuck on a live demo, but we can run this manually. Oh my God. What the? I'm setting a variable. It's doing this. What am I missing? Still stuck? Oh, okay. It's done. Okay. So here we got, we got, let's see how many tokens, how many tokens did you spend on getting these? So even if you're not building pipelines for any big companies, then you don't need to get for personal use. I always have the MCP connected. I don't ask the LLM to go do. I always ask, build a script that can later on be used by it. It's using its own script to save the tokens. Literally, we're talking about a thousand tokens where if it needs to go through the JSON, it's like 10,000 tokens. We're talking about literally pennies compared to what it would be to actually scrape it. Any more questions? How about this? Anybody has problems with actually getting data? No, no question. Everybody is like, oh yeah, I'm not sharing my problems. I have no problems. And then they go home and like, oh, I do have a problem. But in any case, listen, we have a booth. If you don't, if you want to talk in private and you have questions about data or access or LLM, feel free. I'm always willing to help. We can connect on LinkedIn, by the way, guys, if you want. Let me just summarize a few things, what we did here. So self-healing pipeline. When you have, when your agent actually has access to a blocked website and you're working, the MCP is mostly useful in about 20% of the domains, right? The ones that have Akamai, DataDome, Cloudflare, those heavy protected domains. They're usually the most juiciest one, right? Like real estate or big e-commerce places. With RMCP, it has access to it. If it has access to it, it can build a scraper. If it can build a scraper, it can maintain the scraper. This is basically the whole thing that I wanted to show you that in Bright Data, we have all the tools that your agent might need to explore, to build and maintain a pipeline, right? So if you want to set up a listener for an apartment and you're useful in about 20% of the domains, right? The ones that have Akamai, DadaDone, Cloudflare, those heavy protected domains. They're usually the most juicy ones, right? Real estate or big e-commerce places. With RMCP, it has access to it. If it has access to it, it can build a scraper. If it can build a scraper, it can maintain the scraper. This is the whole thing that I wanted to show you—that in Bright Data, we have all the tools that your agent might need to explore, to build and maintain a pipeline, right? So if you want to set up a listener for an apartment and you're looking to move into a cheaper place, or maybe you want to book a table in a restaurant where it's always packed, and I have a listener right now. I'm waiting for two months already. Everybody's booking right away. So as soon as this spot opens, it will automatically book a spot for me. It can be very useful in just small things, personal things, right? I'm not talking about just enterprise scale where I need to download a million records, okay? But having access to all the websites, it's important. So apart from scraping, can you also do actions on the website? Yes, of course. Fill up a form. Fill up a form, submit it as well, yes. The only thing you can do is log in, right? So if you have, let's say, you need to perform a search, right? You cannot generate the URL, you need to click buttons, let's say flights. You want to check flights availability, Skyscanner or something like that, right? The URL is usually a hash and you can't do anything with it. So yes, LLM can spool a browser, remote browser. Even if it's a geo-restricted site, it can be like, okay, I want IP from the United States. So it opens a browser with the United States IP and then it goes and clicks things, fills it up and so on. The beauty of it is that our browser will mimic real human behavior. So when your agent clicks, it's not a teleportation. There's a mouse pre-recorded, like a real human being moving it. When it types, it will type a little slower, speed up, maybe even mistake and so on. So we have pre-recorded typing, we have pre-recorded mouse movements. So if the website has a tracker that's constantly sending to the server what the user is doing, it will look like a real human being. It doesn't matter what your agent is, even if it's a low model. For browsing agents, I've never used the top models, right? For example, with a high-eco model, there's more than enough for browsing. If it's being masked that it's a real human, it works just fine. And so that's pretty much it. If you guys want to connect on LinkedIn, feel free. I'm always willing to help. If you have any questions regarding how to get data, if you have problems with accessing anything, feel free to message me. I feel like there's been so much more than 15 minutes. I love it. It's the second time I'm doing a speech and I got like half an hour instead of 15 minutes. It's amazing. Now's the time if you have any questions, discussions. No, nothing? Great. Okay, guys, I guess I will conclude this session. Feel free to come to the booth on the third floor if you have more questions. And let's keep the public data public. Thank you. Thank you. Yeah, let's do a different website. Pick a website. What's a popular marketplace that's aggressive and blocking in UK? Amazon. Amazon, it's not that aggressive. You'll be surprised. I can scrape it with data center IPs. What's a popular website that everybody uses here? Very. Very. V-E-R-Y.com. V-E-R-Y.com. V-E-R-Y.com. Like that? Y. Y.com. Yeah. Okay. Oh, that's the other K? Okay, perfect. So this way, I don't want to... Stop it. New session. Let's do that. Let's do that. Where's the new session? Oh, there it is. Really? What is wrong with my paste? Okay. Uh, it's clothing star? Okay, so let's do headphones again. The same thing, right? So I don't know. I've never used this website. I've... Clean test. No cheating. Um... So what else does this give you? So again, market research. Um... For example, I was looking for a new apartment, right? I wanted to move a house. So with a Cloud Code, with Bright Data, I set up a listener. Literally just telling, Hey, listen, build me a scraper that will run every half an hour. When in this area, there's a house, private house under this price, notify me. That's all I did. In a few days, I got a notification and now I live there. So these things are not only useful on a scale where, you know, I need to scrape millions of records, but even for your personal use, it's so easy these days. Cloud Code does an amazing job, right? I know Codex does a great job as well, but I prefer Cloud Code because, I don't know, it just, uh, gives me less headache. Codex sometimes takes me on a wild goose chase. How many of you have Vype code? It's amazing. Why not? Right? Anybody can build anything now. Um... So did you build it? I've got the structure, the search patterns, I don't know, with the version, let me build... Oh, okay. It's building a scraper. Um... This used to take... I don't know, if you're having a good day, maybe a full day, maybe a day and a half, right? You need to go check out the selectors, figure it out. It used to be interesting and fun, but you don't need to do that anymore. So what I'm trying to show you guys is that with our MCP, with our infrastructure, we have over 150 million IPs, uh, without unlocking technology, where we, even if it needs to run a remote browser, so if you want to, it can actually write you a browser automation, uh, the browsers are running on our system. So I can literally right now open a thousand browsers on this laptop that are running on our servers, and they can go do things whatever you need it to. So it already got 90 products. Okay, it just needs to fix the Unicode because it's in pounds. Eh, eh, eh, eh, eh, eh, almost done. Can you also, uh, go into the ingredients of all the MCP? Sure, I did it in a previous session, but while it's loading, let me tell you. So basically, the MCP gives you, uh, uh, agent 66 tools. Uh, some of them is basically, uh, we have a system where it can send a curl to any URL, and our system will literally get the HTML back. It will solve a capture if needed and send it with the token. It will, uh, it knows exactly what headers and cookies the website needs. So basically, it will make sure that the server thinks it's a browser and serve the HTML back. So your agent can literally send curls, pull data without any questions. It can pull a full HTML. It can pull just the scrape as a markdown. Markdown, it's, it's to save tokens, right? You just want the text of the page. You don't care about the HTML tags. It doesn't like the pound. Come on, you can do it. Um, so that's number one. Second, we have about 500 different APIs pre-built for different domains. So instead of actually getting a markdown, you can actually get a JSON of the product. For example, for Amazon, we have pre-built API. You can listen, uh, when you add your agent, it can be like, okay, go and check on Amazon something. It doesn't even need to build a scraper. It can literally just send the, uh, the request and get the data back. Um, on top of that, remote browser infrastructure and, uh, anything that needs, uh, anything that your agent needs to access the web, it has, right? So you don't get blocked. While this is running, I want to show you, for example, uh, tell, uh, what, to Walmart, do a search for headphones without MCP. So I'm telling it, go to Walmart, do a search for headphones and tell me what is the first result without the bright data MCP. Um, it's going to do a fetch, which is going to get blocked. Product security verification screen, robot or human, obviously, right? That's, that's the first thing. Now do the same with Brett. Um, oh my God. So now it's using scrape as a markdown. It's doing a search. It's only pulling the text. It's not pulling the HTML, just the text itself. Always when you do a live, it slows down. The time is like, Oh, actually it could be. It could be opening a browser and actually holding the button. You know, the Walmart has like one of those click and hold and it needs to hold it for like 30 seconds. So it is like, we have built a capture solving solution. Like we literally have in house, we have an AI solving capture. Moving, clicking things and like, wow, it's perfect. So it already got first result blue headphone. Basically there you go. It already has the results of the headphones without robot or human with full product listing name. So we here already have output. Tell me if you were to do this manually. How much more tokens have you used? How many did you save? Let's just give it a breakdown. And this is probably the biggest issue, right? Tokens is expensive. I'm going through millions and millions of tokens a day. I used to go more, but I optimized it. So, um, all of you are looking, how do we save money on the LLM? How do we waste less tokens of web access? And, um, Bright Data has the solution. Instead of parsing the full HTML, create a scraper. Instead of using, um, you know, okay, it pulls the HTML and it extracts the data. It builds the parser. But it's working very slow right now. I feel like everybody's vibe coding. So basically here's the breakdown, right? What do we have here? The big swings output, parsing any products, input tokens, output tokens, total save. So it's about a 62% save of tokens. This is not a high number. This website, I guess, it has, uh, maybe a structured, um, HTML. I'm not sure. But this is what I wanted to show you. And the best thing about it is that it will maintain it. If it breaks it, it will fix it. I can set up a loop, right? So for example, in close, I can set a schedule. Every 30 minutes, you do and run and check something. And this is what we are giving you guys. Okay? With our MCP, your agent has access to all the web, uh, without getting blocked. No capture will stop it. No robots, anti-bot systems will stop it. And, uh, you have no headaches of that. I don't know how many of you have, you guys probably dealt with blocking, right? When you build a scraper, it gets really complicated. So this kind of removes all the headache for it. Um, any questions? Where is the code? Excuse me? The code, the code to generate the scraper, right? Uh, what is it? Oh, I can, uh, if you want to see the code? I mean, it's here. We have the scraper. Let's check it out. Let's check it out. So it's using API Web Unlocker. This is the same thing as a scraper's markdown with an MCP. It sends a simple request and gets the data back. This is the parser that it built. I mean, it looks pretty good. This is the, uh, the schema that it built for the output. It's doing a keyword search. It has the max page, right? So I also, for two inputs, the keyword search and how many pages to scrape. And hold on. Let's see. Let's see. We can run it through here as well. Where's the variable? API key, zone key. Let's run it. Do another run on laptops. What are the two pages results? And now you kind of control it. Like you don't even need to run it, right? It will trigger it for you. It will maintain it for you. Um, you can create a script that will automatically do all of this. I prefer the visual interface for the demonstration. Sometimes they, you know, like looking at our code, it's a little more complicated. But right now, basically I can ask it any questions. It will actually run the script. So it's not wasting tokens. To execute the script, maybe like 60 tokens, right? Right now. So it's going to literally get all this data for a hundred tokens. And then you can do whatever. You can ask it questions about it. But now we're working with a JSON. JSON is much more token efficient than any markdown or any other file, right? It's a little bit than usual. Can you do this on all, with authorization as well? And no, we only talk, we only deal with public data. So behind login, it's, um, it's private data. And, uh, just, you know, just a heads up. If you create an account, you accept terms and conditions of the website. Always check terms and conditions of the website. If they say, don't scrape, don't use robots. And you do, then that company can actually sue you. And this is happening, you probably heard in the news all the time. LinkedIn is suing them. Everybody's suing each other because they, the data right now is like the new goal, the oil, right? Everybody's trying to like, oh, this is my data. Elon Musk took over Twitter, locked it down. That's it. Literally, like, uh, it used to be so open and now it's all locked in. A few accounts that you can actually scrape. So we're seeing this on every single level. Very slow today. And so, yes, also same thing. If you accept, uh, there's like a terms and like, if you do a search and there's a checkpoint, I accept terms and conditions. Always be careful with that. So we only deal with public data. So nothing behind login. We don't accept terms and conditions. And we actually won a few lawsuits. We were sued by Meta. We were so sued by Elon Musk a month after he took over. And, uh, the judge said it's very simple. Public data is public data. It doesn't matter how you collect it. Doesn't matter what you do with it. It's public. You know, it's like walking on the street. You write down the prices on the counter and then you sell it to somebody. It's public data. It's available. You can do whatever you want with it. Okay. So as usual, it gets stuck on a live demo, but we can we can run this manually. Oh my God. What the? I'm setting a variable. It's doing this. What am I missing? Still stuck? Oh, okay. It's done. Okay. So here we got, we got, let's see how many tokens, how many tokens did you spend on getting these? So even if you're not building pipelines for any, uh, you know, big companies, then you don't need to get for personal use. I always have the MCP connected. I don't ask the, uh, you know, LLM to go do. I always ask, build a script that can later on be used by it. It's using its own script to save the tokens. Literally, like we're talking about a thousand tokens where if, uh, you know, if, if it needs to go through the Jason, it's like 10,000 tokens. We're talking about like literally pennies compared to what it would be to actually scrape it. Any more questions? How about this? Anybody has problems with actually getting data? No, no question. Everybody is like, oh, yeah, I'm not sharing my problems. I have no problems. And then they go home and like, oh, I do have a problem. But in any case, listen, we have a booth. If you don't, if you want to talk in private and you have questions about data or access or LLM, feel free. I'm always willing to help. Um, we can connect on LinkedIn, by the way, guys, if you want. Let me just summarize a few things, what we did here. So, um, self-healing pipeline. When you have, when your agent actually has access to a blocked website and you're working, I mean, the MCP is mostly useful in about 20% of the, of the domains, right? The ones that have Akamai, DadaDone, Cloudflare, those heavy protected domains. They're the, usually the most juiciest one, right? Like real estate or big e-commerce places. Uh, with RMCP, it has access to it. If it has access to it, it can build a scraper. If it can build a scraper, it can maintain the scraper. This is basically the whole thing that I wanted to show you that in Bright Data, we have all the tools that your agent might need to explore, to build and maintain a pipeline, right? So if you want to set up a listener for an apartment and you're looking to move into a cheaper place, or maybe you want to book a table in a restaurant where it's always packed, and literally, I, I, I have a listener right now. I'm waiting for two months already. Everybody's like booking right away. So as soon as this spot opens, it will automatically book a spot for me. It can be very useful in just kind of like, even the small things, personal things, right? I'm not talking about just enterprise scale where I need to download a million records, okay? But having access to all the websites, it's, it's, it's important. So apart from scraping, can you also do actions on the website? Yes, of course. Fill up a form. Fill up a form, submit it as well, yes. The only thing you can do is log in, right? So if you have, let's say, you need to perform a search, right? You cannot generate the URL, you need to click buttons, let's say flights. You want to check flights availability, Skyscanner or something like that, right? The URL is usually a hash and you can't do anything with it. So yes, LLM can spool a browser, remote browser. Even if it's a geo-restricted site, it can be like, okay, I want IP from the United States. So it's open a browser with the United States IP and then it goes and clicks things, fills it up and so on. The beauty of it is that our browser will mimic real human behavior. So when your agent clicks, it's not a teleportation. There's a mouse pre-recorded, like a real human being moving it. When it types, it will type a little slower, speed up, like maybe even mistake and so on. So we have pre-recorded typing, we have pre-recorded mouse movements. So if the website has a tracker that's constantly sending to the server what the user is doing, it will look like a real human being. It doesn't matter what your agent, even if it's a low, like for browsing agents, I've never used the top models, right? For example, if with a high-eco model, there's more than enough for browsing. If it's being masked that it's a real human, it works just fine. And so that's pretty much it. If you guys want to connect on LinkedIn, feel free. I'm always willing to help. If you have any questions regarding how to get data, if you have problems with accessing anything, feel free to message me. I feel like there's been so much more than 15 minutes. I love it. It's the second time I'm doing a speech and I got like half an hour instead of 15 minutes. It's amazing. Now's the time if you have any questions, discussions. No, nothing? Great. Okay, guys, I guess I will conclude this session. Feel free to come to the booth on the third floor if you have more questions. And let's keep the public data public. Thank you. Thank you.