SPEAKER_01
Okay, so let's just work with it like this, a small room. Hi everybody, welcome. My name is Rafael, representing Bright Data. Bright Data is a web access platform to help agents or anybody collect public data at scale. And I'm here to talk about LLMs misleading people all the time, convincing them, "Hey, I did a search, hey, I did this," while they didn't. Why? Because LLMs are programmed to please people, please users, so they make up things. And this is the biggest issue right now I'm seeing with LLMs. I'm building applications all the time. I would rather an LLM tell me no, I can't, but it never does. It always tries to make things up. So currently, the web is actually fighting robots and automations. And it's been something that's going on for years. Everybody knows CAPTCHAs. The first CAPTCHAs showed up a decade ago. And it just keeps growing. And now we have AI blocking AI, and there's a whole world going on. And web access is actually not as simple as it looks. So they're getting CAPTCHAs, and they don't actually report CAPTCHAs. So it tries to find the data a different way. Sometimes it goes into training data, and this is where the worst thing is, when it uses training data and tells you that this is the current situation. Training data is from 2024, we're in 2026, and it doesn't add up, right? So these are some of the new things that I was just literally checking out, right? So Cloudflare blocks AI crawling for about 20% of the web, right? So 20% of the web is literally not accessible by AI by default fetch that's built into it. And Cloudflare also right now released an AI labyrinth to actually trap bots and mislead them and provide to them fake data. So then your results are getting even worse, okay? So the invisible failure route, right? There's no error, no warning, just wrong answer, right? So the agent sends the request, it gets a CAPTCHA, even an empty page. It doesn't tell you, "Hey, I got an empty page." It will try to make up something, and this is where most of the hallucinations come from. The need to please and lack of data. So it literally makes things up. I've seen it literally make numbers up, provide fake citations. You click on the citation, it's a 404, the page doesn't exist, and I'm sure all of you have seen that happening recently. I mean, literally, 60% of citations on ChatGPT is not working. How many of you have tried to purchase a product, "Hey, find me this product online, I want to buy it and give me a link to the product." You click on the link to the product and there's no product. Like, so what is this product that you're talking about? The URL doesn't exist, the product doesn't exist, so where do I buy this product for the 50 bucks? It doesn't exist. And so what I want to show you, and I'm going to show you in code, how many of you are familiar with coding like VS code? No. Okay, perfect. So nobody's going to get lost if I'm going to switch to that. So I'm going to do a demo basically with MCP and without the Bright Data MCP, and we're going to compare what is going on there. So first I'm going to show you that I have exactly identical prompts for both of the scripts, right? So without MCP and with MCP. So I give it five tasks. Property, right, move.co. Let's go check out some of the properties. LinkedIn, Instagram, Amazon, and TikTok. These are the five basic sites that I wanted to access. They are very heavy on anti-bot systems. And I want to show you the difference. So first I'm going to run it without an MCP. And I'm going to literally let the AI talk for itself. And GPT-5 is a bit slow, so it's going to take some time. But basically what we're trying to do is we're trying to access this URL. It's some local properties that I did literally half an hour ago. LinkedIn company in Israel. Let's check it out. Instagram account, some Amazon product, and some TikTok. And this is not limited to just these five websites. It's just something that I picked that I know for sure will not work without MCP. So as you can see, without MCP, I don't have live web data access. It doesn't have any browsing tools, right? So this is with tools not available, just by default, out of the box GPT-5. I mean, it's a strong LLM. And zero success, five failed. Same exact thing, exactly the same prompt. I'm running with our MCP. Our MCP has 66 tools. While it's running, I just want to go over some of the tools that it has. Search engine. Search engine is basically the LLM is able to do Google search, Bing search, DuckDuckGo search. Real searches, not just search the web that it's doing in the background. It has scrape as markdown. That's a very strong one. Basically, it can send a curl to any URL and get just the markdown without HTML tags. So you're not wasting tokens on parsing the HTML. Search engine batch, if you want to do like a hundred keywords, you can literally send a hundred keywords and get the hundred keyword backed results. So the scaling is also huge. Discover. It also has pre-built APIs for many websites, as you can see. And of course, it has a scraping browser infrastructure. So it's a remote browser that your LLM can open and navigate. The remote browser solves CAPTCHA by itself. It can open a hundred browsers, navigate the same website without getting blocked. So the whole idea is that with our MCP, not only can it do a single session, it can do multiple sessions in parallel. And as you can see, let's see, we have success for the Rightmove. We have success for the LinkedIn in Israel. Instagram also worked. Amazon product, it has the information for the product itself. And then what I did is on the second part, I asked the LLM to compare the results from no MCP with MCP. So that you don't take my word for it. Let's see what ChatGPT will actually say. If it didn't get stuck, it looks like it's stuck for some reason.
SPEAKER_01
Can I ask a question? Of course, please. I love questions. For social profiles, do you have accounts that you run? No, we only work with public data. Only publicly available data. Collecting data behind login is not really legal. Why? Because when you sign up and create an account, you accept terms and conditions. When you accept terms and conditions, you need to really check if it says can you scrape? Can you, do you allow, do they allow robots to access the website? And that's why maybe some of you heard there's a lot of lawsuits going on. LinkedIn suing these people. Everybody's suing you. Amazon suing, yes. Can I ask a question? Of course, please. I love questions.
SPEAKER_01
For social profiles, do you have accounts that you run?
SPEAKER_01
No, we only work with public data. Only publicly available data. Collecting data behind login is not really legal. Why? Because when you sign up and create an account, you accept terms and conditions. When you accept terms and conditions, you need to really check if it says, can you scrape? Can you, do they allow robots to access the website? And that's why maybe some of you heard there's a lot of lawsuits going on. LinkedIn suing these people. Everybody's suing you. Amazon suing, yes. But I think even, for example, LinkedIn and Instagram, I think we don't even show you really any public data, even if you're on the... That's plenty of public data for LinkedIn, of course. If you take, for example, I will take this URL. I might get blocked. I don't know how the local IP is, but... And I open an incognito window, right? Usually, just like a video. There you go. So this is public data that can be collected. So if you put a person instead of... Company, it's the same thing. It's the same thing. The only thing is they're very critical, right? So if you are using, let's say, a Wi-Fi IP of a big event like this, it probably will block you because it's a data center IP. It's low quality IP. From home, you could probably access maybe 5-10 profiles, but eventually it will also ask you to log in. So we only deal with public data. So when you guys are using us, you are safe in the sense of nobody's going to come knocking on your doors to sue you. If that makes sense.
SPEAKER_01
Can we provide credentials to our own social networks? No. Again, we don't deal with data behind login. We consider it to be illegal. So we don't deal with accepting terms and conditions and data behind login. Only publicly available. So I imagine you would cache some of these so that you don't have to constantly go and request it. Do you have an idea of how...
SPEAKER_01
We have a whole data set. So if you guys don't want to do live data and you don't care if it's a few months old, we have data sets, literally, that you can just filter by, let's say, that you're looking... If we're talking about LinkedIn people, you're looking for AI engineers in a certain area, you can filter it and just get the data set right away. And your agent will actually have access to that. So your agent can filter the data set and get you the data if we want to, right? So it has access to all the tools. So he has basically had to get comparison by the LLM itself. Without MCP failed, no live web access. With it listed, failed successful, failed successful. So that's basically it. Anti-bot bypass, capture solving, right? So again, our system automatically solves capture. So if your bot navigates to a website with a browser and it has a capture, our browser has built-in capture solving solution. So it will automatically solve the capture and your bot can continue on browsing without getting blocked.
SPEAKER_01
How much time do I have? I don't know.
SPEAKER_01
So just to summarize it, right? This is the biggest hallucinations that you guys see. The agent gets blocked, it needs to please you, and it makes things up. And fake content also, right? So now, if you want to Google what is Cloudflare AI Labyrinth, it's basically a system. Once it detects a bot, it doesn't block it. It literally feeds it fake data. So bigger hallucinations, right? And the easiest fix for this is just to make sure that your agent doesn't get blocked. And it's as easy as to implement our MCP. Our MCP has a free tier of 5,000 requests. So if you guys want to try it out, you can connect with me on LinkedIn if you want to. Or is this the... Hold on a second. Is this the sign up for the MCP? I'm lost in these QR codes. One second.
SPEAKER_01
Internet. Yes, please. How does it detect if there's Cloudflare Labyrinth or not?
SPEAKER_01
So the way we approach it is that we make your agent look like a human being. Literally, there's mouse movement recorded. There's typing. When it types, it's like... Mimics real human behavior. So Cloudflare literally just doesn't even ask, are you a robot or not? Right? So this is our approach. Instead of trying to understand how they detect, we make the agent look as human as possible so that it doesn't trigger the actual blockage. Misleading data is one of the toughest things that you can actually encounter. A lot of websites right now in Asia are doing that, right? Hotels, they're literally providing different prices. You check out on your phone, you get one price. You check from your computer, you get a different price. You add through proxy, you get a third price. Which one is correct? It's really hard to tell. When it comes into the domain of misleading, the best bet is to make sure that your agent looks like a human and hope for the best. That's basically what the approach is right now. AI Labyrinth was literally released a month ago. I don't have much statistics on exactly how it's working. All I know is that it didn't really affect us. We don't see any change in data and we are collecting petabytes of data on the daily. We have so many customers always scraping. We're caching data so we're always comparing the results. We don't see any degradation in the results. So I think we're doing a good job in that case.
SPEAKER_01
This is a QR code that you guys can sign up for. We have a GitHub page as well. GitHub bright data.com. GitHub bright data. And another thing that I would recommend for you guys to check out is the skills page. Right? So what we did is we created skills. And I'm going to have another session in a couple of hours if you're interested in seeing what the skills does. It's basically you can take any agent, tell it to go here and this page will teach it on how to build a scraper or how to build a pipeline that will collect you the data. It has all the information it needs, all the APIs. And if you come back to my next session which is at one o'clock I think? Something like that. I'm going to literally demonstrate how it builds the pipeline. Literally in front of you I'm going to tell it, hey listen let's build a Walmart collector for ABC and it's going to build it and it's going to scrape it. And instead of parsing each individual in ChiptML it's going to build a parser and it saves about 99% of the tokens. Because I see a lot of people like, hey I need to parse 10,000 pages but it's so token heavy. Don't parse with the LLM. LLM builds the parser and then the script runs it. But that's the next session. Any questions?
SPEAKER_01
session which is at one o'clock I think? Something like that. I'm going to literally demonstrate how it builds the pipeline. Literally in front of you I'm going to tell it, hey listen let's build a Walmart collector for ABC and it's going to build it and it's going to scrape it. And instead of parsing each individual in ChiptML it's going to build a parser and it saves about 99% of the tokens. Because I see a lot of people say, hey I need to parse 10,000 pages but it's so token heavy. Don't parse with the LLM. LLM builds the parser and then the script runs it. But that's the next session. Any questions?
SPEAKER_01
Question on the performance. I saw that you have CP exposes 69 tools which means if I need to say search capability for my agent, do you need to load all the 69 tools? No. Of course, filter it. I just showed 69 tools because just to show it, if I need just a scrape markdown and search I would just literally load two tools. Otherwise you're flooding contacts with relevant data. Of course not. The experiment at the beginning, does that use the Web search tool? I'm sorry, I didn't hear at the beginning again. The experiment at the beginning, does that use the Web search tool through the OpenAI API? Or is it like for the . And so how does it compare? I'm not that experienced.
SPEAKER_01
[SPEAKER_04] . So we didn't do any searches. I literally told it, hey go to this URL, see if you can load it. Right? So I didn't use the search in this demo. [SPEAKER_04] Sure. But of course, again, even with our MCP, they can actually do a Google search. Yeah. And that's one of the biggest benefits is because we're used to Google results. Right? So by default, when you're asking LLM, you're expected to do a Google search, but it doesn't. Right. So the results with the MCP, much, much better. And I recommend, sign up, try it out, it's free. See the results, compare what you guys get. Is that 5,000 requests per day or...? Per month. Per month.
SPEAKER_01
Per month. Which is, you know, for an MVP, for a little experiment, it's more than enough. And we also have pay as you go, so if you do need a little more, it's nothing. Okay. Just for running a set, just for paying a run. Yeah, yeah, it's for pro-type, it's perfect. I always do a lot of hackathons and I'm always recommending, hey listen, set up an MCP and tell your agent to go build whatever you need. [SPEAKER_03] It does a much better job than without. Forget it. Thank you. trying to access this URL. It's some local properties that I did literally half an hour ago. LinkedIn company in Israel. Let's check it out. Instagram account, some Amazon product, and some TikTok.
SPEAKER_01
And this is not limited to just these five websites. It's just something that I picked that is, I know for sure, will not work without MCP. So as you can see, without MCP, I don't have live web data access. It doesn't have any browsing tools, right? So this is with tools not available, just by default, out of the box. GPT-5. I mean, it's a strong LLM. And zero success, five failed. Same exact thing, exactly the same prompt. I'm running with our MCP.
SPEAKER_01
Our MCP has 66 tools. While it's running, I just want to go off some of the tools that it has. Search engine. Search engine is basically the LLM is able to do Google search, Bing search, DuckDuckGo search. Real searches, not just like, you know, search the web that it's doing in the background. It has...
SPEAKER_01
It also has scrape as a markdown. That's a very strong one. Basically, it can send a curl to any URL and get just the markdown without HTML tags. So you're now wasting tokens on parsing the HTML. Search engine batch, if you want to do like, you know, like a hundred keywords, you can literally... It can literally send a hundred keywords and get the hundred keyword backed results. Like, so the scaling is also huge. Discover, it also has pre-built APIs for many websites, as you can see. And of course, it has a scraping browser infrastructure. So it's a remote browser that your LLM can open and navigate. The remote browser solves capture by itself. It
SPEAKER_01
can open a hundred browsers, navigate the same website without getting blocked. So the whole idea is that with our MCP, not only can it do a single session, it can do multiple sessions in parallel. And as you can see, let's see, we have success for the right move. We have success for the LinkedIn Misrael. Instagram also worked. Amazon product, it has the information for the product itself. And then what I did is on the second part, I asked the LLM to compare the results from no MCP with MCP. So that you don't take my word for it. Let's see what the chat GPT will actually say. If it didn't get stuck, it looks like it's stuck for some reason.
SPEAKER_01
Can I ask a question? Of course, please. I love questions. For social profiles, do you have accounts that you run? No, we only work with public data. Only publicly available data. Collecting data behind login is not really legal. Why? Because when you sign up and create an account, you accept terms and conditions. When you accept terms and conditions, you need to really check it if it says, can you scrape? Can you, do you allow, do they allow robots to access the website? And that's why maybe some of you heard there's a lot of lawsuits going on. LinkedIn suing these people. Everybody's suing you. Amazon suing, yes.
SPEAKER_01
But I think even, for example, LinkedIn and Instagram, I think we don't even show you really any public data, even if you're on the... That's plenty of public data for LinkedIn, of course. If you take, for example, I will take this URL. I might get blocked. I don't know how is the local IP, but... And I open an incognito window, right? I mean, usually, just just like a video. There you go. So this is a public data that can be collected. So if you put a person instead of... Company, it's the same thing. It's the same thing. The only thing is they're very critical, right? So if you are using, let's say, a Wi-Fi IP of a big event like this, it probably will block you
SPEAKER_01
because it's a data center IP. It's low quality IP. From home, you could probably access maybe 510 profiles, but eventually it will also ask you to log in. So we only deal with public data. So when you guys are using us, you are safe in the sense of nobody's going to come knocking on your doors to sue you. If that makes sense. Can we provide credentials to our own social networks? No. Again, we don't deal with data behind login. We consider it to be illegal. So we don't deal with accepting terms and conditions and data behind login. Only publicly available. So I imagine you would cache some of these so that you don't have to constantly go and request it.
SPEAKER_01
Do you have an idea of how... We have a whole data set. So if you guys don't want to do live data and you don't care if it's a few months old, we have data sets, literally, that you can just filter by, let's say, that you're looking... If we're talking about LinkedIn people, you're looking for AI engineers in a certain area, you can filter it and just get the data set right away. And your agent will actually have access to that. So your agent can filter the data set and get you the data if we want to, right? So it has access to all the tools. So he has basically had to get comparison by the LLM itself.
SPEAKER_01
Without MCP failed, no live web access. With it listed, failed successful, failed successful. So that's basically it. Anti-bot bypass, capture solving, right? So again, our system automatically solves capture. So if your bot navigates to a website with a browser and it has a capture, our browser has built-in capture solving solution. So it will automatically solve the capture and your bot can continue on browsing without getting blocked. How much time I got? I don't know. So just to kind of summarize it, right? This is the biggest hallucinations that you guys see. The agent gets blocked, it needs to please you, and it makes things up.
SPEAKER_01
And fake content also, right? So now, if you want to Google what is Cloudflare AI Labyrinth, it's basically a system. Once it detects a bot, it doesn't block it. It literally feeds it fake data. So bigger hallucinations, right? And the easiest fix for this is just to make sure that your agent doesn't get blocked. And it's as easy as to implement our MCP. Our MCP has a free tier of 5,000 requests. So if you guys want to try it out, you can connect with me on LinkedIn if you want to. Or is this the... Hold on a second. Is this the sign up for the MCP? I'm lost in these QR codes. One second.
SPEAKER_01
Internet. Yes, please. How does it detect if there's Cloudflare Labyrinth or not? So the way we approach it is that we make your agent look like a human being. Literally, like there's mouse movement free recorded. There's typing. When it types, it's like it's... Mimics a real human behavior. So the Cloudflare literally just doesn't even ask, are you a robot or not? Right? So this is our approach. Instead of trying to understand how they detect, we make the agent look as human as possible so that it doesn't trigger the actual blockage. Misleading data is one of the toughest things that you can actually encounter. A lot of websites right now in Asia is
SPEAKER_01
doing that, right? Hotels, they're literally providing different prices. You go check out on your phone, you get one price. You can check from your computer, you get a different price. You can add through proxy, you get a third price. Which one is correct? It's really hard to tell. When it comes into the domain of misleading, the best bet is to make sure that your agent looks like a human and hope for the best. That's basically what the approach is right now. AI Labyrinth was literally released a month ago. I don't have much of statistics on exactly how it's working. All I know is that it didn't really
SPEAKER_01
affect us. We don't see any kind of change in data and we are collecting petabytes of data on the daily. We have so many customers always scraping. We're caching data so we're always comparing the results. We don't see any degradation in the results. So I think we're doing a good job in that case. This is a QR code that you guys can sign up for. We have a GitHub page as well. GitHub bright data.com. GitHub bright data. And another thing that I would recommend for you guys to check out is the skills page. Right? So what we did is we created skills. And I'm going to have another session in a couple of hours if
SPEAKER_01
you're interested in seeing what the skills does. Is basically you can take any agent, tell it to go here and this page will teach it on how to build a scraper or how to build a pipeline that will collect you the data. It has all the information it needs, all the APIs. And if you will come back to my next session which is at one o'clock I think? Something like that. I'm going to literally demonstrate how it builds the pipeline. Literally in front of you I'm going to tell it, hey listen let's build a Walmart collector for ABC and it's going to build it and it's going to scrape it. And instead of parsing
SPEAKER_01
each individual in ChiptML it's going to build a parser and it saves about 99% of the tokens. Because I see a lot of people is like, hey I need to parse 10,000 pages but it's so token heavy. Don't parse with the LLM. LLM builds the parser and then the script runs it. But that's the next session. Any questions? Question on the performance. I saw that you have CP exposes 69 tools which means if I need to say search capability for my agent, do you need to load all the 69 tools? No. Of course, filter it. I just showed 69 tools because just to show it, if I need just a scrape markdown and search I would just literally load two tools. Otherwise you're flooding contacts with
SPEAKER_01
relevant data. Of course not.
SPEAKER_01
The experiment at the beginning, does that use the Web search tool? I'm sorry, I didn't hear at the beginning again. The experiment at the beginning, does that use the Web search tool through the OpenAI API? Or is it like for the . And so how does it compare? I'm not that experienced
SPEAKER_04
.
SPEAKER_01
So we didn't do any searches. I literally told it, hey go to this URL, see if you can load it. Right? So I didn't use the search in this demo.
SPEAKER_04
Sure.
SPEAKER_01
But of course, again, even with our MCP, they can actually do a Google search. Yeah. And that's one of the biggest benefits is because we're used to Google results. Right? So by default, when you're asking LLM, you're expected to do a Google search, but it doesn't. Right. So the results with the MCP, much, much better. And I recommend, sign up, try it out, it's free. See the results, compare what you guys get. Is that 5,000 requests per day or...? Per month. Per month. Per month. Which is, you know, like for an MVP, for a little experiment, it's more than enough. And we also have pay as you go, so if you do need a little more, it's nothing like, you know...
SPEAKER_01
Okay. Just for running a set, just for paying a run. Yeah, yeah, it's for pro-type, it's perfect. I always, I do a lot of hackathons and I'm always recommending, hey listen, set up an MCP and tell your agent to go build whatever you need.
SPEAKER_03
It does a much better job than without.
SPEAKER_01
Forget it.
SPEAKER_01
Forget it. Forget it. Forget it. Forget it. Forget it. Thank you.