AI Engineer

How Web Data Infrastructure Powers the Next Generation of AI — Patricija Žemaitytė, Oxylabs

1852 summary words 8 min summary Watch video

Start with the signal

8 min read

Summary

At-a-Glance

  • Verdict: Skim
  • Core thesis: The next generation of AI depends less on models alone than on an adaptive web-data infrastructure layer that can retrieve, structure, and deliver current external information with low latency and at production scale.
  • Why it matters: Agentic and retrieval-based systems are only as useful as their access to timely, reliable outside data; web extraction introduces hard operational problems in anti-bot resistance, browser execution, latency, observability, and scale.
  • Best use: Use this as a vendor-side case study for framing live-web retrieval as a control-plane and infrastructure problem, while treating Oxylabs' performance claims as directional rather than independently validated.

Executive Summary

Patricija Žemaitytė argues that AI is shifting from static, training-bound knowledge toward systems connected to live retrieval. In her framing, the differentiator is not merely access to public web data, but the infrastructure required to collect it reliably, transform it into usable formats, and feed it into models, databases, and agents without forcing customers to operate the underlying scraping, proxy, browser, and anti-bot stack.

The talk uses a rapid video-data project to illustrate how AI-data requirements expand beyond an initial request. A client initially asked for a video API capable of collecting at least five petabytes monthly within two weeks; the engagement then expanded from downloads into transcripts, subtitles, language discovery, metadata, channel information, storage integration, and delivery. Oxylabs turned those successive adaptations into a video API suite in roughly three months.

The strongest operational section concerns search retrieval latency. Oxylabs' conventional SERP scraper averaged roughly four seconds because it captured a full results-page payload. For AI workflows, the company instead designed a narrower fast-search product around organic results, top stories, and news. A first redesign reportedly reached 650 ms p90 but failed under real blocking conditions; a subsequent version required browser-heavy execution and optimization across layouts, parsers, sessions, and proxies, ultimately reaching a claimed 550 ms average latency.

At scale, the speaker stresses that capacity is not solved by adding servers. Scaling its Web Unblocker from about 10,000 to 60,000 end-to-end scraping jobs per second in under two months exposed a more fundamental problem: realistic testing and truthful observability. Synthetic load was insufficient, telemetry itself became part of system load, and the organization ultimately had to increase exposure gradually using production traffic. The useful lesson is that live-web access is an ongoing adaptation problem, not a one-time implementation.

Key Takeaways

  • Claim: Live retrieval is becoming a necessary infrastructure layer around AI models because training alone cannot keep models current or grounded in external reality. | Evidence: The speaker cites SERP data moving from SEO, analytics, and market intelligence into retrieval pipelines, answer grounding, assistants, and agents; she also references Google's grounding documentation as positioning Google Search as a connection to current public knowledge. | Implication: For agent systems, web retrieval should be treated as a production dependency with explicit reliability, freshness, provenance, and fallback requirements—not as an optional search API call. | Caveat: Fresh web access improves recency, but the talk does not address retrieval quality, source trustworthiness, licensing, compliance, or how systems should handle conflicting sources.
  • Claim: AI training-data and multimodal-data products tend to expand from a simple capture request into a pipeline for rich contextual assets. | Evidence: A client request for a video API capable of at least 5 PB per month in two weeks evolved from video downloading into transcript support, then subtitles after the client clarified its actual need, followed by language search, metadata, and channel information; Oxylabs built a suite in roughly three months. | Implication: When scoping multimodal ingestion, define the full downstream artifact contract early: raw media, captions versus transcripts, metadata, discovery/search, storage, lineage, and delivery format. | Caveat: The anecdote demonstrates product adaptation, but it does not establish that every AI workload needs all modalities or that a generalized API is preferable to a task-specific pipeline.
  • Claim: Sub-second web-search retrieval requires a purpose-built product and narrower output contract, not a simple optimization of a comprehensive SERP scraper. | Evidence: Oxylabs says its traditional SERP scraper averaged about four seconds because it collected ads, widgets, rich results, AI-generated results, and varying layouts. Its fast-search design retained primarily organic results, top stories, and news. | Implication: Route requests by workflow: use minimal, low-latency search outputs for interactive agent loops, and reserve richer/full-page extraction for tasks where completeness outweighs response time. | Caveat: Reducing payload and page fidelity can omit useful ranking context, commercial results, structured widgets, and other signals that may matter for some research or agent tasks.
  • Claim: Performance demonstrated in development or controlled tests is not evidence that a web-data system will survive real target behavior and anti-bot defenses. | Evidence: A redesigned sub-second SERP product reportedly achieved roughly 650 ms p90 in under two weeks, then was heavily blocked during a live client test and had to be rebuilt. The second approach relied more on browsers, which the speaker describes as slow, expensive, and operationally complex. | Implication: Before embedding a web provider into autonomous workflows, test against representative targets and geographies under real conditions; require measurement of success rate, freshness, latency distribution, block recovery, and unit economics—not just benchmark latency. | Caveat: The speaker offers no quantitative success-rate, block-rate, cost, or latency-percentile breakdown for the final production system, so the trade-offs cannot be independently assessed.
  • Claim: Low latency is achieved through accumulated system-level reductions across the retrieval path rather than one breakthrough. | Evidence: The speaker describes systematically reviewing layouts, parsers, sessions, and proxies after browser requirements conflicted with a sub-second target. Oxylabs claims the resulting fast-search API now delivers fresh data into AI workflows at 550 ms average latency. | Implication: Design latency budgets end-to-end across query routing, acquisition, rendering, parsing, normalization, safety checks, model inference, and retries; optimizing only one layer will not make an agent interactive. | Caveat: The 550 ms figure is a vendor-reported average and does not specify query mix, geography, success criteria, percentile performance, or whether browser-heavy cases are included.
  • Claim: At high scraping volume, the primary scaling constraint becomes architecture, realistic testability, and observability rather than raw server count. | Evidence: Oxylabs scaled its Web Unblocker from around 10,000 to 60,000 end-to-end requests per second in less than two months, where each job included routing, rendering, proxy handling, browser execution, parsing, retries, normalization, and delivery. Load testing hit a wall around 20,000 RPS; the company says adding 2,000 servers alone would not solve the problem. | Implication: For any high-volume agent data plane, invest before scale-out in realistic traffic replay, bounded retries, failure-domain isolation, and telemetry systems whose own cost and cardinality are controlled. | Caveat: The account is a single vendor case study, and the reported scale does not disclose workload composition, reliability outcomes, or the operational risks of production testing.
  • Claim: Web-data infrastructure is an 'adapt-forever' operating model because target sites, layouts, detection techniques, and client requirements continuously change. | Evidence: The speaker frames Oxylabs' role as taking on changing targets, anti-bot systems, browser handling, data structuring, and delivery so AI companies can focus on their intelligence layer. She characterizes innovation as adapting quickly enough that new requirements become infrastructure. | Implication: If live-web access is strategic, decide explicitly whether it is a core competency to build or a managed dependency; either choice needs provider redundancy, contractual controls, and a fallback path for critical workflows. | Caveat: Outsourcing this layer reduces maintenance burden but creates supplier dependence and may complicate governance, data policy enforcement, and continuity planning.

Detailed Brief

Vendor positioning and operating-model signal

  • Claims: Oxylabs positions itself beyond commodity proxy supply, emphasizing an integrated infrastructure layer for reaching the open web, overcoming anti-bot systems, running browsers when necessary, structuring results, and delivering them into AI applications.; The speaker treats data accessibility as operationally distinct from data availability: public data may be theoretically accessible, but reliably acquiring it at scale is a separate engineering function.
  • Evidence: Oxylabs was established in 2015 and is described by the speaker as a web-intelligence platform and premium proxy provider.; The presentation cites a growth claim from 400 million daily requests to almost 6 billion daily requests and says the former 'Project 60' objective is now shifting toward 'Project 150,' with observed demand around 100,000 RPS.
  • Caveats: The presentation is an Oxylabs conference-style talk and is principally a positioning narrative; operational performance, scale, and customer outcomes are self-reported.; The transcript repeats much of the latter half of the talk, so it contains less distinct material than its 4,172-word count suggests.
  • Implications: The relevant procurement question is not simply whether a provider has proxies, but whether it owns the full operational loop for acquisition reliability, extraction quality, data normalization, incident response, and changing target behavior.; Separate strategic value from vendor marketing by conducting workload-specific proof-of-concept tests with defined performance and compliance acceptance criteria.

Notable Concepts & Terms

  • Live retrieval layer: The infrastructure surrounding a model that retrieves current external information so answers and agent actions are not limited to training-time knowledge.
  • SERP data: Search-engine results page data; presented as a feed for retrieval, grounding, assistants, and agents rather than only SEO and market intelligence.
  • Fast Search API: Oxylabs' narrower SERP product designed for AI workflows by prioritizing organic results, top stories, and news over full-page fidelity.
  • Web Unblocker: Oxylabs' proxy-integrated scraping product, used in the talk as the example of scaling an end-to-end web extraction job rather than simple HTTP traffic.
  • p90 latency: The latency threshold under which 90% of requests complete; the speaker cites roughly 650 ms p90 for an early fast-search iteration.
  • Organic data testing: Testing with traffic that resembles actual customer behavior, contrasted with synthetic load generation that may fail to expose real system limits.
  • Observability as load: At large scale, logs and metrics are not free diagnostics; their collection, storage, processing, and cardinality become part of the system's capacity and complexity.

Operator Notes / Why Ken Should Care

  • Establish separate retrieval service tiers for interactive agent loops versus deep research: define different latency, completeness, cost, and fallback targets rather than using one web-search path for all workloads.
  • For any external web-data vendor evaluation, require a representative proof of concept covering target sites, geographies, block rates, success rates, p50/p90/p99 latency, freshness, payload structure, retry behavior, and cost per successful result.
  • Define a normalized evidence contract for web-fed agents: URL/source identity, retrieval timestamp, extraction method, content hash, structured fields, and confidence or failure state should travel with retrieved data.
  • Avoid making production traffic the first meaningful scalability test; build safe canarying, traffic replay, per-target rate limits, circuit breakers, and degradations to cached or alternate sources.
  • Review data-rights, terms-of-service, privacy, jurisdiction, and supplier-continuity requirements before treating public-web extraction as a foundational data source.

Source/Metadata

  • Title: How Web Data Infrastructure Powers the Next Generation of AI — Patricija Žemaitytė, Oxylabs
  • Transcript words: 4172
  • Duration seconds: 1143
  • Timestamp note: No timestamps or chapters were present in the provided transcript; the latter portion is substantially duplicated.
Full transcript 2492 words · 18 min read
0:00

Hello everyone.

0:13

Most AI talks today start with models. This one starts somewhere less glamorous: with infrastructure that decides whether those models get fresh, usable, real-time data at all. I work at Oxylabs, and Oxylabs was established in 2015 and describes itself as a web intelligence platform and a premium proxy provider. In simple terms, we build infrastructure that allows companies to extract public web data at scale. As we all know, public web data theoretically is available for everyone. But in practice, if you want to connect your AI models, agents, and databases, you need an infrastructure layer. This is what we do. And this is what matters now more than ever.

0:28

Because the industry is shifting away from static knowledge, and training itself still matters, of course. But training alone is no longer enough. To stay useful, models need access to fresh information, live search, and real external data. Without that, even the smartest model is limited by what it knows. This is where my story begins. My name is Patricia. As I mentioned, I work at Oxylabs as a product manager now. But I actually started closer to engineering. I was leading teams, dealing with core services. The first squad that actually taught me one thing was what we called UX.

0:45

And what is UX? UX usually means user experience. That is completely correct. But for us, that often meant something closer to this: the client needs something really unusual. There is no ready-made product. The timeline is extremely painful. And somehow, we need to build everything fast and make it work beautifully. The lesson that I learned with that team was that innovation never comes as a neat roadmap. It comes as pressure, as a deadline, and sometimes, quite often, as a trip report from San Francisco. And this is how the first story started. One day, our sales team came back from San Francisco and said there is a demand for a video API for AI training.

1:00

There is one question that you're actually really scared to ask the sales team: What's the deadline? Two weeks. What's the scale? At least five petabytes per month. At that point, we had never built anything like that. And it seemed like a lot. Actually, this is also a moment when the feature stops sounding like a product feature. It sounds like infrastructure. Because what the client is actually asking us to build is not just to download some videos. They are asking for a pipeline: collection, transfer, storage, delivery. And to do it with enough reliability that it would be compatible with AI training workloads.

1:16

That story actually aged surprisingly well. Because the market has moved exactly in that direction. And AI infrastructure is becoming increasingly multimodal. It's no longer about text. Companies now need pipelines for video, metadata, transcripts, subtitles, and other structural context around the content itself. So what did we do? In two weeks, we had to build a new dedicated scraper with brand-new logic, new storage integrations, and a delivery flow for something that we had actually never built before. And we actually made it. Somehow, we even made it on time. But this is not where the story actually ended. That was only version one.

1:32

The client asked, great that you have a downloader, but what about the transcripts? So, we built transcript support. The client tested it out, and we saw that all of the requests were failing. Then we started talking with the client, and we saw that it wasn't that we had done something wrong. The client actually didn't need a transcript. They needed subtitles. So, we adapted again. We built subtitle support. Then another request came: We're struggling to find videos in the languages that we actually need. Can you build a search so that we could gather those ideas? So, we did it again. What about metadata? Of course, we did it once again.

1:54

This is the part of the story that I really loved. Because what started as one product feature request actually became the seed of a whole product. We started thinking that we were building just a downloader. Then we realized that we were building transcript support, subtitle support, adding metadata, channel information, and we ended up building our own internal library that glues everything together. After enough iterations, as I mentioned, what started as a one-off request became a product family. In roughly three months, we actually ended up having a whole video API suite that supported downloaders, transcripts, subtitles, and channel information.

1:58

And after all of this, the final twist came. So, it's 2026. The client already gathered 30 petabytes of data. And we're still waiting for a payment. So yes, the first lesson is really technical, but also very human: innovation is actually repeated adaptation under high pressure. Because once you learn that the client actually doesn't buy the first product iteration, they buy your ability to adapt, the next question becomes: can you actually make it under extreme latency constraints too? This is the part where I tell you a little about SERP data.

2:10

Search data has always mattered, but AI changed the role it plays. Before, SERP was often used for analytics, SEO, monitoring, and market intelligence. But now it's a huge part of AI systems. It feeds retrieval pipelines. It grounds answers. It powers assistants. It helps agents interact with live information instead of stale training memory. And that shift is not hypothetical. Google's grounding documentation explicitly positions Google Search as a way to connect models to current public knowledge. In simple terms, the model layer is increasingly expected to work with a live retrieval layer around it. And that's why the next request matters so much.

2:32

Back in 2024, a client came and asked for SERP delivery with sub-second delivery. At that time, our traditional regular search scraper was around four seconds average latency. So, the gap was huge. But we still decided to go for it, just to see if it was possible. And we actually did it. But the story doesn't have a happy ending here because the client did not test it out. And to be honest, the market wasn't ready for that. So, we just put it on a shelf.

2:49

But what became clear later on was that this was never about making the old scraper faster. Because the regular scraper, what it does, is built to retrieve as much information as possible. We're talking ads, widgets, rich results, AI-generated results, different layouts. And when we're thinking about a fast search API, it takes a different approach. It focuses only on the things that actually matter for AI systems. So, it's mostly organic results, top stories, and news. It cuts away all the heavy layout. So, even with this smaller scope, it's already something that starts to make lower latency possible.

2:52

Fast forward. It's 2025. Another client comes in, and the request was simple: zero data retention, sub-second latency, and two weeks. For us, that meant supporting different geolocation and query parameters, having a system that is capable of delivering results under 800 milliseconds, and having a solution that is ready to be tested in less than two weeks. When your baseline is at four seconds, we are not talking about optimization. We are talking about redesign. So, we started from scratch. And actually, the first version worked. In less than two weeks, we got around 650 milliseconds p90.

3:07

That alone would be a great story. But the real story actually happened on the next call. We're sitting on a call with the client, getting ready to test out our new product. And while we were on the call, we got blocked. And we got blocked really badly. To be honest, this is a really honest moment when you think about infrastructure and systems. Because this is a reminder that there is a difference between a system that works in development, a system that works in a test, and a system that actually survives reality. So, we had to start over, because nothing worked.

3:25

This second iteration was the hardest one because we actually had to rely a lot on browsers. And don't get me wrong, browsers are amazing. They are extremely useful. But browsers also are slow, expensive, complex, and deeply incompatible with dreams about low latency. So, we had a contradiction. The client wanted sub-second. Reality needed browsers. And browsers really wanted to give us four seconds. At this point, there is no magic trick. You just go hunting for time. You review everything: layouts, parsers, sessions, proxies. Every place where you can cut off a second, two, three, or four.

3:37

And this is how systems become fast. Not by giant breakthroughs, as we thought at first, but by small decisions that add up. That work paid off and actually evolved into something new. Today, we have a fast search API that delivers fresh data directly into AI workflows with 550 milliseconds average latency. And our scale moved from 400 million daily requests to almost 6 billion daily requests. That number matters. Because going from 400 million daily requests to 6 billion daily requests is not just growth. It's a change in operating model. It changes how you think about costs, observability, and failure domains.

3:50

So, the lesson of this part is that in the AI era, speed is not just performance. Speed actually defines what product can exist. Because in four seconds, you have a slow pipeline. In sub-second delivery, you have something that can sit and interact in your AI workflows. So, when speed becomes the product, what's next? Next is when scale actually becomes the real test. The first story was about adapting product scope. The second was adapting architecture for latency. The third one is about adapting systems for scale. And scale is where infrastructure becomes really humbling.

4:14

At one point, another demand forced us to scale our Web Unblocker quite aggressively. I added a slide just to show how it works. In simple terms, it's similar to a scraper but has proxy integration. We were operating at around 10,000 requests per second. Demand forced us to scale to 60,000 requests per second in less than two months. That number alone sounds impressive, but it can also be misleading if you are thinking about it as a simple HTTP request. In our world, that means an end-to-end scraping job. It includes routing, rendering, proxy handling, browser execution, parsing, retries, normalization, and delivery itself.

4:25

When you scale to that workload, even adding an additional 2,000 servers doesn't solve the problem. You need architecture. You need central components that are actually reliable. You need observability that still tells you the truth. You need testing that resembles reality enough to matter. And this is where our main bottleneck showed up. Not in dramatic outages. In load testing. The hardest part was not generating synthetic traffic. Synthetic traffic is relatively easy compared to reality. The hardest part was organic data testing. That means processing traffic that behaves enough like real client usage to tell us something useful.

4:32

During one of those load tests, we hit a wall at around 20,000 requests per second. At that point, there is no question of whether the system is actually working. It is working. The question becomes: do we actually know that it can go further? And that uncertainty was the real bottleneck. So were all the pain points around metrics and logs, and generating and processing everything at scale.

4:39

Everybody loves observability in theory. But observability at scale becomes real work. Because collecting logs is hard. Processing logs is harder. And the same applies to metrics. They are essential, but when you scale up to that kind of load, telemetry itself becomes part of the load and part of the complexity. So, what did we do with it? We scaled gradually. And eventually, we had to accept one unavoidable truth: the real testing was going to be with production traffic. Thankfully, that part actually went completely fine.

4:51

But the story doesn't end here, because the drama is still happening right now. Internally, we call this Project 60, because we had to scale up to 60,000 requests per second. Now, it's already becoming Project 150. While we were scaling our infrastructure to 60,000 requests per second, now we are talking about and seeing results that scale up to about 100,000 requests per second. So, the lesson from this part is also simple: scale is never a finish line. At least not for us. Probably, when you reach one target number, the next one will appear. So anyway, what does Oxylabs do in this whole thing?

5:01

I guess the stories make one thing quite clear: we are not just a proxy provider. Proxies are essential. They are important. But the hardest part, and the larger job, is building the infrastructure layer that allows companies to extract public web data and operate it at scale. That means reaching the open web, collecting data reliably, dealing with anti-bot systems, handling browsers when they are needed, structuring and delivering data, and doing it in a manner that AI companies can actually plug into their systems.

5:13

And this is exactly why it matters. Because the best thing we can offer is not just data access. It is this: you build the intelligence, and we take the messy maintenance underneath. Because the messy part is real. The targets change. Layouts change. Detection changes. The market itself changes. Client needs change. So, this is not a build-once business. This is an adapt-forever business. And honestly, that may be the most useful definition of innovation that I know: innovation is the ability to keep adapting fast enough that changes in requirements become new infrastructure.

5:48

If I need you to leave with one thought today, I will probably go back to where I started: the next generation of AI will not be powered by better models. It will be powered by better infrastructure around them. Infrastructure that can connect models to reality. Infrastructure that can push web data directly into your pipelines, databases, agents, and AI tools. Infrastructure that can scale from 400 million daily requests to 6 billion daily requests. Because this is really the story. Not just scale. Not just scraping. Not just speed. Adaptation. Adapting products. Adapting architecture. Adapting systems.

6:01

And doing it fast enough that AI companies and you can keep building while the maintenance burden stays with us. This is what it actually means for me in the AI world. It means that the model is not alone anymore. It already has a bridge to reality. Thank you. The next question becomes, can you actually make it under extreme latency constraints too? And this is a part where I tell you a little about SERP data. And search data has always mattered, but AI changed the role it plays. Before, SERP was often used for analytics, SEO, monitoring, market intelligence, but now it's a huge part of AI systems. It feeds retrieval pipelines. It grounds, it powers assistance.

6:54

It grounds answers. It helps agents interact with live information instead of sales training memory. And that shift is not hypothetical. Google's grounding documentation explicitly positions Google search as a way to connect models to current public knowledge. In simple terms, the model layer is increasingly expected to work with live retrieval layer around it. And that's why the next request matters so much. So, back in 2024, client came and asked for SERP delivery with, sub-second SERP delivery. At that time, our traditional regular search scraper was around four seconds average latency. So, the gap was huge.

7:43

But we still decided to go for it just to see if it's possible. And we actually did it. But the story doesn't have a happy ending here because client did not test it out. And to be honest, the market wasn't ready for that. So, we just put it on a shelf. But what became clear later on, that was never about making the old scraper faster. Because the regular scraper, what he does, it's built to retrieve as much information as possible. So, we're talking ads, widgets, rich results, AI-generated results, different layouts. And when we're thinking about fast search API, it takes a different approach. It focuses on the things that actually matters only for AI systems.

8:28

So, it's mostly organic results, top stories, news. And it cuts away all the heavy layout. So, even this small scope, it's already something to start thinking about lower latency. So, fast forward. It's 2025. Another client comes in. And the request was simple. Zero data retention, sub-second latency, and two weeks. For us, that meant to support different geolocation and query parameters, to have a system that is capable to deliver results under 800 milliseconds, and to have a solution that is ready to be tested out in less than two weeks. So, when your baseline is at four seconds, we are not talking about optimization. We are talking about redesign.

9:20

So, we started from scratch. And actually, the first version worked. In less than two weeks, we got around 650 milliseconds p90. So, that alone would be a great story. But the real story actually happened on the next call. So, we're sitting on a call with the client, getting ready to test out our new product. And while we were on the call, we got blocked. And we got blocked really bad. And to be honest, this is really honest moment about when you think about infrastructure and systems. Because this is a kind reminder that there is a difference between system that works in development, system that works in a test, and system that actually survives reality.

10:10

So, we had to start over because nothing worked. And at this second iteration was the hardest one because we actually had to rely a lot on browsers. And don't get me wrong, browsers are amazing. They are extremely useful. But browsers also are slow, expensive, complex, and deeply incompatible with dreams about low latency. So, there is... So, we had a contradiction. That reality... The client wanted sub-second. The reality needed browsers. And browsers really wanted to give us four seconds. So, at this point, there is no magic trick. You just go hunting for a time. So, you review everything. Layouts, parsers, sessions, proxies.

11:00

Every place when you can cut off a second, a two, a three, or four. And this is how systems become fast. Not by giant breakthroughs as we thought at first. But by small decision that adds up. And that work paid off and actually evolved into something new. So, today, we have fast search API that delivers results in front of the system. We have fresh data directly into AI workflows with 550 milliseconds average latency. And our scale moved from 400 million daily requests to almost 6 billion daily requests. So, that number matters. So, that number matters. Because going from 400 million daily requests to 6 billion daily requests is not just a change. Not just a growth.

11:49

It's a change in operating model. It changes how you think about costs, observability, and failure of domains. So, the lesson of this part. That in AI era, speed is not just performance. Speed actually defines what product can exist. Because in 4 seconds, you have a slow pipeline. In sub-second delivery, you have something that can sit and interact in your AI workflows. So, when the speed becomes product, what's next? Next is then scale actually becomes the real test. So, the first story was about adapting product scope. The second was adapting architecture for latency. The third one is going to be adapting systems for scale.

12:40

And the scale is where infrastructure becomes really humbling. At one point, another demand was forced us to scale our web on blocker quite aggressively. I added just a slide just to see how it works. It's in simple terms. It's similar to Scraper but has proxy integration. So, we are working our way around 10,000 requests per second. Demand is forced to scale to 60,000 requests per second. And in less than two months. So, now that number alone sounds impressive. But it might be also misleading if you are thinking about it as a simple HTTP request. In our world, that means the end-to-end scraping job.

13:26

It will be routing, rendering, proxy handling, browser execution, parsing, retries, normalization, and delivery itself. So, when you kind of scale to that workload, even adding up additional 2,000 servers doesn't solve the problem. You need an architecture. You need a central component that actually are reliable. You need observability that still tells you the truth. You need testing that resembles reality enough to matter. And this is where our main bottleneck showed up. Not in dramatic output. In load testing. The hardest part was not generating synthetic traffic. Synthetic traffic is relatively easy compared to reality. But the hardest part, organic data testing.

14:17

That means processing traffic that behaves enough like real client usage to tell us something useful. And during one of those load tests, we hit the wall at around 20,000 requests per second. And that point, there is no question if the system is actually working. It is working. The question becomes, do we actually know that it can go further? And that uncertainty was the real bottleneck. So are all the pain points. Metrics, logs, and generating and processing everything at scale. So, everybody loves observability in theory. But observability at scale becomes true work. Because collecting logs is hard. Processing logs is harder. And the same applies to metrics.

15:07

They are essential. But when you scale up to that kind of load, telemetry itself becomes a part of the load and a part of the complexity. So what with it? We scale gradually. And eventually, we had to accept one unavoidable truth. That the real testing is going to be with production traffic. And thankfully, that part actually went completely fine. But the story doesn't end up here. Because the drama is still happening right now. Internally, we call this project 60. Because we had to scale up to 60,000 requests per second. Now, it's already becoming project 150. So while we were scaling our infrastructure to 60,000 requests per second,

15:54

Now we're talking and seeing results and scale up to about 100,000 requests per second. So the lesson from this part is also simple. That the scale is never a finish line. Well, at least not for us. And probably when you reach one target number, the next one will appear. So anyways, what does Oxlabs do in this whole thing? I guess the stories make one thing quite clear. That we are not just a proxy provider. Proxies are essential. They are important. But the hardest part and the larger job is building the infrastructure layer that allows companies to extract public web data and operate it at scale.

16:39

That means reaching the open web, collecting data reliably, dealing with ant bot systems, handling browsers when they are needed, instruction and deliver data, and doing in that manner that AI companies can actually plug into their systems. And this is exactly why it matters. Because the best thing we can offer is not just data access. It is this. That you build the intelligence and we take the messy maintenance underneath. Because the messy part is real. The targets change. Layouts change. Detection changes. Market itself changes. Client needs changes. So this is not a build one's business. This is an adapt forever business.

17:25

And honestly that may be the most useful definition of innovation that I know. That innovation is the ability to keep adapting fast enough that the change in requirements becomes a new infrastructure. So if I need you to leave with one thought today, I will probably get back where I started. That the next generation of AI will not be powered by better models. It will be powered by better infrastructure around it. Infrastructure that can connect models to reality. Infrastructure that can push the web data directly into your pipelines, databases, agents, AI tools. Infrastructure that can scale from 400 million daily requests to 6 billion daily requests.

18:13

Because this is really the story. Not just scale. Not just scraping. Not just speed. Adaptation. Adapting products. Adapting architecture. Adapting systems. And doing it fast enough that AI companies and you can keep on building while the maintenance burden stays with us. So, and this is what it actually means for me in AI world. It means that the model is not alone anymore. It already has a bridge to it. Thank you.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note