AI Engineer

The Rise of CaaS: Context-as-a-Service for Agentic AI — Omer Primor, Bright Data

1867 summary words 8 min summary Watch video

Start with the signal

8 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Agent systems should treat web context as a costed, continuously refreshed asset: use AI search for ad hoc discovery, but own and maintain a purpose-built context layer once queries over known entities or sources become frequent.
  • Why it matters: This provides a practical architecture and economic framing for routing agent retrieval between search, vertical Context-as-a-Service providers, and owned data pipelines rather than paying token and API costs repeatedly for the same web-derived facts.
  • Best use: Use it to pressure-test OpenClaw or other agent workflows that repeatedly research companies, people, markets, products, or other changing web entities, then define a build-versus-buy retrieval threshold.

Executive Summary

Omer Primor argues that the web has shifted from being merely a data source to being an agent context source. But the web is unstructured and fast-decaying: social content becomes stale in under a day, while news, finance, and retail data are often largely irrelevant after roughly 30 days. Therefore, useful web context is not a snapshot or a periodic export; it must be treated as an ongoing acquisition, normalization, and refresh process.

He distinguishes AI search from a rising Context-as-a-Service (CaaS) layer. Search is flexible discovery: it can answer unanticipated questions by exploring the live web. CaaS vendors instead crawl, structure, deduplicate, enrich, and expose domain-specific entities through APIs, MCP, or CLIs—effectively vertical search/knowledge-graph products for agent tasks in areas such as finance, GTM, e-commerce, travel, HR, and real estate.

His central operating argument is economic. In a small experiment enriching 100 event sponsors across 25 company fields with an Opus 4.8-based agent, AI search, Google SERPs, Claude-native search, and one major CaaS provider showed broadly comparable coverage, while costs diverged because search-based paths also required LLM token spend to structure retrieved results. CaaS can have coverage gaps when a requested field is outside its pre-collected dataset, whereas search can continue exploring.

For repeated research over stable, known source sites, Primor says teams should calculate a build-versus-rent tipping point. A quick internal pipeline using dedicated scrapers for LinkedIn company/job pages and Crunchbase, merged with basic entity-resolution rules, achieved respectable coverage with no per-query AI cost after setup. The presentation is vendor-promotional and explicitly not a benchmark, but its reusable insight is to architect a hybrid retrieval control plane: search unknown or novel questions; materialize, own, refresh, and reuse high-frequency context.

Key Takeaways

  • Claim: Web-derived agent context must be continuously refreshed because its value decays rapidly rather than remaining useful as a static corpus. | Evidence: Bright Data's internal data-decay analysis is cited as showing social-media relevance lasting less than a day; news, finance, and retail data are described as mostly no longer relevant 30 days after collection. | Implication: Design entity-level freshness policies and change detection rather than relying on one-time crawls, static RAG indexes, or monthly refreshes for operational web intelligence. | Caveat: The transcript does not provide the methodology, definitions of relevance, or source-specific breakdown behind the decay analysis.
  • Claim: AI search and CaaS solve different retrieval problems: search is broad and adaptive, while CaaS is structured, domain-specific context for repeatable agent work. | Evidence: Primor describes CaaS providers as indexing, deduplicating, enriching, and creating knowledge graphs over entities, then exposing them through MCP, CLI, or API; examples include finance, market research, retail/e-commerce, GTM, HR, travel, and real estate. | Implication: Do not route every agent question to a generic web-search tool. Classify requests by known-versus-unknown sources, expected recurrence, structured-output needs, and freshness requirements. | Caveat: A vertical context provider can only answer from its accumulated schema and corpus; it may miss newly requested fields that an open-web search can discover.
  • Claim: Coverage alone is an insufficient selection criterion for agent context providers because fit depends on the exact fields and task being evaluated. | Evidence: In the speaker's 100-company, 25-field enrichment test, search, Google SERPs, Claude-native search, and one major CaaS provider performed similarly on coverage, while two CaaS products underperformed on fields they had not collected, such as recent hiring data. | Implication: Benchmark retrieval at the field, source, freshness, and downstream-decision level of the actual workflow—not through a generic vendor coverage score. | Caveat: This was explicitly a limited experiment rather than a benchmark: it used 100 event sponsors, a simple field-by-field agent loop, and guardrails set by the presenter.
  • Claim: At scale, query frequency—not just the number of entities—is the main cost driver of rented web context. | Evidence: The speaker frames a million-record workload as repeated revisits to entities for changes in news, hires, departures, or other signals. Each search query can incur the same API and token costs even when the result has not changed, leading teams to reduce cadence or truncate results. | Implication: Measure retrieval economics as cost per entity per refresh cycle and per downstream decision, including repeated LLM extraction/normalization tokens—not simply cost per initial answer. | Caveat: The cost comparison omits exact vendor prices, workload distributions, maintenance costs, and the value of avoiding scraper operations.
  • Claim: For persistent, predictable research against known web sources, an owned context layer can cross a build-versus-buy threshold sooner than teams expect. | Evidence: The proposed prototype searched for company URLs, used dedicated scrapers for sources such as LinkedIn Companies, LinkedIn Jobs, and Crunchbase, merged outputs into entities with simple conflict heuristics, and achieved fairly good coverage over 100 companies. Primor estimates setup at roughly one week or $5,000, after which agent retrieval has no comparable per-query AI charge. | Implication: Materialize and maintain canonical entity records for repeated workflows, then allow agents to query the owned store freely while spending external-search budget only on discovery, exception handling, and refresh. | Caveat: The prototype covered a narrow company-enrichment scenario, required initial engineering, and does not account for long-term source breakage, legal/compliance constraints, storage, monitoring, or data-quality operations. "Free" retrieval means no added vendor/LLM retrieval charge, not zero total cost.
  • Claim: The strongest architecture is a hybrid web-context strategy rather than an all-search or all-CaaS choice. | Evidence: Primor recommends AI search for ad hoc, constantly changing questions and suggests mixing methods by task; he argues that owned context compounds through reuse whereas rented context incurs costs on every repeated call. | Implication: Build a retrieval router that selects live search for novel questions, CaaS for immediate structured vertical coverage, and owned data/knowledge graphs for high-frequency known-source workloads. | Caveat: The speaker does not give a quantitative formula for routing decisions or a validated tipping-point model.

Detailed Brief

Market framing: search is becoming agent infrastructure

  • Claims: The search market is moving from human-facing Google-style query interfaces to machine-consumable retrieval used directly by LLMs and agents.; CaaS is presented as a distinct market category rather than simply traditional data-as-a-service relabeled for AI.
  • Evidence: The speaker names AI-search-oriented companies including Exa, Parallel, You.com, and Tavily.; He says Amazon had announced web retrieval/indexing for agents on AgentCore and that Microsoft had packaged web retrieval as WebIQ within an agent-development/orchestration suite.; ZoomInfo is cited as launching GTM.ai, positioning GTM research and prospecting workflows for use from tools such as Claude Code or Codex.
  • Caveats: These market examples are claims made in a conference presentation and were not independently substantiated in the transcript.; The presentation is delivered by Bright Data, whose scraping and web-data infrastructure benefits if organizations elect to build owned context pipelines.
  • Implications: Agent platform differentiation may increasingly reside in data access, entity resolution, freshness, and source-specific retrieval—not only in the underlying model.; Traditional data vendors are likely to make their datasets agent-callable, so API access and MCP compatibility alone should not be treated as a moat.

Implementation pattern for owned company intelligence

  • Claims: When sources for a recurring task are already known, search is primarily a discovery bootstrap rather than the runtime retrieval layer.; Some websites already expose a usable implicit ontology: company, person, and job entities plus their relationships.
  • Evidence: The prototype starts with only a company name, uses Google to identify relevant URLs, then directly retrieves from known sites.; It creates dedicated source scrapers, merges source records into a company entity, and resolves conflicts with simple source-precedence heuristics.; Bright Data's Scraper Studio is presented as enabling AI-generated scrapers in under five minutes with self-healing when a site changes.
  • Caveats: Source-specific scraping can fail operationally or create legal, contractual, access-control, and anti-bot risks; none of these were addressed in the talk.; Simple heuristic merging is inadequate where provenance, confidence, temporal versioning, or contested attributes materially affect decisions.
  • Implications: A durable owned-context system needs provenance per field, source trust ranking, timestamps, refresh schedules, change logs, and fallbacks—not merely a vector database.; Prefer structured entity storage and explicit relationships where source data already has an ontology; embeddings can support discovery but should not replace canonical records for factual operational workflows.

Notable Concepts & Terms

  • Context-as-a-Service (CaaS/CAS): A proposed category of agent-facing, vertical context providers that turn web data into structured, enriched, deduplicated entity knowledge exposed through APIs, MCP, or CLIs.
  • Web context engineering: The speaker's framing for designing the acquisition, structuring, routing, freshness, and cost model for web-derived context used by agents.
  • Rented versus owned context: Rented context is re-obtained through search or vendor calls and incurs recurring costs; owned context is materialized internally and can be reused, refreshed, and combined with proprietary data.
  • Data decay: The rate at which web information loses relevance, driving the need for ongoing refresh rather than static indexing.
  • Vertical search engine: The analogy for CaaS: specialized retrieval and entity knowledge for a narrow domain rather than general-purpose web discovery.
  • Entity enrichment: Filling a canonical record—here, a company—across multiple fields such as domain, headquarters, hiring, and people data from web sources.
  • Build-versus-rent tipping point: The query volume and recurrence level at which upfront engineering and maintenance for an owned context pipeline become cheaper or strategically superior to recurring vendor/search costs.
  • Self-healing scraper: A scraper intended to adapt to website changes automatically; it is presented as a way to reduce maintenance for owned data acquisition.

Operator Notes / Why Ken Should Care

  • Inventory agent workflows that revisit the same external entities and calculate their current query cadence, external API cost, LLM extraction-token cost, and freshness SLA.
  • Implement a three-way retrieval policy: live search for novel/unbounded discovery, vertical data APIs for immediate structured coverage, and owned entity stores for recurring known-source tasks.
  • Before building an owned web-context pipeline, define source permissions and compliance posture, field-level provenance requirements, data retention, refresh cadence, and scraper-failure escalation.
  • Run a narrow build-versus-buy pilot on one recurring enrichment workflow; compare end-to-end accuracy, freshness, operating burden, and total cost over multiple refresh cycles rather than one-time retrieval cost.
  • Avoid using a vector database as the sole system of record for recurring factual intelligence; preserve structured entities, relationship edges, timestamps, source lineage, and confidence.

Source/Metadata

  • Title: The Rise of CaaS: Context-as-a-Service for Agentic AI — Omer Primor, Bright Data
  • Transcript words: 7663
  • Duration seconds: 1339
  • Timestamp note: No timestamps or chapter markers were present. The transcript contains substantial duplicated sections, likely from extraction.
Full transcript 3989 words · 34 min read
0:12

So hi, everyone. Thank you so much for taking the time to join this session. I hope I'll, or at least I can guarantee I'll do whatever it takes to make it worth your time. My name is Omer. I lead the product marketing team over at Bright Data. Just by maybe a quick show of hands, who here is familiar with Bright Data? Okay, we can do better. I'll pass it on to our brand team.

0:17

Bright Data is a web data company. We help more than 20,000 teams around the world, including more than 70% of the world's biggest AI labs, to extract data from the web. Just to put this in perspective of what scale we're talking about, we're talking well over 50 billion HTML pages every day, more than 20 petabytes of video, audio, and other media data. So that's just the perspective of the type of work that we do at Bright Data.

0:23

But enough about us. Personally, I joined Bright Data about three years ago, which essentially gave me front-row seats to everything around AI and the web and how they started to actually connect. It sounds very old, but if you think about it, only maybe less than two years ago, we were able to start using, access the web, and search the web through cloud or through ChatGPT. That option didn't even exist in the earlier versions. So that connection, that way in which both AI and the web are starting to converge, is something that is still evolving and evolving rapidly. And that's part of what I want to try and shed some light on today and talk about a new emerging breed of companies that's coming out of this connection.

0:27

I think we can agree that the web is by far the world's greatest source of data. Historically, when it comes to Bright Data, that's all we cared about. It's helping our customers extract data from the web. But with the emergence of AI, and the emergence more recently of AI agents that need to do knowledge work, the web is no longer just a source of data. We can actually start looking at it as a source of context. Context in the sense that if I do knowledge work and I have knowledge agents that support my work, I want to go out to the web, find the information that I need, use it as context, but keep on working. So the data itself is only a step in the process for something bigger, for the actions I need to take, for the conclusions I need to draw, and for every downstream application that follows.

0:34

The first one is to figure this one out. Oh, sorry, even before that.

0:40

But one thing that I want all of us to bear in mind, because this is going to follow us through the rest of this conversation: the web is messy, it's unstructured, and most importantly, it changes all the time. This is a chart that shows data decay. It's analysis done by our team. It shows data decay, how long after a new page, a new piece of content goes live, it is no longer relevant. So social media, it's easy for us to understand, it's far less than a day. But also news, finance, retail, 30 days later, data that was collected is mostly no longer relevant. And when we acknowledge that, this simple notion, we understand that extracting context from the web or relying on the web is not a snapshot. It's not a one-time effort. It's not even a monthly effort. It's something that we need to keep on doing. It's something we need to look at as an ongoing process and something that we need to be mindful of.

0:46

So the first ones to figure it out were, of course, search companies. Only, what, three years ago, we were on the far left. This is in our lifetime. Three years ago, we were on the far left. Everything was Google. There was complete and total dominance up until three years ago, for the past 20 or so years. That's what we're talking about. Purely for humans: search something, go, collect the information you need, and carry on. Then fast forward maybe one and a half years ago, two years ago, search began to appear within the LLMs, within the chatbots, which already started blurring the line between humans and agents, because now the same bots also have the same web search available, so the same LLMs have the same access through API for the bots. So we started seeing that convergence happening. And so, for the first time, we're seeing more and more traffic flowing down, search traffic, search intent flowing down these channels, not only to Google.

0:51

And last but not least, we now see a whole breed of companies, the AI search companies. I'm sure you're familiar with them. I caught a talk yesterday by Will, the CEO of Exxon, and Parallel, and New.com, and Tavili, and a bunch of others. They are purely built on indexing the web, especially for agents, not even looking at the humans involved anymore. So Google's dominance, when it comes to search, if Google was synonymous with web search, that is very much shaky. And when there's blood in the water, the sharks come.

0:57

Just last week, Amazon announced, I don't know how many of you saw it, that they developed their own index and started allowing the ability to retrieve data from the web, to retrieve context for agents on AgentCore. Amazon developed their own search engine. Two weeks before that, it was Microsoft. Microsoft always had skin in the game. I don't know, one or two percent of the world search traffic went to Microsoft. But they have now repackaged it and launched it again as part of WebIQ, as part of their suite for agentic development and orchestration. So we're seeing more and more this space becoming crowded.

1:02

But when we're talking about context, and we're looking at this through the lens of search, I believe it only tells us part of the story. I can search for what's the cost of a certain pair of sneakers this morning on a certain website. I cannot really search for how has that price changed over the last six months, what discounts it had. I can search for what open job positions we have at Bright Data. We do. I urge you to go have a look. But I can't see how that was a chart and how that changed over time and how the headcount of the company changed over time. And all of that information existed on the web, simply back then. So when we actually start to think about it, we understand that there's much more context in the web than what web search allows us to extract.

1:10

And this is what we started seeing in the recent years, a whole new breed of companies rising. We like to call them internally CAS, context as a service, because that's what they do. They allow agents to tap into them, MCP, CLI, just pure good old API, and actually start extracting data to retrieve data so they can reason over it for whatever knowledge work they are responsible for. We see this happening in e-commerce. We see this happening in travel. We see this happening in finance, in market research, in HR, in real estate, in a bunch of other domains. I'll show a few examples in a second.

1:16

What all of these have in common is that they don't just discover the web in terms of think crawling, think searching, think all of that, accessing, extracting the data, and indexing it. They take it a step further. They actually develop knowledge graphs to start structuring all of the entities and to dedupe them. And they start enriching them with a lot of different sources. So they actually start merging all of that data. If you think about it, they kind of behave like vertical search engines. They are a very, very, very good search engine for something very specific.

1:24

And it's already in full motion. So as I said, we see this in finance and in market research and retail and e-commerce and GTM and sales intelligence. What all of these companies, by the way, have in common, they're all part of Bright Data's startup program. If you are a builder, and this is a hot space to go in because I think we're only tapping the surface, I invite you to scan this and apply up to $20,000 in credits and all sorts of co-marketing. But that's enough self-promotion.

1:29

So CAS as an industry is already in full bloom. And as always with these situations, the traditional players aren't left too much behind. They're at the bottom. You see good old data as a service. You see ZoomInfo. By researching for this presentation today, I also saw that they launched that thing at the top. It's called GTM.ai. You can only imagine how much they paid for that domain. But they launched a secondary brand for ZoomInfo that is catering specifically for the need of agents. Look at the wording. They talk about GTM work, that knowledge work, that research that you do when you need to prospect, when you need to do headhunting, whatever it is you need to do that involves people mostly. Straight from cloud code, straight from codex, or any other agent.

1:36

apply up to $20,000 in credits and all sorts of co-marketing. But that's enough self-promotion.

1:40

So CAS as an industry is already in full bloom. And as always with these situations, the traditional players aren't left too much behind. They're at the bottom. You see good old data as a service. You see ZoomInfo. By researching for this presentation today, I also saw that they launched that thing at the top. It's called GTM.ai. You can only imagine how much they paid for that domain. But they launched a secondary brand for ZoomInfo that is catering specifically for the need of agents. Look at the wording. They talk about GTM work. That knowledge work. That research that you do when you need to prospect, when you need to do headhunting, whatever it is you need to do that involves people mostly. Straight from cloud code, straight from codex, or any other agent. They understand the gap. So yeah, it's fun to think of CAS as an evolution of this. And it is, in a way. But it's catering for a very specific need, as much as the AI search engines are different than Google. When agents need them, it's different than people. When we let this sink in, at the very least, we have two different types of paths to complete knowledge work as an agent. We can start thinking about this in terms of web context engineering. We can start thinking about this in terms of how do I optimize for the specific task. More importantly, when things come as they are, how do I optimize this for breeds of tasks? How do I do this for various parts of the organization that I'm building for? If I'm an AI engineer, I need to serve different teams. They may have different needs. It's very tempting to throw AI search at all of them, but maybe that's not optimal. Maybe I need a combination of both. Maybe I can start seeing all sorts of cost efficiencies emerge from that.

1:46

So for the second half of this presentation, we actually went ahead and created a test. This is not a benchmark. You won't see anything concrete that I can say with great confidence other than the actual research that we did, because we wanted to start unraveling the different considerations and how these two stack up against each other. So we designed a test. We went for something basic. We said, okay, let's take a company, an entity, and try and enrich it across 25 different fields. Some of them are very easy, the company domain, the name, the headquarters, but some are more challenging. Things about hiring and people and something. And we built a simple agent, a loop in a loop, that uses Opus 4.8 as the harness, and it starts to go field by field, go out, search for it or retrieve it from the CAS, do it again and again and again until it completes and brings back, set some guardrails, like budget and stuff, just to keep it fair. And I'm going to share with you the results.

1:54

So the first thing that we would care about, being knowledge work, would be, sorry, we ran it 100 times on all of the sponsors of today's event.

1:57

So the first thing that we saw in terms of coverage is that there's pretty good convergence. They all did fairly well. I'll get to the two at the bottom in a second. So search showed consistent performance. One of the major CAS providers did also very well. The third one, by the way, you can see Unlocker and SERP. SERP is good old data, good old Google. We did the same thing just with Google, and it performed pretty well in extracting that information. Native is Claude's own search, and you see that they converge really well. I was originally surprised about the two CAS solutions at the bottom. It was counterintuitive. I expected CAS to dominate this thing because you had one job, to map out these companies. But after diving into it a bit more, you understand that they are limited in the sense that they know what they have about an entity. If I asked it a question that is beyond that, they will never have that data. Unlike a search that can go out and continue searching and exploring it, if they didn't collect data about the recent job hiring, it will never be there. So it makes sense that they are a bit behind, but I'm sure at the same time that they have a lot of other advantages that we simply didn't ask for, a lot of other fields that they didn't have that aren't represented. So again, it creates some complexities in how do we measure coverage when it relates to the specific job that we need to do rather than in general.

2:03

The second thing we looked at was cost, of course. Here we started seeing it spread out a bit. So you can see that massive bulk in the center. Most of the search and the cost, and even using Google, converged to pretty much the same cost, only different. The CAS was just about the service itself, what you pay the vendor. All of the other search solutions, you also needed a lot of token burn to actually structure that data so you can act on it and use it as something retrievable. So it's the same output. Native, obscenely expensive, and the CAS on the right, I'm sure you're all familiar with, they're by far the most expensive in the industry. I will not name and shame them. Interestingly, you see that small CAS there at the left, that CAS number two, that was very cheap. They're also the ones that are here at the bottom, which is funny because what I believe is happening there is that we're seeing, even within this industry, niche players that have lower-quality data but are much cheaper. They're already carving that niche of the long tail, of small shops or small usage that don't want to pay as much and don't need as much data, and we're all seeing them branch out there.

2:08

Most of you here, I presume, are engineers. So there's a very evident question that we did not ask here, which is, what is the one thing that an engineer would care about? Thank you. Let's talk about scale.

2:13

What happens if we need a million? Now, yes, a million records will not fit in the context. Obviously, we're not talking about a single run that needs a million. You can think about million in terms of the frequency. If I am a market researcher, I do due diligence for private equity. I revisit these companies all the time. I ask more questions about them as time goes by. Was there any new news about them? Was there anything that changed? Did somebody join? Did somebody leave? Do they have new hires? I keep on asking the same thing. So when I'm talking about this, the multiply by a million, it's not just about the number of companies. It's the frequency in which I'm asking it. Frequency is the cost killer when we talk about these. And we need to acknowledge that. We're thinking about this in terms of web context engineering. We're starting to look at it differently. Every repeated query costs the same as the first. Even if it brought back the exact same answers. Nothing changed? Pay up. False positives? For sure, go in. Token costs. We saw that there's very high token. We know that doesn't shrink well over time. There's always some volume element in terms of the cost, but it's not the same as flatlining. And if we bring this back to knowledge work, this is where we see teams that are starting to cut corners. So I won't research this company every day. I'll look at it once a week or once a month. I won't ask that question now. I don't want all the results. I'll only take 10 results, 20 results. So we already have the setup. We have what we need to do the knowledge work. But at the same time, we're not extracting all of the value because we're starting to be conscious about cost. We're renting context. We're not owning the context that we use. That is a very important distinction.

2:18

Again, if we're good engineers and we ask ourselves, what about scale? The second most obvious thing that will come to mind now, so how about we build it? What if we take all of that web data ourselves and stick it in some vector database and try and see what comes out of it? So I asked my engineer to do exactly that. Again, this is a test. This is not a benchmark or a full-blown operation. This is a day's work at best just to illustrate the concept and to show something about the cost efficiencies that you can generate potentially by doing it yourself, potentially in specific scenarios.

2:24

The test? Simple. Take the company name, nothing but, run it through Google, find the relevant entries, the relevant URLs of that company in various websites that have all of that information. You use search when you don't know the source. But when we're talking about company enrichment, we all know these sources. We all know where that data comes from. what about scale? The second most obvious thing that will come to mind now, so how about we build it?

2:34

What if we take all of that web data ourselves and stick it in some vector database and try and see what comes out of it? So I asked my engineer to do exactly that. Again, this is a test. This is not a benchmark or a full-blown operation. This is a day's work at best, just to illustrate the concept and to show something about the cost efficiencies that you can generate potentially by doing it yourself, potentially in specific scenarios. The test? Simple. Take the company name, nothing but, run it through Google, find the relevant entries, the relevant URLs of that company in various websites that have all of that information, right?

2:46

You use search when you don't know the source. But when we're talking about company enrichment, we all know these sources. We all know where that data comes from. The cast also bring it from them. Zoom info bring it from them. It's the same thing over and over again. Why not just go straight to the source? Why are we doing that middleman thing? LinkedIn companies, LinkedIn jobs, Crunchbase. Right there, we have scrapers for those who just tap in, and you start paying as it pays as you go.

2:52

We build two dedicated scrapers. We have a new AI tool called Scraper Studio. It lets you build a scraper for any website in less than five minutes, all powered by AI. And then it also has a self-healing function, right? So if the website changes, it fixes itself and keeps on going. Merge it all into one entity. Basic heuristics, if there's conflict, choose that over that. And eventually, we have a data set of these 100 companies.

2:58

Zero AI cost involved. There's no tokens. Coverage? Fairly well. Not amazing. Not the best that we saw here. But stacking up pretty well. And again, this is just a day's experiment, probably not even as much. Okay? So again and again, very specific tasks, very limited context, very limited situation. Tread lightly and proceed with caution when it comes to conclusions. The real story is not this. The real story is this.

3:05

That's what it costs to just go and fetch that data that is out there. We think about knowledge graphs. We think about entities. But if you think about, for example, LinkedIn, the data is already structured in form of entities. There's an entity for a company. There's an entity for a person. There's an entity for a job. And they're connected between them. Sometimes the ontology is already there. Again, this is not the most complicated of scenarios. But this is pretty damn good.

3:11

Now, yes, it took time to set up. So it's not really fair to compare apples to apples when it comes to the cost, because these are out of the box. You can just tap into the API. That one that I just showed you required some setup. Let's say it's a week. Let's price it at $5,000 just to give us some perspective. We can actually start thinking about this in terms of a tipping point. We can actually start thinking about what is that tipping point in which it makes more sense for me to build it myself, right, than keep on renting it.

3:16

So it makes sense to do it at this point. Maybe it's not 15. Maybe it's 30. Maybe it's 100,000. Maybe it's 10,000. It really depends on the use case. But there is a tipping point in which it actually makes sense to do it yourself. Which leads us to the fact that both AI search and CAS and all these solutions, they're very good in the sense that you can just plug and play. But if your knowledge work needs, right, are persistent, inconsistent, and to a certain degree may even continue escalating and growing, then this is perhaps a direction to start considering. Maybe I can just go ahead and build my own.

3:21

Because the nice thing about it is that all of the things that we see on the left, up until the third part, is upfront investment. And the most important thing, that whatever retrieval happens later on from the agents, right, is free. Not really free, but you get what I mean. There's no added cost. I can just ask that question over and over again. I did not like the first answer. I'll ask it again. I'll ask it 100 times until I get what I need. I have no more fear, no more cutting corners, which is maybe the most important thing.

3:29

And I'm leaving aside the fact that this is also custom business logic. I can connect it with my own data. There's all sorts of other advantages of owning it. We'll keep it to the imagination. Remember, we asked about a million, not about 15,000. This compounds. This compounds greatly, right? We need that horizon. Remember, the web keeps changing. We saw the staleness of the data and how the data decays. So we need to be thinking about this in the long run and how this will evolve when we keep on asking the questions about the entities that we care about.

3:42

Just to wrap it up. So AI search costs, they can get you very far when what you need is ad hoc and what you need is always changing. When sometimes you look at different things, even the mix and match of them, for certain tasks use this, for certain tasks use that. You can, I'm sure, again, I just showed a test. There's a lot of ways to optimize it, just like any other context engineering, and make and use lighter models and use other stuff. There's a lot of great stuff to be done. But eventually, the frequency will come and bite you in the ass when it comes to cost. And that's something to be mindful of. And there's a fair chance that that tipping point is much lower than you think. And that's something that, as we design these systems, when we think about web context engineering, we need to be mindful of.

3:48

And last but not least, the last slide we showed. Owned context compounds while rented decays, right? It's not a one-time task. Again, if it's a one-time question, use AI search. It will be amazing. When you need to do it over and over again, there's a fair chance that it will not, that you're missing out on potential compounding effect and you are losing out. Thank you very much. two years ago, search began to appear within the LLMs, within the chatbots, right? Which already started blurring the line between humans and agents, because now the same bots also have the same web

4:12

search available, so the same LLMs have the same access through API for the bots. So we started seeing that convergence happening, right? And so for the first time, we're seeing more and more traffic flowing down, search traffic, search intent flowing down these channels, not only to Google. And last but not least, right, we now see a whole breed of companies, the AI search companies. I'm sure you're familiar with them. I caught a talk yesterday by Will, the CEO of Exxon, and Parallel, and New.com, and Tavili, and a bunch of others. They are purely built and indexing the web, especially for agents, not even looking at the humans involved anymore. So Google's dominance,

4:49

when it comes to search. If Google was synonymous of web search, that is very much shaky. And when there's blood in the water, the sharks come. Just last week, Amazon, Amazon announced, I don't know how many of you saw it, that they developed their own index and started allowing the ability to retrieve data from the web to retrieve context for agents on AgentCore. Amazon developed their own search engine. Two weeks before that, it was Microsoft. Microsoft always had skin in the game. I don't know, one or two percent of the world search traffic went to Microsoft. But they have now repackaged it and

5:27

launched it again as part of WebIQ, as part of their suite for agentic development and orchestration. So we're seeing more and more this space becoming crowded. But when we're talking about context, and we're looking at this through the lens of search, I believe it only tells us part of the story. Right? I can search for, I don't know, what's the cost of a certain pair of sneakers this morning, right, on a certain website. I cannot really search for how has that price changed over the last six months, what discounts it had. Right? I can search for what open job positions we have at Bright Data.

6:08

We do. I urge you to go have a look. But I can't see how that was a chart and how that changed over time and how the headcount of the company changed over time. And all of that information existed on the web simply back then. So when we actually start to think about it, we understand that there's much more context in the web than what web search allows us to extract. And this is what we started seeing in the recent years, a whole new breed of companies rising. We like to call them internally CAS, context as a service, because that's what they do. They allow agents to tap into them, MCP, CLI, just pure good old API,

6:46

and actually start extracting data to retrieve data so they can reason over for whatever knowledge work they are responsible for. We see this happening in e-commerce. We see this happening in travel. We see this happening in in finance, in market research, in HR, in real estate, in a bunch of other domains. I'll show a few examples in a second. Right? What all of these have in common is that they don't only just discover the web, you know, in terms of think crawling, think searching, think all of that, accessing, extracting the data, and indexing it. They take it a step further. They actually develop knowledge graphs to start structuring all of the

7:20

entities and to dedupe them. And they start enriching them with a lot of different sources. So they actually start merging all of that data. If you think about it, they kind of behave like vertical search engines. Right? They are a very, very, very good search engine for something very specific.

7:38

And it's already in full motion. Right? So as I said, we see this in finance and in market research and retail and e-commerce and GTM and sales intelligence. What all of these companies, by the way, have in common, they're all part of Bright Data's startup program. If you are a builder, and this is a hot space to go in because I think we're only tapping the surface, I invite you to scan this and apply up to $20,000 in credits and all sorts of co-marketing. But that's enough self-promotion.

8:06

So CAS as an industry is already in full bloom. And as always with these situations, the traditional players aren't left too much behind. They're at the bottom. You see good old data as a service. You see ZoomInfo. Right? By researching for this presentation today, I also saw that they launched that thing at the top. It's called GTM.ai. You can only imagine how much they paid for that domain. But they launched a secondary brand for ZoomInfo that is catering specifically for the need of agents. Look at the wording. Right? They talk about GTM work. Right? That knowledge work. That research

8:43

that you do when you need to prospect, when you need to do headhunting, whatever it is you need to do that involves people mostly. Straight from cloud code, straight from codex, or any other agent. They understand the gap. Right? So yeah, it's fun to think of CAS as an evolution of this. And it is in a way. But it's catering for a very specific need as much as the AI search engines are different than Google. When agents need them, it's different than people. When we let this one sink, that at the very least, we have two different types of paths to complete knowledge work as an agent. We can start thinking

9:18

about this in terms of web context engineering. We can start thinking about this in terms of how do I optimize for the specific task. More importantly, when things come as it is, how do I optimize this for breeds of tasks? How do I do this for various parts of the organization that I'm building for? If I'm an AI engineer, I need to serve different teams. They may have different needs. It's very tempting to throw AI search at all of them, but maybe that's not optimal. Maybe I need a combination of both. Maybe I can start seeing all sorts of cost efficiencies emerge from that. So for the second half of this presentation,

9:54

we actually went ahead and created a test. This is not a benchmark. You won't see any something concrete that I can say with great confidence other than the actual research that we did because we wanted to start unraveling the different considerations and how do these two stack up against each other. So we designed a test. We went for something basic. We said, okay, let's take a company, an entity, and try and enrich it across 25 different fields. Some of them are very easy, you know, the company domain, the name, the headquarters, but some are more challenging, right? Things about hiring and people and something.

10:26

And we build a simple agent, a loop in a loop that uses Opus 4.8 as the harness, and it starts to go field by field, go out, search for it or retrieve it from the cast, do it again and again and again until it completes and brings back, set some guardrails, you know, like budget and stuff just to keep it fair. And I'm going to share with you the results. So the first thing that we would care about, right, being knowledge work would be, sorry, we ran it 100 times on all of the sponsors of today's event. So the first thing that we saw in terms of coverage is that there's pretty good convergence.

11:04

They all did fairly well, right? I'll get to the two at the bottom in a second. So search were consistent performance. One of the major cast providers were also very well. The third one, by the way, you can see Unlocker and SERP. SERP is good old data, good old Google. We basically did the same thing just with Google and it performed pretty well in extracting that information. Native is Claude's own search and you see that they converge really well. I was originally surprised about the two cast solutions at the bottom. It was counterintuitive. I expected cast to dominate this thing because that's,

11:39

you know, you had one job, right, to map out these companies. But after diving into it a bit more, you understand that, well, they are limited in the sense that they know what they have about an entity. If I asked it a question that is beyond that, they will never have that data, right? Unlike a search that can go out and continue searching and exploring it, if they didn't collect data about the recent job hiring, it will never be there, right? So it makes sense that they are a bit behind, but I'm sure at the same time that they have a lot of other advantages that we simply didn't ask for, a lot of other fields that

12:12

they didn't have that aren't represented. So again, it creates some complexities when how do we measure coverage when it relates to the specific job that we need to do rather than in general. The second thing we looked at was cost, of course. Here we started seeing it spread out a bit. So you can see that massive bulk in the center, most of the search and the cost and even using Google, right, converged to pretty much the same cost, only different, right? The CAS was just about the service itself, what you pay the vendor, right? All of the other search solutions, you also needed a

12:48

lot of token burn to actually structure that data so you can actually act on it and use it as something retrievable, right? So it's the same output. Native, obscenely expensive, and the CAS on the right, I'm sure you're all familiar with, they're about far the most expensive in the industry. I will not name and shame them. Interesting, you see that small CAS there at the left, that CAS number two, that were very cheap. They're also the ones that are here at the bottom, which is funny because what I believe is happening there is that we're seeing even within this industry niche players that have,

13:24

you know, lower quality data but much cheaper. They're already carving that niche of the long tail, right, of small shops or small usage that don't want to pay as much and don't need as much data, data, and we're all seeing them branch out there. Most of you here, I presume, are engineers. So there's a very evident question that we did not ask here, which is, what is the one thing that an engineer would care about? Thank you. Let's talk about scale.

14:05

What happens if we need a million? Now, yes, a million records will not fit in the context with obviously, we're not talking about a single run that needs a million. You can think about million in terms of the frequency, right? If I am a market researcher, I do due diligence for private equity. I revisit these companies all the time. I ask more questions about them as the time goes by. Was there any new news about them? Was there anything that changed? Did somebody join? Did somebody leave? Do they have new hires? I keep on asking the same thing. So when I'm talking about this, the multiply by a million, it's not

14:34

just about the number of companies. It's the frequency in which I'm asking it. Frequency is the cost killer when we talk about these. And we need to acknowledge that, right? We're thinking about this in terms of web context engineering. We're starting to look at it differently. Every repeated query costs the same as the first. Even if it brought back the exact same answers. Nothing changed? Pay up, right? No, it's false positives? For sure, go in. Token costs, right? We saw that there's very high token. We know that doesn't shrink well over time. There's always some volume element in terms of the cost, but it's not the

15:11

same as flatlining, right? And if we bring this back to knowledge work, this is where we see teams that are starting to cut corners. So I won't research this company every day. I'll look at it once a week or once a month. I won't ask that question now. I don't want all the results. I'll only take 10 results, 20 results. So we already have the setup. We have what we need to do the knowledge work. But at the same time, we're not extracting all of the value because we're starting to be conscious about cost, right? Basically, we're renting context. We're not owning the context that we use.

15:45

That is a very important distinction. Again, if we're good engineers and we ask ourselves, what about scale? The second most obvious thing that will come to mind now, so how about we build it? What if we take all of that web data ourselves and stick it in some vector database and try and see what comes out of it? So I asked my engineer to do exactly that. Again, this is a test. This is not a benchmark or a full-blown operation. This is a day's work at best just to illustrate the concept and to show something about the cost efficiencies that you can generate potentially by doing it yourself. Potentially in specific scenarios.

16:26

The test? Simple. Take the company name, nothing but, run it through Google, find the relevant entries, the relevant URLs of that company in various websites that have all of that information, right? You use search when you don't know the source. But when we're talking about company enrichment, we all know these sources. We all know where that data comes from. The cast also bring it from them. Zoom info bring it from them. It's the same thing over and over again. Why not just go straight to the source?

16:49

Why are we doing that middleman thing? LinkedIn companies, LinkedIn jobs, Crunchbase. Right there, we have scrapers for those who just tap in and you start paying as it pays as you go. We build two dedicated scrapers. We have a new AI tool called Scraper Studio. It basically lets you build a scraper for any website in less than five minutes, all powered by AI. And then it also has a self-healing function, right? So if the website changes, it fixes itself and keeps on going. Merge it all into one entity. Basic heuristics, if there's conflict, choose that over that. And eventually, we have a data set of these 100 companies.

17:28

Zero AI cost involved. There's no tokens. Coverage? Fairly well. Not amazing. Not the best that we saw here. But stacking up pretty well. And again, this is just a day's experiment. Probably not even as much. Okay? So again and again, very specific tasks, very limited context, very limited situation. Tread lightly and proceed with caution when it comes to conclusions. The real story is not this. The real story is this.

18:03

That's what it costs. To just go and fetch that data that is out there. We think about knowledge graphs. We think about entities. But if you think about, for example, LinkedIn, the data is already structured in form of entities. There's an entity for a company. There's an entity for a person. There's an entity for a job. And they're connected between them. Sometimes the ontology is already there. Again, this is not the most complicated of scenarios. But this is pretty damn good. Now, yes, it took time to set up. So it's not really fair to compare apples to apples when it comes to the cost, because these are out of the box.

18:39

You can just tap into the API. That one that I just showed you required some setup. Let's say it's a week. Let's price it at $5,000 just to give us some perspective. We can actually start thinking about this in terms of a tipping point. We can actually start thinking about what is that tipping point in which it makes more sense for me to build it myself, right, than keep on renting it.

19:04

So it makes sense to do it at this point. Maybe it's not 15. Maybe it's 30. Maybe it's 100,000. Maybe it's 10,000. It really depends on the use case. But there is a tipping point in which it actually makes sense to do it yourself. Which leads us to the fact that both AI search and CAS and all these solutions, they're very good in the sense that you can just plug and play. But if your knowledge work needs, right, are persistent, inconsistent, and to a certain degree may even continue escalating and growing, then this is perhaps a direction to start considering.

19:42

Maybe I can just go ahead and build my own. Because the nice thing about it is that all of the things that we see on the left, up until the third part, is upfront investment. And the most important thing that whatever retrieval happens later on from the agents, right, is free. Not really free, but you get what I mean. There's no added cost. I can just ask that question over and over again. I did not like the first answer. I'll ask it again. I'll ask it 100 times until I get what I need. I have no more fear, no more cutting corners, which is maybe the most important thing.

20:16

And I'm leaving aside the fact that this is also custom business logic. I can connect it with my own data. There's all sorts of other advantages of owning it. You know, we'll keep it to the imagination.

20:28

Remember, we asked about a million, not about 15,000. This compounds. This compounds greatly, right? We need that horizon. Remember, the web keeps changing. We saw the staleness of the data and how the data decays. So we need to be thinking about this in the long run and how this will evolve when we keep on asking the questions about the entities that we care about.

20:50

Just to wrap it up. So AI search costs, they can get you very far when what you need is ad hoc and what you need is always changing. When sometimes you look at different things, even the mix and match of them for certain tasks use this, for certain tasks use that. You can, I'm sure, again, I just showed a test. There's a lot of ways to optimize it just like any other context engineering and make and use lighter models and use other stuff. There's a lot of great stuff to be done. But eventually, the frequency will come and bite you in the ass when it comes to cost.

21:19

And that's something to be mindful of. And there's a fair chance that that tipping point is much lower than you think. And that's something that, as we design these systems, when we think about web context engineering, we need to be mindful of that. And last but not least, the last slide we showed. Owned context compounds while rented decays, right? It's not a one-time task. Again, if it's a one-time question, use AI search. It will be amazing. When you need to do it over and over again, there's a fair chance that it will not, that you're missing out on potential compounding effect and you are losing out. Thank you very much. .

22:14

.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note