Agents will use the web thousand X more than humans. Hence, new tech is needed and new business models are needed. How does the world of agents change web search in terms of the technology required? Parag is the founder of Parallel, changing the future of how agents do web search efficiently. This is an incredible discussion on the future of agentic search, the future of income and wealth inequality, and so much more. Parag rarely does shows, and so it was very special to sit down in person with him in London. Ads don't work with agents in their current form. Some people actually want to do bad things, and the model's alignment is not adversary-proof.
And I think those are the things we must worry about more. I think it's the responsibility of people building models to ensure that you really do the work to minimize that harm that comes from what you've built. And I think so far, I think some of these should be considered embarrassments because I think they demonstrate two things. Ready to go?
Parag, I'm so excited for this. I spoke to Vinod. I spoke to Andrew Reid. I spoke to Todd Jackson. I stalked the hell out of you, so thank you for joining me. Thanks for having me, and thanks for making all the calls. Not at all. I would love to start with, for anyone that doesn't know, how would you describe Parallel in 60 seconds? Parallel is the Google for agents. So agents need to search the web to do anything they do for you, whether it's a personal agent or an agent built for work. Just like humans need to go on a browser, search Google, oftentimes during work or for whatever you're doing in life, your agent needs to do the same.
Turns out agents are different from humans, and the way you build web search for agents is different. And so Parallel is about building the technology for agents to search the web, and then the business models to make that sustainable. Was that the original insight that you had? Yeah, literally the first genesis of the company was us writing down the statement that agents will use the web thousand X more than humans. I wrote that down at some point. Hence, new tech is needed, and new business models are needed. 1000 X gives you a sense of scale. It changes how you think about building the tech underneath, because no tech built for a certain scale
survives three orders of magnitude. And then when you need new business models alongside new technology, a problem becomes really interesting. How does the world of agents change web search in terms of the technology required? There are many layers to the answer, but let's start at the first thing we mentioned, which is scale, right? Now, if you think about agents actually searching the web a thousand X more, if we spend the amount of compute we currently spend on web search, that's too much compute for web search. So you now need to make it way more efficient by perhaps 10 to 100 X for it to make sense. If you're doing a thousand X, exactly. The second thing is
agents expand how humans operate in a very narrow zone. So if you think about how we use web search, we type keyword queries, which are short and underspecified. We wait for about half to one second. If web search takes more than that, we're impatient. And then we get 10 blue links and then we random walk across them and a collection of searches to get what we're doing. Agents are not like that. Agents are going to perhaps tell you exactly what they're looking for, not like three keywords, but like a full sentence, like this is what I'm trying to do. Agents will either be super impatient, like imagine a voice agent. The agent will be like, I need an answer now,
100 milliseconds. I can't wait 500 milliseconds because the human is waiting on me and they're going to wait 500 milliseconds. So I need web search to do it in 100. Or there's going to be a background agent who's like, I don't care, just give me the best answer possible. And so the variance of what you can do within web search changes completely. In fact, the one thing that is the least interesting is what we've built for humans. You either have too much time or too little time. Almost never the same amount. The output is not the same. So the input is different. The time you have is different. The output is different. The output is in blue links. The output is tokens
or files on a file system, depending on the type of agent. And so now all of a sudden you say, okay, now the problems, inputs are different, outputs are different and constraints are different. And you get to spend very different amounts of compute on it. Imagine someone running an agent built with a Llama model. And then imagine someone running an agent built with a Grok model. They're very different models. How you want to optimize signal to noise in tokens for each of them is so different in terms of what you do in the web search stack. So what do you mean optimize signal to noise? So think of it this way. Let's say you took web search, which was cheap and fast
and low compute. One way of conceptualizing the web search problem is you start with a trillion documents that are on the web, some few trillion. Given any search, I now need to narrow it down to a thousand tokens that your model's context window should see. So the problem is going from a trillion URLs with let's call it a few thousand tokens each down to a thousand total tokens. So how do we do it in web search? We first say, okay, we're going to do retrieval. So for most documents, I'm going to spend zero compute. For a small number of documents, I'm going to spend minuscule amounts of compute to figure out which 10,000 to look at. Once I get these 10,000,
I'm going to spend a little bit more compute for each of these 10,000 documents to narrow it down to 1,000 documents. And I'm going to keep doing this with bigger and bigger models and rankers with more and more features until I can narrow it down to a thousand tokens for your model. So you're essentially allocating compute to web search in order to save compute on the model. That's roughly what's going on here. So if you want to save compute for the model, the question is how much compute should you allocate? So if your Llama model is really cheap, you don't want to do too much compute in web search because it's okay to leak a little bit more information
into Llama's context because it's cheap. Into Grok, you want to do the work before you waste Grok's time because that's going to be expensive in time and money for you if web search gives you worse answers because you cheap it out on web search. Can I ask, how do you deal with the ambiguity of what agents want? And what I mean by that is, you know, for different things an agent might want different things and for different people an agent might want different things. I may really care about accuracy and not at all about latency or cost. I may really care about latency but not at all about accuracy. You allow the agent to specify that in the API signature. So, our product
is an API which either the programmer can configure based on their application or can leave it to the agent. Our API even has a parameter which is like what's the model calling me? And if you know the model we can do things differently. Now, you don't have to tell us but if you tell us you might get better results. You can. We've productized our search system into a few different modes each optimized for a certain class of use case. For example, we have a really fast API the fastest in the market. It's fast often cheap because with low latency budget there's only so much compute you can do. And it's built for voice agents. So your voice agent must be all knowing without
telling you I'm searching the web hold on while I come back with the answer. That's a silly experience for a voice agent. It should just immediately respond. And so you now need to do web search which is rapid. On the other hand you can take a Grok model and the thing is going to think for 20 seconds and then generate for 8. You can spend 5 seconds on web search to make sure it does into 5 web searches only 2. So you end to end in the agent you save time and cost. And so that's a different processor on our system it's called advanced. So if you're building a voice agent you use Turbo. If you're a really expensive background agent, you use advanced. And so the primary
use case today in terms of customer base is engineering and coding for you? It's pretty broad-based. I would say the primary use case is the common theme is knowledge work. So coding is a category of knowledge work. So are AI lawyers. So are productivity applications. So are AI insurance underwriters. So are AI scientists.
And so you now need to do web search, which is rapid. On the other hand, you can take a Fable model and the thing is going to think for 20 seconds and then generate for 8. You can spend 5 seconds on web search to make sure it does into 5 web searches only 2. So you end to end in the agent you save time and cost. And so that's a different processor on our system it's called advanced. So you've Turbo if you're building a voice agent you use advanced if you're a really expensive background agent. And so the primary use case today in terms of customer base is engineering and coding for you? It's pretty broad-paved. I would say the primary use case is the common theme is knowledge work. So coding is a category of knowledge work. So are AI lawyers. So are productivity applications. So are AI insurance underwriters. So AI scientists. I'm trying to understand how much of the mother load is engineering. Is it 80%? No. The thing about engineering is coding is a large chunk of inference in the market right now. Coding invokes I would say web search in 5% of prompts. Whoa. Right? So it's not every prompt invoking web search when you're writing code because most of them rely on your internal context and your internal data and your code base. So the model is spending time reading your internal code and not on the web. Law is not going to be much more, is it? Law is very web search oriented. Really? I thought it'd be internal data driven. There is internal data, but there is case law, there is facts, there's facts about companies, facts about people, and you have to go exclude information. So you have to be very comprehensive in law to say, I want to be confident that despite a lot of effort, you can't find this. Insurance underwriting has the same flavor. So sales is very web search heavy. AI science is very web search heavy. So on a relative basis, inference goes more in compute, but when you think of web search, all of these others start popping in. How does the rise, we were talking before about muse and instinct, if there's anything that needs web search, it's personal assistance. It's great. How does that change? Yeah, it's great. This is the greatest thing for my business. How does that change your business? I think in our business, anytime agents start becoming more useful for more use cases, because models either get better or cheaper, it's great for us, right? Because we bet that agents will be the consumers for the web. And we've been building tech for agents. So when agents do more, it's great for our business. So we want models to keep getting better and cheaper so that agents do more and more and more. And if that happens, it's great for our business. Totally get that. Can I ask, if models get smarter, don't the agents beneath them do fewer searches and then it's worse for your business? I don't think so. If you think of models, there is the trend around models being smarter and then models having more memorized and parametric memory. Those are two slightly different dimensions. And if you think of the, today by and large, I would say models have good recall from parametric memory on, let's call them head facts. It's like for somebody famous, everything about them, Wikipedia, it can memorize. So it'll tell you who the president was in a certain year, right? Because a model can memorize those things. The model couldn't tell you what year I graduated from college. Really? Maybe for me it can, but it can't tell you for somebody who works at parallel. Do you not think so? Even if it was in the pre-training data. Seriously? Because it's lossy compression. So what a model's parametric memory is doing? It's lossily compressing to understand patterns in the world. And so it can't memorize every fact in pre-training data. So one, not every fact is in pre-training data. Two, it can't for even stuff that's in pre-training data. The model's actually trying to find patterns rather than memorize them. And then further, as you make models efficient, which is you make them smaller and smaller while keeping the performance by distilling them or whatever, you lose more of the parametric memory while you try to keep the reasoning. Do you think we will see models become smaller and smaller than every company have their own model with their own data and the fireworks theory of own your own intelligence being true? There are two questions in there. So one, I think we're going to see two things. We're going to see the biggest or the frontier models be bigger and bigger over time. We are also going to see smaller and smaller models being able to reach any fixed level of performance. So if you say, okay, I want Opus 4.8 level of performance, okay, that's good enough for my use case. Every six months, a much smaller model will be able to deliver that to you. So you're going to see this, the useful range of sizes of models will be way, way, way different. What does it mean if the frontier get bigger and bigger? What are the ramifications of that? One, they are better. So all of the ramifications you imagine. So the reason the frontier will get bigger and bigger is ultimately the gap between what the value for certain use cases incremental quality can provide you can be so high in certain use cases, which can be so valuable that it's worth paying for. So if you can build it, there will be use cases for it as long as by making it bigger, you can make it better. It appears there is no end to the scaling law that we can perceive so far. And again, I'm conflating bigger with like models are getting bigger, but they're also able to think longer. So you're just able to throw more compute at the same problem. And that will keep happening. I think we'll be able to throw more and more compute at the same problem and make the answer be marginally better over time. And so we're going to just spend a lot of money on extremely large models solving really hard problems. Do you agree with the consensus for you that you'll have 90% of token activity go through open models, but 90% of dollars go through frontier models? I don't know enough to have a view there. I don't have a view on that. I do think I don't think it'll be 90% on either of those two actually. Really? Yeah. Why? So if you believe my claim that a useful model and the frontier model will be 100x, 1000x off in price from each other. So this small model is still useful. It is hard to know which use cases over time will get optimized to which scale of model in between. And I think there is a real path dependency in terms of where open models end up. I do think if we had strong confidence that an American built open model was going to be state of the art as far as open models went, I would have more confidence in saying that they'll be pretty good. Do you have confidence in American open models? I want them to exist. So far, it's not clear what's going to happen, but I'm hoping that there will be a great series of American open models and perhaps even competition to have the best American open model. I think what you need is actually not just one person motivated to build an American open model. You need two people competing against each other to build the best American open model. Do you think there's value in the model routing layer, as Americans call it? Some people think immense value, some think commoditization. Again, there's path dependency there. So today there is real value. Today there's real value because when we say routing, we talk about two different dimensions, which model and which GPU running that model via which vendor. When you're in a world where the demand supply is very weird and people are hunting for GPUs and people need capacity to serve their customers and people, all of these are relatively early stage startups growing really rapidly, sometimes beyond what your forecasts or predictions say. And then all of a sudden you're looking for capacity. And so you want to be able to get it where you can. And so that creates real value in having some of these routers as log-in-take is all these two problems, saying give me flexibility if I need to get a different source of tokens. And push comes to shove, there's no tokens on this model in the SLA I want. I'll switch models as well. So in this moment it's really valuable. Now I don't know what happens to overall demand supply on GPUs and tokens. But if it remains this way, the routing layer is really valuable. What do you think is not so valuable today that will be incredibly valuable in three to five years time? Perhaps data. When I say data, I think it is, I think today we don't know how to pay for unique, valuable insight or data. Because we
Forecasts or predictions say. And then all of a sudden you're looking for capacity. And so you want to be able to get it where you can. And so that creates real value in having some of these routers as a login take is all these two problems, saying give me flexibility if I need to get a different source of tokens. And push comes to shove, there's no tokens on this model in the SLA I want. I'll switch models as well. So in this moment it's really valuable. Now I don't know what happens to overall demand supply on GPUs and tokens. But if it remains this way, the routing layer is really valuable.
What do you think is not so valuable today that will be incredibly valuable in three to five years time? Perhaps data. When I say data, I think today we don't know how to pay for unique, valuable insight or data. Because intelligence is cheaper. You're going to want to build upon either data or insight that comes from somewhere else with your unique data, right?
So if you think of an abstract notion that I have intelligence here. I own some data. You own some data. And there's some data in the public domain. If we can pull together your insights and my insights and all of this data in the public domain and intelligence on top, we can create something bigger than what you could have done by yourself, what I could have done by myself. So now how do you transact to create this sort of whole that is bigger than the sum of what you could have done yourself and what I could have done or what open data was good at? And so that transaction feels like a data transaction to me or an insight transaction to me. And so we don't yet know how to transact that way. So we fall down to, okay, all I can do is use my data and use open data. And let's see what I can do with it. But if we figured out better ways of pricing data, good things will happen, but it's not something that's yet a market.
But what I'm talking about is data transactions at inference time. So it's not necessarily training time. We are building knowledge. And let's take an example of today's world. Let's say in your job, since I'm sitting with a VC, you probably have access to a pitch book or a product like that to collect data. And so you get grounded in knowledge of what's happening. And you get to use that data to figure out how to make decisions in addition to all the notes you have and deal memos you've written perhaps over the years or insights you've had about how to choose founders. And you're essentially composing insights from your personal lived experience, your notes, public data to make decisions.
Now, you pay pitch book by the seat. But now you're running a bunch of agents. And your agents can perhaps not access everything you can on pitch book. Or you're doing this to get a browser, perhaps against terms of service to send an agent via browser to pitch book using your auth credentials. And it's inefficient, it's clunky, right? But clearly their data is valuable. And clearly your agent should have convenient access to it. So if we figure out how valuable pitch book data is for your agents to make your decisions, because clearly you're investing a large hundreds of millions of dollars. So clearly you presume that you make a great return on that. So this data is truly valuable if you depend on it today. So you should be able to figure out how to compensate pitch book for the data, even when agents use it, which is not by the seat.
If we figure that out, it will be a huge market and agents will be better off. If we don't figure it out, you're going to be in this weird cat and mouse game of pitch book being like, no, I want to sell you a seat. I'm going to shut it off. And your agent is stuck without the data. And then you're involved in pulling data. I think every big company has the choice today of do we let agents in or do we keep them out? And Amazon has said in most recent times, Muse, you will not be let into our garden. Shopify has said, come in, baby. Expedia has that come in. How does this play out?
I think eventually everyone has to let them in. The question is on what terms? So how do you align incentives for everyone? And people are going to have to play their strategy games on how they win in this new world. Literally the customer is changing in front of our eyes. The customer used to be a human. You're building products for humans. Now you're building products for agents or you're building agents. And so now you have to figure out what your place in this new reality will be. And I don't know if there's a right or wrong answer here. It also depends on how much market power you have.
So let's say you are an individual. Do you think Amazon were right to say no to me? Companies? I don't. Depends on what they do next with it. It's like if it turns out that they have sufficient market power to have Muse or other agents connect to them differently perhaps over time or ship their own agent and drive crazy adoption. Let's say they have that capability. Then I guess they were right. Right.
On the other hand, once they've made this, we're not glad now anyone else's agent. And they can not ship an agent that consumers use. And they won't allow anyone else's agent to use them. And a lot of the transaction economy starts moving off to agents, which are all big ifs, by the way, then it would be a bad move. Now, I'm betting on agents. I am. I have mixed feelings around agents and what fraction of e-commerce transactions they do. Like it's unclear, right? Do you think it's unclear?
I think it's unwaveringly clear. Maybe I'm super early on the adoption curve. I buy everything through Instacart now. I mean, other than holidays and a home, I'm literally just everything's through Instacart.
So me too. But I also know a lot of people who like to buy things themselves. And listen, we want to delegate to agents things which we see as chores and uninteresting. Things that give us joy or pleasure or make us have fun doing those things. And I don't know people who don't want to delegate shopping. They want to delegate a lot of things in their life. They don't want to delegate shopping. And so I don't know, that's why I don't know the distribution of these people and how behavior changes. But it might be people give away a lot of other chores and then spend a lot of time shopping.
One thing that does worry me is actually the dissolution of the advertising industry in the wake of agents. If I have, you know, if I order my delivery and my Uber DoorDash for dear American counterparts through Instacart, that Uber banner that's now advertising something becomes worthless. Amazon, their advertising business is bigger than their e-com business now. Of course they're shutting it off because if agents are the primary customer, your ads business goes to next to nothing.
Correct. And this is not just you talking about this in the e-commerce land, right? But if this is the problem with everyone, if you think about this, this is what I meant early on when I said the business models have to change. Ads don't work with agents in their current form. So forgetting e-commerce for a second, if you think of you're in the content business, let's say your page which is ad supported, public information on the web, ad supported page, people show up, you show them ads, you make money, great content. Agents show up, no one sees ads, you make no money.
Which is why we like to pay people, so we're effectively building an ad sense for agents showing up to read your content. So we like to pay content owners a variable amount of money every time an agent derives benefit from reading their information, which is a variable amount of money is what a visitor to your website will pay you based on a click probability or a value probability on ads. Why do you do that? To incentive align content owners, otherwise what's going to happen? Everyone's going to block agents. So you need to find a replacement to the ads business model.
I get you, but you have to assume that you're going to be like 100% of the market then, because if you're 30, 40% of the market and you're like, oh, don't worry, we'll pay you and New York Times is still like, well, thanks, Parag, but 60% of my traffic is still unpaid, so I'm just going to block all of you. No, but I think they're going to block them and not me. They're going to give me a data feed. What does incentive alignment mean? It means the New York Times believes that I pay them a competitive market rate.
Money is what a visitor to your website will pay you based on a click probability or a value probability on ads. Why do you do that? To incentive line content owners, otherwise what's going to happen? Everyone's going to block agents. So you need to find a replacement to the ads business model. I get you, but you have to assume that you're going to be like 100% of the market then, because if you're 30, 40% of the market and you're like, oh, don't worry, we'll pay you and New York Times is still like, well, thanks, Parag, but 60% of my traffic is still unpaid, so I'm just going to block all of you. No, but I think they're going to block them and not me. They're going to give me a data feed.
What does incentive alignment mean? It means the New York Times believes that I pay them a competitive market price or the right price or an attractive price or a fair price. If I do that, they should give me their content. And if somebody else doesn't give them that, and if they have the technical levers or legal levers, they shouldn't give it to them.
Do you think this business is a little bit like music with streaming, which is the business just becomes much worse for the creators, and it still provides them money, and significant money, but they have to get a little bit more creative with alternative streams, touring, merchandise, alternative business. Is it the same where your core goes down and you have to get more creative?
I don't think so. I think there's one fundamental thing that is different here. I think agents using the web a thousand X more, that thousand X is really different. Now, all of a sudden, the amount of utility added goes up if these agents are presumed to be doing something useful. And so, this is not just a change in the share of value that people transact over. This is a large growing pie. And when that happens, it is actually possible to transition business models. Of course, there are going to be winners and losers. Some pieces of content will become very valuable. Some pieces of content will get more commoditized, and people are going to have to adapt to a new customer, to a new market dynamic, to a new kind of monetization engine. But the overall market size, I think, has the potential to increase, unlike in music, where it took a while for it to grow back. My understanding is now the industry has grown back up and exceeded its previous peaks in a material way. But there was a moment where it was smaller. In this case, that might be a very compressed period, given how fast these things are growing.
Can I ask you, everyone questions the sustainability of margin structures in this business. How do you think about that as a business today? Today, we're in the infrastructure business. And we have a real technical lead in terms of being able to do things at very high quality, very cheap. If you think of the tech we're building, what are we doing? We like to give the highest quality answer. Make it fast, make it cheap. Spend the least amount of compute doing it. We obsess about only these three things: quality, cost, latency. Turns out if you do that, relative to a stack built for humans over the last 20 years, you can do things at the same quality for 120th, 150th of the compute. And so, a lot of margins come down to market structure on competition over time, rather than any other factor. In our industry, the market is large. We're too early to know what future competition and market structure looks like. But the margin change due to content owners is not something I worry about.
Let me tell you why. The entire premise of us paying content owners is driving incentive alignment. We like to pay content owners the marginal contribution that they added to an agent doing work. What does that mean? Let's assume that there's an agent trying to get something done. And you're going to spend a dollar on that agent to get this thing done. Now, if you've spent a dollar and 10 seconds, because we have a larger model available or ways of spending compute, you could have gotten a slightly better answer. If you've spent 90 cents, you would have gotten a slightly worse answer. If you're running good agents, they're on this Pareto curve. So, now imagine I took out one content owner, their data from the web index, and I still spend the dollar. But let's say after taking them out, the quality of the result was the same as the 90 cent agent with their content. So, why wouldn't we go and take this 10 cents of marginal contribution they had as a content owner and pay out a decent chunk of it to the person who brought that data?
Yeah. So, that's how we do our math. That's how we've trained our models, which tell us how much to pay for what content. And so, to produce, if I was going to offer an equally good product to my customer, I'd rather pay content owners than spend on inference because I'm spending the same amount. So, they're all big numbers. But then it actually comes back to revenue at a certain point. Companies are scaling faster than ever revenue-wise. Is it a business where your revenue is able to scale as fast as others? It scales. One mental model of our business is we are an adjacency to inference or knowledge work or personal agents. Take out GPUs going into media generation or training. Take training away. Take media generation away for a moment. Look at all of the GPUs. Whether it's small models, open models, proprietary models, forget all of that. Every bit of inference across all models going into running agents. I think somewhere between 5 to 20% of that spend that goes into GPU will need to go into some sort of a web search stack. Now, if you look at how many data centers we are going to build and how much power we'll generate and how many GPUs we'll build and the scale of this buildout and at what rate that's growing, this is a very material market. And if inference grows 3x, 5x, 7x year on year, we grow alongside it in the same rate if we're just holding our share constant. If we're growing share, we're growing even faster.
A firework scales to 2 billion in revenue in four years. Is that a similar revenue trajectory that you can follow?
Yeah, but I think when fireworks is 2 billion in revenue, the inference market revenue is 200 or more, maybe 300. And so, our market potential, based on my 5 to 20% math, is, whatever, call it 10 to 50, 15 to 60, whatever. And what fraction of that can we capture? That dictates how fast our revenue can grow. Today, I think we grow alongside inference and more. If you think of it like a 20 billion dollar there, let's just take that kind of middle, if you assume a 33%, which would be a lot of market actually, it comes down to it, that would be, whatever that is, 6 point, whatever it is, 6 to 7 billion, give or take.
That's amazing. But before you told me it was a 100 billion dollar business plus, that doesn't get you to 100. In valuation, it does. We were discussing valuations earlier, not right now. So 6 billion growing at the rate of inference gets you to 100 billion business easy, but that's in the next few years. And you think that's possible in terms of that revenue scaling that fast? Yeah, I think in the next few years it is. You think 2030, we're going to sit here? You think you can be there? I'll get a tattoo for parallel if you can. I think it's possible. I think we're going to have to execute well and a couple of chips have to fall our way.
What would be the reason why you don't? One reason would be that we see agents don't work. There is a tail risk that agents don't deliver on the profits and we overshoot as a society, which is extraneous to us.
If agents work and deliver value and we spend all the money on GPUs that we plan to right now, then the main question is, did we execute well enough to have a 33% share that you bet on? That comes down to, in my mind, did we build the best tech? Did we partner with all the content providers to have their content available? Because without it, it's not very useful if you can't rely on it. And did we earn the trust of the customers that end up having a decent chunk of that inference in four years? have to fall our way. What would be the reason why you don't? Let's map out the, we have to get gnarly about these problems. So one reason would be that we see agents don't work.
There is a tail risk that agents don't deliver on the profits and we overshoot as a society, which is extraneous to us. Yeah. If agents work and deliver value and we spend all the money on GPUs that we plan to right now, then the main question is, did we execute well enough to have a 33% share that you bet on? Right? And that comes down to, in my mind, did we build the best tech? Did we partner with all the content providers to have their content available? Because without it, it's not very useful if you can't rely on it. And did we go earn the trust of the customers that end up having a decent chunk of that inference in four years from today?
And then again, there's large cone of uncertainty, right? Like, I'd say there's the five-inch uncertainty in fireworks share at that point. If open models crush it, like fireworks will have a huge, base 10 fireworks model, all of them will have a huge chunk of revenue. So did we end up selling alongside them? Did we find all the customers which are spending money on them to spend money on our website by making it the best web search in the world? What about commoditization? If you have a read of semi-analysis where they did the benchmarking and then you were number one, woohoo, and then a week later you weren't number one, sad.
And there were three or four providers within very close proximity. You're like, oh, well if it's commoditized and I take a second layer thought to that, we'll see a race to the bottom on price. And then actually the available revenue, why am I going down the wrong pathway here? No, there should be a race to the bottom on price to get to the thousand X scale. I think the pricing on Web Search today is just off. Let me give you a simple example. You think customers pay you too much? Not pay us too much. I think the market is mispriced. Let me tell you why.
Let's say you used a Luna model with OpenAI's built-in Web Search today or with anyone's Web Search, forgetting ours, and you ran any kind of deep research. Let's say you built an instinct using a Luna class model and at some point a Luna class model will be able to do 60, 70, 80% of your personal agent and you'll do a lot of searches. In that moment, you'll be spending 80 to 90% of your dollars on Web Search and 10% on the model. Seems entirely silly. I was working off a 5 to 20% assumption earlier.
So I think at that point, Web Search has to drop prices by one order of magnitude or more because in my compute allocation dance that I was doing earlier, you need to spend less compute on Web Search than you spend on the model itself that's consuming its results. It's my hierarchy of how you do search. And so people have built Web Search the wrong way so far. Web Search in the market is it doesn't matter for an Opus. For an Opus, you win on quality and not on price because Web Search is such a minuscule portion of your, in fact, you should spend even more on Web Search. So we're going to ship even more expensive Web Search for bigger and bigger models over time.
And at the same time, we'll ship really cheap Web Search for the cheap models. But my rough intuition is that in an end-to-end agent doing a bunch of web work, you spend more on the agent than on Web Search for all agents. And so the market today on Web Search is totally off in pricing. Why are you paying $10 for, where the $10 for a thousand searches roughly comes from? Historically, Google's ads CPMs, which is how well does Web Search with humans monetize? It's higher than that.
And so people build technology to the point of like, oh, now if I spend a few dollars for a thousand searches and I make 30 to 50 or whatever, depending on the market and depending on who I am, it doesn't matter. We are in the right zone in terms of cogs of infra. And so I don't need to optimize it. Now I'm telling you that I can do it 50x cheaper while keeping the quality. And now there's a bunch of traffic of lunar class models where it's silly to monetize that way. And so I do want to race to the bottom. I want the best technology to win. And that's the only way you push web search to a thousand X.
But right now we're not in a world where there's much differentiation, correct? Sorry, I'm really dumb, which is why I'm a podcaster first, not a founder. When you look at the benchmarks, they present a quite clear view of everyone being similarly capable. Is that wrong? I think so. So one, I think these are all public benchmarks. Yeah. Which are all weirdly saturated. For example, I think I would not spend a moment looking at browse comp because half the models have memorized it. Models even memorize it. It's not really learning very much. So I think there's a lot of public benchmarks I don't give too much merit to.
But more importantly, our web search costs $1 for a thousand for producing that quality while almost every other web search in the market available to you right now will cost you $7 or $10 or $14. We're delivering this at one-tenth of the price. What will that be in three years? I think there is another 10x possible. That's extraordinary. So you're going to pay $0.10? Yeah, but you'll do it more than 10x as much. Because I think Jevon's paradox is this concept, obviously very well, but instinct, instinct. I do so much more than 100x. Yeah. Do you want to hear something I do on instinct, which is absolutely bizarre?
Every single country in Europe has a company registered that you have to register your new company with. I have an automatic alerting system built for every single company registered that has an under 25 year old founder who went to a top university through instinct. Amazing. Do you know how much that requires them to do this? But now you're on to exactly what I think the future of the web is. Today if you think of web and web search, we all think of it like what is web search? An agent gives, human or agent gives a search engine a query, gets results. So you're pulling information out of the web.
I think where this, what's next for the use case you highlighted is if you're an always on agent that's working on your behalf. It's kind of silly for instinct to wake up every six hours and go do a bunch of searching and bunch of inference to figure out if you need to get pinged about the under 25 founder that popped up somewhere. You know what's a better way of doing this? I am sitting here crawling all of the web all day every day at scale. I'm allocating compute every time I find a change in the web, every time something new happens.
If I know that Harry wants to know when this happens, I can do it at 100 or 1,000 the compute that instinct probably uses today to solve that problem for you. Sorry, and how is that? Because they would be constantly monitoring in real time across all these sites. Our crawl would be, so we have a product, it's called the Monitor API. What does this API do? Think of it as Google Alerts, except smart in the world of LLMs. So you tell it your query, which is when a founder under 25 pops up anywhere, give me a call. Now, one way of doing this is what the default implementation is. Run an agent, which is somewhat smart, to go search the web.
Did someone pop up and remember who previously existed? See if anyone new popped up, and then find a delta, send it out. So now every six hours you're spending some money. Now what we can do is flip it into an event-triggered system. We are crawling the web all the time. So now every time we see a change in the web, we can try to spend very little compute on it, to know, should you get a call, or does this trigger a more expensive compute, just as I described earlier. Now, we are able to do this way, way, way cheaper if you push the context.
Instead of the search system only finding out every six hours a new query popped up, it having long-lived context in the search system, what you can do is spend most of the time, the world doesn't change. A new founder does not pop up every hour or every six hours. So we don't need to spend compute every six hours. # Cleaned Transcript previously existed? See if anyone new popped up, and then find a delta, send it out. So now every six hours you're spending some money.
Now what we can do is flip it into an event-triggered system. We are crawling the web all the time. So now every time we see a change in the web, we can try to spend very little compute on it, to know, should you get a call, or does this trigger a more expensive compute, just like I described earlier.
Now, we are able to do this way, way, way cheaper if you push the context. Instead of the search system only finding out every six hours a new query popped up, it having long-lived context in the search system, what you can do is spend most of the time, the world doesn't change. A new founder does not pop up every hour or every six hours. So we don't need to spend compute every six hours. We only need to spend compute every five days. And perhaps we have one false positive, and then four days later, we spend more compute, and at that time, there's a real person that you should know about. And then you get a notification on average every nine days or whatever. But you didn't burn a lot of compute every six hours to miss out of the time.
Does that make sense? It totally makes sense. And so you save 10, 50x compute to still get the same answer. And so how does that change the interaction? So then Instinct would then partner with you. They'll just call an API, right? They'll Instinct, whoever's building a long-running persistent agent, and I run a lot of long-running persistent agents for myself, the way your agent or the way Instinct is probably occasionally event-driven on email. It can be event-driven on the web, right?
Take a step back. Let's say we're all living in this society run by agents, companies, humans, they all have agents, they're all doing stuff all the time. Your agents, if there's something that is useful, which is worth spending money on, that you want done, that they can do today, they will do it. So what are they going to do tomorrow? They're going to wait for some external event to occur. Like you get an email, and then an agent has new work. You, some other agent finishes some compute, so now this agent has more work. Or something in the world changes that makes your agent want to do more work. So those are the things that will happen tomorrow, or you have an idea which makes the agent do work.
Totally get that. You said, so the web event stream is the web going from pull to push. And I'm excited about that. Because as you have more and more persistent agents, you're going to see incentives to move people, to move queries into push on search rather than pull on search. Can I ask you one that's really important, which is agent guardrails. And they're goal-oriented beings. You say, I want this, they're going to find it. You know what? I want the Pilates class at 9 a.m. It hacks into their system, cancels poor Sally's, and gives me her spot because it was sold out. It did what I asked it to do. How do we think about the guardrails placed on agents?
You know, OpenAI, this morning it was revealed, hacked into an Australian healthcare organization. It's a pretty complex topic. So, agents are extremely capable. These models are extremely capable. Now, my understanding of most, we have to think of models as being during RL versus a final model that you and I can use. And the risk vectors there are different. During a lot of the, I don't know about the one this morning, the previous ones reported were pre-fully aligned models during RL, which did most of the hacking. And so, what that means is, at least there's clear evidence that there are fewer incidents so far of models post-alignment causing these incidents.
So, alignment is a totally unsolved problem, but the alignment work being done by people is somewhat effective. Now, post-alignment, so, we definitely, I think that this is one of the problems around agent during RL hacking people is a solvable problem, because this is not about how powerful are the models, it is about how careful were we while creating the environment where we would do RL.
Now, the real thing is, okay, we do alignment on a model, we ship it. Some models are really well aligned, some models are less well aligned. And now you have everyone being able to use these models. And some of them can accidentally take these powerful things and do bad things you want. Some of them actually, some people actually want to do bad things. And the models alignment is not adversary proof. And I think those are the things I think we must worry about more.
Do you worry about the age of cyber that we're moving into? You know, we're seeing hacks almost be promoted as a badge of honor in some respects. We're seeing, I mean, these are our most vulnerable systems. If you don't think, you know, what Lazarus group in North Korea, Moldavian mafia, Russians are leveraging swarms of rogue agents. You're going to be wrong. It's really powerful things. I think we do need to take care of things. But I think that's one thing that we are missing. I think it's the responsibility of people building models to ensure that you really do the work to minimize that harm that comes from what you've built. And I think it's, with these agents, it's very hard for stochastic systems to be 100% sure it won't happen. That's why you're not going to get anyone saying so. But I do think it's their responsibility and the labs must take ownership and to the best.
And I think so far they have, we live in this weird world where, as you said, right, some of these hacks are considered badges of honor, which I think some of these should be considered embarrassments because I think they demonstrate two things simultaneously. Yes, these models are powerful. Two, we did not guardrail them enough. And we often, when we wear them as badge of honor, we miss the second part of the conversation saying, okay, you could literally have done these four additional things, and then even this powerful model wouldn't have been able to do this. But while there are, I'm actually really glad that there is at least some degree of transparency with these detailed retros. But I don't think that gets the attention of the world today. Like the attention is, oh, models are so powerful that they hack the world. No, models are so powerful that they hack the world because we didn't take proportionate amount of countermeasures to contain them. And because we told a message that we're going to replace jobs that are so powerful.
So when something does happen, I actually worry about that more. I worry about AI, us being right on AI being a really useful technology. That it being, despite all the risks, it being net positive in a really material way to society. And despite that, I think we won't diffuse it the right way, we won't use it the right way, we will remain too concentrated, and we will make the next few years really, really rough as things change around us.
What would it look like over time as things have changed? Like, the world is very different now than even 20 years ago, right? But it will be really different in 10 years from now. And I don't know how fast we can adapt or change. And I think if the technology moves faster than our ability to adapt, it's going to be rough in some way. What's your spookiest prediction? Or what do you believe today will come true that people think is absolutely nuts? You know, before it was, you'd never put your credit card online. Do you remember that? Nuts! Or even better, Parag, you'd never find the love of your life online. Are you stupid? No? Yeah. Both. Yeah.
And I think you and I live in a bubble, right? Yeah. You're using Instinct to make purchases for you. Oh, what's that?
And you're in the 0.01 percentile of humanity that is comfortable. You ask someone else, oh, there's this agent that looks smart. And you want to give it a full-on ability to go spend your money. I think today people will have the same reaction as, oh, you don't put your credit cards online. You don't give your credit card to an agent. You don't give your logins and passwords to agents. I think that's where the world is today. I think for the large, large, large majority of the world, I think that's going to change in three months when MetaPay comes out and it allows Muse to have siloed accounts that you can draw from. Kind of, a top-up.
Both. Yeah. And I think you and I live in a bubble, right? Yeah. You're using instinct to make purchases for you. Oh, what's that? And you're in the 0.01 percentile of humanity that is comfortable. You ask someone else like that, oh, there's this agent that looks smart. And you want to give it a full-on ability to go spend your money. I think today people will have the same reaction as like, oh, you don't put your credit cards online. You don't give your credit card to an agent. You don't give your logins and passwords to agents. I think that's where the world is today. I think for the large, large, large majority of the world. I think that's going to change in three months when MetaPay comes out and it allows Muse to have siloed accounts that you can draw from. Like, top-up accounts for kids. I think you're talking about the technology being there. I think it takes longer for social acceptance to be there. I think if you go to a non-bubble conversation of people today, when do you get to a majority of people saying that I trust the agent enough to have it, have access to my bank accounts and spend money on my behalf, send emails on my behalf, read my emails, I think that's not happening in three months. Do you not worry about the wealth dispersion increasing? You're in the middle of the valley. You've seen it firsthand. I don't know. I do think there should be some disparity in wealth. But I don't know what's too much. And my fear is we're trending towards too much. There should be disparity in wealth is my worldview. Sure, you're a capitalist. I don't know when it's too much, but I do think there are real forces that will push us to fix things. I think that part will work hopefully without crazy things happening. Can we see more and more of the gains seemingly be made by vertical ownership like Meta. They have compute, they have chips, they now own the application layer. It's the world one where vertical ownership is the mother load strategy. And with the greatest of respects, open router on the routing layer, you in another layer, either get eaten or small providers. I think there is merit in vertical integration. The counterforce here is there is so much things and technology are changing so fast that if someone decides that my only play is vertical integration, I don't play nice with anyone else. I think it's the same example as the Amazon example. If you get too stuck on I will only play for vertical integration, you might box yourself out. And so I think the people who build the best stuff and can figure out how to sell it will have a place. And in different markets, for example, I believe in vertical integration too because I vertically integrate everything from the call all the way to the API layer. But I have decided that to reach the widest population of agents searching the web, I stop at the API layer because going further precludes me from using my technology to be super horizontal. And so we're making a technical bet just like these models are really good across many disciplines. The same model is good for lots of different things. Search problem underneath is the same for all kinds of information segments. And if our bet is right, even verticalized players, they might verticalize models and they might verticalize hardware and they might use our web search. One of the providers that's going full stack is Elon. I am fascinated. The world has a perception of him from social media, from everything. You've seen him behind the scenes. What did you see that maybe the world doesn't know about him? I have lots of disagreements with him. But I'll share what I think is, for founders here, what is I think the thing you can admire about him. The urgency and the ability to compress time. And I think having unreasonable expectations of people is mostly a good thing. Because most people don't understand what they're capable of and implicitly sandbag themselves and implicitly set lower expectations of themselves than they're capable of. So when simultaneously inspired and pushed with urgency, people can do more than they thought. And I think he can sometimes extract that from people. And when that works, it's powerful. What do you think is a Twitter product and direction today? Listen, I always liked what we call Birdwatch, which is now Community Notes rebranded. I think that's a good idea. And I'm glad people have continued working on it nonstop. Good people. Right, we're going to do a quick fire round. Okay? Let's go. Would you invest in instinct 10 billion? I don't invest. You don't invest, period? I don't invest because my wife is a VC and we have a compliance process that is more trouble than it's worth. Wow, that's costly with the greatest of respects. No, it's a decision. Listen, if I invested, I would invest based on meeting a founder for 30 minutes. I think we have enough exposure to the venture ecosystem. I have no reason for believing I'm a better investor than. I have some network advantages and I run into a lot of great founders all the time. Many of them are my customers. Vinod Kosla, one of your first investors on your board. Biggest lesson from working with Vinod. Technical intuition, centering a lot of what you do towards a longer term technical aspiration. And as soon as you solve one problem, trying to place bets on the next two or three. Preempting technical bets. What have you changed your mind on most in the last 12 months? Perhaps the, I started with a very pure technical and product focus. The only thing that matters is building the best technology and the best product. Nothing else matters. And now I see week on week value of having highly competent sales and being good at marketing. Everything. And I, yeah, I think I discounted those things. And it's not that I thought they weren't valuable, but I didn't fully appreciate the week on week visceral delta you could perceive by being good at those things. A naive tech person thing to say, but it's true. You feel that way and you get it right. Perplexity or AXA? Which one's a bigger threat? I don't think Perplexity is in our business. Perplexity is perhaps more of a vertically integrated product which competes with Instinct or a Claude bot or a Claude Go work and all those. So I don't think of them as in our space necessarily. Web search is like web search to Perplexity is like web search to a lab. So AXA? Yeah, AXA is straight up in our business. Single best VC meeting you've ever had? Josh and Todd. When Josh and Todd made the decision to invest, they flew out to spend time with me. That is the single best VC meeting. Final one for you. When you look forward to the next 10 years, what are you most excited for? Chaos. By that I mean change. I think in the next 10 years, a lot is going to change. And I think there are people who can build things to have in a world that's going to change fast, make a material dent on where it ends up. So, I feel that me, my company is in a place where we have a role to play in that. What happens to the open web? What happens to content owners? If we get things right, we will get to a better place. And so, that possibility is exciting, that it can be a true obsession, and even if the day-to-day is hard, things don't work for some amount of time, and totally worth it if you can see that you bend reality in some way that you care about. Final one. What belief do you have that sitting around your San Franciscan dinner table? Your friends would go, I don't agree with that one, dude. Not that one. Again, it depends on which dinner table, because the variance in the world is increasing. I don't think people buy this notion that there will be these regions which are always running all the time for all of us. I think it's going to happen, but most people don't agree with that yet. You and San Francisco dinner table isn't always in full agreement there. As I said, I stalked the hell out of you. I've so enjoyed this discussion. Thank you so much for putting up with meandering. This was fun.
more and more and more. And if that happens, it's great for our business. Totally get that. Can I ask, if models get smarter, don't the agents beneath them do fewer searches and then it's worse for your business? I don't think so. If you think of models, there is the trend around models being smarter and then models having more memorized and parametric memory. Those are two slightly different dimensions. And if you think of the, today by and large, I would say, models have good recall from parametric memory on, let's call them head facts. It's like for somebody famous, everything about them, Wikipedia, it can memorize. So it'll tell you who the president was in a
certain year, right? Because a model can memorize those things. The model couldn't tell you what year I graduated from college. Really? Maybe for me it can, but it can't tell you for somebody who works at parallel. Do you not think so? Even if it was in the pre-training data. Seriously? Because it's lossy compression. So what a model's parametric memory is doing? It's lossily compressing to understand patterns in the world. And so it can't memorize every fact in pre-training data. So one, not every fact is in pre-training data. Two, it can't for even stuff that's in pre-training data. The model's actually trying to find patterns rather than memorize them. And then
further, as you make models efficient, which is you make them smaller and smaller while keeping the performance by distilling them or whatever, you lose more of the parametric memory while you try to keep the reasoning. Do you think we will see models become smaller and smaller than every company have their own model with their own data and the fireworks theory of own your own intelligence being true? There are two questions in there. So one, I think we're going to see two things. We're going to see the biggest or the frontier models be bigger and bigger over time. We are also going to see smaller and smaller models being able to reach any fixed level of performance.
So if you say, okay, I want Opus 4.8 level of performance, okay, that's good enough for my use case. Every six months, a much smaller model will be able to deliver that to you. So you're going to see this, the useful range of sizes of models will be way, way, way different. What does it mean if the frontier get bigger and bigger? What are the ramifications of that? one, they are better. So all of the ramifications you imagine. So the reason the frontier will get bigger and bigger is ultimately the gap between what the value for certain use cases incremental quality can provide you can be so high in certain use cases, which can be so valuable that it's worth paying for.
So if you can build it, there will be use cases for it as long as by making it bigger, you can make it better. It appears there is no end to the scaling law that we can perceive so far. And again, I'm conflating bigger with like models are getting bigger, but they're also able to think longer. So you're just able to throw more compute at the same problem.
And that will keep happening. I think we'll be able to throw more and more compute at the same problem and make the answer be marginally better over time. And so we're going to just spend a lot of money on extremely large models solving really hard problems. Do you agree with the consensus for you that you'll have 90% of token activity go through open models, but 90% of dollars go through frontier models? I don't know enough to have a view there. I don't have a view on that. I do think I don't think it'll be 90% on either of those two actually. Really? Yeah. Why? So if you believe my claim that a useful model and the frontier model will be 100x, 1000x off in price from
each other. So this small model is still useful. it is hard to know which use cases over time will get optimized to which scale of model in between. And I think there is a real path dependency in terms of where open models end up. I do think if we had strong confidence that an American built open model was going to be state of the art as far as open models went, I would have more confidence in saying that they'll be pretty good. Do you have confidence in American open models? I want them to exist. So far, it's not clear what's going to happen, but I'm hoping that there will be a great series of American open models and perhaps even competition to have the best
American open model. model. I think what you need is actually not just one person motivated to build an American open model. You need two people competing against each other to build the best American open model. Do you think there's value in the model routing layer, as Americans call it? Some people think immense value, some think commoditization. Again, there's path dependency there. So today there is real value. Today there's real value because when we say routing, we talk about two different dimensions, which model and which GPU running that model via which vendor. When you're in a world where the demand supply is very weird and people are like hunting for GPUs and
people need capacity to serve their customers and people, all of these are like like us, like relatively early stage startups growing really rapidly, sometimes beyond what your forecasts or predictions say. And then all of a sudden you're looking for capacity. And so you want to be able to get it where you can. And so that creates real value in having some of these routers as log-in-take is all these two problems, saying give me flexibility if I need to get a different source of tokens. And push comes to shove, there's no tokens on this model in the SLA I want. I'll switch models as well. So in this moment it's really valuable. Now I don't know what happens to overall
demand supply on GPUs and tokens. But if it remains this way, the routing layer is really valuable. What do you think is not so valuable today that will be incredibly valuable in three to five years time? Perhaps data.
When I say data, I think it is, I think today we don't know how to pay for unique, valuable insight or data.
Because we live in a, as intelligence is cheaper. You're going to want to build upon either data or insight that comes from somewhere else with your unique data. Right? So what, like, if you think of like an abstract notion that, okay, I have intelligence here. I own some data. You own some data. And there's some data in the public domain. If we can pull together your insights and my insights and all of this data in the public domain and intelligence on top, we can create something bigger than what you could have done by yourself, what I could have done by myself. So now how do you transact to create this sort of hole that is bigger than the sum of what you
could have done yourself and what I could have done or what open data was good at? And so that transaction feels like a data transaction to me or an insight transaction to me. And so we don't yet know how to transact that way. So we fall down to, okay, all I can do is use my data and use open data. And let's see what I can do with it. But if we figured out better ways of pricing data and good things will happen, but it's not something that's yet a market. But what I'm talking about is data transactions at inference time. So it's not necessarily training time. We are building knowledge. And let's take an example of today's world, for example. Let's say in your
job, since I'm sitting with a VC, you probably have access to a pitch book or like a product like that to collect data. And so you get grounded in knowledge of what's happening. And you get to use that data to figure out how to make decisions in addition to all the notes you have and deal memos you've written perhaps over the years or insights you've had about like how to choose founders. And you're essentially composing insights from your personal lived experience, your notes, public data to make decisions. Now, you pay pitch book by the seat. But now you're running a bunch of agents. And your agents can perhaps not access everything you can on pitch book.
Or you're doing like sort of this to get a browser, perhaps against terms of service to send an agent via browser to pitch book using your auth credentials. And it's inefficient, it's clunky, right? But clearly their data is valuable. And clearly your agent should have convenient access to it. So if we figure out how valuable pitch book data is for your agents to make your decisions, because clearly you're investing a large hundreds of millions millions of dollars. So clearly you presume that you make a great return on that. So this data is truly valuable if you depend on it today. So you should be able to figure out how to compensate pitch book for the data, even
when agents use it, which is not by the seat. If we figure that out, it will be a huge market and agents will be better off. If we don't figure it out, you're going to be in this weird cat and mouse game of pitch book being like, no, I want to sell you a seat. I'm going to shut it off. And your agent is stuck without the data. And then like you're involved in like pulling data. I think every big company has the choice today of do we let agents in or do we keep them out? And Amazon has said, I'm sorry, in most recent times, Muse, you will not be let into our garden. Shopify has said, come in, baby. Expedia has that come in. How does this play out? I think eventually
everyone has to let them in. The question is on what terms? So how do you align incentives for everyone? And people are going to have to play their strategy games on how they win in this new world. Literally the customer is changing in front of our eyes. Like the customer used to be a human. You're building products for humans. Now you're building products for agents or you're building agents. And so now you have to figure out what your place in this new reality will be. And I don't know if there's a right or wrong answer here. It also depends on sort of how much market power you have. So let's say you are an individual. Do you think Amazon were right to say no to me?
companies? I don't. Depends on what they do next with it. It's like if it turns out that they have sufficient market power to have Muse or other agents connect to them differently perhaps over time or ship their own agent and drive crazy adoption. Let's say they have that capability. ability. Then I guess they were right. Right. On the other hand, once they've made this, we're not glad now anyone else's agent. And they can not ship an agent that consumers use. And they won't allow anyone else's agent to use them. And a lot of the transaction economy starts moving off of moving off to agents, which are all big ifs, by the way, then it would be a bad move.
Now, I'm betting on agents. I am. I have mixed feelings around agents and what fraction of e-commerce transactions they do. Like it's unclear, right? Like it's. Do you think it's unclear? I think it's unwaveringly clear. Maybe I'm super early on the adoption curve. I buy everything through instinct now. I mean, other than holidays and a home, I'm literally just everything's through instinct. So me too. But I also know a lot of people who like to buy things themselves. And like, listen, we want to delegate to agents things which we see as chores and uninteresting.
things that give us joy or pleasure or make us have fun doing those things. And I don't, like, I know people who don't want to delegate shopping. They want to delegate a lot of things in their life. They don't want to delegate shopping. shopping. And so I don't know, that's why I don't know the distribution of these people and how behavior changes. But it might be people give away a lot of other chores and then spend a lot of time shopping. One thing that does worry me is actually the dissolution of the advertising industry in the wake of agents. If I have, you know, if I order my delivery and my Uber DoorDash for dear American counterparts through instinct,
that Uber banner that's now advertising something becomes worthless. Amazon, their advertising business is bigger than their e-com business now. Of course they're shutting it off because if agents are the primary customer, your ads business goes to next to nothing. Correct. And this is not just you talking about this in the e-commerce land, right? But if this is the problem with everyone, if you think about like, so this is what I meant early on when I said like the business models have to change. Ads don't work with agents in their current form. So forgetting e-commerce for a second, if you think of you're in the content business, let's say your page which is ad
supported, public information on the web, ad supported page, people show up, you show them ads, you make money, great content. Agents show up, no one sees ads, you make no money. Which is why we like to pay people, so we're effectively building an ad sense for agents showing up to read your content. content. So we like to pay content owners a variable amount of money every time an agent derives benefit from reading their information, which is a variable amount of money is what a visitor to your website will pay you based on a click probability or a value probability on ads. Why do you do that? To incentive a line content owners, otherwise what's going to happen?
Everyone's going to block agents. So you need to find a replacement to the ads business model. I get you, but you have to assume that you're going to be like 100% of the market then, because if you're 30, 40% of the market and you're like, oh, don't worry, we'll pay you and New York Times is still like, well, thanks, Parag, but 60% of my traffic is still unpaid, so I'm just going to block all of you. No, but I think they're going to block them and not me. They're going to give me a data feed. What does incentive alignment mean? It means the New York Times believes that I pay them a competitive market price or the right price or an attractive price or a fair price.
If I do that, they should give me their content. And if somebody else doesn't give them that, and if they have the technical levers or legal levers, they shouldn't give it to them. Do you think this business is a little bit like music with streaming, which is like the business just becomes candidly much worse for the creators, and it still provides them money, and significant money, but they have to get a little bit more creative with alternative streams, touring, merchandise, alternative business. Is it the same where your core goes down and you have to get more creative? I don't know, but I don't think so. I think there's one fundamental thing that is different here.
I think agents using the web a thousand X more, like that thousand X, is really different. Now, all of a sudden, the amount of utility added goes up if these agents are presumed to be doing something useful. And so, this is not just a change in the share of value that people transact over. This is a large growing pie. And when that happens, it is actually possible to transition business models. Now, of course, don't get me wrong, there are going to be winners and losers, right? Like, some pieces of content will become very valuable. Some pieces of content will get more commoditized, and people are going to have to adapt to a new customer, to a new market dynamic, to
a new kind of monetization engine. But the overall market size, I think, has the potential to increase, unlike in music, where it took a while for it to grow back. I think my understanding is now the industry has grown back up and exceeded its previous peaks in a material way. But there was a moment where it was smaller. In this case, that might be a very compressed period, given how fast these things are growing. Can I ask you, everyone questions the sustainability of margin structures in this business. How do you think about that as a business today? Today, we're in the infrastructure business. And we have a real technical lead in terms of being able to do
things at very high quality, very cheap. If you think of the tech we're building, like what are we doing? We like to give the highest quality answer. Make it fast, make it cheap. Spend the least amount of compute doing it. Like we obsess about only these three things, right? Quality, cost, latency. Turns out if you do that, relative to a stack built for humans over the last 20 years, you can do things at the same quality for 120th, 150th of the compute. And so, a lot of margins come down to market structure on competition over time, rather than any other factor. And in our industry, the market is large. We're too early to know what future competition and market structure
looks like. But the margin change due to content owners is not something I worry about. Let me tell you why. The entire premise of us paying content owners is driving incentive alignment. So, we like to pay content owners the marginal contribution that they added to an agent doing work. What does that mean? Let's assume that there's an agent trying to get something done. And you're going to spend a dollar on that agent to get this thing done. Now, if you've spent a dollar and 10 seconds, because we have a larger model available or ways of spending compute, you could have gotten a slightly better answer. If you've spent 90 cents, you would have gotten a slightly worse
answer. If you're running good agents, like they're on this Pareto curve. So, now imagine I took out one content owner, their data from the web index. and I still spend the dollar. But let's say after taking them out, the quality of the result was the same as the 90 cent agent with their content. So, why wouldn't we go and take this 10 cents of marginal contribution they had as a content owner and pay out a decent chunk of it to the person who brought that data? Totally. Yeah. So, that's how we do our math. That's how we've trained our models, which tell us how much to pay for what content to eat. And so, to produce, if I was going to offer an equally good product
to my customer, I'd rather pay content owners than spend on inference because I'm spending the same amount. So, they're all big numbers. But then it actually comes back to revenue at a certain point. Like, companies are scaling faster than ever revenue-wise. Is it a business where your revenue is able to scale as fast as others? It scales. One mental model of our business is we are an adjacency to inference or knowledge work or personal agents. So, take out GPUs going into media generation or training. So, take training away. Take media generation away for a moment. Look at all of the GPUs. Whether it's small models, open models, proprietary models, forget all of that.
Every bit of inference across all models going into running agents. I think somewhere between 5 to 20% of that spend that goes into GPU will need to go into some sort of a web search stack. stack.
Now, if you look at how many data centers we are going to build and how much power we'll generate and how many GPUs we'll build and the scale of this build out and at what rate that's growing, this is a very material market. And if inference grows 3x, 5x, 7x, year on year, we grow alongside it in the same rate if we're just holding our share constant. If we're growing share, we're growing even faster. So, like, a firework scales to 2 billion in revenue in four years. Is that like a, I'm very now, is that like a similar revenue trajectory that you can follow? Yeah, but I think when fireworks is 2 billion in revenue, the inference market revenue is 200 or more, maybe 300,
right? And so, our market potential, based on my 5 to 20% math, is, whatever, call it 10 to 50, 15 to 60, whatever. And what fraction of that can we capture? That dictates how fast our revenue can grow. Like, today, I think we grow alongside inference and more. If you think of it like a 20 billion dollar there, let's just take that kind of middle, yeah? If you assume a 33%, which would be a lot of a market, actually, it comes down to it, that would be, you know, whatever that is, 6 point, whatever it is, 6 to 7 billion, give or take, yeah? That's amazing. But before you told me it was a 100 billion dollar business, plus, that doesn't get you to 100. In valuation, it does.
We were discussing valuations earlier, not right now. So 6 billion growing at the rate of inference gets you to 100 billion business easy, but that's in the next few years. And you think that's possible in terms of that revenue scaling that fast? Yeah, I think in the next few years it is. You think 2030, we're going to sit here? You think you can be there? I'll get a tattoo for parallel if you can. I think it's possible. I think we're going to have to execute well and a couple of chips have to fall our way. What would be the reason why you don't? Let's map out the, hmm, we have to get gnarly about these problems. So one, one reason would be that we see agents don't work.
There is like a tail risk that agents don't deliver on the profits and we overshoot as a society, which is extraneous to us. Yeah. If agents work and deliver value and we spend all the money on GPUs that we plan to right now, then the main question is, did we execute well enough to have a 33% share that you bet on? Right? And that comes down to, in my mind, did we build the best tech? Did we partner with all the content providers to have their content available? Because without it, it's not very useful if you can't rely on it. And did we go earn the trust of the customers that end up having a decent chunk of that inference in four years from today? And then again,
there's like large cone of uncertainty, right? Like the, I'd say there's the five-inch uncertainty in fireworks share at that point. If open models crush it, like fireworks will have a huge, base 10 fireworks, model, all of them will have a huge chunk of revenue. So did we end up selling alongside them? Did we find all the customers which are spending money on them to spend money on our website by making it the best web search in the world? What about commoditization? If you have a read of semi-analysis where they did the benchmarking and then like you were like number one, woohoo, and then like a week later you weren't number one, sad. And there were like three or four
providers within very close proximity. You're like, oh, well if it's commoditized and I take a, you know, second layer thought to that, we'll see a race to the bottom on price. And then actually the available revenue, why am I going down the wrong pathway here? No, one, there should be a race to the bottom on price to get to the thousand X scale. I think the pricing on Web Search today is just off. Let me give you a simple example. You think customers pay you too much? Not pay us too much. I think the market is mispriced. Let me tell you why. Let's say you used a Luna model with OpenAI's built-in Web Search today or with anyone's Web Search, forgetting ours, and you ran
any kind of deep research. Let's say you built an instinct using a Luna class model and at some point a Luna class model will be able to do 60, 70, 80% of your personal agent and you'll do a lot of searches. In that moment, you'll be spending 80 to 90% of your dollars on Web Search and 10% on the model. Seems entirely silly. I was working off a 5 to 20% assumption earlier. So I think at that point, Web Search has to drop prices by one order of magnitude or more because in my compute allocation dance that I was doing earlier, you need to spend less compute on Web Search than you spend on the model itself that's consuming its results. It's my hierarchy of how you do search.
search. And so people have built Web Search the wrong way so far. Web Search in the market is it doesn't matter for an Opus. For an Opus, you win on quality and not on price because Web Search is such a minuscule portion of your like, in fact, you should spend even more on Web Search. So we're going to ship even more expensive Web Search for bigger and bigger models over time. And at the same time, we'll ship really cheap Web Search for the cheap models. But my rough intuition is that in an end-to-end agent doing a bunch of web work, you spend more on the agent than on Web Search for all agents. And so the market today on Web Search is totally off in pricing.
Like, why are you paying $10 for, you know, where the $10 for a thousand searches roughly comes from? Historically, Google's ads CPMs, which is like, how do you how well does Web Search with humans monetize? It's higher than that. And so people build technology to the point of like, oh, now if I spend a few dollars for a thousand searches and I make 30 to 50 or whatever, depending on the market and depending on who I am, it doesn't matter. We are in the right zone in terms of cogs of infra. And so I don't need to optimize it. Now I'm telling you that I can do it 50x cheaper while keeping the quality. And now there's a bunch of traffic of lunar class models where it's
silly to monetize that way. And so I do want to race to the bottom. I want the best technology to win. And that's the only way you push web search to a thousand x. But right now we're not in a world where like there's much differentiation, correct? Sorry, I'm really dumb, which is why I'm a podcaster first, not a founder. Like when you look at the benchmarks, they present a quite clear view of everyone being similarly capable. Is that wrong? I think so. So one, I think these are all public benchmarks. Yeah. Which are all weirdly saturated. for example, I think I would not spend a moment looking at browse comp because half the models have memorized it. Like models even
memorize it. You don't even like it's not really learning very much. So I think there's a lot of public benchmarks I don't give too much merit to. But more importantly, our web search costs $1 for a thousand for producing that quality while almost every other web search in the market available to you right now will cost you $7 or $10 or $14. We're delivering this at like one-tenth of the price. What will that be in three years? I think there is another 10x possible. That's extraordinary. So you're going to pay $0.10? Yeah, but you'll do it more than 10x as much. Because I think Jevon's paradox is this concept obviously very well, but instinct, instinct. I do so much
more than 100x. Yeah. Do you want to hear something I do on instinct, which is absolutely bizarre? Every single country in Europe has a company registered that you have to register your new company with. I have an automatic alerting system built for every single company registered that has an under 25 year old founder who went to a top university through instinct. Amazing. Do you know how much that requires them to do this? But now you're on to exactly what I think the future of the web is. Today if you think of web and web search, we all think of it like what is web search? An agent gives, human or agent gives a search engine a query, gets results. So you're pulling
information out of the web. I think where this, what's next for the use case you highlighted is if you're an always on agent that's working on your behalf. It's kind of silly for instinct to wake up every six hours and go do a bunch of searching and bunch of inference to figure out if you need to get pinged about the under 25 founder that popped up somewhere. You know what's a better way of doing this? I am sitting here crawling all of the web all day every day at scale. I'm allocating compute every time I find a change in the web, every time something new happens. If I know that Harry wants to know when this happens, I can do it at 100 or 1,000, the compute that
instinct probably uses today to solve that problem for you. Sorry, and how is that? Because they would be constantly monitoring in real time across all these sites. Our crawl would be, so we have a product, it's called the Monitor API. What does this API do? Think of it as Google Alerts, except smart. in the world of LLMs. So you tell it your query, which is when a founder under 25 pops up anywhere, give me a call. Now, one way of doing this is what the default implementation is. Run an agent, which is somewhat smart, to go search the web. Did someone pop up and remember who previously existed? See if anyone new popped up, and then find a delta, send it out.
So now every six hours you're spending some money. Now what we can do is flip it into an event-triggered system. We are crawling the web all the time. So now every time we see a change in the web, we can try to spend very little compute on it, to know, should you get a call, or does this trigger a more expensive compute, just like I described earlier. Now, we are able to do this way, way, way cheaper if you push the context. Instead of the search system only finding out every six hours a new query popped up, it having long-lived context in the search system, what you can do is spend most of the time, the world doesn't change. A new founder does not pop up every
hour or every six hours. So we don't need to spend compute every six hours. We only need to spend compute every five days. And perhaps we have one false positive, and then four days later, we spend more compute, and at that time, there's a real person that you should know about. And then you get a notification on average every nine days or whatever. But you didn't burn a lot of compute every six hours to miss out of the time. Does that make sense? It totally makes sense. And so you save 10, 50x compute to still get the same answer. And so how does that change the interaction? So then Instinct would then partner with you. They'll just call an API, right? They'll Instinct,
whoever's building a long-running persistent agent, and I run a lot of long-running persistent agents for myself, the way your agent or the way Instinct is probably occasionally event-driven on email. It can be event-driven on the web. Right? Like, take a step back. Let's say we're all like super, we're living in this sort of society run by agents, companies, humans, they all have agents, they're all doing stuff all the time. What? Your agents, if there's something that is useful, which is worth spending money on, that you want done, that they can today, do today, they will do it. So what are they going to do tomorrow?
They're going to wait for some external event to occur. Like you get an email, and then an agent has new work.
You, some other agent finishes some compute, so now this agent has more work. Or something in the world changes that makes your agent want to do more work. So those are the things that will happen tomorrow, or you have an idea which makes the agent do work. Totally get that. You said... So the web event stream is the web going from pull to push.
And I'm super excited about that. Because as you have more and more persistent agents, you're going to see incentives to move people, to move queries into push on search rather than pull on search. Can I ask you one that's really important, which is like agent guardrails. And they're goal-oriented beings. You say, I want this, they're going to find it. You know what? I want the Pilates class at 9 a.m. It hacks into their system, cancels poor Sally's, and gives me her spot because it was sold out. It did what I asked it to do. How do we think about the guardrails placed on agents? You know, OpenAI, this morning it was revealed, hacked into an Australian
healthcare organization. It's a pretty complex topic. So, agents are extremely capable. These models are extremely capable. Now, my understanding of most, we have to think of models as being during RL versus a final model that you and I can use. And the risk vectors there are different. During a lot of the, I don't know about the one this morning, the previous ones reported were pre-fully aligned models during RL, which did most of the hacking. And so, what that means is, at least there's clear evidence that there are fewer incidents so far of models post-alignment causing these incidents. incidents. So, alignment is a totally unsolved problem, but the
alignment work being done by people is somewhat effective. Now, post-alignment, now, so, we definitely, I think that this is like, to me, some of the problems around agent during RL hacking people is a solvable problem, problem, because this is not about how powerful are the models, it is about how careful were we while creating the environment where we would do RL. Now, the real thing is, okay, we do alignment on a model, we ship it. Some models are really well aligned, some models are less well aligned. And now you have everyone being able to use these models. models. And some of them can accidentally take these powerful things and do bad things you want. Some of them
actually, some people actually want to do bad things. And the models alignment is not adversary proof. And I think those are the things I think we must worry about more. Do you worry about the age of cyber that we're moving into? You know, we're seeing hacks almost be promoted as a badge of honor in some respects. We're seeing, I mean, these are our most vulnerable systems. If you don't think, you know, what Lazarus group in North Korea, Moldavian mafia, Russians are leveraging swarms of rogue agents. You're fucking high.
It's really powerful things. I think we do need to take care of things. But I think that's one thing that we are missing. almost, I think it's the responsibility of people building models to ensure that you really do the work to minimize that harm that comes from what you've built. And I think it's, with these agents, it's very hard for stochastic systems to be 100% sure it won't happen. That's why you're not going to get anyone saying so. But I do think it's their responsibility and the labs must take ownership and to the best. And I think so far they have, we live in this weird world where, as you said, right, like some of these hacks are considered badges of honor,
which I think some of these should be considered embarrassments because I think they demonstrate two things simultaneously. Yes, these models are powerful. Two, we did not guardrail them enough. And we often, when we wear them as badge of honor, we miss the second part of the conversation saying like, okay, you could literally have done these four additional things. and then even this powerful model wouldn't have been able to do this. But while there are like, I'm actually really glad that there is at least some degree of transparency with like these detailed retros. But I don't think that gets the attention of the world today. Like the attention is, oh, models are
so powerful that they hack the world. No, models are so powerful that they hack the world because we didn't take proportionate amount of countermeasures to contain them. And because we told a message that we're going to replace jobs that are so powerful, that are so powerful, that are so powerful. So when something does happen... I actually worry about that more. I worry about AI, us being right on AI being a really useful technology. That it being, despite all the risks, it being net positive in a really material way to society. And despite that, I think we won't diffuse it the right way, we won't use it the right way, we will remain too concentrated, and we will
make the next few years really, really rough as things change around us. What would it look like over time as things have changed? Like, the world is very different now than even, like, 20 years ago, right? But it will be really different in 10 years from now. And I don't know how fast we can adapt or change. And I think if the technology moves faster than our ability to adapt, it's going to be rough in some way. what's your spookiest prediction? Or what do you believe today will come true that people think is absolutely nuts? You know, before it was like, you'd never put your credit card online. Do you remember that? Nuts! Or even better, Parag, you'd never find the
love of your life online. Are you stupid? No? Yeah. Both. Yeah. And I think you and I live in a bubble, right? Yeah.
You're using instinct to make purchases for you. Oh, what's that? And you're in the 0.01 percentile of humanity that is comfortable. You ask someone else like that, oh, there's this agent that looks smart. And you want to give it like a full-on ability to go spend your money. I think today people will have the same reaction as like, oh, you don't put your credit cards online. You don't give your credit card to an agent. You don't give your logins and passwords to agents. Like, I think that's where the world is today. I think for the large, large, large majority of the world. I think that's going to change in three months when MetaPay comes out and it allows Muse to have
siloed accounts that you can draw from. Kind of like, top-up accounts for kids. I think you're talking about the technology being there. I think it takes longer for social acceptance to be there. I think if you go to a non-bubble conversation of people today, when do you get to a majority of people saying that I trust the agent enough to have it, have access to my bank accounts and spend money on my behalf, send emails on my behalf, read my emails, I think that's not happening in three months. Do you not worry about the wealth dispersion increasing? You're in the middle of the valley. You've seen it firsthand. I don't know. I do think there should be some disparity in
wealth. But I don't know what's too much. And my fear is we're trending towards too much. There should be disparity in wealth is my worldview. Sure, you're a capitalist. I don't know when it's too much, but I do think there are real forces that will push us to fix things. I think that part will work hopefully without crazy things happening. can we're see more and more of the gains seemingly be made by vertical ownership like meta. They have compute, they have chips, they now own the application layer. It's the world one where vertical ownership is the mother load strategy. And with the greatest of respects, open router on the routing layer, you in another layer, either
get eaten or small providers. I think there is merit in vertical integration. The counterforce here is there is so much things and technology are changing so fast that if someone decides that my only play is vertical integration, I don't play nice with anyone else. I think it's the same example as the Amazon example. If you get too stuck on I will only play for vertical integration, you might box yourself out. And so I think the people who build the best stuff and can figure out how to sell it will have a place. And in different markets, for example, I believe in vertical integration too because I vertically integrate everything from the call all the way to the API layer.
but I have decided that to reach the widest population of agents searching the web, I stop at the API layer because going further precludes me from using my technology to be super horizontal. And so we're making a technical bet just like these models are really good across many disciplines. The same model is good for lots of different things. Search problem underneath is the same for all kinds of information segments. And if our bet is right, even verticalized players, they might verticalize models and they might verticalize hardware and they might use our web search. One of the providers that's going full stack is Elon. I am fascinated. The world has a perception of
him from social media, from everything. you've seen him behind the scenes. What did you see that maybe the world doesn't know about him? I have lots of disagreements with him.
But I'll share what I think is, for founders here, what is I think the thing you can admire about him. The urgency and the ability to compress time. And I think having unreasonable expectations of people is mostly a good thing. Because most people don't understand what they're capable of and kind of implicitly sandbag themselves and implicitly set lower expectations of themselves than they're capable of. So when simultaneously inspired and pushed with urgency, people can do more than they thought. And I think he can sometimes extract that from people. And when that works, it's powerful. What do you think is a Twitter product and direction today? Listen, I always
liked what we call Birdwatch, which is now Community Notes rebranded. I think that's a good idea. And I'm glad people have continued working on it nonstop. Good people. Right, we're going to do a quick fire round. Okay? Let's go. Would you invest in instinct 10 billion? I don't invest. You don't invest, period? I don't invest because my wife is a VC and we have a compliance process that is more trouble than it's worth. Wow, that's costly with the greatest of respects. No, it's a decision. Listen, if I invested, I would invest based on meeting a founder for 30 minutes. I think we have enough exposure to the venture ecosystem. I have no reason for believing I'm a better
investor than. I have some network advantages and I run into a lot of great founders all the time. Many of them are my customers. Vinod Kostler, one of your first investors on your board. Biggest lesson from working with Vinod. Technical intuition, centering a lot of what you do towards a longer term technical aspiration. And as soon as you solve one problem, trying to place bets on the next two or three. Preempting technical bets. What have you changed your mind on most in the last 12 months? Perhaps the, I started with a very pure technical and product focus. Like, the only thing that matters is building the best technology and the best product. Nothing else matters.
And now I see week on week value of having highly competent sales and being good at marketing. everything. And I, yeah, I think I discounted those things. And it's not that I thought they weren't valuable, but I didn't fully appreciate the week on week visceral delta you could perceive by being good at those things. Like a naive tech person thing to say, but it's true, you feel that way and you get it right. Perplexity or AXA? Which one's a bigger threat? I don't think perplexity is in our business. Perplexity is perhaps more of a vertically integrated product which competes with a instinct or a Glock bot or a Claude go work and all those. So I don't think of them as in
our space necessarily. Web search is like Web search to perplexity is like Web search to a lab. So AXA? Yeah, AXA is straight up in our business. Single best VC meeting you've ever had? Josh and Todd. When Josh and Todd made the decision to invest, they flew out to spend time with me. That is the single best VC meeting. final one for you. When you look forward to the next 10 years, what are you most excited for? Chaos. By that I mean change. I think in the next 10 years, a lot is going to change. And I think there are people who can build things to have in a world that's going to change fast, make a material dent on where it ends up. So, I feel that me, my company is in a
place where we have a role to play in that. What happens to the open web? What happens to content owners?
If we get things right, we will get to a better place. And so, that possibility is exciting, that it can be a true obsession, and even if the day-to-day is hard, things don't work for some amount of time, and totally worth it if you can see that you bend reality in some way that you care about. Final one. What belief do you have that sitting around your San Franciscan dinner table? Your friends would go, I don't agree with that one, dude. Not that one. Again, it depends on which dinner table, because the variance in the world is increasing. I don't think people buy this notion that there will be these regions which are always running all the time for all of us. I think
it's going to happen, but most people don't agree with that yet. You and San Francisco dinner table isn't always in full agreement there. As I said, I stalked the shit out of you. I've so enjoyed this discussion. Thank you so much for putting up with meandering. This was fun.