It's the biggest strategic mistake I made in the history of Speechify. How do you reflect on that? So today is a real discussion. Cliff Weitzman, founder and CEO at Speechify, one of the fastest growing text-to-speech startups in the world on the show. The best way to lose is not to be in the race. Be in the race. You don't want to be a fat manager who's a general sitting in the back saying, take that hill. You want to be the warrior who runs up with your sword and engages the enemy first. Ready to go?
Cliff, it is so good to have you back in the studio, dude. I was looking forward to this one because when I was writing it up, it's a very different thread of conversation to how I'd normally go. And so thank you so much for joining me again today, dude. My pleasure. Glad to be here as always. Now, I wanted to start with you're spending tens of millions of dollars on NVIDIA GPUs and you're paying an additional $100,000 per GPU to receive them four months early. Why? What do you know that the market doesn't know? So in 2022, we bought a huge rack of GPUs from NVIDIA. And the reason we bought them is for training. We have a bunch of models.
The newest Speechify Simba 3.2 model is ranked number one in the world for quality. Above all the frontier labs, 10x more affordable than 11 Labs. We used to rent GPUs. And we found that engineers at Speechify would be parsimonious with how they use the GPUs because they were worried about costing the company tens of thousands of dollars. And the analogy my brother and I came up with is imagine you're Michael Jordan and you want to be in the NBA. It's the only thing you care about. And you need to pay $20 an hour just to train in a basketball center. Well, that sucks. You want one that you can go to whenever you want to. In fact, you want a hoop in your house.
And so our initial idea was we want a hoop in our house. And so we bought a bunch of our own GPUs. And that deal ended up being really good for us. And we ended up training really good models. So with time, we invested more and more. So that's the first part. The second part is how the economics work out. If you look at it, the Transformer was invented inside Google in 2017. NVIDIA came out with A100 GPUs in 2019. Shortly after, they came out with H100 GPUs. The original ChatGPT was trained on A100s. And then they came out with Blackwells. B200s, B300s. And now they came out with Rubens, which is the GPUs that Elon is sending to space.
They're liquid cooled. They're very cool. And we're thinking, okay. One, every class of GPU is more affordable per 1 trillion flops. A flop is addition, subtraction, multiplication, any mathematical operation. And you measure them in how many trillion operations happen per second in a GPU. And they're more affordable as it relates to this. If I was to buy an H100 for $30,000, that's how much a single card would cost. If I wanted to rent an H100 for one hour spot instance from GCP, it could cost me $5. If I rented it from Azure or AWS, maybe it would cost me $3.5 per hour. If I multiply that times 24 hours, and then times 365 days in a year,
I'm actually going to end up paying $35,000 to $50,000 to rent that GPU for one year, but I could buy it for $30,000. So it's 1.5X the cost of owning the hardware to rent the hardware for a year. Now, the hardware is typically warranted to work properly for three years, but it'll keep working beyond the warranty for, I imagine, 10 years. So the math shows it makes way more sense to buy them. The other big part is if you want to do large scale training like we do, you need the memory to be co-located with a large cluster of GPUs. I can't just rent from Google or Microsoft or Base 10 and run the size of training that I want,
because I need a gigantic memory card next to it with all of my data that all the GPUs are accessing. So that's why we first started buying them. The next thing that we found is if you run open source models for coding, you could pay Anthropic, and then you're paying for all the tokens and the fact that you're doing the branded thing. Or you can run an open source model, and instead of running it on a spot instance from Azure, you run it on your own hardware, and then you're paying a fraction of a cent per token. And so for all those reasons, it made a ton of sense. I just want to dig in. The first thought that I have is I completely understand the rationale there,
but chips depreciate. You have chip cycles, and they are accelerating. We are seeing newer and newer chips being created. We're seeing specialization within chips. By locking yourself in to one chip architecture, how do you think about that? At Speechify, we still use K80s for a lot of specific operations for inference. And we use older models of GPUs constantly. And there's essentially a difference between inference and training. For training, I have this hypothesis. I want to know the answer to this hypothesis as soon as possible. Every minute that it doesn't come out, I'm in competition with everybody else.
And so having a GPU architecture that is much faster by orders of magnitude is a huge advantage. But if you do speech-to-text or text-to-speech with Speechify, I can afford to give you a lower quality GPU, and it'll give you what you need in 100 milliseconds. So it's totally good. And so I can always use these older GPU models for inference. That's number one. Number two, we have so many experiments running at every single point in time. Not all of them need to run on the newest hardware. So the analogy I always give: let's say you bought an iPhone in 2011, an iPhone 3G. And then you bought another iPhone and another iPhone and another iPhone.
You could have a drawer in your house with five iPhones collecting dust, because you can only use one iPhone at a time. But if I own 100,000 GPUs, I'm still going to use all of them at the same time. And so I'm not losing anything by having more GPUs. Not only do I own a bunch, I still rent from the hyperscalers all the time. And I rent both dedicated instances that I prepaid for, and spot instances. For example, more people use Speechify in September, because everybody goes back to school. So I need to level out the load. And so the parts of that load that I know for sure I'm always going to use, whether training or inference, I might as well own it.
And then on top of that, I have so many other friends who are running training and running inference. I can always rent it out to other people if I have excess capacity, which I don't expect to have. But occasionally, you have an interesting situation. So for all those reasons, it just makes mathematical financial sense. Lastly, if you have excess capital, you either stick it in the bank, or you buy a bond. The best bond you can buy yields you 5%, or you can buy a GPU. And because renting would cost me 1.5X buying it for the year, the return is way higher. So how many GPUs do you buy then? So let's talk about Rubens, for example. Rubens come in 72 cards in one rack.
So we'll buy multiple racks of Rubens. And on top of that, we'll buy B300s, the newest form of Blackwells, because we can get them earlier. And when we bought our first instances of DGX H100 GPUs, we bought racks of those. And then those get delivered by truck to the data center. We rent the data center space. And so the data center provides the networking capability. It provides the energy, which is actually the largest constraint now. And it provides physical engineers that take it off the truck. They install it. If it has an issue, they fix it. And then it just runs. Does 11 Labs do this? 11 Labs is amazing at this. So let's talk about Rubens, for example.
Rubens come in the form of 72 cards in one rack. So we'll buy multiple racks of Rubens. And on top of that, we'll buy B300s, which are the newest form of Blackwells, because we can get them earlier. And then the same thing, when we bought our first instances of DGX, H100 GPUs, we bought just a bunch of racks of those. And then those get delivered in a truck to the data center. We rent the data center space. And so the data center provides the networking capability. It provides the energy, which is actually the largest constraint now. And it provides physical engineers that take it off the truck. They install it. If it has an issue, they fix it. And then it just runs.
Does 11 Labs do this? 11 Labs is amazing at this. 11 Labs, I think Piotr had 11 Labs, literally bought a bunch of GPUs early on and set them up in his house. And then they just kept building bigger and bigger and bigger clusters. They do the same thing that we do. How do you think about forecasting chip buying? It's incredibly difficult to know A, demand, but also B, supply of chips. How do you think about forecasting chip purchasing? Yeah. So number one, I want to explain again, it's very different than buying an iPhone or buying a MacBook. I can only use one MacBook at a time, one iPhone at a time, but I can use all the chips I have at any given point in time.
And I still will have more demand, especially when I have multiple teammates and 60 million users who are using inference on my Speechify software that's providing text-to-speech and helping them read their work and dictate their work and use Speechify work, which is our newest product that's a GenTech, similar to Jarvis from Iron Man. And so I go, okay, let's imagine I have 100% capacity that is the average usage per month that I need for GPUs for training of my AI models and for inference on my AI models. Inference is when you actually make a call to Speechify and you give me text and I give you back audio. There's math that happens in the background. That's inference.
Training is I take a gigantic amount of data. I take all the architecture and software engineering that we're doing. And I go, I think that this will give me a better model. I'm baking that model in the oven and I'm going to come out with a new black box. And then when I give you text, that black box is what calculates it and gives you back the audio. So those are the two usages. Let's say I have 100%, which is what I would have in, a month like November. In October, I'll have 140% because it's a big month for us. In December, everyone's at home. They're not necessarily studying or working. So I might have 80% utilization. Okay. So I go, cool.
Well, I can take 20% of the usage that is normal and buy it because it's the best deal. I'll take another 25% of the usage and I do long-term contracts with hyperscalers. The rest, I'll rent what's called spot instance from the hyperscalers. And then I'm still not even close to overcommitting myself. And so that's how we think about the math. And then we go, okay, well, also we have 45 engineers, but we want the team to be 150 engineers. And even inside of my 45 engineering person team, there's a couple of people who are rock stars. They have dedicated DGX racks just for that one person. And 25% of my team are almost waiting. And I want to double the size of the team.
They just need it's like you have a football team and you just need another field because they don't have enough field to practice on. And so that's how I think about how to allocate. And then in terms of depreciation of the asset over time, I go, okay, well, these are still amazing GPUs. Even A100s, you can run amazing experiments on. So it's completely valid to use that as long as it's hooked up and as long as it's not stopping to work. And so think about the mileage of a car, right? If a car gets to 250 miles, you know it's going to break at this point. That's not necessarily true for a GPU because it doesn't have as much wear and tear.
Yes, it's moving and yes, all these things, but it's in a very clean environment. It's very much cooled. It has constant maintenance because it's not moving around. It's very expensive. And NVIDIA just does a really good job. And so that asset is going to stay for a very long time. And let's say it got so outdated that I no longer can run training on it. Cool. Now I'll use it for inference. There's one more thing that's very interesting that just happened. So I believe earlier this month, NVIDIA did a huge deal with Blackstone, BlackRock, Apollo, and Goldman Sachs. And they said, listen, we want more people to buy more GPUs. We're going to underwrite for you
up to 25% the value of a GPU that if you lend money to someone who buys a GPU, let's say Google or a startup, Corweave, and that startup goes out of business and you have that GPU as collateral against that investment, we'll buy back the GPU for up to 25% of the value of the GPU. And so they're succeeding in creating a liquid secondary market for GPUs that they're underwriting. So now the large banks have an incentive to loan money at much better interest rates. This is actually exactly what Elon did in the beginning of SolarCity. He went to Morgan Stanley and Merrill Lynch and got them to amortize the price of a solar panel over 30 years.
So the whole invention behind SolarCity was the fact that you could take a loan against the collateral of your solar panel. So NVIDIA has done an amazing job now in creating a clear floor for the value of the GPU over time. Do you think the circular economy fears that people often cast against NVIDIA are justified or not? We saw their CFO push back on them and say, enough of this. Do you think that justified or not? I think that a lot of the things that about a year ago, like what was going on between Oracle and OpenAI, that was way too much. That was ridiculous. I think the NVIDIA stuff is not because you're talking about a real asset. So if you think, for example,
about the logic behind the value of Bitcoin, right, Bitcoin, what is the intrinsic value of Bitcoin? I can't really tell you, right? What's the intrinsic value of gold? Well, gold, you can use it for some medical stuff because it's a really amazing metal and it's jewelry, whatever. But a GPU, it has intrinsic value. You can actually use that asset for something that's really, really valuable. And it doesn't matter where that GPU is. It could be in Iceland. It's still useful to anybody all over the world as long as it's networked. And so actually it has a pretty good store of value. Even if new GPUs come online, really my one question,
and this is the math for everybody to come back to, is how many teraflops per second can this device do? And that is essentially a token. That's the value. And so there is intrinsic value. So yes, you can have all these circular things, but at the end of the day, NVIDIA is making a product that's real. It's not complete tulip mania. There's a real, real intrinsic value here. What does no one know about buying chips that they should know? What's the, oh my God, people are so naive about this. I mean, it's not that people are naive. It's just they haven't been in the space. So I'll give you an example. Imagine you're buying a GPU. Well, you're going to buy it.
You know, it's an NVIDIA produced product, but NVIDIA is not going to waste the time talking to Cliff Weitzman. So who do I buy it from? Well, one of the best rated vendors is Dell. So everybody thinks Dell is a personal computer company. No, Dell is a GPU rack supplier at this point. And then, okay, I want to buy it from Dell. Well, Dell has a constraint because there's not a lot of Blackwells out there. Well, it happens to be that they have some in France. All right, well, I'm going to order mine from France. Okay. Shoot, it was supposed to come a month ago and it's still not here, right? Because of demand. So then you need to negotiate to make sure that you get it,
which is why we're very willing to pay a hundred K per month extra to get them earlier. So you'll call up Pierre in France and say, you know, hey, we'll give you an extra hundred K kicker So who do I buy it from? Well, one of the best rated vendors is Dell. So everybody thinks Dell is a personal computer company. No, Dell is a GPU rack supplier at this point. And then, okay, I want to buy it from Dell. Well, Dell has a constraint because there's not a lot of Blackwells out there. Well, it happens to be that they have some in France. All right, well, I'm going to order mine from France. Okay.
Shoot, it was supposed to come a month ago and it's still not here, right? Because of whatever, there's demand. So then you need to negotiate to make sure that you get it, which is why we're very willing to pay a hundred K per month extra to get them earlier. So you'll call up Pierre in France and say, hey, we'll give you an extra hundred K kicker if you get them here in a month? Even more than that. So in that France situation, which is something that happened to me, I was like, Pierre, what the heck? We have a contract. You're not delivering on time.
And so it is the case that we had a contract with another company beforehand. And they were a few weeks late. And I called them and I was like, listen, I've got a better deal. I'm canceling our contract because you're out. You didn't deliver. So I'm going to go with this other contract. But if you have a better price, we'll go with you. But I just need the GPU now. And remember, I'm paying for the renting space of my data center. So the most expensive part of a delivery of a GPU is if it's late, I'm still paying rent for that data center space.
Now that GPU, and so you put pressure on Pierre to send you the thing when he said he was going to send it to you. And then you go to NVIDIA or Dell or whatever. And you're like, oh, it's a market. Hey, can I pay more to get it earlier? Skip the queue. Yeah, you can. Cool. Now there's a truck somewhere in the United States with a GPU whose value is the value of a house that's coming to my data center, right? Well, I should have insurance from that, because if that truck gets hit or there's too much humidity or the GPU gets flipped, I lost multiple houses worth of GPUs.
So, okay, the insurance is really important. And then also the value of the amortization is really important. And there's all these nuances of how to do the math through. Then there's the cooling, right? So you're not only paying for the physical space and the networking and the energy, and the energy is the biggest constraint. We'll talk about it in a second. Well, how do you cool that thing? Because you have a thing that's moving and moving, moving, moving, moving. And the thing that's most new now is liquid cooling, because air is just not enough. And the thermal load of water is much better. And there's other liquids that are even better than water.
And so Rubens are liquid cooled. But most of these data centers don't have liquid cooling installations already approved. So we had to do a bunch of research and we find, okay, we could buy what's called a side cart of liquid cooling that you enter into the data center. Then you pay someone at the data center to install it for you. Cool. Now you can have the rack that you want. And so there's a big difference between running a purely software company and running a company that includes hardware.
But when I listen to all of this, I'm now more sure than ever that it is a mistake to price optimize and to spend the money to buy it versus to rent it. Because I get you on the optimization, but you're not saving ten times more. It's 0.5x more per year. No, no. Per year. Exactly. Yeah, but you have the flexibility to tailor it up and down. You don't have any of the logistical nightmares of insurance, transportation, security, water cooling, logistics. And then you can build your product, actually what matters most, against Eleven Labs who are running fast. I don't want to worry about water cooling and insurance for a freight truck.
Eleven Labs worries about the same thing because for them to train excellent models, they need to have co-located GPUs with a lot of memory available. You can't do it if you rent it. You can. It just becomes, one, ridiculously expensive. Two, you need to commit for many, many years ahead of time because you need to build a co-located cluster. And then you don't have as much control because you don't own it. So it's difficult to...
Like you suddenly need an InfiniBand cable which allows for the memory to flow from one DGX to the other one. And the answer then is, well, now my ability to train is so much bigger. I can have a larger AI team. Every person in the AI team is leveraged and I could just, I could shoot ahead of everybody so much faster. And let me just make one thing clear. If I want a Ruben, which is these much faster GPUs, I'll get it faster if I buy it than if I'd wait for Google to buy it and then there's other people in front of me in line. So I'm going to skip the queue by a lot. And then I'm going to have a year of access to Rubens before everybody else does.
The way I think about it is the following. How do you build an amazing company in a world where there's so much competition today? The team is the most important part, but the team is the most important part because the team gets you the other resources. And so what are the missing pieces? The missing pieces are data, compute, and architecture. In a world where intelligence is commodified and no one needs to hand write code at all anymore, right? Our engineers, really what I'm looking for is ten really good decisions per day, which is very tiring, not optimizing the random parts of the code.
And each one has five to eighteen agents running at any point in time, doing long horizon tasks on these GPUs, coming up with theses, testing them, going back and forth, back and forth, back and forth. If they don't have the capacity to train, the team is limited. If they don't have the data to train, the team is limited. And by the way, a lot of data, it's cleaning the data, right? You get this raw data in the beginning. Well, you need to organize it in the data sets. And so you need the GPUs to also organize data sets too. Like I have one of my best engineers right now who is not even writing models. He's making synthetic data sets to train models.
And so it really becomes an indispensable asset. Would you ever buy data? We have, but small data sets. So I suggest you use Fireworks. Okay. But I mean, Fireworks is amazing. Lynn and the founder is one of the co-founders of PyTorch. But I had the very obvious realization that you'd have every company having their own specialized models of a certain size, trained on their own data. But you would need supplemental data, like this synthetic data, or real world data that you just don't have. And that you would buy that from data providers like Mercor, which is why I was—
Cronon, Surge, all these companies are amazing. And they shorten the cycle, by the way, to getting to revenue. A hundred percent. Because if you are a Mercor, shout out to Brendan Foody, Eleven Labs or OpenAI or Anthropic is going to make money for the next decade or two on the data that they bought from you. And so they're willing to pay a fraction of that ten years of revenue to you today to supply them the data. And so they're willing to pay a fraction of that ten years of revenue to you today to supply them the data. And again, it's all a speed thing.
Yes, Eleven Labs or OpenAI can go and make a team that will get the data, but they don't want to manage it. And so the data is key. All this, you need all three things. You need compute, you need data, and you need team that writes great products. And ideally, you need users that use you a lot, and a lot of them, to have a feedback loop of whether this stuff is good or not. So benchmarking. And so for me, when I was doing the as a venture investor, we do outcome scenario planning, which is the most bullshit exercise to pretend you're smart predicting the future. We do it because it makes us feel important.
But they've predominantly sold to Frontier Labs today. And that's where ninety percent of their revenue is from. With the rise of specialized models on a per company basis with their own data, I believe that you move that customer base from purely Frontier Labs to every large scale enterprise who needs supplemental data. If that is the case, how big an outcome is the data marketplace? So the first problem to understand about the data marketplace is it's not ARR, right? It's not annual recurring revenue. And ideally, you need users that use you a lot, and a lot of them, to have a feedback loop of whether this stuff is good or not. So benchmarking.
And so for me, when I was doing the, as a venture investor, we do outcome scenario planning, which is the most bullshit exercise to pretend like you're smart predicting the future. We do it because it makes us feel important. But they've predominantly sell to Frontier Labs today. And that's where 90% of their revenue is from. With the rise of specialized models on a per company basis with their own data, I believe that you move that customer base from purely Frontier Labs to every large scale enterprise who needs supplemental data. If that is the case, how big an outcome is the data marketplace?
So the first problem to understand about the data marketplace is it's not ARR, right? It's not annual recurring revenue. It's one-time deals every single time. So the buyer of the data is not required to buy it from you again. So it's a very risky business. And if you look at early days of companies like Mercore, they didn't raise significant funding off the bat because investors were very skittish about that fact.
Let's put that aside. Well, it's very important that the T's are crossed and I's are dotted about how you got that data, right? So you need to indemnify the companies who are using you and that's part of why they buy it from you as opposed to sourcing it themselves. We've seen the lawsuits. But it's a great business. And if you could do it well, but you need to be an ops monster. You need to be really, really good at operations. You need to be very fast. And really the key is the company training on your data needs to actually see improvements in their model at the end of the day.
The thing that has always been challenging for Speechify compared to other companies is B2C customers pay a lot less than B2B customers. So 11 Labs, huge credit to them, leapfrogged us because they sell to B2B. Historically, we've only sold to B2C. And so our big constraint is we need to do this on a cost basis of less than $10 per million characters. 11 charges $100 per million characters. The OpenAI model on the benchmarks cost $196 per million characters. So ours, when we sell it to other B2B companies now, we just launched our API, Simba 3.2, it costs $10 per million characters.
Dude, I am too old to not ask the painful questions. And I think the joy is when you ask them. You said 11 Labs kind of leapfrogged you. Is that on you for not doing B2B? Yeah, 100% is on me. 100% is on me. It's the biggest strategic mistake I made in the history of Speechify. How do you reflect on that?
So I met Piotrek and Mati. I was living in London at the time in my house in London. I think it was 2022. And we were very impressed by them. And we wanted to use the model, by the way. It was just too expensive for us to use. And I looked at it and my thought to myself was, they're very smart. They're going to do well, but I don't like their strategy because I think that an API for text-to-speech feature will become commoditized with time, right? You're going to get to the point that you can run that API on your computer and then on your phone and then what are they selling anymore? So I don't want to go into that business. And I made a critical error. What I didn't understand is that the point of an AI lab like Speechify or 11 Labs is to continuously innovate. And the first product that you release is your wedge that gets other people to then later use your other technology.
So for example, if you're in text-to-speech, you built the best text-to-speech model in the world for one specific voice. Cool. Well, now you can do other voices. Now you can add emotional prosody. Now you can add voice cloning. Now you can add speech-to-text. Now you build duplex models where it makes the ah, laughter, interruption handling, turn-taking. You add a harness for voice conversations. Then you optimize it for sales and you optimize it for customer support and you optimize it for all these things.
And so what they did is they first built an amazing API. They were great at launches. They built a really great product for creators. Then they built their best product ever, which was agents. Agents is amazing because the buyer is no longer a software engineer. The buyer is a CTO, CIO, CEO, executive in the company. Sierra has this concept called outcome-based pricing. Brett Taylor is amazing. And so you can start fighting on the outcome. And having an AI agent is like having an AI co-worker. But it was my mistake to think that an API product was a bad strategy because I thought it was something that would become commoditizable. And I forgot the central thesis about Silicon Valley, which is constantly innovate. Get the user to start using your product. I don't care if it's free. Then you sell them other things. And so that was my big mistake.
How possible do you think it is? I think people underestimate the complexity of building out a B2B GTM. Super hard. I think it's a strategic mistake for Speechify to go to B2B. Hmm. A lot of people think that. Tell me your position. You are now competing against Eleven Labs and Sierra, really. And those two are competing. Whether they like to admit it or not, they absolutely are competing. And they will. I'm sure if you ask them off camera. That's Brett Taylor.
Yeah. You don't want to compete against Brett Taylor. I don't want to compete against Brett Taylor. And that is the tidal wave of Eleven Labs. Now, Eleven Labs is an unstoppable machine at this point, to the point where it has government buy-in across all of the large, major Western democracies. Actually, it's insane the government buy-in they have. And they started three months ago. I just said the key thing. They started three months ago.
Yeah. And so you know the graph. But I think they've reached a tipping point where actually they've just taken the market. I think Sierra is running behind them chasing. And they're doing a decent job of it, but they've got Brett and they've got Sequoia and Green Oaks and every royalty of Silicon Valley behind them. And they're still running behind chasing Eleven Labs with Sequoia kind of pretending to be neutral because they're in both of them, which is incredibly challenging. And I just think being third, the Postmates effect is never a good market to be in when I could be the dominant consumer brand that leads with a really different and compelling story.
So here's the two things to consider. The first one is if you go to the App Store and you search text-to-speech, Speechify has 98% of the installs in text-to-speech for B2C. Speechify has served more than 770 billion words to users over the last few years, which in terms of times of listening, it's like 6,000 years of listening, right? If you go from today to zero BC and back, you still have thousands of years left. So we've completely dominated that market and it's still a business that's growing really, really fast. And we're constantly adding more features into that product.
The thing is, we have a pretty big engineering team and now everybody is capable of doing 10x what they did before. So I have extra staff. I have a huge AI engineering team with ability to make amazing models. So where is the highest ROI for that to go? Well, it needs to go both B2C, but it should also go B2B. And one thing that I will never be is a person who doesn't learn. So I might as well just learn B2B.
Now, to your point about competing against giants like Sierra or Eleven Labs, hey, Anthropic came into the market as a second to OpenAI and they were second for a very long time and now they're not second. Facebook came as a second to Friendster and MySpace and now they're not second. And so the nice part is this space is not a monopolistic space, it's an oligopolistic space. And if you look at what happened with Eleven Labs, I'm going to exclude Sierra because Brett Taylor effect is huge. It's amazing to see how good of a business that is. And so it might very well be that for the core offering that they're currently winning on, I will not win. But what did I learn last time? It's fine if I offer my product essentially for free because I'm an AI research lab. And as long as people start to use me with time, I'll be embedded in the system and I'll keep coming out with more and more and more innovations that are useful to them.
Anthropic came into the market as a second to OpenAI and they were second for a very long time and now they're not second. Facebook came as a second to Friendster and MySpace and now they're not second. And so the nice part is this space is not a monopolistic space, it's an oligopical space. And if you look at what happened with Eleven Labs, I'm going to exclude Sierra because Brett Taylor effect is huge. It's just amazing to see how good of a business that is. And so it might very well be that for the core offering that they're currently winning on, I will not win. But what did I learn last time? It's fine if I offer my product essentially for free because I'm an AI research lab. And as long as people start to use me with time, I'll be embedded in the system and I'll keep coming out with more and more and more innovations that are useful to them. And so there's an unbelievable demand from all these companies and governments and everybody else for great tools, whether they be AI agents or APIs or products. I just want to be on your phone if you're a user or in your stack if you're a company and supply you with the best front-deploying engineer experience and AI orchestration experience and API experience to give you an amazing experience and there's room for everybody.
I agree there's room for everybody. I think value accrues to top one player. I agree. I think it's the inference market where Fireworks will be a multi-hundred billion-dollar company and then I think a genuine base time will be a hundred billion-dollar company and then together and a load of the others will be 15, which is amazing. It is completely true. It's hugely amazing, valuable companies. But you would then think that OpenAI would be the place where value accrues for voice AI, right? That's what you would have thought three years ago and that's not what ended up happening. So you can't not go into the race because there's a big incumbent.
Well, I think with all candor that's because of incredibly poor management. I agree. And hiring. And that was theirs to take and they fumbled the bag across every spectrum. And for every company in the world, no matter how exceptional the leadership team is, niches get fumbled, right? So voice AI was a niche for OpenAI, right? LLMs are the core. And by the way, they also fumbled AI coding. Now they're trying to cash because it's such a big space. All respect to Piotr and Mati. I think they're absolutely amazing and I love working adjacently to them. I just don't think they're going to fumble the bag. That's my trouble. But they have so much in their net right now.
That's true. And so much is getting added to the net constantly. That's true. And so you just have to go where the football is going. Yeah. And so I think that it's too expensive for Speechify not to be playing in B2B as well as playing in B2C. The best way to lose is not to be in the race. Be in the race. In terms of products that we build, we were chatting earlier and you said that every startup stage has to be a compound startup. Can you talk to me about that and how you think about that? It's not that every startup has to be a compound startup. It's at a certain point you can't afford not to be that.
Do you not think there are a few companies that are just absolutely running rings around everyone else?
Yeah, absolutely. Those are the winners, right? Eleven Labs is an example. Anthropic is an example. Ramp is an example. Speechify is an example. All the companies that have absolutely maniacal leadership teams and engineering teams. That's why people care about team more than almost anything else. Because the right team will iterate fast, get there, and then figure it out. And now, when everything can be turned into a reinforcement learning problem where you can have long horizon agents and orchestrating agents thinking about the problem for like two weeks at a time, if you set that up, of course you're going to win.
I got into a lot of trouble as I always do with most of my social posts. I used to be quite a sweet little boy, actually. No, really. I used to be the Harry Potter of Ange Capital and now I'm more like... Yeah, you lost the glasses. Lost the glasses and it kind of became more like Piers Morgan, if you know Piers Morgan in the UK. Highly despised figure. Very opinionated.
But a question that I have is, I said, if you're a startup, it's never been harder to hire great talent because OpenAI and Anthropic, candidly, have such a carrot reward mechanism in front of you, especially with impending IPOs, that the best talent just wants to go there and talent follows talent. And you're seeing the founder of Monzo, a multi-billion dollar bank in the UK go there from YC as a partner. Matt Clifford, the founder of EF, which is a multi-billion dollar company. I mean, he should be prime minister and he's going to join Anthropic. Am I wrong that this is the hardest time of startups to hire because the prizes of Anthropic and OpenAI are so great?
My favorite type of person to hire is a CTO of another company. We have, well, we were 21 people at Speechify. 18 of the folks at the company were previously either CEO, CTO, or VP of engineering at the last company. Anthropic, I have never seen a company like this, hires so many CTOs of publicly traded companies and other successful startups. Workday, one of them. The reason is they build the best, most beloved product for engineers in the history of the world. So it's easy to hire CTOs. By the way, they hire many more CTOs than CEOs because CTOs are the ones who get the most excited about this product. And you're right. They're the fastest growing company ever, especially at the scale that they are. So they're going to keep growing. OpenAI is going to keep growing. You had this very condensed period, like fireworks of growth in both of those companies.
Yeah, it's very hard to hire. But remember, they're hiring people that their annual compensation needs to be $15 million a year minimum. What startup is hiring someone and paying them $15 million a year? You're not. Your seed founder, that was not something that you were going to hire. And so I will push back against it. The competition for growth stage companies hiring exceptional leadership talent is more difficult. For seed companies, I would say it's the easiest time ever because the impact of even just the founder on their own is bigger because they can orchestrate agents. And so one thing that we have changed about our hiring in the last even six months is we really cared that you read a ton of textbooks about software engineering and that your handcrafted code was amazing. I still care that you read a lot of textbooks about software engineering and you understand it. But the thing I care about the most today is technical aptitude and raw technical intelligence because I know that we could teach you everything else and in six months you could be a machine. And so we hire a lot of math olympiads and LeetCode coders and Kaggle award winners and people who studied physics and math. They might have not coded before because I just need the hunger and the work ethic and the intelligence and anyone can become so good so fast now. And so the pool for hiring exceptional talent is bigger than ever before and Duolingo did this really well. They love hiring college grads and coaching them. And so I wouldn't say that it's harder to hire than ever before for seed companies. Seed companies now almost anyone can be someone that you hire if they're smart and hardworking because you can teach them very fast. What is more challenging to hire is for growth companies because you're fighting with just absolute juggernauts.
Are you not a growth company? So it's challenging for us. Why do you think it's hard to hire a really good salesperson? I totally get that and I completely agree. I will see CRO packages in the 50 million plus range by the way. Yeah, exactly. 15 is kids play. By the way, with the greatest of respects I will even see $15 million on the table for comp packages for seed companies today. That is the dislocation that I think with the greatest of respects... Wait, wait, wait. Sorry, sorry. But this is a seed company that has raised how much money at what evaluation?
I mean, you've got to understand that seed round today will be $150, $200 million. And there are several of them. I mean, there's 30, 40 companies that at seed have raised $100 to $300 million. And this is a company of a guy who's one year out of university? No, no, no. This is a guy who's probably spent four years at OpenAI or spent four years at TMI. So then what about the company that's, you know, the guy who's been in university... By the way. Yeah, exactly. 15 is kids play. By the way, with the greatest of respects I will even see $15 million on the table for comp packages for seed companies today. That is the dislocation that I think with the greatest of respects.
Wait, wait, wait. Sorry, sorry. But this is a seed company that has raised how much money at what valuation? You've got to understand that seed round today will be $150, $200 million. And there are several of them. There's 30, 40 companies that at seed have raised $100 to $300 million. And this is a company of a guy who's one year out of university? No, no, no. This is a guy who's probably spent four years at OpenAI or spent four years at TMI. So then what about the company that's, the guy who's been in university for two, three, four years and now they're starting a company? Or do you think that those people are out of the water now?
No, I think that's a very different world. And so, yeah, they'll raise $10 million seed rounds. Yeah. So for the company that you just described, they raised a seed round at $150 valuation and they raised, I don't know, $20 million? No, I said it was $150 million raise. Oh, I wouldn't call that a seed round. Maybe that's the name. But my point is, and that's my point though, which is the talent is concentrated. The people who really get AI and systems and have seen the magic inside OpenAI, Anthropic.
I agree with you that if you have a company that's raised $150 million and a $500 to $2 billion valuation, definitely that company should give $15 million comp package. And there's a lot of them. Yeah, that makes perfect sense. But there's a lot of them, there's 30. And those 30 take 30 people. And there is a thousand people now. That is hard.
And so, but what you just described is exactly what used to happen with Google and Meta, six years ago, which is if you were really cracked, there was essentially a maximum amount that you can get paid at a company like Google or Meta. And the best way for you to make a life-changing amount of money is to go to a company that is small and ride from the beginning all the way through and be a really solid founding engineer at that company. I think people want more certainty of cash today than upside, which sounds more.
No, I think that the equation is the same as always, which is each person has their own equation in their head of how much certainty and how much risk they're willing to take. It hasn't changed. It's the same. Humans are still humans. But I think people would rather know that the certainty of a $10 million from Anthropic versus a 60 from that quirky startup they could make. This is the reason why companies IPO, right? There's two reasons. Either you want a ton of money or you want a lot of credibility in B2B like Zoom did or you're hiring and the value of the package that you offer is so much better when your stock is liquid.
When we look at that dev team for you today, you said, hey, I wanted to go in. I want to see how we're orchestrating agents. What did you find? What did you learn in that discovery process around agent orchestration internally?
So inside of our AI research team, everybody's orchestrating agents. It's when you go lower, if you go then into the product-facing things that we build, for example, the platform team or the iOS team or the Mac team or the Chrome team or the web team or the Android team, these are super smart folks who have been working in those domains for 10 years and they know iOS like the back of their hand. They know Kotlin, JetBrains like the back of their hand and so it's very easy for them to hand code things because you're not dealing with something that's super, super new. So why change? People, it's hard to change, right? And so you just need to force them to change.
So one, the best thing is to inspire. So you do a Zoom screen share and you show them how the best engineer in the team is orchestrating agents and they're like, oh wow, I didn't know you could even do that. And then you go, yeah, please do it. You recommend blog posts for them to read, books for them to read, Twitter threads for them to read. What's the team using? Cloud Code, Cursor, Codex? Cursor and Cloud Code. Those are the two most popular. Yeah. It's a little bit of Codex usage. It's not that big. I would say Cloud Code is number one, then Cursor, then Codex. We want you to use as many tokens as possible in whatever harness way is the best for you.
You mentioned Linear. Linear is amazing. Automatically cutting tickets from Linear is fantastic. And just being able to go into your agents and be like, okay, I have these six Linear tickets, start on them. And then really a good engineer today is just an exceptional QA, right? The AI will make them feature. You will test the feature, see if it's good. You'll figure out where the edge cases are. You'll prompt it to fix it. And then you try to make it as efficient as possible, which is hard to do. And then you need to make essentially roughly 10 really good product and engineering architecture decisions a day.
How do you think about token allocation internally? I mean, we've seen leaderboards be used, which is I think the most fucked up form of incentive playing. You don't want to prevent people. There's a lot of people who are a lot of talk and I'll ask for examples and I'll read the examples and you look and I'm like, eh. And so I think about it in terms of demos. Can we hop on a Zoom call and you'll show me what you built and then I use it myself and I'm like, wow, that's amazing. Or you send me a screen recording of a feature or technology that you built and I'm like, wow, that's so good.
And so we give credit when things get shipped to production to users. So even inside of the AI team, if you build a really amazing, and this is part of why Speechify ended up winning, you asked, how did you build bigger labs? The answer is, we ship to production all the time. That's how we won. We are not in the theory space. We are an applied AI company. That's why we win.
And so if you're an engineer at Speechify, the analogy I always give people is imagine that you are in the milk delivery business and you make me a beautiful bottle of milk and you leave it down the road. The milk will spoil. You have to get it to my door. Knock. If you didn't do that, you get no credit. If you carry the football all the way to the line, but you don't cross over to the end zone, if you don't kick it into the goal, you get no credit. If you bring the ball just to the rim, you don't put it in the rim, you get no credit. And in the rim means push to production with no bugs and users are actually using it. And then we get feedback.
How many companies do that iteration cycle fast? Almost no one. Definitely not with a user-based number that Speechify has. And so in the AI team at Speechify, you make some amazing discovery. We're like, great. Push it to production. And then you go, oh wait, there's this QA problem and this QA problem. And if you have this many people use it on the AI serving layer, then you have this other issue. Cool. You get no credit from me. It's not in production. I can't use it on my phone. When I can use it on my phone, I will give you credit.
And so this morning, actually not yesterday, yesterday I had a call with our AI engineering team. And I said, listen, the project that we have running for duplex models and for AI conversational harnesses is something I'm really excited about. And it's been moving fast. I want it to move faster. Here's 14 notes that I want. And then what I do always is I'm on a Zoom call. I flip my computer around to face my phone and I use the product in front of them and we record it. And so then they see all the bugs. And then I send the recording in the chat.
Someone on our team, he's 19 years old, sent me a demo this morning off of that conversation that solved all of my problems. And he was like, hey, I was waiting for three training runs to finish. So I had a little bit of time while I was waiting. So I implemented everything that you asked and it blew my mind. It was so good. That's using AI correctly. So it's not a token leaderboard.
The project that we have running for duplex models and for AI conversational harnesses is something I'm really excited about. And it's been moving fast. I want it to move faster. Here's 14 notes that I want. And then what I do always is I'm on a Zoom call. I flip my computer around to face my phone and I use the product in front of them and we record it. And so then they see all the bugs. And then I send the recording in the chat. Someone on our team, he's 19 years old, sent me a demo this morning off of that conversation that solved all of my problems. And he was like, hey, I was waiting for three training runs to finish. So I had a little bit of time while I was waiting. So I implemented everything that you asked and it blew my mind. It was so good. That's using AI correctly. So it's not a token leaderboard. It's what did you show in production that was good?
How many companies do you think are actually as token-pilled AI-centric as we think in terms of devs? There's a guy, Jason Yeager, who used to work at Speechify and now he has MyTech CEO on Instagram. He's super funny. And so he makes a lot of videos about crazy CEOs who all are like, use tokens, use tokens. I think all founders in some way have that animal inside of them because you know that it's the right path. But there is a difference between reality and theory and you need to make sure that you don't overdo it.
Do you have any price sensitivity on tokens? Yeah, of course. Absolutely. I mean, I'll lose my mind if to implement a tiny feature you use 15,000 tokens. Why did you do that? And we will let people go if they just go bananas with something for no reason.
Are you able to accurately budget tokens? Not accurately, but within bounds. The other thing is a lot of engineers are, look, you go into engineering because you like optimization. Most engineers are not blind and it physically hurts them to overspend tokens. And I always think that the best way to interact with AI is you are chatting in the chat or actually doing it verbally and you're essentially pseudocoding with your words constantly and you're explaining architecture. And a great example would be I know someone who has no engineering background and they wanted to build an app and they built exactly what they want. It took them two hours. But they needed an API call and they needed to scrape this website and they basically scraped every single page of the website, every single part of the website. And so the bill that they got for the scraping was gigantic. And then I was like, why are you doing that? Why aren't you going into the database to this exact URL and then scraping that from the URL? So the amount of nodes they needed to hit became like 20 instead of 25,000. And so an engineer will spend their time making sure that the thing is optimized. So that's how you build a good database or a good architecture system. You do the same thing when you're interfacing with the agent. You want the agent to take the path of least resistance, not the path of most resistance.
I think one of the biggest problems is that agents are goal seeking. It's all about the target. You need to be good at picking the right target. And I think Anthropic published this paper when Claude 1 came out about long horizon tasks with Claude. So the first thing is it was much better at running a two-week task. And it could burn $12,500 worth of tokens in two weeks and basically make a better model with that. That's a perfect, amazing way of using tokens. That's exactly what you want. And what you don't want is burning 12,000 tokens in the span of five hours doing something that's totally unnecessary and doesn't make any sense. You need the loops to happen and then you need to check the result. So what you want to build, and Boris, who's the inventor of Cloud Code, talks about this all the time, it's all about the loop. You say, here is the target. Here's how you measure the target. Now iterate against the target over and over and over again until you get it.
What did you not know about building an AI-centric dev team that you wish you had known? How useful is it to own your own GPUs? What was that realization moment? Did you see a build one there? The realization moment was when we realized that we had really talented engineers who were essentially moving at one-seventh of the speed they could have if they had the compute one-to-one with their creativity and ideas.
If you're a founder listening to this, how should I change my hiring process in a new AI world? Number one, functional interviews. Build this and then you see if they can build the thing and then you run it through unit tests. The second one is give them a large code base, even an open source repository and have them understand the code base, make changes and then check what they broke. And then they have to be able to orchestrate agents well. And if they're not doing that, it's not worth having the person. And then the next thing I'll say is it is more fun to have a smaller team. Having a big team is great as long as everyone is carrying their weight. But the way I think about it is I can have multiple agents running on my computer or I can have several Slack chats with really smart people who are bigger domain experts than I am and basically that human being is the outcome owner for that task and they have the agents. And so I can run as a founder multiple projects at the same time to a really amazing level of granularity.
And so I think about moments earlier in the year when my brother Tyler would literally have an alarm to wake up at 3 in the morning because he needed to check what the agent was doing at 3 in the morning and then you wake up, make sure it's good, go back to sleep. You want to babysit your agent basically every three hours and the beautiful thing now is you can go work out and the agent will tell you the answer and then you voice note back with Speechify what you wanted to do next and that'll happen. And so you want people who are essentially that level of addicted. Obviously that creates massive AI fatigue so make sure your teams don't burn out. But you want someone who is that level of excited. And so I think hiring for slope more than intercept is more important today than ever before. Said another way, I look for the potential the person has more than I look for where they are today.
When I look at Whisperflow and Willow and I did this tweet and I deleted it because I don't ever want to be snarky and miserable and it's an amazing thing to build a company and you should be incredibly credited for doing so as an entrepreneur. But I found Whisperflow's product was just getting worse. And I said it on Twitter just because I honestly wanted alternatives. I really need this product and I wanted alternatives. I got 500 different alternatives and I was like, we'll talk about the commoditization of a market. That is not one that I want to be in.
Can you help me understand? We've seen the complete commoditization of that. Whisperflow, Willow, speech to text for productivity. What they came out to the market with first was not necessarily their own model. Part of the reason they got worse is they switched their own model because it's a lot more affordable. And so they had a harness that ties together a bunch of other things. Probably it was DeepL or DeepGram under the hood with a bunch of optimizations and more products. Well now they're trying to do notes and they're trying to move more into productivity. And I think they're being successful with it.
Yeah, 100%. Yeah. So that's to your point of the compound startup. One of my biggest mentors, when you look at them, do you not reflect on what we said before about not announcing fundraisers, not announcing anything? They've announced everything. They're the opposite of me. They're the opposite of me. They're the opposite of me. They're the opposite of me. They're the opposite of me. They're the opposite of me.
Correct. And hence they have I would say a bigger brand. Not in terms of users. If you walk down the street in New York City way more people will know Speechify than know Whisperflow just by virtue of the fact we have way more users. But in the tech world, way bigger brand, right? Investors know who Whisperflow is because they announce. We intentionally don't announce. But we don't have any competitors. Who are you going to use instead of Speechify to do text-to-speech for your models? The closest thing is 11 Labs and we're so much bigger than 11 Labs for B2C. What's Gradium? What's Gradium? A Brazilian company that does text-to-speech, competitors to 11 Labs.
Okay but they do B2C? Yeah I know. Yeah so there's unlimited numbers of companies doing B2B text-to-speech but we are unique in our market. So because...
Correct. And hence they have a bigger brand. Not in terms of users. If you walk down the street in New York City, way more people will know Speechify than know Whisperflow just by virtue of the fact we have way more users. But in the tech world, way bigger brand, right? Investors know who Whisperflow is because they announce. We intentionally don't announce. But we don't have any competitors. Who are you going to use instead of Speechify to do text-to-speech for your models? The closest thing is 11 Labs and we're so much bigger than 11 Labs for B2C.
What's Gradium? A Brazilian company that does text-to-speech competitors to 11 Labs. Okay but they do B2C? Yeah. So there's unlimited numbers of companies doing B2B text-to-speech but we are unique in our market. So because Whisperflow was so public about it they now have a lot of competition and so Peter Thiel, right? Only losers compete. Try to not compete. And so yes, what happens in that market? Whisperflow take majority and then there's thousands of ankle biters? I don't know. I mean I want them in that market right? I think that market becomes oligopolical as well obviously and this, by the way, I think it's another mistake that I made. I built my own speech-to-text experience that I've been using on my computer for the last seven years. Sideloaded on my iPhone and on my computer but I figured it's a commoditized product, right? Apple's going to release it instead of the button it'll be great and there you go but Apple keeps not doing it. If you remember two years ago Apple announced a partnership with ChatGPT that will improve Siri nothing happened and so that's also the reason why I never went after Siri and so now we've launched a product to compete with Siri and we've launched a product to compete with Whisperflow and we launched a product to compete with 11 Labs because I learned a lesson that I should have learned before which is the same lesson from 11 Labs. The way that you win is you offer an excellent product for free and then you have a wedge and then you add more and more things. So I don't know what happens with the Whisperflow space I just know that if you're a founder you should also always try.
Final one, another one that I stand by strongly is I just think the customer support market is a challenging market to really get behind. You have Sierra and Decagon out in front with the majority of funding and attention but to say that there are 18 companies that have now raised over 100 million in the last 18 months there is what I would call the mid tier which is your intercoms and your talk desks and your crescendos all these ones whereas they're not old but they're old enough, 8-10 years old and they're pretty good and then you've got Salesforce, Atlassian and the much older ones and then the worst thing about this market is that for any sophisticated buyer, an Airwallex, a Klarna, a Navan, a technology facing company, everyone has built their own because they need a sophisticated built-in. Why would you pay a tax for it?
So what am I missing?
Yeah, so the first thing you're missing is the core product that we're offering B2B is the API not the agents, right? So Sierra doesn't have their own model team they use other people's models right because the value of Sierra is the go-to-market, it's Brett Taylor and so that's why if you talk to Monty and Piotrek they'll tell you we're not competitive with Sierra because their main business historically has been the API. So that's the first thing. In the API business you have Speechify, Eleven Labs, Gemini, Grok, so SpaceX is now in the race and Cartesia and that's kind of it and so that's not that competitive of a space compared to the B2B customer support thing everybody's in that space, Finn, everybody and so I'm not building that product. What can I offer you that's 10x better than the next person? Not much. And so in the core API side I can offer you better quality, faster speed and 10x cheaper, good offering but then I have to also offer agents because there are so many pockets of value that have not been unlocked and unless I am, I have this model for leadership. You don't want to be a fat manager who's a general sitting in the back saying take that hill. You want to be the warrior who runs up with their sword and engages the enemy first. You need to be the same thing with your product. You need to be the number one user of your B2C product and you need to help your customers use your product better and if you do that you will learn their problems and then you will figure out what the next product is that you need to offer them. So unless I have front deployed engineers working with my B2B customers building agents for them using our technology I will not figure out what the really amazing next innovation across the hill is and so you mentioned the right thing which is Eleven Labs now has all these partnerships with governments. Governments is not exactly customer support. They would have never gotten to governments had they not done a great job on the private sector first. I agree. Eleven Labs, in addition to OpenAI, is the most integrated company right now, AI company with governments. That means they figured something out but you got to start in something like customer support. Now we support the models so if you're a person building a customer support product and you're a CFO frustrated with the size of your Eleven Labs bill and you want to cut it by 10x go to speechify.ai or hit me up cliff at speechify.com. I'll give you some discounts but again if you're a founder you need to try. You cannot not try. You cannot give up before you're even in the race.
What will be a bigger company in five years, Sierra or Eleven Labs?
Brett Taylor has the best resume I think of anyone in the world, right? I think he started Google Maps then he was CTO of Meta then he was co-CEO of Salesforce, he's on the board of OpenAI and now he founded Sierra. I would never try to fight Brett Taylor and I think the field is so large. No, we don't understand how big the space for AI agents is. Not even just AI voice agents, not even close. In the same way that people didn't understand how big the field was for LLMs in 2019 and the same way people didn't understand how big the space was for AI coding agents in 2021, this is the next huge space and so both those companies are going to be massive. I think they're playing very different games. I think Brett Taylor is actually trying to recreate the next generation of Salesforce. He is absolutely not playing the customer support game. He's moving into pre-sales, everything. He's moving post-sales but neither is Eleven Labs. Eleven Labs has a product that also does customer support but they do everything else too. That's why I call it AI agents not customer support. But I think Matt is building a very opinionated voice centric company, it's voice and I think Brett Taylor is doing all of it. Put another way, if you use a tool like Sierra the wedge right now is voice but the important part is tool calling. Eleven Labs lets you do some tool calling but that's not the bread and butter. There was a really good presentation that Brett Taylor did a screen share of him building a guitar store on Shopify and how he uses Sierra to do customer support and sales and everything else. It was extremely impressive. If you haven't searched this you should search this. Brett Taylor is a big guitar guy. That is a very different product than Eleven Labs is doing and so they're both going to crush. I agree with you on the Sierra conclusion.
What crazy thing today will be incredibly common, and this is a quick fire my friend because I could talk to you all day, what crazy thing today will be very common in five years time?
You know before it was finding your partner online, duh, no weird. Put your credit card online, no. What today is no and in five years time we'll be like yeah of course. Human computer interface is going to become primarily voice as opposed to a screen. So part of the reason why Google succeeded is it is a very simple interface. There's a text box and a button. That's it. Anyone can learn how to use it. The reason why ChatGPT worked as opposed to GPT 3 is because it was also a very simple interface, just chat. There's a text box and a button. You get a response. That's it. The simpler version of that is just having a conversation. I say something, I hear something in response. If you use voice AI from ChatGPT right now it sucks. It's too slow. The LLM is much dumber than the core LLM. The escalation to the higher quality LLM is pretty weak. I think what will happen and Meta has the right idea by the way, so go Chris Cox, is people are going to be talking to their computer and phone and some wearable constantly throughout the day and using screens a lot less.
You can buy one, SpaceX or Meta, which you buy, Meta. Why? Elon's distracted. Is he distracted or is he building full stack because actually I think he's never been more strategically positioned and he has an outlet for each of the different products that he's built and each one feeds the next. When you look at Zuck and Meta, bluntly the compute spend
It was also a very simple interface. Just chat. There's a text box and a button. You get a response. That's it. The simpler version of that is just having a conversation. I say something. I hear something in response. If you use voice AI from ChatGPT right now, it sucks. It's too slow. The LLM is much dumber than the core LLM. The escalation to the higher quality LLM is pretty weak. I think what will happen—and Meta has the right idea, by the way, so go Chris Cox—is people are going to be talking to their computer and phone and some wearable constantly throughout the day and using screens a lot less.
You can buy one: SpaceX or Meta. Which you buy? Meta. Why? Elon's distracted. Is he distracted, or is he building full stack? Because actually I think he's never been more strategically positioned, and he has an outlet for each of the different products that he's built, and each one feeds the next. When you look at Zuck and Meta, bluntly, the compute spend that he's producing—the outlet is increased conversion on an ads business, which is the biggest ads business in the world. So 7% on 240 billion dollars is a lot of money, but it's actually not in the same quantum league as doing space data centers.
Yeah, so let's take the space data centers out for a second. I think the space data centers is a very interesting idea, and what it does really well is it lets me underwrite a gigantic TAM from my expectation for SpaceX. It ruins all estimates, right. And so that's like that was a great rabbit out of the hat by Elon in order to pitch investors really well. Let's take that out for a second, and I'm going to talk to you about SpaceX and Tesla like they're one company because really I'm assessing Elon, not assessing SpaceX as an individual stock.
For data centers, the biggest constraint right now is memory cards, and then very soon it's going to be energy. And it's energy a lot of the times. So what do you need for energy? You need energy supply and you need energy storage. And the best energy storage right now actually comes from Tesla. Tesla also has a chip manufacturer that they're doing—basically competing with everyone else—and that's going to do really well.
And if you saw that Joe Rogan interview with Elon maybe two years ago, he was explaining that the hard part is not building the product. The hard part is building the manufacturing for the physical product. So Elon is number one in the world for manufacturing complex items like that. So that's very exciting. And so the TAM for Elon's companies are bigger. However, I think the Meta trades—what does Meta trade at right now? Less than SpaceX. It's less than SpaceX. And so I think Meta has more data than anybody else in the world.
I think Meta is actually super hampered by laws like GDPR. If GDPR didn't exist and the other laws in the US didn't exist, Meta would be ripping. They just can't train on their data properly. And so they'll figure that out at some point in some way. I don't know how, but I believe in Zuck. And at the end of the day, I'm a huge believer in founder-led companies. And so both—we're talking about two of the best founders in the world.
And the last thing I'll say: look at Zuck's age and look at Elon's age. And Zuck's not going to stop and Elon's not going to stop. But at a certain point one of them will expire. And so Zuck has 20 extra years. And so depending on how long you're investing—I'm younger than Zuck. Let's see what happens if Zuck expired. Meta's stock price would increase. What? I disagree completely. I know. Because you'd have a CEO who comes in and understands—and this may be short term—but that's saying we're going to invest more and more and more and more and more in CapEx when we don't have an outlet for it. You'll actually see a stock price appreciation in the short term.
Every time Zuck steps out on the podium and says CapEx, CapEx, CapEx, he's hammered for it. But that's why Meta is a good investment right now because what Meta doesn't have is what Palantir has, which Palantir has the Alex Karp effect. Alex is really good at pumping up the PE ratio of the stock. And Zuck, I agree, is the opposite. Same as Elon. It's the Elon premium. If Elon were to be removed, he loses 70% of that value. Exactly. If Zuck is removed, you definitely don't lose 70%. You maybe lose—I don't think you lose anything. I think you get an experienced exec in who says we're an ads business.
Charlie Munger and Warren Buffett actually—no, it's Benjamin Graham has this concept of the cigar butt, right? What's the intrinsic value of a company? And they approach it from an accounting perspective. I think about it as from an underlying technology and business perspective. So the underlying asset, the intrinsic value of Meta, is so large in relation to how it's valued in the market today. And you're correct. What's the PE ratio of Meta? 32, something like that. And SpaceX is insane, right. Tesla is also in multiple hundreds. I think that there has to be a correction that happens unless Elon succeeds with a big, big, big, big vision, in which case then he wins.
Final one for you: what are you most excited by? We talked about Jeff Dean. I'm—Jeff Dean? Yeah. I'm most excited by applications of AI to pharmacology and biology. So I have a family member who has very severe autoimmune neural inflammation. He's had it for six years. I took a blood sample from him every week for 15 weeks, sent it to a lab, sequenced his genome, did proteomics on it to figure out how the proteins are expressing in his body, and ran an RNA analysis in each one of those weeks. And then I compared that to self-reporting data on what the quality of life is and what his mood effect was every day.
I have six years worth of data on him and ran it on a GPU cluster, and I found so many things that no doctor could ever tell me. And he has a very rare disease called an orphan disease because there's not that many people. There's a Facebook group for this disease. I'm buying now a five thousand dollar device you can fit in your pocket. But if you put a piece of hair or saliva or blood into it, it can sequence your entire genome. And so I'm organizing meetups with all the people who have this disease to sequence all of their genomes and then compare them all on a gigantic GPU cluster to figure out what epigenetic common thread there is between them.
And I know I'm going to solve this disease. I would have never had an edge to do that in the past. And it gets even more beautiful because I can then take all the conclusions that I have about it and put it into AlphaFold from Isomorphic, and I can design not just the protein that is creating these issues, but I can design the molecule that needs to bind to that protein to either turn it on or off. I can use CRISPR to do the same thing. I can use a lab like Twist where I can tell it I want you to make me this RNA sequence or this DNA sequence, and it can make it for me and ship it to my lab or my house. And I can create amazing outcomes with it.
And I can simulate all of it on my computer that's SSH-ing into my GPU cluster in Scottsdale, Arizona. And I can cure my brother. And so my experience is I'm a kid who when I was eight years old couldn't learn how to read, and my dad had to open a book and read Harry Potter to me, and that's how I learned how to read. And then when I was 13, I moved to the United States of America, and I didn't speak English, and I listened to Harry Potter audiobooks 22 times in a row, and I still had the first chapter memorized.
And then I couldn't get into the private high school that my brother went to and my sister went to, and I was really bummed. I went to a lower quality high school. And I didn't get into AP US History because I made a bunch of spelling mistakes in my essay and couldn't read the passage in time. And I needed to train myself to read the SAT English portion. I wouldn't read the passage. I would read the answers and then I'd go and hunt for the answer.
And then when I got to college, somehow by the grace of God, I ended up going to Brown and starting a major for renewable energy engineering because I couldn't do literature. And I built a text-to-speech tool that would read out all my books to me, and that's why I graduated. Technology solved my dyslexia and it solved my ADHD, and it's going to solve my brother's disease. And it's already solved my dad's prostate cancer because I figured out, with a bunch of help from other people, how to use GPUs to identify where in his body the lesion was.
That's what I'm excited for: there's better quality of life for literally everybody because you have this magical machine that can run a trillion operations per second on as many GPUs as you want, and it can solve problems that we can't.
I find it staggering that still today we have orphan diseases, which is like, "Oh, there's too few people to make it economically viable for us to try and solve," and there are thousands, low thousands. But again, it's the same thing. You just need data, you need compute, and you need to ask good questions. Like I said, 10 good decisions per day—either hypotheses or actual product decisions—and you can solve these problems. Freaking amazing. Cliff, it's been so great to have you on the show. I much prefer it when it's a discussion.
My brother's disease, and it's already solved—my dad's prostate cancer—because I figured out, with help from other people, how to use GPUs to identify where in his body the lesion was. That's what I'm excited for: better quality of life for literally everybody, because you have this magical machine that can run a trillion operations per second on as many GPUs as you want, and it can solve problems that we can't.
I find it staggering that still today we have orphan diseases where there are too few people to make it economically viable for us to try and solve. And there are hundreds, thousands, low thousands, but low thousands. Again, you just need data, you need compute, and you need to ask good questions. As I said, 10 good decisions per day, either hypotheses or actual product decisions, and you can solve these problems. Freaking amazing. Cliff, it's been so great to have you on the show. I much prefer it when it's a discussion; it's been an amazing discussion, so thank you so much for putting up with me. My pleasure. Thank you so much for your attention.
Thank you so much for your attention. Thank you so much for your attention. Thank you so much for your attention. Thank you so much for your attention. Thank you so much for your attention. Thank you so much for your attention. Thank you so much for your attention. Thank you so much for your attention. Thank you so much for your attention. Thank you so much for your attention. Thank you so much for your attention. Thank you so much for your attention. Thank you so much for your attention. Thank you so much for your attention. Thank you so much for your attention. Thank you so much for your attention. Thank you so much for your attention.
Thank you so much for your attention. Thank you so much for your attention. Thank you. So like, you know, like you need to indemnify the companies who are using you and that's part of why they buy it from you as opposed to sourcing it themselves. We've seen the lawsuits.
But it's a great business. And if you could do it well, but like you need to be an ops monster. Like you need to be really, really good at operations. You need to be very fast. And really the key is the company training on your data needs to actually see improvements in their model at the end of the day. The thing that has always been challenging for speechifying compared to other companies is B2C customers pay a lot less than B2B customers. So 11 Labs, huge credit to them, leapfrogged us because they sell to B2B. Well, historically, we've only sold to B2C. And so our big constraint is we need to do this on a cost basis of it needed to cost us, you know,
less than $10 per million characters. 11 charges, $100 per million characters. The OpenAI model on the benchmarks cost $196 per million characters. So ours, when we sell it to other B2B companies now, we just launched our API, Simba 3.2, it costs $10 per million characters. Dude, I am, I'm too old to not ask the painful questions. And I think the joy is when you ask them and kind of less, you know, why are you about asking? You said 11 Labs kind of leapfrogged you. Is that on you for not doing B2B? Yeah, 100% is on me. 100% is on me. It's the biggest strategic mistake I made in the history of speechifying. How do you reflect on that? So I met Piotrek and Mati.
I was living in London at the time in my house in London. I think it was 2022. And we were very impressed by them. And we wanted to use the model, by the way. It was just too expensive for us to use. And I looked at it and my thought to myself was, they're very smart. They're going to do well, but I don't like their strategy because I think that an API for Texas Feature will become commoditized with the time, right? You're going to get to the point that you can run that API on your computer and then on your phone and then like, what are they selling anymore? So I don't want to go into that business. And I made a critical error. What I didn't understand
is that the point of an AI lab like Speechify or like 11 Labs is to continuously innovate. And the first product that you release is your wedge that gets other people to then later use your other technology. So for example, if you're in text-to-speech, you built the best text-to-speech model in the world for one specific voice. Cool. Well, now you can do other voices. Now you can add emotional porosity. Now you can add voice cloning. Now you can add speech-to-text. Now you build duplex models where it makes the um, ah, laughter, interruption handling, turn-taking. You add a harness for voice conversations. Then you optimize it for sales and you optimize it
for customer support and you optimize it for all these things. And so what they did is they first built an amazing API. They were great at launches. They built a really great product for creators. Then they built their best product ever, which was agents. Agents is amazing because the buyer is no longer a software engineer. The buyer is a CTO, CIO, CEO, executive in the company. Sierra has this concept called outcome-based pricing. Brett Taylor is amazing. And so you can start fighting on the outcome. And having a AI agent is like having an AI co-worker. But it was my mistake to think that an API product was a bad strategy because I thought it was something
that would become commoditizable. And I forgot the central thesis about Silicon Valley, which is constantly innovate. Get the user to start using your product. I don't care if it's free. Then you sell them other things. And so that was my big, big, big, big mistake. How possible do you think it is? I think people underestimate the complexity of building out a B2B GTM. Super hard. I think it's a strategic mistake for Speechify to go to B2B. Hmm. A lot of people think that. Tell me your position. You are now competing against Eleven Labs and Sierra, really. And those two are competing. Whether they like to admit it or not, they absolutely are competing. And they will.
I'm sure if you ask them off camera. That's Brett Taylor. Yeah. You don't want to compete against Brett Taylor. Motherfucker, I don't want to compete against Brett Taylor. And that is the tidal wave of Eleven Labs. Now, Eleven Labs is an unstoppable machine at this point, to the point where it has government buy-in across all of the large, major Western democracies. Actually, it's insane the government buy-in they have. And they started three months ago. I just said the key thing. They started three months ago. Yeah. And so, you know the graph. But I think they've reached a tipping point where actually they've just taken the market. I think Sierra are running
behind them chasing. And they're doing a decent job of it, but they've got Brett and they've got Sequoia and Green Oaks and every royalty of Silicon Valley behind them. And they're still running behind chasing Eleven Labs with Sequoia kind of pretending to be neutral because they're in both of them, which is incredibly challenging. And I just think being third, the Postmates effect is never a good market to be in when I could be the dominant consumer brand that leads with a really different and compelling story. So here's the two things to consider. The first one is if you go to the App Store and you search text-to-speech, Speechify has 98% of the installs
in text-to-speech for B2C. Speechify has served more than 770 billion words to users over the last few years, which in terms of times of listening, it's like 6,000 years of listening, right? If you go from today to zero BC and back, you still have like thousands of years left. So we've like completely dominated that market and it's still a business that's growing really, really fast. And we're constantly adding more features into that product. The thing is, we have a pretty big engineering team and now everybody is capable of doing 10x what they did before. So I have extra staff. I have a huge AI engineering team with ability to make amazing models. So like,
where is the highest ROI for that to go? Well, it needs to go both B2C, but it should also go B2B. And one thing that I will never be is a person who doesn't learn. So I might as well just freaking learn B2B. Now, to your point about competing against giants like Sierra or Eleven Labs, hey, Anthropic came into the market as a second to OpenAI and they were second for a very long time and now they're not second. Facebook came as a second to Frontster and MySpace and now they're not second. And so the nice part is this space is not a monopolistic space, it's an oligopical space. And if you look at what happened with Eleven Labs, I'm going to exclude Sierra
because Brett Taylor effect is huge. It's just amazing to see how good of a business that is. And so it might very well be that for the core offering that they're currently winning on, I will not win. But what did I learn last time? It's fine if I offer my product essentially for free because I'm an AI research lab. And as long as people start to use me with time, I'll be embedded in the system and I'll keep coming out with more and more and more innovations that are useful to them. And so there's an unbelievable demand from all these companies and governments and everybody else for great tools, whether they be AI agents or APIs or products. I just want to be
on your phone if you're a user or in your stack if you're a company and supply you with the best front-deploying engineer experience and AI orchestration experience and API experience to give you an amazing experience and there's room for everybody. I agree there's room for everybody. I think value accrues to top one player. I agree. I think, you know, it's kind of like the inference market where Fireworks will be a multi-hundred billion-dollar company and then I think a genuine base time will be a hundred billion-dollar company and then together and a load of the others will be 15, which is amazing. It is completely true. It's hugely amazing, valuable companies.
But you would then think that OpenAI would be the place where value accrues for voice AI, right? That's what you would have thought three years ago and that's not what ended up happening. So you can't not go into the race because there's a big incumbent. Well, I think with all candor that's because of incredibly poor management. I agree. And hiring. And that was theirs to take and they fumbled the bag across every spectrum. And for every company in the world, no matter how exceptional the leadership team is, niches get fumbled, right? So voice AI was a niche for OpenAI, right? LLMs are the core. And by the way, they also fumbled AI coding. Now they're trying to cash
because it's such a big space. All respect to Piotr and Mati. I think they're absolutely amazing and I love working adjacently to them. I just don't think they're going to fumble the bag. That's my trouble. But they have so much in their net right now. That's true. And so much is getting added to the net constantly. That's true. And so you just, you have to like go where the football is going. Yeah. And so I think that it's too expensive for speechifying not to be playing in B2B as well as playing in B2C. The best way to lose is not to be in the race. Be in the race. In terms of products that we build, we were chatting earlier and you said that every startup stay
has to be a compound startup. Can you talk to me about that and how you think about that? It's not that every startup has to be a compound startup. It's at a certain point you can't afford not to be that. Do you not think there are a few companies that are just absolutely fucking running rings around everyone else? Yeah, absolutely. Those are the winners, right? Eleven Labs is an example. Anthropic is an example. Ramp is an example. Speechify is an example. All the companies that have absolutely maniacal leadership teams and engineering teams. Like that's why people care about team more than almost anything else. Because the right team will iterate fast, get there,
and then figure it out. And now, when everything can be turned into a reinforcement learning problem where you can have long horizon agents and orchestrating agents thinking about the problem for like two weeks at a time, if you set that up, of course you're going to win. I got into a lot of trouble as I always do with most of my social posts. I used to be quite a sweet little boy, actually. No, really. I used to be like the Harry Potter of Ange Capital and now I'm more like... Yeah, you lost the glasses. Lost the glasses and it kind of became more like Piers Morgan, if you know Piers Morgan in the UK. Highly despised figure. Very opinionated. But a question that I have
is like, I said, if you're a startup, it's never been harder to hire great talent because OpenAI and Anthropic, candidly, have such a carrot reward mechanism in front of you, especially with impending IPOs, that the best talent just wants to go there and talent follows talent. And you're seeing the fucking founder of Monzo, a multi-billion dollar bank in the UK go there from YC as a partner. Matt Clifford, the founder of EF, which is a multi-billion dollar company. I mean, he should be fucking prime minister and he's going to join Anthropic. Am I wrong that this is the hardest time of startups to hire because the prizes of Anthropic and OpenAI are so great?
My favorite type of person to hire is a CTO of another company. We have, well, we were 21 people at Speechify. 18 of the folks at the company were previously either CEO, CTO, or VP of engineering at the last company. Anthropic, I have never seen a company like this, hires so many CTOs of publicly traded companies and other successful startups. Workday, one of them. The reason is they build the best, most beloved product for engineers in the history of the world. So it's easy to hire CTOs. By the way, they hire much more CTOs than CEOs because CTOs are the ones who get the most excited about this product. And you're right. They're the fastest growing company ever,
especially at the scale that they are. So they're going to keep growing. OpenAI is going to keep growing. You had this like very condensed period, like fireworks of growth in both of those companies. Yeah, it's very hard to hire. But remember, they're hiring people that their annual compensation needs to be $15 million a year minimum. What startup is hiring someone and paying them $15 million a year? You're not. Like, your seed founder, that was not something that you were going to hire. And so I will push back against it. The competition for growth stage companies hiring exceptional leadership talent is more difficult. For seed companies, I would say
it's the easiest time ever because the impact of even just the founder on their own is bigger because they can orchestrate agents. But the same thing for hiring. So one thing that we have changed about our hiring in the last even six months is we really cared that you read a ton of textbooks about software engineering and that your handcrafted code was amazing. I still care that you read a lot of textbooks about software engineering and you understand it. But the thing I care about the most today is technical aptitude and just like raw technical intelligence because I know that we could teach you everything else and in six months you could be a machine. And so we hire
a lot of math olympiads and leak coders and like Kaggle award winners and people who like studied physics and math. Like they might have been not coded before because I just need the hunger and the work ethic and the intelligence and anyone can become so good so fast now. And so the pool for hiring exceptional talent is bigger than ever before and Duolingo did this really well. They love hiring college grads and coaching them. And so I wouldn't say that it's harder to hire than ever before for seed companies. Seed companies now almost anyone can be someone that you hire if they're smart and hardworking because you can teach them very fast. What is more challenging to hire
is for growth companies because you're fighting with just absolute juggernauts. Are you not a growth company? So it's challenging for us. Why do you think it's hard to hire a really good salesperson? I totally get that and I completely agree. I will see CRO packages in the 50 million plus range by the way. Yeah, exactly. 15 is like kids play. By the way, with the greatest of respects I will even see $15 million on the table for comp packages for seed companies today. That is the dislocation that I think with the greatest of respects Wait, wait, wait. Sorry, sorry. But this is a seed company that has raised how much money at what evaluation?
I mean, you've got to understand that seed round today will be $150, $200 million. And there are several of them. I mean, there's 30, 40 companies that at seed have raised $100 to $300 million. And this is a company of like a guy who's like one year out of university? No, no, no. This is a guy who's probably spent four years at OpenAI or spent four years at TMI. So then what about the company that's like, you know, the guy who's been in university for like two, three, four years and now they're starting a company? Or do you think that those people are out of the water now? No, I think that that's just a very different world. And so, yeah, they'll raise $10 million
seed rounds. Yeah. So for the company that you just described, they raised a seed round at $150 valuation and they raised, I don't know, $20 million? No, I said it was $150 million raise. Oh, I wouldn't call that a seed round. Maybe that's the name. But my point is, and that's my point though, which is like the talent is concentrated. The people who really fucking get AI and systems and have seen the magic inside OpenAI, Anthropic, I agree with you that if you have a company that's raised $150 million and a $500 to $2 billion valuation, definitely that company should give $15 million comp package. And there's a lot of them. Yeah, that makes perfect sense.
But there's a lot of them, there's 30. And those 30 take 30 people. And there is a thousand people now. That is fucking hard. And so, but what you just described is exactly what used to happen with Google and Meta, let's call it six years ago, which is if you were really cracked, there was essentially a maximum amount that you can get paid at a company like Google or Meta. And the best way for you to make a life-changing amount of money is to go to a company that is small and ride from the beginning all the way through and be a really solid founding engineer at that company. I think people want more certainty of cash today than upside, which sounds more... No,
I think that the equation is the same as always, which is each person has their own equation in their head of how much certainty and how much risk they're willing to take. It hasn't changed. It's the same. Humans are still humans. But I think people would rather know that the certainty of a $10 million from Anthropic versus a 60 from that quirky startup they could make. This is the reason why companies IPO, right? There's two reasons. Either you want a ton of money or you want a lot of credibility in B2B like Zoom did or you're hiring and the value of the package that you offer is so much better when your stock is liquid. When we look at that dev team for you today,
you said, hey, I wanted to go in. I want to see how we're orchestrating agents. What did you find? What did you learn in that discovery process around agent orchestration internally? So inside of our AI research team, everybody's orchestrating agents. It's when you go lower, not lower, if you go then into the product-facing things that we build, for example, the platform team or the iOS team or the Mac team or the Chrome team or the web team or the Android team, these are super smart folks who have been working in those domains for like 10 years and they know iOS like the back of their hand. They know Kotlin, JetBrains like the back of their hand and so it's very easy
for them to hand code things because you're not dealing with something that's like super, super new. So why change? People, you know, it's hard to change, right? And so you just need to force them to change. So one, the best thing is to inspire. So you do a Zoom screen share and you show them how the best engineer in the team is orchestrating agents and they're like, oh wow, I didn't know you could even do that. And then you go, yeah, like please do it. You recommend blog posts for them to read, books for them to read, Twitter threads for them to read. What's the team using? Cloud Code, Cursor, Codex? Cursor and Cloud Code. Those are the two most popular. Yeah.
It's a little bit of Codex usage. It's not that big. I would say Cloud Code is number one, then Cursor, then Codex. We want you to use as many tokens as possible in whatever harness way is the best for you. You mentioned linear. Linear is amazing. Like automatically cutting tickets from linear is fantastic. And just like being able to go into your agents and be like, okay, I have these like six linear tickets, start on them. And then really a good engineer today is just an exceptional QA, right? The AI will make them feature. You will test the feature, see if it's good. You'll figure out where the edge cases are. You'll prompt it to fix it. And then you try to make it
as efficient as possible, which is hard to do. And then you need to make essentially like, you know, roughly 10 really good product and engineering architecture decisions a day. How do you think about token allocation internally? I mean, we've seen leaderboards be used, which is I think the most fucked up form of incentive kind of playing. You don't want to prevent people. There's a lot of people who are a lot of talk and I'll ask for examples and I'll read the examples and like, you know, I'm doing this, I'm doing this, I'm doing this, I'm doing this, and then you look and I'm like, eh. And so I think about it in terms of demos. Can we hop on a Zoom call
and you'll show me what you built and then I use it myself and I'm like, wow, that's amazing. Or you send me a screen recording of a feature or technology that you built and I'm like, wow, that's so good. And so we give credit when things get shipped to production to users. So even inside of the AI team, if you build a really amazing, and this is part of why Speechify ended up winning, you asked, how did you build bigger labs? The answer is, we ship to production all the time. That's how we won. We are not in the theory space. We are an applied AI company. That's why we win. And so if you're an engineer at Speechify, the analogy I always give people
is imagine that you are in the milk delivery business and you make me a beautiful bottle of milk and you leave it down the road. The milk will spoil. You have to get it to my door. Knock. If you didn't do that, you get no credit. If you carry the football all the way to the line, but you don't cross over to the end zone, if you don't kick it into the goal, you get no credit. If you bring the ball just to the rim, you don't put it in the rim, you get no credit. And in the rim means push to production with no bugs and users are actually using it. And then we get feedback. How many companies do that iteration cycle fast? Almost no one. Definitely not with a user-based
number that Speechify has. And so in the AI team at Speechify, you make some amazing discovery. We're like, great. Push it to production. And then you go, oh wait, there's this QA problem and this QA problem. And if you have this many people use it on the AI serving layer, then you have this other issue. Cool. You get no credit from me. It's not in production. I can't use it on my phone. When I can use it on my phone, I will give you credit. And so this morning, actually not yesterday, yesterday I had a call with our AI engineering team. And I said, listen, the project that we have running for duplex models and for AI conversational harnesses is something
I'm really excited about. And it's been moving fast. I want it to move faster. Here's like 14 notes that I want. And then what I do always is I'm on a Zoom call. I flip my computer around to face my phone and I use the product in front of them and we record it. And so then they see all the bugs. And then I send the recording in the chat. Someone on our team, he's 19 years old, sent me a demo this morning off of that conversation that solved all of my problems. And he was like, hey, I was waiting for like three training runs to finish. So I had a little bit of time while I was waiting. So I implemented everything that you asked and it blew my mind. It was so good.
That's using AI correctly. So it's not a token leaderboard. It's what did you show in production that was good? How many companies do you think are actually as token-pilled AI-centric as we think in terms of devs? There's a guy, Jason Yeager, who used to work at Speechify and now he has MyTech CEO on Instagram. He's super funny. And so he makes a lot of videos about like, you know, crazy CEOs who all are like, use tokens, use tokens. I think all founders in some way have that animal inside of them because you know that it's the right path. But there is a difference between reality and theory and you need to make sure that you don't overdo it.
Do you have any price sensitivity on tokens? Yeah, of course. Absolutely. I mean, I'll lose my mind if to implement a tiny feature you use 15,000 tokens. Like, why did you do that? And like, we will let people go if they just go bananas with something for no reason. Are you able to accurately budget tokens? Not accurately, but within bounds. The other thing is like a lot of engineers are, look, you go into engineering because you like optimization. Most engineers are not blind and it physically hurts them to overspend tokens. And I, again, I always think that the best way to interact with AI is you are chatting in the chat or actually doing it verbally
and you're essentially pseudocoding with your words constantly and you're explaining architecture. And a great example would be I know someone who has no engineering background and they wanted to build an app and they built exactly what they want. It took them two hours. But they needed an API call and they needed to scrape this website and they basically scraped every single page of the website, every single part of the website. And so the bill that they got for the scraping was gigantic. And then I was like, why are you doing like that? Why aren't you going into the database to this exact URL and then scraping that from the URL? So the amount of nodes they needed to hit
became like 20 instead of 25,000. And so an engineer will spend their time making sure that the thing is optimized like that. So that's how you build like a good database or a good architecture system, whatever. You do the same thing when you're interfacing with the agent. You want the agent to take the path of least resistance, not the path of most resistance. I think one of the biggest problems is that agents are goal seeking. And so they are like... It's all about the target. You need to be good at picking the right target. And I think Anthropic published this paper when Fable 1 came out about long horizon tasks with Fable. So the first thing is it was much better
at like running a two-week task. And it could burn $12,500 worth of tokens in two weeks and basically make a better model with that. That's a perfect, amazing way of using tokens. That's exactly what you want. And what you don't want is burning 12,000 tokens in the span of five hours doing something that's like totally unnecessary and doesn't make any sense. You need the loops to happen and then you need to check the result. So what you want to build, and Boris, who's the inventor of Cloud Code, talks about this all the time, it's all about the loop. You say, here is the target. Here's how you measure the target. Now iterate against the target over and over and over again
until you get it. What did you not know about building an AI-centric dev team that you wish you had known? How useful is it to own your own GPUs? What was that realization moment? Just did you see a build one there? The realization moment was when we realized that we had really talented engineers who were essentially moving at one-seventh of the speed they could have if they had the compute one-to-one with their creativity and ideas. How, if you're a founder listening to this, how should I change my hiring process in a new AI world? Number one, functional interviews. Build this and then you see if they can build the thing and then you run it through unit tests.
The second one is give them a large code base, even an open source repository and have them understand the code base, make changes and then check what they broke. And then, yeah, like they have to be able to orchestrate agents well. And if they're not doing that, it's kind of not worth to have the person. And then the next thing I'll say is it is more fun to have a smaller team. Like having a big team is great as long as everyone is carrying their weight. But the way I kind of think about it is yes, I can have multiple agents running on my computer or I can have several Slack chats with really smart people who are bigger domain experts than I am
and basically that human being is the outcome owner for that task and they have the agents. And so I can run as a founder multiple projects at the same time to a really amazing level of granularity. And so I think about moments earlier in the year when my brother Tyler would literally have an alarm to wake up at 3 in the morning because he needed to check what the agent was doing at 3 in the morning and then you wake up, make sure it's good, go back to sleep. Like you want to babysit your agent basically every three hours and the beautiful thing now is you can go work out and the agent will tell you the answer and then you like voice note back with Speechify
what you wanted to do next and like that'll happen. And so you want people who are essentially that level of addicted. Obviously that creates massive AI fatigue so make sure your teams don't burn out. But you want someone who is that level of excited. And so I think hiring for slope more than intercept is more important today than ever before. Said another way, I look for the potential the person has more than I look for where they are today. When I look at Whisperflow and Willow and I did this tweet and I deleted it because I don't ever want to be sarky and miserable and it's an amazing thing to build a company and you should be incredibly credited for doing so
as an entrepreneur. But I found Whisperflow's product was just getting worse. And I said it on Twitter just because I honestly wanted alternatives. I really need this product and I wanted alternatives. I got 500 different alternatives and I was like motherfucker we'll talk about the commoditization of a market. That is not one that I want to be in. Can you help me understand? And we've seen the complete commoditization of that Whisperflow Willow speech to text for productivity. What they came out to the market with first was not necessarily their own model. Part of the reason they got worse is they switched their own model because it's a lot more affordable.
And so they had a harness that ties together a bunch of other things. Probably it was DeepL or DeepGram under the hood with a bunch of optimizations and more products. Well now they're trying to do notes and they're trying to move more into productivity. And I think they're being successful with it. Yeah. 100%. Yeah. So that's to your point of the compound startup. One of my biggest mentors When you look at them do you not reflect on your we said before not announcing fundraisers not announcing anything They've announced everything. They're the opposite of me. They're the opposite of me. They're the opposite of me. They're the opposite of me. They're the opposite of me.
They're the opposite of me. Correct. And hence they have I would say a bigger brand. Not in terms of users. Like if you walk down the street in New York City way more people will know Speechify than know Whisperflow just by virtue of the fact we have way more users. But in the tech world way bigger brand right? Investors know who Whisperflow is because they announce. We intentionally don't announce. But we don't have any competitors. Who are you going to use instead of Speechify to do text-to-speech for your models? Like the closest thing is 11 Labs and we're like so much bigger than 11 Labs for B2C. What's Gradium? What's Gradium? A Brazilian company
that does text-to-speech competitors to 11 Labs. Okay but they do B2C? Yeah I know. Yeah so there's unlimited numbers of companies doing B2B text-to-speech but like we are unique in our market. So because Whisperflow was so public about it they now have a lot of competition and so Peter Thiel right? Only losers compete. Try to not compete. And so yes they What happens in that market? Whisperflow take majority and then there's thousands of ankle biters? I don't know. I mean I want them in that market right? I think that market becomes oligopical as well obviously and this by the way I think it's another mistake that I made. I built my own speech-to-text experience
that I've been using on my computer for the last like seven years. Sideloaded on my iPhone and on my computer but I figured it's a commoditized product right? Apple's going to release it instead of the button it'll be great and like there you go but Apple keeps not doing it. If you remember two years ago Apple announced a partnership with ChatGPT that will improve Siri nothing happened and so that's also the reason why I never went after Siri and so now we've launched a product to compete with Siri and we've launched a product to compete with Whisperflow and we launched a product to compete with 11 Labs because I learned a lesson that I should have learned before
which is the same lesson from 11 Labs. The way that you win is you offer an excellent product for free and then you have a wedge and then you add more and more and more things. So I don't know what happens with the Whisperflow space I just know that if you're a founder you should also always try. Final one another one that I stand by strongly is I just think the customer support market is a challenging market to really get behind. You have Sierra and Decagon out in front with the majority of funding and attention but to say that there are 18 companies that have now raised over 100 million in the last 18 months there is kind of what I would call like the mid tier
which is like your intercoms and your talk desks and your crescendos all these ones whereas like you know they're not old but they're old enough 8-10 years old and they're pretty good and then you've got Salesforce Atlassian and the much older ones and then the worst thing about this market is that for any sophisticated buyer an Airwallex a Klarna a Navan a technology facing company everyone has built their own because they need a sophisticated built-in Why would you pay a tax for it? So what am I missing? Yeah so the first thing you're missing is the core product that we're offering B2B is the API not the agents right so Sierra doesn't have their own model team
they use other people's models right because the value of Sierra is the go-to-market it's Brett Taylor and so that's why if you talk to Monty and Piotrek they'll tell you we're not competitive with Sierra because their main business historically has been the API so that's the first thing in the API business you have Speechify Eleven Labs Gemini Grok so SpaceX is now in the race and Cartesia and that's kind of it and so that's not that competitive of a space compared to yeah the B2B customer support thing everybody's in that space Finn everybody and so I'm not building that product what can I offer you that's 10x better than the next person not much and so in the core
API side I can offer you better quality faster speed and 10x cheaper good offering but then I have to also offer agents because there are so many pockets of value that have not been unlocked and unless I am again I have this model for leadership you don't want to be a fat manager who's like a general sitting in the back saying take that hill you want to be the warrior who runs up with their sword and engages the enemy first you need to be the same thing with your product you need to be the number one user of your B2C product and you need to help your customers use your product better and if you do that you will learn their problems and then you will figure out
what the next product is that you need to offer them so unless I have front deployed engineers working with my B2B customers building agents for them using our technology I will not figure out what the really amazing next innovation across the hill is and so you mentioned the right thing which is Eleven Labs now has all these partnerships with governments governments is not exactly customer support they would have never gotten to governments had they not done a great job on the private sector first I agree Eleven Labs is in addition to OpenAI is the most integrated company right now AI company with governments that means they figured something out but you got to start
in something like customer support now we support the models so if you're a person building a customer support product and you're a CFO frustrated with the size of your Eleven Labs bill and you want to cut it by 10x go to speechify.ai or hit me up cliff at speechify.com I'll give you some discounts but again if you're a founder you need to try you cannot not try you cannot give up before you're even in the race what will be a bigger company in five years Sierra or Eleven Labs? Brett Taylor has the best resume I think of anyone in the world right? I think he started Google Maps then he was CTO of Meta then he was co-CEO of Salesforce he's on the board of OpenAI
and now he founded Sierra I would never try to fight Brett Taylor and I think the field is so large like no like we don't understand how big the space for AI agents is like not even like AI voice agents not even close in the same way that people didn't understand how big the field was for LLMs in 2019 and the same way people didn't understand how big the space was for AI coding agents in 2021 like this is the next huge space and so both those companies are going to be massive I think they're playing very different games I think I think Brett Taylor is actually trying to recreate the next generation of Salesforce he is absolutely not playing the customer support game
he's moving into pre-sales everything he's moving post-sales but neither is Eleven Labs Eleven Labs has a product that also does customer support but they do everything else too that's why I call it AI agents not customer support but I think Matt is building a very opinionated voice centric company it's voice and I think Brett Taylor is doing all of it put another way if you use a tool like Sierra the wedge right now is voice but the important part is tool calling Eleven Labs lets you do some tool calling but that's not the bread and butter there was a really good presentation that Brett Taylor did a screen share of him building a guitar store on Shopify
and how he uses Sierra to do customer support and sales and everything else it was extremely impressive if you haven't searched this you should search this Brett Taylor is a big guitar guy that is a very different product than Eleven Labs is doing and so they're both going to crush I agree with you on the Sierra conclusion what crazy thing today will be incredibly and this is a quick fire my friend because I could talk to you all day what crazy thing today will be very common in five years time you know before it was like find your partner online duh no weird put your credit card online fuck no what today is no and in five years time we'll be like yeah of course
human computer interface is going to become primarily voice as opposed to a screen so part of the reason why Google succeeded it is a very simple interface there's a text box and a button that's it anyone can learn how to use it the reason why chat GPT worked as opposed to GPT 3 is because it was also a very simple interface just chat there's a text box and a button you get a response that's it the simpler version of that is just having a conversation I say something I hear something in response if you use voice AI from chat GPT right now it sucks it's too slow the LLM is much dumber than the core LLM the escalation to the higher quality LLM is pretty weak
I think what will happen and Meta has the right idea by the way so go Chris Cox is people are going to be talking to their computer and phone and some wearable constantly throughout the day and using screens a lot less you can buy one SpaceX or Meta which you buy Meta why Elon's distracted is he distracted or is he building full stack because actually I think he's never been more strategically positioned and he has an outlet for each of the different products that he's built and each one feeds the next when you look at Zuck and Meta you know bluntly the compute spend that he's producing the outlet is increased conversion on an ads business which is the biggest
ads business in the world so 7% on 240 billion dollars is a lot of fucking money but it's actually not in the same quantum league as doing space data centers yeah so let's take the space data centers out for a second I think the space data centers is a very interesting idea and what it does really well is it lets me underwrite a gigantic TAM from my expectation for SpaceX it ruins all estimates right and so that's like that was a great rabbit out of the hat by Elon in order to pitch investors really well let's take that out for a second and I'm going to talk to you about SpaceX and Tesla like they're one company because really I'm assessing Elon I'm not assessing
like you know SpaceX is an individual stock you know for data centers the biggest constraint right now right now is memory cards and then very soon it's going to be energy and it's energy a lot of the times so what do you need for energy you need energy supply and you need energy storage and so the best energy storage right now actually comes from Tesla Tesla also has a chip manufacturer that they're doing basically competing with everyone else and like that's going to do really well and if you saw that Joe Rogan interview with Elon maybe two years ago he was explaining that the hard part is not building the product the hard part is building the manufacturing
for the physical product so Elon is number one in the world for manufacturing complex items like that so like that's very exciting and so the TAM for Elon's companies are bigger however I think the Meta trades what does Meta trade at right now less than SpaceX it's less than SpaceX it's fucking and so I think Meta has more data than anybody else in the world I think Meta is actually super hampered by laws like GDPR like if GDPR didn't exist and the other laws in the US didn't exist Meta would be ripping they just can't train on their data properly and so they'll figure that out at some point in some way I don't know how but I believe in Zuck and at the end of the day
I'm a huge believer in founder-led companies and so both we're talking about two of the best founders in the world and the last thing I'll say look at Zuck's age and look at Elon's age and Zuck's not going to stop and Elon's not going to stop but at a certain point one of them will expire and so Zuck has like 20 extra years and so depending on how long you're investing I'm younger than Zuck let's see what happens if Zuck expired Meta's stock price would increase what? I disagree completely I know because you'd have a CEO who comes in and understands and this may be a short term but that's saying we're going to invest more and more and more and more and more in CapEx
when we don't have an outlet for it you'll actually see a stock price appreciation in the short term every time Zuck steps out on the podium and says CapEx CapEx CapEx he's like fucking hammered for it but that's why Meta is a good investment right now because what Meta doesn't have is what Palantir has which Palantir has the Alex Karp effect Alex is really good at pumping up the PE ratio of the stock and Zuck I agree is the opposite same as Elon it's the Elon premium if Elon were to be removed he loses 70% of that value exactly if Zuck is removed you definitely don't lose 70% you maybe lose I don't think you lose anything I think you get an experienced exec in who says
we're an arts business Charlie Munger and Warren Buffett actually no it's Benjamin Graham has this concept of the cigar bot right like what's the intrinsic value of a company and they approach it from an accounting perspective I think about it as from an underlying technology and business perspective so the underlying asset the intrinsic value of Meta is so large in relation to how it's valued in the market today and you're correct what's the PE ratio of Meta 32 something like that and SpaceX is insane right Tesla is also in multiple hundreds I think that there has to be a correction that happens unless Elon succeeds with a big big big big vision in which case
then he wins final one for you what are you most excited by we talked about Jeff Dean I'm Jeff Dean yeah I'm most excited by applications of AI to pharmacology and biology so I have a family member who has very severe autoimmune neural inflammation he's had it for six years I took a blood sample from him every week for 15 weeks sent it to a lab sequenced his genome did proteomics on it to figure out how the proteins are expressing in his body and run an RNA analysis in each one of those weeks and then I compared that to self-reporting data on what the quality of life is and what his mood effect was every day I have like six years worth of data on him and ran it
on a GPU cluster and like I found so many things that no doctor could ever tell me and he has a very rare disease it was called like an orphan disease because there's not that many people there's a Facebook group for this disease I'm buying now basically like you know it's like a five thousand dollar device you can fit in your pocket but if you put a piece of hair or saliva or blood into it it can sequence your entire genome and so I'm organizing meetups with all the people who have this disease to sequence all of their genomes and then compare them all on a gigantic GPU cluster to figure out what epigenetic common thread there is between them and I am I know I'm
going to solve this disease I would have never had an edge to do that in the past and like it gets even more beautiful because I can then take all the conclusions that I have about it and put it into alpha fold from isomorphic and I can design not just the protein that is creating these issues but I can design the molecule that needs to bind to that protein to either turn it on or off I can use CRISPR to do the same thing I can use a lab like Twist where I can tell it I want you to make me this RNA sequence or this DNA sequence and it can make it for me and ship it to my lab or my house and I can create amazing outcomes with it and I can simulate all of it on my computer
that's SSH into my GPU cluster in Scottsdale, Arizona and I can cure my brother and so my experience is I'm a kid who when I was eight years old I couldn't learn how to read and my dad had to open a book and read Harry Potter to me and that's how I learned how to read and then when I was 13 I moved to the United States of America and I didn't speak English and I listened to Harry Potter audiobooks 22 times in a row and I still had the first chapter memorized and then I couldn't get into the private high school that my brother went to and then my sister went to and I was really bummed I went to like you know a lower quality high school whatever and I didn't get
into AP US history because I made a bunch of spelling mistakes in my essay and I couldn't read the passage in time and I needed to train myself to read the SAT English portion I wouldn't read the passage I would read the answers and then I'd go and hunt for the answer and then when I got to college somehow by the grace of God I ended up going to Brown and like starting a major for renewable energy engineering because I couldn't do literature and I built a text-to-speech tool that would read out all my books to me and that's why I graduated technology solved my dyslexia and it solved my ADHD and it's going to solve my brother's disease and it's already solved
my dad's prostate cancer because I figured out with a bunch of help from other people how to use GPUs to identify where in his body the lesion was that's what I'm excited for is there's better quality of life for literally everybody because you have this magical machine that can run a trillion operations per second on as many GPUs as you want and it can solve problems that we can't I find it staggering that still today we have orphan diseases which is like oh there's too few people to make it economically viable for us to try and solve and there are hundreds thousands low thousands but low thousands and again it's the same thing you just need data you need compute
and you need to ask good questions like I said 10 good decisions per day either hypotheses or actual product decisions and you can solve these problems freaking amazing Cliff it's been so great to have you on the show I much prefer it when it's a discussion it's been an amazing discussion so thank you so much for putting up with me my pleasure thank you so much for your attention thank you so much for your attention thank you so much for your attention thank you so much for your attention thank you so much for your attention thank you so much for your attention thank you so much for your attention thank you so much for your attention thank you so much for your attention
thank you so much for your attention thank you so much for your attention thank you so much for your attention thank you so much for your attention thank you so much for your attention thank you so much for your attention thank you so much for your attention thank you so much for your attention thank you so much for your attention thank you so much for your attention thank you so much for your attention Thank you.