SPEAKER_01
MATTHEW STANISCHEVSKY MATTHEW STANISCHEVSKY MATTHEW STANISCHEVSKY co-founded Eleven Labs in 2022 and has since scaled it to the $11 billion leader in AI audio. He's credited with capturing the human-ness of speech through realistic emotional inflection, and they're now expanding into everything from agentic workflows to music.
SPEAKER_00
Nice to have done this.
SPEAKER_01
Thanks for having me. Maybe a little place to start is, describe to me how an audio model works. If we were Carpathi-style looking to build a toy one from scratch, how does it work? [SPEAKER_00] In the early days, you try to replicate it [SPEAKER_00] exactly like you would replicate it with the human body.
SPEAKER_00
So you would try to completely try to reproduce a machine, analog machine, that would create a vocal tract effectively. Then that progressed into trying to create effectively a digital signal for speech. Bell Labs was one of the first to try to create a structured set of signals that would represent the speech. And that is a first precursor to what we would do today. Then you would try to stitch in phonemes, effectively different sounds of how we would speak as humans, and then try to concatenate them together. It's another important part in that equation where you would, based on the most probabilistic approach
SPEAKER_00
of the next word, you would effectively try to bring the phonemes and bring them together. And then down to the modern history, where now we effectively do similar neural nets in other domains. So you predict the next sound based on, of course, the context of the previous sounds. If it's a streaming speech, if it's, let's say, a context of audio, you will use combination of predicting of the phonemes, but you also use the contextual text element of that work. And here, credit to my co-founder, Piotr, who effectively came with that new idea of how you can now create voice models, which are both reliable, high-quality, quick,
SPEAKER_00
where you would bring a lot of the ideas from transformer models, from diffusion models into the speech space. Yes. So that prediction of the next token in the phonemes space wasn't something that was possible. You might always talk briefly about this, of how you kind of operate on the text, on the waveform space. There's also a mel spectrogram space. So usually you do text, mel spectrogram, waveform. Sorry, what's the spectrogram space? It's a visual representation of how the speech sounds, across pitch, across energy, and then you transform that into waveform. Got it. So when WaveNet came along and Tacotron models, they would effectively use text to mel spectrograms,
SPEAKER_00
so that visual representation, and then how you decode and encode that into the waveform to bring it across. And Piotr figured out how to abstract some of those steps and decode and encode them much better. So that predicting of the next phoneme was one of the big pieces. And second big piece was, how do you bring that context into the equation? So what I mean by context is, if a voice actor was reading a textual copy, you would know that this is a dialect sequence, I need to produce a dialect. If it's a happy sentence, I might need to pronounce it as a happy sentence. What happens before and after comes into the equation, and you need to bring that across.
SPEAKER_00
And then there's a last big piece. So voice model has the sound of how you intonate the given fragment. But the second big part is the voice itself, the characteristics of accents, of style, of prosody across that voice. So when you actually try to vocalize something, when you create that voice model, you turn text into audio, you need the text, you also need the voice reference of how you want it to be spoken. So here is the second big innovation. Part of context is how you decode and encode those features. So when Bell Labs came with their initial representation of speech, the big piece there was you would have effectively hard-coded parameters for that speech.
SPEAKER_00
With 11Labs models... [SPEAKER_01] Hard-coded parameters for enthusiastic speaker, British accents. Exactly, exactly.
SPEAKER_01
[SPEAKER_00] That kind of stuff, the set of pitch elements that you can select,
SPEAKER_00
set of energies, spectrograms you can select from. And in our approach, effectively, you would give the model open-ended ability to select what those parameters should be. So it's not going to be British, Polish, Spanish, English speaker, but the model will deduce them themselves. The same for other set of parameters that are not hard-coded, whether it's the enthusiasm, whether it's the sadness, etc. [SPEAKER_01] You're saying Britishness is an emergent property... [SPEAKER_01] Exactly.
SPEAKER_01
...in your voice models. Exactly. [SPEAKER_00] But okay. [SPEAKER_00] Yeah, and those two big parts,
SPEAKER_00
so it's encoding and decoding of how you create the voice. Yep. Super hard problem before, and figured out too. How you then construct that in the sentence, how you get the context across, so you can predict the next phonemes. Yes. So how you bring them together in a reliable and stable way, while doing it quick. And these were the two first big innovations in the voice models that continue to today. [SPEAKER_01] But okay. [SPEAKER_01] So if LLMs reason about text and words and parts, tokens,
SPEAKER_01
as the way they think about the world, what is the equivalent of a token in the voice model? You mentioned phonemes a bunch. [SPEAKER_00] How you then construct that in the sentence, how you get the context across, so you can predict the next phonemes. [SPEAKER_00] Yes. [SPEAKER_00] So, how you bring them together in a reliable and stable way, while doing it quick.
SPEAKER_00
And these were the two first big innovations in the voice models that continue to today. [SPEAKER_01] But, okay. So, if LLMs reason about text and words and parts, tokens, as the way they think about the world, what is the equivalent of a token in the voice model? [SPEAKER_01] You mentioned phonemes a bunch. What is that representation? So, we do store the voice embedding effectively for the speaker. So, you need that reference when you produce and create the speech. Yes. Of course, in the input to the voice model, you still get the text, and you bring the speaker encoding.
SPEAKER_00
And then when you produce speech, you do operate on the waveform or effectively on the phoneme level of that speech. And then when we kind of go the opposite, so of course, when we— [SPEAKER_01] Sorry, what is a phoneme?
SPEAKER_01
[SPEAKER_00] Fill in my understanding.
SPEAKER_00
It's a syllable deconstructed even to smaller elements. [SPEAKER_01] Okay. And these are effectively the human sounds you can produce. [SPEAKER_01] Got it. So, this would be the most close to that representation. But of course, in our models now, it's going to be a combination of not only operating on phoneme level, you also operate on the text level. You operate in both in sync, because when you are predicting the context, you need to understand how that sentence will get constructed. And especially if it's more of a streaming real-time use case and a voice agent setting, you need both parts to work across.
SPEAKER_00
So, it's similar to how you would operate on the token level on the text side. We operate on the token level on the audio side. [SPEAKER_01] It feels like a big part of the magic of Eleven was your voices were much more human sounding. How did you accomplish that? So, give you a quick synopsis of what we think about the models on the text to speech side today. In any model, you need the architecture, you need compute, you need data. So, architecture innovations were one thing. The data part was the second big thing.
SPEAKER_01
[SPEAKER_00] With audio, you will have a lot of audio data available, but frequently you will not have it annotated in the right way. [SPEAKER_00] You won't have which speaker is speaking when. Some of what is annotated, but the how isn't.
SPEAKER_00
So, as we are speaking now, what are the emotions that we use? What are the actions that we use? So, we would invest a lot internally on effectively creating our own data labelers, our own team, to be able to create those data sets that will be better. And that was a combination of semi-automatic techniques and manual techniques. And actually, a lot of the models that we did afterwards actually spun out from a lot of that research too. So, speech to text model initially was a model we did for ourselves because the models on the market just weren't good to annotate that.
SPEAKER_00
And then, another brilliant researcher on our team was able to construct it so we could spin it out as a model that we brought to the customers. [SPEAKER_01] So, you've just been doing useful stuff in voice, and that has emerged with a whole bunch of products that you might not have expected, because you find you're building useful stuff. Exactly, exactly. And that combination of data of being able to do it automatically and creating a team that's coached on voice on how to describe it. [SPEAKER_01] Yes. Because most of the labelers out there just aren't as well versed on understanding the audio and voice, helped us a lot to bring that back.
SPEAKER_00
And then, of course, deploying those models in production, seeing how customers interact with them, having them annotate all of the data helped us refine those models over time. A very interesting thing on the side. So, we spoke about the speech representation. The first guy who created the speech representation is a guy called Von Kempelen. So, he created an analog machine that would represent effectively a human vocal tract and try to produce that sound. He had to spend decades on that and that started producing vowels.
SPEAKER_01
[SPEAKER_00] But that's the same person that created a chess machine, the first viral chess machine, that would simulate playing chess. Is this the Mechanical Turk?
SPEAKER_00
It was called Turk. [SPEAKER_01] Yeah, yeah. But exactly. The crazy thing behind it, it was operated by a human. [SPEAKER_01] Yeah, yeah, yeah. And it was all a hoax. [SPEAKER_01] Yes, yes. And that's where the Mechanical Turk comes from, which actually we use in that data labeling production to make that work there. [SPEAKER_01] Yeah, yeah. And sorry, we kind of jumped right in. But if you describe the Eleven business today, people think of you as the speech company. How should they actually think of your business to the extent you can describe the big areas? Text to speech, speech to text, voice agents—just break down the business for us.
SPEAKER_00
Cool. So, in a nutshell, I'll describe Eleven Labs as a research and product deployment company. We built foundational audio and voice models and then build a platform for businesses to transform how they communicate with their customers, with their employees. And that applies through AI agents from customer support, sales, hiring, training, all the way through to marketing and storytelling for our creative tools. And in that set, we've created all types of foundational audio models. So, text to speech models for producing speech, speech to text models that work over 100 languages and happily beat others on benchmarks.
SPEAKER_00
All the way through to conversational models of how you loop them together, to music, to other domains of audio. And then, of course, beyond the models, when you actually bring them to production, that's where the second level of the platform comes in, where that meets the businesses on the specific use case.
SPEAKER_01
[SPEAKER_00] So, on the agent specific example, it would be how you now connect those models to the knowledge base, to telephony, to the integrations that you need to perform the actions.
SPEAKER_00
How you evaluate and monitor the agent so it behaves in the right way, how you build the right safeguards. On the creative side, on the marketing side, it's how do you create a good ad so you can create a good video voiceover for one of the campaigns. All the way through to conversational models of how you loop them together, to music, to other domains of audio. And then, of course, beyond the models, when you actually bring them to production, that's where the second level of the platform comes in, where that meets the businesses on the specific use case.
SPEAKER_00
So, on the agent specific example, it would be how you now connect those models to the knowledge base, to telephony, to the integrations that you need to perform the actions. How you evaluate and monitor the agent that behaves in the right way, how you build the right safeguards.
SPEAKER_01
[SPEAKER_00] On the creative side, on the marketing side, it's how do you create a good ad so you can create a good video voiceover for one of the campaigns. [SPEAKER_00] How you create an article that's narrated with a specific voice that represents the brand in a good way. [SPEAKER_00] So, that's where we combine the models and understanding of the customers we work with into one policy platform. Every platform company has this question about how far they go into applications. So, how do you think about where you go horizontal and power the whole ecosystem versus where you develop applications?
SPEAKER_00
[SPEAKER_01] Because you can imagine there being a whole ecosystem of closed captioning tools that grow up that, again, are built on the 11 Labs tech. [SPEAKER_01] It's not necessarily a space that you would have to go after yourself. I think the big difference between your question is today we see ourselves as a platform where if you're building a horizontal use case in your business, a great place to come. If you have a lot of domain specificity, that's where I see a lot of application companies forming over time. That's specifically not the spaces we will go into. [SPEAKER_01] And I think it also is interesting when the tech is moving as quickly as it is here.
SPEAKER_00
[SPEAKER_01] It's one thing with SaaS where you get these vertical specific providers. [SPEAKER_01] But I would imagine one of the biggest risks for you guys in being intermediated is if there is a closed captioning service that is on a two versions old version of 11 Labs and hasn't upgraded. [SPEAKER_01] That's a problem because you want people to be using the latest and greatest model that you've developed. [SPEAKER_01] And you'll be deploying new capabilities every week. [SPEAKER_01] And I presume that's part of your thinking is that when it's moving that quickly, you need to go direct in a lot of cases.
SPEAKER_00
That's right. In the closed captioning, we already know that our services is going to be able to tackle 99.9% of the cases that customers have. And then there's the added benefit of we work with healthcare customers where we will create custom models for those customers where we'll get that transcription perfectly.
SPEAKER_01
[SPEAKER_00] The context is a tricky thing in closed captions where we talk a lot about technical stuff on this. Yeah, for sure. And that's where you need effectively a dictionary of words that you detect beforehand, which as we work with the businesses, we know we need to embed in that creation process. [SPEAKER_00] We're talking a bit about products here. And one thing I note is that LLMs are amazing. And you have the usage stats of ChatGPT and Gemini and all the popular LLMs where they're working and people use them a ton. [SPEAKER_00] It feels like there's a big product overhang when it comes to voice where the leading edge voice models are incredibly capable.
SPEAKER_00
[SPEAKER_01] And yet I was driving home the other day and I needed to read a PDF while I was driving. And so I said I'll just have my phone read the PDF to me. [SPEAKER_01] And you can try and hack it with iOS screen reader, but it doesn't really work with the scrolling. [SPEAKER_01] And then in theory, you can upload to Gemini, but you're trying to get it to not summarize it and it actually hung when I tried to press the read this to me button.
SPEAKER_01
And so there was no way I could get my phone to read me something, which seemed like a fairly basic feature. And all cars advertise voice control. And yes, it sucks. If you want to input something to the navigation, no car has a good version of that. Maybe Tesla does. And so why does it seem like with LLMs and cloud code and everything, we are using all the capabilities of the intelligence, whereas with voice, we're living 10 years ago somehow? Well, I'm thinking I agree with the premise that we are 10 years behind. In the lived experience of people day to day, they're using Siri's transcription, which has gotten better. Yeah, yeah.
SPEAKER_00
[SPEAKER_01] And it's still way behind the leading edge. [SPEAKER_01] Yeah. There is definitely a piece like we, I think the technology in many of those cases is ready, there's a deployment gap to what you are saying. [SPEAKER_01] Yeah, exactly.
SPEAKER_01
It's like automotive or some of the big companies are not adopting that quickly enough or bringing that into the production.
SPEAKER_00
[SPEAKER_01] But there are plenty of different problems that you need to fix along the way. I mean, the quality of voice models for them to actually sound good. [SPEAKER_01] This is only a last three years thing.
SPEAKER_01
[SPEAKER_00] Yeah, that's three years. It's a three years thing. [SPEAKER_00] I think so, that's three years for the first voice model that can narrate. Yes. [SPEAKER_00] Text async. [SPEAKER_00] Two years ago, you can start seeing the real time version of that. [SPEAKER_00] And not really, I think that real big was a year ago. [SPEAKER_00] Okay. [SPEAKER_00] Where you can start seeing that in production. [SPEAKER_00] And then I think over 2025, the big piece that hasn't been possible is how you connect now the real time voice interaction with something which I think you are referring to.
SPEAKER_01
[SPEAKER_00] It has context of what you want to do. What is the material that you want to read? How does it connect to set up your preferences from the past? [SPEAKER_00] Yes. [SPEAKER_00] And gets that across. [SPEAKER_00] I think that's only recently became possible and where we've seen the big adoption across the enterprises leading on the technical side. [SPEAKER_00] I think this year it should be in the automotive side too or some of the applications.
SPEAKER_00
[SPEAKER_01] Okay. So you think we'll start seeing great voice models in cars this year?
SPEAKER_01
[SPEAKER_00] For sure.
SPEAKER_00
This year for their own cloud use cases, in car.
SPEAKER_01
[SPEAKER_00] So without connectivity, not yet.
SPEAKER_00
2026 is deployment of course gap of how you bring that into the gaps. But I think the next two, three years. [SPEAKER_01] How about the PDF reading use case? That should work. I think, yeah. [SPEAKER_01] Well, how should I have done it? We've, so back in the day, I'll preempt this with a story to care 11 reader. But we had this problem. We have so many audio book authors come into 11 Labs. So 2023 released first software. We had a lot of creators and then a lot of audio book authors or book authors that tried to, couldn't afford professional narration and wanted to create an audio book. How about the PDF reading use case? That should work. I think, yeah.
SPEAKER_00
Yeah. Well, how should I have done it? We had this problem. Back in the day, I'll preempt this with a story to cure 11 reader. We had so many audio book authors come into 11 apps. In 2023 we released first software. We had a lot of creators and then a lot of audio book authors or book authors that tried to, couldn't afford professional narration and wanted to create an audio book. However, none of the companies accepted AI audio books. And you can't sell an audio book on Audible or something. Exactly. So Audible would block AI content. [SPEAKER_01] Yeah. [SPEAKER_01] So we had no choice. We need to create an avenue for them to bring them.
SPEAKER_00
[SPEAKER_01] Because there was no distribution. [SPEAKER_01] Exactly. [SPEAKER_01] For AI audio books. Exactly. So we created an 11 reader. Okay.
SPEAKER_01
[SPEAKER_00] And that came with functionality where you can upload your PDF. [SPEAKER_00] You can upload your text and have it read out loud with a number of incredible voices.
SPEAKER_00
So whether it's Sir Michael Caine all the way through to a state and working together with Sir Richard Feynman where you can have that.
SPEAKER_01
[SPEAKER_00] And so this is you're working with the Sir Michael Caines of the world? [SPEAKER_00] Exactly. [SPEAKER_00] And then you can actually read it out loud.
SPEAKER_00
And that works extremely well. So that works. Yeah. Now how can you do it? [SPEAKER_01] I think. Actually I do want everything read to me by Michael Caine. It's a great voice. Yeah. Shouldn't you guys have a consumer app where I can just do the common voice things? I want to be able to have an 11 app on my phone. And then if I upload a PDF to it, it can do the common things that I would want such as have it read it to me. Yeah. That's exactly your own reader. So that works.
SPEAKER_01
[SPEAKER_00] Okay.
SPEAKER_00
The phone makers allow third party keyboards. Do you think they allow third party transcription engines? Will they, do you think? The phone makers you said, right? [SPEAKER_00] Yeah. [SPEAKER_01] Apple and Google. Yeah. [SPEAKER_01] They are not. [SPEAKER_01] So they are the OS makers. [SPEAKER_01] Yeah. Not all of them. Android. With Android you can work through it. [SPEAKER_01] There are variations like Nothing, the Tag, and others. [SPEAKER_01] But again, if you had a popular 11 app that allowed for transcription, people would use it a bunch.
SPEAKER_01
[SPEAKER_00] And maybe eventually Apple would say, oh, we should allow third party transcription engines, if that's what people want.
SPEAKER_00
I mean, there seems like they might be going in that direction, right? Yeah.
SPEAKER_01
[SPEAKER_00] And recently they announced that we'll open up the LLM ecosystem. Hopefully they will do the same with voice ecosystem, which is similar. [SPEAKER_00] But again, I think it's rational to do when it's moving so quickly.
SPEAKER_00
Yeah. The voice assistant paradigm is one of the oldest UI paradigms in computing. Open the pod bay doors, Hal, from 1969. Yeah.
SPEAKER_01
Yeah. [SPEAKER_00] I will claim it's not working yet. So Siri doesn't have the intelligence.
SPEAKER_00
[SPEAKER_01] And on Gemini and ChatGPT and those apps, I want to use the voice mode, but I don't know about you, it just doesn't work. [SPEAKER_01] So sometimes I'll be using my phone and I'll use the iOS keyboard transcription to type in the field and then say a bunch of stuff and then send it off. [SPEAKER_01] But this suggests to me that consumers really want voice mode that works. And yet it's just not working for the major LLM apps or for anyone. [SPEAKER_00] Why isn't it working yet? It is pretty hard to do because you want two things.
SPEAKER_01
[SPEAKER_00] You want to be able to say things that you want, but you want sometimes for it to execute it, sometimes to wait for you to finish and add something in a sentence.
SPEAKER_00
[SPEAKER_01] Sometimes you want it to be interactive. [SPEAKER_01] So it asks you questions back to clarify and get some additional detail. [SPEAKER_01] And all of that is actually pretty hard. [SPEAKER_01] That's where the magical, ideal version of a voice agent for us comes through where you need the speech to text element. [SPEAKER_00] You need the transcription side. [SPEAKER_00] You need then the turn taking mechanism. [SPEAKER_01] So when do you finish a sentence?
SPEAKER_01
When is it likely based on silence, based on context? And then sometimes you want it to speak back and clarify or at least give you the text back to clarify and then maybe execute a set of instructions. So that problem is still very hard research. So I agree with the claim that this orchestration side has not passed a true conversational agent Turing test where it behaves as you would expect from another person.
SPEAKER_00
[SPEAKER_01] Yeah. [SPEAKER_01] Maybe that's the simpler way of saying what I'm saying is that we have passed the Turing test with text LLMs a long time ago. And we're actually nowhere near that on voice LLMs. So it's interesting how that's a final frontier.
SPEAKER_01
[SPEAKER_00] Yeah. [SPEAKER_00] I feel like it's going to work in specific domains, like in customer support calls. [SPEAKER_00] Yes. [SPEAKER_00] Passes the voice Turing test.
SPEAKER_00
Works well. In another spectrum of that, an interactive gaming experience, a truly interactive as you would have with another human in that game.
SPEAKER_01
[SPEAKER_00] It's hard and further out there. [SPEAKER_00] We haven't passed it yet there. [SPEAKER_00] Yes. [SPEAKER_00] Yes. [SPEAKER_00] Yeah. [SPEAKER_00] But I think that's a combination of even a simpler version of that. Sometimes you might give a response immediately back.
SPEAKER_00
[SPEAKER_01] Sometimes you need a tool call to get additional information from the database. [SPEAKER_01] So how you orchestrate that. [SPEAKER_01] So that's probably the most common thing we see as we work with some of the companies out there. You want those systems to orchestrate extremely well, where if it's a conversational use case, pretty simple. You can route the agent to speak with, but if you need to authenticate, if you need to pull additional information from the database, what do you do? Yes. Yes. Yeah. But I think that's a combination of even a simpler version of within that. [SPEAKER_01] Sometimes you might give a response immediately back.
SPEAKER_00
[SPEAKER_01] Sometimes you need a tool call to get additional information from the database. [SPEAKER_01] So how you orchestrate that. [SPEAKER_01] That's probably the most common thing we see as we work with some of the companies out there is you want those systems to orchestrate extremely well, where if it's a conversational use case, pretty simple. You can route the agent to speak with, but if you need to authenticate, if you need to pull additional information from the database, what do you do? [SPEAKER_01] How do you handle that graciously?
SPEAKER_01
[SPEAKER_00] That's really true. [SPEAKER_00] And yeah. [SPEAKER_00] And to that extent, I agree. [SPEAKER_00] That's just getting there.
SPEAKER_00
Yes.
SPEAKER_01
And we'll hopefully see that.
SPEAKER_00
[SPEAKER_01] Our goal is to pass the voice Turing test in all those cases, or the Turing test for all conversational agents outside of voice too. And I hope we will all be there in the next year or so. For subscription businesses, a lot of revenue is lost in that last few seconds before checkout. Someone has to get up and find their wallet, or they mistyped their card number, or they hit an error. They just give up and you lose the sale.
SPEAKER_01
[SPEAKER_00] For a company like Eleven Labs, adding hundreds of thousands of subscribers, even a tiny bit of friction like that would really add up. [SPEAKER_00] But that's why Eleven Labs uses Link from Stripe.
SPEAKER_00
Customers save their details once, and then they can check out in seconds across more than a million businesses with saved credentials.
SPEAKER_00
So if you want a faster checkout for your customers, you should turn on Link from Stripe. Are you guys working on personalized voice transcription? I think that's a good question, where it feels like part of the way we're making it hard for ourselves is when I speak to Siri, I have a bit of an accent. And so it sometimes has a hard time understanding me, but my accent doesn't change. And so it could just get good at listening to John. But my understanding is it's not. It's just running the global voice recognition model.
SPEAKER_00
And I'm guessing it's the same for Eleven Labs where you're running the global voice recognition model. [SPEAKER_01] But again, you have an accent. [SPEAKER_01] And so if someone's understanding, if you walked up to someone in a coffee shop and said two words, they might have a hard time understanding it because they're not putting it through their Matty Polish accent filter. [SPEAKER_01] And so where is this going with actually interpreting the person that you know to exist on the other side? [SPEAKER_01] Yeah, I have a very tricky one to detect. [SPEAKER_01] So my voice is frequently used in the test. [SPEAKER_01] Ah, you're part of the test suite. [SPEAKER_01] Yeah.
SPEAKER_01
For text to speech, for speech to text, for everything. Yeah, yeah, yeah. It's pretty tricky. But again, trying to parse your voice in a global model is just making life hard. It's having a Matty specific model. Yeah. So on the speech to text, transcription, exactly. The big part now that we are bringing in is you have two parts. [SPEAKER_01] One, effectively a person or voice specific detection, which is true for the accent side, but it's also true for a crowded room. So that's where we have an incredible research team that's able to continually do both keeping accuracy high, but also adding things like speaker detection, of course, noise reduction.
SPEAKER_01
But then the second part is also keyword detection. [SPEAKER_00] So there are specific words that you would want to say in those settings that you want to effectively monitor for. [SPEAKER_00] So we spoke about, let's say I'm going to the coffee shop and order things. [SPEAKER_00] The set of actions I would like the coffee shop would expect me to do is pretty limited. [SPEAKER_00] There's information theory. [SPEAKER_00] They can just listen out for the coffee words. [SPEAKER_00] Exactly. And then try to match it to the closest proximity.
SPEAKER_00
[SPEAKER_01] Yeah. [SPEAKER_01] So both things will help. And in a setup where you have my voice, perfect. Yeah. You can decode it and code it on that. Yeah.
SPEAKER_01
[SPEAKER_00] If you don't have my voice or even if you want to double amplify it, we already support effectively a keyword detection. [SPEAKER_00] Which is useful for real time setting and async setting. [SPEAKER_00] So back to cheeky pine transcription.
SPEAKER_00
Yeah. You could effectively pre-generate that from the previous podcast and look for a set of words that you would use traditionally in that. And so how hard, okay, so you do the keyword detection already, but how hard are the, I want to get superhuman transcription performance by feeding it an hour of Matty audio before it listens to Matty and then it should be able to do a much better job transcribing. [SPEAKER_01] Is that just a really hard research problem? No, solvable. Solvable. We think we can roll it out in one of the next versions, which is hopefully in the next months. Oh, so you think this year you're doing person-specific transcription? Yeah, for sure.
SPEAKER_00
Person-specific transcription.
SPEAKER_01
[SPEAKER_00] We can already diarize speakers extremely well.
SPEAKER_00
So if we are speaking, we can of course dissimilate who is speaking when. Yes. Which is on the transcription side, operative accuracy, diarization is one of the harder problems. Yes. And we do that extremely well. And now it's going to be effectively what you're saying, fine tuning based on the speaker that I want to listen to. [SPEAKER_01] Yes. [SPEAKER_01] Which we know will be important. [SPEAKER_01] In healthcare setup, such an important part. You're in an operating room, you're a doctor, you want to say a command, then you want to really be able to listen to that one person specific piece. Yes.
SPEAKER_00
[SPEAKER_01] You have a hardware device at home, let's say it's a pilot, that helps you control the TV. Here too, you will want that to listen to you versus the family roaming around.
SPEAKER_01
[SPEAKER_00] Or maybe you want it to everyone. [SPEAKER_00] So you could decide that, but in many cases you want to be able to specify that. [SPEAKER_00] Hmm.
SPEAKER_00
Okay, that's really exciting. It's great because there's still so many unsolved research problems.
SPEAKER_01
[SPEAKER_00] There's just breakthrough after breakthrough coming in the domain of voice models.
SPEAKER_00
How about on the flip side when it comes to speech generation, in the Zoom touch up my appearance feature? I've always thought about that in the context of voice, where should you offer a deaccenting filter for voices? Or maybe you want it to everyone. So you could decide that, but in many cases you want to be able to specify that. Hmm. Okay, that's really exciting. It's great because there's still so many unsolved research problems. There's just breakthrough after breakthrough coming in the domain of voice models. How about on the flip side when it comes to speech generation, in the Zoom touch up my appearance feature?
SPEAKER_00
I've always thought about that in the context of voice, where should you offer a deaccenting filter for voices? Or even this one podcast that I like to listen to, but the voice is a little mumbly. And I always thought they should put it through a de-mumbling filter, just to make the enunciation a little better. [SPEAKER_01] But all these things, again, Photoshopping an image, there's no reason that the—have you thought about voice to voice, rather than voice to text or text to voice?
SPEAKER_01
[SPEAKER_00] Yeah, so there are two big parts.
SPEAKER_00
One on the speech generation side, similar, so many innovations still there. There's a wider piece, and we released the vFree model that kind of we're solving that for the first time. It's, can you control speech? So you can have the text to speech, you generate something that sounds emotionally great. [SPEAKER_01] Previously, until the end of last year, effectively, it would rely on model to decide what's the best performance.
SPEAKER_01
You could regenerate it, but that ultimately model decides the best performance. [SPEAKER_00] So that's where the controllability came in, where we can finally give it cues of, say it in a slower way, or change how you deliver the dramatic pause, or any cues that you give. And to be able to do that, you need the architectural changes and the data that we kind of created over time,
[SPEAKER_01] where you annotated what was said and how it was said, so you can actually train the model to do that. [SPEAKER_01] So today, finally, you can have both speech generation or entire voice agent experience with what we call expressive mode, where the agent knows the emotions on the other side. So if the person is stressed, it can react and be reassuring, and that's generating a limb response on the reassuring side and response in that set of emotions too. And that breakthrough was super hard to do. And that, of course, stretches to a lot of what you said.
SPEAKER_01
[SPEAKER_01] It could be some version of speech enhancement, either real-time or in a post setup to change how that's delivered. And that's relatively recent innovation, and we know it can still be so much better. [SPEAKER_00] The edge cases of how you want to describe it is pretty large. [SPEAKER_00] So that's one. And then the second part of the question, which is a huge question, is speech-to-speech models. [SPEAKER_00] So as you said, our approach, as you think about voice agent, conversational agent, is effectively a cascaded approach.
SPEAKER_00
They use transcriptional speech-to-text, LLAM, text-to-speech, and orchestrates all of that together. And then you have speech-to-speech, which goes directly from speech, and there's a speech response on the other side. When I say speech-to-speech, is that the idea that it doesn't go through text as an encoding in the intermediate set? Oh, interesting. For performance reasons, for accuracy reasons? You usually do it for latency. Okay. For latency. So it's faster to run a model that does not have to transcribe and then generate. Exactly. It's quicker, but on the flip side, you lose reliability. Yes. You lose all visibility into the parts of the pipeline.
SPEAKER_00
And emotionality, we think you can deliver both on both sides extremely well, and maybe you can make it more controllable, too. So today, we are optimizing heavily on a cascaded approach. I'm sorry, a cascaded approach is? Is the speech-to-text, going through the text layer. Oh, okay. Going through the text layer. And as we work with all of the businesses and enterprises, they will need that visibility into what happens. They will want to execute certain tasks on top of that. [SPEAKER_01] They want good visibility into each of the steps and great accuracy of all the models.
SPEAKER_00
But beyond that, they can abstract away what's the LM layer, what's the intelligence layer. The integrations are easier in that system. So that's where we are betting a lot of the research work of how you can make that great. And we think we can make that great.
SPEAKER_01
And speech-to-speech, as you think about maybe more of a companion version of the applications, that's where that will flourish because maybe the hallucinations aren't as important, but the latency is a little bit more.
SPEAKER_00
[SPEAKER_01] And maybe hallucinations are even a feature. [SPEAKER_00] And maybe in the future, just to finish that part, you will have some version of combination of the models. That for low complexity, easy models, you will have speech-to-speech. And for higher complexity, you will have the cascaded.
SPEAKER_01
[SPEAKER_00] Okay. So I was going to ask about this. [SPEAKER_00] You know there is research on how the invention of writing changed humans' brains and just changed the neural pathways in ways beyond the actual written language. Do you observe that speech-to-speech models think differently than cascaded models?
SPEAKER_00
It sounds to me they're dumber. They are definitely dumber. You need smaller model, you cannot... But that's interesting, right? That forcing models to reason about text. I mean, I know they just have much more in there as well, but they're smarter.
SPEAKER_01
[SPEAKER_00] Yeah.
SPEAKER_00
But it's that if you are going speech-to-speech, usually you will use smaller models, so it's still quick. Yeah, yeah. I see. So it's also just a model size thing. Yeah, yeah. Okay, but are there interesting differences beyond correlates like size? What I can say is slightly different to your question. The people interacting through voice and the performance we see for how they interact with the business changes just by nature of interacting with voice. [SPEAKER_01] A good example, you can contact Eleven Labs and register for your interest, you go through the form.
SPEAKER_00
[SPEAKER_01] And at the end of that, we supplemented that instead of going through the form process, you can speak with our agent and leave more details. [SPEAKER_01] And what happened are two things. [SPEAKER_01] One, people were actually much more keen to leave the forms through speaking with the agent. [SPEAKER_01] So we would go through the form a lot easier. But second, they would be a lot more open-ended in terms of what the use cases are. So they would start giving us information about the wider set of use cases, the complexity of the use case. So the writing out was tedious and tricky. [SPEAKER_01] It's an open-ended adventure game.
SPEAKER_01
Open-ended, you could ask follow-up questions, you could clarify. And at the end of that, we supplemented that instead of going through the form process, you can speak with our agent and leave more details. And what happened are two things. One, people were actually much more keen to leave the forms through speaking with the agent. So we would go through the form a lot easier.
SPEAKER_00
But second, they would be a lot more open-ended in terms of what the use cases are. So they would start giving us information about the wider set of use cases, the complexity of the use case. So the writing out was tedious and tricky.
SPEAKER_01
It's an open-ended adventure game. Open-ended, you could ask follow-up questions, you could clarify. But people were just more at ease and could trust the system while doing that, that it's working.
SPEAKER_00
And that helped us a lot. And then third, which maybe is more of a technological barrier, it also works across all languages. So now we have leads from all parts of the world coming in and leaving their details. So we did that use case and now we have a few different companies building their SDR versions of that too.
SPEAKER_01
Yes. To help them capture the leads coming in from banks all the way to one of the automotive companies that leaves that,
SPEAKER_00
where people are just more keen to speak through voice.
SPEAKER_01
So I want to ask about this kind of second order effect.
SPEAKER_00
You've talked in the past about how growing up in Poland, the dubbing of TV shows was cheap and so they would only have one voice actor for a TV show. So no matter all the parts, male and female, they're like, "I love you." "I love you too." You know, there's one voice actor doing all of them. And now, thanks to better voice models, you'll be able to have really good voices, AI generated for all the dubbing. [SPEAKER_00] Because it's not like it's taking jobs from great dubbing that was happening previously. It was awful dubbing happening in Poland previously. So that's one example of the second order effects.
SPEAKER_00
What are the other second order effects you're seeing of ubiquitous good text to speech, speech to text across a broad array of languages, because whatever about in English, this didn't exist in Polish or Irish or pick your language. One, breaking down the language barrier. The inspiration came from the movie side. But it also applies in any communication setup. Could in the future, could I travel to another country and speak Polish or speak English, and that language isn't being understood in the local native language? Like from Hitchhiker's Guide to the Galaxy, this version of the Babel fish. Exactly. That you can actually understand the world.
SPEAKER_00
And voice, of course, will be an interaction layer. But similarly, all of us will have our own extension and voice agents that can help on our behalf. [SPEAKER_01] And there are very clear and great examples of that of people that lost their voice and can get it back. [SPEAKER_01] We see that everywhere, whether that's people that lost it due to ALS or throat cancer that can get it back. [SPEAKER_01] Just recently, there was an example of a patient that had Neuralink. [SPEAKER_01] I worked with them to bring the voice so that person could speak with their own voice back with the family around. [SPEAKER_01] We worked with a lady that lost her voice before she got married.
SPEAKER_00
[SPEAKER_01] And then finally technology became possible. [SPEAKER_01] We were able to recreate that voice.
SPEAKER_01
And for the first time, she could replicate the marriage ceremony and speak the vows together, which was such a heartfelt moment. Probably the most important from all the work that we do. When you guys talk about voice agents, is a voice agent just the idea that you're going to be able to do that? You have some long running or persistent agent that is going out and interacting with the world through voice. And so customer service would be one example of it. [SPEAKER_00] In the other direction, your AI going and making you a restaurant reservation and actually calling up the restaurant. [SPEAKER_00] Is that how I should think about voice agents? [SPEAKER_00] That's right.
SPEAKER_01
[SPEAKER_00] It's exactly what it is—the reactive side of being able to interact with the customer or the proactive, to call it that. [SPEAKER_00] We recently had a very interesting one, topical, because it was a Guinness-related one, where there was a developer developing a Gindex effectively. [SPEAKER_00] Oh, I saw that. They were calling all the pubs in Ireland to check the price of a pint. [SPEAKER_00] Yeah, you could ask that or report information. [SPEAKER_00] The Gindex was built with 11 labs technology.
SPEAKER_00
It was built with 11 labs, too. So people could actually do both sides, could proactively reach out, reactively reach out, all captured through voice. And then 3,000 different entities could report their prices and get that across. Have you, by the way, hooked up your Open AI API to 11 labs? Is the Open AI 11 labs combo something that a lot of people at 11 are doing? So, as you know, the Open AI API will look for the most popular tools frequently where it tries to hook up. So 11 labs is one of the recommended ones. It's the top option for voice.
SPEAKER_00
[SPEAKER_01] Can you tell me a bit about the business of voice models, where I think people have an intuition around big LLMs, where there are these very expensive training runs. [SPEAKER_01] And yes, they depreciate quickly, but there's so much usage that all of the models trained to date have paid off their training runs and then some. [SPEAKER_01] And then there's this ever larger capex going into, a lot of it is inference these days, but also training. [SPEAKER_01] And so people have some intuitions from the LLM world. [SPEAKER_01] I'm curious how I should think about voice. For one, how expensive is training the voice models? Is the expense in the researchers?
SPEAKER_00
Is the expense in the training runs? And the economics is presumably straightforward, but it's just per usage. But yeah, talk us through the business.
SPEAKER_01
Yeah, definitely cheaper than the LLM and image video models. [SPEAKER_00] So you can use smaller models. Yeah. [SPEAKER_00] Okay, so the models are smaller. [SPEAKER_00] Smaller. [SPEAKER_00] What's the parameter count for a leading edge voice model?
SPEAKER_00
Few billion to low tens of billion parameter models. Yeah. [SPEAKER_01] And for context, I think CPUs eventually moved away from gigahertz as the metric as they move to more cores. I think we've mostly moved away from just raw parameter count, but I think the leading edge LLMs are in the hundreds of billions of parameters. I think the leading ones, yes.
SPEAKER_01
But of course, you have the variations that you will use at lower scale.
SPEAKER_00
[SPEAKER_01] Yeah.
SPEAKER_01
[SPEAKER_00] Okay, so the models are smaller.
SPEAKER_00
Smaller. What's the parameter count for a leading edge voice model? Few billion to low tens of billion parameter models. Yeah. [SPEAKER_01] And for context, I think CPUs moved away eventually from gigahertz as the metric as they moved to more cores.
SPEAKER_01
[SPEAKER_00] I think we've mostly moved away from just raw parameter count, but I think the leading edge LLMs are in the hundreds of billions of parameters.
SPEAKER_00
I think the leading ones, yes. [SPEAKER_01] But of course, you have the variations that you will use at lower scale.
SPEAKER_01
Yeah.
SPEAKER_01
We've raised recently half a billion, that's 11 billion valuation. Makes sense. Makes sense. To continue being able to build the best models in the world. Researchers, you want the best people in the world. I think we have those people working in audio and my co-founder who is leading that work. So that's definitely a big piece of how you keep ambitious about deployment. You continue building leading models, which helps you attract more talent. And then on how we serve it. Of course, inference is correlated with how the models are used. [SPEAKER_00] And for us, we've seen incredible growth across the work.
SPEAKER_00
Mostly this is charged per input text or text to speech, it's usually per text token. If it's voice agent or transcription, then it's per minute. And we see that being the bigger part, but usually it's on a per token basis. And of course, as we work with businesses, it's an annual agreement. [SPEAKER_01] The bigger the spend, the bigger the discount to get that gross margin.
SPEAKER_01
[SPEAKER_00] The way we usually do is when we have a new model, we try to give it at cost to a lot of the customers so they can experience the best.
SPEAKER_00
It's still usually not as reliable. [SPEAKER_01] The newest thing is often the most expensive, whereas you make the newest thing the most economically attractive one.
SPEAKER_01
We try to make it attractive so the customers know it's more expensive for us than any previous generation. [SPEAKER_00] We don't like the quality is higher, so we try to keep the prices still competitive.
SPEAKER_00
I see. You subsidize it, but it's inherently more expensive for the bigger model. Exactly. Exactly. And over time, we might do some tricks to optimize it, but we want the customers to experience it. Because of research, the big thing that we've seen is the reliability of the model in the early days might not be there. And two, people don't even know what's possible with that model. So you want the widest set of distribution so people can show the world what's possible. So you can have it as a distribution mechanism, learn yourself what to improve, what to change, and then get it out there. Are the voice models just getting bigger and bigger?
SPEAKER_00
Like, will we have voice models in the hundreds of billions of parameters? Or have we found that for certain types of model architectures, there's an upper limit on the natural size. Have we found that upper limit for voice models? It feels like for specific use cases, like audiobook narration, you probably found that size. You probably don't need to stretch it too much bigger to make the quality much higher. But for certain use cases, that will probably grow. The thing is, I hesitated on the question because in a cascaded approach, you probably will not see dramatic size changes. You inherently want the models to be quick and reliable.
SPEAKER_00
You want to orchestrate them in a smart way. In a fused approach, probably that will get into tens, hundreds of billion-parameter models, because you combine the LLM side and the voice side. So that will get bigger. But on the just voice, I think it will keep being small. [SPEAKER_01] Okay.
SPEAKER_01
[SPEAKER_00] But there are certain domains where we'll see bigger models.
SPEAKER_00
That's so interesting. Yeah, yeah, yeah. [SPEAKER_01] It is amazing how there are still these various unsolved aspects and how you guys are just making technical breakthroughs and then releasing them down the product pipeline.
SPEAKER_01
[SPEAKER_00] That's a really fun stage of a company's life cycle.
SPEAKER_00
For sure. It's fun because we feel like we can do innovations on both sides. There's so much on the research side, so much on the product side. And ultimately the biggest part is how we deploy it to the customers. Where SMB will have very different dynamics than the enterprise. It's not a vendor SaaS relationship where you just give the product out there for the biggest companies. [SPEAKER_01] But you are more of a partner in their AI transformation part.
SPEAKER_01
So you want the resources to work alongside them on frequently very new use cases that were impossible, to help create and bring those voice agents to production. So that's a big shift. But the biggest focus is how we bring conversational agents out there to the businesses around the world. [SPEAKER_00] So when you say bring conversational agents, is this for customer service type use cases?
SPEAKER_00
Like what are the most popular use cases for conversational agents? Yeah, we want to be a partner for full interactions between businesses and their customers or their audience. I'm saying their audience because that will apply in support. Support is the easiest one because that's where it's most ready. But that's maybe the big difference in how we see ourselves to some of the other companies in the space. This can also apply to sales. You can have the proactive side of reaching back. You can have AI SDR versions of that. [SPEAKER_01] Yes.
SPEAKER_00
[SPEAKER_01] And then you can have all the way to the marketing use cases where we are your partner for working on even outside of the conversational agent space on how you create a great marketing campaign.
SPEAKER_01
Yes. [SPEAKER_00] And so how does this break down between, we had Dev Trainer from Intercom on here and they have Finn, their agent, and it's a thing on the website that you can go talk to. And he described a very similar phenomenon that you described, which is you start maybe thinking, oh, this will help me answer customer support queries.
SPEAKER_00
[SPEAKER_01] But it becomes a generic UI for the website where it's a box you can type in to go do things and understand things.
SPEAKER_01
[SPEAKER_00] And so why wouldn't you read the docs and design your integration by whatever. [SPEAKER_00] And so will I have one for text and then one for voice?
SPEAKER_00
Will you guys do text too? Will just, how does that break down, because it seems like this is also succeeding at the text level with Finn and Sierra and all these things. [SPEAKER_01] And he described a very similar phenomenon that you described, which is you start maybe thinking, oh, this will help me answer customer support queries. But it becomes like a generic UI for the website where it's a box you can type in to go do things and understand things.
SPEAKER_00
And so why wouldn't you read the docs and design your integration by whatever. And so will I have like one for text and then one for voice? Will you guys do text too? Will just, how does that, because it seems like this is also succeeding at the text level with Finn and Sierra and all these things.
SPEAKER_00
The places where we know we will be able to provide the biggest value is where ultimately today you will have more, either a big portion or most of the interactions coming through voice. So if that kind of intersection is there, that's where we can provide higher value. And of course, if you need a text chatbot there, that's like, if you fix the voice agent, you'll have fixed text piece or inherently as well. But the place where we optimize today is going to be like, how do you select the right voice for the right customer interaction? How you pull that in the pretty complex case of what you mentioned earlier of like how you orchestrate that to pause or look for something deeper into the docs, how it can be extension of entirety of the business. So not only in support, but across entire of the user journey.
SPEAKER_00
[SPEAKER_01] But the bottom line is like, we want to be able to provide you across entirety of the interactions. Yeah. Voice is usually a big part of those interactions. Yes. And yes, we need to solve the integrations. We need to solve the knowledge. We need to solve text as part of that. But we wouldn't, for example, go into what I think will happen in a lot of those cases, very deeply into reasoning version of those use cases where you maybe need to like the multi-touch. Yeah, yeah, yeah. And a lot of complex actions. A lot of like financial analysis of like, is like that would be not something we optimize for.
SPEAKER_00
[SPEAKER_01] Can we talk about your revenue ramp where you're just one of the fastest, fastest growing startups period of the past few years? What's your most recently announced revenue figure? Most recently announced was end of 2025. Or whatever number you want to give us. [SPEAKER_01] So most recently announced was 350 at the end of 2025. Yeah.
SPEAKER_00
But the best proof of the technology working. So recently we were announced, I worked with Deutsche Telekom and T-Mobile with Revolut, with Klarna, with Meta, with IBM, a wide set of use cases. And this quarter was kind of one of the best for enterprise growth where we had the first quarter hit 100 million in additional ARR growth. Which is crazy. In net new ARR.
SPEAKER_01
[SPEAKER_00] In net new ARR. Okay. So if you're thinking this quarter was 100 million in net new ARR and 350 million at the end of the year. Yeah. I'm no mathematician, but it's up in the 450 million range. So versus this time last year, that's a several fold increase. Just what's working? Like from the outside, I would assume that there is really strong cohort growth within accounts. And then you seem to have self-serve and enterprise businesses that both contribute a lot. I don't know how big self-serve is, but as a user, I like to be able to fiddle with 11 labs and not have to go talk to sales.
SPEAKER_01
But maybe you can just talk about what works to reach 450 million plus of ARR so quickly.
SPEAKER_00
Yeah, so exactly. So we are over 50% is now sales led on enterprise. Yeah. And I think largely that the technology that powers a lot of their agentic interactions just became reliable at the same time as high quality over the last year, year and a half. So that's, frequently, you know this extremely well. You will start the account and then, of course, it continues expanding. And we see there's a land and expand motion across 11 labs. And what does that expand look like? Is it like new departments? Is it just the usage starts taking off? When a customer expands?
SPEAKER_00
Both. But usually the first part too, it's like we try to make it very easy for our customers. Maybe that kind of against ourselves where we give the technology a pretty attractive economics. Because we so much believe in the technology providing value. So you can actually try it and test it. And then within that one department. And you think you'll make it up in usage basically. Exactly. That usage, the kind of content continues increasing because you know it's providing value. And then it's so much easier to make that a choice.
SPEAKER_00
Yes. And then of course cross department pollination is there too. And it's like, our work of digital comes sort of marketing side. So we did Magenta work and Pogta's generation. Then it kind of expanded to customer support. And then it expanded to us working on the agent across the entirety of the network so people can call in and have the agent. So you could see those step changes across. [SPEAKER_01] But we are now 470 people as a company. So we keep on growing. But some of the things that stay consistent is small teams. So we have less than 10 people teams for each of the product or research initiatives.
SPEAKER_00
Or even as you think about sharding some of our go to market strategy. Those will be smaller teams understanding the industry in depth, understanding the market in depth, and going independently and going quickly. So that definitely contributed largely to that. Two, especially on the biggest enterprises, what we found works is. And it's like we have the full spectrum. Self-serve, PLG motion that helps drive distribution, drive awareness of 11 labs. And on the completely other spectrum, we have the high touch for deployed engineering working side by side with the customers to customize the entirety of their work together.
SPEAKER_00
Why did you guys do self-serve? Because I presume you have a lot of competitors where they have tech and it's behind a contact sales form, and you have to go talk to an SDR and then talk to an AE, blah, blah, blah, blah, blah. And you guys just offer the tech available on this side. And I'm a huge believer in this. I mean, a huge part of Stripe's growth has been driven by the fact that we just made Stripe available to anyone. And on the completely other spectrum, we have the high touch for deployed engineering working side by side with the customers to customize the entirety of their work together.
SPEAKER_01
[SPEAKER_00] Why did you guys do self-serve? [SPEAKER_00] Because I presume you have a lot of competitors where they have tech and it's behind a contact sales form, and you have to go talk to an SDR and then talk to an AE, blah, blah, blah, blah, blah. [SPEAKER_00] And you guys just offer the tech available on this side. [SPEAKER_00] And I'm a huge believer in this. [SPEAKER_00] I mean, a huge part of Stripe's growth has been driven by the fact that we just made Stripe available to anyone and built a lot of product around that adoption pattern. [SPEAKER_00] But so many companies seem to skip it. [SPEAKER_00] So I'm curious how you guys can...
SPEAKER_00
So many reasons. So many reasons. I think the quick ones that come to mind is feedback loop. You just have immediate understanding of how good your technology is. Two, which is an extension of that. We stand behind our tech. We believe it's the best in the world for models, for voice agents, for deployment. So we want people to experience that. And I think you do that the same in Stripe where the best version of the technology is available to everyone, which is so attractive to actually try it out. We always try to make everything we build for the highest end use cases, bring it back to the ecosystem free.
SPEAKER_00
Frequently, the newest of the use cases, for enterprise, you will need reliability, you need compliance, you need the scale which we deliver. So frequently, as you develop new technology, it might not be ready for a lot of those parameters, but it's definitely ready for developers and SMBs. And we love what they are doing because they are showing us the future and effectively helping us find a trajectory of where Eleven Labs should go. I'm totally convinced, I'm just always amazed that more companies don't pursue it where it feels like they're really shooting themselves in the foot by not... [SPEAKER_01] Like, did you guys self-serve on Stripe or did you...?
SPEAKER_00
We self-serve on Stripe.
SPEAKER_01
Yeah. For example, Eleven Labs is a huge company and yet you started on Stripe on a self-serving basis.
SPEAKER_00
[SPEAKER_01] You kind of initially, and it's like, you know, we were two of us at the beginning.
SPEAKER_01
You try to see what's working in the industry, but you try to think from first principles. [SPEAKER_00] So you want to try it out, you want to understand how it works. [SPEAKER_00] So the more friction elements before you're trying it out, the less you trust whether it's available, whether it'll be additional payment that's hidden behind some of those steps.
SPEAKER_00
So you don't want to go through them. So it's so much... Speaking of Stripe, do you have any Stripe feedback for us? Anything you want us to fix? My most common feedback until recently is, why don't you give us pay-as-you-go user-based billing type version? But one of our finance leaders, Maciek, I know I was speaking with your team and that was the day before. Yeah, yeah. He was thinking about it for a long time. He's great. He said you guys should buy Metronome.
SPEAKER_01
[SPEAKER_00] You should buy Metronome. And then the next day, Metronome acquisition was announced.
SPEAKER_00
So now you have it. So that was my most common feedback and we'll be launching.
SPEAKER_01
[SPEAKER_00] That's a good announcement for this podcast. [SPEAKER_00] We'll be launching user-based billing to everyone.
SPEAKER_00
[SPEAKER_01] Sorry, I'm shocked you... Oh, as in previously...
SPEAKER_01
So pay-as-you-go, pay-as-you-go. Okay. Previously you had it on an enterprise basis, but everything on the self-serve basis was like plans? [SPEAKER_00] So we had the subscriptions, yeah. Subscription plans, you can go over them. [SPEAKER_00] Yeah. But now we are launching a full pay-as-you-go experience. [SPEAKER_00] So you can just try out voice engine, which is effectively this all orchestration loop all the way through to any of the models directly.
SPEAKER_00
Going back to self-serve, I think a new thing in AI is that all self-serve products should have pay-as-you-go as an option. Maybe you want to have a subscription with some unlimited tiers, but I don't know if you had the experience of using Claude and you're typing away your queries and eventually you hit some rate limit. And it's sorry, you've hit your usage limit. And you want to be able to do the thing that you can do with Claude, which is just pay per API. It's fine, I'll pay for it. And it's kind of very funny as a consumer to not have the option to pay more to use the product more.
SPEAKER_00
And so, yeah, I think every AI product will need, they probably want to have some all-you-can-eat. Most of all-you-can-eat is subscription with limits and then the ability to pay for overages. [SPEAKER_01] So it sounds like that's what you're doing. [SPEAKER_01] Yeah, exactly. That's what we're doing.
SPEAKER_01
[SPEAKER_00] The only thing I want to ask you about is I feel like all CEOs of larger companies today are trying to figure out how do all these AI advancements change the nature of the organization? [SPEAKER_00] And how do you redesign your organization around all this new intelligence? [SPEAKER_00] And so that could be about what the scaling factor is of the number of people you need to do the work. [SPEAKER_00] But it also should be, do you need more senior people because they're better able to direct the AIs, or maybe you can do the work of what previously would have been junior people?
SPEAKER_01
[SPEAKER_00] Do you need more junior people because they're going to be more AI native in how they work?
SPEAKER_00
Do you want smaller teams? Do you want bigger teams? How do you actually go do the process engineering of, you know, your finance team should be using Claude extensively, but finance teams did not historically have a lot of homemade software. And so there's all these questions that are floating around and you have very rapidly built a much more AI native company. And so I'm curious what lessons we should all be learning from Eleven Labs as a large business recently built and without the baggage of decades of how we've always done it. [SPEAKER_00] Yeah. Yeah. We started in 2022, which is a year when the two topics of the day were crypto and metaverse.
SPEAKER_00
So just before, and then of course AI took off. [SPEAKER_00] You scaled the AI era. Exactly. Exactly. But we could have had the privilege of scaling through the world when it was all happening. For us, it works, and we really believe in that being the big part of the future. The first is small teams, keeping the teams small and super flat. So can you have both me and my co-founder will have over 15 direct reports each that we'll work with. And most of those people will have that same scale of direct reports. [SPEAKER_01] Okay. So your span of control is way larger than the traditional company. So just before, and then of course AI flow started. You scaled the AI era.
SPEAKER_00
Exactly. Exactly. But we had the privilege of scaling through the world when it was all happening. For us, Woodworks, and we really believe in that being the big part of the future. The first is small teams, keeping the teams small and super flat. So can you have both me and my co-founder will have over 15 direct reports each that we'll work with. And most of those people will have that same scale of direct reports. [SPEAKER_01] Okay. So your span of control is way larger than the traditional company. Exactly. And if you have double that, you have double that. And obviously that's an exponential.
SPEAKER_00
[SPEAKER_01] Exactly. And of course, there are some teams which in the short term might not do that. But ultimately that's where we think it's going to be headed. It's roughly 10 team size within each of those work items. And startups, no offense, but startups often have pretty wacko management ideas. [SPEAKER_01] There's a funny tweet, "Grant me the confidence of an early stage startup founder blogging about their management theories." [SPEAKER_01] But you think this is not a startup effect. This is an AI effect where basically... [SPEAKER_01] No, it's definitely a little bit of startup effect too. I figured out it's the hindsight benefit.
SPEAKER_00
[SPEAKER_01] I'm canceling our Stripe changes. Yeah. [SPEAKER_01] No, it's the hindsight of this may be working. [SPEAKER_01] We'll see in the next five to 10 years. And there are places... [SPEAKER_01] Much flatter org. [SPEAKER_01] Much flatter org. So it works for us, might not work for all the companies. [SPEAKER_01] And there are some parts where go to market, we still are trying to figure out what's the best way. But smaller teams, flatter org. And I think there are two paradigms, but generally people being more technical. Or if not technical, even in non-technical teams, having a technical resource.
SPEAKER_00
[SPEAKER_01] So we will have a person in ops or in talent that will, we have effectively a tech lead for that team.
SPEAKER_01
Yes. That helps them automate a lot of that work and helps up level the rest of the team too. Yes. So there are two parts that are helping. Okay. So talk me through this in talent or something like that. Is it that you are building your own software where other companies might have bought software like a Workday or a Greenhouse or something? Is it that they are using the existing software you have better? Is it the process that spreadsheets in a traditional company are built with software? How do you use the software in these sorts of organizations?
SPEAKER_00
[SPEAKER_01] Yeah. Sometimes, but we still use a lot of the traditional vendors. [SPEAKER_01] One pattern is, of course, elementifying everything, making the data explorable for you to be able to interact with it. [SPEAKER_01] Yes. [SPEAKER_01] Like who's in the pipeline? What worked? Who does the best references? All of that works. So you can double down on that. But two, it's frequently things that you manually do that a lot of the current vendors have a gap between where the agents are today versus what you could do if you have the technical skill set. And a good example is how do you scrape all the right profiles to be able to reach out to the right candidates?
SPEAKER_00
So you analyze whether it's, how much I should want to say, but about the, try to detect specific things that we know worked. So you'll bring that across to the people. On go to market side, there's so many things you can do with additional amplifiers. It goes from understanding what case studies are relevant and creating a good pre-read for you before you go to the meeting, through creating the AI SDR experience that we spoke about, to creating an entire deck experience. So you have a pre-populated deck with the right numbers that is customized to that customer, which you want still the person to go through and develop, but ultimately is in there.
SPEAKER_00
So there's plenty of those additional things that will amplify the work of the people around, potentially replace some of those easier tasks that are done. And then there's, we wanted for people to explore the culture at Eleven Labs. So we created a voice agent that people can speak with and see what's the culture, but also get prepped for the interviews. I think across many of those teams, there's additional benefit of what they can do. Interesting piece. So of course, in Ukraine, with ongoing work, they need to rethink a lot of how their development, their systems, their support works for the citizens across the country.
SPEAKER_00
[SPEAKER_01] And people are in the war zone. They don't have the same access to the information. They cannot rely on the same phone lines. [SPEAKER_01] They cannot rely on the same physical services around the country. So they've developed effectively a central...
SPEAKER_01
[SPEAKER_00] We had a few, but they reached out because they were developing their central map called Dia. They developed it over the years, but now with war, they were doubling down on how this can be a way of supporting the citizens. [SPEAKER_00] And of course, there's an easy part of how you create a digital government where you have help with the benefits and what's happening on the front line or education. So that's delivered to everyone. [SPEAKER_00] Or healthcare, so you can book your checkup or appointment. So how you create all of that.
SPEAKER_00
And of course, we traveled to Kyiv. We worked with them on bringing that and making that available for voice so everybody can access it.
SPEAKER_01
But the thing we've learned while being there was that model of what we speak about where you have technical resources in each of the teams.
SPEAKER_00
[SPEAKER_01] They actually have the same in every one of the ministries. [SPEAKER_01] So every ministry had technical resources working on creating that agentic version of their work. [SPEAKER_01] And then it was a central digital transformation team that would assemble this all together to deliver that for the central citizen support. Which I thought was brilliant. That's very tech forward by Ukraine. So tech forward, the most advanced set of work we've seen. [SPEAKER_01] So we got validated, okay, maybe technical resources in each of the teams is a good idea.
SPEAKER_01
[SPEAKER_00] And that works happily for us. [SPEAKER_00] And you mentioned some of the other parts, do you hire the senior or younger?
SPEAKER_00
The main thing we try to filter for, of course, the culture piece is so important. You can scale people, but scaling culture is much harder. [SPEAKER_00] So you want to optimize for that being right. And in our case, it's first principles, taking ownership, striving for excellence, but staying humble. And the main thing that's in that ownership part that I think works well for the AI world is agency. If you have that agency to explore, regardless of where you are in the experience cycle, it's going to be a tremendous amplifier to your work.
SPEAKER_00
You mentioned some of the other parts—do you hire senior or younger? The main thing we try to filter for is the culture piece, which is so important. You can scale people, but scaling culture is much harder. So you want to optimize for that being right. In our case, it's first principles, taking ownership, striving for excellence, but staying humble. The main thing in that ownership part that I think works well for the AI world is agency. If you have that agency to explore, regardless of where you are in the experience cycle, it's going to be a tremendous amplifier to your work. My biggest takeaway from all this has been that around agency—high agency people are the winners of the advances in AI and within organizations. Low agency people will lose out.
SPEAKER_00
Completely agree. Probably the most proud thing that Piotr and I are is as we scale 11 Labs, the people at 11 Labs—it's been the culture and seeing the expansion of the culture where culture builds the company now, rather than any single person or any single product. That was probably the biggest validation and happiness. There's the other angle of that where I think people are striving to be incredible in their craft and their work, but at the same time have fun in their work. That combination of agency and just enjoying what you do is probably the best thing we've been able to do today at 11 Labs.
SPEAKER_00
[SPEAKER_01] It sounds like a really fun stage. We were saying interesting research breakthroughs, really fast growing business. So I'm sure you're enjoying it. Andy, thank you. John, thank you so much. Thank you. Yeah. Yeah. We started in 2022, which is a year when the two topics of the day were crypto and metaverse. So just before, and then of course AI flow started. You scaled the AI era. Exactly. Exactly. But we could like had the privilege of like kind of scaling through the world when it was all happening. For us, Woodworks, and we like really believe in that being the big part of the future.
SPEAKER_00
The first is small teams, like keeping the teams small and super flat. So like, can you have both me and my co-founder will have over 15 direct reports each that we'll work with. And most of those people will have that same scale of direct reports.
SPEAKER_01
Okay. So your span of control is way larger than the traditional company.
SPEAKER_00
Exactly. And if you have double that, you have double that. And obviously that's an exponential.
SPEAKER_01
Exactly. And of course, you know, there are some teams which in the short term might not do that.
SPEAKER_00
But ultimately that's where we think it's going to be headed. It's like roughly 10 team size within each of those work items. And startups, no offense, but like startups often have pretty wacko management ideas.
SPEAKER_01
Like there's a funny tweet, Laura grant me the confidence of a, you know, early stage startup founder blogging about their management theories. But like, you think this is not a startup effect. This is an AI effect where basically... No, it's definitely a little bit of startup effect too. I figured out it's like, it's the hindsight benefit. I'm canceling our stripe changes. Yeah. No, no, it's like, I need to pre-end that. I kind of, you know, it's the hindsight of this may be working. We'll see in the next five to 10 years. And there are places... Much flatter org. Much flatter org. So it works for us, might not work for all the companies.
SPEAKER_01
And there are some parts where like go to market, we still are trying to figure out what's the best way.
SPEAKER_00
But smaller teams, flatter org. And I think there are two paradigms, but like generally people being more technical. Or if not technical, even in non-technical teams, having a technical resource.
SPEAKER_01
So, you know, we will have a person in ops or in talent that will, we have effectively a tech lead for that team. Yes. That helps them automate a lot of that work and helps up level the rest of the team too. Yes. So there are kind of two parts that are helping. Okay. So talk me through this in talent or something like that. Is it that you are building your own software where other companies might have bought software like a workday or a greenhouse or something? Is it that they are using the existing software you have better? Is the process that will be spreadsheets in a traditional company are built with software?
SPEAKER_01
How do you kind of use the software in these sorts of organizations? Yeah. Like, sometimes, but we still use a lot of like the traditional vendors. Like one pattern is, of course, elementifying everything, like making the data explorable for you to be able to interact with it. Yes. Like who's in the pipeline? What worked? Who does the best references? Like all of that. All of that works.
SPEAKER_00
So you can double down on that. But two, it's frequently things that you manually do that a lot of the current, like there is a gap between where the agents are today versus what you could do if you have the technical skill set. And a good example is like, how do you scrape all the right profiles to be able to reach out to the right candidates? So you like analyze whether it's, you know, how much I should want to say, but about the, like, try to detect specific things that we know worked. So you'll bring that across to the people. On go to market side, like there's just so many things you can do with additional amplifiers.
SPEAKER_00
You know, it goes from understanding what case studies are relevant and creating a good pre-read for you before you go to the meeting, through creating the AISDR experience that we spoke about, to creating an entire deck experience. So you have like a pre-populated deck with the right numbers that is customized to that customer, which you want still the person to go through and develop, but ultimately is in there. So there's plenty of those additional things that, you know, will amplify the work of the people around, potentially replace some of those, those easier tasks that are done.
SPEAKER_00
And then there's like, you know, we wanted for people to explore the culture at Eleven Labs. So we created a voice agent that people can speak with and see what's the culture, but also get prepped for the interviews. I think across many of those teams, like additional, additional benefit of what they can do. Interesting piece. So of course, in Ukraine, with ongoing work, they need to rethink a lot of how their development, their systems, their support works for the citizens across the country.
SPEAKER_01
And people are in the war zone. They don't have the same access to the information. They cannot rely on the same phone lines. They cannot rely on the same physical services around the country. So they've developed effectively a central…
SPEAKER_00
We had a few, but they reached out because they were developing their central map called Dia.
SPEAKER_01
They developed it over the years, but now with war, they were double downing of how this can… doubling down on how this can be a way of supporting the citizens.
SPEAKER_00
And of course, there's an easy part of like how you create a first agenda government where you have help with the benefits and what's happening on the front line or education. So that's delivered to everyone. Or healthcare, so you can book your checkup or appointment. So like how you create all of that. And of course, we traveled to Kiev. We worked with them on bringing that and making that available for voice so everybody can access it.
SPEAKER_01
But the thing we've learned while being there was that model of what we speak about where you have technical resources in each of the teams. They actually have the same in every of the ministries. So every ministry had technical resources working on creating that agentic version of their work. And then it was like a central digital transformation team that would like assemble this all together to deliver that for the central citizen support,
SPEAKER_00
which I thought was brilliant. That's very tech forward by Ukraine. So tech forward, like the most advanced set of work we've seen.
SPEAKER_01
So we got a little bit validated like, okay, maybe technical resources in each of the teams is a good idea.
SPEAKER_00
And that works, ah, happily for us. And, you know, you mentioned some of the other parts, like do you hire the senior or younger? Like main thing we try to filter for, of course, the culture piece is so important. You can scale people, but it's scaling culture is much harder. So like you want to optimize for that being right. And in our case, it's first principles, taking ownership, striving for excellence, but staying humble. And the main thing that's kind of in that ownership part that I think works well for the AI world is agency. Like if you are, if you have that agency to explore, regardless of where you are in the experience cycle,
SPEAKER_00
it's going to be a tremendous amplifier to your work. My biggest takeaway from all this has been that around agency where I feel like high agency people are the winners of the advances in AI and within organizations, low agency people will lose out. Yeah, completely agree. Probably the most proud thing that Piotr and I are is as we scale the 11 labs, the people that are at 11 labs, it's been like just the culture and seeing the expansion of the culture where culture builds the company now, rather than any single person or any single product builds the company. That was probably the biggest validation and happiness.
SPEAKER_00
And there's a kind of the other angle of that where I think people are like striving to be incredible in their craft and their work, but at the same time have fun and a lot of their work and that kind of combination of agency and just enjoying what you do is probably the best thing we've been able to do today at 11 labs.
SPEAKER_01
Well, it sounds like a really fun stage, like we were saying, interesting research breakthroughs, really fast growing business. So I'm sure you're enjoying it. Andy, thank you. John, thank you so much. Thank you.