SPEAKER_02
So I'm from Mistral AI and we are going to talk about speech generation and text-to-speech. There is an occasion, we released last week our first text-to-speech model and it's open source, so I really encourage you to check it out. It's an extremely strong text-to-speech model. We are very proud of it. And for this occasion, I thought we could review some of the recent trends in text-to-speech architecture since there is a dominant trend emerging these days, although this can change very quickly. And so this talk is slightly academic and addressed to people who want to know a bit more about how you do text-to-speech.
SPEAKER_02
This being said, we have a few years before the machines do all the science for us, so we might enjoy it today. I'm Sam.
SPEAKER_02
Yeah, I work at Mistral as an AI scientist. Before I was at Facebook Fair when it was called Facebook. Mistral, a few words about the company. We are a frontier lab.
SPEAKER_02
We have been founded a couple of years ago. We produce frontier models, but we are also a B2B business. We help organizations in their AI transformation, which is a buzzword, but literally every company is transforming with AI. We help them by providing them tools, products, and dedicated people to help them in their custom needs. Back to the text-to-speech. So there are a few offline use cases of speech generation, like the famous listen to the blog or listen to the article. But nowadays the king use case for text-to-speech is its usage within agents.
SPEAKER_02
And in particular, it's used to interface with a chat agent, typically in a pipe like this, where you have a central chat agent that does text-to-text, but does it extremely well. And you want to talk to it, so you add a speech-to-text, and you want it to speak to you, so you add a text-to-speech. As everybody in this conf will tell you, the latency is key here. So you can reduce the latency on the left by having the speech-to-text done in real-time, so that when you detect the end of turn, you already have the transcript, it's already done. And we are going to focus a bit on the right side today.
SPEAKER_02
It's also very important that as soon as you have the first audio packets, you start to voice them out. This way the perceived latency is lower. In fact, since your LLM can stream some text to you, actually what you ultimately want is something like this, if you're going to interface a chat assistant, which is a real-time text input text-to-speech, where as soon as you have the first token of the LLM, the machine starts to speak. We are going to talk about it at the end of the talk. I want to focus a bit at the beginning on the output side, and what it means to stream audio. So to illustrate this, I have this app that I've coded for the occasion.
SPEAKER_02
And so we are going to use this text-to-speech model that we released, that I mentioned, and we are going to hear Paul. So Paul is an actual human being that sounds like this. [SPEAKER_03] The persistent anxiety that fills the rest of my life is calmed for as long as I have the flavor of something good in my mouth. So this is like some actual recording of some actual person named Paul, and we are copying his voice. [SPEAKER_03] So with the sunshine and the great bursts of leaves growing on the trees, just as things grow in fast movies, I had that familiar conviction that life was beginning over again with the summer. So let's focus first on what's happening here.
SPEAKER_02
As you can see, we are copying the voice, and the first audio packet happens first, and we can start to emit audio, which greatly reduces the perceived latency, even though the full computation of the audio happens a few seconds later. So if you use it in an agent, so here I crafted a small agent using a speech-to-text, one of our LLMs and this very text-to-speech. So we can speak to Paul. Hey Paul, can you tell me what's the title of the session at 12:20, please? [SPEAKER_03] The session at 12:20 PM is titled RICCI Mini, Giving a Body to AI by Andres Marafioti. And what was the session at 11:15, please?
SPEAKER_02
[SPEAKER_03] The session at 11:15 AM is Beyond Transcription, Building Voice AI that Actually Understands Conversations by Hervé Braden. Did you enjoy it as much as I did? [SPEAKER_03] I don't have personal experiences or emotions, but I'm glad you enjoyed it. That's all I can do. So the important thing here is that since the audio packet arrived first, you still have a decent latency and you can enjoy the conversation with the agent, despite the fact that the audio is still not generated fully. And so we're going to dig... Oh, yeah, sorry. I want to make one digression.
SPEAKER_03
[SPEAKER_02] So I mentioned the voice cloning here. [SPEAKER_02] This model can only need a few seconds to clone the voice of someone.
SPEAKER_02
So again, this is how it sounded. All day long I can...
SPEAKER_03
[SPEAKER_02] And this is how we generate text. And so with the sunshine and the great bursts of leaves growing on the tree. [SPEAKER_02] It really sounds alike.
SPEAKER_02
It's also very good at inferring how a person would speak in another language. So for example, this is a voice, a French voice. Maybe. [SPEAKER_00] Ryan Kugler, réalisateur et scénariste des deux Black Panthers. And if we generate... [SPEAKER_00] And so with the sunshine and the great bursts of leaves growing on the trees, just as things grow in fast movies. Yeah. So we can clearly recognize her and we can clearly recognize a very strong French accent, which as a French person myself, I do enjoy. Again, I can even clone my own voice. So this is how I sounded during the recording.
SPEAKER_02
[SPEAKER_01] Hi, this is Sam. And this is how I sound. And so with the sunshine and the great bursts of leaves growing on the trees.
SPEAKER_03
[SPEAKER_02] It works pretty well. [SPEAKER_02] So this way in my time of delusion and at the peak of my ego, I can discuss with myself on a complicated problem, which is nice.
SPEAKER_02
And so it's becoming so easy to impersonate a voice that it's becoming very easy to configure.
SPEAKER_03
[SPEAKER_02] So we can clearly recognize her and we can clearly recognize a very strong French accent, which as a French myself, I do enjoy. [SPEAKER_02] Again, I can even clone my own voice.
SPEAKER_02
So this is how I sounded during the recording.
SPEAKER_03
[SPEAKER_01] Hi, this is Sam. [SPEAKER_01] And this is how I sound. [SPEAKER_01] And so with the sunshine and the great bursts of leaves growing on the trees.
SPEAKER_02
It works pretty well. So this way in my time of delusion and at the peak of my ego, I can discuss with myself on a complicated problem, which is nice. And so it's becoming so easy to impersonate a voice that it's becoming very easy to configure. So this is a small digression, but currently a lot of large companies that do have a concept of vocal identity and they do care in their branding about how they sound in their advertisement in particular. But I think this concept will be becoming more mainstream. And just as a lot of companies define how their website appears as their brand identity, it would be the same for the voice identity. Back to how we do it in general.
SPEAKER_02
So, right, this is an audio. Physically, it's the pressure of the microphone that we measure from time to time, several thousands of times per second. So it looks like this. And historically, to generate the audio, there have been a lot of attempts, a lot of systems. In the prehistoric time, you have stitching of words that were spoken, in the French train system, SNCF, for those who know.
SPEAKER_03
[SPEAKER_02] And then neural generation arrived.
SPEAKER_02
At some point, the train was to generate each sample one after the other. Then another era was generating the whole audio at once. But as we saw, it's very interesting to have the beginning of the audio generated first so that we can start to play it out. So it seems that most labs have converged to some common patterns.
SPEAKER_00
[SPEAKER_02] And obviously, the first one is inspired by a large language model. [SPEAKER_02] We are trying to transform the problem as a language modeling problem because humanity is extremely good at modeling sequences of token.
SPEAKER_02
So pretty much everybody is using an autoregressive decoder backbone and generate audio one piece after the other.
SPEAKER_00
[SPEAKER_02] Now, as I said, we really don't want to generate one sample after the other. [SPEAKER_02] So what we want to do is generate patches of audio one after the other.
SPEAKER_02
So the expected system looks like this, where you have an encoder that transforms a frame of audio, something like this, 80 milliseconds into something that ideally is a token because humanity is very good at modeling sequences of token. And then you have a decoder that does the opposite. Now for text, it's pretty easy because transforming the text into tokens is easy, right? You can take words as token and it works pretty well, even though we do much better. For audio, it's much harder because one token doesn't have a lot of information.
SPEAKER_01
[SPEAKER_02] One token of a vocabulary of thousands has 10 bits of information. [SPEAKER_02] And the audio requires much, much more, a much larger bit rate. [SPEAKER_02] For example, a standard quality MP3 is 200 kilobits per second.
SPEAKER_02
And so in order to transform this into a sequence of token and not have thousands and thousands of tokens, we need to somehow compress it and reduce the size of it, maybe drop what's not needed. An interesting point of comparison is text captioning. Because if you drop all the acoustic information and you just focus on the text, with a subtitle track, you actually drop most of the information.
SPEAKER_02
It's a massive reduction and you only have a few bits per second remaining. So here in this demo, I use real-time speech text to measure my bit rate in terms of tokens per second of text. And I'm a very competitive person and very good at speaking. Yet, I'm barely 15 bits per second of actual information. You can try to beat that. But in the grand scheme of things, compared to 200,000 bits of information per second, that's not a lot. Obviously, we want to use something that allows to recover acoustic features, the voice and other aspects, and not just the semantic information as in the text.
SPEAKER_02
So the codecs that are used typically reduce the audio to a few thousand bits per second. In our case, for instance, we cut the audio with pieces of 80 milliseconds, so 12 frames per second. And we transform each frame into several tokens, 37 in our case. So we reduce the problem to about 500 tokens per second. I'm not going to dig too much on how we train these codecs, but obviously first we train them. We train them by reconstructing a very large set of audio and using a bottleneck here. So here, the training procedure constrains the reconstruction to go through a step where each frame is decomposed into several tokens.
SPEAKER_02
To do this, typically it's guided so that the model drops the information that is useless and only retains the one that is useful. And so we guided via some losses, reconstruction losses, adversarial losses, and particularly for some of the tokens, we try to make sure that they contain the text information so that you can reconstruct the text from it. So, still 500 tokens per second is a lot of tokens. And you could put them one by one, aligned as a sequence like this. But it would make a lot of steps of the main transformer that is at the core of the system, which is huge.
SPEAKER_02
In our case, it's 4 billion parameters, which is still a lot, even though it's not extremely big now. What most people do then is have one step of the backbone per frame and a smaller model here, typically a dev transformer, that recomputes all the tokens of one frame at each step.
SPEAKER_02
So this way you still have a lot of tokens. You still carry a lot of information, but the computation is much faster. So this is the main pattern that we see. Even though for that last bit, the model we released does not follow this pattern, we actually defer on that part. I'm not going to dig too much on it, but just so you know, each frame, which is represented by 37 tokens in our case, we do generate these 37 tokens at once using a diffusion model. So it's slightly different from the vanilla text-to-speech nowadays. I encourage you, by the way, to read our technical report, which contains all this information.
SPEAKER_02
Also, it's a pretty cool use case of flow matching models, which is similar to diffusion models. Now, the main part is conditioning, actually, because so far we are just generating audio, but we are not conditioning it on text. So it's not really a text-to-speech, it's just speech. And for conditioning, there is way more variance across labs and papers and implementations. You have typically two categories. There are the people who focus on producing the audio once you have the text, and some of them focus on having a stream of text. I encourage you, by the way, to read our technical report, which contains all this information.
SPEAKER_02
Also, it's a pretty cool use case of flow matching models, which is similar to diffusion models. Now, the main part is conditioning, actually, because so far we are just generating audio, but we are not conditioning it on text. So it's not really a text-to-speech, it's just speech. And for conditioning, there is much more variance across labs and papers and implementations. You have typically two categories. There are the people who focus on producing the audio once you have the text, and some of them focus on having a stream of text. Typically, the first category will tend to provide all the context at the beginning and then produce the audio as we saw.
SPEAKER_02
Typically, the second category will also add some context as the audio is produced. The model we release is in the first category. So what we do is we provide the audio of the voice we want to clone, so a few seconds, then the text to pronounce, and that's our context in our case. Regarding the latency, it's pretty fast if you remove the network. And with a single GPU, you have 70 milliseconds between the moment where you input your text and the moment where you have the first audio you can play. Regarding real-time text input stream, which is our next step for us, there is not a clear winner.
SPEAKER_02
First, it's still possible to generate independently the text and stitch them out, but obviously you will have a lot of continuity problems. There are several patterns. The two main ones are people who interleave audio and text. So as soon as there is new text, they put the text in the same layer. And some others who have a dual-stream architecture where you have a stream of audio and a stream of text, and you blend them together during the inference. How am I doing on time? There's two minutes remaining. Thank you.
SPEAKER_02
The takeaway is check our open-source model, please. Read the technical paper, and I hope you learned a few things today. We do have two minutes.
SPEAKER_02
How you spend them is up to you. Yeah. [SPEAKER_01] You said that your model first takes hold of the text and then produces the audio, but on the example that you showed of the voice agent, it seemed like it was generating the text and the audio at the same time. No. So the question was, on the demo, it looks like we are generating the text and the audio at the same time. No, for the voice agent. Hello, Paul, it's me again. Can you say anything, I don't know, a poem? [SPEAKER_03] I'm afraid I can't recite poems. It's just, the text is produced in one go. It's just that I'm using a small LLM that is very fast, so it's nearly immediate, and then the audio is produced later.
SPEAKER_02
Yeah. Yeah. Thanks for the weight model, by the way. I know that the weights are open. [SPEAKER_01] Is the voice cloning encoder also open? Yeah, there is a small asterisk here. We didn't release this part, the encoder part, which means it's the only thing that is missing for you to clone your own voice. It's a feature that we only serve in a proprietary fashion for now. So what you can do is use the text-to-speech model, use one of the open voices that we provide. We may provide more in the future. Yeah, so far we just didn't want to give everybody the ability to clone any voice.
SPEAKER_02
Yeah. [SPEAKER_04] So big labs like Google and OpenAI are working more on native voice-to-voice models, let's say. Yeah.
SPEAKER_02
[SPEAKER_04] So your lab, you have a lab, they're doing more cascading architecture. [SPEAKER_04] What's your take on both and how you're seeing things? What's my take? So on the consumer side, you will always have the impression to speak to a single system that hears what you say and outputs something, right? So it's purely an architectural model here. My take on this is that we can go very far by just using speech as an interface, especially because these central LLMs are extremely capable, but they also do a lot of things. So just for the sake of being able to use any agent that has been released with the same interface, it has an advantage to interface.
SPEAKER_02
So you can go very far with just interface, especially if you are doing this kind of thing where you stream the text tokens that are output by the LLM. Yeah, last question because I'm out of time, sorry. [SPEAKER_05] Yeah, so the next steps on the interleaved function, which accepts both audio and text, sounds really interesting. [SPEAKER_05] So if you do that real-time, what do you see as the possibilities with that feature? So I didn't say that our next step would be this, right? I just said that there are several patterns to handle a stream of text as input as opposed to a finite amount of text.
SPEAKER_02
We actually don't know which one we'll choose, whether it's interleaved or another solution like delayed sequence modeling, for instance. So it's unclear which architecture is best, at least to us, at least to me. What it allows is lower latency because as soon as you have the first bit of text that is produced by the LLM, you can start voicing them out.
SPEAKER_01
[SPEAKER_02] So in this agent, that was not clear because the utterances were very short. [SPEAKER_02] But imagine I ask Paul to generate a full page of text.
SPEAKER_02
It would be nice if I don't have to wait for the end of the text generation to voice it out. [SPEAKER_05] Yeah, absolutely. Thank you. Hello, Paul, it's me again. Can you say anything like, I don't know, a poem?
SPEAKER_03
I'm afraid I can't recite poems.
It's just, so the text is produced in one go. It's just that I'm using a small LLM that is very fast, so it's nearly immediate, and then the audio is produced later. Yeah. Yeah. Thanks for the weight model, by the way. I know that the weights are open. Is the voice cloning encoder also open?
SPEAKER_02
Yeah, there is a small asterisk here.
SPEAKER_02
We didn't release this part, like the encoder part, which means it's the only thing that is missing for you to clone your own voice. It's a feature that we only serve in a proprietary fashion for now. So what you can do is use the text-to-speech model, use one of the open voices that we provide. We may provide more in the future. Yeah, so far we just didn't want to give everybody the ability to clone any voice.
SPEAKER_02
Yeah.
SPEAKER_04
So big labs like Google and OpenAI are working more on native voice-to-voice, let's say, models.
SPEAKER_02
Yeah.
SPEAKER_04
So your lab, you have a lab, they're doing more cascading architecture. What's your take on both like driving and how you're seeing things about?
SPEAKER_02
What's my take? So on the consumer side, you will always have the impression to speak to a single system that hear what you say and output something, right? So it's purely an architectural model here.
SPEAKER_02
My take on this is that we can go very, very far by just using speech as an interface, especially because these central LLM, they are extremely capable, but they also do a lot of things. So just for the sake of being able to use any agent that has been released with the same interface, it has an advantage to interface. So you can go very far with just interface, especially if you are doing this kind of thing where you stream the text token that's output by the LLM.
SPEAKER_02
Yeah, last question because I'm out of time, sorry.
SPEAKER_05
Yeah, so the next steps on the interleaved function, which accepts both audio and text, sounds really interesting. So if you do that real-time, what do you see as the possibilities with that feature?
SPEAKER_02
So I didn't say that our next step would be this, right? I just said that there are several patterns to handle a stream of text as input as opposed to like a finite amount of text. We actually don't know which one we'll choose, whether it's interleaved or another solution like delayed sequence modeling, for instance. So it's unclear which architecture is best, at least to us, at least to me. What it allows is lower latency because as soon as you have the first bit of text that are produced by the LLM, you can start voicing them out. So in this agent, that was not clear because the utterance were very short. But imagine I ask Paul to generate a full page of text.
SPEAKER_02
It would be nice if I don't have to wait the end of the text generation to voice it out.
SPEAKER_05
Yeah, absolutely.
SPEAKER_02
Thank you.