Open Reader

Why TTS Models Now Look Like LLMs — Samuel Humeau, Mistral

completed 22:26 May 09, 2026 Watch on YouTube

Current Status

completed

Video ID

3jGAU2sbAyY

RAG / Chat

Enabled
Why TTS Models Now Look Like LLMs — Samuel Humeau, Mistral
Description

The dominant architecture pattern for text-to-speech in 2026 looks a lot like an LLM — an autoregressive transformer generating sequences of tokens, one frame of audio at a time. Samuel Humeau from Mistral walks through why the field converged there, how neural audio codecs solve the information-density problem (audio carries ~200kbps of signal; you can't feed that raw to a transformer), and what the streaming trick actually is that makes voice agents feel responsive before the full audio has even finished generating. The talk uses Mistral's just-released open-weight TTS model as a running example — live demos of voice cloning from a few seconds of reference audio, a voice agent answering real conference schedule questions, and a breakdown of the codec-to-backbone-to-decoder pipeline that produces it all. There's also a frank section on what's still unsettled: how to handle streaming text input (tokens arriving from an LLM in real time rather than a fixed block of text) and why getting that right is the next meaningful latency win in agent pipelines. It's the kind of talk that makes the system feel less like a black box — not by oversimplifying, but by showing exactly which engineering choices are load-bearing and which are still open problems. Speaker info: - https://x.com/DrSamuelBHume - https://www.linkedin.com/in/samuelhumeau/

Summary

Generated by claude-haiku-4-5-20251001

Why TTS Models Now Look Like LLMs — Samuel Humeau, Mistral

Main Topics

  • Modern Text-to-Speech (TTS) Architecture: How TTS models are converging toward LLM-like designs
  • Real-time Speech Generation: Techniques for streaming audio with minimal latency
  • Voice Cloning: Rapid voice adaptation with minimal samples
  • Integration with AI Agents: Using TTS as a conversational interface
  • Audio Tokenization: Converting continuous audio into discrete tokens for efficient processing

Key Points

TTS in Modern AI Systems

  • TTS is now primarily used as an interface for conversational AI agents, rather than offline use cases like article-to-speech
  • The dominant pipeline combines speech-to-text input, LLM processing, and TTS output
  • Latency is critical—users expect responses within milliseconds

Architectural Convergence to LLM-Style Models

  • Autoregressive decoder backbone: Most labs now use LLM-inspired architectures with autoregressive decoding
  • Encoder-decoder pattern: Audio frames (e.g., 80ms chunks) are encoded into tokens, then decoded back to audio
  • Frame-based generation: Rather than generating individual audio samples sequentially, systems generate "patches" of audio (multiple tokens per frame)
  • Mistral's model uses 37 tokens per 80ms frame, reducing the problem to ~500 tokens/second instead of 200,000 bits/second of raw audio

Audio Compression & Tokenization

  • Raw audio requires massive bitrate (e.g., MP3: 200 kilobits/second)
  • Text carries only ~15 bits/second of semantic information
  • Codec training uses reconstruction losses and adversarial losses to compress audio while preserving voice characteristics
  • Some tokens are explicitly trained to contain text information for reconstruction

Streaming & Latency Optimization

  • First packet first: Audio generation starts immediately; first audio packets are output before full generation completes
  • Perceived latency: Playing early audio packets drastically reduces perceived latency for users
  • Single GPU performance: 70ms latency between text input and first playable audio output
  • Real-time text streaming: Still evolving—options include interleaved audio-text generation or dual-stream architectures

Voice Cloning Capabilities

  • Requires only a few seconds of voice sample
  • Works across languages, preserving speaker identity and accent
  • Enables voice identity as a branding element for companies

Notable Quotes

> "The king use case for text-to-speech is its usage within agents."

> "Humanity is extremely good at modeling sequences of tokens."

> "We really don't want to generate one sample after the other. We want to generate patches of audio one after the other."

> "As soon as you have the first token of the LLM, the machine starts to speak."

> "We can go very, very far by just using speech as an interface, especially because these central LLMs are extremely capable."

> "It's becoming so easy to impersonate a voice that it's becoming very easy to configure."

Takeaways

Technical Insights

  • TTS models are adopting LLM design patterns because sequence modeling of tokens is highly effective
  • Audio tokenization is the key challenge—compressing 200kbps of audio into a tractable token sequence while preserving acoustic information
  • Streaming matters more than batch generation—early audio output dramatically reduces perceived latency
  • Multiple architectural approaches coexist for handling real-time text streams, with no clear winner yet

Implementation Considerations

  • Frame-based generation (one transformer step per frame) with smaller models handling token-level details is becoming standard
  • Flow matching/diffusion models offer alternatives to vanilla autoregressive approaches
  • Voice cloning requires both a voice encoder and the main TTS model; some labs keep encoders proprietary

Future Directions

  • Real-time streaming of both text and audio simultaneously (not yet resolved)
  • Voice identity as mainstream branding, beyond just visual design
  • Trade-offs between cascading (modular) vs. end-to-end native architectures for voice agents

Actionable Items

  • Check Mistral's open-source TTS model (weights available; voice cloning encoder proprietary for now)
  • Read the technical report for detailed codec and training methodology
  • Consider latency optimization for voice agent deployment (streaming first packets is critical)

Transcript

3045 words en Processed in 111.3s

So I'm from Mistral AI and we are going to talk about speech generation and text-to-speech. There is an occasion, we released last week our first text-to-speech model and it's open source, so I really encourage you to check it out. It's an extremely strong text-to-speech model. We are very proud of it. And for this occasion, I thought we could review some of the recent trends in text-to-speech architecture since there is a dominant trend emerging these days, although this can change very quickly. And so this talk is slightly academic and addressed to people who want to know a bit more about how you do text-to-speech. This being said, we have a few years before the machines do all the science for us, so we might enjoy it today. I'm Sam. Yeah, I work at Mistral as an AI scientist. Before I was at Facebook Fair when it was called Facebook. Mistral, a few words about the company. We are a frontier lab. We have been founded a couple of years ago. We produce frontier models, but we are also a B2B business. We help organizations in their AI transformation, which is a buzzword, but literally every company is transforming with AI. We help them by providing them tools, products, and dedicated people to help them in their custom needs. Back to the text-to-speech. So there are a few offline use cases of speech generation, like the famous listen to the blog or listen to the article. But nowadays the king use case for text-to-speech is its usage within agents. And in particular, it's used to interface with a chat agent, typically in a pipe like this, where you have a central chat agent that does text-to-text, but does it extremely well. And you want to talk to it, so you add a speech-to-text, and you want it to speak to you, so you add a text-to-speech. As everybody in this conf will tell you, the latency is key here. So you can reduce the latency on the left by having the speech-to-text done in real-time, so that when you detect the end of turn, you already have the transcript, it's already done. And we are going to focus a bit on the right side today. It's also very important that as soon as you have the first audio packets, you start to voice them out. This way the perceived latency is lower. In fact, since your LLM can stream some text to you, actually what you ultimately want is something like this, if you're going to interface a chat assistant, which is a real-time text input text-to-speech, where as soon as you have the first token of the LLM, the machine starts to speak. We are going to talk about it at the end of the talk. I want to focus a bit at the beginning on the output side, and what it means to stream audio. So to illustrate this, I have this app that I've coded for the occasion. And so we are going to use this text-to-speech model that we released, that I mentioned, and we are going to hear Paul. So Paul is an actual human being that sounds like this. [SPEAKER_03] The persistent anxiety that fills the rest of my life is calmed for as long as I have the flavor of something good in my mouth. So this is like some actual recording of some actual person named Paul, and we are copying his voice. [SPEAKER_03] So with the sunshine and the great bursts of leaves growing on the trees, just as things grow in fast movies, I had that familiar conviction that life was beginning over again with the summer. So let's focus first on what's happening here. As you can see, we are copying the voice, and the first audio packet happens first, and we can start to emit audio, which greatly reduces the perceived latency, even though the full computation of the audio happens a few seconds later. So if you use it in an agent, so here I crafted a small agent using a speech-to-text, one of our LLMs and this very text-to-speech. So we can speak to Paul. Hey Paul, can you tell me what's the title of the session at 12:20, please? [SPEAKER_03] The session at 12:20 PM is titled RICCI Mini, Giving a Body to AI by Andres Marafioti. And what was the session at 11:15, please? [SPEAKER_03] The session at 11:15 AM is Beyond Transcription, Building Voice AI that Actually Understands Conversations by Hervé Braden. Did you enjoy it as much as I did? [SPEAKER_03] I don't have personal experiences or emotions, but I'm glad you enjoyed it. That's all I can do. So the important thing here is that since the audio packet arrived first, you still have a decent latency and you can enjoy the conversation with the agent, despite the fact that the audio is still not generated fully. And so we're going to dig... Oh, yeah, sorry. I want to make one digression. [SPEAKER_02] So I mentioned the voice cloning here. [SPEAKER_02] This model can only need a few seconds to clone the voice of someone. So again, this is how it sounded. All day long I can... [SPEAKER_02] And this is how we generate text. And so with the sunshine and the great bursts of leaves growing on the tree. [SPEAKER_02] It really sounds alike. It's also very good at inferring how a person would speak in another language. So for example, this is a voice, a French voice. Maybe. [SPEAKER_00] Ryan Kugler, réalisateur et scénariste des deux Black Panthers. And if we generate... [SPEAKER_00] And so with the sunshine and the great bursts of leaves growing on the trees, just as things grow in fast movies. Yeah. So we can clearly recognize her and we can clearly recognize a very strong French accent, which as a French person myself, I do enjoy. Again, I can even clone my own voice. So this is how I sounded during the recording. [SPEAKER_01] Hi, this is Sam. And this is how I sound. And so with the sunshine and the great bursts of leaves growing on the trees. [SPEAKER_02] It works pretty well. [SPEAKER_02] So this way in my time of delusion and at the peak of my ego, I can discuss with myself on a complicated problem, which is nice. And so it's becoming so easy to impersonate a voice that it's becoming very easy to configure. [SPEAKER_02] So we can clearly recognize her and we can clearly recognize a very strong French accent, which as a French myself, I do enjoy. [SPEAKER_02] Again, I can even clone my own voice. So this is how I sounded during the recording. [SPEAKER_01] Hi, this is Sam. [SPEAKER_01] And this is how I sound. [SPEAKER_01] And so with the sunshine and the great bursts of leaves growing on the trees. It works pretty well. So this way in my time of delusion and at the peak of my ego, I can discuss with myself on a complicated problem, which is nice. And so it's becoming so easy to impersonate a voice that it's becoming very easy to configure. So this is a small digression, but currently a lot of large companies that do have a concept of vocal identity and they do care in their branding about how they sound in their advertisement in particular. But I think this concept will be becoming more mainstream. And just as a lot of companies define how their website appears as their brand identity, it would be the same for the voice identity. Back to how we do it in general. So, right, this is an audio. Physically, it's the pressure of the microphone that we measure from time to time, several thousands of times per second. So it looks like this. And historically, to generate the audio, there have been a lot of attempts, a lot of systems. In the prehistoric time, you have stitching of words that were spoken, in the French train system, SNCF, for those who know. [SPEAKER_02] And then neural generation arrived. At some point, the train was to generate each sample one after the other. Then another era was generating the whole audio at once. But as we saw, it's very interesting to have the beginning of the audio generated first so that we can start to play it out. So it seems that most labs have converged to some common patterns. [SPEAKER_02] And obviously, the first one is inspired by a large language model. [SPEAKER_02] We are trying to transform the problem as a language modeling problem because humanity is extremely good at modeling sequences of token. So pretty much everybody is using an autoregressive decoder backbone and generate audio one piece after the other. [SPEAKER_02] Now, as I said, we really don't want to generate one sample after the other. [SPEAKER_02] So what we want to do is generate patches of audio one after the other. So the expected system looks like this, where you have an encoder that transforms a frame of audio, something like this, 80 milliseconds into something that ideally is a token because humanity is very good at modeling sequences of token. And then you have a decoder that does the opposite. Now for text, it's pretty easy because transforming the text into tokens is easy, right? You can take words as token and it works pretty well, even though we do much better. For audio, it's much harder because one token doesn't have a lot of information. [SPEAKER_02] One token of a vocabulary of thousands has 10 bits of information. [SPEAKER_02] And the audio requires much, much more, a much larger bit rate. [SPEAKER_02] For example, a standard quality MP3 is 200 kilobits per second. And so in order to transform this into a sequence of token and not have thousands and thousands of tokens, we need to somehow compress it and reduce the size of it, maybe drop what's not needed. An interesting point of comparison is text captioning. Because if you drop all the acoustic information and you just focus on the text, with a subtitle track, you actually drop most of the information. It's a massive reduction and you only have a few bits per second remaining. So here in this demo, I use real-time speech text to measure my bit rate in terms of tokens per second of text. And I'm a very competitive person and very good at speaking. Yet, I'm barely 15 bits per second of actual information. You can try to beat that. But in the grand scheme of things, compared to 200,000 bits of information per second, that's not a lot. Obviously, we want to use something that allows to recover acoustic features, the voice and other aspects, and not just the semantic information as in the text. So the codecs that are used typically reduce the audio to a few thousand bits per second. In our case, for instance, we cut the audio with pieces of 80 milliseconds, so 12 frames per second. And we transform each frame into several tokens, 37 in our case. So we reduce the problem to about 500 tokens per second. I'm not going to dig too much on how we train these codecs, but obviously first we train them. We train them by reconstructing a very large set of audio and using a bottleneck here. So here, the training procedure constrains the reconstruction to go through a step where each frame is decomposed into several tokens. To do this, typically it's guided so that the model drops the information that is useless and only retains the one that is useful. And so we guided via some losses, reconstruction losses, adversarial losses, and particularly for some of the tokens, we try to make sure that they contain the text information so that you can reconstruct the text from it. So, still 500 tokens per second is a lot of tokens. And you could put them one by one, aligned as a sequence like this. But it would make a lot of steps of the main transformer that is at the core of the system, which is huge. In our case, it's 4 billion parameters, which is still a lot, even though it's not extremely big now. What most people do then is have one step of the backbone per frame and a smaller model here, typically a dev transformer, that recomputes all the tokens of one frame at each step. So this way you still have a lot of tokens. You still carry a lot of information, but the computation is much faster. So this is the main pattern that we see. Even though for that last bit, the model we released does not follow this pattern, we actually defer on that part. I'm not going to dig too much on it, but just so you know, each frame, which is represented by 37 tokens in our case, we do generate these 37 tokens at once using a diffusion model. So it's slightly different from the vanilla text-to-speech nowadays. I encourage you, by the way, to read our technical report, which contains all this information. Also, it's a pretty cool use case of flow matching models, which is similar to diffusion models. Now, the main part is conditioning, actually, because so far we are just generating audio, but we are not conditioning it on text. So it's not really a text-to-speech, it's just speech. And for conditioning, there is way more variance across labs and papers and implementations. You have typically two categories. There are the people who focus on producing the audio once you have the text, and some of them focus on having a stream of text. I encourage you, by the way, to read our technical report, which contains all this information. Also, it's a pretty cool use case of flow matching models, which is similar to diffusion models. Now, the main part is conditioning, actually, because so far we are just generating audio, but we are not conditioning it on text. So it's not really a text-to-speech, it's just speech. And for conditioning, there is much more variance across labs and papers and implementations. You have typically two categories. There are the people who focus on producing the audio once you have the text, and some of them focus on having a stream of text. Typically, the first category will tend to provide all the context at the beginning and then produce the audio as we saw. Typically, the second category will also add some context as the audio is produced. The model we release is in the first category. So what we do is we provide the audio of the voice we want to clone, so a few seconds, then the text to pronounce, and that's our context in our case. Regarding the latency, it's pretty fast if you remove the network. And with a single GPU, you have 70 milliseconds between the moment where you input your text and the moment where you have the first audio you can play. Regarding real-time text input stream, which is our next step for us, there is not a clear winner. First, it's still possible to generate independently the text and stitch them out, but obviously you will have a lot of continuity problems. There are several patterns. The two main ones are people who interleave audio and text. So as soon as there is new text, they put the text in the same layer. And some others who have a dual-stream architecture where you have a stream of audio and a stream of text, and you blend them together during the inference. How am I doing on time? There's two minutes remaining. Thank you. The takeaway is check our open-source model, please. Read the technical paper, and I hope you learned a few things today. We do have two minutes. How you spend them is up to you. Yeah. [SPEAKER_01] You said that your model first takes hold of the text and then produces the audio, but on the example that you showed of the voice agent, it seemed like it was generating the text and the audio at the same time. No. So the question was, on the demo, it looks like we are generating the text and the audio at the same time. No, for the voice agent. Hello, Paul, it's me again. Can you say anything, I don't know, a poem? [SPEAKER_03] I'm afraid I can't recite poems. It's just, the text is produced in one go. It's just that I'm using a small LLM that is very fast, so it's nearly immediate, and then the audio is produced later. Yeah. Yeah. Thanks for the weight model, by the way. I know that the weights are open. [SPEAKER_01] Is the voice cloning encoder also open? Yeah, there is a small asterisk here. We didn't release this part, the encoder part, which means it's the only thing that is missing for you to clone your own voice. It's a feature that we only serve in a proprietary fashion for now. So what you can do is use the text-to-speech model, use one of the open voices that we provide. We may provide more in the future. Yeah, so far we just didn't want to give everybody the ability to clone any voice. Yeah. [SPEAKER_04] So big labs like Google and OpenAI are working more on native voice-to-voice models, let's say. Yeah. [SPEAKER_04] So your lab, you have a lab, they're doing more cascading architecture. [SPEAKER_04] What's your take on both and how you're seeing things? What's my take? So on the consumer side, you will always have the impression to speak to a single system that hears what you say and outputs something, right? So it's purely an architectural model here. My take on this is that we can go very far by just using speech as an interface, especially because these central LLMs are extremely capable, but they also do a lot of things. So just for the sake of being able to use any agent that has been released with the same interface, it has an advantage to interface. So you can go very far with just interface, especially if you are doing this kind of thing where you stream the text tokens that are output by the LLM. Yeah, last question because I'm out of time, sorry. [SPEAKER_05] Yeah, so the next steps on the interleaved function, which accepts both audio and text, sounds really interesting. [SPEAKER_05] So if you do that real-time, what do you see as the possibilities with that feature? So I didn't say that our next step would be this, right? I just said that there are several patterns to handle a stream of text as input as opposed to a finite amount of text. We actually don't know which one we'll choose, whether it's interleaved or another solution like delayed sequence modeling, for instance. So it's unclear which architecture is best, at least to us, at least to me. What it allows is lower latency because as soon as you have the first bit of text that is produced by the LLM, you can start voicing them out. [SPEAKER_02] So in this agent, that was not clear because the utterances were very short. [SPEAKER_02] But imagine I ask Paul to generate a full page of text. It would be nice if I don't have to wait for the end of the text generation to voice it out. [SPEAKER_05] Yeah, absolutely. Thank you. Hello, Paul, it's me again. Can you say anything like, I don't know, a poem? I'm afraid I can't recite poems. It's just, so the text is produced in one go. It's just that I'm using a small LLM that is very fast, so it's nearly immediate, and then the audio is produced later. Yeah. Yeah. Thanks for the weight model, by the way. I know that the weights are open. Is the voice cloning encoder also open? Yeah, there is a small asterisk here. We didn't release this part, like the encoder part, which means it's the only thing that is missing for you to clone your own voice. It's a feature that we only serve in a proprietary fashion for now. So what you can do is use the text-to-speech model, use one of the open voices that we provide. We may provide more in the future. Yeah, so far we just didn't want to give everybody the ability to clone any voice. Yeah. So big labs like Google and OpenAI are working more on native voice-to-voice, let's say, models. Yeah. So your lab, you have a lab, they're doing more cascading architecture. What's your take on both like driving and how you're seeing things about? What's my take? So on the consumer side, you will always have the impression to speak to a single system that hear what you say and output something, right? So it's purely an architectural model here. My take on this is that we can go very, very far by just using speech as an interface, especially because these central LLM, they are extremely capable, but they also do a lot of things. So just for the sake of being able to use any agent that has been released with the same interface, it has an advantage to interface. So you can go very far with just interface, especially if you are doing this kind of thing where you stream the text token that's output by the LLM. Yeah, last question because I'm out of time, sorry. Yeah, so the next steps on the interleaved function, which accepts both audio and text, sounds really interesting. So if you do that real-time, what do you see as the possibilities with that feature? So I didn't say that our next step would be this, right? I just said that there are several patterns to handle a stream of text as input as opposed to like a finite amount of text. We actually don't know which one we'll choose, whether it's interleaved or another solution like delayed sequence modeling, for instance. So it's unclear which architecture is best, at least to us, at least to me. What it allows is lower latency because as soon as you have the first bit of text that are produced by the LLM, you can start voicing them out. So in this agent, that was not clear because the utterance were very short. But imagine I ask Paul to generate a full page of text. It would be nice if I don't have to wait the end of the text generation to voice it out. Yeah, absolutely. Thank you.