SPEAKER_10
Sous-titrage Société Radio-Canada pour les gens qui ont construit des agents et des solutions. Nous ne travaillons pas sur orchestration, nous ne travaillons pas sur les verticules. Nous faisons juste des blocs pour les gens qui veulent construire des AI.
SPEAKER_10
Je peux laisser ça par un podcast que j'ai cloné ce matin.
SPEAKER_10
Il faut avoir audio qui est en train de faire.
SPEAKER_02
Vous avez vu ce que Gradium est en train de faire avec voice cloning? C'est un peu crazy, sérieusement.
SPEAKER_10
Vous avez 10 secondes de votre voix, c'est... Vous avez beaucoup plus de volume, s'il vous plaît.
SPEAKER_02
Vous avez vu ce que Gradium est en train de faire avec voice cloning? C'est un peu crazy, sérieusement. Vous avez vu 10 secondes de votre voix, c'est ça. Et le système analyse le tone, le pitch, l'accent, les petits quirks qui font votre voix. Et vous typez text et il parle...
SPEAKER_10
Ok, je pense que vous reconnaissez Joe. Je pense que vous avez vu. Donc, c'est un spin-off d'un lab que nous avons créé il y a deux ans, avec des fondations d'un philanthropiste, avec Eric Schmidt, Rodolphe Sadie et Xavier Niel. The main idea was to create a lab that has open research. And we focused mostly on speech. So we developed Moshi, which was the first speech-to-speech model for conversation, speech-to-speech translation, Pocket TTS, most recently a CPU model. And we decided to also create this for-profit structure to make products that can be used in production beyond open source. So the goal of the talk is based on the Her movie.
SPEAKER_10
So this has been the most overused, most annoying analogy, I think, in the field. At the same time, it's extremely relevant, because it's 13 years old. And if we look at one of the introduction scenes, so that's when the main character meets his AI voice, Samantha, for the first time, it still sounds like it was anticipating, I think, what interaction could look like.
SPEAKER_01
Oh, what do I call you? Do you have a name? Yes. Samantha.
SPEAKER_04
Where did you get that name from? I gave it to myself, actually.
SPEAKER_01
How come?
SPEAKER_04
Because I like the sound of it. Samantha.
SPEAKER_01
And then came the trend of the Her moment.
SPEAKER_04
And we got so many Her moments on Twitter. And in real life, again, I don't want to be mean to anyone,
SPEAKER_10
so I will also make fun of myself. It's more trying to be pragmatic about what was the promise and where we are right now. So this is a very recent demo from, for me, what is the best voice AI company in the world, which is Eleven Labs.
SPEAKER_10
Hello, you're speaking to the government's AI helper. How can I help you today? I would like to start a new business.
SPEAKER_03
To start a new business, you'll need to choose a business structure and name, complete the registration and upload required documents. And this one is a demo that I did a few months ago with Richemini. So our presenter this morning, but it was very risky. I recognize. I was very scared when I did it.
SPEAKER_10
But basically putting our streaming voice models into Richemini. Action. Just say the word. Okay. Maybe you can do something a bit fun. I want to improve my health overall. And so I'm looking for a bro to go to the gym.
SPEAKER_09
Can you take the personality and the voice of a gym addict? Hey, I'm Logan, your gym bro. Let's crush those gains together. You ready to lift, sweat and feel awesome?
SPEAKER_10
I've got your back. No excuses, just results. So, okay. This is fine. It sounds more natural than it used to.
SPEAKER_07
But in both cases, we're still not there, right? The latency is still quite high. The ability to handle simultaneous speaking between the user and the system is not there. Intelligence starts to become much better.
SPEAKER_10
And I think that's also why we see all this traction around voice agents because there are agents to whom we can give a voice and they can be useful. And also, this is mostly just glorified text model with a voice around it. And so, anything that is not in the text will not be able to be leveraged. And maybe I could ask one last time to increase the volume a bit. I think that could be even better. And so what does it take to get there? So basically, we have a very nice presentation this morning by Samuel about cascaded systems. So I will go quickly. Speech-to-text, LLM text-to-speech, that's the classic cascade.
SPEAKER_10
In our case, we do streaming speech-to-text, streaming text-to-speech with voice cloning, semantic VAD, the classic stuff. So latency, getting to a fast conversation. So we have a very fast TTS. That's the latency for our TTS compared to a few other models. But, that means that just the TTS is still more than 200 milliseconds, while in a human conversation, you need the entire stack of understanding, producing an answer and pronouncing it to be around 200 milliseconds. So none of this will allow conversations that sound human. And this is just latency for text conversation. There is no tool call, no actual TTS that is performed. Now that's another thing from the movie.
SPEAKER_10
OK, let's start with your emails. You have several thousand emails regarding LA Weekly, but it looks like you haven't worked there in many years. Oh, yeah. OK, so here it just went instantly into all the emails and gathered information and so on and so forth. Obviously, we are not there at all. Today, if you have a voice agent that is supposed to use a tool,
SPEAKER_04
call, you're going to wait. The tool call, we are fighting for latency of the TTS, [SPEAKER_10] Now that's another thing from the movie.
SPEAKER_10
OK, let's start with your emails. You have several thousand emails regarding LA Weekly, but it looks like you haven't worked there in many years. Oh, yeah. OK, so here it just went instantly into all the emails and gathered information and so on and so forth. Obviously, we are not there at all. Today, if you have a voice agent that is supposed to use a tool call, you're going to wait. [SPEAKER_04] The tool call, we are fighting for latency of the TTS, trying to grab 10 milliseconds, 20 milliseconds. And then you have a tool call or open router that is going to have a latency between 500 milliseconds and four seconds.
SPEAKER_10
So in a way, sometimes you feel like we're fighting for something that is not the most relevant anymore. I think now the main bottleneck is becoming the tool call, which is very unpredictable. And we should have models that are resilient to very complex tool calls. So what we are, one solution to do that is to have fillers. So URLLM, it splits into two things. So it sends a tool call. And while it waits for getting the result back, it can keep the conversation going in a natural way. And then it retrieves the result and it tries to insert it back naturally in the conversation.
SPEAKER_10
So just because I always do live demos, I will do a very short improvised live demo with something that was very quickly so it may go very wrong. But the idea here is to have a vibe coded travel agent. And that can, when I ask it to retrieve places, it's going to try to find a nice thing to say about the location I'm going to. Hello. This is Colin from Wanderlust Travel. I see you're looking to book a trip for two people from April 10th to the 13th, 2026. Where would you like to travel to for your getaway? I want to go to Tokyo. Tokyo is such an incredible choice. It's a fascinating mix of ultra-modern skyscrapers and beautiful, peaceful shrines.
SPEAKER_10
I found some fantastic options for you. You could stay at the Fairmont for a serene... [SPEAKER_13] So this needs polishing.
SPEAKER_13
The main idea is that while it was trying to gather things and doesn't really know how long it will take to get the data, it tries to say a nice thing about the place you're going to. So that's one way to get latency that is going to be more controlled and more reliable, even despite the complexity of the thing. But then we know that latency of cascaded system is inherently high, right? And so then what I hear a lot from customers, from investors and so on, is what about speech to speech? [SPEAKER_10] So I think a big confusion is what is speech to speech?
SPEAKER_13
[SPEAKER_10] So speech to speech is the idea that now instead of having the three blocks, you only have one that does everything together, right? [SPEAKER_10] So instead of having audio that goes in speech to text and then text to speech, you have a model that takes speech as input and is going to output speech. [SPEAKER_10] And that reduces latency a lot, but that's still not a human conversation. [SPEAKER_10] In particular because every single speech to speech model except Moshi is half duplex.
SPEAKER_10
What that means is even the best speech to speech model, we could argue, maybe that's the advanced voice model of OpenAI or SESAME. I'm a huge fan of the SESAME voice model. It's still half duplex. What it means is that the model is either listening or it's speaking. And it cannot handle the ambiguity of a human conversation where you can have overlap between people speaking on one another, a lot of things happening, you cough, you just do, which is called backchanneling, and then it breaks completely. Full duplex, that's a human conversation.
SPEAKER_10
It varies between cultures and languages, but for example in Japanese, it's a sign of politeness and that you are actively listening to do a lot of backchanneling. So you say, mm, mm, ah, ah, ah, ah, ah, constantly when the other person is speaking. And you get up to 20% of the time that is overlapped between the people, right? So that's what makes a conversation human. And so I shot this video in my hotel room this morning just to show that that's a speech to speech model, but that's not a full duplex model. And you see how it can become annoying. Hey, how's it going? Hey, I'm doing really well. Thanks for asking. How about you? How's your day going so far? I'm great.
SPEAKER_10
You know, I'm preparing to talk about voice AI and how far we are from the Her movie. And, yeah, I'd just like to brainstorm a bit about it with you. [SPEAKER_09] Oh, that's a great topic. [SPEAKER_05] Yeah, I'd love to help you brainstorm. [SPEAKER_05] Are you thinking more about it? [SPEAKER_05] Exactly. [SPEAKER_05] Yeah. [SPEAKER_09] No, I didn't mean to interrupt. [SPEAKER_09] You know, I was just saying, yeah, that. [SPEAKER_09] You can just keep going. [SPEAKER_09] You know, don't mind me. [SPEAKER_05] That's something I typically do. [SPEAKER_05] No worries at all. [SPEAKER_05] Yeah. [SPEAKER_05] I was just going to say we could break it down into a few aspects.
SPEAKER_09
Yeah, exactly.
SPEAKER_05
[SPEAKER_09] No, please stop interrupting. [SPEAKER_09] You know, it's called backchanneling. [SPEAKER_09] Humans do it all the time. [SPEAKER_09] It shows that you're just following the conversation, that you don't interrupt you in your flow.
SPEAKER_09
[SPEAKER_05] Just keep going. [SPEAKER_05] You're going to be. [SPEAKER_05] Ah, got it. [SPEAKER_05] Thanks for letting me know.
SPEAKER_05
[SPEAKER_09] No problem. [SPEAKER_09] Oh, come on. [SPEAKER_09] Okay, so I was a bit mean, right? [SPEAKER_09] That's my point. [SPEAKER_09] It's not an actual conversation, right?
SPEAKER_09
And it's going to become very annoying in particular. [SPEAKER_05] That's also why, you know, most of the voice AI demos, they are shot in empty, a quiet room next to the phone and so on. [SPEAKER_05] So a lot of things can break. [SPEAKER_05] Just keep going. [SPEAKER_05] You're going to be.
SPEAKER_05
Ah, got it. Thanks for letting me know. [SPEAKER_09] No problem. [SPEAKER_09] Oh, come on. [SPEAKER_09] Okay, so I was a bit mean, right?
SPEAKER_09
That's my point. It's not an actual conversation, right? And it's going to become very annoying in particular. [SPEAKER_05] That's also why most of the voice AI demos are shot in empty, a quiet room next to the phone and so on. [SPEAKER_05] So a lot of things can break. So, what we did instead, in our case was, sorry, where am I in my presentation? [SPEAKER_05] I'm right here with Moshi, was the forceful duplex system.
SPEAKER_05
[SPEAKER_10] So here you'll see my co-founder Alex talking to it. [SPEAKER_10] It's almost two years old now. [SPEAKER_10] I think it's still aged well because what you'll see is they are going to talk on one another constantly and it's just fine. [SPEAKER_10] So, the planet is Sirius 22. Can you plot a trajectory course to it, please? [SPEAKER_10] Yes, sir.
SPEAKER_10
Okay. How long is it going to take us to get there? I've mapped it out. It's approximately five months to get there. Okay, that's not too bad. Do you think we have all we need on board the ship to start the mission? Yes, sir. We have everything we need. Okay. So, even when the model has guessed what you're going to say, it starts answering before you're done. At the same time, you can talk over it and it's not ignoring what you're saying. It really considers it afterwards and so on. [SPEAKER_08] So, you have what is the most robust conversational experience to this day, robust to noise, to a lot of people speaking and so on and so forth.
SPEAKER_10
[SPEAKER_08] But now, if we compare it to the Her movie. [SPEAKER_08] So, do you know what I'm thinking right now? [SPEAKER_08] Well, I take it from your tone that you're challenging me. Maybe because you're curious how I work? Do you want to know how I work? [SPEAKER_08] Yeah, actually. [SPEAKER_00] So, maybe this was not very clear, but this snippet here, it's the AI understanding that the character is a bit uncomfortable.
SPEAKER_08
[SPEAKER_00] So, that's paralinguistic understanding. It's interesting all the cues that come from the way people speak. Technically, that is in Moshi, that is in any speech-to-speech model because this information is not lost. However, if you don't exploit this information to make your model say irrelevant things, it's never going to exploit it. [SPEAKER_00] If you train it on the audio version of an Instruct dataset and it's just factual question answering, why would it even try to capture this information? [SPEAKER_00] So, Moshi, I think we saw the only full duplex model. Recently, NVIDIA published the personal duplex model based on it.
SPEAKER_00
[SPEAKER_08] What was great is the flow of it is just honestly impossible to match. [SPEAKER_10] It's conversational and very robust. At the same time, the model was very stupid.
SPEAKER_08
[SPEAKER_10] So, it was just useless. You could take it for a few minutes and then it was a bit pointless, because it was not an agent. [SPEAKER_10] It had no tool code, no ability to do anything. It's impossible to use in production something that has no observability.
SPEAKER_00
[SPEAKER_10] You don't know if it's very hard to detect if someone said something that should not be accepted and so on and so forth. [SPEAKER_10] And there was no real paralinguistic understanding. So, the main takeaway for me is, we know that this nature of interactive models that are going to be full duplex, that are going to be really the way to get an interaction that is as natural as you would have with a human.
SPEAKER_08
[SPEAKER_04] But as long as we are not able to give to this kind of very natural sounding models the same level of reliability, intelligence, and personalization as cascaded systems, I don't see a path towards them replacing cascaded systems.
SPEAKER_10
[SPEAKER_04] So, I used to be really at war against cascaded systems. I think they are so practical and so convenient. [SPEAKER_04] The main challenge, honestly, I think we solved that with Moshi. Anybody who implements it and trains it on beta data will have something that sounds just indistinguishable from humans. But all of this is going to happen here. The last point is the scalability. So, now, let's say you have the best speech-to-speech model. OK. It solves all the stuff that I've talked about.
SPEAKER_04
[SPEAKER_10] If we take the analogy from the movie again, you will talk to it maybe several hours a day. [SPEAKER_10] Or it will always be on because it's on your computer when you work and you asynchronously ask things to it and so on and so forth. [SPEAKER_10] I'm not going to mention the cost of the API of our competitors, but voice is very expensive.
SPEAKER_10
Everybody in this room probably knows it. The voice mode of most hyperscalers is run at a loss. It's a gigantic multimodal model and they lose money every time you use it. But it's a marketing thing and so on. It's fine. But now, if we want to make it an actual profitable product that people are using at a massive scale, it's just not going to work. And in particular, anyone who tried developing a consumer app with voice realized that LLM now is almost nothing in terms of cost. All the bill is TTS. Speech-to-text is very cheap as well. Diarization is affordable. TTS is really what is going to consume most of the...
SPEAKER_10
And I saw people burning their fundraising in TTS bills and they don't even get the opportunity to get their user base to grow. So another aspect is privacy. The more you're going to open to your AI, the more you'll want it to be more controlled and private and feel more comfortable that things are not shared publicly. In particular, we see now people with mythos being afraid that any single database is going to be hacked in a few months. And so you will be more comfortable if all your private data is local. And so to solve that, our first step is Gradium Phonon. It's an on-device TTS.
SPEAKER_10
So on-device means a lot of things. For some people, on-device means it runs on a gamer GPU. For us, on-device means it runs on a smartphone CPU. And so it's a very small model, less than 100 million parameters. And for its size, it works quite well. It's better than all the existing on-device models. Kokoro is a good one, but it doesn't have voice cloning. And I'm out of time, so I will just play a short demo if I can. Yes. Ah, geez, Morty stopped looking for a signal. Gradium Phonon runs locally on the CPU, which means high fidelity neural speech right here without those intergalactic cloud government hacks.
SPEAKER_10
It is about simplicity with no servers and no waiting, just a smooth, quiet performance that stays on the device with local processing and total privacy. What else? Local processing, that sounds the machine is making its own donut. Wait, if the CPU is doing all the work, can it sprinkle some? And for its size, it works quite well. It's better than all the existing on-device models. Kokoro is a good one, but it doesn't have voice cloning. And I'm out of time, so I will just play a short demo if I can. Yes.
SPEAKER_10
Ah, geez, Morty stopped looking for a signal. Gradium Phonon runs locally on the CPU, which means high fidelity neural speech right here without those intergalactic cloud government hacks. It is about simplicity with no servers and no waiting, just a smooth, quiet performance that stays on the device with local processing and total privacy. What else? Local processing, that sounds like the machine is making its own donut. Wait, if the CPU is doing all the work, can it sprinkle some?
SPEAKER_10
Yeah, so this runs on a smartphone CPU, which means that you can use that to power any kind of voice application without paying a single cent of API fee. So we opened the private beta for this model. The goal for us is to allow people to create consumer apps with voice, and they are able to scale the usage without having to lose money on API. So the conclusion, the path forward is for us, I have a strict opposition to some of our competitors who say that voice is a commodity now. I think it's completely false. Voice is very challenging. The last minute is going to be the most difficult to solve. And for us, it's really about science and engineering.
SPEAKER_10
And it will get her to us, to the her movie, sorry. So you can use us on Gradium.ai. And if you want to join us to bring her to life. Yes, I'm using this analogy. No, just if you want to work on very exciting voice models, you can apply and join us. Thanks a lot for your attention. Thank you. And it will get her to us, to the her movie, sorry. So you can use us on Gradium.ai. And if you want to join us to bring her to life. Yes, I'm using this analogy. No, just if you want to work on very exciting voice models, you can apply and join us. Thanks a lot for your attention. Thank you.