SPEAKER_01
Good morning everyone. Thanks for being here to the Voice and Vision session. So I'm Hervé Brodin, Chief Science Officer and Co-Founder at Pyannote AI. So I'm going to talk to you today about conversations, understanding conversations, and what you can do on top of transcription.
SPEAKER_01
So a quick word about myself. So I've been an academic researcher all my life until two years ago when I started this company. So I worked on this topic called speaker diarization, which I'll introduce a bit later. For those of you who don't know this weird word that is tricky for me to pronounce, over the years I built an open source toolkit called Pyannote, which focuses on speaker diarization. And that became quite popular over the years, in particular since OpenAI released Whisper speech to text open source models. Whisper was some kind of revolution in terms of STT, the fact that it was free, that it was very good, but it didn't provide the actual names and tags of the speaker. So people naturally turned to Pyannote to combine the two. And you can see that on the inflection points on the GitHub Star history. And just a few words, I happened to turn 45 in just one week from now. And we are almost at 10k stars on the GitHub. So please give me a birthday gift by just going to the website, to the GitHub website and add your star. That will make my birthday.
SPEAKER_01
So let's go to the core of the presentation. So you all know what transcription or a speech to text is. Basically, that's answering the question, what was said from the recording or streaming of a conversation. Basically, you get from the audio and you get a sequence of words as output. That's, for instance, what Whisper does. But without the actual names of the speaker, usually the conversations on the transcript are not really understandable. So the next step after transcription is basically to attribute a speaker tag to each of the words. So that's what I call here a speaker attributed transcription. And so that answers the question, who said what? But for some applications, that's enough. Like, for instance, maybe for meeting note takers to make a summary of a meeting, that might be enough to assign an action point to the right people to know exactly that Hervé said that and John said this other thing. But to understand the conversation a bit better, we might want to go slightly further.
SPEAKER_01
So basically, there are many cases where knowing who said what is as important as what was said, actually. So I have here a few examples, like for instance, video dubbing, automatic video dubbing. Knowing who said what is actually really important to put the actual correct voice to the right speaker. So if you want to automatically translate a video from one language to another, you want the voices to be consistent as well. So for that, you need to know who spoke when. Same for meeting note takers. And there are other applications like, for instance, what I call podcast intelligence, like being able to track a speaker across different episodes of a podcast or across different podcasts, finding the same guest across multiple podcasts.
SPEAKER_01
But yeah, so just knowing who said what is sometimes not enough to really understand the conversation, how it goes. So knowing who said what and when actually brings even more information. For instance, in this example, without knowing exactly when each word was pronounced, you can't detect that actually the black speaker interrupted the green one. You can't really detect that maybe this small back channel that the black speaker does during the speech of the green speaker. If you miss it, basically you don't really understand that maybe the black speaker is actually agreeing, noting at what the green speaker says. So knowing exactly when this small word has been pronounced is very important. And same, without precise timestamps, you can't really know whether I make a pause in between two speech terms, which might convey additional information about my state of mind, the point that I'm trying to make in the conversation.
SPEAKER_01
And we can go even a step further. Like, we might want to know who said what and when and how they said it. I have a few examples here. If some of you are laughing at some point, is it because I said something stupid or I was actually funny? Same for coughing. Maybe there's something in the conversation that makes me uncomfortable and I'm coughing for some reason to hide behind that. And yes, I also mentioned stress, discrepancy, prosody. The fact that I stress on particular words in a sentence might actually change completely the meaning of the sentence. I have this example in mind, in the sentence, the dog ate the cake. If you're talking to a child and the child tells you, the dog ate the cake, or the dog ate the cake, or the dog ate the cake, the meaning might be slightly different. So stressing particular words, being able to have a voice AI actually understand this stress and this kind of low-level details might bring some more information to any downstream LLMs or whatever tools that you use after this first enriched transcription step.
SPEAKER_01
And if we take even a step back, knowing who's talking to who, am I here in this example? And probably it can be seen as this gray speaker addressing all of you. But at some point, if at the end of this talk, you do have questions, then I will probably be the green one talking to one particular person answering the particular question. So knowing that, having a higher level view of the conversation, it also opens up a new bunch of applications.
SPEAKER_01
And all of this also happens in an acoustic environment that gives you information about the context, whether it's in this quiet room for now, or if I'm in the street. This brings more context to the game and you can actually take a better decision about what's going on, why a person is actually saying what they are saying, and so forth.
SPEAKER_01
So all of this is basically what we are trying to work on at Pyannote AI, but we started with a smaller problem called speaker diarization. And that's really what I want to talk about today. So speaker diarization is usually what I wanted to say about this slide is that who said what is usually just as important as what was said. And you can tell by, for instance, going through to the Hugging Face model repository and filter all the models there with the ones that have the audio tag. Here. So this is what I did. And then you saw them by number of downloads. And basically, among the top here, seven models that are at the top of this list, three of them, the first three are actually related to speaker identity and speaker diarization. And the other ones are STT. So that's one way of showing that speaker diarization is actually quite important in the community.
SPEAKER_01
So going into a bit more detail about what speaker diarization is. So as I said, it's answering the question, who speaks when? So starting from the recording of a conversation, basically, the first step that people usually do is a basic voice activity detection. So you basically the audio tag. Where is Marmas? Yeah, here. So this is what I did. And then you saw them by number of downloads.
SPEAKER_01
Among the top here, seven models that are at the top of this list, three of them, the first three are actually related to speaker identity and speaker diarization. And the other ones are STT. So that's one way of showing that speaker diarization is actually quite important in the community. So going into a bit more detail about what speaker diarization is. So as I said, it's answering the question, who speaks when? So starting from the recording of a conversation, the first step that people usually do is a basic voice activity detection. So you tell whether someone is speaking at one time or not, anyone is speaking at one time or not. And then you can go a step further, you can actually segment those speech regions into smaller speech turns. So this is what I call segmentation. And it includes finding speaker change points between into those longer in those long speech regions, you might find a speaker change point. And also find regions where people are actually interrupting each other, like in the first example here, there's definitely here someone interrupting someone else. And it's possibly here some kind of back channel where someone is actually nodding or saying, OK, this kind of stuff. And this is the kind of small speech turn that you don't want to miss because sometimes they actually convey the most important information of the conversation. If it's a small yes here, and if you miss it, then you have no idea what's the state of mind of this speaker. But that's not yet speaker diarization. Speaker diarization goes all the way to actually assigning a speaker identity to each speech turn. Speaker diarization. And so here in this example, the speaker diarization system automatically detected that there are two speakers, the green one and the black one. But we usually don't have any prior knowledge on the actual number of speakers in the conversation. You might have some kind of guidance on, I don't know, if you are a meeting notetaker, for instance, you have the list of attendees. So you might have an idea of the number of people that were invited, but that doesn't mean that someone, two people joined from the same channel or that an attendee that was not invited actually joined in finally. And so that makes the problem difficult in the sense that we don't know in advance the number of classes that we're supposed to detect, contrary to other, let's say, classical machine learning problems. And also we don't know the identity of the speaker that we're supposed to diarize in the sense that speaker diarization does not really output John, Hervé, and Jack. But it's more like speaker one, speaker two, speaker three, and they can be permutated and it's still correct. For example, in this example, if I just change green into black and black into green, the diarization is still exactly the same from the point of view of the evaluation metrics that I'm going to talk about and from the point of view of the actual speaker diarization task.
SPEAKER_01
So that makes this problem difficult. And that's why even though the community has been working on this topic for a long time now, it's still not solved. There are other reasons like the fact that we have to detect overlapping speech to handle overlapping speech that we have to take into account very short speech turns that have to take into account imbalance between the speech time of multiple speakers in the conversation, all of that makes the problem difficult. On top of the usual acoustic conditions, problems and everything that any speech processing task has to deal with.
SPEAKER_01
So I'm going to switch to a demo to show you a bit more how the diarization systems are evaluated. Maybe when you stumble into benchmark, you find this DER acronym, which stands for diarization error rate. So hopefully this will work. So I have here a Python notebook that I've prepared for this talk. It's actually already available on GitHub. I'll show you the QR code at the end of the presentation so you can play with it. But so in this example, I have a conversation between two women talking over the phone.
SPEAKER_01
Hello? Hello? Hello? Oh, hello. I didn't know you were there. Neither did I. Okay, I thought, you know, I heard a beep. This is Diane in New Jersey. And I'm Sheila in Texas, originally from Chicago. Oh, I'm originally from Chicago also. I'm in New Jersey now though. Okay, so you get the idea. A 30 seconds conversation between two women. So this is the expected output of a perfect speaker attributed transcription. This was manually labeled. And now what I'm going to do here is actually run a Pyannote open source community one model on the same file. So it's actually running right now on my Mac. So that's why here I'm using this MPS PyTorch backend. So how it works is that you first download the model from Hugging Face with this line of code and simply apply it on the audio file. So now it's finished and there's this prediction variable that contains... Sorry, I should have done that on the other tab because this was the one that was already processed. I wanted to do it live. So let's do it live now.
SPEAKER_01
So it's running. And so the next step is to actually visualize the errors that this system made. So at the top here, you have the reference annotation. So the expected output of the diarization. And here is the output of a Pyannote community one speaker diarization pipeline that is open source and free to use. So it makes three kinds of mistakes, as I was saying, like a confusion here. It got the speaker wrong. It can also make false alarms. So false alarm is when it detects speech when there's actually no speech in the ground truth. And misdetection is the other way around. And misdetection can actually happen, for instance, here during overlapping speech when we actually detect only one of the two speakers here.
SPEAKER_01
And then once we have these errors of false alarm confusion and misdetection, we can actually compute the diarization error rate. So this is with this library called Pyannote metrics. And basically it gives us in this example, a 5% diarization error rate, which is basically the sum of the confusion, false alarm and misdetection, divided by the total duration of speech in this file.
SPEAKER_01
And so we have this other model that I'm currently running. Hopefully it works. So it's running right now on our cloud API. Basically, it's a better model than community one that we call Precision two that is running here. Hopefully it will work. But as for every demo, it might fail. If no, it worked. And this is the kind of output that you can get. So you see that it got the speaker right here and makes slightly less mistakes in terms of false alarm and misdetection. And overall, in this example, I think we got a 3% diarization error rate here.
SPEAKER_01
So let me switch back to the presentation. So this was about benchmarking speaker diarization models. And I'm often asked how well the state of the art speaker diarization works today. And it's a difficult question to answer because really depending on the use case, it might vary completely. For instance, in this example, we have here conversation telephone speech like the one we just listened to like 2% over the phone. Hopefully it will work. But as for every demo, it might fail. If not, it worked. And this is the kind of output that you can get.
SPEAKER_01
So you see that it got the speaker right here and makes slightly fewer mistakes in terms of false alarm and misdetection. And overall, in this example, I think we got a 3% diarization error rate here. So let me switch back to the presentation. So this was about benchmarking speaker diarization model. And I'm often asked how well the state of the art speaker diarization works today. And it's a difficult question to answer because it really depends on the use case. It might vary completely. For instance, in this example, we have a here conversation telephone speech like the one we just listened to, like 2% over the phone.
SPEAKER_01
And we can go down to 8% diarization error rate. The best system does that. But if you are now in a restaurant with many friends with lots of background noise, even the best system reaches 41% diarization error rate. So it's far from being a solved problem. But we are working on that. And so once you have speaker diarization and you have transcription, it shouldn't be that hard to actually find a speaker attributed transcription, right? It's just a matter of assigning a word to a speaker, a speaker to a word. But actually, that's not that easy.
SPEAKER_01
And there are many reasons why that's not that easy. The first reason is that most speech to text models are actually trained on single speaker data. As soon as you apply them on multi speaker data with overlap, with speaker change, with all kinds of mess, they fail miserably. So for instance, when you look at the open ASR leaderboard from Hugging Face, if we look, for instance, at Nvidia Parakeet that we are using, that I will be using in a demo right after that, they report 11.4% word error rate. When we apply the very same model on the same AMI on our side, we get 26%. So the question is, why is there a difference?
SPEAKER_01
Actually, it's in the way that those benchmarks are set up here. For example, in this AMI data set, that's a data set with meetings between four to five people, I think, in meeting rooms. And they basically have microphones like the one I'm having here, so a headset microphone, as well as one microphone in the middle of the table. And those numbers here on the open ASR leaderboard are based on the headset microphone. Those numbers here are based on the microphone that is in the middle. And so on one side, you have a single speaker speech. On the other side, you have multi speaker and even distant microphone speech. So that's why it degrades a lot.
SPEAKER_01
And so that was my point of this slide. So when doing speaker attributed transcription, the reason why it might go wrong is either because STT doesn't work great, and it's usually the case that they don't generalize very well to multi speaker recordings for, as I was saying, distant microphone, speaker change, crosstalk, interruptions, you name it. Maybe code switching as well when you change language in the middle of a sentence. But also it might be because of the actual reconciliation between diarization and STT timestamps. Though it may sound obvious how to do that, as I was saying, I will hopefully show in this live demo once again that it's not such an easy problem.
SPEAKER_01
Because STT does not transcribe overlapping speech well, because the timestamps disagree between STT and diarization, and because sometimes diarization will detect speech that the transcription will not transcribe, and the other way around. So, second part of the demo about speaker attributed transcription. So, first I'm going to apply this Parakeet model from NVIDIA that does, by the way, a great job at transcribing this sample. And so Parakeet gives us this kind of output. So we have this sequence of words, and for each word we actually have the corresponding timestamps. Okay? So I can play it quickly. Hello. Oh, hello. I didn't know you were there.
SPEAKER_01
Neither did I. Okay. I thought, you know, I heard a beep. This is Diane in New Jersey. Well, you get the idea. And then the question is, now we have our best diarization. This is Pyannote output here.
SPEAKER_01
This is Parakeet output here. And the question is, now we need to assign a speaker to each word. So let's do it step by step, for instance. So in this example, even though the timestamps of the words are a bit shifted on the left here, it's obvious that it's probably the yellow speaker who spoke the gray word. But if you go just the third word. So now we have this. So we have this word, the word O here, that is in between those two speech turns according to the diarization. Which one do you assign it to? So we, I can listen to it. Oh, hello.
SPEAKER_01
Oh, hello. Oh, hello. So here I just submitted to our cloud API a job to transcribe and use both Pyannote diarization model and Parakeet transcription model and take care of the, what I call this reconciliation between the two. And you end up with this kind of output. What's nice is that, now you, it did the job for you. But in particular, I like to go to this particular example. Even where there was overlapping speech before it actually managed to interleave the words from the two speakers. Let me play just this part and focus on this area maybe. In New Jersey. And I'm Sheila and tech.
SPEAKER_01
And there's actually an end that, or this is the hum here that is actually a, the, the, the, the yellow lady actually has this hum at the end of their speech. And the blue one actually interrupts her. But they are actually overlapping and we managed to get them right. Yeah. So this is the end of my talk. [SPEAKER_02] I wanted to give you a few minutes. [SPEAKER_02] I have two minutes left for questions, but I wanted to say that this demo that I did is already available on the piano.ai slash tutorials GitHub repo.
SPEAKER_01
[SPEAKER_02] So in there you'll be able to play with Pyannote open source toolkit for diarization, with piano.metrics for evaluation, with IPyNOT for this nice visualization widget that we've implemented. [SPEAKER_02] I stands for, not iOS, iMac or whatever, but really for interactive. [SPEAKER_02] And also the SDK to play with our premium models. [SPEAKER_02] And I'll stop here and I'm happy to take a few questions. We have one more minute left. [SPEAKER_02] Yes, please. [SPEAKER_02] So what's the trick that you use to resolve the confidence? Is it part of the model training or is it some sort of heuristics that you have around to identify who to attribute the word to?
SPEAKER_01
Yeah. [SPEAKER_02] With IPNOT for this nice visualization widget that we've implemented. [SPEAKER_02] I stands for interactive, not iOS, iMac or whatever. [SPEAKER_02] And also the SDK to play with our premium models. [SPEAKER_02] I'll stop here and I'm happy to take a few questions. We have one more minute left. [SPEAKER_02] Yes, please. [SPEAKER_02] So what's the trick that you use to resolve the confidence? Is it part of the model training or is it some heuristics that you have around to identify who to attribute the word to? Yeah. So the question is what's the trick to actually solve this problem of reconciliation?
SPEAKER_01
That's a proprietary trick, but what I can say is that there's already part of this trick that is available in the community one model, which we call exclusive diarization. Basically, what we do is find a way when there is overlap to actually select the most likely of the two speakers that will be transcribed by the STT model. So that simplifies the reconciliation between the two. That's one major part of the approach. The model training is like a couple of heuristics you have to. Yeah. So the question is about, I'm repeating because I've been told I need to repeat questions.
SPEAKER_01
It's not part of the model training in the sense that we really plan to support any kind of STT without having to change the STT model itself. So really it's supposed to work with any STT, even fine-tuned ones that you might have internally because you fine-tuned them for your particular use case and nobody has it. You can combine it like that. All right. Yeah. 18 seconds. So I guess we need to stop. I'm sorry. Maybe we can talk offline. Otherwise they'll beat me, I guess. Thank you very much.
SPEAKER_01
I'm sorry. Thank you. here, you have the reference annotation. So the expected output of the realization. And here is the output of a community one speaker authorization pipeline that is open source and free to use. So it makes three kinds of mistakes, as I was saying, like a confusion here. It got the speaker wrong. It can also make force alarms. So force alarm is when it detects speech when there's actually no speech in the ground truth. And misdetection is the other way around. And misdetection can actually happen, for instance, here during overlapping speech when we actually detect only one of the two speakers here.
SPEAKER_01
And then once we have these basically errors of force alarm confusion and misdetection, we can actually compute the diarization error rate. So this is with this library called Pianot metrics. And basically it gives us in this example, a 5% diarization error rate, which is basically the sum of the confusion, force alarm and misdetection, and misdetection, divided by the total duration of speech in this file. And so we have this other model that I'm currently running. Hopefully it works. So it's running right now on our cloud API. Basically, it's a better model than community one that we call precision two that is running here.
SPEAKER_01
Hopefully it will work. But as for every demo, it might fail. If no, it worked. And this is the kind of output that you can get. So you see that it got the speaker right here and makes slightly less mistakes in terms of force alarm and misdetection. And overall, in this example, I think we got a 3% diarization error rate here. So let me switch back to the presentation. So this was about benchmarking speaker diarization model. And I'm often asked how well the state of the art speaker diarization works today.
SPEAKER_01
And it's a difficult question to answer because really depending on the use case, it might vary completely. For instance, in this example, we have a here conversation telephone speech like the one we just listened to like 2% over the phone. And we can go down to 8% diarization error rate. The best system does that. But if you are now in a restaurant with many friends with lots of background noise, even the best system reaches like 41% diarization error rate. So it's far from being a solved problem. But we are working on that.
SPEAKER_01
And so once you have speaker diarization and you have transcription, shouldn't be that hard to actually find a speaker attributed transcription, right? It's just a matter of assigning a word to a speaker, a speaker to a word. But actually, that's not that easy. And there are many reasons why that's not that easy. The first reason is that most speech to text models are actually trained on single speaker data. As soon as you apply them on multi speaker data with overlap, with speaker change, with all kinds of mess, they fail miserably.
SPEAKER_01
So for instance, when you look at the open ASR leaderboard from hugging face, if we look, for instance, at Nvidia Paraket that we are using, that I will be using in a demo right after that, they report 11.4% word error rate. When we apply the very same model on the same AMI on our side, we get 26%. So the question is, why is there a difference? Actually, it's in the way that those benchmarks are set up here. For example, in this AMI data set, that's a data set with meetings between four to five people, I think, in meeting rooms. And they are basically microphones like the one I'm having here, so a headset microphone, as well as one microphone in the middle of the table.
SPEAKER_01
And those numbers here on the open ASR leaderboard are based on the headset microphone. Those numbers here are based on the microphone that is in the middle. And so on one side, you have a single speaker speech. On the other side, you have multi speaker and even distant microphone speech. So that's why it degrades a lot. And so that was my point of this slide. So when doing speaker attributed transcription, the reason why it might go wrong is either because STT doesn't work great, and it's usually the case that they don't generalize very well to multi speaker recordings for, as I was saying, distant microphone, speaker change, crosstalk, interruptions, you name it.
SPEAKER_01
Maybe code switching as well when you change language in the middle of a sentence. But also it might be because of the actual reconciliation between diarization and STT timestamps. Though it may sound obvious how to do that, as I was saying, I will hopefully show in this live demo once again that it's not such an easy problem. Because STT does not transcribe overlapping speech well, because the timestamps disagree between STT and diarization, and because sometimes diarization will detect speech that the transcription will not transcribe, and the other way around.
SPEAKER_01
So, second part of the demo about speaker attributed transcription. So, first I'm going to apply this Paraket model from NVIDIA that does, by the way, a great job at transcribing this sample. And so Paraket gives us this kind of output. So we have this sequence of words, and for each word we actually have the corresponding timestamps. Okay? So I can play it quickly. Hello. Oh, hello. I didn't know you were there. Neither did I. Okay. I thought, you know, I heard a beep. This is Diane in New Jersey. Well, you get the idea. And then the question is, now we have our best diarization. This is Precision 2 output here. This is Paraket output here.
SPEAKER_01
And the question is, now we need to assign a speaker to each word. So let's do it step by step, for instance. So in this example, even though the timestamps of the words are a bit shifted on the left here, it's obvious that it's probably the yellow speaker who spoke the gray word. But if you go just the third word. So now we have this. So we have this word, the word O here, that is in between those two speech turns according to the diarization. Which one do you assign it to? So we, I can listen to it. Oh, hello. Oh, hello. Oh, hello.
SPEAKER_01
So here I just submitted to our cloud API a job to transcribe and use both Precision 2 diarization model and Paraket transcription model. And take care of the, what I call this reconciliation between the two. And you end up with this kind of output. What's nice is that, so now you, it did the job for you. But in particular, I like to go to this particular example. Even where there was overlapping speech before it actually managed to interleave the words from the two speakers. Let me play just this part and focus on this area maybe. In New Jersey. And I'm Sheila and tech. And there's actually a end that, or this is the um here that is actually a,
SPEAKER_01
the, the, the, the, the yellow lady actually has this hum at the end of their speech. And the, the blue one actually interrupts her. But they, they, they are actually overlapping and we managed to get them right.
SPEAKER_01
Yeah. So this is the end of my talk.
SPEAKER_02
I wanted to give you a few minutes. I have two minutes left for questions, but I want, just wanted to say that this demo that I did is already available on the, in the piano.ai slash tutorials, uh, GitHub repo. So in there you, you, you'll be able to play with piano.to.do open source toolkit for dialization, with piano.metrics for evaluation. Uh, with IPNOT for this nice, uh, visualization widget that we've, uh, implemented. I stands for, uh, not iOS, iMac or whatever, but really for interactive. And, uh, also the SDK to, to play with our premium models. And, uh, I'll stop here and, uh, I'm happy to take, uh, a few questions.
SPEAKER_01
We have one more minute left.
SPEAKER_02
Yes, please. So what's the trick, uh, that you use to resolve the confidence? So part of the model training or is it some sort of like heuristics that you have around to identify who to attribute the word to?
SPEAKER_01
Yeah. So the question is, uh, what's the trick to, uh, actually solve this problem of a reconciliation? Uh, so that, that's a, a proprietary trick, but what I can, what I can say is that we, uh, there's already part of this trick that is available in the community one model, which we call exclusive diarization. Basically, uh, what we do is that we find a way when there is overlap to actually select, uh, the, the most, uh, likely of the two speaker that will be transcribed by the STT model. So that simplifies actually the reconciliation between the two. So that, that's one major part of the, uh, of the approach.
SPEAKER_01
The model training is like a couple of heuristics you have to. Yeah. So, so the question is about, uh, I repeating because I've been told I need to repeat questions. Uh, so it's not part of the model training. In the sense that we really plan to support any kind of STT without having to change the STT model itself. So really it's supposed to work with any STT, uh, even fine tune ones that you might have internally for, because you, you fine tune them for your particular use case and nobody has it. Uh, you can combine it like that. All right. Yeah. 18 seconds. So I guess we need to stop. I'm sorry. Maybe we can talk offline.
SPEAKER_01
Otherwise, uh, they'll, they'll beat me, I guess. Thank you very much. I'm sorry. Thank you.