Open Reader

Beyond Transcription: Building Voice AI That Understands Conversations — Hervé Bredin, pyannoteAI

completed 25:19 Jun 05, 2026 Watch on YouTube

Current Status

completed

Video ID

mFLlVpnGpds

RAG / Chat

Enabled
Beyond Transcription: Building Voice AI That Understands Conversations — Hervé Bredin, pyannoteAI
Description

The open ASR leaderboard reports Nvidia Parakeet at 11.4% word error rate on AMI meeting data. Hervé Bredin runs the same model on the same dataset and gets 26%. Same model, same recordings, different microphone: the leaderboard uses headset audio, he uses the table mic. Most voice AI benchmarks are measuring single speaker speech and calling it solved. The talk covers speaker diarization (who speaks when), why combining it with transcription is harder than it looks, and what breaks at the word level when two speakers overlap. Bredin demos live on a two speaker phone call, walks through the word that falls between two speaker boundaries with no clean owner, and runs pyannoteAI's Precision 2 model down to 3% diarization error against the open source baseline at 5%. State of the art today: 2% on clean telephone calls, 41% in a noisy restaurant. Speaker info: - https://x.com/hbredin - https://www.linkedin.com/in/herve-bredin/ - https://github.com/hbredin

Summary

Generated by claude-sonnet-4-5

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Speaker diarization (who spoke when) is critical infrastructure for multi-speaker transcription, but state-of-the-art still fails badly in noisy/overlapping speech; pyannote's proprietary reconciliation approach can interleave overlapping speakers where open-source STT+diarization fail
  • Why it matters: Voice AI beyond single-speaker transcription is foundational for meeting notes, dubbing, podcast intelligence, and any multi-speaker workflow; understanding where SOTA breaks informs build-vs-buy decisions and realistic expectations
  • Best use: Watch to understand diarization error rates, why transcript+diarization reconciliation is hard, and where proprietary models (pyannote Precision 2) outperform open-source (Community 1); skip if you only care about single-speaker STT

Executive Summary

Hervé Bredin, co-founder/CSO of pyannoteAI, presents speaker diarization as the missing layer between transcription and understanding multi-speaker conversations. Transcription alone answers 'what was said'; diarization adds 'who spoke when', unlocking video dubbing, meeting action-item attribution, podcast guest tracking, and context-rich conversational AI. He demonstrates that knowing precise timestamps reveals interruptions, back-channels, and pauses that convey meaning beyond words.

He walks through diarization evaluation: Diarization Error Rate (DER) measures confusion (wrong speaker), false alarms (speech where there's none), and missed detection (especially overlapping speech). SOTA performance varies wildly by domain: 8% DER on clean telephone calls, but 41% DER in noisy restaurant settings. Open-source pyannote Community 1 achieves ~5% DER on his demo; proprietary Precision 2 drops to ~3% DER by better handling overlap and edge cases.

The hardest problem is reconciliation: merging STT word timestamps with diarization speaker timestamps. Most STT models (including Nvidia Parakeet) are trained on single-speaker data; they degrade from 11% WER on headset mics to 26% WER on distant multi-speaker mics (same AMI dataset). STT fails on overlapping speech, code-switching, and crosstalk. Reconciliation must decide which speaker to assign when word timestamps fall between diarization segments or during overlap. Bredin demos pyannote's proprietary approach interleaving overlapping speakers correctly where naive timestamp matching fails.

He emphasizes pyannote's open-source toolkit (10k GitHub stars, inflection after Whisper release) and commercial API (Precision 2 model). The reconciliation trick is partly 'exclusive diarization' (selecting the most likely speaker during overlap to simplify alignment) and proprietary heuristics. The SDK supports any STT model, including customer-fine-tuned ones, so users can plug in domain-specific transcription without retraining.

Key Takeaways

  • Claim: Speaker diarization (who spoke when) is as important as transcription (what was said) for multi-speaker use cases | Evidence: Hugging Face audio models by downloads: top 3 of 7 are diarization-related, rest are STT; video dubbing requires consistent voice mapping; meeting notes need action-item attribution; podcast intelligence tracks guests across episodes | Caveat: Importance depends on use case—single-speaker scenarios (voicemail, audiobooks) don't need diarization at all | Implication: Voice AI products for meetings, calls, podcasts, or interviews must budget for diarization quality, not just STT accuracy | Timestamp: 01:30
  • Claim: Diarization Error Rate (DER) is the sum of confusion, false alarms, and missed detections divided by total speech duration | Evidence: Demo: Community 1 model achieves 5% DER on a 30-second two-woman phone call (confusion where it tagged wrong speaker, false alarms where it detected speech in silence, missed detections during overlap); Precision 2 model achieves 3% DER on the same file by reducing all three error types | Caveat: DER doesn't penalize speaker label permutation (calling speaker A 'speaker 1' vs 'speaker 2' is equivalent), so it measures who-said-when accuracy, not identity naming | Implication: Operators should track DER segmented by overlap rate, speaker balance, and acoustic conditions; a global DER hides where the system breaks | Timestamp: 07:45
  • Claim: State-of-the-art diarization ranges from 8% DER (clean phone calls) to 41% DER (noisy restaurant with many speakers) | Evidence: Benchmarks cited: 8% DER on conversational telephone speech (CTS), 41% DER on restaurant multi-speaker scenarios; far from solved, especially in challenging acoustic environments | Caveat: No benchmark for streaming/real-time latency or adaptation to speaker drift (e.g., voice changes due to emotion, fatigue); quoted numbers are offline batch processing | Implication: Voice AI products targeting noisy, multi-party scenarios (retail, hospitality, open-plan offices) should plan for 30-40% error rates or costly custom training; clean use cases (1-on-1 calls, meeting rooms) are viable today | Timestamp: 10:20
  • Claim: STT models trained on single-speaker data fail badly on multi-speaker recordings, even when single-speaker benchmarks look good | Evidence: Nvidia Parakeet reports 11.4% WER on AMI dataset (meeting transcription) when evaluated on headset mics (single-speaker signal), but pyannote measures 26% WER on the same dataset using a distant mic (multi-speaker, overlapping speech, crosstalk) | Caveat: No mention of whether fine-tuning STT on multi-speaker data closes the gap or whether architectural changes (e.g., attention masking for overlap) are needed | Implication: Operators using Whisper, Parakeet, or other open STT for multi-speaker calls should measure WER on their actual use case, not Hugging Face leaderboards; expect 2-3× degradation in multi-speaker settings | Timestamp: 13:00
  • Claim: Reconciliation (aligning STT word timestamps with diarization speaker timestamps) is non-trivial because timestamps disagree, overlap is poorly transcribed, and one system may detect speech the other misses | Evidence: Demo: word 'oh' falls between two diarization segments; system must decide which speaker said it; during overlap, STT may transcribe only one speaker or interleave words incorrectly; pyannote's Precision 2 + Parakeet API successfully interleaves overlapping speakers ('in New Jersey' / 'and I'm Sheila') where naive timestamp matching would fail | Caveat: The reconciliation algorithm is proprietary; open-source Community 1 includes 'exclusive diarization' (selecting the most likely speaker during overlap to simplify alignment), but full logic is not disclosed | Implication: Building speaker-attributed transcription in-house requires solving a hard alignment problem on top of STT+diarization; pyannote's API abstracts this, but operators lose visibility into failure modes and can't debug edge cases | Timestamp: 16:30
  • Claim: Knowing precise speech turn boundaries reveals conversational cues (interruptions, back-channels, pauses) that LLMs miss from text-only transcripts | Evidence: Example: detecting that a speaker interrupted another vs. a natural turn-taking; a small 'yes' back-channel during another speaker's turn conveys agreement; pauses between speech turns signal hesitation or emphasis; stress on specific words ('the DOG ate the cake' vs 'the dog ATE the cake') changes meaning | Caveat: No discussion of how downstream LLMs or agents should consume these cues—does pyannote's API return structured annotations (interruption=true, back-channel=true) or just timestamps? | Implication: Voice AI agents for negotiation, therapy, sales coaching, or sentiment analysis should ingest diarization boundaries and prosody, not just flattened text; operators should design prompts/tools that surface these features to LLMs | Timestamp: 03:45
  • Claim: Pyannote's open-source toolkit (Community 1 model) is freely available, integrated with Hugging Face, and saw adoption spike after OpenAI Whisper's release | Evidence: GitHub star inflection points correlate with Whisper launch; Whisper provided free, high-quality STT but no speaker tags, so users combined it with pyannote; approaching 10k stars; demo uses pyannote.audio (diarization), pyannote.metrics (DER calculation), and ipyannotate (interactive visualization widget) | Caveat: Community 1 is less accurate than Precision 2 (5% vs 3% DER in demo); no mention of commercial licensing terms for Precision 2 API or whether Community 1 will remain free if pyannote scales | Implication: Operators can prototype multi-speaker transcription for free (Whisper + pyannote Community 1); production quality may require Precision 2 API; track pyannote's roadmap for open-source vs. commercial feature divergence | Timestamp: 02:00

Detailed Brief

Speaker diarization architecture and evaluation

  • Claims: Diarization answers 'who spoke when' by (1) voice activity detection (where speech exists), (2) segmentation (finding speaker change points and overlaps), (3) assigning speaker IDs to segments; Diarization does not output real names (John, Hervé) but permutable labels (speaker 1, speaker 2); number of speakers is unknown in advance, unlike typical ML classification; Evaluation metric is Diarization Error Rate (DER) = (confusion + false alarms + missed detections) / total speech duration; Challenges include overlapping speech, very short speech turns (back-channels), imbalanced speaker time, acoustic conditions (distant mic, noise)
  • Evidence: Demo: Community 1 model on 30-second two-woman phone call yields 5% DER (confusion: tagged wrong speaker; false alarms: detected speech where none exists; missed detections: missed one speaker during overlap); Precision 2 model yields 3% DER; Visualization widget (ipyannotate) shows reference annotation (ground truth) vs. system output, color-coded by error type; DER computed via pyannote.metrics library: sums errors, divides by total speech time
  • Caveats: DER doesn't penalize label permutation, so absolute speaker identity is not measured; Benchmarks shown are offline batch; no latency, streaming, or real-time adaptation numbers; Evaluation datasets (CTS, AMI, restaurant) may not generalize to specific domains (medical, legal, contact center)
  • Implications: Operators should eval diarization on domain-specific test sets, not just public benchmarks; Track DER by overlap percentage, speaker count, and acoustic condition to identify failure modes; Short back-channels and overlaps are high-value targets for improvement—missing a 'yes' can change action-item attribution

State-of-the-art performance and domain brittleness

  • Claims: Best systems achieve 8% DER on conversational telephone speech (clean, two-speaker); DER jumps to 41% in noisy restaurant scenarios with many speakers and background noise; Diarization is far from solved, especially in challenging acoustic environments
  • Evidence: Cited benchmarks: 8% DER on CTS, 41% DER on restaurant multi-speaker; Demo file (phone call) is a favorable case; real-world meeting rooms, call centers, open offices are harder
  • Caveats: No breakdown of which error type dominates in restaurant scenario (overlap? false alarms? confusion?); No comparison to human annotator agreement—what is the ceiling?
  • Implications: Voice AI products for clean environments (1-on-1 calls, meeting rooms) can rely on SOTA today; Noisy, multi-party scenarios (retail, hospitality, events) require domain-specific training or manual fallback; Operators should measure DER on production data, not rely on vendor-reported benchmarks

STT degradation on multi-speaker recordings

  • Claims: Most STT models are trained on single-speaker data; they fail on multi-speaker recordings with overlap, speaker change, crosstalk, code-switching; Nvidia Parakeet reports 11.4% WER on AMI (meeting transcription) dataset when evaluated on headset mics, but 26% WER on distant mic (same dataset); Open ASR leaderboard numbers are not representative of multi-speaker use cases
  • Evidence: AMI dataset has both headset mics (single-speaker signal per mic) and distant mic (multi-speaker signal); leaderboard uses headset, pyannote uses distant; Demo uses Parakeet, which transcribes the phone call sample well, but Bredin notes it struggles with overlap and crosstalk in general
  • Caveats: No mention of whether fine-tuning STT on multi-speaker data closes the gap; No comparison to STT models explicitly designed for multi-speaker (e.g., Whisper-large vs. Parakeet for overlap)
  • Implications: Operators should measure WER on their multi-speaker use case, not trust leaderboard numbers; Expect 2-3× WER degradation when moving from single-speaker benchmarks to multi-speaker production; Consider domain-specific STT fine-tuning or models trained on multi-speaker corpora

Reconciliation: the hard problem of aligning STT and diarization

  • Claims: Reconciliation is assigning a speaker to each transcribed word by merging STT word timestamps with diarization speaker timestamps; It's non-trivial because (1) STT doesn't transcribe overlap well, (2) timestamps disagree between STT and diarization, (3) one system may detect speech the other misses; Pyannote's proprietary reconciliation includes 'exclusive diarization' (selecting the most likely speaker during overlap to simplify alignment) and additional heuristics
  • Evidence: Demo: word 'oh' falls between two diarization segments; system must guess which speaker; naive timestamp matching fails; Overlapping speech example: 'in New Jersey' (speaker A) / 'and I'm Sheila' (speaker B) are correctly interleaved by Precision 2 + Parakeet API, with a hum at the end of A's utterance that overlaps B's start; Bredin states the reconciliation algorithm is proprietary but hints at 'exclusive diarization' in Community 1 model
  • Caveats: No open-source code or paper reference for reconciliation algorithm; No discussion of how the system handles cases where STT transcribes overlap as garbled text or omits one speaker entirely; No mention of confidence scores or uncertainty estimation for borderline word assignments
  • Implications: Building in-house speaker-attributed transcription requires solving a hard alignment problem on top of STT+diarization; not just a simple timestamp merge; Pyannote's API abstracts reconciliation, but operators lose visibility into failure modes and can't debug edge cases; Operators should test reconciliation on overlapping speech samples, not just clean alternating speakers

Conversational cues beyond text: interruptions, back-channels, pauses, prosody

  • Claims: Knowing precise speech turn boundaries reveals interruptions, back-channels (e.g., 'yes', 'uh-huh'), and pauses that convey meaning beyond words; Prosody (stress, intonation) can change sentence meaning: 'the DOG ate the cake' vs. 'the dog ATE the cake' vs. 'the dog ate THE cake'; Laughter, coughing, and other vocal cues signal emotional state (humor, discomfort, stress) that text-only transcripts miss
  • Evidence: Example: missing a small 'yes' back-channel during another speaker's turn loses the fact that the listener agrees; Example: detecting an interruption vs. natural turn-taking requires precise timestamps and diarization; Example: 'the dog ate the cake' with stress on different words conveys different emphasis (who, what, or which action)
  • Caveats: No discussion of how pyannote's API surfaces these cues—are interruptions/back-channels tagged as structured metadata or just implicit in timestamps?; No mention of prosody models or integration with TTS for dubbing use cases; No guidance on how downstream LLMs should consume these features (special tokens? structured JSON fields?)
  • Implications: Voice AI agents for coaching, therapy, negotiation, or sentiment analysis should ingest diarization boundaries and prosody, not just flattened text; Operators should design prompts or tools that surface interruptions, back-channels, and pauses to LLMs; Pyannote's roadmap may include prosody/emotion tagging, but it's not mentioned here

Pyannote open-source vs. commercial offering

  • Claims: Pyannote open-source toolkit (Community 1 model) is free, on Hugging Face, and saw adoption spike after OpenAI Whisper's release; GitHub star inflection points correlate with Whisper launch; approaching 10k stars; Bredin asks audience to star the repo as a 45th birthday gift; Commercial offering: Precision 2 model (cloud API, proprietary reconciliation, better accuracy than Community 1); Pyannote's SDK supports any STT model, including customer-fine-tuned ones, without retraining
  • Evidence: Demo uses pyannote.audio (diarization), pyannote.metrics (DER), ipyannotate (visualization widget, 'i' stands for interactive, not iOS); Community 1: 5% DER on demo file; Precision 2: 3% DER on demo file; GitHub repo: piano.ai/tutorials contains Jupyter notebook for the demo
  • Caveats: No mention of commercial licensing terms, pricing, or whether Community 1 will remain free if pyannote scales; No discussion of Precision 2's training data, model size, or latency vs. Community 1; No roadmap for future open-source releases or feature parity between Community and Precision models
  • Implications: Operators can prototype multi-speaker transcription for free (Whisper + pyannote Community 1); production quality may require Precision 2 API; Track pyannote's roadmap for open-source vs. commercial feature divergence; Community 1 may lag behind Precision 2 in accuracy and feature support; Pyannote's 'any STT' compatibility is a key differentiator vs. vertically integrated offerings (e.g., Deepgram, AssemblyAI) that bundle STT+diarization

Notable Concepts & Terms

  • Speaker diarization: The task of answering 'who spoke when' in a multi-speaker recording by segmenting audio into speech turns and assigning speaker IDs to each turn; does not output real names, just permutable labels (speaker 1, speaker 2, etc.)
  • Diarization Error Rate (DER): Sum of confusion (wrong speaker tagged), false alarms (speech detected where none exists), and missed detections (speech not detected, especially during overlap), divided by total speech duration; industry-standard metric for diarization quality
  • Speaker-attributed transcription: Transcription output that assigns a speaker label to each word or utterance, combining STT (what was said) with diarization (who said it, when); requires reconciliation between STT and diarization timestamps
  • Reconciliation: The process of aligning STT word timestamps with diarization speaker timestamps to produce speaker-attributed transcription; non-trivial because STT and diarization may disagree on boundaries, overlap, and speech presence
  • Exclusive diarization: Pyannote's technique (available in Community 1 model) that selects the most likely speaker during overlapping speech to simplify reconciliation with STT, which typically transcribes only one speaker during overlap
  • Back-channel: Short vocal cues (e.g., 'yes', 'uh-huh', 'mm-hmm') that a listener produces during another speaker's turn to signal agreement, attention, or encouragement; often very short and easily missed by diarization systems
  • AMI corpus: Academic dataset of meeting recordings with 4-5 participants, annotated with both headset microphones (single-speaker signal per mic) and distant microphones (multi-speaker signal); used to benchmark STT and diarization, but headset vs. distant mic results differ significantly
  • Pyannote open-source toolkit: Academic/community project by Hervé Bredin, now ~10k GitHub stars, providing diarization models, evaluation metrics (pyannote.metrics), and visualization tools (ipyannotate); Community 1 model is free on Hugging Face
  • Precision 2 model: Pyannote's proprietary, commercial-grade diarization model (cloud API); achieves lower DER than Community 1 (3% vs. 5% in demo) via better overlap handling and reconciliation; supports any STT model without retraining
  • Nvidia Parakeet: Open-source STT model from Nvidia; reports 11.4% WER on AMI headset mics but degrades to 26% WER on AMI distant mic (multi-speaker, overlapping speech); used in pyannote's demo for speaker-attributed transcription

Operator Notes / Why Ken Should Care

  • Multi-speaker transcription is foundational for meeting notes, call analytics, podcast workflows, and dubbing—but STT+diarization are not plug-and-play; reconciliation is a hard alignment problem that pyannote's API abstracts
  • Diarization quality varies wildly by domain (8% DER on clean calls, 41% DER in noisy restaurants); operators should measure DER on production data, segmented by overlap rate and acoustic conditions, not rely on vendor benchmarks
  • Open-source pyannote Community 1 + Whisper is a viable free prototype; production quality likely requires pyannote Precision 2 API or equivalent commercial service; track open-source vs. commercial feature divergence
  • STT models degrade 2-3× on multi-speaker recordings vs. single-speaker benchmarks (Parakeet: 11.4% → 26% WER); operators should measure WER on their use case, not Hugging Face leaderboards, and consider domain-specific fine-tuning
  • Voice AI agents for coaching, negotiation, therapy, or sentiment analysis should ingest diarization boundaries, back-channels, interruptions, and pauses—not just flattened text—to capture conversational dynamics; design prompts/tools that surface these cues to LLMs
  • Pyannote's 'any STT' compatibility (including customer-fine-tuned models) is a differentiator vs. vertically integrated services (Deepgram, AssemblyAI); useful if you need domain-specific STT (medical, legal, accented speech) but still want SOTA diarization
  • For agentic workflows: diarization output (speaker turns, overlap flags, back-channels) should be structured JSON consumed by downstream agents, not buried in raw timestamps; pyannote's API returns structured output, but reconciliation logic is proprietary—no debugging edge cases

Watch Map

  • 00:00: Intro: Hervé Bredin, pyannoteAI CSO; academic researcher → startup founder; 10k GitHub stars, birthday gift plea
  • 01:30: Why diarization matters: video dubbing, meeting notes, podcast intelligence; Hugging Face downloads show 3 of top 7 audio models are diarization-related
  • 03:45: Beyond 'who said what': precise timestamps reveal interruptions, back-channels, pauses, prosody (stress changes meaning: 'the DOG ate the cake' vs. 'the dog ATE the cake')
  • 05:00: What is speaker diarization: (1) voice activity detection, (2) segmentation (speaker change points, overlap), (3) assigning speaker IDs; output is permutable labels, not real names; number of speakers unknown in advance
  • 07:45: Demo 1: Diarization evaluation with DER (confusion + false alarms + missed detections / total speech duration); Community 1 model: 5% DER, Precision 2: 3% DER on 30-second phone call
  • 10:20: SOTA performance: 8% DER on clean phone calls, 41% DER in noisy restaurant; far from solved in challenging environments
  • 13:00: STT degradation: Nvidia Parakeet 11.4% WER on AMI headset mics (single-speaker) vs. 26% WER on AMI distant mic (multi-speaker); open ASR leaderboards not representative of multi-speaker use cases
  • 16:30: Demo 2: Reconciliation (aligning STT + diarization timestamps); word 'oh' falls between segments, overlap interleaving ('in New Jersey' / 'and I'm Sheila'); proprietary algorithm, partly 'exclusive diarization' (select most likely speaker during overlap)
  • 21:00: Q&A: reconciliation is proprietary heuristics + exclusive diarization, not part of STT model training; SDK supports any STT (including fine-tuned) without retraining; demo notebook at piano.ai/tutorials

Source/Metadata

  • Title: Beyond Transcription: Building Voice AI That Understands Conversations — Hervé Bredin, pyannoteAI
  • Transcript words: 5697
  • Duration seconds: 1519
  • Timestamp note: Timestamps manually estimated from transcript structure and demo flow; actual chapter markers not provided in transcript

Transcript

3711 words en Processed in 308.7s

Good morning everyone. Thanks for being here to the Voice and Vision session. So I'm Hervé Brodin, Chief Science Officer and Co-Founder at Pyannote AI. So I'm going to talk to you today about conversations, understanding conversations, and what you can do on top of transcription. So a quick word about myself. So I've been an academic researcher all my life until two years ago when I started this company. So I worked on this topic called speaker diarization, which I'll introduce a bit later. For those of you who don't know this weird word that is tricky for me to pronounce, over the years I built an open source toolkit called Pyannote, which focuses on speaker diarization. And that became quite popular over the years, in particular since OpenAI released Whisper speech to text open source models. Whisper was some kind of revolution in terms of STT, the fact that it was free, that it was very good, but it didn't provide the actual names and tags of the speaker. So people naturally turned to Pyannote to combine the two. And you can see that on the inflection points on the GitHub Star history. And just a few words, I happened to turn 45 in just one week from now. And we are almost at 10k stars on the GitHub. So please give me a birthday gift by just going to the website, to the GitHub website and add your star. That will make my birthday. So let's go to the core of the presentation. So you all know what transcription or a speech to text is. Basically, that's answering the question, what was said from the recording or streaming of a conversation. Basically, you get from the audio and you get a sequence of words as output. That's, for instance, what Whisper does. But without the actual names of the speaker, usually the conversations on the transcript are not really understandable. So the next step after transcription is basically to attribute a speaker tag to each of the words. So that's what I call here a speaker attributed transcription. And so that answers the question, who said what? But for some applications, that's enough. Like, for instance, maybe for meeting note takers to make a summary of a meeting, that might be enough to assign an action point to the right people to know exactly that Hervé said that and John said this other thing. But to understand the conversation a bit better, we might want to go slightly further. So basically, there are many cases where knowing who said what is as important as what was said, actually. So I have here a few examples, like for instance, video dubbing, automatic video dubbing. Knowing who said what is actually really important to put the actual correct voice to the right speaker. So if you want to automatically translate a video from one language to another, you want the voices to be consistent as well. So for that, you need to know who spoke when. Same for meeting note takers. And there are other applications like, for instance, what I call podcast intelligence, like being able to track a speaker across different episodes of a podcast or across different podcasts, finding the same guest across multiple podcasts. But yeah, so just knowing who said what is sometimes not enough to really understand the conversation, how it goes. So knowing who said what and when actually brings even more information. For instance, in this example, without knowing exactly when each word was pronounced, you can't detect that actually the black speaker interrupted the green one. You can't really detect that maybe this small back channel that the black speaker does during the speech of the green speaker. If you miss it, basically you don't really understand that maybe the black speaker is actually agreeing, noting at what the green speaker says. So knowing exactly when this small word has been pronounced is very important. And same, without precise timestamps, you can't really know whether I make a pause in between two speech terms, which might convey additional information about my state of mind, the point that I'm trying to make in the conversation. And we can go even a step further. Like, we might want to know who said what and when and how they said it. I have a few examples here. If some of you are laughing at some point, is it because I said something stupid or I was actually funny? Same for coughing. Maybe there's something in the conversation that makes me uncomfortable and I'm coughing for some reason to hide behind that. And yes, I also mentioned stress, discrepancy, prosody. The fact that I stress on particular words in a sentence might actually change completely the meaning of the sentence. I have this example in mind, in the sentence, the dog ate the cake. If you're talking to a child and the child tells you, the dog ate the cake, or the dog ate the cake, or the dog ate the cake, the meaning might be slightly different. So stressing particular words, being able to have a voice AI actually understand this stress and this kind of low-level details might bring some more information to any downstream LLMs or whatever tools that you use after this first enriched transcription step. And if we take even a step back, knowing who's talking to who, am I here in this example? And probably it can be seen as this gray speaker addressing all of you. But at some point, if at the end of this talk, you do have questions, then I will probably be the green one talking to one particular person answering the particular question. So knowing that, having a higher level view of the conversation, it also opens up a new bunch of applications. And all of this also happens in an acoustic environment that gives you information about the context, whether it's in this quiet room for now, or if I'm in the street. This brings more context to the game and you can actually take a better decision about what's going on, why a person is actually saying what they are saying, and so forth. So all of this is basically what we are trying to work on at Pyannote AI, but we started with a smaller problem called speaker diarization. And that's really what I want to talk about today. So speaker diarization is usually what I wanted to say about this slide is that who said what is usually just as important as what was said. And you can tell by, for instance, going through to the Hugging Face model repository and filter all the models there with the ones that have the audio tag. Here. So this is what I did. And then you saw them by number of downloads. And basically, among the top here, seven models that are at the top of this list, three of them, the first three are actually related to speaker identity and speaker diarization. And the other ones are STT. So that's one way of showing that speaker diarization is actually quite important in the community. So going into a bit more detail about what speaker diarization is. So as I said, it's answering the question, who speaks when? So starting from the recording of a conversation, basically, the first step that people usually do is a basic voice activity detection. So you basically the audio tag. Where is Marmas? Yeah, here. So this is what I did. And then you saw them by number of downloads. Among the top here, seven models that are at the top of this list, three of them, the first three are actually related to speaker identity and speaker diarization. And the other ones are STT. So that's one way of showing that speaker diarization is actually quite important in the community. So going into a bit more detail about what speaker diarization is. So as I said, it's answering the question, who speaks when? So starting from the recording of a conversation, the first step that people usually do is a basic voice activity detection. So you tell whether someone is speaking at one time or not, anyone is speaking at one time or not. And then you can go a step further, you can actually segment those speech regions into smaller speech turns. So this is what I call segmentation. And it includes finding speaker change points between into those longer in those long speech regions, you might find a speaker change point. And also find regions where people are actually interrupting each other, like in the first example here, there's definitely here someone interrupting someone else. And it's possibly here some kind of back channel where someone is actually nodding or saying, OK, this kind of stuff. And this is the kind of small speech turn that you don't want to miss because sometimes they actually convey the most important information of the conversation. If it's a small yes here, and if you miss it, then you have no idea what's the state of mind of this speaker. But that's not yet speaker diarization. Speaker diarization goes all the way to actually assigning a speaker identity to each speech turn. Speaker diarization. And so here in this example, the speaker diarization system automatically detected that there are two speakers, the green one and the black one. But we usually don't have any prior knowledge on the actual number of speakers in the conversation. You might have some kind of guidance on, I don't know, if you are a meeting notetaker, for instance, you have the list of attendees. So you might have an idea of the number of people that were invited, but that doesn't mean that someone, two people joined from the same channel or that an attendee that was not invited actually joined in finally. And so that makes the problem difficult in the sense that we don't know in advance the number of classes that we're supposed to detect, contrary to other, let's say, classical machine learning problems. And also we don't know the identity of the speaker that we're supposed to diarize in the sense that speaker diarization does not really output John, Hervé, and Jack. But it's more like speaker one, speaker two, speaker three, and they can be permutated and it's still correct. For example, in this example, if I just change green into black and black into green, the diarization is still exactly the same from the point of view of the evaluation metrics that I'm going to talk about and from the point of view of the actual speaker diarization task. So that makes this problem difficult. And that's why even though the community has been working on this topic for a long time now, it's still not solved. There are other reasons like the fact that we have to detect overlapping speech to handle overlapping speech that we have to take into account very short speech turns that have to take into account imbalance between the speech time of multiple speakers in the conversation, all of that makes the problem difficult. On top of the usual acoustic conditions, problems and everything that any speech processing task has to deal with. So I'm going to switch to a demo to show you a bit more how the diarization systems are evaluated. Maybe when you stumble into benchmark, you find this DER acronym, which stands for diarization error rate. So hopefully this will work. So I have here a Python notebook that I've prepared for this talk. It's actually already available on GitHub. I'll show you the QR code at the end of the presentation so you can play with it. But so in this example, I have a conversation between two women talking over the phone. Hello? Hello? Hello? Oh, hello. I didn't know you were there. Neither did I. Okay, I thought, you know, I heard a beep. This is Diane in New Jersey. And I'm Sheila in Texas, originally from Chicago. Oh, I'm originally from Chicago also. I'm in New Jersey now though. Okay, so you get the idea. A 30 seconds conversation between two women. So this is the expected output of a perfect speaker attributed transcription. This was manually labeled. And now what I'm going to do here is actually run a Pyannote open source community one model on the same file. So it's actually running right now on my Mac. So that's why here I'm using this MPS PyTorch backend. So how it works is that you first download the model from Hugging Face with this line of code and simply apply it on the audio file. So now it's finished and there's this prediction variable that contains... Sorry, I should have done that on the other tab because this was the one that was already processed. I wanted to do it live. So let's do it live now. So it's running. And so the next step is to actually visualize the errors that this system made. So at the top here, you have the reference annotation. So the expected output of the diarization. And here is the output of a Pyannote community one speaker diarization pipeline that is open source and free to use. So it makes three kinds of mistakes, as I was saying, like a confusion here. It got the speaker wrong. It can also make false alarms. So false alarm is when it detects speech when there's actually no speech in the ground truth. And misdetection is the other way around. And misdetection can actually happen, for instance, here during overlapping speech when we actually detect only one of the two speakers here. And then once we have these errors of false alarm confusion and misdetection, we can actually compute the diarization error rate. So this is with this library called Pyannote metrics. And basically it gives us in this example, a 5% diarization error rate, which is basically the sum of the confusion, false alarm and misdetection, divided by the total duration of speech in this file. And so we have this other model that I'm currently running. Hopefully it works. So it's running right now on our cloud API. Basically, it's a better model than community one that we call Precision two that is running here. Hopefully it will work. But as for every demo, it might fail. If no, it worked. And this is the kind of output that you can get. So you see that it got the speaker right here and makes slightly less mistakes in terms of false alarm and misdetection. And overall, in this example, I think we got a 3% diarization error rate here. So let me switch back to the presentation. So this was about benchmarking speaker diarization models. And I'm often asked how well the state of the art speaker diarization works today. And it's a difficult question to answer because really depending on the use case, it might vary completely. For instance, in this example, we have here conversation telephone speech like the one we just listened to like 2% over the phone. Hopefully it will work. But as for every demo, it might fail. If not, it worked. And this is the kind of output that you can get. So you see that it got the speaker right here and makes slightly fewer mistakes in terms of false alarm and misdetection. And overall, in this example, I think we got a 3% diarization error rate here. So let me switch back to the presentation. So this was about benchmarking speaker diarization model. And I'm often asked how well the state of the art speaker diarization works today. And it's a difficult question to answer because it really depends on the use case. It might vary completely. For instance, in this example, we have a here conversation telephone speech like the one we just listened to, like 2% over the phone. And we can go down to 8% diarization error rate. The best system does that. But if you are now in a restaurant with many friends with lots of background noise, even the best system reaches 41% diarization error rate. So it's far from being a solved problem. But we are working on that. And so once you have speaker diarization and you have transcription, it shouldn't be that hard to actually find a speaker attributed transcription, right? It's just a matter of assigning a word to a speaker, a speaker to a word. But actually, that's not that easy. And there are many reasons why that's not that easy. The first reason is that most speech to text models are actually trained on single speaker data. As soon as you apply them on multi speaker data with overlap, with speaker change, with all kinds of mess, they fail miserably. So for instance, when you look at the open ASR leaderboard from Hugging Face, if we look, for instance, at Nvidia Parakeet that we are using, that I will be using in a demo right after that, they report 11.4% word error rate. When we apply the very same model on the same AMI on our side, we get 26%. So the question is, why is there a difference? Actually, it's in the way that those benchmarks are set up here. For example, in this AMI data set, that's a data set with meetings between four to five people, I think, in meeting rooms. And they basically have microphones like the one I'm having here, so a headset microphone, as well as one microphone in the middle of the table. And those numbers here on the open ASR leaderboard are based on the headset microphone. Those numbers here are based on the microphone that is in the middle. And so on one side, you have a single speaker speech. On the other side, you have multi speaker and even distant microphone speech. So that's why it degrades a lot. And so that was my point of this slide. So when doing speaker attributed transcription, the reason why it might go wrong is either because STT doesn't work great, and it's usually the case that they don't generalize very well to multi speaker recordings for, as I was saying, distant microphone, speaker change, crosstalk, interruptions, you name it. Maybe code switching as well when you change language in the middle of a sentence. But also it might be because of the actual reconciliation between diarization and STT timestamps. Though it may sound obvious how to do that, as I was saying, I will hopefully show in this live demo once again that it's not such an easy problem. Because STT does not transcribe overlapping speech well, because the timestamps disagree between STT and diarization, and because sometimes diarization will detect speech that the transcription will not transcribe, and the other way around. So, second part of the demo about speaker attributed transcription. So, first I'm going to apply this Parakeet model from NVIDIA that does, by the way, a great job at transcribing this sample. And so Parakeet gives us this kind of output. So we have this sequence of words, and for each word we actually have the corresponding timestamps. Okay? So I can play it quickly. Hello. Oh, hello. I didn't know you were there. Neither did I. Okay. I thought, you know, I heard a beep. This is Diane in New Jersey. Well, you get the idea. And then the question is, now we have our best diarization. This is Pyannote output here. This is Parakeet output here. And the question is, now we need to assign a speaker to each word. So let's do it step by step, for instance. So in this example, even though the timestamps of the words are a bit shifted on the left here, it's obvious that it's probably the yellow speaker who spoke the gray word. But if you go just the third word. So now we have this. So we have this word, the word O here, that is in between those two speech turns according to the diarization. Which one do you assign it to? So we, I can listen to it. Oh, hello. Oh, hello. Oh, hello. So here I just submitted to our cloud API a job to transcribe and use both Pyannote diarization model and Parakeet transcription model and take care of the, what I call this reconciliation between the two. And you end up with this kind of output. What's nice is that, now you, it did the job for you. But in particular, I like to go to this particular example. Even where there was overlapping speech before it actually managed to interleave the words from the two speakers. Let me play just this part and focus on this area maybe. In New Jersey. And I'm Sheila and tech. And there's actually an end that, or this is the hum here that is actually a, the, the, the, the yellow lady actually has this hum at the end of their speech. And the blue one actually interrupts her. But they are actually overlapping and we managed to get them right. Yeah. So this is the end of my talk. [SPEAKER_02] I wanted to give you a few minutes. [SPEAKER_02] I have two minutes left for questions, but I wanted to say that this demo that I did is already available on the piano.ai slash tutorials GitHub repo. [SPEAKER_02] So in there you'll be able to play with Pyannote open source toolkit for diarization, with piano.metrics for evaluation, with IPyNOT for this nice visualization widget that we've implemented. [SPEAKER_02] I stands for, not iOS, iMac or whatever, but really for interactive. [SPEAKER_02] And also the SDK to play with our premium models. [SPEAKER_02] And I'll stop here and I'm happy to take a few questions. We have one more minute left. [SPEAKER_02] Yes, please. [SPEAKER_02] So what's the trick that you use to resolve the confidence? Is it part of the model training or is it some sort of heuristics that you have around to identify who to attribute the word to? Yeah. [SPEAKER_02] With IPNOT for this nice visualization widget that we've implemented. [SPEAKER_02] I stands for interactive, not iOS, iMac or whatever. [SPEAKER_02] And also the SDK to play with our premium models. [SPEAKER_02] I'll stop here and I'm happy to take a few questions. We have one more minute left. [SPEAKER_02] Yes, please. [SPEAKER_02] So what's the trick that you use to resolve the confidence? Is it part of the model training or is it some heuristics that you have around to identify who to attribute the word to? Yeah. So the question is what's the trick to actually solve this problem of reconciliation? That's a proprietary trick, but what I can say is that there's already part of this trick that is available in the community one model, which we call exclusive diarization. Basically, what we do is find a way when there is overlap to actually select the most likely of the two speakers that will be transcribed by the STT model. So that simplifies the reconciliation between the two. That's one major part of the approach. The model training is like a couple of heuristics you have to. Yeah. So the question is about, I'm repeating because I've been told I need to repeat questions. It's not part of the model training in the sense that we really plan to support any kind of STT without having to change the STT model itself. So really it's supposed to work with any STT, even fine-tuned ones that you might have internally because you fine-tuned them for your particular use case and nobody has it. You can combine it like that. All right. Yeah. 18 seconds. So I guess we need to stop. I'm sorry. Maybe we can talk offline. Otherwise they'll beat me, I guess. Thank you very much. I'm sorry. Thank you. here, you have the reference annotation. So the expected output of the realization. And here is the output of a community one speaker authorization pipeline that is open source and free to use. So it makes three kinds of mistakes, as I was saying, like a confusion here. It got the speaker wrong. It can also make force alarms. So force alarm is when it detects speech when there's actually no speech in the ground truth. And misdetection is the other way around. And misdetection can actually happen, for instance, here during overlapping speech when we actually detect only one of the two speakers here. And then once we have these basically errors of force alarm confusion and misdetection, we can actually compute the diarization error rate. So this is with this library called Pianot metrics. And basically it gives us in this example, a 5% diarization error rate, which is basically the sum of the confusion, force alarm and misdetection, and misdetection, divided by the total duration of speech in this file. And so we have this other model that I'm currently running. Hopefully it works. So it's running right now on our cloud API. Basically, it's a better model than community one that we call precision two that is running here. Hopefully it will work. But as for every demo, it might fail. If no, it worked. And this is the kind of output that you can get. So you see that it got the speaker right here and makes slightly less mistakes in terms of force alarm and misdetection. And overall, in this example, I think we got a 3% diarization error rate here. So let me switch back to the presentation. So this was about benchmarking speaker diarization model. And I'm often asked how well the state of the art speaker diarization works today. And it's a difficult question to answer because really depending on the use case, it might vary completely. For instance, in this example, we have a here conversation telephone speech like the one we just listened to like 2% over the phone. And we can go down to 8% diarization error rate. The best system does that. But if you are now in a restaurant with many friends with lots of background noise, even the best system reaches like 41% diarization error rate. So it's far from being a solved problem. But we are working on that. And so once you have speaker diarization and you have transcription, shouldn't be that hard to actually find a speaker attributed transcription, right? It's just a matter of assigning a word to a speaker, a speaker to a word. But actually, that's not that easy. And there are many reasons why that's not that easy. The first reason is that most speech to text models are actually trained on single speaker data. As soon as you apply them on multi speaker data with overlap, with speaker change, with all kinds of mess, they fail miserably. So for instance, when you look at the open ASR leaderboard from hugging face, if we look, for instance, at Nvidia Paraket that we are using, that I will be using in a demo right after that, they report 11.4% word error rate. When we apply the very same model on the same AMI on our side, we get 26%. So the question is, why is there a difference? Actually, it's in the way that those benchmarks are set up here. For example, in this AMI data set, that's a data set with meetings between four to five people, I think, in meeting rooms. And they are basically microphones like the one I'm having here, so a headset microphone, as well as one microphone in the middle of the table. And those numbers here on the open ASR leaderboard are based on the headset microphone. Those numbers here are based on the microphone that is in the middle. And so on one side, you have a single speaker speech. On the other side, you have multi speaker and even distant microphone speech. So that's why it degrades a lot. And so that was my point of this slide. So when doing speaker attributed transcription, the reason why it might go wrong is either because STT doesn't work great, and it's usually the case that they don't generalize very well to multi speaker recordings for, as I was saying, distant microphone, speaker change, crosstalk, interruptions, you name it. Maybe code switching as well when you change language in the middle of a sentence. But also it might be because of the actual reconciliation between diarization and STT timestamps. Though it may sound obvious how to do that, as I was saying, I will hopefully show in this live demo once again that it's not such an easy problem. Because STT does not transcribe overlapping speech well, because the timestamps disagree between STT and diarization, and because sometimes diarization will detect speech that the transcription will not transcribe, and the other way around. So, second part of the demo about speaker attributed transcription. So, first I'm going to apply this Paraket model from NVIDIA that does, by the way, a great job at transcribing this sample. And so Paraket gives us this kind of output. So we have this sequence of words, and for each word we actually have the corresponding timestamps. Okay? So I can play it quickly. Hello. Oh, hello. I didn't know you were there. Neither did I. Okay. I thought, you know, I heard a beep. This is Diane in New Jersey. Well, you get the idea. And then the question is, now we have our best diarization. This is Precision 2 output here. This is Paraket output here. And the question is, now we need to assign a speaker to each word. So let's do it step by step, for instance. So in this example, even though the timestamps of the words are a bit shifted on the left here, it's obvious that it's probably the yellow speaker who spoke the gray word. But if you go just the third word. So now we have this. So we have this word, the word O here, that is in between those two speech turns according to the diarization. Which one do you assign it to? So we, I can listen to it. Oh, hello. Oh, hello. Oh, hello. So here I just submitted to our cloud API a job to transcribe and use both Precision 2 diarization model and Paraket transcription model. And take care of the, what I call this reconciliation between the two. And you end up with this kind of output. What's nice is that, so now you, it did the job for you. But in particular, I like to go to this particular example. Even where there was overlapping speech before it actually managed to interleave the words from the two speakers. Let me play just this part and focus on this area maybe. In New Jersey. And I'm Sheila and tech. And there's actually a end that, or this is the um here that is actually a, the, the, the, the, the yellow lady actually has this hum at the end of their speech. And the, the blue one actually interrupts her. But they, they, they are actually overlapping and we managed to get them right. Yeah. So this is the end of my talk. I wanted to give you a few minutes. I have two minutes left for questions, but I want, just wanted to say that this demo that I did is already available on the, in the piano.ai slash tutorials, uh, GitHub repo. So in there you, you, you'll be able to play with piano.to.do open source toolkit for dialization, with piano.metrics for evaluation. Uh, with IPNOT for this nice, uh, visualization widget that we've, uh, implemented. I stands for, uh, not iOS, iMac or whatever, but really for interactive. And, uh, also the SDK to, to play with our premium models. And, uh, I'll stop here and, uh, I'm happy to take, uh, a few questions. We have one more minute left. Yes, please. So what's the trick, uh, that you use to resolve the confidence? So part of the model training or is it some sort of like heuristics that you have around to identify who to attribute the word to? Yeah. So the question is, uh, what's the trick to, uh, actually solve this problem of a reconciliation? Uh, so that, that's a, a proprietary trick, but what I can, what I can say is that we, uh, there's already part of this trick that is available in the community one model, which we call exclusive diarization. Basically, uh, what we do is that we find a way when there is overlap to actually select, uh, the, the most, uh, likely of the two speaker that will be transcribed by the STT model. So that simplifies actually the reconciliation between the two. So that, that's one major part of the, uh, of the approach. The model training is like a couple of heuristics you have to. Yeah. So, so the question is about, uh, I repeating because I've been told I need to repeat questions. Uh, so it's not part of the model training. In the sense that we really plan to support any kind of STT without having to change the STT model itself. So really it's supposed to work with any STT, uh, even fine tune ones that you might have internally for, because you, you fine tune them for your particular use case and nobody has it. Uh, you can combine it like that. All right. Yeah. 18 seconds. So I guess we need to stop. I'm sorry. Maybe we can talk offline. Otherwise, uh, they'll, they'll beat me, I guess. Thank you very much. I'm sorry. Thank you.