Open Reader

5 Voice Agent Failure Modes You'll Hit in Week One — Venky B, Plivo

completed 26:46 Sep 15, 2026 Watch on YouTube

Current Status

completed

Video ID

vblnYHzBgS4

RAG / Chat

Enabled
5 Voice Agent Failure Modes You'll Hit in Week One — Venky B, Plivo
Description

Nearly every intelligence gain in language models over the past year has come from letting them think longer. Voice agents have to turn thinking off, because the budget between a user finishing a sentence and the agent starting to speak is measured in hundreds of milliseconds. Venky B is founder and CEO of Plivo, which carries over a billion voice calls a month and has been building telephony infrastructure since 2011, and this talk is a tour of what breaks when a voice agent leaves the demo and meets production. On latency his numbers are blunt. Teams aim for under 550 milliseconds and most land between 750 and 1,200, and past that users simply hang up. His team's answer is smaller open source models hosted themselves, targeting under 300 milliseconds, chosen partly on how many tokens a language needs per word. The failure he says wrecks half of all deployments is data collection, and his fix is to stop treating it as transcription at all. Decide the shape before you ask. A phone number is a typed field with a length and a validator, so a stray letter in the middle is either corrected with confidence or sent back to the caller, and evaluation happens per field as a unit test rather than end to end. That reframing took his accuracy from roughly 30 percent to the mid nineties with no fine tuning. He is equally specific about transcription being brittle by default, especially with proper nouns and code switched languages, and about never feeding model output straight into speech synthesis. His own benchmark for a vendor is whether it can pronounce his surname and his company's name. Speaker info: - https://x.com/bevenky - https://www.linkedin.com/in/bevenky/ Timestamps: 0:00 - Where voice agents break between demo and production 2:20 - A billion calls a month 4:28 - The pipeline everyone builds first 5:34 - Failure one: latency and time to first audio 6:42 - Balancing cost, intelligence, and latency 7:48 - Why thinking models do not fit 9:56 - Choosing and sizing o

Summary

Generated by claude-sonnet-4-5-20250929

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Production voice AI agents fail predictably in five areas—latency, transcription brittleness, data collection, TTS normalization, and turn detection/barging—and Plivo has battle-tested patterns from processing 1B+ calls/month that let you architect around these failure modes before they break your agent.
  • Why it matters: This is rare production playbook material: Venky B shares Plivo's accumulated failure modes and fixes at scale (1B+ monthly calls, 14 years telephony), not conceptual architecture. The patterns—open-source MOE models hitting <300ms TTFT, dynamic keyword boosting in STT, treating data collection as typed fields, LLM post-processing for transcripts, and normalization layers before TTS—map directly to OpenClaw's agent orchestration and control plane needs.
  • Best use: Use as an implementation checklist when building or refactoring voice agent pipelines: latency optimization tactics (open-source models, dedicated capacity tradeoffs), transcription robustness patterns (dynamic keyword boosting, LLM post-processing), data collection as typed fields with unit-test evals, and TTS normalization layers. Also reference for GTM when positioning agent reliability or debugging customer production issues.

Executive Summary

Venky B, founder and CAO of Plivo (14-year telephony API platform, 1B+ calls/month, 50M ARR/profitable), walks through the five failure modes that break voice AI agents when they move from POC to production. Plivo's vantage point—owning the full stack from carrier layer to AI agent studio—lets them see patterns across customer deployments that single-layer providers miss. The talk is structured as a practitioner's playbook: what breaks, why it breaks at scale, and the specific architectural patterns Plivo uses to fix it.

The five failure modes are: (1) Latency—frontier models (OpenAI, Claude) hit 450–500ms P50 TTFT and spike to 1.2s+ at P95, killing conversational flow; Cerebras/Grok offer speed but require 12-month committed capacity at prohibitive cost. Plivo's solution: open-source MOE models (Qwen 3.5, Gemma 4 for multilingual) on self-hosted GPUs, consistently hitting <300ms TTFT. (2) Transcription brittleness—state-of-the-art STT engines claim 4–6% WER on eval sets but hit double-digit WER in production (accents, noise, domain vocab, code-switched languages). Patterns that break: proper nouns, phone numbers with substitutions (E→3), addresses, multilingual script mixing (English in Hindi Devanagari). Plivo's fix: dynamic keyword boosting (context-specific, not global pollution), LLM post-processing of transcripts (domain context corrects substitutions), and neural transliteration layers to normalize multilingual input before the LLM. (3) Data collection—50–60% of agents fail here by treating collection as unstructured transcript parsing instead of typed field validation. Plivo's pattern: think Pydantic/Zod for voice—decide field shape before asking (phone number as phone_number type, date as datetime with current-date anchoring), validate allowed values, confirm letter-by-letter for hard names, and eval at unit-test/field level instead of end-to-end. This pattern moves accuracy from 30% to 95% without fine-tuning. (4) TTS normalization—LLM output sent directly to TTS breaks on emojis, markdown, proper nouns, acronyms, speeds. Plivo inserts a normalization layer: strip non-speech tokens, apply custom pronunciation dictionaries, dynamically slow speed (0.7–0.8x) for entities like emails/phone numbers, normalize messy formats (currency, dates) in-house so TTS switching doesn't break the agent. (5) Turn detection and barging/back-channeling—covered at high level only due to time, but Plivo's position is you don't need speech-to-speech models to handle these; speech-to-text pipelines with proper state management suffice.

Throughout, Venky emphasizes architectural principles: balance cost/intelligence/latency (thinking must be turned off for voice agents, so 2023–2024 LLM intelligence advances don't apply here); build normalization/validation layers so your agent is decoupled from underlying model/STT/TTS provider failures; treat evals as unit tests on fields, not end-to-end black-box runs; and assume brittleness at every layer. Plivo runs two model flavors in production: fine-tuned 8B–12B models for domain-specific use cases (healthcare, etc.) and out-of-box MOE models for generic conversational + tool-calling tasks. The talk is dense with implementation detail (Gemma 4's token fertility is 2.5–3x better than Qwen 3.5 for multilingual, MOE models hit 90% accuracy out-of-box but are hard to fine-tune, phone number transcription example with 'E' substitution, Balasubramanian pronunciation as a TTS test) and positions these patterns as the difference between 30% and 95%+ production reliability.

Key Takeaways

  • Claim: Frontier LLMs (OpenAI, Claude, Gemini) deliver P50 TTFT of 450–500ms but spike to 1.2s+ at P95, breaking conversational flow; Cerebras/Grok offer speed but require 12-month committed capacity at prohibitive cost. | Evidence: Plivo benchmarked frontier models in production across 1B+ monthly calls. P95 latency spikes make agents feel unresponsive (>1.2s is 'annoying' tier). Cerebras/Grok teams confirm 12-month advance booking required for dedicated capacity. | Implication: For production voice agents targeting <550ms (natural feel) or <750ms (acceptable), frontier models are too spiky and expensive fast models are too locked-in. You need a third path.
  • Claim: Open-source MOE models (Qwen 3.5, Gemma 4) on self-hosted GPUs consistently hit <300ms TTFT, balance cost/intelligence/latency, and work out-of-box for 90% of use cases without fine-tuning if they do instruction-following and high tool-calling success well. | Evidence: Plivo runs two production flavors: MOE for generic conversational + tool calling (3–4B MOE), fine-tuned 8–12B for domain-specific (healthcare). Gemma 4 has 2.5–3x better token fertility than Qwen 3.5 for multilingual (fewer tokens per word = faster time-to-words). MOE models hit 90% accuracy out-of-box. | Implication: For OpenClaw voice agents, self-hosted open-source MOE models are the production-grade path: faster, cheaper, no 12-month lock-in, and reliable enough for most tool-calling/conversational tasks. Reserve fine-tuning for domain-specific verticals where out-of-box fails. | Caveat: MOE models are hard to fine-tune (risk breaking the model); if fine-tuning is required, start with 8B–12B dense models. English-only can use Qwen 3.5; multilingual needs Gemma 4.
  • Claim: State-of-the-art STT engines claim 4–6% WER on eval sets but hit double-digit WER in production due to accents, noise, domain vocab, and code-switched languages (e.g., English in Hindi Devanagari script). | Evidence: Plivo sees consistent patterns of breakage: proper nouns, phone numbers with substitutions (E→3, one→1), missing address parts, multilingual script mixing that breaks STT→LLM→TTS chain. Example: 'E' in phone number transcript corrected to '3' by LLM post-processing. | Implication: Never send raw STT output to your LLM. Build a normalization layer: dynamic keyword boosting (context-specific, not global), LLM post-processing with domain context, and neural transliteration for multilingual. Assume transcription is brittle and architect around it.
  • Claim: Dynamic keyword boosting (adding keywords only during specific call states when needed, not globally) prevents STT hallucination and improves accuracy without polluting the transcription context. | Evidence: Plivo recommends context-aware keyword injection instead of static lists. Many STT engines offer keyword boosting, but static lists cause hallucination. Dynamic state-specific boosting solves this. | Implication: For OpenClaw orchestration, implement stateful keyword injection: pass current conversation state to STT layer and update boosted terms dynamically. This is a control-plane coordination task, not a prompt-engineering fix.
  • Claim: 50–60% of voice AI agents fail at data collection because they treat it as unstructured transcript parsing instead of typed field validation (think Pydantic/Zod for voice). | Evidence: Plivo's pattern: decide field shape before asking (phone_number type, datetime type), validate allowed values, confirm letter-by-letter for hard names. Accuracy jumps from 30% to 95% without model fine-tuning. Example: 'E' in phone number instantly flagged as error or smart-guessed to '3' with confirmation. | Implication: Architect data collection as typed fields with validation rules, not open-ended LLM extraction. Eval at field/unit-test level, not end-to-end. This shifts reliability from prompt engineering to structured data modeling and is directly applicable to OpenClaw's form/workflow logic.
  • Claim: Relative values (dates, times) break dramatically in unstructured collection: 'next week Wednesday 8' is ambiguous (8am vs 8pm, which Wednesday). Treat as constrained datetime field anchored to current date. | Evidence: Plivo example: if field is datetime, agent takes current date, calculates 'next Wednesday,' and validates 8am vs 8pm with confirmation. Tool calling does heavy lifting for field validation. | Implication: For scheduling/booking agents, never let LLM parse relative dates freeform. Use tool calling with datetime fields + current-date context. This is a structured workflow pattern, not a prompt pattern.
  • Claim: LLM output sent directly to TTS breaks on emojis, markdown, proper nouns, acronyms, and entity speed. Insert a normalization layer between LLM and TTS to control pronunciation, speed, and format. | Evidence: Plivo's normalization layer: strip emojis/markdown (most orchestration frameworks do this with flags), apply custom pronunciation dictionaries (proper nouns, brands, acronyms), dynamically slow speed to 0.7–0.8x for entities (email, phone number, letter-by-letter spelling), normalize messy formats (currency, dates) in-house. Test: TTS must pronounce 'Balasubramanian' and 'Plivo' correctly or it fails baseline. | Implication: Don't rely on TTS engines to normalize—build in-house so you can switch TTS providers without breaking pronunciation. For OpenClaw, this is another normalization layer in the pipeline: LLM → normalization → TTS, with provider-agnostic logic.
  • Claim: All LLM intelligence advancements in 2023–2024 (thinking, reinforcement learning, extended reasoning) are turned off by default for voice agents because latency kills conversational flow. | Evidence: Venky notes the irony: LLMs got smarter via thinking layers, but voice agents must disable thinking to stay fast. Only instruction-following and tool-calling matter for voice. | Implication: When evaluating LLMs for voice agents, ignore reasoning benchmarks and focus on TTFT, tool-calling success, and instruction-following. This changes OpenClaw's model selection criteria for voice vs. text agents.

Detailed Brief

Latency optimization tactics and model selection

  • Claims: Some customers use dual-model architecture: small conversational model (3B) for talking, larger model for tool calling to improve success ratio.; For English-only, Qwen 3.5 or Gemma 4 MOE work; for multilingual, Gemma 4's token fertility (tokens per word) is 2.5–3x better, yielding faster time-to-words.; If fine-tuning is required for domain-specific use cases (healthcare, etc.), start with 8–12B dense models, not MOE, because fine-tuning MOE is fragile and risks breaking the model.
  • Evidence: Plivo runs production agents with both architectures: generic MOE out-of-box and fine-tuned 8–12B for verticals.; Qwen 3.5 and Gemma 4 are currently cutting-edge open-source; 'maybe six months from now a 4B beats 8B hands down, but today minimum 8–12B for fine-tuning.'; MOE gets you '90% of the way there' without fine-tuning if instruction-following and tool-calling are strong.
  • Caveats: MOE fine-tuning is difficult and can break the model; dense models are safer for custom training.; Model performance evolves rapidly; 4B models may catch up to 8B in six months, so architecture should be model-agnostic.
  • Implications: For OpenClaw, default to MOE for generic voice agents and reserve fine-tuning (8–12B dense) for vertical-specific deployments where domain accuracy is critical and customers will pay for it.; Design model router logic that can swap models without rewriting orchestration layer, anticipating model performance shifts every 6 months.

Transcription robustness patterns and multilingual handling

  • Claims: Code-switched languages (e.g., English in Hindi Devanagari, Hindi in Latin script) break STT, LLM, and TTS sequentially because each layer misinterprets script.; Neural transliteration engines (many open-source options) normalize multilingual STT output before sending to LLM, preventing cascading failures.; LLM post-processing corrects domain-specific transcription errors because LLM has domain context that STT does not.
  • Evidence: Example: 'Namaste, how are you?' in Devanagari script vs. Hindi words in Latin script—both break the pipeline if not normalized.; Phone number example: STT outputs 'E' in '555-E678'; LLM with domain context corrects to '555-3678.'; Plivo recommends transliteration layer or LLM normalization before LLM reasoning step.
  • Implications: For OpenClaw agents serving international markets or multilingual users, add neural transliteration + LLM post-processing as standard pipeline stages before LLM reasoning.; This is not a prompt fix—it's an architectural stage in the agent pipeline, likely a pre-processing microservice or edge function.

Data collection as typed fields with validation and confirmation

  • Claims: Treat data collection as UX problem for voice: decide field shape (phone_number, datetime, name) before asking, constrain allowed values, validate, and confirm when uncertain.; Unit-test evals at field level, not end-to-end: if unit tests pass, agent is reliable; if one field breaks, you don't re-run 100 end-to-end cases.; Hard-to-spell names (example: Balasubramanian) require letter-by-letter confirmation and field-level validation rules; no transcription engine will get it right first try.
  • Evidence: Plivo saw 30% → 95% accuracy jump by switching from unstructured LLM extraction to typed field validation.; Name example: agent asks for spelling letter-by-letter only after field validation detects potential error, not by default.; Datetime example: 'next week Wednesday 8' parsed as datetime field with current-date anchor and am/pm confirmation.
  • Implications: For OpenClaw workflow engine, model data collection steps as strongly-typed field nodes with validation logic, not as LLM prompt steps. This is closer to form-builder logic than agent prompting.; Build eval harness that runs field-level unit tests independently, not just end-to-end agent runs. This speeds up debugging and CI/CD for agent changes.

TTS normalization layer and provider decoupling

  • Claims: Custom pronunciation dictionaries (proper nouns, brands, acronyms) should be managed in your normalization layer, not left to TTS engine, so switching TTS providers doesn't break pronunciation.; Dynamic speed control (0.7–0.8x) for entities (emails, phone numbers, letter-by-letter spelling) improves clarity without slowing entire conversation.; Plivo's TTS test: if engine can't pronounce 'Balasubramanian' and 'Plivo' correctly, it fails baseline and will break on customer-specific terms.
  • Evidence: Most TTS engines (Cartesia, PlayHT, ElevenLabs, etc.) provide custom dictionaries and speed controls, but Plivo recommends abstracting this into in-house layer for portability.; Plivo normalizes emojis, markdown, currency, dates in-house before sending to TTS.
  • Implications: For OpenClaw, treat TTS as swappable provider with normalization logic in control plane, not in TTS config. Store customer-specific pronunciation dictionaries and speed rules as workflow metadata.; If building white-label agent platform, expose pronunciation dictionary and speed control as customer-configurable settings in agent studio.

Notable Concepts & Terms

  • Token fertility: Number of tokens required to generate one word in a language; lower is better for multilingual voice agents because time-to-words is faster. Gemma 4 has 2.5–3x better token fertility than Qwen 3.5.
  • Dynamic keyword boosting: Context-aware injection of keywords into STT engine only during specific call states when those terms are expected, avoiding hallucination from static keyword pollution. Plivo's preferred STT accuracy pattern.
  • Mixture of Experts (MOE) models: LLM architecture where subsets of parameters activate for each token; enables smaller effective model size with frontier-like performance. Plivo uses 3–4B MOE for generic voice agents. Hard to fine-tune but works out-of-box for 90% of cases.
  • TTFT (Time To First Token): Latency metric for LLM response start; critical for voice agents because it determines how quickly agent begins speaking after user stops. <300ms is Plivo's production target.
  • Word Error Rate (WER): STT accuracy metric; state-of-art claims 4–6% on eval sets but Plivo sees double-digit in production due to noise, accents, domain vocab. Drives need for post-processing layers.
  • Code-switched languages: Mixing languages or scripts in speech (e.g., English in Hindi Devanagari, Hindi in Latin script). Breaks STT→LLM→TTS pipeline if not normalized via transliteration layer.
  • Neural transliteration engine: Open-source tool that converts text from one script to another (e.g., Devanagari to Latin) to normalize multilingual STT output before LLM reasoning.
  • Normalization layer (TTS): Intermediate processing stage between LLM output and TTS input that strips emojis/markdown, applies custom pronunciation, adjusts speed for entities, and normalizes formats. Decouples agent from TTS provider.
  • Field-level unit tests (evals): Eval strategy where each data collection field (phone_number, datetime, name) is tested independently with validation rules, rather than running full end-to-end agent conversations. Speeds up debugging and improves reliability.
  • Plivo: 14-year-old telephony API platform ($50M revenue, profitable, no external VC) processing 1B+ voice calls/month. Now focused on full-stack AI agent offering (programmable pipeline + no-code studio + carrier layer).

Operator Notes / Why Ken Should Care

  • Benchmark open-source MOE models (Qwen 3.5, Gemma 4) on self-hosted GPUs for OpenClaw voice agents; target <300ms TTFT and measure tool-calling success rate as primary selection criteria, not reasoning benchmarks.
  • Implement dynamic keyword boosting in STT layer: pass conversation state to STT provider and inject context-specific keywords per call stage, not static global lists.
  • Add LLM post-processing stage between STT and LLM reasoning to correct domain-specific transcription errors (phone numbers, proper nouns, addresses) using domain context.
  • For multilingual agents, integrate neural transliteration layer to normalize code-switched language output from STT before LLM reasoning, preventing cascading failures in LLM and TTS.
  • Architect data collection as typed fields with validation logic (Pydantic/Zod pattern for voice): phone_number, datetime, name fields with constrained allowed values, confirmation logic, and letter-by-letter spelling for hard names.
  • Build field-level eval harness for data collection: test each field's validation and accuracy independently with unit tests, not end-to-end agent runs, to speed debugging and CI/CD.
  • Insert TTS normalization layer between LLM and TTS: strip emojis/markdown, apply custom pronunciation dictionaries, dynamically adjust speed (0.7–0.8x for entities), normalize currency/dates in-house to decouple from TTS provider.
  • Store customer-specific pronunciation dictionaries and speed rules as workflow metadata in OpenClaw agent studio if building white-label platform; expose as configurable settings.
  • Test TTS baseline with hard-to-pronounce names (e.g., Balasubramanian) and company/brand terms; if TTS fails, assume it will fail on customer-specific terms and build pronunciation layer.
  • Reserve fine-tuning for domain-specific verticals (healthcare, legal, etc.) where out-of-box MOE fails; use 8–12B dense models for fine-tuning, not MOE, to avoid model breakage.
  • Design model router in OpenClaw to swap LLMs without rewriting orchestration layer; assume model performance shifts every 6 months (today's 8B may be beaten by 4B in six months).
  • For GTM, position Plivo's five failure modes as reliability checklist when selling OpenClaw to customers deploying production voice agents; use as diagnostic framework when debugging customer issues.
  • Monitor P50, P90, P95 TTFT in production; if P95 exceeds 750ms consistently, investigate LLM provider spikes or fallback to self-hosted open-source models.
  • Do not rely on 2023–2024 LLM intelligence advances (thinking, reasoning, RL) for voice agents; they are disabled for latency. Focus model selection on instruction-following and tool-calling only.

Source/Metadata

  • Title: 5 Voice Agent Failure Modes You'll Hit in Week One — Venky B, Plivo
  • Transcript words: 7495
  • Duration seconds: 1606
  • Timestamp note: Timestamps not provided in transcript; speaker referenced time constraints during talk but no MM:SS markers available.

Transcript

4180 words en Processed in 164.6s

Let's just do a couple of quick questions and then we'll jump right in. How many of us in the room here have built Voice AI agents? Okay, that's a pretty good audience here. And how many of you guys have built AI agents that have been deployed in production? Not bad. Okay, cool. So we'll talk about what typically happens, right? Everyone's talking about Voice AI agents. The one pill solution to pretty much everything in the world today is Voice AI agents. So everyone's building one and trying to deploy that. They sound great when you're building that in your dev landscape. And then the moment you take this from a proof of concept to production, things start failing. So we'll walk through these five different angles of how or what we have seen at Plevo with Voice AI agents. But just before that, a quick intro from my side. I am Venki, the founder and CAO. She used the title agent engineering manager. I'm calling myself chief agent officer from a title standpoint. Okay, so what is, you know, why are we even qualified for this discussion? And what are we seeing that a lot of companies don't get to see? I'll talk a bit about our journey in terms of how we have come along so far and then jump right in. You know, we've been around for about 14 years. Our journey has been a developer API platform. And then now in an AI agent business, we started with voice and SMS APIs back in 2011. And then now we are primarily focused on our AI agent offering, the full stack. On our platform, we see over a billion voice calls each month across the globe. And that is where we have seen a lot of these patterns emerge in terms of how, when we work with our customers, what happens with their voice AI agents in production. We're a 90 member team. And we have 50 million funding in the bank. Fun fact, this is not from external VC investors. This is all from being a profitable company, having put that cash in the bank over these years. Some customers we power across the globe. We've just left some logos in there. But primarily from an offering standpoint, I would cohort this into three different buckets. One is a programmable AI agent offering. We call it a speech pipeline, not a true speech-to-speech product yet, but that's a programmable offering. We also have an AI agent studio. It's a no-code visual builder. And then, like I said, we started with voice APIs. So we obviously have built this out over the last 14 years, the SIP trunking and the audio streaming layer. So we don't rely on other folks for the telephony or the carrier layer. That's the bread and butter business we've built over all these years. And that's on top of which our AI agent platform sits. Okay, with that, let's get into this, right? Which I'm sure, since you guys have all built AI agents, you've all seen this or built this in one manner or another. And we'll spend more time on this in terms of how the entire pipeline looks, right? What we see with customers is, and I'm sure you guys can all relate to this, is anyone thinking about AI agents, what they do is they pick up a bunch of these orchestration frameworks and they do a pretty good job, LiveKit or PipeCat, you know, build their AI agent on top of that. They think they can just orchestrate these different four layers, speech to text, LLM, and TTS with turn detection in between, and we're off to the races. My AI agent works in a POC and it's good to work in production. Typically that's what happens. They sort of measure their latencies and you can see some indicative latencies on the slide at each layer. And they're like, yeah, this seems good for me for what I need, so let's position production. And then the production woes start to kick in and you see all sorts of failure modes, which we are going to spend most of the time on in this talk. I've kept some time at the end for Q&A if you want to have questions, but we'll jump right in to different failure modes we see. Let's start with the first one, which everyone talks about. This is the most spoken about failure mode, which is latency. I think we have a few AI agent talks today, or voice AI agent talks today. I'm pretty sure everyone's going to touch upon this specific failure mode, which is why I'm bringing this right up in terms of how this entire experience is for users, right? Typically, most folks measure this by time to first audio, so the time when your users stop speaking to your agent starts speaking, right? And I think you've probably seen this if you guys have built voice AI agents on what good or natural feels like, what annoying feels like or noticeable feels like, and then what annoying feels like, which is different tiered steps. We notice most people want to be under 550 because that's what's advertised by platforms or solutions or layers, but I think most end up between 750 to 1.2. That's where most of the folks end up at. The really bad performing ones end up more than 1.2, and then you start to see users hang up. Now, I'll share with you what we've seen practically in production with customers using this at different layers, and then solutions to some of these. The way we want to think about this layer is a balance between these three, which is cost, intelligence, and latency, right? And why do I bring these three up? Because they're interrelated. I think one of the things, I was just chatting with a couple of folks outside. One of the things, last year, we've seen a lot of innovations, a lot of intelligence spike on the LLM side of things, right? And most of the intelligence has come in terms of thinking or reinforcement learning and so on and so forth. The irony with voice AI agents is almost always your LLM or the agent that's talking has to have thinking turned off, right? So all the advancements we've had in the LLM layer in the last year, none of that even applies here now, right? Obviously, you have better models that can do better instruction following or tool calling, but pretty much all of your intelligence that's been built in on the thinking layer is all off by default if you want it to be fast enough. So that's one of the ironies that we come up with. So then how do you sort of balance intelligent cost and latency? Let's look at some of these options that are out there in the market, right? So I'm specifically picking LLM because if you looked at the previous chart, LLM is your highest latency bucket that adds to this, right? And if you look at frontier models, which I think most folks start by default, your OpenAI, your Claude, your Gemini's, you know, P50, TTFT is roughly around 450 to 500 on a good day, and it can get spiky, right? P90, P95 can go easily upwards of 1.2, 1.3 seconds, and that's not good for the overall agent experience. So that's your frontier model. Now, there's another option, which is your Cerebris or the Grok that is famous and popular for spitting out a lot of tokens or tokens very fast, right? These work, but for you to get dedicated latency or time-to-first token on these, you need dedicated capacity, and that is really expensive. That's where I spoke about the cost as being one of the things to balance, right? It's really expensive, and then, you talk to anyone from the Grok team or the Cerebris team, they'll tell you, you need to book 12 months in advance for dedicated capacity. They're booked out for the next 12 months. So that's a pretty expensive option, and then you really need to be sure that the model you're deploying on some of these infrastructure layers will be here 12 months from now, and it's a big investment and a big unknown. So what's a realistic option for production-grade agents that are good quality and end up balancing three of these? This is what has worked for us, which is the open-source models. There are obviously a lot of them in terms of the variety and variations you can pick. I'm specifically talking about the two we work with, Qwen 3.5 and Gemma 4. These are cutting-edge open-source models out in the market right now. You talk to anyone from the Grok team or the Cerebris team, they'll tell you, you need to book 12 months in advance for dedicated capacity. They're booked out for the next 12 months. So that's a pretty expensive option, and then you really need to be sure that the model you're deploying on some of these infralayers will be here 12 months from now, and it's a big investment and a big unknown. So what's a realistic option for production-grade agents that are good quality and end up balancing three of these? This is what has worked for us, which is the open-source models. There are obviously a lot of them in terms of the variety and variations you can pick. I'm specifically talking about the two we work with, QEN 3.5 and Gemma 4. These are cutting-edge open-source models out in the market right now, and we've done a lot of benchmarking around this in how they work. It can be scary to think, okay, I have the models. Now I have to host them, run them on my own GPUs and so on and so forth, but if you are consistently targeting under 300 MS, we've seen this to be a great option to balance between latency, cost, and intelligence. Now, some more deep dive here. If you're doing only English, QEN 3.5 or Gemma both work fine, but if you're doing multilingual, right, international audiences, different languages, Gemma 4 is a much better model for that. We have seen token fertility evals. Essentially what that means is if I were to de-jargonize that is, how many tokens does it take to generate one word in that language? Okay, so Gemma is much, much better, at least 2.5 to 3x better than QEN 3.5 from that perspective, so your time to words is much faster on Gemma 4, everything else equal, right, on a multilingual basis. Now, what sizes do you pick at the LLM layer? The mixture of experts usually works fine. The three or four billion mixture of experts usually works fine. The problem with mixture of experts is if anyone goes down, wants to go down the direction of fine-tuning, that can be a challenge, because fine-tuning mixture of experts models are not easy. You can end up breaking the model a lot of times, so that's one challenge we see with mixture experts, but usually, out of the box, it gets you 90% closer to where you want to be, even without any fine-tuning or custom work done on the model. So that's the advantage of mixture experts. Now, if you want to fine-tune and you want to go deeper and say, look, I'm working for a specific domain, healthcare, what have you, right, and I want to make sure I'm able to fine-tune my model, you want to start at least with the $8 billion, $12 billion, at least from where we are today. Maybe six months from now, a $4 billion model beats the $8 billion model, hands down, but for today, what we've seen is you minimum need an $8 billion or $12 billion model, because you're looking for two things in these models. One, obviously, fast tokens, but good instruction following, okay, and the second thing is very high success ratio in tool calling, because if you can do these two things well, then you are on to 70%, 80% there for not even having to fine-tune any model. Models will work out of the box, right? So that's been our recipe. We've actually run two flavors, one a fine-tune model for specific industries, and then for most generic use cases, a MOE model just works out of the box. There are a few more tips and tricks we'll talk about in the upcoming slides where we see failure models, but that's where we stand from a latency LLM standpoint. All right, I'm running short on time, so I'm going to fast-track this. Now, there are a couple of other flavors in this. People build agents with a mixture of models. What they do is, for the talking part of it, they have a conversational model, which is a much lower, smaller model, and then maybe even a three billion model, and then for tool calling, they have a much larger model, so they have an improved tool calling success ratio there. The second one is, assume your transcriptions are going to be brittle. That's something you want to live by when you're building AI agents, even if you have the best transcription engine out there, and I'll show you why, right? The state-of-the-art transcription engines out in the market, you know, get you to four to six percent word error rate, right? And this is on known eval sets. On real world, noisy calls with people having different accents, domain vocabulary, and so on and so forth, those usually end up in the double digits from a word error rate perspective, right? Now, obviously, you can fine-tune, pick up an open source model and fine-tune, but we see typically what breaks here often, and there are patterns here in terms of what breaks. So, proper nouns, jargons, phone numbers, random missing digits with phone numbers, wrong substitutions. I'll walk through some examples of how you solve for these. Addresses, when you're trying to collect a long address, the transcription engine could just end up missing some parts of it. Code-switch languages. I'll just take an example of a language I speak, because that was easy for me to put on the slide, where if you were to take English, but written in a different script, that's what's used for Hindi, right? This is English written in that script, right? Whereas the actual English version of this is, hello, how are you? So if I'm addressing an audience in a different country where I have code-switched languages and I start getting my English in a different script, everything starts breaking from the transcription engine to the LLM layer and then beyond, because your LLM starts then producing output in that script a lot of times, and then your TTS messes up. Okay, so this is very important to be careful about, and if you want to build your agent independent of the transcription engine, you need to build a layer that normalizes all of this, right? We'll talk about solutions in a minute, and there is the other case, which is Hindi, and just Latin or Roman, right? Which is this is Hindi, but it reads English, which again messes up everything downstream. Those are just examples. This applies to Arabic, Mandarin, Japanese, what have you, pretty much any language. So what actually moves the needle at the transcription layer? For proper nouns, we recommend you using not just keyword boosting. I think a lot of transcription engines provide you keyword boosting where you can put in specific words into their engine, but doing dynamic keyword boosting. What that means is don't keep the keyword for the entire state of the call. Just add that dynamically when you think you need that as an answer so that you get the highest accuracy, meaning at different states of the call, the transcription engine will have different keywords boosted during different phases, right? And that's what we've seen works best because if you just pollute your context of the transcription engine with tons of keywords, it'll start hallucinating again, right? So that's what we see typically working best. Yeah, post-process. Post-process your transcripts with an LLM, right? Because your LLM has domain context, your transcription engine does not. So a lot of words that it would say, I'll give you some examples, may not make sense. This is transcription like a phone number from a transcription engine, right? Like what do you think that E is, right? If you give it to an LLM, it knows that's a three. Similarly, what that one is, it's a digit one. So your transcription engine a lot of times could mess that up, but when you post-process it with an LLM layer, it'll instantly correct that from a collection standpoint. And the last one, transliteration is your STT output, that's multilingual, also gets normalized using either an LLM, you first transliterated, or use some kind of a neural transliteration engine. There are a lot of them open source, you can just pick one of them, right? That will do all of that work for you, send cleaned transcripts consistently, independent of the transcription engine to your LLM. All right, the third one we typically see is collecting data. This is where I think That E is, right? If you give it to an LLM, it knows that's a three. Similarly, what that one is, it's a digit one. So your transcription engine a lot of times could mess that up, but when you post-process it with an LLM layer, it'll instantly correct that from a collection standpoint. And the last one, I said, transliteration is your STT output, that's multilingual, also gets normalized using either an LLM, you first transliterated, or use some kind of a neural transliteration engine. There are a lot of them open source, you can just pick one of them, right? That will do all of that work for you, send cleaned transcripts consistently, independent of the transcription engine to your LLM. All right, the third one we typically see is collecting data. This is where I think 50 to 60% of AI agents mess up pretty badly. And we to think of it as a UX problem, but just for voice. So think data models and not a transcript coming into an LLM and trying to figure out what the transcript said. So let's take some inspiration from, I'm assuming most of us are developers here, take inspiration from Python's data classes, PyDantic, Zod from TypeScript, or form fields in the UI, right? If you start thinking of it from that problem statement, we have seen accuracy grow from 30% to 95% from a data collection standpoint when you start thinking in that manner. So, decide your shape before you ask, right? Instead of keeping it open-ended, can you keep it constrained? So, can a phone number be a phone number type field? The moment you do that, right, you know, how many digits it needs to have. You can do validation on top of that, right? And then what sort of allowed values can even be there? So, in the previous example we saw, if an E comes in in the middle of a phone number and you know it's a phone number, you instantly know, either you smart guess that to three and confirm that with the user, or you know that's an error and then you validate that and ask the user to repeat again, right? So that's, I think, one of the common patterns we've seen here from a collection pattern. Name, I think, is the interesting one. I've just picked a hard to pronounce name. Like, there is no way a human is going to get this right and no way a transcription engine will get this right, how many ever times you do this, right? So the moment you start thinking of this as fields and then have rules and then confirmation mechanisms on spelling this, letter by letter, only then you kind of get it right. Otherwise, it's going to mess up pretty badly in terms of how you collect this on a voice call. And that's just an example of what I'm talking about in terms of the data collection piece of it. Another place where it goes badly, dramatically, is relative values, date being one of the examples. If somebody says, next week, Wednesday, eight, it could mean 8 a.m., 8 p.m., and then figuring out what that date actually is, again, now becomes a very constrained problem. If you knew this was a date time field and I'm collecting a date time field and then you take the current date and then figure out what this value would be based on that, right? So that's how you want to make sure like you do this with a combination of the LLM with the tool calling and the tool calling is doing a lot of this heavy lifting for you from a field standpoint. Yeah, and then you make, you run like this from a unit test perspective. So, all of your evals need to start treating these fields as unit tests and as long as your unit tests sort of validate and pass, you know your agent is going to be sort of reliable and repeatable. You don't run hundreds of end-to-end agent test cases just to find out one field collection is broken. You do your evals at a field level and a unit test level. And then, yeah, I think this mindset makes everything more structured instead of hoping I'll put a ton of prompt, keep changing the prompt by a few characters every time and somehow my prompt engineering is going to make LLM much more instruction tuned and sort of magically start following some of these things. So, in fact, right, we have seen us get to 95, 97% accuracy without having to fine-tune a model, right? And the trick is basically just breaking down your context of what the agent is doing at that point with specific states of what the agent is going through. All right, I'm just going to quickly skip through this from a time standpoint. I just see I got three more minutes. Hopefully, that's a bug, but we'll leave it at that. So, this is the fourth area where we see issues coming in. Most folks take the LLM output and then we send it to a TTS. Obviously, I think there are a lot of good TTSs in the market that take care of a lot of heavy lifting, but a lot of times it messes up. What we recommend and what we've seen is you usually want to have a normalization layer between your LLM and what is fed to a TTS. You don't send your LLM output directly to a TTS, right? And we'll just walk through some examples. The basics, which is strip emojis, markdown before any synthesis into the TTS. Most orchestration pipelines do this, a LiveCAD or a PipeCAD will do that for you if you just set a few flags. But just make sure if you're not using them or building from scratch that you've set this explicitly because you don't want an emoji showing up on something read out or markdown showing up there. I think some more common ones, custom dictionaries. Most TTS engines provide this to you, how to pronounce custom words, whether it's proper nouns, brands, acronyms, and so on and so forth. So set those in when you go from your LLM to your TTS output because if you don't, you're going to mess that up. And I'll show you an example of how we test that. The other one is most engines also give you speed. So if you know you're pronouncing an entity, slow down, have your agent slow down at 0.8x or 0.7x so that it's able to enunciate on that specific entity and doesn't mess up how it's pronouncing an email or a phone number or a name letter by letter. And yeah, just normalize all the messy stuff, right? Like emails, currency, dates. Don't leave it to the TTS to do it. Most of them do it, but don't leave it to the TTS to do it. Build your normalization layer at your end so that tomorrow you think you need to switch TTS or for whatever reason the first one's down and you want to use another TTS, you're able to not rely natively on the TTS's engine but you're building this in-house for this to be managed. And then yeah, I don't have my batch here but I don't have my last name on that. So my first test is if it cannot pronounce my last name or my company's name, it's already dropping the ball. So my last name is Balasobramanian and if you cannot pronounce that using a voice AI agent, that's a check for me. I know the agent will mess up a lot of words that need to be spelled out day by day. The second one is our company named Plivo. So a lot of engines pronounce it Plivo or Plivo and so on and so forth but I think specifically being able to control this in your pipeline is super critical and then if you're building a customer facing product then you know, sort of give this option to your customers. I'm just going to skim through the last two slides. I'm running badly over time. End of turn detection, I think this is a separate topic but I'm just going to quickly pull up all the points so you guys can skim through that and if you need a chat after this, we can talk about this. I'm quite over time and then the last one is barging and back channeling. I think there's a lot of talk around speech-to-speech models that do some of this but we've been able to see how we could do all of this in speech-to-speech pipelines. You really don't need a speech-to-speech model to do all of this. Again, I just put this up on the slide and sort of close at that. All right. I don't think we have Give this option to your customers. I'm just going to skim through the last two slides. I'm running badly over time. End of turn detection, I think this is a separate topic but I'm just going to quickly pull up all the points so you guys can skim through that and if you need a chat after this, we can talk about this. I'm just going to leave that for five seconds and then we can chat about this offline. I'm quite over time and then the last one is barging and back channeling. I think there's a lot of talk around speech-to-speech models that do some of this but we've been able to see how we could do all of this in speech-to-speech pipelines. You really don't need a speech-to-speech model to do all of this. Again, I just put this up on the slide and close at that. All right. I don't think we have time for questions. We can take them offline if you have any time but hopefully this was helpful and gave you some insights on what we are seeing in productions with billions of calls at scale. All right. Thanks. We'll see you next time. We'll see you next time. Thank you. is, like, how many tokens does it take to generate one word in that language? Okay, so Gemma is much, much better, at least 2.5 to 3x better than QEN 3.5 from that perspective, so your time to words is much faster on Gemma 4, everything else equal, right, on a multilingual basis. Now, what sizes do you pick at the LLM layer? The mixture of expert usually works fine. The three or four billion mixture of expert usually works fine. The problem with mixture of expert is, like, if anyone goes down, wants to go down the direction of fine-tuning, that can be a challenge, because fine-tuning, mixture of experts models are not easy. You can end up breaking the model a lot of times, so that's one challenge we see with mixture experts, but usually, out of the box, it gets you 90% closer to where you want to be, like, even without any fine-tuning or custom work done on the model. So that's the advantage of mixture experts. Now, if you want to fine-tune and you want to go deeper and say, like, look, I'm working for a specific domain, healthcare, what have you, right, and I want to make sure I'm able to fine-tune my model, you want to start at least with the $8 billion, $12 billion, at least from where we are today. Maybe six months from now, a $4 billion model beats the $8 billion model, hands down, but for today, what we've seen is you minimum need an $8 billion or $12 billion model, because you're looking for two things in these models. One, obviously, fast tokens, but good instruction following, okay, and the second thing is, like, very high success ratio in tool calling, because if you can do these two things well, then you are on to, like, 70%, 80% there for not even having to fine-tune it, fine-tune any model. Like, models will work out of the box, right? So that's been our recipe. We've actually, we run two flavors, one a fine-tune model for specific industries, and then for, you know, most generic use cases, a MOE model just works out of the box. There are a few more tips and tricks we'll talk about in the upcoming slides where we see failure models, but that's where we stand from a latency LLM standpoint. All right, I'm running tired on time, so I'm going to fast-track this. Now, there are a couple of other flavors in this. People build agents with a mixture of models. What they do is, you know, for the talking part of it, they have a conversational model, which is a much lower, smaller model, and then, you know, maybe even a three billion model, and then for tool calling, they have a much larger model, so they have a improved tool calling success ratio there. Sorry. The second one is, assume your transcriptions are going to be brittle. Like, that's something you want to sort of live by when you're building AI agents, even if you have the best transcription engine out there, and I'll show you why, right? Like, the state-of-the-art transcription engines out in the market, you know, sort of get you to four to six percent word error rate, right? And this is on known eval sets. On real world, noisy calls with, you know, sort of accents, like people having different sort of accents, domain vocabulary, and so on and so forth, like those usually end up in the double digits from a word error rate perspective, right? Now, obviously, you can fine-tune, you know, pick up an open source model and fine-tune, but we see typically, like, what breaks here often, and there are patterns here in terms of what breaks. So, proper nouns, jargons, phone numbers, like random missing digits with phone numbers, wrong substitutions. I'll walk through some examples of, like, how you solve for these. Addresses, when you're trying to collect a long address, you know, the transcription engine could just end up missing some parts of it. Code-switch languages. I'll just take an example of a language I speak, because that was easy for me to put on the slide, where, you know, like, if you were to sort of take English, but written in a different script, that's what's used for Hindi, right? Like, this is English written in that script, right? Whereas, like, the actual English version of this is, hello, how are you? So if I'm addressing an audience in a different country where I have code-switched languages and I start getting my English in a different sort of script, everything starts breaking from the transcription engine to the LLM layer and then beyond, because your LLM starts then producing output in that sort of script a lot of times, and then your TTS messes up. Okay, so this is very important to be careful about, and if you want to build your agent independent of the transcription engine, you need to build a layer that normalizes all of this, right? We'll talk about solutions in a minute, and there is the other case, which is Hindi, and just Latin or, you know, Roman, right? Which is, like, this is Hindi, but it reads English, which again messes up everything, you know, downstream. Those are just examples. This applies to, you know, Arabic, Mandarin, Japanese, what have you, pretty much any language. So what actually moves the needle with a, at the transcription layer? For proper nouns, we recommend you using not just keyword boosting. I think a lot of transcription engine engines provide you keyword boosting where you can put in specific words into their engine, but doing dynamic keyword boosting. What that means is don't keep the keyword for the entire state of the call. Just add that dynamically when you think you need that as an answer so that you get the highest accuracy, meaning at different states of the call, the transcription engine will have different keywords boosted during different phases, right? And that's what we've seen works best because if you just pollute your context of the transcription engine with tons of keywords, it'll start hallucinating again, right? So that's what we see typically working best. Yeah, post-process. Post-process your transcripts with an LLM, right? Because your LLM has domain context, your transcription engine does not. So a lot of words that it would say, I'll give you some examples, may not make sense. This is transcription, like a phone number from a transcription engine, right? Like what do you think that E is, right? If you give it to an LLM, it knows that's a three. Similarly, like what that one is, it's a digit one. So your transcription engine a lot of times could mess that up, but when you post-process it with an LLM layer, it'll instantly correct that from a collection standpoint. I mean, and the last one, like I said, transliteration is your STT output, that's sort of, you know, multilingual, also gets normalized using either an LLM, you first transliterated, or, you know, use some kind of a neural transliteration engine. There are a lot of them open source, you can just pick one of them, right? That will do all of that work for you, send cleaned transcripts consistently, independent of the transcription engine to your LLM. All right, the third one we typically see is collecting data. This is where I think 50 to 60% of AI agents mess up pretty badly. And like, we like to think of it as a UX problem, but just for voice. So think data models and not a transcript coming into an LLM and trying to figure out what the transcript said. So let's take some inspiration from, I'm assuming most of us are developers here, you know, take inspiration from Python's data classes, PyDantic, Zod from TypeScript, or form fields in the UI, right? Like, if you start thinking of it from that problem statement, we have seen accuracy grow from 30% to like 95% from a data collection standpoint when you start thinking in that manner. So, like, decide your shape before you ask, right? Like, instead of keeping it open-ended, can you keep it constrained? So, can a phone number be a phone number type field? The moment you do that, right, you know, like, how many digits it needs to have. You can do validation on top of that, right? And then what sort of allowed values can even be there? So, in the previous example we saw, if an E comes in in the middle of a phone number and you know it's a phone number, you instantly know, like, either you smart guess that to three and confirm that with the user, or, you know that's an error and then you validate that and ask the user to repeat again, right? So, so that's, I think, one of the common patterns we've seen here from, from a collection pattern. Name, I think, is the, is the interesting one. I've just picked a, you know, a, a hard to pronounce name. Like, there is no way a human is going to get this right and, and no way a transcription engine will get this right. How many ever times you do this, right? So the moment you start thinking of this as fields and then have rules and then confirmation mechanisms on, on spelling this, you know, sort of letter by letter, only then you kind of get it right. Otherwise, it's going to mess up pretty badly in terms of how you collect this on a voice call. And, and that's just an example of, you know, what I'm talking about in terms of the, the data collection piece of it. Another place where it goes badly, dramatically, is relative values, date being one of the examples. If somebody says, next week, Wednesday, eight, it could mean 8 a.m., 8 p.m., and then figuring out what that date actually is, again, now becomes a very constrained problem. If you knew this was a date time field and I'm collecting a date time field and then you take the current date and then figure out what this value would be based as that, right? So, so that's how you want to make sure like you do this with a combination of the LLM with the tool calling and the tool calling is doing a lot of this heavy lifting for you from a, from a field standpoint. Yeah, and then you make, you, you run like this from a unit test perspective. So, all of your evals need to start treating these fields as unit tests and as long as your unit tests sort of validate and pass, you know your agent is going to be sort of reliable and repeatable. You don't, you know, run hundreds of end-to-end agent test cases just to find out, you know, one field collection is broken. You do your evals at a field level and a unit test level. And then, yeah, like I said, I think this mindset makes everything more structured instead of hoping I'll put a ton of prompt, keep changing, you know, the prompt by a few characters every time and somehow my prompt engineering is going to make LLM much more instruction tuned and sort of magically start following some of these things. So, in fact, like I said, right, like we have seen us get to 95, 97% accuracy without having to fine-tune a model, right? And then, and the trick is basically like just breaking down your context of what the agent is doing at that point with specific states of what the agent is going through. All right, I'm just going to quickly skip through this from a time standpoint. I just see I got three more minutes. Hopefully, that's a bug, but we'll leave it at that. Okay. So, so this is the fourth area where we see issues coming in. Most folks take the LLM output and then we send it to a TTS. Obviously, I think there are a lot of good TTSs in the market that take care of a lot of heavy lifting, but a lot of times it messes up. What we recommend and what we've seen is you usually want to have a normalization layer between your LLM and what is fed to a TTS. You don't send your LLM output directly to a TTS, right? And we'll just walk through some examples. The basics, which is strip emojis, markdown before any synthesis into the TTS. Most orchestration pipelines do this, like a LiveCAD or a PipeCAD will do that for you if you just set a few flags. But just make sure if you're not using them or buildings from scratch that you've set this explicitly because you don't want an emoji showing up on something read out or markdown showing up there. Okay. I think some more common ones, custom dictionaries. Most TTS engines provide this to you, like how to pronounce custom words, whether it's proper nouns, brands, acronyms, and so on and so forth. So set those in when you go from your LLM to your TTS output because if you don't, you're going to mess that up. And I'll show you an example of how we test that. The other one is most engines also give you speed. So if you know you're pronouncing an entity, slow down, have your agent slow down. So at 0.8x or 0.7x so that it's able to like enunciate on that specific entity and doesn't mess up how it's pronouncing an email or a phone number or a name letter by letter. And yeah, just normalize all the messy stuff, right? Like emails, currency, dates. Don't leave it to the TTS to do it. Most of them do it, but don't leave it to the TTS to do it. Like build your normalization layer at your end so that tomorrow you think you need to switch TTS or you know, for whatever reason the first one's down and you want to use another TTS, you're able to sort of not rely natively on the TTS's engine but you're building this in-house for this to be managed. And then yeah, I think I don't have my batch here but I don't have my last name on that. So my first test is if it cannot pronounce my last name or my company's name, it's already dropping the ball. So my last name is Balasobramanian and if you cannot pronounce that using a voice AI agent, like that's a check for me. I know like you know, the agent will mess up a lot of words that you know, need to be spelled out day by day. The second one is our company named Plivo. So a lot of engines pronounce it Plivo or Plivo and so on and so forth but I think specifically being able to control this in your pipeline is super critical and then if you're building a customer facing product then you know, sort of give this option to your customers. I'm just going to skim through the last two slides. I'm running badly over time. End of turn detection, I think this is a separate topic but I'm just going to quickly pull up all the points so you guys can skim through that and if you need a chat after this, we can talk about this. I'm just going to leave that for like five seconds and then we can chat about this offline. I'm quite over time and then the last one is barging and back channeling. I think there's a lot of talk around speech-to-speech models that do some of this but we've been able to see how we could do all of this in speech-to-speech pipelines. You really don't need a speech-to-speech model to do all of this up. Again, I just put this up on the slide and sort of close at that. All right. I don't think we have time for questions. We can take them offline if you have any time but hopefully this was helpful and gave you some insights on what we are seeing in productions with billions of calls at scale. All right. Thanks. We'll see you next time. We'll see you next time. Thank you.