Let's just do a couple of quick questions and then we'll jump right in. How many of us in the room here have built Voice AI agents? Okay, that's a pretty good audience here. And how many of you guys have built AI agents that have been deployed in production? Not bad.
Okay, cool. So we'll talk about what typically happens, right? Everyone's talking about Voice AI agents. The one pill solution to pretty much everything in the world today is Voice AI agents. So everyone's building one and trying to deploy that. They sound great when you're building that in your dev landscape. And then the moment you take this from a proof of concept to production, things start failing. So we'll walk through these five different angles of how or what we have seen at Plevo with Voice AI agents. But just before that, a quick intro from my side. I am Venki, the founder and CAO. She used the title agent engineering manager.
I'm calling myself chief agent officer from a title standpoint. Okay, so what is, you know, why are we even qualified for this discussion? And what are we seeing that a lot of companies don't get to see? I'll talk a bit about our journey in terms of how we have come along so far and then jump right in. You know, we've been around for about 14 years. Our journey has been a developer API platform. And then now in an AI agent business, we started with voice and SMS APIs back in 2011. And then now we are primarily focused on our AI agent offering, the full stack. On our platform, we see over a billion voice calls each month across the globe.
And that is where we have seen a lot of these patterns emerge in terms of how, when we work with our customers, what happens with their voice AI agents in production. We're a 90 member team. And we have 50 million funding in the bank. Fun fact, this is not from external VC investors. This is all from being a profitable company, having put that cash in the bank over these years. Some customers we power across the globe.
We've just left some logos in there. But primarily from an offering standpoint, I would cohort this into three different buckets. One is a programmable AI agent offering. We call it a speech pipeline, not a true speech-to-speech product yet, but that's a programmable offering. We also have an AI agent studio. It's a no-code visual builder. And then, like I said, we started with voice APIs. So we obviously have built this out over the last 14 years, the SIP trunking and the audio streaming layer. So we don't rely on other folks for the telephony or the carrier layer. That's the bread and butter business we've built over all these years.
And that's on top of which our AI agent platform sits. Okay, with that, let's get into this, right? Which I'm sure, since you guys have all built AI agents, you've all seen this or built this in one manner or another. And we'll spend more time on this in terms of how the entire pipeline looks, right? What we see with customers is, and I'm sure you guys can all relate to this, is anyone thinking about AI agents, what they do is they pick up a bunch of these orchestration frameworks and they do a pretty good job, LiveKit or PipeCat, you know, build their AI agent on top of that. They think they can just orchestrate these different four layers, speech to text, LLM, and TTS
with turn detection in between, and we're off to the races. My AI agent works in a POC and it's good to work in production. Typically that's what happens. They sort of measure their latencies and you can see some indicative latencies on the slide at each layer. And they're like, yeah, this seems good for me for what I need, so let's position production. And then the production woes start to kick in and you see all sorts of failure modes, which we are going to spend most of the time on in this talk. I've kept some time at the end for Q&A if you want to have questions, but we'll jump right in to different failure modes we see.
Let's start with the first one, which everyone talks about. This is the most spoken about failure mode, which is latency. I think we have a few AI agent talks today, or voice AI agent talks today. I'm pretty sure everyone's going to touch upon this specific failure mode, which is why I'm bringing this right up in terms of how this entire experience is for users, right? Typically, most folks measure this by time to first audio, so the time when your users stop speaking to your agent starts speaking, right? And I think you've probably seen this if you guys have built voice AI agents on what good or natural feels like, what annoying feels like or noticeable feels like,
and then what annoying feels like, which is different tiered steps. We notice most people want to be under 550 because that's what's advertised by platforms or solutions or layers, but I think most end up between 750 to 1.2. That's where most of the folks end up at. The really bad performing ones end up more than 1.2, and then you start to see users hang up. Now, I'll share with you what we've seen practically in production with customers using this at different layers, and then solutions to some of these. The way we want to think about this layer is a balance between these three, which is cost, intelligence, and latency, right? And why do I bring these three up?
Because they're interrelated. I think one of the things, I was just chatting with a couple of folks outside. One of the things, last year, we've seen a lot of innovations, a lot of intelligence spike on the LLM side of things, right? And most of the intelligence has come in terms of thinking or reinforcement learning and so on and so forth. The irony with voice AI agents is almost always your LLM or the agent that's talking has to have thinking turned off, right? So all the advancements we've had in the LLM layer in the last year, none of that even applies here now, right? Obviously, you have better models that can do better instruction following or tool calling,
but pretty much all of your intelligence that's been built in on the thinking layer is all off by default if you want it to be fast enough. So that's one of the ironies that we come up with. So then how do you sort of balance intelligent cost and latency? Let's look at some of these options that are out there in the market, right? So I'm specifically picking LLM because if you looked at the previous chart, LLM is your highest latency bucket that adds to this, right? And if you look at frontier models, which I think most folks start by default, your OpenAI, your Claude, your Gemini's, you know, P50, TTFT is roughly around 450 to 500
on a good day, and it can get spiky, right? P90, P95 can go easily upwards of 1.2, 1.3 seconds, and that's not good for the overall agent experience. So that's your frontier model. Now, there's another option, which is your Cerebris or the Grok that is famous and popular for spitting out a lot of tokens or tokens very fast, right? These work, but for you to get dedicated latency or time-to-first token on these, you need dedicated capacity, and that is really expensive. That's where I spoke about the cost as being one of the things to balance, right? It's really expensive, and then, you talk to anyone from the Grok team or the Cerebris team, they'll tell you,
you need to book 12 months in advance for dedicated capacity. They're booked out for the next 12 months. So that's a pretty expensive option, and then you really need to be sure that the model you're deploying on some of these infrastructure layers will be here 12 months from now, and it's a big investment and a big unknown. So what's a realistic option for production-grade agents that are good quality and end up balancing three of these? This is what has worked for us, which is the open-source models. There are obviously a lot of them in terms of the variety and variations you can pick. I'm specifically talking about the two we work with, Qwen 3.5 and Gemma 4.
These are cutting-edge open-source models out in the market right now. You talk to anyone from the Grok team or the Cerebris team, they'll tell you, you need to book 12 months in advance for dedicated capacity. They're booked out for the next 12 months. So that's a pretty expensive option, and then you really need to be sure that the model you're deploying on some of these infralayers will be here 12 months from now, and it's a big investment and a big unknown.
So what's a realistic option for production-grade agents that are good quality and end up balancing three of these? This is what has worked for us, which is the open-source models. There are obviously a lot of them in terms of the variety and variations you can pick. I'm specifically talking about the two we work with, QEN 3.5 and Gemma 4. These are cutting-edge open-source models out in the market right now, and we've done a lot of benchmarking around this in how they work.
It can be scary to think, okay, I have the models. Now I have to host them, run them on my own GPUs and so on and so forth, but if you are consistently targeting under 300 MS, we've seen this to be a great option to balance between latency, cost, and intelligence.
Now, some more deep dive here. If you're doing only English, QEN 3.5 or Gemma both work fine, but if you're doing multilingual, right, international audiences, different languages, Gemma 4 is a much better model for that. We have seen token fertility evals. Essentially what that means is if I were to de-jargonize that is, how many tokens does it take to generate one word in that language? Okay, so Gemma is much, much better, at least 2.5 to 3x better than QEN 3.5 from that perspective, so your time to words is much faster on Gemma 4, everything else equal, right, on a multilingual basis.
Now, what sizes do you pick at the LLM layer? The mixture of experts usually works fine. The three or four billion mixture of experts usually works fine. The problem with mixture of experts is if anyone goes down, wants to go down the direction of fine-tuning, that can be a challenge, because fine-tuning mixture of experts models are not easy. You can end up breaking the model a lot of times, so that's one challenge we see with mixture experts, but usually, out of the box, it gets you 90% closer to where you want to be, even without any fine-tuning or custom work done on the model. So that's the advantage of mixture experts.
Now, if you want to fine-tune and you want to go deeper and say, look, I'm working for a specific domain, healthcare, what have you, right, and I want to make sure I'm able to fine-tune my model, you want to start at least with the $8 billion, $12 billion, at least from where we are today. Maybe six months from now, a $4 billion model beats the $8 billion model, hands down, but for today, what we've seen is you minimum need an $8 billion or $12 billion model, because you're looking for two things in these models. One, obviously, fast tokens, but good instruction following, okay, and the second thing is very high success ratio in tool calling, because if you can do these two things well, then you are on to 70%, 80% there for not even having to fine-tune any model. Models will work out of the box, right? So that's been our recipe. We've actually run two flavors, one a fine-tune model for specific industries, and then for most generic use cases, a MOE model just works out of the box.
There are a few more tips and tricks we'll talk about in the upcoming slides where we see failure models, but that's where we stand from a latency LLM standpoint. All right, I'm running short on time, so I'm going to fast-track this. Now, there are a couple of other flavors in this. People build agents with a mixture of models. What they do is, for the talking part of it, they have a conversational model, which is a much lower, smaller model, and then maybe even a three billion model, and then for tool calling, they have a much larger model, so they have an improved tool calling success ratio there.
The second one is, assume your transcriptions are going to be brittle. That's something you want to live by when you're building AI agents, even if you have the best transcription engine out there, and I'll show you why, right? The state-of-the-art transcription engines out in the market, you know, get you to four to six percent word error rate, right? And this is on known eval sets. On real world, noisy calls with people having different accents, domain vocabulary, and so on and so forth, those usually end up in the double digits from a word error rate perspective, right?
Now, obviously, you can fine-tune, pick up an open source model and fine-tune, but we see typically what breaks here often, and there are patterns here in terms of what breaks. So, proper nouns, jargons, phone numbers, random missing digits with phone numbers, wrong substitutions. I'll walk through some examples of how you solve for these. Addresses, when you're trying to collect a long address, the transcription engine could just end up missing some parts of it.
Code-switch languages. I'll just take an example of a language I speak, because that was easy for me to put on the slide, where if you were to take English, but written in a different script, that's what's used for Hindi, right? This is English written in that script, right? Whereas the actual English version of this is, hello, how are you? So if I'm addressing an audience in a different country where I have code-switched languages and I start getting my English in a different script, everything starts breaking from the transcription engine to the LLM layer and then beyond, because your LLM starts then producing output in that script a lot of times, and then your TTS messes up. Okay, so this is very important to be careful about, and if you want to build your agent independent of the transcription engine, you need to build a layer that normalizes all of this, right? We'll talk about solutions in a minute, and there is the other case, which is Hindi, and just Latin or Roman, right? Which is this is Hindi, but it reads English, which again messes up everything downstream. Those are just examples. This applies to Arabic, Mandarin, Japanese, what have you, pretty much any language.
So what actually moves the needle at the transcription layer? For proper nouns, we recommend you using not just keyword boosting. I think a lot of transcription engines provide you keyword boosting where you can put in specific words into their engine, but doing dynamic keyword boosting. What that means is don't keep the keyword for the entire state of the call. Just add that dynamically when you think you need that as an answer so that you get the highest accuracy, meaning at different states of the call, the transcription engine will have different keywords boosted during different phases, right? And that's what we've seen works best because if you just pollute your context of the transcription engine with tons of keywords, it'll start hallucinating again, right? So that's what we see typically working best.
Yeah, post-process. Post-process your transcripts with an LLM, right? Because your LLM has domain context, your transcription engine does not. So a lot of words that it would say, I'll give you some examples, may not make sense. This is transcription like a phone number from a transcription engine, right? Like what do you think that E is, right? If you give it to an LLM, it knows that's a three. Similarly, what that one is, it's a digit one. So your transcription engine a lot of times could mess that up, but when you post-process it with an LLM layer, it'll instantly correct that from a collection standpoint.
And the last one, transliteration is your STT output, that's multilingual, also gets normalized using either an LLM, you first transliterated, or use some kind of a neural transliteration engine. There are a lot of them open source, you can just pick one of them, right? That will do all of that work for you, send cleaned transcripts consistently, independent of the transcription engine to your LLM. All right, the third one we typically see is collecting data. This is where I think
That E is, right? If you give it to an LLM, it knows that's a three. Similarly, what that one is, it's a digit one. So your transcription engine a lot of times could mess that up, but when you post-process it with an LLM layer, it'll instantly correct that from a collection standpoint. And the last one, I said, transliteration is your STT output, that's multilingual, also gets normalized using either an LLM, you first transliterated, or use some kind of a neural transliteration engine. There are a lot of them open source, you can just pick one of them, right? That will do all of that work for you, send cleaned transcripts consistently, independent of the transcription engine to your LLM.
All right, the third one we typically see is collecting data. This is where I think 50 to 60% of AI agents mess up pretty badly. And we to think of it as a UX problem, but just for voice. So think data models and not a transcript coming into an LLM and trying to figure out what the transcript said. So let's take some inspiration from, I'm assuming most of us are developers here, take inspiration from Python's data classes, PyDantic, Zod from TypeScript, or form fields in the UI, right? If you start thinking of it from that problem statement, we have seen accuracy grow from 30% to 95% from a data collection standpoint when you start thinking in that manner.
So, decide your shape before you ask, right? Instead of keeping it open-ended, can you keep it constrained? So, can a phone number be a phone number type field? The moment you do that, right, you know, how many digits it needs to have. You can do validation on top of that, right? And then what sort of allowed values can even be there? So, in the previous example we saw, if an E comes in in the middle of a phone number and you know it's a phone number, you instantly know, either you smart guess that to three and confirm that with the user, or you know that's an error and then you validate that and ask the user to repeat again, right?
So that's, I think, one of the common patterns we've seen here from a collection pattern. Name, I think, is the interesting one. I've just picked a hard to pronounce name. Like, there is no way a human is going to get this right and no way a transcription engine will get this right, how many ever times you do this, right? So the moment you start thinking of this as fields and then have rules and then confirmation mechanisms on spelling this, letter by letter, only then you kind of get it right. Otherwise, it's going to mess up pretty badly in terms of how you collect this on a voice call. And that's just an example of what I'm talking about in terms of the data collection piece of it.
Another place where it goes badly, dramatically, is relative values, date being one of the examples. If somebody says, next week, Wednesday, eight, it could mean 8 a.m., 8 p.m., and then figuring out what that date actually is, again, now becomes a very constrained problem. If you knew this was a date time field and I'm collecting a date time field and then you take the current date and then figure out what this value would be based on that, right? So that's how you want to make sure like you do this with a combination of the LLM with the tool calling and the tool calling is doing a lot of this heavy lifting for you from a field standpoint.
Yeah, and then you make, you run like this from a unit test perspective. So, all of your evals need to start treating these fields as unit tests and as long as your unit tests sort of validate and pass, you know your agent is going to be sort of reliable and repeatable. You don't run hundreds of end-to-end agent test cases just to find out one field collection is broken. You do your evals at a field level and a unit test level.
And then, yeah, I think this mindset makes everything more structured instead of hoping I'll put a ton of prompt, keep changing the prompt by a few characters every time and somehow my prompt engineering is going to make LLM much more instruction tuned and sort of magically start following some of these things. So, in fact, right, we have seen us get to 95, 97% accuracy without having to fine-tune a model, right? And the trick is basically just breaking down your context of what the agent is doing at that point with specific states of what the agent is going through.
All right, I'm just going to quickly skip through this from a time standpoint. I just see I got three more minutes. Hopefully, that's a bug, but we'll leave it at that. So, this is the fourth area where we see issues coming in. Most folks take the LLM output and then we send it to a TTS. Obviously, I think there are a lot of good TTSs in the market that take care of a lot of heavy lifting, but a lot of times it messes up. What we recommend and what we've seen is you usually want to have a normalization layer between your LLM and what is fed to a TTS. You don't send your LLM output directly to a TTS, right? And we'll just walk through some examples.
The basics, which is strip emojis, markdown before any synthesis into the TTS. Most orchestration pipelines do this, a LiveCAD or a PipeCAD will do that for you if you just set a few flags. But just make sure if you're not using them or building from scratch that you've set this explicitly because you don't want an emoji showing up on something read out or markdown showing up there.
I think some more common ones, custom dictionaries. Most TTS engines provide this to you, how to pronounce custom words, whether it's proper nouns, brands, acronyms, and so on and so forth. So set those in when you go from your LLM to your TTS output because if you don't, you're going to mess that up. And I'll show you an example of how we test that. The other one is most engines also give you speed. So if you know you're pronouncing an entity, slow down, have your agent slow down at 0.8x or 0.7x so that it's able to enunciate on that specific entity and doesn't mess up how it's pronouncing an email or a phone number or a name letter by letter.
And yeah, just normalize all the messy stuff, right? Like emails, currency, dates. Don't leave it to the TTS to do it. Most of them do it, but don't leave it to the TTS to do it. Build your normalization layer at your end so that tomorrow you think you need to switch TTS or for whatever reason the first one's down and you want to use another TTS, you're able to not rely natively on the TTS's engine but you're building this in-house for this to be managed.
And then yeah, I don't have my batch here but I don't have my last name on that. So my first test is if it cannot pronounce my last name or my company's name, it's already dropping the ball. So my last name is Balasobramanian and if you cannot pronounce that using a voice AI agent, that's a check for me. I know the agent will mess up a lot of words that need to be spelled out day by day.
The second one is our company named Plivo. So a lot of engines pronounce it Plivo or Plivo and so on and so forth but I think specifically being able to control this in your pipeline is super critical and then if you're building a customer facing product then you know, sort of give this option to your customers.
I'm just going to skim through the last two slides. I'm running badly over time. End of turn detection, I think this is a separate topic but I'm just going to quickly pull up all the points so you guys can skim through that and if you need a chat after this, we can talk about this. I'm quite over time and then the last one is barging and back channeling. I think there's a lot of talk around speech-to-speech models that do some of this but we've been able to see how we could do all of this in speech-to-speech pipelines. You really don't need a speech-to-speech model to do all of this. Again, I just put this up on the slide and sort of close at that.
All right. I don't think we have Give this option to your customers. I'm just going to skim through the last two slides. I'm running badly over time. End of turn detection, I think this is a separate topic but I'm just going to quickly pull up all the points so you guys can skim through that and if you need a chat after this, we can talk about this. I'm just going to leave that for five seconds and then we can chat about this offline.
I'm quite over time and then the last one is barging and back channeling. I think there's a lot of talk around speech-to-speech models that do some of this but we've been able to see how we could do all of this in speech-to-speech pipelines. You really don't need a speech-to-speech model to do all of this. Again, I just put this up on the slide and close at that. All right. I don't think we have time for questions. We can take them offline if you have any time but hopefully this was helpful and gave you some insights on what we are seeing in productions with billions of calls at scale. All right. Thanks. We'll see you next time. We'll see you next time. Thank you.
is, like, how many tokens does it take to generate one word in that language? Okay, so Gemma is much, much better, at least 2.5 to 3x better than QEN 3.5 from that perspective, so your time to words is much faster on Gemma 4, everything else equal, right, on a multilingual basis. Now, what sizes do you pick at the LLM layer? The mixture of expert usually works fine. The three or four billion mixture of expert usually works fine. The problem with mixture of expert is, like, if anyone goes down, wants to go down the direction of fine-tuning, that can be a challenge, because fine-tuning, mixture of experts models are not easy. You can end up breaking the model
a lot of times, so that's one challenge we see with mixture experts, but usually, out of the box, it gets you 90% closer to where you want to be, like, even without any fine-tuning or custom work done on the model. So that's the advantage of mixture experts. Now, if you want to fine-tune and you want to go deeper and say, like, look, I'm working for a specific domain, healthcare, what have you, right, and I want to make sure I'm able to fine-tune my model, you want to start at least with the $8 billion, $12 billion, at least from where we are today. Maybe six months from now, a $4 billion model beats the $8 billion model, hands down, but for today, what we've seen is
you minimum need an $8 billion or $12 billion model, because you're looking for two things in these models. One, obviously, fast tokens, but good instruction following, okay, and the second thing is, like, very high success ratio in tool calling, because if you can do these two things well, then you are on to, like, 70%, 80% there for not even having to fine-tune it, fine-tune any model. Like, models will work out of the box, right? So that's been our recipe. We've actually, we run two flavors, one a fine-tune model for specific industries, and then for, you know, most generic use cases, a MOE model just works out of the box. There are a few more tips
and tricks we'll talk about in the upcoming slides where we see failure models, but that's where we stand from a latency LLM standpoint. All right, I'm running tired on time, so I'm going to fast-track this. Now, there are a couple of other flavors in this. People build agents with a mixture of models. What they do is, you know, for the talking part of it, they have a conversational model, which is a much lower, smaller model, and then, you know, maybe even a three billion model, and then for tool calling, they have a much larger model, so they have a improved tool calling success ratio there.
Sorry. The second one is, assume your transcriptions are going to be brittle. Like, that's something you want to sort of live by when you're building AI agents, even if you have the best transcription engine out there, and I'll show you why, right? Like, the state-of-the-art transcription engines out in the market, you know, sort of get you to four to six percent word error rate, right? And this is on known eval sets. On real world, noisy calls with, you know, sort of accents, like people having different sort of accents, domain vocabulary, and so on and so forth, like those usually end up in the double digits from a word error rate perspective, right? Now, obviously,
you can fine-tune, you know, pick up an open source model and fine-tune, but we see typically, like, what breaks here often, and there are patterns here in terms of what breaks. So, proper nouns, jargons, phone numbers, like random missing digits with phone numbers, wrong substitutions. I'll walk through some examples of, like, how you solve for these. Addresses, when you're trying to collect a long address, you know, the transcription engine could just end up missing some parts of it. Code-switch languages. I'll just take an example of a language I speak, because that was easy for me to put on the slide, where, you know, like, if you were to sort of take English,
but written in a different script, that's what's used for Hindi, right? Like, this is English written in that script, right? Whereas, like, the actual English version of this is, hello, how are you? So if I'm addressing an audience in a different country where I have code-switched languages and I start getting my English in a different sort of script, everything starts breaking from the transcription engine to the LLM layer and then beyond, because your LLM starts then producing output in that sort of script a lot of times, and then your TTS messes up. Okay, so this is very important to be careful about, and if you want to build your agent independent of the
transcription engine, you need to build a layer that normalizes all of this, right? We'll talk about solutions in a minute, and there is the other case, which is Hindi, and just Latin or, you know, Roman, right? Which is, like, this is Hindi, but it reads English, which again messes up everything, you know, downstream. Those are just examples. This applies to, you know, Arabic, Mandarin, Japanese, what have you, pretty much any language. So what actually moves the needle with a, at the transcription layer? For proper nouns, we recommend you using not just keyword boosting. I think a lot of transcription engine engines provide you keyword boosting where you can put in
specific words into their engine, but doing dynamic keyword boosting. What that means is don't keep the keyword for the entire state of the call. Just add that dynamically when you think you need that as an answer so that you get the highest accuracy, meaning at different states of the call, the transcription engine will have different keywords boosted during different phases, right? And that's what we've seen works best because if you just pollute your context of the transcription engine with tons of keywords, it'll start hallucinating again, right? So that's what we see typically working best. Yeah, post-process. Post-process your transcripts with an LLM, right?
Because your LLM has domain context, your transcription engine does not. So a lot of words that it would say, I'll give you some examples, may not make sense. This is transcription, like a phone number from a transcription engine, right? Like what do you think that E is, right? If you give it to an LLM, it knows that's a three. Similarly, like what that one is, it's a digit one. So your transcription engine a lot of times could mess that up, but when you post-process it with an LLM layer, it'll instantly correct that from a collection standpoint. I mean, and the last one, like I said, transliteration is your STT output, that's sort of, you know, multilingual,
also gets normalized using either an LLM, you first transliterated, or, you know, use some kind of a neural transliteration engine. There are a lot of them open source, you can just pick one of them, right? That will do all of that work for you, send cleaned transcripts consistently, independent of the transcription engine to your LLM.
All right, the third one we typically see is collecting data. This is where I think 50 to 60% of AI agents mess up pretty badly. And like, we like to think of it as a UX problem, but just for voice. So think data models and not a transcript coming into an LLM and trying to figure out what the transcript said. So let's take some inspiration from, I'm assuming most of us are developers here, you know, take inspiration from Python's data classes, PyDantic, Zod from TypeScript, or form fields in the UI, right? Like, if you start thinking of it from that problem statement, we have seen accuracy grow from 30% to like 95% from a data collection standpoint when you start thinking
in that manner. So, like, decide your shape before you ask, right? Like, instead of keeping it open-ended, can you keep it constrained? So, can a phone number be a phone number type field? The moment you do that, right, you know, like, how many digits it needs to have. You can do validation on top of that, right? And then what sort of allowed values can even be there? So, in the previous example we saw, if an E comes in in the middle of a phone number and you know it's a phone number, you instantly know, like, either you smart guess that to three and confirm that with the user, or, you know that's an error and then you validate that and ask the user to repeat again,
right? So, so that's, I think, one of the common patterns we've seen here from, from a collection pattern. Name, I think, is the, is the interesting one. I've just picked a, you know, a, a hard to pronounce name. Like, there is no way a human is going to get this right and, and no way a transcription engine will get this right. How many ever times you do this, right? So the moment you start thinking of this as fields and then have rules and then confirmation mechanisms on, on spelling this, you know, sort of letter by letter, only then you kind of get it right. Otherwise, it's going to mess up pretty badly in terms of how you collect this on a voice call. And,
and that's just an example of, you know, what I'm talking about in terms of the, the data collection piece of it. Another place where it goes badly, dramatically, is relative values, date being one of the examples. If somebody says, next week, Wednesday, eight, it could mean 8 a.m., 8 p.m., and then figuring out what that date actually is, again, now becomes a very constrained problem. If you knew this was a date time field and I'm collecting a date time field and then you take the current date and then figure out what this value would be based as that, right? So, so that's how you want to make sure like you do this with a combination of the LLM with the tool calling
and the tool calling is doing a lot of this heavy lifting for you from a, from a field standpoint.
Yeah, and then you make, you, you run like this from a unit test perspective. So, all of your evals need to start treating these fields as unit tests and as long as your unit tests sort of validate and pass, you know your agent is going to be sort of reliable and repeatable. You don't, you know, run hundreds of end-to-end agent test cases just to find out, you know, one field collection is broken. You do your evals at a field level and a unit test level.
And then, yeah, like I said, I think this mindset makes everything more structured instead of hoping I'll put a ton of prompt, keep changing, you know, the prompt by a few characters every time and somehow my prompt engineering is going to make LLM much more instruction tuned and sort of magically start following some of these things. So, in fact, like I said, right, like we have seen us get to 95, 97% accuracy without having to fine-tune a model, right? And then, and the trick is basically like just breaking down your context of what the agent is doing at that point with specific states of what the agent is going through. All right, I'm just going to quickly skip through
this from a time standpoint. I just see I got three more minutes. Hopefully, that's a bug, but we'll leave it at that. Okay. So, so this is the fourth area where we see issues coming in. Most folks take the LLM output and then we send it to a TTS. Obviously, I think there are a lot of good TTSs in the market that take care of a lot of heavy lifting, but a lot of times it messes up. What we recommend and what we've seen is you usually want to have a normalization layer between your LLM and what is fed to a TTS. You don't send your LLM output directly to a TTS, right? And we'll just walk through some examples. The basics, which is strip emojis, markdown before any synthesis
into the TTS. Most orchestration pipelines do this, like a LiveCAD or a PipeCAD will do that for you if you just set a few flags. But just make sure if you're not using them or buildings from scratch that you've set this explicitly because you don't want an emoji showing up on something read out or markdown showing up there. Okay. I think some more common ones, custom dictionaries. Most TTS engines provide this to you, like how to pronounce custom words, whether it's proper nouns, brands, acronyms, and so on and so forth. So set those in when you go from your LLM to your TTS output because if you don't, you're going to mess that up. And I'll show you an example
of how we test that. The other one is most engines also give you speed. So if you know you're pronouncing an entity, slow down, have your agent slow down. So at 0.8x or 0.7x so that it's able to like enunciate on that specific entity and doesn't mess up how it's pronouncing an email or a phone number or a name letter by letter. And yeah, just normalize all the messy stuff, right? Like emails, currency, dates. Don't leave it to the TTS to do it. Most of them do it, but don't leave it to the TTS to do it. Like build your normalization layer at your end so that tomorrow you think you need to switch TTS or you know, for whatever reason the first one's down and you want to use
another TTS, you're able to sort of not rely natively on the TTS's engine but you're building this in-house for this to be managed. And then yeah, I think I don't have my batch here but I don't have my last name on that. So my first test is if it cannot pronounce my last name or my company's name, it's already dropping the ball. So my last name is Balasobramanian and if you cannot pronounce that using a voice AI agent, like that's a check for me. I know like you know, the agent will mess up a lot of words that you know, need to be spelled out day by day. The second one is our company named Plivo. So a lot of engines pronounce it Plivo or Plivo and so on and so forth
but I think specifically being able to control this in your pipeline is super critical and then if you're building a customer facing product then you know, sort of give this option to your customers. I'm just going to skim through the last two slides. I'm running badly over time. End of turn detection, I think this is a separate topic but I'm just going to quickly pull up all the points so you guys can skim through that and if you need a chat after this, we can talk about this.
I'm just going to leave that for like five seconds and then we can chat about this offline. I'm quite over time and then the last one is barging and back channeling. I think there's a lot of talk around speech-to-speech models that do some of this but we've been able to see how we could do all of this in speech-to-speech pipelines. You really don't need a speech-to-speech model to do all of this up. Again, I just put this up on the slide and sort of close at that. All right. I don't think we have time for questions. We can take them offline if you have any time but hopefully this was helpful and gave you some insights on what we are seeing in productions
with billions of calls at scale. All right. Thanks.
We'll see you next time. We'll see you next time. Thank you.