Okay, hello everyone.
So my name is Midam Kim.
I am an ML engineer from ServiceNow, and I'll be talking about a linguistic framework for VoiceAI. So, quick background on me, so you know where I'm coming from.
I'm an ML engineer at ServiceNow, but I'm also a researcher, lifelong researcher, of speech communication in the wild. So my motto is doing linguistics, and what I'm going to be doing today is to hand you that lens of linguistics. Have you experienced VoiceAI failures?
Yeah, every one. So I'm going to introduce an example that I experienced myself.
The bot asked me, could you please spell your first name? And then I slowly start to spell my name. Yes, it is M-I-D-A-M. And the bot says, confirming with you, is it M-I-D-A-N? And then I say, no, it is M-I-D-A-M. And the bot says, thank you for your correction.
Happy to help you today, Midam. And then I get slightly annoyed, more annoyed, because my name is Midam, not Midam. And then it asks me, now, what is your account number? And then I start getting confused. What is that account number thing? And then I try to find information about that. So which one? It must be, and I slowly start spelling the account number. So it is A-X-4-5-1. And then I take time, because I'm not used to reading this strange number. And then the bot cuts me off. And then it says, I couldn't find your record. And then, without even trying, it asks me to repeat that again. Could you please repeat that?
And then I get super annoyed, and then I say, can I talk to a person? I just don't want to deal with you anymore. So this is a very typical pattern of voice AI, unfortunately, at this point. So I just want to navigate how we can solve this problem with linguistics. Voice AI is booming, but users are still often preferring human agents over voice agents.
How can we mitigate this issue? And then, in the first place, what are the actual problems? So I think we can think about a fundamental framework to understand this end-to-end architecture of voice AI, which is called linguistics.
So as all of us already know, human communication is a joint activity like the thing that we're doing right now. So I give you my sounds and words, you hear them, and then if it is a conversation, you're going to give me your sounds and your words, and then this is going back and forth through interaction. And then in this process, we're continuously processing and updating our mental models. So that's a joint activity for human communication, and I would like to say in voice AI human communication, it also has to be a joint activity like this, because that's the only thing that we know about communication as a human being. We have been evolving thousands of years as communicators, and this is what we know. So we expect the same thing from bots.
So let me go over the failure scene of my call with the voice agent in this framework. So you see there's listen and speak for each party. So I start spelling my first name, and then the bot did not hear the difference between M and N correctly, so it's an ASR failure in the listening level. And then the TTS applies only English-centric reading rules to my name. M-I-D-A-M would read as M-I-D-A-M in the American English version. So I'm confused, but at this time I'm kind of generous because that happens a lot, even with human beings. So I'm okay. But then when it brought up account number thing, because I don't know what that is, I'm confused again. But I'm adaptive. I can find it. So I found the number, start reading it, but the ASR did not recognize the word unit correctly, so it cuts me off, and finally it eventually talked over me. So I get really irritated. And then when it asks me for the repetition of the same information, and then it is clear that the bot is not tracking the mental model with me. And then very rudely, it does not even try interactive clarification, which is a common strategy by human beings. So I don't want to deal with this anymore. So I say, can I talk to a person?
So let's go over the framework again.
So these are the linguistic components that are expected and well maintained in human-to-human conversation. So there are listening channels and speaking channels, and there are different components like sounds, words, interaction, and mental model. So the first component is does a bot recognize the user's speech well? And all of these technical terms will fall under this. And then there's going to be this second component, which is words in the listening channel. So does a bot understand the user's words? And then the third one is does a bot wait until the right timing for its turn? It's going to be about listening channel interaction. And then the last part is mental model. So does a bot understand the user's intention in the listening part?
And then we can also go to the speaking channel. So it's going to be about pronunciation for the sound. And also does a bot choose the words the user can understand? And in the interaction part, does a bot speak with the right timing? And lastly, does a bot speak with the information the user actually needs? So there are a lot of engineering or linguistic or cognitive science terms that are here. You can now see that all of those have their right spots in this linguistic framework.
And importantly, these components are interdependent, not separate or independent from each other. They're interdependent and they're aligned. So when you want to do good things about sounds, you have to think about words level. And then when you want to do good things about these sounds and words, you also have to account for interaction, so turn taking or turn detection. And then finally, you want to have good task completion, which is the goal of these mental model layer. Then you have to have all of these. Without any of those components, your voice agent will fail. And then finally, it has to be well aligned. All of these have to be well aligned.
And additionally, you have to keep in mind that this is happening on the timeline. What I mean by that is it is silently tracked. Unlike in chat, in chat you see the history of what was said as text. But in voice agent experience, you say something and the bot says something, you go back and forth, and then see all these waveforms, the air via the vibration in the air, they're all gone. And only the user's mental model is a thing that's left and that matters. So sounds, words, interactions vanish the moment they're spoken, but the mental model perceives and grows over the timeline. So this is what you have to target for user satisfaction.
And then what can we do for the bot to meet the standard of the user?
So what we can do would include choosing good ASR models or configurations and doing some post-processing. Choosing good TTS models, configurations and pre-processing. And carefully curate the vocabulary that can be shared between the bot and the user. And do a good job of turn detection, latency and turn-taking. And very importantly, we have to, it would be great if we can do good emotion detection and handling and context retention. And by context, what I mean is context about all of these. And importantly, it has to be dynamic because things are always changing throughout the course of the call. So we would have to do this management dynamically along the timeline for different kinds of people. So kids or different kinds of people like these will have different expectations that we have to satisfy. Not just when they're happy, but also when they're not happy.
So only then you can pursue a dynamic and truly scalable orchestration of voice AI. So it's a very difficult job to do. We always say that voice is the most natural way of communication, but it is actually not easy. Behind the scene, it is thanks to this linguistic orchestration. When your bot is not good at it, it's a catastrophic failure. So paying attention to this linguistic framework would have lots of business implications, because then you can decrease all of these user frustration, task failures, live agent escalation, abandoned calls, or silent failures.
So in ServiceNow, we have made a good benchmark, a multimodal benchmark called EvaBench. So you can try that to diagnose your voice agent status. Key takeaways. Voice AI is a joint activity between the bot and the user, not just the pipeline. And we must serve users' needs on multiple layers, real time. So it's not that I have given you a fix today, because there's nothing like that. It's just, the fix is in you and your system. But what I have given you today is the linguistic framework you can try to diagnose your system and to build your system upon. You can try Eva, but also you can learn linguistics and hire linguists.
Another thing I want to remind you of is that business implications are linguistic implications and vice versa in this voice AI scene, because voice is fundamentally a linguistic and very human and cognitive experience.
I would like to ask you a longer term question. Speakers adapt. So I'm pretty sure that in this talk with you guys today, you have learned something about me, about my speaking style, what kind of accents I speak, what kind of words I'm using. So next time I see you guys in person, you would find it more comfortable to talk to me, because you have paid attention to me. Right? So speakers are always adapting. So the user will be adapting to your voice agent throughout the call. So is your system ready for them to use you better, use your voice agent better the next time? And languages always change. So is your voice agent ready for language change in one year or six months even?
So thank you. it does not even try interactive clarification, which is a common strategy by human beings. So I don't want to deal with this anymore. So I say, can I talk to a person? So let's go over the framework again. So these are the linguistic components that are expected and well maintained in human-to-human conversation. So there are listening channels, listening channel and speaking channel, and there are different components like sounds, words, interaction, and mental model. So the first component is does a bot recognize the user's speech well? And all of these technical terms will fall under this.
And then there's going to be this second component, which is words in the listening channel. So does a bot understand the user's words? And then the third one is does a bot wait until the right timing for its turn? It's about, it's going to be about listening channel interaction. And then the last part is mental model. So does a bot understand the user's intention in the listening part? And then we can also go to the speaking channel. So it's going to be about pronunciation for the sound. And also does a bot understand the words users are, does a bot choose the words the user can understand? And in the interaction part, does a bot speak with the right timing?
And lastly, does a bot speak with the information the user actually needs?
So there are a lot of engineering or linguistic or cognitive science terms that are here.
You can now see that all of those have their right spots in this linguistic framework.
And importantly, these components are interdependent, not separate or independent from each other. They're interdependent and they're aligned. So when you want to do good things about sounds, you have to think about words level. And then when you want to do good things about these sounds and words, you also have to account for interaction. So turn taking or turn detection. And then finally, you want to have good task completion, which is the goal of these mental model layer. Then you have to have all of these. Without all of those, without any of those, any of those components, your voice agent will fail.
And then finally, it has to be well aligned. All of these have to be well aligned. And additionally, you have to keep in mind that this is happening on the timeline. What I mean by that is it is silently tracked. Unlike in chat, in chat you see the history of what was said that as text. But in voice agent experience, you say something and the bot says something, you go back and forth, and then see all these waveforms, the air via the vibration in the air, they're all gone. And only that user's mental model is a thing that's left and that matters. So sounds, words, interactions vanish the moment they're spoken, but the mental model perceives
and grows over the timeline. So this is what you have to target for user satisfaction. And then what can we do for the bot to meet the standard of the user? So what we can do would include, of course, choosing good ASR models or configurations and do some post-processing. Choosing good TTS models, configurations and pre-processing. And carefully curate the vocabulary that can be shared between the bot and the user. And do a good job of turn-to-turn detection, latency and turn-taking. And very importantly, we have to, it would be great if we can do good emotion detection and handling and context retention. And by context, what I mean is context about all of these.
And importantly, it has to be dynamic because things are always changing throughout over the course of the call. So we would have to do this management dynamically along the timeline for different kinds of people. So kids or different kinds of people like these will have different expectations that we have to satisfy. Not just when they're happy, but also when they're not happy. So only then you can pursue a dynamic and truly scalable orchestration of voice AI. So it's a very difficult job to do.
We always say that voice is the most natural way of communication, but it is actually not easy. Behind the scene, it is thanks to this linguistic orchestration. When your bot is not good at it, it's a catastrophic failure.
So paying attention to this linguistic framework would have lots of business implications, because then you can decrease all of these user frustration, task failures, live agent escalation, or abandoned calls, or silent failures.
So in ServiceNow, we have made a good benchmark, a 2N benchmark called EvaBench. So you can try that to diagnose your voice agent status. Key takeaways. So voice AI is a joint activity between the bot and the user, not just the pipeline. And we must serve users' needs on multiple layers, real time. So it's not that I have given you a fix today, because there's nothing like that. It's just, the fix is in you and your system. But what I have given you is, today is, the linguistic framework you can try to diagnose your system and to build your system upon. You can try Eva, but also you can learn
linguistics and hire linguists. Another thing I want to remind you of is that business implications are linguistic implications and vice versa in this voice AI scene, because voice is fundamentally a linguistic and very human and cognitive experience. I would like to ask you a longer term question. Speaker's adapt. So I'm pretty sure that in this talk, in my talk with you guys today, you have learned something about me, about my speaking style, what kind of accents I speak, what kind of words I'm using. So next time I see you guys in person, you would find it more comfortable to talk to me, because
you have paid attention to me. Right? So speakers are always adapting. So the user will be adapting to your voice agent throughout the call. So is your system ready for them to use you better, use your voice agent better the next time? And languages always change. So is your voice agent ready for language change change in one year or six months even? So thank you.