"My name is... my name is...": A Linguistic Map for Voice Agents — Midam Kim, ServiceNow
Description
Midam Kim spells her first name for a voice agent, letter by letter. The bot reads it back with an N on the end. She corrects it. The bot replies thank you for your correction, then says her name wrong anyway, because its pronunciation rules are English and her name is not. Moments later it asks for an account number she has to go looking for, cuts her off while she reads the unfamiliar string aloud, announces it cannot find her record, and asks her to repeat the whole thing without trying anything different. She asks for a human. Kim is an ML engineer at ServiceNow and a longtime researcher of speech in the wild, and she uses that call to argue that voice AI failures are not a scattered list of bugs but a structured one, and that linguistics already has the structure. Her map is a grid. Each party has a listening channel and a speaking channel, and each channel operates at four levels: sounds, words, interaction, and mental model. Recognition failures sit in one cell, pronunciation in another, turn taking in a third, intent tracking in a fourth. Every familiar problem lands somewhere, which is the point. The cells are interdependent rather than separable, so fixing sounds without words, or words without timing, does not hold. Her sharpest observation concerns what survives a call. In chat the history sits on screen as text. In speech the sounds vanish as they are spoken, and the only thing that accumulates is the user's mental model, which is what you are really designing for. Speaker info: - https://www.linkedin.com/in/midamkim/ Timestamps: 0:00 - A call that goes wrong, step by step 3:39 - Conversation as a joint activity 5:25 - Replaying the failure through that lens 7:14 - Eight cells: two channels, four levels 8:56 - Why the layers cannot be fixed separately 9:49 - What vanishes, and what accumulates 10:42 - What to actually work on 12:28 - Diagnosing your own agent 14:13 - Speakers adapt, and languages change
Summary
Generated by claude-sonnet-4-5-20250929At-a-Glance
- Verdict: Watch fully
- Core thesis: Voice AI failures stem from treating conversation as a pipeline instead of a joint activity requiring alignment across sounds, words, interaction, and mental model layers—all tracked dynamically and silently over time
- Why it matters: Ken's agent systems will need voice interfaces; this framework diagnoses where voice orchestration fails and why interdependence/timeline/context management are critical, not just ASR/TTS quality
- Best use: Use as design checklist when scoping voice agent architecture; apply the 4-layer framework (sounds/words/interaction/mental-model) to evaluate vendor claims, diagnose escalation/abandonment patterns, and identify gaps in context retention or turn-taking that cause catastrophic UX failures
Executive Summary
Midam Kim, an ML engineer and speech communication researcher at ServiceNow, presents a linguistic framework for diagnosing and building voice AI systems that don't enrage users. She opens with a failure story: a voice bot misrecognizes her name (M-I-D-A-M as M-I-D-A-N), mispronounces it using English-centric TTS rules, cuts her off mid-account-number, and demands repetition without clarification—a cascade that ends with 'can I talk to a person?' This is not a pipeline problem; it is a failure to treat conversation as a joint activity.
The framework maps human conversation to voice AI across two channels (listening and speaking) and four layers: sounds (ASR/TTS accuracy), words (vocabulary understanding/choice), interaction (turn-taking/timing/latency), and mental model (intention/context/emotion tracking). These layers are interdependent—good ASR means nothing if turn-detection cuts users off, and good TTS means nothing if the bot doesn't retain context. Critically, sounds and words vanish after utterance; only the user's mental model persists and grows over time, so that mental model is the UX target.
Kim emphasizes that voice is 'silently tracked'—unlike chat, there is no scrollback. Users adapt to the bot's style during the call, and the bot must dynamically manage context, emotion, and vocabulary in real time for diverse users (kids, accents, frustration states). Without alignment across all four layers, voice agents hit catastrophic failure modes: user frustration, task failure, live agent escalation, abandoned calls, or silent failures where users give up without escalating. ServiceNow built EvaBench, a multimodal benchmark to diagnose voice agent status. Kim's final questions: Is your system ready for users to adapt and use it better next time? Is your agent ready for language change in six months?
Key Takeaways
- Claim: Voice AI must be treated as a joint activity between bot and user, not a pipeline, because human communication evolved as continuous mutual mental model updating | Evidence: Kim's call failure: bot misrecognizes M vs N (ASR), mispronounces name (TTS), cuts user off mid-number (turn-detection), demands repetition without clarification (no interactive repair), causing escalation | Implication: Voice agent orchestration must coordinate ASR, TTS, turn-taking, and context retention as interdependent systems; optimizing one layer in isolation will still produce catastrophic UX failures
- Claim: Voice conversation is 'silently tracked': sounds and words vanish after utterance, leaving only the user's mental model, which persists and grows over time | Evidence: Unlike chat (scrollback visible), voice is waveforms in air that disappear; the user's mental model is the only artifact that matters for satisfaction | Implication: Ken's voice agents must target mental model continuity—context retention, emotion handling, and dynamic vocabulary—because users judge the entire interaction by whether the bot 'understood' them over time, not by per-turn accuracy
- Claim: The linguistic framework has four interdependent layers (sounds, words, interaction, mental model) across two channels (listening, speaking), and failure in any layer causes system failure | Evidence: Listening channel: ASR (sounds), NLU (words), turn-detection (interaction), intent/context (mental model); Speaking channel: TTS pronunciation (sounds), vocabulary choice (words), turn-timing (interaction), information relevance (mental model) | Implication: When evaluating voice vendors or building voice systems, Ken should audit all eight components (4 layers × 2 channels) for alignment and interdependence, not just ASR/TTS quality in isolation
- Claim: Without dynamic context management along the timeline, voice agents cannot handle emotion, adaptation, or diverse user types (kids, accents, frustration states) | Evidence: Kim states context must cover all layers and be managed dynamically throughout the call; users adapt to the bot's style, and the bot must adapt to language change over months | Implication: Ken's voice orchestration must support stateful, evolving context across calls (user learning, emotion state) and anticipate vocabulary/accent drift, not just per-session context
- Claim: Business metrics (user frustration, task failure, live agent escalation, abandoned calls, silent failures) are direct linguistic failures in the four-layer framework | Evidence: Kim explicitly maps business implications to linguistic orchestration failures; paying attention to the framework decreases all five failure modes | Implication: Ken should instrument voice agents to trace escalations/abandonments back to specific layer failures (ASR errors, turn-detection cutoffs, missing clarification, context loss) for ROI-driven optimization
- Claim: ServiceNow built EvaBench, a multimodal benchmark to diagnose voice agent status across the linguistic framework | Evidence: Kim recommends trying EvaBench to assess voice agent quality; also suggests learning linguistics and hiring linguists | Implication: Ken should evaluate EvaBench for voice agent diagnostic tooling and consider whether to hire linguists or speech researchers for voice system design/audit | Caveat: EvaBench details (availability, metrics, open-source status) not provided in the talk
- Claim: Users adapt to voice agents during calls and expect the agent to remember and improve for next time; language itself changes every 6 months to 1 year | Evidence: Kim asks: 'Is your system ready for users to use your voice agent better the next time?' and 'Is your voice agent ready for language change in one year or six months even?' | Implication: Ken's voice systems must support cross-session user modeling (user learns the bot's quirks, bot should learn the user's) and model/vocabulary refresh cycles under 6 months to avoid drift
Detailed Brief
Listening Channel Components and Interdependence
- Claims: Listening channel layer 1 (sounds): does the bot recognize speech well (ASR quality); Listening channel layer 2 (words): does the bot understand the user's words (NLU); Listening channel layer 3 (interaction): does the bot wait until the right timing for its turn (turn-detection, latency); Listening channel layer 4 (mental model): does the bot understand the user's intention (context, emotion)
- Evidence: Kim's call failure traced to ASR confusion (M vs N), turn-detection cutting her off mid-account-number, and no interactive clarification (mental model layer); She emphasizes that improving sounds requires accounting for words, and improving sounds/words requires good turn-taking, and all three require mental model alignment for task completion
- Implications: Ken's listening pipeline must not optimize ASR in isolation; turn-detection and context retention are equally critical for avoiding user escalation
Speaking Channel Components and Interdependence
- Claims: Speaking channel layer 1 (sounds): does the bot pronounce correctly (TTS quality, accent handling); Speaking channel layer 2 (words): does the bot choose words the user can understand (vocabulary curation); Speaking channel layer 3 (interaction): does the bot speak with the right timing (turn-taking rules); Speaking channel layer 4 (mental model): does the bot speak with the information the user actually needs (relevance, context)
- Evidence: Kim's name mispronounced by TTS applying English-centric rules (M-I-D-A-M → Midam), causing confusion but not immediate escalation until compounded by later failures; Vocabulary mismatch: bot asks for 'account number' without clarifying which field, causing user confusion and slower response (which then triggers turn-detection cutoff)
- Implications: Ken's TTS must support non-English phonetics and user-provided pronunciation corrections; vocabulary must be shared/curated with the user upfront or clarified interactively
Implementation and Diagnostic Approach
- Claims: Kim offers no single fix; the fix is in the system's orchestration and adherence to the linguistic framework; Recommended actions: choose good ASR/TTS models and configurations, do post-processing (ASR) and pre-processing (TTS), curate shared vocabulary, tune turn-detection/latency, implement emotion detection/context retention dynamically; EvaBench is ServiceNow's multimodal benchmark for diagnosing voice agent status
- Evidence: Kim lists model selection, configuration, vocabulary curation, turn-detection tuning, emotion/context handling as concrete steps; She states the framework is for diagnosis and building, not a plug-and-play solution
- Caveats: EvaBench specifics (open-source, pricing, metrics) not disclosed; unclear if it covers all four layers or is ASR/TTS focused
- Implications: Ken should use the framework as a design and post-mortem checklist, instrumenting failure modes by layer (ASR error rate, turn-detection false positives, context loss events, emotion misclassification) to prioritize orchestration improvements
Notable Concepts & Terms
- Joint activity: Human communication as continuous mutual mental model updating; voice AI must replicate this, not just process inputs sequentially
- Silently tracked timeline: In voice (unlike chat), sounds/words vanish after utterance; only the user's mental model persists, so UX targets mental model continuity over time
- Four-layer linguistic framework: Sounds (ASR/TTS), Words (NLU/vocabulary), Interaction (turn-taking/latency), Mental Model (context/emotion/intent)—all interdependent and must align
- Interactive clarification: Human conversational repair strategy (e.g., 'did you mean X or Y?') that bots often skip, causing user frustration when they instead demand repetition
- EvaBench: ServiceNow's multimodal benchmark for diagnosing voice agent quality across the linguistic framework
- Context retention (dynamic): Bot must track context about sounds, words, interaction, and mental model dynamically throughout the call and across calls, adapting to user learning and language change
- Catastrophic failure modes: User frustration, task failure, live agent escalation, abandoned calls, silent failures (user gives up without escalating)—all traceable to linguistic layer breakdowns
Operator Notes / Why Ken Should Care
- Instrument voice agent failures by linguistic layer (ASR error rate, turn-detection false positives, missing clarifications, context loss events) to prioritize orchestration fixes over model swaps
- Evaluate EvaBench for voice diagnostic tooling; consider hiring a linguist or speech researcher for voice system design/audit if building production voice agents
- When scoping voice agent vendors, audit all eight components (4 layers × 2 channels) for interdependence and alignment, not just ASR/TTS accuracy in isolation
- Design cross-session user modeling: track user adaptations (learned quirks) and bot adaptations (user pronunciation, vocabulary) to reduce repeat escalations
- Plan model/vocabulary refresh cycles under 6 months to handle language drift and accent evolution
- For customer support voice agents, add interactive clarification prompts (e.g., 'did you mean field X or Y?') instead of demanding repetition after confusion
- Map business metrics (escalation rate, abandonment rate, silent failure rate) to specific linguistic layer failures in post-mortems to quantify ROI of orchestration improvements
Source/Metadata
- Title: My name is... my name is...: A Linguistic Map for Voice Agents — Midam Kim, ServiceNow
- Transcript words: 2852
- Duration seconds: 904
- Timestamp note: Timestamps unavailable in transcript
Transcript
Okay, hello everyone. So my name is Midam Kim. I am an ML engineer from ServiceNow, and I'll be talking about a linguistic framework for VoiceAI. So, quick background on me, so you know where I'm coming from. I'm an ML engineer at ServiceNow, but I'm also a researcher, lifelong researcher, of speech communication in the wild. So my motto is doing linguistics, and what I'm going to be doing today is to hand you that lens of linguistics. Have you experienced VoiceAI failures? Yeah, every one. So I'm going to introduce an example that I experienced myself. The bot asked me, could you please spell your first name? And then I slowly start to spell my name. Yes, it is M-I-D-A-M. And the bot says, confirming with you, is it M-I-D-A-N? And then I say, no, it is M-I-D-A-M. And the bot says, thank you for your correction. Happy to help you today, Midam. And then I get slightly annoyed, more annoyed, because my name is Midam, not Midam. And then it asks me, now, what is your account number? And then I start getting confused. What is that account number thing? And then I try to find information about that. So which one? It must be, and I slowly start spelling the account number. So it is A-X-4-5-1. And then I take time, because I'm not used to reading this strange number. And then the bot cuts me off. And then it says, I couldn't find your record. And then, without even trying, it asks me to repeat that again. Could you please repeat that? And then I get super annoyed, and then I say, can I talk to a person? I just don't want to deal with you anymore. So this is a very typical pattern of voice AI, unfortunately, at this point. So I just want to navigate how we can solve this problem with linguistics. Voice AI is booming, but users are still often preferring human agents over voice agents. How can we mitigate this issue? And then, in the first place, what are the actual problems? So I think we can think about a fundamental framework to understand this end-to-end architecture of voice AI, which is called linguistics. So as all of us already know, human communication is a joint activity like the thing that we're doing right now. So I give you my sounds and words, you hear them, and then if it is a conversation, you're going to give me your sounds and your words, and then this is going back and forth through interaction. And then in this process, we're continuously processing and updating our mental models. So that's a joint activity for human communication, and I would like to say in voice AI human communication, it also has to be a joint activity like this, because that's the only thing that we know about communication as a human being. We have been evolving thousands of years as communicators, and this is what we know. So we expect the same thing from bots. So let me go over the failure scene of my call with the voice agent in this framework. So you see there's listen and speak for each party. So I start spelling my first name, and then the bot did not hear the difference between M and N correctly, so it's an ASR failure in the listening level. And then the TTS applies only English-centric reading rules to my name. M-I-D-A-M would read as M-I-D-A-M in the American English version. So I'm confused, but at this time I'm kind of generous because that happens a lot, even with human beings. So I'm okay. But then when it brought up account number thing, because I don't know what that is, I'm confused again. But I'm adaptive. I can find it. So I found the number, start reading it, but the ASR did not recognize the word unit correctly, so it cuts me off, and finally it eventually talked over me. So I get really irritated. And then when it asks me for the repetition of the same information, and then it is clear that the bot is not tracking the mental model with me. And then very rudely, it does not even try interactive clarification, which is a common strategy by human beings. So I don't want to deal with this anymore. So I say, can I talk to a person? So let's go over the framework again. So these are the linguistic components that are expected and well maintained in human-to-human conversation. So there are listening channels and speaking channels, and there are different components like sounds, words, interaction, and mental model. So the first component is does a bot recognize the user's speech well? And all of these technical terms will fall under this. And then there's going to be this second component, which is words in the listening channel. So does a bot understand the user's words? And then the third one is does a bot wait until the right timing for its turn? It's going to be about listening channel interaction. And then the last part is mental model. So does a bot understand the user's intention in the listening part? And then we can also go to the speaking channel. So it's going to be about pronunciation for the sound. And also does a bot choose the words the user can understand? And in the interaction part, does a bot speak with the right timing? And lastly, does a bot speak with the information the user actually needs? So there are a lot of engineering or linguistic or cognitive science terms that are here. You can now see that all of those have their right spots in this linguistic framework. And importantly, these components are interdependent, not separate or independent from each other. They're interdependent and they're aligned. So when you want to do good things about sounds, you have to think about words level. And then when you want to do good things about these sounds and words, you also have to account for interaction, so turn taking or turn detection. And then finally, you want to have good task completion, which is the goal of these mental model layer. Then you have to have all of these. Without any of those components, your voice agent will fail. And then finally, it has to be well aligned. All of these have to be well aligned. And additionally, you have to keep in mind that this is happening on the timeline. What I mean by that is it is silently tracked. Unlike in chat, in chat you see the history of what was said as text. But in voice agent experience, you say something and the bot says something, you go back and forth, and then see all these waveforms, the air via the vibration in the air, they're all gone. And only the user's mental model is a thing that's left and that matters. So sounds, words, interactions vanish the moment they're spoken, but the mental model perceives and grows over the timeline. So this is what you have to target for user satisfaction. And then what can we do for the bot to meet the standard of the user? So what we can do would include choosing good ASR models or configurations and doing some post-processing. Choosing good TTS models, configurations and pre-processing. And carefully curate the vocabulary that can be shared between the bot and the user. And do a good job of turn detection, latency and turn-taking. And very importantly, we have to, it would be great if we can do good emotion detection and handling and context retention. And by context, what I mean is context about all of these. And importantly, it has to be dynamic because things are always changing throughout the course of the call. So we would have to do this management dynamically along the timeline for different kinds of people. So kids or different kinds of people like these will have different expectations that we have to satisfy. Not just when they're happy, but also when they're not happy. So only then you can pursue a dynamic and truly scalable orchestration of voice AI. So it's a very difficult job to do. We always say that voice is the most natural way of communication, but it is actually not easy. Behind the scene, it is thanks to this linguistic orchestration. When your bot is not good at it, it's a catastrophic failure. So paying attention to this linguistic framework would have lots of business implications, because then you can decrease all of these user frustration, task failures, live agent escalation, abandoned calls, or silent failures. So in ServiceNow, we have made a good benchmark, a multimodal benchmark called EvaBench. So you can try that to diagnose your voice agent status. Key takeaways. Voice AI is a joint activity between the bot and the user, not just the pipeline. And we must serve users' needs on multiple layers, real time. So it's not that I have given you a fix today, because there's nothing like that. It's just, the fix is in you and your system. But what I have given you today is the linguistic framework you can try to diagnose your system and to build your system upon. You can try Eva, but also you can learn linguistics and hire linguists. Another thing I want to remind you of is that business implications are linguistic implications and vice versa in this voice AI scene, because voice is fundamentally a linguistic and very human and cognitive experience. I would like to ask you a longer term question. Speakers adapt. So I'm pretty sure that in this talk with you guys today, you have learned something about me, about my speaking style, what kind of accents I speak, what kind of words I'm using. So next time I see you guys in person, you would find it more comfortable to talk to me, because you have paid attention to me. Right? So speakers are always adapting. So the user will be adapting to your voice agent throughout the call. So is your system ready for them to use you better, use your voice agent better the next time? And languages always change. So is your voice agent ready for language change in one year or six months even? So thank you. it does not even try interactive clarification, which is a common strategy by human beings. So I don't want to deal with this anymore. So I say, can I talk to a person? So let's go over the framework again. So these are the linguistic components that are expected and well maintained in human-to-human conversation. So there are listening channels, listening channel and speaking channel, and there are different components like sounds, words, interaction, and mental model. So the first component is does a bot recognize the user's speech well? And all of these technical terms will fall under this. And then there's going to be this second component, which is words in the listening channel. So does a bot understand the user's words? And then the third one is does a bot wait until the right timing for its turn? It's about, it's going to be about listening channel interaction. And then the last part is mental model. So does a bot understand the user's intention in the listening part? And then we can also go to the speaking channel. So it's going to be about pronunciation for the sound. And also does a bot understand the words users are, does a bot choose the words the user can understand? And in the interaction part, does a bot speak with the right timing? And lastly, does a bot speak with the information the user actually needs? So there are a lot of engineering or linguistic or cognitive science terms that are here. You can now see that all of those have their right spots in this linguistic framework. And importantly, these components are interdependent, not separate or independent from each other. They're interdependent and they're aligned. So when you want to do good things about sounds, you have to think about words level. And then when you want to do good things about these sounds and words, you also have to account for interaction. So turn taking or turn detection. And then finally, you want to have good task completion, which is the goal of these mental model layer. Then you have to have all of these. Without all of those, without any of those, any of those components, your voice agent will fail. And then finally, it has to be well aligned. All of these have to be well aligned. And additionally, you have to keep in mind that this is happening on the timeline. What I mean by that is it is silently tracked. Unlike in chat, in chat you see the history of what was said that as text. But in voice agent experience, you say something and the bot says something, you go back and forth, and then see all these waveforms, the air via the vibration in the air, they're all gone. And only that user's mental model is a thing that's left and that matters. So sounds, words, interactions vanish the moment they're spoken, but the mental model perceives and grows over the timeline. So this is what you have to target for user satisfaction. And then what can we do for the bot to meet the standard of the user? So what we can do would include, of course, choosing good ASR models or configurations and do some post-processing. Choosing good TTS models, configurations and pre-processing. And carefully curate the vocabulary that can be shared between the bot and the user. And do a good job of turn-to-turn detection, latency and turn-taking. And very importantly, we have to, it would be great if we can do good emotion detection and handling and context retention. And by context, what I mean is context about all of these. And importantly, it has to be dynamic because things are always changing throughout over the course of the call. So we would have to do this management dynamically along the timeline for different kinds of people. So kids or different kinds of people like these will have different expectations that we have to satisfy. Not just when they're happy, but also when they're not happy. So only then you can pursue a dynamic and truly scalable orchestration of voice AI. So it's a very difficult job to do. We always say that voice is the most natural way of communication, but it is actually not easy. Behind the scene, it is thanks to this linguistic orchestration. When your bot is not good at it, it's a catastrophic failure. So paying attention to this linguistic framework would have lots of business implications, because then you can decrease all of these user frustration, task failures, live agent escalation, or abandoned calls, or silent failures. So in ServiceNow, we have made a good benchmark, a 2N benchmark called EvaBench. So you can try that to diagnose your voice agent status. Key takeaways. So voice AI is a joint activity between the bot and the user, not just the pipeline. And we must serve users' needs on multiple layers, real time. So it's not that I have given you a fix today, because there's nothing like that. It's just, the fix is in you and your system. But what I have given you is, today is, the linguistic framework you can try to diagnose your system and to build your system upon. You can try Eva, but also you can learn linguistics and hire linguists. Another thing I want to remind you of is that business implications are linguistic implications and vice versa in this voice AI scene, because voice is fundamentally a linguistic and very human and cognitive experience. I would like to ask you a longer term question. Speaker's adapt. So I'm pretty sure that in this talk, in my talk with you guys today, you have learned something about me, about my speaking style, what kind of accents I speak, what kind of words I'm using. So next time I see you guys in person, you would find it more comfortable to talk to me, because you have paid attention to me. Right? So speakers are always adapting. So the user will be adapting to your voice agent throughout the call. So is your system ready for them to use you better, use your voice agent better the next time? And languages always change. So is your voice agent ready for language change change in one year or six months even? So thank you.