Suman Yu Reviewer Reviewer My name is Suman Yu, and I'm the founder and CEO of Hamming. Before working on voice agent reliability and safety, I worked at a company called Citizen out of New York. Anybody here use Citizen app? Awesome. Thank you. At Citizen, we listened to crime, thousands of hours of police radio station data, and sent millions of alerts to users in San Francisco, New York, LA, Chicago, Baltimore, and so on. Some obviously gory and pretty sad, but others more funny, like a person stealing bags of ice cream from Safeway. A report of a man hanging off the side of the house after a woman stole his ladder.
If I actually take a look at the Citizen app right now, for those who are customers or users, I can see that there is a man yelling at person. There is indecent exposure. This is real. This is real time. This is a couple hours ago. These are real time alerts that we're sending.
Now, voice agents scare me more because they're finally graduating from demos and POCs to production. We should be super excited, but I'm nervous. I'm personally nervous. They're talking to users at a scale that would make Gary Tan and Paul Graham proud.
When I got started in voice agent reliability in early 2024, voice was just starting to work. It was not quite good yet, but it was just starting to work. You would have to pay me a lot of money for me to stop using Aqua Voice, Super Whisper, Whisper Flow, and so on. These products are just getting super good. And a big reason is because the underlying infrastructure is getting better and the orchestration layer is getting meaningfully better. It's getting much faster to build products and voice experiences that maybe are 60% good in a pretty short period of time, but the long tail is still wide away. I think speech-to-speech models are getting better.
Teams are experimenting with hybrid architectures of combining voice-to-voice modalities and cascading stacks to make the experience reliable but still pretty low latency. Things are obviously getting better. Agents are being connected to calendars, CRMs, EHRs, reservation systems, and so on. Voice agents can now take actions. However, reliability is still the number one problem holding back most voice agent deployments at scale. This is still the number one problem. This is an example I found on Twitter two weeks ago. A person is trying to get information for a trade-in and gets absolutely confused with the information that they're receiving.
Alex now has to correct for this loss of trust by trying to call the person and see what happened and fix the situation. Let me see if audio works here. I'm screwed up with another customer. We're getting it fixed, but I got to call them and see if I can work it out. I'm like, dude, half the time I'm like, I don't know if I'm talking to AI. I don't know if I'm talking to a person. It was just confusing, but we got there. It probably is AI and human. So I think voices sound very confident.
They sound very natural, but the information provided is often not correct. That's the biggest problem here. This example is more personal. I had booked an appointment with a physician a couple weeks ago, or I thought I did. I showed up to the appointment and turns out I was not actually on the schedule. So the front desk made it turn me away. I wasted two hours. For me, this was a waste of time. But what if this was actually your parent? What if this was your grandparent? What if this appointment was for a procedure instead of a regular checkup? The costs for these different permutations of the same failure mode can actually be super high.
Now let's compare crime to voice agents. I think observation number one is crime is actually decreasing over time. This is a good thing. And I hope it crosses the x-axis at some point in the future. Voice, on the other hand, is generally taking off, right? We're seeing a pretty fast takeoff of voice agents being deployed in production. There's at least a trillion calls that are done every single year. And majority of these will be done by conversational voice agents over the next five years. If you assume a one-person error rate, that is still 10 billion incidents per year. That's a lot.
In practice, we currently monitor 10,000 agents, and the error rate is closer to 10% in practice. These range from agents saying they found the right policy when they actually skipped the eligibility or verification steps, or applying discounts when they were not really supposed to, mishearing what the person said, providing incorrect information, or claiming they booked an appointment when they actually did not, just like it happened for me. Now, not every single call has an equally bad cost. Some range in the crime land, some range from trash fires, which are funny, annoying, not really hurting somebody. For a voice equivalent, that would be annoyances like repetition,
or not quite understanding what the user is saying. All the way to safety risks like mass shootings, or in the voice agent equivalent, it would be a drive-through that's deploying voice agents at scale, like a Taco Bell or McDonald's, and a person orders a vegan burger with peanut allergies. If one of those two situations are not handled correctly, that is definitely a safety concern at scale. The other big difference between crime and voice agent deployments is crime generally tends to be pretty hyperlocal. It tends to be very decentralized, right? Things like robbery, or motor vehicle theft, or larceny.
They're impacting a finite set of individuals that are involved in that situation. On the other hand, voice agents are much more centralized.
A single prompt change, or an architecture change, can have pretty massive implications downstream for all of the millions of users that are in the crossfire. So the blast radius is quite massive. So the natural question is, how do you make these incidents much more visible and obvious? That's the obvious question here. I'll borrow a framework from a couple of my friends who were OG growth folks at Facebook. Step one is to identify, okay, what are all the challenges and problems that exist in your conversation experience? Step two is to prioritize impact size. There's a frequency and severity analysis that's pretty important.
Step three is to understand, okay, how do we actually fix this? Step four, execute. Step five, okay, did my change actually work? And did it cause any regressions somewhere else? And lastly, we continue to monitor in production. On the y-axis, I think it's important to highlight there are known problems that already exist. Things like turnover latency, interruptions, maybe some ASR problems you're aware of. And these are known problems that exist that the team should track over time. On the other axis is actually emerging behavior or patterns that are only obvious across lots of conversations. On the x-axis, you have coverage, just like insurance.
Are you analyzing few conversations? Are you analyzing many conversations? Most teams will typically start by listening to calls manually. And I think that's the best place to start. I don't think you should skip that step. There's a lot of depth and insights you get by actually listening to specific conversations and building that texture that comes from that intuition. However, it's obviously not scalable. So most teams end up having a spreadsheet of five or ten different rubrics around greetings, closing, validation, core logic, and so on. To scale that up even further, you then end up investing in some evals product, right?
You might run some LMs as a judge and compute classic metrics and also more deterministic and stochastic scoring logic. But there, you're still stuck with checking for consistency of known problems, but you're not really discovering novel insights that are actually happening across conversations. We're spending a ton of time on performing cross-conversation analysis, not a pattern on a single call, but across conversations. And some of the best teams that we work with are doing the same. Now, to prioritize impact size, I think there's problems that are one-off that are low impact. I mean, who cares?
Even low impact and systematic problems in the crime world, that would be a trash fire. In a voice agent world, it could be some repetitions the team is experiencing. They're still annoying at scale, and if you are doing a bake-off, it's still worth solving for them. I would not ignore these kinds of problems. One-off and high impact? Well, hope it is a big chronic. And I think systematic and high impact are obviously the P0 target areas for the team to solve. An example of that would be in a FinServe capacity, there's a voice agent that helps users freeze their credit cards. And if it doesn't do that, well, that's a massive fail. All right, so understand and execute.
I'm pretty sure everyone's doing this. Please fix my agent. I think fixing, or rather attempting to make a fix, is the simplest and the lowest effort component of this debugging pipeline and loop. The next step is, all right, I made a change to my system. How do I actually know this thing works for real? A great way that's naive is to take a real call. For example, in my case, I booked an appointment and it didn't get scheduled, and replay that exact conversation and run that maybe 5, 10, 20, 50 times and see, okay, what is my probability of passing this type of issue?
A better way is to keep the same intent, but change the wordings, change the patterns, change the accents, change the style, add one more intent to the mix. And that gives teams much more coverage to feel confident that yes, I actually made a change, and my changes are net positive instead of net negative. There are certain fixes and hypothesis that are very difficult to test in a pre-deployment synthetic setting. And so A-B testing ends up being pretty critical for those circumstances. For example, if you have an outbound agent, the first five seconds of a conversation tends to be the most important.
And so the vocal quality and the specific words you end up using, they matter the most. And so A-B testing that is the only way in real life setting to get results. You can't really do it through simulations alone. And so there we have the loop. Identify, prioritize impact size, understand the fix, execute, check, make sure it didn't break anything, and then continue monitoring. So I think making voice agents useful is already hard as it is, even when dealing with earnest users on the other line, right?
These are people who just want their problem solved. They're not trying to mess with you. These are normal people. Now what happens when Mythos learns how to dial? So if we can extract trade secrets from the NSA, it can certainly seduce you into revealing PHI and PII data as well.
And I think both voice agents and humans will be targeted here.
Voice agents because there's a pressure to make these more capable, give them access to more data, give them access to more tools, deploy them quickly. The more the capability, the bigger the surface area.
This is common sense. And the more the voice agents become natural and human sounding, the more humans will be tricked along the way as well for those who are weaponizing. We shipped our red teaming product back in April just to test out this hypothesis for how many agents can we actually break from an adversarial capacity.
And we can probably break one in five agents at this point. We've tested this across financial services, healthcare, consumer, and so on. We've bypassed verification. We've definitely had agents, and we've been able to reject several agents and gotten data we should not have. So this is not theoretical. This is actually a real concern. I think the only real defense against the dark arts is step one to invest deeply in pre-deployment testing.
This could be text to text. This could be voice to voice. There's pros and cons to both. Have to chat offline if folks are interested. And this is just making sure you're not self-owning when you're talking to real people who just want to get their problem solved. Step two is to have a great monitoring system of all kinds. And I've highlighted different flavors of monitoring. Per-call scoring, manual evals, listening to conversations and cross-call analysis. And this is helpful both for monitoring what the agent is saying and behaving and how it's actually doing, but also the users. Are the users being adversarial? Are they being annoying?
Are they trying to trick the agent into doing things it's not supposed to be doing? And I think our new recommendation now is to run 24-7 red teaming for your agents, especially if you believe the cost of bad interactions can be rather large. So I think voice agents have this awesome potential of making the world feel much more human compared to interacting with clunky IVR trees or chatbots or worse, being stuck on a hold. And when we think about crime, we often think of crime happening to somebody else. Crime does not happen to you, typically.
With voice agents, especially bad actors, as these agents are deployed and as bad actors start to exploit a lot of the vulnerabilities, the number of incidents is about to go way up. And so the reason I fear voice agents more than crime is that one of these incidents is going to impact you. It already did for me. Awesome. So it's time for me to shill. We brought a lot of tokens. If you are interested in working in this space, please come and talk to us. And if you are deploying voice agents and want to validate whether your architecture or your evals are set up correctly, please come and talk to us. We'll be outside. And here's my number. Here's my WhatsApp.
Thanks, everyone. Thanks, everyone. Thank you. Thank you. Thank you.
Thank you. Thank you. Thank you. Thank you. Thank you. Thank you.