My name is Bo.
I'm going to be here presenting real-time voice agents with Frontier Intelligence. I'm going to be talking a little bit about how we at Elise AI architected our voice agent harness to get real-time voice with the Frontier level of intelligence that we need. So before I start, I think I wanted to draw some parallels about why we decided to go with Cascaded Voice Agents, and especially comparing that to self-driving cars, which I was working in before.
So, to me, Cascaded Voice Agents makes sense when you view it through the lens of breaking it down into perception, which for self-driving cars is the bounding boxes, the camera, the LIDAR. For voice, it's going to be the transcription, effectively turning these signals from the real world into elements of data that the language model or whatever brain you're working on can process. Second is the planning stack, which is straightforward. This is where the language model will take in the outputs from the perception stage and produce the outputs that you want to produce back out into the real world.
And finally, there's the controls layer where in self-driving, you would be taking the trajectory that the planner would output and turn it into the real controls to drive the car. Here, we're turning the text into audio that we use to express our voice agent's thoughts. So here I'll be diving into each one of these elements. We've made a few interesting tricks on each of these areas to improve the speed of our voice agents without sacrificing the intelligence.
So the first one is going to be the transcriber layer. We came up with this concept called the streaming speculative transcriber, where effectively we are layering a fast streaming transcriber, like Flux, on top of or below a Scribe V2 or an accurate batch transcription, which takes in more context. It's a little bit slower, but it will give you more accurate detections. So we're going to walk through a setting now. In this case, the agent just asked, "Can you write your name and date of birth?" and the user is going to say this, and we'll see how that plays out timing-wise.
First we're going to get the short detection from the streaming layer. The accurate layer, the corrective layer is not going to fire because it's the same text. We're going to get some more streaming text detections. And in this case, the corrective layer is actually cancelled because we got new text. So more context, more audio is going to beat the old accurate one. And here's where the first correction comes in. Because the Scribe V2 layer understands the context of the question, it's able to understand that this is talking about a name and this is a date of birth. And then a couple more detections. These are just punctuation. We don't care.
And so in the end, we release this text over to the agent. Moving on to the language model. So here, since we're using these slow but intelligent LLMs, we really want to reduce the number of round trips. And the thing that causes us to do a lot of inferences is tool calling. So one way to get rid of that is by having background agents do the tool calling for you and push the tools back into the context of the main agent so that it thinks it made the tool call, but it really didn't.
So we remember from detections from before. What will happen is each one of these detections is going to trigger an early generation of the agent. But we won't actually emit this out until we're confirming that the user has finished speaking. So in this case, the user says "sure." The agent knows that the user is about to say something else. Our background tool calling here, which is going to be helping us figure out the name and the date of birth from the user detection, is not firing. So nothing much there. The next instant detection comes in. It says that, you know, still not really a name. Our agent plays along and continues there.
Now more context comes back. The agent feels like there should be a name. It's going to ask to spell it out because it's probably thinking there's some transcription error here. Still no name or date of birth. And then finally this—remember this is our corrected final instant detection from the transcriber from the Scribe V2. Here, our eager agent generation that was made without any tool calls is going to get cancelled because the background agent finally is able to find the name and date of birth it's looking for. So it's going to re-trigger and now the agent actually has the context it needs.
And you see here we're doing it. The tool call here has some intelligence there. We're correcting mistranscriptions of name, doing some phonetic matching here. And then, once we've understood that this is the end of the user utterance, we'll emit it out. So pretty standard.
Okay, and then the next layer here is going to be text-to-speech. So with text-to-speech, the goal is to take what the agent said and the agent is going to be emitting this in a streaming fashion. So we're going to need to produce audio as quickly as possible. And ideally what you can do is before the agent has even finished generating the full text, you can have the audio play. So it's hiding the latency of finishing the generation. So I'm going to play the streaming agent output now. It starts with "you." And actually before I dive further, there's this new concept that we're introducing here called the prefix cache.
So the prefix cache is going to be looking at the agent stream and seeing if we already have generated audio for that sequence of words from a prior generation or maybe the same generation in this call as well. So it sees the word "you." For this prefix cache, we don't want to immediately hit on every single word. We're going to be waiting for a little bit more words. So after three words, the prefix cache gets our first hit.
And over here on the right, this is our text-to-speech standard provider. Cartesia is a text-to-speech engine with WebSocket support. So we're piping the agent through the cache and also piping it through WebSocket. More tokens come in, more cache, more sending through WebSocket. Not much to say here. And okay, so now we get our first unique thing, which is we found a token that actually causes a cache miss. It makes sense. If we're caching previous generations, "you said your name is" is a pretty common thing. But once we add in the name, suddenly that's going to result in a cache miss.
At this point, we're actually going to yield out our cached audio. So "you said your name is" is going to be emitted as the rest of the streaming text is coming back. So at this point, the user hears the agent. The user doesn't really know what's going on. It just looks like really fast response times to them. And now the remaining text flows through. And at this point, we've already emitted from the cache. The cache has done its job. The rest we can throw into Cartesia.
And here's the trick where Cartesia has seen the entire transcript up to this point. To Cartesia, it doesn't know about the existence of this prefix cache. It's just going to generate this full sentence with standard natural prosody. But what we do is when the generation comes back, since we've already played the audio here, we can actually suppress the audio from Cartesia here and just play out the remaining stuff.
So the user, there might be a tiny bit of a hiccup. I'll play some audio later. And you'll know that you probably won't be able to notice. And effectively, we're taking this part and just emitting it directly after this audio has finished playing. So it looks seamless to the user. Yeah, so at this point, this is a lot of technical details. I think the best way is just to play a call and see this in action. So here I recorded a call and let's just hear what happens. [CALL RECORDING] This is Bo, OBGYN, Elise speaking. Hi, I think I might be pregnant and I want to schedule an ultrasound to confirm. That's exciting news. Can you provide your name and date of birth?
Sure, it's Elise's trial, date of birth 2303. Thank you. Are you a new patient with us? Yes. Welcome to Bo, OBGYN. Can I text you a link to upload your insurance information? Sure. Okay, sent. Let me know when you've received it. I've received your insurance information. The earliest appointment I have is Thursday, July 2nd at 10 a.m. Actually, give me a second. Let me check my calendar. Sure. Take your time. Do you have something for next week? For next week, I can offer an ultrasound with Dr. Avery Stone at North Clinic on Tuesday, July 7th at 2 p.m. or 3 p.m. Do either of those work for you? Yeah, 2 p.m. works.
Great. Your appointment has been booked. We look forward to seeing you then. Thanks. Bye-bye. [END CALL RECORDING] All right. Yeah, that's pretty much it. Yeah, you can see all this streaming and a lot of things happening in the background. And this is what really makes voice agents interesting. And there's a lot of effort that can be done in the harness to really get a natural conversation, which is what we're after.
Okay, so briefly, in the last part, I want to talk a little bit about Elise. So Elise, our headquarters are in New York and we're trying to expand our presence here in the Bay Area. We're a different style of company than what people think of when they think about AI startups in San Francisco, where we're actually very focused on helping people where they need it, in life's most critical areas. We work on housing, healthcare, and we're doing really well. And here's a link to join our team, and we're going to be posting a lot on Twitter, so you can follow us at Elise.ai as well. Yeah, that's it. Thank you.
So what will happen is each one of these detections is going to trigger an early kind of generation of the agent. And we, but we won't actually admit this out until we're confirming that the user has finished speaking. So in this case, the user says sure. The agent kind of knows that the user is about to say something else. Our background tool calling here, which is going to be helping us figure out the name and the date of birth from the user detection, is not firing. So nothing much there. The next instant detection comes in. It says that, you know, still not really a name. Our agent kind of plays along and continues there. Now kind of more context comes back.
The agent kind of feels like there should be a name. It's going to ask to spell it out because it's probably thinking there's some transcription error here. Still no name or date of birth. And then finally this, you remember this is kind of our corrected final instant detection from the transcriber from the Scribev2. Here, our eager kind of agent generation that was made without any tool calls is going to get cancelled because the background agent finally is able to find the name and date of birth it's looking for. So it's going to re-trigger and now the agent actually has the context it needs. And you see here it's kind of we're doing it.
The tool call here is a little bit some intelligence there. We're going to be, you know, correcting mistranscriptions of name, kind of doing some like phonetic matching here. And yeah, and then we'll kind of, once we've understood that this is the end of the user utterance, we'll kind of emit it out. So pretty standard. Okay, and then the next layer here is going to be text-to-speech. So with text-to-speech, the goal is to kind of take what the agent said and the agent's going to be admitting this in a streaming fashion. So we're going to need to produce audio as quickly as possible.
And ideally what you can do is before the agent has even finished generating the full text, you can have the audio play. So it's kind of hiding the latency of finishing the generation. So I'm going to kind of play the streaming agent output now. So it starts with U. And actually before I dive further, there's this new concept that we're introducing here called the prefix cache. So the prefix cache is going to be looking at the agent stream and seeing if we already have generated audio for that sequence of words. From like a prior generation or maybe like the same generation in this call as well. So it sees the word you.
For this prefix cache, we're going to be, you know, we don't want to like immediately hit on every single word. We're going to be waiting for a little bit more words. So after three words, the prefix cache gets our first hit. And over here on the right, this is kind of our text-to-speech standard provider. You know, Cartesia is a text-to-speech engine with WebSocket support. So we're piping the agent through the cache and also piping it through WebSocket.
More tokens come in, more cache, more sending through WebSocket. Not much to say here. And okay, so now we get our first unique thing, which is we found a token that actually causes a cache miss. It makes sense. If we're kind of caching previous generations, you said your name is is a pretty common thing. But once we add in the name, suddenly that's going to result in a cache miss. At this point, we're actually going to yield out our cached audio. So you said your name is is going to be admitted as the rest of the streaming text is coming back. So at this point, the user hears the agent. The user doesn't really know what's going on.
It just looks like really fast response times to them. And now the kind of remaining text flows through. And at this point, we've already admitted from the cache. The cache has done its job. The rest we can kind of throw into Cartesia. And here's kind of the trick where Cartesia has seen the entire transcript up to this point. To Cartesia, it doesn't know about the existence of this prefix cache. It's just going to generate this full sentence with standard natural prosody. But what we do is when the generation comes back, since we've already played the audio here, we can actually suppress the audio from Cartesia here and just play out the remaining stuff.
So the user, there might be a tiny bit of a hiccup. You know, I'll play some audio later. And you'll know that you probably won't be able to notice. And effectively, we're kind of taking this part and just emitting it directly after this audio has finished playing. So it looks seamless to the user. And yeah, so I guess at this point, kind of this is a lot of technical details. I think the best way is just kind of play a call and see this in action. So here I recorded a call and let's just hear what happens. This is Bo, OBGYN, Elise speaking. Hi, I think I might be pregnant and I want to schedule an ultrasound to confirm. That's exciting news.
Can you provide your name and date of birth? Sure, it's Elise's trial, date of birth 2303.
Thank you. Are you a new patient with us? Yes. Welcome to Bo, OBGYN. Can I text you a link to upload your insurance information? Sure. Okay, sent. Let me know when you've received it.
I've received your insurance information. The earliest appointment I have is Thursday, July 2nd at 10 a.m. Actually, give me a second. Let me check my calendar.
Sure. Take your time. Do you have something for next week? For next week, I can offer an ultrasound with Dr. Avery Stone at North Clinic on Tuesday, July 7th at 2 p.m. or 3 p.m. Do either of those work for you? Yeah, 2 p.m. works. Great. Your appointment has been booked. We look forward to seeing you then. Thanks. Bye-bye. All right. Yeah, that's pretty much it. Yeah, you can kind of see all this streaming and a lot of things happening in the background. And yeah, this is what really makes voice agents interesting. And there's a lot of effort that can be done in the harness to really kind of get a natural conversation, which is what we're after.
Okay, yeah, so I guess briefly, you know, in the last part, I want to just talk a little bit about Elise. So I think Elise, you know, our headquarters are in New York and kind of we're trying to expand our presence here in the Bay Area. We've, I think it's maybe like a different style of company that I think people are, like, think of when I think about AI startups in San Francisco, where we're actually very focused on just, like, helping people and helping people where they need it, like kind of the life's most critical areas. We work on housing, healthcare, and we're doing really well.
And, you know, here's a link here to kind of join our team, and there's going to, we're going to be posting a lot on Twitter, so you can follow us at Elise.ai as well. Yeah, that's it. Thank you.