SPEAKER_00
All right, well folks, thanks for being here. Excited to chat more about how to engineer voice agents. High quality, low latency, at scale. First, maybe a little bit about me. My name is Rishabh. I work at a company called Together AI. I lead the voice AI team there. Prior to Together, I was the co-founder CEO of a company called Refuel that was acquired by Together last year, but generally been building AI and machine running infrastructure for about a decade. For folks who maybe don't know about Together, Together is building the AI native cloud.
SPEAKER_00
What that really means is for companies that are looking to train models and need access to reliable compute, or you want to do inference at scale, we're probably a very good fit for you. We work with a million plus developers. We closely work with hundreds of companies and we're very proud to be working with companies like Cursor and Decagon. Okay, here is the agenda for today. We're going to first talk about why are we talking about voice. Although, the previous speaker alluded to a lot of interesting things that he's doing with voice. But why does voice really matter? What does it actually take to build voice agents at scale? What are the challenges?
SPEAKER_00
We'll talk about this pipeline architecture, which is becoming the dominant way to build agents in production today. We'll deep dive a bit and look at all the components, we'll look at the system and trade-offs. And then finally, we'll chat about maybe what might be the next generation of building voice agents in the coming months and years. Okay, so starting with why voice matters. Well, there's billions of phone calls a year that are still handled by humans. I am pretty sure all of us have the experience of calling customer support, asking about the status of our order, looking to change the reservation.
SPEAKER_00
We've also probably had the experience of calling a doctor's office because we've got to book an appointment for ourself or a loved one. And pretty much everybody here probably has had the experience of being on hold. Now, it would be amazing for AI agents to be able to handle some of these calls for us. But really, one of the more exciting directions is, frankly, voice is just this brand new interface to interact with systems and computers. Look, humans, we learn how to talk before we learn how to read, right? So this comes very naturally to us. And obviously, we're seeing this with ChatGPT's advanced voice mode.
SPEAKER_00
And we're seeing this with folks who are starting to directly talk to Cursor, talk to Cloud Code, in order to get their work done. And this is, frankly, just the beginning. And one of the exciting pieces of 2026 is building these rich, high-quality conversations. This is not the domain of science fiction or research anymore. This is primarily an engineering problem today. Now, why is it hard? Well, there's a few things that you've got to solve first. First, building voice AI and building voice agents, this has to be real-time. When humans are having a conversation, we respond to each other's cues in something like 300 milliseconds.
SPEAKER_00
And so if you're talking to an AI and it's taking more than 500 milliseconds to respond, you'll start to notice. If it takes a second, if it takes two seconds, people will just hang up. So you've got to get latency down. The next thing that matters is you want it to be a reasonably smart call. You want to get the work, the job done. And so for real-world complex workflows, the instructions are complicated. There's a lot of ambiguity. You have to be good at tool calling because that's the way you give agents access to the real world. So you have a baseline level of intelligence that you've got to need.
SPEAKER_00
The third piece that you've got to solve for is the voice has to be natural enough. It has to sound pleasant enough. And this means a lot of different things. It means, can it talk to you in your own language with the right accent, potentially? Can it pronounce your name? Can it deliver the right emotion that is needed for a particular situation? A lot of things fall into that bucket. And finally, you could stitch together a nice demo with one person calling. But what happens when you're doing 100 calls, 1,000 calls, 10,000 calls concurrently? Reliability really starts to matter. And this is an and problem.
SPEAKER_00
You have to solve every single one of them at the same time, or you're going to be in a little bit of trouble. So, at least today, the dominant way of building these voice agents is this pipeline architecture, or this cascading architecture, which attempts to solve all of the problems that I outlined earlier. Now, there's a few boxes that are going on, but conceptually, it's relatively simple to understand, which is audio chunks from an end user that are being streamed in, potentially to an agent orchestrator, something like a pipecat or live kit or something that is homegrown.
SPEAKER_00
And then, essentially, this audio is being fed into a speech-to-text system that converts it into text. That's being fed into an LLM that then decides, do I do a tool call? What is the output? Produces text that is then fed into a text-to-speech model, which starts to produce audio chunks that are then streamed back to the end user. So, that's a rough architecture. Let's maybe look at each of the components one by one, the components that matter here. The first is speech-to-text, very much like the ears of your agent here. The performance metrics that matter here, the first is quality, word error rate.
SPEAKER_00
Now, depending on use case, the numbers might look different, but state-of-the-art models are typically in the 6% word error rate on open benchmarks. What that really means is the transcript that is produced by your model, comparing it with the reference transcript, 6% of words have an error in them. Now, you can imagine why this might matter, right? Because if you don't get the transcript right, you don't get somebody's name right, you don't get the name of, let's say, a driver. You don't get the rug, right?
SPEAKER_00
Essentially, there's no way to fix this. The performance metrics that matter here, the first is quality, word error rate. Now, depending on use case, the numbers might look different, but state-of-the-art models are typically in the 6% word error rate on open benchmarks. What that really means is the transcript that is produced by your model, comparing it with the reference transcript, the 6% of words have an error in them. Now, you can imagine why this might matter, right? Because if you don't get the transcript right, you don't get somebody's name right, you don't get the name of, let's say, a driver. You don't get the rug, right?
SPEAKER_00
Essentially, there's no way to fix this. Your LLM will carry forward the mistake, the TTS model will carry forward the mistake, so you have to get it right for the important keywords. And then the second metric that matters often, which is latency-driven, is time to complete a transcript. The way to understand this is when somebody completes an utterance, they stop speaking, how many milliseconds does it take for you to complete the transcript and have that be ready for the LLM? And so, as an example, for some of the models that we run on together, we get consistently P90 of 100 milliseconds, which is pretty fast.
SPEAKER_00
Aside from just raw performance, there's a few other capabilities that matter. Turn detection, very important, still somewhat unsolved problem. Frankly, that could be a 20-minute talk in itself. But really, the best way to understand this is you've got people who are talking, maybe they pause for a second. But do you actually know? Does that pause mean their turn has ended? Are they going to continue talking? Because really, the last thing that you want here is for the agent to start sending audio back and talking at this person, even though their turn hasn't ended. We don't enjoy this in human conversations, and we will certainly not enjoy this in AI conversations.
SPEAKER_00
Depending on who your customers are, language matters. And so being able to do this for a wide variety of language and getting it right there, it's important. And the final piece that I'll mention, this is somewhat new, is we're also starting to see architectures, model architectures that are streaming native. A little sidebar, we won't spend too much time on this, but there's an architectural evolution for speech-to-text models that is in progress, which is going from batch models to streaming models. Whisper is the canonical model, came out a few years ago. It was trained on 30-second audio clips. 30 seconds is way too much.
SPEAKER_00
You can't wait 30 seconds to start transcription. So people have had to build all sorts of complicated logic around models like Whisper to do chunking and to pad it with silences, and then make multiple calls, stitch that together to produce the final transcript in streaming mode. But recently, and this is a fairly new model from the NVIDIA team, instead you have the encoder of the model have two interesting characteristics. The first is it's trained with different amounts of look-ahead time. So it only looks at perhaps 80 milliseconds or maybe up to a second of audio instead of 30 seconds.
SPEAKER_00
And it's also able to cache these activations so that as you make small steps in audio frames, you're actually only doing the heavy computation once. So again, stepping out, but it's an interesting direction that we're seeing to be able to handle streaming conversations for these voice-agent use cases. Okay, so that's speech-to-text. Jumping into the next part of the pipeline, which is LLMs, the brains of your agent. The performance metric that matters here first and foremost is streaming latency. And so, TTFT is the metric here. And a rough metric is it's usually pretty good if you can get to 300 milliseconds of TTFT.
SPEAKER_00
Because you want to start producing tokens, start feeding that into the TTS model as fast as possible. That number, 200 to 300 milliseconds, has implications for what models you can use. And so a good size model typically ends up being in this 8 to 30 billion range. If you go any bigger, you'll burn through your latency budgets. If you go too small, that has implications for the intelligence of the model and frankly, the tool calling that is needed, which are both pretty critical if you want to build a voice agent that does meaningful stuff in the world. Okay, text-to-speech. This is the voice of your agent.
SPEAKER_00
There's a few interesting things on sort of performance and capabilities. Performance, again, the trend continues. What is the time to first audio, right? As you get a transcript, how long does it take to produce the first audio chunk that can start to be streamed back? And aside from TTFA, what does the real-time factor look like? Real-time factor is, and this is generally the case for most TTS models, but what it means is how much audio can you produce in a certain number of seconds of processing time? So if you can produce 10 seconds of audio in five seconds, your RTF is 0.5. And so you typically want that to be less than one so that you're not buffering.
SPEAKER_00
Quality is one of the hard ones with TTS because there are some objective measures, but frankly, nothing quite beats listening to audio samples for the models, for the voices that you care about, and getting a feel for whether this is the right experience that you want your end customers to have. Some of the other capabilities, it's naturalness across a number of different voices, being able to pronounce things exactly right, whether it's customer names, whether it's product names, being able to have some amount of control over emotions.
SPEAKER_00
And so you might see TTS models that allow you to add these different tags, which say this is happy or angry or sad, and it's a start, but these models are getting pretty good at emotional control. And of course, coverage over language continues to matter. Okay, so those are the main components. But to zoom out a little bit, these components are part of this larger architecture, which is multiple models being orchestrated. And so there's a few things that we should always keep in mind, which is first, there's a latency and cost budget across these models that we're thinking about.
SPEAKER_00
A rough rubric is the LLM is going to take up a majority of it, followed by TTS, followed by speech-to-text, both from a latency and a cost perspective. And so again, rough rules to think about. And one piece that we didn't mention, and we'll come back to this in a second, is so far a lot of the numbers that we're looking at is just engine latency. How much time does it take the model to produce an output? But actually, when you're calling models that might be sitting in different data centers, there's network latency as well. And that starts to have an impact. But again, we'll come to that in one more slide.
SPEAKER_00
A rough rubric is the LLM is going to take up a majority of it, followed by TTS, followed by Speciotex, both from a latency and a cost perspective. And so again, just rough rules to think about. And one piece that we didn't mention, and we'll come back to this in a second, is so far a lot of the numbers that we're looking at is just engine latency. How much time does it take the model to produce an output? But actually, when you're calling models that might be sitting in different data centers, there's network latency as well. And that starts to have an impact. But again, we'll come to that in one more slide.
SPEAKER_00
Autoscaling is also somewhat interesting and tricky to get right for agent systems. Of course, you want to be doing autoscaling to scale up. As demand goes up, scale down, potentially night times or weekends, you want to scale down seamlessly. Scaling up, you know, what we've typically seen is people are much more aggressive about scaling up because the last thing you want is requests to be slowed down or backed up. So you typically might autoscale earlier than you might do with somewhat more asynchronous systems. And scaling down is also tricky because you might actually have these stateful long lived connections to your models. And so you can't just arbitrarily kill a pod. You might want to wait for conversations to kind of fully finish. So some interesting nuances with autoscaling.
SPEAKER_00
And finally, global deployments are important because you want your models, you want your system to be as close to your end users to shave off latency as much as possible. And of course, if you're building models in Europe or in places where residency matters, you want to be making sure that you have the ability to deploy wherever you absolutely need.
SPEAKER_00
I know I reference co-location. But here's one way to kind of understand this problem. So the chart on the left hand side, this is a very optimized sort of system where you're doing a pretty good job with your speech to text and text to speech and LLM models, where the engine latency is in exactly the right ballpark. You're doing 100 to 200 milliseconds of time to first token or audio. But you might actually end up having your models being sufficiently far away from your agent orchestrator that it's taking 75 milliseconds of network latency. 75 milliseconds is really not that much. Even US West to Europe would certainly be 75 milliseconds. But depending on networking, it can be much higher as well. And so an interesting direction that we're seeing folks go is how can you co-locate all your models and potentially your agent orchestrator to either be in the same data center or be very, very close to each other? How can you get them literally in the same building? Because that drop from 75 milliseconds to five basically gets you a 30% reduction in already a fairly optimized voice agent setup. So some of these things, especially with real-time systems, it's pretty important to have fairly deep observability and every 10 milliseconds matters.
SPEAKER_00
Okay, so hopefully that's an interesting picture on this pipeline architecture. But that's not the only way people do it. One of the other directions that is becoming interesting is a pure speech-to-speech model. And so instead of having speech-to-text followed by LLM followed by text-to-speech where you're coordinating and orchestrating across a number of different models, it's way simpler if you could have a pure speech-to-speech model that still is responsible for function calling, still handles all the complicated instructions, but just a single model doing it. And of course, for folks who've played around with OpenAI's real-time API, they have a single model behind the scenes. NVIDIA recently launched a model called VoiceChat. Again, very similar ideas. The reason why most of these models are not used in production too much is because they still have trouble with instruction following and tool calling. So the real-world experience often looks like you'll try them, and then you spend a lot of time prompt engineering and hoping to fix issues and eventually move to a pipeline architecture. But as these models get better, which I'm confident they will, it has some pretty incredible benefits. Because suddenly, you don't lose anything about the nuances of speech when that speech is getting converted into text. So the model will natively understand what was the tone, what was the emotion, was the user hesitant. That stuff will still remain with the model to make the next decision. And this type of model also allows for more full duplex communication, which basically means that the model can start producing audio while it's still receiving audio. And this means that as a customer is talking to this model, you can back channel. You can say, I see or uh-huh, similar to what a human conversation would look like. And these models become much better at handling interruptions and barges, which again, with the pipeline architecture, you have to do a lot more complicated engineering work around. So hopefully this points in the direction of what might be coming in the future. But frankly, we have a lot of engineering work ahead for all of these voice interfaces that still have to be built.
SPEAKER_00
If you want to learn more about what Together does, here are a couple of links. We are hiring. We also have a booth G1 downstairs. Happy to chat more and happy to answer questions. Thank you. [SPEAKER_03] For voice-to-function calling use cases, what sort of evals do you use? And then what score do you need to get in those evals in order to have something that's good enough for production? [SPEAKER_03] That's the classic answer is it depends.
SPEAKER_00
But I think in terms of evals, there's of course component by component evals. So if we're talking about the pipeline architecture and assuming that speech-to-text is good, text-to-speech is good, and the only thing that we're caring about is function calling or tool calling evals in the LLM, then it's very similar to how one might do evals for tool calling for LLMs broadly, which is was the tool call correct? Was the output actually parsable? There's a bunch of those. One would imagine that you'd want the tool call structure to at least be very close to 100%. And then the correctness, again, it depends a little bit on what the use case eventually demands. One of the other things that we are seeing is especially to get around, to make models better at tool calling, and because we have to stay within that LLM budget of the models have to be relatively small.
SPEAKER_00
Then it's very similar to how one might do evals for tool calling for LLMs broadly, which is: was the tool call correct? Was the output actually parsable? There's a bunch of those. One would imagine that you'd want the tool call structure to at least be very close to 100%. And then the correctness, again, it depends a little bit on what the use case eventually demands.
SPEAKER_00
One of the other things that we are seeing is especially to make models better at tool calling, and because we have to stay within that LLM budget of the models have to be relatively small. We do see customers fine-tune smaller LLMs with their use case specific data, so that they can get tool calling quality to go up while remaining a model that is relatively small. Yeah. When you mentioned about co-location, what I understand from that is you're decreasing the latency in the network. [SPEAKER_04] Yeah, it's literally because the machines are closer to each other. [SPEAKER_04] What does that mean, if I'm using cloud providers?
SPEAKER_00
Good question. So, for example, let's say you're building a voice agent here in London, right? You have servers here. But you're using perhaps OpenAI's models for the LLM. And now, odds are that OpenAI servers might be somewhere in the US. So literally the data has to go from your server here all the way there and back. Instead, if you were able to run an open source model in the data center that you're running your voice agent, now it's basically going to be intra the data center rather than going over the Atlantic. That's one way to think about it, and distance literally has that big of an impact, aside from other networking related concerns. Thank you. Yes.
SPEAKER_00
One more question. [SPEAKER_02] So we've had an issue where we had to introduce some guard railing. So let's say we have a classifier model in between, which checks that the model's not operating an 8% discount, which is not authorized to do. How does it fit into the pipeline? And how do you do that without compromising the latency and the experience? [SPEAKER_02] Yeah, it's a great question. So the question is: what if we have other models like guardrail models or other classifiers in the mix? How does this fit into this architecture?
SPEAKER_00
You're absolutely right. This is the most basic reference architecture that one might have. But in many production settings, there's actually multiple models that might be in the mix. The guardrail or the classifier, we definitely see people introduce that right before the main LLM as well, because maybe you want to check: is this something that goes to an LLM that handles refunds versus something that handles order tracking, perhaps? And so you might have a classifier there. Guardrails at the end of the LLM generation before you produce a response, that also makes sense. And so often this ends up growing as the scope of what you're hoping to achieve grows. And it puts real pressure on latency concerns and so forth. So no easy answers, except that it becomes one or two more components to think about, have very clear guidelines and SLAs on how much budget you can really associate with them. And then independently scale them as needed. But unfortunately, no easy answers.
SPEAKER_00
Yeah, especially because when you have things where the agent already answered something which was not what it should be, but the classifier, the guardrail, for example, catches it later. You can't take things that are spoken. You might have to say something like, "Hey, sorry, I shouldn't have said that. I need to revolve that or something." So this is really a problem.
SPEAKER_00
[SPEAKER_02] Yeah, I think that's spot on. Catching all of those before the TTS model gets invoked is certainly important. One of the other patterns that we've seen is this thinker talker pattern, where you might have a small LLM that is handling all of the conversation. And so as it gets text from speech to text, it produces a response and the response might look like, "Let me think about it or let me get back to you." And then it basically issues one tool call to a much bigger model that then has much better instructions, has all the tools associated, maybe more guardrails. And then it produces a much cleaner response that you're much more comfortable with. And so that gets fed into the TTS model to produce a response. But in some ways, this is the beauty of all of these architectures. The components can only add components. And so it pushes more on reliability, having detailed observability on every single component.
SPEAKER_00
Yes. I'm curious about the upcoming voice to voice agents coming out. And I know that this is upcoming. At the end of the day, all the surrounding infrastructure that we have, the professional systems, and now with the pipeline approach, how do you do evals for observability? Do you still need to describe everything?
SPEAKER_00
Yeah, very good question. So how does observability logging and evals change? In some ways, some parts of it can still remain the same. Sometimes what you might have is, aside from pure speech to speech, you might actually have a transcription model that is running so that you could see the transcription at the same time as the audio is being generated. So that gives you some amount of auditability in terms of what audio is coming in and what audio is being produced. But yes, some evals are going to change. There's no real concept of text to speech anymore. There's no concept of pure text to text anymore. And so the evals become much more full duplex conversation evals, which is a much longer conversation. And then evals and metrics that are focused on that entire conversation.
SPEAKER_00
Right. What I mean is: is the nature of those models able to output the correct aspects of the conversation in the way that you can evaluate? Typically, a lot of the eval stuff would happen on top of the base inference API. I think we're running out of time, but okay. Well, thank you everybody for being here. [SPEAKER_01] There is no real concept of text to speech anymore. There's no concept of pure text to text here anymore. [SPEAKER_01] And so the evals become much more full duplex conversation evals, which is a much longer conversation. [SPEAKER_01] And then evals and metrics that are focused on that entire conversation.
SPEAKER_00
Right. Well, what I mean is, is the nature of those models able to output the correct respects of the conversation in the way that you can evaluate? Typically, a lot of the eval stuff would happen on top of the base inference API. I think we're running out of time, but okay. Well, thank you everybody for being here. Thank you. But, you know, depending on networking, it can be much higher as well. And so an interesting kind of direction that we're seeing folks go is how can you co-locate all your models and potentially your agent orchestrator to either be in the same data center or be very, very close to each other? How can you get them literally in the same building?
SPEAKER_00
Because that drop from 75 milliseconds to five basically gets you a 30% reduction in already a fairly optimized voice agent setup, right? So some of these things, especially again with real-time systems, it's just pretty important to have like fairly deep observability and, you know, every 10 milliseconds matters. Okay, so hopefully that's an interesting picture on sort of this pipeline architecture. But that's not the only way people do it. One of the other kind of directions that is becoming interesting is a pure speech-to-speech model. And so instead of having speech-to-text followed by LLM followed by text-to-speech
SPEAKER_00
where you're sort of coordinating and orchestrating across a number of different models, it's just way simpler if you could have a pure speech-to-speech model that still is responsible for function calling, still handles all the complicated instructions, but just a single model doing it. And of course, you know, for folks who've played around with OpenAI's real-time API, they have a single model behind the scenes. NVIDIA recently launched a model called VoiceChat. Again, very similar ideas. The reason why most of these models are not used in production too much is because they still have trouble with instruction following and tool calling.
SPEAKER_00
So, you know, the real-world experience often looks like you'll try them, and then you spend a lot of time just prompt engineering and hoping to kind of fix issues and, you know, eventually move to a pipeline architecture. But as these models get better, which I'm confident they will, it has some pretty incredible benefits. Because suddenly, you don't lose anything about the nuances of speech when that speech is getting converted into text. So the model will natively understand what was the tone, what was the emotion, was the user hesitant. That stuff will still remain with the model to make the next kind of decision.
SPEAKER_00
And this type of model also allows for sort of more full duplex communication, which basically means that the model can start producing audio while it's still receiving audio. And this means that, you know, as a customer is talking to this model, you can back channel. You can say, you know, IC or, uh-huh, like similar to what a human conversation would look like. And these models become much better at handling interruptions and bargens, which again, with the pipeline architecture, you have to do a lot more complicated engineering work around.
SPEAKER_00
So, hopefully this kind of points in the direction of like, you know, what might be coming in the future. But frankly, we have a lot of engineering work ahead for all of these kind of voice interfaces that still have to be built. If you want to learn more about what Together does, you know, here are a couple of links we are hiring. We also have a booth G1 downstairs. Happy to chat more and happy to answer questions. Thank you.
For voice-to-function calling use cases, what sort of evals do you use? And then what score do you need to get in those evals in order to have something that's good enough for production? That's like the classic answer is like, it depends, you know. But I think in terms of evals, like, you know, there's of course like component by component evals.
SPEAKER_00
So, if we're talking about sort of like the pipeline architecture and assuming that speech-to-text is good, text-to-speech is good, and the only thing that we're caring about is sort of function calling or tool calling evals in the LLM, then it's very similar to how one might do evals for tool calling for LLMs broadly, which is, you know, was the tool call correct? Was the output actually parsable? There's a bunch of those. One would imagine that, you know, you'd want the tool call structure to at least be very close to 100%. And then the correctness, again, it depends a little bit on what does the use case kind of eventually demand.
SPEAKER_00
One of the other things that we are seeing is especially to get around, to make models better at tool calling, and because we have to stay within that LLM budget of like, you know, the models have to be relatively small. We do see customers fine-tune smaller LLMs with their kind of use case specific data, so that they can get tool calling quality to go up while remaining a model that is relatively small.
SPEAKER_00
Yeah. When you mentioned about co-location, so in that, what I understand from that is you're using, you're decreasing the latency in the network of latency.
SPEAKER_04
Yeah. It's literally because the machines are closer to each other. What does that mean, like, if I'm using cloud providers?
SPEAKER_00
Good question. So, for example, you know, let's say you're using, let's say you're building a voice agent here in London, right? You have servers here. But you're using perhaps OpenAI's models for the LLM. And now, odds are that OpenAI servers might be somewhere in the US. So literally the data has, the network hop has to be from your server here all the way there and back. Instead, if you had, if you were able to run, let's say, an open source model in the data center that you're running your voice agent, now it's basically going to be intra sort of like the data center rather than going, let's say, over the Atlantic.
SPEAKER_00
That's one way to kind of think about like, and just distance literally kind of like has that big of an impact, aside from other kind of networking related concerns. Thank you. Yes. One more question.
SPEAKER_02
So we've had, when trying this issue that we had to introduce some guard rating.
SPEAKER_00
So like, let's say we have a classifier model in between, which checks that the model's not operating 8% discount,
SPEAKER_02
which is not authorized to do. How does it fit into the pipeline? And how you do that without compromising the latency and the experience? Yeah, it's a great question. So, you know, the question is like, what if we have other models like guardrail models or other classifiers in the mix? How does this fit into this architecture?
SPEAKER_00
You're absolutely right. Like this is the most kind of basic reference architecture that one might have. But in many production settings, there's actually multiple models that might be in the mix. The guardrail or the classifier, we definitely see people kind of like start to introduce that right before the main LLM as well, because maybe you want to check, is this something that goes to an LLM that handles refunds versus something that handles order tracking, perhaps? And so you might have a classifier there. Guardrails at the end of the LLM generation before you produce a response, that also makes sense.
SPEAKER_00
And so often this ends up growing as the sort of like scope of what you're hoping to achieve grows. And it puts real pressure on sort of like latency concerns and so forth. So, you know, no easy answers, except that it becomes, you know, one more or two more components to think about, have very clear sort of guidelines and SLAs on, you know, how much budget can you really kind of associate with them. And then sort of independently scaling them as needed. But unfortunately, no, no easy answers. Yeah, especially because when you have things like, like the agent already answered,
SPEAKER_00
an answer which was already not what something it should be, but the classifier, the guardrail, for example, catches later. You can't take things that are spoken, you might have to be someone like, hey, sorry, I shouldn't have said that.
SPEAKER_02
I need to revolve that or something. So this is really a problem. Yeah, I think that's spot on. Like catching all of those before the TTS model gets invoked is certainly important. One of the other things, patterns that at least like we've seen is sort of this thinker talker pattern, where you might have a small LLM that is handling all of the conversation. And so as it gets sort of text from speech to text, it produces a response and the response might look like,
SPEAKER_00
let me think about it or let me get back to you. And then it basically issues one big tool, one tool called to a much bigger model that then has, you know, much, you know, has better instructions, has all the tools associated, maybe more guardrails. And then it produces a much cleaner response that, you know, the model is much more, the architecture is like, you know, you're much more comfortable with. And so that gets fed into the TTS model to produce a response. But, you know, in some ways, this is kind of the, this is the kind of beauty of like all of these architectures is, you know,
SPEAKER_00
the components are, you know, you can only add components. And so it kind of, you know, pushes more on sort of reliability, having like, you know, sort of detailed kind of observability on every single component. Yes. I was just, I'm curious about the upcoming voice to voice kind of agents coming out. And I know that this is upcoming that at the end of the day all the surrounding infrastructure that we have in the, the international, well, we have the professional systems now and now I guess with the pipeline approach, how do you need, like, how do you do evals, you know, for observability? Like, you still need to describe everything.
Yeah. Very good question. So how does sort of observability logging evals change? So in some ways, so some parts of it can still remain the same. So sometimes what you might have is, aside from pure speech to speech, you might actually have a transcription model that is running so that you could at least seeing the transcription a lot at the same time as the audio is being generated. So that gives you some amount of auditability in terms of what audio is coming in and what audio is being produced. But yes, you know, some evals are going to change.
There is no, there's no sort of real concept of sort of text to speech anymore. There's no concept of pure kind of text to text here anymore. And so the evals become much more full duplex kind of conversation evals, which is like a much longer conversation. And then, you know, evals and metrics that are sort of focused on that entire conversation. Right. Well, what I mean is like, is the nature of those models like able to output the correct respects of the conversation in the way that you can evaluate?
SPEAKER_00
Typically, a lot of the eval stuff would happen on top, on top of the base kind of inference API. I think we're running out of time, but okay. Well, thank you everybody for being here. Thank you.