Hi everyone, my name is Nick, and I'm here with Akshay to give a talk about eValve. We are from Lyft, and we've been building Lyft customer support AI agent for a year and two now, and gave a lot of thought about how to build eValve that actually matters and scale our AI agents, multi-AI agent system. Just a bit of quick introductions. My name is Nick. I'm a data science manager, being at Lyft for six years, a longtime Lyfter. I really decided to talk to you a little bit more about eValve. I'll head on to Akshay. Hi everyone, I'm Akshay.
I'm on Nick's team, and we've been working together on customer support agents for Lyft, improving the hardness, improving the evals, things like that. And I've been at Lyft for almost four years now. I'm very excited to be here and talk about building evals that actually matter. I'm super excited to be here and super honored to be on the online track for AI engineer warfare. Yeah, and let's dive in. For the agenda of today, I will talk primarily focused on eValve. We will start by sharing how we think about the end-to-end pipeline for our evaluations for building customer support AI agent system.
We will go into a deep dive into each component more deeply as we go. We'll start by talking about offline evaluations, online evaluations, eval hardness, as well as what we are planning to build going forward. All right, let's dive in. I want to quickly explain the high-level system of how we think about evaluation system for AI agents. So here you can see we have the development phase and the production phase. So during development, if you're building agents, you should be very familiar with managing contacts, building red pipelines to give your agents educational contacts, defining your tool, building your agentic graph, as well as writing a system prompt.
So once all of that agent engineering process is done, you have an AI agent. So once again, the way we think about this is, before we launch this AI agent to production, we want to go through a rigorous offline evaluation process to make sure that this agent actually has sufficient performance before we launch this to a live user. So, coming from a data science and machine learning background, we've been building a machine learning model for a while. And I think the way that we think about agent development is very similar to building machine learning models as well.
So if we are running offline evaluations for our machine learning model before that goes to production, I think we should do the same for AI agents as well, or any other agent applications. But what I think offline evaluation, how that is different than traditional machine learning model, is that we typically, we're building specifically for customer support, our AI use case, we're building an agent that's multi-turn. So for offline evaluation, there will be a component of simulated conversations.
So you typically want to have a dataset, synthetic dataset, that's representative of your production traffic, have our user that plays out the complete multi-turn simulated conversation, and as well as having a grader such as the LLM as a judge, to be able to evaluate how good that interaction was. And then we have a launch gate, right? We want to make sure that we have certain criteria on our offline eval, and we're meeting that criteria before we decide to launch this AI agent to production.
And so the real imperative here really is that we don't want to use our live user as test data for our AI agents. And I think in any cases that is not good practice. So we really want to emphasize the importance of having an offline evaluation process.
So once the agent hits production, we also have an online evaluation pipeline as well. We have our own favorite tracing tools to trace all the executions and contacts that the AI agent uses to respond to a real user in a production environment. We have our online grader as well that grades how well our AI agent is doing in production, as far as having a human in the loop pipeline to do error analysis, identify failure mode, and feedback that insights to the development teams to continuously improve our AI agents.
I want to quickly go over, I think, three of the most common reasons why we think evaluation typically fails for different teams. So the first reason is that the grader that we create, the scores that we create, need to be meaningfully gating something. This is what we really emphasized on in the previous slide, that we need to have a launch gate. If your LLM as a judge is just floating out there, there's a score, but no one is really using that score as a meaningful gate for your development and production environment, then that LLM as a judge is not valuable. We've also seen a lot of mishaps people have when they're creating the LLM as a judge. There's a lot of different opinions out there in terms of how do you create a good LLM as a judge. And typically, and unfortunately, also very early on in our journey, the LLM judges that we created are very noisy, too generic. It will output a score, but people don't really believe in what the LLM judge is doing, or they don't think the LLM judge insight is actionable. And finally, I think when something regresses in production, we need to have a clear mechanism to be able to catch that regression, as well as identify clear owners to be able to take actions on the insights of our graders and regression gate. Very cool. I want to sequence into talking about our offline evaluation system. And as I touched on early on, I think this is the most critical piece of going from development cycle to production. We really want to have a robust offline evaluation system to be able to get more confidence in the AI agents that we are shipping to production. We took a lot of inspiration from this paper called cow bench, which is developed by the wonderful people at Sierra AI. And this is specifically for customer support, customer support AI agent. But we also think this is applicable for any user-facing authentic applications.
So, here, as you can see, this is offline simulations where you have the AI agent, as well as a user LLM, that are interacting with each other to produce the multi-turn traces. You have agent domain policy, which is essentially what, for each customer support use case, there is instruction, policy on how to handle different customer support issues. So, taking inspiration from towel bench, this is sort of a high-level approach that we have in creating our offline simulator.
So, here, as we mentioned earlier, we have our land graph agents that we've built, and we have defined instructions for our user LLM. As you can see, for simulation-wise, we define the user intent, we define what this intent is supposed to represent, and we define the user data point, or the world state of our user. For example, the driver that is coming to us might be a luxury driver that might have been driving for us for a couple of years, and so forth. And finally, to define the user behavior, user personas, as we would typically see with our real LLM user as well, we also created these different personas for our user.
One example here is this can be a loyal longtime LLM customer, but they are frustrated with LLM earning systems. So, in offline simulator, we have this land graph agent that is interacting with our user LLM model and generating this multi-turn trajectory, of this multi-turn agentic trajectory. And we also built our offline grader or offline evaluator. So, LLM judge is a big component of that. We will dive a lot more deeply into how to build a great LLM as a judge. And apart from LLM as a judge, we also have more deterministic evaluator as well. And this usually looks like a code assertion, as you will see in traditional unit tests.
And, for example, here, some of the deterministic criteria that we have created so far looks something like this, right? If in this specific interaction, the AI agent is supposed to grant concession, we will write rules like this, whether or not the AI agent has indeed granted concessions, and we will compare the agent tool calls with the expected outcome to make sure that we can measure the accuracy of the agent instruction-following behavior. So one of the big challenges that we face with running our offline evaluation is creating synthetic data is one of the key challenges in making sure our offline dataset is representative of our production data.
So here we have a meme here. Ideally, what you don't want to be doing is to simply prompt an LLM model to generate 50 different test queries for your offline datasets. So here's a couple of suggestions where you can approach this much more realistically and assemble a dataset that closely resembles your production data. So this is also something that we would try to do for customer support AI agent as well. We take some samples from our production data. Lyft users have been reaching out to support for a long time. We take a sample of our real production examples and supplement our offline dataset with that. And the second thing that we can,
So one of the big dot-shot that we face with running our offline evaluation is creating syntactic data. It is one of the key challenges in making sure our offline dataset is representative of our productions data. So here we have a meme here. Ideally, what you don't want to be doing is simply prompting an LLM model to generate 50 different test queries for your offline datasets. So here's a couple of suggestions where you can approach this much more realistically and assemble a dataset that closely resemble to your production data.
So this is also something that we would try to do for customer support AI agent as well. We take some sample from our production data. Lyft user has been reaching out to support for a long time. We take a sample of our real production example and supplement our offline dataset with that. And the second thing that we can, we do is we mutate different criteria for now offline datasets to be able to cover different golden path and edge cases. So one problem with our offline evaluator is that we're using a frontier lab. As we mentioned in the previous slide, if you were simply sampling different 50, 50 test uses, test query from by using an LLM model.
So another thing, another big godshot that we would face in building this offline simulator was that we were using a frontier lab LLM model to role play Lyft user in our offline evaluations. And for the most part, frontier lab model are trained to be helpful assistant rather than an LLM user that might not sound as nice as always. So I think for everyone that have tried customer support before, you don't typically reach out to customer support agent with a very nice verbatim.
So in our first past at running our offline evaluation, what we noticed is that our LLM user sounds almost too nice. And as you can see on the right here, these verbatims are very, very, very, very complete. The use, the LLM user are very patiently explaining the issues that they are facing in productions. And our first attempt at our offline evaluation gave us 90 plus post rate or accuracy rate, right? This almost sounds too good to be true. And I think it indeed is too good to be true.
So in reality, these are the real user verbatim that we get in productions. As you can see here, most user, they are, they're in patients, they're already frustrated. So the verbatim, they, they, they, they don't want to explain their issues like LLM user will. So typically what we see in production is something like this. And in fact, in reality, this makes AI agents much more difficult to evaluate.
So what can we do here? How can we make sure our LLM user LLM simulate this real life LLM user much closely, right? So what we, what we, how we approach this is we fine tune a LLM model with LLM user verbatim. So instead of speaking like this, very verbose and very nicely and patiently explaining their issues, our LLM user will produce verbatim that's resembled as much closely. And the benefit of this is, well, we did see our evaluation score goes down after we fine tune a LLM model that speaks more like our LLM user, and therefore making our evaluation more difficult. But in reality, this is really what you want when you're building this user simulator, right? If you have an edile that's too easy, that doesn't give you any real production insights into how your AI agent is actually going to perform.
And this also gives you a lot more room to be able to tweak your AI agents to deal with quote unquote difficult user. And as I gave a little bit of sneak peek earlier as well, we also did a lot of work to define live user personas. This can really ground our LLM user to adopt a specific live user persona and therefore simulate our real life user much more closely.
A couple of different user persona that we've defined here are bypasser user who just want to escalate to agent, regardless of and not giving AI a chance. We found secret AI skeptics. So we really also took inspiration from this paper from Microsoft user LL, Microsoft paper. And I think they adopt a very similar, similar approach as well. They fine tune a user LL model and saw evaluation score goes down. But I think in reality, this is really, really what's expected and good for your applications. All right. I'll hand it off to Akshay to talk more about LLM judge.
All right. Hi everyone. So I'm going to take from here. I'll talk about second problem, which we usually face when we are doing evals with LLM as a judge. And here you can see this is how pretty much everyone is using LLM as a judge to evaluate their agents. The focus can be slightly different based on the use case. Someone can focus more on safety. Someone can focus more on cost and latency. Someone can focus more on quality. But people are going to measure these sort of metrics more or less. So we want to detect leaks. We want to detect safety issues and things like that. So some parts are deterministic, which can be evaluated by code, and some are not, which are evaluated by the LLM.
Now, the problem with this approach is that these metrics are too generic and not actionable. So for example, we also started with our evaluation using prebuilt metrics from deep eval, which were measuring tool usage appropriateness, response helpfulness, conversation naturalness, completeness, and things like that. And we did see those metrics, but the problem was these metrics were not actionable. They were not giving us any actionable insights. If something, if let's say response helpfulness is 0.5, then what do we do with it? So things like and other scores like toxicity score, bias, fairness, conciseness, all these are kind of relevant, but if the metrics are just scores, we don't know what to do with them.
So we can use these prebuilt eval metrics as a baseline, but we shouldn't use them as our core eval metrics because we want eval metrics to be actionable and tied to the business outcome or the product which we are focusing on.
Okay. So what LLM as a judge should be: we collaborate very closely with domain experts and utilize their insights. So eval should be framed around a task, success or failure. And a binary outcome is very easy to calibrate and train LLM jets that can consistently score your agentic trajectory. When we partner with domain experts and data scientists to build metrics which are actionable and aligned with business goals, we can see much more meaningful results and actionable insights. Not only this is more consistent, but when an agent fails and interaction, we can systematically analyze the error pattern and actually know what we can do to fix this.
Here is an example of an actionable metric for our use case, which is called education rubric. Here the AI, we define the metric, how it should be, and we define success and fail criteria for LLM judge. So for example, the AI agent tries too many times to educate the user if it could have escalated the issue, or if it's escalating too soon without giving any chance to educate the user. So things like that comes under failure category. So we mark this as fail, but if it's an expected behavior, we mark it as a spouse.
Okay. So with this, how do we validate if our LLM judge is working as expected? First thing is we need to treat it as a classifier. So how we train our classification models and machine, traditional machine learning, we can also treat our evaluation judges as those traditional ML classifiers with binary output. Once we have binary outputs for every metric or a task based on our business goals and functional requirements, we can handle around hundred examples with past, fail labels and then split the data into train dev and validation sets, like how we used to do with machine learning models.
Then we score precision and recall for our judge based on human labeled ground truths, which will give us an actual report on how good our judges performing. And how do we split our data to do, to calculate precision and recall is this. So we split the data similar to how we used to do in training machine learning models, but the difference is the percentage of splits. And here we are not actually training model weights. So we are just using the data to inform judges from. Once we have binary outputs for every metric or a task based on our business goals and functional requirements, we can handle around hundred examples.
With pass, fail labels, and then split the data into train, dev, and validation sets, how we used to do with machine learning models. Then we score precision and recall for our judge based on human-labeled ground truths, which will give us an actual report on how good our judge is performing. And how do we split our data to calculate precision and recall? It is this. So we split the data similar to how we used to do in training machine learning models, but the difference is the percentage of splits. And here we are not actually training model weights. So we are just using the data to inform judges from. So the percentages are a little bit different.
So we are using the data to inform judges from the judge's prompt. So we are using the data to inform judges from the judge's prompt. And then we iterate the prompt against the dev set and improve our harness or our prompt. And then finally we validate against the test set to see that we didn't overfit on the dev examples. This is the practical way of splitting the data and then calculating precision and recall scores for your judge to actually know that the judge is working as expected.
Okay. So this is another thing which is another thing which most of us ignore when we are doing evaluations, which is criteria drift and validating the validators. The key idea is that we actually discover what our evaluation criteria is by looking at the data and grading our outputs. And our sense of quality will also evolve with new data we see and more examples we grade. So the evaluation should not be decoupled from model observations. In fact, they should be developed; they should be co-developed with the model when we are testing the evaluator and calculating the precision and recall scores.
So there's always a gap when we talk about LLM as a judge; we cannot define the criteria beforehand and then evaluate agents against them. Okay. Okay. Okay. Okay. Okay. Okay. Okay.
Okay. Okay. Okay.
Okay. To add some statistical rigor to it, right? We report alignment rates as bare point estimates. So if we add confidence intervals, we do some calibration and proper sampling, the same numbers can become more meaningful. We should definitely reserve the expensive rigor for the moment a number actually gates something or a shipping decision, or we are reporting numbers to company leaders. But depending on the use case, we should definitely have confidence intervals for the numbers we report, because every score needs an interval. So this is just a small example to give you an insight into what this actually means.
Let's say we have two evaluators and one scores 84% and the other scores 88%. And the number of samples, the number of traces which we have used, is let's say 50. So to show that this is a very small gain, and we need much more than 50 examples to actually show that this gain is real. With just 50 examples and only 4 percentage point gains, we don't actually know if this gain is real or not.
So bigger gains and paired designs need far less. And we can reserve the rigor, statistical rigor, for things which matter the most. So this is a non-exhaustive list of observed eval antipatterns. I'm not going to read all of them, but these definitely contain some low-hanging fruits. We need to put in time and effort if we need meaningful evals. So we cannot rely on LLMs for everything yet. And what I can say is ignoring the data is one of the most important things which we shouldn't do and which we sometimes don't focus on due to lack of time or resourcing, but it acts as the foundation for meaningful evaluations.
If you don't look at the data, you won't be able to create meaningful criteria or labels. And if you don't have labels, you won't be able to evaluate your judges. And if you are not evaluating your judges, you don't know if your agentic pipeline is working as expected. So this acts as a base. So one of the most important things we should not ignore. Okay. We said that we want to make metrics more actionable and standardize the pipeline, but how actually should we do it? So this gives you a template to do an error analysis loop. And it is important to know that this loop is something which runs continuously. It's not a one-off audit.
So we deep dive into raw traces. So once we have logged our traces for our agentic flows, multi-agent systems, or whatever we have in our use case, then we pinpoint failure modes. So we try to identify what exactly is failing. And we only keep the metrics that change a decision. We remove all the noise. We only prioritize the metrics which are tied to business use cases, functional requirements, and actually something which changes a decision.
Then we form a fresh premise to reevaluate. And then we repeat it. So we can have a regular cadence of doing this pipeline. It can be weekly. It can be biweekly.
But this is something which needs to run continuously. And it's not a one-off audit. Okay. So tracing. Tracing, as we said, is one of the most important things. And everything depends on it. For diving deep into raw traces, we definitely need to log them first. So we can use tools like Lang Smith, Lang Fuse, etc. to log and view the traces. And here each trace captures the full graph execution. Which nodes ran? What LLM saw? Which tools were called? What was the token usage? What was the latency for every call? And things like that. We can also enrich traces with metadata if we want, which gives you more insights than the actual data.
Okay. And we also have annotation queues. Annotation queues are nothing but an interface, which is very helpful for domain experts to label or give feedback to evaluators in an easy-to-understand UI. So they don't have to look at the raw traces, JSONs, and stuff like that to figure out what to focus on. They can use this annotation queue and they can give feedback or label examples easily. We can then add these traces to data sets for offline evaluation. Or we can use this for calculating our precision and recall for our judges. So this forms the basis to validate the evaluation with ground truth labels.
Okay. Now I'll hand it over to Nick to close the evaluation loop. Thank you, Akshay. Thank you, Akshay. And I think the goal of having Eval is to be able to feed our evaluation insights back into improving the model performance or the agent's performance. So here I'll introduce a couple of ways that we think about continual learning for our AI agent and closing the evaluation loop. And here you have model learning, context learning, and harness learning.
Model learning is really about post-training, updating the underlying model weights, and training a custom LLM model. Context learning and harness learning are really more about improving everything else other than the model. So context is improving what information the agent actually sees. This can be document. This can be the stored memories of the user, tool outputs, and so on and so forth. Harnessed. That means updating the model system prompt and tool schemas, control flow, routing, retries, and so on and so forth.
So I think from the error analysis that I actually shared earlier, I think that has really helped us to understand how can we improve our agent prompt as well as updating our knowledge base and tune the context management strategy that we have for our AI agent. The identified failure mode really helped us be able to feed those insights into actual improvement for the AI agents. So I want to quickly talk about what's next for us in our journey of building customer support AI agent at Lyft. We have some level of ability to run our offline simulator. But in fact, I think it's not repeatable.
These are currently stored as scattered scripts across different notebooks and different analysis repo. I think one thing that we're really looking into investing in is a systematic eval harness and having a harness system that can help run our offline evaluation in a systematic and standardized manner and allow different people to contribute to our evaluation suite. With predefined primitive and conflict-based workflow. And another thing that we've been talking, thinking a lot about is post training. As we mentioned earlier, identify model failure mode has really helped us tune the agent context and as far as the agent harness.
But over the years we have gathered a lot of real user signals on our agent performance as well. So we were really starting to think about how do we fine-tune a model that does different tasks for our customer support AI agent, as well as framing a reward modeling problem to enable reinforcement learning. I want to quickly share a little bit about the work that we're doing with eval harness and how we think about building eval harness for agent-facing, user-facing agentic applications.
So again, eval is really that scaffolding that we need to be able to run eval efficiently in a standardized format across all the different agents and sub-agents that we have for customer support AI agents. And as you can see from my offline simulator slide earlier And another thing that we've been talking and thinking a lot about is post-training.
As we mentioned earlier, identifying model failure mode has really helped us tune the agent context and the agent harness. But over the years, we have gathered a lot of real user signals on our agent performance as well. So we were really starting to think about how do we fine-tune a model that does different tasks for our customer support AI agent, as well as framing a reward modeling problem to enable reinforcement learning. I want to quickly share a little bit about the work that we're doing with eval harness and how we think about building eval harness for agent-facing, user-facing agentic applications.
So again, eval is really that scaffolding that we need to be able to run eval efficiently in a standardized format across all the different agents and sub-agents that we have for customer support AI agents. And as you can see from my offline simulator slide earlier, our eval harness is conflict-driven, and these are typically stored as YAML file that's easily editable by different contributor, and not just by engineers. Analysts and data scientists can contribute to this evaluation suite as well.
And with thousands, if not tens of thousands, of examples in our evaluation suite, we need parallelisms and throughputs to be able to run our offline evaluation in a reasonable amount of times.
And to enable a user to be a different user to be able to contribute to our evaluation suite, we also define primitive around our eval harness. These are high-level things like tasks, data sets, personas, LM adapter, and evaluator. And after all, the benefit of having an eval harness is that you can define the config once and run these eval indefinitely many times across different touch lines or different gates of your agent development process. This can be locally when you're developing this agent. At any point when you tune a prompt, you can run the evaluation suite and get immediate feedback on how your agent is doing compared to the previous versions.
We can run these at pre-commit hook to make sure that our performance doesn't degrade before we push a change to our agent service. Another area that we are looking at is also at CI/CD and how we can use our eval harness to build our regression test suite, acceptance test suite as well. So that wraps up our presentation today. We've gone through a lot of different topics for evaluations. Really, I think this is an end-to-end journey for building an evaluation pipeline that works for customer support AI agent or any user-facing agentic applications.
Action and I are very interested to hear about what you all have been working on and share any learnings that you have for eval, so feel free to contact us if you have any questions or you just want to share a brainstorm about how to improve your evaluation system.
All right, thank you, and I hope you all enjoy the AI engineer welfare. Thank you, thanks.
you
değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш өзгерიშ تغییر পরিবর্তন পরিবর্তন পরিবর্তন পরিবর্তন পরিবর্তন পরিবর্তন পরিবর্তন পরিবর্তন পরিবর্তন পরিবর্তন পরিবর্তন পরিবর্তন পরিবর্তন পরিবর্তন পরিবর্তন পরিবর্তন փոփոխন փոփոխন փոփոխন փոփոխন փոփոխন փոփոխन փոփոխन փոփոխन փոփոխन փոփոխন փոփոխন փոփոխन փոփոխन փոփոխन փոփոխन փոփոխन փոփոխन փոփոխन փոփոխन փոփոխਨոփոխന փոփոխनոփոխন 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경
to add some statistical rigor to it, right? We report alignment rates as bare point estimates. So if we add confidence intervals, we do some calibration and proper sampling, the same numbers can become more meaningful. We should definitely reserve the expensive rigor for the moment a number actually gates something or like a shipping decision, or we are reporting numbers to company leaders. But depending on the use case, we should definitely have confidence intervals for the numbers we report, because every score needs an interval.
So this is just a small example to give you an insight on what this actually means. Let's say we have two evaluators and one scores 84% and the other scores 88%. And the number of samples, the number of traces which we have used is let's say 50. So to show that like this is a very small gain and we need much more than 50 examples to actually show that this gain is real. With just 50 examples and only 4% point percentage gains, we don't actually know if this gain is real or not. So bigger gains and bigger gains and paired designs need far less. And we can reserve the rigor, statistical rigor for things which matter the most.
So this is a non-exhaustive list of observed eval antipatterns. I'm not going to read all of them, but these definitely contain some low-hanging fruits. We need to put in time and effort if we need meaningful evals. So we cannot rely on LLMs or everything yet. And what I can say is ignoring the data is one of the most important things which we shouldn't do and which we sometimes don't focus on due to lack of time or resourcing, but it acts as the foundation for meaningful evaluations. If you don't look at the data, you won't be able to create meaningful criteria or labels. And if you don't have labels, you won't be able to evaluate your judges.
And if you are not evaluating your judges, you don't know if your agentic pipeline is working as expected. So this acts as a base. So one of the most important things we should not ignore. Okay. We said that we want to make metrics more actionable and standardize the pipeline, but how actually we should do it. So this gives you like a template to do an error analysis loop. And it is important to know that this loop is something which runs continuously. It's not in one-off audit. So we deep dive into raw traces. So once we have logged our traces for our agentic flows, multi-agent systems, or whatever we have in our use case, then we pinpoint failure modes.
So basically we try to identify what exactly is failing. And we only keep the metrics that change a decision. We remove all the noise. We only prioritize on the metrics which are tied to business use cases, functional requirements, and actually something which changes a decision. Then we form a fresh premise to reevaluate. And then we repeat it. So we can have like a regular cadence of doing this pipeline. It can be weekly. It can be biweekly. But this is something which needs to run continuously. And it's not a one-off audit.
Okay. So tracing. Tracing, as we said, is one of the most important things. And everything kind of depends on it. For diving deep into raw traces, we definitely need to log them first. So we can use tools like Lange Smith, Lange Fuse, etc. to log and view the traces. And here each trace captures the full graph execution. Which nodes ran? What LLM saw? Which tools were called? What was the token usage? What was the latency for every call? And things like that. We can also enrich traces with metadata if we want, which gives you more insights than the actual data. Okay. And we also have annotation queues. Annotation queues are nothing but an interface,
which is very helpful for domain experts to label or give feedback to evaluators in an easy to understand UI. So they don't have to look at the raw traces, JSONs and stuff like that to figure out what to focus on. They can use this annotation queue and they can give feedback or label examples easily. We can then add these traces to data sets for offline evaluation. Or we can use this for calculating our precision and recall for our judges. So this forms the basis to validate the evaluation with ground truth labels.
Okay. Now I'll hand it over to Nick to close the evaluation loop. Thank you, Akshay. Thank you, Akshay. And I think the goal of having Eval is to be able to feed our evaluation insights back into improving the model performance or the agent's performance. So here I'll introduce a couple of ways that we think about continual learning for our AI agent and closing the evaluation loop. And here you have model learning, context learning and harness learning. Model learning is really about post-training, updating the underlying model weights and training a custom LLM model.
Context learning and harness learning is really more about improving everything else other than the model. So context is improving what information the agent actually sees. This can be document. This can be the store memories of the user, two outputs and so and so forth. Harnessed. That means updating the model system prompt. And two schemas, control flow, routing, retries and so and so forth. So I think from the error analysis that I actually shared earlier, I think that has really helped us to understand how can we improve our agent prompt as well as updating our knowledge base and tune the context management strategy that we have for our AI agent.
The identify failure mode really helped us be able to feed that insights into actual improvement for the AI agents. So I want to quickly talk about what's next for us in our journey of building customer support AI agent at Lyft. We have a we have we have some some level of ability to run our offline simulator. But in fact, I think, you know, it's not repeatable. These are currently stored as scatter scripts across different notebooks and different analysis repo.
I think one thing that we're looking really really looking into investing is a systematic eval harness and having a harness system that can help run our offline evaluation in a systematic and standardized manner and allow different people to contribute to our evaluation suite. We're predefined primitive and conflict based workflow. And another thing that we've been talking, thinking a lot about is post training. As we mentioned earlier, you know, identify model failure mode has really helped us tune the agent context and as far as the agent harness. but over the years we have gathered a lot of real user signals on our agent performance as well
so we were really starting to think about how do we fine-tune a model that does different tasks for our customer support ai agent as well as framing a reward modeling problem to to enable reinforcement learning i want to quickly share a little bit about the work that we're doing with eval hardest and how we think about building building building eval harness for agent facing user facing agentic applications so again eval is really that scaffolding that you that we need to be able to run eval efficiently in a standardized format across all the different agents and sub-agents that we have
for customer support ai agents and as you can see from my offline simulator slide earlier our eval harness is conflict-driven and these are typically stored as yaml file that's easily editable by different contributor and not just by engineers analysts and data scientists can can contribute to this evaluation suite as well and you know with thousands if not tens of thousands of examples in our evaluation suite we need you know parallelisms and throughputs to be able to run our offline evaluation in a reasonable amount of times and to enable a user to be a different user to be able to
contribute to our evaluation suite we also define you know primitive around our eval either harness these are high level things like tasks data sets personas lm adapter and evaluator and after all you know the benefit of having an evil hardest is that you can define the config once and run these eval indefinitely many times across different at different touch lines or different gates of your agent development process this can be you know locally when you're developing this agent uh at any point when you tune a prompt you can run the evaluation suite and get an immediate in immediate feedback on how your agent is doing compared to the previous versions
we can run these at pre-committ hook uh to make sure that our performance doesn't degrade before we push a change to our agent service another area that we are looking at is also at ci cd and how we can use our eval harness to build our regression test suite acceptance acceptance test suite as well so that wraps up our presentation today uh we've gone through a lot of we've gone through a lot of uh a lot of different topics uh for evaluations uh really i think this is an end-to-end journey uh for building a evaluation pipeline that works for customer support ai agent or any user-facing agentic applications
uh action and i we are very interested to hear about you know what you all have been working on and share any learnings that you have for eval so feel free to contact us uh if you have any questions or you just want to share a brainstorm about how to improve your evaluation system all right thank you and i hope you all enjoy the ai engineer welfare thank you thanks you
değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş