Build Evals That Actually Matter - Nick Ung, Lyft
Description
Your agent passes offline evals at 90%. You ship. Production immediately finds failure modes your eval never saw. Sound familiar? The culprit is almost always the same: the "customer" in your offline eval is an off-the-shelf LLM that sounds nothing like your real users, and your synthetic test set doesn't capture how messy, angry, or off-topic real conversations get. Your eval was too easy. At Lyft, our customer-care agents resolve roughly a third of all customer issues — millions of conversations a month. To trust them at that scale, we built an adversarial user simulator: a fine-tuned LLM trained on real Lyft rider and driver transcripts that can role-play frustrated, confused, and adversarial users with the same distribution as production. It found regressions our synthetic dataset missed for months. This talk walks through the full eval lifecycle that surrounds it: the harness primitives that let any engineer write a benchmark in 20 lines, how we calibrate LLM-judge rubrics against human labels until they match inter-rater agreement, how we route failed production traces back into the offline test set, and the continual-learning loop that feeds improvements into prompts, harness, and the model. Speakers: - Nick Ung (Lyft): Nick Ung leads Data Science for Safety & Customer Care at Lyft, where his team built and operates the multi-agent platform that powers AI agents resolving roughly a third of all Lyft customer issues. LinkedIn: https://www.linkedin.com/in/unglikteng
Summary
Generated by gpt-5.6-terraAt-a-Glance
- Verdict: Watch fully
- Core thesis: Meaningful agent evaluations require realistic multi-turn offline simulations, task-specific and human-validated judges, production tracing, and gates that directly control launch and regression decisions.
- Why it matters: For user-facing agents, generic LLM-judge scores and overly polite synthetic users can create false confidence; Lyft's approach turns evaluations into an operational control loop for safely improving and shipping agents.
- Best use: Use this as a practical blueprint for designing an eval control plane: simulation data, deterministic checks, calibrated LLM judges, trace review, annotation workflows, and CI/CD regression gates.
Executive Summary
Nick Ung and Akshay from Lyft describe how their customer-support agent team evaluates a multi-turn, multi-agent system before and after production deployment. Their central operating principle is that live customers must not function as test data: agents should clear a rigorous offline launch gate, then be monitored through production tracing, online grading, human review, and a recurring error-analysis loop.
The most important technical lesson is that evaluation quality depends on realism and actionability. Lyft simulates full conversations between its agent and an LLM user, but found an off-the-shelf frontier model was implausibly cooperative and produced 90%+ pass rates. They improved hardness by grounding scenarios in real support examples, mutating cases for edge coverage, fine-tuning the user simulator on real customer language, and defining difficult personas such as users who demand escalation or distrust AI.
They reject generic judge metrics such as helpfulness, naturalness, or tool-use appropriateness as core decision metrics because a score alone does not identify what to fix. Instead, domain experts define binary, task-specific pass/fail rubrics tied to product behavior—for example, whether an agent improperly persists in educating a customer when escalation is appropriate. The LLM judge is treated like a conventional classifier: calibrated against human labels, iterated on a development set, and tested for precision and recall on held-out examples.
The presentation's broader operating model is an eval flywheel: trace full executions, collect expert annotations, isolate failure modes, retain only metrics that alter a shipping or product decision, then use findings to improve model weights, context, prompts, tools, routing, retries, and control flow. Lyft is moving from scattered notebooks toward a config-driven YAML eval harness with reusable primitives and execution at local development, pre-commit, and CI/CD gates.
Key Takeaways
- Claim: Offline evaluation must act as a real launch gate rather than a dashboard of advisory scores. | Evidence: Lyft places an offline evaluation phase between agent development and production, with explicit criteria an agent must meet before launch; the speakers argue that an LLM-as-judge score is not valuable if it does not gate a development or production decision. | Implication: Define concrete release thresholds for each agent change and make failing them block deployment rather than merely generate a report. | Caveat: Offline success is necessary but not sufficient because production behavior still requires tracing, online grading, and human error analysis.
- Claim: For multi-turn customer-facing agents, the offline test set must simulate realistic conversations, not just isolated prompts. | Evidence: Lyft runs its LangGraph agent against a user LLM configured with an intent, world state/user data, behavior, and persona; the resulting trajectory is evaluated by LLM judges and deterministic assertions such as whether the agent actually called a tool to grant an expected concession. | Implication: Test whole agent trajectories, including state, tool calls, policy adherence, and user reactions—not only the final answer quality. | Caveat: Synthetic scenarios must be continuously checked against production traffic; a simulator can otherwise encode assumptions that diverge from real users.
- Claim: Eval hardness is a feature: easy, cooperative simulated users produce misleadingly high scores and weak production readiness. | Evidence: Lyft's first simulator used a frontier LLM trained to be a helpful assistant, yielding unusually complete and patient user messages and 90%+ pass rates. The team fine-tuned its simulated user model on real Lyft customer verbatims and added personas such as frustrated long-time customers, escalation-seeking bypassers, and AI skeptics; scores fell, which the speakers treat as a healthier signal. | Implication: Build test-data generation around real interaction distributions, terse/ambiguous inputs, adversarial behaviors, and intentionally mutated golden-path and edge cases. | Caveat: Lower scores alone do not prove better evaluation; realism must be established by comparison with actual production interactions and observed failure modes.
- Claim: LLM judges should use narrow, binary, domain-expert-defined rubrics tied to task success or failure, not generic quality metrics. | Evidence: Lyft began with prebuilt DeepEval measures such as tool-use appropriateness, helpfulness, naturalness, and completeness, but found them non-actionable. Its alternative is a specific 'education rubric' that marks failure when an agent keeps educating a user when it should escalate, or escalates before reasonably attempting education. | Implication: Replace broad quality scores with a small set of failure-mode metrics where each outcome maps to an owner and a known remediation path. | Caveat: Some deterministic criteria should remain code assertions; LLM judges are for evaluative dimensions that cannot reliably be reduced to fixed rules.
- Claim: An LLM judge is itself a model component that needs validation against human ground truth. | Evidence: The speakers recommend labeling roughly 100 examples as pass/fail, splitting them into train, development, and validation sets, using the development set to improve the judge prompt, and reporting judge precision and recall against held-out human labels. They stress that the split informs prompts rather than trains model weights. | Implication: Version, test, and monitor evaluators like classifiers; do not treat an unvalidated LLM prompt as authoritative measurement. | Caveat: Evaluation criteria will drift as teams observe new failures and refine what quality means, so a one-time calibration is insufficient.
- Claim: Reported eval gains need statistical discipline when they influence shipping or leadership decisions. | Evidence: The talk contrasts 84% versus 88% scores measured on only 50 traces, arguing that a four-point difference at that sample size may not be real. The recommendation is confidence intervals, calibration, and appropriate sampling, with the most expensive rigor reserved for launch gates and executive reporting. | Implication: Do not promote small score movements as improvements without uncertainty estimates and a comparison design suited to the decision's stakes. | Caveat: The speakers do not prescribe a universal sample size; required sample size depends on effect size and whether the comparison uses paired designs.
- Claim: The durable evaluation system is a continuous trace-to-remediation loop, supported by a standardized harness rather than scattered scripts. | Evidence: Lyft's loop is: inspect raw traces, identify failure modes, keep only metrics that change decisions, form a new evaluation premise, and repeat weekly or biweekly. It logs graph execution, model inputs, tool calls, token usage, and latency; annotation queues let domain experts label examples. Lyft is building a YAML/config-driven harness with primitives for tasks, datasets, personas, LLM adapters, and evaluators, runnable locally, at pre-commit, and in CI/CD. | Implication: Treat eval infrastructure as shared product infrastructure: make cases editable by non-engineers, support high-throughput parallel runs, and wire regression/acceptance tests into the delivery pipeline. | Caveat: Lyft acknowledges its own current offline simulator is not yet fully repeatable because it remains distributed across notebooks and analysis repositories.
Detailed Brief
Closing the loop: choose the right remediation layer
- Claims: Evaluation findings should not automatically lead to model fine-tuning; they should identify the layer actually responsible for the failure.; Lyft separates improvement work into model learning, context learning, and harness learning.
- Evidence: Model learning means post-training or updating underlying model weights for customer-support tasks.; Context learning changes what the agent sees, including documents, user memories, and tool outputs.; Harness learning changes system prompts, tool schemas, routing, retries, and control flow.; Lyft reports that error analysis has already informed prompt changes, knowledge-base updates, and context-management tuning; it is also exploring task-specific fine-tuning, reward modeling, and reinforcement learning.
- Caveats: The talk presents reward modeling and reinforcement learning as future directions, not established results or demonstrated production outcomes.; Model post-training should follow evidence that harness and context changes cannot address the observed failure mode.
- Implications: Maintain a failure taxonomy that routes each issue to model, retrieval/context, tool/interface, or orchestration owners.; Preserve production signals and human labels as potential assets for future post-training rather than viewing them only as QA artifacts.
Data and observability are the foundations of evaluator quality
- Claims: Criteria should be co-developed with model observations because teams discover their real quality bar by reviewing data.; Ignoring raw data breaks the chain from criteria definition to labels to judge validation.
- Evidence: The speakers recommend trace systems such as LangSmith or Langfuse, with each trace capturing graph nodes, model-visible context, tool calls, token usage, latency, and optional business metadata.; Annotation queues provide domain experts an understandable interface for feedback rather than requiring them to inspect raw JSON; annotated production traces can become offline datasets and judge-validation data.
- Caveats: More telemetry is not automatically useful; the team explicitly recommends retaining only metrics that change a decision.
- Implications: Design trace schemas so evaluation, debugging, cost/latency analysis, and expert labeling all operate on the same canonical execution record.; Give domain specialists direct participation in labeling and rubric evolution instead of bottlenecking all quality definition through engineering.
Notable Concepts & Terms
- Offline multi-turn simulation: A pre-production test method in which an agent interacts with a simulated user across a complete trajectory, allowing evaluation of policy, state, tool calls, and final resolution.
- Eval hardness: How difficult and production-representative an evaluation is; Lyft argues that a score decline after making users more realistic can indicate a more useful benchmark.
- LLM-as-a-judge: An LLM used to grade dimensions that are not deterministic, but one that must be narrowly rubriced, human-calibrated, and measured for precision and recall.
- Deterministic evaluator: A code-level assertion for observable requirements, such as confirming an expected tool call or concession outcome.
- Criteria drift: The quality standard evolves as a team observes more model behavior and failure examples, requiring continued rubric and judge maintenance.
- Annotation queue: A domain-expert-friendly labeling interface that converts production traces into ground truth for judge validation and new offline test cases.
- Context learning: Improving agent performance by changing available information—documents, memory, or tool outputs—rather than changing model weights.
- Harness learning: Improving the agent's orchestration layer: prompts, tool schemas, routing, retries, and control flow.
Operator Notes / Why Ken Should Care
- Establish a release policy in which every agent change must pass a defined offline regression and acceptance suite; designate an accountable owner for each gate and failure category.
- Create a seed corpus from real interaction traces, then explicitly generate terse, incomplete, hostile, escalation-demanding, and policy-boundary variants rather than relying on generic synthetic prompts.
- Audit current evaluator metrics: remove or demote any score that does not change a release, routing, or product decision; replace it with binary rubrics owned by domain experts.
- Build a judge-validation dataset with human pass/fail labels and require held-out precision/recall reporting before using an LLM judge for a high-stakes gate.
- Add confidence intervals and paired comparisons to release scorecards; block claims of improvement when sample sizes cannot distinguish the observed change from noise.
- Standardize trace capture and annotation into a shared eval data pipeline, then promote high-value production failures into versioned regression cases.
- Implement a config-driven eval harness with reusable task, dataset, persona, model-adapter, and evaluator primitives; run it locally, before merge, and in CI/CD.
Source/Metadata
- Title: Build Evals That Actually Matter - Nick Ung, Lyft
- Transcript words: 7630
- Duration seconds: 2264
- Timestamp note: No usable timestamps or chapters were provided. The transcript contains duplicated sections and trailing extraction noise.
Transcript
Hi everyone, my name is Nick, and I'm here with Akshay to give a talk about eValve. We are from Lyft, and we've been building Lyft customer support AI agent for a year and two now, and gave a lot of thought about how to build eValve that actually matters and scale our AI agents, multi-AI agent system. Just a bit of quick introductions. My name is Nick. I'm a data science manager, being at Lyft for six years, a longtime Lyfter. I really decided to talk to you a little bit more about eValve. I'll head on to Akshay. Hi everyone, I'm Akshay. I'm on Nick's team, and we've been working together on customer support agents for Lyft, improving the hardness, improving the evals, things like that. And I've been at Lyft for almost four years now. I'm very excited to be here and talk about building evals that actually matter. I'm super excited to be here and super honored to be on the online track for AI engineer warfare. Yeah, and let's dive in. For the agenda of today, I will talk primarily focused on eValve. We will start by sharing how we think about the end-to-end pipeline for our evaluations for building customer support AI agent system. We will go into a deep dive into each component more deeply as we go. We'll start by talking about offline evaluations, online evaluations, eval hardness, as well as what we are planning to build going forward. All right, let's dive in. I want to quickly explain the high-level system of how we think about evaluation system for AI agents. So here you can see we have the development phase and the production phase. So during development, if you're building agents, you should be very familiar with managing contacts, building red pipelines to give your agents educational contacts, defining your tool, building your agentic graph, as well as writing a system prompt. So once all of that agent engineering process is done, you have an AI agent. So once again, the way we think about this is, before we launch this AI agent to production, we want to go through a rigorous offline evaluation process to make sure that this agent actually has sufficient performance before we launch this to a live user. So, coming from a data science and machine learning background, we've been building a machine learning model for a while. And I think the way that we think about agent development is very similar to building machine learning models as well. So if we are running offline evaluations for our machine learning model before that goes to production, I think we should do the same for AI agents as well, or any other agent applications. But what I think offline evaluation, how that is different than traditional machine learning model, is that we typically, we're building specifically for customer support, our AI use case, we're building an agent that's multi-turn. So for offline evaluation, there will be a component of simulated conversations. So you typically want to have a dataset, synthetic dataset, that's representative of your production traffic, have our user that plays out the complete multi-turn simulated conversation, and as well as having a grader such as the LLM as a judge, to be able to evaluate how good that interaction was. And then we have a launch gate, right? We want to make sure that we have certain criteria on our offline eval, and we're meeting that criteria before we decide to launch this AI agent to production. And so the real imperative here really is that we don't want to use our live user as test data for our AI agents. And I think in any cases that is not good practice. So we really want to emphasize the importance of having an offline evaluation process. So once the agent hits production, we also have an online evaluation pipeline as well. We have our own favorite tracing tools to trace all the executions and contacts that the AI agent uses to respond to a real user in a production environment. We have our online grader as well that grades how well our AI agent is doing in production, as far as having a human in the loop pipeline to do error analysis, identify failure mode, and feedback that insights to the development teams to continuously improve our AI agents. I want to quickly go over, I think, three of the most common reasons why we think evaluation typically fails for different teams. So the first reason is that the grader that we create, the scores that we create, need to be meaningfully gating something. This is what we really emphasized on in the previous slide, that we need to have a launch gate. If your LLM as a judge is just floating out there, there's a score, but no one is really using that score as a meaningful gate for your development and production environment, then that LLM as a judge is not valuable. We've also seen a lot of mishaps people have when they're creating the LLM as a judge. There's a lot of different opinions out there in terms of how do you create a good LLM as a judge. And typically, and unfortunately, also very early on in our journey, the LLM judges that we created are very noisy, too generic. It will output a score, but people don't really believe in what the LLM judge is doing, or they don't think the LLM judge insight is actionable. And finally, I think when something regresses in production, we need to have a clear mechanism to be able to catch that regression, as well as identify clear owners to be able to take actions on the insights of our graders and regression gate. Very cool. I want to sequence into talking about our offline evaluation system. And as I touched on early on, I think this is the most critical piece of going from development cycle to production. We really want to have a robust offline evaluation system to be able to get more confidence in the AI agents that we are shipping to production. We took a lot of inspiration from this paper called cow bench, which is developed by the wonderful people at Sierra AI. And this is specifically for customer support, customer support AI agent. But we also think this is applicable for any user-facing authentic applications. So, here, as you can see, this is offline simulations where you have the AI agent, as well as a user LLM, that are interacting with each other to produce the multi-turn traces. You have agent domain policy, which is essentially what, for each customer support use case, there is instruction, policy on how to handle different customer support issues. So, taking inspiration from towel bench, this is sort of a high-level approach that we have in creating our offline simulator. So, here, as we mentioned earlier, we have our land graph agents that we've built, and we have defined instructions for our user LLM. As you can see, for simulation-wise, we define the user intent, we define what this intent is supposed to represent, and we define the user data point, or the world state of our user. For example, the driver that is coming to us might be a luxury driver that might have been driving for us for a couple of years, and so forth. And finally, to define the user behavior, user personas, as we would typically see with our real LLM user as well, we also created these different personas for our user. One example here is this can be a loyal longtime LLM customer, but they are frustrated with LLM earning systems. So, in offline simulator, we have this land graph agent that is interacting with our user LLM model and generating this multi-turn trajectory, of this multi-turn agentic trajectory. And we also built our offline grader or offline evaluator. So, LLM judge is a big component of that. We will dive a lot more deeply into how to build a great LLM as a judge. And apart from LLM as a judge, we also have more deterministic evaluator as well. And this usually looks like a code assertion, as you will see in traditional unit tests. And, for example, here, some of the deterministic criteria that we have created so far looks something like this, right? If in this specific interaction, the AI agent is supposed to grant concession, we will write rules like this, whether or not the AI agent has indeed granted concessions, and we will compare the agent tool calls with the expected outcome to make sure that we can measure the accuracy of the agent instruction-following behavior. So one of the big challenges that we face with running our offline evaluation is creating synthetic data is one of the key challenges in making sure our offline dataset is representative of our production data. So here we have a meme here. Ideally, what you don't want to be doing is to simply prompt an LLM model to generate 50 different test queries for your offline datasets. So here's a couple of suggestions where you can approach this much more realistically and assemble a dataset that closely resembles your production data. So this is also something that we would try to do for customer support AI agent as well. We take some samples from our production data. Lyft users have been reaching out to support for a long time. We take a sample of our real production examples and supplement our offline dataset with that. And the second thing that we can, So one of the big dot-shot that we face with running our offline evaluation is creating syntactic data. It is one of the key challenges in making sure our offline dataset is representative of our productions data. So here we have a meme here. Ideally, what you don't want to be doing is simply prompting an LLM model to generate 50 different test queries for your offline datasets. So here's a couple of suggestions where you can approach this much more realistically and assemble a dataset that closely resemble to your production data. So this is also something that we would try to do for customer support AI agent as well. We take some sample from our production data. Lyft user has been reaching out to support for a long time. We take a sample of our real production example and supplement our offline dataset with that. And the second thing that we can, we do is we mutate different criteria for now offline datasets to be able to cover different golden path and edge cases. So one problem with our offline evaluator is that we're using a frontier lab. As we mentioned in the previous slide, if you were simply sampling different 50, 50 test uses, test query from by using an LLM model. So another thing, another big godshot that we would face in building this offline simulator was that we were using a frontier lab LLM model to role play Lyft user in our offline evaluations. And for the most part, frontier lab model are trained to be helpful assistant rather than an LLM user that might not sound as nice as always. So I think for everyone that have tried customer support before, you don't typically reach out to customer support agent with a very nice verbatim. So in our first past at running our offline evaluation, what we noticed is that our LLM user sounds almost too nice. And as you can see on the right here, these verbatims are very, very, very, very complete. The use, the LLM user are very patiently explaining the issues that they are facing in productions. And our first attempt at our offline evaluation gave us 90 plus post rate or accuracy rate, right? This almost sounds too good to be true. And I think it indeed is too good to be true. So in reality, these are the real user verbatim that we get in productions. As you can see here, most user, they are, they're in patients, they're already frustrated. So the verbatim, they, they, they, they don't want to explain their issues like LLM user will. So typically what we see in production is something like this. And in fact, in reality, this makes AI agents much more difficult to evaluate. So what can we do here? How can we make sure our LLM user LLM simulate this real life LLM user much closely, right? So what we, what we, how we approach this is we fine tune a LLM model with LLM user verbatim. So instead of speaking like this, very verbose and very nicely and patiently explaining their issues, our LLM user will produce verbatim that's resembled as much closely. And the benefit of this is, well, we did see our evaluation score goes down after we fine tune a LLM model that speaks more like our LLM user, and therefore making our evaluation more difficult. But in reality, this is really what you want when you're building this user simulator, right? If you have an edile that's too easy, that doesn't give you any real production insights into how your AI agent is actually going to perform. And this also gives you a lot more room to be able to tweak your AI agents to deal with quote unquote difficult user. And as I gave a little bit of sneak peek earlier as well, we also did a lot of work to define live user personas. This can really ground our LLM user to adopt a specific live user persona and therefore simulate our real life user much more closely. A couple of different user persona that we've defined here are bypasser user who just want to escalate to agent, regardless of and not giving AI a chance. We found secret AI skeptics. So we really also took inspiration from this paper from Microsoft user LL, Microsoft paper. And I think they adopt a very similar, similar approach as well. They fine tune a user LL model and saw evaluation score goes down. But I think in reality, this is really, really what's expected and good for your applications. All right. I'll hand it off to Akshay to talk more about LLM judge. All right. Hi everyone. So I'm going to take from here. I'll talk about second problem, which we usually face when we are doing evals with LLM as a judge. And here you can see this is how pretty much everyone is using LLM as a judge to evaluate their agents. The focus can be slightly different based on the use case. Someone can focus more on safety. Someone can focus more on cost and latency. Someone can focus more on quality. But people are going to measure these sort of metrics more or less. So we want to detect leaks. We want to detect safety issues and things like that. So some parts are deterministic, which can be evaluated by code, and some are not, which are evaluated by the LLM. Now, the problem with this approach is that these metrics are too generic and not actionable. So for example, we also started with our evaluation using prebuilt metrics from deep eval, which were measuring tool usage appropriateness, response helpfulness, conversation naturalness, completeness, and things like that. And we did see those metrics, but the problem was these metrics were not actionable. They were not giving us any actionable insights. If something, if let's say response helpfulness is 0.5, then what do we do with it? So things like and other scores like toxicity score, bias, fairness, conciseness, all these are kind of relevant, but if the metrics are just scores, we don't know what to do with them. So we can use these prebuilt eval metrics as a baseline, but we shouldn't use them as our core eval metrics because we want eval metrics to be actionable and tied to the business outcome or the product which we are focusing on. Okay. So what LLM as a judge should be: we collaborate very closely with domain experts and utilize their insights. So eval should be framed around a task, success or failure. And a binary outcome is very easy to calibrate and train LLM jets that can consistently score your agentic trajectory. When we partner with domain experts and data scientists to build metrics which are actionable and aligned with business goals, we can see much more meaningful results and actionable insights. Not only this is more consistent, but when an agent fails and interaction, we can systematically analyze the error pattern and actually know what we can do to fix this. Here is an example of an actionable metric for our use case, which is called education rubric. Here the AI, we define the metric, how it should be, and we define success and fail criteria for LLM judge. So for example, the AI agent tries too many times to educate the user if it could have escalated the issue, or if it's escalating too soon without giving any chance to educate the user. So things like that comes under failure category. So we mark this as fail, but if it's an expected behavior, we mark it as a spouse. Okay. So with this, how do we validate if our LLM judge is working as expected? First thing is we need to treat it as a classifier. So how we train our classification models and machine, traditional machine learning, we can also treat our evaluation judges as those traditional ML classifiers with binary output. Once we have binary outputs for every metric or a task based on our business goals and functional requirements, we can handle around hundred examples with past, fail labels and then split the data into train dev and validation sets, like how we used to do with machine learning models. Then we score precision and recall for our judge based on human labeled ground truths, which will give us an actual report on how good our judges performing. And how do we split our data to do, to calculate precision and recall is this. So we split the data similar to how we used to do in training machine learning models, but the difference is the percentage of splits. And here we are not actually training model weights. So we are just using the data to inform judges from. Once we have binary outputs for every metric or a task based on our business goals and functional requirements, we can handle around hundred examples. With pass, fail labels, and then split the data into train, dev, and validation sets, how we used to do with machine learning models. Then we score precision and recall for our judge based on human-labeled ground truths, which will give us an actual report on how good our judge is performing. And how do we split our data to calculate precision and recall? It is this. So we split the data similar to how we used to do in training machine learning models, but the difference is the percentage of splits. And here we are not actually training model weights. So we are just using the data to inform judges from. So the percentages are a little bit different. So we are using the data to inform judges from the judge's prompt. So we are using the data to inform judges from the judge's prompt. And then we iterate the prompt against the dev set and improve our harness or our prompt. And then finally we validate against the test set to see that we didn't overfit on the dev examples. This is the practical way of splitting the data and then calculating precision and recall scores for your judge to actually know that the judge is working as expected. Okay. So this is another thing which is another thing which most of us ignore when we are doing evaluations, which is criteria drift and validating the validators. The key idea is that we actually discover what our evaluation criteria is by looking at the data and grading our outputs. And our sense of quality will also evolve with new data we see and more examples we grade. So the evaluation should not be decoupled from model observations. In fact, they should be developed; they should be co-developed with the model when we are testing the evaluator and calculating the precision and recall scores. So there's always a gap when we talk about LLM as a judge; we cannot define the criteria beforehand and then evaluate agents against them. Okay. Okay. Okay. Okay. Okay. Okay. Okay. Okay. Okay. Okay. Okay. To add some statistical rigor to it, right? We report alignment rates as bare point estimates. So if we add confidence intervals, we do some calibration and proper sampling, the same numbers can become more meaningful. We should definitely reserve the expensive rigor for the moment a number actually gates something or a shipping decision, or we are reporting numbers to company leaders. But depending on the use case, we should definitely have confidence intervals for the numbers we report, because every score needs an interval. So this is just a small example to give you an insight into what this actually means. Let's say we have two evaluators and one scores 84% and the other scores 88%. And the number of samples, the number of traces which we have used, is let's say 50. So to show that this is a very small gain, and we need much more than 50 examples to actually show that this gain is real. With just 50 examples and only 4 percentage point gains, we don't actually know if this gain is real or not. So bigger gains and paired designs need far less. And we can reserve the rigor, statistical rigor, for things which matter the most. So this is a non-exhaustive list of observed eval antipatterns. I'm not going to read all of them, but these definitely contain some low-hanging fruits. We need to put in time and effort if we need meaningful evals. So we cannot rely on LLMs for everything yet. And what I can say is ignoring the data is one of the most important things which we shouldn't do and which we sometimes don't focus on due to lack of time or resourcing, but it acts as the foundation for meaningful evaluations. If you don't look at the data, you won't be able to create meaningful criteria or labels. And if you don't have labels, you won't be able to evaluate your judges. And if you are not evaluating your judges, you don't know if your agentic pipeline is working as expected. So this acts as a base. So one of the most important things we should not ignore. Okay. We said that we want to make metrics more actionable and standardize the pipeline, but how actually should we do it? So this gives you a template to do an error analysis loop. And it is important to know that this loop is something which runs continuously. It's not a one-off audit. So we deep dive into raw traces. So once we have logged our traces for our agentic flows, multi-agent systems, or whatever we have in our use case, then we pinpoint failure modes. So we try to identify what exactly is failing. And we only keep the metrics that change a decision. We remove all the noise. We only prioritize the metrics which are tied to business use cases, functional requirements, and actually something which changes a decision. Then we form a fresh premise to reevaluate. And then we repeat it. So we can have a regular cadence of doing this pipeline. It can be weekly. It can be biweekly. But this is something which needs to run continuously. And it's not a one-off audit. Okay. So tracing. Tracing, as we said, is one of the most important things. And everything depends on it. For diving deep into raw traces, we definitely need to log them first. So we can use tools like Lang Smith, Lang Fuse, etc. to log and view the traces. And here each trace captures the full graph execution. Which nodes ran? What LLM saw? Which tools were called? What was the token usage? What was the latency for every call? And things like that. We can also enrich traces with metadata if we want, which gives you more insights than the actual data. Okay. And we also have annotation queues. Annotation queues are nothing but an interface, which is very helpful for domain experts to label or give feedback to evaluators in an easy-to-understand UI. So they don't have to look at the raw traces, JSONs, and stuff like that to figure out what to focus on. They can use this annotation queue and they can give feedback or label examples easily. We can then add these traces to data sets for offline evaluation. Or we can use this for calculating our precision and recall for our judges. So this forms the basis to validate the evaluation with ground truth labels. Okay. Now I'll hand it over to Nick to close the evaluation loop. Thank you, Akshay. Thank you, Akshay. And I think the goal of having Eval is to be able to feed our evaluation insights back into improving the model performance or the agent's performance. So here I'll introduce a couple of ways that we think about continual learning for our AI agent and closing the evaluation loop. And here you have model learning, context learning, and harness learning. Model learning is really about post-training, updating the underlying model weights, and training a custom LLM model. Context learning and harness learning are really more about improving everything else other than the model. So context is improving what information the agent actually sees. This can be document. This can be the stored memories of the user, tool outputs, and so on and so forth. Harnessed. That means updating the model system prompt and tool schemas, control flow, routing, retries, and so on and so forth. So I think from the error analysis that I actually shared earlier, I think that has really helped us to understand how can we improve our agent prompt as well as updating our knowledge base and tune the context management strategy that we have for our AI agent. The identified failure mode really helped us be able to feed those insights into actual improvement for the AI agents. So I want to quickly talk about what's next for us in our journey of building customer support AI agent at Lyft. We have some level of ability to run our offline simulator. But in fact, I think it's not repeatable. These are currently stored as scattered scripts across different notebooks and different analysis repo. I think one thing that we're really looking into investing in is a systematic eval harness and having a harness system that can help run our offline evaluation in a systematic and standardized manner and allow different people to contribute to our evaluation suite. With predefined primitive and conflict-based workflow. And another thing that we've been talking, thinking a lot about is post training. As we mentioned earlier, identify model failure mode has really helped us tune the agent context and as far as the agent harness. But over the years we have gathered a lot of real user signals on our agent performance as well. So we were really starting to think about how do we fine-tune a model that does different tasks for our customer support AI agent, as well as framing a reward modeling problem to enable reinforcement learning. I want to quickly share a little bit about the work that we're doing with eval harness and how we think about building eval harness for agent-facing, user-facing agentic applications. So again, eval is really that scaffolding that we need to be able to run eval efficiently in a standardized format across all the different agents and sub-agents that we have for customer support AI agents. And as you can see from my offline simulator slide earlier And another thing that we've been talking and thinking a lot about is post-training. As we mentioned earlier, identifying model failure mode has really helped us tune the agent context and the agent harness. But over the years, we have gathered a lot of real user signals on our agent performance as well. So we were really starting to think about how do we fine-tune a model that does different tasks for our customer support AI agent, as well as framing a reward modeling problem to enable reinforcement learning. I want to quickly share a little bit about the work that we're doing with eval harness and how we think about building eval harness for agent-facing, user-facing agentic applications. So again, eval is really that scaffolding that we need to be able to run eval efficiently in a standardized format across all the different agents and sub-agents that we have for customer support AI agents. And as you can see from my offline simulator slide earlier, our eval harness is conflict-driven, and these are typically stored as YAML file that's easily editable by different contributor, and not just by engineers. Analysts and data scientists can contribute to this evaluation suite as well. And with thousands, if not tens of thousands, of examples in our evaluation suite, we need parallelisms and throughputs to be able to run our offline evaluation in a reasonable amount of times. And to enable a user to be a different user to be able to contribute to our evaluation suite, we also define primitive around our eval harness. These are high-level things like tasks, data sets, personas, LM adapter, and evaluator. And after all, the benefit of having an eval harness is that you can define the config once and run these eval indefinitely many times across different touch lines or different gates of your agent development process. This can be locally when you're developing this agent. At any point when you tune a prompt, you can run the evaluation suite and get immediate feedback on how your agent is doing compared to the previous versions. We can run these at pre-commit hook to make sure that our performance doesn't degrade before we push a change to our agent service. Another area that we are looking at is also at CI/CD and how we can use our eval harness to build our regression test suite, acceptance test suite as well. So that wraps up our presentation today. We've gone through a lot of different topics for evaluations. Really, I think this is an end-to-end journey for building an evaluation pipeline that works for customer support AI agent or any user-facing agentic applications. Action and I are very interested to hear about what you all have been working on and share any learnings that you have for eval, so feel free to contact us if you have any questions or you just want to share a brainstorm about how to improve your evaluation system. All right, thank you, and I hope you all enjoy the AI engineer welfare. Thank you, thanks. you değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш değiş өзгериш өзгерიშ تغییر পরিবর্তন পরিবর্তন পরিবর্তন পরিবর্তন পরিবর্তন পরিবর্তন পরিবর্তন পরিবর্তন পরিবর্তন পরিবর্তন পরিবর্তন পরিবর্তন পরিবর্তন পরিবর্তন পরিবর্তন পরিবর্তন փոփոխন փոփոխন փոփոխন փոփոխন փոփոխন փոփոխन փոփոխन փոփոխन փոփոխन փոփոխন փոփոխন փոփոխन փոփոխन փոփոխन փոփոխन փոփոխन փոփոխन փոփոխन փոփոխन փոփոխਨոփոխന փոփոխनոփոխন 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 변경 to add some statistical rigor to it, right? We report alignment rates as bare point estimates. So if we add confidence intervals, we do some calibration and proper sampling, the same numbers can become more meaningful. We should definitely reserve the expensive rigor for the moment a number actually gates something or like a shipping decision, or we are reporting numbers to company leaders. But depending on the use case, we should definitely have confidence intervals for the numbers we report, because every score needs an interval. So this is just a small example to give you an insight on what this actually means. Let's say we have two evaluators and one scores 84% and the other scores 88%. And the number of samples, the number of traces which we have used is let's say 50. So to show that like this is a very small gain and we need much more than 50 examples to actually show that this gain is real. With just 50 examples and only 4% point percentage gains, we don't actually know if this gain is real or not. So bigger gains and bigger gains and paired designs need far less. And we can reserve the rigor, statistical rigor for things which matter the most. So this is a non-exhaustive list of observed eval antipatterns. I'm not going to read all of them, but these definitely contain some low-hanging fruits. We need to put in time and effort if we need meaningful evals. So we cannot rely on LLMs or everything yet. And what I can say is ignoring the data is one of the most important things which we shouldn't do and which we sometimes don't focus on due to lack of time or resourcing, but it acts as the foundation for meaningful evaluations. If you don't look at the data, you won't be able to create meaningful criteria or labels. And if you don't have labels, you won't be able to evaluate your judges. And if you are not evaluating your judges, you don't know if your agentic pipeline is working as expected. So this acts as a base. So one of the most important things we should not ignore. Okay. We said that we want to make metrics more actionable and standardize the pipeline, but how actually we should do it. So this gives you like a template to do an error analysis loop. And it is important to know that this loop is something which runs continuously. It's not in one-off audit. So we deep dive into raw traces. So once we have logged our traces for our agentic flows, multi-agent systems, or whatever we have in our use case, then we pinpoint failure modes. So basically we try to identify what exactly is failing. And we only keep the metrics that change a decision. We remove all the noise. We only prioritize on the metrics which are tied to business use cases, functional requirements, and actually something which changes a decision. Then we form a fresh premise to reevaluate. And then we repeat it. So we can have like a regular cadence of doing this pipeline. It can be weekly. It can be biweekly. But this is something which needs to run continuously. And it's not a one-off audit. Okay. So tracing. Tracing, as we said, is one of the most important things. And everything kind of depends on it. For diving deep into raw traces, we definitely need to log them first. So we can use tools like Lange Smith, Lange Fuse, etc. to log and view the traces. And here each trace captures the full graph execution. Which nodes ran? What LLM saw? Which tools were called? What was the token usage? What was the latency for every call? And things like that. We can also enrich traces with metadata if we want, which gives you more insights than the actual data. Okay. And we also have annotation queues. Annotation queues are nothing but an interface, which is very helpful for domain experts to label or give feedback to evaluators in an easy to understand UI. So they don't have to look at the raw traces, JSONs and stuff like that to figure out what to focus on. They can use this annotation queue and they can give feedback or label examples easily. We can then add these traces to data sets for offline evaluation. Or we can use this for calculating our precision and recall for our judges. So this forms the basis to validate the evaluation with ground truth labels. Okay. Now I'll hand it over to Nick to close the evaluation loop. Thank you, Akshay. Thank you, Akshay. And I think the goal of having Eval is to be able to feed our evaluation insights back into improving the model performance or the agent's performance. So here I'll introduce a couple of ways that we think about continual learning for our AI agent and closing the evaluation loop. And here you have model learning, context learning and harness learning. Model learning is really about post-training, updating the underlying model weights and training a custom LLM model. Context learning and harness learning is really more about improving everything else other than the model. So context is improving what information the agent actually sees. This can be document. This can be the store memories of the user, two outputs and so and so forth. Harnessed. That means updating the model system prompt. And two schemas, control flow, routing, retries and so and so forth. So I think from the error analysis that I actually shared earlier, I think that has really helped us to understand how can we improve our agent prompt as well as updating our knowledge base and tune the context management strategy that we have for our AI agent. The identify failure mode really helped us be able to feed that insights into actual improvement for the AI agents. So I want to quickly talk about what's next for us in our journey of building customer support AI agent at Lyft. We have a we have we have some some level of ability to run our offline simulator. But in fact, I think, you know, it's not repeatable. These are currently stored as scatter scripts across different notebooks and different analysis repo. I think one thing that we're looking really really looking into investing is a systematic eval harness and having a harness system that can help run our offline evaluation in a systematic and standardized manner and allow different people to contribute to our evaluation suite. We're predefined primitive and conflict based workflow. And another thing that we've been talking, thinking a lot about is post training. As we mentioned earlier, you know, identify model failure mode has really helped us tune the agent context and as far as the agent harness. but over the years we have gathered a lot of real user signals on our agent performance as well so we were really starting to think about how do we fine-tune a model that does different tasks for our customer support ai agent as well as framing a reward modeling problem to to enable reinforcement learning i want to quickly share a little bit about the work that we're doing with eval hardest and how we think about building building building eval harness for agent facing user facing agentic applications so again eval is really that scaffolding that you that we need to be able to run eval efficiently in a standardized format across all the different agents and sub-agents that we have for customer support ai agents and as you can see from my offline simulator slide earlier our eval harness is conflict-driven and these are typically stored as yaml file that's easily editable by different contributor and not just by engineers analysts and data scientists can can contribute to this evaluation suite as well and you know with thousands if not tens of thousands of examples in our evaluation suite we need you know parallelisms and throughputs to be able to run our offline evaluation in a reasonable amount of times and to enable a user to be a different user to be able to contribute to our evaluation suite we also define you know primitive around our eval either harness these are high level things like tasks data sets personas lm adapter and evaluator and after all you know the benefit of having an evil hardest is that you can define the config once and run these eval indefinitely many times across different at different touch lines or different gates of your agent development process this can be you know locally when you're developing this agent uh at any point when you tune a prompt you can run the evaluation suite and get an immediate in immediate feedback on how your agent is doing compared to the previous versions we can run these at pre-committ hook uh to make sure that our performance doesn't degrade before we push a change to our agent service another area that we are looking at is also at ci cd and how we can use our eval harness to build our regression test suite acceptance acceptance test suite as well so that wraps up our presentation today uh we've gone through a lot of we've gone through a lot of uh a lot of different topics uh for evaluations uh really i think this is an end-to-end journey uh for building a evaluation pipeline that works for customer support ai agent or any user-facing agentic applications uh action and i we are very interested to hear about you know what you all have been working on and share any learnings that you have for eval so feel free to contact us uh if you have any questions or you just want to share a brainstorm about how to improve your evaluation system all right thank you and i hope you all enjoy the ai engineer welfare thank you thanks you değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş değiş