Open Reader

The maturity phases of running evals — Phil Hetzel, Braintrust

completed 18:33 May 27, 2026 Watch on YouTube

Current Status

completed

Video ID

FB-MLPhL9Ms

RAG / Chat

Enabled
The maturity phases of running evals — Phil Hetzel, Braintrust
Description

Most teams approach evals like unit tests and try to cover every possible failure. Phil Hetzel from Braintrust argues that is the wrong frame: enumerate your known failure modes, cover those specifically, and ship. The goal is a flywheel where production traces surface what is going wrong, feed back into offline experimentation, and guide the next improvement. The session walks four maturity stages: vibe checking with documented human justifications not just thumbs up or down, LLM as judge built from those justifications at scale, then the hard part, tool calls that touch external systems. Context gathering tools are manageable. CRUD tools are not, because you have to represent the state of external systems at the exact moment the original trace ran. Timestamp queries against a vector database and injecting captured system state directly into the trace are two approaches for getting there. Speaker info: - https://www.linkedin.com/in/philliphetzel/

Summary

Generated by claude-sonnet-4-5

30-second take

Phil Hetzel (Braintrust solutions engineering lead) presents a four-stage maturity model for AI agent evaluations: starting with human annotation, scaling with LLM-as-judge, handling external tool complexity, and advanced automation. His core thesis is that evals should rerun production rather than exhaustively test every edge case, focusing on known failure modes extracted from domain experts. The talk is practical but high-level, aimed at teams still vibe-checking agents. Key insight: capture production traces + human justifications, then automate those judgments. Caveats matter: LLM-as-judge needs its own evals, and stateful systems (CRUD operations) create unsolved eval challenges around representing external system state offline.

Key takes

  • Evals defend against reputational/compliance/cost risk and enable offensive iteration: knowing which tweaks improve agents requires systematic measurement, not just preventing catastrophic failures.
  • Start with vibed human annotation + justification extraction: have domain experts thumbs-up/down outputs and explain why, then use those justifications (via Cursor/Claude) to derive actual failure modes for automation.
  • The "flywheel" is capture production → identify failures → rerun offline → guide improvements: evals should approximate production workloads, not exhaustively cover hypothetical edge cases like unit tests would.
  • LLM-as-judge must be evaluated itself: "putting a robe on an LLM doesn't make it trustworthy"—align judge outputs to human ground truth datasets before trusting them at scale.
  • Tool-calling agents create trace-level eval complexity: once agents do CRUD on external systems, you must evaluate entire multi-step traces and solve for offline state representation (timestamp queries, mock APIs, embedding system state in traces).
  • Deterministic scorers (code-based) are fine for objective failures: token count, tool call count, or other measurable metrics don't require LLM judges—use code when appropriate.

Useful details

  • Eval primitives: task (agent/prompt under test), dataset of example inputs, scoring functions (judge quality).
  • Four maturity stages: (1) Just getting started (human annotation), (2) Measuring to manage (LLM-as-judge + production traces), (3) Accounting for complexity (tool calls, external systems), (4) Advanced (topic modeling, automated CI).
  • Braintrust platform: lets users custom-code annotation views for domain-specific agent output review; flywheel concept is core to product philosophy.
  • Stateful eval challenges (unsolved): representing external system state when an eval input was created, and interacting with those systems offline without corrupting production data. Partial solutions: embed state in arbitrarily large traces, use timestamp-based vector DB queries.
  • Emerging patterns: topic modeling at scale to auto-discover production failure modes; quad-code / eval provider CLI for CI/CD automation.
  • Speaker background: 12 years consulting (KPMG, Slalom Databricks lead), saw customers struggle to productionize GenAI POCs, joined Braintrust a year ago after using it as a customer.

Caveats / counterpoints

  • Stateful system evals are "not completely solved": Phil admits CRUD-based tool agents remain hard to eval offline; mock APIs and trace-embedded state are workarounds, not solutions.
  • LLM-as-judge isn't 100% accurate: directional trends are acceptable, but teams must continuously eval the judge itself—no process shown for that.
  • Advanced techniques (topic modeling, CI) were glossed over: only 2 minutes left, so stage 4 wasn't explained; must visit booth for detail.
  • Audience was "way more advanced" than expected: most attendees were beyond stage 1, suggesting the maturity model may skew toward greenfield teams.
  • No discussion of cost/latency tradeoffs: running large-scale LLM-as-judge evals or rerunning production traces could be expensive; not addressed.

Ken relevance

High relevance for Ken's agent systems work. The flywheel concept (production traces → failure extraction → offline reruns) maps directly to how Ken should operationalize evals for any production agent he ships. The "justification extraction" step is a tactical way to scale domain expertise into automated scorers, useful for Ken's consulting/ops workflows. Stateful eval challenges are critical if Ken's agents do CRUD (e.g., calendar writes, database updates)—embedding system state in traces or using timestamp queries could unblock testing. LLM-as-judge needing its own evals is a non-obvious ops requirement Ken should plan for. The "evals ≠ unit tests" framing helps prioritize known failure modes over exhaustive coverage, saving time. If Ken's building agent tooling, Braintrust's approach (custom annotation views, trace introspection) offers product design patterns worth stealing or integrating.

Watch verdict

Skim. The maturity model is a useful mental framework and the flywheel + justification extraction tactics are actionable, but the talk is introductory and rushed. Stage 3/4 details were cut short, and the stateful eval problem (most relevant for complex agents) is acknowledged but unsolved. Read the transcript for the structure; visit Braintrust docs or Phil's booth content for depth. If Ken's already past stage 2, the marginal insight is low—better to watch advanced eval talks or read Braintrust's technical blog.

Transcript

2952 words en Processed in 360.6s

[SPEAKER_01] It's always a challenge to be a presenter directly after lunch because that's typical when the energy level goes from right around here to around here. But I'm going to try to make this session worth your while today. We've got 18 very quick minutes together. And during that time, I'm going to be talking about the different maturity levels that I see people go through as they perform evals for their agents. Before we get into that, just roughly a quick agenda today. I'll explain a little bit about myself, the company that I work for. We'll spend most of the time today on more theoretical concepts, not product concepts. And then we'll talk about where I think this field is going in the future. I'll also make sure to leave enough time, hopefully a couple minutes, for questions as well. I didn't over-prepare the content in hopes that we could have a little bit more of a discussion at the end of this. [SPEAKER_00] First of all, this is me. My name is Phil Hetzel. I lead solutions engineering for a company called Braintrust. Effectively, what that means is that it is me and my team's job to make sure that people are getting the most value out of the platform as quickly as possible. Prior to Braintrust, I spent 12 years in consulting and systems implementation. First four years with KPMG, last eight years in consulting with a company called Slalom Consulting. And with Slalom, I led their global Databricks business unit. And I noticed that a lot of my customers were prolific at creating generative AI proofs of concepts. They were not as prolific at bringing those proofs of concepts to production. So I started using Braintrust first as a user because I wanted to help bridge that gap for my customers. And I liked the product so much that I ended up joining the company and I've been here for about a year. Outside of work, I like to play chess, but I'm not very good at it. And I like to spend time with my wife and my dachshund, his name's Pistol Pete. He's the one in brown, not the one in black. What is Braintrust? The company that I work for. Braintrust is an agent quality company. One of the main ways that we contribute to agent quality are evals and observability, which we consider to be very much the same problem from a systems perspective. Evals, of course, being the thing that you're doing in order to gain confidence in your agent as you want to bring it to production. And then observability being the practice of once that agent is in production, remaining confident in it. It is a growing space, very fast moving space. And when you build an evals platform, you really have to grow with the technology, the underlying technology as it changes. So it's a very fun place to be in. Let me give a quick overview of the problem. We talked a little bit about why we do evals in the first place. How many of you are doing evals today? Hopefully as you build, every single hand should be up. And certainly when I give this talk next year at this conference, all you're going to come back, of course, to this session and every hand is going to be up. Evals are very important. The reason why we do evals is wholly in service to agent quality. That's the most important thing. We want to make sure that our agents are doing what we expect when confronted with real usage and real users. This is really important from a risk perspective and a brand perspective. We don't want the reputational risk of an agent being unkind or unhelpful to a customer. We don't want the systems risk of an agent costing us too much money as it operates. And there could even be compliance and legal risks if your agent goes too far off the rails. So evals are a defense against those types of risks, but they're also they can play offense with evals in knowing with each tweak that you make to your agent, how it's improving, and how much it's improving your application. A couple of primitives here. Evals are not unit tests. Whereas unit tests are very exhaustive in how you perform them, with evals you want to make sure that you start very high level with the failure modes of your agent. Either you or a subject matter expert can educate about the specific failure modes of an agent, and you build evals around those very specifically. What you don't do, like you would with unit tests, is think about exhaustively every single thing that could potentially go wrong with your agent and try to make an eval for it. Why can't we do that? Because it's infinite. You would spend all of your time writing tests and none of your time shipping, which is not productive. Eval results don't need to be perfect. Sometimes they can be, sometimes they can be directional. Using LLMs to judge other LLMs, LLM-as-judge techniques, you're probably not going to get 100% every time. That's okay. As long as you're trending in the right directions with those more non-deterministic techniques, that is completely fine. Different primitives with the eval itself, how it's constructed. You have three things. You have a task. That's the agent under test or the prompt under test. You have some data set of examples that initiate that task. How do you invoke that task? As judge techniques, you're probably not gonna get 100% every time. That's okay. As long as you're trending in the right directions with those more non-deterministic techniques, that is completely fine. Different primitives with the eval itself, how it's constructed. You have three things. You have a task. That's the agent under test or the prompt under test. You have some data set of examples that initiate that task. How do you invoke that task? You use some example that you give to an LLM or give to an agent to start that workflow. And then you have certain scoring functions, which you're using to judge the utility or the quality of that task. There are a couple of different maturity areas that I've noticed some of our customers go through. I've listed four here. This is probably more of a continuum than being very discrete. But suffice to say that these stages, you will traverse these stages as and when you create more complexity within your agent, just by necessity. The more complex agent you're building, the more vectors there are for failure, the more failure modes that you may need to account for. We're only gonna be focusing on the eval theory itself today. We're not gonna really talk about the platform surrounding evals. We've got a booth for that downstairs. If you're interested, you can come find me. So we'll go through these four. Just getting started, measuring to manage, accounting for complexity, and then some advanced eval techniques. Okay. Just getting started. It's not wrong to just get started with vibes. I know vibe checking is a very nasty phrase here at this conference. I actually think it's okay. It's certainly better than nothing. When you're first starting out, you can't help but start with vibes. I think the only thing that I would really recommend is that as you are vibe checking, you're also documenting. So when you have an agent under test, you give that agent maybe 10 different example inputs and loop through those inputs to see what the output is. You should probably have some human, whether it's the person who built the agent or even better, a subject matter expert that really knows what a quality response would look like. You should really have them analyze these outputs and give two pieces of information. You should give a thumbs up or thumbs down. Is it, was this response good? Was it bad? But more importantly, you should make that human annotator perform a justification for why they chose that thumbs up or thumbs down. Reason being is that you need to extract a lot of this domain specific knowledge out of that human annotator's head so that eventually you can scale that type of knowledge through a technique like LMS judge. But this is a great first step performing human annotation. Who in this room is at this step? We, okay, this is a way more advanced group. That's okay. That's a good place to start. Are you using human expert annotators? I just got, I haven't got any infrastructure set up. Yeah. Just, yeah, it's pure run an agent on my data. See, you gotta see how it looks. Yeah. Yeah, totally. It's, you have to start somewhere. It's a great place to start. This is how that workflow is gonna look. You have a trace come in, thumbs up or thumbs down, and then you add some justification to that so that eventually you can use it as an LMS judge score down the line. This is what this might look like in a platform like Braintrust. We have a human annotator view built into the platform. We actually let you build code your own annotation views. Important point. Don't give a generic annotation platform to users. Really make it very specific to them. They're gonna have an idea of how these agent traces should look. So you should deliver that to them and it'll encourage them to evaluate these appropriately. Okay. The next part is expanding upon that a bit where now I don't have only some human grader giving thumbs up and thumbs down and justification. Now I'm starting to use those justifications and I'm probably running those justifications through cursor or cloud code or codex to try to derive the actual failure modes of why, when they gave a thumbs down, why they delivered a thumbs down. These, you now know and understand the failure modes of your agent. Now that you understand the failure modes of your agent, you want to be able to scale that human knowledge and be able to automate it so that you're not dependent on a few people with expertise to judge agent outputs. A couple ways to do this. One of which using LLMs to judge other LLMs. LLMs judge, that concept's been around for quite some time. Very effective. Important here is that whenever you use an LLM judge, just because you put a robe and a cloak on an LLM, that doesn't make it inherently more trustworthy. You should be evaluating LLMs as judge outputs as well. That's not really covered in this presentation, but you should not just judge LLM judges blindly in that regard. There also might be some objective failure modes where you can deterministically encounter them just through code. That's okay too. You don't have to use LLMs to judge other LLMs. You can use code to understand if you're using too many tool calls, you might wanna fail that eval as an example. If you're using too many tokens, you might wanna fail that eval. You should be evaluating LLMs as judge outputs as well. That's not really covered in this presentation, but you should not just judge LLM judges blindly in that regard. There also might be some objective failure modes where you can deterministically encounter them just through code. That's okay too. You don't have to use LLMs to judge other LLMs. You can use code to understand if you're using too many tool calls, you might want to fail that eval as an example. If you're using too many tokens, you might want to fail that eval. I think the most important point here is that this dataset that's on the right-hand side of this slide at this point you should probably be gathering production traces or at least UAT level traces into that evaluation dataset. We want it to be very—don't think about evals as running tests. [SPEAKER_01] Think about evals like rerunning production because ultimately we want to be confident as we run these workloads in production. Great way to do that is just to capture production data. Most important point is we call it the flywheel internally. We want to be able to capture these traces, these agent traces in production, understand what's going wrong with them, either through a human or through automated tooling. And then bring those examples back to some offline experimentation environment, rerun production through an eval, and then use that to guide us to which direction we should be improving our agent. So evals—that's more playing offense with your evals. It's an example of setting up an LLM as a judge scoring function to expand your abilities. So evaluate at scale rather than using just a human. Okay, level two. Now we're starting to not just do simple model calls. We might be performing work with external systems. I think of tool calls in two different ways. There is context gathering tools that are gathering data and injecting that into the LLM. And then there is CRUD based tools where you're creating, reading, updating, or deleting information from a database or an external system. Both of these can have a lot of lift in terms of whether your agent is quality or not. It also means that there are a lot of other things that can go wrong with your agent when you're starting to interact with external systems. Often now instead of just evaluating one specific part—the output of an agent—now you might be having to evaluate the entire trace of an agent. So in that sense, this is where tooling starts to come into play. You'll need some way to capture these large traces, understand each and every step that an agent took, to be able to introspect and eventually target evals towards maybe even individual tool or MCP calls that your agent is creating. The other problem here that we might have is when you're performing CRUD on a system, you don't really want to do that when you're offline, of course. There might not be a way to do that when you're offline. So when you run an eval, there are two things that are problem areas. One, it's really challenging to represent the state that other external systems were in at the time that eval input was created. And then two, it makes it really challenging to interact with those systems that the agent could be interacting with because you don't want to overwrite any production data. These are real challenges that we have to solve for. I would say it's not completely solved right now. However, there are some ways where you can represent external system state and interact with mock level APIs so that you can approximate a real production environment as you're running evals. The idea for this is that these agent traces can be arbitrarily large. In that sense, it's a lot different than application tracing. So if a trace can be arbitrarily large, you can actually cram in a ton of context—system state, the state that the external systems were in at the time—into these traces and inject that into the task that you're running the eval upon. In that way, instead of having to create entire test structures and infrastructure, you can represent a lot of that stuff within the trace itself and encapsulate it there. The other thing that you can do is you can use really specific querying techniques to perform timestamp queries to systems that support them. So if an input came in and you added it to your dataset at a certain point in time, perhaps the way that you've set up your vector database, you can run a version query to query the vector database at a certain point in time. [SPEAKER_00] So that way you're adequately representing the state when that task ran originally. These are more complex techniques, but ones that are a bit more emerging. I only have about two minutes left to go. What's next? Performing topic modeling at scale to make sure that you're uncovering those failure modes automatically in production. That's something that I'm more than happy to talk about at the booth downstairs. And then, of course, performing evals in a way where you're using quad code and the eval provider CLI to be able to do this in an automated way. These are two other patterns that I see emerging in this space. I want to be conscious of time. I probably have time for one question before I have to jump here. Is anyone curious about anything specifically? Otherwise, you can find me at the booth. Yes, sir? In our sphere, it's normal to put a bit more respect on deterministic evaluation of deterministic readers. Do you agree with it? Do you think that we should push for more deterministic readers in this, on platforms? Or do we embrace LLM as judge? Some things are subjective. That's why we love agents so much. I would embrace LLM as judge, but also perform a lot of evals on the LLM as judge so that it's very aligned with what a human would decide in the same circumstance. You would eval the eval. [SPEAKER_00] Yeah. It's easier to do that because LLM judge outputs are going to be discrete. [SPEAKER_00] So you can create a ground truth dataset for that. Yeah. [SPEAKER_00] All right, everyone. I have to jump. I'm at my time. [SPEAKER_00] It was a pleasure to be with you all. [SPEAKER_00] Yeah, feel free to find me in the booth downstairs. And, and certainly when I give this talk next year at this conference, all you're gonna come back, of course, to this session and every hand is gonna be up. Evals are very important. The reason why we do evals is wholly in service to agent quality. That's the most important thing. We wanna make sure that our agents are doing what we expect when confronted with real usage and, and, and real users. Um, this is really important from a risk perspective and, and a brand perspective. We don't want, um, the reputational risk of an agent being unkind or unhelpful to a customer. We don't want the systems risk of an agent costing us too much money as it, as it operates. Um, and there could even be compliance and legal risks if your agent goes too far off the rails. So evals are a de, both a defense against those types of risks, but they're also, uh, they can play offense with evals in knowing with each tweak that you make to your agent, how it's improving, and how much it's improving, improving your application. Um, a couple of primitives here. Evals are not unit tests, where, whereas unit tests are very exhaustive in, in how you perform them. With evals, you wanna make sure that you start very high level with the failure modes of your agent. Either, either you or a subject matter expert can educate about the specific failure modes of an agent, and you build evals around those very specifically. What you don't do, like you would with unit tests, is think about exhaustively every single thing that could potentially go wrong with your agent and try to make an eval for it. Why can't we do that? Because it's, it's infinite. You would spend all of your time writing tests and none of your time shipping, which is, which is not productive. Uh, eval results don't need to be perfect. Sometimes they can be, sometimes they can be directional. Um, using LLMs to judge other LLMs, LLMs, as judge techniques, you're probably not gonna get 100% every time. That's okay. As, as long as you're trending in the right directions with those more non-deterministic techniques, um, that, that, that is completely fine. Um, different primitives with the eval itself, how it's constructed. You have three things. You have a task. That's the agent under test or the prompt under test. You have some data set of examples that initiate that task. How do you invoke that task? You use some example that you give to an LLM or give to an agent to, um, to start that workflow. And then you have certain scoring functions, which you're using to judge the utility or the quality of that task. Um, there are a couple of different maturity, uh, areas that I've noticed some of our customers go through. Um, I've listed four here. This is probably more of a continuum than, than being very discreet. But suffice to say that these stages, um, uh, you know, uh, you will, you will traverse these stages as and when you create more complexity within your agent, just by, just by necessity. Uh, the more complex, uh, agent you're, uh, that you're building, the more vectors there are for failure, the more failure modes you, that you may need to account for. Um, we're only gonna be focusing on the, uh, like eval theory itself today. We're not gonna really talk about the platform surrounding evals. Uh, we've got a booth for that downstairs. Uh, if, if you're interested, you can come find me. So we'll go through these four. Just getting started, uh, measuring to manage, accounting for complexity, and then, um, some advanced eval techniques. Okay. Just getting started. It's not wrong to just get started with, with vibes. I know like, like vibe checking is a, a very, uh, uh, nasty phrase here at this conference. I actually think it's okay. It's, it's certainly better than nothing. Um, when you're first starting out, you can't help but start with vibes. I think the only thing that I would really recommend is that as you are vibe checking, you're also documenting. So when you have an agent under test, you give that agent maybe 10 different example inputs and, and, and loop through those inputs to see what the output is. You should probably have some human, whether it's the person who built the agent or even better, a subject matter expert that really knows what a quality response would look like. You should really have them analyze these outputs and, and, uh, give two pieces of information. You should give a thumbs up or thumbs down. Is it, was this response, uh, good? Was it bad? But more importantly, you should make that human annotator, um, perform a justification for why they chose that thumbs up or thumbs down. Reason being is that you, you're, you need to extract a lot of this domain specific knowledge out of that human annotators head so that eventually you can scale that type of knowledge through a, through a technique like, like LMS judge. But this is a great first step performing human, human annotation. Who, who in this room is like at, at, at this step? We're, okay, this is way more advanced group. That's a, that's okay. That's a good place to start. Are you using like human expert, uh, annotators? I, I just got like a, I mean, I, I haven't got any infrastructure set up. Yeah. Just, uh, yeah, it's pure run an agent on my data. See, You gotta see how it looks. Yeah. Yeah, totally. It's, it's, it's, it's, you have to start somewhere. It's a great place to start. Um, uh, this is, this is how like that, that workflow is gonna, is gonna look. You have a trace come in, thumb up, thumbs up or thumbs down, and then you add some justification, um, to that so that eventually you can use it as, as an LMS judge score down the line. Um, this is like what this might look like in a, in a platform like Braintrust. Um, uh, we have like a, like a human annotator view, uh, built into the platform. We actually let you vibe code your own annotation views. Um, important point. Don't give a generic, uh, annotation platform to users. Really make it very specific to them. They're gonna have an idea of how these agent traces should look. So you should deliver that to them and it'll encourage them to, um, evaluate these, uh, these appropriately. Um, okay. The, the next part is expanding upon that a bit where now I just don't, I don't have only some human grader giving thumbs up and, and thumbs down and justification. Now I'm starting to use those justifications and I'm, I'm probably, um, running those justifications through, uh, uh, cursor or cloud code or, or codex to try to derive the actual failure modes of why, when they gave a thumbs down, why they delivered a thumbs down. Those, these, you, you're, you now know and understand the failure modes of your agent. Um, now that you understand the failure modes of your agent, you want to be able to scale that human knowledge and be able to, to automate it so that you're not dependent on, uh, a few people with expertise to judge, uh, agent, uh, agent outputs. Uh, a couple ways to, a couple ways to do this. One of which using LLMs to judge, uh, other LLMs. LLMs judge, we, we've, that, that concept's been around for, for quite some time. Very effective. Um, important here is that whenever you, uh, use an LLM judge just because you put a robe and a cloak on an LLM, that doesn't make it inherently more trustworthy. You should be evaluating LLMs as judge outputs as well. Um, that's not really covered in this presentation, but, um, you should not just judge, uh, LLM judges blindly in that regard. Uh, there also might be some objective failure modes where you can deterministically, um, encounter them just through code. That's okay too. You don't have to use LLMs to judge other LLMs. You can use code to understand, um, if you're using too many tool calls, you might wanna fail that eval as an example. If you're using too many tokens, you might wanna fail that eval, uh, eval. Um, I think the most important point here is that this dataset, uh, that's, that's on, on the right-hand side of this slide, at this point you should probably be gathering production traces or at least, uh, UAT level traces into that evaluation dataset. We want it to be very- like, don't think about evals as running tests. Think about evals like rerunning production because ultimately we want to be confident as we run, uh, run these workloads in, in production. Great way to do that is just to capture production data. Um, most important point is, is this, uh, we, we call it like the, the flywheel internally. We want to be able to capture these traces, these agent traces in production, understand what's going wrong with them, either through a human or, or through automated tooling. Um, and then bring those examples back to some offline experimentation environment, rerun production through an eval, and then use that to guide us to which direction we should be improving our agent. So evals, that, that's like more playing offense with, with your evals. Um, it's just like a, an example of, uh, setting up and, setting up an LOM as, as, uh, as judge scoring function to expand, um, your, your abilities. So evaluate at scale rather than using just a human. Um, okay, uh, level two. Now we're starting to not just do simple model calls. We might be performing work with external systems. I think of, uh, tool calls in two different ways. There is context gathering tools that are just, uh, gathering data and injecting that into the LOM. And then there is CRUD based tools where you're creating, reading, updating, or deleting information, information from a database or an external system. Um, both of these are, uh, can have a lot of lift in terms of whether your agent is quality or not. It also means that there is a lot of other things that can go wrong with your agent when you're starting to interact with external, uh, external systems. Often now instead of just having one, uh, um, uh, of evaluating one specific part, i.e the output of an agent, now you might be having to evaluate the entire trace of an agent. So in that sense, this is where tooling starts to come into play. You'll need some way to capture these large traces, understand each and every step that an agent took to be able to introspect and eventually target evals towards maybe even individual tool or MCP calls that your agent is creating. The other problem here that we might have is when you're performing CRUD on a system, you don't really want to do that when you're offline, of course. There might not be a way to do that when you're offline. So when you run an eval, there's two things that are problem areas. One, really challenging to represent the state that other external systems were in at the time that eval input was created. And then two, it makes it really challenging to interact with those systems that the agent could be interacting with because you don't want to overwrite any production data. These are real challenges that we have to solve for. I would say it's not completely solved right now. However, there does need- there are some ways where you can represent external system state and interact with like mock level APIs so that you can approximate a real production environment in- as you're running evals. The idea for this is that these- these agent traces can be arbitrarily large. In that sense, it's a lot different than application tracing. So if a trace can be arbitrarily large, you can actually cram in a ton of context, i.e. system state- the state that the external systems were in at the time- into these traces and inject that into- into the task that you're running the eval upon. In that way, instead of having to create an entire test structures and- and infrastructure, you can represent a lot of that stuff within the trace itself and encapsulate it there. The other thing that you can do is you can use really- really specific querying techniques to perform timestamp queries to systems that support them. So if an input came and- and you added it to your data set at a certain point in time, perhaps the way that you've set up your vector database, you can run a version query to query the vector database at a certain point in time. So that way you're adequately representing the state of- of- of when that task ran originally. These are more complex techniques, um, but ones that- ones that are- that are a little bit more emerging. Um, I only have about two minutes left to go. Um, uh, what's next? Performing topic modeling at scale to make sure that you're uncovering those failure modes automatically in production. Um, that's something that, like, more than happy to talk about at- uh, at the booth downstairs. And then, of course, performing evals in a way where you're, um, using quad code and the eval provider CLI to be able to do this in an automated- automated way. These are two other patterns that- that I see emerging in this space. I want to be conscious of time. I probably have time for, like, one question, uh, before- before I have to jump here. Is anyone curious about anything specifically? Otherwise, uh, you can find me at- at the booth. Yes, sir? In our sphere, like, it's kind of normal to put a bit more respect on deterministic evaluation- Yeah. ...of deterministic readers. Do you agree with it? Do you think that we should push for more deterministic readers in this, you know, about platforms? Or do we embrace LM as judge? I would- some things are subjective. That's why we love agents so much. I would embrace LM as judge, but also perform a lot of evals on the LM as judge so that, like, it's- it's very aligned with what a human would decide in the same circumstance. You would eval the eval. Yeah. It's easier to do that because LM judge outputs are gonna- are going to be, uh, discrete. So you can create a ground truth dataset for that. Yeah. All right, everyone. I have to jump. I'm at my time. Um, it was a pleasure to be with you all. And, yeah, feel free to find me in the booth downstairs. ...