Open Reader

How agent o11y differs from traditional o11y — Phil Hetzel, Braintrust

completed 20:43 May 28, 2026 Watch on YouTube

Current Status

completed

Video ID

XBaznoTRDFI

RAG / Chat

Enabled
How agent o11y differs from traditional o11y — Phil Hetzel, Braintrust
Description

Traditional observability answers one question: is the system up? Phil Hetzel from Braintrust argues that question is not the right one for agents. An individual agent trace can exceed a gigabyte. A single span can hit 20 megabytes. The data is semistructured, packed with unstructured text, and still arrives in real time. None of the systems built for uptime monitoring were designed to ingest, index, and actually use that. Braintrust built a custom database from scratch for this problem: a write ahead log for instant visibility, analytical indexes for fast filtering, and a forked version of Tantivy (a Rust based full text search library similar to Apache Lucene) so an engineer can query every trace that mentioned a specific word. The other difference is who does this work: clinicians, lawyers, and wealth advisers now open traces directly to grade whether an agent responded correctly, and their written justifications become the training signal for automated scoring functions. The human annotations surface the failure modes. The scoring functions scale them. Speaker info: - https://www.linkedin.com/in/philliphetzel/

Summary

Generated by claude-sonnet-4-5

30-second take

Phil Hetzel from Braintrust argues that agent observability is fundamentally different from traditional observability—not just a subset of the same problem. Traditional o11y focuses on uptime and technical metrics (latency, errors) for deterministic applications; agent o11y must handle non-deterministic behavior, massive semi-structured traces (up to 1GB per trace, 20MB per span), qualitative evaluation (grounding, tool use, brand alignment), and serve both technical and non-technical users (clinicians, lawyers, wealth advisors). Braintrust built a custom database from scratch to handle these requirements, treating observability and evals as the same system (batch vs. real-time). This talk is valuable for understanding the systems-level constraints of production AI agents, not for product features.

Key takes

  • Agent traces are a different data problem: Individual traces can exceed 1GB, spans can hit 20MB, and they're semi-structured with massive unstructured text. This requires custom indexing (Tantivy for full-text search across traces), write-ahead logs for instant reads, and SQL-like query interfaces—far beyond what ClickHouse or traditional OLAP tools handle well.
  • Non-determinism demands qualitative metrics: Unlike deterministic apps with known code paths, agents require measuring grounding in context, expected tool use, brand alignment, and user intent—not just latency and errors. Traditional o11y tools can't compute these because they don't ingest the voluminous trace data needed.
  • Non-technical experts must be in the loop: The best agent teams include clinicians, lawyers, and wealth advisors who review traces, annotate failures, and write prompts in natural language. This human annotation feeds into automated scoring functions and failure mode detection—a workflow that doesn't exist in traditional o11y, which serves only systems engineers.
  • Observability and evals are the same system: The only difference is batch (known inputs, evals) vs. real-time (unknown inputs, observability). Braintrust unifies both workflows in one platform, letting users add production traces to offline datasets for experimentation.
  • Auto-insights from traces are emerging: Braintrust recently rolled out LLM-based embedding and clustering on production traces to auto-detect user intent, sentiment, and issue patterns. This shortens the iteration loop from production problem to experimental fix.

Useful details

  • Phil spent 12 years in consulting, led Databricks practice at Slalom, saw customers build many AI POCs but struggle to ship them to production. He joined Braintrust a year ago.
  • Braintrust uses Datadog internally for traditional 400/500-level errors and uptime monitoring—agent o11y doesn't replace traditional o11y for infrastructure.
  • Tantivy is a Rust-based text indexing framework (similar to Apache Lucene) enabling queries like "show me every trace containing 'Amazon.'"
  • Braintrust's founder was an early employee at SingleStore and built a custom database for agent traces after finding ClickHouse inadequate for text-based indexing.
  • Real customer traces: Some customers have individual spans at 20MB, full traces over 1GB.
  • Read patterns must support both instant UI updates for engineers and CLI/SQL batch queries for automation.
  • The iceberg slide referenced human annotation as a key "below the waterline" activity: experts grade traces, justify their scores, and those justifications become scalable scoring functions via LLM-based automation.

Caveats / counterpoints

  • This talk is explicitly "theoretical, not product forward," so Phil avoids detailed Braintrust feature demos. The insights are conceptual, not operational.
  • No hard data on query speed, cost per trace, or ROI from agent o11y vs. traditional o11y—just assertions that users want it "faster."
  • Phil doesn't address how to handle agent traces in resource-constrained environments or what happens if you don't want to build/buy a custom database.
  • The claim that non-technical users (clinicians, lawyers) participate in trace review is based on Braintrust customer anecdotes, not quantified adoption rates or retention metrics.
  • No discussion of privacy/compliance risks when storing gigabyte-scale traces containing user interactions, especially in regulated domains (healthcare, finance).

Ken relevance

High relevance for agent infrastructure decisions. If you're building or investing in agent systems, this clarifies why off-the-shelf observability tools (Datadog, Grafana) won't suffice for production AI agents—you need specialized storage, indexing, and UX for non-engineers. This could inform:

  • Ops strategy: You may need Braintrust or a similar platform if your agents ship to users; Datadog alone won't catch hallucinations, tool misuse, or brand misalignment.
  • Investment lens: Agent o11y is a distinct market from traditional o11y. Companies solving this (Braintrust, competitors) face hard systems problems (custom DBs, text indexing, LLM-based auto-insights) but also have moats.
  • Product/GTM: If you're selling AI agents to enterprises, expect them to ask how they'll monitor quality and iterate. The insight that evals and observability are the same system (batch vs. real-time) is a useful framing for your own agent products.
  • Content opportunity: This is a strong "category creation" narrative—could be a blog post or whitepaper on "Why agent observability is not traditional observability."

Watch verdict

Skim. The transcript captures the core argument (3 problems: non-determinism, nasty traces, multi-persona users) and key details (Tantivy, custom DB, human annotation). The slides Phil references (the iceberg, the trace screenshots) would add clarity, but the verbal explanation is sufficient for understanding. Watch only if you want to see the trace UI or hear the Q&A nuances live.

Transcript

3189 words en Processed in 297.2s

Thanks for joining me today towards the end of the day here. So I hope everyone has enough energy left for maybe what is your most excited topic of the day. Remains to be seen how traditional observability differs from agent observability. Quick agenda, I'll do a quick intro about myself and the company that I work for. This is not going to be a very product forward talk. It's going to be more theoretical, so I won't drown you in sales slides, I promise. And then we'll get into how these two ideas differ and also talk about what's next in the space. My name is Phil Hetzel. I lead solutions engineering for BrainTrust. What that means effectively is that me and my team are the folks that are charged with making sure that our customers are getting the most value out of the platform as quickly as possible. Prior to BrainTrust, I spent 12 years in consulting and systems implementation. I led the global Databricks practice for Slalom Consulting before I came here, and I noticed that a lot of my customers were prolific at creating generative AI proofs of concepts, but not nearly as good at bringing those proofs of concepts to production. So I started using BrainTrust as a user first, and I really liked it, and I applied for a job, and I've been here for about a year. I like to play chess outside of work. I like to spend time with my wife and Dachshund. That's Pistol Pete right there. That's my dog. He's the person in brown and not black. Those of you who have been to my sessions before didn't laugh at that joke because you've heard it at least once already. What is BrainTrust? BrainTrust is an agent quality platform. We mainly look at agent quality in two different ways. Is your agent performing as well as you thought it would when it's in production, i.e., can you remain confident in your agent? And then on the other side of that is as you're experimenting with new versions of your agent, do you feel like you can become confident as you tweak it and change it over time? Those are a couple of things that BrainTrust does, obviously relevant to today's discussion because agent observability is a massive part of what we do. Anyone has heard of BrainTrust? Show of hands. Before this week, did you hear of BrainTrust? Yeah? Okay. A couple folks. Well, welcome back for the folks that have heard of this before. I'm going to go pretty quickly through the slides. Hopefully you have enough time for questions as well. I don't have a ton of content. Traditional observability is established. So even when folks come to us, they'll say, well, we already have open source tools like Grafana, as an example, why wouldn't this be the same problem that we're solving with perhaps either an implementation or a contract that we already have? It's very established and we know that these applications can operate at scale. So the case that I'll be making is that the scope of traditional observability is actually quite different from the scope of agent observability and I'll explain why. Scope of traditional observability is all about uptime and technical performance. Is the application up and is the application giving a user experience from a technical lens that we would expect? So latency, duration of interactions, 400 and 500 level errors. These are all things that we're measuring with very established tools like Grafana, like Datadog. I will even say that at BrainTrust, though we are an agent observability platform, we're happy users of Datadog. It's great for this specific type of use case for us to understand if people are running into 500 or 400 level errors on our website as an example. Is the system operational? Are we up or are we down? That's what traditional observability is. The building blocks of this are a couple different things. Metrics—these are the things that you're measuring. I gave a couple examples before, but latency is the most obvious one. Error count is another. The things that you can aggregate and measure over time. And then traces and spans. Is everyone here familiar with observability? Does everyone know what a trace is? Okay. I don't want to take that for granted. A trace is just a full interaction of some workflow. And a span is just one step within that interaction. All of these things would apply to agent observability as well. So we have the same building blocks. Problem one for why agent observability is different. Agents are non-deterministic whereas applications are deterministic. The reason why we love LLMs so much is because they have high variety. They can do a lot of different things. They are abstracted. So because of that, while typical applications have very deterministic code paths and it's on purpose that they do that, where they're performing some type of known control flow, agent applications are very non-deterministic. We're curious about why an agent might take one path versus the other. This also means that traditional observability is going to focus on very constrained and known metrics. Whereas agent observability needs to be a little bit broader in terms of the things that it needs to measure. This is just an example of that. So at the bottom, let's start there. Agent observability can measure some of these more traditional metrics, albeit with more of an AI flavor. Time to first token, total tokens, duration, latency. These are all things that you would think be very traditional observability level metrics. But also you might want to understand more qualitative things about your application. So it's not just how long did I take to start responding to my user, which is more traditional observability. I want to know, was the information that I gave grounded in the context that I gathered with my application? Did I use the tools that I would have expected as I was reasoning towards my response? Is a response aligned to the brand standard that I set for this agent in the system prompt? These are all things that are not really able to be tested by traditional observability tools. Because if you think about it, the trace—the information in the trace that's necessary for us to compute these things up at the top is far larger than the volume that a traditional observability trace would handle. That goes to the next point here. Agent traces are really nasty in a variety of different ways. They're nasty because they're highly semi-structured. Even within those semi-structured, there's a ton of unstructured text data that we need to chew through. They're voluminous. So they can be an agent trace could be over a gigabyte in size. We've seen that even with our own customers. An individual span can be 20 megabytes in size. So it's just a far different systems problem that you have to solve in order to ingest, process, and most importantly use that type of data. And also it's just as fast as traditional observability data. So hopefully your agent that you're putting in production gets product market fit and you have a ton of users and usage associated with it. You as the AI engineer or as the product manager for that agent, you're going to want to see that observability in real time, in true real time. Trust me, we know that's the case because we always get the feedback. even with our own customers. An individual span can be 20 megabytes in size. So it's just a far different systems problem that you have to solve in order to ingest, process, and most importantly use that type of data. And also it's just as fast as traditional observability data. So hopefully your agent that you're putting in production gets product market fit and you have a ton of users and usage associated with it. You as the AI engineer or as the product manager for that agent, you're going to want to see that observability in real time, in true real time. Trust me, we know that's the case because we always get the feedback. We're talking about, can you just make it faster? We're always trying to make it faster. People always want it to be faster. Tough to do when the agent traces look like this. This is just an example of an agent trace in BrainTrust, where not only does it have a bunch of spans here encompassing the model calls and tool calls, but even within those spans, you saw the amount of unstructured text that's in there as well. Very different problem to solve. A little bit more here. Maybe I'll just dive into the read pattern piece specifically. We need to do two things simultaneously. We need to be able to perform the very fast read, ingest and read style workflows that are common with observability, i.e. if someone does an action with my agent, I need to be able to see that interaction instantaneously. We also have to commit to read patterns where someone wants to use our CLI and fire off SQL commands to us so that they can incorporate either observability or eval traces to improve their application automatically. There are just a lot of different mediums that people use now in order to query these very large trace shapes. This is a completely new systems problem. At least at BrainTrust, we designed a database from the ground up specifically for agent traces. I'm not going to really go into depth about this. We have a blog on our website. I think it was the last blog that we published if you're really interested in diving deep. But just very quickly, there are a lot of different components that we have to build into this database in order to make it work. For example, we need to immediately get data into a write-ahead log so that people can instantly see these traces as soon as they expect. We need to be able to perform indexing on these data so that whenever someone is performing a filter or analytical query that it's fast. We have this thing called a Tantivy index. Tantivy is an open source framework that we've worked. Anyone know what Tantivy does? Any guesses? Tantivy is how we perform text-style indexing. If you remember when I was showing this trace, it makes so much sense for someone to want to perform the workflow of, okay, I just want to know every trace that had the word Amazon into it. Well, it turns out it's really hard to do that unless you perform a full text-based index across your traces. That's another reason why agent observability is far different than traditional observability. You really don't have to think about the text problems in traditional observability. Is it the same as OpenSearch? Sorry? It's an OpenSearch? Tantivy is most similar to Apache Lucene, except it's written in Rust. Yeah. And then all of these things come together and have to be unified through a SQL or SQL-similar language. That's what we've gone to at BrainTrust. Problem three, this is a whereas there's a very specific type of persona for traditional observability. It's a systems engineer. Maybe it's a product engineer. It's probably not a subject matter expert, or if it's a medical application, it's not a clinician or a registered nurse. It's very technical people that align with traditional observability. That could not be further from the truth for agent observability if you're doing it well. We noticed that the best teams that are building agents have both technical and non-technical people in the fold performing this work because it's the non-technical people that are either A, closest to the users, or B, have knowledge that is closest to the problem space. And what can they do now with prompts? They can write it in natural language. So they can add real value into being able to participate in agents. We have folks that are clinicians or registered nurses or wealth advisors or lawyers. We have seen them operate in our platform looking through traces and using that information to improve their agents. That is a workflow that you simply don't see in traditional observability where you're more worried about uptime. I think in general, people don't realize that in order to perform observability and evals well, we think of observability and evals as the same problem. The only difference between evals is that you're running them in batch and you know the inputs ahead of time. It's incredible. The depth that you end up going into when you create a platform like this looks like this because of the reasons that I've described, the nuances with the data, and the amount and types of people that you have to bring into the fold. Those are some of the reasons why it's so different to perform in this space. Where is the space going? We've done a lot of work in this area. I think the natural question that we used to get asked was, if you're, if BrainTrust are collecting all of our agent traces and all of our agent traces have all of these valuable data in them, can't you just tell me how people are using my agent? It's the simple questions that usually need the most complex systems behind them. This is something that we are starting to do. We just rolled it out, I think about a month ago in our software as a service offering where we see agent observability traces come in and then we'll run a very lightweight LLM on top of them to perform embedding and then clustering on those traces to see how we can perform elevate topic modeling to see, for example, how people are using your traces, their intent, how people are feeling about interacting with your agent, the sentiment, or if they're running into issues, what those issues potentially are. The whole idea is there is that you can make the iteration loop between a problem that you're seeing in production and the fix that you perform experimentation on. The whole idea is to just make that faster and a little bit more direct. I promise I would go through that really fast. I've got about three minutes for questions if there are any, and I'd be really happy to answer them. Anyone curious about this? Yes. So I'll bring us this clearly about the functional observability of agents. Would you say is that also good for non-functional agent performance or is traditional observability good for that? That's a good question. Yeah. I would, well, I think traditional observability can do that. seeing in production and the fix that you perform experimentation on. The whole idea is to just make that faster and a little bit more direct. I promise I would go through that really fast. I've got about three minutes for questions if there are any, and I'd be really happy to answer them. Anyone curious about this? Yes. So I'll bring us this clearly about the functional observability of agents. Would you say— Is that also good for non-functional agent performance or is traditional observability good for that? That's a good question. Yeah. I would, well, I think traditional observability can do that. Braintrust specifically does do. I like the way that you put that functional observability, what's the quality of my agent, how I've defined it, and then the technical observability just comes on the house. Like when you trace the application, you automatically get prompts, duration, time to first token, et cetera, cache hits, et cetera. Yeah. In your iceberg slide, you have human annotation. Yeah. You have the water line there. Yeah. Could you explain what the human annotation part is and get this cloud on? Yeah. So let's think about it this way. Actually, if I can, I'll go on a high wire act here and just show and not tell. So let's say that you have a trace come in and you want your product manager to be able to opine on whether that agent did a good job or a bad job. It's really valuable for you to have an expert come in, grade the agents, but then also justify why they're grading the agents the way that they are. Because eventually you're going to take those justifications. You're going to probably run an LLM over it and you're going to make more scalable scoring functions from those justifications. You're finding the failure modes that you can then implement and automated scores through that. Yeah. Human annotation is a really key part of this process. Okay. Thanks. Uh, yes. In the second row. Yeah. I just have a question, because you just focus on the ability today, but I'm interested actually how you also integrate to the other agent framework to for the offline optimization. Mm-hmm. And then also, I guess the main difference I feel that I see in this, your database is that we're not using OLAP or Quickhouse. We used to use Quickhouse actually. Yeah. We moved away from it. Yeah. I just wanted to, curious to like, why build your own, like, what is the efficiency? Well, the funny answer there is that our founder is kind of an insane person. Like, only an insane person would build their own database, but he is cut from that cloth. He was one of the first employees at SingleStore, so he's kind of used to doing that. But what he found was when he was performing some of these workloads, he just needed more of the text-based indexes, which Quickhouse wasn't really able to do, at least at that time. So we built our own. And then the first part of your question, observability and evals to us, it's like we solve it with the same system. The only difference is that with evals, we know the inputs ahead of time with obs, and we're doing it in batch. With observability, we don't know what the inputs are ahead of time and we're doing them in real time. Oh, right. But I mean the experiment functionality, like, how easy is it to integrate with, like? Oh yeah, it should be pretty easy. Like, when once you've traced, and I apologize and I'm making this about the product. But when you have a trace come in, you've traced it, and then you just add it to an offline dataset, basically, so that you can experiment upon it. Yeah. Do we have, do we have, I'm not sure if there's anyone after us in this room. Do we have to? Yeah, I'm not sure. Okay. According to agenda, yes. Oh, is it? Okay. I'm happy to go on then. Yeah. Yeah. Oh, great. Perfect. Yeah. Do you always measure it quantitatively, or do you also sometimes have some kind of qualitative piece of prose as the result? Like, user satisfaction can be a number. And also, I mentioned certain metrics to be just a vibe check or something. Um, we can, do you want to talk specifically about Braintrust, like, for that answer? Yeah. So, like, there is, the online scoring piece here, where it's a known unknown, where you can very much put a score behind that. But also, there are ways where more, these are not scores. This is the unknown unknowns piece. Got it. Yeah, thanks. Where we can, in a more open-ended way derive insight from it. Yeah. Yeah. Probably time for one question. If not, great. I appreciate everyone's attention today. Thank you. where they're performing some type of known control flow. Agent applications are very much non-deterministic. We're curious about why an agent might take one path versus the other. This also means that traditional observability is going to really have to focus on very constrained and known metrics. Whereas agent observability needs to be a little bit broader in terms of the things that that it needs to measure. This is just an example of that. So at the bottom, let's start there. Agent observability can measure some of these more traditional metrics, albeit with more of an AI flare. Time to first token, total tokens, duration, latency. These are all things that you would think be very traditional observability level metrics. But also you might want to understand more qualitative things about your application. So it's not just how long did I take to start responding to my user, which is more traditional observability. I want to know, was the information that I gave grounded in the context that I gathered with my application? Did I use the tools that I would have expected in this as I was reasoning towards my response? Is a response aligned to the brand standard that I set for this agent in the system prompt? These are all things that are not really able to be tested by traditional observability tools. Because if you think about it, like the trace necessary, the information in the trace that's necessary for us to compute these things up at the top is far larger than the volume that a traditional observability trace would handle. That kind of goes to the next point here. Agent traces are really nasty. They're in a variety of different ways. They're nasty because they're highly semi-structured. Even within those semi-structured, there's a ton of unstructured text data that we need to chew through. They're voluminous. So they can be, an agent trace could be over a gigabyte in size. We've seen that even with our own customers. An individual span can be 20 megabytes in size. So it's just a far different systems problem that you have to solve in order to ingest, process, and most importantly use that type of data. And also it's just as fast as traditional observability data. So hopefully your agent that you're putting in production gets product market fit and you have a ton of users and and usage associated with it. You as the AI engineer or as the product manager for that agent, you're going to want to see that observability in real time, in true real time. Trust me, we know that's the case because we always get the feedback. We're talking about, can you just make it faster? We're always trying to make it faster. People always want it to be faster. Tough to do when the agent traces look like this, basically. This is just like an example of an agent trace in BrainTrust, where not only does it have a bunch of spans here encompassing the model calls and tool calls, but even within those spans, you saw the amount of unstructured text that's in there as well. Very different problem to solve. A little bit more here. Like, maybe I'll just dive into the read pattern piece specifically. We need to do two things simultaneously. We need to be able to perform like the very fast read, ingest and read style workflows that are common with observability, i.e. if someone does an action with my agent, I need to be able to see that interaction basically instantaneously. We also have to commit to read patterns where someone wants to use our CLI and fire off SQL commands to us so that they can incorporate either observability or eval traces to improve their application automatically. There are just a lot of different mediums that people use now in order to query these very large trace shapes. This is a completely new systems problem. At least at BrainTrust, we designed a database from the ground up specifically for agent traces. I'm not going to really go into depth about this. We have a blog on our website. I think it was the last blog that we published if you're really interested in diving deep. But just very quickly, there are a lot of different components that we have to build into this database in order to make it work. For example, we need to immediately get data into a write-ahead log so that people can instantly see these traces as soon as they expect. We need to be able to perform indexing on these data so that whenever someone is performing a filter filtering or analytical query that it's fast. We have this thing called a Tantivy index. Tantivy is an open source framework that we've worked. Anyone know what Tantivy does? Any guesses? Tantivy is how we perform text-style indexing. If you remember when I was showing this trace, it makes so much sense for someone to want to perform the workflow of, okay, I just want to know every trace that had the word Amazon into it. Well, it turns out it's really hard to do that unless you perform a full text-based index across your traces. That's another reason why agent observability is far different than traditional observability. You really don't have to think about the text problems in traditional observability. Is it the same as OpenSearch? Sorry? It's kind of like an OpenSearch? Tantivy is most similar to like an Apache Lucene, except it's written in Rust. Yeah. And then all of these things come together and have to be unified through a SQL or SQL-similar language. That's what we've, that's the route that we've gone to at BrainTrust. Problem three, this is a, whereas there's a very specific type of persona for traditional observability. It's a systems engineer. Maybe it's a product engineer. It's probably not a subject matter expert, or if it's a medical application, it's not a, not a clinician or, or, or a registered nurse. It's very technical people that align with traditional observability. That could not be further from the truth for agent observability if you're doing it well. We noticed that the best teams that are building agent have both technical and non-technical people in the fold performing this work because it's the non-technical people that are either A, closest to the users, or B, have knowledge that is closest to the problem space. And what can they do now with prompts? They can write it in natural language. So they can add real value into being able to participate in agents. We have folks that are clinicians or registered nurses or wealth advisors or, or, or lawyers. We have seen them operate in our platform looking through traces and using that information to improve their agents. That is a workflow that you, that you simply don't see in traditional observability where you're more worried about uptime. I think it, like in general, people don't realize that in order to perform observability and, and also evals well, we kind of think of observability and evals as the same problem. The only difference between evals is that you're running them in batch and you know the inputs ahead of time. It's, it's incredible. The depth that you end up going into when you create a platform like this, it looks like, it looks like this. It looks like this, because of the, the reasons that I've described, the nuances with the data, the amount of, and, and types of people that you have to bring into the fold. Those are some of the reasons why it's so different to perform in this space. Where is the space going? We, we've done a lot of work in this area. I think the, the natural question that we used to get asked was, if you're, if, if you brain trust are collecting all of our agent traces and all of our agent traces have all of these valuable data in them, can't you just tell me how people are using my agent? And it's a, it's the simple questions that usually need the, the most complex systems behind them. Um, this is something that we are starting to do. Um, we, we just rolled it out, I think about a month ago in, in our software as a service offering where we see agent observability traces come in and then we'll run like a very lightweight LLM on top of them to perform embedding and then clustering on those traces to see how we can perform like elevate topic, uh, elevate topic modeling to see, for example, how people are using, uh, your traces, their intent, how people are feeling about interacting with your agent, the sentiment, or if they're running into issues, what those issues potentially are. The whole idea is there is that you can, um, make the iteration loop between a problem that you're seeing in production and the fix that you perform experimentation on. The whole idea is to just make that faster and a little bit more direct. Um, I promise I would go through that really fast. I've got about three minutes for questions if there are any, and I'd be really happy to answer them. Anyone curious about this? Yes. So I'll bring us this clearly about, uh, the, the functional observability of agents. Um, would you say- Is that also good for non-functional agent performance or is traditional observability good for that? That's a good question. Yeah. I would, well, I think traditional observability can do that. Um, Braintrust specifically does do, I like the way that you put that functional observability, um, what's the quality of my agent, how I've defined it, and then the technical observability just kind of comes on the house. Like when, when, when you, when you trace the application, you automatically get prompts, duration, time to first token, et cetera, cache hits, et cetera. Yeah. In your iceberg slide, you have human annotation. Yeah. You have the water line there. Yeah. Could you explain what the human annotation part is and get this cloud on? Yeah. So let's think about it this way. Um, actually, if, if I can, I'll go on a high wire act here and, um, and just show and, and, and not tell. So let's say that you have a trace come in and you want your product manager to be able to opine on whether that agent did a good job or a bad job. Um, it's really valuable for you to have an expert come in, grade the agents, but then also like justify why they're grading the agents the way that they are. Cause eventually you're going to take those justifications. You're going to probably run an LLM over it and you're going to make more, um, scalable scoring functions from those justifications. You're finding the failure modes that you can then implement and, and automated scores through that. Yeah. Cumin annotation is a really key part of this process. Okay. Thanks. Uh, yes. In the second row. Um, yeah. I just have a question, um, because I, uh, I, I, you know, you just focus on the ability today, but I'm interested actually how you also integrate to the other agent framework to, uh, uh, for the offline optimization. Mm-hmm. And then also, um, I guess the, yeah, I think the main, uh, difference I feel that I see in this, uh, your, uh, your database is that we're not using OLAP or Quickhouse. We used to use Quickhouse actually. Yeah. We moved away from it. Yeah. I just wanted to, curious to like, why build your own, like, what is the efficiency? Well, the, the funny answer there is that our, um, our founder is kind of an insane person. Like he, like only an insane person would build their own database, but he, he is, he is cut from that cloth. Um, he, he was one of the first employees at single stores, so he's kind of used to doing that. Um, but what he found was when, um, I think it was, let's see this slide. He found that when, um, he was performing some of these workloads, he just needed the, more of the text-based, uh, indexes, which Quickhouse wasn't really able to do, at least at that time. So we, we built our own. Um, and then the first part of your question, observability and evals to us, it's like we solve it with the same system. The only difference is that with evals, we know the inputs ahead of time with obs, and we're doing it in batch. With observability, we, we don't know what the inputs are ahead of time and we're doing them in real time. Oh, right. Uh, but I mean the experiment functionality, like, um, how, how easy is it to integrate with, like, uh, uh, . Oh yeah, it should be pretty easy. Like, when, once you've traced, and I, I apologize and I'm making this, like, about the product. Um, but when you, when you have a, a trace come in, you've traced it, and then you just, like, add it to an offline dataset, basically, so that you can experiment upon it. Yeah. Do we have, do we have, I'm not sure if there's anyone, uh, after us in this room. Do we have to? Yeah, I'm not sure. Okay. According to agenda, yes. Oh, is it? Okay. I'm, I'm happy to go on then. Yeah. Yeah. Oh, great. Perfect. Yeah. Do you always, um, measure it quantitatively, or do you also sometimes have some kind of qualitative piece of prose as the result? Like, user satisfaction can be a number. And also, I mentioned certain metrics to be just, yeah, a vibe check or something. Um, we can, do you want to talk specifically about brain trust, like, for that answer? Yeah. So, like, there is, like, the online scoring piece here, where it's, like, a known unknown, where you can, like, very much put a score behind that. Uh, but also, there are, there are ways where more, like, like, this is, these are not scores. This is, like, the unknown unknowns piece. Got it. Yeah, thanks. Um, where we can, in a, in a more, like, open-ended way derive insight from it. Yeah. Yeah. Probably time for, for one question. If not, great. I appreciate everyone's attention today. Thank you.