Thank you all for coming, first of all. And I want to talk today about raising the floor. So it's this term we use a lot. Mainly, I want to talk about very, very practical: what do we actually see working in the real world? How are people making their agents better? So the first thing I want to say is I could just say a bunch of stuff. I think the title of this track is continual learning. I think it's notable that in the real world, there's really not that much continual learning. If you look at the labs, if you look at products that are in the real world, you really don't see a lot of continual learning.
So I think it's very easy to, I could spend 20 minutes just talking about, hey, here's a bunch of frameworks, here's a bunch of really nice terms. But what I'd rather do is actually turn this a little bit into a dialogue. This is not just because I procrastinated making a bunch of slides, and because Fable was delayed, and I was counting on that to make the slides. But also because the reality is that there aren't really good standards for these things, right? There's not some one-size-fits-all solution.
And so I do have slides, believe it or not. But what we're also going to do is I'd like to hear from you guys, people actually building agents, where have the Twitter eval discourse failed you, right? Where is it not working? What are the things you're actually hitting in real life? I'd like to talk about that. So please, right now, start thinking about your questions. Start thinking about the annoying parts of your flow when you're building agents. And I'd like to keep a lot of time for Q&A.
So the reality is, a year ago, agents barely existed. I remember being at a speaker dinner here a year ago, and we're like, yeah, do you think agents will keep getting better? Will they not? And I think the crazy thing is, we were at this point in time a year ago where everything was a chatbot, mostly, right? I think people here, people in this room probably were a little further ahead.
It was a lot simpler then. If you remember evals, the eval discourse for a year or two ago, it would be like, oh, what is the capital of the United States? And you're like, yeah, you want to make sure that it returns Washington, D.C. And that was an easier time, right? It was like chatbots were so much more limited in their flexibility that it was easy to do a bunch of fact-checking things. You knew the answer to most questions your users would ask, almost, is another way of saying that. At least 80% of them or 90% of them.
But, yeah, we have agents being deployed in finance, healthcare, defense. And I was actually really against the word. One of my worst takes is early agents. I was like, I hate the word agent. I think there were a lot of people that felt the same. It was like, come on, it's an LLM, it's whatever. But I actually think it's valuable because I think we're seeing that agents are this almost self-aware entity, right? They run around their environment, they have these tools they're using. When they hit roadblocks, they start getting really creative, right? And that's what makes agents really powerful.
But that's also what makes them catastrophic. It's like, oh, well, I'll just decompile this. And I'll just do this thing that you had no idea, that you could have never imagined. Sometimes those solutions are helpful, right? Sometimes they're actually pretty harmful. But, yeah, what's very certain is we've come a very long way from next token prediction. If you think about chatbots, it was literally just, oh, it's good. What is the next likely word? It was very easy to reason about.
And I think the thing that, on the evaluation front, I think the reality is most of the things you'd read online about evals are really still stuck in this chatbot era. It's very, well, come up with your 1,000-eval data set. And the reality is nobody's doing that, very few people anyway. Sorry if you are.
And I think what teams have seen over and over again is, yeah, you can do that. But those evals break as soon as you have a new model, as soon as you switch harnesses. You have a bunch of tools that you're like, oh, yeah, I'm going to make sure that I'm going to write an eval where it has to call this tool if I ask it this question. And it's like, oh, then you switch to Cloud Code CLI, and now 80% of your evals suck. And it's like, okay, you could keep doing that. But the reality is the one thing I could promise you is that things are going to keep changing. We're not done.
And so I'd be very careful about investing months in some sort of eval set that's going to slow you down, right? I think the whole thing here is you want more safety, but you don't want theater. And I think that, again, the evals as has been prescribed by what I call big eval, I think that there's this reality where it's kind of like, oh, you really should eval, but then do you actually delay upgrading to the new model in your product, in your products? Do you actually delay it two weeks to update your evals or not, right? I think most people would say no.
So I'm the CTO and co-founder of this company called Raindrop. Very quickly, it's like we find critical issues in production agents. We verify those fixes actually work without unexpected side effects. And we also simulate changes before they land in production based on past behavior. We're used by the best AI companies in the world and Fortune 100s. A lot of logos I'm not allowed to put on here yet. But companies like Vercel, Speak, Framer. And I think what it means is we get this amazing peek into, again, what is actually working in the real world.
I think one of our tenants as a company is that things are changing constantly, and we have to change what we're doing constantly. And so I think we try to be very, very honest with both ourselves and our customers about what works and what does not work. And we try not to sell things that don't work.
Again, two things. We have this open source tool that thousands of people use. Maybe you use it yourself. It's called Workshop. So that's made by us. It's an open source tracing tool. It's really, really cool if you're trying to experiment with self-healing loops. I think it's the best way to do that because if there's anything it can't do, your agent could just add it, which is pretty cool. And so highly, highly recommend it. Again, I know thousands of people use it. I bump into people all the time that use it. Raindrop is our hosted offering that does issue detection. Think of it like Sentry, but it detects issues for agents. Sorry.
And we also make howtoeval.com. And so I think it's one of the most popular resources on how to evaluate AI agents. Again, the link is literally in the name. It's howtoeval.com. And it's our attempt at a very, very no-bullshit guide at what actually works. And I'll be talking a little bit about it today.
I think the root question that we're trying to figure out today together is how do you make your agent better, right? It's not even what issues does your agent have? It is actually how to make your agent better. Because your agent will have issues that potentially you can't solve or are not exactly worth solving, right? I think that we saw this over and over again, where it's, you can Think of it like Sentry, but it detects issues for agents. Sorry. And we also make howtoeval.com. And so I think it's one of the most popular resources on how to evaluate AI agents. It is, again, the link is literally in the name. It's howtoeval.com.
And it's our attempt at a very, very no-bullshit guide on what actually works. And I'll be talking a little bit about it today.
I think the root question that we're trying to figure out today together is, how do you make your agent better, right? It's not even what issues does your agent have? It is actually how to make your agent better. Because your agent will have issues that potentially you can't solve or are not exactly worth solving, right? I think that we saw this over and over again, where you can imagine that there are some things that you're better off waiting for. We know Fable exists now. Maybe should you train your own Fable-level model? It's probably not, right? And there will be benefits when you can just incorporate that into your product.
And so there's this actual balance: how do I actually make my agent better with the tools that I have? The way that we start thinking about it with customers is something like this, which is, are you a benchmark maxer or a floor raiser? I think that one of the problems when we talk about evaluating agents is that the terms are really confused. You hear OpenAI has a new eval benchmark. And they have evals. They run evals. And then you hear, oh, well, companies have evals. There's online evals. And the word eval is more or less a meaningless word. It literally is just you're evaluating something, right? It's a test in some cases. So it's a little confusing.
I think it's helpful. I think what it means is that companies start borrowing the language that labs are using and even copying similar benchmarks. But they're doing completely different things, right? They have completely different tools at their disposal. The companies that are downstream of models just have very, very different responsibilities than labs. Labs are trying to make these super general-purpose things. When they fail, at least at an API level, when I say that, if they get something wrong, it's just different. Companies are trying to imbue all this company-specific domain knowledge. Like, oh, here's the shape of the data. And here's what all this data means.
And here's how to access it. And so it's very, very different. We have this funny quiz on the howtoeval site. And it's interesting, right? One of the questions we would think about is, oh, are your users domain experts in the thing they're doing? Again, is it almost replacing someone or is it augmenting them? Because if you think about Copilot, autocomplete style, or Cursor, tab complete now, if it gets something wrong, you can just delete it, right? Even Claude Code CLI or Codex, if you're an engineer, it does do things wrong all the time. But then when you think about products like Devon, it actually gets more interesting, right?
If something messes up on the Claude Code side, it could be that you don't have something installed correctly on your computer. There's a lot more user error. You leave a lot more up to the users to get correct. And I think when you start thinking about things like AI doctors, for example, it's a very, very different shape of responsibility as far as how much responsibility the user has in actually getting things correct. So, anyways, it's kind of funny, just breaking this down. So, we think about the ceiling as what is the best thing, craziest capability, immersion capability that your product or agent is capable of?
Things that people would just not expect that it could do. And then the floor is, what is the worst thing your agent can do? Like, recommend a competitor, or delete a bunch of data, or accidentally send an AI slop email to a customer because it technically had access to your email or something. And, again, I think that the floor is very interesting because I think that that is the thing that breaks user trust.
This is the reason why people, if you think about the worst things that could start happening in society, and things we've already seen, whether that's the 4.0 psychophancy or things in that vein, a lot of it is more on the floor side rather than the capability side. Anyway, so I could talk about that for a long time. I'll skip this. So, the talk, obviously, is going to be about floor raising. And the first thing that we're going to talk about is offline evals. Again, we'll keep it very simple. I think that, as we said before, things have changed a lot since this chatbot era. The oh, you just look at string contains on the text output or something.
Or even the style of eval tools that have a prompt playground, this sort of thing. I actually don't know many companies that use some sort of managed prompt in the cloud anymore. There's one or two I can think of. And the reality is just that the prompt is actually the whole thing now. It's all the code. It's your whole harness. It's everything you're connected to. It's not just some string where you tell the agent what to do. And so what I think that means is the evals themselves actually should look a lot more like code. In other words, a lot more like tests, whether that's unit tests, whether that's end-to-end tests. They should look a lot more like tests.
Sentry has this really cool package called vitest evals. It's literally just vitest with some syntactic sugar on top. OpenAI calls this macro evals. And, again, I don't think it really matters what you call it, but essentially, run tests on your agent locally is the advice. And keep these evals as code. And, again, as much as possible, I don't see a lot of companies using the prompt playground stuff anymore because the shape of agents has really changed. When we think about raising the floor, we think about really three things. One is discovering all these unknown issues that you have in your app, things you're just not seeing. That's one.
Two is that for each issue, you really need to know two things. You need to know when it actually started, and you need to know how many people it affects. It sounds obvious, but I promise you that in the day-to-day of actually having an agent, you're going to get thousands of people saying, oh, I saw this weird thing. I saw this weird thing. So, again, the first thing is, is this new? Because if it's not new, I probably am going to care about it less. If I tell you, hey, look, this issue started yesterday, or this issue started three or four days ago, suddenly your mind starts turning, and you're like, oh, what did I do? Well, what changed, right? Did we change model?
Did we change something else downstream? And, again, the second one is percent of users. Knowing that it happened to three users versus 100,000 users just is critical. Because, again, I think agents will have an infinite number of problems. That's the great and terrible thing about them. They're these little stochastic, crazy things exploring everywhere. So, again, the first thing is, is this new? Because if it's not new, I probably am going to care about it less. If I tell you, hey, look, this issue started yesterday or this issue started three or four days ago, suddenly, your mind starts turning, and you're like, oh, what did I do? Well, what changed, right?
Did we change model? Did we change something else downstream? And again, the second one is percent of users. If I'm knowing that it happened to three users versus 100,000 users, it is critical. Because, again, I think agents will have an infinite number of problems. That's the great and terrible thing about them is they're these little stochastic, crazy things exploring everywhere. And so, in order to even start making things better, you really need to know these two things: when it started and percent of users. And I think also, another question we get a lot is around, oh, what should I be doing? And the first question I always ask people is, how many users do you have?
We have customers with millions of users, and we have customers with five.
And, to be clear, customers with five users, especially if, let's say it's an internal app in an enterprise context where it's giving very critical information, it could be very, very important to get well, or, sorry, to get correctly. But it does mean you should be taking a radically different approach. So, for example, on the 10, 20, 100 million messages a day side of things, experiments become extremely valuable. If you have a free tier, you can run experiments on a very small sample of your free tier, and that can just be extremely, extremely useful. Obviously, if you have five or 10 users, I would not recommend experiments or A-B tests, et cetera.
So this is one of those things that really, really depends on the person. What I want to talk about now before we get into Q&A are three very, very, very, very tactical lessons on the issue discovery and analysis side. There's three things that I've never heard anyone talk about, three things that we've just discovered from first principles as we do stuff at Raindrop. So any competitors in the audience, please pay attention. This is very important. The first one is that clusters are not issues. So the naive approach that we've seen either customers or also sometimes competitors take is, well, you just take all the traces and you just cluster it, right?
And you get these clusters. It could be useful from an analysis, Hamill calls this error analysis, this finding these clusters of things. It could be useful to see, whoa, what's going on in your data, right? Going from a bunch of logs to something. The problem is that, and again, I have here, it's useful for one-off analysis, but it just doesn't really scale well. And there's a very good reason why we also, if you think about normal telemetry, we try to think a lot about normal telemetry. What are the analogies? There's a reason why you don't take all of your normal logs and just start clustering it, right?
Because when you're building software, you need to know when something started. You need to know how much it's grown. Those things really matter. So, again, with clusters, it's very, very hard to reliably track over time. Again, it's called temporal clustering, and there's research in this. But it's pretty hard to do reliably. You also just don't have control of boundaries. And this also changes a lot depending on your product. What you consider to be the same issue or not is actually very, very unique to every company. And so you will get these weird clusters.
You can imagine each of these is, oh, wrong price quoted and wrong refund calculated are actually, you'll get a cluster, price issues or something. And it's, yeah, sort of. But actually, these could have extremely different root causes, right? So price issues or issues calculating as a cluster, it's not really that useful. And it also, again, doesn't really tell you the things that we talked about needing. So, yes, last one here is going to be code mode actually really scales. You've heard about code mode in the context of MCPs. I highly recommend just trying to apply this to traces.
You can just write these classifiers, and you can write them and you can run them in a sandbox, and you can run them at production scale. We have a feature that makes this easier, but you can do this. So I highly recommend it. The last lesson here is that agents are very, very bad at anomaly detection. So don't ask your agent to find anomalies. Ask it to investigate anomalies you've already found. So what I mean is, pull out as many deterministic things as you can, like keyword frequency, right? So if you see a spike in a keyword, it doesn't necessarily mean that there's an issue, but it does mean that it's
something more tangible, tractable that you can have an agent actually investigate. And I'm going to skip through the rest because we're tight on time, and I lied to you, which is that we're not going to have enough time for Q&A because I only have a minute left. But what I'd love if you could do is find me after. I'll be around for the next hour, and let's just talk. It's probably a better format than standing up here, and it'll be hard to hear your questions anyway. So, yeah, thank you guys so much. it's a very, very different shape of responsibility as far as, like, how much responsibility the user has in actually getting things correctly.
So, anyways, it's kind of funny. Just kind of breaking this down. So, we think about the ceiling as, like, what is the best thing, like, craziest capability, immersion capability that your product or agent is capable of? Like, things that people would just not expect that it could do. And then the floor is, like, what is the worst thing your agent can do? Like, recommend a competitor, or, like, delete a bunch of data, or, like, accidentally send an, you know, AI slop email to a customer because it, like, technically had access to, like, your email or something.
And, again, I think that, like, the floor is very interesting because I think that that is the thing that, like, breaks user trust. This is the thing that, like, the reason why people, like, if you think about the worst things that could start happening in society, whether that's, and things we've already seen, whether that's the, you know, like, 4.0 kind of psychophancy or things in that vein, a lot of it is more on, like, the floor side rather than the capability side. And, anyway, so I could talk about that for a long time. I'll kind of skip this.
So, the talk, obviously, is going to be about floor raising. And the first thing that we're going to talk about is, like, offline evals. Again, we'll keep it very simple. I think that, like we said before, things have changed a lot sort of since this chatbot era. The sort of, like, oh, you just, you know, like, look at, you know, string contains, you know, on the, like, the text output or something. Or even, like, the style of eval tools that have, like, a prompt playground, like, this sort of thing. Like, I actually don't know many companies that use some sort of, like, managed prompt, like, in the cloud anymore. There's, like, one or two I can think of.
And the reality is just, like, the prompt is actually, like, the whole thing now. It's, like, all the code. It's your whole harness. It's, like, everything you're connected. Like, it's not just, like, some string where you tell the agent what to do. And so what I think that means is the evals themselves actually should look a lot more like code. In other words, like, a lot more like tests, whether that's unit tests, whether that's end-to-end tests. They should look a lot more like tests. Sentry has this really cool package called, like, vitest evals. It's literally just, like, vitest with, like, some syntactic sugar on top. OpenAI calls this, like, macro evals.
And, again, I don't think it really matters what you call it, but, like, essentially run tests on your agent locally is the advice. And keep these evals as code. And, again, as much as possible, like, the, I don't see a lot of companies using the sort of, like, prompt playground stuff anymore because of how the shape of agents has really changed.
When we think about raising the floor, we think about really, like, three things. One is it, like, discovering all these, like, unknown issues that you have in your app, like, things you're just, like, not seeing. That's one. Two is that for each issue, you really need to know two things. You need to know when it actually started, and you need to know how many people it affects. It sounds, like, obvious, but I promise you that, like, in the day-to-day of actually, like, having an agent, you're going to get, like, you know, you already get thousands of people, like, oh, I saw this weird thing. I saw this weird thing. So, again, the first thing is, like, is this new?
Because if it's not new, like, I probably, like, am going to care about it less. If I tell you, like, hey, look, this issue started yesterday or this issue started, like, three or four days ago, suddenly, like, your mind starts turning, and you're like, oh, what did I do? Like, well, what changed, right? Did we change model? Did we change, you know, something else downstream? And again, the second one is, like, percent of users. Like, if I'm knowing that it happened to three users versus 100,000 users just is critical. Because, again, I think agents will have an infinite number of problems.
That's sort of, like, the great and terrible thing about them is, like, they're, like, these little stochastic, you know, crazy things exploring everywhere. And so you just, in order to even start making things better, you really need to know these two things, when it started and percent of users. And I think also, another question we get a lot is, like, around, like, oh, like, I, you know, what should I be doing? And, like, the first question I always ask people is, like, how many users do you have? Like, we have customers with millions of users, and we have customers with, like, five.
And, like, to be clear, like, customers with five users, like, especially if, let's say it's, like, an internal app in an enterprise context where it's, like, you know, giving, like, very critical information. Like, it could be very, very important to get well, or, sorry, to get correctly. But it does mean you just, like, should be taking a radically different approach. Like, so, for example, on the, like, you know, let's say, like, 10, 20, 100 million, you know, messages a day side of things, like, experiments become extremely valuable.
If you have a free tier, you can, like, run experiments on a very small sample of your free tier, and that can just be extremely, extremely useful. Obviously, if you have five or 10 users, like, I would not recommend, you know, experiments or A-B tests, et cetera. So this is one of those things that, like, really, really depends on the person.
What I want to talk about now before we get into Q&A are, like, three very, very, very, very tactical lessons on the sort of, like, issue discovery and analysis side. There's, like, three things that I've never heard anyone talk about, like, three things that we've sort of just discovered from first principles as we do stuff at Raindrop. So any competitors in the audience, please pay attention. This is very important. The first one is that clusters are not issues. So the sort of, like, naive approach that we've seen either customers or also sometimes competitors take is, like, well, you just take all the traces and you just cluster it, right?
And you get these, like, clusters. It could be, like, useful from, like, an analysis, you know, like, Hamill calls this, like, error analysis, like, this, you know, finding these clusters of things. It could be useful to see, like, whoa, what's going on in your data, right? Going from, like, a bunch of logs to, like, something.
The problem is that, like, and again, I have here, it's useful for one-off analysis, but it just doesn't really scale well. And there's, like, a very good reason why we also, you know, if you think about, like, normal telemetry, we try to think a lot about, like, normal telemetry. What are the analogies? There's a reason why you don't sort of, like, take all of your, you know, normal logs and just, like, start clustering it, right? Because when you're building software, you need to know, like, when something started. You need to know how much it's grown. Those things really matter. So, again, with clusters, it's very, very hard to reliably track over time.
Like, again, it's, like, called, like, temporal clustering, and there's, like, research in this. But, like, it's pretty hard to do reliably. You also just, like, don't have control of boundaries. And this also changes a lot depending on your product. Like, what you consider to be, like, you know, the same issue or not is actually very, very unique to every company. And so you sort of will get these, like, kind of weird clusters. Like, you know, you can imagine each of these is, like, oh, wrong price quoted and wrong refund calculated are, like, actually, like, you'll get a cluster, like, you know, price issues or something. And it's, like, yeah, sort of.
But, like, actually, these could have, like, extremely different root causes, right? So price issues or, like, you know, issues calculating as, like, you know, as a cluster, it's not really that useful. And it also, again, doesn't really tell you the things that, you know, we talked about needing.
So, yes, last one here is going to be code mode actually really scales. Like, you've heard about code mode in the context of MCPs. I highly recommend just trying to apply this to traces. Like, you can just write these classifiers and you can write them and you can run them in a sandbox and you can run them at production scale. We have a, you know, feature that makes this easier, but, like, you can do this. So I highly recommend it. The last lesson here is that agents are very, very bad at anomaly detection. So don't ask your agent to find anomalies. Ask it to investigate anomalies you've already found.
So what I mean is, like, pull out as many deterministic things as you can, like keyword frequency, right? So if you see a spike in, like, a keyword, it doesn't necessarily mean that there's an issue, but it does mean that you can, it's, like, something more tangible, tractable that you can have an agent actually investigate.
And I'm going to skip through the rest because we're tight on time, and I lied to you, which is that we're not going to have enough time for Q&A because I only have a minute left. But what I'd love if you could do is find me after. I'll be around for the next hour, and let's just talk. It's probably a better format than standing up here, and it'll be hard to hear your questions anyway. So, yeah, thank you guys so much. .