[SPEAKER_00] All right, all right.
SPEAKER_00
First of all, thank you so much for coming. I'm actually rather surprised. A lot of times, you're working on this stuff, and you're cooked up in a room, and you think no one cares. And then so many people showed up. So I suppose someone cares. So anyway, the title of my talk today is, evals are broken, and you should use them anyway. A lot of this talk is just a straight up critique of the way we do evals these days. And I want to help you out. I want to give you a way out of this. You have this interesting technology, and you can use it, but there's so many ways to mess it up. So I want to help you out. My first claim is that most people are wrong about evals.
SPEAKER_00
And I want you to be right about evals. I want you to use them. I want you to be able to build with them, interpret them, use evals in your own agentic flows, leverage them in any way that makes sense. So that's the point of the conversation. To be right about something that has a lot of nuances that can go in many different directions, the fundamental question is, how are people wrong about that thing? There are basically two camps of people who are wrong about evals. The first camp is the objective metrics camp.
SPEAKER_00
The objective metrics camp is this: there are people who would look at this dashboard and interpret it as something meaningful, as in, GPT 5.4 is effectively the same as Gemini 3.1 Pro Preview. Believe me, they're not the same. There are a lot of these models which show up with similar numbers. And at a certain point, the whole thing is a hoax. You won't believe it at all. So there was this tweet that came out just this morning. It was a critique of Meta, where Meta came out with classic benchmark maxing. It was like, we're doing best on the benchmark. Everything's great.
SPEAKER_00
And I assure you, if you try a lot of these models, you just won't hold the test of actual real world evidence. The other camp, the other way where people are wrong, is that they go too far the other way. This is the vibes camp. The people in the vibes camp think that everything is about vibes. If you ask them why they like a model, they'll say things like, I like talking to it. They anthropomorphize it. And there's no right answer either. I think the truth is somewhere in the middle: evals are not the end-all be-all, but they're also not completely useless. The right way is to use them. The wrong way is to use them poorly.
SPEAKER_00
So to do that, I'll give you three stages that will help you use them really well. The first stage is leverage evals from other people. The second stage is use evals to improve your own agents. And the third stage is build your own evals for specific use cases. In the interest of time, I could talk about evals for hours, but I can only talk about level one and two. And I think those would be most helpful for most people in the audience. So I'm going to give you a few heuristics to interpret evals. The first heuristic is that whenever a model app comes out with a number, just don't believe them. These are approximations.
SPEAKER_00
Coming back to the tweet, just don't believe the model app eval numbers. They're somewhat of an approximation. Sometimes they're good. Sometimes they're not. There was this tweet that's a pretty cool one where Nikon said that a lot of AI researchers and engineers routinely dismiss evals. They don't really think of it as something where the numbers should be taken that seriously. And to some extent, it's a matter of actual trying and preferences. And I think that is somewhat more accurate. So the second heuristic is that you want to stay current, but you don't want to be the earliest adopter. Why am I saying this? This is the Epoch Index.
SPEAKER_00
It's basically the aggregate score of different models on evals. If you notice, in the last two years, every couple of months, the frontier model is changing. And it's changing so fast. It's so hard to keep up with this stuff. I've worked on this. I've been doing this for a living for years at this point. And even for me, my preferences are changing so fast. When you're working through these things, the way I would recommend is that you let the thing come out first. Let things set on fire for a couple of weeks.
SPEAKER_00
And then if the thing still stands the test of time, at that point you should do your model switch and try something, rather than always trying to be on the cutting edge. The people who have to always try the new model and try the new things—they'll be me. And even me, I have preferences changing so fast. So, I think that when you're working through these things, the way I would recommend is that let the thing come out first. Let things set on fire for a couple weeks. And then if the thing still stands the test of time, I think at that point you should do your model switch and try something rather than always trying to be on the cutting edge.
SPEAKER_00
The people who have to always try the new model and try the new things, they'll be me. But I do this for a living. And you don't have to. And I think to a lot of people in the AI research community, that was very obvious. It was very obvious that SdbBench doesn't measure frontier coding capabilities. Because it had problems like solve this Fibonacci sequence. It would have problems like do matrix multiplication or something. And it's just it doesn't apply to real old software engineering. So you want to have something that's very new, but also actually legit. And it takes some discernment to figure that out. So that's the first part.
SPEAKER_00
But the second part is okay, now that we know that we have a few heuristics of how to use evals, how do you use evals to improve your agent upon them? And I think this is the part where I lean into the core philosophy of this conversation where you want to think of evals as an engineering problem, but also as a philosophy problem. Right? So the engineering problem is obviously hard, but the philosophy problem is also very hard. The philosophy problem is that you have a problem, and you can't exactly approximate the search space of where the problem could go, where the problems could fail. It's somewhat easier to do it for coding problems.
SPEAKER_00
But even then, coding problems have an infinite source space. They can go in any direction. So you want to build evals that are somewhat more approximate representation of the actual thing that you're dealing with. And for us, to give some context in Client's Journey. So I work at Client. Client is an open source coding agent company. We have a very interesting product. I encourage you to try it out. So in Client's Journey, one of the things that we dealt with is in the last year, one of the things we found is that there were a few evals available. At the time, we were very rudimentary. Every other company was very rudimentary as well.
SPEAKER_00
And our thinking was okay, if there's so few standardized evals available, and also they're not effective, they really are not measuring what it is that you're trying to do in your day-to-day programming job, what do you do? So our stance was, and this was the stance of the Codex team and a lot of other teams that we've talked to, we were saying, just this eval just completely ignore them. They're completely unnecessary. You're probably wasting your time, and I don't know who will be appeased by them. And then last year, we came to this idea that okay, listen, I think we got to up the ante and we got to have some measure. We got to try evals.
SPEAKER_00
And if no one else is doing it, we'll do it ourselves. We'll build actual evals from scratch that would actually test real-world programming problems of users. So we went through a lot of our massive data sets of people who had opted in to share their coding usage of client with us. And we offered them money and we got a lot of this data set of okay, this is what the problems that people are actually doing. Then spent a lot of time parsing through that, figured out an actual data set, these are the problems that people are solving.
SPEAKER_00
And then completely cleaned it all up, doing a lot of really hard manual labor, trying to make very decent problems that can be solved with, say, a client or any other coding agent. The hardest part for us when we were building evals is that if you're building evals for anything that's rudimentary, if you're building eval for, say, an LLM model, you have a very simple one-shot use case of how many toes does a cat have? And then the LLM can just be I don't know, 11 or whatever. I don't know how many toes a cat has. But a single turn eval is very easy to do because it has a binary answer. And it has a very limited search page of what the answer could be.
SPEAKER_00
But when you're working with an agent, that can't be the case. You're working with an agent, you can give an agent a problem like hey, I have this new MCP server. It's probably not working. How do you make it work for me? And that's usually how a lot of you guys talk to Cloud Coda, whatever agent you're using, and myself as well. So in this, it's very hard to gauge because the agent reads through files, searches through docs, installs the environment, sets things up, runs Python scripts, does all of that. And then in the end, runs some tests and then maybe the whole thing works. So we're trying to grade the second thing.
SPEAKER_00
We're trying to grade all these things that will take a lot of time. And then figure out oh, did it actually work or did it not? Did it work but broke other things? So that's why it was harder. So in the same time, some very awesome, smart, bright people from Stanford University Files, searches through docs, installs the environment, sets things up, runs Python scripts. Does all of that. And then in the end, runs some tests and then maybe the whole thing works. So we're trying to grade the second thing. We're trying to grade all these things that will take a lot of time. And then figure out, did it actually work or did it not? Did it work but broke other things?
SPEAKER_00
So that's why it was harder. So in the same time, some very, very awesome, smart, bright people from Stanford University came up with Terminal Bench, which does the same thing, where they came up with 89 coding problems, which are just very approximate, decent representation of real-world programming problems. So these could be things like race conditions, database issues, other stuff. Figure out this infra issue. And I think that Terminal Bench was built to use with any coding agent CLI so you can actually test and run things really fast with CLI. And then it will take a couple minutes to run. So some of these tasks would take up to 30 to 40 minutes.
SPEAKER_00
And that's how you know they're legit, because the agent does a lot of things and just runs in circles and sometimes goes crazy. And yeah. So we started using that. And the way to use Terminal Bench is that you think of an eval problem as an evaluation suite, which has a set of problems. This one has 89 tasks. Some others would have more. And what you want to do is you want to give it an environment. You want to give it an isolated environment where you let's say you have a task, hey, figure out this race condition for me in this repo. And the race condition is that this thing is not working.
SPEAKER_00
So you want to be able to give the eval run an isolated environment, a virtual machine. And that virtual machine has the whole setup. It has the repo. It has everything. And then you install whatever agent you have. In our case, client, clock code, codex, whatever you want to use, you can do that. To do that, it's not hard, but it is also not trivial. And Harbor is another software that came from Lott Institute where they made this thing where let's say you have 89 tasks. One way to do evals is that learn each of these tasks in sequence and then do all the setups.
SPEAKER_00
Another way to do it is have a very standardized configuration defined in infrastructure where each of these 89 tasks have the proper Linux machine, the proper RAM CPU usage. And then being able to isolate those environments and then run those 89 tasks in parallel on infrastructure. So you could use a couple different things for the infrastructure here. You could use Daytona. You could run it on your Docker machine if you have very powerful machines. I'm sure if you can handle that much compute, sure, but I wouldn't. We use Modal. Modal, we're very thankful to Modal. They've helped us a lot. So shout out to them.
SPEAKER_00
And yeah, so in this case, Harbor basically lets you split up the 89 tasks and then they all run in parallel. So that way, your limit, the limiting factor is basically the slowest task. Yeah. So the slowest task is the limiting factor. So the process is this. You get a score. You first do a run on the 89 task. You get a score. You evaluate all the failures. So let's say you get 50 failures, right, out of the 89 task. What you want to be able to do is you want to portfolio allocate those failures.
SPEAKER_00
You want to say, you want to run another agent which goes through the traces of all the failures, so the trace would be a massive file which has every single LLM call that the agent did. And then be like, OK, this one, this specific problem failed because it didn't run tests. This failed because the read file tool was broken. And once you portfolio allocate those failures, you figure out, OK, these are the small levers that I can pull. If I pull those levers, I can make massive improvements to my AI agent. So what you're testing is you're basically testing three things. You are testing the model itself.
SPEAKER_00
If you have a very decent model, somehow you could have a horrible harness, you could have a horrible agent, but the model just overshoots so hard that you just get a great score. You're testing the harness. You're testing your coding harness. So you're testing, say, if you're using clock code codex. So sometimes you'll find, I'm sure, I guarantee you, some of you have noticed that let's say Anthropics models could potentially work with the cursor, could potentially work with Droid, could work with other coding agents. But for some reason, it just seems to work so much better with clock code, right?
SPEAKER_00
And I think that is testing the harness, that is the harness actually really leveraging the best of the model? And the third problem, whether the problem is sane. If you're solving stupid problems, it doesn't matter if you score 100% all the time. So you really got to make sure that the problems are sane, which the Law Institute has done a pretty great job of. So for us, it was a case like this. We basically had this original score, which was much lower, 43%. We made changes to CPU. We made changes to memory in front of the containers. We raised timeouts. We improved the thinking behavior. Sometimes we would ask the model to think more.
SPEAKER_00
If you're solving stupid problems, it doesn't matter if you score 100% all the time. So you really got to make sure that the problems are sane, which the Law Institute has done a pretty great job of. So for us, it was a case like this. We basically had this original score, which was much lower, like 43%. We made changes to CPU. We made changes to memory in front of the containers. We raised timeouts. We improved the thinking behavior. Sometimes we would ask the model to think more. Sometimes asking the model to think more actually interferes with the quality of the response, because it goes in a stroke. And it just goes in circles. And it was just I am a model. I am a model. I was just keep doing it for like 2,000 tokens. So you got to think through all of that. And so for us, we have a huge internal benchmark for all kinds of models, open source models. So we just keep a list of trying different versions and stuff. We encourage other people to try that as well if you, pretty helpful. So whenever you get zones of improvements, you get basically three zones of improvements when you get an original score. The first one is the obvious flaws. Sometimes your harness really has very obvious flaws of there's this bug that straight up caches the harness. And those obvious bugs you've got to fix, right? Sometimes you are not getting rate limited or whatever. Fix those. That's fine. I think zone two is the most critical one where you actually do nuance improvements. And these nuance improvements are things like there are certain prompt engineering techniques that apply to anthropic model families that just straight up would not apply to codex model family that would be very different from Gemini model family. And those are the nuance of why is it that this is a model that's so good that so many people are saying it's so good but for some reason it just isn't working for me? I think that is the essence of working with agents and hill climbing. That you figure out those nuance improvements of tweaking your prompt, making it larger, making it smaller. And then zone three is the danger zone where it's you're straight up overfitting. So you're overfitting in the sense that you're just straight up cheating so you get the highest score and then you can make a tweet about it. Don't, a lot of people have done it. Don't do it. I wouldn't do it. I mean, never mind. Anyway. So this was basically the rough outline. So the final wording for me would be basically, regardless of the kind of problem that you have, I want you to find a benchmark and build an eval and hill climb on it. So hill climbing means that you get a score and then you improve the score off your harness on the eval. And you have to do both. You can't just have a good number and be happy with it. You got to both pass the vibe check. Does it actually feel good to use this product in this model? And at the same time, you also have a very, very, very decent score, hopefully. If a new thing comes out, you do your absolute best to give it the right judgment. For us, one of the things that we learned was that we were very decent on anthropic model families, not so much on, say, Gemini model family, not so much on, say, Ki Mini model family, which, again, are very decent models. So when we started hill climbing, we learned that oh, if we support these models, we have this entire swath of people who love these models and they can start using us. And I think that some reflection of that would also apply with you. So my final note to you guys is that I've done some hot takes or whatever. And if you work for some of the companies that I've said not so nice things about, I still love you and everything. And it was, I work at Klein. So if you find these problems fascinating, if you want to learn more about these, this is my Twitter. So you can feel free to reach out to me, DM me about, if you want to work on problems like these, by all means, I put a word for you. If you want to learn more about evals, if you want to learn, how do you, I have a problem that's completely orthogonal to everything you're defining for coding agents. How do we work on that? So feel free to reach out to me and I'll respond to you. And once again, it's very kind of you to give me your time. Thank you so much. Thank you.
SPEAKER_00
And the way to use that, like, the way to use Terminal Bench is that, like, you think of an eval problem as, like, an evaluation suite, which has a set of problems. This one has 89 tasks. Some others would have more. And what you want to do is, like, you want to give it an environment. You want to give it an isolated environment where you just, like, you, let's say you have a task, like, hey, figure out this race condition for me in this repo. And the race condition is that, like, this thing is not working. So you want to be able to give the eval run an isolated environment, like a virtual machine. And that virtual machine, it has the whole setup. It has the repo.
SPEAKER_00
It has everything. And then you install whatever agent you have. In our case, client, clock code, codex, whatever you want to use, you can do that. To do that, it's not that it's, like, hard, but it is also not trivial. And Harbor is another software that came from Lott Institute where they made this thing where, let's say you have 89 tasks. One way to do evals is that learn, like, each of these tasks in sequence and then, like, do all the setups. Another way to do it is, like, have a very standardized configuration defined in infrastructure where each of these 89 tasks have, like, the proper Linux machine, the proper RAM CPU usage.
SPEAKER_00
And then, like, being able to, like, isolate those environments and then run those 89 tasks in parallel on infrastructure. So you could use a couple different things for the infrastructure here. You could use Daytona. You could run it on your Docker machine if you have, like, very powerful machines. I'm sure if you can handle those that much compute, sure, but I wouldn't. We use Modal. Modal, we're very thankful to Modal. They've helped us a lot. So shout out to them. And yeah, so in this case, like, Harbor basically lets you split up, like, the 89 tasks and then they all run in parallel. So that way, your limit, the limiting factor is basically the slowest task. Yeah.
SPEAKER_00
So, yeah, so the slowest task is the limiting factor. So the process is this. You get a score. You first do a run on the 89 task. You get a score. You evaluate all the failures. So let's say you get, like, say, 50 failures, right, out of the 89 task. What you want to be able to do is you want to portfolio allocate those failures. You want to say, you want to run, like, another agent which goes through the traces of all the failures, so the trace would be, like, this massive file which has, like, every single LLM call that the agent did. And then be like, OK, this one, this specific problem failed because it didn't run tests.
SPEAKER_00
This failed because the read file tool was broken. And once you portfolio allocate those failures, you figure out, OK, these are the small levers that I can pull. If I pull those levers, like, I can make, like, massive improvements to my AI agent. So what you're testing is, like, you're basically testing, like, three things. You are testing the model itself. Like, if you have a very decent model, like, somehow, like, you could have a horrible harness, you could have a horrible agent, but, like, the model just, like, overshoots so hard that just, like, you get a great score. You're testing the harness. You're testing your coding harness.
SPEAKER_00
So, like, you're testing, say, if you're using clock code codex. So sometimes you'll find, I'm sure, I guarantee you, some of you have noticed that, like, let's say, Anthropics models could potentially work with the cursor, could potentially work with Droid, could work with other coding agents. But for some reason, it just seems to work so much better with clock code, right? And I think that that is, like, the testing the harness, that, like, is the harness actually really leveraging the best of the model? And the third problem, whether the problem is sane. If you're solving stupid problems, it doesn't matter if you score 100% all the time.
SPEAKER_00
So you really got to make sure that, like, the problems are sane, which the Law Institute has done a pretty great job of. So for us, it was a case like this. Like, we basically, like, had, like, this original score, which was, like, much lower, like 43%. We made changes to, like, CPU. We made changes to memory in front of the containers. We raised timeouts. We improved the thinking behavior. Sometimes we would ask the model to think more. Sometimes asking model to think more actually interferes with the quality of the response, because it goes in, like, it gets, like, a stroke. And it just, like, goes in, like, circles. And it's, like, it was just, like, I am a model.
SPEAKER_00
I am a model. I was just, like, keep doing it for, like, like, 2,000 tokens. So, yeah. So, like, you got to think through all of that. And, yeah. So for us, like, we have, like, a huge, like, internal benchmark for all kinds of models, open source models. So, like, we just, like, keep, like, a list of, like, trying different versions and stuff. We encourage other people to try that as well if you, pretty helpful. So whenever you get, like, zones of improvements, you get, like, basically three zones of improvements when you get an original score. The first one is the obvious flaws. Like, sometimes your harness really has, like, very obvious flaws of, like, there's this bug
SPEAKER_00
that straight up caches the harness. And those obvious bugs you've got to fix, right? Sometimes you are not, like, you're getting rate limited or whatever. Fix those. That's fine. I think the zone two is the most critical one where you actually do nuance improvements. And these nuance improvements are things like there are certain prompt engineering techniques that apply to anthropic model families that just straight up would not apply to codex model family that would be very different from Gemini model family. And those are the nuance of, like, why is it that this is a model that's so good that
SPEAKER_00
so many people are saying it's so good but for some reason it just isn't working for me? I think that is the essence of, like, working with agents and hill climbing. That, like, you figure out those nuance improvements of, like, tweaking your prompt, making it larger, making it smaller. And then zone three is the danger zone where it's, like, you're straight up overfitting. So you're overfitting in the sense that, like, you're just straight up cheating so you get the highest score and then you can, like, make a tweet about it. Don't, like, a lot of people have done it. Don't do it. Like, I wouldn't do it. I mean, never mind. Anyway.
SPEAKER_00
So, yeah, so, anyway, so this was, like, this was, like, basically the rough outline. So the final, the final wording for me would be, like, basically, regardless of the kind of problem that you have, I want you to, like, find a benchmark and, like, like, build an eval and just, like, hill climb on it. So hill climbing means that, like, you get a score and then you improve the score off your harness on the eval. And you have to do both. Like, you can't just, like, have, like, a good number and be happy with it. Like, you got to both pass the vibe chip. Like, does it actually feel good to use this product in this model?
SPEAKER_00
And at the same time, you also have, like, a very, very, very decent score, hopefully. If a new thing comes out, you do your absolute best to, like, give it the right judgment. For us, like, one of the things that we learned was that, like, we were very decent on anthropic model families, not so much on, say, Gemini model family, not so much on, say, Ki Mini model family, which, again, are very decent models. So when we started hill climbing, we learned that, like, oh, if we support these models, we have this, like, entire swath of people who love these models and they can start using us. And I think that some reflection of that would also apply with you.
SPEAKER_00
So my final note to you guys is that, like, you know, I've done some hot takes or whatever. And if you work for some of the companies that I've said not so nice things about, I still love you and everything. And it was, it was, I work at Klein. So if you find these problems fascinating, if you want to learn more about these, like, this is my Twitter. So, like, you can feel free to reach out to me, DM me about, like, if you want to work on problems like these, like, by all means, like, I put a word for you.
SPEAKER_00
If you want to learn more about evals, if you want to learn, like, how do you, like, I have a problem that's, like, completely orthogonal to everything you're defining for coding agents. Like, how do we work on that? So feel free to reach out to me and I'll respond to you. And once again, it's very, very kind of you to give me your time. Thank you so much. Thank you.