AI Engineer

Evals Are Broken, Use Them Anyway — Ara Khan, Cline

2961 summary words 13 min summary Watch video

Start with the signal

13 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Traditional benchmark evals are misleading at best and fraudulent at worst, but when combined with real-world agent testing, continuous hill climbing, and vibes validation, they become the only practical way to improve AI agents systematically.
  • Why it matters: Anyone building AI agents or choosing models is navigating a benchmark-maxing minefield where published numbers are often meaningless, yet abandoning evals entirely leaves you with pure vibes and no improvement pathway
  • Best use: Watch for Ara's three-zone hill climbing framework (obvious flaws → nuanced prompt tuning → overfitting danger zone) and the portfolio allocation technique for diagnosing agent failures at scale

Executive Summary

Ara Khan from Cline argues that the AI eval ecosystem splits into two dysfunctional camps: the objective metrics camp, which treats benchmarks like SWEBench as gospel despite rampant benchmark-maxing by model labs, and the vibes camp, which rejects all quantitative measurement in favor of anthropomorphized subjective preference. Both are wrong. The truth is nuanced: evals are broken—Meta tweeted this morning about beating benchmarks while the actual model disappoints in practice—but you must use them anyway because they're the only scalable way to improve agents systematically.

The talk centers on Cline's internal journey. Initially, they dismissed evals entirely because early coding benchmarks (SWEBench classic) measured trivial problems like Fibonacci sequences rather than real-world software engineering. When no legitimate benchmarks existed, Cline built their own from opt-in user data, then discovered TerminalBench from Stanford—89 real-world coding tasks including race conditions and database issues that take agents 30-40 minutes to solve. Using Harbor (from Lott Institute) to parallelize 89 isolated VM environments on Modal infrastructure, they score agent runs, portfolio-allocate failures (LLM trace analysis reveals 'didn't run tests' vs 'read file tool broken'), and hill climb by tweaking CPU/memory/timeouts/prompts. Original score: 43%. After systematic improvement across three zones (obvious flaws, nuanced tuning, overfitting danger), scores rose meaningfully.

The core methodology tests three things simultaneously: the model itself (a strong model can compensate for poor harness design), the harness (why Claude works better in Claude Code than Cursor despite being the same model), and problem sanity (100% on trivial tasks is worthless). Ara emphasizes prompt engineering is model-family-specific: techniques that work for Anthropic models fail on Codex or Gemini families. The danger zone (zone 3) is overfitting to leaderboards for tweets. The correct approach combines quantitative hill climbing with qualitative vibes validation—both must pass. Cline discovered they were strong on Anthropic models but weak on Gemini/DeepSeek families; fixing that unlocked new user segments.

Key Takeaways

  • Claim: Model labs routinely publish fraudulent or benchmark-maxed eval scores that don't reflect real-world performance | Evidence: Meta tweeted this morning claiming benchmark superiority; Ara says trying the actual model reveals the numbers are a hoax. GPT 5.4 and Gemini 3.1 Pro Preview show similar dashboard numbers but behave completely differently in practice. Nikon (AI researcher) tweeted that many AI engineers dismiss evals entirely because published numbers are untrustworthy approximations. | Caveat: Ara doesn't provide specific Meta model names or quantitative evidence of the discrepancy; this is informed critique from a practitioner, not a controlled study | Implication: Ken should never trust vendor-published benchmark scores at face value for agent/model selection decisions; wait 2-3 weeks post-launch for community vetting before switching models | Timestamp: 00:30 / 02:15
  • Claim: Early coding benchmarks like SWEBench classic are useless for frontier agent evaluation because they test toy problems (Fibonacci, matrix multiplication) rather than real software engineering | Evidence: SWEBench had problems like 'solve this Fibonacci sequence'—obvious to AI research community that it doesn't measure real coding capability. TerminalBench from Stanford uses 89 tasks including race conditions and database issues that take agents 30-40 minutes to complete, proving task difficulty by agent behavior (runs in circles, goes crazy). | Caveat: Ara doesn't specify exactly when SWEBench classic became obsolete or whether newer versions (e.g., SWEBench Verified) address this; TerminalBench may itself become outdated as models improve | Implication: For Ken's agent ops work: task duration (30-40 min agent runtime) and failure mode diversity are better proxies for benchmark legitimacy than task count or published scores | Timestamp: 03:45 / 08:20
  • Claim: Portfolio allocation of failures—using an LLM to categorize why each eval task failed by analyzing agent traces—is the most critical step in systematic agent improvement | Evidence: After scoring 50 failures out of 89 TerminalBench tasks, Cline runs another agent over massive trace files (every LLM call recorded). Analysis reveals specific failure modes: 'didn't run tests,' 'read file tool was broken,' etc. Categorizing failures identifies small levers that yield massive improvements when pulled. | Caveat: Ara doesn't share the prompt or architecture of the trace-analysis agent, nor quantify how much improvement came from each lever; this is a qualitative framework, not a cookbook | Implication: Ken should instrument every agent run to produce detailed traces, then batch-analyze failures with a separate LLM to find systematic issues rather than treating each failure as unique—this is the path from 43% to competitive scores | Timestamp: 12:30 / 13:45
  • Claim: Prompt engineering techniques are model-family-specific: what works for Anthropic models straight-up fails for Codex or Gemini families, which is why the same model performs differently across coding agent harnesses | Evidence: Cline discovered they were 'very decent on Anthropic model families, not so much on Gemini or DeepSeek.' After hill climbing with family-specific prompt tuning, they unlocked new user segments. Ara notes Claude feels better in Claude Code than Cursor despite being the same underlying model—this is harness optimization exploiting model-family quirks. | Caveat: No specific prompt differences disclosed (likely proprietary); unclear if this applies to frontier models that may be more robust to prompt variation | Implication: Ken should maintain separate prompt templates per model family and A/B test cross-family when evaluating new models, rather than assuming a one-size-fits-all agent harness | Timestamp: 15:20 / 16:10
  • Claim: The three-zone hill climbing framework prevents both stagnation and overfitting: Zone 1 (obvious flaws like caching bugs), Zone 2 (nuanced prompt/timeout/resource tuning), Zone 3 (overfitting for leaderboard tweets) | Evidence: Cline's score progression from 43%: Zone 1 fixes included CPU/memory/timeout adjustments and rate limit handling. Zone 2 involved prompt size tuning and discovering that asking models to 'think more' sometimes causes 2,000-token loops of 'I am a model' rather than better reasoning. Zone 3 is explicitly labeled 'the danger zone'—Ara warns many teams cheat for Twitter clout. | Caveat: No quantitative breakdown of which zone contributed how much improvement; Zone 2/3 boundary is subjective and requires judgment to avoid self-deception | Implication: Ken should explicitly categorize every proposed agent improvement into one of these three zones before implementation to maintain discipline and avoid the overfitting trap that destroys generalization | Timestamp: 16:45 / 17:50
  • Claim: You must pass both the quantitative eval score AND the qualitative vibes check; either alone is insufficient for shipping agent improvements | Evidence: Ara's final guidance: 'You can't just have a good number and be happy with it. You got to both pass the vibe check—does it actually feel good to use this product/model?—and at the same time have a very decent score.' This dual requirement prevents both benchmark-maxing fraud and unjustified vibes-based decisions. | Caveat: No operational definition of what constitutes 'passing the vibe check' or how to reconcile conflicts when evals improve but vibes degrade (or vice versa) | Implication: Ken's agent deployment pipeline should require both automated eval score gates AND manual user testing/feedback loops before production release to catch divergence between metrics and reality | Timestamp: 18:25

Detailed Brief

The Two Camps Problem: Why Most People Are Wrong About Evals

  • Claims: The objective metrics camp treats dashboards showing GPT 5.4 ≈ Gemini 3.1 Pro Preview as meaningful truth when models behave completely differently; The vibes camp anthropomorphizes models ('I like talking to it') and rejects all quantitative measurement; Meta's benchmark-maxing tweet this morning exemplifies the objective metrics failure mode; Nikon's tweet: AI researchers/engineers routinely dismiss evals as untrustworthy approximations
  • Evidence: Real-world usage reveals models with similar benchmark numbers are 'not the same—believe me'; SWEBench classic had Fibonacci/matrix multiplication problems that don't test real software engineering; The Epoch Index shows frontier model rankings change every few months, making it hard to stay current
  • Caveats: Ara provides anecdotal rather than controlled evidence for benchmark fraud claims; No specific quantitative comparison of published vs. real-world model performance gaps; The 'vibes camp' characterization may be a strawman; serious practitioners likely use informal heuristics rather than pure anthropomorphization
  • Implications: Ken should treat all vendor-published scores as upper bounds requiring 2-3 weeks of community validation; Agent selection decisions need multi-week trials on realistic tasks, not day-one adoption of frontier releases; The right philosophy is nuanced middle ground: evals are broken approximations but still the best improvement tool available

Cline's Eval Journey: From Dismissal to TerminalBench Hill Climbing

  • Claims: Initially, Cline (and Codex team, other coding agent teams) completely ignored evals as ineffective and unnecessary; Cline built custom evals from opt-in user data by cleaning and curating real programming problems users submitted; Stanford's TerminalBench (89 tasks) became the standard because tasks are real-world (race conditions, DB issues, infra problems) and take 30-40 minutes per agent run; Harbor (Lott Institute) parallelizes eval runs by giving each of 89 tasks an isolated Linux VM with proper RAM/CPU, eliminating sequential bottlenecks
  • Evidence: Single-turn LLM evals ('how many toes does a cat have?') have binary answers and limited search space—trivial to build; Agent evals are hard because agents read files, search docs, install environments, run scripts, execute tests across many turns with massive search space; Cline uses Modal for infrastructure (shout-out mentioned); alternatives include Daytona or powerful Docker setups; The slowest task becomes the limiting factor when running 89 tasks in parallel
  • Caveats: No disclosure of Cline's custom eval dataset composition or whether it's public; TerminalBench may become obsolete as models improve (like SWEBench classic did); Harbor setup is 'not hard but not trivial'—Ara doesn't provide implementation details or cost estimates
  • Implications: Ken should prioritize multi-turn agent evals over single-turn LLM evals for any agentic use case; Task realism is validated by agent behavior: if tasks complete in seconds, they're too easy; 30-40 min runtime suggests appropriate difficulty; Infrastructure choice (Modal, Daytona, Docker) matters for eval velocity; parallelization is mandatory to iterate quickly

The Three-Thing Test and Portfolio Allocation Debugging Technique

  • Claims: Every eval run tests three things: (1) the model itself, (2) the harness/agent wrapper, (3) problem sanity; A very strong model can compensate for horrible harness design and still achieve good scores; Harness quality explains why Anthropic models feel better in Claude Code than Cursor despite being identical models; Portfolio allocation of failures: run a second agent over trace files (logs of every LLM call) to categorize failure modes like 'didn't run tests' vs 'read file tool broken'; Identifying small levers via portfolio allocation enables massive agent improvements
  • Evidence: Cline scored 43% initially, then improved via CPU/memory/timeout changes, thinking behavior adjustments (asking model to think more sometimes causes 2,000-token 'I am a model' loops); Cline maintains huge internal benchmark across model families and versions; Discovered they were strong on Anthropic, weak on Gemini/DeepSeek families; fixing that unlocked new user segments
  • Caveats: No shared trace-analysis agent prompt or architecture details; No breakdown of which specific levers contributed how much improvement from 43% baseline; Problem sanity is subjective; even TerminalBench from Stanford could have issues
  • Implications: Ken's agent systems should log exhaustive traces (every LLM call, tool invocation, decision point) for post-hoc failure analysis; Building a meta-agent that categorizes failure modes at scale is a critical infrastructure investment for rapid iteration; Harness optimization per model family is as important as model selection itself

The Three-Zone Hill Climbing Framework and Overfitting Danger

  • Claims: Zone 1 (obvious flaws): caching bugs, rate limiting issues, infrastructure failures—easy wins that must be fixed immediately; Zone 2 (nuanced improvements): prompt engineering tuning per model family, timeout/resource optimization, thinking behavior adjustments—the critical zone for real progress; Zone 3 (danger zone): overfitting to benchmarks to juice leaderboard scores for Twitter clout—destroys generalization and credibility; Prompt techniques that work for Anthropic straight-up fail for Codex or Gemini families; Asking models to think more can backfire into 2,000-token degenerative loops
  • Evidence: Cline's improvements included CPU/memory/timeout tweaks (Zone 1), prompt size/structure changes (Zone 2); Many teams have fallen into Zone 3 overfitting—Ara warns 'a lot of people have done it, don't do it'; Model-family-specific prompt tuning is 'the essence of working with agents and hill climbing'
  • Caveats: No clear operational boundary between Zone 2 legitimate optimization and Zone 3 overfitting—requires judgment; No disclosure of specific prompt engineering techniques that differ across model families; Risk of self-deception: teams may believe they're in Zone 2 while actually in Zone 3
  • Implications: Ken should establish explicit overfitting detection: hold-out eval sets, out-of-distribution task testing, or qualitative review before every agent release; Zone 2 is where competitive advantage lives—Ken should invest heavily in model-family-specific harness tuning rather than treating all models identically; The thinking behavior example (more ≠ better) suggests Ken should A/B test reasoning token budgets rather than assuming more inference compute always helps

Notable Concepts & Terms

  • TerminalBench: Stanford-developed benchmark of 89 real-world coding tasks (race conditions, DB issues, infra problems) designed for CLI coding agents with 30-40 min task runtimes proving difficulty
  • Harbor: Lott Institute tool that parallelizes eval runs by provisioning isolated Linux VMs per task with standardized RAM/CPU configs, eliminating sequential bottlenecks so slowest task is the limiting factor
  • Portfolio allocation of failures: Ara's technique: run meta-agent over agent trace files (logs of every LLM call) to categorize failure modes systematically (e.g., 'didn't run tests' vs 'tool broken'), identifying small levers for massive improvements
  • Hill climbing: Iterative process of scoring agent on eval, diagnosing failures, making targeted improvements, re-scoring—used to systematically improve from 43% baseline without overfitting
  • The three-thing test: Every eval run tests: (1) model capability, (2) harness quality (why Claude feels better in Claude Code than Cursor), (3) problem sanity (100% on trivial tasks is worthless)
  • Benchmark-maxing: Labs optimizing models for specific benchmark performance rather than real-world capability, resulting in fraudulent scores that don't reflect actual usage quality (Meta's tweet example)
  • The three zones of improvement: Zone 1: obvious flaws (caching bugs, rate limits). Zone 2: nuanced tuning (model-family-specific prompts, timeouts). Zone 3: overfitting danger (cheating for leaderboard tweets)
  • Vibes check: Qualitative user experience validation that must be passed alongside quantitative eval scores before shipping agent improvements—neither alone is sufficient

Operator Notes / Why Ken Should Care

  • For Ken's agent systems: implement exhaustive trace logging (every LLM call, tool use, decision) as foundational infrastructure for portfolio allocation debugging—this is the unlock for systematic improvement
  • Model selection heuristic: wait 2-3 weeks post-launch for community vetting before production adoption, rather than chasing frontier releases on day one—let others get burned first
  • Harness optimization per model family is as important as model selection: maintain separate prompt templates for Anthropic/OpenAI/Google/DeepSeek families and A/B test cross-family rather than one-size-fits-all
  • Establish explicit overfitting detection before claiming eval improvements: hold-out sets, out-of-distribution tasks, qualitative review—Zone 2/3 boundary requires discipline to avoid self-deception
  • Infrastructure for parallel eval runs (Modal, Daytona, Docker) is mandatory for iteration velocity; sequential eval runs create unacceptable feedback loop latency for agent development
  • The dual requirement (quantitative score + qualitative vibes) prevents both benchmark fraud and unjustified vibes-based decisions—build both gates into deployment pipeline
  • Task realism proxy: if agent completes tasks in seconds, they're too easy; 30-40 min runtimes with visible struggle (runs in circles, goes crazy) suggest appropriate difficulty level
  • Thinking behavior counter-intuition: asking models to reason more can backfire into degenerative loops; A/B test reasoning token budgets rather than assuming more inference compute always helps
  • Meta-agent for failure categorization is worth building: scales analysis of hundreds of failures to find systematic issues rather than treating each as unique debugging session

Watch Map

  • 00:00: Introduction: evals are broken, use them anyway thesis
  • 00:30: The two camps: objective metrics vs vibes, both wrong; Meta benchmark-maxing tweet example
  • 03:45: SWEBench classic failure mode: Fibonacci/matrix multiplication don't test real engineering
  • 05:20: Heuristic 1: don't believe model lab eval numbers; Heuristic 2: stay current but not earliest adopter
  • 07:15: Epoch Index: frontier model changes every few months, hard to keep up
  • 08:20: Cline's journey: initially dismissed evals entirely, then built custom from user data
  • 09:45: Single-turn LLM evals vs multi-turn agent evals: search space and difficulty difference
  • 11:00: TerminalBench: Stanford's 89 real-world tasks (race conditions, DB issues) with 30-40 min runtimes
  • 12:00: Harbor parallelization: isolated VMs per task, slowest task is limiting factor; Modal infrastructure
  • 12:30: Portfolio allocation of failures: meta-agent analyzes traces to categorize failure modes
  • 14:00: The three-thing test: model, harness, problem sanity; why Claude feels better in Claude Code
  • 15:20: Prompt engineering is model-family-specific; Cline strong on Anthropic, weak on Gemini/DeepSeek initially
  • 16:10: 43% baseline improvement via CPU/memory/timeout tweaks and thinking behavior tuning
  • 16:45: Three zones: obvious flaws, nuanced improvements (the critical zone), overfitting danger
  • 17:50: Asking models to think more can cause 2,000-token 'I am a model' loops rather than better reasoning
  • 18:25: Final guidance: must pass both quantitative score AND qualitative vibes check before shipping
  • 18:50: Closing: reach out on Twitter for eval questions, Cline hiring if you find these problems fascinating

Source/Metadata

  • Title: Evals Are Broken, Use Them Anyway — Ara Khan, Cline
  • Transcript words: 5217
  • Duration seconds: 1144
  • Timestamp note: No explicit timestamps in transcript; approximate timestamps generated from talk flow and duration (19:04 total)
Full transcript 3446 words · 23 min read
0:00

[SPEAKER_00] All right, all right.

0:15

SPEAKER_00

First of all, thank you so much for coming. I'm actually rather surprised. A lot of times, you're working on this stuff, and you're cooked up in a room, and you think no one cares. And then so many people showed up. So I suppose someone cares. So anyway, the title of my talk today is, evals are broken, and you should use them anyway. A lot of this talk is just a straight up critique of the way we do evals these days. And I want to help you out. I want to give you a way out of this. You have this interesting technology, and you can use it, but there's so many ways to mess it up. So I want to help you out. My first claim is that most people are wrong about evals.

0:43

SPEAKER_00

And I want you to be right about evals. I want you to use them. I want you to be able to build with them, interpret them, use evals in your own agentic flows, leverage them in any way that makes sense. So that's the point of the conversation. To be right about something that has a lot of nuances that can go in many different directions, the fundamental question is, how are people wrong about that thing? There are basically two camps of people who are wrong about evals. The first camp is the objective metrics camp.

0:57

SPEAKER_00

The objective metrics camp is this: there are people who would look at this dashboard and interpret it as something meaningful, as in, GPT 5.4 is effectively the same as Gemini 3.1 Pro Preview. Believe me, they're not the same. There are a lot of these models which show up with similar numbers. And at a certain point, the whole thing is a hoax. You won't believe it at all. So there was this tweet that came out just this morning. It was a critique of Meta, where Meta came out with classic benchmark maxing. It was like, we're doing best on the benchmark. Everything's great.

1:26

SPEAKER_00

And I assure you, if you try a lot of these models, you just won't hold the test of actual real world evidence. The other camp, the other way where people are wrong, is that they go too far the other way. This is the vibes camp. The people in the vibes camp think that everything is about vibes. If you ask them why they like a model, they'll say things like, I like talking to it. They anthropomorphize it. And there's no right answer either. I think the truth is somewhere in the middle: evals are not the end-all be-all, but they're also not completely useless. The right way is to use them. The wrong way is to use them poorly.

1:55

SPEAKER_00

So to do that, I'll give you three stages that will help you use them really well. The first stage is leverage evals from other people. The second stage is use evals to improve your own agents. And the third stage is build your own evals for specific use cases. In the interest of time, I could talk about evals for hours, but I can only talk about level one and two. And I think those would be most helpful for most people in the audience. So I'm going to give you a few heuristics to interpret evals. The first heuristic is that whenever a model app comes out with a number, just don't believe them. These are approximations.

2:23

SPEAKER_00

Coming back to the tweet, just don't believe the model app eval numbers. They're somewhat of an approximation. Sometimes they're good. Sometimes they're not. There was this tweet that's a pretty cool one where Nikon said that a lot of AI researchers and engineers routinely dismiss evals. They don't really think of it as something where the numbers should be taken that seriously. And to some extent, it's a matter of actual trying and preferences. And I think that is somewhat more accurate. So the second heuristic is that you want to stay current, but you don't want to be the earliest adopter. Why am I saying this? This is the Epoch Index.

2:56

SPEAKER_00

It's basically the aggregate score of different models on evals. If you notice, in the last two years, every couple of months, the frontier model is changing. And it's changing so fast. It's so hard to keep up with this stuff. I've worked on this. I've been doing this for a living for years at this point. And even for me, my preferences are changing so fast. When you're working through these things, the way I would recommend is that you let the thing come out first. Let things set on fire for a couple of weeks.

3:16

SPEAKER_00

And then if the thing still stands the test of time, at that point you should do your model switch and try something, rather than always trying to be on the cutting edge. The people who have to always try the new model and try the new things—they'll be me. And even me, I have preferences changing so fast. So, I think that when you're working through these things, the way I would recommend is that let the thing come out first. Let things set on fire for a couple weeks. And then if the thing still stands the test of time, I think at that point you should do your model switch and try something rather than always trying to be on the cutting edge.

3:33

SPEAKER_00

The people who have to always try the new model and try the new things, they'll be me. But I do this for a living. And you don't have to. And I think to a lot of people in the AI research community, that was very obvious. It was very obvious that SdbBench doesn't measure frontier coding capabilities. Because it had problems like solve this Fibonacci sequence. It would have problems like do matrix multiplication or something. And it's just it doesn't apply to real old software engineering. So you want to have something that's very new, but also actually legit. And it takes some discernment to figure that out. So that's the first part.

4:02

SPEAKER_00

But the second part is okay, now that we know that we have a few heuristics of how to use evals, how do you use evals to improve your agent upon them? And I think this is the part where I lean into the core philosophy of this conversation where you want to think of evals as an engineering problem, but also as a philosophy problem. Right? So the engineering problem is obviously hard, but the philosophy problem is also very hard. The philosophy problem is that you have a problem, and you can't exactly approximate the search space of where the problem could go, where the problems could fail. It's somewhat easier to do it for coding problems.

4:22

SPEAKER_00

But even then, coding problems have an infinite source space. They can go in any direction. So you want to build evals that are somewhat more approximate representation of the actual thing that you're dealing with. And for us, to give some context in Client's Journey. So I work at Client. Client is an open source coding agent company. We have a very interesting product. I encourage you to try it out. So in Client's Journey, one of the things that we dealt with is in the last year, one of the things we found is that there were a few evals available. At the time, we were very rudimentary. Every other company was very rudimentary as well.

4:55

SPEAKER_00

And our thinking was okay, if there's so few standardized evals available, and also they're not effective, they really are not measuring what it is that you're trying to do in your day-to-day programming job, what do you do? So our stance was, and this was the stance of the Codex team and a lot of other teams that we've talked to, we were saying, just this eval just completely ignore them. They're completely unnecessary. You're probably wasting your time, and I don't know who will be appeased by them. And then last year, we came to this idea that okay, listen, I think we got to up the ante and we got to have some measure. We got to try evals.

5:15

SPEAKER_00

And if no one else is doing it, we'll do it ourselves. We'll build actual evals from scratch that would actually test real-world programming problems of users. So we went through a lot of our massive data sets of people who had opted in to share their coding usage of client with us. And we offered them money and we got a lot of this data set of okay, this is what the problems that people are actually doing. Then spent a lot of time parsing through that, figured out an actual data set, these are the problems that people are solving.

5:36

SPEAKER_00

And then completely cleaned it all up, doing a lot of really hard manual labor, trying to make very decent problems that can be solved with, say, a client or any other coding agent. The hardest part for us when we were building evals is that if you're building evals for anything that's rudimentary, if you're building eval for, say, an LLM model, you have a very simple one-shot use case of how many toes does a cat have? And then the LLM can just be I don't know, 11 or whatever. I don't know how many toes a cat has. But a single turn eval is very easy to do because it has a binary answer. And it has a very limited search page of what the answer could be.

6:10

SPEAKER_00

But when you're working with an agent, that can't be the case. You're working with an agent, you can give an agent a problem like hey, I have this new MCP server. It's probably not working. How do you make it work for me? And that's usually how a lot of you guys talk to Cloud Coda, whatever agent you're using, and myself as well. So in this, it's very hard to gauge because the agent reads through files, searches through docs, installs the environment, sets things up, runs Python scripts, does all of that. And then in the end, runs some tests and then maybe the whole thing works. So we're trying to grade the second thing.

6:36

SPEAKER_00

We're trying to grade all these things that will take a lot of time. And then figure out oh, did it actually work or did it not? Did it work but broke other things? So that's why it was harder. So in the same time, some very awesome, smart, bright people from Stanford University Files, searches through docs, installs the environment, sets things up, runs Python scripts. Does all of that. And then in the end, runs some tests and then maybe the whole thing works. So we're trying to grade the second thing. We're trying to grade all these things that will take a lot of time. And then figure out, did it actually work or did it not? Did it work but broke other things?

7:17

SPEAKER_00

So that's why it was harder. So in the same time, some very, very awesome, smart, bright people from Stanford University came up with Terminal Bench, which does the same thing, where they came up with 89 coding problems, which are just very approximate, decent representation of real-world programming problems. So these could be things like race conditions, database issues, other stuff. Figure out this infra issue. And I think that Terminal Bench was built to use with any coding agent CLI so you can actually test and run things really fast with CLI. And then it will take a couple minutes to run. So some of these tasks would take up to 30 to 40 minutes.

7:35

SPEAKER_00

And that's how you know they're legit, because the agent does a lot of things and just runs in circles and sometimes goes crazy. And yeah. So we started using that. And the way to use Terminal Bench is that you think of an eval problem as an evaluation suite, which has a set of problems. This one has 89 tasks. Some others would have more. And what you want to do is you want to give it an environment. You want to give it an isolated environment where you let's say you have a task, hey, figure out this race condition for me in this repo. And the race condition is that this thing is not working.

8:08

SPEAKER_00

So you want to be able to give the eval run an isolated environment, a virtual machine. And that virtual machine has the whole setup. It has the repo. It has everything. And then you install whatever agent you have. In our case, client, clock code, codex, whatever you want to use, you can do that. To do that, it's not hard, but it is also not trivial. And Harbor is another software that came from Lott Institute where they made this thing where let's say you have 89 tasks. One way to do evals is that learn each of these tasks in sequence and then do all the setups.

8:42

SPEAKER_00

Another way to do it is have a very standardized configuration defined in infrastructure where each of these 89 tasks have the proper Linux machine, the proper RAM CPU usage. And then being able to isolate those environments and then run those 89 tasks in parallel on infrastructure. So you could use a couple different things for the infrastructure here. You could use Daytona. You could run it on your Docker machine if you have very powerful machines. I'm sure if you can handle that much compute, sure, but I wouldn't. We use Modal. Modal, we're very thankful to Modal. They've helped us a lot. So shout out to them.

9:17

SPEAKER_00

And yeah, so in this case, Harbor basically lets you split up the 89 tasks and then they all run in parallel. So that way, your limit, the limiting factor is basically the slowest task. Yeah. So the slowest task is the limiting factor. So the process is this. You get a score. You first do a run on the 89 task. You get a score. You evaluate all the failures. So let's say you get 50 failures, right, out of the 89 task. What you want to be able to do is you want to portfolio allocate those failures.

9:49

SPEAKER_00

You want to say, you want to run another agent which goes through the traces of all the failures, so the trace would be a massive file which has every single LLM call that the agent did. And then be like, OK, this one, this specific problem failed because it didn't run tests. This failed because the read file tool was broken. And once you portfolio allocate those failures, you figure out, OK, these are the small levers that I can pull. If I pull those levers, I can make massive improvements to my AI agent. So what you're testing is you're basically testing three things. You are testing the model itself.

10:11

SPEAKER_00

If you have a very decent model, somehow you could have a horrible harness, you could have a horrible agent, but the model just overshoots so hard that you just get a great score. You're testing the harness. You're testing your coding harness. So you're testing, say, if you're using clock code codex. So sometimes you'll find, I'm sure, I guarantee you, some of you have noticed that let's say Anthropics models could potentially work with the cursor, could potentially work with Droid, could work with other coding agents. But for some reason, it just seems to work so much better with clock code, right?

10:33

SPEAKER_00

And I think that is testing the harness, that is the harness actually really leveraging the best of the model? And the third problem, whether the problem is sane. If you're solving stupid problems, it doesn't matter if you score 100% all the time. So you really got to make sure that the problems are sane, which the Law Institute has done a pretty great job of. So for us, it was a case like this. We basically had this original score, which was much lower, 43%. We made changes to CPU. We made changes to memory in front of the containers. We raised timeouts. We improved the thinking behavior. Sometimes we would ask the model to think more.

11:09

SPEAKER_00

If you're solving stupid problems, it doesn't matter if you score 100% all the time. So you really got to make sure that the problems are sane, which the Law Institute has done a pretty great job of. So for us, it was a case like this. We basically had this original score, which was much lower, like 43%. We made changes to CPU. We made changes to memory in front of the containers. We raised timeouts. We improved the thinking behavior. Sometimes we would ask the model to think more. Sometimes asking the model to think more actually interferes with the quality of the response, because it goes in a stroke. And it just goes in circles. And it was just I am a model. I am a model. I was just keep doing it for like 2,000 tokens. So you got to think through all of that. And so for us, we have a huge internal benchmark for all kinds of models, open source models. So we just keep a list of trying different versions and stuff. We encourage other people to try that as well if you, pretty helpful. So whenever you get zones of improvements, you get basically three zones of improvements when you get an original score. The first one is the obvious flaws. Sometimes your harness really has very obvious flaws of there's this bug that straight up caches the harness. And those obvious bugs you've got to fix, right? Sometimes you are not getting rate limited or whatever. Fix those. That's fine. I think zone two is the most critical one where you actually do nuance improvements. And these nuance improvements are things like there are certain prompt engineering techniques that apply to anthropic model families that just straight up would not apply to codex model family that would be very different from Gemini model family. And those are the nuance of why is it that this is a model that's so good that so many people are saying it's so good but for some reason it just isn't working for me? I think that is the essence of working with agents and hill climbing. That you figure out those nuance improvements of tweaking your prompt, making it larger, making it smaller. And then zone three is the danger zone where it's you're straight up overfitting. So you're overfitting in the sense that you're just straight up cheating so you get the highest score and then you can make a tweet about it. Don't, a lot of people have done it. Don't do it. I wouldn't do it. I mean, never mind. Anyway. So this was basically the rough outline. So the final wording for me would be basically, regardless of the kind of problem that you have, I want you to find a benchmark and build an eval and hill climb on it. So hill climbing means that you get a score and then you improve the score off your harness on the eval. And you have to do both. You can't just have a good number and be happy with it. You got to both pass the vibe check. Does it actually feel good to use this product in this model? And at the same time, you also have a very, very, very decent score, hopefully. If a new thing comes out, you do your absolute best to give it the right judgment. For us, one of the things that we learned was that we were very decent on anthropic model families, not so much on, say, Gemini model family, not so much on, say, Ki Mini model family, which, again, are very decent models. So when we started hill climbing, we learned that oh, if we support these models, we have this entire swath of people who love these models and they can start using us. And I think that some reflection of that would also apply with you. So my final note to you guys is that I've done some hot takes or whatever. And if you work for some of the companies that I've said not so nice things about, I still love you and everything. And it was, I work at Klein. So if you find these problems fascinating, if you want to learn more about these, this is my Twitter. So you can feel free to reach out to me, DM me about, if you want to work on problems like these, by all means, I put a word for you. If you want to learn more about evals, if you want to learn, how do you, I have a problem that's completely orthogonal to everything you're defining for coding agents. How do we work on that? So feel free to reach out to me and I'll respond to you. And once again, it's very kind of you to give me your time. Thank you so much. Thank you.

11:10

SPEAKER_00

And the way to use that, like, the way to use Terminal Bench is that, like, you think of an eval problem as, like, an evaluation suite, which has a set of problems. This one has 89 tasks. Some others would have more. And what you want to do is, like, you want to give it an environment. You want to give it an isolated environment where you just, like, you, let's say you have a task, like, hey, figure out this race condition for me in this repo. And the race condition is that, like, this thing is not working. So you want to be able to give the eval run an isolated environment, like a virtual machine. And that virtual machine, it has the whole setup. It has the repo.

11:46

SPEAKER_00

It has everything. And then you install whatever agent you have. In our case, client, clock code, codex, whatever you want to use, you can do that. To do that, it's not that it's, like, hard, but it is also not trivial. And Harbor is another software that came from Lott Institute where they made this thing where, let's say you have 89 tasks. One way to do evals is that learn, like, each of these tasks in sequence and then, like, do all the setups. Another way to do it is, like, have a very standardized configuration defined in infrastructure where each of these 89 tasks have, like, the proper Linux machine, the proper RAM CPU usage.

12:22

SPEAKER_00

And then, like, being able to, like, isolate those environments and then run those 89 tasks in parallel on infrastructure. So you could use a couple different things for the infrastructure here. You could use Daytona. You could run it on your Docker machine if you have, like, very powerful machines. I'm sure if you can handle those that much compute, sure, but I wouldn't. We use Modal. Modal, we're very thankful to Modal. They've helped us a lot. So shout out to them. And yeah, so in this case, like, Harbor basically lets you split up, like, the 89 tasks and then they all run in parallel. So that way, your limit, the limiting factor is basically the slowest task. Yeah.

13:00

SPEAKER_00

So, yeah, so the slowest task is the limiting factor. So the process is this. You get a score. You first do a run on the 89 task. You get a score. You evaluate all the failures. So let's say you get, like, say, 50 failures, right, out of the 89 task. What you want to be able to do is you want to portfolio allocate those failures. You want to say, you want to run, like, another agent which goes through the traces of all the failures, so the trace would be, like, this massive file which has, like, every single LLM call that the agent did. And then be like, OK, this one, this specific problem failed because it didn't run tests.

13:38

SPEAKER_00

This failed because the read file tool was broken. And once you portfolio allocate those failures, you figure out, OK, these are the small levers that I can pull. If I pull those levers, like, I can make, like, massive improvements to my AI agent. So what you're testing is, like, you're basically testing, like, three things. You are testing the model itself. Like, if you have a very decent model, like, somehow, like, you could have a horrible harness, you could have a horrible agent, but, like, the model just, like, overshoots so hard that just, like, you get a great score. You're testing the harness. You're testing your coding harness.

14:12

SPEAKER_00

So, like, you're testing, say, if you're using clock code codex. So sometimes you'll find, I'm sure, I guarantee you, some of you have noticed that, like, let's say, Anthropics models could potentially work with the cursor, could potentially work with Droid, could work with other coding agents. But for some reason, it just seems to work so much better with clock code, right? And I think that that is, like, the testing the harness, that, like, is the harness actually really leveraging the best of the model? And the third problem, whether the problem is sane. If you're solving stupid problems, it doesn't matter if you score 100% all the time.

14:43

SPEAKER_00

So you really got to make sure that, like, the problems are sane, which the Law Institute has done a pretty great job of. So for us, it was a case like this. Like, we basically, like, had, like, this original score, which was, like, much lower, like 43%. We made changes to, like, CPU. We made changes to memory in front of the containers. We raised timeouts. We improved the thinking behavior. Sometimes we would ask the model to think more. Sometimes asking model to think more actually interferes with the quality of the response, because it goes in, like, it gets, like, a stroke. And it just, like, goes in, like, circles. And it's, like, it was just, like, I am a model.

15:18

SPEAKER_00

I am a model. I was just, like, keep doing it for, like, like, 2,000 tokens. So, yeah. So, like, you got to think through all of that. And, yeah. So for us, like, we have, like, a huge, like, internal benchmark for all kinds of models, open source models. So, like, we just, like, keep, like, a list of, like, trying different versions and stuff. We encourage other people to try that as well if you, pretty helpful. So whenever you get, like, zones of improvements, you get, like, basically three zones of improvements when you get an original score. The first one is the obvious flaws. Like, sometimes your harness really has, like, very obvious flaws of, like, there's this bug

15:54

SPEAKER_00

that straight up caches the harness. And those obvious bugs you've got to fix, right? Sometimes you are not, like, you're getting rate limited or whatever. Fix those. That's fine. I think the zone two is the most critical one where you actually do nuance improvements. And these nuance improvements are things like there are certain prompt engineering techniques that apply to anthropic model families that just straight up would not apply to codex model family that would be very different from Gemini model family. And those are the nuance of, like, why is it that this is a model that's so good that

16:25

SPEAKER_00

so many people are saying it's so good but for some reason it just isn't working for me? I think that is the essence of, like, working with agents and hill climbing. That, like, you figure out those nuance improvements of, like, tweaking your prompt, making it larger, making it smaller. And then zone three is the danger zone where it's, like, you're straight up overfitting. So you're overfitting in the sense that, like, you're just straight up cheating so you get the highest score and then you can, like, make a tweet about it. Don't, like, a lot of people have done it. Don't do it. Like, I wouldn't do it. I mean, never mind. Anyway.

16:52

SPEAKER_00

So, yeah, so, anyway, so this was, like, this was, like, basically the rough outline. So the final, the final wording for me would be, like, basically, regardless of the kind of problem that you have, I want you to, like, find a benchmark and, like, like, build an eval and just, like, hill climb on it. So hill climbing means that, like, you get a score and then you improve the score off your harness on the eval. And you have to do both. Like, you can't just, like, have, like, a good number and be happy with it. Like, you got to both pass the vibe chip. Like, does it actually feel good to use this product in this model?

17:25

SPEAKER_00

And at the same time, you also have, like, a very, very, very decent score, hopefully. If a new thing comes out, you do your absolute best to, like, give it the right judgment. For us, like, one of the things that we learned was that, like, we were very decent on anthropic model families, not so much on, say, Gemini model family, not so much on, say, Ki Mini model family, which, again, are very decent models. So when we started hill climbing, we learned that, like, oh, if we support these models, we have this, like, entire swath of people who love these models and they can start using us. And I think that some reflection of that would also apply with you.

18:00

SPEAKER_00

So my final note to you guys is that, like, you know, I've done some hot takes or whatever. And if you work for some of the companies that I've said not so nice things about, I still love you and everything. And it was, it was, I work at Klein. So if you find these problems fascinating, if you want to learn more about these, like, this is my Twitter. So, like, you can feel free to reach out to me, DM me about, like, if you want to work on problems like these, like, by all means, like, I put a word for you.

18:25

SPEAKER_00

If you want to learn more about evals, if you want to learn, like, how do you, like, I have a problem that's, like, completely orthogonal to everything you're defining for coding agents. Like, how do we work on that? So feel free to reach out to me and I'll respond to you. And once again, it's very, very kind of you to give me your time. Thank you so much. Thank you.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note