Open Reader

Agentic Evaluations at Scale, For Everybody — Nicholas Kang & Michael Aaron, Google DeepMind

completed 20:02 May 25, 2026 Watch on YouTube

Current Status

completed

Video ID

Ubwb6NzegyA

RAG / Chat

Enabled
Agentic Evaluations at Scale, For Everybody — Nicholas Kang & Michael Aaron, Google DeepMind
Description

On SWE-Bench Pro, six frontier models land within a couple of percentage points of each other. The harness they run inside shifts performance by 22%. A competing lab once took a Kaggle benchmark, reran it with their own compaction settings, and published much better results. Neither number was wrong. Both were useless. The talk is from Nicholas Kang and Michael Aaron at Google DeepMind's Kaggle team, who are building the infrastructure to fix evals at the community level: an open benchmark platform anyone can contribute to, a PvP Game Arena where models play poker and chess for an ELO rating that cannot saturate, and a standardized agent exam that returned 500 plus submissions in its first week without any promotion. The wastewater treatment plant engineer from Turkey who built a novel safety benchmark from 20 years of field experience, data that does not exist anywhere else, is the use case they keep coming back to. Speaker info: - https://www.linkedin.com/in/nicholaskangjj

Summary

Generated by claude-sonnet-4-5

30-second take

Kaggle/DeepMind is building four open-source products to fix AI evaluation's broken ecosystem: scattered benchmarks that go stale, opaque corporate testing, and research-only eval creation. They argue that 30,000 AI researchers can't capture what 30M+ technical professionals need measured—leading to superhuman AI in some areas but mediocrity in critical domains like wastewater safety. Their approach: democratize eval creation through hackathons, standardized agent exams, a PvP Game Arena with ELO ratings that never saturate, and a community benchmarks platform. They're candid about hard problems: statistical significance costs (400K poker hands), harness vs. model ambiguity, and incentivizing domain experts to build evals. This is not a production eval tool—it's community infrastructure for open, verifiable testing.

Key takes

  • Eval creation is dangerously centralized: Only ~30K AI researchers build benchmarks, leaving massive blind spots in domains that matter (e.g., a Turkish wastewater engineer created a proprietary safety eval after fatal incidents because no lab cares about economically unproductive niches). This creates "cognitive jaggedness"—superhuman performance in researched areas, mediocrity elsewhere.
  • Published benchmarks die immediately: 10+ benchmarks drop daily on arXiv, then go stale as authors move on. Leaderboards in papers become irrelevant because there's no maintenance or standardization—just a publish-and-forget cycle that wastes the community's ability to track real progress.
  • Corporate benchmarks are gamed and opaque: Competing labs re-run the same benchmark with model-specific optimizations (e.g., quantization tricks) and publish wildly different results. You can't verify configurations, test setups, or what's actually measured—charts in model release notes are marketing, not science.
  • Consumer agents ship untested into the wild: Hundreds of people build OpenClaw-style agents but don't evaluate safety or correctness before deploying them to manage email/Amazon accounts. Kaggle launched a one-line "standardized agent exam" last week and got 500+ agents tested without promotion—revealing huge unmet demand for accessible safety baselines.
  • Game Arena attacks saturation with PvP: Traditional benchmarks plateau when models max out scores. Having models play Werewolf (deception), poker (risk/randomness), and chess against each other creates an evergreen ELO leaderboard that's forever hill-climbable. Grok "loves to go all in," newer models are weirdly more risk-averse—personality differences emerge.
  • Harness matters more than the model: MorphLM's blog claims frontier models score within 2% of each other on SWE-Bench Pro, but harness choice causes a 22% performance swing. The eval infrastructure is the real variable—you're testing the wrapper, not the AI.
  • Cost and ambiguity block agentic evals: Running 400K poker hands for statistical significance gets expensive fast on Claude/Sonnet pricing. Add rapid model deprecation, opaque API versioning, and the harness-vs-model confusion, and longitudinal comparisons become nearly impossible.

Useful details

  • Kaggle Benchmarks platform: Open-source system where anyone can write assertions (e.g., "contains towel?"), use LLM judges, group tests into tasks, run against model collections, and publish results. Paige (previous speaker) built an XKCD SVG parsing task as example.
  • Agent Exams: Paste a one-line prompt → agent takes exam → get leaderboard score. Launched one week ago, already 500+ agents evaluated. Reddit spinoff: someone created an "SAE prep course" meme post.
  • AGI Faculty Hackathon: DeepMind published a paper on measuring 10 cognitive faculties; Kaggle launched a hackathon focused on 5 of them to crowdsource novel benchmarks from domain experts.
  • Wastewater benchmark: Turkish engineer with 20 years' experience built a proprietary dataset after fatal safety incidents in his country—knowledge that doesn't exist anywhere on the web or in lab priorities.
  • Bradley-Terry pairwise scheduling: Used in Game Arena to minimize the number of games needed for statistical significance, but still hit cost/scale issues.
  • LLM Model Proxy: Available on Colab—standardized way to talk to all models consistently. Game conversations exported as open datasets.
  • Original platform roots: Kaggle's simulation infrastructure was built for RL competitions pre-LLM era, now repurposed for agentic game evals.

Caveats / counterpoints

  • Not a production tool: Explicitly positioned as community infrastructure, not for enterprise CI/CD pipelines or shipping features. If you need prod evals, use BraidTrust or similar.
  • Human judgment still required: AI isn't good at evaluating innovation/creativity. Even expert alignment on what makes a good benchmark is hard—hackathon judging requires manual curation at scale.
  • Incentivization is unsolved: Getting domain experts to invest time building quality evals is difficult. Kaggle's points/medals system helps, but writing rigorous evaluations is labor-intensive with unclear payoff for participants.
  • Model endpoint opacity: Some providers aren't transparent about what model version is actually running behind the API, making apples-to-apples comparisons sketchy over time.
  • Statistical significance is expensive: Even with Bradley-Terry optimization, they needed 400K poker hands. Scaling to more games/models could become prohibitively costly without better sampling techniques.
  • Prompt fairness is iterative work: They spend "a lot of time" ensuring prompts are fair across models in Game Arena—not a solved problem, just ongoing tuning.

Ken relevance

High relevance for AI ops/eval infrastructure thinking. If you're building agent systems or evaluating frontier models for investment/product decisions, this talk exposes critical eval validity problems: harness effects dominating model differences, corporate benchmark gaming, and the impossibility of keeping pace with daily benchmark churn. The 22% harness swing on SWE-Bench Pro is a direct warning for any coding agent evaluation you're running—you might be optimizing the wrong thing.

Opportunity angle: The wastewater engineer story + 500 agent exams in one week signal massive latent demand for domain-specific and consumer-grade eval tooling. If you're productizing agents, there's a clear gap between "BraidTrust for enterprises" and "nothing for prosumers." Kaggle's open-source approach could become the de facto standard if they solve incentivization—worth tracking for GTM/ecosystem plays.

Content/investment angle: The saturation problem (benchmarks maxing out) vs. Game Arena's PvP approach is a useful framing for "how do we measure AGI progress when traditional metrics plateau?" Could inform content on post-benchmark AI evaluation or where to look for differentiation signals in model capabilities.

Watch verdict

Skim. The conceptual framing (eval centralization risk, harness vs. model ambiguity, consumer agent testing gap) is valuable and the wastewater/poker anecdotes are memorable. But the product walkthroughs are shallow—you'd learn more from the GitHub repos or live demos. Watch at 1.5x for the problem articulation; skip the "here's how our platform works" sections unless you're actively building eval infra.

Transcript

3813 words en Processed in 139.0s

All right, hi everybody. Let me just try to stand straight so I don't have to crouch over. Thank you all for coming. This is our talk on energetic evaluations at scale for everybody. I hope everyone's in the right room and if you are, thank you for coming. We were expecting 20 people so this is way more than what we expected. So all right, who are we? So I'm Nick. I'm a product manager in Kaggle Benchmarks and I run and build our benchmarks platform alongside a couple of our engineers and I also focus on our agentic eval solutions. I'm originally from Singapore but I live in the San Francisco Bay Area and I flew in to do this talk and attend all the great talks at this conference today. Hi, I'm Michael. I'm a software engineer on Kaggle. I've been working at Google for about a third of the time that I've been alive and Kaggle for about half of that. So mostly working on evaluations and benchmarks for Kaggle at the moment. All right. Has anyone here heard of Kaggle? Put your hands up if you have. Okay, great. So a lot of people know us for competitions but we don't just do that. We're the world's largest AI ML community of 30 plus million users and we've been working a lot in the Gen AI eval space over the course of the past two years. And we think there are lots of interesting problems in the industry that not many people are trying to solve and we feel positioned well to solve them. And we want to share more of the work that we've been doing and also invite contributions if you want to get involved in the space. So we have a simple agenda for today. First is AI evals today are broken and we'll talk about why. And then step two is we're trying to solve it. Not saying we're the all cure and we have everything sorted out. We'll talk about what we're trying to do, the challenges we're running into, and also maybe it might inspire some of you in terms of how you think you might be able to help contribute to this very important problem that we're trying to solve. With that, I'll jump into the first section. So first problem, evals are scattered, decentralized, and get stale fast. I don't know how many of you have tried to keep track of AI benchmarks but basically 10 plus of them drop every single day. And the best way to find out what they are is go to arXiv and spend hours scrolling through them reading every paper. That doesn't make sense. We don't think it makes sense. I can't even do it even though it's my full-time job. And I think what happens after that paper gets published is that you see some of the leaderboards in the papers. And what happens after that? It just gets stale. The authors move on to the next best benchmark because they just want to publish lots of papers. No fault of their own, but these leaderboards no longer become relevant as time goes on. The second issue is that evals aren't always transparent, accessible, and verifiable. I'm sure many of us have seen these charts on these model publisher notes when they release a new model. But what's the problem with that? We don't actually know how these benchmarks are set up. It's a lot of configurations you could use for the models themselves. And also how the benchmark is orchestrated and facilitated. And we don't always know what's actually being tested here. I'll give you one real anecdote, which is that we have published a benchmark with one of these AI labs. And another competing AI lab came to us and said, hey, we don't like the results of this particular benchmark you published. Let's run it on our own. And so they ran it and then they published it with much higher, much better results. And the difference was that they were optimizing it for their model. So they had used quantization that they had provided through their API. And we did it for all the models we ran. So the results you're seeing don't always reflect the actual state of things. And that's a problem. Number three, big circle represents all the world and its knowledge. Small circles represent AI researchers and technical professionals. There are something like 30,000 AI researchers. At least that's what Google AI search told me. And I think there's 30 million software engineers, data scientists, technical folks. We expect AI to help most of humanity. But then a very small percentage of people are creating all these evals. And if something's not being evaluated, not being benchmarked, we cannot hill climb on it. We cannot know how good we are at those things. And what will this lead to? More of these cognitive edges or jaggedness with these models as we're already seeing now. And that's only going to get more exacerbated over time as we see superhuman intelligence in some areas and just very mediocre performance in other areas. That's not equitable AI that we want that will benefit all of humanity. This is a fun anecdote, but not many AI researchers are also wastewater treatment plant engineers. I give you this example because this is an actual benchmark built by one of our users. He lives in Turkey. He's been a wastewater plant engineer for 20 years. I don't know what that entails, but it's clearly a very important job. Because he built this benchmark because he cares. And he recounted a story where he's been doing this for 20 years. There have been severe incidents in his country where there was some incident. People didn't follow the safety protocols. And people ended up dying as a result of that. So he built this benchmark to evaluate how AI could help him in his job and help avoid these incidents in the future. So this is a proprietary novel data set that he's created from his own experience. Doesn't live anywhere else on the web. Doesn't live in any of the AI lab focus area because that's not something that's economically productive for them at the moment. So it's very important why we think we should work on open source contributions to the eval space. All right. Enough of that. Talking about the solutions ahead. We're working on a couple solutions, but it's tough. So I'll very quickly cover the first two at the top. And then I'll hand it over to Michael to deep dive into the remaining two products at the bottom. So at the top left, we have hackathons. We have a platform that lets anybody host a hackathon. And I'll talk about why I think it's relevant to the problem of evals. Second, on the right-hand side, we have agent exams. We want to democratize the process of taking evals. We heard a lot about open claw and consumer agents being a big thing. Talking about the solutions ahead. We're working on a couple solutions, but it's tough. So I'll very quickly cover the first two at the top. And then I'll hand it over to Michael to deep dive into the remaining two products at the bottom. So at the top left, we have hackathons. We have a platform that lets anybody host a hackathon. And I'll talk about why I think it's relevant to the problem of evals. Second, on the right-hand side, we have agent exams. We want to democratize the process of taking evals. We heard a lot about open claw and consumer agents being a big thing. But the problem is that most people don't care about evaluating their own open claw agents, which is crazy. Number three, at the bottom left, we have Game Arena. That's where we have this evergreen benchmark when models are playing PvP games against each other. So it's an ELO score type rating, and it's forever hill climbable and unsaturated because you're just fighting against each other. And there must be one winner and one loser. At the bottom right, we have benchmarks. So it's the product I run, which is basically a platform that enables anybody to build, run, and share evals to the open community. So very quickly, hackathons. Why I think it's important. Hackathons are a great way to channel people's energy and expertise to solving a problem. I think we've seen that with the right energy, investment, and time, we can do a lot with very little. I think a great example is the world galvanized over the past three years to make Gen AI happen. And this is a very small form of how we want to help facilitate that process towards solving the right problems in this space. It's important that we put guardrails around the problems that we're trying to solve so that people don't go crazy, but also give them enough space so that they can flourish and their creativity can show. And the results of everything will be open source for the benefit of everybody and not just a small group of people. But there are some challenges. And actually, before I dive into the challenges, so on the screenshot on the right is a hackathon that we're actually running right now with the Google DeepMind AGI team. So Google DeepMind a couple weeks ago published a paper on how we can measure the cognitive faculties of AGI. And so we started this hackathon to focus on five particular faculties of the 10. And we want people to build benchmarks in those areas. And we want to give everybody the chance to contribute to AI research and not just a few people. And also knowing that everyone has something unique to contribute that these AI labs couldn't do themselves. But running a hackathon platform isn't all that easy either. I think these are fairly self-explanatory. Maybe I'll talk about the second and the third one in particular. The second one being that we need to provide them the right tools in order for them to do their best work. It might sound trivial, but things like, if you have a thousand participants, everyone's operating globally online. How can we give them the tools like hosting their own data sets to have access to AI models? That's something I never thought about before. But a lot of people coming from a poorer background might not have money to pay for these five API keys to access all these state-of-the-art models. And how can we let them share their work in a way that's understandable through write-ups so others can see, perceive, understand, learn, and build on the work that they've been doing? And then the third point is that as much as AI agents and everything you hear about this conference are very good at a lot of things, they're not very good at judging innovation and creativity. So a lot of the work still requires human experts. And even alignment within experts is difficult and not trivial. And that's something that we have to facilitate as part of this platform too. The second thing that we're working on right now is what we're calling standardized agent exams. I originally called it SATs, standardized agent tests. But there was a trademark issue, so I had to change the name. It's a true story. So how it worked is that you just paste a one-line prompt to your agent, and essentially it takes an exam, and we return a score for you on a leaderboard that you can compare its performance against. This was a very experimental MVP that we just launched last week. And I think this is important because if you look at AI evals that people are doing, it's at two ends of the spectrum. You have research labs and enterprises using BraidTrust, using all these state-of-the-art technology to set up to measure their agents and their models. And then on the other end, you have consumer agents, people who are building open claw and then filing 1,100 security advisories this morning, as we heard. But most of them aren't actually testing their agents before they're sending them out to the real world, which I think is a huge problem, as we've seen, and will become even more important. So one conversation that came up this week was how can we maybe do more safety-focused exams so you can do a quick baseline of your agent before you send it out into the world to run your inbox, to run your Amazon accounts, and to do stuff for you. So I talked about the first point. I think the second and third one are quite interesting. On the second one is that when something's accessible, we want to make sure that it's also challenging enough. And so we have this spectrum. If we make something too difficult, people can take the exam, but no one finishes it because it runs for too long. It's too difficult. But if we make it too easy, then it doesn't give you the right signal for what you want to measure. And then finally, the chart at the bottom shows there is maybe a market for agent consumers. We only launched a week ago, and we have hundreds, like 500-plus agents already evaluated on our exam without us even really promoting it very much. So I think that's an interesting insight that we've gleaned from this experience. On the right-hand side, on the screenshot that you see there, we posted in MoteBook, and then we started seeing these weird spin-off posts about people sharing agents sharing their exam results, and even an SAE prep course that came up on MoteBook. So that's always interesting to see. With that, I'll hand over to Michael. Cool. Yeah, thanks, Nick. Yeah, so I really wanted to talk about Game Arena and benchmarks with you all, AI engineering, and mostly give an overview of how it works, some of the really cool things we've seen in it, [SPEAKER_00] but mostly I want to talk about the challenges that we've been having with them. [SPEAKER_00] And so please come find me at the DeepMind booth after this to talk about them. [SPEAKER_00] and even an SAE prep course that came up on MoteBook. [SPEAKER_00] So that's always interesting to see. [SPEAKER_00] With that, I'll hand over to Michael. [SPEAKER_00] Cool. [SPEAKER_00] Yeah, thanks, Nick. [SPEAKER_00] Yeah, so I really selfishly wanted to talk about Game Arena and benchmarks with y'all, AI engineering, and mostly give an overview of how it works, some of the really cool things we've seen in it, but mostly I want to talk about the challenges that we've been having with them. [SPEAKER_00] And so please come find me at the DeepMind booth after this to talk about them. [SPEAKER_00] I would love some more insight on these things, too. [SPEAKER_00] So for Game Arena, it's a benchmarking platform. [SPEAKER_00] One of the problems with benchmarks is they very quickly get saturated. [SPEAKER_00] We see this with community benchmarks. [SPEAKER_00] We see it with AI and researching benchmarks as well. So Game Arena is an approach to help us work against the saturation by just having PvP, and so you can never have saturation because you'll always have one model able to compete against others. So saturation might just be, for a while, a model is the best. Really quick engineering slide here of how is all this set up, and how does it work? So when we're trying to figure out games to put into Game Arena, we want to analyze separate capabilities of AI models, and so we try to pick good, varied games. So far, the ones we've invested a lot in are Werewolf to play around with what's best at deception, poker for the randomization, also some of the deception, and how good Grok loves to go all in on poker. Less models are a little bit less crazy or a little bit more conservative. Most interesting, some of the newer generation of models are worse at poker because they are more risk-averse, and so you just see these personalities start to emerge over time. And then chess, because any time you're analyzing ML things, you have to be analyzing chess. And so for the quick overview of how all this works, we design and iterate on a game, figure out a good game we want to do, make sure models can actually play it, and then spend a lot of time on iterating prompts to try to make sure that our prompts are fair. And this is all open source and very viewable for what's happening. So if you want to check it out, I put the GitHub link there. Really, all of the things that we've talked about in this presentation so far are live on Kaggle. A lot of them are open source. Please come look at it, play around, give us feedback. We love it. So after that, we work on building this harness. Mostly we've done OpenSpiel games so far, which is an RL framework. And so we test, are these models better at playing at random? Sometimes they are. Usually they are, but sometimes not that much better. Are there any other interesting properties that emerge? Finally, we end up running the simulations. We use LLM model proxy. This is actually available on Colab if anybody uses that to just talk in a consistent way to all of the models that we want to run these games against. And so it runs on top of the Kaggle simulation platform, which was initially an RL platform that we had for Kaggle before LLMs became a huge thing. We schedule game runs. It uses Bradley Terry pairwise to try to not have too many games we have to run. I'll talk about that in a second. And then finally we publish the results. We have all these LLM conversations. We stick them in a data set. That's available on Kaggle. People can check it out and learn things from that. We put it on a benchmarks to show the ELO scores. And then we also have a game visualizer for all these things so you can go and see Grok go all in on PokerHands. That's a little demo to the left of that. So now, as promised, some of the really big challenges that we've had with this. And so love to hear all y'all's ideas and different ways to approach this. You can imagine it gets very expensive very quick. I'm sure you all have seen your Claude 4.6 bills are something along those lines. So you can imagine that for poker, in order to get statistical significance, we had to run about 400,000 poker hands. And there's many turns inside of each of those hands. You can imagine what those bills start to look like. And so the Bradley Terry pairing is part of this. But any way that we can get statistical significance without having to run millions of games is great. And we're always trying to think of new ways to be able to be sure that the models are best at the things that we're claiming they're best at. But running as quickly as possible. It gets a little boring to just watch LLMs play against each other all the time. For some of the initial games, it's pretty fun. But as we're trying to build this out, it might get a little bit repetitive. And so trying to figure out ways to engage Kaggle's community to be able to participate in this process. And so, even things like, could we have a hackathon that somebody provides a prompt for, for example. And then we would have prompts given by our community as part of the competition to play these games and see who prompts the model the best and climbs on a leaderboard. And then comparison over time is difficult. Old models disappear. New models come along. Sometimes when you're talking to a model endpoint, if you're not talking to yourself, they're not exactly honest on what model is happening in the background. So that's always a little bit of a difficulty. But yeah, go check it out. Pretty fun. prompts given by our community as part of the competition to play these games and see who prompts the model the best and climbs on a leaderboard. And then comparison over time is difficult. Old models disappear. New models come along. Sometimes when you're talking to a model endpoint, if you're not talking to yourself, they're not exactly honest on what model is happening in the background. So that's always a little bit of a difficulty. But, yeah, go check it out. Pretty fun. So, yeah, moving on to benchmarks with my limited time here. So what is this not? Is this not a production evaluation platform? There's plenty of people talking about that over in Moore if you want to go and run this for your production code. Very cool things. This is much more about community involvement. Anyone can build and run and share evals in a hopefully open and verifiable way. For time's sake, I'll just talk about this really a second. But we basically, it looks very similar to the production evaluations platform where you write some assertions. So this thing worked. You know, what gets wetter as it dries? You can say, okay, does this contain a towel? We also do LLM judging, similar to the production platforms. These all get grouped together in a task. They then get evaluated against a collection of models that the users want to run against. And then all these tasks get aggregated together in a benchmark, such as the wastewater treatment one that Nick was talking about previously. So Paige, who just presented in this room before this, actually made this a nice little task for us. It was parsing an SVG from XKCD. And it's like, can you recreate this SVG? And so you can kind of see the code for this is over on the left. One of the models is a little bit outdated, so Sonic 4 created this nice reproduction below it. And then Paige created a number of assertions. So, you know, can it generate an SVG at all? Does it have the correct text? And some other checks. And then an easy way to compare these things side by side. So, yeah, some of the challenges with this, of that, inspiration and incentivization are hard. If you want to, it's not hard for a production evaluations platform because you're shipping a thing to consumers. They care about, does this model work or not? But to just inspire people in the community to create benchmarks that other people find interesting. We've had good luck with hackathons. And, you know, Kaggle has a points and metal system and things like that. So we have some things baked in the platform to help inspiration and incentivization. But it just takes a lot of work to write a good evaluation. For agentic benchmark, yeah. So when we started this, people were really interested about analyzing just models. As we've moved more on to what are agents doing, it gets really hard to figure out what we're actually testing against. So I pulled out this little thing from MorphLM paper, blog post that they published on March 16th that basically called out that, you know, against SweetBench Pro, the six frontier models are within a couple of percentage points of each other. Definitely go check out this blog post. I haven't specifically verified it, but it does seem likely to me. The thing that really matters a lot for coding performance is what harness is it running inside of with a 22% difference depending on the harness. And so that can get really tricky of, are you testing the harness? Are you checking the model? Things like actual ambiguity under test. And then, again, fast release and deprecation cycles of models, it gets a little bit tricky to figure out what we're, or to be able to do comparisons over time. So, yeah, I think that's about time. But as I mentioned, we'll all be at the Dmine booth for the in-between times. Or you can email either me or Nick here. But thank you so much. Thank you. Thank you. Thank you. And so trying to figure out ways to engage Kaggle's community to be able to participate in this process. And so, you know, even things like, could we have a hackathon that somebody provides a prompt for, for example. And then, like, you know, we would have, like, prompts given by our community as part of the competition to, like, play these games and see who prompts, like, the model the best and, like, climbs on a leaderboard. And then comparison over time is difficult. Old models disappear. New models come along. Sometimes when you're talking to a model endpoint, if you're not talking to yourself, they're not exactly honest on what model is happening in the background. So that's always a little bit of a difficulty. But, yeah, go check it out. Pretty fun. So, yeah, moving on to benchmarks with my limited time here. So what is this not? Is this not, like, a production evaluation platform? There's plenty of people talking about that over in Moore if you want to go and run this for your production code. Very cool things. This is much more about community involvement. Anyone can, like, build and run and share evals in a hopefully open and verifiable way. For time's sake, I'll just talk about this really a second. But we basically, it looks very similar to the production evaluations platform where you write some assertions. So, like, this thing worked. You know, what gets wetter as it dries? Like, you can say, okay, does this contain a towel? We also do LLM judging, similar to the production platforms. These all get grouped together in a task. They then get evaluated against a collection of models that the users want to run against. And then all these tasks get aggregated together in a benchmark, such as the wastewater treatment one that Nick was talking about previously. So Paige, who just presented in this room before this, actually made this a nice little task for us. It was parsing an SVG from XKCD. And it's like, can you recreate this SVG? And so you can kind of see the code for this is over on the left. One of the models is a little bit outdated, so Sonic 4 created this nice reproduction below it. And then Paige created a number of assertions. So, like, you know, can it generate an SVG at all? Does it have the correct text? And some other checks. And then an easy way to compare these things side by side. So, yeah, some of the challenges with this, of that, like, inspiration and incentivization are hard. If you want to, like, it's not hard for a production evaluations platform because you're, you know, shipping a thing to consumers. They care about, like, does this model work or not? But to just inspire people in the community to create benchmarks that, like, other people find interesting. We've had good luck with hackathons. And, like, you know, Kaggle has a points and, like, metal system and things like that. So we have some things baked in the platform to help inspiration and incentivization. But it just takes a lot of work to write a good evaluation. For agentic benchmark, like, oh, yeah. So when we started this, people were really interested about analyzing just models. As we've moved more on to what are agents doing, it gets really hard to figure out, like, what we're actually testing against. So I pulled out this little thing from MorphLM paper, like, blog post that they published on March 16th that basically called out that, you know, against SweetBench Pro, the six frontier models are within a couple of percentage points of each other. Definitely, like, go check out this blog post. I haven't specifically verified it, but it does seem likely to me. The thing that really matters a lot for coding performance is what harness is it running inside of with, like, a 22% difference depending on the harness. And so that can get, you know, really tricky of, like, are you testing the harness? Are you checking the model? Things like actual ambiguity under test. And then, again, fast release and deprecation cycles of models, it gets a little bit tricky to figure out what we're, or, like, to be able to do comparisons over time. So, yeah, I think that's about time. But as I mentioned, we'll all be at the Dmine booth for the in-between times. Or you can email either me or Nick here. But thank you so much. Thank you. Thank you. Thank you.