Task Fidelity Scaling Laws — Kobie Crawdord, Snorkel
Description
Same model. Same compute. Same number of tasks. Fine-tuning on low quality tasks improved the base model by 1%. Fine-tuning on high quality tasks improved it by 6%. Kobe Crawford from Snorkel ran that experiment on TerminalBench style agentic tasks and got a 5x difference in training uplift from task quality alone. The talk breaks down what separates the two buckets. Accepted tasks averaged twice as many tool calls, lower pass rates, and more output tokens. Genuinely harder problems. More importantly, their failure modes were cleaner: when a model failed on a well specified task, it failed for a real reason. Rejected tasks tended to fail because of mismatches between what was requested and what the tests actually checked, or because the task never gave the model the context needed to satisfy implicit dependencies. Ambiguous specs do not produce harder tasks. They produce noise. Speaker info: - https://www.linkedin.com/in/kobie-crawford - https://snorkel.ai/author/kobie-crawford/
Summary
Generated by claude-sonnet-4-530-second take
Snorkel's research demonstrates that task quality in agentic RL training produces a 5x performance differential (6% improvement vs. 1%) when training with high-quality versus low-quality terminal-bench-style tasks, holding compute and task count constant. Kobe Crawford argues that task quality = data quality in agent training contexts. Their method: define acceptance criteria (achievable, non-trivial, functionally correct, reliable environment), categorize failure modes (logic errors vs. environmental issues), and show accepted tasks produce "cleaner failures" that models can actually learn from. The core thesis: expert-in-the-loop data generation plus rigorous task verification yields meaningful training signal, not just harder benchmarks.
Key takes
- 5x training efficiency gap: High-quality tasks improved base model performance by ~6% after RL training; low-quality tasks yielded only ~1% improvement with identical compute budget and task count. This suggests most benchmark tasks may be training on noise rather than signal.
- Accepted tasks are meaningfully harder, not just broken differently: High-quality tasks averaged 2x more tool calls, lower pass rates, and more output tokens. Critically, their failures skewed toward logic errors and incomplete reasoning rather than environmental bugs—failures that teach models useful behaviors.
- Underspecification is the primary task-quality killer: Rejected tasks often had mismatches between task descriptions and backend tests (e.g., implicit dependencies not mentioned in prompts, success criteria never stated). Models fail these for "wrong" reasons that don't improve performance.
- Snorkel's acceptance criteria: Tasks must be (1) achievable, (2) non-trivial, (3) functionally correct, (4) environmentally reliable. Only tasks passing all four tests enter training sets. This filters ~half of terminal-bench-style tasks in public benchmarks.
- Inter-annotator agreement via rubrics + LLM judges: Snorkel scales human expert labeling by building detailed rubrics that both humans and LLM judges use, then validates agreement across human-human and human-LLM pairs. This maintains quality at scale for less verifiable domains (e.g., emotional reasoning benchmarks).
Useful details
- Models tested: Claude Sonnet 3.5 and OpenAI Codex (GPT-4, 4.0-turbo variants) were used for task completion and failure mode analysis.
- Failure taxonomy: Logic errors, incomplete tasks, environmental failures (degenerate cases no model could solve), test-definition mismatches. Accepted tasks had higher logic-error rates; rejected tasks had higher environmental failure rates.
- Benchmark work: Snorkel built/maintains "agentic coding" benchmark (terminal-bench-style) and is analyzing public benchmarks (SWE-bench variants, Terminal Bench v1 vs. v2) for task quality issues. Found some tasks "never get completed" due to fundamental defects, masking model improvement signals.
- Open benchmark grants program: Partnering with orgs building evals in fuzzy domains (e.g., emotional reasoning, multi-outcome tasks) where correctness isn't binary.
- Containerized environments: Tasks run in isolated, parallelizable containers for reproducibility and rollout efficiency.
Caveats / counterpoints
- Verification bias: The research explicitly focuses on domains with verifiable outcomes (coding, math). Crawford acknowledges that "fuzzy" domains (emotional reasoning, multi-step open-ended tasks) are harder and require different approaches (rubric-based scoring, multiple valid outcomes). The 5x result may not transfer.
- Single training run per condition: Only two training runs (one low-quality, one high-quality) were mentioned. No variance estimates or multiple runs reported, so reproducibility/confidence bounds are unclear.
- Rejected tasks as "pure noise" assumption: The study treats rejected tasks as uniformly low-quality, but Crawford admits in Q&A that the effect of mixing rejected tasks with accepted ones is unclear—whether rejected tasks actively harm training or just add noise wasn't tested.
- Benchmark saturation masking: Crawford notes that poor task quality in public benchmarks can mask whether models are actually improving, but doesn't quantify how much of observed "saturation" in leaderboards is task-quality artifacts vs. real model plateaus.
- No comparison to synthetic task generation: All discussion assumes human-expert-in-the-loop task creation. No comparison to automated task generation or LLM-generated tasks.
Ken relevance
- Agent training pipelines: If you're building or fine-tuning agents (especially for coding, ops, or multi-step workflows), this suggests curating high-quality eval/training tasks may matter more than adding compute or task volume. A 5x efficiency difference is worth investing in task QA infrastructure.
- Data vendors / RL data: Snorkel's positioning as "Frontier AI Data Lab" selling curated datasets for foundation model training. If you're evaluating data vendors or building data pipelines for agent training, their acceptance criteria (achievable, non-trivial, functionally correct, reliable) are a useful rubric.
- Benchmark literacy: When evaluating agent models via public benchmarks (Terminal Bench, SWE-bench, etc.), be aware that task quality varies widely. Snorkel found many "never completable" tasks. If you're using benchmarks for investment/product decisions, dig into task-level failure modes, not just aggregate scores.
- Content/GTM: The "expert-in-the-loop" positioning and rubric-based scaling approach could inform how you think about human-AI workflows in content production or GTM ops—how to maintain quality while scaling with LLM judges.
- Low relevance for: Non-agentic use cases, domains without verifiable ground truth, or if you're not training/fine-tuning models.
Watch verdict
Skim. The 5x training efficiency finding is striking and the task-quality framework is useful if you're evaluating agent benchmarks or building RL pipelines. However, the talk is light on methodology details (single runs, no variance), and the Q&A reveals more caveats than the main presentation. The core insight—task quality = data quality in agent training—is valuable, but you can extract it from this summary without watching 20+ minutes of repetitive explanation and placeholder content.
Transcript
My name is Kobe Crawford. I'm a developer advocate at Snorkel. We are the Frontier AI Data Lab and what that means is that we produce data sets for foundation models to hill climb on. So our research team is highly integrated with the work that we do in terms of our production work and we put a lot of emphasis on how we integrate research in that. This company's origins actually begin from a Stanford University AI research lab and the work that they were doing there actually was part of one of the CEO's PhD thesis and then that became a library that was used open source for a while and then we've grown into focusing on delivering things with data sets for our customers. One of the things that's been a consistent through line for Snorkel since they got started as a company in 2019 is that the core thesis has been that the quality of data is critical and that the data that you're looking at you want to make sure is top quality in all those cases. So we look at how that applies to the data sets that we provide as well as as things move into the agentic space how that applies to agentic tasks and what we wanted to show in this context is how data quality impacts things in the context of task quality and the task quality and data quality are largely the same thing. So we're going to talk about the particular research objective here was looking at how task quality affects the training outcomes that you get when you're trying to improve models and then from there we're going to talk about the techniques and the path that we chose to verify that these behaviors were actually happening for us. So does the task quality actually matter? You can tell from where we are from our thesis about it that we actually expect that of course it does and the way that we're going to break down talking about that is to make sure we understand where we're approaching this. We're talking specifically in the context of agentic terminal bench style tasks, so we're working with a flow that is going to be a containerized environment and then a task definition within that. We want to show that when you're looking at how you look at the tasks themselves and how they're built that what we do in the agentic context is still also governed by the same data quality premise, so that applies in terms of talking about task quality as well as data quality and so if your architecture changes if the harness that you're using changes these kinds of things all of those things are also obviously impactful things but underlying all of that is still the data quality is at the center of it. So here what we're doing actually puts some specific kinds of rigor and delivering empirical evidence to validate that this is true. We want to not just say we accept this as a thing that we like to say is true. We actually want to verify that that's the case. So in the definition of task quality we're talking about four core things. If you've worked with these kinds of environments the harbor framework, open-end, when you've built tasks for agentic purposes what we're talking about in the context of evaluation benchmarking RL is that we are creating an environment in which that is going to run. We have it containerized for reproducibility and isolation and that also allows us to do parallelizing for rollouts of these kinds of practical elements of how that works. Inside of that you have looking at what's in the logic of the task. You want to talk about that the task is achievable, that is non-trivial, that is functionally correct, that the logic actually plays as expected and then the environment itself is reliable and so that environment reliability is as key as well. Those four criteria for us as we work on this the Snorkel team has built in our research harnesses a setup where we verify all four of those criteria and our tests to verify those criteria are the tests that we use. If a task passes all of those tests then we consider it an accepted task and then become something we can use for our training and research purposes and then if it's not accepted then it would be put in the rejected bucket and we use those two buckets as a basis for talking about how we're going to compare what is a high quality task—the accepted ones—to what is a low quality task. So let's look at those comparisons for just a moment and make sure that when we actually take a look at it we level set. Does the acceptance criteria that we use tend to correlate with actual performance behaviors that we think we want to see and how they differ? So the way that we did that was we used Sonnet 4.5, so obviously some months ago, and Codex which had GPT 5.2, 5.1 and sometimes 4.0 included in terms of the tests that were run but we used those two for running these tests and we compared how those tasks were completed and in the completions we found that our accepted tasks averaged twice as many tool calls, demonstrating more difficulty, more steps needed, and more engagement with external tools, a lower pass rate so higher difficulty intrinsically and then also more output tokens needed so there was more reasoning that was done by the models to actually do that. The failure modes however—so if we let's go back to the pass rate for a second—it's possible in the context of a lower pass rate that you could have failure modes that actually don't show meaningful signals so we actually wanted to dig into the failure modes a little bit as well. And so the next step here is to say okay what does it mean to be talking about those failures? Failures. To do that we broke down the failures into categories to identify the kinds of failures that represent something meaningful like this model is not performing the task completely because the model is not achieving a logical conclusion that it needs to versus a failure that's more of a degenerate case where you have a problem that is an environmental problem, something that literally makes it so no model would be able to solve that problem and not complete that task in this particular flow. So we have a breakdown of those and then given those again we wanted to compare the accepted versus rejected tasks and see where the failures are occurring in those tasks and see what we get from that. So here's a summary breakdown of each of these things across the percentage of failures that we saw in each of these different categories and you can see I like to highlight in particular the logic error and the incomplete task bars and just observe that in each case you can see a reversal of which one had a higher percentage appear in those there and then this breakdown even with this comparative analysis based on percentages of failures you can see the difference in terms of where you see the over representation or under representation of these kinds of failures across the types of failures that we categorized and again between where the rejected tasks versus where the accepted tasks had errors. The general tendency that we are taking away from this is that the accepted tasks are producing cleaner failures. These are failures due to the task itself being more difficult, truly more difficult, that the steps that it needs to accomplish are more difficult and that means that this is See a reversal of which one had a higher percentage appear in those there and then this breakdown. Even with this comparative analysis based on percentages of failures you can see that you can see the difference in terms of where you see the overrepresentation or underrepresentation of these kinds of failures across the types of failures that we categorized and again between where the rejected tasks versus where the accepted tasks had errors. The general tendency that we are taking away from this is that the accepted tasks are producing cleaner failures. These are failures due to the task itself being more difficult, truly more difficult, that the steps that it needs to accomplish are more difficult and that means that this is a test that would be actually very useful for the model to be able to hill climb on, provide some data samples that could help it actually be improved in terms of those performance behavior patterns versus something where it's just a failure that's not super meaningful in terms of it being just a tactical thing that's happening inside of the context that's not working. With that in mind we then take it as we've accepted that we've actually put together enough analysis that gives us a pretty strong sense that the accepted tasks are also higher quality tasks in the main and the same thing again that given that we have a differentiation between the higher quality tasks and the lower quality tasks now we want to actually see can we see an impact on model performance when we use it. Following forward from that we just actually run a training run, an RL training run with the same model, the same compute budget, the same number of tasks in each case and then look at the difference there. So that's where we've level set that we have a set of test tasks that we consider high quality, level set of tasks that we consider lower quality and that usage is going to help us say something about that. We trained it twice and we wanted to see what we got and the performance uplift is actually very meaningful. We're talking about a one percent improvement with using the low quality tasks. So after the RL training was done, the low quality tasks only improved the base model by about one percent improvement but the improvement was about a six percent improvement with the higher quality tasks. So that uplift of the 5x uplift difference based on just quality is really striking from our point of view. We think that that cements the intuition that it really is important for the data quality to be high. The way Snorkel generates datasets and the way that we put together RL environments we're using human expertise and having experts in the loop for generating the data and we have a strong feeling that the expert in the loop is an important element of delivering data quality and between those that ultimately gives you, we talk more about how our platform, how we use our platform to help make that work by having our experts be something that we scale and we can deliver quality at scale but the key is that we have, we want to put the emphasis on making sure that quality is the first thing that people think about, about what you need your data to, to what you need from your data to make sure that you're getting good results. So that's the end of what we got out of that. Hopefully if you have additional questions we have just a couple of additional minutes but here's a couple of quick links to the research page as a suffering summary, how our research team works and what we do and what we put our emphasis on, and then the leaderboard is pointing to a couple of benchmarks that Snorkel built and curates ourselves that are similar to things that you've seen. We have one called agentic coding which is focused specifically on these terminal bench style tasks and we're doing the same kinds of evaluation but we put apply a certain kind of rigor to how we're going about it that we want to make sure that people could see the difference. Thanks very much and then please can I take any follow-up questions. I'll start here. I'm just wondering if you would look at another, I guess, of all the tasks projected accepted to see if it's like, is there an effect of the rejected task pulling it back down or actually the model can get over it as long as all the tasks are good? A good question. The interesting thing about, so I don't have a specific quick direct answer about that analysis that we did in this particular case. But one of the things that we saw in surveying and working with, for example, working with the terminal bench team and looking at the tasks that were in terminal bench one versus what we did for terminal bench two and in some other contexts around some of the sweep bench, some of the variants of the sweep bench, we've been doing some analysis internally to compare the various public benchmarks that are out there and looking at what we see across those things. And when we look at that, you certainly see that the failure rates and where the models have been improving over time, whether they're getting saturation faster or the benchmarks themselves are getting saturation faster, sometimes we ended up seeing some noise about that because you end up with a certain number of tasks that never get completed so then we start to see that these tasks will never be completed literally just because they actually can't be right. And because of that we ended up finding that that was actually a source of noise in the process of actually evaluating whether the model improvement was actually happening. So in that way I can't say it's whether or not the models are actually improving ended up being something that ended up being more so masked by the quality of task issue as opposed to be something where we could tell did the models actually improve or not despite their presence. It'll be some more source of noise. I saw a couple other questions right here. Do you have any sense of how the input is making this task? You know there's a correlation sometimes and input can be a little better, less prescriptive could make it harder or it could be very prescriptive could make it easier. I don't know but I'm just wondering if you encountered something. Yes absolutely. In fact the way that we've been looking at it, a lot of times what makes a task one of the rejected tasks is it being underspecified in terms of when the task is defined in a way that the desired testable outcome is not clearly specified in the task definition up front but then on the back end the tests themselves expect certain things to pass that were never actually requested. [SPEAKER_03] Those kinds of mismatches are, so some of the places where you can see where the task becomes or at least appears to be harder because the tests don't match what the requested task setup is. [SPEAKER_03] Also sometimes there are implicit dependencies in the testing that the task doesn't specify be there and then without knowing that the dependency is required in the first place and what hasn't been fed into the context of the model then the model doesn't even have the right context to be able to approach those dependencies. Just a follow-up? [SPEAKER_03] the task definition up front but then on the back end the tests themselves expect certain things to pass [SPEAKER_03] that were never actually requested those kinds of mismatches are some of the places [SPEAKER_03] where you can see where the task becomes harder because the tests don't match what the requested task set up also sometimes there are implicit dependencies in the testing that the task doesn't specify be there and then without knowing that the dependency is required in the first place and what hasn't been fed into the context of the model then the model doesn't even have the right context to be able to approach those dependencies just a follow-up or yeah classifying them as a failure could be an issue because not every task needs to be complete there's iteration there's a journey typically when we solve problems it's never a one shot kind of approach in most of the problems in the world so differentiating that but the under specified could maybe certainly 100% agreed that those kinds of things about what we ultimately want the models to do tends to go like that in the context of building benchmarking tasks we work [SPEAKER_03] to make it so that we have something that's verifiable on the back end and that in principle if we [SPEAKER_03] do it right that the skills that are being learned are still going to be more applicable in the context [SPEAKER_03] of unverifiable results or things where there's going to be an iteration that needs to occur following [SPEAKER_03] the step that you're working on currently yeah so thank you sure yes in terms of future challenges next steps are you working on tasks that are not as straightforward verifiable and then maybe you're working on the right to be more to what's very involved in the horizon we certainly are looking at all of those and different projects of ours are in those different spaces especially once you get outside of the places where verification isn't easy and coding and math make it straightforward and then things that are more fuzzy are different we have an open benchmark grants program that we're working with we're partnering with folks who are developing benchmarks and evaluations in [SPEAKER_04] some more of the less verifiable areas and there's a lot of very interesting stuff that we're [SPEAKER_04] doing that has to do with one of them that there's an interesting organization that's working on trying to remember the name of it but it was about really looking at things that involve emotional level things and stuff like that's very lots of very human centric thinking and so there the even the notion of what is correct or not is something where we want to have actually multiple possible outcomes and then score them differently but have them all fit on the spectrum somewhere and so there's a lot that we're trying to do in the things in a lot of different dimensions so yeah it's very interesting space one last question in the back yeah I was just putting an expansion of this conversation that you would recommend? I'll speak to inter-annotator agreement centrally. And then the complexity you're talking about 100% is something that the longer the horizon, the multiple steps involved, and the different dimensions are an issue. This is the last question. Yeah, or are we done the time? Do I still have one minute left? You only need. OK, brilliant. So the way that our platform works, we actually do a number of things to bring together human annotators as well as using LLM judges. And that's partly to help us replicate and scale [SPEAKER_02] what our human annotators are delivering, but also these kinds of agreement. We feel like the way that we are doing things with rubrics these days and providing a longer list of data points and criteria that need to be met, that as we build out a set of rubrics, that then can be used both by LLM judges and people, we're actually looking at high-level qualitative things as well as individual more quantitative comparisons. And so through a longer list of things built on a rubric, and then using the human annotators and the experts to help us give us the information, the ground truth information that we can inform LLM judges to look at, we're actually looking to make sure, for example, that we actually test and get inter-annotator agreement very high between both individual humans as well as between the LLM judges and humans, and then use all of those comparisons to do quality assessment. So it's part of our assessment process, and we use that for each of these kinds of tests. So in the context here where it's explicitly verifiable, tests will pass or tests will fail, it's still obviously an easier domain than others, but we still keep using that guiding principle across all these domains. All right. Well, thank you very much. I really appreciate your time. It's great to have you here. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. We'll see you next time. the way the snorkel generates uh data sets uh and and the way that we put together rl environments we're using uh human expertise all uh and having experts in the loop for generating the data and we have a strong feeling that the expert in the loop is an important element of delivering data quality and between those that ultimately gives you uh you know we talk more about how our platform uh how we use our platform to help uh make that work by our experts be something that we scale and we can deliver of quality uh at scale but the but the key is that we have we want to put the emphasis on making sure that quality is the first thing that people think about about what you need your data to uh to what you need from your data to make sure that you're getting good results so that's the that's the end of what we got out of that um hopefully uh if you have uh additional questions uh we have just a couple of additional minutes but um here's a couple of quick links to the research page as a suffering summary how our summarizing how our research team works and what we do and what we put our emphasis and then the leaderboard uh is pointing to a couple of benchmarks that snorkel uh built and curates uh ourselves that are similar to things that you've seen uh we have one called agentic coding which is focused specifically on these terminal bench style tasks um and uh we're doing the same kinds of of of evaluation but we want we put apply a certain kind of rigor to how we're going about it that we want to make sure that people could see the difference thanks very much and then please can i take any follow-up questions i'll start here i'm just kind of wondering if you would look at um another i could be see i guess of like all the tasks projected accepted to see if it's like is there uh an effect of the rejected task is pulling it back down or actually the model can get over it as long as all the tasks um a good question um the interesting thing about so i don't have a specific quick direct answer about like from that analysis that we did in this particular case um but um one of the things that we saw in surveying and working with for example working with the terminal bench team um and looking at the tasks that were uh in terminal bench one versus what we did for terminal terminal bench two and in some other contexts around uh some of the sweep bench um some of the variants of the sweep bench we've been some doing some analysis internally to uh compare the various public benchmarks that are out there and looking at what what we see across those things and when we look at that um you certainly see that the uh the the the failure rates and like which you know where the where the model models have been improving over time um whether they're getting the saturation faster or the benchmarks themselves are getting saturation faster uh sometimes we ended up seeing sort of some noise about uh that because you end up with a certain number of tasks that never get completed so then we start to see that like these tasks will never be completed literally just because they actually can't be right yeah and because because of that we ended up like finding that that was actually sort of a source of noise in the in the process of actually evaluating whether the model improvement was actually happening so in that way i can't say it's like sort of like whether or not the models are actually improving ended up being something that ended up being more so masked by like the quality of task issue as opposed to be something where we could tell did the models actually improve or not uh despite their presence um it'll be some more source of noise um i saw a couple other questions right here um do you do you have any sense of how the input is making this task uh you know there's a correlation sometimes and input can be a little better less prescriptive could make it harder or it could be very prescriptive could make it easier i don't know but i'm just wondering if you encountered something yes absolutely in fact uh the way that we've been looking at it uh a lot of times what makes a task uh sort of a one of the rejected tasks is it being underspecified in terms of like you know when you when the task is defined in a way that the desired testable outcome is not uh clearly specified in the task definition up front but then on the back end the tests themselves expect certain things to pass that were never actually requested those kinds of mismatches are so the sort of some of the places where you can see where the task becomes uh or at least appears to be harder because the tests don't match what the the requested uh task set up um also sometimes there are like implicit dependencies in the in the testing that uh that the task doesn't specify uh be there and then in the without knowing that the dependency uh is required in the first place and what hasn't been fed into the context of the model then the model doesn't even have the the right context to be able to to approach those dependencies just a follow-up or yeah classifying them as a failure could be an issue because not every task needs to be complete you know there's a iteration there's a journey typically when we solve problems it's never a one shot kind of in most of the problem in the world so yeah uh differentiating that but the under specified could maybe i don't know certainly certainly uh 100 agreed that that those kinds of things about like what we ultimately want the models to do tends to go like that in the context of building benchmarking tasks we work to make it so that we have something that's verifiable on the back end and that that in principle if we do it right that the skills that are being learned are still going to be more applicable in the context of unverifiable results or things where there's going to be an iteration that needs to occur following the the step that you're working on currently yeah yeah so thank you sure yes um in terms of like future challenges next steps like are you working on tasks that are not as straightforward verifiable and then maybe you're working on the right to be more to what's the very involved in the horizon we certainly are looking at all of those uh and different projects of of ours are we're in those different spaces especially once you get outside of the places where verification isn't easy and coding and and math make it straightforward and then things that are more fuzzy are are different we're uh we have an open benchmark grants program that we're working with we're partnering with folks who are developing uh benchmarks and evaluations in in some more of the less verifiable areas and uh there's a lot of very interesting stuff that we're doing that has to do with uh one of them that uh there's a a an interesting organization uh that's working on um trying to remember the name of it but it was about sort of like uh really looking at things that involve like sort of emotional like level things and stuff like that's very lots of very human centric uh thinking and so there uh the even even the notion of like you know what is correct or not is something where we want to have like actually sort of multiple possible outcomes and then score them differently but you know but have them all sort of fit on the spectrum somewhere and so there's there's a lot that we're trying to do in the the very things in a lot of different dimensions um but yeah so it's very interesting space one last question in the back yeah i was just putting an expansion of um this conversation uh um uh uh uh uh uh uh uh uh uh uh uh uh uh uh that you would recommend? I'll speak to inter-annotator agreement sort of centrally. And then the complexity you're talking about 100% is something that the longer the horizon, the multiple steps involved, and the different dimensions are an issue. This is the last question. Yeah, or are we done the time? Do I still have one minute left? You only need. OK, brilliant. So the way that our platform works, we actually do a number of things to bring together human annotators as well as using LLM judges. And that's partly to help us replicate and scale what our human annotators are delivering, but also these kinds of agreement. We feel like the way that we are doing things with rubrics these days and providing sort of like a longer list of data points and criteria that need to be met, that as we sort of build out a set of rubrics, that then can be used both by LLM judges and people, that we're actually sort of like looking at high-level qualitative things as well as individual sort of like more quantitative comparisons. And so through a longer list of things built on a rubric, and then using the human annotators and the experts to help us give us the information, the ground truth information that we can inform LLM judges to look at, we're actually looking to make sure, for example, that we actually test and get inter-annotator agreement very high between both individual humans as well as between the LLM judges and humans, and then use all of those comparisons to do quality assessment. So it's part of our assessment process, and we use that for each of these kinds of tests. So in the context here where it's explicitly verifiable, tests will pass or tests will fail, it's still obviously an easier domain than others, but we still keep using that sort of guiding principle across all these domains. All right. Well, thank you very much. I really appreciate your time. It's great to have you here. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. We'll see you next time.