AI Engineer

The Art & Science of Benchmarking Agents — Vincent Chen, Snorkel AI

2863 summary words 13 min summary Watch video

Start with the signal

13 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Effective agent benchmarks require rigorous scientific foundations (task quality, distributional control, model headroom, robust methodology) plus strategic art (clear thesis, field roadmaps, researcher UX) to shape AI progress rather than merely measure it
  • Why it matters: Ken needs to understand what makes benchmarks actually useful for evaluating agent systems in production vs. academic hill-climbing, especially as Snorkel deploys $3M in grants and filters 120+ benchmark proposals
  • Best use: Extract the evaluation framework for assessing agent benchmarks Ken encounters, understanding the gap between capability vibes and production-ready measurement, and identifying frontier directions (environment complexity, autonomy horizon, output complexity)

Executive Summary

Vincent Chen from Snorkel AI presents a meta-framework for building benchmarks that actually shape agent capabilities rather than just measure them backward-looking. The core insight: there's a widening gap between agent capabilities (which are genuinely improving, especially in coding) and enterprises' willingness to deploy them, driven by inadequate measurement tools. Snorkel sits at the intersection of frontier research and real enterprise deployments (finance, insurance, healthcare), giving them unique visibility into what measurement gaps block production adoption.

The framework splits into 'science' (empirical rigor) and 'art' (field-shaping strategy). Science requires: (1) individual task quality validated by domain experts with adversarial QC (example: GPQA's multi-reviewer protocol with incentive structures), (2) intentional distributional diversity covering real-world task distributions or paradoxically rare but critical failure modes, (3) unsaturated difficulty with model headroom (Arc AGI remained sub-1% even for frontier models), and (4) robust evaluation beyond accuracy to measure cost, latency, policy adherence, reasoning quality.

The art differentiators for frontier-shaping benchmarks: (1) a thesis on future capabilities (TerminalBench betting on CLI as general computer-use interface proved prescient), (2) spawning research roadmaps (SWE-bench inspiring families of derivative benchmarks across multilingual/multimodal coding), and (3) researcher UX that makes adoption frictionless (Helm's standardized harness, TerminalBench 2.0's Harbor infrastructure). Chen argues this product-thinking around benchmark interfaces is severely underrated but critical for adoption.

Snorkel's forward-looking bet identifies three frontier dimensions: environment complexity (capturing org-specific policies, distributed state, human-in-loop context that current benchmarks miss), autonomy horizon (long-running agents with continual learning and mid-stream requirement changes), and output complexity (moving beyond text to artifacts, uncertainty signals, nuanced rewards for training). The $3M Open Benchmarks Grant program (120+ applications reviewed) actively funds work in these directions, with Chen emphasizing benchmarks should define progress, not just snapshot it.

Key Takeaways

  • Claim: Enterprises hesitate to deploy agents despite improving capabilities because measurement lags behind actual capability | Evidence: Snorkel works with finance, insurance, healthcare deployments where 'stakes are high' and sees reluctance to let agents loose despite vibes improving and model cards showing progress, particularly in coding | Caveat: Chen doesn't quantify the measurement gap or provide specific deployment failure examples; assertion based on Snorkel's client conversations rather than controlled data | Implication: For Ken: the bottleneck to agent deployment at scale isn't purely capability—it's trusted measurement. Building or selecting evaluation tools becomes a deployment prerequisite, not an academic exercise | Timestamp: 02:30
  • Claim: GPQA introduced adversarial quality control with multi-reviewer protocols and incentive payouts to ensure graduate-level tasks were tractable yet rigorous | Evidence: Original author, multiple reviewers, adjudicators, revision cycles; payouts tied to inter-rater agreement; inspired by peer review process; remained on model cards years later as unsaturated benchmark | Caveat: Chen calls academic peer review 'flawed as it is,' suggesting the mechanism has limits; doesn't specify GPQA's exact agreement thresholds or rejection rates | Implication: High-stakes agent evals need adversarial validation where single-expert judgment insufficient; incentive design matters for quality at scale; appendix innovations often more valuable than headline metrics | Timestamp: 07:45
  • Claim: Arc AGI's intentional unsaturation and model headroom design captured reasoning leaps invisible to other benchmarks | Evidence: Remained unsaturated for months/years; when O1-style reasoning emerged 18-24 months ago, Arc showed massive capability jumps correlating with actual model improvements; Arc AGI 3 launched with frontier models under 1% while all tasks human-solvable | Caveat: Chen doesn't address whether Arc's specific task types (visual pattern reasoning) generalize to other agent domains or if it's a narrow capability proxy | Implication: Benchmarks designed for headroom act as early warning systems for capability shifts; for Ken's agent work, picking evals that aren't already saturated reveals when architectural changes (like reasoning loops) actually matter | Timestamp: 11:20
  • Claim: TerminalBench's bet on CLI as general computer-use interface proved prescient as Claude/Codex built enterprise capabilities on CLI-based tools | Evidence: Benchmark made early thesis that CLI would be core abstraction for agents; now appears on all recent model cards; teams at Claude and Codex building general-purpose enterprise capabilities on coding/CLI tools | Caveat: Unclear whether TerminalBench shaped these architectural choices or merely predicted them; doesn't address GUI-based agent approaches or whether CLI limits certain use cases | Implication: Benchmarks with strong theses on future interfaces can accelerate field convergence; for Ken: evaluating agents against TerminalBench implicitly tests CLI-first assumptions worth questioning per use case | Timestamp: 16:50
  • Claim: SWE-bench spawned a family of derivative benchmarks (Lite, Verified, Pro, Multilingual, Multimodal) and new research directions because it provided a roadmap | Evidence: Simple idea leveraging existing PRs inspired multiple benchmark variants; evolved thinking on coding agents; continues to be relevant and extended | Caveat: Chen doesn't discuss whether proliferation of SWE-bench variants fragments evaluation or whether Lite/Verified/Pro measure meaningfully different capabilities vs. just difficulty levels | Implication: Benchmark extensibility and forkability signal health; for Ken's work, watching which benchmarks spawn families indicates where research momentum concentrates and which eval paradigms have legs | Timestamp: 18:25
  • Claim: Researcher UX is severely underrated: Helm pioneered standardized harnesses, TerminalBench 2.0 shipped Harbor as de facto agent evaluation infrastructure | Evidence: Helm (Stanford CRFM) provided modular scenarios for reproducible evaluation; Harbor from TerminalBench 2.0 became standard for teams building agents; ease of extension drives adoption | Caveat: Doesn't compare adoption metrics pre/post infrastructure improvements or address whether standardization constrains innovation in evaluation methodology | Implication: For Ken: benchmark infrastructure choices lock in evaluation paradigms; prioritize benchmarks with good harnesses/APIs; if building internal evals, invest in developer experience or adoption will stall | Timestamp: 20:15
  • Claim: Real coding environments have org-specific policies, Slack context, flaky CI, human reviewer preferences, parallel contributors—benchmarks capture only a fraction of this complexity | Evidence: Lists specific production realities (screenshots, distributed CI, tribal knowledge, concurrent work) absent from current benchmarks as typical failure points for agents | Caveat: Doesn't propose how to benchmark these soft factors (human preferences, Slack context) in reproducible ways or whether simulating them is tractable | Implication: Current coding benchmarks (including SWE-bench family) systematically underestimate production difficulty; for Ken: internal evals must inject org-specific messiness or agent performance will degrade post-deployment | Timestamp: 23:40
  • Claim: Autonomy horizon benchmarks should test agents operating for weeks with changing requirements, reorgs shifting priorities, and continual learning under distributional drift | Evidence: Customer experience agents losing context from weeks prior; product spec changes mid-task; organizational priority shifts—real-world settings involve state changes over time | Caveat: No existing benchmarks mentioned that successfully test these long horizons; unclear if simulation environments can capture reorg/priority-shift complexity or if they require live production traces | Implication: Short-horizon benchmarks (minutes/hours) miss critical failure modes; Ken should distinguish copilot (short-loop) vs. autonomous (long-loop) agent evaluation needs and budget for longitudinal testing | Timestamp: 25:10

Detailed Brief

The Measurement Gap Blocking Agent Deployment

  • Claims: Capabilities are genuinely improving (vibes shifting, especially in coding); Enterprises hesitant to deploy agents in high-stakes environments despite progress; Measurement/evaluation capabilities lag behind actual model capabilities; Closing evaluation gap requires field deployments, red teaming, human evals, crowdsourcing, AND open benchmarks
  • Evidence: Snorkel deploys in finance, insurance, healthcare where 'stakes are high, not just about a number'; Model cards show hill-climbing progress; practitioners feel improvement; Open benchmarks like TerminalBench, Meters (long horizon), Arc AGI set critical guideposts; Best benchmarks define progress forward, not just snapshot backward
  • Caveats: No quantification of measurement gap or deployment failure rates provided; Claims based on Snorkel client conversations, not controlled studies; Doesn't separate technical measurement challenges from organizational risk tolerance
  • Implications: Agent deployment bottleneck is trusted measurement, not pure capability; Ken needs evaluation strategy as deployment prerequisite; Open benchmarks complement but don't replace private evals/red teaming; Benchmark selection should prioritize field-defining over backward-looking metrics

Science: Four Pillars of Rigorous Benchmark Construction

  • Claims: Individual task quality requires domain expert validation with adversarial QC; Distributional diversity demands intentional taxonomy and task distribution; Difficulty/model headroom ensures benchmarks remain unsaturated and separate frontier models; Robust methodology measures what actually matters (cost, latency, policy adherence, reasoning quality) beyond accuracy
  • Evidence: GPQA: multi-reviewer protocol, revision cycles, payout incentives for agreement, remained on model cards years later; MMLU: 57-domain taxonomy across STEM/humanities for graduate/professional knowledge; Arc AGI: unsaturated for months, captured reasoning leap when O1-style models emerged, Arc AGI 3 launched <1% for frontier models; TauBench: measured task completion AND policy constraint adherence (booking right flight with wrong fare class still fails)
  • Caveats: GPQA's peer review mechanism acknowledged as 'flawed'; details on rejection rates/agreement thresholds not provided; MMLU taxonomy now several years old; unclear if still representative of frontier professional knowledge; Arc AGI's visual pattern tasks may not generalize to all agent reasoning domains; Measuring multiple dimensions (cost, latency, quality) increases eval complexity and infrastructure requirements
  • Implications: Single-expert validation insufficient for high-stakes benchmarks; adversarial review catches blind spots; Taxonomy design phase often determines whether benchmark captures real-world distribution; Headroom design = early warning system for architectural capability shifts; For Ken: multi-dimensional evaluation (not just accuracy) required for production-grade agents; budget infrastructure accordingly

Art: Three Differentiators for Frontier-Shaping Benchmarks

  • Claims: Great benchmarks embed thesis on future capabilities/subspace; They spawn research roadmaps and derivative benchmarks; Researcher UX (ease of running models, contributing tasks, leveraging signals for RL) severely underrated but critical for adoption
  • Evidence: TerminalBench thesis: CLI as general computer-use interface; proved correct as Claude/Codex built on CLI tools; SWE-bench: simple PR-based idea spawned Lite/Verified/Pro/Multilingual/Multimodal variants, new research directions in coding agents; Helm (Stanford CRFM): standardized modular harness for reproducible evaluation; TerminalBench 2.0 Harbor: became de facto evaluation infrastructure for agent builders
  • Caveats: Unclear if TerminalBench shaped architectural choices or merely predicted them; doesn't address GUI-based alternatives; SWE-bench variant proliferation could fragment evaluation; not discussed whether Lite/Verified/Pro test different capabilities or just difficulty; Researcher UX investment competes with research novelty; no adoption metrics comparing pre/post infrastructure improvements; Standardization via harnesses may constrain evaluation methodology innovation
  • Implications: Benchmarks without clear thesis become measurement exercises, not field-drivers; Extensibility/forkability signals benchmark health and research momentum; For Ken: prioritize benchmarks with good APIs/harnesses; if building internal evals, product-thinking on developer experience determines adoption; Thesis clarity helps Ken filter which benchmarks align with his architectural bets

Frontier Directions: Three Axes for Next-Generation Benchmarks

  • Claims: Environment complexity: benchmarks must capture real-world messiness (org policies, distributed state, human-in-loop, concurrent contributors); Autonomy horizon: test long-running agents with continual learning, mid-stream requirement changes, weeks-long context; Output complexity: move beyond text to artifacts, uncertainty signals, nuanced rewards usable for training
  • Evidence: Coding example: real codebases have Slack context, flaky CI, human reviewer tribal knowledge, parallel contributors—current benchmarks miss this; Customer experience agents: lose context from weeks prior, product specs change mid-task, reorgs shift priorities; Strategic recommendations: non-trivial to verify 'good' proposals; requires capturing org context and human judgment nuance; Trustworthy outputs: agents need to surface uncertainty, ask for clarification, stop when unsure
  • Caveats: No existing benchmarks successfully test these dimensions at scale; Simulating org-specific policies/reorgs/human preferences in reproducible ways—tractability unclear; Long-horizon evaluation increases cost and infrastructure requirements significantly; Nuanced reward signals for complex outputs (strategic recommendations, proposals) subjective and hard to automate
  • Implications: Current benchmarks (including SWE-bench) systematically underestimate production complexity; For Ken: internal evals must inject org-specific messiness or expect post-deployment degradation; Distinguish short-loop (copilot) vs long-loop (autonomous) evaluation needs in system design; Output complexity dimension opens research on agents producing structured artifacts, not just text responses; Snorkel's $3M grant program signals where academic/industry focus will concentrate; Ken should track funded projects

Notable Concepts & Terms

  • Evaluation gap: The widening distance between actual agent capabilities (which are improving) and ability to measure/trust those capabilities in production, blocking enterprise deployment despite technical progress
  • Model headroom: Intentional design of benchmark difficulty to remain unsaturated, ensuring frontier models show meaningful separation and benchmarks capture capability jumps (e.g., Arc AGI staying <1% while human-solvable)
  • Distributional control/diversity: Intentional taxonomy design and task distribution to represent either real-world task frequencies OR paradoxically rare but disproportionately critical failure modes (e.g., pedestrians in self-driving despite low base rate)
  • Researcher UX: Product-quality developer experience for benchmarks: ease of running models against them, contributing tasks, extending, and leveraging signals for RL/tuning—severely underrated adoption driver per Chen
  • Adversarial quality control: Multi-reviewer validation protocols with revision cycles and incentive structures (e.g., GPQA payouts tied to agreement) to ensure tasks are well-posed when single experts insufficient
  • Autonomy horizon: Time span agent operates before reliability breaks down; frontier benchmarks should test weeks-long operation with continual learning, distributional drift, requirement changes mid-stream
  • Environment complexity: Real-world messiness missing from benchmarks: org-specific policies, tribal knowledge, Slack context, flaky tooling, human reviewers, parallel contributors—typical agent failure points in production
  • Output complexity: Moving beyond text/document outputs to artifacts, structured work products, uncertainty signals, and nuanced rewards usable for RL—third frontier dimension Snorkel sees as underexplored
  • Benchmark thesis: Explicit research question or bet on future capabilities/interfaces embedded in benchmark design (e.g., TerminalBench betting CLI becomes general computer-use abstraction)—differentiates field-shaping from measurement-only evals
  • Harbor (TerminalBench 2.0): De facto evaluation harness/infrastructure for agent builders; shipped with TerminalBench 2.0 update; example of researcher UX investment driving adoption

Operator Notes / Why Ken Should Care

  • For Ken's agent systems work: the science framework (task quality, distributional control, headroom, robust methodology) provides selection criteria when choosing external benchmarks or designing internal evals—prioritize benchmarks hitting all four
  • TerminalBench's CLI thesis and resulting adoption by Claude/Codex suggests CLI-first agent architectures have momentum; Ken should evaluate whether GUI-based approaches or other interfaces matter for his use cases before defaulting to CLI
  • SWE-bench family's extensibility signals coding agent evaluation is mature but may miss production realities (org policies, human reviewers, concurrent work); if deploying coding agents, budget for internal eval layers capturing these
  • The researcher UX insight applies to Ken's internal tooling: if building custom eval harnesses for his team, invest in developer experience (easy task contribution, simple model runs, signal extraction for RL) or adoption will stall regardless of eval quality
  • Snorkel's $3M Open Benchmarks Grant reviewing 120+ proposals = leading indicator of where academic/industry eval efforts concentrate; Ken should monitor funded projects at benchmarks.snorkel.ai for early visibility into frontier eval directions
  • Autonomy horizon dimension critical for Ken's long-running agent designs: distinguish copilot (minutes/hours) vs autonomous (days/weeks) evaluation needs upfront; long-horizon testing requires infrastructure for continual learning, state tracking, requirement drift
  • Output complexity direction (artifacts, uncertainty signals, nuanced rewards) opens research on agents producing structured work products; if Ken's agents generate anything beyond text (code, proposals, strategic plans), need multi-dimensional output eval beyond accuracy
  • The measurement gap framing (capabilities ahead of evaluation) suggests deployment blocker isn't technical—it's trust/measurement; for Ken's GTM/sales, positioning around transparent evaluation and safety cases may matter more than capability demos
  • Environment complexity gap (Slack context, tribal knowledge, org policies) means agent performance on public benchmarks likely overestimates production performance; Ken's deployment strategy should budget for performance degradation and incremental rollout with monitoring
  • Chen's emphasis on benchmarks shaping the field (not just measuring it) suggests strategic benchmark contribution could influence research direction; if Ken identifies eval gaps in his domain, publishing benchmark may attract research solving his problems

Watch Map

  • 00:00: Introduction: Vincent Chen, Snorkel AI co-founder, meta-evaluations for building benchmarks
  • 01:30: Snorkel context: frontier AI data lab, intersection of academic research and enterprise deployment
  • 02:30: Core asymmetry: capabilities improving but enterprises hesitant to deploy due to measurement gap
  • 04:00: Open Benchmarks Grant: $3M commitment, 120+ applications reviewed
  • 05:30: Framework overview: science (measuring sticks) vs art (frontier differentiators)
  • 06:45: Science pillar 1: Individual task quality, GPQA example with adversarial QC
  • 09:00: Science pillar 2: Distributional diversity, MMLU 57-domain taxonomy example
  • 11:20: Science pillar 3: Difficulty/model headroom, Arc AGI unsaturation capturing reasoning leaps
  • 13:30: Science pillar 4: Robust eval methodology, TauBench measuring policy adherence beyond accuracy
  • 15:15: Art differentiator 1: Benchmark thesis on future capabilities, TerminalBench CLI bet
  • 17:45: Art differentiator 2: Spawning research roadmaps, SWE-bench family proliferation
  • 19:50: Art differentiator 3: Researcher UX, Helm/Harbor infrastructure examples
  • 21:30: Frontier direction 1: Environment complexity, coding agent realism gaps
  • 23:40: Frontier direction 2: Autonomy horizon, long-running agents with continual learning
  • 25:30: Frontier direction 3: Output complexity, artifacts and uncertainty signals beyond text
  • 27:00: Conclusion: Open Benchmarks Grant still accepting proposals at benchmarks.snorkel.ai

Source/Metadata

  • Title: The Art & Science of Benchmarking Agents — Vincent Chen, Snorkel AI
  • Transcript words: 4937
  • Duration seconds: 1405
  • Timestamp note: Timestamps estimated based on 1405-second duration and transcript structure; not explicitly provided in source
Full transcript 3759 words · 22 min read
0:00

[SPEAKER_00] Hey, everybody. [SPEAKER_00] How's it going? [SPEAKER_00] Lovely. [SPEAKER_00] Well, I'm very excited to be here hailing from San Francisco. [SPEAKER_00] So a little bit of a trek over, but I'm super excited to chat with you all today after the talk and beyond. My name is Vincent. I'm a research fellow and co-founder at Snorkel AI. [SPEAKER_00] And today I'm going to be talking about some meta-evaluations for building benchmarks, the art and science of what we found to be really useful when building effective benchmarks.

0:00

[SPEAKER_00] I have the great privilege at Snorkel of working with both our researchers, great collaborators in academia, industry, and the open source community to build great benchmarks. [SPEAKER_00] And I wanted to share some of the learnings that we've had over the last few years on what really makes for benchmarks that shape the field and move it forward. [SPEAKER_00] So a little bit about us. We're a frontier AI data lab. We have labs both at academic settings. [SPEAKER_00] Our co-founders have labs at Stanford, UW, Wisconsin. [SPEAKER_00] We have an internal team for deployed engineers, applied engineers.

0:02

[SPEAKER_00] And I mention this because we get a lot of exposure to both the academic frontier, work with frontier labs, and also real enterprises and companies who are deploying in practice. [SPEAKER_00] So our focus as a company is on building the best data sets and environments to define and advance future AI capabilities. [SPEAKER_00] And we are in a fortunate spot where we get to play at this unique intersection of both the frontier academically, but also to contact reality with our deployments in enterprises.

0:14

SPEAKER_00

So today I wanted to talk about an asymmetry that we see in our real world deployments. There's real excitement around agents today. Every person in this room, I'm sure, has played with these agents. And we see real progress marked by hill climbing on model cards. We see the vibes are improving, right, and truly shifting, especially in coding.

0:22

SPEAKER_00

But when you ask individuals and enterprises or these large scale organizations if they're fully ready to let these agents loose and deploy them in high stakes environments, you get a little bit of hesitation. And that's not to say the capabilities aren't there, but our ability to actually measure these agents in practice is falling behind where the capabilities actually are. This is one of the challenges and research questions that I think are actually one of the most important in the field and one of the ones that we're very interested in here at Snorkel. So closing that gap, that evaluation gap, we believe requires a toolkit, right?

0:39

SPEAKER_00

As I mentioned earlier, we're strong believers in field deployments, right? So this is actually deploying engineers, researchers on our teams to contact reality and work with folks to deploy these models in real production settings where the stakes are high, right? These are finance settings, insurance, healthcare settings, where it's not just about a number, it's about real outcomes. And we're also big fans of other eval tools, right? This is red teaming, private human evals, crowdsourced labeling, a lot of the themes that we saw talked about today, which has been awesome to see.

1:12

SPEAKER_00

But one of the things that we feel most strongly about is that benchmarks, open benchmarks in particular, remain a really critical piece of the measurement toolkit. The best open benchmarks aren't just about taking a snapshot of progress looking backwards. They're actually about defining progress and shaping the field and setting a goalpost about where capabilities need to go. And even looking at the last few months, right, benchmarks like Terminal Bench, Meters Long Horizon Benchmark, Arc AGI, these are really exciting and critical guideposts for where the field is going.

1:31

SPEAKER_00

And as a result, the path to safe, trustworthy agents will really depend on more of these benchmarks in practice. So what are we doing at Snorkel? One of the things that I'm very fortunate to be able to be a part of is the open benchmarks grants. We recently, a few weeks ago, a month ago, deployed $3 million to commit to open benchmarks. And this is a really fun job. I get to work with the best academic teams, builders, to really accelerate and fund the next wave of benchmarks that's going to really steer and guide where the field is going. We've had a wild reception so far. I'm admittedly behind on some reviews.

2:14

SPEAKER_00

But we've been really excited to see what the community has come up with so far. And in this talk in particular, we've reviewed over 120 applications so far, spanning academia and industry labs. We wanted to share a few perspectives, a few learnings over the past few months about what we view as, one, table stakes for useful benchmarks. Right? How do you actually build good empirical measuring sticks that are actually useful to measuring progress? And two, what really separates those benchmarks that are shaping the frontier? Right? What is the art and the science, if you will, of building really effective benchmarks at the end of the day?

2:52

SPEAKER_00

So, as I'm doing this, I'll have a fun opportunity, maybe this is a little too American, but to pull out Timothy Chalamet and honor some of the greats, some of the great benchmarks over the last few years that have really shaped the field in talking about some of these axes. And I hope that these themes resonate with you and also inspire a bit of new thinking about how can we actually deploy some of the learnings we're all driving towards in our day-to-day work to shape the field and move it forward. So, two themes here, again, on the science side, how do we actually build effective measuring sticks?

3:08

SPEAKER_00

We'll talk about task quality, distributional control, robust evals in general. And on the art side, right, really the differentiators for great benchmarks. How do you build benchmarks with a thesis on where the field is going that inspire new roadmaps? And that critically are built for this audience, right, a researcher audience, a builder audience, so that adoption is something that is way smoother and a first-class citizen for a bunch of these benchmarks. So, let's start with the science. This is what makes for really effective measuring sticks, as we've seen them in practice, in the deployments that we see in industry and academia and with Frontier Labs.

3:27

SPEAKER_00

So, the first thing I want to talk about is individual task quality, right? This is the idea that individual tasks need to be exceptionally rigorously validated, right? They need to represent real-world complexity. They need well-posed, well-structured instructions. They need verifiable solutions that ideally have been actually validated by real-world domain experts. One of the benchmarks here, GPQA, is one of my favorites, not just because it's been a very lasting and enduring benchmark that captures graduate-level and professional knowledge, even to this date, right?

3:54

SPEAKER_00

As we've seen them in practice, in the deployments that we see in industry and academia and with Frontier Labs. So, the first thing I want to talk about is individual task quality, right? This is the idea that individual tasks need to be exceptionally rigorously validated, right? They need to represent real-world complexity. They need well-posed, well-structured instructions. They need verifiable solutions that ideally have been actually validated by real-world domain experts.

4:25

SPEAKER_00

One of the benchmarks here, GPQA, is one of my favorites, not just because it's been a very lasting and enduring benchmark that captures graduate-level and professional knowledge, even to this date, right? You still see this on model cards. But one of my favorite contributions is actually tucked away in the appendix. GPQA introduced one new adversarial quality control mechanism. So, the idea was that not only do these tasks need to be well-posed, they need to be tractable for other experts to solve.

4:39

SPEAKER_00

So, they had a very rigorous multi-reviewer protocol where there was an original author, there were reviewers and adjudicators in the loop, there was opportunity for revision, right? These were tasks that were really pushing the frontier of knowledge and it was non-trivial for any single expert to say, yeah, this is actually a good task or not. And so, developing this sort of rigorous adversarial quality control mechanism was one of the contributions I was most excited about here. And if you read the appendix, you also see that they introduced new incentive mechanisms, right?

5:02

SPEAKER_00

Payouts were actually based on whether there was certain agreement and, coming from academia, there's some inspiration here from the peer review process as flawed as that is. But, this type of innovation around how you actually get really rigorous, multi-expert quality control leads to the type of outcomes that we see around individual task quality that we see as a key foundation for any benchmark that matters at the end of the day. Two is distributional diversity, right? This is the idea that for any benchmark that really matters, you want to define a clear taxonomy for the domain, for real-world tasks, and distribute those tasks intentionally.

5:13

SPEAKER_00

So this might be, hey, I captured a trace or a kind of real-world stream of traffic along how my agent is operating in the real world, and I want to really represent that distribution. It could also mean, hey, I'm specifically characterizing and taxonomizing the failure modes that are paradoxically rare but disproportionately important in production, right? If you take classic self-driving settings, right? Yellow lights or pedestrians or motorcyclists, right? Might actually show up way less than other types of scenarios but are disproportionately important to get right in these settings.

5:48

SPEAKER_00

And so defining that taxonomy, being really intentional about distributing tasks across it, is one of the hallmarks of great benchmarks in our view. MMLU, a few years old now, constructed a quite ambitious taxonomy of 57 academic and professional domains across STEM, humanities, etc. It's remained one of the lasting benchmarks for understanding graduate and professional level knowledge. And again, a lot of this was, as we believe, a result of really thoughtful and intentional taxonomy designed and building towards that. The third axis here is around difficulty of individual tasks and model headroom, right?

6:13

SPEAKER_00

It's really important that the benchmark is unsaturated, that it exposes real soft spots in capabilities and reliably separates where models sit at the frontier. One of my favorite plots is the one on the top right. This was, all credit to the ArcPrize Foundation team. Arc AGI 2, for a very long time was unsaturated, right? For several months and years. And when there was the big reasoning push, maybe 18, 24 months ago, we saw a massive leap in capabilities that actually corresponded to a real leap in model capabilities, right? This was a benchmark that was intentionally designed to represent a type of efficiency or capability that humans have, but models didn't have.

6:57

SPEAKER_00

And they really captured, well, hey, there's a lot of model headroom here. Humans can do this. Where's that gap? And again, it correlated quite well with the recent, 01-style reasoning push that has really dominated the field in the past 18, 24 months. Just a few weeks ago, the Arc team just launched Arc AGI 3. And again, at launch, they had frontier models under 1%. Every single task was human solvable to some degree. And so it remains one of, I think, the most meaningfully exciting benchmarks in the space where any new model, people are awaiting, hey, how does it do on Arc? And I think they did this quite well, right? The model headroom here is really exciting.

7:54

SPEAKER_00

This last axis here I want to talk about on the empirical measurement side is all about robust eval methodologies. Now, this goes really deep. So capturing some of the high-level ideas. Benchmarks need to ideally go beyond accuracy to capture real-world dimensions that matter, right? This is everything from cost, latency, the quality of the reasoning traces, some of the intermediate steps and tool use. Whatever dimensions actually matter for the capability at hand, capturing those as reward or supervision signals is really critical. And measuring what it claims to is actually a non-trivial feat in building robust and reproducible benchmarks.

8:29

SPEAKER_00

So TauBench is a benchmark that we're a big fan of, it's had multiple evolutions over the years, but it was a benchmark that was built to evaluate both task completion of these multi-turn agents. They built a clever kind of user simulator. But also, not just accuracy and completion, but adherence to policy constraints. So a model, for example, on the right-hand side, this was one of the examples from the paper, a model that books the right flight, but violates fair class rules, still fails. It's still a kind of no-go at the end of the day.

8:57

SPEAKER_00

So this notion of, hey, being intentional about what axes we actually care about, what do we actually want to measure, and measuring that rigorously is one of the hallmarks that matters when you're building these frontier evals. So I want to shift a little bit now to the differentiators, right? What actually leads to the benchmarks that push the frontier? And that's not to say anything I mentioned on the last few slides have not pushed the frontier. These are just special characteristics that I view as critical to the benchmarks that are real research contributions, that are really shaping where's the field going, where are all the labs going to hill climb next?

9:15

SPEAKER_00

And this is the art, the special sauce that helps push us forward. So one of the key hallmarks here is these benchmarks should have a thesis, right? So I want to shift a little bit now to the differentiators, right? What actually leads to the benchmarks that push the frontier? And that's not to say anything I mentioned on the last few slides have not pushed the frontier. These are just special characteristics that I view as critical to the benchmarks that are real research contributions, that are really shaping where the field is going, where are all the labs going to hill climb next? And this is the art, the special sauce that helps push us forward.

10:04

SPEAKER_00

So one of the key hallmarks here is these benchmarks should have a thesis, right? They should have a research question about a subspace of capabilities, about where the field is going. It should revisit previous capabilities. And the most ambitious benchmarks are really a statement about where the world is going. Terminal Bench is one of these bets, right? It was a bet on the CLI, not just for coding agents, but for general purpose computer use. And in many ways I think this has turned out to be a largely correct and consequential bet, right?

10:33

SPEAKER_00

As we're seeing teams at Claude and Codex build their general purpose enterprise capabilities on top of these coding and CLI-based tools, we're seeing this bet pan out. And again, Terminal Bench remains one of the most robust and most important benchmarks that are measured on all the recent model cards. So again, this was a bet early on to say, hey, we think the CLI is going to be really important as a core interface, a core abstraction and affordance for agents to interact with the real world in a general purpose way. And by measuring those capabilities, I'd argue that it actually helped accelerate how the field is operating in this way.

10:47

SPEAKER_00

The second piece here I think worth mentioning is the ability to roadmap for the field. A great benchmark, one that really shapes where all of us are going is producing new roadmaps, right? It's inspiring new attacks against research problems. It's helping folks ideate and come up with new ways for thinking about benchmarks and methods in general. And I think Sweebench is a really phenomenal example of this, right? It was a simple idea. Often the best ones are quite simple, right? How do you leverage existing coding type capabilities via PRs? And it spawned a new family of benchmarks, right? All the way from Sweebench Lite, Verified, Pro, Multilingual, Multimodal, etc.

11:27

SPEAKER_00

And its evolution, I think, is still very relevant today. It's evolved how we think about coding agents. And one of the things that's been awesome to see with the Sweebench team is how many new research directions and inspired benchmarks have come after it in this coding space. And arguably, there's a lot more room to innovate on top of this as well. What are the new ways of coding look like? How do these types of workflows apply to Vibe coding and this new layer of abstraction that software developers are applying?

11:52

SPEAKER_00

I think it's been really exciting to see the foundation that the Sweebench team sets and how that's going to shape how we think about coding agents moving forward. So this theme here, I think, is severely underrated. And this is the notion of researcher UX, right? I think the most prescient benchmark builders are committed to the researcher and builder experience. This is to say, it's really simple to run models and agents against your benchmark. It's really simple to contribute new tasks, to extend. And also, it's really simple to leverage some of the signals that you're getting from the benchmark for RL or tuning post hoc. I think this is really underrated.

12:37

SPEAKER_00

It's a classic product principle to make what you're building and putting out there easy to use by the community or by your core users. And in this case, benchmarks have core users, which are other builders or researchers. And so really putting in time and attention to building those interfaces has been important for the adoption of some of the most important benchmarks. Right, to call out, I think the Stanford team at CRFM built Helm several years ago, which I'd argue pioneered a standardized modular harness for evaluating reproducible different scenarios as well as models against the standard testbed of models.

12:58

SPEAKER_00

Terminal Bench 2.0, just a few months ago, again, shipped with Harbor, which has been in many ways the de facto harness and evaluation infrastructure for teams who are building agents more broadly. And so, thankfully, we have a bunch of open source software out there today based on this principle. But as you're building your benchmarks, right, considering, hey, how easy is this to extend? How easy is it for the community to adopt and eventually hill climb against this, I think is a severely underrated factor for what makes for really high adoption of these frontier benchmarks. So this is the full framework. Again, can go into more detail and please find me afterwards.

13:28

SPEAKER_00

But what makes for really empirically meaningful measuring sticks, right? It's task quality and attention to distributional control and diversity. It's difficulty in model headroom. And of course, a robust eval methodology that measures the concrete axes that actually matter in practice and is intentional about it. And of course, on the art side, right, these great benchmarks really have a thesis on where the frontier is going. They set roadmaps for the field and they really prioritize researcher UX.

13:54

SPEAKER_00

Now, before I wrap up, I want to propose a few dimensions that we're really excited about at Snorkel that we think are really going to encapsulate the next wave of benchmarks. Tried to leave some more degrees of freedom here for creativity, but these are areas where we think there's a lot of room to push complexity, to push realism in benchmarks. And so I wanted to share a little bit of our internal roadmap and thinking around where the field is going and where we need more benchmarks and more contributions. So this is our point of view. We think that the axes for the next great benchmarks are threefold.

14:25

SPEAKER_00

And I'll go into a little bit more detail about what I mean here in just a second.

14:36

SPEAKER_00

But one, it's environment complexity, right? How complex, how realistic, how dynamic is the operating environment that these benchmarks are working in? Are they representative of real world settings that are professional, that scientists, that someone using these tools could actually use? Two is autonomy horizon. Do these benchmarks represent realistic and frontier horizon lengths that these agents are operating against? Are they capturing different points on the autonomy slider that are representative of how users are using them as copilots versus fully autonomous agents? Is this an intentional design in the benchmark?

15:00

SPEAKER_00

And three, capturing the wide range of output complexity. I think this is very underexplored today, right? Lots of chat-based or document-based outputs, not as much around nuanced differentiated reward signals, right? Real artifacts that show up in day-to-day work that we represent. Are they representative of real world settings that are professional, that scientists, that someone using these tools could actually use? Two is autonomy horizon. Do these benchmarks represent realistic and frontier horizon lengths that these agents are operating against?

15:18

SPEAKER_00

Are they capturing different points on the autonomy slider that are representative of how users are using them as copilots versus fully autonomous agents? Is this an intentional design in the benchmark? And three, capturing the wide range of output complexity. I think this is very underexplored today. Lots of chat-based or document-based outputs, not as much around nuanced differentiated reward signals. Real artifacts that show up in day-to-day work that we represent. And critically, new artifacts. New types of form factors that we haven't even imagined, how agents interact with humans, how agents interact with each other.

15:40

SPEAKER_00

So a little bit about each one. The first one here, I won't go into all of these in significant detail. Environment complexity is all about capturing the real world complexity that is in our day-to-day working environments. And this gap is often where agents fail today. Consider coding agents. A real code base has org-specific policies. Lots of slack context, screenshots, flaky tool chains, CI that's distributed.

15:59

SPEAKER_00

Human reviewers with knowledge in their heads about what they like and what they prefer. Many contributors in parallel. Benchmarks today capture a fraction of this complexity, and not just in coding but other domains. There's a lot of excitement and opportunity to up the level of complexity and continue to drive what these models can do to represent real world uses. Two, again on autonomy horizon. This is all about how long an agent can operate before reliability breaks down. Let's take a customer experience agent. In many cases, these agents may lose track of context that was delivered a few weeks ago.

16:19

SPEAKER_00

Different integrations or product specs might change the actual spec or requirements for a particular model. Reorgs can shift priorities midstream. Real world settings actually represent a lot more complexity that is represented in these long term continual learning type settings that represent changes in state and environment. So we think that there's a lot of room to contribute and build out new benchmarks that represent very long horizon and autonomous agents. And lastly, this axis is all about producing more complex work, more representative work, and also nuanced signals that can be used for not just evaluation but reward signals during training.

16:47

SPEAKER_00

This also has to be complex and this gap is growing as well. Let's take again a software example, or a complex report for making strategic recommendations. It's non-trivial and subjective to define what is verifiable about a good recommendation, a good strategic proposal, a good roadmap in general. The nuances of this need to be captured well. They need to capture organizational context, really good human judgment. And tomorrow's benchmarks, we're excited about signals that capture all of these settings. Trustworthy outputs, the ability for agents to actually capture their own uncertainty and define, I'm actually not sure about this.

17:45

SPEAKER_00

I actually need to stop or ask for more information. Different types of outputs that aren't just plain text answers is something we're really excited about. So, hopefully, this inspired a little bit of thinking around what am I working on? How can I turn this into a meaningful benchmark? We are still accepting benchmarks in the Open Benchmarks Grant. So if you're excited, please reach us at benchmarks.storcle.ai. Feel free to reach me directly. And we're really excited to see where the field is going and to use benchmarks to not just measure progress looking backwards, but really shape where things are going moving forward.

18:37

SPEAKER_00

Thanks for your time and excited to catch up soon. So this is our point of view. We think that the axes for the next great benchmarks are threefold. And I'll go into a little bit more detail about what I mean here in just a second. But one, it's environment complexity, right? How complex, how realistic, how dynamic is the operating environment that these benchmarks are working? Are they representative of real world settings that are professional, that, you know, scientists, that someone using these tools could actually use? Two is autonomy horizon. Do these benchmarks represent realistic and frontier horizon lengths that these agents are operating against?

19:19

SPEAKER_00

Are they capturing different points on the autonomy slider that are, again, representative of how users are using them as copylons versus fully autonomous agents? Is this an intentional design in the benchmark? And three, capturing the wide range of output complexity. I think this is very underexplored today, right? Lots of chat-based or document-based outputs, not as much around nuanced kind of differentiated reward signals, right? Real artifacts that show up, you know, in day-to-day work that we represent.

19:48

SPEAKER_00

And critically, new artifacts, right? New types of form factors that we haven't even imagined about, you know, how agents interact with humans, how agents interact with each other. So a little bit about each one. The first one here, I won't go into all of these in significant detail, right? Environment complexity is all about capturing the real world complexity that is in our day-to-day working environments, right? And this gap is often where agents fail today. Consider coding agents, right? A real code base has org-specific policies. You know, lots of slack context, screenshots, flaky tool chains, you know, CI that's kind of distributed.

20:24

SPEAKER_00

Human reviewers with knowledge in their heads about what they like and what they prefer. Many contributors in parallel. Benchmarks today capture a fraction of this complexity and, you know, not just in coding but other domains. There's a lot of excitement and opportunity to up the level of complexity and continue to drive what these models can do to represent real world uses. Two, again on autonomy horizon, right? This is all about how long an agent can operate before reliability breaks down. Right, let's take a customer experience agent. You know, in many cases, right, these agents may lose track, you know, of context that was, you know, delivered a few weeks ago.

21:06

SPEAKER_00

Different integrations or product specs might, you know, change the actual spec or requirements for a particular model. Reorgs can kind of shift, you know, priorities midstream. Real world settings actually represent a lot more complexity that, again, is represented in these kind of long term continual learning type settings that represent kind of changes in state and environment. So we, again, think that there's a lot of room to contribute and build out new benchmarks that represent very, very long horizon and autonomous agents.

21:37

SPEAKER_00

And lastly, this axis is all about producing more complex work, more representative work, and also nuance signals that can be used for not just evaluation but reward signals during training. This also has to be complex and this gap is growing as well, right? Let's take again a software example, right, or a complex report for making strategic recommendations. It's non-tribial and subjective, you know, to define, hey, what is verifiable about a good recommendation, a good strategic proposal, a good roadmap in general. The nuances of this need to be captured well. They need to capture organizational context, really good human judgment.

22:11

SPEAKER_00

And tomorrow's benchmarks really, you know, we're excited about signals that capture all of these settings. Trustworthy outputs, right, the ability for agents to actually capture their own uncertainty and define, hey, I'm actually not sure about this. I actually need to stop or kind of ask for more information. Again, different types of outputs that aren't just, you know, a kind of plain text answer is something we're really excited about. So, again, hopefully, you know, this inspired a little bit of thinking around, hey, what am I working on? How can I turn this into a meaningful benchmark? We are still accepting, you know, benchmarks in the Open Benchmarks Grant.

22:49

SPEAKER_00

So if you're excited, please reach us at benchmarks.storcle.ai. Feel free to reach me directly. And we're really excited to see where the field is going and, again, to use benchmarks to not just measure, you know, progress looking backwards, but really shape where things are going moving forward. Thanks for your time and excited to catch up soon. .

23:11

. . . .

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note