AI Engineer

Computer Use at the Edge of the Statistical Precipice — Pierluca D'Oro, Programma Labs

1943 summary words 9 min summary Watch video

Start with the signal

9 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Computer-use benchmarks and model-selection decisions are unreliable when static environments permit replay scripts and when confidence intervals ignore variation across environment configurations.
  • Why it matters: Ken needs agent evaluations that predict production behavior rather than benchmark memorization, especially when model-routing or deployment choices carry meaningful error costs at scale.
  • Best use: Use this as a design review for computer-use evals: stress-test replayability, add verified environmental variation, and revise statistical reporting before trusting comparative agent scores.

Executive Summary

Pierluca D'Oro argues that many computer-use-agent benchmarks measure the wrong thing because their environments are static and deterministic. A “replay agent” can simply store successful click/type/scroll trajectories from a frontier model and blindly replay them; on standard benchmarks such as OSWorld or MobileWorld, this script can match or exceed the originating model. The speaker treats this as evidence that benchmark scores can reward memorization of a fixed environment rather than general computer-use ability.

The critique extends to pass@K. In deterministic computer-use environments, the paper formally shows that pass@K—the probability that at least one of K attempts succeeds—effectively evaluates the success probability of the replay-agent exploit. D'Oro's remedy is not merely more generated tasks, but multifactorial task variation whose combinations are mechanically verified as valid.

His proposed environment-design framework is PRISM: environments should be parameterized/multifactorial, realistic, isolated or sandboxed, support privileged-information-based verifiers, and ensure generated configurations are valid. DigiWorld is the implementation example: 15 Android apps, 387 verified scenarios, and 3.2 million verified configurations varying task instance, underlying data, visual theme, and initial UI state. A compiler-like system combines parameterized templates, verifiers, mock data, and base UI state to generate and filter valid configurations.

The other major argument is statistical: repeated model rollouts capture only action randomness, while real deployment performance also varies across task and environment configurations. Ignoring the latter can yield severely overconfident intervals—reported 95% intervals may achieve only roughly 17–20% empirical coverage. The practical consequence is potentially costly false confidence when choosing between models; a properly hierarchical estimate may instead tell an operator that more evaluation is necessary before deployment.

Key Takeaways

  • Claim: Static computer-use benchmarks can be defeated by a blind replay agent, so high benchmark performance need not indicate genuine generalization or interface understanding. | Evidence: The replay agent records one successful frontier-model trajectory per task—taps, typing, scrolling—and replays it blindly. D'Oro says that on standard benchmarks including OSWorld and MobileWorld, such an agent can achieve the same or better success rate than the frontier model that generated the traces; for hundreds of tasks, the script would be under 1 MB. | Implication: Before using a computer-use benchmark for procurement, routing, or release gating, Ken should test a trace-replay baseline. If it performs materially well, the benchmark is likely overstating agent capability. | Caveat: Some individual tasks are inherently repeatable, so replay success is not automatically invalid; the issue is strong replay performance on average across a benchmark intended to measure adaptable computer use.
  • Claim: Pass@K is a fragile metric for deterministic computer-use tasks because it operationalizes the same replayability problem rather than robust task competence. | Evidence: The speaker defines pass@K as the probability that at least one of K attempts succeeds and states that the paper formally proves that, in deterministic environments, it evaluates the replay agent's success rate. | Implication: Ken should not treat pass@K as a standalone deployment-readiness metric for UI agents, particularly on fixed task sets. It needs to be paired with held-out environment configurations and metrics that reflect configuration-level uncertainty. | Caveat: This critique is specifically tied to deterministic or insufficiently varied environments; the transcript does not claim pass@K is universally invalid in all evaluation settings.
  • Claim: The core requirement for robust computer-use environments is verified multifactorial variation, not simply a larger volume of AI-generated software or tasks. | Evidence: DigiWorld varies task instance (for example, transfer amount), data profile (contacts or emails), visual theme, and starting screen. It contains 15 Android apps, 387 verified scenarios, and 3.2 million verified configurations. D'Oro emphasizes that coding agents can generate abundant software, but cannot by themselves ensure that all task/UI/data combinations are valid. | Implication: For Ken's agent evals, build variation into task state, data, and UI conditions, then make validity checking a first-class component. A large task count without verified configuration integrity is not evidence of a trustworthy eval. | Caveat: Broad combinatorial variation can introduce invalid or incoherent tasks unless each generated configuration has a verification path.
  • Claim: PRISM provides a practical design checklist for trustworthy computer-use environments: parameterized variation, realism, isolation/sandboxing, support for privileged verifiers, and validation of generated configurations. | Evidence: The talk derives the principles from the replay-agent failure mode and standard environment needs: vary data, appearance, and initial state; verify combinations; sandbox the environment; provide privileged information for verifiers; and faithfully reproduce relevant real systems. | Implication: Ken can use PRISM as an acceptance checklist for internal or vendor-provided computer-use benchmarks, with particular scrutiny on initial-state variation and verifier coverage. | Caveat: The transcript presents PRISM as design guidance rather than a demonstrated guarantee that an environment is production-predictive.
  • Claim: Frontier computer-use models can be brittle under seemingly superficial environmental changes, including initial screen and app theme changes. | Evidence: Using the configuration axes in DigiWorld, D'Oro reports that frontier models are weak in worst-case robustness: a model that appears competent on a task may not retain performance when the starting screen or app theme changes. | Implication: Production readiness should be evaluated as a distribution and worst-case robustness question, not as a single average score from one canonical app state. | Caveat: No model-by-model results or exact degradation figures are provided in the transcript, so the magnitude and universality of this effect cannot be assessed from the talk alone.
  • Claim: Confidence intervals that account only for repeated agent rollouts can be radically overconfident because they omit environment-level variation. | Evidence: The speaker distinguishes action randomness across model trajectories from variation across benchmark configurations. He says nominal 95% intervals based on rollouts/base cases can achieve only about 17–20% coverage in realistic cases, whereas a method that respects the benchmark hierarchy restores roughly 95% coverage. | Implication: When comparing agents or models, Ken should require uncertainty estimates that resample or otherwise model both task/configuration variation and stochastic agent execution, rather than accepting narrow rollout-only error bars. | Caveat: The transcript does not specify the exact hierarchical estimator or its sample-size requirements; those details are deferred to the paper.
  • Claim: Overconfident evaluation can create material operating losses by causing teams to deploy the wrong model despite apparently decisive benchmark results. | Evidence: D'Oro gives a scale example: at 1 million tasks, a real 4% performance mismatch between two models, with mistakes costing roughly $20 each and $12 on average, can create hundreds of thousands of dollars of monthly cost. A reliable interval may instead indicate that evidence is insufficient and additional evaluation is warranted. | Implication: Ken should define a decision threshold that ties evaluation uncertainty to business downside: when intervals overlap at a consequential level, pay for additional evaluation rather than force a model-selection decision. | Caveat: This is an illustrative scenario rather than a documented customer case, and its economics depend on task volume and cost of failure.

Detailed Brief

Compiler-like generation architecture for valid task configurations

  • Claims: The scalable unit of benchmark construction is a parameterized task plus a verifier, rather than a manually authored static test case.; Configuration generation should be a reject/filter process: enumerate combinations across variation axes, retain only those that satisfy validity conditions, and expose the resulting valid configurations to evaluation.
  • Evidence: DigiWorld's example task template is conceptually “send a certain amount to a certain recipient.”; Its generation pipeline combines a parameterized task template, corresponding verifier, task-specific mock data, and a base UI state in a system D'Oro compares to a visual compiler.; The speaker notes that even modest numbers of values per axis create millions of combinations, and that scaling the same approach can reach billions.
  • Caveats: Generation scale is not the bottleneck by itself; the engineering burden moves to reliable task verifiers, coherent mock data, and valid UI-state composition.
  • Implications: Treat eval infrastructure as a software-engineering system with typed task specifications, state generators, and executable success conditions—not as a collection of screenshots and prompts.; Coding agents may accelerate content generation, but they should be used behind a verification gate rather than trusted as the source of benchmark validity.

Decision discipline under uncertainty

  • Claims: A benchmark may be informative at the task level yet still be inadequate for a model-selection decision if its reported uncertainty does not represent the task distribution.; The correct outcome of an eval is sometimes a deferred decision rather than a declared winner.
  • Evidence: The talk contrasts deceptively narrow confidence intervals for two candidate models with the unobserved “real performance” bars, where the apparently justified decision is wrong.; The recommended alternative is to spend more time or money evaluating when the statistically appropriate method indicates insufficient confidence.
  • Caveats: The talk does not provide a universal minimum number of tasks, configurations, or rollouts; required sample sizes depend on the desired confidence and cost of an incorrect choice.
  • Implications: Connect evaluation budgets to expected downside from incorrect deployment decisions, instead of optimizing for the fastest possible leaderboard comparison.; Make uncertainty reporting a release requirement for consequential agent changes, especially when a small aggregate score gap will be used to select a model.

Notable Concepts & Terms

  • Replay agent: A script that blindly replays previously successful UI action traces; it is used as a diagnostic for whether a benchmark can be solved through deterministic memorization.
  • Pass@K: Probability of at least one success in K attempts; the speaker argues it becomes equivalent to measuring replay success in deterministic computer-use environments.
  • PRISM principles: The speaker's environment-design mnemonic: multifactorial/parameterized variation, realistic reproduction, isolated/sandboxed execution, privileged-information-capable verification, and validity checks for generated configurations.
  • DigiWorld: The proposed Android computer-use benchmark built around verified variation: 15 apps, 387 scenarios, and 3.2 million verified configurations.
  • Configuration-level variation: Changes in task instance, data profile, visual theme, and initial UI state that test whether an agent generalizes beyond a canonical benchmark setup.
  • Coverage: The frequency with which a stated confidence interval actually contains true performance; used here to expose rollout-only intervals as severely miscalibrated.
  • Hierarchical confidence interval: An uncertainty estimate that accounts for both stochastic agent trajectories and variation across benchmark tasks/configurations, rather than treating all uncertainty as rollout noise.

Operator Notes / Why Ken Should Care

  • Add a replay-trace baseline to every fixed-state UI-agent benchmark and flag any benchmark where blind replay attains meaningful aggregate success.
  • Require each computer-use eval task to declare its mutable axes—data, task parameters, visual state, and starting state—and its executable verifier before it is admitted to a release suite.
  • Revise model-comparison reporting to include configuration-aware uncertainty; do not approve a routing or deployment change from mean score plus rollout-only error bars.
  • Set an evaluation escalation rule based on expected failure cost: if uncertainty leaves a materially costly model choice unresolved, collect more configurations and runs rather than selecting on a marginal score lead.
  • Run explicit robustness slices for non-semantic UI changes and alternate valid entry states, since average canonical-state success can conceal operational brittleness.

Source/Metadata

  • Title: Computer Use at the Edge of the Statistical Precipice — Pierluca D'Oro, Programma Labs
  • Transcript words: 4475
  • Duration seconds: 1048
  • Timestamp note: No timestamps or chapters were present in the supplied transcript.
Full transcript 2653 words · 20 min read
0:12

Hi everyone, I'm Pierluca Loro, and I'm the founder of Programma Labs. Today I'm going to talk about computer usage and evaluation, and most of the work and details about it are in a paper with this title. I did this work while at Manta Superintelligence Labs with the collaborators you see on this slide.

0:18

To start, I want to introduce this type of agent. It's a weird type of agent that I call a replay agent. Imagine we run this process: we run our frontier model, a good one, on a benchmark we like, and then for every task we collect a successful trace or a successful trajectory, and we have our recorded table of these types. The actions might be tapping, typing, scrolling, and we record this. Then we do this for all the tasks in the benchmark, and we compile these into a replay agent that, when the tasks arrive, replays that sequence of actions blindly.

0:25

If you do this for a common benchmark with hundreds of tasks, this is going to be a script that is less than a megabyte, and this is a completely valid type of agent that you can evaluate on the benchmark. If you try to evaluate this kind of agent on standard benchmarks such as OS Word or Mobile Word, you will see that the success rate of this agent, compared to the frontier model from which the agent was extracted, is actually the same or even better.

0:32

This is a weird but maybe trivial phenomenon, but I would argue that we shouldn't accept this kind of blind script beating the frontier models. The reason why this happens is the determinism of most existing benchmarks. If the benchmark is static and deterministic, then it is somehow gameable by this sort of strategy.

0:41

It goes even deeper than this. If you look at one of the metrics that people have been using in the past for evaluating computer use agents, it's pass at K. This metric is defined as the probability of at least one of K attempts succeeding, but if you look into the details of how this metric works on a deterministic environment, you will see that it's literally, and we prove it formally in the paper, evaluating the success rate of the replay agent that I've shown to you. If that replay agent felt weird to you, also pass at K on computer use tasks should somehow feel weird to you. In other words, pass at K is a metrification of that exploit of the replay agent.

0:49

These are two specific problems, but they point at two general classes of problems in Kua benchmarks. These problems are around environments, building environments, and evaluation, building good metrics to know if your agent is good or not. In particular, we want to have environments that don't have exploitable structure, and we want to have metrics that are not fragile or based on fragile statistics. I'm going to talk about both of the aspects now, so let's talk about building principled environments first.

0:56

The first aspect that I worked on while working on environments is to try to design a set of principles that could be guiding principles when building environments, so that we build robust environments and trustworthy environments. If you think about the problem that I was describing with replay agents, the first thing that you could think about as a solution, to not have a replay agent hack your benchmark, is to have your benchmark be multifactorial. That means generating variation for your benchmark, having stochasticity in the benchmark. For computer use environments, that means varying stuff like data or appearance or simply the initial state.

1:03

But if you want to make sure that all the combinations that you generate are valid, you want to have, as a design principle in your environment, a system for checking and verifying that everything is working as intended for every combination. Of course, you want the usual things for your environment, so you want your environment to be sandboxed, and you want your environment to support verifiers. You want your environment to provide privileged information, and you want your environment to be realistic. If it's a reproduction of a real system, you want that reproduction to be faithful so that the score that you get out is a good one.

1:11

If you sort them out, you can remember these principles as the PRISM principles for environment design. We tried that method to build a benchmark that would satisfy all of these principles. If you look at existing benchmarks, some of them do some things in a good way, some others do other things in a good way, but there is no unified benchmark that matches all these boxes. We built one that's called DigiWord.

1:20

The way DigiWord was built in practice is as a set of mobile apps for Android devices. It's 15 apps spanning different domains, with 387 verified scenarios and a number of configurations. These configurations are large in number, 3.2 million, but the important thing is that they are verified. The axes are the ones that I was mentioning before, so you can imagine for each one of the tasks you can vary things like the instance, what is the exact amount of money that you're sending, for instance, or the data profile, which kind of contacts or emails you have in the data for your tasks, or the theme, or the starting screen. Do you start from the login page, or do you start from another valid page?

1:26

If you do the math, even if you start from a relatively low number of base cases for each one of these variables, you end up having many combinations, so you can get to millions of combinations. If you scale this up, you can get easily to billions of combinations. All of these different axes can be manipulated by coding agents because, in the end, they are forms of software, so you can have a coding agent generate different instances, different themes, and such.

1:32

You might think maybe it's easy to build an environment. You just generate as much software as you can with a coding agent, and then you have a diverse environment. But it's a little trickier than that. Coding agents can generate a lot of software, but a lot of software is not the same as an effective KUA environment. The reason for this is that you need to verify the correctness of your combinations. The key to scaling this up is to have a verification strategy for the variations of your tasks.

1:39

The kind of verification strategy to follow is this one. You can generate many configs, all the combinations of the different factors that I've explained before, and you can then have a system that rejects the ones that are not valid and just keeps the valid configs. In the case of DigiWord, we did this by building a system that looks a little bit like a compiler and works in the following way.

1:45

You start from a parameterized task template, and this might look like something like this: send certain amount to certain recipient. Then you have a verifier that corresponds to that template, and then you have mock data for that task, data that you need for that specific task to happen. Then we have a system that is like a visual compiler that takes all of this and, given a base case of data and a base case of UI state, puts all of this together and creates a valid configuration. You can build systems like this in which the main craft is software engineering, to make sure that the combinations that you have are both diverse and valid.

1:54

If you do the same process we did before, you evaluate your frontier model and then you evaluate the corresponding replay agent, you will see that the replay agent doesn't get a lot of performance. It gets low performance, which is probably what you want. Sometimes some tasks may be repeatable by nature, but on average you shouldn't expect a replay agent to have good performance on the benchmark.

2:00

Once you build these diverse combinations, you also can do other things, like measuring the robustness of frontier models over different axes of variation. The axes of variation I described before are represented here, and you can see that, in the worst case, frontier models are actually pretty bad at being robust to these variations. If you have a model that seems to be good at a given task, you would expect that if you just vary which screen the task is starting from or what the theme of the app is, the models should pretty much have the same performance. But this is actually not the case for most frontier models. If you have infrastructure like this, you can actually measure that and tailor your expectations about this kind of robustness.

2:08

This was about the first aspect, building an environment that supports diversity and is robust enough to evaluate models. But the second aspect is as important as the first one, which is to measure uncertainty honestly. Once you have all of this variation, how do you handle computing the real performance of your agent?

2:14

There are two sources of stochasticity or variation, and they are not exactly the same, but they are equally important. The one that we usually think about is the one about the actions. You run your model multiple times, and in many cases you can have quite different trajectories out of it because the action on each step will be different. But if you have a benchmark like the one that I've described, with multiple combinations and multiple variations, then also the variability from the environment becomes important, and we want to capture that because that is what we are going to find in the real world.

2:22

We need a methodology that captures both of these types of variation, and in the paper there are the details, but basically we build a methodology that can accurately capture these two types of variation, taking into account the structure of the benchmark. In practice, it is useful to use this concept of coverage. When you compute a confidence interval, you have some confidence that the performance of the model is inside that range, and so you would expect that a 95% confidence interval would say that 95% of the time the performance of the model is in that range.

2:27

But if you only use rollouts, so you only use the base case and what people would use normally, actually in realistic cases you have something like 17% or 20% coverage. That means that basically only 20% of the time you guess the right performance of the agent, which can be pretty bad. But if you take into account the hierarchy and you use the proper way of computing confidence intervals, you can get to the full confidence interval and be 95% accurate.

2:37

If this seems quite abstract, in practice that means that if you want to make a decision about which models to deploy, maybe you have model A and model B, and you do an eval for those two models, you can have cases in which the confidence intervals seem really small, and so you make a decision based on those small confidence intervals, but they are actually overconfident. So this was the wrong decision. The orange bars are the real performance here.

2:45

If a mistake is pretty costly for you and you have many tasks, if you have one million tasks and there is a 4% mismatching performance for real in the models, and each mistake is $20, like $12 on average, it can cost you hundreds of thousands of dollars in a single month. It can be super costly as a mistake, just from confidence intervals being overconfident. But if you have a reliable way of computing the confidence interval, the method would tell you, I'm not confident enough to make an informed decision. Then you can choose to spend more money or more time evaluating models and avoid the costly mistake. We don't want to delude ourselves with wrong confidence intervals because there's money on the table, essentially.

2:53

This is a final checklist of the things that I've discussed so far. To recap, some of the things that are important in building a benchmark are about the environment, and some other things are about the metrics. On the environment, you can follow the principles that I described before, the PRISM principles. Some of these things are rather common, but some things, like varying initial state across runs, are pretty rare across existing benchmarks, but they are very important, so I would suggest you try to incorporate these into your evals.

3:04

On the metrics side, you can read the paper for the details, but essentially it's very important to avoid replayability as something that you can have in your benchmark, and also to focus on having accurate confidence intervals, so respecting the benchmark structure and trying to avoid underestimating the uncertainty overall.

3:10

I've heard many times sentences like, this benchmark can be gamed, but everybody's using it, or, there is no error bar, but I don't see people using them. These are things that we can think when we don't have enough time, but actually a non-rigorous benchmark can be misleading. It can be misleading for the field because everybody could be maximizing a score on a benchmark that maybe is not capturing what we can do. But especially, it can be misleading for your own decisions. If you are deluding yourself into thinking that a score is confident and that it is confidently telling you that your model is good, actually you are going to pay for those mistakes. I think it's usually very good to be honest with yourself and to try to be rigorous in the evaluations that you have.

3:24

As a last slide, I just started this company, Programma, and we are building the best infrastructure for QA and Ableware verification, and so we are hiring if you are interested or want to chat. This is our website. Thank you very much. Thank you very much.

4:10

and they're always describing with replay agents the first thing that you could think about as a solution not to have a replay agent to like hack your benchmark is to have your benchmark to be multifactorial so that means varying generating variation for your benchmark so having stochasticity into the benchmark and for computer use environment that means varying stuff like data or appearance or simply the initial

4:40

the initial state but if you want to make sure you want to make sure you want to make sure you want to make sure that all the combinations that you generate are valid and so you want to have as a design principle in your environment also a system for checking and verifying that everything is working as intended for every combination and of course you want the usual things for your environment so you want your environment to be sandboxed and you want your environment to support like verifiers so

5:10

you want your environment to be privileged information and you want your environment to be realistic so if it's a reproduction of a real system you want that reproduction to be faithful so that the score that you get out is a good one and so if you sort them out you can remember these sort of principles as the prism principles for environment design and we tried that method to build a benchmark that would be satisfying all of these principles

5:40

and if you look at existing benchmarks some of them do some things in a good way some others do other things in a good way but there is no unified benchmark that sort of matches all these boxes and we built one that's called DigiWord so the way DigiWord in practice was built is as a set of like mobile apps for Android devices so it's 15 apps spanning in different domains with 387 verified

6:10

scenarios and a number of scenarios and a number of configurations so these configurations they are in large numbers of 3.2 million but the important thing is that they are verified and indeed the axes are the ones that I was mentioning before so you can imagine for each one of the tasks you can vary things like the instance so what is the exact amount of money that you're sending for instance or the data profile like which kind of contacts or emails

6:40

you have in the data for your tasks or like the theme or the starting screen so do you start from the login page or do you start from another valid page so if you do the math even if you start from a relatively low number of base cases for each one of these variables you end up having many many combinations so you can get to like millions of combinations and if you scale this up you can get to easily to billions of combinations and all of these you know different axes can be manipulated by coding agents because in the end they are like forms of software so you can have a coding agent to generate different instances different teams and such

7:24

so you might think maybe it's easy to build an environment you just generate as much software as you can with a coding agent and then you have like a diverse environment but it's a little bit trickier than that and indeed coding agents can generate a lot of software but a lot of software is not the same as an effective KUA environment and the reason for this is that you need to verify the correctness of your combination right

7:52

and so the key to scale these up is to have a verification strategy for the variations of your tasks and so the kind of verification strategy to follow is this one so you can generate many configs all the combinations of the different factors that I've explained before and you can then have a system that rejects the ones that are not valid and just keeps the valid configs and so in the case of DigiWord we did this by building a system that looks a little bit like a compiler and that works in the following way

8:32

so you start from a parameterized task template and so this might look like something like this so you have send certain amounts or certain recipient and then you have a verifier that corresponds to that template and then you have mock data for that task so data that you need for that specific task to happen

8:54

and then you have a system and then you have a system and then you have a system and then we have a system that is like a visual compiler that takes all of this and given a base case of data base case of UI state puts all of this together and creates like a valid configuration and so you can build systems like this in which the main craft is software engineering to make sure that actually the combinations that you have are both diverse and valid

9:21

and so if you have a benchmark and so if you have a benchmark and so if you have a benchmark and so if you do the same process we did before you evaluate your frontier model and then you evaluate the corresponding replay agent you will see that the replay agent doesn't get a lot of performance it gets a lot of performance it gets a lot of performance it's probably what you want sometimes some tasks maybe are repeatable by nature but on average you shouldn't expect a replay agent to have good performance on the benchmark

9:58

once you build like these diverse combinations you also can do other things like measuring the robustness of frontier models over different axes of variation so the axis of variation I described before are here represented there and you can see that in the worst case frontier models are pretty bad actually at being robust to these variations and so for instance if you have a model that seems to be good at a given task you would expect that if you just vary you know which screen the task is starting from or like what is the theme of the app the models should pretty much have the same performance but this is

10:43

actually not the case for most frontier models and so if you have infrastructure like this you can actually measure measure that and like tailor your expectation about this kind of robustness so this was about the first aspect that was building an environment that supports diversity and that is robust enough to evaluate models but the second aspect is as important as the first one is to measure uncertainty

11:13

honestly so once you have all of this variation how do you handle like computing the real performance of your agent and basically there are two sources of stochasticity of variation and they are not exactly the same but they are equally important and so the one that we usually think about is the one about the actions right and so you you run your model multiple times and in many cases you can have even quite

11:42

different trajectories out of it because the action on each step will be different but if you have a benchmark like the one that I've described with multiple combinations with multiple variations then also the variability from the environment becomes important and we want to capture that because that is what we are going to find in the real world and so we need a methodology that captures both of these types of variation and in the paper there are the details but basically

12:12

we build a methodology that can accurately capture these two types of variation taking into account the structure of the benchmark and so if you start like in practice it is useful to use this concept of coverage so when you when you compute a confidence interval basically you have some confidence that the performance of the model is inside of that range and so you would expect that a 95% confidence

12:42

confidence interval would say that you know 95% of the time the performance of the model is on that range but if you only use rollout so you only use the base case and what people would use normally actually in realistic cases you have something like 17% or 20% coverage and that means that basically you only 20% of the time you guess the right performance of the agent which can be pretty bad but like if you take it into account the hierarchy and you use the proper

13:12

way of computing confidence intervals you can get to the full confidence interval and be 95% accurate and so if this seems quite abstract you know in practice that means that if you want to make a decision about which models to deploy maybe you have model A model B and you do an eval for those two models you can have cases in which the confidence intervals seem really really small

13:41

and so you make a decision based on those small confidence intervals but they are actually overconfident and so this was the wrong decision so that the orange bars are the real performance here so you make this decision and if a mistake is pretty costly for you and you have many tasks like if you have one million tasks and there is a 4% mismatching performance for real in the models and each mistake is like $20 like $12 on average

14:11

it can cost you like hundreds of thousands of thousands of dollars in a single month so it can be like super costly as a mistake just for confidence interval being overconfident but if you have like a reliable way of computing the confidence interval the method would tell to you I'm not confident enough to make an informed decision and so you can choose like to spend more money to spend more time on evaluating models and avoid the costly mistake so we don't want to like elude ourselves with like elude ourselves with like wrong confidence intervals because there's you know money on the table essentially

14:49

and so this is sort of a final checklist of the things that I've that I've discussed so far so again to recap some of the things that are important in building a benchmark are about the environment and some other things are about the metrics so on the environment you can follow the principles that I described before

15:11

like prism principles so some of these things are rather common but some things like varying initial state across runs they are pretty rare across existing benchmarks but they are very important and so I would suggest you to try to incorporate these into your evals and things about the metrics you can of course read the paper for the details but essentially it's very important to avoid replacing the ability as something that you can have in your benchmark and also to focus on having accurate confidence intervals so respecting the benchmark structure and trying to avoid underestimating the uncertainty overall

15:56

and so I've heard many times sentences like this benchmark can be game but everybody's using it or like there is no error bar but I don't see people using them so they are sort of things that we can think when we don't have enough time but actually a non-regurous benchmark is misleading you know it can be misleading for the field because everybody could be seeking you know maximizing a score on a benchmark that maybe is not capturing what we can do

16:24

but especially it can be misleading for you know your own decisions and so if you are deluding yourself on thinking that a score is like confident and that is confidently telling that your model is good actually you are going to pay for those mistakes and so I think it's very good usually to be honest with yourself and to try to be rigorous in the evaluations that you have as a last slide I just started this company program and we are building the best infrastructure for QA and Ableware verification and so we are hiring if you are interested or want to chat and this is our website thank you very much thank you very much

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note