AI Engineer

Mousepower: agents that can’t be measured, can’t be managed. — Maximillian Piras, Yutori

1793 summary words 8 min summary Watch video

Start with the signal

8 min read

Summary

At-a-Glance

  • Verdict: Skim
  • Core thesis: Agent adoption and sustainable spend depend less on measuring tokens than on defining outcome-based, user-legible verification rubrics for each agent task.
  • Why it matters: Without a credible way to connect agent cost to validated business outcomes, teams will overrun budgets, lose trust, and move the human bottleneck from execution to review.
  • Best use: Use this as a product and systems-design framing for prioritizing agent workflows and designing their evaluation/control layer, rather than as an implementation guide.

Executive Summary

Maximillian Piras argues that agents have a measurement problem: early adopters may tolerate parallel experiments, unclear quality, and high token bills, but mainstream customers will not adopt agents unless they can understand and calculate their value. Token usage is an operational input, not proof of value. The relevant unit of measurement is a validated outcome—such as bugs resolved or support requests closed—tied cleanly to an organizational objective.

He uses James Watt's introduction of “horsepower” as an analogy. Horsepower was not a perfectly scientific measure, but it translated the unfamiliar value of a steam engine into the incumbent mental model of people accustomed to horses. “Mousepower” is Piras's proposed equivalent idea for agents: not a literal cursor-speed metric, but a product principle that every agent should come with an understandable rubric for judging whether its work was worthwhile.

The practical constraint is verification. Coding agents can generate work at compute speed, but human code review becomes the limiting factor; Anthropic is cited as having acknowledged that code review remains unsolved even amid major progress in code generation. Piras argues that agent builders must design both execution and verification, ideally enabling measurement at compute speed as well.

His task-selection heuristic maps two kinds of uncertainty: uncertainty in the steps needed to execute a task, and uncertainty in the acceptance criteria used to judge it. Fully predictable tasks should be scripts; highly unpredictable tasks may be poorly suited to current agents; tasks whose quality is as hard to verify as to perform are poor automation candidates. The attractive middle consists of tasks with meaningful but bounded execution uncertainty and repeatable, cheaper verification—an NP-style “easier to verify than execute” shape that may support separate verifier agents.

Key Takeaways

  • Claim: Token consumption is not the right primary measure of agent value; it must be traced to validated business outcomes. | Evidence: Piras contrasts token counts with outcome measures such as how many bugs were fixed or support requests closed, and notes examples of teams exhausting an annual token budget in one quarter. | Implication: For every deployed workflow, Ken should require an outcome metric, an attribution method linking spend to that outcome, and a quality threshold before scaling concurrency or model usage. | Caveat: Tokens still matter as an internal operational measurement and cost-control signal; the argument is that they are insufficient as the customer-facing value metric.
  • Claim: Agent adoption requires a mental model that makes the efficiency gain legible to non-experts, not merely technical claims about capability. | Evidence: The James Watt analogy frames horsepower as a useful adoption metric because it let buyers who understood horse-driven mills estimate the value of steam engines, despite horsepower being imperfect and not fully scientific. | Implication: Product positioning and internal rollout should express agents in the incumbent workflow's terms—time-to-resolution, cases processed, revenue protected, errors avoided, or approved deliverables—not in model, token, or agent-count terms. | Caveat: The speaker does not propose a universal replacement metric equivalent to horsepower; “mousepower” is explicitly a design idea rather than a standardized metric.
  • Claim: The main bottleneck created by coding agents is often verification, not generation. | Evidence: Piras describes teams as “dying by a thousand pull requests” and cites Anthropic's acknowledgement that code review has not been solved even as coding capability has advanced. | Implication: Increasing agent throughput without proportionally improving tests, policy checks, evaluators, and reviewer workflow can reduce rather than increase net engineering throughput. | Caveat: The talk does not offer a concrete solution for agentic code review; it uses code review mainly to illustrate the broader verification problem.
  • Claim: Building an agent is incomplete unless its builder also designs a reliable method for verifying that the output is good. | Evidence: Piras says the goal is to move from execution at the speed of compute to measurement at the speed of compute, using shared assumptions and clear rubrics analogous to those underlying conventional code review. | Implication: Agent architecture should include evaluation and control-plane components from the outset: acceptance criteria, test or audit artifacts, escalation conditions, and where feasible an automated verifier separate from the executor. | Caveat: Verification methods must fit the customer's actual mental model and objective; a technically sophisticated internal metric may not establish customer trust or ROI.
  • Claim: The best agent tasks are neither fully deterministic nor too open-ended, and they must have relatively low uncertainty in acceptance criteria. | Evidence: The proposed two-axis matrix uses uncertainty in task steps and uncertainty in acceptance criteria. Booking a flight is presented as more structurally constrained than painting a masterpiece because it has required fields and completion conditions. | Implication: Prioritize workflows where an agent must reason through some variation but success can be checked cheaply and consistently; avoid automating tasks where reviewers must effectively redo the work to decide whether it is acceptable. | Caveat: This is explicitly a thought starter, not an empirically validated task-prioritization framework or quantitative scoring system.
  • Claim: At the extremes of execution uncertainty, agents are usually the wrong tool: use scripts for predictable work and avoid highly unpredictable, out-of-distribution work. | Evidence: Piras argues that low-uncertainty task paths should be scripted rather than consume tokens, while high-uncertainty tasks risk being out of distribution for pretraining and offering sparse reinforcement-learning rewards. | Implication: Maintain a routing policy: deterministic processes to conventional automation, bounded-variable workflows to agents, and high-ambiguity/high-stakes work to human-led processes until evaluation and training signals improve. | Caveat: A task's position can change as tooling, data, models, and structured interfaces improve; this is not a permanent classification.

Detailed Brief

The overspending-and-underuse loop

  • Claims: Piras borrows Ramp's framing that organizations can enter a cycle of overspending on AI and then underusing it after budget or trust backlash.; FOMO and enthusiasm can restart experimentation, but without better value measurement the organization repeats the same spend-and-retreat cycle.; Model routing is a useful partial control: start with cheaper/default models and reserve frontier models for genuinely hard tasks.
  • Evidence: He references a Coinbase CEO-posted chart in which changing default model selection and reserving frontier models for harder tasks caused AI spend to diverge from token usage.; He characterizes early agent users as inclined to run many parallel agents while focusing elsewhere, then later question whether the result justified the bill.
  • Caveats: Cheaper-model-first routing controls spend but does not itself demonstrate business value; outcome attribution and quality verification remain necessary.; The transcript provides no detailed routing policy, thresholds, or measured Coinbase savings.
  • Implications: Treat budget management and task-quality measurement as a single system rather than separate FinOps and product concerns.; Monitor for vanity incentives such as token leaderboards, which reward consumption rather than validated output.

“Mousepower” as a product-design requirement

  • Claims: Mousepower is not intended to measure physical computer interaction efficiency, such as cursor movement or clicks.; Its useful meaning is the comparative, task-specific story and rubric that makes an agent's value auditable to its buyer or operator.; A verifier agent can be justified when verification follows a repeatable pattern that is less costly than the original task.
  • Evidence: Piras says he tried having Claude create a cursor-movement measurement device, then rejects the approach because information space is too high-dimensional.; He describes the desired task shape as an NP-style problem: easier to verify than to execute.
  • Caveats: An automated verifier can inherit model errors, create false confidence, or be gamed by the executor unless it uses independent evidence, diverse checks, or human audit sampling.; The speaker does not address adversarial behavior, security boundaries, or how to validate verifier quality.
  • Implications: A credible agent offering should expose evidence of completion and quality, not simply return a final answer or task-complete status.; The strongest near-term agent products may be those whose outputs naturally generate inspectable artifacts: structured records, citations, test results, reconciliations, or policy-check logs.

Notable Concepts & Terms

  • Mousepower: Piras's metaphor for a task-specific, user-legible way to express and verify agent value, analogous to how horsepower translated steam-engine capability into a familiar frame.
  • Horsepower / James Watt analogy: A lesson that adoption metrics need not be perfectly scientific if they credibly help customers compare a new technology with their current workflow and estimate ROI.
  • Overspending and underusing: The Ramp-derived cycle in which organizations spend aggressively on AI experimentation, retrench after costs disappoint, then resume due to FOMO without solving value measurement.
  • Execution uncertainty: Uncertainty in the sequence of steps needed to finish a task; low uncertainty favors scripting, while excessive uncertainty makes agent performance and training difficult.
  • Acceptance-criteria uncertainty: Uncertainty in how to judge whether the agent's result is correct or valuable; high uncertainty makes verification approach the cost of doing the work manually.
  • NP-style task shape: A desirable automation shape in which producing a solution is harder than checking it, enabling scalable evaluators or verifier agents.
  • Measurement at the speed of compute: The objective of scaling validation alongside agent execution so human review does not become the throughput bottleneck.

Operator Notes / Why Ken Should Care

  • Add a required evaluation design review to every new agent workflow: define the business outcome, acceptance criteria, evidence artifact, failure/escalation path, and cost attribution before rollout.
  • Create a task-routing inventory with three lanes: deterministic work to scripts, bounded-variable work to agents, and high-ambiguity or high-verification-cost work to human-led handling.
  • Instrument agent runs with outcome-level telemetry rather than only tokens, model calls, latency, and completion status; include accepted versus rejected output and reviewer effort.
  • Deploy automated verification only with independent checks and periodic human audit samples; do not equate a second model's agreement with reliable quality assurance.
  • Set model-routing defaults by task difficulty and require an explicit reason to use frontier models, then measure whether quality-adjusted outcomes improve enough to justify the incremental cost.
  • For coding-agent workflows, invest in test coverage, invariant checks, CI gates, and review triage before increasing parallel code-generation capacity.

Source/Metadata

  • Title: Mousepower: agents that can’t be measured, can’t be managed. — Maximillian Piras, Yutori
  • Transcript words: 7875
  • Duration seconds: 1237
  • Timestamp note: No timestamps or chapters were present in the supplied transcript. The transcript substantially repeats the talk's content.
Full transcript 3956 words · 35 min read
0:00

[SPEAKER_00] Well, thanks a lot for your time. Really appreciate you dropping by, and it's always a great honor to speak at the World's Fair, so I'll do my best to give you guys some valuable insights and hopefully make it worth your time. So my name is Maximilian Puros, and today I'll be talking about mouse power, and this is a talk about measuring agents through mental models. But before I get into talking about measuring agents, I'm going to talk through a bit about how I use them every day, and it might seem familiar to you, but just to level set, we'll go through it.

0:12

So I tend to background them, as I'm sure a lot of you people are as well. So while my active attention is focusing on one thing, such as giving this talk to you, I still want to make some progress on peripheral tasks, so I'll keep my attention focused on giving this talk while my agents can help me explore some designs in the background, because I think that my slides need a bit of work. So I've got my design system already set up, I've got some guidance given to my agents, and so I'll kick off an agent to try to explore some different directions on the type treatment and the layout, and try to get as many explorations as possible.

0:14

But of course, one agent's never enough, so I like to kick off a bunch in parallel. I've got a lot of slides to get through, so I need all of my agents exploring it in different directions, and hopefully I can get some interesting things to make my slides a bit better, and hopefully they can finish the job soon, because we're obviously up against the deadline here.

0:16

So this is generally how I work. I'm sure it's probably familiar to a lot of you, where we're trying to kick off agents for as much as possible in parallel, because it always feels like there's just way more research to do, we want it to be as thorough as possible. There's way more design explorations to do, so whenever our main focus is on one thing, why not kick a bunch of agents off in parallel and try to maximize your time? And it's a lot of fun, of course, until you get the bill. And then you start to wonder, was it all worth it, right? Did you vibecode too hard? Were you token maxing too much? Could you have been more efficient in how you approached sequencing your agents?

0:19

And so this is what I'm going to get into today. It's how do we value the token cost, and specifically how do we help our customers value it? So for the past year and a half, I've had the pleasure of working as the founding designer at a company called Utori, and we focus on computer use models. These are models that learn to use a computer like a human would, and the use case for them is when you can't get information from an API or an MCP, why not send an agent out to use a computer like a human would, and then we can extract all types of data and manipulate it in ways that let us access all the stuff that wasn't accessible previously. So obviously less efficient than APIs and MCPs, but as a last resort, have an agent go use the computer and try to get the information.

0:22

Here's the Utori agent using the Utori website. It's checking out its own benchmark, so in a way it's admiring itself.

0:25

So it gets a bit weird, and a lot of what I do as a founding designer there is talk to customers, try to understand how can we make agents as intuitive as possible, how do we figure out the mental models they're using to value the use cases they want to send out agents for, and a lot of them do seem pretty confused so far. A lot of people are excited about agents, but the phrase that comes up quite often is that they feel like they're just scratching the surface. It seems like it's not quite intuitive how we can best use them yet, and so in a lot of my customer discussions, it always comes down to a question of what is the best way to use agents, what are the best use cases for them, and how do I think about the trade-offs with regards to token cost relative to value?

0:27

So I think we're still building this muscle today, and this leads me to the thesis of the talk, which is that I think agents have a measurement problem. And as an example, here's me at work trying to measure some agents, and one of my coworkers took this photo and told me it looked like I was trying to solve the mystery of Pepe Silva.

0:30

So as you can see, it's not an easy task to measure agents, but I'm sure some of you are saying, hold on a sec, what is this guy talking about? I've got a fleet of agents working for me right now. We're building our next million-dollar app as we speak, and I'm having a totally fine time measuring my agents, to which I will agree with you, but then I will point you to the mandatory Upton Sinclair quote to remind us all that everybody in this room is very biased, and we're early adopters, and we're very excited to explore this new technology, but it doesn't mean that we represent the people that ultimately we're going to be trying to help adopt this technology.

0:35

And so, I think it's important to remind ourselves that in some way or another, we probably are selling tokens, whether it's indirectly or directly, and so when we think about our own token usage, is it really representative of all the people out there who have never touched an agent yet? Some people are still copy and pasting into ChatGPT. I may be married to one of these people, and despite how much I tried to get her to try out agents, she's not let me set her up with it yet.

0:38

And so, as a reminder, when we think about helping people adopt agents, all the people across the world that we think could get as much excitement and value as we do when we run off parallel agents, let's remember this quote. And so it really boils down to the age-old problem of a new technology. And, of course, there's tons of history we can go to to study how people saw this in the past. We have this really exciting new thing, but we haven't quite figured out the right ways to communicate it. And so for this talk, I'll go back to the 1700s, and we can take some notes from when James Watt was trying to sell steam engines.

0:43

And at the time, he decided that a great use case for his steam engines was trying to replace a horse gin. These were the power source of a mill at the time, so when you're, for instance, a brewery, and you need some power source to grind your barley or whatever. I don't know, I'm not a big brewery guy, so I don't know exactly how it's made, but you need a power source. And the power source at the time that was common was you hooked a horse up to a rotary arm, and the horse walked in a circle, and that's how it generated your power.

0:44

And it seems crazy today, but at the time it was commonplace. And Watt thought, you know, it would be much better than a horse is this very efficient machine. Although he rightfully acknowledged that one of the big barriers to adopting it would be this cognitive dissonance of trying to tell people who think in horses, how do you adapt to this old machine that's intimidating and scary, and perhaps somebody's going to say it's going to solve all your problems, but you can't go out and see the vision yet, so perhaps that sounds familiar to any of us working in agents today.

0:46

And Watt's solution was that he needed to understand the mental model of these people, and specifically to create a metric that would help him give some baseline of the relative improvement in efficiency. And so he literally studied horse gins and tried to get some kind of armchair measurements of how the mechanics and the average performance of it worked, and eventually came to a metric called horsepower, which may sound familiar. And he used this measure to quantify the general power that the horses were creating at the time, and then he could use it as a basis to show the multiplier of efficiency that a steam engine could provide.

0:49

And this metric was not very scientific at the time. It was not necessarily even accurate, you could say, but the main thing it did was it communicated an increase in value, and so this let people who love horses calibrate their efficiency gains that they could get by attempting to adopt a steam engine. So not necessarily what you would get when you use it, but what would get you over the limit of trying it out in the first place.

0:52

Which may sound familiar, and he used this measure to quantify the general power that the horses were creating at the time, and then he could use it as a basis to show the multiplier of efficiency that a steam engine could provide. And this metric was not very scientific at the time. It was not necessarily even accurate, you could say, but the main thing it did was it communicated an increase in value, and so this let people who love horses let them calibrate their efficiency gains that they could get by attempting to adopt a steam engine. So not even necessarily what you would get when you use it, but what would get you over the limit of trying it out in the first place. And it's a pretty big feat because although he had efficiency on his side with regards to this metric, you know, let's be honest, regardless of how efficient this was, horses just have great vibes, so it's hard to beat the vibes of horses, and so he knew he had to overcome the emotion and actually speak to something that gave him an ability to calculate the ROI.

0:54

And the lesson being, if we're not able to give something that is a tangible ROI for our customers, then it's very hard for us to communicate value. And I think we only need to look to our own industry to see all the examples where other people in the technology sector are failing to calculate good ROIs as well. And so we might, in this room, think this is somewhat of a solved problem, but if you look to the other engineers in the world who are perhaps not as AI-pilled, they're theoretically very smart and should be able to figure out how to calculate this quite well, but then you get these scenarios where people are blowing through their entire token budget for a year and they're going through it in a quarter, or they're dealing with token leaderboards and such. And so obviously the incentives haven't quite aligned and we haven't perhaps got the right measure of value in terms of the technology sector itself.

0:57

And so how then do we end up scaling past that and talk to people who have no idea what we're talking about, but still try to provide them a measure of increased efficiency with agents. And so right now I think we're in this doom loop where we're overspending and we're underusing. This is a term I borrowed from RAMP, and they have a great blog post on this. And so it's this vicious cycle where we're token maxing ourselves into austerity and then dropping out of the loop until we get more FOMO to get activated enough to try it again.

1:01

And so I think we have to break this loop and I think the way we do that is by getting better measures that will communicate value. Some people are obviously on the right track. There was this chart floating around on X recently that the Coinbase CEO posted where they had internally started changing the defaults of what models they will start with and trying to only save the frontier models for the hardest tasks. And as a result, saw AI spend start to diverge from token usage. And this is a good start. RAMP also, as I mentioned, has a great blog post about this. But I think the problem is still that it's too focused on tokens. And tokens are of course useful as a measurement of an internal system, but at the end of the day, they're just an output. And so the tokens need to then be traced very cleanly to an outcome.

1:05

How many bugs got squashed with our token spend? How many support requests got closed, et cetera? So clean outcomes and then cleanly tying those to progress on our objectives. And so without a very tight measure of ROI, this becomes very hard to do. And I think I'll take this further and say that it need not even be the broader technology industry where it's encountering this problem, but also many of us in this room perhaps are. And although we're all probably enjoying coding with various agents and feeling that there's something there in terms of the increase in ability and efficiency, the problem of course is that we're all dying by a thousand pull requests.

1:08

And so even Anthropic, who has some people on the team who have claimed to have solved coding, they have also admitted that they've not solved code review. And so as a result, the bottleneck is now shifted to human review where the efficiency gains from coding agents aren't quite seen yet, because we spend most of the time reviewing the code and we've not figured out how to scale that in tandem with the generation of the code itself.

1:10

And so the bottleneck ends up shifting to the verification side and thus we don't have a way to measure value at scale and to judge quality at the same speed. And so again, going back to the ROI calculations, we generate all this code, but how do we know we don't know that enough of it is good to justify the spend. And of course maybe code review was always flawed, but it's just that agents are now exposing it, are exposing the actual problem.

1:13

And I like this quote by Noah Hine from a post about how to solve code review, where he's mentioning specifically that the assumptions underneath code review are what needs to be revisited. So we have to check our priors to try to figure out a new basis for how we can code review in the age of agents.

1:16

And I'm not going to go into how to solve code review. I think that's definitely better a talk that's better given by somebody else and is a totally different subject. But what I think is important for this talk is why does code review feel like it is solvable? And I think that Noah is sitting on something important here, which is that as a culture, code review has a very good convergence on shared assumptions. And that lets you measure things at scale when we can all converge on the measurements and it becomes somewhat of a clear rubric.

1:19

And so the task at hand now is we have to adapt those assumptions for the agentic age. And so if we're able to do that, then we can go from execution at the speed of compute to measurement at the speed of compute. And of course, the measurements need to fit the mental models of the customers using it. And I think the lesson here being that if you're going to think of how to build an agent for something, you also have to think about how do you help the customers build or at least create a method for verifying that the output is good. And so it's not enough to build it. We also have to help them get to clear ROI calculations to justify their spend.

1:25

And so this brings me to the idea of mouse power, which could be the equivalent of horsepower for the agentic age, just as James Watt was able to show a measure of efficiency relative to the horses in the horse gins that were the source of power at the time. We perhaps can also figure out how do we create a baseline of efficiency for the way we use computers today, and can then demonstrate how much better or perhaps more performant on certain vectors an agent could be at that task.

1:27

And of course, it's not as easy a task as he had back then, where he could just study the horse gin, because it's not as if we can create some method to measure our cursor movements and figure out the delta of how much more efficient an agent could move them, and thus we can say agents are this much more performant than humans at these tasks.

1:29

Just as James Watt was able to show a measure of efficiency relative to the horses in the horse gins that were the source of power at the time, we perhaps can also figure out how to create a baseline of efficiency for the way we use computers today and can then demonstrate how much better or perhaps more performant on certain vectors an agent could be at that task. And of course, it's not as easy a task as he had back then, where he could just study the horse gin, because it's not as if we can create some method to measure our cursor movements and figure out the delta of how much more efficient an agent could move them, and thus we can say agents are this much more performant than humans at these tasks. Trust me, I've tried. I had Claude write code for me, this measurement device, and I thought maybe if I can figure out the movement, the potential movement across the screen and measure how fast it went, I could get some clean measure of mouse power. But of course, it's only joking. This is a fool's errand, because information space is just way too high dimensional, and so I think mouse power is never going to be a metric, but it's more so an idea, which the idea being if you're going to sell somebody an agent, you also have to help them with the rubric of how do we actually verify that this agent is doing good work, and thus we can have a good measure of saying that these tokens are worth it.

1:33

So how to do that is really up to you, and I won't be able to tell you how to do it. I don't have any good frameworks for how to figure out the right measurements to help provide anybody you're building an agent for. But what I can do is give a principle, an idea that I've been kicking around, which is based in information theory, so going back to Claude Shannon's ideas about measuring entropy in information, entropy being the uncertainty of a probability distribution, and of course, very much the basis of how we train agents today, things like cross entropy and such being a big factor in determining how capable an agent is, I think that entropy's an interesting idea to think through with regards to not just the performance of an agent, but also the tasks that we're sending them out to perform on.

1:38

And so I put together this matrix, which maps on the x-axis the uncertainty in the steps it takes to perform a task. And so when we're thinking of building an agent, I think it's not enough to just think what would be a valuable task for the agent to do, but also thinking about how much uncertainty is in the steps to perform that task itself.

1:41

So an example would be booking a flight has much less uncertainty than, let's say, painting a masterpiece, right, because you know there's certain information that has to happen in the flight purchase. There has to be a departing destination, arriving destination, there's going to be a seat chosen, it might be by the person, it might just be random, but these things have to happen for that task to be completed. And on the other hand, there is the task of painting a masterpiece, right, and who knows what the steps are to that. Maybe you can get an agent to do it, but it would be very hard to figure out how we can actually create a relatively predictable pathway to that.

1:47

But then on the other axis is the uncertainty in the acceptance criteria itself. So not just can the agent perform the task, but can we help somebody actually, or is there actually a clean rubric for how it's graded? And so thinking about ideas on these two axes and where they intersect perhaps gives us a better guide for how to build agents, and we can run through a few examples.

1:50

So if we look at the left side, your right side, yes, your left as well. Then I know the last speaker was also confused by that. So yeah, on the left side, when uncertainty in the task steps are low, then it's a very predictable outcome, or it's a very predictable pathway to achieve that goal. And so then, you know, why would you waste tokens? Just write a script.

1:53

On the other side, when the steps to perform the task are very high in uncertainty, then you have very unpredictable information, and so it's probably at risk of being out of distribution from pre-training, and probably has very sparse rewards for reinforcement learning, and so perhaps it's not a good task for an agent, because it's just much harder to figure out how to actually model that data.

1:57

And so obviously in the middle is probably the sweet spot, but then on the other axis, what's the uncertainty in verifying that this is actually valuable? So when you have high uncertainty in the acceptance criteria, you pretty much have a spot where verification is indistinguishable from execution. So why would you build an agent for something that to verify it was useful, a person pretty much has to do the work again? So waste of tokens, obviously.

1:59

And then it leaves that middle area where you have this interesting intersection of tasks that are not too uncertain in that they have a degree of uncertainty where they're not just a script, or they're not out of distribution for training, but they have enough uncertainty to be interesting, but at the same time, they also have a property of being relatively easy to validate the value of them. And so they become in this place where they kind of become the shape of an NP style problem, which means they're easier to verify than to execute.

2:01

And the reason I say that is because if you can figure out a pretty repeatable pattern for verifying their work, you can actually just throw agents at that problem as well. And so of course, you don't just build the agent, you perhaps build the agent that verifies the work of the agent. And so yeah, this is perhaps a thought starter mostly, still in the works. So happy to hear any thoughts on it. But with this guidance, I hope when you're building your next agent, you can also figure out how to also build its mouse power. And thanks very much. And it's a lot of fun, of course, until you get the bill. And then you start to wonder, was it all worth it, right?

2:17

Did you Vibecode too hard? Were you token maxing too much? Like, could you have been more efficient in how you approached your sequencing your agents? And so this is what I'm going to get into today. It's how do we value the token cost, and specifically how do we help our customers value it? So for the past year and a half, I've had the pleasure of working as the founding designer at a company called Utori, and we focus on computer use models. These are models that learn to use a computer like a human would, and the use case for them is when you can't get information from an API or an MCP, why not just send an agent out to use a computer like a human would,

2:56

and then we can extract all types of data and manipulate it in ways that let us access all the stuff that wasn't accessible previously. So obviously less efficient than APIs and MCPs, but as a last resort, just have an agent go use the computer and try to get the information. Here's the Utori agent using the Utori website. It's checking out its own benchmark, so kind of it's admiring itself in a way. So yeah, it gets a bit weird, like, and a lot of what I do as a founding designer there is talk to customers, try to understand how can we make agents as intuitive as possible,

3:30

how do we figure out the mental models they're using to value the use cases they want to send out agents for, and a lot of them do seem pretty confused so far. A lot of people are excited about agents, but the phrase that comes up quite often is that they feel like they're just scratching the surface. It seems like it's not quite intuitive how we can best use them yet, and so in a lot of my customer discussions, it always comes down to a question of, like, what is the best way to use agents, what are the best use cases for them, and how do I think about the trade-offs with regards to token cost relative to value? So I think we're still kind of building this muscle today,

4:06

and this leads me to the thesis of the talk, which is that I think agents have a measurement problem. And as an example, here's me at work trying to measure some agents, and one of my coworkers took this photo and told me it looked like I was trying to solve the mystery of Pepe Silva. So as you can see, it's not an easy task to measure agents, but I'm sure some of you are saying, hold on a sec, like, what is this guy talking about? I've got a fleet of agents working for me right now. We're building our next million-dollar app as we speak, and I'm having a totally fine time measuring my agents,

4:43

to which I will agree with you, but then I will point you to the mandatory Upton Sinclair quote to remind us all that everybody in this room is very biased, and we're early adopters, and we're very excited to explore this new technology, but it doesn't mean that we represent the people that ultimately we're going to be trying to help adopt this technology. And so, you know, I think it's important to remind ourselves that in some way or another, we probably are selling tokens, whether it's indirectly or directly, and so when we think about our own token usage, is it really representative of all the people out there who have never touched an agent yet?

5:17

Some people are still copy and pasting into ChatGPT. I may be married to one of these people, and despite how much I tried to get her to try out agents, she's not let me set her up with it yet. And so, as a reminder, when we think about helping people adopt agents, you know, all the people across the world that we think could get as much excitement and values as we do when we run off parallel agents, let's just remember this quote. And so it really boils down to the age-old problem of a new technology. And, of course, there's tons of history we can go to to study how people saw this in the past.

5:52

We have this really exciting new thing, but we haven't quite figured out the right ways to communicate it. And so for this talk, I'll go back to the 1700s, and we can take some notes from when James Watt was trying to sell steam engines. And at the time, he decided that a great use case for his steam engines was trying to replace a horse gin. And these are the, was the power source of a mill at the time, so when you're, for let's say a brewery, and you need some power source to, to grind your barley or whatever. I don't know, I'm not like a big brewery guy, so I don't know exactly how it's made, but you need a power source.

6:26

And the power source at the time that was common was you hooked a horse up to a rotary arm, and the horse walked in a circle, and that's how he generated your power. And seems crazy today, maybe, but at the time was commonplace. And Watt thought, you know, it would be much better than a horse is like a very efficient machine. Although he, um, rightfully acknowledged that one of the big barriers to adopting it would be this cognitive dissonance of trying to tell people who kind of think in horses, how do you adapt to this, to this, uh, old machine that's kind of intimidating and scary, and perhaps, uh, somebody's gonna say it's gonna solve all your problems,

7:02

but you, you can't, can't go out and see the vision yet, so, uh, perhaps that sounds familiar to any of us working in ages today. And Watt's solution was that he needed to understand, um, the mental model of these people, and specifically to create a metric that would help him, uh, give some baseline of the relative improvement in efficiency. And so he literally studied, uh, horse gins and tried to get some kind of armchair measurements of, of how is, uh, like, what are the mechanics and the average, um, performance of it, and eventually came to a metric called horsepower,

7:35

which may sound familiar, and, uh, he used this measure to, you know, this was to quantify the general power that the horses were, um, creating at the time, and then he could use it as a basis to show the multiplier of efficiency that a steam engine could provide. And, uh, this metric was, uh, not very scientific at the time. It was not, not necessarily even accurate, you could say, uh, but the main thing it did was it communicated, uh, an increase in value, and so this let people who, who love horses, uh, let them kind of calibrate their, um, the, the efficiency gains that they could get by, by, uh, attempting to adopt a steam engine.

8:12

So not even necessarily, um, what you would get when you use it, but what would get you over the limit of trying it out in the first place. And, uh, you know, it's, it's, it's a pretty big feat because, like, although, um, he had efficiency on his side with regards to this metric, um, you know, let's be honest, regardless of how efficient this was, um, horses just have great vibes, so, like, it's kind of hard to beat the vibes of horses, and so he knew he had to kind of overcome the emotion, uh, and actually speak to, to something that gave him an ability to calculate the ROI. And, oh, sorry, skipped something.

8:48

And so, yeah, the, the lesson being, um, if we're not able to give something that is a tangible ROI, um, for our customers, then it's very hard for us to communicate value. And, um, I think we only need to look to our own industry to see all the examples where other people in the technology sector are failing to calculate good ROIs as well. And so we might, in this room, think this is somewhat of a solved problem, uh, but if you look to the other engineers in the world who are perhaps not as AI-pilled,

9:16

uh, they're theoretically very smart and, uh, should be able to figure out how to calculate this quite well, but then you get these scenarios where people are blowing through their entire, um, um, token budget, uh, for a year and they're going through it in, in a quarter, or they're, like, dealing with token leaderboards and such. And so, obviously the, uh, incentives haven't quite aligned and we haven't perhaps got the right measure of value in terms of the technology sector itself. And so how then do we end up scaling past, past that and talk to people who have no idea what we're talking about,

9:45

but still try to provide them, um, a measure of, like, uh, increased efficiency with agents. And so right now I think we're kind of in this doom loop where we're, we're overspending and we're underusing, uh, this is a term I borrowed from RAMP, um, and they have a great blog post on this. And so it's kind of this vicious cycle where we're just token maxing and ourselves into austerity and then kind of dropping out of the loop until we get more FOMO to, to get activated enough to try it again. And so I think we have to break this loop and I think the way we do that is by getting better measures of, that will communicate value. Some people are obviously on the right track.

10:21

There was this chart floating around on X recently that the Coinbase, Coinbase CEO posted where they had internally started changing the defaults, uh, of what models they will start with and trying to only save the frontier models for the hardest tasks. And as a result, saw some good, um, saw AI spend start to diverge from token usage. And, uh, this is a good start, uh, RAMP also, as I mentioned, has a great blog post about this. Uh, but I think the problem is still that it's too focused on tokens. And tokens are, of course, uh, useful as a measurement of an internal system, but at, at the end of the day, they're just an output.

10:56

And so the tokens need to then be traced very cleanly to an outcome. So how many, uh, bug, uh, how many bugs did the tokens, uh, sorry, how many, um, bugs squashed to the tokens that we bought, um, sorry, totally butchered that. Um, how many, uh, bugs got squashed with the token, with our token spend? How many, uh, support requests got closed, et cetera? So clean outcomes and then cleanly tying those to, to progress on our objectives. And so without, uh, a very tight measure of ROI, this becomes very hard to do. And I think I'll take this further and, um, say that it need not even be the, the broader, um, technology industry where it's encountering this problem.

11:37

But also many of us in this room perhaps are. And, uh, although we're all probably enjoying, uh, coding with, uh, with various agents and feeling like it, it's, it does feel like there's something there in terms of the increase in ability and efficiency. Um, the problem of course is that we're all kind of dying by a thousand pull requests. And so, uh, even Anthropic who has, uh, some people on the team have claimed to have solved coding, uh, they have also admitted that they've not solved code review. And so as a result, um, the, the bottleneck is now shifted to human review where the efficiency gains from coding agents aren't quite, aren't seen yet.

12:12

Because we spend most of the time reviewing the code and we've not figured out how to scale that in tandem with, uh, the generation of the code itself. And so the bottleneck ends up shifting to the verification side and thus we don't have a way to, uh, measure value at scale and, uh, to judge quality at, at the same speed. And so again, going back to the ROI calculations, we generate all this code, but how do we know, uh, we don't know that enough of it is good to justify the spend. And, and of course maybe, uh, code review was always flawed, uh, but it's just that agents are now exposing it for, uh, the, are exposing the actual problem.

12:48

Um, and I like this quote by Noah Hine who, from a post about how to solve code review, where he's mentioning specifically that the assumptions underneath code review are what's now being, uh, what needs to be revisited. So we have to, uh, check our priors to try to figure out a new basis for, um, how we can code review in the age of agents. And I'm not gonna go into how to solve code review. I think that's definitely better, uh, a talk that's better given by somebody else and, um, is totally different subject. But, uh, what I think is important for this talk is why does code review feel like it is solvable?

13:19

And I think that Noah is sitting on something important here, which is that as a, as a culture, uh, code review has a very good, uh, convergence on shared assumptions. And that lets you, um, that lets you, uh, measure things at scale when we can all kind of converge on the measurements and it becomes somewhat of, of a, uh, clear rubric. And so, uh, the task at hand now is we have to adopt, uh, we, we have to, sorry, adapt those assumptions for the agentic age.

13:49

And so, uh, we can, we need to go, um, if we're able to do that, then we can go from execution at the speed of compute to measurement at the speed of compute. And of course, the measurements need to fit the mental models of the customers using it. And, um, I think the lesson here being that if you're going to, um, think of how to build an agent for something, you also have to think about how do you help the customers bill or build, or at least create a method for verifying that the output is good. And so it's not enough to build it. We also have to help them. Uh, we also, we also have to help them get to, uh, clear ROI calculations to justify their spend.

14:25

And so, um, this brings me to the idea of mouse power, which could be the equivalent of horsepower for the agentic age, just as James Watt was able to show a measure of efficiency relative to the horses in, in the gins, uh, in the horse gins, uh, that were the source of power at the time. We perhaps can also figure out how do we create a baseline of efficiency for the way we use computers today, and can then demonstrate how much, uh, better or perhaps more performant on certain vectors an agent could be at that task. Um, and, of course, it's not, uh, as easy, perhaps as easy a task as he had back then,

15:01

where he could just study the horse gin, uh, because it's not as if we can create some method to measure our cursor movements and, like, figure out the delta of how much more efficient an agent could move them, and thus we can say, yeah, agents are this much more performant than humans at these tasks. Uh, trust me, I've, I've tried. I had Claude, uh, Vibecode me, this measurement device, and I thought maybe if I can figure out the movement, uh, like, the potential movement across the screen and measure how fast it went, I could get some clean measure of mouse power. Uh, but, of course, it's, it's only joking. Um, this is, of course, um, like a fool's errand,

15:34

because information space is just way too high dimensional, and so I think mouse power is, is never going to be a metric, of course, but it's more so an idea, which the idea being, if you're going to sell somebody an agent, you also have to help them with the, with the rubric of, how do we actually verify that this agent is doing good work, and thus we can, uh, have a, a good measure of saying that these tokens are worth it. Um, so how to do that, of course, is, is really up to you, and I won't be able to tell you, uh, how do you, I don't have any good frameworks for how do you figure out the right measurements

16:05

to, to help provide, uh, anybody you're building an agent for, uh, but what I can do is give a principle, uh, give an idea that I've been kicking around, which is based, um, in information theory, so, going back to, uh, Claude Shannon's ideas about measuring entropy in information, uh, entropy being, uh, the uncertainty of a probability distribution, and of course, very much the basis of how we train agents today, things like cross entropy and such, uh, being a big factor in determining how capable an agent is, um, I think that entropy's an interesting idea to think through with regards to, not just the performance of an agent,

16:38

but also the tasks that we're sending them out to, to perform on, and so, uh, I put together this matrix, which, uh, it maps on the x-axis, the uncertainty in the steps it takes to perform a task, and so, when we're thinking of building an agent, I think it's not enough to just think, what would be a valuable task for the agent to do, but also thinking about how, um, how much uncertainty are in the steps to perform that task itself. So, an example would be, uh, booking a flight has, uh, much less uncertainty than, let's say, painting a masterpiece, right, because you know there's certain information that has to be, that has to happen in the flight purchase,

17:15

there has to be a departing destination, arriving destination, there's gonna be a seat chosen, it might be by the person, it might just be random, uh, but these things have to happen for that task to be completed, and on the other hand, there is the task of, like, painting a masterpiece, right, and who knows what the steps are to that, uh, maybe you can get an agent to do it, but it would be very hard to figure out, uh, how we can actually create a, a, a, a relatively, um, predictable pathway to that. Uh, but then on the other axis, uh, is the, uh, the uncertainty in the acceptance criteria itself.

17:45

So, not just can the agent perform the task, but can we help somebody actually, or is, is there actually a, a clean rubric for how it's graded? And so, thinking about ideas on, on these two axes and where they intersect, uh, perhaps gives us a better guide for how to build agents, and we can run through a few examples. Um, so, if we look at the, uh, at the left side, your right side, um, yes, um, no, your left as well. Um, then, uh, I know the last speaker was also, uh, confused by that. Um, so, uh, yeah, on the left side, uh, when uncertainty in the task steps are low, then it's a very, it's a very predictable outcome,

18:24

or it's a very predictable pathway to achieve that goal. And so, then, you know, why would you waste tokens? Just write a script. On the other side, when, uh, the, the steps to do, perform the task are very high in, in uncertainty, then you, you have very unpredictable information, and so it's probably at risk of being out of distribution and pre-training, and probably has very sparse rewards for reinforcement learning, and so, perhaps it's not a, a good task for an agent, because it's just much harder to figure out how to actually model that data. And, uh, so, obviously in the middle is, is, um, is, uh, so I'm, I think I'm out of time, but I'm not getting kicked off yet.

19:02

Uh, so I'll just finish this up quickly. Um, so yeah, in the middle is, is probably the sweet spot, but then on the other axis, uh, what's the uncertainty in verifying that this is actually valuable? So, when you have high uncertainty in the acceptance criteria, you pretty much have a spot where verification is indistinguishable from execution. So, why would you build an agent for something that, to verify it was useful, a person pretty much has to do the work again. So, like, waste of tokens, obviously. And then it leaves that, that middle area where you have this interesting intersection of tasks that are, um,

19:33

they're not too uncertain in that they, or they, they have, um, a degree of uncertainty where they're not great, they're not just a script, or they're not out of distribution for training, but they have enough uncertainty to be interesting, but at, at the same time, they also have a property being relatively easy to, uh, validate the, uh, to check the value of them. And so they become in this place where they kind of become the shape of an MP style problem, which means they're easier to verify than to execute. And the reason I say that is because if you can figure out a pretty repeatable pattern for verifying their, their work,

20:08

you can actually just throw agents at that problem as well. And so, of course, you don't just build the agent, you perhaps build the agent that verifies the work of the agent. Um, and so, yeah, this is, uh, perhaps, uh, this is a thought starter mostly, kind of, uh, kind of, um, still in the works. So, uh, happy to hear any thoughts on it. But if, uh, with this guidance, I hope, uh, when you're building your next agent, you can also figure out how to also build its mouse power. And thanks very much.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note