SPEAKER_06
Ash Khandari - Nice to meet you guys. I'm Ash. This is Andrew. We both work as engineers in our Applied AI team here at Anthropic. The topic for this session was inspired by a blog post we put out just a couple of weeks ago about how to think about building agents that can actually run for really long extended periods of time. We're talking five, six hour plus runs. I think we've all seen these demos of companies saying, hey, we've one-shotted a browser, for example, but not necessarily sharing the details into what goes into the harness. That's what we want to talk about today. First, my amazing colleague Andrew will talk about how we've got here, some of the primitives that we've shipped in Claude, and where we are today. Then I'll hop back on stage to talk about some of the more experimental stuff we're playing with in harnesses, as well as a few examples of what we've seen. Over to you.
SPEAKER_06
[SPEAKER_08] Andrew - Sounds good. Thank you, Ash. And thanks everyone for joining our first session of the AI Engineer Conference. Glad you're spending it with us. My name's Andrew. I'm on the Applied AI team based out of London, working as a solution architect with a lot of our digital native industries customers. I'm going to give a little bit of a history tour, but really with the focus on all the things that we've shipped that lead to agents being able to run for multiple hours or even days at a time. Then I'll hand over to Ash to do more of the state of the art.
SPEAKER_06
Okay. So, a little quote from Boris, the creator of Claude Code. This was on the one year anniversary of Claude Code. Basically saying a year ago, Claude was struggling just to write bash commands and escaping strings and it could run for maybe 20 minutes at a time. And we're now at the point where almost all of Claude Code is being written by Claude Code and it can run effectively for days at a time. A big swing over just the course of a year and I'll walk through that history now. Just to frame the problem up. Why is it really difficult for these agents to run for extended periods of time? Broadly, there are three big buckets. Some are more intuitive than others.
SPEAKER_06
First, context. Context window is very much finite. You start a new session and there's amnesia. The agent has to start from scratch. So you need some sort of memory components. Also, as you're working through a context window, there's this notion of context rot. There's less coherence as you're getting deeper into that session. Also, you might get to the point where the model actually exhibits what's called context anxiety. It gets nervous as it reaches the end of its context window and it just quickly hurries up to finish what it's doing.
SPEAKER_06
This leads into planning. In general, models are not that great at planning out of the box. They might try and do everything in just one shot. Or they might build half a feature and then stop. Or they might just run out of context altogether and leave a half-finished app built.
SPEAKER_06
But models are really bad at judging their own output. Models can be sycophantic and tell you what you want to hear. But this applies as well to coding tasks. It might look at a feature and see that it's half-baked or only partially implemented and say, yeah, okay, that looks done and then it'll move on to the next thing. Or it might build a feature like a button but the back end doesn't exist for it. There's nothing behind that. But it looks like the feature is done. Ash will talk quite extensively about some of the new techniques we have to help with this. Specifically, models can become better at judging their own output.
SPEAKER_06
There are two ways really we can fix these things. The first one is the model. So baking it all into the model weights themselves. I'm sure you've all seen this meter chart. It's basically how long can an agent run for with a minimal scaffold where it's completing 50% of the tasks. You'll see from Opus 3.7, it's around one hour. Up to Opus 4.6, one year later, it's at 12 hours. An entire day. We've managed to get that running much longer. Other people have as well. But this is just a very minimal scaffold.
SPEAKER_06
The second thing you can do is make changes to the harness itself. This is the scaffolding around the model. We have the agent SDK, which ships with all the primitives that we've been building over time. There's the core agent loop itself. You have the Claude model that's determining what to do, what tools to run. Maybe it's pulling in tools from MCP servers. It might delegate some tasks to a sub-agent. It's bringing in all the context from things like Claude.md or the skills that are loaded or slash commands. There's a whole permission system. This will change over time as the models get better and improve. But these are the core primitives that we're working with.
SPEAKER_06
Then, of course, you use this framework to build your own harness for whatever it is you're trying to do. Such as some of the things that Ash will show later when we get into more long-running agents. What's also interesting is looking back at the last year of releases. When we've released a model, we've always also released a lot of harness changes alongside the models. Really, these things are co-evolving together. So we'll just look back. [SPEAKER_08] And there's a whole permission system. [SPEAKER_08] And this will change over time as the models get better and improve. [SPEAKER_08] But these are the core primitives that we're working with.
SPEAKER_06
[SPEAKER_08] And then, of course, you use this framework to build your own harness for whatever it is you're trying to do. [SPEAKER_08] Such as some of the things that Ash will show later on when we get into more long-running agents. [SPEAKER_08] I think what's also interesting is just looking back at the last year of releases. [SPEAKER_08] Is that when we've released a model, we've always also released a lot of harness changes alongside the models.
SPEAKER_08
So, really, these things are co-evolving together. So, we'll just look back. I suppose, firstly, just prehistory beyond one year ago. I think we all remember that period where Claude had the artifact section of Claude.ai. And Sonnet 3.5 was the first model that really showed promise when it came to coding. And it can now verify that. It can look at what it had built and iterate from there. And that was quite an aha moment, pre-Claude code. But then also, we shipped computer use. So, it could start clicking around, taking screenshots, testing its own code. As well as MCP spec, which enabled it to use tools. So, then getting into Claude code. This is February 2025.
SPEAKER_08
So, this is about just over a year ago. Sonnet 3.7 was released. And this was state of the art on Sweebench. And Claude code was released in research preview. And I think an interesting quote that I pulled from this release actually. Is that the goal of Claude code was to better understand how developers use Claude. For coding to inform future model improvements. So, essentially when we released Claude code. The whole idea was for it to be somewhat experimental. To inform how we actually improve the base model itself. And you'll see this trend that over time the models become better. The harness, certain aspects of it might become less necessary. Or it will evolve.
SPEAKER_08
Just in terms of these slides as well. In the bottom left corner, these are some of the things that are the focus of these releases. Whether it's context or planning or verification. And then some stats. But I'm not going to read everything. So, yeah. Next, this was around May time of last year. Opus 4 and Sonnet 4 were released. And just in general, these tools got much better at managing their own context. And getting to task completion. Without reward hacking or anything like that. And then Claude code became GA as well. And we released the Claude code SDK. So, the harness powering Claude code. A little interlude here from the timeline.
SPEAKER_08
I think everybody now knows about this Ralph Wiggum technique. You might not know that it was actually last July. That this came out. When Geoffrey Huntley initially released the paper. Because it really gained a lot of traction around December or so of last year. When, for example, people started playing around with it themselves. Claude also released our own Ralph loop within the Claude code harness itself. But essentially, it's a simple technique. That you're just taking a prompt. And you're feeding it into Claude code CLI, for example. And then you're just running that on a loop until all the tasks are complete. It's a little bit deeper than that.
SPEAKER_08
I think people tend to simplify it. There's actually a few phases where it first would have some kind of planning. Where it breaks down that prompt into a few different features. And then it would pick one task from that. And start a new session. And then work with a fresh context window. So a lot of those concepts were applied in the Ralph loop. But I think why it caught so much attention is because it seems really simplistic. And he put it deterministically bad in an undeterministic world. So the idea being that it's better to fail predictably than it is to succeed unpredictably. When we actually created our own plugin for this in Claude code.
SPEAKER_08
You'll see, well, I don't know if people can recognize what the major difference is. There's some people say that's not a real Ralph loop. The idea is that this is just running within a single Claude code session. So it's not creating a fresh context window. It's just relying on compaction to happen over time. So, maybe it's not considered a real Ralph loop. But you'd set the max iterations. You'd set a safe word. And then essentially a stop hook would intercept when Claude would typically stop. And if it's not finished, it would just continue until it hits one of those exit criteria. Okay, so on to Sonnet 4.5.
SPEAKER_08
This was when the model just generally started getting better at handling its own context. So this is when it became more context aware. Tracking how many tokens had been consumed. So as it got towards the end of the context window, it understood that. And it could manage its own context. Claude code 2.0. So this is where we also shipped. This is where we introduced checkpoints. So actually keeping track of the code over time. Being able to rewind to previous parts of the session. And then we released, we just renamed the Claude code SDK to the agent SDK. And that's because we realized it's much more general purpose than actually just for coding.
SPEAKER_08
So you'll see we're talking about coding a lot right now. But I think what's very interesting is applying these long running harnesses to other domains as well. At this point, we could run for about 30 hours or so with Claude Sonnet 4.5. But then completing the family with Haiku 4.5 and Opus 4.5. This is where it got really interesting because all of a sudden running many sub-agents became really economical. And Opus 4.5 became really good at planning. So we could start doing things like using Opus 4.5 for planning. And then using Sonnet 4.5 as the workhorse for really executing all of that code.
SPEAKER_08
And then there's a big couple months as well because this is when we released skills. Which again, very good at making effective use of the context with this notion of progressive disclosure. So just the front matter of the skill is loaded in. But then completing the family with Haiku 4.5 and Opus 4.5. This is where it got really interesting because all of a sudden running many sub-agents became really economical. And Opus 4.5 became really good at planning. So we could start doing things like using Opus 4.5 for planning. And then using Sonnet 4.5 as the workhorse for really executing all of that code.
SPEAKER_08
And then there's a big couple months as well because this is when we release skills. Which again, very good at making effective use of the context with this notion of progressive disclosure. So just the front matter of the skill is loaded in. Instead of all of your tool descriptions, which can consume quite a lot of the context window up front. And then the entire rest of the body of the skill is loaded in if it's instantiated. Followed by some references to even code that could run more deterministically. And then more context improvements. Things like programmatic tool calling.
SPEAKER_08
So instead of running a bunch of tools, pulling all of that into context and then trying to process it. Actually just writing code on the fly and being able to run a series of tool calls. And then just get the final result back. And again, this is all just to improve the usage of the context window. Okay, so a lot going on on this slide. But at this point, this is around November time. We released our first blog post on long running agents and how you would go about building these. So a lot of the concepts I've already described should make this fairly easy to understand actually. Where a human would write something like, write me a browser.
SPEAKER_08
Or create a Slack clone or a Salesforce clone. Just something really, really vague. And the first thing that would happen in this harness that we built is there's an initializer agent. That would take that simple prompt and it would break it down into a series of persistent artifacts. The first being a feature list of x number of features. Feature list dot JSON because we actually found the models might overwrite markdown files. Whereas they're less likely to just overwrite JSON files. Which is interesting. It would also write a progress file. Of course, start the git repo. Build an init script. And then just have a flag for whether the features are complete or not.
SPEAKER_08
If they pass all the tests. From there it would go into this harness loop where there's multiple different steps here. So the first one is, in a fresh context window. Just getting the bearings. What's the present working directory? What's the progress file say? Okay. And then doing a smoke test or running the init script. So it didn't have to figure out how to do that every time. Get the server up and running, et cetera. And then picking one feature. Only one feature. That hasn't passed all of the tests. Implementing that feature. Doing some actual tests. Much verification loop. Much like a human would do. Using puppeteer in this case. And then if everything passes.
SPEAKER_08
Actually writing the git commit. And changing the state of this particular feature. It passes. And then if there are any features that are unfinished. Just continuing that loop in a fresh context window. So we're starting to layer in a lot of these concepts here. Fresh context windows. These persistent artifacts. Verification loops. Really good planning up front. You'll see this is the first iteration of these long running harnesses here. Okay. So continue with the history tour. So then Opus 4.6. Sonnet 4.6. These models were really great. Because Sonnet 4.6 was basically offering that Opus level intelligence more at the Sonnet price.
SPEAKER_08
And it became very much a workhorse for a lot of cloud code. And Opus 4.6 just became really good at planning. We called it very much an agentic model. So Opus 4.6 was great at deciding which tools to use. And just being able to run for much longer. If you recall that meter chart. You'll see that this was a jump from about four hours up to 12 hours. With that very simple harness. So this model is very agentic. And then along with that. With some of the research we had done. We released agent teams. Which the idea being in cloud code. This more general purpose way for you to scaffold out your own set of custom agents.
SPEAKER_08
And the innovation with agent teams is that instead of everything reporting back into the main agent. The actual sub-agents could communicate with each other. So they sort of had their own way to coordinate. And then report back to the main agent. Only when it was required. We also introduced server side compaction. Which basically means that these models can now just run indefinitely. And compaction could just happen on the server side. And then this one million context GA. So now we have one big context window. You see the models are getting better. Maybe you can just run a lot within a single context window even.
SPEAKER_08
Instead of necessarily needing new sessions all the time. You see how things start shifting over time. So that's the whole overview. You can see all of the different releases that I shared here on this table. And you can see how it's changed from Sonnet 3.7 at one hour. To 12 hours with Opus 4.6. And then we have our own anecdotes as well. Where tasks would take 20 minutes when it was Opus 3.5. And now we're building fully fledged apps. That you don't have to run for 30 hours. They can run typically we're seeing 3 to 5 hours. You can build a really fully featured application. That runs out of the box.
SPEAKER_08
So what's really interesting is the harness doesn't just disappear as the models get better. It's really evolving as the models change over time. To 12 hours with Opus 4.6. And then we have our own anecdotes as well. Where tasks would take say 20 minutes when it was Opus 3.5. And now we're building fully fledged apps. That you don't have to run for 30 hours. They can run typically we're seeing 3 to 5 hours. You can build a really fully featured application. That runs out of the box. So what's really interesting is the harness doesn't just disappear as the models get better. It's really evolving as the models change over time.
SPEAKER_08
And it's really fascinating to find the gaps in the model. And then fill that in with the harness. And then you train the model on using that aspect of the harness. Then maybe at some point you actually remove that entirely. And this iterative loop just keeps happening over time. With more and more of these co-releases that we have. So hopefully that was an interesting trip. Back through the Claude evolution. And how it applies to the long running agents. And so I'll hand over to Ash to continue with where we are today. In terms of the state of the art. Thank you. Thank you. Thank you. [SPEAKER_06] All right. [SPEAKER_06] Quick question.
SPEAKER_08
[SPEAKER_06] Any of you guys have any agents running at the moment in the background? [SPEAKER_06] Doing work while you're here? [SPEAKER_06] Just one, two, three? [SPEAKER_06] Okay. [SPEAKER_06] Probably should be more of you. [SPEAKER_06] Hopefully by the end of this you'll have some ideas to take away and actually put into practice. [SPEAKER_06] So, that's the history. [SPEAKER_06] And I quite like that quote that Andrew talked about, [SPEAKER_06] where the frontier doesn't really shrink, [SPEAKER_06] it just moves. [SPEAKER_06] And so, what I wanted to talk about is some very simple harness patterns [SPEAKER_06] that we've been playing around with internally
SPEAKER_08
[SPEAKER_06] that we use to build these very fancy one-shot demo apps. [SPEAKER_06] But also, we're experimenting with this stuff [SPEAKER_06] in post-training, in RL, how do we make our models [SPEAKER_06] and the general behaviors more adept at autonomous work. [SPEAKER_06] So, if you've ever tried to get an agent [SPEAKER_06] to review its own PR, [SPEAKER_06] you'll understand where this is going. [SPEAKER_06] So, this general idea is shamelessly stolen [SPEAKER_06] from GANs, generative adversarial networks. [SPEAKER_06] So, you have this generator model [SPEAKER_06] and then you have some sort of discriminator
SPEAKER_08
[SPEAKER_06] and you have some sort of adversarial pressure between them. [SPEAKER_06] The generator builds, the evaluator grades, [SPEAKER_06] and the whole idea here is we're splitting up [SPEAKER_06] the context windows, system prompts, [SPEAKER_06] the jobs entirely. [SPEAKER_06] The evaluator here isn't just reading diffs, [SPEAKER_06] but it's actually using Playwright to open live pages, [SPEAKER_06] click around, try things out, [SPEAKER_06] and then it eventually hands back [SPEAKER_06] whatever critique gets decided back to the actual generator [SPEAKER_06] and you continue that loop. [SPEAKER_06] Contrast that with what most people today are doing,
SPEAKER_08
[SPEAKER_06] which is using one core code session, [SPEAKER_06] telling it to check its own work, [SPEAKER_06] and looping that way. [SPEAKER_06] So, the obvious question for me at least is, [SPEAKER_06] if the evaluator is also just an LLM, [SPEAKER_06] why doesn't it just rubber stamp it too? [SPEAKER_06] And so, the key idea that we're exploiting here is, [SPEAKER_06] yes, the evaluator is still a large language model, [SPEAKER_06] and yes, it's still going to be biased towards liking [SPEAKER_06] large language model style outputs, [SPEAKER_06] but tuning a standalone critic to be harsh [SPEAKER_06] is actually very tractable,
SPEAKER_08
[SPEAKER_06] but tuning a builder to be somewhat self-critical is not. [SPEAKER_06] I think a really good analogy for this is the same as humans. [SPEAKER_06] It's very easy for me to critique a lovely piece of artwork [SPEAKER_06] or a fine meal, [SPEAKER_06] much harder for me to actually paint that [SPEAKER_06] or cook that meal myself.
SPEAKER_08
[SPEAKER_06] So, what we're doing here is exploiting the gap [SPEAKER_06] between the ability of an LLM [SPEAKER_06] to be a critic versus a generator. [SPEAKER_06] So, the next thing I want to talk about is [SPEAKER_06] how do you actually think [SPEAKER_06] about designing these critics? [SPEAKER_06] It's very similar to the process of creating good evals, [SPEAKER_06] but in the context of full-stack apps, [SPEAKER_06] there are a lot of fuzzy areas [SPEAKER_06] which go into what makes something good. [SPEAKER_06] It's not just does it work, [SPEAKER_06] but does it look good, does it feel good, [SPEAKER_06] is there an element of taste
SPEAKER_06
[SPEAKER_06] in these products as well? [SPEAKER_06] So, this is where we've been doing [SPEAKER_06] a lot of experimental work, [SPEAKER_06] especially when trying to imbue Claude with design taste in post-training, but also create these front-end design skills that we put out there and generally improve the front-end design ability of our models. So, the way we think about this is, most people say you can't grade taste, but we think you can if you have a strong enough opinion on it and you just write it down.
SPEAKER_06
And so, the way we do this is with creating a rubric with four criteria: design, originality, craft, and functionality. We actually weight this towards design and originality. We've shifted the weightings between these four things depending on which model's in play, but at the moment, Opus 4.6 is pretty good at functionality already, so the problem that we're trying to overcome is how do we prevent things like purple gradients, general AI slop type aesthetics in general. And we just go ahead and calibrate this with a few short examples on reference sites, so the evaluator's taste converges on our own. And let me show you an example.
SPEAKER_06
between these four things depending on which model's in play, but at the moment, Opus 4.6 is pretty good at functionality already, so the problem that we're trying to overcome is how do we prevent things like purple gradients, general AI slop type aesthetics in general. And we just go ahead and calibrate this with a few short examples on reference sites, so the evaluator's taste converges on our own. And let me show you an example of what this actually looks like in practice. So, this is an example of a model going through this similar loop: generator, evaluator launches playwright, navigates, screenshots, scores on those four criteria, writes critique, and then hands back to generator. So, all of these examples are HTML and CSS only that I've gone through for maybe four to five hours, 5 to 15 rounds. I think the interesting thing here, which is quite unique and something which you wouldn't necessarily get if you're just using a single agent loop, is that the thing pivots, right? So, imagine the generator gets stuck on one of the four criteria. Let's say it's really struggling and constantly scoring low on originality. This GAN style harness which we're using will just throw the whole thing out and try again from scratch. Whereas, in a single pass generation or a RALF loop, it keeps trying to patch the same thing. And this ability to course correct over very long time horizons is something which is quite unique to breaking down different roles that go into building something. So, that was just an insight into how we think about the frontend component. But how to go from nice pages to fully working apps, we added one more role: a planner. And so, again, it sounds very simple. It's ultimately just taking a one line prompt and then breaking it down into a very deliberately high level spec. So, what it does is actually spec the granular—sorry, it specs the general workflow into a series of sprints. What it doesn't do and what most harnesses do today is necessarily try and plan the granular technical details of the product. The reason being is, one, it's very likely to still make an error. But when it does make an error, it's going to cascade through every single one of these sprints and magnify errors over a multi-hour time horizon. If you squint at this, this is just a very simple PM, IC, and QA kind of org structure, right? We didn't invent this. We just gave each role its own context window. And the bit which is interesting, I think, to talk about is the glue between the generator and the evaluator in this setup. So, before the generator actually goes ahead and writes a single line, we have the two agents basically negotiate what done actually means. And so, let's say the generator proposes, I'm going to build X feature and you should verify it by testing Y. The evaluator might push back and be, actually, the scope is too big, and those tests that you propose are a bit too weak and you've missed XYZ edge case. And you basically have this back and forth via files on disk. One writes the markdown, the other reads it, and you iterate until both agree. And then once you reach that condition, then you actually start building. And then the evaluator grades against the contract that those two agents have decided between themselves, not the original spec, which the planner has one-shotted at the beginning. And why this matters is it bridges this idea of user stories—the spec—and converts it into slightly more tangible, testable assertions, some sort of contract, without the planner having to over-specify up front. And I think this is the key innovation that the RALF loop never really had. It had a fixed plan, .md style, but nobody on the other side
SPEAKER_06
i.e. the spec, and converts it into slightly more tangible, testable assertions, some sort of contract, without the planner having to over-specify up front. And I think this is the key innovation that the RALF loop never really had. It had a fixed plan, .md style thing, but nobody on the other side is necessarily arguing with the main loop. And again, it comes back to having these separate context windows and adversarial pressure. So let me show you an example of a very simple prompt that we had in a solo loop versus the harness that we just discussed. So the prompt was basically build a retro game maker. And that was it. And I'm not going to try and convince you that this is necessarily the most cost-effective or most efficient way to try and build an app. As you can see, one, it takes, at the moment, an extremely long amount of time. Two, it's very expensive. But also, as we'll see in a second, a lot of the stuff actually starts working only with this harness when it didn't in a solo loop. So this is what it looked like, the opening screen, at least when we didn't have the harness. Pretty simplistic, a little bit boring. But it still looks nice, right? If this were the whole app, you'd ship it. But it's the bait, I guess, if you will. This was the sprite editor, if you will. Again, it still looks fine. The canvas is there, the palette, the frame timeline, live preview. Maybe it's a little bit cramped. And the color picker is just black swatches, but it works. Clearly, the agent actually did understand what it was trying to do. And then, the one thing that you actually have to do, which is play mode. Entities rendered, score, health, all the other things which go into an actual game. Pressing an arrow key does nothing. Pressing a space key did nothing. The agent really didn't have any idea how to test itself, what it actually meant to play a game and actually succeed. And, yeah, this is the same prompt, same model, and this is the breaking point. It looks done on the surface, but when you try and actually push it to its limits, it just failed. And then, if we ran the same prompt with the same model, this is what it looked like when we ran the harness. So this was about, yeah, 200 bucks, six hours. First up, it decided to name itself Retroforge. It decided to create a new project dialogue, have a very nice canvas. None of that was in our prompt. So this was all the planner deciding, okay, here's what the product decisions should look like. And then, the two other agents deciding, right, how am I going to test this? If we look at the sprite editor, we have a full 54 color palette, the 8-bit preset from the project dialogue flowing through. You see the sprite at actual game scale. It's a lot more complete as a product in general. We had a whole new AI level assistant. This is where it started to get recursive. The planner had decided, right, we should have some AI features, which is just a vague line in the spec. And the harness turned that into a full AI level assistant inside the app that it was building. So, someone could come in and say, hey, create a castle with sprites guarding it, let's say. This is something which a solo run would never even attempt to look at. Without the planner, that phrase just becomes a work item to look at. Then finally, I guess, the actual results kind of applied. So, play modes. You know, you have this whole debug HUD in the top left, which you can clearly tell is to make life easier for the evaluator.
SPEAKER_06
This is something which a solo run would never even attempt to look at. Without the planner, that phrase just becomes a work item to look at. Then finally, the actual results applied. So, play modes. You have this whole debug HUD in the top left, which you can clearly tell is to make life easier for the evaluator, for example. Those numbers are live. The physics loop is actually running. Arrow keys work. The player moves. Collider's Castle Wars. Because the evaluator actually launched the game, tried to play it, knew what features need to be tested to make this game real and successful. And the difference between this output and the previous output is entirely just scaffolding. It's a very simple loop, ultimately, but the results are quite startlingly different, at least. And so, in case you're curious, the things which the evaluator did catch are pretty basic stuff. It's things like fast IPA route ordering, passes every unit test but might actually break in prod. The evaluator catching things like the delete key having some kind of Boolean logic bug. Again, these are things which were only caught because the evaluator is actually using the app. It's things which might get through CI in a RALF loop, but this isn't that level of specificity which isn't something which happened by accident. And so, this is the kind of level of detail which these models are going at this point in time. So, we talked about the contracts that the generator and the evaluator would write between themselves. For this app, it decided that there were 27 contract criteria. That's a level of granularity which we found that you really need to make findings actionable. If you have vague criteria, you have vague critiques. The generator just shrugs and does things. Whereas, if you have granular criteria, the agent knows I need to fix this exact line. What's interesting, I thought, and I want to be honest about this part is that out of the box, Claude is a really bad general QA agent. Andrew talked about this a little bit in his bit, right? But the same sycophancy and generosity bias that everyone hits with general elements of judge systems also applies here. Most of the time in early runs, the QA agent would find a bug and be like, fix it later, might take two weeks and then just be done with it. So we actually had to spend an exorbitant amount of time going through, trying to tune small layout bugs, edge cases, and feeding that into the prompts. I wish there was some secret to actually doing this, but realistically, the whole art to building this system and making it good was reading the traces. The primary debugging loop was this and not necessarily running more experiments. It was reading what the agent actually did, finding where its judgment diverged from ours as humans and then tuning the prompt for that. It was the same muscle as reading a stack trace. One tooling tip that we had was piping agent transcripts into files, grabbing them with another agent or having another agent go through them and then update the prompts itself. So you have some sort of closing the loop even on just building this harness out. So the last thing I wanted to talk about was how to think about adjusting your harness as these models get better in time. I think there's a lot of discussion around whether harness design is dead or null, especially with models that, I mean, when I wrote this, it was just Opus 4.6.
SPEAKER_06
the prompts itself. So you have some sort of closing the loop even on just building this harness out. So the last thing I wanted to talk about was how to think about adjusting your harness as these models get better in time. I think there's a lot of discussion around whether harness design is dead or null, especially with models that, when I wrote this, it was just Opus 4.6, but even like mythos level models. And I think the key thing that we noted is it's really important to get a feel for what the spiky behaviors of any individual model are and then try and adapt your harness to fill the gaps.
SPEAKER_06
So Andrew talked about this a bit, but context resetting between sessions. We dropped that entirely. Opus 4.5 used to have really bad context anxiety, whereas Opus 4.6 just didn't as part of post-training there. And so one continuous session and compaction was more than enough to handle very long sessions. Sprint decomposition. We don't have a very strong opinion on this, but it was something which was really critical to getting Opus 4.5 to work, but Opus 4.6 was able to hold a two-hour continuous build coherently in a way without necessarily having to be force-fed one feature at a time.
SPEAKER_06
The cadence at which the evaluator should run. Previously, we were running at every single sprint per se, whereas now, we were just running at the end of a one-shot generation from the model and then passing back. So the harness is still the same. We're just simplifying the specific loops and the recipe that goes into it. The lesson isn't necessarily that our harness was wrong, but rather, it was right for 4.5, the frontier moved, and we ran a simplified version to see how it'd work.
SPEAKER_06
So this is what the final setup looks like today. Having that planner generator evaluator loop is still the core of our system, but you can see we ditched a bunch of the other components that made this slightly more complicated than it had to be. We also, as mentioned, are a big fan of just using a file system for shared state instead of leaning on context windows for very long-running agents in general.
SPEAKER_06
And this is an example of the simplified harness running with one of our latest models. Again, very expensive, but you can see it's actually roughly half the cost of the previous runs just because we're doing things in a slightly more simplified manner, but it's still running over a very extended period of time. And so this is an example of a DAW, which is basically a music creating app. The agent sets the tempo, a key, it lays down the melody, it builds the drum tracks. This is the evaluator actually going and testing the app itself.
SPEAKER_06
We did actually listen to the music in this. Obviously, Claude can't hear at the moment, and so the music was pretty bad, but the app was really good in general and pretty fleshed out, which a model ago, this is something which would never have worked. But this is something which was possible with just a couple rounds. And this is that meter curve which Andrew was talking about in action.
SPEAKER_06
And so I wanted to close just by saying you don't necessarily need our internal harness to go away and start thinking about this. We are constantly trying to ship bits of these primitives into Code Code directly, but also there's nothing stopping you from going ahead and building something similar to this on your own. So we just shipped auto mode, which is probably my favorite thing for slightly more safe yellow, if you will, instead of running dangerous C skip permissions all the time. We already have custom subagents as a primitive, right? Your evaluator, your QA role, give it a harsh system prompt and a very detailed rubric.
SPEAKER_06
Playwright MCP or Claude for Chrome MCP, already extremely good at web app stuff.
SPEAKER_06
My favorite thing for slightly more safe yellow, if you will, instead of running dangerous C skip permissions all the time. We already have custom subagents as a primitive, right? Your evaluator, your QA role, give it a harsh system prompt and a very detailed rubric. Playwright MCP or Claude for Chrome MCP, already extremely, extremely good at web app stuff or just use computer use if you're building native apps. And skills, again, a very nice way to package your grading rubrics into your general development flow. So, yeah, five things if you're taking a photo, this is the slide I would say to remember. Self-evaluation, very much a trap, just use an adversarial evaluator. Compaction does not equal coherence, right? Lossy summaries really drift. Structured handoffs and clean context are a very good pattern that we've seen. Don't think that subjective quality isn't gradable. If you have a strong view on what something should look like, then force yourself to write it down. We found this made a really massive difference to the quality of apps that a model was able to generate. And then finally was really just sit with them all, read the traces. Only then can you really know what bits of a scaffold to delete, what bits to keep, especially as the frontier moves. But, yeah, that's it from me. Thank you very much for listening. And check out our blog post. But I wanted to just open up for Q&A in general because we've been talking for close to an hour now. So if you have any questions for me and Andrew, just fire away. We'll do our best to try and answer them. Yeah. Thank you.
SPEAKER_06
[SPEAKER_03] Johan from Paulside. One question for you. When you improve the evaluator by reading the logs and improving it, is that on a per-project basis or more of a secret source that you reuse across project?
SPEAKER_06
The goal was very much to try and do this in a way which was reusable, right? I think anyone can tune this in a way that's creating a very specific type of app. That's fine. At that point, it's not that different from going ahead and just prompting tool code yourself and doing it, right? I think the key was, what are the common patterns that you can draw across the model weak points, right? So talking to that front-end design piece, we knew what we thought good design would be, right? You could give examples like this is what a read-view-for-pro looks like. This is what AI slop looks like, right? And that generalizes quite well. So, yeah, this was all around web apps, but it could quite easily apply to other things as well. Thanks for.
SPEAKER_06
[SPEAKER_09] Thank you for the presentation, very interesting. I was just wondering, what is your view on concept of dump zone and smart zone of a model? So I understand before it was around 40%, now with one million context, it's about 100k is what I understand. And the way I understood Ralph Loop was designed is to negotiate this problem. So basically, we're keeping the model always in the smart zone, basically trying to slice the task below 100, so it executes the task within the 100 context zone. And what I understand from your presentation, you can advocate.
SPEAKER_06
[SPEAKER_04] Context, it's about 100k is what I understand. And the way I understood Ralph Loop was designed is to negotiate this problem. We're keeping the model always in the smart zone, trying to slice the task below 100, so it executes the task within the 100 context zone. And what I understand from your presentation, you can advocating not to use it anymore because we can now rely on a compaction and so on. Is it something you're suggesting to do or with still with a Ralph Loop model still has its own place given the smart and dump zone concept?
SPEAKER_06
[SPEAKER_08] Well, I suppose from Ash's presentation and mine, you see that the one million context window is now GA and so you have a bigger context to use. The models are more agentic, so they can maintain coherence for a longer period of time within that context window. And that actually, with the release of 4.6, we decided to move from new context windows to just a single, long-running, continuous session with compaction. So I think, whether or not you use multiple fresh sessions or just one long-running one is probably still up to your use case and your evals depending on what you're seeing is working best, but at least for this general purpose generator evaluator pattern with Opus 4.6, we saw that it was possible to use a single session. I don't know if you want to add to that.
SPEAKER_06
I think it's also just a temporary problem, right? Context rot is a failing of today's models to some extent and much less so than even just one model generation ago. So is there a place for the type of thing which you're discussing? I think yes, depending on your use case, but it's not a piece which I'd look at as okay, I'd be hunting for the model release where I can strip it out. Let's put it that way.
SPEAKER_06
[SPEAKER_01] I always have a lot of FOMO around playwright. I mean, you said playwright MCP, there's playwright skills. Can you speak to how to improve the playwright? Because I imagine I would like to have my browser open and then I can see the model working through it and then maybe I could steer it, the few tabs open. But what is there some innovation there I'm missing out on or is playwright MCP really, would you recommend people use?
SPEAKER_06
I mean, playwright MCP or just use the called for Chrome MCP which is a slightly more robust thing, I guess, around browser control. I mean, I don't know why you want to watch it do things. I mean, you can, but I think that's a trust gap, right, today. The whole point of what we're trying to get to do here is you set something off, you trust it to do it, do the work and test it and you have the confidence that it's doing it correctly and you come back to it. And that's where, yes, there's going to be some iteration at the beginning where you're watching, reading the traces until you get to a point you can trust it. But at least internally, right, when I'm doing full stack app dev, I have got to a point now where I'm okay, with Opus 4.6 I can reliably trust the model to go ahead, read, network errors, console errors, actually navigate an app, zoom in where it needs to. The vision is now good enough on these models that it can identify overlapping text on elements and things like that.
SPEAKER_06
Now where I'm okay, with Opus 4.6 I can reliably trust the model to go ahead, read, network errors, console errors, actually navigate an app, zoom in where it needs to. The vision is now good enough on these models that it can identify overlapping text on elements and things like that, whereas that just wasn't the case until realistically the last generation of models. So, yeah, I would recommend. [SPEAKER_00] I'm curious, with the generator evaluator pattern, what happens? Can you throw unlimited tokens at it or will it stop because the evaluator is not good enough? Can you tell me more about that? Sorry. Do you mind clarifying? I messed up.
SPEAKER_06
[SPEAKER_00] Yeah, so okay, let's say I say, okay, create a very cool game with some features. You have to generate an evaluator pattern that creates contracts, builds the apps. If I, then it will give me back something, right? Can I restart it again and say, okay, make it better, I'm not happy about it and generate the evaluator will pick, the pattern will pick it up and make it better? Yeah. Or will the evaluator not go to at one point and just say, this is it?
SPEAKER_06
I think that's. One, first of all, if you want some level of human in the loop in this process, that's just implement hooks at some point in this loop. I think the bit which was surprising to us was with this general pattern and especially with the kind of 4.6 liner models, both Sonnet and Opus, it was extremely willing to throw away everything, even if it had done kind of ten passes at something, it was very happy to just throw it all away and start from scratch if, for some reason, it wasn't able to hill climb against the rubric of the evaluator in an effective way. And so, that's why, when we're playing with this kind of thing, we didn't naturally lean towards having some kind of resume or human-in-the-loop type intervention system. And we didn't really observe, we expected to, but we didn't really observe that kind of behavior which you were talking about, where it just evaluators, I just give up, let's just pass it on, shall we say? Yeah, it was just much more willing to throw away everything and restart. And that was just a behavior which we never saw when it was the generator itself almost being proud of its own work and being, I'm not going to restart this whole thing. So, yeah. I mean, there's been an example which I've seen where the evaluator is it kind of gets fed up and is, right, this approach you're taking just obviously isn't working. Can you just delete everything and restart? Which, I don't know about you guys, but vibe coding regularly I often do as a human to benefit from fresh context windows, not have to deal with an already messy code base, et cetera. So it's quite neat seeing models now also get to that point.
SPEAKER_06
Delete everything and restart? Which, I don't know about you guys, but vibe coding regularly I often do as a human to just benefit from fresh context windows, not have to deal with an already messy code base, et cetera. So it's quite neat seeing models now also get to that point. And also,
SPEAKER_06
[SPEAKER_08] just briefly add, obviously, you can then open that code base in Cloud Code and continue where you left off. It goes without saying. And I think we're generally thinking about what the workflow looks like if it's more back and forth because there's the extreme of build me a really complex gaming application that you don't know is it going to take three hours, is it going to take 20 hours? It's a bit unclear, so maybe there's something in the middle that's more of a feedback loop. Hi. [SPEAKER_11] I really like the idea that you have,
SPEAKER_06
[SPEAKER_10] there's a human element here where it's PM, engineer, evaluator. PM role is a lot of the time it's scope creep and keeping the time going and stuff. But you're just letting this off. You're letting engineers go play in the sandbox for ages. Is there a harness loop that needs to go back to the planner eventually? Does it need to move again? Well, maybe
SPEAKER_06
because we're engineers we just decided, ah, screw the PM, we'll just stuff it to the side. You guys need a PM. We actually, well, this is where the contracting piece between the evaluator and the builder worked quite well. For context, in that, we typically insert the main spec that was generated by the PM per se into those sessions regularly. So that it's always a reference point for okay, this is what we're still actually trying to build. And the main function of then the builder and the evaluator is just to figure out the exact feature set and tests and contracts per se that actually satisfy that spec.
SPEAKER_06
But the reason we don't is because we don't want the planner to be a core part of this loop. It should be very high level. Its purpose really is just to set out the hard outer lines of what this product could be. But its job is not necessary to come in and intervene and be like, actually, this is an impossible feature. We should not do this and edit itself. We wanted to keep that context relationship between just the builder and the generator.
SPEAKER_06
That being said, this loop, I've applied it in lots of different ways. It doesn't have to just be one generator and one builder, right? That adversarial trade-off can be applied to a workflow consisting of multiple separate agents, right? It could, if you're trying to do generate evals, let's say, you could use a similar harness to be like, hey, generate it could be planner, generate a synthetic data sets, right? With a QA agent. Then hand off to an integrator, which actually wires up something. Also has a QA agent. Then has a final, you can basically add this generator, evaluator thing into a multi-step workflow. We each build a, maybe has a slightly different
SPEAKER_06
synthetic data sets, right? With a QA agent. Then hand off to an integrator, which actually wires up something. Also has a QA agent. Then has a final kind of generator, evaluator thing into a multi-step workflow. We each build a, maybe has a slightly different function, per se, as part of a longer workflow. So there are different ways in which you keep things on track depending on the task and break down this general pattern into slightly more specified tasks or workflows, if that makes sense.
SPEAKER_06
[SPEAKER_05] You mentioned that some of the later tasks could not possibly be done by an earlier model. Can you talk a little bit about your process comparing the tasks on the different models? Like, do you fire off the same task on Opus 4.6, Opus 4.5, Sonnet, or is this artisanal, co-evolving harness model set up operating that?
SPEAKER_06
[SPEAKER_08] Yeah, I mean, I suppose we walk through the history a little bit, and if you look at, say, the first blog post on long-running agents versus the more recent one, there are some pretty significant differences there. One being what we were just discussing, that the initializer agent would build this super comprehensive spec of, say, 200 different features, and then the loop would have to actually go and execute against every single one of those features, which may lead to incorrect design decisions, but it's forced into that behavior, whereas I think now you're able to have a more generic creative direction set with, say, Opus 4.6, and then just having this loop of the generator evaluator. But it does, yeah, your model selection does inform your harness design very much so. Of course, in a perfect world, you could just throw everything and say Opus 4.6, but if you have cost concerns, for example, maybe you do use Opus 4.6 for planning and then Sonnet 4.6 for the coding or the execution, that's something that we tend to see quite frequently. But again, if you're building specific sub-agents for each of these, you probably want to have some evaluations to be able to understand for that model and that prompt how it's performing against that task and then just optimize.
SPEAKER_06
[SPEAKER_09] Do you have any advice on moving beyond these one-shot applications to long-lived products where you're looking to make changes days, weeks later? What sort of artifacts you need to persist to future instances to be able to know what has come before, what can I change, what should I change?
SPEAKER_06
Yeah, it's something which we're working on. Right now, we use similar patterns for just a bunch of random stuff internally, shall we say. And so, at the moment, it's set this thing off. It's running on a remote server somewhere and I'll just come back and check it like after this talk, let's say. And then I kind of iterate on it manually in whole code directly, polish any rough edges, that kind of thing. I think in terms of the way that you're actually setting up this harness, just having, this is why we kind of default to using a file system state for this kind of loop. One, because it's just very easy for another model to come up and grip through and pick up what's been going. But one thing which I like to do is embed little bits of prompting throughout this kind of loop which basically tells it to write learnings and states to some kind of JSON file because the model doesn't overwrite that too much. And so, the nice thing about that is you're basically just leaving breadcrumbs.
SPEAKER_06
and grip through and pick up what's been going. But one thing which I like to do is embed little bits of prompting throughout this loop which tells it to write learnings and states to some JSON file because the model doesn't overwrite that too much. And so, the nice thing about that is you're just leaving breadcrumbs for another model to come and pick up. So, honestly, the key thing for me is how do I instruct this harness to leave crumbs for a human to come in and then use code on top of.
SPEAKER_06
So, generally, it's like, hey, the shape of that file might be like tried this, evaluator found this bug, implemented this fix, this fix worked, yes, tick. And then continue. And you have a time stamped time log, if you will, of everything the model has tried, the fix it's made and the final state. And then also having some sort of live updating set of docs, if you will, just very high level. Here's the file structure. And then those two files, to be honest, are more than enough for code and a human to come in and start iterating on the app with. But that's all we're doing at the moment.
SPEAKER_06
[SPEAKER_02] It's very interesting [SPEAKER_02] to hear the, [SPEAKER_02] oh yeah, [SPEAKER_02] first of all, [SPEAKER_02] congrats on the presentation. [SPEAKER_02] Thank you. [SPEAKER_02] And then [SPEAKER_02] I was wondering, [SPEAKER_02] there are two approaches. [SPEAKER_02] You have the agent team [SPEAKER_02] where multiple agents [SPEAKER_02] interact with each other. [SPEAKER_02] And then [SPEAKER_02] the explicit [SPEAKER_02] generator critic [SPEAKER_02] setup. [SPEAKER_02] Because in sense, [SPEAKER_02] the agent team [SPEAKER_02] has the same setup [SPEAKER_02] where the main agent [SPEAKER_02] instructs someone [SPEAKER_02] and then can act
SPEAKER_06
[SPEAKER_02] as a critic [SPEAKER_02] for the subagent. [SPEAKER_02] But what are [SPEAKER_02] the current failure modes [SPEAKER_02] that causes us [SPEAKER_02] to still need [SPEAKER_02] the specific generator [SPEAKER_02] critic harness [SPEAKER_02] instead of just [SPEAKER_02] the agent team [SPEAKER_02] itself? [SPEAKER_02] And then [SPEAKER_02] what's your estimate [SPEAKER_02] of how many [SPEAKER_02] model generations [SPEAKER_02] we would need [SPEAKER_02] to just [SPEAKER_02] completely rely [SPEAKER_02] on the agent team? [SPEAKER_02] Maybe, [SPEAKER_08] I mean, [SPEAKER_08] I can address [SPEAKER_08] the first aspect [SPEAKER_08] of that. [SPEAKER_08] So,
SPEAKER_06
[SPEAKER_08] I mean, [SPEAKER_08] one of the limitations [SPEAKER_08] of, [SPEAKER_08] firstly, [SPEAKER_08] cloud code [SPEAKER_08] is using [SPEAKER_08] the same harness [SPEAKER_08] that is the agent SDK. [SPEAKER_08] So, [SPEAKER_08] you can, [SPEAKER_08] technically, [SPEAKER_08] you should be able [SPEAKER_08] to build this type [SPEAKER_08] of a pattern [SPEAKER_08] into cloud code. [SPEAKER_08] Agent Teams [SPEAKER_08] is a useful framework [SPEAKER_08] for potentially [SPEAKER_08] doing that [SPEAKER_08] because you could, [SPEAKER_08] say, [SPEAKER_08] have the generator [SPEAKER_08] and the evaluator [SPEAKER_08] intercommunicating [SPEAKER_08] or maybe it's
SPEAKER_06
[SPEAKER_08] the generator [SPEAKER_08] is [SPEAKER_08] the main agent [SPEAKER_08] and the evaluator [SPEAKER_08] is, [SPEAKER_08] say, [SPEAKER_08] a member [SPEAKER_08] of the agent team. [SPEAKER_08] But I think [SPEAKER_08] it's evolved [SPEAKER_08] more so from [SPEAKER_08] that first blog post [SPEAKER_08] that I shared. [SPEAKER_08] I think that was [SPEAKER_08] the results [SPEAKER_08] of that to some extent [SPEAKER_08] to try and make that [SPEAKER_08] more generally available. [SPEAKER_08] But one of the things [SPEAKER_08] you're limited by, [SPEAKER_08] obviously, [SPEAKER_08] is cloud code [SPEAKER_08] would just have to run [SPEAKER_08] on your machine.
SPEAKER_06
[SPEAKER_08] I think with the agent SDK, [SPEAKER_08] you can also just run it [SPEAKER_08] in more of a cloud [SPEAKER_08] environment [SPEAKER_08] and a sandbox environment [SPEAKER_08] for long periods [SPEAKER_08] of time [SPEAKER_08] and without it failing [SPEAKER_08] or you having to run [SPEAKER_08] the caffeinate [SPEAKER_08] on your machine. [SPEAKER_08] But I think, [SPEAKER_08] yeah, [SPEAKER_08] cloud code is a good [SPEAKER_08] testing ground [SPEAKER_08] for building out [SPEAKER_08] any of these types [SPEAKER_08] of harnesses [SPEAKER_08] to experiment [SPEAKER_08] and explore [SPEAKER_08] and see what works [SPEAKER_08] before maybe you build
SPEAKER_06
[SPEAKER_08] it into the agent SDK [SPEAKER_08] and then actually deploy [SPEAKER_08] it as its own application. [SPEAKER_08] And yeah, [SPEAKER_08] I mean, [SPEAKER_08] again, [SPEAKER_08] I would just experiment [SPEAKER_08] and see if agent teams [SPEAKER_08] is something that makes [SPEAKER_08] sense for you [SPEAKER_08] or if maybe just using [SPEAKER_08] regular subagents [SPEAKER_08] or some other [SPEAKER_08] framing of it [SPEAKER_08] that works better. [SPEAKER_08] But yeah, [SPEAKER_08] I think people are using [SPEAKER_08] and explore [SPEAKER_08] and see what works [SPEAKER_08] before maybe you build [SPEAKER_08] it into the agent SDK
SPEAKER_06
[SPEAKER_08] and then actually deploy [SPEAKER_08] it as its own application. [SPEAKER_08] Again, [SPEAKER_08] I would just experiment [SPEAKER_08] and see if agent teams [SPEAKER_08] is something that makes [SPEAKER_08] sense for you [SPEAKER_08] or if maybe just using [SPEAKER_08] regular subagents [SPEAKER_08] or some other [SPEAKER_08] framing of it [SPEAKER_08] that works better. [SPEAKER_08] But I think people are using [SPEAKER_08] agent teams a ton. This is the thing. I don't think we have a super strongly opinioned viewpoint on what is the best set up at any given moment in time. So Boris always updates his tweets like this is what I'm doing now. Agent teams
SPEAKER_06
is something which a bunch of people loved internally and so we were like, okay, let's ship it. Let's see what people think about it in the field. I'm not saying we will, but we regularly unship things as well. And I do see the generator evaluator kind of pattern and there's a subset of that like teams approach to thinking about subagent design. Not necessarily contradictory to it per se. You can imagine, classically in which teams breaks down is front end, back end, some sort of integration between them like subagents. Each of those probably deserve their own kind of critic kind of agent pairing with them, for example. So you can see how the two concepts overlap.
SPEAKER_06
Just the general idea behind this is most people when they're running called code, at the moment, their goal isn't to one shot an app over six hours.
SPEAKER_06
And so that isn't that's very primitive which we by default ship there. [SPEAKER_12] One thing I was [SPEAKER_02] wondering, [SPEAKER_02] have you also tried [SPEAKER_02] a critic [SPEAKER_02] that gets the context [SPEAKER_02] of the generator? [SPEAKER_02] Then you, [SPEAKER_02] if it has some clue [SPEAKER_02] about the traces [SPEAKER_02] of the agent [SPEAKER_02] or the executor? Is that currently [SPEAKER_02] the case [SPEAKER_02] in the critic? We use a handoff pen. I would be very hesitant of that. We did try this, but this is the whole muddying of thoughts between the two model streams. I think it's actually much more effective to just let it judge the output
SPEAKER_06
and just provide, instead of being like, hey, you made a misstep when building this by doing X and that's what's resulting in this issue, it's much more effective to just have the value be like, this is an issue and then let the generator purely reflect on its own work and then try and figure out how to fix that issue. Otherwise, you just see, we found that it's very easy for the model to kill itself that something is working or not and that feed into the value agent as well. [SPEAKER_02] I think it would then [SPEAKER_02] be interesting [SPEAKER_02] if you, [SPEAKER_02] for the training team, [SPEAKER_02] if you could train [SPEAKER_02] the generator
SPEAKER_06
[SPEAKER_02] to predict [SPEAKER_02] what a critic [SPEAKER_02] currently said. [SPEAKER_02] To have it [SPEAKER_02] be more honest [SPEAKER_02] about what it did [SPEAKER_02] and stuff. [SPEAKER_02] Maybe we'll work on that. [SPEAKER_07] I want to ask more [SPEAKER_07] about traceability. [SPEAKER_07] I use [SPEAKER_07] superpowers [SPEAKER_07] or my own prompts [SPEAKER_07] to generate [SPEAKER_07] multiple sub-agents [SPEAKER_07] to implement my, [SPEAKER_07] my software or app. [SPEAKER_07] But what happens [SPEAKER_07] is I don't really know. [SPEAKER_07] I want to go back [SPEAKER_07] and see where [SPEAKER_07] it actually went wrong.
SPEAKER_06
[SPEAKER_07] But then I'm not able [SPEAKER_07] to figure out [SPEAKER_07] how to find those traces. [SPEAKER_07] What do you use [SPEAKER_07] for traceability [SPEAKER_07] is my question. [SPEAKER_07] When you have so many, [SPEAKER_07] five, six agents [SPEAKER_07] running in background. To be honest, a lot of it is just reading [SPEAKER_07] I don't really know. [SPEAKER_07] I want to go back [SPEAKER_07] and see where [SPEAKER_07] it actually went wrong. [SPEAKER_07] But then I'm not able [SPEAKER_07] to figure out [SPEAKER_07] how to find those traces. [SPEAKER_07] What do you use [SPEAKER_07] for traceability [SPEAKER_07] is my question.
SPEAKER_06
[SPEAKER_07] When you have so many, [SPEAKER_07] five, six agents [SPEAKER_07] running in background. To be honest, a lot of it is just reading through traces by hand. We do a lot of that, I would say, at Anthropic in general, just reading through traces by hand. We also have hacked together various things where we point Claude at a bunch of traces with some custom prompts to try and identify issues with the loop, this is where it veered off and whatnot. We use that as a first pass, I would say, to see where something might have gone wrong.
SPEAKER_06
But to be honest, by far and away, the best approach we use is just reading the traces by hand. Only then do you truly get to relate to what the model is trying to actually do. [SPEAKER_11] I have a few questions. [SPEAKER_11] First of all, [SPEAKER_11] how do you measure [SPEAKER_11] the quality [SPEAKER_11] of a harness agent pair? [SPEAKER_11] It feels like a vibe check, [SPEAKER_11] like it's a green field, [SPEAKER_11] let's build an app. [SPEAKER_11] But let's say [SPEAKER_11] you're going into [SPEAKER_11] a new project, [SPEAKER_11] maybe Brownfield. [SPEAKER_11] It feels like a vibe check [SPEAKER_11] or some kind of art. [SPEAKER_11] Can you make it
SPEAKER_06
[SPEAKER_11] more scientific [SPEAKER_11] or is that just [SPEAKER_11] not feasible? The way that we've thought about it, at least, is we specify the rubrics in extreme detail at the generator and evaluator level. So we talked about, for example, those four criteria. That's very high level, the rubric, which we use for design taste, let's say. And we set those up for various bits of this app. So that can be just for the design element, maybe another piece for how we think about API design, let's say, code quality, whatever. And we use those as the various rubrics which we're hill climbing against. And the evaluator's job encourages the builder to hill climb against those.
SPEAKER_06
And so for any given app or output, we have a signal of this is where the model started on those criteria, and this is where we ended up. Now, that's less useful for working on newer code bases, but it still applies. You could point the evaluator at a given code base and say, this is where we are now, and then give it the spec of what you're trying to achieve, and then let the loop iterate against those criteria. So it isn't necessarily one set of evals at the very end. It's like, here are the criteria for what we think good looks like, then letting the evaluator and the generator come up with a set of tests or contracts that needs to satisfy, and then letting it,
SPEAKER_06
as the harness, hill climb against those. That's not super comparable across different products and runs, but it's very useful within a product or run. [SPEAKER_08] This particular pattern [SPEAKER_08] is great [SPEAKER_08] for Greenfield, [SPEAKER_08] as you said, [SPEAKER_08] but it's quite opinionated.
SPEAKER_06
[SPEAKER_08] It might be using React, [SPEAKER_08] Postgres [SPEAKER_08] as a database, [SPEAKER_08] and Node on the back end, [SPEAKER_08] but your Brownfield app [SPEAKER_08] might be using [SPEAKER_08] something totally different, [SPEAKER_08] or the rubric [SPEAKER_08] that we've created [SPEAKER_08] for what we think within a product or run. Also,
SPEAKER_06
[SPEAKER_08] this particular pattern is great for Greenfield, as you said, but it's quite opinionated. It might be using React, Postgres as a database, and Node on the back end, but your Brownfield app might be using something totally different, or the rubric that we've created for what we think good design patterns are might be totally different in your project. So I think that's why we're proposing this as more of a pattern that you would then tailor towards your application. Thanks. One follow-up question. [SPEAKER_11] Does it work? [SPEAKER_08] Yeah.
SPEAKER_06
[SPEAKER_11] Do you use direct the harness individually, and how do you cooperate as a team on that? I find it's very hard when I share my screen and I'm working conversationally, it's very hard for people to keep up, and the other way around, I find it cumbersome to dictate what to prompt. How do you cooperate as a team? Do you have team-owned harnesses? Is it maybe a good feature for cloud code? Yeah, maybe.
SPEAKER_06
I think we probably do have work to do on that, right? Quite often what happens internally is people come up with these ideas, and then they're generally adopted from the bottom-up by different teams, and it's then the job of the original idea holder, shall we say, which was Prithby in this case, to maintain it and make it composable and generalizable for different teams, and different teams will adapt it and make it useful for their section of the code base. But we don't have anything good in that sense. Even just observability, as some of the people talked about, is generally a thing which is not fully solved yet for these ultra-long-running agents. Yeah, an interesting area of greenfield software to explore. Yeah,
SPEAKER_06
[SPEAKER_08] that is an interesting one, whether it should be a collaborative experience in Cloud Code or even Cloud.ai. I think just leveraging software engineering best practices with version control and making your commits and pull requests, or if you're working on your own using something like Git work trees so that you're not overriding the file system on multiple different features, all make sense. But yeah, I think when it comes to collaboration, maybe it's something that doesn't happen quite as much because people just build these projects as Ash said from the ground up and then present them to the rest of the company.
SPEAKER_06
[SPEAKER_12] Hi. Jose from Mercedes-Benz Research and Development here. Hi. Thanks for the talk. While looking at it, I thought, okay, it looks a lot like a scrum team, a feature team working for longer times on a product. And I was thinking, how does human in the loop look like in that scenario? Because you have this kind of sprint. Have you thought about a sprint review kind of moment where you, as a human, get asked, hey, here's what we built the last two hours. Yeah. How's it looking like for you?
SPEAKER_06
Yeah. Should we subject our agents to the same trauma that engineers go through of scrum review? I mean, the whole point of this, the general idea behind this talk and also what we're trying to do is trying to be as AGI-peeled as possible, right? How do we build harnesses where we don't need a human in the loop, right? What does that look like? Are we using this today for everything? Obviously not, right? But the goal is, this is a technique or a pattern which should extend very nicely
SPEAKER_06
is trying to be as AGI-peeled as possible, right? How do we build harnesses where we don't need a human in the loop, right? What does that look like? Are we using this today for everything? Obviously not, right? But the goal is this is a technique or a pattern which should extend very nicely such that you don't have a human in the loop for most things. If you did, right, hooks is probably the main primitive to inject given some sort of specific type of stop condition, let's say, with an evaluator to hand back to human, allow some kind of developer message input and then continue the loop would be a simple way to implement it. But, yeah, to be honest, we're exploring this from a what can we do fully autonomously approach as opposed to thinking this as here's called code and how do we make this more powerful. It's very much a more greenfield exploration of agent design.
SPEAKER_06
[SPEAKER_12] No, of course, it's just if I would get the chance to review it maybe a few hours in, then I might be able to steer it in a better way so that eight hours later it's more like the one kind of project I would like to have. Yeah.
SPEAKER_06
I mean, I get what you're saying. I think the question then is should that be a permanent feature of the harness or is that just a thing which you should have, basically prompted around when building the harness, right? So we would have that, right? We would run this harness on loop and we would have we might spin up ten generations of different things and three of them succeed and seven of them fail in random ways and then we would just sit down with those seven, read through them, adjust the prompting of the main harness and then try again until we get to a point where we're quite happy leaving it to run fully autonomously. So, ultimately, that's still the end goal for us as opposed to basically giving up on the harness and being okay, we just insert a human here to cover for any kind of steerability issues instead, but rather embed that and bake that into the harness itself in the first place.
SPEAKER_06
[SPEAKER_03] Have you used this to build anything, like non-greenfield or, I guess, production, like anything in code code itself or have you used it for actual features and seeing it to the end?
SPEAKER_06
[SPEAKER_08] I mean, this does mostly extend to greenfield projects. I think for brownfield, maybe you do need a little bit more control as you're starting to build out your own rubrics and patterns. I mean, what we're seeing in brownfield is that if you look at the whole software development life cycle, it's not just the coding aspects that people are starting to use something like cloud code for. It might be, say, there's autonomous monitoring happening and then that could feed into generating some kind of issue or feature request that could then just feed into an agent that would then go through to make the pull request and then there's sort of a pull review already happening and then maybe you're just reviewing that before you actually merge. So I think there are other ways to automate the whole software development lifecycle in a brownfield project, but I think that this particular pattern, maybe without a lot of testing within
SPEAKER_06
[SPEAKER_08] that would then go through to make the pull request and then there's a pull review already happening and then maybe you're just reviewing that before you actually merge. So I think there are other ways to automate the whole software development lifecycle in a brownfield project, but I think that this particular pattern, maybe without a lot of testing within your project and building, customizing it for your project, it's probably more suited towards brand new applications. [SPEAKER_03] Have you built any greenfield apps that, I don't know, an internal tooling or anything like that that you've been using, not just a demo? [SPEAKER_03] Yeah.
SPEAKER_06
To be frank, I can't really talk to internal tooling too much, but a good anecdote of this was, a lot of the new and fun stuff that you see in Code Code will, when I'm speaking to the team and working with them on stuff, use a lot of the lessons from this, per se, in, even just general hands-on Code Code usage, the way that they prompt the main model to spin up a subagent, let's say, and go after something, or as Andrew said, right, in monitoring and bug-fixing loops, when generating a fix, should you have a separate evaluation generator and a generator go after the same thing? So a lot of these principles apply. Is it one-for-one this? Maybe not, but it's taking the good bits of this or whatever you think is applicable to a certain space and field and then running with it in your own way. Thank you.
SPEAKER_06
[SPEAKER_03] You say reading the traces. [SPEAKER_09] Is that literally just the raw output, or is there something more specific you've prompted it to, write this to file, these are the sorts of things I care about and I want to see?
SPEAKER_06
No, you've got to read the whole thing. Read the whole thing. I do think it's a really important skill when building agents in general is to empathize as much with the model. This was, there's an interesting anecdote which we used when we were building, for example, the agent harness for Codefor Chrome, which is our browser use thing. And we would run this experiment where, imagine if you were trying to navigate a web page and click around where you're effectively doing with your eyes closed and every 10 seconds you just opened it to see a static page and then closed it again and then had to do things. And really putting yourself in the shoes of the model is, this kind of empathetic skill set which you need to develop and the only way to really do that is to spend as much time with these models but also, yeah, reading through line by line being, oh, why did it think this? Oh, I can see why I did that and then adjusting the way you instruct it next time to do better. But that's why I think Claude of Kremie is very good with just spending a lot of time as a team closing our eyes and trying to navigate web pages, for example. So, yeah.
SPEAKER_06
[SPEAKER_08] Yeah, and I think then actually taking those learnings and putting them into, say, your prompt templates or your Claude.md or building a skill or just generally understanding how to avoid that type of behavior in the future. I know Claude Code now has auto-memory for sessions as well so it's constantly memorizing little things as it goes. But, yeah. You can learn quite quickly from reading some traces like where things might be going wrong.
SPEAKER_06
[SPEAKER_08] your prompt templates or your Claude.md or building a skill or just generally understanding how to avoid that type of behavior in the future. I know Claude Code now has auto-memory for sessions as well so it's constantly memorizing little things as it goes. But, yeah. You can learn quite quickly from reading some traces like where things might be going wrong. Should we wrap up there? I think we have a few minutes left but we'll be around in general in case you guys want to ask any questions or just chat. But, otherwise, thanks for coming down for our session today. Thank you. Thank you. to handle very long sessions. Sprint decomposition.
SPEAKER_06
We don't have a very strong opinion on this, but it was something which was really, really critical to getting Opus 4.5 to work, but Opus 4.6 was able to kind of hold a two-hour continuous build coherently in a way without necessarily having to be force-fed one feature at a time. The cadence at which the evaluator should run. Previously, we were running at every single sprint per se, whereas now, we were just running at the end of a one-shot generation from the model and then passing back. So the harness is still the same. We're just kind of simplifying the specific kind of loops and the kind of recipe that kind of goes into it. The lesson isn't necessarily
SPEAKER_06
our harness was wrong, but rather, it was right for 4.5, the frontier moved, and we ran a simplified version to see how it'd work. So this is kind of what the final kind of setup kind of looks like today. Having that planner generator, evaluator loop is still the kind of core of our system, but you can see we kind of ditched a bunch of the other kind of components that made this slightly more complicated than it had to be. We also, as kind of mentioned, big fan of just using a file system for shared state instead of kind of leaning on context windows for very long-running agents in general. And this is an example of the simplified harness running
SPEAKER_06
with one of our latest models. Again, very, very expensive, but you can see it's actually roughly like half the cost of the previous runs just because we're kind of doing things in a slightly more simplified manner, but it's still running over a very extended period of time. And so this is an example of a DAW, which is basically just like a music creating app, if you will. The agent sets the tempo, a key, it lays down the melody, it builds the drum tracks. This is the evaluator actually going and testing the app itself. We did actually listen to the music in this. Obviously, Claude can't hear at the moment, and so the music was pretty trash, but the app was really good
SPEAKER_06
in general and pretty fleshed out, which, you know, a model ago, this is something which would never have worked. But this is something which was possible with just a couple rounds. And this is kind of that meter curve which Andrew was talking about kind of really in action. And so I kind of wanted to close just by saying you don't necessarily need our internal harness to go away and start thinking about this. We are constantly trying to ship bits of these primitives into code code directly, but also there's nothing stopping you from just going ahead and building something similar to this kind of on your own. So we just shipped, you know, auto mode is probably
SPEAKER_06
my favorite thing for slightly more, you know, safe yellow, if you will, instead of running dangerous C skip permissions all the time. We already have custom subagents as a primitive, right? Your evaluator, your QA role, give it a harsh system prompt and a very detailed rubric. Playwright MCP or Claude for Chrome MCP, already extremely, extremely good at web app stuff or just use computer use if you're building kind of native apps. And skills, again, a very nice way to package your kind of grading rubrics into your kind of general development flow.
SPEAKER_06
So, yeah, five things if you're kind of taking a photo, this is the slide I would say to kind of remember. Self-evaluation, very much a trap, just use an adversarial evaluator. Compaction, doesn't necessarily, does not equal kind of coherence, right? Lossy summaries really drift. Structured handoffs and clean context are a very good pattern that we've seen. Don't think that subjective quality isn't gradable. If you have a strong view on what something should look like, then kind of force yourself to write it down. We found this made kind of a really massive difference to the quality of kind of apps that a model was able to generate. And then kind of
SPEAKER_06
finally was really just, you know, sit with them all, read the traces. Only then can you kind of really know what bits of a scaffold to delete, what bits to keep, especially as the kind of frontier moves. But, yeah, that's it from me. Thank you very much for listening.
SPEAKER_06
And, yeah, check out our blog post. But, I wanted to just open up for Q&A in general because we've been yapping for like close to an hour now. So, if you have any questions for me and Andrew, just fire away. We'll do our best to try and answer them. Yeah.
SPEAKER_06
Thank you.
SPEAKER_03
Johan from Paulside. One question for you. When you improve the evaluator by like reading the logs and improving it, is that sort of like on a per-project basis or more of a secret source that you reuse across project? The goal
SPEAKER_06
was very much to try and do this in a way which was reusable, right? Like, I think anyone can tune this in a way that's creating, you know, a very specific type of app. That's fine. At that point, it's not that different from, you know, going ahead and just prompting tool code yourself and doing it, right? I think they were just, the key was like, what are the common patterns that you can kind of draw across the model weak points, right? So, talking to that kind of front-end design piece, we knew like what we thought good design would be, right? You could give examples like this is what, you know, a read-view-for-pro looks like. This is what AI slop looks like, right?
SPEAKER_06
And that generalizes quite well. So, yeah, this was all around web apps, but it could quite easily apply to other kind of things as well.
SPEAKER_06
Thanks for,
SPEAKER_09
Thank you for
SPEAKER_04
presentation, very interesting. I was just wondering, what is your view on concept of dump zone and smart zone of a model? So, I understand like before it was around 40%, now with one million context, it's about 100k is what I understand. And the way I understood Ralph Loop was designed is to kind of negotiate this problem. So, basically, we're keeping the model always in the smart zone. basically, trying to slice the task below 100, so it executes the task within the 100 context zone. And what I understand from your presentation, you can like advocating not to use it anymore because we can now rely on a compaction and so on. Is it like something you're suggesting to do
SPEAKER_04
or with still with like a Ralph Loop model still has its own place given the smart and dump zone concept? Yeah. well, I suppose from
SPEAKER_08
Ash's presentation and mine, you see that the one million context window is now GA and so you have sort of a bigger context to use. The models are more agentic, so they can sort of maintain coherence for a longer period of time within that context window. And that actually, with the release of 4.6, we decided to move from new context windows to just a single, long-running, continuous session with compaction. So I think, I mean, whether or not you use multiple fresh sessions or just one long-running one is probably still up to your use case and your evals depending on what you're seeing is working best, but at least for sort of this general purpose generator
SPEAKER_08
evaluator pattern with Opus 4.6, we saw that it was possible to use a single session. I don't know if you want to add to that. I think it's also
SPEAKER_06
just like a temporary problem, right? Like, context rots is a failing of today's models to some extent and much less so than even just one model generation ago. So is there a place for the type of thing which you're discussing? I think yes, depending on your use case, but it's not like a, it's one of those pieces which I'd look at as like, okay, I'd be kind of hunting for the model release where I can kind of strip it out. Let's put it that way.
SPEAKER_06
I always have a lot
SPEAKER_01
of FOMO around playwright. I mean, you said playwright MCP, there's playwright skills. Can you speak to how to improve the playwright? Because like, I imagine I would like to like have my browser open and then I can see the model working through it and then maybe I could steer it, you know,
SPEAKER_00
the few tabs open.
SPEAKER_01
But like, yeah, what, is there some innovation there I'm missing out on or is, is playwright MCP really, would you recommend people use? I mean, playwright MCP
SPEAKER_06
or just use the called for Chrome MCP which is like a slightly more robust thing, I guess, around browser control. I mean, I don't know why you want to watch it do things. I mean, you can, but I think that's like a trust gap, right, today. Like, you know, the whole point of what we're trying to get to do here is like, you set something off, you trust it to do it, do the work and test it and you have the confidence that it's doing it correctly and you come back to it. And that's where, you know, yes, there's going to be some iteration at the beginning where you're like watching, reading the traces until you get to a point you can trust it. But, at least internally, right,
SPEAKER_06
like when I'm doing full stack app dev, I have got to a point now where I'm like, okay, with Opus 4.6 I can like reliably trust the model to go ahead, read, network errors, console errors, actually navigate an app, zoom in where it needs to. The vision is now good enough on these models that it can like identify overlapping text on elements and things like that, whereas that just wasn't the case until realistically the last, you know, generation of models. So, yeah, I would recommend.
SPEAKER_06
I'm curious,
SPEAKER_00
like, with the generator evaluator pattern, what happens? Can you throw unlimited tokens at it or will it stop because the evaluator is not good enough? Like, can you tell me more about that? Sorry,
SPEAKER_06
do you mind clarifying? I kind of messed up. Yeah,
SPEAKER_00
so, okay, let's say I say, okay, create, like, a very cool game with some features. You have to generate an evaluator pattern that creates, like, contracts, builds the apps. If I, then it will give me back something, right? Can I restart it again and say, like, okay, make it better, I'm not happy about it and generate the evaluator will pick, the pattern will pick it up and make it better? Yeah. Or will the evaluator not go to at one point and just say, like, this is it? Um,
SPEAKER_06
I think that's, I mean, one, first of all, like, if you want, like, some level of human in the loop in this process, that's, you know, just implement hooks at some point in this loop. I think the bit which was kind of surprising to us was with this, this general pattern and especially with the kind of 4.6 liner models, both Sonnet and Opus, it was extremely willing to, like, throw away everything, you know, even if it had done kind of 10 passes at something, it was kind of very happy to just, like, throw it all away and start from scratch if, for some reason, it wasn't able to, like, hill climb against the rubric of the evaluator in a kind of effective way. And so,
SPEAKER_06
that's why, kind of, when we're kind of, when we're playing with this kind of thing, we didn't naturally, like, lean towards having some kind of resume or human-the-loop type intervention system, I guess. And we didn't really observe, we expected to, but we didn't really observe that kind of behavior which you were talking about, where it kind of just, like, evaluators, like, I just give up, let's just, like, pass it on, shall we say? Yeah, it was just much more willing to, like, throw away everything and restart. And that was just a behavior which we never saw when it was the generator itself almost kind of being proud of its own work and being, like, I'm not going to
SPEAKER_06
restart this whole thing. So, yeah. I mean, there's been an example which I've seen where the evaluator is, like, it kind of gets fed up and is, like, right, this approach you're taking just obviously isn't working. Can you just, like, delete everything and restart? Which, I don't know about you guys, but vibe coding regularly I often do as a human to, like, you know, just benefit from fresh context windows, not have to deal with an already messy code base, et cetera. So it's quite neat seeing models now also kind of get to that point. And also,
SPEAKER_08
just briefly add, obviously, you can then open that code base in Cloud Code and continue where you left off. Sort of goes without saying. And I think we're generally thinking about what the workflow looks like if it's sort of more back and forth because there's sort of the extreme of build me a really complex gaming application that you don't know is it going to take three hours, is it going to take 20 hours? It's a bit unclear, so maybe there's sort of something in the middle that's, like, more of a, yeah, feedback loop.
SPEAKER_08
Hi.
SPEAKER_11
I really like the idea
SPEAKER_10
that you have, like, you know, there's a human element here where it's, like, you know, PM, engineer, evaluator. PM role is a lot of the time it's, like, scope creep and keeping the time going and stuff like that. But you're just, like, letting this off. You're letting engineers go play in the sandbox for ages. Is there a harness loop that needs to go back to the planner eventually? Does it need to move again? Well, maybe
SPEAKER_06
because we're engineers we just decided, like, ah, screw the PM, we'll just stuff it to the side. You guys need a PM. We actually, well, this is where the kind of, like, that kind of contracting piece between the, the kind of evaluator and the builder worked quite well. For context, in that, we typically, like, insert the main spec that was generated by, like, the PM per se into those sessions regularly. So that, you know, it's always a reference point for, like, okay, this is what we're still actually trying to build. And the main function of then the builder and the generator, sorry, the builder and the evaluator is just to, like, figure out the exact
SPEAKER_06
feature set and tests and contracts per se that actually satisfy that spec. But the reason we don't is because we don't want the planner to be, like, a core part of this loop. It should be very high level. It should, its purpose really is just kind of set out, like, kind of the hard outer lines of what this product could be. But its job is not necessary to come in and intervene and be, like, actually, this is, like, an impossible feature. We should not do this and edit itself. We kind of wanted to keep that context relationship between just the builder and the generator. That being said, like, this loop, I've applied it in lots of different ways. It doesn't have to just
SPEAKER_06
be, you know, one generator and one builder, right? Like, that adversarial kind of trade-off can be applied to, like, a workflow consisting of multiple separate agents, right? I don't know. It could, if you're trying to do, I don't know, generate evals, let's say, you could use a similar harness to be, like, hey, generate a, it could be, like, planner, generate a synthetic, a generator for synthetic data sets, right? With a QA agent. Then hand off to, like, an integrator, which, like, actually wires up something. Also has a QA agent. Then has, like, a final kind of a, you can basically add this kind of generator, evaluator thing into a multi-step workflow. We each, like,
SPEAKER_06
build a, maybe has, like, a slightly different function, per se, as part of a longer workflow. So there are different ways in which you kind of keep things on track depending on the task and break down this general pattern into slightly more specified, you know, tasks or workflows, if that makes sense.
SPEAKER_06
You mentioned that
SPEAKER_05
some of the later tasks could not possibly be done by an earlier model. Can you talk a little bit about your process comparing the tasks on the different models? Like, do you fire off the same task on Opus 4.6, Opus 4.5, Sonnet, or is this sort of artisanal, co-evolving harness model set up operating that? Yeah, I mean,
SPEAKER_08
I suppose we walk through the history a little bit, and if you look at, say, the first blog post on long-running agents versus the more recent one, there are some pretty significant differences there. You know, one being what we were just discussing, that the initializer agent would build this super comprehensive spec of, say, 200 different features, and then the loop would have to actually go and execute against every single one of those features, which may lead to, say, incorrect design decisions, but it's sort of forced into that behavior, whereas I think now you're able to have sort of a more generic creative direction set with, say, Opus 4.6, and then just having
SPEAKER_08
this loop of the generator evaluator. But it does, yeah, your model selection does inform your harness design very much so. Of course, in a perfect world, you could just sort of throw everything and say Opus 4.6, but if you have cost concerns, for example, maybe you do use Opus 4.6 for planning and then Sonnet 4.6 for the coding or the execution, that's something that we tend to see quite frequently. But again, if you're building specific sub-agents for each of these, you probably want to have some evaluations to be able to understand for that model and that prompt how it's performing against that task and then just optimize.
SPEAKER_08
Do you have any advice
SPEAKER_09
on moving beyond sort of these one-shot applications to long-lived products where you're looking to make changes days, weeks later? What sort of artifacts you need to persist to future instances to be able to know what has come before, what can I change, what should I change? Yeah,
SPEAKER_06
it's something which we're working on. Like right now, like we use like similar patterns for just a bunch of random stuff internally, shall we say. And so, at the moment, it's like set this thing off. It's running, you know, on a remote server somewhere and I'll just come back and check it like after this talk, let's say. And then I kind of iterate on it kind of manually in whole code directly like polish any rough edges, that kind of thing. I think in terms of the way that you're actually like setting up this harness, just having, this is why we kind of default to kind of using a file system estate for this kind of loop. One, because it's just very easy for another model
SPEAKER_06
to come up and grip through and pick up what's been going. But one thing which I like to do is kind of embed little bits of prompting throughout this kind of loop which basically tells it to write kind of learnings and states to some kind of JSON file because the model doesn't kind of overwrite that too much. And so, the nice thing about that is you're basically just leaving like breadcrumbs for another model to come and pick up. So, honestly, the key thing for me is like how do I instruct this harness to leave crumbs for a human to come in and then use code code on top of. So, generally, it's like, hey, the shape of that file might be like tried this,
SPEAKER_06
evaluator found this bug, implemented this fix, this fix worked, yes, tick. And then continue. And you have kind of like a time stamped kind of time log, if you will, of like everything the model has tried, the fix it's made and the final state. And then also having some sort of live updating kind of set of docs, if you will, just very high level. Here's the file structure. And then those two files, to be honest, are more than enough for code and a human to come in and start iterating on the app with. But that's all we're doing at the moment.
SPEAKER_06
So,
SPEAKER_02
it's very interesting to hear the, oh yeah, first of all, congrats on the presentation. Thank you. And then I was wondering, there are kind of two approaches. Like, you have the agent team where multiple agents interact with each other. And then the explicit generator critic setup.
SPEAKER_02
because in sense, like the agent team has the same setup where the main agent instructs someone and then can act as a critic for the subagent. But what are the current failure modes that causes us to still need the specific generator critic harness instead of just the agent team itself? And then what's your estimate of how many model generations we would need to just completely rely on the agent team?
SPEAKER_02
Maybe,
SPEAKER_08
I mean, I can sort of address the first aspect of that. So, I mean, one of the limitations of, firstly, cloud code is using the same harness that is the agent SDK. So, you can, technically, you should be able to build this type of a pattern into cloud code. Agent Teams is a useful framework for potentially doing that because you could, say, have the generator and the evaluator sort of intercommunicating or maybe it's the generator is sort of the main agent and the evaluator is, say, a member of the agent team. But I think it's sort of evolved more so from that first blog post that I shared. I think that was like the results of that to some extent to try and make that
SPEAKER_08
more generally available. But one of the things you're limited by, obviously, is like cloud code would just have to run on your machine. I think with the agent SDK, you can also just run it in more of a cloud environment and a sandbox environment for long periods of time and without it failing or you having to run the caffeinate on your machine. But I think, yeah, cloud code is a good testing ground for building out any of these types of harnesses to experiment and explore and see what works before maybe you build it into the agent SDK and then actually deploy it as its own application. And yeah, I mean, again, I would just experiment and see if agent teams
SPEAKER_08
is something that makes sense for you or if maybe just using regular subagents or some other framing of it that works better. But yeah, I think people are using agent teams like a ton. I don't know if you... Yeah, well,
SPEAKER_06
this is the thing. I don't think we have like a super strongly opinioned viewpoint on like what is the best at any, you know, set up at any given moment in time. So Boris always like updates his tweets like this is what I'm doing now. Like agent teams is something which a bunch of people loved internally and so we were like, okay, let's ship it. Let's see what people think about it in the field. I'm not saying we will, but you know, we regularly on ship things as well. And I do see the generator evaluator kind of pattern and there's like a subset of that like teams approach to thinking about subagent design. Not necessarily like contradictory to it per se. You know,
SPEAKER_06
you can imagine like, you know, classical in which teams breaks down is like, you know, front end, back end, some sort of integrated between them like subagents. Each of those probably deserve their own kind of critic kind of agent pairing with them, for example. So you can kind of see how the two concepts like overlap. Just the general idea behind this is, you know, most people when they're running called code, at the moment, their goal isn't to like one shot an app over like six hours. And so that isn't that's very primitive which we like by default like ship there. So yeah.
SPEAKER_06
Yeah, sorry.
SPEAKER_12
One thing I was
SPEAKER_02
wondering, have you also tried like a critic that gets the context of the generator? Then you, I feel like if it has some clue about the traces of the agent or like the executor? Yeah.
SPEAKER_06
Is that currently
SPEAKER_02
the case and the critic? We use like a handoff
SPEAKER_06
pen. I would be very hesitant of that. We did try this, but this is the whole like muddying of like thoughts between the two model streams. I think it's actually much more effective to just let it judge the output and just provide, instead of being like, hey, you made a misstep when building this by doing X and that's what's resulting in this issue, it's much more effective to just have the value to be like, this is an issue and then let the generator purely reflect on its own work and then try and figure out how to fix that issue. Otherwise, you kind of just see, we found that it's very easy for the model to like kill itself that something is working or not and that feed
SPEAKER_06
into the value agent as well. Last note on that.
SPEAKER_02
I think it would then be interesting if you, for the training team, if you could train like the generator to predict what a critic currently said. Yeah. Like to have it be more honest about what it did and stuff. Maybe we'll work on that.
SPEAKER_02
I want to ask more
SPEAKER_07
about traceability. Like I use superpowers or like my own prompts to generate like multiple sub-agents to implement my, let's say, my software or app. But what happens is like, I don't really know. I want to go back and see where it actually went wrong. But then I'm not able to figure out how to find those traces. What do you use for traceability is my question. When you have so many, like five, six agents running in background, like, yeah. To be honest,
SPEAKER_06
a lot of it is just reading through traces by hand. We do a lot of that, I would say, at Anthropic in general, just like reading through traces by hand. We also just like have, you know, hacked together various things where we, you know, point Claude at a bunch of traces with some custom prompts to try and identify like issues with the loop, like this is where it veered off and whatnot. We kind of use that as like a first pass, I would say, maybe to just kind of like see where something where, like where something might have gone wrong. But to be honest, by far and away, the best approach release that we use in time is just reading, reading the traces by hand. Only then
SPEAKER_06
do you kind of like truly get to kind of relate to what the model is trying to actually do. Yeah.
SPEAKER_06
Thanks for the talk.
SPEAKER_11
I have a few questions. First of all, how do you measure the quality of a harness agent pair? Is it, it feels like a vibe check, like it's a green field, let's build an app. But let's say you're going into a new project, maybe Brownfield. It feels like a vibe check or some kind of art. Can you make it more scientific or is that just not feasible? I mean,
SPEAKER_06
the way that we've thought about it, at least, right, is like we specify the rubrics in kind of extreme detail at the kind of generator and evaluator level, right? So we talked about, for example, those four kind of criteria. That's very high level, the rubric, which we use for like design taste, let's say. And so we set those up for various bits of this app, right? So that can be just for the design element, maybe another piece for like how we think about kind of API design, let's say, code quality, whatever. And we kind of use those as the kind of various rubrics which we're hill climbing against, right? And then the evaluator's jobs that, you know,
SPEAKER_06
encourage the builder to hill climb against those. And so for any given app or output, we have like a signal of this is where the model started on those kind of criteria, and this is where we kind of ended up. Now, that's less useful for like kind of, like you said, working on kind of newer code bases, but it still applies, right? Like you could point, you just have to start the loop in a different way. You would just point the evaluator at a given code base and be like, this is where we are now, and then give it the spec of what you're trying to achieve, and then let the loop kind of iterate against those kind of criteria. So it isn't like necessarily a one set of evals
SPEAKER_06
at the very end. It's kind of like, here are the criteria for what we think good looks like, then letting the evaluator and the generator come up with a set of kind of tests or contracts that needs to satisfy, and then letting it, just as the harness, he'll climb against those. That's not super comparable across different products and runs, but it's very useful for, yeah, within a product or run. Also,
this particular pattern is it's great for Greenfield, like you said, but it's quite opinionated. You know, it might be using React, like Postgres as a database, and Node on the back end, but your Brownfield app might be using something totally different, or the rubric that we've created for what we think, you know, good sort of design patterns are, might be totally different in your project. So I think that's why we're proposing this as more of a pattern that you would then tailor towards, you know, your application. Thanks. One follow-up question. Does it work?
SPEAKER_08
Yeah.
SPEAKER_11
Do you use, do you like direct the harness individually, and how do you cooperate as a team on that? I find it's very hard to like, when I share my screen and I'm working conversationally, it's very hard for people to keep up, and the other way around, I find it cumbersome to dictate what to prompt. How do you cooperate as a team? Do you have like team-owned harnesses?
SPEAKER_11
Is it maybe a good feature for cloud code?
SPEAKER_11
Yeah, maybe.
SPEAKER_06
I think we probably do have work to do on that, right? Like, I think like, quite often what happens internally is like, you know, people come up with these ideas, and then they're generally quite bottoms-up adopted by different teams, and it's then the job of the kind of original, you know, idea holder, shall we say, which was Prithby in this case, to kind of maintain it and make it kind of composable and generalizable for different teams, and different teams will adapt it and, you know, make it useful for like, their section of the code base, let's say. But we don't have any like, good things in that sense. I think like, you know, even just observability,
SPEAKER_06
like some of the people talked about, right, is like a, generally speaking, a thing which is not fully solved yet for these like, ultra-long-running agents. Yeah, interesting area of kind of greenfield software to explore. Yeah,
SPEAKER_08
that is an interesting one, whether it should be sort of a collaborative experience in Cloud Code or even Cloud.ai. I think just leveraging software engineering best practices with version control and making your commits and pull requests or if you're working on your own using something like Git work trees so that you're not overriding the file system on multiple different features all make sense. But, yeah, I think when it comes to collaboration, maybe it's something that doesn't happen quite as much because people just build these projects as Ash said from the ground up and then sort of, you know, present them to the rest of the company.
SPEAKER_08
Hi.
SPEAKER_12
Jose from Mercedes-Benz Research and Development here. Hi. Thanks for the talk. While looking at it, I thought, okay, it looks a lot like a scrum team, a feature team working for longer times on a product. And I was thinking, how does human in the loop look like in that scenario? Because you have this kind of sprint. Have you thought about a sprint review kind of moment where you, as a human, get asked, hey, here's what we built the last two hours. Yeah. How's it looking like for you?
SPEAKER_06
Yeah. Should we subject our agents to the same trauma that, like, engineers go through of scrum review? I mean, like, the whole point of this, the general idea behind this talk and also what we're trying to do is trying to be as AGI-peeled as possible, right? how do we build harnesses where we don't need a human in the loop, right? Like, what does that look like? Are we using this today for everything? Obviously not, right? But the goal is, you know, this is a technique or a pattern which should extend very nicely such that you don't have a human in the loop for most things. If you did, right, it's like, you know, hooks is probably the main primitive
SPEAKER_06
to just basically inject given some sort of specific type of stop condition, let's say, with an evaluator to basically, like, hand back to human, allow some kind of developer message input and then continue the loop would be, like, kind of a simple way to implement it. But, yeah, to be honest, we're kind of exploring this from a what can we do fully autonomously kind of approach as opposed to thinking this as, like, here's called code and, like, how do we make this, like, you know, more powerful per se. It's very much like a kind of more greenfield exploration of agent design. No, of course,
SPEAKER_12
it's just, like, if I would get the chance to review it maybe a few hours in, then I might be able to steer it in a way better way so that eight hours later it's more like the one kind of project I would like to have. Yeah.
SPEAKER_06
I mean, I get what you're saying. I think the question then is, like, should that be, like, a permanent feature of the harness or is that just, like, a thing which you should have, like, kind of basically prompted around when building the harness, right? So, we would have that, right? We would run this harness on loop and we would have, you know, we might spin up, like, ten generations of different things and, like, three of them succeed and seven of them fail in, like, random ways and then we would just sit down with those seven, read through them, adjust the prompting of the main harness and then try again and then until we get to a point where we're, like, quite happy
SPEAKER_06
leaving it to run fully autonomously. So, ultimately, that's still the end goal for us as opposed to being, like, basically giving up on the harness and being, like, okay, we just insert a human here to, like, cover for any kind of sterability issues instead, but rather embed that and bake that into the harness itself in the first place.
SPEAKER_06
Have you used this
SPEAKER_03
to build anything, like, sort of, like, non-greenfield or, I guess, like, production, like, anything in code code itself or have you used it for actual features and seeing it to the end?
SPEAKER_03
I think,
SPEAKER_08
I mean, this does mostly extend to greenfield. projects. I think for brownfield, maybe you do need a little bit more control as you're starting to build out your own rubrics and patterns. I mean, what we're seeing in brownfield is that if you look at the whole software development life cycle, it's not just the coding aspects that people are starting to use something like cloud code for. It might be, say, there's, like, autonomous monitoring happening and then that could feed into, say, generating some kind of, like, issue or feature request that could then just feed into an agent that would then go through to make the pull request and then there's sort of a pull review
SPEAKER_08
already happening and then maybe you're just reviewing that before you actually merge. So I think there are other ways to automate the whole software development lifecycle in a brownfield project, but I think that this particular pattern, maybe without a lot of testing within your project and building, like, customizing it for your project, it's probably more suited towards brand new applications. Have you built
SPEAKER_03
any greenfield apps that, like, I don't know, like, an internal tooling or anything like that that you've been using, like, not just a demo sort of, like... Yeah.
SPEAKER_06
To be frank, I can't really, like, talk to, like, internal tooling too much, but a good anecdote of this was, like, a lot of the new and fun stuff that you see in Code Code will, like, when I'm speaking to the team and working with them on stuff, use a lot of the lessons from this, per se, like, in, like, even just general hands-on Code Code usage, the way that they prompt, you know, the main model to spin up a subagent, let's say, and go after something, or as kind of Andrew said, right, in kind of monitoring and bug-fixing loops, like, you know, when generating a fix, like, should you have a separate evaluation generator and a generator go after the same thing?
SPEAKER_06
So a lot of these principles apply. Is it, like, you know, one-for-one this? Maybe not, but it's, like, taking the good bits of this or whatever you think is kind of applicable to a certain space and field and then kind of running with it in your own way. Thank you.
SPEAKER_03
You say reading the traces.
SPEAKER_09
Is that literally just, like, the raw output, or is there something more specific you've prompted it to, like, write this to file, these are the sorts of things I care about and I want to see? No, you've got to read
SPEAKER_06
the whole thing. Read the whole thing. I do think it's, like, a really important skill when building agents in general is to, like, empathize as much with the model. This was, like, there's an interesting anecdote which we used when we were building, for example, the agent harness for called for Chrome, which is our, kind of, browser use thing.
SPEAKER_06
And we would run this, like, experiment where, like, imagine if, you know, you were trying to navigate a web page and click around where, like, you know, you're effectively doing with your eyes closed and, like, every 10 seconds you just opened it to see, like, a static page and then closed it again and then had to, like, do things. And, like, really putting yourself in the shoes of the model is, kind of, like, this kind of empathetic skill set which you need to develop and the only way to really do that is to, like, spend as much time with these models but also, yeah, reading through line by line being, like, oh, why did it think this? Oh, I can, kind of,
SPEAKER_06
see why I did that and then, kind of, adjusting the way you instruct it next time to do better. But that's why I think Claude of Kremie is very good with just really just, like, spending a lot of time as a team closing our eyes and trying to navigate web pages, for example. So, yeah.
SPEAKER_06
Yeah, and I think
SPEAKER_08
then actually taking those learnings and putting them into, say, your prompt templates or your Claude.md or building a skill or just generally understanding how to, sort of, avoid that type of behavior in the future. I know Claude Code now has auto-memory for sessions as well so it's, sort of, constantly memorizing little things as it goes. But, yeah. You can learn quite quickly from reading some traces like where things might be going wrong. Cool.
SPEAKER_06
Should we wrap up there? I think we have a few minutes left but we'll be around in general in case you guys want to ask any questions or just chat. But, otherwise, thanks for coming down for our session today. Thank you. Thank you.