Open Reader

Evolution of agentic surfaces — Gagan Bhat & Isabella Kai He, Anthropic

completed 31:23 Aug 11, 2026 Watch on YouTube

Current Status

completed

Video ID

K0X9QDRkIdg

RAG / Chat

Enabled
Evolution of agentic surfaces — Gagan Bhat & Isabella Kai He, Anthropic
Description

Sonnet 4.5 developed what Anthropic's Applied AI team came to call context anxiety: approaching its context window limit, it would wrap work up early and stop with room to spare. They built context resets into the harness to compensate. Then Opus 4.5 shipped without the behavior, and the fix turned into pure overhead, adding latency and discarding cache it should have kept. That is the principle Gagan Bhat and Isabella Kai He build the whole session on: a harness encodes assumptions about what the model cannot do on its own, and those assumptions go stale as models improve. The architectural consequence is decoupling the brain, meaning the agent loop, from the hands, meaning the tool execution environment. Both started in one container, so the model could not begin reasoning until setup finished and either half failing took the whole agent down. Splitting them lets reasoning start while the container builds in parallel, which they measured at 60% faster time to first token at P50 and over 90% at P95. It also changes failure into something recoverable: a dead sandbox is simply retried, and a dead brain resumes from a durable session log. That log ends up doing triple duty, providing observability, letting the harness read context slices back in after Claude discards them mid run, and feeding a periodic batch process they call dreaming that rewrites the agent's memory so the next day's sessions start smarter. Speaker info: Gagan Bhat (Anthropic): - https://www.linkedin.com/in/gagan-bhat/ Isabella Kai He (Anthropic): - https://x.com/IsabellaKHe - https://www.linkedin.com/in/isabella-kai-he/ Timestamps: 0:00 - Who the Applied AI team is 1:52 - From simple questions to owning outcomes 2:45 - The Messages API and the hand rolled agentic loop 4:29 - Six production infrastructure problems 5:20 - The Claude Agent SDK 6:13 - What managed agents takes off your plate 7:02 - Harnesses encode assumptions that go stale 7:51 - Context anxiety, and the fix that outlived its need

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: As agents shift from simple request-response tasks to long-running, production-owned outcomes, the agent harness and operational control plane—not merely the model—become the critical constraint, so Anthropic argues they should be managed, modular, durable, and continuously adapted to model changes.
  • Why it matters: The talk offers a concrete production architecture for agent systems: separate reasoning from execution, persist sessions as an event log, isolate tools and credentials, and use evaluation and memory loops to improve over time.
  • Best use: Use it as an architecture review for OpenClaw-style long-running agents and as a checklist for deciding which infrastructure to own versus delegate to a managed agent platform.

Executive Summary

Anthropic frames agent development as an evolution of surfaces: the Messages API provided model calls; the Agent SDK packaged an agentic loop, tools, filesystem access, and sandboxing; Claude-managed agents aim to own the remaining production layer. Their central argument is that teams should retain the task definition, domain context, prompts, tools, and product experience, while a managed platform handles the fragile operational machinery: sessions, hosting, execution isolation, observability, credentials, and the agent loop.

The most useful design lesson is that a harness embeds assumptions about model limitations, and those assumptions can become harmful as models improve. Anthropic cites a Sonnet 4.5 behavior it calls “context anxiety,” for which it added context-reset logic; with Opus 4.5, the behavior disappeared, leaving the fix as latency and cache-management overhead. Their implication is that agent systems must be modular enough to revise or remove harness behavior quickly rather than locking in model-specific workarounds.

The proposed architecture decouples the agent’s “brain” (the model-driven loop) from its “hands” (the sandboxed tool-execution environment). This lets reasoning begin before container setup, allows a failed sandbox to be replaced without losing the agent loop, and permits execution in a customer VPC. A durable session log records messages, tool calls, and results, allowing a failed agent process to reconstruct context and resume, while also serving as the source for observability, memory, and post-run improvement.

The demo is an SRE incident investigator: an agent with an Opus model, system prompt, standard shell-like tools, and MCP access to metrics/deploy data investigates a 10x P99 latency spike, searches logs, correlates deployment timing and code diffs, and returns a root cause. The final section describes emerging control-plane features: vault-backed credentials, private MCP tunnels, periodic “dreaming” to revise memory from prior session transcripts, and an “outcomes” grader agent that retries work against a defined success rubric. The claims are product-oriented and mostly lack independent benchmarks, but the underlying design patterns are highly reusable.

Key Takeaways

  • Claim: Production agent infrastructure is becoming a larger bottleneck than the basic model call or agent loop, particularly as agents become asynchronous, long-running, and responsible for complete outcomes. | Evidence: The speakers enumerate the infrastructure teams otherwise build themselves: hosting and scaling, session persistence, concurrent runs, filesystem access, secure code execution, credential handling, and observability. They contrast this with the earlier Messages API and Agent SDK surfaces, which left meaningful production work to customers. | Implication: For any serious agent program, treat the runtime/control plane as a first-class architecture decision; a strong prompt and tools alone do not constitute a production agent. | Caveat: This is Anthropic's rationale for Claude-managed agents, so the framing is necessarily product-led rather than a neutral comparison of build-versus-buy options.
  • Claim: Harness logic must be treated as perishable because it encodes assumptions about model weaknesses that may disappear or reverse with the next model release. | Evidence: Anthropic says Sonnet 4.5 showed “context anxiety”: it prematurely wrapped up work near its context limit even when space remained. The team added context resets, but says Opus 4.5 no longer exhibited that behavior, making the reset mechanism counterproductive through extra latency and incorrect cache discards. | Implication: Version agent harnesses independently from application logic, run regression evaluations on every model upgrade, and be willing to delete safeguards that no longer improve outcomes. | Caveat: The example is specific to Anthropic model generations and is not evidence that all context-management mechanisms are unnecessary; it supports frequent revalidation of them.
  • Claim: Decoupling reasoning from tool execution is the foundational reliability, latency, and deployment pattern for long-running agents. | Evidence: Anthropic initially placed the agent loop and tools in one container, then found that model reasoning could not begin until container setup completed and that failure of either component brought down the entire agent. In the split design, the brain starts immediately while the sandbox is provisioned in parallel; if a sandbox dies, the brain can create a replacement and retry. | Implication: Design OpenClaw-style systems as a durable orchestrator plus replaceable execution workers, rather than a single stateful container running both planning and tool use.
  • Claim: A durable, append-like session log is a multipurpose primitive: it enables recovery, observability, context reconstruction, and learning across sessions. | Evidence: A managed-agent session persists each user message, model response, tool execution, and result. If the brain fails, it can reread the log and resume. The same log can be exposed as an event trace for debugging and reused as history for memory updates. | Implication: Make event-sourced session state central to agent architecture, but pair it with explicit trace governance, redaction, retention, and replay-cost controls. | Caveat: The talk does not address retention policies, PII handling, data residency, access controls, or the cost of storing and replaying potentially large traces.
  • Claim: Credentials should never be exposed to the model; tool execution should receive narrowly scoped secrets only at runtime. | Evidence: Anthropic introduces vaults in which credentials are stored securely and decrypted only when needed for tool execution, stating that the model never sees the underlying security tokens. This builds on separating the loop from the execution environment. | Implication: Adopt a brokered-credential design: agents request capabilities, an execution layer obtains scoped credentials just in time, and raw secrets stay outside both model context and ordinary sandbox filesystems. | Caveat: Keeping tokens out of the model does not by itself prevent harmful tool actions; authorization scope, approval flows, egress controls, and auditability remain necessary.
  • Claim: The split brain/hands architecture can materially reduce perceived latency because reasoning does not wait for sandbox provisioning. | Evidence: Anthropic reports 60% faster P50 time to first token and more than 90% improvement at P95 after decoupling the model loop from container setup; the system can also skip sandbox setup entirely when a task does not require it. | Implication: Classify requests by whether they actually need code/filesystem execution, initiate model reasoning immediately, and provision execution environments lazily or in parallel. | Caveat: No test methodology, absolute latency values, workload mix, or external comparison is provided, so the figures should be read as internal product results rather than general performance expectations.
  • Claim: Agent reliability can increasingly be expressed as outcome attainment rather than a single execution attempt, using a separate grader to evaluate a defined success rubric and trigger retries. | Evidence: Anthropic's “outcomes” feature has developers define success criteria and failure cases; a grader agent runs alongside the main loop, checks the work against the rubric, and causes the worker to continue trying if criteria are not met. Its “dreaming” process periodically processes session transcripts plus current memory to extract insights and revise future memory. | Implication: For high-value workflows, define machine-checkable success criteria, use independent grading where feasible, and place explicit retry budgets, escalation conditions, and human review gates around iterative execution. | Caveat: A grader only improves reliability to the degree that its rubric and judgments are valid; repeated retries can amplify cost, latency, or unwanted actions without budgets and stopping conditions.

Detailed Brief

Resource model and recovery semantics

  • Claims: Anthropic organizes managed agents around three primitives: agent, environment, and session.; The agent definition contains the model, prompts, tools, skills, and other use-case-specific behavior.; An environment defines where execution occurs; multiple sessions may share an environment definition while receiving isolated container instances.; A session combines an agent and environment into a durable cloud resource and can be idle, running, rescheduling after an error, or terminated when recovery is not possible.
  • Evidence: The SRE demo defines an “SRE investigator” agent using Claude Opus 4.8, a system prompt, a standard toolset including Bash and grep, and an MCP toolset connected to a metrics/deploy dashboard.; Its environment restricts networking to allowed hosts, specifically the intended MCP server.; The demo uploads application logs and skills, then starts a session that searches logs, retrieves metrics and deploy data, identifies timing, examines a code diff, and synthesizes a root-cause explanation.
  • Caveats: The demonstration is a curated incident-investigation flow; the transcript does not show failure cases, inaccurate MCP data, unsafe remediation, or the quality of the final root-cause finding.
  • Implications: Separate reusable runtime/environment specifications from individual user or job runs so that concurrency and isolation are explicit rather than accidental.; For operational agents, constrain allowed network destinations at the environment level rather than relying only on prompt instructions.

Private enterprise connectivity and the emerging memory layer

  • Claims: Security-conscious enterprises may want the agent loop managed in the cloud while retaining tool execution in their own VPC.; MCP tunnels are intended to let private MCP servers remain inside a private network and make only outbound connections to the cloud agent loop.; Anthropic expects memory to evolve beyond per-user facts toward organizational memory containing items such as team runbooks and operational details.
  • Evidence: The speakers position self-hosted sandboxes as a way for customers to control the sandbox control plane and apply their own execution policies.; They describe dreaming as a periodic batch process over daily transcripts and current memory state, extracting organized insights that edit memory for subsequent sessions.; They mention, without elaboration, scheduled deployments and multi-agent orchestration as additional experimental/frontier capabilities.
  • Caveats: The transcript provides no implementation details for memory conflict resolution, provenance, permissions, stale-runbook detection, or how organizational memory is isolated among users and teams.; Features characterized as experiments or frontier capabilities should not be assumed mature solely from this presentation.
  • Implications: Treat organizational memory as a governed knowledge system with provenance and permission boundaries, not simply a larger prompt or undifferentiated vector store.; Private connectivity can preserve enterprise network posture, but it does not eliminate the need for tool-level authorization and outbound-action controls.

Notable Concepts & Terms

  • Agentic surface: The developer interface and abstraction layer for building agents, presented as evolving from raw model calls to SDK-provided harnesses to a fully managed production runtime.
  • Harness: The operational logic around a model—agent loop, context management, tool execution, memory, and reliability behavior—that can either unlock or constrain model capability.
  • Brain / hands separation: Separating the model-driven reasoning loop from the sandbox that executes tools; the talk treats this as the central architectural decision for resilience, lower startup latency, and flexible deployment.
  • Durable session log: A persisted record of messages, model responses, tool calls, and results that supports recovery, traceability, context replay, and downstream memory processing.
  • Context anxiety: Anthropic's label for Sonnet 4.5 prematurely ending work as it approached its context limit; used to illustrate why model-specific harness fixes must be continuously reconsidered.
  • Vaults: A credential mechanism where secrets are decrypted only for tool execution, so the model does not receive raw tokens.
  • Dreaming: A periodic batch process that analyzes session transcripts plus existing memory, extracts new insights, and updates memory to improve later agent runs.
  • Outcomes: A success-rubric mechanism in which a separate grader agent assesses task completion and can make the main agent continue until success criteria are met.

Operator Notes / Why Ken Should Care

  • Audit the current agent stack for coupled orchestration and execution processes; prioritize a durable orchestrator plus disposable, isolated tool workers.
  • Establish a model-upgrade harness regression suite that tests context policy, retries, prompt scaffolding, tool selection, cache behavior, and latency before adopting each new model.
  • Make session traces an event-sourced system of record, then define retention, redaction, tenant isolation, replay permissions, and cost limits before using them for memory.
  • Move all secrets behind just-in-time, capability-scoped credential brokering; prohibit raw credentials from prompts, model context, generic environment files, and persisted traces.
  • For autonomous workflows, introduce explicit outcome rubrics, grader validation, retry ceilings, spend/time budgets, and escalation-to-human conditions.
  • Evaluate whether private MCP connectivity and self-hosted execution are requirements for target enterprise deployments, especially where internal systems cannot be exposed publicly.

Source/Metadata

  • Title: Evolution of agentic surfaces — Gagan Bhat & Isabella Kai He, Anthropic
  • Transcript words: 7339
  • Duration seconds: 1883
  • Timestamp note: No timestamps or chapters were present in the supplied transcript. The latter security, dreaming, and outcomes segment is duplicated in the transcript.

Transcript

5609 words en Processed in 209.6s

Gagan Reviewer All right. Hi, everyone. Thank you for joining us today. I'm Gagan. And I'm Isabella. We're both members of technical staff here at Anthropic, at the Applied AI team. Our team sits at the intersection of product, research, and go-to-market, and we spend a lot of our daytime building agents, evaluating Claude, and finding ways to make it better in different use cases. We're here today to talk about how the surfaces for building agents have evolved in the last three years, and what our teams have learned building agents both internally at Anthropic and externally with our enterprise customers along the way. I'll hand it off to Gagan to kick us off. So here's the plan. We'll first start off by talking about how we've seen agentic surfaces evolve all the way from the Messages API to Claude-managed agents. Isabella will then cover the engineering principles behind Claude-managed agents and how it works under the hood. We'll share a demo that shows what it feels like to build with this latest surface, and then we'll talk about lessons that we've learned from the field by taking this to customers, and finally we'll close with what it feels like to build at the frontier as model capabilities evolve dramatically. Okay, with that, let's talk about how agentic surfaces have evolved. To set the stage, it's important to realize that AI progress is accelerating. From back in the days when the transformer architecture was first defined, to the scaling laws discovered by our founders, to the models that have been released today. Every model release, and we all feel it, improves capabilities that the previous model did not have. And we see this translate over to the tasks that we give models as well. The complexity of these has grown dramatically. Initially, we used to only give them questions, simple Q&A. We then started delegating tasks to them, and now we let agents own entire outcomes. What this means is that as task complexity and model capability have improved, the agentic surfaces have to evolve with them, and we're going to talk about that in this section. So, go back in time to when we first launched Claude 3. We launched it alongside the first agentic surface called the Messages API. This Messages API was simply tokens in and tokens out. You gave information in, and you got text completions out. As I mentioned earlier, task complexity grew over time, and there was a need for the model to fetch information and manage this context as it runs for longer and longer. And so we saw the invention of the agentic loop. This loop is something that every customer started building manually from scratch, that calls Claude, runs the tools, and manages its context. And it was painstaking. And on top of that lived their product, where it had the AI feature, and it would call the agentic loop to accomplish some sort of task. But this was not all. In order to take this to production, there was a whole slew of production infrastructure challenges, like session management, observability, credentials, hosting infrastructure, sandboxing, and more. These challenges were tedious to deal with, and they didn't allow teams to focus on building what mattered most: their product. One of the core challenges here was the agentic loop. It was surprisingly complex to maintain, which brings me to the second part. The second evolution of the agentic surface was the Claude Agent SDK. This agent SDK essentially packaged the harness that we know and love, which is Claude Code, and it shipped with a built-in agentic loop, along with file system access, tools, and a system for doing sandboxing. And then your product would embed the SDK. And there are some primitives already provided for session management and observability, but you still had to hand-roll things like credentials and hosting infrastructure, and much more. You had to figure out how to put this in a box and scale it for your customers. And I want to double-tap on these production infrastructure challenges. And I'm going to enumerate some of these as questions to provoke thought. First, hosting and scaling. Where does the agent run? And how long does the process live for? What scales under load, and what doesn't? Session management. Where does the history and progress of sessions live? How can you have multiple concurrent agents running at scale? File system. How does Claude actually have access to create files and edit files? Fourth, execution isolation. Where does Claude actually run the code that it writes? And how do we keep it secure? Fifth, credentials. How does Claude reach into your sensitive systems without actually getting exposed to the security tokens that you want to protect? And finally, observability. With all the complex agent orchestration going on, how do you figure out what's actually happening under the hood? These were production infrastructure challenges that most teams spend significant portions of their time doing instead of being able to focus on just their product, their task, and their context. Which is why we built Claude-managed agents. The idea of Claude-managed agents is simple. You own the product. You own the task. You own your context. And you call Claude-managed agents to get production-grade infrastructure for your agents. So your product would call Claude-managed agents and get a brain, which is the agentic loop and Claude itself, and all the custom harnesses and evolutions built inside of it. Along with the hands, which is a sandbox that spins up just in time for things like file system access and code execution. And of course, all the bells and whistles that were difficult to maintain before with production infrastructure, like credentials, session management, observability, and hosting infrastructure, all of this is run by Anthropic. And what's yours to build and run is your task, your context, and your domain knowledge. So if you step back and look at how this has evolved, we can see the evolution over time. The Messages API introduced the ability for you to interact with the model with tokens in and tokens out. The Claude Agent SDK brought in a built-in agent harness for tasks. And Claude-managed agents covers everything that you need from your product and everything under that stack. So now I will hand it over to Isabella to talk a bit more about the engineering principles that drove how we built Claude-managed agents. Over to you. Perfect. Thank you. And as Gagan walked through, a takeaway from that entire last section is that models evolve quickly. And so when our team set out to build managed agents, we drew inspiration from the lessons that our teams learned building effective agents and harnesses for these models to capture that evolution as it rapidly advances alongside us. So for the next few minutes here, I want to talk you through some of the fundamentals that underpin the design of Claude-managed agents and the lessons that our team learned along the way as we went about engineering this harness built for model evolution. Let's start with one of the very core principles that inspired our team to build managed agents. And that is that harnesses encode assumptions about what Claude cannot do on its own. This is things like resetting context and managing compaction, and some of the other core primitives that you see here on the screen. The thing about these assumptions is that they have to be questioned frequently because they go stale as models improve. Let's dive into one concrete example. Back when Sonnet 4.5 came out, it exhibited an interesting behavior that came to be known as context anxiety. What this is is that the agent literally got anxious as it approached its context window limit. It started to wrap up tasks early, it started to terminate work, even when it actually had room left to spare in its context window. In order to accommodate for this behavior, what our team did was build in fixes into the harness itself, adding in context resets so that Sonnet 4.5 would be able to reset its context and continue working. But when Opus 4.5 came out, the interesting thing here was that this behavior went away entirely. Opus 4.5 no longer exhibited context anxiety, which means that the fixes that we had added into the harness itself became dead weight. In fact, it became pure overhead, adding things like latency and causing issues with the cache being discarded incorrectly at times. The takeaway here is that we saw that the harness fixes were no longer needed and were actually detracting from model performance with Opus 4.5. So when the model moves and the harness doesn't, it degrades the agent. What we've seen across these last couple of examples is that there's significant maintenance burden that comes with maintaining a harness that can keep up with Claude's rapid evolution. As we work with a range of enterprise teams that are building on top of Claude, we also see a range in the harnesses that are ready for this level of adaptation. Some of the harnesses that we see from customers are more agile, and others are more rigid because they were built around older Claude models. What you don't want to do is have a stale harness that takes weeks or even months to migrate to a new model, especially with how model release cycles have been coming out shorter and shorter. Now, to build an effective harness, what this means is that you have to be designing for the model capabilities of tomorrow, anticipating what future Claude models will be able to accomplish and building your harnesses for that capability. It also means that your harnesses have to be agile, making it easy to iterate for the model capabilities to capture them quickly as soon as they're ready. As we work with a range of enterprise teams that are building on top of Claude, we also see a range in the harnesses that are ready for this level of adaptation. Some of the harnesses that we see from customers are more agile, and others are more rigid because they were built around older quad models. What you don't want to do is have a stale harness that takes weeks or even months to migrate to a new model, especially with how model release cycles have been coming out shorter and shorter. Now, to build an effective harness, what this means is that you have to be designing for the model capabilities of tomorrow, anticipating what the future Claude models will be able to accomplish, and building your harnesses for that capability. It also means that your harnesses have to be agile, making it easy to iterate for the model capabilities to capture them quickly as soon as they're ready. That brings me to Claude Manage Agents, which is a harness designed around a small set of primitives with individual components that you see here on the screen that are independent, making it easy to swap them out and iterate upon them as individual pieces while keeping the overall architecture stable. Another key thing that Claude Manage Agents is designed around is long-running agents. Now, when I look to internal products that are exciting and anthropic, like Claude Code and Claude Tag, which our team is really excited about, and as I work with other enterprise teams who are also building exciting, truly agentic products, a common pattern that I see is that these agents are becoming increasingly asynchronous and are tackling tasks that are increasingly complex and challenging. In order to design a harness that's actually able to capture those levels of work, it needs to have a couple of things. It needs to be good at context engineering, as the context will accumulate over those long horizons and bodies of work. It needs to be good at giving the agent a sandbox that's secure so the agent can actually take action within an environment. It has to be reliable so the agent can run for hours or even days at a time. And it also has to be able to do things like parallelized workflows so the agent can tackle multiple parts of a complex problem at once, and many, many more. So now what I want to do is dive into some of the engineering fundamentals that go into Manage Agents to make it possible to tackle some of those challenges. One of the core architectural decisions that went into Manage Agents that sets the foundation for the rest of the slides that we're going to walk through is the decision to decouple the brain from the hands of the agent. When our team first set out to build Manage Agents, we started by putting the agent loop and the tool execution in the same box, in the same environment. What this meant was that the agent loop would be easily able to call tools and read in results because it had it right there in the same container. But then we ran into a series of limitations. This being that the container was blocking the agent from being able to start its model reasoning, so the agent wouldn't be able to kick off until the container was fully set up. It also meant challenges for reliability because if one part of this component went down, the entire box of the agent would go down. Our solution to this was to decouple the two elements, separating the brain, or the agent loop, from the hands, or the tool execution environment, of the agent. This meant several things. It meant improved reliability, and it also meant that the brain could only spin up sessions when it actually needed them on demand. Now let's dive into how some of these replaceable components meant keeping long-running agents safe. First of all, if the sandbox, or the hands of the agent, died, because the brain was in a separate component, the brain could just spin up a new sandbox and retry, and then continue as it left off. If the brain of the agent dies, we're actually going to walk through something in just a moment here about how everything that the agent does is logged into a durable session resource in a session log. Which means that the brain of the agent can actually just read from that session log, go back into context, and resume exactly where it left off as well. This means that Manage Agents is designed around three core primitives, and that is the agent, and that is what your agent is, especially defining what your agent does for your use case. This is things like the model that goes into your agent, the prompts, the tools, the skills, everything that makes your agent work for your particular use case. Next up is the environment, and this is the container that the agent actually runs in. You can actually have multiple sessions run on the same environment definition, and you can even attach multiple sessions to run the same environment at once, but each with its own isolated container instance. When you combine an agent with an environment, you get a session. What a session is, is a durable resource persisted in the cloud of every single interaction that you have with the agent, which unlocks several things like observability, a long-running instance, reliability, all through this core architectural decision. Now, when I work with many teams that are bringing an agent from a prototyping phase all the way to a production phase, one of the main challenges that we see is that reliability is a core concern. It's a different story to build an agent that runs on your laptop and serves you as a single user compared to when you actually want to deploy it in production and run it at scale for hundreds of thousands or even millions of users. You need to make sure, especially if your agent is going to run for long hours at a time, that it's going to be reliable and can actually recover from tool failures. Manage Agents, because of the way it's designed around those three primitives that we just walked over, is able to have four distinct session states. And that is idle, when your agent is waiting on user input; running, when it's actually executing; rescheduling, when it encounters an error and it's going to retry; or terminated, if it's unrecoverable. This means that the agent can always go back to an existing session and resume where it left off, and it also has a mechanism for it to recover from those failures in production. Another common theme that we see with designing effective agents, especially as model capabilities evolve, is context engineering. Now, context engineering is something that our team has done a ton of research into because it is one of the things that separates an effective agent from an agent that gets lost in context rot. Context engineering is also difficult, and with many traditional harness implementations, the context window and the session are one and the same, which means that quad, if it wants to come in and discard portions of the context in its current session run, doesn't have a mechanism to be able to recover pieces of that context back into its window if it loses them at one point in its current session run. However, because everything in managed agents is logged to a durable, persisted session log resource, what this unlocks is that the harness can actually just read in slices of that context from the session log into its current window. If it then has quad coming in and editing or discarding portions of that run, it can simply recover them by just rereading them from the session log because everything is persisted in that log resource. What this means is that what we see is increasingly developers are able to rely on portions of the managed agent harness that come with the harness itself. These are things like the agent loop, memory, observability. All that comes alongside building with quad managed agents. It also exposes key areas for the developer to be able to customize. And this is context management and domain expertise. This is what separates a coding agent from a legal agent or go-to-market agent. And again, it's what makes your agent truly ready for your users. For instance, with quad code, quad code uses a set of tools like Bash and grep on your laptop, just like how developers do when they open up their terminal. But a go-to-market agent or legal agent would need a vastly different set of tools. So by having this part of the managed agent harness managed by Anthropic, what that means is that the developer can focus their time on designing the right system prompts, the right skills, and the right tools to make their agent truly work for their users. Now what I want to do is hand it back over to Goggin to walk you through one example in a live demo where you can see a customized agent for a production use case and how simple it is to build a production-ready agent with managed agents. All right. Thank you, Isabella. Let's switch over. All right. So now that you heard from Isabella how managed agents works under the hood, let's look at what it feels like to actually build with it. What does it feel like to actually create your own production-grade agent from scratch? This is a semi-interactive demo, so please bear with me here and follow along. Imagine that you're an engineer and you own a dashboard that contains all the key metrics for the services that you own. It's called Atlas. And one day you're just enjoying life, and you start to see that the P99 latency suddenly starts spiking. It's 10x over baseline. It's 10x over baseline. So you have an incident on your hands. You see your logs. You see a bunch of text. And you have to figure out what exactly is going wrong. Wouldn't it be nice if there was a site-reliability engineering agent that could investigate all these tedious data points and come back to you with the root cause before even you open your dashboard? Well, that's what we're going to implement today. What does it feel like to actually create your own production-grade agent from scratch? This is a semi-interactive demo, so please bear with me here and follow along. Imagine that you're an engineer and you own a dashboard that contains all the key metrics for the services that you own. It's called Atlas. And one day you're just enjoying life, and you start to see that the P99 latency suddenly starts spiking. It's 10x over baseline. It's 10x over baseline. So you have an incident on your hands. You see your logs. You see a bunch of text. And you have to figure out what exactly is going wrong. Wouldn't it be nice if there was a site-reliability engineering agent that could investigate all these tedious data points and come back to you with the root cause before you even open your dashboard? Well, that's what we're going to implement today. We're going to build this from scratch using Claude-managed agents, and it'll walk you through the steps using the primitives that Isabella shared earlier. Okay, so let's build it. The first primitive, as mentioned earlier, is the agent definition. You define what the agent does and everything that it needs to accomplish that task. So in this case, I define the name as the SRE investigator. I give a model as Claude Opus 4.8. I have a system prompt that defines the instructions for the agent on how to behave. And tools. The agent toolset gives it a standard set of tools like Bash, Grep, and Blob, et cetera. And there's an MCP toolset that connects to my dashboard and allows it to pull specific things like deploys and the metrics that I showed earlier. That's step one: agent definition. Step number two. Let's now define where does it run. Claude. This is the environment that we were mentioning before. Here, we create an environment that is SRE Sandbox. And I configure it to run on the Anthropic cloud with the networking limited and allow those hosts only, being the MCP server that I wanted to communicate to. This environment effectively stops Claude from doing things that you didn't intend. So you can control boundaries here. Next, let's give it the logistics and the details that it needs to solve this task. And one of the key things is the application logs I showed earlier. You can upload files and skills just like this. And you can set it up so that it reads from those. So now we have the agent definition, so what it is. We have the environment, which is where it runs. We have relevant evidence. And so now we set it up so we can kickstart a session. The session combines these durable resources into a new session. And it specifies the log point as a resource. And it allows the agent to kick off. That's it. What you have defined here is now living in the Anthropic cloud. And this can be dynamic as needed in your application. So once we do this, we kickstart the session. And let's go back to the investigator agent earlier and say, hey, I have an incident. My checkout is super high. Can you please investigate? Claude immediately spins up. The brain spins up in Claude Managed Agents in the cloud. It uses the hands, which is a sandbox, to grep and find specific details in the application logs. It uses relevant metrics using MCP tools and finds the recent deploys and isolates where the incident started. It does further investigation, finds the code diff, and synthesizes this information to figure out a final root cause. Just like that, Claude Managed Agents was able to run in the cloud. We were able to define it in code and have everything running end to end. And to clarify, this is just one session. All of this is production infrastructure. So you can imagine multiple sessions kicked off by all of your users, ready to go immediately. And we have a beautiful observability dashboard that allows you to see all of these sessions on demand. If I click into one of the observability dashboards on the Claude console, you will find the exact event trace, the session logs for that event, including the tools that are used, the results of it, and any other agent messages that came into picture. With that, we will walk through what it feels like to build a Claude Managed Agent from scratch with just a few lines of code. All right. Perfect. So with that, let's go into the next section. We've learned now how Claude Managed Agents works under the hood. We've learned what it feels like to build a production-grade agent at scale using it. Now let's talk about what we've learned from taking this to the field, what we've heard from customers, and what we've learned generally building production-scale agents. There's four lessons here, and I'll start off with the first one. The first lesson is to keep the credentials away from your agent. A lot of customers ask me, how do I make sure my agent doesn't read or see the environment file that contains all my security tokens? This is a very important aspect. We already get some of this because, as Isabel mentioned earlier, we separated the brain from the hands. So where the agentic loop runs is separate from where the tool execution happens. We took it a step further by introducing the concept of vaults, where you can store security credentials in a secure way, and they're decrypted only when needed at tool execution runtime. This way, you can effectively keep credentials away from your agent, and the model never sees your security tokens. Perfect. Now time for lesson two. And I'm sure if any of you in this room have built a production-ready agent, latency has been one of the things that's top of mind for you. What we realized when we decoupled the brain from the hands of the agent, as we talked about, is that this actually unlocked a key benefit that really mattered for a lot of our customers building on managed agents out in the wild. And that is that it improved latency significantly because the agent was no longer blocked on reasoning based on container setup. So we go back to the first version with the coupled design where we had the harness in the same container in one single box. Essentially, the model wouldn't be able to start reasoning or outputting its first token until that container setup was fully complete, which meant that it had delays in the latency, especially for time to first token. When we then decoupled the brain from the hands of the agent, this is what we get. Now we can have model reasons start immediately, and we can run container setup in parallel. What this means is that the model can then run container setup so that the brain of the agent actually has the hands when it needs it, or we can actually skip the container setup entirely if, for this particular task, we actually don't need the container setup in the first place. What we then saw when we tested this is that we saw 60% faster time to first token for P50 use cases or median use cases, and over 90% improvements in latency for time to first token in P95 use cases. Lesson three is about session logs. A lot of customers asked us the question, how can I figure out what's actually going on in my agent under the hood, and how can I make my agent better over time? Turns out the answer to both of these questions lies in something that we call the session log or traces. The session log essentially contains events of everything that happened during an agent execution. So the user message, the model response, the tool executions, the results, everything is written play by play. Now if you surface the session log in a UI that users can see, it provides observability. It turns out that the same session log also improves memory and provides self-improvement for the agent. Memory essentially allows the agent to remember things about the user. And session logs gives a history of past executions. And if you combine that with something that we call dreaming, it allows memory to be updated and improved over time. So the next time your agent runs, it gets better. We'll talk a bit more about this later. And now for the last lesson that we have for you today, that is security for tool execution. And this is something that we heard from a lot of enterprise teams that were wanting to build on manage agents, is that it really mattered to them how they were able to control the environment where they ran tool execution. For a lot of teams that were very security conscious, they wanted to be able to have everything controlled in their own virtual private cloud. And because of the key decision that we made decoupling the brain from the hands of the agent, what we actually get is that the hands can run anywhere, including in your virtual private cloud. So the feature that we released called self-hosted sandboxes, we built this from an engineering perspective because of the feedback that we heard and essentially made it available to have customers control their sandbox control plane exactly for their own execution environments and to have tools run exactly under their own policies. Another feature that we unblocked is MCP tunnels. And this is from teams that were saying that they wanted to expose MCP servers to their agent, but didn't want to have their MCP servers running over the public internet. For those teams, essentially with MCP tunnels, they can have their MCP servers run only within their private network and only making outbound calls to the cloud agent loop. And so with everything that we've talked through and those four lessons that Gagan and I just walked over, we talked about how manage agents is helping you build for this iterative capability what we actually get is that the hands can run anywhere, including in your virtual private cloud. So the feature that we released, called self-hosted sandboxes, we built this from an engineering perspective because of the feedback that we heard and essentially made it available to have customers control their sandbox control plane exactly for their own execution environments and to have tools run exactly under their own policies. Another feature that we unblocked is MCP tunnels. And this is from teams that were saying that they wanted to expose MCP servers to their agent, but didn't want to have their MCP servers running over the public internet. For those teams, essentially with MCP tunnels, they can have their MCP servers run only within their private network and only making outbound calls to the cloud agent loop. And so with everything that we've talked through and those four lessons that Gagan and I just walked over, we talked about how manage agents is helping you build for this iterative capability and helping you follow along with model evolution. What I now also want to talk you through is how those fundamentals are built and expanded upon with some of our most exciting frontier features to show you how cloud manage agents will continue to evolve and capture the model capabilities of tomorrow. Cloud manage agents, we touched on just the tip of the iceberg of the features that are available today, and these cover the fundamentals. But we're also excited to see how the harness evolves as model capabilities evolve. And I'm excited for some of the new features that we've been experimenting with, like scheduled deployments, self-hosted sandboxes, multi-agent orchestration, dreaming, outcomes, memory, and more. We won't have time to cover all of them, so I'm going to talk about two of our favorites, which is dreaming and outcomes. Let's start with dreaming. As I mentioned before, we have access to the transcripts or the session logs from the agent's daily sessions, and the agent also has a current memory state. What we found is that as models have evolved and become more capable, if you feed the transcripts and the memory state as a periodic batch process with what we call dreaming, it allows us to extract new insights and new organized structures that essentially feed back and edit the memory as needed to make the next day's agent sessions automatically much more intelligent. This is how we're seeing self-improving agents as they execute more and more over time. Dreaming and memory, we feel, are just two cornerstones of a new frontier unified memory system. Memory gives the agent the ability to remember things across the user that's specific to its use case. Dreaming allows agents to self-improve, but we see a new form of memory emerge that is organizational scale, and that kind of illustrates and stores the team's runbooks and details. And we believe that this is just the initial areas in which we can see harnesses evolve towards as models become more capable. I also love dreaming, but one of my other favorite features is something called outcomes. What you get with an outcome is that what we allow users to essentially define is success criteria for their agents. You define a rubric. You say what means that the agent is actually able to complete the task successfully. You define failure cases. And then outcomes essentially starts a separate grader agent that runs alongside your agent loop and looks at whether that agent was actually able to accomplish the task based on your defined success criteria. What this then does is that the agent will execute the task at hand. It will then look at that grader and check across the rubric that you defined. If the grader determines that the agent was not able to complete the task, it will keep trying until it reaches that success criteria that you have defined for your agent. What really excites me about outcomes is that we're moving more towards a world where we can have an agent understand what success actually means for a task and have a mechanism to keep iterating it, which gives us more reliability that the agent can actually complete the outcomes. And we can start to unlock a new set of tasks that were not possible just a couple of months ago as models continue to evolve and can accomplish outcomes or tasks that are increasingly complex, especially as Claude achieves new levels of intelligence. So across everything that Goggin and I have talked about today, what Manage Agents is trying to do is to close the gap between what products offer today on many surfaces with static harnesses and what models can actually do. What we see as quad models and other models essentially evolve alongside this exponential trajectory is that harnesses have become the limiting factor to what models can achieve. And so with Manage Agents, with the core architectural foundation that we offer to developers, along with all of these new exciting features that we're continuing to build, what we're trying to do is to close that gap so that products can get closer to what models can actually achieve today. And so with that, I hope all of you walk out of this room learning something new about how our team went about building Manage Agents and how Quad Manage Agents is structured to capture frontier intelligence as models continue to evolve and be a harness that's production-ready for real workloads. Thank you all so much today for being here and for listening to our talk. Thank you. Thank you. A lot of customers asked us the question, how can I figure out what's actually going on in my agent under the hood, and how can I make my agent better over time? Turns out the answer to both of these questions lies in something that we call the session log or traces. The session log essentially contains events of everything that happened during an agent execution. So the user message, the model response, the tool executions, the results, everything is written play by play. Now if you surface the session log in a UI that users can see, it provides observability. It turns out that the same session log also improves memory and provides self-improvement for the agent. Memory essentially allows the agent to remember things about the user. And session logs gives a history of past executions. And if you combine that with something that we call dreaming, it allows memory to be updated and improved over time. So the next time your agent runs, it gets better. We'll talk a bit more about this later. And now for the last lesson that we have for you today, that is security for tool execution. And this is something that we heard from a lot of enterprise teams that were wanting to build on manage agents, is that it really mattered to them how they were able to control the environment where they ran tool execution. For a lot of teams that were very security conscious, they wanted to be able to have everything controlled in their own virtual private cloud. And because of the key decision that we made decoupling the brain from the hands of the agent, what we actually get is that the hands can run anywhere, including in your virtual private cloud. So the feature that we released called self-hosted sandboxes, we built this from an engineering perspective because of the feedback that we heard and essentially made it available to have customers control their sandbox control plane exactly for their own execution environments and to have tools run exactly under their own policies. Another feature that we unblocked is MCP tunnels. And this is from teams that were saying that they wanted to expose MCP servers to their agent, but didn't want to have their MCP servers running over the public internet. For those teams, essentially with MCP tunnels, they can have their MCP servers run only within their private network and only making outbound calls to the cloud agent loop. And so with everything that we've talked through and those four lessons that Gagan and I just walked over, we talked about how manage agents is helping you build for this iterative capability and helping you follow along with model evolution. What I now also want to talk you through is how those fundamentals are built and expanded upon with some of our most exciting frontier features to show you how cloud manage agents will continue to evolve and capture the model capabilities of tomorrow. Cloud manage agents, we touched on just the tip of the iceberg of the features that are available today, and these cover the fundamentals. But we're also excited to see how the harness evolves as model capabilities evolve. And I'm excited for some of the new features that we've been experimenting with, like scheduled deployments, self-hosted sandboxes, multi-agent orchestration, dreaming, outcomes, memory, and more. We won't have time to cover all of them, so I'm going to talk about two of our favorites, which is dreaming and outcomes. Let's start with dreaming. As I mentioned before, we have access to the transcripts or the session logs from the agent's daily sessions, and the agent also has a current memory state. What we found is that as models have evolved and become more capable, if you feed the transcripts and the memory state as a periodic batch process with what we call dreaming, it allows us to extract new insights and new organized structures that essentially feed back and edit the memory as needed to make the next day's agent sessions automatically much more intelligent. This is how we're seeing self-improving agents as they execute more and more over time. Dreaming and memory, we feel, are just two cornerstones of a new frontier unified memory system. Memory gives the agent the ability to remember things across the user that's specific to its use case. Dreaming allows agents to self-improve, but we see a new form of memory emerge that is organizational scale, and that kind of illustrates and stores the team's runbooks and details. And we believe that this is just the initial areas in which we can see harnesses evolve towards as models become more capable. I also love dreaming, but one of my other favorite features is something called outcomes. What you get with an outcome is that what we allow users to essentially define is success criteria for their agents. You define a rubric. You say what means that the agent is actually able to complete the task successfully. You define failure cases. And then outcomes essentially starts a separate grader agent that runs alongside your agent loop and looks at whether that agent was actually able to accomplish the task based on your defined success criteria. What this then does is that the agent will execute the task at hand. It will then look at that grader and check across the rubric that you defined. If the grader determines that the agent was not able to complete the task, it will keep trying until it reaches that success criteria that you have defined for your agent. What really excites me about outcomes is that we're moving more towards a world where we can have an agent understand what success actually means for a task and have a mechanism to keep iterating it, which gives us more reliability that the agent can actually complete the outcomes. And we can start to unlock a new set of tasks that were not possible just a couple of months ago as models continue to evolve and can accomplish outcomes or tasks that are increasingly complex, especially as Claude achieves new levels of intelligence. So across everything that Goggin and I have talked about today, what Manage Agents is trying to do is to close the gap between what products offer today on many surfaces with static harnesses and what models can actually do. What we see as quad models and other models essentially evolve alongside this exponential trajectory is that harnesses have become the limiting factor to what models can achieve. And so with Manage Agents, with the core architectural foundation that we offer to developers, as long as with all of these new exciting features that we're continuing to build, what we're trying to do is to close that gap so that products can get closer to what models can actually achieve today. And so with that, I hope all of you walk out of this room learning something new about how our team went about building Manage Agents and how Quad Manage Agents is structured to capture frontier intelligence as models continue to evolve and be a harness that's production-ready for real workloads. Thank you all so much today for being here and for listening to our talk. Thank you. Thank you.