Open Reader

Claude for Long-Horizon Tasks — Lance Martin, Anthropic

completed 25:18 Jul 22, 2026 Watch on YouTube

Current Status

completed

Video ID

9QebvrrY3KY

RAG / Chat

Enabled
Claude for Long-Horizon Tasks — Lance Martin, Anthropic
Description

Claude is capable of long horizon tasks. In this talk, we'll share lessons learned about building agent harnesses for reliable and secure long-horizon work. This include decoupling the brain and hands, self-verification, self-learning, and design for evolving agent harnesses. ### Lance Martin Member of Technical Staff · Anthropic [X/Twitter](https://x.com/RLanceMartin) · [LinkedIn](https://www.linkedin.com/in/lance-martin-64a33b5) · [Website](https://rlancemartin.github.io) Member of technical staff at Anthropic. Working on the Claude Platform, including Claude Managed Agents and the claude-api skill in Claude Code. Prior to Anthropic, was one of the early team at LangChain. Prior to LangChain, spent several years focused on vision for self-driving cars (Uber ATG, Ike, Nuro) and got a PhD from Stanford.

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Long-horizon asynchronous agents require more than a capable model: they need durable state, decoupled execution and credentials, independent verification loops, self-improving memory, and an organization-wide harness.
  • Why it matters: This is a concrete architecture and operating model for moving from locally supervised coding agents to reliable, secure agents that can work for many hours across shared organizational systems.
  • Best use: Use it as a design review for Ken's agent control plane, especially session durability, credential isolation, verifier-driven workflows, memory maintenance, and multi-user organizational agent access.

Executive Summary

Lance Martin frames agent-product evolution as a function of model task horizon. When models could act autonomously for only 10–20 minutes, autocomplete and chat were the appropriate human-in-the-loop interfaces. Around one hour of autonomy enabled synchronous local coding agents such as Claude Code. As frontier models enter a claimed 12-plus-hour regime on METR-style long-horizon tasks, Anthropic believes asynchronous managed agents become viable—but only with infrastructure designed for failure recovery, security, memory, and evaluation.

The architecture centerpiece is a separation of the agent "brain" from its "hands." Rather than colocating the harness, execution sandbox, session state, and credentials in one container, Managed Agents treats the harness as stateless, stores the session as an append-only event log, runs tools in separate containers, and keeps credentials in a vault. This permits recovery from harness or sandbox failure, allows one agent to coordinate multiple execution environments, and avoids leaving sensitive credentials exposed to an autonomous model for hours.

Martin's main workflow pattern is a builder-verifier loop: one context performs work while a separate, independently prompted context checks it against a defined outcome or rubric. The agent continues iterating until the verifier passes rather than relying on the original working context to grade itself. He argues that this shifts steering from a human repeatedly correcting the model to environmental feedback that enables self-correction, and demonstrates the pattern with OpenAI's Parameter Golf ML-research benchmark.

The talk also argues that memory should be model-managed but not trusted blindly. Better models write more useful strategic abstractions during a task, while a separate offline "dreaming" process can inspect prior sessions and memories, consolidate them, and correct accumulated mistakes. Finally, Martin positions Claude Tag as an early example of an "org-level harness": a shared, credentialed, organizational agent with persistent context, shared access, research and de-duplication capabilities, and proactive alerting—not merely a Slack bot.

Key Takeaways

  • Claim: Asynchronous agents become a good user experience only when autonomous task horizons are long enough that agents can make meaningful progress without repeatedly returning to the user. | Evidence: Martin cites METR-style horizons of roughly 10–20 minutes for Opus 3-era models in 2024, around an hour for the synchronous coding-agent era, and says leading frontier models are now in a 12-plus-hour regime; he says prior attempts at async agents failed when agents hit errors and returned too quickly. | Implication: Do not add async execution merely as a deployment mode for short-loop agents; reserve it for jobs with clear outcomes, meaningful unattended work, and enough duration to justify asynchronous handoffs. | Caveat: The long-horizon capability claims are discussed at a high level and are not accompanied by benchmark methodology or product-level reliability rates in the transcript.
  • Claim: A long-running agent should decouple its reasoning harness from execution sandboxes, session persistence, and credentials. | Evidence: Managed Agents uses a stateless harness connected to an append-only session event log; the harness invokes separate containerized "hands" for execution, while secrets remain in a separate vault. If a harness or sandbox dies, the session survives, and one harness can coordinate multiple containers. | Implication: Treat session state, tool execution, and identity as independent control-plane components. Design for restart/replay, scoped tool access, and credentials that are never simply injected into an agent-controlled runtime. | Caveat: This architecture improves recoverability and credential isolation, but it does not eliminate risks from unsafe tool authorization, malicious instructions, or poor outcome definitions.
  • Claim: Independent verifier loops are a stronger primitive for long-horizon work than asking the same agent context to both execute and judge its own work. | Evidence: Martin says self-grading in the same context produces confabulation and weak critical evaluation; the proposed loop separates a build context from a verifier context with a goal or rubric and exits only after the verifier confirms the required outcome. | Implication: For consequential agent workflows, encode acceptance criteria in executable tests, rubrics, or external checks, then use an independent verifier context as the stopping gate rather than trusting an agent's completion claim. | Caveat: Verification quality is bounded by the measurability of the outcome and the verifier's own capability; ambiguous or gameable rubrics can still yield false completion.
  • Claim: High-capability models benefit disproportionately from outcome-driven iterative loops because feedback can replace continual human steering. | Evidence: On OpenAI's Parameter Golf benchmark—training a small model with eight H100 GPUs in under 10 minutes—Martin ran Managed-Agent outcome loops for Opus 4.7 and an unnamed frontier-class model, allowing iteration until 20 experimental iterations and benchmark criteria were satisfied. He describes strong results from frontier models under this pattern. | Implication: Invest effort in the environment signal—tests, telemetry, experiment constraints, and outcome contracts—because that is what lets capable agents iterate productively without a human micromanaging every step. | Caveat: The talk provides no comparative score table, baseline setup, or independent replication for this demonstration.
  • Claim: Memory systems should combine in-session model-authored notes with offline consolidation and correction. | Evidence: In Claude Plays Pokémon, Sonnet 3.5 produced tactical, low-value memory notes, while newer 4.6-era models wrote more strategic notes and progressed further. In five of five raw-memory replicates, an incorrect location memory led the agent into a trap door; an offline "dreaming" process reviewing traces and memory corrected the error and enabled progress. | Implication: Do not regard persistent memory as ground truth. Maintain provenance and versioning, run periodic consolidation or critique jobs, and measure whether memory updates improve downstream task success rather than merely producing cleaner-looking notes. | Caveat: Offline memory rewriting consumes additional compute and must be evaluated in the target workflow; Martin explicitly says benchmarks/evals are needed to verify that it improves results enough to justify the cost.
  • Claim: Use flexible, model-programmable memory substrates rather than a prescriptive schema that dictates memory categories in advance. | Evidence: Martin says file systems are not uniquely necessary—databases can work—but the substrate should expose simple primitives that allow the model to write and manage memory. He reports performance declines when developers predefine rigid memory types and schemas, arguing that stronger models can infer useful abstractions themselves. | Implication: Give agents general read/write/search primitives over a governed memory layer, then evaluate their self-organized structures before hard-coding taxonomy. Avoid prematurely constraining memory to developer-designed fields. | Caveat: A flexible substrate does not remove governance requirements; production systems still need access controls, retention rules, observability, and safeguards against uncontrolled memory growth.
  • Claim: The next significant product surface is an organization-level, multi-user agent harness with its own identity, shared context, and proactive behavior. | Evidence: Martin describes Claude Tag as more than a Slack interface: it has organization-scoped identity and credentials, shared organizational context, concurrent steering by many users, the ability to check prior work and de-duplicate findings, and configurable alerts for information the organization may need. | Implication: Build agent systems as shared organizational infrastructure rather than collections of personal copilots: centralize connectors and context, define agent identity and authority, and create explicit policies for proactive notifications and multi-user requests. | Caveat: Shared organizational context materially raises permissioning, prompt-injection, data-boundary, ownership, and accountability requirements, which the talk identifies as important but does not detail operationally.

Detailed Brief

Product/API progression: from model calls to managed deployment

  • Claims: Anthropic presents three progressively higher-level surfaces: the Messages API, Agent SDK, and Managed Agents.; The Messages API is a prompt-response primitive for teams that want to build and deploy their own harness.; Agent SDK exposes Claude Code programmatically, providing Anthropic's coding-agent harness while leaving broader deployment concerns to the implementer.; Managed Agents bundles both an agent harness and managed deployment infrastructure for long-running asynchronous work.
  • Evidence: Martin dates the Messages API to roughly two years before the talk and says Managed Agents was released over the preceding months, since April.; He characterizes the key transition as moving from a simple message exchange to a harness, then to a harness plus durable managed infrastructure.
  • Caveats: The transcript does not specify API capabilities, pricing, availability constraints, or migration paths between these products.
  • Implications: Choose the abstraction layer based on how much control-plane responsibility Ken wants to own: custom harnessing, reusable agent harnessing, or managed long-running operations.; A raw LLM API should not be mistaken for an agent platform; persistence, execution isolation, and deployment semantics are separate engineering responsibilities.

Why frontier-agent products differ from frontier models alone

  • Claims: Martin attributes the gap between frontier and non-frontier agent products to a combined system rather than model intelligence alone.; Memory needs to capture user preferences and escalation behavior over a multi-hour task.; Security, particularly prompt-injection resistance, and failure-tolerant agent architecture are required alongside model capability.
  • Evidence: In Q&A, Martin says that for a 12-hour agent, productivity preferences must be encoded so the agent knows when to reach out if stuck.; He explicitly names architecture, infrastructure, security, and memory as areas in which frontier labs have invested to make long-duration agents practical.
  • Caveats: Martin says he is not certain why non-frontier models are not in the same long-horizon regime, so this is a system-design hypothesis rather than a demonstrated causal account.
  • Implications: Benchmarking base-model task performance is insufficient for selecting an autonomous-agent stack; evaluate the full system's memory, security, recovery, and escalation behavior under extended runs.

Notable Concepts & Terms

  • Task horizon: The amount of autonomous work a model can complete before needing intervention; it determines whether chat, synchronous local agents, or asynchronous managed agents are viable product interfaces.
  • Managed Agents: Anthropic's managed surface combining an agent harness with deployment infrastructure intended for long-running asynchronous tasks.
  • Brain / hands decoupling: Separating the reasoning harness from execution containers, persistent session state, and secrets to improve recovery, scale across multiple execution environments, and limit credential exposure.
  • Append-only session event log: A durable, non-destructive record of agent activity that survives runtime failure and can serve as persistent external context for later retrieval.
  • Recursive language models: Referenced as a related architecture in which a model interrogates a persistent external context object rather than relying only on a destructively compacted active context window.
  • Builder-verifier loop: A workflow in which a work-producing agent/context is independently checked against a rubric or measurable outcome until the result passes.
  • Dreaming: An offline, out-of-band memory-consolidation process that reviews sessions and memory traces to correct local errors and preserve more globally useful knowledge.
  • Org-level harness: A shared, multi-user agent environment with organization-scoped identity, credentials, context, connectors, and proactive behavior rather than a personal local agent.

Operator Notes / Why Ken Should Care

  • Audit Ken's agent stack against the brain/hands/session/secrets separation: persistent event log outside runtimes, isolated execution sandboxes, vault-mediated credentials, and recoverable runs should be explicit requirements.
  • Adopt a standard outcome contract for any async workflow: measurable completion conditions, independent verification path, maximum iteration/cost limits, and an escalation trigger when the verifier cannot pass.
  • Run a memory experiment before committing to a schema: compare an agent-managed general store against a fixed taxonomy on repeated tasks, and evaluate task success, error persistence, retrieval quality, and maintenance cost.
  • Add an offline memory-review job with trace provenance and rollback rather than allowing unattended agents to overwrite long-term memory irreversibly.
  • Define the authorization model for a shared org agent before expanding connectors: agent identity, per-tool scopes, data partitions, approval boundaries, notification ownership, and audit trails are prerequisites for proactive operation.
  • Evaluate long-running agents end-to-end under injected container failures, stale/incorrect memories, ambiguous acceptance criteria, and prompt-injection attempts; model benchmark results alone are not an operational readiness signal.

Source/Metadata

  • Title: Claude for Long-Horizon Tasks — Lance Martin, Anthropic
  • Transcript words: 5308
  • Duration seconds: 1518
  • Timestamp note: No usable timestamps or chapter markers were present in the supplied transcript; the latter Q&A material is duplicated in the source.

Transcript

4123 words en Processed in 163.8s

I'm good to go. All right, well, take a quick sip and then let's start. It is great to be here. This is my third year coming to this conference, and I always really enjoy it. Thank you for coming to this workshop. I know there are many interesting talks. Let me talk a little bit about our view of asynchronous agents at Anthropic and some things we've been up to lately. This is a way I think about models and product. You can think about Claude as a light source, and you can think about products as windows that allow the light to pass through. What's interesting is, over time, the window that you need to actually see the light of the model shifts, and we've seen this over the past few years. I'm plotting here different Claude models and their task horizon, so how much autonomous work can they do over time. You might recall back in the Opus 3 days, this was 2024. Models could only do maybe 10 to 20 minutes of autonomous work. This is measured by METR. In that regime, only certain product surfaces made sense, things like autocomplete, things like chat, where your human is very in the loop, because the model is really only doing a very short amount of work before you're steering it. Now, the past year, we saw the rise of synchronous coding agents like Claude Code, and this is a shift because then models could do maybe an hour of work, so it made sense to have them run, but typically locally, where you could still steer them easily. It's interesting because during this regime, I remember efforts, and I was involved in some efforts, to build asynchronous agents, but when models can only do an hour of work, async as an experience is bad. The model goes off, it hits an error, and it comes back to you over a short period of time. In order to really unlock async, we needed longer task horizons, and we're starting to see that now. With this shift in capability and time horizon came a shift in the API surfaces. If you look at the lower left, Message API came out two years ago. It's prompt-response. It's great for building harnesses, but it's a very simple API. Again, you're passing a set of messages, you get a response out. There's no sense of deployment with that. So you basically take Messages API, and you can roll your own harness. You can deploy that harness, and you have an agent. Now, over the past year, we saw the rise of coding agents in particular, so we released Agent SDK. That's a way to programmatically call Claude Code, and that's us giving you a harness. Over the past few months, since April, as we've seen longer task horizons, we released a new API called Managed Agents, which packages both the harness as well as all the managed deployment infrastructure for you. I'll talk about some of the themes that underpin this new surface, Claude Managed Agents, and some of the themes that extend beyond just Managed Agents broadly, to think about this new type of asynchronous agents, which can apply, of course, to Claude and other types of longer-running, long-horizon agents. Theme one is decoupling the brain from the hands. When we first set out to build Managed Agents, we started with the container. We put the harness in the sandbox in the same container. Now, the problem here is what happens if the harness dies or the container dies. What we saw is you actually lose the session. Basically, this architecture is tricky for long-horizon agents because what can happen is your agent is running, and if that container dies, you lose everything with it. Also, as models get more capable, putting the credentials in the same container with the agent itself can be problematic. For example, giving Claude access to a bunch of your secrets and letting it run for 10 hours while you're not watching it can be a little bit spooky and has some security concerns, especially as models get extremely capable. For this reason, we decouple what we call the brain, that's the harness, from the hands, the execution environments, and Managed Agents is set up like this. The story here is that the harness becomes a stateless process that talks to a session. The session is an append-only event log, and that can reach out to hands, which are just containers. Those are sandboxes where work is done. One thing that's interesting is Claude is increasingly capable of managing many hands, so you can give one harness access to many different containers to perform tool execution, and Claude can manage this very easily and effectively. If the harness dies or the sandbox dies, it's completely fine because the session is always backed up in this append-only log, and credentials are never actually added to the sandbox. They're stored in a separate vault. So this decoupling actually makes it quite reliable and safe, particularly for long-horizon tasks, and this is one of the core ideas that underpins Managed Agents architecturally. I think an interesting thing that falls out of this is related to, you guys may have seen or come across, the recursive language models work. This session becomes an external context object that the model interrogates, and this has all sorts of benefits for context management. Think about it. When you're doing something like compaction, you're choosing some logic to retain some amount of context, and naively, in a typical step, you're discarding all the context that you didn't compact. In this architecture, and also more broadly with recursive language models, the idea is that the context object is persistent and is unadulterated, so it's append-only, and the model can always go back and fetch old context. So it basically creates a very nice architecture for context engineering because the core context object is immutable in the sense that it's non-destructive, and you only append to it over time. We've seen this to be quite nice in terms of long-horizon context engineering as well. The second theme is use verifiers. One of the problems that we've seen with Claude and other models in general is that when you ask them to do a bunch of work and then say, okay, grade your work, if that same context is being used to both do the work and grade, you can get lots of odd artifacts and confabulation and basically odd behavior. For example, this is an image showing, you can think about the context window as filled with lots of different information, and the model is grading itself. Often it's not properly tuned to do critical verification, and so what we found is it's quite effective to separate verification into a separate context window. There's a very general trend we talk about in a number of different engineering blogs, and the reason is the verifier context can be tuned very specifically for the critiquer verification task. So how this works in practice is, when you build loops, you can have a loop of a build context and a verifier context, and this can be a build agent, verifier agent. What happens is the verifier has some goal or rubric, and it's verifying the result or work of the build agent. This continues in a loop until verification is complete. This is really the big idea behind this whole loops trend that you might have heard about, and we found it to be a very powerful paradigm, especially for working with some of the higher-capacity models. Here are some of the primitives. In Claude Code, you have goal. In Managed Agents, you have outcomes. The principles are really the same. You're setting up a measurable end state. In both cases, you're using an independent context model, you're using independent context to grade over the course of this loop. The loop can run, and you only exit the loop once this independent verifier has verified that it has the outcomes or outputs that you want. That's the key idea. Now let me tell you a story about how I've used this. This is a fun and interesting challenge called parameter golf. It's a benchmark that was put up by OpenAI, and it tests models' ability to effectively do ML research. It asks the model to basically take a small model and train it with eight H100 GPUs in less than 10 minutes, and what you see on the Y is basically loss, so lower is better. What I did was I set up a verifier loop using Managed Agents and outcomes to test the ability for Opus 4.7 and one of our frontier models, mythos-class models, on this task. What you see is basically I allow the model to continue to iterate until the outcome that I specify is satisfied, which is it finished exactly 20 iterations and met all the experimental criteria as defined by the benchmark. What you see is the frontier capability models are extremely good with this pattern of loops and verification because what happens is instead of encoding steering into me as the human, you're encoding the signal into the environment so the model can self-correct when it receives feedback from, for example, the verifier. Using this kind of paradigm with very high-capacity models, you can get very strong results. So the main point I'm trying to make here is that this paradigm of loops, which a lot of people are talking about today, paired with very high-capacity models, is a very good general primitive for long-running asynchronous work. That's really the key point here. Now let me talk about another theme of self-learning. The human brain has two interesting systems for memory. One is, as you go about your day, the hippocampus is writing traces of short-term, very fast experiential memory. You might remember what you had for lunch today. You had lunch an hour ago, you can remember. That's written to short-term memory. When you go to bed at encoding, steering me and into me as the human. You're encoding the signal into the environment so the model can self-correct when it receives feedback from, for example, the verifier. Using this kind of paradigm with very high-capacity models, you can get very strong results. The main point I'm trying to make here is that this paradigm of loops, which a lot of people are talking about today, paired with very high-capacity models, is a very good general primitive for long-running asynchronous work. That's really the key point here. Now let me talk about another theme of self-learning. The human brain has two interesting systems for memory. One is, as you go about your day, the hippocampus is writing traces of short-term, very fast experiential memory. You might remember what you had for lunch today. You had lunch an hour ago; you can remember it. That's written to short-term memory. When you go to bed at night, though, an offline process, or out-of-band process, dreams and dreaming stores certain important details to long-term memory in the cortex. For example, tomorrow you might not remember what you ate for lunch today. That's a local trace. But if you had a very important experience today, maybe this talk, maybe not this talk, but if you had an interesting experience today, that might be written to long-term memory. That's the point: these two subsystems in humans work like this. We actually found memory systems with Claude can employ these same two principles. This is showing Claude's capacity as an in-band memory writer. You basically give Claude memory tools, and when I say memory tools, I mean the ability to write to a file system that is basically a memory directory. That's really it. This is showing some work on Claude Plays Pokemon with Claude Sonnet 3.5, and here's the key point: when Sonnet 3.5 is given access to a memory directory and it can write memory, quote-unquote, in-band as it progresses through this game, it's not very good. The memories it writes are pretty crappy. It's tactical notes. It's not very strategic, and the game progress is quite limited. But with more recent models, like this is looking at 4.6, the notes are much more strategic and game progress is much further. The key point I'm making here is that Claude has gotten much better at this in-band memory writing across model generations. This is another way to show that same result. This is a benchmark that I ran called Continued Learning Bench. It's an open-source benchmark. I took one of the tasks. This is a task that basically asks the model to perform sequential question answering with a SQL database, and it can write memory in between each step. What you see is the performance improves across models. What this is showing is that models get natively better at this in-band memory writing with respect to model capability. Some of the most interesting things I found from this are that the main differentiation between a lower-capacity model and a high-capacity model is this distillation step. Basically, higher-capacity models have a better sense of what abstraction to save to memory that will be useful later. They're not just writing a specific fact; they're writing, how does this generalize to future sessions? That's the key difference that I found higher-capacity models have when they're writing memory. This is a very important thing to keep in mind: models are getting better and better at this kind of in-band memory writing across model generations. Now, there's a little trick here which is very important. We talked about in-band memory, and we talked about dreaming. At night I dream, and I write things to long-term memory. Dreaming is very important because when I'm writing memory in-band over the course of a day, over the course of the session, sometimes you can write incorrect memories, and you're writing things that are locally optimal but not globally optimal. You're writing over the course of a task to help you solve that task, but not necessarily looking forward to future tasks. This is a very important nuance, and this process of dreaming is an offline or out-of-band process that we've used to consolidate and improve memory. I want to show you a fun example that I've used dreaming for. This is again Pokemon, and I played a lot of games of Pokemon with Claude to find this. This was a very hard-won lesson, so I hope you appreciate it. Here's the point: basically what happened is Claude wrote an incorrect memory. What happened is this incorrect memory was related to the location. The details don't necessarily matter. The point is that this incorrect memory causes Claude, or the Pokemon, to mislocalize itself, and it falls through this trap door. That's the key point. It writes this incorrect memory, this incorrect memory causes it to mislocalize in the game, and it falls down this trap. This is very consistent, so I saw this in five replicates. Five out of five replicates with raw memory store fell down this trap. With dreaming, this error is corrected, and it's able to properly localize itself and not fall down this trap. I'll show you a fun visualization of this. This is looking at memory traces, or basically traces of game progress. Going upwards on the y-axis is improvement; that's moving to the next level. Going down is backtracking. What's interesting here, the no-memory baseline, which is that gray bar, doesn't make much progress at all. It's stuck. It's a particularly hard level. The memory, which is the orange, actually keeps falling down this trap door and falls back, so it backtracks. The dreaming traces, though, consistently fix this error in its memory and proceed to the next level. This is a very practical example of how dreaming can work out-of-band on your memory store to fix corrections, because what it does is it looks at your memory store and it looks at all your prior traces or sessions and can find and correct errors. That's the key point, and that's why the dreaming process can be very helpful, because in-band, while Claude is writing to memory, it can make mistakes, and those mistakes get stuck in memory unless you have an offline process to correct them. The key intuition and theme for what I'll open up for questions after this is what I think is this trend that we're going to see moving towards org-level harnesses with async agents. We released Claude Tag, and a lot of reaction was like, ah, Slack bot. Look, I have actually created a lot of Slack bots myself. I understand not every Slack bot is particularly interesting and great. In fact, I've created many Slack bots that are quite bad. But what's interesting about Claude Tag is not the fact that it's accessible through Slack. What's interesting about it is the fact that it has a very, very rich system underneath it, which I want to touch on briefly. In particular, what's interesting about it is it represents what I'm calling an org-level harness. Agents historically have been single-player. You have an agent like Claude Code on your machine with your local context that you've tuned and configured for yourself. What's interesting about Claude Tag is it is a harness that everyone in the organization has access to and can use, so it is a multiplayer harness. What's nice about that is it has its own identity. Its identity and credentials are not tied to a given user, and it has access to organizational context, not just my local context. This has many interesting and useful implications, including the ability to check others' work before you do an experiment, the ability to de-duplicate findings, the ability to do internal research, the ability to give everyone access to a very well-developed harness on day one, whereas with your own personal harness, often new employees, it takes them weeks or maybe even months to ramp up fully, to configure all the right connectors, and so forth. Org-level harnesses are a real leveler of the playing field. I think people saw the Slack bot piece, but they didn't really appreciate the depth of benefit you get from building out org-level harnesses. I do think that was an important thing to note, and I think we're going to see the rise of the kind of harnesses that operate across orgs, across many different users, that can operate increasingly on longer async agents, that can operate on longer time frames. That's one clear follow-up that I think we're going to see from this. Another thing where I think we're going to see is that asynchronous agents are going to be increasingly proactive. Typically, for example, with locally scoped agents, they tend to be reactive. They're responsive to how you steer it, versus async agents increasingly have the ability to steer proactivity. That's one very nice thing about Claude Tag, where you can configure it to tell you things when, looking at this org-level context, it should alert me with things I might need to know about. This is a very important new UX that I think is going to be more and more common with async agents that have access to organizational context, and of course multiplayer. The ability for a single harness to be steered by many, many different people concurrently is an important shift in agent UX that I think will be quite interesting going forward. Yeah, let me just open up for questions, and thank you for listening. sure and and and and and and and and and and and and and and and and and and and and ! and and and and So the question was on the gap between the frontier models, for example, on a benchmark like Meter on long-horizon tasks. Yeah. Okay. steer it versus async agents increasingly have the ability to steer proactivity, and that's one very nice thing about Claude Tagg, where you can configure it to tell you things when looking at this org level context: alert me with things I might need to know about. This is a very important new kind of UX that I think is going to be more and more common with async agents, that access to organizational context. And of course, multiplayer, so the ability for a single harness to be steered by many, many different people concurrently is an important shift in agent UX that I think will be quite interesting going forward. Yeah, let me just open up for questions, and thank you for listening. sure and and and and and and and and and and and and and and and and and and and and and ! and and and and and So the question was on the gap between the frontier models, for example, on a benchmark like Meter on long horizon tasks. Yeah. Okay. So in the latest results that I saw from, for example, Codex 5.6, I think it also is in that 12-plus-hour regime on Meter. So I think you're right that the frontier models, the Mythos class models, strong models from OpenAI, are in this 12-plus-hour regime. Why is it that non-frontier models are not in that regime? I am actually not necessarily sure. I do think, I do think in order to build agents that can effectively operate in this regime, it's important to know that it's not just the model capability. For example, with the Cloud Tag product, it's actually a combination of improvement in memory, because memory is very important. If you have agents working for, for example, 12 hours, you want to make sure that your productivity preferences are well encoded in memory, so it knows when to reach out to you if it gets stuck, for example. So memory is very important. Security is very important, so resistance to prompt injection. Also, model architecture, the agent architecture, is very important, that decoupling in brain and hand, so it's secure and safe and resistant to failure. So actually, I think to build real agents that can operate in these long time horizons, a bunch of things need to come together in terms of architecture, infrastructure, security, memory. That might be why. Frontier Labs invested in all these areas, so that might be why you see a gap in terms of the agent products that we've released. And so we spent a lot of time, for example, building managed agents to have these considerations baked in. Yeah. Sure. Yeah. Okay. This is interesting. The question was about the best memory substrate, so why file systems versus, for example, databases. This is a subtle point that actually I want to think about carefully. So I don't necessarily think that it has to be the case that you use a file system for memory. I think what's quite important that we've seen is that you want something that is highly programmable with simple primitives that the model can manipulate to write and manage its own memory. So, for example, a database could work fine relative to the file system. But what I've seen doesn't work is when you specify the structure of memory for the model very explicitly, whether that's in a file system or database or whatever, like a memory schema: I pre-populate, here's the types of memories you need to save. So that ends up being not very bitter lesson filled in the sense that models can learn to manage your own memory much better than you can intuit these memory types for the model ahead of time. So I think what we've seen is that very general substrates for memory, be it just your database or file system, are good because the model can manage them freely, versus a very, very prescriptive memory schema that you're trying to pigeonhole the model into. That's when you see performance drop. That's the key differentiation. So you're seeing there needs to be that the model is going to be meaningless. Right. That's the key point. Let the model structure and maintain its own memory. Don't give it a prescribed memory schema. And that's a common failure because models are getting good enough that they can manage their own memory much more effectively. This is very classically bitter lesson filled. Models can reason about their own memory and context structure much better than you can prescribe for them a way to structure their own memories. That's the key observation. Yep. Yes. That's right. Exactly. General substrates for memory management. Yep. Yes. Okay. So that's a good point. The question is, so you do this dreaming thing. You look at the sessions, you look at the memory store, you update the memory store. How do you know those are correct? So evaluations obviously are one way to do it. This is a fun anecdotal example from Pokemon showing that you can perform corrections via dreaming. The key point is that we've actually run a lot of different evals showing that dreaming can indeed improve performance for very intuitive reasons, as you see here. But of course, evals are important in your own context to confirm it's actually worth the offline compute. Yep. I guess we're done. Thank you all. Thank you. proactive so typically with for example like local locally scoped agents they tend to be reactive they're responsive to how you steer it versus async agents increasingly have the ability to steer proactivity and that's one very nice thing about Claude Tagg where basically you can configure it to tell you things when like looking at this org level context alert me with things I might need to know about and this is a very important kind of new kind of UX that I think is going to be more and more common with async agents that kind of access to organizational context and of course multiplayer so the ability for a single harness to be steered by many many different people kind of concurrently is an important shift in agent UX that I think will be quite interesting going forward so um yeah let me let me just open up for questions and thank you for listening sure and and and and and and and and and and and and and and and ! and and and So the question was kind of on the gap between the frontier models kind of on, for example, a benchmark like Meter on like long horizon tasks. Yeah. Okay. So in the latest results that I saw from like, for example, Codex 5.6, I think it also kind of is in that 12 plus hour regime on Meter. So I think you're right that like the frontier models, like the, you know, Mythos class models, strong models from OpenAI kind of are in this like 12 plus hour regime. Why is it that non frontier models are not kind of in that regime? I am actually not necessarily sure. I do think, I do think in order to build agents that can effectively operate in this regime, it's important to know that it's not just the model capability. Like for example, with the Cloud Tag product, it's actually a combination of improvement in memory, because memory is very important. If you have agents working for, for example, 12 hours, you want to make sure that your productivity preferences are well encoded in memory. So it knows when to reach out to you if it gets stuck, for example. So memory is very important. Security is very important. So resistance to prompt injection. Also like model architecture, like kind of the agent architecture is very important that decoupling in brain and hand. So it's secure and safe and like resistant to failure. So actually, I think to build real agents that can operate in these long time horizons, a bunch of things need to come together in terms of like architecture, infrastructure, security, memory. That might be why, but Frontier Labs invested in all these areas. So that might be why you see kind of a gap in terms of like the agent products that we've released. And so we spent a lot of time, for example, building managed agents to kind of have these kind of considerations baked in. Yeah. Sure. Yeah. Okay. This is interesting. The question was about kind of like the best memory substrate. So like why file systems versus for example, databases. This is kind of a subtle point that actually I want to think about carefully. So I don't necessarily think that it has to be the case that you use a file system for memory. I think what's quite important that we've seen is that it's, you want something that is highly programmable with simple primitives that the model can manipulate to like write manage its own memory. So for example, a database could work fine relative to the file system. But what I've seen doesn't work is when you specify the structure of memory for the model very explicitly, whether that's in a file system or database or whatever, like a memory schema, I kind of pre-populate, here's the types of memories you need to save. So that ends up being not very bitter lesson filled in the sense that models can learn to manage your own memory much better than you can intuit these memory types for the model ahead of time. So I think what we've seen is that very general substrates for memory, be it just your database or file system are good because the model can manage them freely versus a very, very kind of like prescriptive memory schema that you're trying to pigeonhole the model into. That's when you see performance drop. That's the key differentiation. So you're seeing there needs to be that the model is going to be meaningless. Right. That's the key point. Let the model structure maintain its own memory. Don't give it a prescribed memory schema. And that's like a common failure because models are getting good enough that they can manage their own memory much more effectively if you can reason about types of, this is like very classically bitter lesson filled. But like you can models can read about their own memory and context structure much better than you can prescribe for them a way to structure their own memories. That's the key observation. Yep. Yes. That's right. Exactly. General substrates for memory management. Yep. Yes. Okay. So that's a good point. Basically the question is, so you do this dreaming thing. You look at the sessions, you look at the memory store, you update the memory store, how do you know those are correct? So evaluations obviously are one way to do it. This is kind of a fun anecdotal example from Pokemon showing that like you can perform corrections via dreaming. The key point is that we've actually run a lot of different evals showing that dreaming can indeed improve performance for very intuitive reasons as you see here. But of course, evals are important in like your own context to confirm it's actually worth the offline compute. Yep. I guess we're done. Thank you all. Thank you.