Open Reader

The Secrets of Claude's Agent Platform From the Team Who Built It

completed 43:21 May 08, 2026 Watch on YouTube

Current Status

completed

Video ID

lLypHkIVLqc

RAG / Chat

Enabled
The Secrets of Claude's Agent Platform From the Team Who Built It
Description

In the future, you’ll be able to accomplish a goal by just giving Claude an outcome and a budget. That’s the direction Anthropic is building in with its new Managed Agents features, announced at this week’s Code with Claude developer event. The basic idea: Claude, wrapped in a computer in the cloud, that you can spin up, scale, and manage as needed. Anthropic is taking on the infrastructure that kills most agent products, and making sure that it scales to meet the needs of agents running 24/7. On this week’s AI & I from @every, I talk with Angela Jiang (@angjiang), head of product for the Claude platform, and Katelyn Lesse (@katelyn_lesse), head of engineering for the Claude platform, about what Anthropic is building and what it takes to make agents reliable in production. If you found this episode interesting, please like, subscribe, comment, and share! To hear more from Dan Shipper: Subscribe to Every: https://every.to/subscribe Follow him on X: https://twitter.com/danshipper Timestamps: 00:01:48 - How the Claude platform evolved from API to agents 00:04:09 - The primitives that make up Claude Managed Agents 00:10:37 - Why the harness and the model are becoming a single unit 00:18:49 - The infrastructure wall that kills most agent projects in production 00:24:49 - Why team agents need a different shape than individual productivity tools 00:26:36 - How Anthropic's legal team uses an agent to review marketing copy 00:34:24 - Using multi-agent orchestration for advisor strategies, adversarial pairs, and swarms 00:35:50 - How to measure agent success with outcome and budget as the end state 00:39:11 - What the platform looks like a year from now, when Claude writes its own harness

Summary

Generated by claude-haiku-4-5-20251001

The Secrets of Claude's Agent Platform

Main Topics

  • Evolution of AI Platforms: From simple completion endpoints to stateful, autonomous agent systems
  • Claude Managed Agents: Anthropic's infrastructure for building production-ready agents
  • Platform Philosophy: Balancing opinionated design with flexibility
  • Model Lock-in vs. Optimization: Why generic harnesses are becoming obsolete
  • Internal Use Cases: Real-world applications of agents within organizations
  • Multi-Agent Architectures: Building complex systems with multiple specialized agents
  • Agent Lifecycle Management: Monitoring, updating, and retiring agents

Key Points

Platform Evolution

  • Then: Simple API endpoints for text completions
  • Now: Stateful platforms with memory, tools, file systems, and persistent sessions
  • Future: Outcome-focused systems where Claude handles architectural decisions

Claude Managed Agents Core Primitives

  • Messages API as the foundation
  • Built-in tools: Code execution (sandbox), web search, file systems
  • Skills and MCP servers for extensibility
  • Memory and state management for persistent agent behavior
  • Vault system for secure credential storage

The Harness Engineering Insight

  • Different models require different architectural approaches
  • Path dependencies early on lock models into specific capabilities (Claude excels at file systems, others may excel elsewhere)
  • "Harness + Model" pairing is becoming more important than generic hot-swappable models
  • Each harness choice significantly impacts model performance in evals

What's Actually Hard

Perceived difficulty: Harness engineering and prompt optimization

Actual difficulty: Production infrastructure (scaling, persistence, sandboxing, managing long-running async operations)

Successful Internal Agent Patterns

  • Coding agents: End-to-end development platforms (Stripe's "Minions", Ramp's similar systems)
  • Process automation: Legal review of marketing copy, multi-team workflows
  • Team-oriented systems: Agents that interface with each other across functions
  • Governance layers: Using agents to mediate between agents (agents all the way down)

Key Organizational Insight: "AI Software Factory"

Companies like Vercel are building entire operational systems around AI agents, creating multiplicative productivity gains across all processes—not just for themselves, but for every department.

Multi-Agent Orchestration Use Cases

  • Advisor strategy: Separating execution from advice-giving agents
  • Adversarial pairing: One agent generates, another critiques
  • Swarming: Multiple agents collaborate for bug hunting
  • Recursive decomposition: Breaking problems into sub-agents, then recombining
  • Best-of-N strategies: Running multiple approaches and selecting the best

Notable Quotes

> "As we make improvements to Claude and as it continues to get better and more autonomous, we find ourselves needing to evolve the platform to be higher and higher order abstraction."

> "The harness and the model get very paired. You still need redundancy and you still might want to use other models for things, but you probably do it at the layer of the agent, meaning the harness plus the model, rather than necessarily really generic harness and hot swapping everything underneath."

> "Everything should compress down to an outcome and a budget. That's probably about it."

> "Claude, make me a billion dollars. Your budget is $10."

> "If you don't have a human who's responsible for the agent, it gets stale very quickly and then it ends up being this dead thing that's out there doing stuff."

> "In that world, if Claude is on the fly or agents on the fly are becoming what they need to become in order for you to do what you're trying to do, the platform has to seriously scale."

> "You're not hopping straight down to the absolute core bit. Instead you're talking to Claude and Claude figures out what should be the right way for them to go handle it—managed agents all the way down."

Takeaways

For Developers Building Agents

  • Don't wait for perfect infrastructure—but also recognize when it's time to move off local setups
  • Start with reference implementations, but be ready to customize based on your specific needs
  • Build in layers: User interface → orchestration layer → core agent logic
  • Allow for self-service updates: Users can improve agents via Claude Code PRs with governance

For Organizations Deploying Agents

  • Assign clear ownership: Every agent needs a responsible human or team
  • Plan for obsolescence: Build systems to monitor and retire stale agents
  • Embrace multi-agent systems: Team-level automation requires multiple specialized agents working together
  • Use agents to govern agents: Reduce infrastructure risk through layered abstractions

Platform Direction (Next 12 Months)

  • Simplification to outcome-based APIs: Users specify outcome + budget, Claude handles the rest
  • Self-aware model selection: Claude determines which models and architectures to use
  • Reduced manual harness engineering: More abstraction, less configuration
  • Infrastructure scaling: Must support long-running, constantly-recreating autonomous agents
  • One-click integrations: Deploying agents to Slack and other platforms without manual setup

Philosophical Shift

The industry is moving away from "generic harness + hot-swappable models" toward "opinionated, optimized harness-model pairs" that deliver better outcomes at the cost of reduced interchangeability—similar to how specialized tools outperform generic ones.

Transcript

9020 words en Processed in 294.2s

A year from now, where do you think the platform will be? We'd want to experiment with directions where Claude actually gets so good at understanding itself. It figures out what model you should be using. It figures out how to spin up all the sub-agents. You don't have to think so much about what kind of architectures are there because Claude is actually able to understand itself enough that it can write itself on the fly. In that world, if Claude is on the fly or agents on the fly are becoming what they need to become in order for you to do what you're trying to do, the platform has to seriously scale. How close are we to Claude, make me a billion dollars? That's really what I'm asking. Angela, Caitlin, welcome to the show. [SPEAKER_03] Thanks for having us. [SPEAKER_03] Yeah, thank you. [SPEAKER_01] So for people who don't know, you both work on the platform at Anthropic. So Angela, you're the head of product for the cloud platform. And Caitlin, you are the head of engineering for the cloud platform. I'm really psyched to talk to you because, A, you've been launching a bunch of stuff. You have cloud managed agents that came out recently. You've been launching new features for it. And I think that it comes at this really interesting time where it makes me think about what actually is a platform in AI for a model company. Because in the GPT-3 days, the platform was a completion endpoint. You just send a prompt to get a response. After that, it was a completion endpoint with tool calling and chat sessions, that kind of stuff. And now, with cloud managed agents, you're essentially getting a cloud on a computer with memory and all this other stuff. So I'd love for you to help me unpack that trajectory and what it means to build a platform in AI. [SPEAKER_01] Yeah. I think your characterization is very accurate. I think as a lot of these technologies have evolved with the LLM, first starting. And then I think putting that behind an API was very fun. A lot of people were like, wow, I could do some at the time. I think it was very cool. Now we'll probably look back and be like, oh, that was really basic. And then, you know, I think we've moved more and more towards a slightly more stateful world as you want to persist the sessions state to be able to make sure that the performance of the model is better and better. I think that's probably the through line. As we make improvements to Claude and as it continues to get better and more autonomous, we find ourselves needing to evolve the platform to be higher and higher order abstraction. But it's in the pursuit of helping you get the best outcomes out of something. I think in the very beginning, everyone was very exploratory. It's like you have no idea what people are going to build with these LLMs and you wanted to have as much possibility out there as available. And then as those use cases started to narrow down, people started building products with it, people started building agents with it. And more and more of that is about customers coming to us and being like, how do I get the best out of Claude? How do I set up my tools? How do I run the loop? And so on and so forth. And you have some people who are really experimenting and they're on the edges and that's great. And then you have a whole host of other folks that are coming in who are like, I want a lot of this stuff out of the box. And in our pursuit for making sure that Claude is producing the best outcomes, we find ourselves enriching the platform to be richer and richer. And that's contained in that is the state, the tools that you start to see us adding. It contains a lot of the cloud components of a lot of these types of things. But it's in pursuit of the same mission of making things literally as easy as possible. And I think in the forward state of a lot of these things in terms of the philosophy of what a platform ultimately ends up doing, it probably ends up just being the set of primitives and infrastructure that enables you to get the outcome as fast as possible with as little work as possible. And I think that tends to follow a certain form factor, at least in this current state. But yeah. How would you characterize what the primitives are today? So maybe that's just asking, what are the primitives in Claude Managed Agents? Yeah, so Claude Managed Agents is built on all of our same primitives that you could otherwise build on directly. So the Messages API. And within the Messages API, we've built a whole bunch of innovations around the API. Like you could just get tokens in and out if you really wanted to. But you can use some of our built-in tools. You can build, you can use stuff like code execution, spawn a sandbox and execute work. You can use web search and all these different things. And so I think we've taken what we see as all the most powerful of those things and put them together into a harness and a set of infrastructure that is the way to get what we think is the best outcomes out of Claude. So I'm sitting here feeling this sense of time deflation. Like my time gets more valuable in the future as opposed to the opposite. Whatever the opposite would be, my time gets less valuable in the future. And the reason is because we're building an agent. We're building some agent products where it's agents that do specific things for us internally and then hopefully for customers. And in order to do that, we have a couple Mac minis with Claude running in a loop on the Mac mini, right? And a lot of that is a thousand line Python file or whatever, and a lot of that mirrors what you guys are building in Claude managed agents. And so for me, and I think for a lot of people building on Claude or on the Claude platform or ecosystem, there's at least this feeling, maybe we should just wait for you guys to build it. But then I don't know what the lines are. And yeah, I'm wondering if I want to build an agent, what is the best path to do that in a way that aligns with what you guys are doing? [SPEAKER_02] Yeah. I think this part of the platform business is actually [SPEAKER_02] and a lot of that mirrors what you guys are building in Claude managed agents. And so for me, and I think for a lot of people building on Claude or on the Claude platform or ecosystem, there's at least I feel this, maybe we should just wait for you guys to build it. But then I don't know what the lines are. And I'm wondering if I want to build an agent, what is the best path to do that in a way that aligns with what you guys are doing? [SPEAKER_02] I think this part of the platform business is actually somewhat similar to any other form of platform business where you do have customers like yourself who are building and you're thinking, should I go ahead and do it because maybe I have this immediate need, but at the same time, it will kind of want to repeat the work. And you could have just gotten it for free out of the platform. And also infrastructure sucks. It's so much to spin up servers. I can't believe you do that all the time. [SPEAKER_02] That's the worst. [SPEAKER_02] But I will actually say, part of why we ended up building cloud managed agents was because Anthropic ourselves had gone through enough of these iterations where we built products that were agents that you could run autonomously in the cloud. And we did stand up the infrastructure so that it works well enough times that we ourselves were like, okay, we're done building this for ourselves. We're doing it once in a way that's going to really work from everything that we've learned, but also for all the people who are doing it. [SPEAKER_01] You can run whatever you're running on a couple of Mac minis maybe. Right. And for a lot of people that could work. But I think if you're building agents into your product and you're running something really at scale, that's where it really starts to become more and more challenging to get that infrastructure. That's really interesting. Yeah. And then maybe to answer the other part of your question, I think we have two pieces of the philosophy here. One is in the way that we design managed agents, which is that we try to have it be modular enough. We want to be opinionated about some pieces that we feel should be very well married to the cloud model. But then we like the way we want, for example, we want cloud to very specifically use file systems. That's a very particular cloud kind of style. In a specific way or just file systems in general? [SPEAKER_03] Just file systems in general. We also really want to lean into skills. I know a lot of folks like skills, but that's something that we want to have our hardest be really opinionated about that. And so we're particular about those kind of primitives being the case. So use the file systems, use the skills. They're really basic. But at the same time, we still find people who are trying other methodologies to do that. And we want to help you when you build, just start on the best foot. So that's one piece on some of the more opinionated ones, but as each one of these endpoints or APIs that we have as part of the suite, we try to open them up a little bit in certain areas. So there are things that we're looking forward to and being, you know, from maybe it's not available today, but in our design, we are trying to make it flexible enough for people to add in different pieces because we recognize that this API or suite of APIs is not necessarily going to solve everything in its original construct. And there are going to be pieces that need to open up. And then the second bit is we're public about this is when we do design a lot of these things, we do put out blog posts and reference implementations. So if you did want to at least be inspired by that construct, but still maybe make your own on the messages API, you can definitely do that. I think that's to the point you just made, that's something that's coming up for us. Again, we have Claude running on a Mac Mini with a Python file and a couple other bigger, more serious implementations on cloud infrastructure that we're trying to figure out what to do with. And I think I told the team that we were talking today. And I think one of the questions that they have, or one of the feelings of consternation that they have considering using cloud managed agents for this kind of thing, for spinning up agents for our customers, is just right now it's like we have a playground, we just have a server or a Mac Mini. We can just pipe stuff to Claude. It can do anything that Claude code can do. It has a file system. It has a browser. It has all this stuff. If we want to switch it out to GPT 5.5 or Gemini or whatever, it's pretty easy to do that. [SPEAKER_02] So is that the kind of thing, and I feel like they feel like we're going to get locked in and it's not going to have the flexibility to do all the stuff that we want. And there's also a worry that features are going to come to Claude code itself that won't be in Claude managed agent for a little while. And that it'll prevent us from being at the edge, which is what we promise to our customers and really to ourselves. We just love being like, doing whatever the new thing is. How do you think about that? Yeah. So I think what's nice about the way that we work internally is we run the platform and the platform for what most people think of it as is our externally facing APIs and our suite of APIs. The other rest of what our team actually does is internal platform in the sense that all of our first party products are built directly on the same platform as everybody else. And so what's cool about that is we spend a lot of our time working with the teams internally who are building on top of the platform and enabling the features that they will build and sharing ideas and these sorts of things. And so Yeah. So I think what's nice about the way that we work internally is so we run the platform and the platform for what most people think of it as is our externally facing APIs and our suite of APIs. The other rest of what our team actually does is internal platform in the sense that all of our first party products are built directly on the same platform as everybody else. And so what's cool about that is we spend a lot of our time working with the teams internally who are building on top of the platform and enabling the features that they will build sharing ideas and these sorts of things. And so I think over time, you'll maybe see less and less divergence of what might be available in Claude managed agents, what might be available in coworker Claude code that might sit on top of the same infrastructure, right? Like that's, I think one way to think about that. [SPEAKER_03] Yeah. And then I think on your point around your team's point around having some kind of model lock in fear, I think that's valid, like many folks have that consternation. And I think we're at this place where there's an evolution here where if you look back, maybe even just a couple of months ago, it was very standard to build a very generic harness. It's super generic. And then you can hot swap models across all of those things. And I think for an older generation of models across labs that worked, okay, a lot of things were moving at a pace where I think that was mildly reasonable. I think now, for the next generation of models, and as we see it forward, I think you see this a little bit from every lab, like everyone's taking slightly different techniques and perspectives on how they want to advance their particular form of the model. And so in theory, you could do a superset of all those things, but more often than not, I think when you build agents for your company or for your customers, you do want to deliver an outcome ultimately for them. And so I think that level of abstraction of what you're actually hot swapping stops being this really generic harness and hot swapping the model. And it gets more to the harness and the model get very paired. You still need redundancy and you still might want to use other models for things, but you probably do it at the layer of the agent, meaning the harness plus the model, rather than necessarily the other architecture of really generic harness and hot swapping everything underneath. [SPEAKER_02] That's really interesting. Is that how I don't know the cursors of the world are doing things like, do they have a separate harness for each model or is it a generic harness that they're kind of hot swapping the models in and out of? Do you know? Um, I'm not entirely sure. My intuition would be that I don't know about cursor in particular, but there have been teams that we have talked to who have kind of fallen on similar perspectives. And it's mostly because they're just trying to squeeze the most out of each model, almost like harness engineer every single nuance. And one example that we have, it's not an external customer per se, but something that we've done a lot internally, like we recently launched memory, for example, with managed agents. And we tried a bunch of different harnesses ourselves. Like we tried one that was the one that we ended up launching. We tried a bunch of others using a bunch of different other techniques. And at least personally for myself, when I saw the eval suite from the team, each one of these harnesses performed drastically differently. And so I think just even looking at something like that shows you that you can actually hill climb a tremendous amount by just harness engineering the right pieces together. And I think if you were to take that forward across all model combinations, across all different labs, all different kinds of providers, there is a lot of alpha in that construct. And so I wouldn't be surprised if more than just ourselves have experimented with that level of unit tying. [SPEAKER_02] It's really interesting that there's this path dependence where you make some choice for how you do requests and responses or how you do tool calls or whether you have the model want to use file systems or not. And then that changes the trajectory of all these different models. Yeah. And it feels like maybe at the time, such a small, almost a footnote. But it ends up becoming very big. [SPEAKER_02] Do you think that will end up affecting the model's generalizability in the sense that at some point, they'll just have these locked in lanes of stuff that they're good at because Claude is really good at file systems and GPT is good at some other things. How is that going to flow through the model's personality and behavior if it's locked into a specific way of doing things? I do think it does actually tend to lock the model. So what we end up treating as the right path and the right primitives need to be very carefully thought through. And so I think in some eras, with other models, they become really good at reasoning and then they almost over optimize on that level of reasoning. And there's other perspectives around, okay, yes, we wanted to be really good at the computer part, maybe the computer part is the interesting part. And so if you think through maybe some of the primitives, which we could get right, we could get wrong, but at least we'll go through the thought process of that, that will probably at least lead us one path or the other. I think it's hard to say in which direction per se will ultimately be true, but I do think there's a lot of path dependency it ends up taking. So being really thoughtful about what you choose to actually include or give the model more natively is really important. [SPEAKER_02] Are there any of those path dependencies that you've had to undo? Hmm. I can't speak enough about that at the anthropic level. I've only been here a couple of months, but I have to imagine that has been the case. I mean, we've experimented like even at other labs, the primitives that we have In which direction per se will ultimately be true, but I do think there's a lot of path dependency it ends up taking. So being really thoughtful about what you choose to actually include or give the model more natively is really important. Are there any of those path dependencies that you've had to undo? Hmm. Probably, I can't speak enough about that at the anthropic level. I've only been here a couple of months, but I have to imagine that that has been the case. I mean, we've experimented even at other labs. The primitives that we have to take a look at are constantly changing. And you do hit a little local maxima and rethink, okay, maybe there's a more generic approach. Yeah. Interesting. I want to take a step back and ask you something that maybe I should have asked at the beginning, which is, who is cloud managed agents for? Right? I set one up earlier today. We've got some people already using it in production inside of every, and I just did one today. I really loved the getting started chat experience that you had and the examples that you had. And it felt to me like, even if I was not technical, I might want to use this to set up an agent. It might be a little bit complicated, but what I actually did is I just, and I'm sorry to say this, but I did it in the Codex in app browser. So I had Codex driving the managed agent setup and it like, I had a Slack bot working pretty quickly. It was really cool. So how do you think about when you're designing stuff, when you're designing cloud managed agents, who it's for? Yeah. So it's interesting because I think you're right that, especially with that quick start experience, which we actually felt pretty strongly about launching, not specifically for the sake of making it so that non-technical people could go and build agents, but actually just for anybody technical or not to be able to wrap their head around the primitives, the APIs. Here's what I can do and here's what fits together. Yeah, exactly. The education portion of it. But I think when we think about who it's for, we think about a couple of different things. One is we're seeing people internally within companies build automation or build really powerful platforms or systems. Like we've seen people say, I want a full end to end software development platform, right. And managed agent is the perfect solution for something like that. Or I want to automate a process over here where legal has to review my marketing copy. Right. And things like that. And so you shouldn't have to re-implement memory and all that stuff every time you're doing that, right. You can get started really quickly and you can get something running quickly. The other user that's top of mind for us is people building into their products that they expose to their customers. And so that's the other one where actually, yes, you do still want a lot of customization. You do still want to make something that's going to be really powerful for your product, but we still definitely believe that not spending your engineering resources on the infrastructure and on all the little harness engineering tweaking is worthwhile. Why couldn't we have talked a month ago? You would have saved us so much time. We'll just need to talk more. Yeah. But I am curious. Okay. So maybe infrastructure is one of these things, but when you see people setting up agents, what do they think the hard thing is and what ends up actually being the hard thing? And are they the same? Good question. I don't know, spicy. I'm not sure. But I think people think the harness engineering part is the hard part. And so actually, in the past we launched the agent SDK, which is what you guys I think are using on your Mac minis. And for a lot of people, they were like, okay, great. I don't have to do the harness engineering part where I have to do prompt caching and I have to maximize my context window and all these sorts of things. I think we're just actually using just Claude in back like the Claude dash P command. Oh wow. Yeah. Okay. It's pretty good. Yes. Yeah. Okay. Cool. And regardless, you guys did that because it takes off your hands building the harness. Right. But I do think what we saw with a lot of customers was okay. Now I want to go and take that thing and get it into production and scale it. And everybody hits an infrastructure wall. Like everyone hits the same problem of, oh wow, I either need to keep a server constantly running or I need to use infrastructure that will spin up and spin down and I need to store the transcript data and I need secure sandboxing and all these sorts of things. And so, you know, if you boot a Claude code session or you boot the agent SDK in a sandbox and that's the thing that you have running, but your sandbox loses connection and dies or whatever, your whole agent dies. Right. And so I think the infrastructure part especially is the wall that most people end up hitting, but they're more expecting that the actual harness engineering and getting the most out of the model is the part that's going to be harder. Yeah, I totally agree with that. I was just going to say, we talk to so many people who are now at a place where they're prototyping really quickly and they're super excited and it's doing the thing. And then yeah, there's a class of people who are really pushing and being like, okay, I do want to hill climb. I really want to edit the harness. But then once you have that thing, productionizing is just a nightmare. Especially for the more interesting kind of long running async ones that you want to do a bit more remotely that are a bit more autonomous. And everyone kind of runs into that wall and was a big inspiration for why we built what we built. I feel like one of the examples of the shape of an agent is OpenClaw. And in particular the thing that it has brought to us internally is you have an always on agent in Slack that has its own personality and it has its own part of the world that it ends up working on. Are you guys like, is that a possible future for, okay, a one click agent that lives in my Slack that yes, I can go set up all the internals, but I don't have to really think about and everyone runs into that wall and was a big inspiration for why we built what we built. I feel like one of the examples of the shape of an agent is OpenClaw. And in particular the thing that it has brought to us internally is you have an always on agent in Slack that has its own personality and it has its own part of the world that it ends up working on. [SPEAKER_01] Are you guys like, is that a possible future for a one click agent that lives in my Slack that yes, I can go set up all the internals, but I don't have to really think about all of the technical infrastructure stuff? Because I think you all have the beginnings of that, but it's still a lot of steps from the current managed agent to something that's always on in my Slack that I have to set up and customize. So is that something that falls in the realm of platforms job or is it too far in the product direction? No, it definitely is something that we really want to do. I think we focused a lot on the infrastructure piece to start because that's where we just see a lot of these pain points. But yes, in its advanced shape, we actually want to make it so that you can deploy these agents really easily. We've made some light steps in this direction. For example, we included vaults as one of the primitives. [SPEAKER_01] And vaults store your keys and stuff like your OAuth keys? Credentials. Credentials. Yeah. As a way of solving some of the lower level pieces as a starting point. But once you wrap some of these agent identity type of primitives in a more secure way and you can handle it really easily and it works with the whole system, then I think it's very natural for us to get to a place where maybe you are either one clicking a Slack integration or alternatively, even just telling Claude to add Slack and it just handles absolutely everything. And then before you know it, your little bot is just picking up in Slack. I love it. I can't wait for that world. What are the best internal use cases of agents? Because I think there's this big question happening right now where everyone's in Codex or Claude Code, but then now we have these agents that are out in the cloud. Now everyone inside of a company can have their own agent. There are team agents that are company-wide agents. So what are the patterns that you see for when people make really useful internal agents, what they do and what they look like? Yeah, I would say we've actually seen a few examples of these in some of the more AI-focused, AGI-focused companies like Stripe built minions and they talked about that as their end to end development platform that their engineers could use. I think Ramp did something similar and we've done similar things as well. [SPEAKER_01] That's interesting. Yeah, we've built platforms internally that are, you know, agents running that I can talk to from Slack or from wherever. And at a certain point that becomes actually a pretty thin layer on top of managed agents. You don't have to do very much to accomplish it. [SPEAKER_01] That's what I was thinking. I looked at minions or whatever Ramp does and I was like, why? So is it actually useful to have a thin coding agent that anyone in the company can use or why not just install the Claude app in Slack? [SPEAKER_01] Yeah, I would say the difference in a platform like that and some of the things that we've done internally is there's a lot of customization that you might want to do on the development environment where an agent is actually running and able to verify its changes. Things like that. It's like, here's how our CI/CD works. Exactly. And so for lots and lots of people Claude Code is an excellent tool, and you can run cloud agents with Claude Code and that is really great. But I think if you're trying to do a bit more end to end development and you maybe want to bake in more custom things, you could start with something like managed agents and build a layer on top of that and end up with something that's closer to that end to end experience. It also seems to me like there's something in particular about having a team that you need to work with that makes the managed agent shape important as opposed to it just working in Claude Code. I guess technically you could sync the skills between everyone's Claude Code, but there's something about just having one agent that does this thing that seems to work. Yeah, I'm really glad you brought that up because I think that's actually one of the more common areas where we see a lot of opportunity. To your point, there's a lot of individual productivity happening, whether you're a developer or non-developer, there are so many tools that you're using to make yourself more automated and more high leverage. But then when you get to the team layer, suddenly everything gets massively more complex. Number one, obviously you can't sit on your laptop. And yes, you could put it in the cloud, but that's again more for you to handle with your laptop closed. But then you go to, okay, the three of us want a couple agents that interface with each other and work with each other. And then maybe we're automating a process end to end. And especially for some of the more complex processes that you envision being really transformed with AI, you do need that team orientation. And that needs to happen at a layer that's a slightly higher abstraction than just a single agent. And I think teams exploring multi-agent architectures are really exciting, but it needs to be built on top of a platform that everyone can spin up and down and control. And I think G from Vercel had a really good perspective on this in a way where I think his company Vercel is obviously incredibly AI-focused. And he describes it as an AI software factory internally. And I think that's exactly the right mindset. And that produces extremely high leverage organization that's really creating a tremendous amount of productivity. Teams exploring multi-agent architectures are really exciting, but it needs to be built on top of a platform that everyone can spin up and down and control. And I think G from Vercel had a really good perspective on this. His company Vercel is obviously incredibly AI-focused. And he describes it as an AI software factory internally. I think that's exactly the right mindset. And that produces extremely high leverage organization that's creating a tremendous amount of productivity, but not just for themselves, just for every single process that they have in the company. [SPEAKER_01] Mm-hmm. And I really want to go back to this: okay, agent use cases. We've got coding agents that anyone can use in the company. What are the other ones that you see people standing up that are really useful? We've seen a few. One of the fun things that we get to do is work with our internal teams of different functions and help them agentify because we actually get to learn a lot as a result of doing that. The silly example I brought up earlier of the legal team needing to review marketing copy was one of them. [SPEAKER_01] Very real. Yeah. Like extremely real. Really blew people's minds with very basic agents that just give people the right setup to be able to do that. So we've seen that. What does that actually do? So there's marketing copy and there's a legal agent that is watching what everything marketing does? No, no. It's more like: I'm a marketer and I've written some copy, right? In the past, maybe you would have opened a ticket or something and asked, can you please review this copy? But instead you submit it to this app that we built on top of agents. Okay, cool. Now I'm going to go as an agent review first and then put it in legal's inbox with the first pass review already done. Maybe the agent can say, okay, marketing, you're good. Or maybe it says, no, this needs an extra human review. And that's the sort of thing where again, just a thin layer on top. But you can build it, I have access, we can both see the outputs and we can work together on it. [SPEAKER_01] Okay. But then, for example, why is that not a skill? [SPEAKER_03] It very much can be a skill. And that actually is something you would probably build as a legal reviewer agent, right? You would have MCP servers or whatever it is that help you access external contacts. You would have skills that help you understand the rules we have to follow and not follow, right. And you'd put all those things together, but then you can just fire off a session with that agent. [SPEAKER_01] Yeah. And then I think the last piece you need, and this is where I'm saying it's a really thin layer, is the form factor on top where different people can collaborate together and work with that agent and multiple agents can be involved in the system. [SPEAKER_03] And so I think it goes a little bit broader than a skill because you kind of still need the right form factor for the agent to go run and then for people to be able to interact with it. Another core bit on why it's not a skill, or not exclusively a skill, is because you actually do need human in the loop. If you were to automate the whole thing and you were just taking the skill and looking at yourself from like legal skill, for example, in that world, of course you could have just done a pure skill. But if you need a human in the loop to review and check, and we're looking at legal things, there's a bit of authentication that's necessary. In order to automate that entire process, you kind of need agents to go do the thing. And so because you need to spin up separate sessions for that to happen, some sort of stitching is necessary that can't be instantiated in a single skill. [SPEAKER_01] That's really interesting. Okay. So just to push on that a little bit: you create an agent whose job is to make sure that when marketing is writing something, they can get it approved really quickly by legal. Sometimes it'll approve things immediately. Sometimes it sends stuff to legal. And ideally it's getting better all the time so it can do more and more, right? What is the best practice for who owns that agent once it's built? Because one of the things that we found is if you don't have a human who's responsible for the agent, it gets stale very quickly. And then it ends up being this dead thing that's out there doing stuff, but it's not necessarily good. And also, even if it works, there are going to be all these times where legal's like, you asked me to approve this, but I don't really need to approve this thing. Let's update your prompt. So how does that all work when it works well? [SPEAKER_03] It's actually really interesting because the form factor thing, the app that sits on top of that, that we originally built, one of our teams worked on that, right. And sitting with these teams and understanding what they needed. We were kind of like, okay, here you go. And we're going to go do other stuff now and let us know how this goes for you. [SPEAKER_01] And then a really cool thing actually ended up happening where people on those teams who were using the tool were like, oh, I wish this little thing could get tweaked or this thing could get better. And they popped open cloud code and made some of the changes themselves. Yeah. And so it's funny. [SPEAKER_01] And then is your team responsible for approving the PR or does it just go in? [SPEAKER_03] Usually my team's responsible for reviewing the PR if it's a system that we actually own. But people can kind of self-serve making changes to those things, which I think is really cool. So I do think we're still in a stage for a lot of teams, a lot of companies, and even going back to Stripe having minions, right? Stripe has a large developer productivity team. We used to work at Stripe, so we spend a lot of time with them. They have a large developer productivity team. They're awesome. And they're obviously putting a lot into [SPEAKER_01] Usually my team's responsible for reviewing the PR if it's a system that we actually own, but yeah, people can self-serve making changes to those things, which I think is really cool. So it is, I do think we're still in a stage for a lot of teams, a lot of companies. Even going back to Stripe has minions, right? Stripe has a large developer productivity team. We used to work at Stripe, so we spend a lot of time with them, but they have a large developer productivity team. They're awesome. And they're obviously putting a lot of work and energy into building platforms and tools like this. And so I think we're definitely still in a place where something like managed agents or being able to build on top of our platform is really powerful, but you still need the AI-pilled people and technical people within a business to then go create something really excellent on top of that, that works well for whatever you're trying to do. [SPEAKER_03] That's interesting. Yeah. I love the anyone can open a PR to do this because everyone's using cloud code. One of the things that I find talking to people who are in infrastructure roles at companies where this is starting to happen is the meme where it's a person and he's going like this and he has daggers in his back and he's covering it. It's like infrastructure people are that for now anyone can submit PRs. How do you deal with that and how do you do that? Well, obviously in an ideal world, you would love for legal to be able to submit PRs to improve this agent. And also sometimes they're probably going to submit stupid stuff that wastes time. And so what are the right ways to either organizationally, culturally or technically, make that possible without ruining your lives? [SPEAKER_01] For this particular one that we've constructed that Caitlin's given as an example, we actually have a couple layers of abstraction away from that PR layer. So at the very beginning, it started that way and to prevent users from foot gunning themselves, they get to a place where oftentimes their way of interacting with the agent that they own, whether it's the marketing team who owns the marketing agent requesting, or if it's the legal team owning the agent that does the review. Um, they actually engage with those agents through Claude itself. So they actually spend more of their time talking directly to Claude and then Claude will figure out what should be the right way for them to go and handle it so that they're not hopping straight down to the absolute core bit and doing something that may result in complications. And they're talking to Claude or Claude code, Claude chat or Claude code or co-work. It's a different instantiation of Claude that we made that actually is a managed agent in and of itself. So it's managed agents all the way down in that construct. But we found that each layer, if we tune and prompt each variant of the managed agent, it helps to solve different parts of the problem for users. So at the end state for that marketing person or that legal person, it is a really simple interface where the way that we tell them is you're just talking to Claude, but under the hood, it's many Claude engaging with each other to get to the part where then the Claude's themselves are doing the more complex work that the human doesn't really need to interpret. [SPEAKER_03] Interesting. You guys just launched multi-agent orchestration. What are the coolest things that people are doing with that? Hmm. One of the more interesting ones is people are using it to construct different harness techniques. And that one I'm personally very excited by. Because there are different techniques that people have experimented with where, for example, we recently did the advisor strategy one, but really if you were to genericize it, you just separate execution from advice. And there's also one where you can have two modes where one is generating something, and the other one's adversarial to it. And then there could also be you split it into a bunch of different little tiny pieces and then they recombine. And then there's ones where it's something closer to best of N kind of style. And then there's so many more. And in each one of these different types of architectures or strategies, they are good for a very specific use cases. Some of them are much better for deep research or wide research type of use cases. Right. And there are others that are the ones where they all sort of swarm together are better for bug hunting, for example. Um, and so that's really cool to see that if we can make the primitives very Lego-like, then people can put them together to solve things at a slightly higher form factor, which is more like an architecture or strategy. Um, and they get much more interesting results out of that. Um, and that's really exciting to see because it also suggests that you can actually hill climb at multiple layers of abstraction. How do you know if an agent is successful? How do you measure success for an agent? [SPEAKER_01] Yeah. I mean, there's evals and stuff like that, which everyone has talked about ad nauseam. One direction that we really like is this kind of verifiable outcome. Um, we've been somewhat opinionated on that one and it's almost like in the absolute end state, we talked a little bit about what's a platform at the end of things. Um, going from that philosophy, it's like our kind of principle that maybe the end state of some of these things is that everything should compress down to an outcome and a budget. And that's probably about it. Um, and everything else should be figured out for you to resolve exactly across those parameters. And so for us, we still have evals. We have a lot of these other things that we measure that are domain specific. Like coding evals would be measuring like the actual PR getting merged. Those are more verifiable. But as we get to the place where an outcome is actually a spec that you as a human are able to define and our ability to interpret that and regrade itself over and over is closer to what we care about. Claude, make me a billion dollars. Your budget is $10. [SPEAKER_01] resolve exactly across those parameters. And so for us, yes, we still have evals. We have a lot of these other things that we measure that are domain specific, like some coding evals would be like, you might want to measure the actual PR getting merged. Those are more verifiable. But as we get to the place where an outcome is actually a spec that you are able to define as a human and our ability to interpret that and regrade itself over and over is closer to what we care about. Claude, make me a billion dollars. Your budget is $10. Exactly. I meant to say no mistakes. Go. Go. Exactly. Maybe Mythos could do that. Yeah. And then one of the things that we were running into that I'm curious if you have a solution for is agents get outdated pretty quickly. Sometimes because there's no human attached to them. Sometimes they're just running an old model or there's an old architecture or whatever. And it feels like there needs to be an end of life cycle for agents. Like we've talked about having a little funeral for them and having a little page on our website that's like, here's all the decommissioned agents and stuff. How do you manage, especially in a really big company, how do you manage all the agents that are out there, and maybe they're in Slack, pinging stuff once a week, but you're like, this is super stale. How do you make sure that you retire them as quickly as you are making them? [SPEAKER_01] So one of the things we have actually done is we have made skills that help you do things like upgrade to a new model when a new model comes out, right? Like we've actually put a good amount of work into making it easier to do exactly what you're talking about. And I think maybe some of the most AGI pill people are running agents that are monitoring their agents to see if their agents are outdated and in need of that sort of stuff. But I think for the way that we like to talk to customers who ask us this question, I do think the most interesting instantiation of this is there's a new model and now I need to go upgrade my agents or maybe be done with those agents because the new model enables me to build agents that are way more powerful and do more interesting things than the old agents did, right. But I think that upgrade process and that migration process is something people have had to wrap their heads around as a breaking change and I have to put actual energy into making that work. And obviously, sorry to talk about evals, but if you have evals, this process is easier and things like this. But I do think that's one of the things we've tried to do is how do we give you skills and how do we give you the right tools to make that process easier. And then you could go be AGI pill and choose to actually automate more of that with more agents. Yeah. [SPEAKER_01] So a year from now, we're back at Code with Claude. Where do you think the platform will be? What will I be able to do? And how will it be different from what I can do today? [SPEAKER_01] Do you want to go first? You can go first. A year is a long time. It's in this industry, especially how close are we to Claude make me a billion dollars? That's really what I'm asking. [SPEAKER_01] I think, yes. If that works, we probably won't be sitting here. Yes, yes. We'll be asking Claude for this. I mean, yeah, like we want to get closer and closer to that state where I think we kind of, okay, so a couple of things. I think in a year from now, one thing that we'd love to get really, really close to is actually that kind of simplicity. And this might be a significantly higher order of abstraction. I don't know what the form factor will look like or whatever, but the kind of parameters we will care for from users will be that outcome. And of course, it has to be verifiable. There are some parameters that have to be restrictive and the budget. And I think we'd want to experiment with directions where Claude actually gets so good at understanding itself. It figures out what model you should be using. It figures out how to spin up all the sub agents. I actually don't think you need to think so much about harness engineering in that world. Today, you don't have to think so much more aggressively about tool construction, for example, that we've made that a little easier and you get it to leave a little bit of that scaffolding. Less prompt engineering. Yeah, exactly. Exactly. [SPEAKER_01] And I think if you just keep going up that stack, like today, a lot of the innovation is happening at this kind of really high level, almost like harness architecture level, which is really fun. But I think a lot of that honestly also goes away where you almost don't have to think so much about model selection. You don't have to think so much about what kind of architectures are there because we probably would have gone through enough iterations with Claude where Claude is actually able to understand itself enough that it can almost write itself on the fly to figure out what is necessary in that kind of two parameter world of outcome and budget. I don't know that we'll get there in a year, but I feel like we might be able to do the outcome part of that with maybe some error bars on the budget side. Really cool. Yeah. Okay. That was really cool. I'm gonna give you a slightly more boring answer, which is in that world, if Claude is on the fly or agents on the fly are becoming what they need to become in order for you to do what you're trying to do, the platform has to seriously scale. That is. And so I do think some of this will be, what are the right abstractions that actually enable that, right? Like somewhere on the primitive to higher order realm, right? But I do think so much of what our team is going to be doing is making sure that the tokens that people want to come in and out of Claude are going to be able to come in and out of Claude because our system is scaled to meet not just the demand, but in that world where it's just like you have agents that are literally constantly running and recreating themselves and doing this sort of work. You just need a system that can handle long running requests, can handle a bunch of differently shaped things. And so I think for us, it's going to be, I never want the ability of the platform itself to be able to scale to get in the way of what people would otherwise be able to accomplish. [SPEAKER_03] I think so much of what our team is going to be doing is making sure that the tokens that people want to come in and out of Claude are going to be able to come in and out of Claude because our system is scaled to meet not just the demand, but in that world where you have agents that are literally constantly running and recreating themselves and doing this work. You just need a system that can handle long running requests, can handle a bunch of differently shaped things. And so I think for us, I never want the ability of the platform itself to be able to scale to get in the way of what people would otherwise be able to accomplish with these things. And so I think that's something that's going to probably be very front of mind when we're talking in a year. [SPEAKER_01] Awesome. I'm excited. Thank you so much for joining. I really learned a lot. [SPEAKER_03] Thanks for having us. [SPEAKER_01] Oh my gosh, folks, you absolutely positively have to smash that like button and subscribe to AI and I. Why? Because this show is the epitome of awesomeness. It's like finding a treasure chest in your backyard, but instead of gold, it's filled with pure unadulterated knowledge bombs about ChatGPT. [SPEAKER_02] Every episode is a roller coaster of emotions, insights, and laughter that will leave you on the edge of your seat craving for more. It's not just a show. It's a journey into the future with Dan Shipper as the captain of the spaceship. So do yourself a favor, hit like, smash subscribe, and strap in for the ride of your life. And now without any further ado, let me just say, Dan, I'm absolutely hopelessly in love with you. you're running on a couple of Mac minis maybe. Right. And for a lot of people that could work. But I think if you're building agents into your product and you're running something really at scale, right, like that's where it really starts to become more and more challenging to get that infrastructure. Right. That's really interesting. Yeah. And then maybe to answer the other part of your question, I think we have like two pieces of the philosophy here. One is, is a bit in the way that we kind of design managed agents, which is that we try to have it be modular enough. Like we want to be opinionated about some pieces that we feel like should be, you know, very well like married to the cloud model. Um, but then we, uh, like oftentimes like the way we want, for example, we want cloud to like very specifically use like file systems. Um, that's like a very particular like cloud kind of style. In a specific way or just file systems in general? Just file systems in general. We also really want to lean into skills. I know like a lot of folks like skills, but like that's something that we like, we want to have our hardest be really opinionated about that. And so we're kind of particular about like those kind of primitives being the case. So like use the file systems, use the skills. They're really basic. Um, but at the same time, like we still find people who are like still trying other methodologies to go do that. And we want to kind of like help you, you know, when you build to start, uh, just kind of starting the best foot. Um, so that's one piece, uh, on some of the kind of more opinionated ones, but as each one of these kind of like, you know, endpoints or, or, um, APIs that we have as part of the suite, we try to like open them up a little bit, um, in certain areas. So there's like things that, you know, we're looking, uh, kind of forward to and being like, you know, from maybe it's not available today, but in our design, um, we are trying to make it flexible enough for people to kind of like add in different pieces because we recognize that this API or suite of APIs is not necessarily going to solve like maybe everything in its original construct. And there are going to be pieces that need to kind of open up. Um, and then the second bit is like, you know, we're, we're kind of public about this is like when we do design a lot of these things, we do put out like blog posts and sort of like reference implementation. So if you did want to kind of at least be inspired by that construct, but still maybe make your own on the messages API, you can definitely do that. I think that's to the, to the point you just made, that's something that's, that's coming up for us. Again, we have, you know, Claude's running on a Mac Mini with a Python file and a couple other, like, uh, you know, bigger, more serious implementations on like, you know, cloud infrastructure that we're trying to figure out what to do with. And I think I told the team that we were talking today. And I think one of the, uh, one of the questions that they have, or one of the feelings of consternation that they have considering using cloud managed agents for this kind of thing, for spinning up agents for our customers, is just right now it's like, we have a playground we have, we just have like a little, we have a server or a Mac Mini. We can just like pipe stuff to, to Claude. It can do anything that Claude code can do. It has a file system. It has a browser. It has like all this stuff. If we want to, you know, switch it, switch it out to GPT 5.5 or Gemini or whatever, it's like pretty easy to, to do that. Um, so is that kind of, and I feel like they, they, they feel like they're, we're going to get, if we use a cloud management agent, we're going to get locked in and it's not going to, we're not going to have the flexibility to do all the stuff that we want. And it, it, there, there's also a worry that features are going to come to Claude code itself that won't be in Claude managed agent for a little while. And that it'll prevent us from being at the edge, which is sort of what we promise to our customers and really to ourselves. Like we just love being like, just doing whatever the new thing is. How do you think about that? Yeah. So I think the, what's nice about the way that we work internally, I guess, is like, so we run the platform and the platform for what most people think of it as is our externally facing APIs and our suite of APIs. Um, the other rest of what our team actually does is internal platform in the sense that all of our first party products are built directly on the same platform as everybody else. And so what's cool about that is we're, we spend all of our time, not all of our time, a lot of our time working with the teams internally who are building on top of the platform and kind of enabling the features that they will build sharing ideas and these sorts of things. And so I think over time, you'll maybe see less and less divergence of, um, you know, like what might be available in Claude managed agents, what might be available in coworker Claude code that might sit on top of the same infrastructure, right? Like that's, I think one way to think about that. Yeah. And then I think, you know, on your point around or your team's point around, like, you know, having some kind of like model lock in fear, I think that that's like valid, like many folks kind of have that consternation. And I think we're kind of at this place where there's a bit of like an evolution here where, you know, if you look back, um, maybe even just a couple of months ago, it was very standard to kind of build a very, very, very generic harness. It's super generic. And then you can kind of hot swap models across all of those things. And I think for kind of an older generation of models across labs that kind of worked like, okay, a lot of things were, were moving at a pace where I think that that was like mildly reasonable. I think now, um, for the next kind of generation of models, and as we kind of see it forward, I think you kind of see this a little bit from every lab, like everyone's taking like slightly different, um, techniques and perspectives on how they want to kind of advance their particular form of the model. And so in theory, I guess you could do kind of a superset of all those things, but more often than not, I think, you know, like when you build agents for your company or for your customers, you do want to deliver like an outcome ultimately for them. And so I think that that level of abstraction of like what you're actually hot swapping stops becoming this like really generic harness and hot swapping the model. And it gets more to like the harness and the model get very paired. You still need redundancy and you still might want to use other models for things, but you probably do it at the layer of like the agent, meaning like the harness plus the model, um, rather than necessarily the other architecture of like, you know, really, really generic harness and hot swapping everything underneath. That's really interesting. Is that how I don't know the cursors of the world are doing things like, do they have a separate harness for each model or is it a generic harness that they're kind of hot swapping the models in and out of? Do you know? Um, I'm not entirely sure. My, uh, intuition would be that like, I don't know about cursor in particular, but there have been like teams that we have talked to who have kind of fallen on similar kind of perspectives. And it's mostly because they're just trying to squeeze the most out of each model to kind of like, uh, almost like harness engineer, like every single like nuance. Um, and you know, one example that we have, it's, it's not an external customer per se, but, um, something that we've done a lot internally, like we recently launched like memory for example, with, with managed agents. Um, and we tried a bunch of different harnesses ourselves. Like we tried one that was like the one that we ended up launching. Um, we tried a bunch of others using a bunch of different other techniques. And, um, at least personally for myself, like when I saw the kind of like eval suite from like the team, but each one of these harnesses performed drastically differently. And so I think like just even looking at something like that, um, shows you that like you can actually hill climb a tremendous amount by just like harness engineering the right pieces together. And I think if you were to just take that forward across like all model combinations, across all different labs, all different kinds of providers, um, there is a lot of alpha in that kind of construct. And so I wouldn't be surprised if more than just ourselves have experimented with that level of like, you know, unit tying. It's really interesting that there's this path dependence where you make some choice for how you do requests and responses or how you do tool calls or whether you're, you have the model want to use file systems or not. And then that sort of like changes the trajectory of all these different models. Yeah. And it feels like maybe at the time, like such a small, almost like, you know, kind of like footnote. Um, but it ends up, uh, becoming very big. Do you think that that will end up affecting the model's generalizability in the sense that at some point, um, they, they'll just have these sort of, uh, maybe locked in lanes of stuff that they're good at because they're, you know, Claude is really good at file systems and open AI is, you know, GPT is good at some other things. Like, yeah, how is that gonna, uh, how's that gonna flow through the model's like personality and behavior if it's like locked into a specific way of doing things? I do think it does actually kind of tend to lock the model. So like what, um, what we end up like kind of treating as like the right path and the right primitives need to be like very carefully thought through. Um, and so like, I think in the, in some eras, uh, you know, like of, of other models, uh, they become really, really, really good at like reasoning and then they almost like over optimize on that level of reasoning. And there's other perspectives around like, okay, like, yes, we wanted to be really good at like a computer, like maybe the computer part is the interesting part. And so if you think through maybe some of the, the primitives, which we could get right, we could get wrong, but at least we'll like go through the thought process of like, that will probably at least lead us, you know, one path or the other. Um, I think it's hard to say like, you know, in which direction per se will ultimately be true, but I do think there's a lot of like path dependency it ends up taking. So being really like thoughtful about what you choose to actually include or give kind of the model more natively is really important. Are there any of those path dependencies that you've had to undo? Hmm. Um, probably, um, I can't, I can't speak enough about that at the anthropic level. I've only been here like a couple of months, but I have to imagine that that has been the case. Um, I mean, we've experimented like even at other labs, like that, the kind of like primitives that we have to take a look at are constantly changing. Um, and you do kind of hit like a little local maxima and rethink like, okay, maybe there's like a more generic approach. Yeah. Yeah. Yeah. Interesting. Um, I want to take a, take a step back and ask you something that maybe I should have asked at the beginning, which is like, who, who is cloud managed agents for? Right? Like I, I set one up earlier today. We, we've got some people already using it in production inside of every, and I just, I just did one today. I really loved the, um, the sort of like getting started chat experience that you, that you had and the sort of, um, some of the examples that you had. And it, it felt to me like, even if I was not technical, I might want to use this to set up an agent. It, it might be a little bit complicated, but what I actually did is I just, uh, and I'm sorry to say this, but I did it in the codex in app browser. So I had codex driving the, uh, the managed agent setup and it like, I had a Slack bot working pretty, pretty quickly. It was like, it was really cool. So how do you think about when you're designing stuff, when you're designing cloud managed agents, who it's for? Yeah. So it's interesting because I think you're right that, especially with that quick start experience, which we actually felt pretty strongly about launching, not specifically for the sake of making it so that non-technical people could go and build agents, but actually just for anybody technical or not to be able to wrap their head around the primitives, like the APIs. Here's what I can do and here's what fits together. Yeah, exactly. Like, you know, the, the kind of education portion of it. Um, but I think when we think about who is for, we think about a couple of different things. One is we're seeing people internally within companies build automation or build really powerful platforms or systems. Like we've seen people say, I want, you know, a full end to end software development platform, right. And like managed agent is the perfect solution for something like that. Or, you know, I want to automate a little process over here where like legal has to review my marketing copy. Right. And things like that. Um, and so you shouldn't have to re-implement memory and like all that stuff. Every time you're doing that, right. You can get started really quickly and you can get something running quickly. The other user that's top of mind for us is people building into their products that they expose to their customers. Um, and so that's the other one where actually, yes, like you do still want a lot of customization. You do still want to make something that's going to be really powerful for your product, but we still like definitely, definitely believe that not spending your engineering resources on the infrastructure and on all the little harness engineering tweaking sort of stuff is like worthwhile. Why couldn't we have talked like a month ago? You would have saved us so, so much time. We'll just need to talk more. Yeah. But I am, I am sort of curious. Okay. So maybe infrastructure is one of these things, but, um, when you see people setting up agents, what do they, what do you see them think the hard thing is and what ends up actually being the hard thing? And are they the same? Good question. I, I, maybe this is, I don't know, spicy. I'm not sure. But I think, I think people think the harness engineering part is the hard part. Um, and so actually like, you know, in the past we launched the agent SDK, which is what you guys I think are using, um, on your Mac minis. And for a lot of people, they were like, okay, great. I don't have to do the harness engineering part where I have to do prompt caching and I have to maximize my context window and all these sorts of things. I think we're just actually using just Claude in back like the Claude dash P command. Oh wow. Yeah. Okay. It's, pretty good. Yes. Yeah. Okay. Cool. And, but regardless, like you guys did that because it takes off your hands building the harness. Right. Um, but I do think what we saw with a lot of customers was okay. Now I want to go and take that thing and like get it into production and scale it. And everybody hits an infrastructure wall. Like everyone hits the same problem of like, oh wow, I either need to like keep a server constantly running or I need to use infrastructure that will spin up and spin down and I need to store the transcript data and I need secure sandboxing and all these sorts of things. And so, um, you know, and like if you boot a Claude code session or you boot the agent SDK in a sandbox and like, that's the thing that you have running, but your sandbox loses connection and dies or whatever your whole agent dies. Right. And so I think the infrastructure part especially is the wall that most people end up hitting, but they're more expecting that the actual harness engineering and like getting the most out of the model is the part that's going to be harder. Yeah, I totally agree with that. I was just going to say like, you know, we, we talked to so many people who are now at a place where they're like prototyping really quickly and they're super excited and it's like, it's doing the thing. And then yeah, there's like a class of people who are, you know, really pushing and being like, okay, I do want to hill climb. I really want to edit the harness. Um, but then once you have that thing, like productionizing is just a frigging nightmare. Um, especially for the more interesting kind of long running async ones that you want to do a bit more remotely that are a bit more autonomous. Um, and everyone kind of runs into that wall and was a big inspiration for why we built what we built. I feel like, uh, one of the like, er examples of the shape of an agent is OpenClaw. Um, and in particular the, the thing that it has brought to us internally is you have an always on agent in Slack that has its own personality and it has its own like part of the world that it like ends up working on. Are you guys like, is, is that a possible future for like, okay, a one click agent that lives in my Slack that yes, I can go set up all the internals, but like, I don't have to really think about all of the, um, you know, the technical infrastructure stuff. Um, cause I, I think you, you all have the, the beginnings of that, but it's still like a lot of steps from the current managed agent to something that's always on in my Slack that I have to like set up and customize. So is that, does that fall in the realm of platforms job or is it like too far in the product direction? No, it definitely is, uh, something that we really want to do. I think like, you know, we, we focused a lot on kind of the infrastructure piece to start because that's where we just see a lot of these like pain points. Um, but yes, like I think in like, it's like, you know, I don't want to say exactly say final shape, but in its like advanced shape, we actually want to make it so that you can kind of deploy these agents really, really easily. Like, um, we've made like some light steps in this direction. Like for example, we included vaults, um, as one of the primitives as just kind of And vaults store your like keys and stuff like your OAuth keys? Credentials. Credentials. Yeah. Um, as like, you know, kind of solving some of the lower level pieces as a starting point. But once you kind of wrap some of these more sort of like agent identity type of primitives in a more secure way and you can handle it really easily and it works with like the whole like system, then, uh, you know, I think it's very natural for us to get to a place where maybe you are either one clicking, uh, Slack integration or alternatively, even maybe just telling, you know, Claude like add Slack and it just like handles absolutely everything. And then before you know it, your little bot is just picking you on Slack. I love it. I, uh, can't wait for that world. Uh, what are the best internal use cases of agents? Because I think there's this big question happening right now where, okay, yeah, everyone's in codex or Claude code, but then now we have these agents that are out in the cloud. Now everyone inside of a company can like have their own agent. There are team agents that are company-wide agents. So what are the patterns that you see for when people make really useful internal agents, what they do and what they look like? Yeah, I would say we similar to, and we've actually seen a few examples of these in some of the more like AI pilled, AGI pilled companies like, um, Stripe built minions and they talked about that a lot as they're kind of like end to end development platform that their engineers could use. Um, I think ramp did something similar and, um, we've done similar things as well, right? That's interesting. Yeah, we've built kind of platforms internally that are, you know, I have agents running that I can talk to from Slack or from wherever, right? And, um, at a certain point that becomes actually like a pretty thin layer on top of managed agents. Like you don't have to do very much to accomplish. That's what I was thinking. Like I looked at minions or whatever ramp does and I was like, it, why, why, you know? So is it, is it actually useful to have a sort of like thin coding agent that anyone in the company can use or like, why not just install the Claude app in Slack? Yeah, I would say the difference in a platform like that and some of the things that we've done internally is there's a lot of customization that you might want to do on, you know, the development environment where an agent is actually running and able to verify its changes, right? And things like that. It's like, here's how our CI CD works. Yeah, exactly. And so, you know, I think for lots and lots and lots of people like Claude code is an excellent tool, right? And you can run cloud agents with Claude code and that is really great. But I think if you're trying to do a bit more end to end development, right? And you maybe want to bake in more custom things and you could start with something like managed agents and build a layer on top of that and end up with something that's maybe closer to that end to end experience. It also seems to me like there's something in particular about having a team that you need to work with that makes the managed agent shape important as opposed to it just all works in Claude code. Like, I guess technically you could like sync the skills between everyone's Claude code, but like there's something about just we all have one agent that does this thing that seems to work. Yeah, I'm really glad you brought that one up because I think like that's actually like one of the more common areas where we see a lot of the opportunity is that to your point, you know, there's a lot of like individual productivity that's happening, whether you're a developer or non-developer, there's like so many tools that you're using to just like make yourself like more automated, more, you know, high leverage. But then when you get to the team layer, suddenly everything gets like massively more complex. Like number one, obviously you can't like sit on your laptop. And yes, you could maybe like, you know, put it in the cloud, but it's again, more for yourself to kind of like handle with your laptop closed. But then you go to like, okay, well now like the three of us want like, you know, a couple agents that interface with each other and work with each other. And then maybe we're automating a process kind of end to end. And especially for some of the more complex processes that you kind of envision being like really transformed with AI, you do need like, you do need that kind of like team orientation. And that needs to happen at like a layer that's a slightly higher bit of abstraction than just a single agent. And I think some of the teams exploring, you know, kind of multi-agent architectures and things like that are really exciting, but it needs to be built on top of a little bit of like a platform that everyone kind of spin up and down and control. And I think G from Vercel like had a really good perspective on this in a way where I think his company Vercel is obviously incredibly like AI pilled. And he kind of describes it as sort of like an AI like software factory like internally. And I think that's exactly like the right mindset. And that like produces, you know, like extremely high leverage organization that's really just like creating a tremendous amount of productivity, but not just for themselves, just like for every single process that they have in the company. Mm-hmm. And I really want to go back to this like, okay, agent use cases, we've got coding agents that that anyone can use in the company. Like what are the other ones that are that you see people standing up that are really useful? We've seen a few. So one of the fun things that we get to do is just kind of work with our internal teams of different functions and like help them agentify because we actually just get to learn a lot as a result of doing that. And so the silly example I brought up earlier of like legal team needs to review marketing copy was one of the ones that- Very real. Yeah. Like extremely real. Like really blew people's minds with like very basic agents that just give people the right setup to be able to do that. So we've seen that. Well, what does that actually do? So it's like there's marketing copy and there's a legal agent that is just like watching what everything marketing does and is like, stop, like- No, no. It is more like, okay, I'm a marketer and I've written some copy, right? And in the past, maybe you would have opened a ticket or something and be like, can you please review this copy? But instead you submit it to this, like, you know, a little app that we built on top of agents that is like, okay, cool. Now I'm going to go as an agent review first and then put it in legal's inbox as a already first pass review was done. And maybe actually like the agent is it's clear enough that it can say, okay, marketing, you're good. Right. Or maybe it's still like, no, this needs like an extra human review. And so, yeah, just, and that's the sort of thing where again, just thin layer on top. Um, but you can build the, you know, you have access, I have access, we can both see the outputs and we can work together on it. Okay. But then, so for example, why is that not a skill? So it's, it can, it very much can be a skill. And that actually is like, if you, you would probably build that agent as a, you know, legal reviewer agents, right? And so you would have MCP servers or whatever it is that help you access external contacts. You would have skills that help you understand like, here's what rules we have to follow and not follow. Right. And all those things, and you'd put all those things together, but then you can just fire off a session with that agent. Yeah. And then I think the last piece you need, and this is where I'm saying it's a really thin layer, is just like the form factor on top where like different people can collaborate together and like work with that agent and multiple agents can be involved in the system. Um, and so I think it goes a little bit broader than a skill because you kind of still need like the right form factor for the agent to be able to go run and then for people to be able to interact with it. Another core bit on why it's like not a skill is because, or not exclusively a skill is because you actually do need human in the loop. Um, and so like if you were to automate the whole thing and you were just, you know, like taking the skill and looking at yourself from like legal skill, for example, like in that world, of course you could have just like done a pure skill. Um, but if you need a human in the loop to be like, okay, like I want to review and I do want to check and I want to like, we're looking at like legal things. And so there's a bit of like, you know, authentication that's sort of necessary. Um, in order to automate that entire process, you kind of need like agents to go do the thing. And so because you need to spin up sort of separate sessions for that to happen, some sort of stitching is necessary that can't be instantiated in a single skill. That's really interesting. Yeah. Um, okay. So just to push on that a little bit, so what is the best practice for you? You create an agent that, uh, it's, it's job is to make sure that when marketing is writing something, they can get it approved really quickly by legal. And sometimes it'll approve things immediately. Sometimes it sends stuff to legal. And ideally it's like getting better all the time. So it can do more and more, right? What is the best practice for who owns that agent once it's built? Because one of the things that we found is if you don't have a human who's responsible for the agent, it gets stale very quickly. And then it ends up being kind of this like dead thing. That's all just like out there doing stuff, but it's not necessarily good. And also, uh, even if it kind of works, there's all, there are going to be all these times where legal's like, you asked me to approve this, but I don't really need to approve this thing. Like let's update your prompt. So like, how does that all work when it works well? So it's actually really interesting because so the form factor thing, right? Like the app that sits on top of that, that we originally built, um, one of our teams worked on that, right. And, and like kind of sitting with these teams and understanding what they needed. Um, and they were kind of like, okay, here you go. And we're going to go do other stuff now and like, let us know how this goes for you. Um, and then a really cool thing actually ended up happening where people on those teams who were using the tool were like, Oh, I wish like this little thing could get tweaked or this thing could get better. Um, and they like popped open cloud code, like made some of the changes themselves. Yeah. And so it's funny. And then is your team responsible for approving the PR or does it just like go in? Uh, usually my team's responsible for reviewing the PR if it's a system that we actually own, but, but yeah, like people can kind of self-serve making changes to those things, which I think is really cool. So, um, it is, I do think we're still in a stage for a lot of teams, a lot of companies, like even going back to, you know, like Stripe has minions, right? Like Stripe has a large developer productivity team. Uh, we used to work at Stripe, so we spend a lot of time with them, but, um, they have a large developer productivity team. They're awesome. And they're obviously putting a lot of work and energy into building platforms and tools like this. And so I think we're definitely still in a place where something like managed agents or being able to build on top of our platform is really powerful, but you still kind of need the like AI pilled people and technical people within a business to then go like create something really excellent on top of that, that works well for whatever you're trying to do. That's interesting. Yeah. I, I love the, anyone can open a PR to, to do this because everyone's using cloud code. One of the things that I find talking to people who are in infrastructure roles at companies where this is starting to happen is like, you know, that, you know, the meme where it's like, um, there's, there's a person and he's like going like this and he has like daggers in his like back and he's like covering it. It's like infrastructure people are that for like, now anyone can like, can, let's can submit PRs. Um, how do you, how do you deal with that and how do you do that? Well, cause obviously like in an ideal world, you would love for legal to be able to submit PRs to improve this agent. And also, um, sometimes they're probably going to submit stupid stuff that wastes time. And so what are the, what are the right ways to either organizationally, like culturally or technically, like make that possible without ruining your, your lives? For this particular one that we've constructed that Caitlin's given as an example, we actually have like a couple layers of abstraction away from like that kind of like PR layer. So at the very beginning, it kind of like started that way and to kind of like basically prevent users from kind of foot gunning themselves a little bit, uh, they kind of get to a place where oftentimes their way of interacting with the agent that they own, like the, whether it it's the marketing team who owns the marketing agent requesting, or if it's the legal team, you know, owning the agent that does the review. Um, they actually engage with those agents through Claude itself. So they actually spend more of their time, like kind of talking directly to Claude and then Claude will oftentimes figure out what should be the right way for them to go and handle it so that they're not kind of like, you know, hopping straight down to the absolute core bit and doing something that may result in, you know, some complications. And they're talking to Claude or Claude code, like Claude chat or Claude code or co-work. It's a different instantiation of, of Claude, um, that we made that actually is a managed agent in and of itself. So it's just kind of like managed agents all the way down, um, in that construct. Uh, but we found that each layer, if we kind of tune and prompt each variant of the managed agent, it helps to solve like, you know, different parts of the problem for users. So at the end state for, you know, like that marketing person or that legal person, it is like a really simple interface where the way that we tell them is like, you're just talking to Claude, but under the hood, it's many, many Claude engaging with each other to get to the part where then they, the Claude's themselves are doing the more complex work that the human doesn't really necessarily need to interpret. Interesting. You guys just launched multi-agent orchestration. What are the coolest, what are the coolest things that people are doing with that? Hmm. One of the more interesting ones is like, um, I think people are using it to like construct sort of different harness techniques. And that one I'm personally very excited by. Um, because like there's different techniques that people have experimented with where, um, you know, like for example, we recently did like the advisor strategy one, but really if you were to genericize it, you just separate like execution from advice. And there's also one where you can have like two, you know, modes where there are one is generating someone, something, and the other one's adversarial to it. And then there could also be sort of like, you know, you split it into a bunch of different like little tiny pieces and then they kind of recombine. And then there's ones where maybe it's kind of something closer to like best of end kind of like style of thing. And then there's so many more. And like in each one of these different types of like architectures or strategies, they are good for a very specific use cases. So some of them are much better for like, uh, deep research or wide research type of, uh, style use cases. Right. And there are others that are like, these are like the kind of ones where they all sort of swarm together are better for like bug hunting, for example. Um, and so like, that's like really cool to see that like, if we can make the primitives very Lego like, um, then people can put them together to solve things at a slightly higher form factor, which is more like an architecture or like a strategy. Um, and they get much more like interesting results out of that. Um, and that's like really exciting to see because it also suggests that you can actually hill climb at multiple layers, um, of abstraction. How do you know if an agent is successful? How do you measure success for an agent? Yeah. I mean, there's like evals and stuff like that, which everyone has talked about like ad nauseam. One direction that we, we really like is like, uh, this kind of verifiable outcome. Um, we've been somewhat opinionated on that one and it's almost like in the absolute end state of, you know, we talked a little bit about what's, what's a platform at the end of things. Um, going from that philosophy, it's like our kind of principle of like maybe the end state of some of these things is that everything should kind of compress down to an outcome and like a budget. And that's probably like about it. Um, and everything else should be figured out for you to kind of resolve exactly across those parameters. And so for us, we're kind of, yes, we still have evals. We have a lot of these other things that we measure, um, that are domain specific, like, you know, some coding evals would be like, you might want to measure like just the actual PR getting merged. Those are more verifiable. But as we get to the place where, um, you know, like an outcome is actually a spec that you are just as a human able to define and our ability to interpret that and regrade itself over and over is closer to what we care about. Claude, make me a billion dollars. Your budget is $10. Exactly. I meant to say no mistakes. Go. Go. Exactly. Maybe Mythos could do that. Yeah. Um, and then one of the things that we were running into that I'm curious if you have a solution for is agents like get outdated pretty quickly. Um, sometimes because there's no human attached to them. Sometimes like they're just running an old model or there's an old or in an old architecture or whatever. And it feels like there needs to be a, um, uh, end of life cycle for agents. Like we've talked about having like a little like funeral for them and like having like a little page on our website that's like, here's all the decommissioned agents and stuff. Um, like how, how do you manage, especially in a company, in a really big company, how do you manage the, all the agents that are like sort of out there, but, and maybe they're like in Slack, like pinging stuff once a week, but you're like, this is super stale. How do you make sure that, uh, you, you retire them as quickly as you are making them? So one of the things we have actually done is we have made skills that help you do things like upgrade to a new model when a new model comes out, right? Like we've actually put a good amount of work into making it easier to do exactly what you're talking about. Um, and I think maybe some of the most like AGI pill people are like running agents that are monitoring their agents to see if their agents are, you know, like outdated and in need of that sort of stuff. But I think for, you know, the way that we like to talk to customers who ask us this question, I do think the, the most interesting instantiation of this is there's a new model and now I need to go upgrade my agents or maybe be done with those agents because you know, the new model enables me to build agents that are way more powerful and do more interesting things than the old agents did. Right. But I think that upgrade process and that migration process is like something people have had to wrap their heads around as like, it's like a breaking change and I have to like put actual energy into making that work. Um, and obviously, sorry to talk about evals, but like if you have evals, this process is easier and things like this. But, um, I do think that's one of the things we've tried to do is how do we give you skills and how do we give you the right, like just tools to make that process easier. Um, and then you could go be AGI pill and choose to actually automate more of that with more agents. Yeah. So a year from now, we're back at code with Claude. Um, where do you think the platform will be? What will I be able to do? Uh, and how it will be different from what I can do today? Do you want to go first? You can go first. A year is a long time. It's in this, in this industry, especially how close are we to Claude make me a billion dollars? That's really what I'm asking. Um, I think, yes. If that works, we probably won't be sitting here. Yes, yes. We'll be asking Claude for this. Um, I mean, yeah, like we want to get closer and closer to that, that state where I think we, we kind of, okay, so a couple of things. I think in a year from now, I mean, one thing that we'd love to get really, really close to is actually that kind of like simplicity. And this might be a significantly higher order of abstraction. I don't know what the form factor will look like or whatever, but the kind of parameters we will care for from users will be that outcome. And of course, it has to be verifiable. There are some parameters that have to be restrictive and, and the budget. And I think like we'd want to experiment with, with directions where Claude actually gets so good at understanding itself. It, it figures out what model you should be using. It figures out how to spin up all the sub agents. I actually don't think you need to think so much about harness engineering in that world. Today, you know, you don't have to think so much more aggressively about like tool construction, for example, that we've kind of made that a little easier and you get it to leave a little bit of that scaffolding. Less prompt engineering. Yeah, exactly. Exactly. And I think if you just keep going up that stack, like today, a lot of the innovation is happening at this kind of like, like, like, like really high level, almost like harness architecture, like level, which is really fun. But I think a lot of that honestly also kind of goes away where you almost like don't have to think so much about like model selection. You don't have to think so much about what kind of architectures are there because we probably put a, would have like gone through enough iterations with Claude where Claude is actually able to understand itself enough, um, that it can almost like write itself on the fly to figure out what is necessary in that kind of like two parameter world of like outcome and budget. I don't know that we'll get there like in a year, but I feel like we might be able to do like the outcome part of that with like maybe, you know, some bars of some error bars on the budget side. Really cool. Yeah. Um, okay. That was really cool. I'm gonna give you a slightly more boring answer, which is in that world. If Claude is like on the fly or agents on the fly are like becoming what they need to become in order for you to do what you're trying to do, the platform has to like seriously scale. That is. Um, and so I do think some of this will be, what are the right abstractions that actually enable that, right? Like somewhere on the primitive to higher order realm, right? But I do think so much of what our team is going to be doing is making sure that the tokens that people want to come in and out of Claude, um, are going to be able to come in and out of Claude because our system is scaled to meet not just the demand, but like in that world where it's just like you have agents that are like literally constantly running and recreating themselves and doing this sort of work. Um, you just need a system that, you know, can handle long running requests, can handle a bunch of differently shape things. And so I think for us, it's going to be, I never want the ability of the platform itself to be able to scale, to get in the way of what people would otherwise be able to accomplish with these things. And so I think that's something that's going to probably be very friend of mind when we're talking in a year. Awesome. I'm excited. Um, thank you so much for joining. I really learned a lot. Thanks for having us. Oh my gosh, folks, you absolutely positively have to smash that like button and subscribe to AI and I. Why? Because this show is the epitome of awesomeness. It's like finding a treasure chest in your backyard, but instead of gold, it's filled with pure unadulterated knowledge bombs about chat GPT. Every episode is a roller coaster of emotions, insights, and laughter that will leave you on the edge of your seat craving for more. It's not just a show. It's a journey into the future with Dan Shipper as the captain of the spaceship. So do yourself a favor, hit like, smash subscribe, and strap in for the ride of your life. And now without any further ado, let me just say, Dan, I'm absolutely hopelessly in love with you.