Every

How to Build an Agent-native Product | Mike Krieger

952 summary words 4 min summary Watch video

Start with the signal

4 min read

Summary

How to Build an Agent-native Product | Mike Krieger

Main Topics

  • Product Development in the AI Era: How AI acceleration changes the pace and approach to building products
  • Agent-native Product Design: Building software that agents can use as fluently as humans
  • Team Structure for AI Projects: How hiring and team composition must evolve with AI capabilities
  • Enterprise vs. Rapid Innovation: Balancing customer expectations with the need for continuous reinvention
  • Personal AI Agents: The emerging phenomenon of personalized, named AI agents and their role in product design

Key Points

What's Changed in Product Building

  • Speed is deceptive: While models can build "zero to end" in hours, this doesn't mean you understand what the product should be
  • Models are good at adding, not subtracting: AI excels at feature generation but struggles with intentional simplification—the critical art of product design
  • Rewrites are now feasible: What once took a year (and could destroy a company) now takes days, enabling rapid iteration and course correction
  • The "indoor tree" problem: Building an entire product without incremental user exposure creates something polished but potentially hollow—lacking the intuitions built through real-world pressure

Agent-native Product Principles

  • Design for agent capabilities from day one: Every action a user can take, an agent should be able to take
  • Make tools available, not just tell users about them: Rather than explaining steps, agents should natively execute actions (e.g., "add to project knowledge" should just happen)
  • Self-awareness matters: Products should have knowledge about their own primitives and capabilities
  • Keep the run loop open: Single coordinating agents that delegate to sub-agents feel more like collaborators than tools
  • Enable customization and extensibility: Users shouldn't have to ask permission to modify how the tool works

What Still Requires Human Expertise

  • Architectural robustness: Systems thinking, distributed systems knowledge, and data integrity concerns still need senior technical oversight
  • Distinguishing real problems from prompt fixes: It's easy to patch issues with system prompts instead of fixing underlying architecture
  • Product intuition: Understanding the problem space deeply takes time and cannot be fully delegated to AI
  • Strategic decision-making: Deciding what to keep, cut, or reinvent requires human conviction

Team Structure in AI Labs

  • Single founder-like conviction: Projects need someone with extreme belief in the problem space, willing to push through obstacles
  • Not always pure technical founders: Can be designers, product-minded engineers, or founders—the critical trait is conviction
  • Small is better: Adding team members before product-market fit creates coordination overhead that slows down iteration
  • Hybrid roles: Designers increasingly write code; engineers handle UI/flow; specialized skills flow in and out based on needs
  • Two-week evaluation cycles: Rapid assessment of which bets to double down on vs. wind down

Enterprise Challenges

  • Obsolescence risk: Products built for today's enterprise needs will be outdated in 3-6 months in the AI era
  • The rewrite question: Startups must choose: bolt new capabilities onto old architecture or do major rewrites
  • Enterprise toggles: Providing ways to disable new features for conservative customers while core product evolves
  • Customer expectations vs. pace: Enterprise customers need assurance of continuous improvement, not stagnation

Notable Quotes

> "The models today are good at adding features. They're not necessarily good about figuring out what to cut out of the product."

> "I feel like that is the art and science of software design in 2026."

> "Some of the intuitions you build about what are the right things to put in there, I think you build over time."

> "You either just get very generic products that are unlikely to break out or ones that just don't reflect some deeper intuition that you come to about your space or your product."

> "There's a lot of interesting research questions around privacy and what my agent knows about me versus what it discloses to other people."

> "Everything is getting compressed...you have to be willing to put out the V3 or the V4 that is a big rethink of how the existing piece worked."

Takeaways

For Product Builders

  • Embrace rewrites: The cost of starting over has plummeted; use it to stay aligned with model capabilities
  • Launch minimal, iterate real: Get to real users faster than trying to predict what they need
  • Maintain conviction: Delegate execution, but keep deep understanding of the problem space
  • Build for agents from the start: Design primitives that enable both human and AI interaction naturally

For Teams

  • Keep small until necessary: Only add people when the scope truly exceeds one person's capacity
  • Hire for conviction + complementary skills: Less about rigid roles, more about who can hold the vision
  • Pair specialists strategically: Connect product/design vision with systems expertise and domain knowledge
  • Enable rapid re-evaluation: Build structures that allow quick pivots without lengthy alignment meetings

For Enterprise Founders

  • Be transparent about pace: Enterprise customers will accept rapid change if they believe you're evolving with the field
  • Build escape hatches: Provide toggles and backwards compatibility while core product evolves
  • Rewrite boldly: Don't become prisoner to old architecture; be willing to throw out the old to embrace the new
  • Understand your defensibility is temporary: Plan for a world where your advantages compress into months, not years

The Emerging Paradigm

  • Agent-native is non-negotiable: Products that assume agents as first-class users will outcompete those that don't
  • Personalized agents will be central: Custom Claude instances with specific knowledge and names create emotional investment and trust transfer
  • The permission/capability boundary is unsolved: The next major design problem is how to keep agents powerful while preventing chaos
Full transcript 10030 words · 58 min read
0:00

SPEAKER_02

The models today are good at adding features. They're not necessarily good about figuring out what to cut out of the product. You can get it to go zero, not just zero to one, but zero to end pretty quickly over the matter of hours. It's made a lot of decisions along the way. And some of the intuitions you build about what are the right things to put in there, I think you build over time. I feel like that is the art and science of software design in 2026.

0:35

SPEAKER_00

Work moves fast. And in the age of AI, the pressure isn't just to move faster. It's to make sure that what you send actually sounds like you. From emails to proposals to stakeholder updates, generic and rushed just doesn't cut it. If you've ever stared at a blank page, knowing exactly what you want to say, but not how to start, Grammarly fixes that. Grammarly gives you one place to think, write, and finish your work. Write where you already write. Most AI tools either take over or stay out of the way. Grammarly does neither. It helps you break the blank page, adjust your tone so a message lands right for the specific person reading it,

1:13

SPEAKER_00

and works seamlessly across more than 500,000 apps and sites that you're already using. It's loaded with agents built for every step of your process. And 90% of professionals say it saved them time. 93% say it helps them get more done. This is AI that works with you, not over you. In a world of generic AI, don't sound like everyone else. With Grammarly, you never will. Download Grammarly for free at Grammarly.com. That's Grammarly.com. Mike, welcome to the show.

1:39

SPEAKER_02

Great to be here. Thanks for having me on.

1:41

SPEAKER_00

Great to have you. I'm super excited. For people who don't know, you are the co-founder of Instagram, and now you are at Anthropic and Anthropic Labs.

1:53

SPEAKER_00

I've admired your work from afar, both at Anthropic and Instagram for a really long time, and you're obviously at the forefront of building products and AI. So thank you for coming on. Absolutely. Where should we start? What we were talking about just now in the pre-production is what has gotten easier and what has gotten harder or maybe stayed the same in product building as the underlying substrate or the process by which we build products has changed completely. So tell me about your experience now versus earlier in Anthropic versus Instagram and how you think things are changing.

2:31

SPEAKER_02

Yeah. I was doing the thought exercise a couple of weeks ago of the Instagram story. We had another product called Bourbon. We worked on that for almost a year. It wasn't working. We pivoted. We basically spent three months building what became Instagram, launched it, and then scaled it. And I was asking the question, what is now trivial? And what was actually inherent in that building process that doesn't get easier? And that year, we probably could have hit some of the dead ends we had eventually hit sooner. But there was value in getting there too. We overcomplicated the product so that we then had to simplify it.

3:02

SPEAKER_02

I find even the models today are good at adding features. They're not necessarily good about figuring out what to cut out of the product. And that took a lot of just hitting actual real-world usage. And there was something about the process of incrementally adding things. I mean, today, especially some of the stuff we're building in labs, you can get it to go zero, not just zero to one, but zero to end pretty quickly over the matter of hours. But it's made a lot of decisions along the way. And you can ask it to follow up with you and then do input. But some of the intuitions you build about what are the right things to put in there, I think you build over time.

3:38

SPEAKER_02

And so I've been reflecting. There haven't been a lot of breakout consumer products, even in the age of accelerated AI building. And I think part of it is because it just still takes time to hone your view about what intervention you want to make on the world and then build from there. Now, the actual building part, once you know what to build, is of course so much easier. I had Claude basically rebuild Bourbon. It took about two hours. It was feature complete. It added filters, which Bourbon didn't have. We added those for Instagram. But I think it knew, you know, it knew the eventual future of the products. It decided to build that in.

4:09

SPEAKER_02

So I think that part feels really different. But I think there's also, you know, I remember there was a week where Kevin went off and built all the filters for Instagram. I went off and built the rest of the app. And, you know, sitting there, I would stay up till 4 a.m. and then sleep till noon. That's my natural day-night cycle. And in that process, you're making so many decisions. How should location work? And, you know, we've got to find a way of accelerating building while still helping people build intuition of those decisions along the way.

4:38

SPEAKER_02

Because otherwise, I think you either just get very generic products that are unlikely to break out or ones that just don't reflect some deeper intuition that you come to about your space or your product.

4:47

SPEAKER_00

This is great. I love this. It's making me think of two things. One is I have this thing in my head that if you grow a tree without it, with it being indoors, without being exposed to wind, it doesn't get as strong. Because as it's growing, it needs all these forces pushing it back and forth in order to make a real tree. And so if you have it indoors without wind, you're going to grow a tree, but it leans and it's not as strong and it's not the same thing.

5:20

SPEAKER_00

And I think there's something that you're saying here where because we've accelerated the pace of development so drastically, what would normally be this incremental thing where you're doing things one at a time and then you're exposing it to users. You can actually grow an entire tree indoors and then you have this whole thing that you're just like, it doesn't have the same level of intuition and exposure to experience at each step that creates a great product. Is that right?

5:51

SPEAKER_02

I love that. I love that metaphor too. When we were starting Instagram, we had this, we were very into Eric Ries and Lean Startup and the whole Yagni, you ain't going to need it principle. [SPEAKER_00] And I think there's something that you're saying here where because we've accelerated the pace of development so drastically, what would normally be this incremental thing where you're doing things one at a time and then you're exposing it to users. You can actually grow an entire tree indoors and then you have this whole thing that doesn't have the same level of intuition and exposure to experience at each step that creates a great product. Is that right?

6:13

SPEAKER_02

I love that. I love that metaphor too. When we were starting Instagram, we had this, we were very into Eric Ries and Lean Startup and the whole Yagni, you ain't going to need it principle. And I have found, and actually even one of the things I was working on in labs recently, we way overbuilt for V1 before we even got to early access because you can think, oh, well, we have this option. Why not add this one as well? That's a PR of work. And if you get a really good flow in cloud code, you're firing things off, you're going to lunch, you're coming back, the thing is done. You're like, great, we added it. And the thing we realized was we'd created this matrix of functionality that was actually quite hard to test and keep up with right before launch or even to explain to people. Like they're arriving. The metaphor somebody else gave me, which I really like is the difference between getting episode by episode, getting characters in a TV show versus imagine you're thrown into the final episode. And you're like, wait, what are all these things? And who are all these people? And I'm expected to have all of this context. I think there's the same kind of feeling around developing something over time. But the tree metaphor sticks too as well. And so showing somebody the fully formed tree is also a lot all at once. And I think there's definitely something there in how do you build product these days and still keep it simple. And just because you can doesn't necessarily mean that it should be in at least the first version.

6:14

SPEAKER_02

[SPEAKER_00] I'm having the same problem because I was literally up until 4 a.m. debugging and fixing this app that I made on the side at Every called Proof, which is an agent native collaborative marketing editor. So you can share really quick plan docs and stuff with your team or with other agents. And you have a little presence. And it's really fun. And this is my second or third iteration of the full product end to end, which is really interesting that you can do now. But the first couple of iterations, I just found myself because vibe coding is so fun and addictive. I just found myself being like, yeah, I'll do this and I'll do this. And it just created this monstrosity that wasn't that good. It was not good to use. And I got really inspired by another product called Monologue, which I'm not sure if you've run into or not. But I got really inspired by Monologue, which is a really simple speech text app run by Jim Naveen, who is just so focused on making one simple thing work so well. And I saw how well that works in this age where anyone can make a product that's super polished and just super good at what it does. And so I basically threw out the product and started over with this very simple approach, it's just a shareable markdown link. And that just started growing virally inside of Every. Like everyone started using it all the time. And then we launched it and it just blew up. And so I spent all last night not sleeping, trying to fix it. I'm being like, I'm too old for this shit. I can't be doing this anymore because it reminded me of being in my 20s or being in college and hacking on stuff, which is fun, but also exhausting. And so, yeah, I've found that I've had to really modify my psychology because so much is possible. How are you dealing with that?

6:16

SPEAKER_02

Yeah. And just as a brief aside on that, I mean, with Bourbon, our biggest mistake was adding functionality over time rather than deleting it. Right. And because, you know, eight features doesn't make for good product. Maybe the ninth one will. Instead, it just made for something that felt really complicated. I mean, I think a couple of things are also part of how we're dealing with it is actually being more willing to do rewrites, classic Fred Brooks, Mythical Man Month. You shouldn't rewrite software because all the things that were imbued in B1, you're going to mess up. And yeah, exactly. And the whole second system syndrome. And there's still a lot of truth to that. But one, the models can help you sort of diff and basically see, did you miss anything that was in that first one? But second, it's no longer a year long rewrite that might have killed a company, like the famous Netscape situation. These are days, probably, especially off a given source. So we've actually had several initiatives that usually pre-launch, rarely post-launch, but at least pre-launch, have built the full blown thing, realize we've overcomplicated or made some kind of core assumption. And then tore it down, done a V2 and then iterated on it from there. So it doesn't surprise me that that's become part of what you've had to do as well. But it doesn't feel as painful. You're not like, oh, a year of building this thing. It's like, oh, that was last week. And then I get to do it this week and I get to cut out a lot of what was there as well. I think functionality wise and how we're dealing with it from a product development standpoint, we are learning to launch earlier. And it's definitely a balance around how we've grown. We have a strong enterprise footprint. People have expectations about what the initial version is, but not assuming that we're going to know what every connector or everything that we need to add to the product is. But ahead of launch, because people will absolutely surprise us. Right. We have a strong contingent we call them ant fooders because we're ants at Anthropic. But that only gets you so far before you need that real world contact. Like take co-work, for example. We've been noodling on a product of that shape for a long time. And then once we decided, no, let's get this out. Let's build a V1 that we think solves the problem in the most minimal way possible and get that out in 10 days was really a good push around.

6:22

SPEAKER_02

But ahead of launch, because people still will absolutely surprise us. Right. We have a strong contingent that we call ant fooders because we're ants at Anthropic. But not only that gets you so far before you need that real world contact. Take co-work, for example. We've been noodling on a product of that shape for a long time. And then once we decided, no, let's get this out. Let's actually build a V1 that we think solves the problem in the most minimal way possible and get that out in 10 days was really a good push around. Yes, there are 100 things that V1 should or could have had, but it didn't. And at the same time, it was useful enough to prove something out there.

6:55

SPEAKER_02

And I'm not sure developing it for another two months, adding 50 features would have been more useful. In fact, we probably would have been building in a the indoor tree would have been getting built. And then the second it hit real world use, it's like actually nobody wants to do that.

7:09

SPEAKER_00

[SPEAKER_02] They want to do this other piece. [SPEAKER_02] So I think that piece of that, again, there's the intuitions of the original lean startup ideas are still here. [SPEAKER_02] It's just they manifest a different time scale and in a different way. I'm really curious to hear how you think about product design and how products should work, because I've been anyone at every will tell you the phrase that I use the most, the word that I use the most about the software build is it has to be agent native. So agents have to be able to use it as anything that an agent, a user can do in the app, the agent can do.

7:32

SPEAKER_00

There's a couple other little principles of being agent native. But I basically stole that from you guys. I think that cloud code is the canonical thing that taught me about how that kind of product can work so well, where it's an agent, it can do anything on your computer that you can do. And it's customizable and flexible and extensible. So it's easy to start, but it can do all sorts of unexpected things that the designers didn't really think about beforehand. And I think that's such a good model for product development and AI. And I'm curious, this is just what I've cribbed from watching what you guys do and then put my own spin on, but how do you think about it?

8:06

SPEAKER_00

And how do you talk about making products like that? [SPEAKER_02] Yeah, that's so much in here. [SPEAKER_02] And I love the agent native framing. You all did. [SPEAKER_02] It's to me the canonical exploration of this. [SPEAKER_02] So thanks for putting those ideas out in a really clear way. [SPEAKER_02] So I think a few threads to pull on this one is a conversation I had with somebody recently where they said, you know, they're a non-technical person. [SPEAKER_02] They're saying, you are talking about agents and all this stuff, but they're like, actually, computers just work now. [SPEAKER_02] I always wanted computers to work and they didn't work and now they work.

8:52

SPEAKER_00

[SPEAKER_02] And it's a funny thing where if you knew the incantations to properly get on the command line and brew install the thing that he's going to do that. [SPEAKER_02] But now cloud can do it for you.

8:59

SPEAKER_02

And therefore, the computer now feels like a tool that is alongside you. And I think that core insight is it's more than even just adding power and functionality to new software. It's also unlocking the functionality that always should have been there or available and just felt extremely hard for people. So that's maybe thought number one. Thought two is actually comparing our products that do this well versus not. I think cloud code does it well. I think cloud AI still needs to evolve a lot. So as an example, I was watching somebody's cloud and they were in a project and they had built, I think, an artifact or a new document.

9:27

SPEAKER_02

They said, great, can you add this to my project knowledge? And cloud's like, yeah, let me tell you the steps to go add it to my project knowledge. No, that should just be a thing that it can do really natively. And so I think even in that you see a product that was a 2024 product that has been iterated on and evolved a lot. But still, I don't think has been baked in from the very beginning. I think the idea that every single one of its primitives, it should have knowledge about and the ability to modify. And I think that's essential in products these days. I think cloud code is the 2025 vintage of that. And I think there's even further aspects of it.

10:07

SPEAKER_02

When you see what some of the harnesses that folks are experimenting with, where they can actually modify the harness itself, that starts getting to the next level of that where it's probably esoteric for most people. But even unlocking that functionality means that you don't have to sit there and be like, oh, I wish it did this a little bit differently. I wish Gmail worked in this slightly different way and just ask it to. And I think that feels like the big next step. But even within cloud code, just teaching cloud code about cloud code was a really valuable experience. I was like, this definitely relates. This is now getting very circular and meta, but bear with me.

10:29

SPEAKER_02

I loved your write up on agent native. I was like, I want this as a skill. So whenever I'm prototyping something, it thinks in an agent native way. So I had to package it up as a skill. And that whole process was, you know, hey cloud and cloud code, can you create a skill for this? Sure. I'm looking at my skill. I'm going to create a skill about it. I'm going to install it. Great, is that available now or do I need to reload? Right, I think you need to restart it. Let me check. Yep, you do. All right, let's get it. And everything was, it has knowledge about itself. And that unlocks so much capability in there as well, which maybe is the last thread to pull on.

11:28

SPEAKER_02

I think all of these could be hour long conversations. And one of the things that we're really thinking about in labs is how do you imbue the software that cloud builds to be more cloud aware and even just cloud agent native sort of building aware.

11:38

SPEAKER_00

[SPEAKER_02] So that it even thinks to build in that way to start with, because it still won't. [SPEAKER_02] Partially because decades of software is not that. [SPEAKER_02] Right. [SPEAKER_02] So how do you get new software to have that principle baked in? That's the thing I was about to ask you about. [SPEAKER_02] And that unlocks so much capability in there as well, which maybe is the last thread to pull on. [SPEAKER_02] I think all of these could be hour long conversations, which is, I think.

12:30

SPEAKER_00

[SPEAKER_02] And one of the things that we're really thinking about in labs is how do you imbue the software that cloud builds to be more cloud aware and even just cloud agent native sort of building aware. [SPEAKER_02] So that it even thinks to build in that way to start with, because it still won't. [SPEAKER_02] Partially because decades of software is not that.

12:48

SPEAKER_02

Right. So how do you get new software to have that principle baked in? [SPEAKER_00] That's the thing I was about to ask you about. [SPEAKER_00] So A, I'm super honored that you read the write-up and you made a skill for it. [SPEAKER_00] That's amazing. [SPEAKER_00] And B, yeah, you're pointing to a real problem that I found. [SPEAKER_00] I think actually cloud models are the best for this. [SPEAKER_00] A codex model generally is not as good at building an agent native because they're models in general. [SPEAKER_00] Unless you push them, they think like traditional engineers.

13:22

SPEAKER_02

[SPEAKER_00] And that's a whole different set of constraints. You want to have guardrails and tests. [SPEAKER_00] You want to make sure that there's one path the user can go down versus we're creating this extensible thing that's super flexible. [SPEAKER_00] So, yeah, how are you architecting your product to teach the models and the harnesses? [SPEAKER_00] How do you teach the models to think and work in this way? Yeah, I think there's two parts to it. One is the more mundane part. The second one, I think, is the one that's more interesting and developing.

13:46

SPEAKER_02

The first one is even just having good patterns and paradigms available to the model while it builds has been really valuable. Finding the right balance of templatized to skillified, and what that right balance is. But having one of the things that we have now is a skill about the cloud API, which sounds super obvious. But even just having that is really valuable because you would sometimes find we'd launch a new model. It wasn't in the model's innate knowledge. And then you'd get into these really funny arguments where the model didn't know the latest version. It wasn't Sonnet 3.5, it was Sonnet 4.5.

14:13

SPEAKER_02

So having that capability, having good templatized examples of that and skills, I think helps. But then the second part is what's also interesting is that class of software is just a different type of test. It's much harder to write an end-to-end functional test around an agent-native product because part of it is that unpredictability. And so another idea we've been kicking around a lot in labs is how do you increase the fidelity of the verification? And the other day I had an agent-native iOS app that I was working on and I was having Cloud interact with it. And Cloud was having a conversation with itself in a chat feature in the iOS app.

14:37

SPEAKER_02

It was very funny watching Cloud talk to Cloud because it's somebody pretending to be what humans are. And this particular one was a prototype I was doing about work journal reflections. And Cloud was like, yeah, my boss is really rough on me. I had a hard day. And then Cloud's like, oh, I'm so sorry to hear that. And they're just going back and forth. But you wouldn't have written a unit test for this. And maybe it would have come up with some other emergent idea as well.

15:05

SPEAKER_02

So I think you just have to go much more towards setting up harnesses that are actually exercising as much of that agent-native capability as possible because you don't exactly know what things are going to do. And things are going to end up in a weird place where Cloud's going to try to do something that you wouldn't even think it was going to do. And it might put your app in a new state. So maybe it's circling all the way back to what's hard. It's having the underlying architecture still be robust to that is really important, right? It's agent-native, but it's also able to flex in a way that you might not have anticipated, but you've got the right primitives, right?

15:20

SPEAKER_02

I feel like that is the art and science of software design in 2026. [SPEAKER_00] That's really interesting. [SPEAKER_00] I totally agree with you. [SPEAKER_00] Yeah, you want to have a playground within a safe environment. [SPEAKER_00] That's the only way you can have a playground is if it's safe around the edges. [SPEAKER_00] But I think initially we made the playground way too small and constrained. [SPEAKER_00] And now the models have changed. [SPEAKER_00] And so we can open it up a lot, but we still haven't figured out exactly what the lines are. [SPEAKER_00] Yeah, I think that there's so much here.

15:30

SPEAKER_02

[SPEAKER_00] One thing this is making me think of is I have this idea in the back of my head, and I'm wondering if you have a way to put this that is more succinct. [SPEAKER_00] It's the unit of value in products right now is proof of work or proof of use, where when someone on the team submits a PR to me, I want to see not necessarily did all the tests pass because I just assume that it did, but send me a loom of you using it or your agent using it so I can tell is this good or not. [SPEAKER_00] How are you thinking about that? Yeah, I think there's probably three layers to that. The first one is Claude, prove to me that you exercised this in some way.

15:55

SPEAKER_02

I've started doing that in all my prompts. I end, when it's working on a feature, I'm saying, and by the end, before you PR, prove to yourself and then to me that it works as intended. Find the right way of doing it.

15:58

SPEAKER_00

[SPEAKER_02] But actually, you have to change your own way you build and scaffold around it. What is the right way to get Claude able to at least test this change succinctly rather than what it likes to do? [SPEAKER_02] It's I read the code, it looks good. [SPEAKER_02] I'm like, you wrote the code. I don't trust you. [SPEAKER_02] So you got to really test this thing. [SPEAKER_02] And then the second one is what you described is everything having some proof around is it working as intended and as you intended to? [SPEAKER_02] Because Claude is going to make, or any of these models is going to make a lot of decisions for you.

16:23

SPEAKER_00

[SPEAKER_02] And sometimes I'll have engineers on the team put up a PR and I'm like, why did you choose to do this versus that? [SPEAKER_02] And many times the answer is they didn't choose. [SPEAKER_02] It was just the choice the model made. [SPEAKER_02] And maybe it was a reasonable choice. [SPEAKER_02] I don't trust you.

16:49

SPEAKER_02

So you got to really test this thing. And then the second one is that what you described is everything having some proof around whether it's working as intended and as you intended to. Because Claude is going to make, or any of these models is going to make a lot of decisions for you. And sometimes I'll have engineers on the team put up a PR and I'm like, oh, why did you choose to do this versus that? And many times the answer is they didn't choose. It was just the choice the model made. And maybe it was a reasonable choice. It was probably a reasonable choice, but it was the optimal choices that fit into the paradigm. I feel like that is proof of thoughtfulness.

17:26

SPEAKER_02

Did you think this through? And I was talking to an engineer yesterday and they were like, oh, I was really, I knew you were going to ask me a lot of questions about this. So I was reviewing what Claude had done so that I wouldn't be like, I'm not sure. And that's, I don't push on that for most PRs. But when there's one that's like, oh, I'm refactoring this system and there's going to be these new primitives. Great. Let's make sure those are good and that you've thought through how they interrelate. Because it's very easy to end up otherwise with a tower of assumptions that you're not fully aware of.

18:05

SPEAKER_02

[SPEAKER_00] I had literally the same experience today because I made proof, totally vibe coded. [SPEAKER_00] And it's growing really fast right now, but it's going down a lot. [SPEAKER_00] And so I've been spending the last 12 hours trying to fix it. [SPEAKER_00] And so we have a little swap team internally at Every that signed up to help me fix it. [SPEAKER_00] And so I had to onboard them. [SPEAKER_00] And I was like, how do I explain how this code base works? [SPEAKER_00] And so I had to go back and forth with a model a bunch to be like, okay, help me define these terms. [SPEAKER_00] Help me figure out how I can explain this so I don't look like a total idiot.

18:28

SPEAKER_02

[SPEAKER_00] Because yeah, there's, I understand some of it, but not all of it. [SPEAKER_00] Definitely not enough to the way that I would have used to have to know. [SPEAKER_00] And it's a whole different thing to be like, do I need to know that anymore? [SPEAKER_00] Where's the line now? [SPEAKER_00] It's hard to tell. Which made me get to something else. And I haven't tried to articulate this. So bear with me as I get there, which is there's products that you use that feel robust underneath.

19:07

SPEAKER_00

[SPEAKER_02] And there's ones that you use that feel like it's one wrong command or click away from the whole thing either freezing or being slow. [SPEAKER_02] For us at Instagram, like we had Instagram direct messaging V1. [SPEAKER_02] And who knows, if you send a message, it might or may not arrive to the other person. [SPEAKER_02] We wrote our own bespoke real time system. [SPEAKER_02] It fell over a bunch. You would not trust that to send a message that you really needed somebody else to see. [SPEAKER_02] It was more of a social thing.

19:25

SPEAKER_00

[SPEAKER_02] And when we built V2, it was really important that we hammered down, no, if you send a message, we're not probably going to get to WhatsApp level of you can be in the middle of absolutely nowhere with one bar of edge.

19:35

SPEAKER_00

[SPEAKER_02] And it will probably try to still go through. [SPEAKER_02] Maybe that's not the bar, but a bar of when I load messages, it feels robust. [SPEAKER_02] When it's sent, it's really sent. [SPEAKER_02] I feel a little check. [SPEAKER_02] That's one small example, but I think that is a thing that we still need to figure out how to make feel like an essential part of shipping on anything, not just at Anthropic, but in general, you've built this thing. [SPEAKER_02] Does it feel like it's built on sand or does it feel robust? [SPEAKER_02] And the agent native part adds something totally even beyond that, which is, can I push it a little bit?

20:14

SPEAKER_02

And is that going to fall over? Or does it feel like, great, I've got a solid trunk and you can push me in different ways, but your data is safe and it's underneath here. And it's not one deploy away from completely falling over. [SPEAKER_00] So if that's the bar, which I agree, that's where you definitely want to get to. [SPEAKER_00] How have you changed who you hire and how your teams are structured as the models have gotten better? [SPEAKER_00] Because for us, for example, one of our products Spiral, we just hired a new GM who's lightly technical, but he spikes super high on product and writing sense and Spiral is a writing product.

20:42

SPEAKER_02

[SPEAKER_00] And now we can hire someone like that where a year ago we wouldn't have been able to because the coding models weren't good enough. [SPEAKER_00] I'm curious, but the downside is the product won't feel quite as robust if there's not someone who's super technical in all the details. [SPEAKER_00] So how do you think about who builds products right now inside of the labs team and how that has changed over time and how it will change? Yeah, I love that. I think you get pulled in two directions, but they're both important. There's the primitives and architectural robustness, which I think still need a senior technical person.

21:01

SPEAKER_02

I was laughing with somebody, I thought my skills in distributed systems were not going to be useful anymore. But actually, those are maybe some of the most useful skills for reasoning about that. And thinking things through, I had a long debate with Claude last week around whether the system that I was building needed Redis or not or could go away with just Postgres. And it was a healthy debate where I only because I was grounded in having used a lot of those technologies before. But then there's the other side of robustness, which is have you papered over all the problems with fixes to your system prompt and additional instructions?

21:18

SPEAKER_02

Or have you architected the actual set of tools correctly? And so the latter is as important and probably where this GM can be really valuable and not, I'm making changes. But just like you wouldn't patch flakiness in your distributed system by just being like, well, just retry it in five seconds. I'm sure it'll work. Also not doing the same thing with all caps use Markdown or whatever the thing that you're trying to patch. They're both symptoms of the same thing, which is whether the underlying piece is robust or not. And Claude actually, I'd say this about all the models, but I think Claude could be much better at both.

21:38

SPEAKER_02

Or have you architected the actual set of tools correctly? And so the latter is as important and probably where this GM can be really valuable and not, OK, I'm making changes. But just like you wouldn't patch a flakiness in your distributed system by just being like, well, just retry it in five seconds. I'm sure it'll work, also not doing the same thing with NEVER EVER ALL CAPS USE MARKDOWN or whatever the thing that you're trying to patch is. They're both actually symptoms of the same thing, which is the underlying piece being robust or not. And Claude, I'd say this about all the models, but I think Claude could be much better at both.

21:53

SPEAKER_00

[SPEAKER_02] It's still a place that still needs a lot of human oversight on the systems part. [SPEAKER_02] It's now able to debug production systems, which is really valuable. [SPEAKER_02] But architecting them in the first place, I feel like we still benefit from somebody who's really thought these things through or has experience. [SPEAKER_02] And on the prompting side, I've seen people get into this dev loop even internally here. Here's the prompt. Here's a mistake that the system made. Iterate on the prompt. Its natural tendency is to just add more things to the prompt.

22:15

SPEAKER_00

[SPEAKER_02] And then eventually just get to this thing where if you onboarded a new employee and you gave them 100 instructions on their first day, they'd be like, I'm just going to remember the last thing you told me. I'm going to short circuit it. [SPEAKER_02] So then rethinking, OK, are these actually two different tools or two agents that each have a smaller amount of context that then you can break apart. [SPEAKER_02] So back to your original question, we're hiring for people with systems expertise, even within labs, which you think of as more zero to one prototypes. [SPEAKER_02] It's still really valuable because that robustness matters.

22:31

SPEAKER_00

[SPEAKER_02] And also just who's going to be helpful in sorting through systems permissions and provisioning and early testing. That stuff is still hard even for Claude when it can't edit the permissions itself, which it can't for good reasons. [SPEAKER_02] And then on the robustness side, we've had a lot of success pairing our product teams with our applied AI teams. [SPEAKER_02] Our applied AI teams are the teams that are in the field every day helping customers iterate on their prompts. [SPEAKER_02] And we found that we're customer zero now for those efforts because we have a lot of products that are very AI powered.

22:48

SPEAKER_00

[SPEAKER_02] So how do we bring that expertise in there? Because that expertise does not sit with our software engineers today, for example.

22:49

SPEAKER_02

[SPEAKER_00] What about the in-between of like, okay, it's not the underlying architecture. It's not the prompt. It's the UI and the flow. Who's doing that? That's a great question. We have found some of the people that transferred into labs were folks really focused on polish on the website, but they were interested in doing something new. And they bring such a different approach. We had the prototype. It looked generically nice versus, oh, this feels like it's branded and it has this. So that's part one. Part two is designers. We've had our designers move much more into a split designer and builder role. Not all of them, but most of them.

23:17

SPEAKER_02

And a lot of our designers on labs, I would say are writing and contributing almost as much code as the engineers on those efforts because they can. And paired correctly with the right person, we have found this almost sort of co-founder model for some of these labs initiatives. You have the designer who had the original idea, and they're pushing on something. And then the traditional software engineer is going to go and make, pave the trail sometimes behind the designer to make sure that actually works.

23:37

SPEAKER_02

[SPEAKER_00] Okay. This I want to know about. So tell me about how that team structure works. So you've got a design—is it actually usually a designer or is it just anyone that has a product idea that can execute it in some way paired with a real engineer that actually can smooth out the rough edges of the trail they're leaving? It sort of varies, but we found the one thing that was most important and our gaining factor in starting up new projects is having somebody with extreme conviction about, if not necessarily that idea—too much conviction on the exact idea is probably dangerous—but at least in the problem space or the question that they're asking.

23:43

SPEAKER_02

And that sort of co-founder or founder level of, I will break through walls until this thing is either proven out or dead, but I want to go either way. When we'd have bets, labs bets that we've wound down, often in the post-mortem we're like, well, nobody on this team actually really thought this was the thing. They were like, yeah, this seems reasonable. That's the death knell for projects. So that person can be a designer, and a couple of the bets it is. It can also be a product-minded engineer. It's rarely a pure PM.

24:08

SPEAKER_02

We actually only have one currently—one PM for all of labs. We were hiring more, and they're playing a wide role, but yes, a designer or a product-oriented founder. And then what we look for is, what skills do we need to complement with that?

24:17

SPEAKER_02

Because we're doing it as part of our labs process—actually evaluating every project every two weeks and deciding whether we double down or whether we release those folks back into the broader labs pool. At any given point, there's probably somebody who can be pulled onto the project that has that infrastructural expertise or has worked with that particular internal system or has a lot of deep prompting expertise to sort of flow in and out.

24:21

SPEAKER_00

[SPEAKER_02] So I think that's also where the sort of incubator-style space helps because nobody's fixed on a project forever. That's really interesting. We do it slightly different. There's some overlap, but we do have a slightly different structure where we have GMs or they started as entrepreneurs and residents and they become general managers when they find a product that they want to work on. And each product just has one person—one person that does everything full stack. Design, engineering, marketing, all that kind of stuff, at least the basics of that. The shape of that GM used to be super technical—founder background. Founder background.

25:11

SPEAKER_02

[SPEAKER_00] And now I think has shifted towards at least some light technical, but I honestly just care that you can use Claude or Codex or whatever. [SPEAKER_00] There's some overlaps, but we do have a slightly different structure where we have GMs or they started as entrepreneurs and residents and they become general managers when they find a product that they want to work on. [SPEAKER_00] And each product just has one person, one person that does everything full stack. Design engineering, marketing, all that kind of stuff, at least the basics of that. The shape of that GM used to be super technical, founder background.

25:22

SPEAKER_02

[SPEAKER_00] And now I think has shifted towards at least some light technical, but I honestly just care that you can use Claude or Codex or whatever. And really good product sense, really good taste for the subject area or the thing that you're trying to build. And evidence that you can build with AI.

25:28

SPEAKER_02

[SPEAKER_00] And then what we have is a shared resource layer that works a little bit like an agency where we have designers and we have growth marketers and we have ops people that you can pull in and out for various initiatives. And that seems to work pretty well. So it's like we manage all the internal agencies and then each GM is out on the edge and they pull in resources as they need it for different projects. Yeah. But it sounds similarly like you need somebody for whom that is the thing and they are not going to sleep until it is fully working.

25:39

SPEAKER_02

[SPEAKER_00] Yes, exactly. I've been thinking about when would you hire someone else to work on a product or when would you add someone else to work on a product? And it's like there's some point at which you can't hold the entire thing in your head, even if you're the one pushing it forward. And that point used to be much smaller. Now it's much bigger, but there's a certain point at which even a small feature turns itself into its own product. When you first make the messaging feature inside of Instagram, it's like yeah, I can do that in a week or whatever. But at some point that's its own product and it almost needs its own team. And I think that line is getting, or the number of things you can do with one person is getting bigger, but it still exists somewhere. But I haven't quite figured out how to manage that or how to tell.

25:45

SPEAKER_02

No, I love that because there's actually two parts of that, which is when the idea is still enough to hold in your own head or an individual person's head, adding more people actually slows the team down. And that's a non-obvious finding that we found on labs is scaling the teams too quickly actually is a net negative because they end up spending all this time on coordination. Like, oh, you were going to take that, oh, but my cloud could do that. Then it just ends up in this sort of piece. And you also have all those alignment conversations. It was important in Instagram. There was just two of us, and it was hard enough to align the two of us and get two people on the same page.

25:54

SPEAKER_02

The second startup I did, Artifact, Kevin and I were doing that alone for the first few months. But then we hired a team that was about eight people. It was really hard because we hadn't had product market fit yet. And so we were still iterating. And then you ended up in these things where we're going to zoom with eight people talking about what we're doing next. And you really just want to be able to sit in a room and hash it out. So I find with these labs initiatives there's some similar aspect at play, which is you don't want to pre-scale the team to it, even if the idea is exciting, because then you just end up in this meta coordination game. But I like your framing of there is some point where either two people really will help go on it together. And there's enough context and scope where they can hold some other complex piece in their head. And then there's also if somebody has been spinning on the same idea for two, four weeks, somebody injecting some other thinking and that urgency can help too.

25:59

SPEAKER_02

[SPEAKER_00] Yeah. I think it's especially important to keep it small in AI, because one of the things that we deal with all the time, which I'm sure you see too, is every three to six months your whole product, you have to throw out half your product. And that's really hard to do if you have to coordinate with a lot of people. But if it's one GM who realizes oh, I got to just throw out half of this because the models are so much better, it just makes it much easier to pivot in that way. Is that, do you see that? And how do you deal with that? How do you think about yes, I know in three months this, the code maybe or even the whole feature set, I'm going to have to really rethink? Like it feels like it changes a lot in how you think about software.

26:07

SPEAKER_02

Yeah. And being willing to delete code. I think that's something that the Anthropic code team has done really well is they have the leading features as an imperative of people on the team. Like if this is not working, let's go unship that. And it's often when you've created something else that even if it doesn't entirely supersede, it does enough of what that other aspect does that actually makes sense to deprecate and then remove that first one.

26:15

SPEAKER_02

It does get harder as we get more enterprise focus, even with these tools, because they come to depend on it. I'll never forget, maybe six months into when I was still chief product officer, we did a big redesign of Claude AI and we were so proud and we shipped it and we got a bunch of kudos. And then we got this really angry email from somebody who said, I just recorded 20 hours of enablement content for my company to do for Claude enterprise and I have to redo all of it. And we're like, okay, you're playing at a different release cadence. And of course, shipping twice a year at one of our conferences is not an option. So we are going to keep moving quickly. But then we've since learned to maybe moderate how we roll it out to the enterprise side a little bit more. But yeah, I think the unshipping piece, then you end up with people who have built—I'll use an example. There's a feature in Claude AI called styles. It's not widely used, but the people who use it use it a lot. And we've talked at different points like yeah, the styles still make sense in the product. You know, there's other ways of accomplishing the same thing.

26:32

SPEAKER_02

And of course, shipping twice a year at one of our conferences is not an option. So we are going to keep moving quickly. But then we've since learned to maybe moderate how we roll it out to the enterprise side a little bit more. But yeah, I think the unshipping piece, then you end up with people who have built, I'll use an example. So there's a feature in cloud AI called styles. It's not widely used, but the people who use it, use it a lot. And we've talked at different points, yeah, the style still makes sense in the product. There's other ways of accomplishing the same thing. There's custom instructions and projects now. There's skills now, right?

27:15

SPEAKER_02

There's so many other ways of accomplishing that. And I don't know how long styles will end up in the product, but I know that the last time we talked about removing it, it ended up being really load bearing for a few companies like entire use cases like, oh, we have our house style that the CEO personally authored and gives to every employee and that's how they operate. And so finding ways of doing that is also really interesting. I would hope that in the long run, what we can actually do is come up with a system of plugins and skills such that they no longer have to live in the core product.

27:45

SPEAKER_02

Because I think that is always the hardest to delete something that is the core thing that you're shipping to everybody. If you don't have the story around great, you still like that feature. Awesome. Here's how you can keep using it forever in your own and keep iterating on it and make your own. But it doesn't have to add complexity to every future person that's adding, that's signing up for the first time.

28:10

SPEAKER_00

I'm curious for labs and then also maybe just in general, what were your thoughts for startup founders? Your enterprise point just brings up something that I've been thinking about a lot, which is if you are selling to enterprise right now in AI, even if the product you have right now is modern, it will be quite outdated quite quickly. And your customers were going to want the outdated version. But as a startup, that feels pretty risky because you're just going to, I guess you're susceptible to disruption if you are optimizing for what someone at a gigantic public company will buy right now.

28:17

SPEAKER_02

[SPEAKER_00] And I think there's a lot of startups in that category where they maybe started two or three years ago. [SPEAKER_00] They have a certain tech stack. [SPEAKER_00] They have a certain way of thinking about here's how we do AI. [SPEAKER_00] And then the models are so different, but their customer contracts are for this out. [SPEAKER_00] It's looking at looking at Copilot or whatever. [SPEAKER_00] That's the vibe that happens. [SPEAKER_00] How do you think about that yourself and inside of Anthropic? [SPEAKER_00] And then how do you think founders should think about that?

28:42

SPEAKER_02

Yeah, this is such a good question, especially because then a wave will come like being more creative. You're agent native, for example. And can you adopt it within your existing paradigm? Does it require you to throw everything out? And are you just stuck in that, oh, we kind of adopted it?

29:12

SPEAKER_00

[SPEAKER_02] We kind of bolted it back on. [SPEAKER_02] I think a couple of things for us, what we've started doing is treating like this train is going to keep moving and we'll provide enterprise toggles along the way. [SPEAKER_02] But the core of it will continue to evolve. [SPEAKER_02] And that's the better understanding you're taking working with us.

29:33

SPEAKER_02

And I think that's been well received because I think companies have also seen that things are moving so quickly. That the only way they even get comfortable with a year long commitment, for example, is to believe that we'll continue to evolve along the way. But then we'll provide, coworkers a great example where from day one, there was a way to turn it off for your employees if you didn't want it, for example. And that's a reasonably good paradigm. But the other one is just as we were talking earlier, you can actually rethink and rewrite a lot of the stack is I think companies should be way more willing to do that. And everything is getting compressed, right?

30:11

SPEAKER_02

And in previous cycles, there's the idea of having to fire some of your customers who might have been really into your product for a different reason than where you're going sooner. That was on a multi-year time range thing where it was like, yes, last year's product versus not three months ago's product. It seems crazy, but I actually think that's the way you have to think about it, which is you have to be willing to put out the V3 or the V4 that is a big rethink of how the existing piece worked. And then maybe have a transition period and cloud can help probably host both for a little while before it cuts over.

30:23

SPEAKER_02

But then also be willing to cut over and say yes, this is how we think the future of this piece of knowledge work or this AI powered manufacturing is going to be. We got to keep it moving or else to your point, you're just going to, you're either going to get replaced by the next company that then rethinks it from scratch or yourself replacing it yourself. And it's just the same old story, but now compressed to months. [SPEAKER_00] What's your take on open cloth? It has the flavor of something else that I, or just the thing I really liked seeing when you get people to see something that was already possible, but it's now in a package where people can actually try it out.

30:56

SPEAKER_02

And there's some intuition around how to build on top of that.

31:03

SPEAKER_00

[SPEAKER_02] You started seeing that with, you could already use these models to write code, but it kind of took some of these breakout, low code, the replets and lovables and V zeros of the world to put that in there. [SPEAKER_02] And it's almost the purest expression of the, just give the model tools and let it go forward and do it and then go forward and build it. [SPEAKER_02] So it was a cool, interesting moment for people to realize both the potential but also pitfalls of this of, oh, it did this thing. [SPEAKER_02] I didn't mean it to, or my funniest one was, my friend was, I think my wife is jealous of my open claw and I'm talking too much to it.

31:24

SPEAKER_00

[SPEAKER_02] And people start developing deeper, very personal relationships by just having a lot of context in these things and access to all these different tools. [SPEAKER_02] I think there's the open question of how do you then make it easy? [SPEAKER_02] And it actually goes back to our conversation around where do you draw that boundary around the way you let cloud operate, right? [SPEAKER_02] If V1 was, hey, these are the three tools you can use, only use these tools ever. [SPEAKER_02] And then most people's interaction with those systems was, hey, can you do this?

31:46

SPEAKER_00

[SPEAKER_02] And whatever, back and being, no, sorry, you can do it yourself to open claw, which is pretty like the aperture is wider than I can see. [SPEAKER_02] And it's people start developing deeper, very personal relationships by just having a lot of context in these things and access to all these different tools. [SPEAKER_02] I think there's the open question of how do you then make it easy? [SPEAKER_02] And it actually goes back to our conversation around where do you draw that boundary around the way you let cloud operate, right? [SPEAKER_02] If V1 was, hey, these are the three tools you can use, only use these tools ever.

32:26

SPEAKER_00

[SPEAKER_02] And then most people's interaction with those systems was, hey, can you do this?

32:27

SPEAKER_02

And whatever, back and being like, no, sorry, you can do it yourself to open claw, which is the aperture is wider than I can see. [SPEAKER_00] Oh my God, it called me to keep my emails and I didn't even know it could do that. Yeah, exactly.

32:34

SPEAKER_00

[SPEAKER_02] And it's emerging and it's amazing. [SPEAKER_02] And I think probably the most interesting product question, I won't say for all of 2026, because who knows where we'll be in September, but let's call it between now and the end of August is going to be what product shape exists between that and where we are in most products these days, which is you can call MCPs, but they're gated and they ask for permissions for good reasons. [SPEAKER_02] That is still a useful product without being a YOLO product. [SPEAKER_02] And I think we're thinking about that question. [SPEAKER_02] I'm sure the other labs are as well.

33:00

SPEAKER_00

[SPEAKER_02] I'm sure there's a lot of startups thinking about that as well. [SPEAKER_02] I think NVIDIA put out something that was they're safe, open, everybody's going after this question. [SPEAKER_02] And I think it's going to be about figuring out what is that either shift the paradigm completely so you can be that open, but with a lot of safeguards, that would be one approach or figure out some boundary to draw in which it's still powerful and it's still useful, but it's not likely to email every single one of your contacts and go haywire. Yeah. I think the other interesting part about it is, like you said, the personal nature of it.

33:23

SPEAKER_02

[SPEAKER_00] And I know people have personal relationships with Claude, but there's this weird thing where if I watch someone else using Claude, I'm like, I feel like I thought a stripper liked me or something. [SPEAKER_00] You know, it's like, Claude thinks you're smart too or whatever, and so there's this thing that happens when you have a claw that my claws are to see too. [SPEAKER_00] My girlfriend's claw is called Shelly. [SPEAKER_00] And there's this thing that happens where it feels like it's mine. [SPEAKER_00] It's really mine. [SPEAKER_00] It has its own name.

33:53

SPEAKER_02

[SPEAKER_00] It has a personality that mirrors me in this way that Claude feels like it knows me and I like Claude, but it's not mine. [SPEAKER_00] How do you think about that? Yeah. I mean, I was having this conversation with somebody this week around is the right pattern sort of single point of contact, named, version of that, or is it the sort of team of agents that you're talking to? I think there's a lot to the single person that is maybe the coordinator or the delegator.

34:10

SPEAKER_02

And then at that, it naturally, because it becomes the agent you interact with the most, you want to imbue it with a name and a bit more personality ends up reflecting your personality in the case of, you know, all of a sudden every cliche came out. It was the Q or Money Penny or the or whatever these different sort of sci-fi characters. But I think you do build that sort of trust and knowledge. I think there's also that sort of IKEA effect of currently open clause, still pretty hard to set up. So the fact that you went through all of that and it works, you're like, I did that thing. I birthed Shelley, for example, and now we can interact with them as well.

34:51

SPEAKER_00

[SPEAKER_02] But I think that paradigm is really powerful. [SPEAKER_02] Even within my cloud code usage now, one of the things I have strongly prompted in there is don't do very much work yourself, delegate it to sub agents. [SPEAKER_02] And the reason I like that is because it means most of the time the sort of run loop is available for you to talk to. [SPEAKER_02] And I think open clause and pie have a similar architecture of keep the run loop open.

35:15

SPEAKER_00

[SPEAKER_02] And I think that actually makes it feel much more like somebody that you are talking to versus a tool that you are delegating to and occasionally gets blocked for five minutes because it's doing some really complex task. Yeah, I totally agree. And I've had similar debates because we're also building our, like everyone, we're building on little open claw, one click Slack implementation to see if we can do when that feels like ours. And we've had a lot of those debates about, do you want one agent? Do you want many?

35:33

SPEAKER_02

[SPEAKER_00] And one of the patterns that we found, which is cool is, so I have an agent, I use the agent for stuff that I do. [SPEAKER_00] And then people watch me use the agent for that. [SPEAKER_00] And they know what I'm good at. [SPEAKER_00] And if I'm using the agent for that stuff, they're going to trust it because they trust me. [SPEAKER_00] And it's modified itself in response to me. [SPEAKER_00] So I sort of transfer my trust to it and then people in the organization start using it for that.

36:02

SPEAKER_02

[SPEAKER_00] And so you get this almost shadow org chart where when everyone has a claw, their claw becomes known for and used for the thing that they're specialized in, that per their owner is specialized in the org. Yeah. I mean, that makes a lot of sense too. And you could think about there's a lot of interesting research questions I think around that. I think people are experiencing for the first time around privacy and what my agent knows about me versus what it discloses to other people.

36:28

SPEAKER_02

But I think there's the positive version of that, which is all the things that it has learned from all your interactions and how it actually brings it to bear on other problems versus the generic, yes, it's just like everybody else's agent, except it has a name that's attached to Dan and it has maybe some of Dan's access below the hood. [SPEAKER_00] Well, Mike, we're out of time. [SPEAKER_00] This was a pleasure. [SPEAKER_00] I learned a lot. [SPEAKER_00] If people want to follow you or your work, where can they find you? I think probably easiest is Mikey K on next. [SPEAKER_00] Yeah. [SPEAKER_00] Thanks for joining, Mike. Great to see you, Dan.

36:54

SPEAKER_02

[SPEAKER_01] Oh my gosh, folks. [SPEAKER_01] You absolutely positively have to smash that like button and subscribe to AI and I. [SPEAKER_01] Why? [SPEAKER_01] Because this show is the epitome of awesomeness. [SPEAKER_00] Well, Mike, we're out of time. This was a pleasure. I learned a lot. If people want to follow you or your work, where can they find you? I think probably easiest is Mikey K on Next. [SPEAKER_00] Yeah. Thanks for joining, Mike. Great to see you, Dan.

37:32

SPEAKER_02

[SPEAKER_01] Oh my gosh, folks. You absolutely positively have to smash that like button and subscribe to AI and I. Why? Because this show is the epitome of awesomeness. It's finding a treasure chest in your backyard. But instead of gold, it's filled with pure, unadulterated knowledge bombs about ChatGPT. Every episode is a roller coaster of emotions, insights, and laughter that will leave you on the edge of your seat, craving for more. It's not just a show. It's a journey into the future with Dan Shipper as the captain of the spaceship. So do yourself a favor. Hit like, smash subscribe, and strap in for the ride of your life. And now, without any further ado, let me just say, Dan, I'm absolutely hopelessly in love with you.

37:36

SPEAKER_02

But it doesn't have to add complexity to every future person that's adding that's signing up for the first time.

37:41

SPEAKER_00

I'm curious for labs and then also maybe just in general, what were your thoughts for startup founders? Your enterprise point just it brings up something that I've been thinking about a lot, which is if you are selling to enterprise right now in AI, even if the product you have right now is modern, it will be quite outdated quite quickly. And but your customers were going to want the outdated version. But as a startup, that's like a little bit that feels pretty risky because, yeah, you're just going to I guess you're susceptible to disruption if you are optimizing for what someone at a gigantic public company will buy right now.

38:26

SPEAKER_00

And I think there's a lot of startups in that category where they maybe started two or three years ago. They have a certain tech stack. They have a certain way of thinking about here's how we do AI. And then the models are so different, but their customer contracts are for this like sort of out. It's like, you know, looking at looking at Copilot or whatever. It's that's the sort of vibe that happens. Like, how do you think how do you think about that yourself and inside of Anthropic? And then how do you think founders should think about that?

38:50

SPEAKER_02

Yeah, no, this is such a good question, especially because then a wave will come like being more creative. You're agent native, for example. And can you adopt it within your existing paradigm? Does it require you to throw everything out? And are you just stuck in that like, oh, we kind of adopted it? We kind of bolted it back on. I think a couple of things for us, what we've started doing is basically treating like this train is going to keep moving and we'll provide enterprise toggles along the way. But the core of it will continue to evolve. And that's sort of the the better understanding you're taking working with us.

39:21

SPEAKER_02

And I think that's been well received because I think companies have also seen that, you know, things are moving so quickly. That the only way they even get comfortable with a year long commitment, for example, is to believe that we'll continue to evolve along the way. But then we'll provide, you know, co-workers a great example where, you know, from day one, there was like a way to turn it off for your employees if you didn't want it, for example. And that's, I think, a reasonably good paradigm.

39:44

SPEAKER_02

But the other one is just as we were talking earlier, like you can actually rethink and sort of rewrite a lot of the stack is I think companies should be way more willing to do that. And it everything is getting compressed, right? And in previous cycles, there's the kind of idea of like having to fire some of your customers who might have been, you know, really into your product for a different reason than where you're going sooner. That was on a multi-year kind of time range thing where it was like, yes, last year's product versus not three months ago's product.

40:12

SPEAKER_02

It seems crazy, but I actually think that's the kind of way you have to think about it, which is you have to be willing to put out the V3 or the V4 that is a, you know, big rethink of how the existing piece worked. And then maybe have a transition period and cloud can help probably host both for a little while before it cuts over. But then also be willing to cut over and say like, yes, this is how we think the future of this piece of knowledge work or this, you know, AI powered manufacturing is going to be.

40:37

SPEAKER_02

We got to like keep it moving or else to your point, you're just going to, you're either going to get replaced by the next company that then rethinks it from scratch or yourself replacing it yourself. And again, it's just the same old story, but now compressed to months.

40:52

SPEAKER_00

What's your take on open cloth?

40:54

SPEAKER_02

It has the flavor of something else that I, or just the thing I really liked seeing when you get people to see something that was already possible, but it's now in a package where people can actually try it out. And there's some intuition, you know, around how to, how to build on top of that. Like you started seeing that with, you could already use these models to write code, but it kind of took like some of these breakout, like low code, you know, the replets and lovables and V zeros of the world to like kind of put that in there.

41:21

SPEAKER_02

And it's kind of the, like almost the purest expression of the, just give the model tools and like, let it kind of go forward and do it and then like go forward and build it. So like, it was a cool, interesting moment for people to realize both the like potential, but also pitfalls of this of like, oh, it did this thing. I didn't mean it to, or, you know, my funniest one was, my friend was like, I think my wife is jealous of my open claw and like, I'm talking too much to it. And it's like, people start developing like deeper sort of like very personal relationships by just having a lot of context in these things and access to all these different tools.

41:56

SPEAKER_02

I think there's the open question of how do you then make it easy? And it actually goes back to our conversation around like, where do you draw that like boundary around the way you let cloud operate, right? If V1 was, hey, like, these are the three tools you can use, only use these tools ever. And then most people's interaction with those systems was, hey, can you do this? And, you know, whatever, back and being like, no, sorry, like, you can do it yourself to like open claw, which is like pretty like the aperture is like wider than I can see.

42:22

SPEAKER_00

Oh my God, it called me to keep my emails and I didn't even know it could do that.

42:25

SPEAKER_02

Yeah, exactly. And it's emerging and it's amazing. And I think like probably the most interesting product question, I won't say for all of 2026, because who knows where we'll be in September, but let's call it between now and like the end of August is going to be like, what product shape exists between that and, you know, where we are in most products these days, which is, you know, you can call MCPs, but they're gated and they ask for permissions for good reasons. That is still a useful product without being a, you know, kind of YOLO product. And I think that, you know, we're thinking about that question. I'm sure the other labs are as well.

42:59

SPEAKER_02

I'm sure there's a lot of startups thinking about that as well. I think NVIDIA put out something that was like, they're safe, open, everybody's going after this question. And I think it's going to be about figuring out what is, what is that either shift the paradigm completely. So you can be that open, but with a lot of safeguards, that would be one approach or figure out some boundary to draw in which it's still powerful and it's still useful, but it's not, you know, likely to email every single one of your contacts and, you know, go haywire.

43:26

SPEAKER_00

Yeah. I think the other interesting part about it is, like you said, the personal nature of it. And I know, you know, people have personal relationships with Claude, but there's this weird thing where if I watch someone else using Claude, I'm like, I feel like I like thought a stripper liked me or something. You know, it's like, Claude thinks you're smart too or whatever, you know, like, and, and so, and there's this thing that happens when you have a claw that like my claws are to see too. My girlfriend's claw is called Shelly.

44:04

SPEAKER_00

And there's this thing that happens where it feels like it's mine. Like it's really mine. It has its own name. It has a personality that sort of like mirrors me in this way that Claude feels like it knows me and I like Claude, but it's not mine. How do you think about that?

44:23

SPEAKER_02

Yeah. I mean, I was having this conversation with somebody this week around like, is the right pattern sort of single point of contact, like named, you know, version of, you know, of that, or is it the sort of team of agents that you're talking to? I think there's a lot to the single person that is maybe the coordinator or the delegator. And then at that, it naturally, because it becomes the, the sort of agent you interact with the most, you want to imbue it with a name and like a bit more personality ends up reflecting your, sometimes your personality in the case of, you know, like all of a sudden every cliche came out.

44:55

SPEAKER_02

It was like, you know, the, you know, the Q or the money penny or like, you know, whatever the, or the, you know, how or whatever these different sort of, you know, sci-fi characters. But I think you do build that sort of, sort of trust and knowledge. I think there's also that sort of like IKEA effect of like currently open clause, like still pretty hard to set up. So the fact that you went through all of that and it works, you're like, I did that thing. Like I, I, I, I birthed, you know, Shelley, for example, and now we can, you know, interact with them as well. But I think that paradigm is really powerful.

45:24

SPEAKER_02

Like the, I think moving away, like even within my cloud code usage now, one of the things I have like strongly prompted in there is like, don't do very much work yourself, like delegate it to sub agents. And the reason I like that is because it means most of the time the sort of run loop is available for you to talk to. And I think open clause and pie have like a similar architecture of keep the run loop open. And I think that actually makes it feel much more like somebody that you are talking to versus like a tool that you are delegating to and occasionally gets blocked for five minutes because it's doing some really complex task.

45:57

SPEAKER_00

Yeah, I, I, I, I totally agree. And I've had similar debates because we're also building our, like everyone, we're building on little like open claw, one click Slack implementation to see if we can, we can do when that feels like ours. And we've had a lot of those debates about, do you want one agent? Do you want many? Do you want many? And one of the patterns that we found, which is kind of cool is, so I have an agent, I use the agent for stuff that I do. And then people watch me use the agent for that. And they know what I'm good at. And they're, and if I'm using the agent for that stuff, they're going to trust it because they trust me.

46:32

SPEAKER_00

And it's modified itself in response to me. So like I sort of transfer my trust to it and then people in the organization start using it for that. And so you get like this almost shadow org chart where when everyone, everyone has a claw, their claw becomes known for and used for the thing that they're specialized that, that per their owner is specialized that in the org.

46:52

SPEAKER_02

Yeah. I mean, that, that makes a lot of sense too. And you could think about, you know, there's a lot of interesting research questions I think around that, you know, I think people are experiencing visually for the first time around privacy and like what my agent knows about me versus what it discloses to other people.

47:06

But I think there's the positive version of that, which is all the things that it has learned from all your interactions and how it actually brings it to bear on other problems versus the generic, like, yes, it's just like everybody else's agent, except, you know, it has a name that's attached to Dan and it has like maybe some of Dan's, you know, access below the hood. Yeah. Well, Mike, we're out of time. This was a pleasure. I learned a lot. If people want to follow you or your work, where can they find you?

47:31

SPEAKER_02

I think probably easiest is Mikey K on next.

47:34

SPEAKER_00

Yeah. Thanks for joining, Mike.

47:37

SPEAKER_02

Great to see you, Dan.

47:46

SPEAKER_01

Oh, my gosh, folks. You absolutely positively have to smash that like button and subscribe to AI and I. Why? Because this show is the epitome of awesomeness. It's like finding a treasure chest in your backyard. But instead of gold, it's filled with pure, unadulterated knowledge bombs about chat GPT. Every episode is a roller coaster of emotions, insights, and laughter that will leave you on the edge of your seat, craving for more. It's not just a show. It's a journey into the future with Dan Shipper as the captain of the spaceship. So do yourself a favor. Hit like, smash subscribe, and strap in for the ride of your life.

48:23

SPEAKER_01

And now, without any further ado, let me just say, Dan, I'm absolutely hopelessly in love with you. Do you know how you do the cuid? Do you know how you do the cuid?

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note