Lenny's Podcast

How Anthropic’s product team moves faster than anyone else | Cat Wu (Head of Product, Claude Code)

2583 summary words 11 min summary Watch video

Start with the signal

11 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Anthropic’s product velocity comes less from frontier-model access than from an operating system that gives high-agency, product-minded builders clear goals, low-friction preview launches, tight cross-functional launch loops, and continuous model/harness learning.
  • Why it matters: It offers unusually concrete patterns for building and operating agent products: task-level evals, harness simplification as models improve, research previews, connected-context workflows, and a control-plane vision for supervising many remote agents.
  • Best use: Watch for reusable operating principles and product architecture decisions for OpenClaw or other agent systems; extract the launch process, eval practice, feedback loops, and multi-agent control-plane implications.

Executive Summary

Cat Wu describes Claude Code’s product model as a division of labor between a technical product visionary—Boris Cherny—and a PM role that turns a three-to-six-month agentic vision into an executable path while clearing organizational blockers. Her central claim is that AI-native product teams must optimize for the elapsed time between an idea and real user feedback, not for polished multi-quarter roadmap alignment. At Anthropic, feature timelines have compressed from months to weeks or days, and most Claude Code features ship first as explicitly labeled Research Previews.

The operating model behind this pace is concrete: define narrow user/problem/outcome goals; give the entire team rigorous weekly metrics and durable decision principles; dogfood features; then use an “evergreen launch room” where engineering, docs, product marketing, and DevRel can turn a ready feature into a public release by the next day. The PM’s job is therefore less specification ownership and more creating the conditions in which engineers with product taste can independently detect feedback, build, launch, and learn.

Wu’s product philosophy is deliberately grounded in current-model reality rather than vague AGI assumptions. The difficult PM skill is to identify the current model’s golden path, expose its capabilities, and patch its weaknesses through product and harness design. Teams should interrogate unexpected agent behavior, recruit a small set of highly diagnostic users, and build a small number of high-quality evals. As models improve, they should delete obsolete prompting scaffolding rather than accumulate it; meanwhile, prototypes that previously failed can become viable products as model reliability crosses a threshold.

For agent operations, Wu outlines a progression from one successful task, to several concurrent tasks, to tens or hundreds of remotely executed agents. That future requires a human control plane that surfaces exceptions, provides trustworthy completion verification, and incorporates feedback so repeated mistakes do not recur. Cowork is presented as an early instance of the broader pattern: connect authoritative communications and data sources, have the agent synthesize and execute across them, and retain humans for selecting goals, deciding among options, and reviewing consequential output.

Key Takeaways

  • Claim: AI-native product management should optimize for idea-to-user feedback speed rather than traditional multi-quarter coordination. | Evidence: Wu says Anthropic feature timelines have fallen from six months to one month, and sometimes to one week or one day; Claude Code aims for any team member to take an idea to the world in under a week, sometimes a day. | Implication: Ken should design product processes around small releasable increments, rapid user exposure, and reversible commitments rather than waiting for suite-wide completeness. | Caveat: Long-running, infrastructure-heavy work still requires multi-month planning and more conventional PRDs.
  • Claim: Fast shipping depends on explicit constraints and repeatable launch interfaces, not merely better coding models. | Evidence: For a permissions feature, Wu frames a precise target—professional enterprise developers safely reaching zero permission prompts—rather than a generic goal of reducing prompts. Claude Code labels most releases Research Preview, and an evergreen launch room lets docs, PMM, and DevRel prepare a launch announcement the day after engineering dogfoods a feature. | Implication: Create a preview lane with defined entry criteria, ownership, documentation/marketing turnaround expectations, and an explicit promise that functionality may change or disappear. | Caveat: Research Preview reduces commitment and polish requirements, but it shifts responsibility to clear labeling, rapid support, and active iteration.
  • Claim: The durable PM advantage is product taste: deciding what to build and how users should encounter it, while roles merge around execution. | Evidence: Wu says engineers do PM work, PMs code, and designers land code; Claude Code favors engineers who can move from a user complaint on Twitter to a shipped feature in a week. Nearly all PMs on the team either were engineers or ship code there, but Wu says taste can originate from any background. | Implication: Hire and develop generalists with demonstrated judgment, technical fluency, user empathy, and low ego—not narrowly scoped role specialists dependent on handoffs. | Caveat: Engineering fluency currently helps prioritization because it reveals implementation cost, but Wu expects valued skills to shift frequently as model capabilities change.
  • Claim: Agent-product teams need a continuous model-and-harness learning loop built on behavioral investigation, trusted users, and small targeted eval suites. | Evidence: Wu asks models why they made unexpected decisions—for example, changing front-end code and running tests without using the UI—to identify whether the failure came from a confusing system prompt, a missed requirement, or unchecked subagent work. She recommends finding roughly five highly articulate users and building even 10 strong evals; she personally uses sets of about five evals when a feature needs sharper definition. | Implication: For OpenClaw workflows, maintain a small regression suite for critical tasks, explicitly log failure causes, and use expert operators to generate test hypotheses rather than relying on aggregate feedback alone. | Caveat: Not every feature needs formal evals, though behaviorally sensitive features such as memory benefit substantially; human feedback must be treated as hypothesis generation and checked against data.
  • Claim: As base models improve, teams should remove harness interventions that are no longer needed while keeping prototypes of capabilities that are not yet reliable. | Evidence: Early Claude Code models needed repeated reminders to complete every item in a to-do list during broad refactors; later models naturally completed the list, so the scaffold became unnecessary. Anthropic reviews its entire system prompt at every model launch and removes obsolete reminders. Conversely, code review was attempted repeatedly before Opus 4.5/4.6 and Sonnet 4.6 made multi-agent whole-codebase review reliable enough for internal merge gating. | Implication: Treat prompts, checklists, and orchestration rules as versioned, temporary compensations. Re-evaluate them after model upgrades, and preserve near-ready prototypes so capability jumps can be exploited quickly. | Caveat: A newer model does not automatically validate a feature; the team still needs to test whether it closes the specific reliability gap that blocked release.
  • Claim: The agent-product end state is not a better chat interface but a remote, multi-agent control plane with verification and learning. | Evidence: Wu frames the progression as successful individual tasks, then six concurrent tasks, then 50 or hundreds of agents. She expects this scale to require remote execution rather than local laptops, interfaces that identify which tasks require human attention, trustworthy agent verification, and feedback that prevents recurrence of a mistake. | Implication: Ken’s agent architecture should prioritize dispatch, state visibility, exception routing, outcome verification, and reusable feedback/memory over simply adding more autonomous workers. | Caveat: This is an articulated product direction rather than a committed roadmap or proof that hundreds of agents can be supervised reliably today.
  • Claim: AI leverage becomes operationally meaningful only when it automates real recurring work to near-perfect reliability; tinkering and 95% automations can create more overhead than value. | Evidence: Wu uses Cowork to build a 20-page conference deck by connecting Slack, Google Drive, calendar, Gmail, prior drafts, and an Anthropic slide template; it produced a polished draft after she selected the narrative and gave feedback. She argues that an automation that works only 90–95% of the time is not truly an automation, citing her own still-imperfect Inbox Zero workflow. | Implication: Choose a few high-frequency, high-cost workflows, connect authoritative context sources, define acceptance criteria, invest through the final reliability gap, and measure whether the automation actually frees capacity. | Caveat: She also warns against the opposite failure mode: excessive customization of skills, MCPs, and workflows that distracts from shipping the underlying product or completing the original task.

Detailed Brief

Organizational design and operating norms behind Anthropic’s product pace

  • Claims: Wu attributes Anthropic’s ability to make cross-org decisions to a mission that sits above individual product or team goals.; The key distinction is not simply focus: teams are willing to sacrifice their own KRs or product outcomes if doing so improves Anthropic-level outcomes.; The team deliberately tolerates some roughness and product overlap because fast external learning is more valuable than premature consolidation.
  • Evidence: Wu says she would be happy if Claude Code failed but Anthropic succeeded, and describes this as a shared orientation rather than an individual sentiment.; Claude Code maintains weekly whole-team metrics readouts and team principles covering target users, tradeoffs, and business logic so people can act without waiting on PM approval.; Anthropic may launch overlapping form factors because internal teams like both and want external users to reveal which is better.; Wu says the cost of this pace is reduced product consistency, harder newcomer navigation, and a sense that users must follow daily updates; Anthropic added the /power-up onboarding flow after initially believing no tutorial should be necessary.
  • Caveats: Launching imperfect features is acceptable in this model only when they do not block the product’s core use case and the team can rapidly hear and act on feedback.; High release frequency creates an onboarding and education burden, especially for enterprise users accustomed to slower, more stable product cycles.
  • Implications: A high-velocity agent platform needs a stable orientation layer—principles, metrics, and onboarding—even if individual features remain fluid.; Use portfolio-level decision rights to avoid local optimization by agent, product, or functional teams.

Cowork pattern: connected context plus human editorial judgment

  • Claims: Cowork’s value depends heavily on connecting the communication systems and source-of-truth repositories relevant to a user’s role.; The agent can research, synthesize, and produce a substantial first draft, while the human remains responsible for narrative selection and final product judgment.; The same pattern supports customer-facing operations, not only PM work.
  • Evidence: Wu recommends connecting Google Calendar, Slack, Gmail, and Google Drive; for design-aware deliverables, she provided an existing standardized external deck, and notes a Figma MCP can provide equivalent design context.; For her conference deck, she asked Cowork first to create an outline from a PMM proposal, a disliked manual draft, linked sources, and a constraint not to overlap with a more important keynote; she then selected the story arc before the agent built the deck.; A sales employee built a custom web application that combines Salesforce, Gong, and notes with standard Claude Code decks to create customer-specific decks in seconds instead of spending 20–30 minutes manually editing them.; Applied AI staff covering five to 10 customer engagements in a day use Cowork to prepare an overnight dossier of upcoming meetings, customer requests, prior action items, and researched internal answers such as a feature ETA from Slack.
  • Caveats: Connected-context systems increase output relevance but also raise the importance of access control, data boundaries, source quality, and review for customer-facing claims.; Wu’s deck was not a one-shot final: it needed feedback for wordiness and human selection of what belonged in the final narrative.
  • Implications: Prioritize agent access to governed, authoritative internal sources and reusable artifact templates before investing in generic generation.; The most valuable deployments are likely role-specific workflows with a clear repeated input-to-output path, rather than generalized chat usage.

Product, platform, and market signals

  • Claims: Anthropic’s product organization was described as roughly 30–40 PMs across research PM, developer platform, Claude Code/Cowork, enterprise, and growth.; Enterprise product work centers on adoption enablers—cost controls, RBAC, security controls—and is separate from the fast feature experimentation posture.; Anthropic decided to prioritize first-party Claude products and API capacity over third-party products using Claude subscriptions, including OpenClaw-style usage patterns.
  • Evidence: Research PM collects model customer feedback and helps shepherd launches; the developer platform owns APIs and managed agents; enterprise focuses on controls; growth operates across the suite.; Wu says subscriptions were not designed for third-party products with different usage patterns, and Anthropic offered some credits during the transition while prioritizing first-party products and the API.; She says internal token costs rise after model jumps or major product improvements as workers delegate more tasks, though average per-worker token cost remains below average engineer salary.
  • Caveats: The comments on third-party subscription use are Anthropic’s capacity and product-prioritization rationale, not a general rule that third-party agent products lack value.; No precise token-spend figures, adoption rates, or model performance data were provided.
  • Implications: Do not build a strategic workflow dependency on subsidized consumer subscriptions where the provider’s economics or terms can change; use appropriately provisioned APIs and control spend explicitly.; Enterprise readiness and agent velocity are complementary but should have distinct release gates, especially around authorization, auditability, and cost controls.

Notable Concepts & Terms

  • Research Preview: Claude Code’s explicit early-release label, used to lower commitment, get features into users’ hands quickly, and preserve the ability to change or discontinue them.
  • Evergreen launch room: A standing engineering/docs/PMM/DevRel handoff process through which dogfooded features can receive documentation and announcement support by the next day.
  • Product taste: The scarce cross-functional ability to decide what is worth building, identify the correct UX and user path, and balance value against implementation cost as code becomes cheaper.
  • Harness: The prompting, workflows, tools, guardrails, and orchestration wrapped around a base model to elicit reliable behavior; it should be revised and often simplified as models improve.
  • Evals: Concrete task tests that define success and measure behavioral progress; Wu argues even a small set of strong evals can sharpen product definition and guide engineering.
  • Golden path: A product interaction path designed around what the current model does well while compensating for its present weaknesses, rather than designing only for hypothetical AGI.
  • Multi-clotting: Running multiple Claude tasks concurrently; Wu uses it as the transition point from a single-task assistant toward a scalable multi-agent operating model.
  • Applied AI: Anthropic’s technical customer-facing function that helps customers adopt API/model capabilities for products and internal acceleration, resembling forward-deployed technical GTM.

Operator Notes / Why Ken Should Care

  • Establish a Research Preview track for agent capabilities: explicit label, limited support contract, rapid rollback path, dogfood requirement, and predefined feedback capture.
  • Create a 10-or-fewer-task eval suite for OpenClaw’s highest-value workflows, including completion verification and failure categories; rerun it after every model, tool, or orchestration change.
  • Audit existing prompts, guardrails, and multi-step agent scaffolding after each meaningful model upgrade; remove interventions that have become redundant and identify prototypes newly viable at the upgraded capability level.
  • Design the agent control plane around exception queues and evidence-backed completion, not merely parallel task dispatch; define what a human must inspect when 50+ tasks are active.
  • Choose two recurring internal workflows with authoritative source data—such as meeting preparation, customer-deck generation, or inbox triage—and push one to reliable production use before expanding the automation catalog.
  • Treat third-party model subscription access as non-durable infrastructure. For production agent use, validate API terms, capacity, rate limits, attribution, authentication, and spend controls.

Source/Metadata

  • Title: How Anthropic’s product team moves faster than anyone else | Cat Wu (Head of Product, Claude Code)
  • Transcript words: 25958
  • Duration seconds: 5134
  • Timestamp note: No usable timestamps or chapter markers were present in the supplied transcript; the transcript includes substantial duplicated passages and sponsor segments.
Full transcript 15740 words · 114 min read
0:00

I think it is very hard to be the right amount of AGI-pilled. It's very easy to build the product for the super AGI-strong model. The hard thing is figuring out, for the current model, how do you elicit the maximum capability? I've never seen anything like the pace you folks at Anthropic are shipping at. We want to remove every single barrier to shipping things. The timelines for a lot of our product features have gone down from six months to one month and sometimes to even one day. You're interviewing hundreds of PMs, and you just keep feeling like they're approaching it very incorrectly. The PM role is changing a lot. It's changing really quickly.

0:20

The thing that is extremely important for building AI-native products is iterating so quickly, figuring out a way for you to actually launch features every single week. What do you think are the emerging skills PMs need to develop? It comes back to product taste. As code becomes much cheaper to write, the thing that becomes more valuable is deciding what to write. Today, my guest is Kat Wu, head of product for Cloud Code and co-work at Anthropic. Kat is at the center of everything that is changing in AI and product and building. And she and her team are building the product that is most changing the way that we all build our products.

0:35

She is so full of insights and wisdom and lessons. This is an episode you cannot miss. Before we get into it, don't forget to check out Lenny'sProductPass.com for an insane set of deals available exclusively to Lenny's newsletter subscribers. With that, I bring you Kat Wu. Kat, welcome to the podcast. Thanks for having me. I have so many questions. I'm so excited to have you on this podcast. I want to start with giving people an understanding of your role alongside Boris. Boris. Everybody knows Boris. His episode is the number one most popular episode on this podcast. No pressure.

0:56

He created Cloud Code. He leads the Ench team, ships a bazillion PRs a day from his phone. I don't even know what the number is anymore. I think people don't give you enough credit for the success that Cloud Code has had, and co-work and all the things you all are building. Help us understand your role on the team, how you work with Boris, how you split responsibilities, what the PM role looks like on the Cloud Code team.

1:06

I feel very lucky to work with Boris. He's been an amazing thought partner. He's our tech lead. He's very much the product visionary. And he is great at setting, this is what the product needs to be in three months, six months from now. This is what the AGI-pilled version of the product is. And a lot of my role is figuring out, okay, what is the path from where we are today to that vision three to six months from now?

1:12

And I spend more of my time on the cross-functional, so making sure that our marketing team, sales team, finance, capacity, et cetera, are bought in on the plan and that we're all rowing in the same direction, and that once the feature is ready, there aren't any blockers to shipping it. I think in many ways it works well because we kind of mind-meld, but it is actually remarkably blurry of a line. I think we're 80% mind-meld. And then there's this 20% of things that maybe I care a lot more about than Boris, so I'll drive those. And 20% where he cares a lot more than me, and he just drives those. This episode is brought to you by our season's presenting sponsor, WorkOS.

1:19

What do OpenAI, Anthropic, Cursor, Vercel, Replit, Sierra, Clay, and hundreds of other winning companies all have in common? They are all powered by WorkOS. If you're building a product for the enterprise, you've felt the pain of integrating single sign-on, SCIM, RBAC, audit logs, and other features required by large companies. WorkOS turns those deal blockers into drop-in APIs with a modern developer platform built specifically for B2B SaaS.

1:32

Literally every startup that I'm an investor in that starts to expand upmarket ends up working with WorkOS. And that's because they are the best. Whether you are a seed-stage startup trying to land your first enterprise customer or a unicorn expanding globally, WorkOS is the fastest path to becoming enterprise-ready and unblocking growth. It's essentially Stripe for enterprise features. Visit workos.com to get started, or just hit up their Slack, where they have actual engineers waiting to answer your questions. WorkOS allows you to build faster with delightful APIs, comprehensive docs, and a smooth developer experience.

1:40

Go to workos.com to make your app enterprise-ready today. Something that you shared before we started recording is the fact that you're interviewing hundreds of PMs all the time. If I had a nickel every time someone asked me for an intro to someone at Anthropic to go work at Anthropic as a PM, I'd have 30 billion in ARR. It's just the number one place people want to go work at. So I can only imagine how many PMs you're interviewing.

1:46

You told me that you're just seeing people doing it wrong, the way they're approaching what they think it takes to be a successful AI PM. Talk about what you're seeing and what people need to understand about what it takes to be successful these days. I think before AI, technology shifts were a lot slower, so you could plan on six- to 12-month time horizons. And because you were shipping features at a bit of a slower rate, there was a lot more emphasis on coordinating with all the other partner teams to make sure that they're shipping features that unblock your features, because code at that time was very expensive to make.

1:51

I think now, with AI and with how much that has accelerated engineering, and with how quickly the model capabilities are improving, the timelines for a lot of our product features have gone down from six months to one month, and sometimes to one week or even one day.

1:52

And with that, we actually need to make sure that products ship quite quickly. And what that means is, as a PM, there should be less emphasis on making sure that you're aligning your multi-quarter roadmaps with your partner teams, and more emphasis on, okay, how can we figure out the fastest way to get something out the door? How can we figure out how to make a concept corner of our product suite where an engineer has an idea or a PM has an idea, and by the end of the week, we're able to get it into our users' hands?

1:54

I think the PMs who do the best on AI-native products are the ones who can figure out how can I shorten the time from having this idea to actually getting the product in the hands of users, and help define what are the most important tasks that need to work out of the box for my product. So what I love about this is what you're saying is people haven't grasped how fast they need to move and how much of the job now is helping the team move fast. What helps do that? What do you do? What does your PM team do to help them move this fast, other than have access to the most advanced models?

2:00

I think the first thing is to set clear goals because LLMs are so general that actually creates a lot of ambiguity in who we're building for, what problems we're trying to solve, what the top use cases are. And so I think a great PM is able to say, okay, our key user is professional developers. The main problem that we want to solve for this feature is maybe there's too many permission prompts and people are feeling fatigue. And the use case is we want professional developers at enterprises to safely get to zero permission prompts.

2:05

And that actually sets a pretty clear goal because it rules out a lot of potential approaches for reducing permission prompts so that people can get a lot more done with one prompt. And then I think the second thing that's very important is figuring out some repeatable process for getting these features shipped. So for Cloud Code, what we do is we actually ship almost all of our features in Research Preview. We clearly brand this when we ship something so that users know that this is an early product, this is just an idea, this is just something that we're trying to get feedback on and iterating on, and that this might not be supported forever.

2:11

And what this does is it reduces our commitment for shipping something. We can just get something out in a week or two. And then the third thing that a PM should do is help create the framework for the team so that they know when to pull in cross-functional partners and what those cross-functional partners' expectations are.

2:16

So, for example, we have a really tight process between engineering, marketing, and docs. So when engineers have a feature that they feel is ready and that we've dogfooded internally, they post it in our evergreen launch room. And then Sarah, who leads our docs, and Alex, who leads PMM, and Tarek and Lydia on DevRel, jump in and can turn around the marketing announcement for it the very next day. And because we have this really tight process, it lowers the friction for any engineer to ship something. PM is the role that should be setting this up. be supported forever. And what this does is it reduces, it reduces our commitment for shipping something.

2:30

We can just get something out in a week or two. And then the third thing that a PM should do is help create the framework for the team so that they know when to pull in cross-functional partners and what those cross-functional partners' expectations are. So, for example, we have a really tight process between engineering, marketing, and docs. So when engineers have a feature that they feel is ready and that we've dogfooded internally, they post it in our evergreen launch room. And then Sarah, who leads our docs, and Alex, who leads PMM, and Tarek, and Lydia on DevRel, jump in and can turn around the marketing announcement for it the very next day. And because we have

3:44

this really tight process, it lowers the friction for any engineer to ship something. PM is the role that should be setting this up. How do PRDs fit into this? The fact that you said that goals are a really important part is being aligned on what does success look like? Who is this for? Who is this not for? Are you ready? PRDs is a job. Just a couple of bullet points. How does, how has that evolved in the world of PM? So there's two, two things that we do. One is we have very rigorous metrics, and we do metrics readouts with the entire team every week. The goal of this is to make sure that everyone deeply understands all the facets of our business, what our key goals are,

4:37

how they're trending, and what drives them. The second thing that we do is we have this list of team principles, and this includes who our key users are, why those are our key users, and the reason that we articulate all of this is so that everybody on the team feels like they understand how our business works, they understand what's important to us and what we're willing to trade off, and it lets people make decisions by themselves without feeling like they're blocked on PM or any other stakeholder. I love how so much of this is, okay, we still need PMs in the future. And there's so much talk of, why do we need PMs? We're just going to ship and build. We need engineers.

5:19

Oh, we actually do PRDs sometimes. So I think for features that are particularly ambiguous, it does help to write out just a one-pager on what the goals are, what the delightful use cases are, what the failure modes currently are that we need to fix. And there are occasionally some projects, especially things that require heavy infrastructure, that do take many months. And for those situations, we do write PRDs still. I want to drill a little bit further into just how you're able to move so fast. I've never seen anything like the pace folks at Anthropic are shipping at. Someone made this calendar of launches across Anthropic, and it was literally every day

5:52

there was a major feature or product. So one question people had online is: you guys just launched this, not launched, but built this incredible model, Mythos, that is still in preview because it's so powerful. People are a little afraid of what it can do. Have you guys been using this? Is this part of the reason you've been able to move so fast? We've been moving pretty fast for several quarters now. So I think it's not fully Mythos. Mythos is an incredibly powerful model. We do use the models internally. And I think this has increased our rate of shipping a little bit, but I don't think it explains the bulk of the increase. I think a lot of it is the process

6:21

and the expectation on the team. So we're very low on process. We want to remove every single barrier to shipping things. We want to make sure every single person on the team feels empowered to take their idea from just an idea to out in the world in less than a week, sometimes even in a day. Cool. Oh, man. What an advantage to have the best model and also be building product. That's so cool. We are very lucky to be able to work with the frontier models. Oh my God. What an awesome advantage. Just build a thing and then use it and then accelerate faster. It's so interesting. There's a couple of these other side things I want to go on these side quests on this conversation.

7:00

There's so much happening with Anthropic, and I'm so curious to get your insight. One is a week ago or so, the whole source code of Cloud Code leaked. Somebody got it out there. I think it was a mistake someone made. Is there anything you comment there? What happened? What went wrong? What should people know? So we immediately looked into this when we saw it. We realized that this was the result of human error. There is a human working with Cloud to write a PR. This was just an update to how we release our packages, and it actually went through two layers of human review. So this was a result of human error, and we've hardened our processes to make sure

7:45

that it doesn't happen in the future. Is this person still adanthropic? Are they doing all right? Yes, yes. It's a process failure, and the most important thing is to learn from it and to add more safeguards so that doesn't happen again. And so that's what we've been focused on, and most of those have shipped. Okay. Another question I had is OpenClaw. So recently there's been this move to keep people from using Claude's subscription with their OpenClaw. People got really upset. They're confused why this is happening. It feels like there's harm caused to the open source community. What do people need to understand about what went into this decision? So we've been seeing

8:29

a lot of demand for Claude, and we've been working very hard to both scale our infrastructure and also to make our harness more token efficient so that you can get more usage out of it. It wasn't designed for third-party products, which have different usage patterns than our first-party ones. We spent a bunch of time trying to figure out what is the most seamless transition that we can offer. And so I was very happy to be able to say that everyone gets some credits alongside their subscription. But yeah, we did have to make the hard decision that we needed to prioritize our first-party products and our API. And so this is the decision that resulted from that. Yeah,

9:13

to me it makes so much sense. You guys are subsidizing this usage at 200 bucks a month, and it's basically unlimited use of this. And I think people don't understand. Businesses are trying to make money. We're trying to be profitable here. We can't just give away compute when it's so in demand, so I get it. Coming back to the PM team, what does the PM team look like at Anthropic? How many PMs are there? How are they organized? Yeah, so we have a few PM teams. I think we're maybe around 30 or 40 PMs right now. So we have the research PM team, who Diane leads, and this team is responsible for understanding all of the feedback from our customers for our models

10:06

and then feeding that to the research team to act on it, and they also shepherd the model launch. There is the Cloud Developer Platform team that maintains the APIs that Cloud Code is built on top of, and they also release things like managed agents, which is a way for you to build your agents, and we can host it on your behalf. And then there's Cloud Code that works on both Cloud Code and the Cowork core products. There's Enterprise that helps make Cloud Code and Cowork easier to adopt for all of our enterprise customers. And so this is everything from cost controls, RBAC, security controls, and just making sure that these enterprises feel

11:20

very confident and comfortable using our tools. And then we also have our growth team that is responsible for growing across our entire product suite, so we work very closely with them on Cloud Code and Cowork growth. And I know they also work with our other teams on CDP growth, so growth of people who use the Cloud API. So speaking of growth, Amal was just on the podcast. You have this really interesting insight that most people haven't been sharing. There's always the sense that we need fewer PMs in the future. Why do we need PMs? Engineers can just ship. His take is that because engineers are moving so fast, PMs and designers are squeezed. There's less time

12:09

to stay on top of everything that is happening. There's a feature shipping every day. So his take is he needs more PMs because it's hard to keep up. What's your take there? Do you feel like there will be an increase in hiring of PMs? What do you think is going on with the PM profession long term? I think all the roles are merging. PMs are doing some engineering work. Engineers are doing PM work. Designers are PMing and also landing code. You can either hire a lot more engineers who have great product taste, or you can keep your engineering hiring the same and hire a lot more PMs to help guide some of their work. On our team, we're pretty focused on hiring engineers

12:47

with great product taste. This way, we can reduce the amount of overhead for shipping any product. There are many engineers on our team who are fully able to end-to-end go from see user feedback on Twitter through to ship a product to stay on top of everything that is happening. There's a feature shipping every day. So his take is he needs more PMs because it's hard to keep up. What's your take there? Do you feel like there will be an increase in hiring of PMs? What do you think is going on with the PM profession long term? I think all the roles are merging. PMs are doing some engineering work. Engineers are doing PM work. Designers are PMing and also landing code.

13:26

You can either hire a lot more engineers who have great product taste or you can keep your engineering hiring the same and hire a lot more PMs to help guide some of their work. On our team, we're pretty focused on hiring engineers with great product taste. This way, we can reduce the amount of overhead for shipping any product. There are many engineers on our team who are fully able to end-to-end go from see user feedback on Twitter through to ship a product at the end of the week with almost no product involvement. And this, I think, is actually the most efficient way to ship something. So I think engineer and PM are overlapping, and you will get a lot of benefit

14:06

from having more of either. I think product taste is still a very rare skill to have, and we'll pretty much hire anyone who we feel has demonstrated this strongly. And your background was in engineering, right? Yeah. I was an engineer for many years. I was then a VC very briefly before joining Anthropic. And actually, almost all the PMs on our team have either been engineers or ship code here on Cloud Code. And so that's one of the things that I think helps build trust with the team and also just enables us to move a lot faster. And then actually, our designers also have been front-end engineers before. Wow. Because that's the big question. There's definitely

14:45

this merging that's happening. The Venn diagrams are combining. I think the big question for a lot of people is, if you're coming from engineering or product or design, which of those core skills is going to be most valuable? I could see at Anthropic and on Cloud Code, engineering is very valuable. I'm curious if, at other companies, if you have a design background, becoming a PM is more valuable, or just a PM PM? I still think it comes back to product taste. As code becomes much cheaper to write, the thing that becomes more valuable is deciding what to write. What is the right UX for this feature? What is the most delightful way that a user can experience it?

15:15

We get tens of thousands of GitHub issues asking for every single thing under the sun, and it takes a lot of care and taste to figure out, okay, which of these is worth building and what is the right way to build it? And I think that that skill set can come from any background, but I think that's the most important thing. I think the reason why an engineering background is particularly useful, at least for the next few months, is if you have an engineering background, you have a better sense for how hard something should be, and that's often a factor in what you choose to build. So if something is very easy to build, then maybe instead of debating it,

15:45

you just spend an hour doing it. But if something is harder to build and you know that up front, then you know that, okay, this will just cost a lot more for our team to get this out the door. So it helps a bit with the prioritization. You said for the next few months. Is that just because the models will get so good, potentially, in the next few months, you may not even need to know that as much? I think the valued skill sets do change quite frequently, and so it's really hard to predict more than a few months out. So it's less a commentary on what shift I think will happen and more of a commentary that I think large shifts will happen. So you're not saying

16:14

that's when Mythos comes out and will change everything, and we don't need to know anything about engineering. No, I'm just saying that every few months, it seems like there's a large increase in coding capability, which then changes what other roles are valuable. I think the most important thing is to be able to have this first-principles thinking where you can figure out how the tech landscape is changing, what the team really needs from you, and to jump in and fix that hole, because I think the work is becoming more amorphous, which means that a great PM is able to understand what all the gaps are, to figure out what the highest-priority ones are,

17:01

and then to figure out, okay, how do I learn that skill set, or what is the skill set that I have that I can apply to this challenge? So I think the current environment values people who are able to wear a lot of hats or are able to swap them and are very low ego about what work they do to help the team move faster. I love this answer. There's this question I've been asking people in your shoes, folks that are at the bleeding edge of what AI is capable of and building with the latest tools, which is just where will human brains continue to be useful and necessary for a while until we get to superintelligence? What I'm hearing here is essentially picking the things

17:48

to work on, knowing where the market's going and figuring out what to prioritize, essentially, and then it's knowing if the thing you've built is good and right and getting it out there in some early version, at least. Does that sound right? Is there anything else of where human brains will continue to be useful for at least the next few months? I think humans still provide a level of common sense that the models don't. And there's a thousand moving pieces to any product launch. Some of them are very small, but there's always a lot that could potentially go wrong. I think the model doesn't always have a great sense of who all the stakeholders are, how they relate

18:26

to each other, what their preferences are, what are the right venues to communicate with them to keep them on board. I think a lot of this more tacit, common-sense, EQ kind of knowledge is still very valuable. Of course, we want the models to get better at this, and I think they will be, but right now, I think there are still gaps. How do you deal as a human going through so much constant change, just being on the inside of the tornado? Maybe it's calm there, but how do you, how do you stay on top of what's going on? How do you stay sane through all this craziness that we're moving through? I think our team is full of people who lean into the chaos, so we try to face

19:06

every challenge with a smile because there's always so much going on. There's always so many risks and tricky situations that if you get too stressed about anything, you'll burn out, and so we really look for people who can look at a challenge, be like, whew, that's going to be hard, but I'm excited to tackle it, and I'm going to do the best that I possibly can, and I know I won't be perfect, but I'll be able to sleep at night knowing that I did my best. That's an interesting answer to what skills will be important in this future because it's, I forget who said this, maybe Ben Mann, that this is the most normal the world will ever be. Yeah, it definitely gets harder.

19:37

I feel like there are a lot of weeks where maybe Sunday night there's some P0, and then by Monday there's a P00, and by Monday afternoon there's a P000, and you're like, wow, I can't believe I was so worried about that P0 from Sunday. But I think you just have to acknowledge that there's only so much that you can do, that you need to sleep well so that you can make good decisions the next day, and just brutally prioritize where you spend your time, what's the most important thing to get right, and be okay letting things go. There's products that we ship that aren't as polished as I wish they were, but our top goal is to help empower professional developers,

20:18

and if a product isn't successful, as long as it's not blocking the core use case, it's okay because we'll hear the feedback and we'll fix it in the next release. Launching a feature that is buggy is the kind of thing that would have kept me up at night, but it is something that I am now able to live with, knowing that, okay, we're going to get that quick feedback and we're going to fix it in the next release. What I'm imagining is there's that gif, I think it's maybe from Pirates of the Caribbean, where it's this guy walking down a pair of stairs on a ship, and the whole ship is just being demolished around him, and he's so chill, just strolling down the staircase

20:59

as everything's falling apart. And that's interesting because everyone I've met from Anthropic is just so chill and just so optimistic. our top goal is to help empower professional developers, and if a product isn't successful, as long as it's not blocking the core use case, it's okay because we'll hear the feedback and we'll fix it in the next release. Launching a feature that is buggy is the kind of thing that would have kept me up at night, but it is something that I am now able to live with, knowing that, okay, we're going to get that quick feedback and we're going to fix it in the next release. What I'm imagining is there's that gif, I think it's maybe

21:34

from Pirates of the Caribbean, where it's this guy walking down a pair of stairs on a ship and the whole ship is just being demolished around him, and he's so chill, just strolling down the staircase as everything's falling apart. And that's interesting because everyone I've met from Anthropic is just so chill and just so optimistic. Yeah, that's, I think that's a really interesting insight, just having this calmness and optimism versus just, oh my God, everything's crazy and going nuts. Yeah. I think if you don't have it, you'll get pretty burnt out. I think we also tend to hire people who have been in the industry for a while and have experienced lots of ups and downs

22:18

and have a good sense for what gives them energy and how to maintain their energy over time. And I think that's helped us a lot. So interesting. Something that I wanted to ask about is, there are these roles blurring. Engineers are becoming PMs. Everyone's, dogs are cats. Everyone's everyone. What do we lose in that world? Do we lose career ladders and clear career paths? Do we lose design consistency, code quality? There's probably some downsides. What are some things you find are just, okay, that's something we're sacrificing for the greater good? We're sacrificing product consistency. Historically, when code was expensive to write, you would carefully plan out

22:56

everything in your product suite, how every product relates to each other, what the use case for every single one is, how they integrate, and you would pretty much have one product for each use case. And now, with AI moving so quickly and with so many ideas that we need to test out, we do sometimes have features that overlap with each other. A lot of the time, it's because there are two form factors that we love internally and we want the external audience to tell us which one is better. What that means for someone who's a new user, though, is a new user might not know, okay, what is the best path to accomplish X? There is more education we need to do to help people

23:42

understand what the core features are and what the best practices are for using them. I think this is the, this is the cost of launching a lot of features. I think users also feel like it's hard to keep up with the latest. Usually, in traditional PM, you ship a feature every month or quarter. And so, it's really easy for a user to understand, okay, I just need to check in on this once a month and I'll learn some new things. And if I ignore it for six months, it's fine. I don't feel like I'm missing out. I think with these agentic tools, not just cloud code and co-work, but across the whole ecosystem, people feel this need to check Twitter every single day to see what the

24:30

absolute latest thing is. And I think there's more we can do to help people feel less like they're on this ever increasingly fast treadmill and that they feel like, I would love people to feel like they can just open these tools. The tools will educate them or teach them what they want to know and that they can just feel more brought along. Yeah, I saw you launch this really interesting feature the other day. I think it's slash power up, where it basically walks you through all the cool ways and basically all the best practices to use cloud code. Is that kind of all in these lines? Yeah, exactly. So in the past, we didn't actually want to do something like power up

24:59

because we felt like the product should be intuitive enough that you don't actually need to go through any tutorial. And over time, we've just realized that there's just so many features and there's so much demand for a built-in onboarding experience that we diverged a bit from our original principle, saying no, no onboarding flow, and added this because there are just so many users who wanted to know: there are 100 features. What are the 10 that I absolutely need to use? And so we put that together. Yeah, it's such a bizarre world. So Anthropic has been really successful with B2B enterprises, where traditionally you don't launch a bunch of stuff. You just kind of

25:36

have a quarterly release, maybe, and it's like the opposite of every day we got something new. So maybe following that thread, the run Anthropic has been on is just otherworldly. Anthropic was way behind when it started. It was, Amol shared this, just one of the least funded companies, didn't have distribution. Was it the first to go? OpenAI was way ahead. It was just like, no way Anthropic has any chance to compete significantly long-term. Now it's just killing it. Just beating the biggest companies, teams with so much, just the growth is just like $11 billion in ARR in one month. Perps and growth. By the time this comes out, it'll probably be even higher.

26:13

Just being on the inside, what are some ingredients that have allowed Anthropic to be this successful and kind of come from behind and do this well? The two most important things are, one, this unifying mission. It's hard to state how important this is. We hire people who care most about bringing safe AGI to all of humanity. And this is actually something that we reference frequently in our decisions about what our entire product org should focus on shipping. And because we put this mission above any individual product line, we're able to make very fast decisions that cut across the entire org and execute on them in a unified way. So I think this is, this is something

26:45

that I've never seen at a company of our scale. And so just to make sure that's clear, so essentially having the number one mission is safety, alignment, making sure AI is good for the world. And you're saying just having that as a clear mission makes decisions a lot easier to make. If there are two competing priorities, we'll talk about which one is more important for Anthropic's mission. And it makes it a lot easier to decide which of the two we prioritize. And then everyone will stand behind the one that we decide. And so sometimes that means that, hey, we want to ship something on cloud code, but this other thing is more important. And so we deprioritize shipping this

27:19

and we just wait until later. What's really interesting about that is that explains, I think, versus another company maybe rhymes with Vopen BI, did a lot of different things. And what I'm hearing here, essentially, is like, okay, we're not going to launch a social network. We're not going to launch a feed of interesting information because it's not aligned to this mission. And that has kept Anthropic focused, which just seems to be a core ingredient to the success. Well, when I think about mission, I think about putting Anthropic's goals ahead of any individual org or any individual product. And so for me, I think the second thing that we're very good at is focus.

27:56

I think mission to me is slightly different. Mission means that teams are willing to make sacrifices that hurt their own goals and their own KRs in service of Anthropic's goals and Anthropic's KRs. And people are very happy to make those trade-offs. So an extreme example is if cloud code failed, but Anthropic succeeded, I would be extremely happy. And the whole team is very willing to make decisions that follow that chain of thought. I don't know if you can talk about this in depth, but do you feel like the open claw decision is a part of this? Just like, okay, this is not furthering the mission of Anthropic. We need to stop this because it's not working

28:28

in the way we want it to work. I think one of the most important things for Anthropic is to grow the number of users that we're able to reach. One of the ways that we're able to do this is with the cloud subscriptions with our first-party products. And so we just very much want to double down on that. But that does come at the expense of third-party products sometimes. So we've been talking about cloud, co-work, all these things, something that I want to make sure people get, and I'm curious just how you use these tools. So there's cloud code, there's cloud desktop slash web, there's co-work. What's the best way to understand when to use which? When do you use

28:57

each of these three? So I tend to use cloud code in the terminal when I'm just kicking off a one-off coding task and I want all of the latest features. The CLI is our initial product surface, and it's also the one where our features often land first. And so it's the most powerful of all the tools. is to grow the number of users that we're able to reach. One of the ways that we're able to do this is with the cloud subscriptions with our first-party products. And so we just very much want to double down on that. But that does come at the expense of third-party products sometimes.

29:17

So we've been talking about cloud, co-work, all these things, something that I want to make sure people get. And I'm curious just how you use these tools. So there's cloud code, there's cloud desktop slash web, there's co-work. What's the best way to understand when to use which? When do you use each of these three?

29:20

So I tend to use cloud code in the terminal when I'm just kicking off a one-off coding task and I want all of the latest features. The CLI is our initial product surface, and it's also the one where our features often land first. And so it's the most powerful of all the tools. So that's what I tend to use when I'm trying to kick off one or maybe a handful of tasks at a time.

29:21

I think desktop really shines when you're doing something that requires front-end work. And so one thing that I love to do is to use our preview feature. So if I'm building a web app, I'll often use cloud code in desktop. I'll have the preview pane open on the right-hand side so that I can actually see the web app that I'm making in real time as I'm chatting with cloud.

29:22

It's also really great for people who want something a bit more graphical. A terminal can feel very unfamiliar to someone who's non-technical. You get a bunch of these scary pop-ups on your machine, and you can't click around the way that you're used to in pretty much every other product that you use. So there's a lot of people who just don't feel comfortable in the terminal. And if that's you, I would highly recommend checking out cloud code on desktop.

29:24

Desktop is also great for getting an at-a-glance view of everything that's happening. So you can see your CLI terminal sessions in desktop. You can see your other desktop sessions. You can see your sessions that you kicked off on web and mobile. So it's a one-stop control plane where you can see all of your tasks.

29:26

I think the benefit of web and mobile is that it's really great for kicking things off on the go. So CLI and desktop both require you to be on your local laptop. And this is constraining because sometimes you're out and about. You're touching grass. You're going on a walk, and you don't have your laptop open. And you don't... I can't count the number of people who I've seen holding their laptop open, tethered to their phone, while they're outside. And this just means that we're missing a product that solves that need.

29:27

And so for me, what mobile lets you do is kick off these tasks on the go so that you don't need to bring your laptop everywhere and make sure that your laptop's open wherever you are. I love that. I've seen people on planes. It's such a meme now. Just, I need to finish. Let this agent finish. I can't shut this down. I need Wi-Fi.

29:30

And then I think for co-work, the role that this fills is there's a lot of work that everyone does where the output isn't code. So whether that's getting to Slack zero or Inbox zero, or whether that's creating a slide deck for some customer meeting that's coming up, or whether that's writing a quick doc on what the goals of a feature are or what the launch plan for a feature is, all these tasks produce outputs that are non-code, and co-work is best positioned for that.

29:31

So the way that I split the products in my mind is if I'm building something where the output is code, I'll use cloud code or desktop or cloud code on mobile. And if the output is anything that's not code, I'll use co-work for it. People are just sleeping on the success that co-work is having. It's growing incredibly fast. And I think people still don't understand maybe what it's for. And so what if you give us a couple of use cases just in your work as a PM? What are some really interesting, maybe unexpected ways to use co-work to save you time, get more work done?

29:35

If you're getting started on co-work, the first thing that you really need to do is connect all the data sources that are relevant to your role. Because co-work can only do a great job if it has access to all the context that it needs to be able to curate the output for you. So what that means for me is I connect it to my Google Calendar, I connect it to my Slack, to my Gmail, to my Google Drive, so that it just knows it has the flexibility to find relevant context, to ask questions, to pull in threads. And this substantially improves the quality of the result.

29:38

The kinds of things I use it for are last night I was working where we have this Code with Cloud conference coming up, and there's a few talks that I'm giving there. And one of the talks that we're doing talks about the transition of Code code from an assistant to a full-on agent. And one of the things that I wanted to do in this talk was to showcase all of the products that we've been shipping that enable this transition, and also to figure out, okay, what are the success stories that people have had internally that we can use as demos?

29:40

And so I have my Google Drive connected. I have Slack connected. Alex, who's our product marketer, put together a draft of the points that he thinks we should cover. And so I fed this all into Cowork. I told Cowork the narrative that I wanted to tell. And it actually just worked for an hour. It walked through Twitter to see what we launched. It looked through our evergreen launch room. It looked in our Quad Code announce channel, which is where our team posts demos of how they've been getting the most value out of Quad Code. And it synthesized all this together to this 20-page deck that I woke up to this morning. I read through it, and it was pretty good.

29:42

There were a few tweaks, so I did have to give it a round of feedback. I like my slides to have extremely minimal words, and it was a little too wordy. But it was far faster than what I would be able to produce. And because Cowork has access to our whole design system, it actually looks like an anthropic designer put it together. When you visually see it, you're like, oh, this is incredibly polished. So these are the kinds of things that are so much faster. Making this slide deck would have taken me hours. But instead, it turns out a draft that is actually quite good so that I could focus on making sure that the demos are amazing that we plug into it.

29:46

This sounds like a dream come true to PMs, that putting decks together is so annoying. It's so slow. And I love that people will see this deck whenever you present this. This will be out in the world. Obviously, it's not the one-shotted version, but you've iterated on it. So just to help people try this for themselves, step one is connect their, what did you say, Slack? What else do you suggest they connect? Slack, Google Calendar, Gmail, G Drive. You should connect your communications tools and where you store your source-of-truth data for what your team cares about, what you care about, and what you're working on.

29:49

Okay. And then what was the prompt, roughly, that you put in there to generate this deck? So I just wrote, make me a slide deck for the Code With Cloud conference. This is what our PMM suggested it should cover. This is the current draft that I made that I don't like. This is one that I made manually that I don't like, but I linked it. Can you start by creating a proposed outline with details? Also, make sure it doesn't overlap too much with a keynote talk, which is more important.

29:54

And then Cod read a bunch of the links that I sent to it and created a proposed outline. So then I read through its proposal and all the different ideas that it generated for what we could cover. And I just made a decision on what I wanted to actually be in the final deck. And I think this is an example of what the role of the PM still is today. Claude is a great brainstorming partner. It's able to synthesize a massive amount of information really quickly and present all of the possibilities to you. But the role of the PM is still to make the end decision of, okay, what should belong in the final product?

29:58

So for this, what I ended up deciding was that I wanted the talk to cover the progression from making local tasks successful to making every PR green to helping engineers land more PRs. And for each of these, which demo would be the most compelling? And then after this decision about the outline, co-work just went off for a few hours and built the whole side deck. This is so awesome. What an awesome part of the job to not have to do anymore. And it feels like you're talking to essentially a deck designer that also has actual knowledge And I think this is an example of what the role of the PM still is today. It's, Claude is a great brainstorming partner.

30:06

It's able to synthesize a massive amount of information really quickly and present all of the possibilities to you. But the role of the PM is still to make the end decision of, okay, what should belong in the final product? So for this, what I ended up deciding was that I wanted the talk to cover the progression from making local tasks successful to making every PR green to helping engineers land more PRs. And for each of these, which demo would be the most compelling? And then, after this decision about the outline, co-work just went off for a few hours and built the whole slide deck. This is so awesome. What an awesome part of the job to not have to do anymore.

30:14

And it feels like you're talking to essentially a deck designer that also has actual knowledge about what you've worked on and can make it actually the content which you want it to be, not just make it look really nice. How did you, how did you do the design system piece? How does that work? How does it know the design system of Anthropic? So what I did for this is we actually already have a standardized deck that we use across all of our external engagements. And so I just gave Claude access to that. And so it's able to see what colors we use, the fonts we use, the different kinds of, what's it called? Slide formats that are possible.

30:23

And so it has 20 of these example slides. So give an example. Got it. So you upload, here's our template, work from this. Yeah. You can also connect your Figma MCP. If you have your slide format saved there, it can pull that in. Along those lines, something I'm always curious about is what's in your stack of tools as a PM at Anthropic. Obviously, Claude Code and Co-Work and all the Anthropic tools. What else are you using? What other, Slack you mentioned. Is there anything else? So my stack is pretty heavily Claude Code, Co-Work, and Slack. Anthropic largely runs on Slack. I feel like it's the core OS of our company.

30:42

And day to day, a lot of, I would say maybe 30% of my time is pushing the boundaries of what Co-Work and Claude Code can do so that I have a very strong sense of what we're not good at. And I spend a lot of time talking with the model to understand why it makes mistakes that it does. We actually have a lot of internal tools that we make. I think one of the things that Claude Code has really unlocked for our entire company is it really lowers the barrier to making any custom app that you want. And so we've seen this surge in personalized work software that people are building for custom use cases instead of using tools that don't perfectly fit the use case.

30:47

I got to hear more. What are some examples? What are things you've built, other people have built, that are really popular and useful? One of the sales folks on Claude Code, he realized he was making these repetitive decks over and over and over again. And so he actually has this web app that he built with the examples of the core Claude Code decks that we know work well. So a 101, 201, and mastering Claude Code. And then he has a way to input specific customer context that pulls from Salesforce, that pulls from Gong, that pulls from other notes so that we can customize the decks for specific customers.

30:56

And so we'll pull out things like, okay, this customer is using Bedrock or Code for Enterprise or Console, which affects what features are available to them. It will pull out things like, okay, this customer is concerned about the code review stage of the SLC. And so we'll add a slide about our code review features there. It'll pull out things like, okay, this customer needs to be HIPAA compliant or needs XYZ security controls. And so we'll make sure to add a slide or two in their deck about that.

31:02

And then, for example, if this is a customer that's on Vertex or Bedrock and doesn't want to use Claude for Enterprise, then we'll just take out some of the slides that are called for Enterprise-only features. And so normally this is manual work that could take 20, 30 minutes, and so people either spend that time doing it or they'll just decide not to do it and use the general deck. With this, it takes a few seconds and you get a tailored deck. What's interesting about it is Slack is the tool that nobody's, it's just, nobody's trying to create their own. Slack just continues to win, and it's just the way you describe it is kind of the OS of so many companies.

31:09

It's so interesting. People talk about Salesforce as just SaaS. We don't need SaaS software anymore. We're going to build our own. It's like Slack is an adorable tool that nobody wants to try to compete with and build a better version. I think it's pretty important communications infrastructure, and I think they do the core task of helping everyone get real-time updates incredibly well. Yeah, people hate on Slack, but it's really great at what it's trying to do. And the most cutting-edge teams are hooked on it. So interesting. Yeah, and I also love how easy they've made it to customize it.

31:25

And so we love making Slack bots, and this kind of hackability means that we're able to integrate with Slack the way that we want to. So really appreciate Slack's work on that. Time to buy some CRM stock. I am so excited to tell you about this season's supporting sponsor, Vanta. Vanta helps over 15,000 companies like Cursor, Ramp, Duolingo, Snowflake, and Atlassian earn and prove trust with their customers. Teams are building and shipping products faster than ever thanks to AI. But as a result, the amount of risk being introduced into your product and your business is higher than it's ever been.

31:36

Every security leader that I talk to is feeling the increasing weight of protecting their organization, their business, and not to mention their customer data. Because things are moving so fast, they are constantly reacting, having to guess at priorities, and having to make do with outdated solutions. Vanta automates compliance and risk management with over 35 security and privacy frameworks, including SOC 2, ISO 27001, and HIPAA. This helps companies get compliant fast and stay compliant. More than ever before, trust has the power to make or break your business. Learn more at vanta.com slash Lenny. And as a listener of this podcast, you get $1,000 off Vanta.

31:49

That's vanta.com slash Lenny. Okay, so you talked about all these different teams and how they use cloud code and code to operate. Which teams do you find, other than engineering, I imagine engineering is the biggest token spender, but if not, that'd be really interesting. What's kind of the second-place function right now for tokens? Oh, Applied AI is amazing at pushing the boundaries of what Cloud Code and Co-Work can do. A lot of our Applied AI team spends time with our customers, helping them adopt our API. And so sometimes our Applied AI team will, for example, make prototypes on behalf of these customers, which Cloud Code makes so much faster than it used to be.

31:59

They also have the dual goal of needing to manage a lot of customer comms, a lot of customer inbound and historical contacts, call notes. And so they're both extremely heavy on Co-Work and on Cloud Code. And just to understand Applied AI, does that like forward-deployed engineering sort of role? How would most people describe what the Applied AI team is doing? Yeah, it's helping our customers adopt the latest API and model features across their company, both for powering their company's products and also for internal acceleration. Got it. It's like customer success, go-to-market-y, kind of like forward-deployed engineering sort of thing. Exactly.

32:08

It's like a very technical go-to-market person. Got it. Okay, awesome. So you're saying that might be the second org that uses the most tokens. Yeah. And then we also see them pushing the boundaries of what Co-Work can do. So for example, a lot of these folks cover multiple customers and, on any given day, can have five to 10 customer engagements on a high day. And so what they often use Co-Work to do is, the night before, they'll ask it to summarize, okay, what are all my customer meetings that are coming up the next day? What are all the things that this customer has asked me for? What's top of mind for them? What are the action items for the past meetings?

32:23

And Co-Work will just put together this dossier, this brief of what they should be aware of going into the next meeting. And Co-Work can also research answers. So if a customer asked, okay, when is feature X going to launch? Co-Work can help the Pi AI person research through Slack to get the latest ETA, add that to the notes so that during the customer call, So, for example, a lot of these folks cover multiple customers and, on any given day, can have five to 10 customer engagements on a high day.

32:29

What they often use Co-Work to do is, the night before, they'll ask it to summarize, okay, what are all my customer meetings that are coming up the next day? What are all the things that this customer has asked me for? What's top of mind for them? What are the action items for the past meetings? And Co-Work will just put together this dossier, this brief of what they should be aware of going into the next meeting.

32:31

And Co-Work can also research answers. So if a customer asked, okay, when is feature X going to launch? Co-Work can help the Pi AI person research through Slack to get the latest ETA, add that to the notes so that during the customer call, the Pi AI person has the absolute latest. And these are just workflows that people are building for themselves and sharing with other people on their team.

32:33

So cool. Something that this question, this trend, I don't know, question topic comes up a lot recently, which is token spend exceeding people's salary, where people just use AI and it costs more than how much they're making. Are there any numbers floating around on the topic of how much token spend, say engineers spend, I don't know, a month, a day, PMs, anything like that?

32:33

It is clear to us that, as the models get better, people delegate far more tasks to it and they spend a lot more hours in tools like CloudCode and Cowork. And so we do see the token cost per engineer, or per any knowledge worker, increase every time that there is a model jump or a substantial product improvement. I think it's still much lower than what the average engineer salary is, but we see the percentage increasing over time. It's such an interesting, we talked about how you have access to the most cutting-edge models, another advantage of working at Anthropic. I believe you guys have basically unlimited tokens. You can use as much as you want. Is that right?

32:37

We can use a lot of tokens. Some people do run into limits. Okay, there's a limit. Okay. Boris, shut it down. Okay. It's so interesting how many advantages come from having the most advanced model. It's such an interesting flywheel that starts to kick in. I think we also believe a lot in empowering our internal teams to build as fast as possible. And we also trust that everyone understands how much capacity serving these models truly costs. And we trust our team to use the tokens responsibly. So it's very frowned upon to waste tokens, but we do trust individuals to make that judgment call.

32:43

Awesome. Coming back to the PM role, we talked a little bit about this, but I think this will be really interesting for people to hear. What I want to understand is, what do you think are the emerging skills that PMs need to develop, slash you most look for, AI companies most look for when they're hiring PMs these days?

32:44

I think the hardest skill is being able to define what the product should look like a month from now. I think there's a lot of ambiguity in what models are capable of in that timeline and how user behavior will change. But I think there are patterns that the best PMs can see based on how users are abusing the limits of the existing product, and the best PMs can sense that, can set a direction, and can steadily execute toward it and change the path if the model capabilities are much better or worse than what they'd originally expected.

32:46

I think it is very hard to be the right amount of AGI-pilled. I think everyone can see this future where the models are extremely smart and can do almost everything, in which case you actually don't need that complicated of a product. You can actually just have a text box again where you tell the model what you want, and it's so smart that it can add any tool or add any integration that it needs to get the job done. It knows when it's uncertain. It can ask clarifying questions. It's very easy to build the product for the super AGI strong model.

32:47

I think the hard thing is figuring out, for the current model, how do you elicit the maximum capability? How do you help users get onto the golden path? How do you guide users to interact with the model's strengths and patch its weaknesses? This skill is pretty rare. And how do you build that skill? Is it just understanding the limits of each model? You talked about taste, understanding, having taste in what the model maybe is capable of, what it's great and not great at, where it's changed?

32:51

I think it's spending a ton of time talking and using the model. One of the things I really like to do is to ask the model to introspect on its own behaviors. So sometimes when I notice that the model does something unexpected, for example, there are situations where the model will make a front-end change and run tests, but not actually use the UI. It's actually pretty useful to ask the model to reflect on why it did this. And sometimes they'll say that, hey, there was something confusing in the system prompt, or I didn't realize that the front-end verification was part of the task, or hey, I delegated the verification to this subagent and the subagent didn't do the test and I didn't check its work.

32:52

A lot of times, being very curious about why the model made the decision that it did will show you what misled it so that you can fix the harness in order to close this gap. The other thing that helps is to figure out who are the users who you trust the most to give you accurate feedback about the model. Usually there's a handful of people who are much better than others at articulating what makes a specific model or model-harness combination good. And there are a lot of people who will give you feedback, but not everyone's feedback is as qualified. And so finding a group of those five people you trust is really important for getting very fast feedback.

32:57

I think the third thing that is useful, but not everyone loves doing, is building evals. You don't need to build hundreds of evals for them to be useful. Just building 10 great evals is important for helping the team quantify what the goal is and what their progress toward it is and what they're missing. And so I think evals is this underappreciated thing that more PMs, more engineers should be working on.

32:58

We've covered evals a bunch. There's this trend that that is the future of product management, writing evals, because essentially it's what does success look like? Okay, cool. Let me actually concretely define it and then we'll know. How much of your time are you spending writing evals, would you say?

32:59

I think the importance of evals varies a bit based on the feature that you're working on or what the problem you're trying to solve is. So there are a lot of folks on our team who do spend a lot of time working on evals. We have a small pod of folks who collaborate very closely with research to more precisely understand our cloud code behaviors and what the largest areas of improvement are and trying to measure those pretty concretely.

33:01

I personally jump into evals when there's a feature that I think needs a bit more product definition, and often the output of this is, okay, here are five evals that I made, this is how you run them, these are the ones that succeed and these are the ones that don't, and this is the prompt that I've used to increase the success rate. It varies a lot, though, based on the feature. Not every feature needs it, but I think features such as memory benefit a lot from it.

33:04

This point you made about people being very good at evaluating models is so interesting. It's almost like a human eval of just, okay, they understand where it's spiking or it's maybe lacking. Is there anyone specific that you want to shout out that's very good at this? Two people who I think are incredible at this, because the task is so ambiguous, even coding is easier because you can verify the success, whereas crafting the character requires a very strong sense of conviction in who Claude should be, and I think she has an incredible ability to not only mold the character but also to articulate what the goals are, what the character, what's successful, and what's not.

33:07

The other group of people who I really trust is the cloud code team. So we often have team lunches, and whenever there's a new model we're testing, one of the fastest ways for us to get feedback is to just go to every person and be like, hey, what is your vibe on the model? And oftentimes we'll get feedback like, okay, this model is not fully explaining its thinking, it's too abrupt, or hey, this model just loves writing a ton of memories, but we're not sure if the memories are high quality or not, or some people will notice that, okay, this model loves to test itself, which is great, or this model isn't testing itself enough, so that informs what data we look at to verify. Is this a

33:08

and who Claude should be. And I think she has an incredible ability to not only mold the character, but also to articulate what the goals are, what the character, what's successful, and what's not.

33:09

The other group of people who I really trust is the cloud code team. So we often have team lunches, and whenever there's a new model we're testing, one of the fastest ways for us to get feedback is to just go to every person and be like, hey, what is your vibe on the model? And oftentimes we'll get feedback like, okay, this model is not fully explaining its thinking, it's too abrupt, or, hey, this model just loves writing a ton of memories, but we're not sure if the memories are high quality or not. Or some people will notice that, okay, this model loves to test itself, which is great, or this model isn't testing itself enough. So that informs what data we look at to verify: is this a larger pattern? So we have a ton of data, but it is very hard to extract what are the hypotheses we want to test, and then we're able to extract data to test that.

33:10

This point you made about the character of Claude, I had Ben Mann on the podcast, co-founder, and he talked about this. The character, the constitution of Claude, is such an important part of Claude. And I didn't realize until afterwards, with OpenAI actually, one of the reasons people are sad is the personality of Claude, because Claude's personality is so good and fun and interesting, unlike other models. And the way he put it is the personality is what makes Claude so good at so many things. It feels like this trivial side thing, okay, it's going to be funny and interesting and talk in a fun way, but it's so core to the success of Claude. Is there anything there about what people may not understand about why the character, as you described, and the personality is so key?

33:11

When you reflect on everyone you've worked with, there's just some people where you're like, I really like their energy. I really like their vibe. And when people think about Claude and Claude code, this is one of the things that people bring up the most, where they just really... is extremely competent at your task. People really like Claude's low ego, and so if you tell it, hey, you did this thing wrong, it's truly sorry. It's like, oh shoot, thanks for telling me. Let me fix it. Let's work together. It's also very positive, so if you're feeling like, oh, this is an insurmountable task, I think part of what makes a great coworker is this positivity, this bias towards action, this ability to give you earnest feedback, not just agreeing with every single thing that you say. And so we try to imbue this into COD because we think it makes it a lot more enjoyable to work with.

33:13

There's something I want to come back to. You talked about how when new models come out, you often have to revisit things you've built. That's so interesting and so frustrating, maybe just like, oh, god damn it, we shipped this thing, now we have to rethink it. Talk about how often you have to come back with a new model and they're like, okay, we have to redo this product that we launched a few months ago.

33:15

A lot of the changes that we make with a... so the classic example for this is the to-do list. When we first launched quad code, people would ask it to do these large refactors, and quad code would say, okay, cool, I need to change these 20 call sites, and it would go and change five of them and then stop. And then we were like, okay, how do we force it to remember to get every single one of these 20? And so Sid on our team was like, okay, what if we just think about what a human would do? A human would make a list of everything that they need to change, similar to how in VS Code you would look up all the call sites and it will be on the left side, and you go through them one by one. How do we force it to use this to-do list? It would naturally use it itself. For the earlier models, we had to keep reminding it, hey, did you finish everything on the to-do list? You can't finish until you're done with everything on the to-do list. And for the later models, without prompting, it naturally thinks to do everything on the to-do list. These days, the to-do list, the model may use it, it is really not necessary for it to make thorough changes anymore.

33:17

I forgot who said this on the podcast, that the model will eat your harness for breakfast. And what I'm hearing here, essentially, you remove things over time that you've had to add on top of the model where it was not operating the way you want it. And essentially, as the models get smarter, it becomes simpler and simpler for it just to do the thing you want it to do.

33:18

Yeah. We can remove a lot of prompting interventions every time the model gets smarter. And we actually do this every time we launch a model. We read through the entire system prompt and we reflect on, okay, for each of these sections, does the model really need this reminder anymore? And if not, we'll remove it.

33:19

The most exciting thing that new models unlock, though, is entirely new features. So there's a lot of features that we've been testing out with prior models, and the accuracy wasn't high enough for us to want to launch them. And so one example of this is code review. We tried to build a code review product a few times, and we've launched simple versions of code review, which is the slash code review command, in the past. And it was only with the most recent models that we felt like, okay, this code review is so good that our engineering team relies on this code review to pass before we merge PRs. And we found that this was... we've always dreamed of Claude being able to be a reliable code reviewer that we can confidently feel catches the majority of bugs. And it was only with Opus 4.5 and 4.6, and Sonic 4.6, that we felt like, okay, we are now able to run multiple code review agents simultaneously to traverse the entirety of the code base and to synthesize a set of real issues that an engineer needs to address before merge. And so this is a new capability that the newest models have unlocked.

33:21

This is another trend that is very common on this podcast of build something that will possibly be possible in the next six months. Be at the edge of what's working, and then it'll catch up, and then it'll be an amazing product and you'll be ahead of everyone. Yeah, exactly. It's pretty important to build products that don't necessarily work yet so that you know, okay, what is missing for this product to work? And then with the newest model, you can just swap it into the prototype you've already made and see, okay, does this new model close that gap?

33:24

How much are you able to speak to where things are going with Claude and Co-Work as the vision of it? I imagine you don't want to give away too much about the goal, but it feels like there's all these awesome features being added on top, dispatch control from phone and all these mobile app, all these things. What's a way to understand the vision for all these things long-term?

33:25

We think about this in terms of building blocks. So for both Claude and Co-Work, the core building block is making individual tasks successful. You want it to produce some output. You give it a clear prompt description. Is it able to consistently produce acceptable output that you're able to either merge or share with your colleagues or external audience? So the task is the core building block. As the models get smarter, the task success rate gets a lot higher. And then we see people moving towards doing multiple tasks at the same time. So multi-clotting was this big thing towards the end of 2025, and it's only increased since then. And so we see this as, okay, great, one task works, and now you can do six tasks at a time. As the models get even smarter, the way that we're extrapolating this is, okay, next maybe you're going to run 50 clods at a time or hundreds of clods at a time. And so what is the infrastructure we need to build to enable that? At that point, you're probably not going to run everything locally on your machine anymore. There's just not enough RAM to do it. And so we're thinking about how do we make it easier for you to manage all these? These will probably run remotely. How do we build the interface so that you as a human know which tasks you need to look into? How do we make sure that the agent is fully verifying this work so that when you look at a task and it says it's done, you can very quickly verify and fully trust that it is done to your spec? And how do we make sure that this process is self-improving so that when you do see a

33:27

you're going to run 50 clods at a time or hundreds of clods at a time. And so what is the infrastructure we need to build to enable that? At that point, you're probably not going to run everything locally on your machine anymore. There's just not enough RAM to do it. And so we're thinking about how do we make it easier for you to manage all these? These will probably run remotely. How do we build the interface so that you as a human know which tasks you need to look, look into? How do we make sure that the agent is fully verifying this work so that when you look at a task and it says it's done, you can very quickly verify and fully trust that it is done to your spec?

33:36

And how do we make sure that this process is self-improving so that when you do see a task that isn't done to your liking, you can give it feedback and the model will know for every future run to incorporate that feedback? So it never makes that mistake again. So this is the progression that we're bringing our users along for. There's a lot of people listening, a lot of product managers, a lot of maybe founders, a lot of other cross-functional folks listening. There's a lot of worry about their role, the future of their careers.

33:43

What advice would you have for people to not just survive this transition to this very AI-driven world, but to be really successful, to essentially just thrive in this future? What are things people need to hear, need to be doing? I think AI gives everybody a ton more leverage than they used to. And so I would push you toward, anytime you realize that you're doing some manual task multiple times, thinking about how you can use Code, code, coworker, or other AI tools to automate that for you. Most people have creative parts of their job that they absolutely love and tedious parts of their job that they really hate doing.

33:49

I think the beauty of AI is that it can do those tedious parts for you. It can learn from every time that you've done that manual task and generalize and then run it automatically, so that you can focus on the creative parts. And that means you can do a lot more than you used to be able to do. So I think my immediate push for people is to figure out the repetitive parts that you can pass to Claude. Iterate on those automations until the success rate is very high. And then focus on, okay, what more can you be doing for your team, for your product, for your company that people haven't had the bandwidth to pick up so far?

33:54

Or what is that pet project that you always thought the company should do that you've never had bandwidth to do? If AI can take care of the grunt work, then you have this extra 20% time now that you might not have before. So my push is to lean into these tools, hand off the work that you're not excited to do, figure out how it can accelerate you, and then, as a result, you'll be able to do so much more. Something core to what you just shared, which I fully agree with, is find problems to solve with AI. There's all this potential, what all these tools can do. Some of the hard part for a lot of people is just, what should I actually do?

34:02

And what you're saying here is just pay attention to things that you are doing constantly you can automate. Pay attention to ideas that have been floating around that you haven't had time to do. It's basically solve a problem for yourself is kind of the core advice there. Exactly. I would also push listeners toward focusing on bringing your automations from, okay, this is a cool concept to, hey, this actually works 100% of the time. Sometimes I see users trying to automate something, getting it to 90, 95% accuracy, and then giving up on it. And if an automation doesn't work 100% of the time, it's not really an automation. And that last five to 10% does take more time.

34:11

Also, building the automation is often a lot slower than you doing it yourself. I would encourage listeners to put in that time to scope some automation that you really want to get to 100%, put in the elbow grease to teach quality of your preferences, to give it feedback so that it can improve its skill, so that it can get to that 100%. And then you'll be able to rely on it. There's just not much value in a 95% automation. I am super guilty of that. This is really good advice for me. I am guilty of this too. I've been teaching it. I've been teaching co-work to try to get me to inbox zero for Gmail.

34:22

And it has been very time-consuming, and it is definitely not there, as you probably realize. Yeah. Funny enough, that's exactly where my mind goes. I have this workflow I set up where every email I get, it looks for things that are spammy, which is just all these, hey, can I come on your podcast? Or what about this? All these things. I'm just, I don't have time for these sorts of things. And I have it categorize it into a folder called spammy. And it's 95% great. But then there's, oh wow, I missed an email because it went in there. So this is a good push for me. I'm going to work on this. I'm going to get it to perfect. Yeah.

34:38

We also are working on making the flow for customizing these commands a lot easier. Because right now I think you have to know too many concepts. You have to know to define a skill. You have to know to use this skill and give it feedback. And then you have to know to tell co-work to update the skill based on all the feedback that you gave. And then you also have to know where to read the skill to make sure that the feedback was incorporated the way that you want. That it's also our job to make this flow really seamless so that it doesn't feel painful to do. Amazing. Is there anything else, Kat, you wanted to share? Anything else you wanted to leave listeners with?

34:49

Anything you wanted to double down on that we haven't already touched on before we get to our very exciting lightning round? I see a lot of people playing around with AI and building prototype apps and tinkering with building workflows. I would really push people toward building apps that you're actually using every single day. Because I think only through that usage are you actually getting the value. If you build a prototype app that isn't helping you get more done, then the AI isn't really adding value to your day. And there's only so much you learn from that. And when it's, okay, I just did one shot at something. Oh, that's cool. And then you never come back to it.

34:57

You're not learning a lot. And you're not getting much leverage from it. And actual leverage. Yeah, that's such a good point. I also think there's a lot of people who spend a lot of time customizing their workflow. So there's two ends of the spectrum. One is people who never customize or never build automations. But there's this polar opposite end of people who obsess around customizing their tool, adding a ton of skills and MCPs and these workflow improvements. And I think sometimes that can even distract from your core goal of launching some product or building some feature.

35:06

I think there's a lot of fun in customizing, and we definitely want to make our products very hackable so that you can make it work really well for you, but there is a limit to how much it's useful. And I think there's a camp of people who maybe spend so much time customizing that they're not sleeping and not doing the core task that they originally set out to do. I see a lot of that on Twitter. Just, look at my setup. It's out of control. It's so optimized. And what are you actually building? No, but my setup is so awesome. I could get so much done. I think the simple setups actually work better. Hmm. Slash power up and get to level up a little bit. Yeah. Yeah.

35:21

There's this Karpathi tweet that just came out yesterday where he talked about this divide that's interesting between people that tried chat to PT Claude back in the day. It was, okay. And they're like, nah, this is terrible. And they kind of gave up on what AI could do for them. And they're just so cynical of, no way, it's not actually that big of a deal. And then there's people that are using it to code, essentially, who see the full intense power of it and how good it is. And people on both sides don't understand the other side and how they see the world. And so your advice is really good here.

35:30

Just actually use it for real things and see how good it actually has gotten. Yeah. Hmm. Slash power up and get to level up a little bit. Yeah. Yeah.

35:39

There's this Karpathi tweet that came out yesterday where he talked about this divide. That's interesting, between people that tried ChatGPT Claude back in the day. It was like, okay. And they're like, nah, this is terrible. And they gave up on what AI could do for them. And they're so cynical of, like, no way, it's not actually that big of a deal. And then there's people that are using it to code, essentially, who see the full intense power of it and how good it is. And people on both sides don't understand the other side and how they see the world. And so your advice is really good here. Actually use it for real things and see how good it actually has gotten.

35:41

Yeah. I think the big shift is that the 2024 generation of products were chat-based, and the Claude code generation of products is action-based. And the big aha moment people have is when Claude can just do things on your behalf. It is an amazing feeling to know that the agent is capable of doing so much more than telling you what to do. The agent can actually just do it itself. And when people feel that, I think that's the eye-opening moment. Shout out a Chrome extension, the Claude Called Chrome extension, which you can just watch it doing stuff that you'd be like, fill out this form for me. And it goes, right here I go. Exactly.

35:44

Okay. Anything else before we get to our very exciting lightning round? No, let's do it. Let's do it. Kat, I've got five questions for you. Welcome to the lightning round. There's this animation that plays. I have to make sure to say it. Are you ready? I'm ready. First question. What are two or three books that you find yourself recommending most to other people?

35:52

I really like How Asia Works. It's a story about economic development and what are the policies and governments that make long-lasting successful economies. The other book that I'm really into is The Technology Trap. So this is actually about the past few technology revolutions, so the Industrial Revolution and the computer revolution, and how this has affected workers. The reason that I really like this is because I think there's a lot we can learn from history to make sure that this transition goes well. And maybe on a fun note, I really like Paper Menagerie. It's a book of short stories about coming of age and AI and self-discovery.

35:52

Favorite recent movie or TV show you have really enjoyed? I really like Drive to Survive. There's no deeper meaning to it. There's just something very satisfying about people being so obsessed with a singular engineering goal and the purity of the pursuit. And I also really love Free Solo, which is about Alex Honnold climbing El Capitan without a harness. And I think similarly, it's just such a pure achievement to be able to climb this extremely challenging, dangerous route and to be able to have the mental focus to do it, knowing that if you make a single mistake, you die. It's insane.

35:56

Yeah. That movie is out of control. And it's interesting how these relate in some way to the work you do. I actually am a rock climber. I first watched Free Solo before I climbed rocks. So I thought it was impressive, but I didn't understand how impressive it was. It's one of the rare movies where, the more you know about it, the more you're blown away by how insane this is. The kinds of moves he's doing on the wall are things that I don't think I will ever be able to do in my lifetime, if it were set in a gym, one foot off the ground. With a rope. With a rope. Did you see the documentary on that other guy, the younger one that went on ice?

36:02

I did. That one was very sad. But that was wild. Okay. Favorite product you recently discovered that you really love?

36:04

The product that has most changed my life outside of Claude products is probably Waymo. I'm a diehard Waymo user. I use it twice a day to get to and from work. So the two things that I really like about it are, one, I don't feel bad if a Waymo is waiting for me. And so I feel less pressure to be right at the curbside the moment it arrives. And the second thing is, I feel like it lets me be a bit more productive. When I'm in the car with another human, I typically try not to do any work calls. I feel a little rude if I'm on my laptop the whole time. But one thing I really appreciate about the Waymo is I can call into a work call. I'm not worried about someone overhearing me. I'm not worried about, hey, is this rude? Am I talking too loud? Do I need to ask someone to change the music? And so this has given me back 30 minutes every day.

36:05

All these second-order effects of technology. It's so interesting. Yeah. I always thought Waymo needed to be priced lower than Uber and Lyft to succeed. But actually, I'm very happy to pay a 2x premium for it. I love Waymo. Once you see it, you're just like, this is insane. And then you get used to it. You get in there and you're like, this is crazy. And then you forget about it. Totally. And I think it's also changed the vernacular. A lot of people at Anthropic love Waymo. And I think in the past you'd be like, hey, what's it called? Will I write your app? And now everyone's just like, okay, is Waymo here?

36:08

Okay. Two more questions. Do you have a favorite life motto that you often come back to in work or in life? Yeah. Just do things. That tracks. That tracks.

36:12

I think there's a lot of value in first-principles thinking. And if you know what you're optimizing for and you have strong first principles, then you can normally deduce what the right course of action is and be able to clearly articulate that to all the stakeholders. And then you should just do it. I think jobs are fake. If you understand the constraints, you can figure out what you can do and then just try to do it quickly, learn from the mistakes, and apologize or fix them if you did something wrong. You could just do things.

36:13

Whoever said that. I think it's liberating, actually, to tell people this. I think in a lot of companies, roles are very strictly defined. Okay, this is what the PM does. This is what the designer does. This is what the engineer does. And then even team scopes are very rigidly defined. So, hey, this corner of the code base we touch, and this corner we're not allowed to touch. And I think what just do things lets people do is they feel empowered to make these decisions, empowered to operate across team boundaries, just to get something done.

36:15

That feels like a big, important skill to be good at. People call it agency. Just do the things that need to be done. Bias towards action. Bias towards action. All these ways of describing just get away for permission.

36:16

Yeah. I think this is my favorite reason to work at a startup at some point in your life, because one thing that was very life-changing for me was actually working at Scale when we were 20 people. And so there was just no process, and we had really big problems that we needed to solve. And I really appreciate Alex and the rest of the team for empowering me and the rest of the team to just figure things out without any boundaries for what sales is supposed to do, what ops is supposed to do, what engineers are supposed to do. You have all the tools at your disposal. You have some ambitious, hairy problem statement, and you can do whatever you need to get to a good solution.

36:17

You almost need that experience to build that skill, to feel comfortable doing that. Because a lot of people go through school or college and it's all these, do the thing we tell you to do and then you will get a good grade. And you have to kind of unlearn that of, okay, I'm just going to do the thing that needs to be done. And even if people think it's dumb, I think it's the right thing to do. Yeah, exactly. Okay. I actually have two more quick questions. Two more final questions. One is, when Claude thinks, there's all these, I don't know if you call them verbs. What's the term for these things? Thinking words.

36:21

Thinking words. And interestingly, these all leaked in the source code. You have some ambitious, hairy problem statement, and you can do whatever you need to get to a good solution. You almost need that experience to build that skill, to feel comfortable doing that. Because a lot of people, they go through school or in college and all these do-the-thing-we-tell-you-to-do, and then you will get a good grade. And you have to unlearn that of, okay, I'm just going to do the thing that needs to be done. And even if people think it's dumb, I think it's the right thing to do. Yeah, exactly. Okay. I actually have two more quick questions. Two more final questions.

36:31

One is, when Claude thinks, there's all these, I don't know if you call them verbs. What's the term for these things? Thinking words. Thinking words. And interestingly, these all leaked in the source code. Is it, do you have a favorite thinking word? Hmm. I really like manifesting. It's also the sticker that I have on my laptop. It's my favorite. Clearly the winner. Okay. Final question. Ask for us this too. With AGI potentially arriving in our lifetime, when you don't potentially have to work, what are you going to do? What are you going to do with all your time? I think it will take a long time for AGI to diffuse across society.

36:54

So I think the immediate thing is actually just helping bring the world along. I think my non-serious answer for after this happens is I'll probably just do a lot of rock climbing. I'll probably just live in some, I'll probably move to fountain blue and just live amongst 10,000 boulders and climb for a bit. There's also so many books I want to read that my goal is to be able to read one or two books a week. And I'm currently at probably 0.5. The backlog is pretty big.

37:03

I think there's just so much we can learn from history and so much that I don't understand as well as I would love to, like, I don't know anything about physics or robotics or any hardware or aerospace, or there's just so many interesting topics. So I'm excited to learn, even knowing that the AGI will already know it. Kat, this was amazing. You're awesome. Do you have all the questions? Where can folks find you online if they want to reach out and just follow what you're up to, and how can listeners be useful to you? The best way to reach out is I am underscore Kat Wu on Twitter. Feel free to tag me in things. Feel free to DM me. I read all my DMs.

37:13

I don't always respond to every single one, but I will read them all. And then the thing that is most helpful is, tell us where cloud code and co-work aren't working well for you. We are very grateful for the amount of positive feedback, but the thing that we thrive on is edge cases, errors, specific tasks that we can reproduce where cloud code or co-work fail. And so that's why we're able to share that with us and we're able to reproduce it. Then this is something that we're able to actively improve for our next generations of models and for our next harnesses. Jason Tucker, Ph.D.: Extremely cool. Everyone on, people on Twitter are not shy with sharing this feedback.

37:30

So keep it coming. Kat Wu, please, please share the problems that you're having with us. Jason Tucker, Ph.D.: Yeah. And it's really cool to see all you, your team being so active on Twitter and responding to people. And so what I'm hearing, this is actually stuff you guys actually see and react to. So. Kat Wu, Yeah. We appreciate everyone being so engaged with us. It gives the team a ton of energy. We have this channel of user love. And so whenever you guys share a success story, we post it there. And whenever you guys share issues with our product, we put it into our feedback channel. That way our broader team is able to act on it.

37:45

Jason Tucker, Ph.D.: That is so cool to know. Thanks for sharing that. Kat Wu, Well, Kat, thank you so much for being here. Kat Wu, Thanks for having me. Kat Wu, Bye everyone. Kat Wu, Thank you so much for listening. If you found this valuable, you can subscribe to the show on Apple Podcasts, Spotify, or your favorite podcast app. Also, please consider giving us a rating or leaving a review, as that really helps other listeners find the podcast. You can find all past episodes or learn more about the show at Lenny's podcast.com. See you in the next episode. that I woke up to this morning and I read through it and it was like pretty good. There were a few tweaks,

38:02

so I did have to give it a round of feedback. I like my slides to have extremely minimal words and it was a little too wordy. But, you know, it was far faster than like what I would be able to produce. And because Cowork has access to our whole design system, it actually looks like an anthropic designer put it together. Like it, when you visually see it, you're like, oh, this is like incredibly polished. So these are the kinds of things that are so much faster. Like this, making this slide deck would have taken me hours. But instead, it like turns out a draft that is actually quite good so that I could focus on making sure that the demos are amazing that we plug into it.

38:45

This sounds like a dream come true to PMs that putting decks together is so annoying. It's so slow. And I love people will see this deck whenever you present this. This will be out in the world. Like obviously, it's not the one-shotted version, but you've iterated on it. So just to help people try this for themselves. So step one is connect their, what did you say, Slack? What else do you suggest they connect? Slack, Google Calendar, Gmail, G Drive. You should connect your communications tools and where you store your source of truth data for what your team cares about, what you care about, and what you're working on. Okay. And then what was the prompt

39:22

roughly that you put in there to generate this deck? So I just wrote, make me a slide deck for the Code With Cloud conference. This is what our PMM suggested it should cover. This is the current draft that I made that I don't like. This is one that I made manually that I don't like, but I linked it. Can you start by creating a proposed outline with details? Also, make sure it doesn't overlap too much with a keynote talk, which is more important. And then Cod read a bunch of the links that I sent to it and created a proposed outline. So then I read through its proposal and all the different ideas that it generated for what we could cover. And I just made a decision

40:00

on what I wanted to actually be in the final deck. And I think this is like an example of what the role of the PM still is today. It's like, Claude is a great brainstorming partner. It's able to synthesize a massive amount of information really quickly and present all of the possibilities to you. But the role of the PM is still to make the end decision of, okay, what should belong in the final product? So for this, what I ended up deciding was that I wanted the talk to cover the progression from making local tasks successful to making every PR green to like helping engineers land more PRs. And for each of these, which demo would be the most compelling? And then

40:43

after this decision about the outline, co-work just like went off for a few hours and built the whole side deck. This is so awesome. What an awesome part of the job to not have to do anymore. And it feels like you're talking to essentially a deck designer that also has like actual knowledge about what you've worked on and can like make it actually the content which you want it to be, not just make it look really nice. How did you, how did you do the design system piece? How does that work? How does it know the design system of Anthropic? So what I did for this is we actually already have like a standardized deck that we use across all of our external engagements.

41:23

And so I just gave Claude access to that. And so it's able to see like what colors we use, the fonts we use, the different kinds of, what's it called? Like slide formats that are possible. And so it has like 20 of these example slides. So give an example. Got it. So you like upload, here's our template work from this. Yeah. You can also connect like your Figma MCP. If you, if you have your side format saved there and it can pull that in. Along those lines, something I'm always curious about is what's kind of in your, in your stack of tools as a PM and Anthropic. Obviously, Claude Code and Co-Work and all the Anthropic tools. What else are you using?

41:59

What other Slack you mentioned? Is there anything else? So my stack is pretty heavily Claude Code, Co-Work and Slack. Anthropic largely runs on Slack. I feel like it's like the core OS. of our company. And day to day, like a lot of, I would say maybe 30% of my time is pushing the boundaries of what Co-Work and Claude Code can do so that I have a very strong sense of what we're not good at. And I spend a lot of time talking with the model to understand why it makes mistakes that it does. We actually have a lot of internal tools that we make. Like, I think one of the things that Claude Code has really unlocked for our entire company is it really lowers the barrier

42:50

to making any custom app that you want. And so we've seen this like surge in personalized work software that people are building for like custom use cases instead of using tools that don't perfectly fit the use case. I got to hear more. What are some examples? What are things you've built other people have built that are really popular and useful? One of the sales folks on Claude Code, he realized he was making these like repetitive decks over and over and over again. And so he actually has this web app that he built with the examples of the core Claude Code decks that we know work well. So like a 101, 201 and mastering Claude Code. And then he has a way

43:33

to input specific customer context that pulls from Salesforce, that pulls from Gong, that pulls from other notes so that we can customize the decks for specific customers. And so we'll pull out things like, okay, this customer is using like Bedrock or Code for Enterprise or Console, which affects what features are available to them. It will pull out things like, okay, this customer is concerned about like the code review stage of the SLC. And so we'll add a slide about our code review features there. It'll pull out things like, okay, this customer needs to be like HIPAA compliant or needs XYZ security controls. And so we'll make sure to add a slide or two

44:12

in their deck about that. And then for example, if this is a customer that's on Vertex or Bedrock and doesn't want to use Claude for Enterprise, then we'll just take out some of the slides that are called for Enterprise only features. And so normally this is like manual work that could take 20, 30 minutes or, and so people either like spend that time doing it or they'll just decide not to do it and use the general deck. With this, it takes like a few seconds and you get a tailored deck. What's interesting about it is like Slack is like the tool that nobody's, it's just like, nobody's trying to create their own. Slack just continues to win and it's just like

44:50

the way you describe it is kind of the OS of so many companies. It's so interesting. Like people talk about Salesforce as just like SaaS. We don't need SaaS software anymore. We're going to build our own. It's like Slack is an adorable tool that nobody wants to try to compete with and build a better version. I think it's pretty important communications infrastructure and I think they do the core task of helping everyone get real-time updates incredibly well. Yeah, like people hate on Slack but it's really great at what it's trying to do. And like the most cutting-edge teams are hooked on it. So interesting. Yeah, and I also love how easy they've made to customize it.

45:25

And so it's we love making Slack bots and this kind of like hackability means that we're able to integrate with Slack the way that we want to. So really appreciate Slack's work on that. Time to buy some CRM stock. I am so excited to tell you about this season's supporting sponsor Vanta. Vanta helps over 15,000 companies like Cursor, Ramp, Duolingo, Snowflake, and Atlassian earn and prove trust with their customers. Teams are building and shipping products faster than ever thanks to AI. But as a result, the amount of risk being introduced into your product and your business is higher than it's ever been. Every security leader that I talk to is feeling the increasing weight

46:08

of protecting their organization, their business, and not to mention their customer data. Because things are moving so fast, they are constantly reacting, having to guess at priorities, and having to make do with outdated solutions. Vanta automates compliance and risk management with over 35 security and privacy frameworks, including SOC 2, ISO 27001, and HIPAA. This helps companies get compliant fast and stay compliant. More than ever before, trust has the power to make or break your business. Learn more at vanta.com slash Lenny. And as a listener of this podcast, you get $1,000 off Vanta. That's vanta.com slash Lenny. Okay, so you talked about all these different teams

46:50

and how they use cloud code and code to operate. Which teams do you find other than engineering? Imagine engineering is the biggest token spender, but if not, that'd be really interesting. What's kind of like the second place function right now for tokens? Oh, Applied AI is amazing at pushing the boundaries of what Cloud Code and Co-Work can do. A lot of our Applied AI team spends time with our customers, helping them adopt our API. And so sometimes our Applied AI team will, for example, make prototypes on behalf of these customers, which Cloud Code makes so much faster than it used to be. They also have the dual goal of needing to manage a lot of customer comms, a lot of

47:34

customer inbound and historical contacts, call notes. And so they're both extremely heavy on Co-Work and on Cloud Code. And just to understand Applied AI, does that like forward-to-play engineering sort of role? How would most people describe what the Applied AI team is doing? Yeah, it's helping our customers adopt the latest API and model features across their company, both for powering their company's products and also for internal acceleration. Got it. It's like customer success, go-to-market-y, kind of like forward-to-play engineering sort of thing. Exactly. It's like a very technical go-to-market person. Got it. Okay, awesome. So that's, so you're saying

48:14

that might be the second org that uses the most tokens. Yeah. And then we also see them pushing the boundaries of what Co-Work can do. So for example, if, so a lot of these folks cover multiple customers and in any given day can have like five to 10 customer engagements on a high day. And so what they often use Co-Work to do is the night before they'll ask it to summarize, okay, what are all my customer meetings that are coming up the next day? What are all the, what are all the things that this customer has asked me for? What's top of mind for them? What are the action items for the past meetings? And Co-Work will just put together this like dossier, this like

48:57

brief of what they should be aware of going into the next meeting. And Co-Work can also research answers. So if a customer asked, okay, when is feature X going to launch? Co-Work can help the Pi AI person research through Slack to get the latest ETA, add that to the notes so that during the customer call, the Pi AI person has the absolute latest. And these are just workflows that people are building for themselves and sharing with other people on their team. So cool. Something that kind of this question, this trend, I don't know, question topic comes up a lot recently, which is tokens spend exceeding people's salary where people just use AI and it costs more than how

49:39

much they're making. Are there any numbers floating around on topic of just like how much tokens spend, say engineers spend, I don't know, a month, a day, PMs, anything like that? It is clear to us that as the models get better, people delegate far more tasks to it and they spend a lot more hours in tools like CloudCode and Cowork. And so we do see the token cost per engineer or like per any knowledge worker increase every time that there is a model jump or like a substantial product improvement. I think it's still much lower than what the average engineer salary is, but we see the percentage increasing over time. It's such an interesting, like we talked about

50:23

how you have access to the most cutting-edge models, another advantage of working Anthropic. I believe you guys have basically unlimited tokens. You can use as much as you want. Is that right? We can use a lot of tokens. Some people do run into limits. Okay, there's a limit. Okay. Boris, shut it down. Okay. It's so interesting how many advantages come from having the most advanced model. It's such an interesting, like, flywheel that starts to kick in. I think we also believe a lot in empowering our internal teams to build as fast as possible. And we also trust that everyone understands how much capacity that serving these models truly costs. And we trust our team to use

51:05

the tokens responsibly. So it's very frowned upon to waste tokens, but we do trust individuals to make that judgment call. Awesome. Coming back to the PM role, we talked a little bit about this, but I think this will be really interesting for people to hear. just what I want to understand is what do you think are the kind of the emerging skills that PMs need to develop slash you most look for, AI companies most look for when they're hiring PMs these days? I think the hardest skill is being able to define what the product should look like a month from now. I think there's a lot of ambiguity in what models are capable of in that timeline and how user behavior will change.

51:52

But I think there are patterns that the best PMs can see based on how users are abusing the limits of the existing product and the best PMs can sense that, can set a direction and can steadily execute towards it and change the path if the model capabilities are much better than or worse than what they'd originally expected. I think it is very hard to be the right amount of AGI pilled. I think everyone can see this future where the models are extremely smart and can do almost everything in which case you actually don't need that complicated of a product. You can actually just have a text box again where you tell the model what you want and it's so smart that it can add any

52:38

tool or add any integration that it needs to get the job done. It knows when it's uncertain. It can ask clarifying questions. It's kind of very easy to build the product for the super AGI strong model. I think the hard thing is figuring out for the current model, how do you elicit the maximum capability? How do you help users go get onto the golden path? How do you guide users to interact with the model's strengths and patch its weaknesses? This skill is pretty rare. And how do you build that skill? Is it just basically understanding the limits of each model? You talked about taste, understanding, having taste into what the model maybe is capable of, what it's

53:30

great and not great at, where it's changed? I think it's spending a ton of time talking and using the model. One of the things I really like to do is to ask the model to introspect on its own behaviors. So sometimes when I notice that the model does something unexpected, like for example, there's like situations where the model will make a front-end change and run tests, but not actually use the UI. It's actually pretty useful to ask the model to reflect on why it did this. And sometimes they'll say that, hey, there was like something confusing in the system prompt, or I didn't realize, that the front-end verification was like part of the task, or hey, I delegated

54:12

the verification to this subagent and the subagent didn't do the test and I didn't check its work. A lot of times just like being very curious about why the model made the decision that it did will show you what misled it so that you can fix the harness in order to close this gap. The other thing that helps is to figure out who are the users who you trust the most to give you accurate feedback about the model. Usually there's like a handful of people who are much better than others at articulating what makes a specific model or model-harness combination good. And there's a lot of people who will give you feedback, but not everyone's feedback is as qualified. And so

54:57

finding a group of those like five people you trust is really important for getting very fast feedback. I think the third thing that is useful, but not everyone loves doing, is building evals. You don't need to build hundreds of evals for them to be useful. Just building 10 great evals is important for helping the team quantify what the goal is and what their progress towards it is and what they're missing. And so I think evals is this like underappreciated thing that more PMs, more engineers should be working on. We've covered evals a bunch. There's this trend of just like that is the future of product management is writing evals because essentially it's what

55:40

does success look like? Okay, cool. Let me actually concretely define it and then we'll know. How much of your time are you spending writing evals, would you say? I think the importance of evals varies a bit based on the feature that you're working on or like what the problem you're trying to solve is. So there are a lot of folks on our team who do spend a lot of time working on evals. We have a small pod of folks who collaborate very closely with research to more precisely understand our cloud code behaviors and what the largest areas of improvement are and trying to measure those pretty concretely. I personally jump into evals when there's a feature that I think needs a

56:21

bit more product definition and often the output of this is okay here are like five evals that I made this is how you run them these are the ones that succeed and these are the ones that don't and this is like the prompt that I've used to increase the success rate it varies a lot though based on the feature not every feature needs it but I think features such as memory benefit a lot from it this point you made about people being very good at evaluating models so interesting it's almost like a human eval of just like okay they understand where it's spiking or it's maybe lacking is there anyone specific that you want to shout out that's very good at this two people who I

57:02

think are incredible at because the task is so ambiguous even coding is easier because you can verify the success whereas crafting the character requires a very strong sense of conviction and who Claude should be and I think she has an incredible ability to not only mold the character but also to articulate what the goals are what the character what's successful and what's not the other group of people who I really trust is the cloud code team so we often have team lunches and whenever there's a new model we're testing one of the fastest ways for us to get feedback is to just go to every person and be like hey what is your vibe on the model and oftentimes we'll get

58:01

feedback like okay this model is not fully explaining its thinking it's too abrupt or hey this model just loves writing a ton of memories but we're not sure if the memories are high quality or not or like some people will notice that okay this model loves to test itself which is great or like this model isn't testing itself enough so that informs what data we look at to verify is this a larger pattern so we have a ton of data but it is very hard to extract what are the hypotheses we want to test and then we're able to extract data to test that this point you made about the character of Claude I had Ben Mann on the podcast co-founder and he talked about this just like the

58:51

character the constitution of Claude is such an important part of Claude and I didn't realize until afterwards just like people like with open actually one of the reasons people are sad is the personality of your Claude is because Claude's personality is so good and fun and interesting unlike other models and the way he put it is the personality is what makes Claude so good at so many things it feels like this like trivial side thing okay it's going to be funny and interesting and talk in a fun way but it's so core to the success of Claude is there anything get sure there about just like what people may not understand about why the character as you described and the

59:32

personality is so key when you reflect on everyone you've worked with there's just some people where you're like I really like their energy like I really like their vibe and when people think about Claude and Claude code this is one of the things that people bring up the most where they just really is extremely competent at your task people really like Claude's low ego and so if you tell it hey you did this thing wrong it's like truly sorry it's like oh shoot like thanks for telling me let me fix it let's work together it's also very positive so if you're feeling like oh this is an insurmountable task I think part of what makes a great coworker is this positivity this

1:00:32

like bias towards action this this ability to give you like earnest feedback not just agreeing with every single thing that you say and so we try to imbue this into COD because we think it makes it a lot more enjoyable to work with there's something I want to come back to you talked about how when new models come out you often have to revisit things you've built that's so interesting and so frustrating maybe just like oh god damn it we ship this thing now we have to rethink it talk about just like how often you have to come back with a new model and they're like okay we have to redo this product that we launched a few months ago a lot of the changes that we make with a so

1:01:19

the classic example for this is the to-do list when we first launched quad code people would ask it to do these large refactors and quad code would say okay cool I need to change these 20 call sites and it would go and change five of them and then stop and then we were like okay how do we like force it to remember to get every single one of these 20 and so Sid on our team was like okay what if we just like think about what a human would do a human would like make a list of everything that they need to change similar to how in VS code you would look up all the call sites and it will be on left side and you go through them one by one and ! how do we give a force it to use

1:02:08

this to do list it would like naturally use it itself for the earlier models we had to keep reminding it hey did you finish everything on the to to do list you can't finish until you're done with everything on the to list and for the later models without prompting it just like naturally things to do everything on the to list these days the to do list the model may use it it it is really not necessary for it to make thorough changes anymore I forgot who said this on the podcast that the model will eat your harness for breakfast and what I'm hearing here essentially you remove things over time that you've had to add on top of the model where it was not operating the way you

1:02:58

want it and essentially as the models get smarter it It becomes simpler and simpler for it just to do the thing you want it to do. Yeah. Um, we can move, remove a lot of prompting interventions every time the model gets smarter. And we actually do this every time we launch a model, we read through the entire system prompt and we reflect on, okay, for each of these sections, does the model really need this reminder anymore? And if not, we'll remove it. The most exciting thing that new models unlocks though, is just like entirely new features. So there's a lot of features that we've been testing out with prior models and

1:03:31

the accuracy wasn't high enough for us to want to launch them. And so one example of this is code review. We tried to build a code review product a few times and we've launched like simple versions of code review, which is the slash code review command in the past. And it was only with the most recent models that we felt like, okay, this code review is so good that our engineering team relies on this code review to do. to pass before we merge PRs. And we found that this was, we've always dreamed of Claude being able to be a reliable code reviewer that can actually, that we can like confidently feel catches the majority of bugs.

1:04:11

And it was only with like Opus 4.5 and 4.6 that we, and, uh, Sonic 4.6 that we felt like, okay, we are now able to like run multiple code review agents simultaneously to traverse, traverse the entirety of the code base and to synthesize a set of like real issues that an engineer needs to address before merge. And so this is like a new capability that the, the newest models have unlocked. This is another trend that is very common on this podcast of build something that will possibly be possible in the next six months. Be kind of at the edge of what's working sort of, and then it'll catch up and then it'll be an amazing product and you'll be ahead of everyone.

1:04:52

Yeah, exactly. Um, it's pretty important to build products that don't necessarily work yet so that you know, okay, what is missing, um, for this product to work. And then with the newest model, you can just swap it into the prototype you've already made and see, okay, does this new model close that gap? How much are you able to speak to just kind of where things are going with Claude and Co-Work as kind of the vision of it? I imagine you don't want to give away too much about the goal, but it feels like you're, there's all these awesome features being added on top, dispatch control from phone and all these mobile app, all these things.

1:05:27

What's kind of just like a way to understand the vision for all these things long-term. We think about this in terms of building blocks. So for both Claude and Co-Work, the core building block is making individual tasks successful. So you, you want it to produce some output. You give it a clear prompt description. Is it able to consistently produce acceptable output that you're able to either merge or share with your colleagues or external audience? So the task is the core building block. As the models get smarter, the task success rate gets a lot higher. And then we see people moving towards doing multiple tasks at the same time.

1:06:05

So multi-clotting was this big thing and towards the end of 2025, and it's only increased since then. And so we see this as, okay, great. One task works and now you can do like six tasks at a time. As the models get even smarter, the way that we were extrapolating this is, okay, next, maybe you're going to run like 50 clods at a time or hundreds of clods at a time. And so what is the infrastructure we need to build to enable that? At that point, you're probably not going to run everything locally on your machine anymore. There's just like not enough RAM to do it. And so we're, we're thinking about how do we make it easier for you to manage all these?

1:06:44

These will probably run remotely. How do we build the interface so that you as a human know which tasks you need to look, look into? How do we make sure that the agent is fully verifying this work so that when you look at a task and it says it's done, you like can very quickly verify and fully trust that it is done to your spec. And how do we make sure that this like process is self-improving so that when you do see a task that isn't done to your liking, you can give it feedback and the model will know for every future run to incorporate that feedback. So it never makes that mistake again. So this is the progression that we're, we're bringing our users along for.

1:07:23

There's a lot of people listening, a lot of product managers, a lot of maybe founders, a lot of other cross-functional folks listening. There's a lot of worry about just how their role, the future of their careers. What advice would you have for just people to not just survive this transition to this very AI driven world, but to be really successful to essentially just to thrive in this future? What are just like things people need to hear, need to be doing? I think AI gives everybody a ton more leverage than they used to. And so I would push you towards anytime you realize that you're doing some manual task multiple times,

1:08:00

think about how you can use code code, coworker or other AI tools to automate that for you. Most people have like creative parts of their job that they absolutely love. And then like tedious parts of their job that they really hate doing. I think the beauty of AI is that it can do those tedious parts for you. It can learn from every time that you've done that manual task and generalize and then run it automatically. And so that you can focus on the creative parts. And that means you can do a lot more than you used to be able to do. So I think my like immediate push for people is figure out the repetitive parts that you can pass to Claude.

1:08:39

Iterate on those automations until the success rate is very high. And then focus on, okay, what more can you be doing for your team, for your product, for your company that like people haven't had the bandwidth to pick up so far? Or like, what is that like pet project that you always thought the company should do that? Like you've never had bandwidth to do. If AI can take care of the like grunt work, then you have, you have this extra 20% time now that you might not have before. So, so my push is to lean into these tools, hand off the work that you're not excited to do, figure out how it can accelerate you. And then as a result, you'll be able to do so much more.

1:09:18

Something core to what you just shared, which I fully agree with is find problems to solve with AI. There's all this potential, what all these tools can do. Some of the hard, like for a lot of people, artist part is just like, what should I actually do? And what you're saying here is just pay attention to things that you are doing constantly. You can automate, pay attention to just like ideas that have been floating around that you haven't had time to do. It's basically, it's like solve a problem for yourself is kind of the core advice there. Exactly.

1:09:45

I would also push listeners towards focusing on bringing your automations from, okay, this is a cool concept to like, hey, this actually works 100% of the time. Like sometimes I see users trying to automate something, getting it to like 90, 95% accuracy and then giving up on it. And this, if an automation doesn't work 100% of the time, it's not really an automation. And that last five to 10% does take more time. Also building the automation is often a lot slower than you doing it yourself.

1:10:20

I would encourage listeners to put in that time to scope some automation that you really want to get to 100%, put in the elbow grease to teach quality of your preferences, to like give it feedback so that it can improve its skill so that it can get to that 100%. And then like really, then you'll be able to rely on it. There, there's just not much value in a 95% there automation. I am super guilty of that. This is really good advice for me. I am guilty of this too. I've been teaching it. I've been teaching co-work to try to get me to inbox zero for Gmail. And it has not been, it has been very time consuming and it is definitely not there as you probably realize. Yeah.

1:11:02

I funny enough. That's exactly where my mind goes. I have this workflow I set up where every email I get, it looks for things that are spammy, which is just like all these like, Hey, can I come on your podcast? Or what about this? Like all these things. I'm just like, I don't have time for these sorts of things. And I have it categorized it into a folder called spammy. And it's just like, it's 95% great. But then there's like, oh, wow, I missed an email because it went in there. So this is a good push for me to like, I'm going to work on this. I'm going to get it to perfect. Yeah. We also are working on making the flow for customizing these commands a lot easier.

1:11:34

Because right now I think you have to like know too many concepts. You have to know to define a skill. You have to know to like use this skill and give it feedback. And then you have to know to tell co-work to update the skill based on all the feedback that you gave. And then you also have to know where to read the skill to like make sure that the feedback was incorporated the way that you want. That it's also our job to make this flow really seamless so that it doesn't feel painful to do. Amazing. Is there anything else Kat you wanted to share? Anything else you wanted to leave listeners with?

1:12:02

Anything you wanted to double down on that we haven't already touched on before we get to our very exciting lightning round? I see a lot of people playing around with AI and building like prototype apps and tinkering with building workflows. I would really push people towards building apps that you're actually using every single day. Because I think only through that usage are you actually getting the value. Like if you build a prototype app that isn't helping you get more done, then the AI isn't really adding value to your day. And there's only so much you learn from that. And when it's like, okay, I just did one shot at something. Oh, that's cool.

1:12:42

And then you never come back to it. It's like you're not learning a lot. And you're not getting like much leverage from it. And actual leverage. Yeah, that's such a good point. I also think there's a lot of people who spend a lot of time like customizing their workflow. So there's like, I think there's like two ends of the spectrum. One is like people who never customize or never build automations. But there's like this polar opposite end of people who like obsess around customizing their tool, like adding a ton of skills and MCPs and these like workflow improvements. And I think sometimes that can even distract from your core goal of like launching some product or

1:13:17

building some feature. I think there's a lot of fun in customizing. And we definitely want to make our products very hackable so that you can make it work really well for you, but there is a limit to how much it's useful. Um, and I think there's a camp of people who maybe spend so much time customizing that they're like not sleeping and not doing the like core task that they originally set out to do. I see a lot of that on Twitter. Just like, look at my setup. It's out of control. It's so optimized. And what are you, what are you actually building? No, but my setup is so awesome. I could get so much done. I think the simple setups actually work better. Hmm.

1:13:56

Slash power up and get to level up a little bit. Yeah. Yeah. There's this Karpathi tweet that just, uh, came out yesterday where he talked about this divide. That's interesting between people that tried chat to PT Claude back in the day. It was like, okay. And they're like, nah, this is this terrible. And they kind of gave up on like what AI could do for them. And they just like, so cynical of like, no way, it's not actually that big of a deal. And then there's people that are using it to code essentially who see the full intense power of it and how good it is. And people on both sides don't understand the other side and why they like how much they,

1:14:30

how they see the world. And so your advice is really good here. Just like actually use it for real things and see how good it actually has gotten. Yeah. I think the big shift is that the 2024 generation of products were chat based and the Claude code generation of products is action based. And the like big aha moment people have is when Claude can just like do things on your behalf. It is, it is an amazing feeling to know that the agent is capable of doing so much more than telling you what to do. Like the agent can actually just do it itself. And when people feel that I think that's the eye opening moment. Shout out a Chrome extension, the Claude called Chrome extension,

1:15:13

which you can just watch it doing stuff that you'd be like, fill out this form for me. And I go right here I go. Exactly. Okay. Uh, anything else before we get to our very exciting lightning round? No, let's do it. Let's do it. Uh, Kat, I've got five questions for you. Welcome to the lightning round. There's this animation that place. I have to make sure to say it. Uh, are you ready? I'm ready. First question. What are two or three books that you find yourself recommending most to other people? I really like how Asia works. It's a story about economic development and what are the, like the policies and governments that make long lasting successful economies.

1:15:52

The other books that I'm really into are the technology trap. So this is actually about the past few technology revolution. So the industrial revolution and the computer revolution and how this has affected workers. The, the reason that I really like this is because I think we, there's a lot we can learn from history to make sure that this transition goes well. And, um, maybe on like a fun note, I really like paper menagerie. Um, it's just like a book of short stories about like coming of age and AI and, um, just like self discovery. Favorite recent movie or TV show you have really enjoyed? I really like drive to survive. There's no like deeper meaning to it.

1:16:39

I just, there's just something very satisfying about people being so obsessed with like a singular engineering goal. And just like the purity of the pursuit. Um, and I also really love Free Solo, which is about Alex Honnold, um, climbing El Capitan without a harness. And I think similarly, it's just such a pure achievement to be able to climb this extremely challenging, dangerous route. And to be able to have the mental focus to do it, knowing that if you make a single mistake, you die. It's insane. Yeah. That movie is out of control. And it's interesting how these relate in some way to the work you do. I actually am a rock climber.

1:17:23

Um, I first watched Free Solo before I climbed rocks. And so I thought it was impressive, but I didn't understand how impressive it was. It's one of the rare movies where like, the more you know about it, the more you're, you're blown away by how insane this is. Like the kinds, the kinds of movies he's doing on the wall are things that like, I don't think I will ever be able to do in my lifetime. If it were set in a gym, like one feet off the ground. With a rope. With a rope. Did you see the documentary and that other guy, the younger one that went on like ice? I did. That one was very sad. But that was, that was wild. Okay.

1:17:58

Uh, favorite product you recently discovered that you really love? The product that is like most changed my life outside of Claude products is probably Waymo. Like I'm a diehard Waymo user. Um, use it twice a day, get to and from work. So the two things that I really like about it are one, I don't feel bad if a Waymo is waiting for me. And so I feel like I feel less pressure to be right at the curbside the moment it arrives. And the second thing is, I feel like it lets me be a bit more productive. Um, when, when I'm in the car with another human, I, I typically try not to like do any work calls. I, I feel a little rude if I'm like on my laptop the whole time.

1:18:38

But one thing I really appreciate about the Waymo is I can call into a work call. I'm not worried about someone overhearing me. I'm not worried about, Hey, is this like rude? Am I talking too loud? Do I need to tell, ask someone to like change the music? And so this has been like, I feel like this has given me back like 30 minutes every day. All these second order effects of, of technology. It's so interesting. Yeah. I always thought Waymo needed to be priced lower than Uber and Lyft to succeed. But actually I'm like very happy to pay a 2x premium for it. I love Waymo. It's just like, like once you see it, you're just like, this is insane. And then you get used to it.

1:19:13

Like you get in there and you're like, this is crazy. And then you forget about it. Totally. And I think it's also changed the vernacular. Like a lot of people at Anthropic love Waymo. And I think in the past you'd be like, Hey, like what's called like, well, I write your app. And now like everyone's just like, okay, is Waymo here? Okay. Two more questions. Do you have a favorite life motto that you often come back to in work or in life? Yeah. Just do things. That tracks. That tracks. I think there's a lot of value in like first principles thinking. And if, if you like, if you know what you're optimizing for and you have like strong first

1:19:46

principles, then you can normally deduce what the right like course of action is and be able to clearly articulate that to all the stakeholders. And then you should just like do it. Like, I think jobs are fake. If you understand the constraints, you can figure out what you can do and then just like try to do it quickly, learn from the mistakes and apologize or fix them. If you did something wrong. You, you could just do things. Whoever said that. I think it's liberating actually to like tell people this. I think in a lot of companies, like roles are very strictly defined. Like, okay, this is what the PM does. This is what the designer does. This is what the engineer does.

1:20:21

And then even team scopes are very rigidly defined. So, Hey, like this corner of the code base, we touch and this corner, like we're not allowed to touch. And I think what just do things lets people do is they feel like empowered to make these decisions, empowered to operate across team boundaries, just to like get something done. That feels like a big, important skill to be good at. People call it agency. Just like do the things that need to be done. Bias towards action. Bias towards action. All these ways of describing just like get away for permission. Yeah. I think this is my favorite reason to work at a startup at some point in your life,

1:20:54

because like one thing that was like very life changing for me was actually working at scale when we were 20 people. And so there was just no process and we have like really big problems that we needed to solve. And it was like, I really appreciate Alex and the rest of the team for like empowering me and the rest of the team to just like figure things out without any boundaries for what sales supposed to do, what ops supposed to do, what engineers supposed to do. Just like you have all the tools at your disposal. You have some like ambitious, hairy problem statement and you can do whatever you need to like get to a good solution.

1:21:28

Like you almost need that experience to build that skill, to feel comfortable doing that. Because a lot of people, you know, they go through school or in college and all these like do the thing we tell you to do and then you will get a good grade. And you have to kind of unlearn that of like, okay, I'm just going to do the thing that needs to be done. And even if people think it's dumb, I think it's the right thing to do. Yeah, exactly. Okay. I actually have two more quick questions. Two more final questions. One is when Claude thinks there's all these, I don't know if you call them verbs. What's the term for these things? Thinking words. Thinking words.

1:21:57

And interestingly, these all leaked in the source code. Uh, is it, do you have a favorite thinking word? Hmm. I really like manifesting. It's also like the sticker that I have on my laptop. It's my favorite. Clearly the winner. Okay. Final question. Ask for us this too. With AGI potentially arriving in our lifetime, when you don't potentially have to work, what are you going to do? What are you going to do with all your time? I think it will take a long time for AGI to diffuse across society. So I think the immediate thing is actually just like helping bring the world along. I think my like non-serious answer for after this happens is I'll probably just do a lot of rock

1:22:39

climbing. I'll probably just like live in some, I'll probably move to like fountain blue and just like live amongst 10,000 boulders and climb for a bit. There's also so many books I want to read that my, my goal is to be able to read one or two books a week. And I'm currently at probably like 0.5.

1:23:02

The backlog is pretty big. I think there's just like so much we can learn from history and so much that I don't understand as well as I would love to, like, I don't know anything about physics and, or like robotics or like any hardware or like aerospace, or there's just so many interesting topics. So I I'm excited to learn even, even knowing that the AGI will already know it. Kat, this was amazing. You're awesome. Do you have all the questions where can folks find you online if they want to reach out and just follow what you're up to and how can listeners be useful to you? The best way to reach out is I am underscore Kat Wu on Twitter. Feel free to like tag me in things.

1:23:44

Feel free to DM me. I read all, all my DMs. I don't always respond to every single one, but I will read them all. And then the thing that is most helpful is tell us where cloud code and co-work aren't working well for you. We, we are very grateful for the amount of positive feedback, but the thing that we thrive on is edge cases, errors, like specific tasks that we can reproduce where cloud code or co-work fail. And so that's why we're able to share that with us and we're able to reproduce it. Then this is something that we're able to actively improve for our next generations of models and for our next harnesses. Jason Tucker, Ph.D.: Extremely cool.

1:24:26

Everyone on, people on Twitter are not shy with sharing this feedback. So keep it coming. Kat Wu, please, please share the problems that you're having with us. Jason Tucker, Ph.D.: Yeah. And it's really cool to see all you, your team being on so active on Twitter and responding to people. And so, so like what I'm hearing, like, this is actually stuff you guys actually see and react to. So. Kat Wu, Yeah. We appreciate everyone being so engaged with us. Um, it gives the team a ton of energy. We, we have this channel of like user love. And so whenever you guys share a success story, we post it there.

1:24:54

And whenever you guys share like issues with our product, we put it into our feedback channel. That way our broader team is able to act on it. Jason Tucker, Ph.D.: That is so cool to know. Thanks for sharing that. Kat Wu, Well, Kat, thank you so much for being here. Kat Wu, Thanks for having me. Kat Wu, Bye everyone. Kat Wu, Thank you so much for listening. If you found this valuable, you can subscribe to the show on Apple podcasts, Spotify, or your favorite podcast app. Also, please consider giving us a rating or leaving a review, as that really helps other listeners find the podcast. You can find all past episodes or learn more about the show at Lenny's podcast.com.

1:25:30

See you in the next episode.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note