Lenny's Podcast

Why AI is going vertical (again) | Dianne Penn (Anthropic)

2231 summary words 10 min summary Watch video

Start with the signal

10 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Anthropic's Diane Penn argues that frontier-AI product work is becoming an eval-driven, hands-on discovery discipline: teams must continuously uncover emergent model capabilities, turn real failures into measurable tests, and build product harnesses that make those capabilities usable.
  • Why it matters: This is a high-signal account of how a leading lab connects research, product, safety, agent UX, and organizational culture—and offers reusable operating patterns for AI-native product and agent teams.
  • Best use: Use it to redesign product discovery around evals and trajectory review, strengthen your AI experimentation culture, and pressure-test whether your products remain useful as model capability jumps.

Executive Summary

Penn frames Anthropic's rise not as a single-model victory but as a sequence of product-and-research feedback loops. Opus 3 gave Anthropic an early coding differentiation after the team noticed users moving from autocomplete toward long-form code generation. Later, Claude Code and Opus 4.5 reinforced each other: the product provided a vehicle for users to experience frontier capability, while the stronger model accelerated adoption. Her central operating premise is that frontier models need frontier products, not merely an API or chatbot.

The practical core is her description of research PM work. At Anthropic, "evals are the new PRDs" because a vague complaint such as "Claude hallucinated" is useless until a PM reconstructs the user trajectory, diagnoses the failure mode, creates a representative test set, and gives researchers a measurable target. Her example: investigation of Claude 2's poor instruction following found that roughly 80% of complaints concerned malformed JSON; a 30–40-example eval set turned that vague problem into a durable model-quality metric. PRDs remain useful for cross-functional alignment and ambiguous product visions, but they are no longer sufficient for model-quality work.

Penn also argues against passive AI adoption. She accepts the underlying point of "token maxing"—those who use frontier models heavily gain an early view of future workflows—but reframes the objective as experimentation rather than spend. The strongest practitioners work directly with each model version, share discoveries publicly inside the organization, and collaborate rather than trying to invent use cases alone. Leaders cannot delegate this learning: even senior PM managers should continue shipping work, reading user feedback, and maintaining direct intuition for model behavior.

The interview's broader implication is that AI systems are discontinuous and jagged-edged: capabilities can emerge abruptly, while adjacent abilities such as tone, writing, safety, or proactive behavior remain uneven. Organizations therefore need adaptable roadmaps, strong evals, safety/red-team processes, fallback UX, and low-ego teams able to revise plans quickly. Penn sees enduring human value in judgment, persistence, original point of view, and deciding what is worth building—not in merely producing artifacts that increasingly capable models can generate.

Key Takeaways

  • Claim: Frontier model adoption depends on a complementary product vehicle; a capable model alone does not reliably create user-visible magic. | Evidence: Penn says Opus 4.5 and Claude Code were mutually reinforcing: Claude Code made frontier intelligence tangible through end-to-end agentic coding, while Opus 4.5 accelerated Claude Code adoption. Earlier, Opus 3 differentiated Anthropic after the team trained for long-form coding rather than just autocomplete. | Implication: For agent products, invest in the harness, workflow, context, and UX that expose model capability; do not treat model upgrades as a complete product strategy. | Caveat: The transcript provides Anthropic's internal interpretation, not comparative adoption data or a causal measurement study.
  • Claim: AI product discovery must assume discontinuous capability jumps and substantial product/user overhang, so fixed roadmaps should yield to rapid evaluation and adaptation. | Evidence: Penn cites scaling-law capability curves where particular tasks jump from failure to reliable success rather than improving smoothly. She says teams often cannot predict the exact model generation when a capability appears and must use evals, prototypes, safety tests, and new products to find out what became possible. | Implication: Maintain scenario-based roadmaps around future capability levels, but allocate recurring capacity to test newly possible workflows immediately after model releases. | Caveat: Emergent capabilities are not presented as universally unpredictable; Penn's point is that their exact thresholds are hard to know without targeted measurement.
  • Claim: For model-centric products, evals are often more operationally important than PRDs because they convert user pain into an actionable and measurable training target. | Evidence: Anthropic's research-PM saying is "evals are the new PRDs." In the Claude 2 era, Penn investigated reports that Claude did not follow instructions, found about 80% referred to incorrect JSON output, assembled 30–40 failing examples, and added them to an eval repository; the issue later reached roughly 99.9–100% performance. | Implication: Instrument your AI product so you can inspect consented trajectories, classify failures by mechanism, create representative regression sets, and use those evals as the shared contract between product, engineering, and model teams. | Caveat: Penn explicitly says PRDs are not dead: Anthropic still creates them for each model, stakeholder alignment, safety/legal/product coordination, and ambiguous zero-to-one opportunities.
  • Claim: Senior product leadership in AI must remain hands-on with model behavior and shipping work; managerial abstraction is especially risky in a fast-moving, non-deterministic domain. | Evidence: Penn says Anthropic uses the same onboarding fundamentals for tenured PMs and junior hires: understand users, read consented feedback, and talk to customers. She personally retains one or two model-release workstreams to preserve her intuition about model progress and decision quality. | Implication: Require AI product leaders to regularly build, test, review agent traces, and own real releases—not just approve strategy, dashboards, or headcount plans.
  • Claim: Heavy token use is valuable only insofar as it produces sustained experimentation, shared discovery, and better ideas—not as a spend target by itself. | Evidence: Responding to Gary Tan's idea that spending $100,000 annually on tokens means living like a 2028 user, Penn says token spend is an input while experimentation is the desired output. She describes Anthropic's early company-wide Slack channel where employees posted model experiments; others iterated on them and sometimes found compelling use cases within roughly ten requests. | Implication: Create a visible internal practice for sharing prompts, agent runs, prototypes, failures, and workflow discoveries; budget experimentation around learning velocity and valuable workflow unlocks rather than raw token consumption. | Caveat: The interview does not offer a spending threshold or ROI model for token-intensive experimentation.
  • Claim: Safety requirements for more capable models must be designed as product systems, including red-teaming, pre-release controls, and graceful fallback experiences. | Evidence: Penn says that as frontier models become more capable, Anthropic must evolve safeguards, testing, and pre-release processes. She gives fallback UX as an example: when access to a newer model is restricted or controlled, users should still receive an immediate useful response from Opus 4.8 rather than a dead end. | Implication: Treat safety restrictions, model routing, permissions, and fallback models as core UX and control-plane design problems; a safe refusal or restriction without a productive alternative can destroy user value. | Caveat: Penn defers policy and access-control specifics to Anthropic's relevant experts, so the interview does not detail its actual control architecture.
  • Claim: The most durable human contribution is not rote production but judgment, persistence, proactive problem selection, and an independent point of view supported—not replaced—by AI. | Evidence: Penn says judgment accumulates from nuanced human experience, while choosing which of the many possible AI-buildable things an organization should pursue requires persistence and proactivity. Personally, she forms her own point of view before using Claude as a sparring partner, while delegating lower-value standardized writing such as business-review drafting for human verification. | Implication: Use models to challenge and extend your reasoning, but preserve human ownership of framing, priorities, judgment calls, and final accountability—especially for consequential decisions. | Caveat: She does not claim these capabilities are permanently exclusive to AI; she describes them as particularly valuable during the current transition.

Detailed Brief

How Anthropic's labs and research-product interface operates

  • Claims: Anthropic Labs is designed to identify discontinuous bets outside the core roadmap and determine whether there is a real opportunity and a potentially 10x, 100x, or 1,000x version of it.; The team can hold a strong conviction about a problem area while remaining weakly attached to a particular prototype, revisiting ideas one or two model generations later when underlying capabilities improve.; Research work is both long-horizon and operational: researchers articulate ambitious futures such as computer use while also inspecting training runs, data, evals, and immediate model deficiencies.
  • Evidence: Penn names Claude Code, MCP, skills, and Claude Design as Labs-associated bets.; She says Lab pods are intentionally small and can begin with a single engineer; the organization selects for people willing to act like founders and tolerate having bets shut down.; For research PMs, a user complaint must be decomposed into specific components such as tool-use failure, retrieval/search failure, incorrect fact synthesis, or alignment/overconfidence failure before it is useful to research.
  • Caveats: The transcript supplies no formal selection rubric, success rate, or resource allocation model for Labs bets.; The labels and products mentioned reflect the transcript and are not independently validated here.
  • Implications: A separate incubation function is useful when it protects small teams from core-roadmap gravity while preserving a path for successful experiments to influence the platform.; Build a translation layer between user language and technical failure modes; without it, customer feedback will not efficiently improve models or agents.

Operating culture: public experimentation, low ego, and sustainable high velocity

  • Claims: Penn attributes Anthropic's speed partly to a bottoms-up culture where engineers, designers, researchers, and product people can initiate and support experiments across organizational lines.; Experimentation is social rather than individual: sharing a promising use case can create rapid iterations and broader adoption.; High-performance AI work is sustainable only when teams combine radical ownership with coverage, mutual review, and low-ego collaboration.
  • Evidence: Golden Gate Claude—an interpretability demonstration in which Claude obsessively referenced the Golden Gate Bridge—was built into Claude.ai in approximately 24 hours by cross-functional volunteers and reached about 2,000 users.; Penn says colleagues who are not the directly responsible individual for a model launch will still help review launch materials, improve demos, and cover each other.; She says the company shipped four model series in all of 2024 but exceeded that volume in Q2 of the subsequent year, according to her account.
  • Caveats: The Golden Gate Claude launch was a small experimental demonstration, not evidence that every cross-functional project can or should be completed at that speed.; A culture of constant launch support can become unsustainable without explicit prioritization, staffing coverage, and real time off.
  • Implications: Institutionalize demo days, shared trace reviews, internal prompt/prototype channels, and cross-functional launch coverage rather than relying on isolated individual experimentation.; Hire and reward contributors for collective impact and decision quality, not primarily empire-building or ownership signaling.

Notable Concepts & Terms

  • Evals are the new PRDs: Penn's shorthand for treating representative, measurable tests of user pain as the core artifact that connects product feedback to model improvement.
  • Sweat the tokens as much as the pixels: AI product teams must inspect prompts, context, tool calls, trajectories, failures, and non-deterministic outputs as closely as traditional PMs inspect screens and flows.
  • Product overhang / user overhang: Current models may already support valuable workflows that products have not exposed and users have not yet discovered.
  • Jagged edge: Model capabilities are uneven: a system can be highly capable in one area while retaining striking weaknesses in adjacent tasks such as writing, tone, tool use, or reliability.
  • Frontier products: Purpose-built product experiences, such as Claude Code, that let users access and trust frontier-model capabilities in end-to-end workflows.
  • Strongly held theme, weakly held prototype: Labs' approach to zero-to-one work: commit to an important opportunity area without becoming attached to the first implementation.
  • Fallback UX: A product response path that preserves useful user outcomes when a more capable or restricted model cannot be used.
  • Think first, then use AI as a sparring partner: Penn's method for avoiding passive dependence: retain an initial human POV, then use the model to challenge, refine, and extend it.

Operator Notes / Why Ken Should Care

  • Establish an eval pipeline for every AI workflow: capture representative failures, label the underlying mechanism, create regression cases, and require releases to report eval movement rather than only qualitative demos.
  • Add a model-capability review to roadmap planning: for each major initiative, ask what changes if the next two model generations can autonomously execute materially more of the workflow.
  • Run a weekly internal AI discovery forum where operators share successful agent runs, prompts, skills, tool configurations, and failure traces; prioritize replication and variation over isolated tinkering.
  • Require senior AI-product and operations leaders to maintain a direct-build cadence, including shipping at least one meaningful workflow or owning live agent-quality reviews each quarter.
  • Design model routing and safety controls with productive degradation: specify what a user receives when a preferred model, tool, permission, or action is unavailable.
  • Use AI for consequential communication as preparation and adversarial coaching, but mandate a human-authored initial perspective and named human verifier for decisions, external communications, and business-critical reports.

Source/Metadata

  • Title: Anthropic’s first technical PM on token maxing, the jagged edge, and living in the future
  • Transcript words: 26718
  • Duration seconds: 5630
  • Timestamp note: No usable timestamps or chapters were present in the supplied transcript; the transcript also contains repeated segments and sponsor material.
Full transcript 15544 words · 118 min read
0:00

In 2023, when I started, nobody said Anthropic and Claude and coding in the same sentence. I want to go back to the beginning of Anthropic. I remember feeling, man, these guys have no chance. OpenAI is so far ahead. At the time, I saw people were starting to use these models not just for code autocomplete, but actually writing long-form code. It's an opportunity for us to train Opus 3 to be better at. That was the inflection. I always think about Opus 4.5 a year later during winter break, when everyone was home, able to code. What was magical about Opus 4.5 is we now not just had a model, but a vehicle, a great product

0:38

experience like Claude Code. Opus 4.5 wouldn't have had that moment without a product like Claude Code, and Claude Code wouldn't have had that type of adoption accelerated without Opus 4.5. I want to talk about how the product role is changing. For my team, the way to drive user value is to figure out the right user feedback. The evals, we actually have a saying on the team: evals are the new PRDs. Something Gary Tan's been talking about: if you're willing to spend $100,000 a year right now in tokens, you are living the way somebody in 2028 is going to live. You have to sweat the tokens as much as you sweat the pixels. You have to be using the models

1:15

to come up with good, then great, then better ideas, and there's no substitute for that. People need to be more ambitious with AI tools these days because they're just capable of so much. One thing I ask the team is, let's say Claude 8 comes around, what changes in what users do? What does that mean for how you're building today? Today, my guest is Diane Penn, head of product for the AI research and labs teams at Anthropic. She joined Anthropic as the first technical product manager over three years ago, which is a lifetime in AI time, when the product team was just five engineers. She's helped ship every

1:51

model at Anthropic from Claude 2 through Fable. She's also helped incubate and launch Claude Code, MCP, skills, Claude design, and also core capabilities like computer use, tool use, and reasoning. It is always such a treat and so mind-expanding to get to talk to someone who's at the very center of AI and product management. It's hard to imagine someone who has seen more of where things are going than the head of product for Anthropic's research and labs teams. Before we get into it, don't forget to check out Lenny'sProductPass.com for a year free of the hottest and most beautifully crafted AI products in the world, available exclusively to Lenny's newsletter subscribers.

2:30

With that, I bring you Diane Penn.

2:36

Diane, thank you so much for being here. Welcome to the podcast. Thank you, Lenny. It's so nice to see you again. I want to go back to the beginning of Anthropic, the early days. I remember when Anthropic first launched. This was, I don't know, the first model, when it launched years ago, three years ago, something like that. It was. Three years. I remember feeling that, man, these guys have no chance. OpenAI is so far ahead. Everyone's just like, what are they thinking? How is this possible? OpenAI has won. It's too late. Things are very different now. The latest number I saw was Anthropic was making, I don't know,

3:14

$50 billion in ARR. That's what companies used to go public at. Very successful companies went public at $50 billion in valuation. Anthropic reportedly is making that every single year. You joined as one of the earliest PMs. There were something like five engineers when you joined. The model hadn't even launched when you joined. What was it like in those early days of Anthropic? What's something that might surprise people about what it was like at the beginning? I think a big part of what's made Anthropic today actually has been very much the core of even the early days. I joined in 2023. Like you said, we had five product engineers. There was one

3:59

engineer for the entirety of our API business, if you can believe it. And I think a big portion of it was the culture was really strong. And I think this is something I emphasize for folks who are interested in the company. We really do walk the walk of the mission and the culture and the values. And the energy was very much like a startup. And I think you're right. We were very much trying to find our identity in the early years. I think there's one piece around the technology, but how does that technology bring value to users, bring value to society, and what can it possibly be? And I think the early years were us

4:42

exploring that in different ways. We did start with Claude.ai, another chatbot, chat assistant, and evolving into things like tool use. I think one of the moments where we really started to get into our groove was shipping things like Golden Gate Claude. I don't know if you remember that. No. So this was actually up for about 24 hours or so. We had just published one of our early interpretability research in early 2024. And one of the examples was essentially you could have what's called features of the model within the layers, which express certain types of thematics. So one of the themes that the researchers were able to identify was, let's say, bullet point

5:33

writing. Another one was people in places. And one that really came up frequently, that resonated, was the Golden Gate Bridge. And so when you actually essentially dialed up that feature, Claude would obsess about the Golden Gate Bridge. So in every one of its responses, it would come back and talk about the Golden Gate Bridge. So if you said, give me a recipe for making spaghetti, it would say, here is a recipe. And the orange color is just like international red that the Golden Gate Bridge looked like. And so it was really quirky. And we very much wanted to, in that situation, bring that user,

6:16

bring it to the masses and bring it to people who were starting to use Claude. And so the entire experience, we spun up on our Claude.ai website within 24 hours. And that took engineering, product design, our research teams all working together. And we were really, really proud of it. I think it maybe reached only 2,000 people, to be honest. But it made us feel like, oh, we can actually bring new user experiences, showcase our research in a way that's different and authentic to us, and in a very startup-y pace. That to me was one of those hidden inflection points of we were

7:00

starting to find our identity, that we could build products, build experiences that were different from what our competitors had seen, what was already out there. And I think that obviously labs, Claude Code, et cetera, we then started to identify ourselves as how we actually think about the world, how to think about AI, how to bring that closer to the public. But it was a very bottoms-up culture. And so that entire experience was very bottoms-up. I see engineers, I see designers donating time to work on it. And so I always use that as an example of what the early days were like. But the

7:40

culture and the values have very much, I think, stayed the same since those early days. This episode is brought to you by our season's presenting sponsor, WorkOS. What do OpenAI, Anthropic, Cursor, Vercel, Replit, Sierra, Clay, and hundreds of other winning companies all have in common? They are all powered by WorkOS. If you're building a product for the enterprise, you've felt the pain of integrating single sign-on, SCIM, RBAC, audit logs, and other features required by large companies. WorkOS turns those deal blockers into drop-in APIs with a modern developer platform built specifically

8:16

for B2B SaaS. Literally every startup that I'm an investor in that starts to expand upmarket ends up working with WorkOS. And that's because they are the best. Whether you are a seed-stage startup trying to land your first enterprise customer or a unicorn expanding globally, WorkOS is the fastest path to becoming enterprise-ready and unblocking growth. It's essentially Stripe for enterprise features. Visit WorkOS.com to get started, or just hit up their Slack, where they have actual engineers waiting to answer your questions. WorkOS allows you to build faster with delightful APIs, comprehensive docs,

8:50

and a smooth developer experience. Go to WorkOS.com to make your app enterprise-ready today. What are some of the other big inflection moments as you think about Anthropic going from just this lab that's trying to compete with this juggernaut of OpenAI at that point to what it is today? What are some moments that stick out of, wow, that really changed things? Definitely when we were training and testing Opus 3. I think that was the moment when the company, I think we were less than 200 people still at that point. And it was very clear that we needed and wanted to create a frontier model. And that was very important in terms of our ability to reach users,

9:36

consumers, and to showcase our research. And we were looking for ways for also why should somebody choose to answer your questions. WorkOS allows you to build faster with delightful APIs, comprehensive docs, and a smooth developer experience. Go to WorkOS.com to make your app enterprise ready today. What are some of the other big inflection moments as you think about Anthropoc going from just this lab that's trying to compete with this juggernaut of OpenAI at that point to what it is today? What are some moments that stick out, like, wow, that really changed things? Definitely when we were training and testing Opus 3. I think that was the moment when the company,

10:32

I think we were less than 200 people still at that point. And it was very clear that we needed and wanted to create a frontier model. And that was very important in terms of our ability to reach users, consumers, and to showcase our research. And we were looking for ways for also why should somebody choose the team's clawed? And that was a core question. And that was a core question we were getting asked in the early days. And I think with Opus 3, it launched, I think, early March 2024. But there were many, many months of various teams across inference, across research, fine tuning, pre-training, that rallied

11:12

at different points toward a common goal. And I think everybody that was involved was really proud. I remember being the PM, us, the research leads, myself, we were all in our, this was around December, so we were all at home in our various parents' homes and seeing everybody's background of their childhood room. And everybody was working really hard to figure out what are we training the model for? Is it showing up the right way? So I think that was really powerful in terms of just building a lot of trust. And a lot of our research leads have actually, from that time, are now leading reinforcement learning, leading

11:58

our character work, alignment work. So that foundational trust, I think, also helped us work well now with any of our production models across product and research, because we were working so much in the trenches together in the early days. And then I think there were things like identifying that coding was important. Right? In 2023, when I started, nobody said anthropic and clod and coding in the same sentence. I think competitor models like GPT-4 at the time were used a bit for coding, but it was one of many use cases. And one thing that, for example, I saw was people were starting to use these models not just for

12:34

code, not just code autocomplete, but actually writing long-form code. And it's an opportunity for us to train Opus 3 to be better at. And it ended up being a relatively smaller change from a training perspective, but it ended up helping us differentiate in the early days competitively for users and actually bring a lot of the very early cloud enthusiasts and developers because we were providing a value that they didn't really think was possible at the time. It's so interesting. You talk about Opus 3 like that's so long ago, and it's hard to think that was a big

13:21

inflection. And so this is really interesting to hear that that was internally a big milestone. It almost feels like this confidence y'all built that, wow, we could really ship a frontier model, which is now today. So not great if you've compared to what we've got today. What I always think about is Opus 4.5, which was interestingly, a year later, also during winter break, when everyone was home, able to code. Was that another big milestone? Yeah. Opus 4.5 was definitely another large moment. I think what was magical about Opus 4.5 is we also now not just had a model, but a vehicle, which is a great product experience like Cloud Code.

14:00

One thing we say a lot on the team is you need frontier products in order to have frontier models and for people to feel the magic of frontier models. And I think we felt the magic of Cloud Code for many months before that. But the fact that the model essentially got to a level of intelligence where the model is, at a very broad level, users can experience both frontier intelligence in new use cases, allow it to run things end to end in an agentic manner. I think that was the inflection. It was actually both. I think Opus 4.5 wouldn't have had that moment without a product like Cloud Code. And Cloud Code, I think, wouldn't have had that type of adoption

14:34

accelerated without Opus 4.5. So speaking on this thread, Dario, interestingly, if you look back at all his predictions, he's just like, okay, coding is going to be solved. There's 100% in a year or something like that. He kept talking about how AI is going to do all our code. And I remember everyone being like, there's no way, this is way too complicated. How is AI ever going to get really good at this very complex thing that humans do? No, this is going to be humans for a long time. He was completely right. Something else that he talks a lot about is this exponential that we're now on. That's the way he describes it. Now we're on the exponential

15:21

curve. I remember not long ago, new models were being released and everybody was like, okay, we're done. There's no more upside. It's plateauing. It's over. There's no more room to grow. And now it's the opposite. Now we're inside. If you think about the curve of the exponential, we're inside of the exponential now, which by definition means every improvement is a massive jump because we're on that hockey stick part. What's it like just being on the inside of this crazy historic moment when AI is improving so fast, so much is being unlocked? What is it like, and how should people prepare for the coming acceleration of more and more improvement from AI? One thing I like

16:18

to say on the team is most of us weren't actively working yet when the internet transitioned from this novelty to something that everyone can use. And it feels like that's just taking humans. I think analogies are helpful. And so the analogy of that is, I think, a couple of things. I think we have evals. We have, on the safety side, safety testing, red teaming, on the capabilities and product side, new prototypes, products like cloud code, tag, and others. But it's very hard to predict the exact moment or the exact model. And so the adaptability of when you're faced with new information, how do you then make better decisions versus keeping the same plan? And so

17:09

that agility is really important. I think another piece is, with that, how do you actually be thinking very first-principles and reason through what's next? What's the so what? How do we invest in new products? How do we invest in explaining the differences to users? So a lot of the experiences I think of being in that exponential is that pace, understanding how you operate and make better decisions, and then applying that first-principles thinking to then do something that maybe we pull up a plan that we were expecting a few months from now, but now the model can actually do and work on and actually bring that to users. So this is things like co-work, skills,

17:46

tag. It's a very positive self-enforcing loop. And I think a big part of it also is just having the trust in each other, making sure we're thinking through the right decision-making, we're bringing folks along. Some teams might see the exponential, feel it faster than others. So how do we have the grace to bring the organization, the growing organization and company, along on that? Yeah. So what I'm hearing here is you almost don't know what will be possible with every model release. And so the important things to focus on are being adaptable as things emerge. To your point, the product itself has to catch up to what is possible. To your point again,

18:32

it can do so much, but people may not understand how to do it and may not be able to do it. So the product making it easy, and even just telling you here's something you could do, feels like an important part. Is that roughly what you're describing? I think so. I think there's some really interesting graphs in the original scaling law papers. And I think folks are very familiar with the scaling laws in the lens of, as you add in more compute and data, what's called loss, AKA the loss from next token prediction, goes down. And so it's a very smooth linear curve of the models get more intelligent as you scale them up. What's actually

19:21

also interesting in that paper is there are these very different emerging capability graphs. And so, for example, as you add in more data and you train the models with more compute, you essentially see these actually discontinuous emerging capabilities jump. So the models go from one plus one being a thing that it can't calculate to a thing that it can reliably calculate. And so these emerging capabilities, this nature of predictability is not necessarily everyone important part. Is that roughly what you're describing? I think so. I think there's some really interesting graphs in the original scaling law papers.

20:03

And I think folks are very familiar with the scaling laws in the lens of, as you add in more compute and data, what's called loss, AKA the loss from next token prediction, goes down. And so it's a very smooth linear curve of the models get more intelligent as you scale them up. What's actually also interesting in that paper is there are these very different emerging capability graphs. And so, for example, as you add in more data and you train the models with more compute, you essentially see these actually discontinuous emerging capabilities jump. So the models go from one plus one being a thing that it can't calculate to a thing that it can reliably calculate. And so these emerging capabilities, some nature of predictability is not necessarily everyone knows the exact moment. You need the evals to be able to assess that. That has actually always been a part of how this technology works, and also what makes things like safety harder because unless you have the evals, unless you have the systems to test, these jumps might actually happen and you don't know.

20:08

Hmm. That's so interesting that you may have developed this AI brain that can do something you're not even aware of. And so part of the job is just uncovering, wow, it just got really good at this thing. What can we do with that? I think there's product overhang and user overhang, to maybe put it in our PM language, even on today's models. And I think there's a lot that we could be exploring on our current opuses and definitely with Fable, for example. And that discovery is actually another part of what's been in the early days of Anthropic's DNA and I think is also continuing to be a big part of how we operate in product and labs and across research.

20:20

This makes me think about something Gary Tan has been talking about, president of YC. I don't know what his title is. He's had this interesting point that if you're willing to spend $100,000 a year right now in tokens, you are living the way somebody in 2028 is going to live. Because by then it'll be really cheap. Everyone can work this way. But if there's this alpha opportunity right now to just live in the future, go crazy on token spend. And so there's a big opportunity for people to learn what the future is like and also just build much faster. Thoughts on this idea and the value of token maxing, let's call it.

20:30

Yeah, I think I take more of an almost product lens. It's almost like token spend is more the input, and really the output is what you described of experimentation. And I think if we were orienting goals around experimentation, I feel like that might be the better framing of the outcomes. And therefore there might be different ways of achieving that outcome. I will say internally, I think it's really some of the most creative thinkers, the best prototypers, do spend a lot of time with Claude with every new version of a research model that we have. And so there is something around, you have to be using the models to then come up with good and great, better ideas. And there's no substitute for that. It's very hard to come up with a perfect strategy without touching the technology when it's moving this quickly. At the same time, I think there's other things that we could be doing. One thing that we do a lot is actually working in public internally within Anthropic. And so in the early days when we had less product surfaces, there was a Slack channel where everyone, almost the entire company, was testing early versions of Claude and trying different use cases. People were not calling them use cases, but you might be asking it to edit an essay or to come up with the right way to send this email. They were all different use cases, but we all worked in public. And then what you would see magically is different users or different folks on the team coming up with an idea and then other people trying different variations of that idea. And then within maybe 10 or so requests, there was something magical or potentially a use case that emerges. And I think there's a lot in not just individuals figuring out by themselves how to use this technology. I think we could be doing more to actually bring that communal discovery when we do experimentation. Experimentation is not always necessarily an individual sport.

20:37

It's so interesting. Yeah. This idea that we're just not sure what this is capable of or what we could do with it. And it takes all this poking around and people trying things, hearing what other people are trying, to figure out what's possible. Such an interesting, I don't know, technology. Or just, okay, here's what, oh, I figured out I could do this thing. What are you going to do with that?

20:50

I think at a broad theme, we know. We know that the models write great essays or can write long-form writing, but individual pain points of what can you actually solve with that and bring it to a user level that people can use, I think is something that is more exploration- or experimentation-based.

20:57

So following this thread, you oversee product for the labs team, which is extremely cool. We've had Ben Mann on the podcast, Mike Krieger, who both work on labs now. Talk about labs. What is labs? What's come out of labs? Many people have heard of these things. And how do they work that enables them to create such innovative ideas outside of even the core Anthropic product team? The thesis of labs in many ways is identifying and pulling the thread on discontinuous large bets that might not be in the core roadmap and figuring out, is there a there there? And also what is the 10x, a hundredths, a thousandx of the data there there? And so, for example, things like cloud code, I think.

21:01

I've heard of it.

21:08

Things like cloud code, things like skills, and most recently, cloud design, MCP. The thing that we really try to emphasize within the teams is, especially right now, there are so many things that could be built. What does it mean then to have a discontinuous bet? And I think one approach we're taking this year is you can have a very strongly held opinion about the theme or the area, and then more weakly held about the exact prototype. And so there is a culture of experimentation. There's a lot of the bottoms up. Engineers on the team are very self-enabled, self-driven, to test out different ideas. And sometimes we have a thesis and it might not work yet. And so we then might revisit it in one to two model generations. And so this idea of these prototypes that actually end up just helping us learn, that's also valuable, even if it doesn't lead to something immediately shipping. And so I think that allows the incubation and the charter of labs to really accelerate and see around corners more broadly for Anthropic.

21:15

It's so funny to think about a labs within an Anthropic, which is already so innovative and creative and just shipping like crazy, that there's value to still creating a labs team within Anthropic. What enables labs to work as well as it has? Because you listed all these products and it's like, what else has Anthropic shipped? It feels like all the biggest wins almost. I'm sure there are many that I'm not thinking about right now. What's core to creating a successful labs org within a larger company?

21:25

I think that team culture, similar to broadly at Anthropic, is very valuable. I think Ben sets an incredible vision and pushes people to think about the 10x, 100x of the idea. And the teams, the pods within labs, are small. Sometimes these ideas start with one engineer, right? And I think sometimes when there's almost really large teams pursuing very ambiguous, large ideas, you end up actually being slowed down because of that. So I think it's culture. I think we actually also select for folks who actually want to do that zero-to-one experimentation. And it's not easy. There's a lot of bets that we end up turning down or turning off. And maybe we revisit them in the future. But that's hard. That's hard when you pour your heart and soul, you're acting as a founder for a bet, and it's not working yet. So I think it's selecting for that type of personality, folks who are really passionate and deep about the zero to one.

21:32

So you lead product for the research team. You work with the researchers at Anthropic. A lot of people kind of get a nice sense of what is research, what researchers do. I think a lot of people don't totally understand these very valuable people at all the AI labs. The way I think about it, and I want to help people understand, help me understand just what are the researchers doing all day. What I imagine is they have a hypothesis for how to improve the model. They find data, they tweak some algorithms, they adjust how it's trained, and they test it, see how it did, keep iterating, and keep trying to find ways to

21:39

acting as a founder for a bet, and it's not working yet. So I think it's that type of selecting for that type of personality, folks who are really passionate and deep about the zero to one.

21:46

So you lead product for the research team. You work with the researchers at Anthropic. A lot of people get a nice sense of what research is, what researchers do. I think a lot of people don't totally understand these very valuable people at all the AI labs. The way I think about it, and I want to help people understand, help me understand, just what are the researchers doing all day? What I imagine is they have a hypothesis for how to improve the model. They find data, they tweak some algorithms, they adjust how it's trained, and they test it, see how it did, keep iterating, and keep trying to find ways to improve the model. Is that roughly right? Slash, help us understand what researchers are doing all day.

21:54

That's really, I think that's a lot of maybe the more day-to-day. I think one piece around researchers and research organizations like Anthropic is there's also a vision of the future more broadly. So, for example, things like, I think even at the founding of the company, researchers were talking about how do we get Claude to use a computer? How do we get AI to navigate a screen? Right? So there's a lot of actually very founder-like energy, is how I describe it, within researchers, really bold and ambitious researchers. And we have a ton of those at Anthropic. So there's one layer of vision of where this technology can go. And then I think on this other side of the loop, there's also, now that this technology or Claude is in people's hands, how do we make it better today? So it's medium and long term, and a lot of energy thinking about that lens of the future, and also in the immediate and short term, what are the improvement areas we can make?

22:00

And so I think you're describing a really good sense of how do we make iterative improvements on different versions of Claude. The way that my team works with researchers is being very integrated and embedded in those loops, particularly areas where there's a lot of impact on users. So this is things like vision, computer use, coding, agentic coding, tool use, test-time compute, things where there's a direct user impact. And then figuring out what are the ways to bring the user feedback and ground it in a level that is understandable for researchers and also actionable for researchers. And I think that's the second piece, is actually a big part of the job and sometimes a hard part of the job. So, for example, we might get feedback on Claude.ai: Claude hallucinated. It's very vague. If you bring that to a researcher and you say, please fix Claude from being hallucinated, it's not very actionable. And so part of the time of the team is understanding, okay, what's the trajectory of why that user gave that feedback? And it's consented. And so we look at, okay, should Claude have called tools in that moment? Or from its current knowledge or tool use, it called the right, it looked at the right document, but it looked at the wrong facts. In the first case, that would have been a failure on tool use. In the second case, it would have been a failure on, let's say, search or knowledge and search and search synthesis. Or it could be something around alignment. And so bringing that level of detail to researchers, coming up with, is this a big enough problem, figuring out things like evals to then describe how we've improved it, those are the levels of actionability. And it's the day-to-day language of the researchers. And so we try to stay very close to how to bring that in an actionable manner between users to the core model training and the research development loop.

22:07

I was talking to someone the other day about how it feels like research, AI research, is the place to be now if you want to be very successful in life. What does it take to become a really successful researcher, from what you can tell? Not everyone can get in, not everyone's brain is going to work this way, but just say people are like, hey, I want to explore this career path. From what you've seen, what does it take to make it there? Lauren Ruffin Researchers generally, or research and product managers working with research, or both.

22:27

Let's do both. But the researchers, PMs working with researchers, are also going to be very successful, but it feels like everyone's trying to poach all the top researchers across every company. So I know you're not an AI researcher, but just from what you've seen, what does it take to make it in that career path? Lauren Ruffin

22:43

Yeah. I think a lot of the most successful researchers and research leadership at Anthropic are folks who are really strong first-principles thinkers about problems. They reason through problems really well, who are just passionate about their research area and have a bold description of what that could look like, and then who are actually close to the details. And so our leadership, our chief scientists, our heads of fine-tuning and RL folks, are actually really close to the training runs and actually look at things like how the training run is going, evals, looking at the underlying data. So actually staying really close and being excited to be in the details, I think, have been a sign of really strong researchers and developing taste. And I think another piece is just their ability to think big over time and be very ambitious, right? Like the Dario, like we can transform software engineering. And I think going in that direction, you learn so much. You have to shoot for the stars in many ways across your ideas, I think, in order to be a successful researcher.

22:50

Jason Wong I love this meme of just be more ambitious. It comes up so often now, which is so hard. It's easy to say that. It's hard to actually, just like, how big can you think? And that's so much of what AI now unlocks, just be more ambitious.

22:59

Yeah. Yeah. I think it's thinking through it once or twice end to end, and then being, I think, stubborn about the area and maybe more loose around the exact approach. It is a question we challenge ourselves with, but the technology is moving so quickly. And so how do you make sure what you're building is actually forward-compatible? And so it's also actually part of the core product development loop to think bigger, right? One thing I ask the team frequently, or how I think about when we're building a product, is let's say Claude 8 comes around, what changes in what users do? And then what should, what does that mean for how you're building today? Is it going to be forward-compatible to that experience? Right? So just grounding it, I think being ambitious is very broad. And so trying to ground it in some ways of describing that. And also, yeah, everything heading in a direction that all is cohesive and makes sense versus just ambitious in a completely different direction.

23:04

Speaking of ambition and Claude 8, Fable slash Mythos recently feels like it hit this very new kind of tipping point with models where it used to be, you have an awesome model, release it. Hey, everyone, welcome, Opus 4-5 is out. Everyone can use it. Mythos went in a very different direction. It got blocked. There was a lot of scrutiny, a lot of concern about what it was capable of. All the companies had to go make sure it wasn't going to hack into all their systems. And it feels like now every model, because they continue to get better, will now have a lot more scrutiny, and there will be more restrictions on who can use them, which feels like a big deal. How do you think about that? How does that change the way you operate?

23:09

I'm going to maybe leave the policy and the expert control side to focus on that and work on that. I think the product question and how we interact with these internally is, I think, as you mentioned, as frontier models become more capable, the safeguards and the ways of red teaming and testing and the pre-release process also need to evolve and adapt quickly to address that. And so one example is, before Fable models, we didn't have strong fallback UXs and systems because our goal is to make sure that there is asymmetrical benefit for this technology and to minimize the downside or a severe risk of it. And so we ended up building fallback systems so that users will still get a great response from Opus 4.8 immediately. And so I think there's a piece around, as we evolve and improve safety systems, how do we continue to develop and deliver great user experiences?

23:13

I think the product question and how we interact with these internally is, as you mentioned, as frontier models become more capable, the safeguards and the ways of red teaming and testing and the pre-release process also need to evolve and adapt quickly to address that. And so one example is, before Fable models, we didn't have strong fallback UXs and systems because our goal is to make sure that there is asymmetrical benefit for this technology and to minimize the downside or a severe risk of it. And so we ended up building fallback systems so that users will still get a great response from Opus 4.8 immediately. And so I think there's a piece around, as we evolve and improve safety systems, how do we continue to develop and deliver great user experiences? I think there's more that we can do on both sides. And so you'll see us innovating, improving on what we call now the model safeguards package more and more in the coming weeks and months.

23:18

What's really interesting and unexpected here is, it creates this really interesting advantage for Anthropic where you have access to the latest stuff. And this is going to happen at every lab. Everyone's going to keep improving, and it creates this unfair advantage within the labs to have access to the best stuff that other people can't get outside of your control. You'd prefer everyone use it. So it's this really interesting new feedback loop that's going to start where models that are so advanced are only accessible to certain companies. And that's going to be a whole new, unexpected second-order effect of all these restrictions.

23:24

Our goal is to develop these systems and the models to be as inclusive as possible. I think our goal is to not have that happen for the general-purpose, general-use technologies and to make it more accessible. I think this is one of our top priorities right now to reduce what we're seeing there. Yeah, that makes sense. I would imagine you'd want as many customers and people using this thing as possible.

23:40

This episode is brought to you by Mercury, radically different banking loved by over 300,000 entrepreneurs. And now with command, I've been a customer of Mercury's for over six years. I have never once thought about leaving. Mercury is what happens when banking is built by product people, not by bankers. They make it so easy, dare I say fun, to send invoices, move money around, and set up virtual cards for folks on my team. Does your bank have an API, a terminal-native CLI, or an AI-ready MCP server? I don't think so. And just recently they launched command, a conversational interface built directly into Mercury, which acts as your financial operator. I've been using command to transfer money around, to figure out what categories I've been spending the most money in, and analyze my cash flows. And just today I used it to find out how much I've made from a specific sponsor over the past year. I just asked, how much have I made from X over the past year? Ten seconds later, I had an answer. It is so freaking cool. Visit mercury.com to learn more and apply online in minutes. Mercury is a fintech company, not an FDIC-insured bank. Banking services provided through Choice Financial Group and Column NA, members FDIC.

23:46

I want to talk a little bit about how the product role is changing and who is doing well in this new world, now that AI is such a core part of our life. When you're hiring PMs, product people, when you're looking at people that do well in today's world, what are some things that you notice? What are you looking for most? What's trending up in what you find is important, and what's trending down?

23:54

We actually, on my team, have not changed our hiring loop for three years now. So what we actually look for and the traits and how we evaluate generalists like PINs, generalists like research product managers, have actually been the same. So I think some of those traits, number one, is first-principles thinking. And this is really rather than pattern matching what you used to do in, let's say, consumer product or B2B SaaS, but actually figuring out in this moment, for this user group, with this technology, what is the user value?

24:01

Is there an example? A lot of people hear first-principles thinking. They're like, yes, I got it. I'm good at this. What's an example of someone having really demonstrated really good first-principles thinking?

24:09

I think one example is, I think you think of a product manager as, I own product strategy and delivering user value, but I demonstrate that day to day by writing a PRD or writing a product vision doc. And for my team, as research product managers, the way to drive user value is to figure out the right user feedback, the evals, that then can be a personification of that user need. So we do write some product documents and PRDs, but we actually have a saying on the team of evals are the new PRDs, because in order to deliver that user value, it's not that exact artifact that people used to write in the last one to two decades. It's a new way of working. And so the first-principles thinking would be, let me figure out what is the thing I should do to achieve my goals, rather than here is a set of activities that I've done, and therefore I will continue to do.

24:18

So the idea here is, it used to be: have an idea, create a PRD, talk to people about it, align on the plan, design it, build it, ship it, see how it goes, iterate. What I'm hearing here is, okay, here's some feedback about something that's wrong or an opportunity. Step one is the eval is now how you define what the work is versus a PRD.

24:25

Maybe step one would be understanding the user pain point. And so the way to even access that user pain point is different. In the past, we might do a user interview. I think if you go deep enough, you might have the user walk you through their user flow, the pixels. Here, you have to sweat the tokens as much as you sweat the pixels. And so one activity we have on the team is reading the transcripts and understanding what were the trajectories that failed very deeply to then say, was this a hallucination? Was this Claude being overconfident? So the theme of the failure actually has a lot of nuance. And then that allows you to build a description, a sustained description of that pain point. So that could be essentially in a new eval. And is an eval on distribution? Is it capturing both the positive situations where this is failing and also areas when it should actually not fail, and then bring that back to, let's say, research so that we can make the improvements and actually measure the quality of, okay, when we have Opus 5.5, is this area improving or not? Is Claude now able to identify the right places in the document and pull the right synthesis out? So it's just the actionability and shortening the distance to actionability for our stakeholders and partner teams like researchers to take action on.

24:28

Is there an example of something like this where you found an issue or opportunity and then wrote the eval? And what is the eval looking like in most cases? When people want to picture an eval, what is that? What do they picture?

24:39

We actually pioneered this concept within Anthropic. So one of the early examples is the early Claude models were not very good at following specific schemas, so things like outputs in JSON. And now that is fundamental to Claude being able to be a good agent. If you can't output a certain format, you don't know how to access APIs, you can't call tools, et cetera. And so the initial end-to-end was, I was hearing feedback around Claude 2 days. Claude was not very good at following instructions. So then digging in with users, what do you mean by Claude is not good at following instructions? Give me situations this was happening. What's the exact paragraph? What did you ask? What was Claude's response? Going to that level of detail. And what I saw was something like 80% of what people meant in the early days for this failure was Claude would not write the right JSON. And so then, okay, let's generate maybe, to start, just 30 to 40 examples of when Claude was not doing this thing correctly. And then that actually is your eval set. And you can have essentially a prompt and a response. And if that is not working in the right golden answer that you might have, then that means that the eval essentially is beneficial because it's identifying a pain point consistently. And so then we added that to our repositories for evals. And when we have versions of Claude, we actually run that eval and just check. I think at this point, it's always 100% or 99.9. And so it's no longer a pain point. But in the early days, it was taking the user feedback, figuring out

24:46

something like 80% of what people meant in the early days for this failure was Claude would not write the right JSON. And so then, okay, let's generate maybe, to start, just 30 to 40 examples of when Claude was not doing this thing correctly. And then that actually is your eval set. And you can have essentially a prompt and a response. And if that is not working in the right golden answer that you might have, then that means that the eval essentially is beneficial because it's identifying a pain point consistently. And so then we added that to our repositories for evals. And when we have versions of Claude, we actually run that eval and just check. I think at this point, it's always 100% or 99.9. And so it's no longer a pain point. But in the early days, it was taking the user feedback, figuring out actually what they mean. Can we reproduce it? Is it consistent? Is it a big issue? And then figuring out how to standardize it in a way that can be consumable for researchers.

24:55

It's test-driven development for PMs, is the world we're living in now, where you write the test first. So is this just a core part of the product management job now at Anthropic, writing evals?

25:02

I think so. I also think it's something I've talked to other PMs at other companies about. And I think it's also more and more of the skill set more broadly. Because a lot of the products that we're building are at the intersection of models with harnesses, with a set of contacts for a set of users. And so having things like evals actually is a way not just for folks working on models, but generally within product to get to better user experiences. Because you can't improve what you can't measure. And a lot of this is still very tactile-based. It's still very judgment-based. And so you have to stay close to the details.

25:10

And also very non-deterministic, which is a big part of this. It's not going to give you the same answer every time. So you got to describe it more broadly. It's not going to be an exact match. So this is a really interesting change in the way product happens and will happen, is evals. Writing evals versus PRDs is a big part of this. Do you guys still do PRDs? Is there still a one-pager describing a problem or is it replay? Okay. Now you're shaking your head yes.

25:19

We are. We do. I think when there's a very defined problem, I think things like evals might be almost a shorthand. I think there are other cases where PRDs are really valuable. PRDs are great vehicles for getting a very large group of people aligned on a set of sources of truth about experience and set of goals. So when we do have a model, we actually, for every model, we do have a PRD. Less necessarily for our researchers, but more for our growing product surfaces, for our engineering teams, for our stakeholders like legal and safety and others, as just a source of truth of putting together what we're aiming to achieve so that a big group of people can row in the same direction. The other place where I do think PRDs are valuable is on the more ambiguous problems and opportunities. Right? So if we haven't shipped a thing like computer use, we don't necessarily have a set of user-specific pain points always. And I think there's value in the product vision portions of a PRD to explore what could, even if a technology is not yet ready to work for everyone, how do you get it to work well for some group so you can explore the value, you can actually bring something that is coherent to a user group. So we do have PRDs. I think the application is a little different now.

25:26

Okay, this is great. I just had Andrew, he's the head of the Codex app at OpenAI, and you guys are aligned. PRD is not dead. Still very useful for specific projects and ideas. Great. Okay. We've closed the book on PRD is still kicking. Okay. So we've been talking a bit about what kind of skills are emerging for product people. Is there anything else that you find has shifted in what patterns are common across people that are doing well in this new AI world, in terms of product managers and folks on the product teams? Is there anything else that you're like, okay, there's something you've got to shift or something you look for more in people?

25:38

I think maybe specifically for folks who might be mid-career or folks who have been more in a managerial product leadership seat, one thing that I think I feel pretty strongly about is, in order to be good managers of teams and PMs working with this technology, you have to be really hands-on yourself and have spent not just time tinkering, but actually shipping the technology and, again, being in the details and sweating the tokens along with your PMs and your engineers and your teams. And so even for folks that I hire who have more tenured PM experience, the onboarding plans are exactly the same as somebody who is more early career. And it's around understanding users, reading consented user feedback, talking to customers. I think there's something around being able to understand what to do with this, what good looks like, and having developed that in a very hands-on manner, that's important. It's not necessarily easy for someone to agree or be able to see what a good or great AI product or AI feature could look like if they haven't experienced building it themselves. So I think there is, I do feel pretty strongly that if you're a manager, you have to be hands-on, you have to spend a portion of your time actually shipping, you have to walk in the shoes of your teams. And I always try to carve out a portion of time to actually own one to two work streams when we have models in order to keep my theory of mind, keep my sense of how the models are moving, how quickly it's improving, so I can help the team make decisions and make better decisions.

25:44

So what I'm hearing here is, no matter where you are in the ladder of hierarchy at a company, if you're not building yourself, if you're not actually talking to Claude, talking to Codex, building stuff, you're not going to make it. And you should have fun working with this technology. I think that's the other piece. I think the folks that would be most successful, regardless of their level, are people who love working with AI and are exploring and experimenting and carving out the time not just for the experimentation, but actually hands-on shipping end to end, getting the user feedback, I think has to be fundamental for everyone.

25:55

I 100% know what you mean there. Me sitting on my newsletter and this podcast, just talking about stuff and, yeah, yeah, that sounds great. Every time I actually build something, and I tinker with all kinds of little projects, you just, okay, I see what's happening here. And you just get so much more. It's hard to exactly describe what you experience actually working with the models and building stuff, but it's a whole different world of, okay, I see. Here's what they're talking about with computer use. Here's what they're talking about with this limitation of this UX situation. Yeah, yeah.

26:01

So it's just like, and you made this really interesting point that you have to have fun with it, which is not easy for a lot of people because they're pushed to use AI or they just don't know exactly what to do with it. For people that are just like, I don't know, it's just so annoying. I just have to do this. I don't know. I hate this frigging thing. Why do I have to work with this? Things are changing so much. I'm tired. Advice for helping people find that joy in this work.

26:06

I think maybe I'll reemphasize something I said earlier around just that experimentation is not an individual sport. Some of the moments where I think I've touched practically every version of research models across 20-plus versions of production Claude at this point, and I think part of the joy comes from seeing other people discover use cases too. And so maybe one idea here would be pairing with somebody who is excited and seeing what, on a use case that you care about, and working together versus identifying or trying to figure out the perfect use case yourself. Because that might feel like work. Working with others feels like joy a lot of the time. And is there more that we could do to bring that, bring other people along? That's something, like a lot of times internally we have somebody who is very curious, and them sharing an idea of a new prototype actually

26:11

find that joy in this work. I think maybe I'll reemphasize something I said earlier around just that experimentation is not an individual sport. Some of the moments where I think I've touched practically every version of research models across 20-plus versions of production Claude at this point. And I think part of the joy comes from seeing other people discover use cases too. And so maybe one idea here would be pairing with somebody who is excited and seeing what on a use case that you care about and working together versus identifying or trying to figure out the perfect use case yourself. Because that might feel like work. Working with others feels like joy a lot of the time. And is there more that we could do to bring that, bring other people along? That's something a lot of times internally we have somebody who is very curious, and them sharing an idea of a new prototype actually brings a ton more people who are like, oh, I didn't know this could work now with Claude. And so there's some virtuous cycles here and ways of continuing to have joy with this technology.

26:12

That's such a good point. I think that's also why Twitter is so useful for a lot of this. You see other people sharing what they've done, and it inspires you to come up with your own little ideas. And also it's just fun to share your own thing that you've done. So that's a really good point. Just find other people to play around with and look for use cases. The thing I've also heard a lot is just find a problem you want to solve in your life or work and just open up Claude. Claude, tell it here's what I want to do. And it's incredible how far you can get just with a vague idea of a problem you want to solve.

26:20

Yeah. I think it gets hard in that there's so many different things that you could try. Yeah. And so just narrowing in on either pairing with someone, working with somebody who has a lot of joy about this technology, or figuring out something that you could immediately find value. Either of those things allow you to go deeper rather than more high-level about too many things. I find it hard to keep pace with the number of prototypes or products that are out there. And so my lens has been, how do I go deep in one to two of them myself?

26:28

That's so interesting you say that, because that's exactly it. We just had the survey that I ran with my colleague, Noam, asking my readers just how they're feeling about all the things going on in tech right now and AI. And one of the most interesting takeaways we had was to find that happiness is exactly what you said: go deep in a couple things versus trying to just ton of little things, find a couple of things to really solve well, and then go deep. And that is a source because a lot of the happiness people feel is when they finally unlocked a way for AI to actually make their lives better versus just a couple of messed up, broken, half-working things.

26:38

Yeah. It's how do you go from this being a check-the-box, right? And so us as product people, it's then an exercise of product prioritization of your time and your energy. And if the goal is to experiment with joy, then what are the inputs that you need for that? Yeah. But yeah, I think a lot of the secret sauce of Anthropic is the culture and the bottoms-up nature of how people work and this experimenting in public. And by doing that, it's very much about how to bring other people along. That ends up being, I think, really valuable.

26:56

Yeah. I've heard this so many times from all the labs, just like, no one's exactly sure how some of this is going to be used. And a lot of it is just putting stuff out early, seeing how people use it, seeing what's possible, and then using that information to build that product, to lean in. Yeah. Yeah. I'm curious, on this thread of finding ways for AI to help you in your work and life, are there any interesting ways you've been using Claude lately in your work as a PM?

27:13

I think there's a lot of things with Fable and things like Tag. So there, I think Tag is in the very early days. I think there's something around how you work in a different paradigm of allowing an agent to go off and work and then bring back product experiences to you. I think one area that is not more recent, but one that I bring up a lot with the team and I think we could do more on using AI, is just how to use it to also be more, to have better conversations with each other, to be better managers. I don't think it's necessarily just about raising the IQ of experiences we build, but also I use it a lot in actually prepping for how to have better conversations in the moment during crucial conversations. So I love that book. And so I actually have a skill that helps me figure out, am I going in the right level of detail given the situation at hand, and actually helping me be a better manager and better supporter for the team? So for managers on the team, that's actually a thing that I've been sharing more with our managers. So, okay, how do you actually use Claude to make you a better coach? Because it's hard sometimes to find the right perfect words, and the models have a lot of perfect and bright words. And I think there is something about how it can actually augment us from an EQ perspective in addition to IQ.

27:18

Oh man, there's so much interesting stuff there. So just to understand what you're doing there. So you built a skill. You just helped Claude build a skill, pulling in lessons from Crucial Conversations, the book, which it knows enough about. You don't have to even give it the content. And then you use that skill to talk to Claude. Hey, I have this very difficult conversation coming up with a colleague. Give me some tips on how to approach it.

27:25

Yeah. And it's a great, it's almost like coaching, individualized, personalized coaching of just how to make you. And there's so much context switching that we do all day, and having Claude help me pair and help me. And maybe there are times where I end up not using suggestions from Claude, but it actually ends up being very helpful for just coming up and brainstorming. Am I thinking about reactions in the right way? How do I actually go a bit deeper, faster, build trust faster, be more direct?

27:30

Yeah, man. I have so many questions here. This is so interesting. One is just like, there's concern people are going to start talking the way AI writes because they're talking to AI so much. And it's going to be like, Diane, it's not this, but it's that. I know that you're not doing that, but that's a concern people have. Let me just ask about that, I guess. Do you fear this? There's this brain rot, atrophy stuff. People talk about it. We're just so reliant on AI now, and we stop learning and thinking and overall AI thoughts on that, being so close to it and being so integrated with AI constantly.

27:34

A lot of actually thinking process and writing process are tied together for me personally. And so I think there are ways where I use Claude to augment my thinking, but what I want to make sure, and maybe this is what you're describing, is Claude doesn't take over all of my thinking for me. And so I think depending on the situation, depending on how much more personal judgment I want to have in a situation, I might come up with my own POV first and then work with Claude through that, and making sure that I maintain my sense and tone throughout. I think there are then other things like updates, right? We have monthly business reviews. And then in those cases, it's much more, you want it to be standard and I want it to be much more like it gets crisp, crystallized information in the right way. And I have a skill and we're augmenting and improving our skill for that. But I want to get to a place where the monthly business review, the writing of that is potentially asymmetrically less valuable than the thinking. And so how do I get that piece delegated to Claude fully, and I'm more of a reviewer and a verifier of that information? So I think it depends on what you're using Claude for and what you're trying to convey. And is there an asymmetrical value in delegating more

27:38

tone throughout, I think there are then other things, updates, right? We have monthly business reviews. And then in those cases, it's much more, I watch, you want it to be standard, and I want it to be much more like it gets crisp, crystallized information in the right way. And maybe I have a skill, and we're augmenting and improving our skill for that. But I want to get to a place where the monthly business review, the writing of that is potentially asymmetrically less valuable than the thinking. And so how do I get that piece delegated to Claude fully? And I'm more of a reviewer and a verifier of that information. So I think it depends on what you're using Claude for and what you're trying to convey. And is there an asymmetrical value in delegating more to Claude? What I'm also hearing, the first tip is really great, which was think first, have a point of view, and then use Claude as a sparring partner almost to evolve the idea, push back on the idea. Yeah. Yeah. And I think this is where things like actually our alignment research and safety research is helpful because what you don't want is an AI that just agrees with you, right? What you want is this technology to actually augment and grow and get to a better outcome. And so sometimes having Claude push back makes me better. And so that's great. Like a coworker, I want somebody to push back when my ideas are not fully formed.

27:45

I want to hear more about that. I've heard that when Ben Mann was on the podcast, he talked about the constitution that is built into Claude and how, unintuitively, the work and the focus on safety and alignment, as you said, and this constitution that describes how Claude should think and operate, actually, you would think that would limit the abilities of Claude and make it less fun and interesting. It's exactly the opposite. Claude is the most interesting personality. I hear that constantly. It's just like, I much prefer talking to like open claw famously was built on Claude. And then people were forced to switch. We won't get into it. Were forced to switch to Chappie T, and they're like, this is so bad. This is not who I'm used to talking to. So that is, I think, a really interesting point. I just want to make sure we spend a little time on why that is the case. Just this focus on alignment, safety, having this clear constitution. Why does that make Claude better and more interesting to talk to also? In order to make Claude as intelligent and as capable as possible, being able to have Claude actually push back at the right points and then add, it's like a yes or no and actually helps you come to a better conclusion. So I've used Claude to help with things like, are we making the right pricing decision on the next version of Claude? It's a little bit meta, but using a research version of Opus, asking it to figure out how it should price and being able to come out with better outcomes is a goal at the end of the day. And so having AI not just be an assistant, not just be a doer and being delegated tasks by figuring out, is it doing the right thing? That's actually very integrated with knowing when to push back, right? That's part of knowing when you should be proactive. Proactivity is not necessarily always doing a thing that you are scheduled to do. It is knowing when to come up with a new idea. And so in order for Claude to be more useful, the general approach has to be that it knows when to push back. It's a core part of the characteristics together of the models. That is so interesting. It's so interesting that that is what a big part of it, like it being less compliant is almost what makes it better and more useful because we need that. I've had so many people where they're like, hey, AI told me I was right. And no, I wish you could listen to other people. Yeah. And it comes back to our earlier point around thinking, right? How do you protect your thinking? If you have an AI that can be a thinking partner, a thinking partner doesn't just agree with you. It should add to you, and you should come away at the end of the day having better ideas because you've worked with Claude. That should be the hero goal, not just making your ideas 10% better. Yeah. I love this. And it used to be think 10x. It used to be the way founders push people. Like what if we 10x this? And I love what I keep hearing. It's like, how do we go 1000x from this idea? What is the most ambitious version of this? I want to come back to something that I was thinking about as we were talking about talking to Claude constantly. It's very clear when AI has written something still. It's funny that it's a large language model. You would think, of all things, it would be very good at writing. And interestingly, no AI is very good at writing. It's always very clear. This was AI-written. Do you think we'll get to a place where we will not know this was AI? I think it depends on what's the goal that you're looking to achieve by knowing or not. What's the eval?

27:49

Yeah. What's the eval? I actually do think there's more that we could be doing on making Claude write better. There's actually very active efforts on my team and on the research side about making Claude write better, just generally. I think it should be clear where an idea is being led by you or by Yuleni or me, Diane. I think it really depends on what's the goal of that writing. For something like a monthly business review, I would actually love to have that end to end be written by Claude. And obviously, not to make it feel like it was written by a human. It's such an interesting point you're making. Is it actually better for us to know that it's AI versus not?

28:07

Yeah. But it's also for maybe the lens is more around verifiability or who's verifying the output, right? Like who's signing off? Maybe less around who's writing, but who's verifying, who's signing off. That becomes more what matters than who's writing it.

28:17

Why do you think AI is not great at writing? My guess is it has studied all of the best writing in all of humanity. It's figured out here's the best way to write. And now that we, and it's just, there's only so many ways to write. And so we've just recognized, okay, this is what AI does. It has these tropes. Is that the core of it? Is there something else that's keeping it from being a great writer? Ironically, being a large language model of all things, you'd think it'd be really great at language. I think part of it is also we need to invest more in training improvements to make AI continuously strong on areas like writing. I think it's also, the technology's jagged-edged, like we mentioned. So sometimes when the models were good at writing but not agentic, our thesis is how do we make the models more agentic or call the right tools? Now that that's improved a bit, then it's, well, now these other areas actually become more of the rough edges. And so I think we're in one of those moments with writing where we need to actually just focus and prioritize on training the models to be great at this area. And that is a very active area for us. So fun, thank you mentioned.

28:27

Okay. I'm glad. I'm glad. And also, it was going to be interesting once AI is so good. We're like, I don't know who wrote that, but to your point, sometimes we actually want to know that it's AI. That's really interesting. I never thought of it that way. The other interesting part of this is that there's that comedian who was joking that we're on a plane and the Wi-Fi is down and we're just like, what the hell? The Wi-Fi is not working on this plane. This sucks, how dare you?

28:34

When you're in a tube in the sky flying like a bird, and how dare you complain that the Wi-Fi doesn't work? Your point is there's so much advancement and so much power. We can't fix it all. We can't make it all work the best possible. And so AI writing has not been the priority, and it feels like there's more investment happening there. Yeah. I think tone and character is a priority. I think this advancement of the technology is a work in progress. And so we see a leap or emergence of a jump in agentic behaviors. And so that is a new normal. And then these other capabilities need to continue improving. Yeah.

28:47

And I think once we improve, let's say writing and tone and character, we probably will say, how do we have Claude be even more proactive? Proactivity is an opportunity, and that's human nature. When you're in a tube in the sky flying like a bird, how dare you complain that the Wi-Fi doesn't work? Your point is there's so much advancement and so much power. We can't fix it all. We can't make it all work the best possible. And so AI writing has not been the priority, and it feels like there's more investment happening there.

29:02

Yeah. I think tone and character is a priority. I think this advancement of the technology is a work in progress. And so we made, we, we see a leap or emergence of a jump in agentic behaviors. And so that is a new normal. And then these other capabilities need to continue improving. Yeah. And I think once we improve, let's say, writing and tone and character, we probably will say, how do we have Claude be even more proactive? Proactivity is an opportunity, and that's human nature. We want to make ourselves better. We want to make this technology better. So yeah, it, I, and I think we're applying it to AI, which is the right thing. We should be making it better.

29:15

I want to ask you a couple of questions. I like to ask folks working at the very center of the future of that is coming. One is where do you think human brains will continue to be most valuable over the years? I know Anthropic's mission and vision is we'll reach AGI, a superintelligence. So in the future, maybe nowhere, but before we get there, where do you think human brains will continue to be most valuable as we approach that timeline?

29:21

We started to talk about making Claude and models better at judgment, especially in the last year or so. I think judgment is one, and is an area where it's accumulation of so much nuance and so much experience, and these systems haven't experienced as much as humans have. And so I think that hard-earned judgment is an area for product leaders, and just generally, will continue to be really critical.

29:28

There are so many things AIs can build. Which are the things that an org like a lab should build, right? A lot of that requires human judgment, persistence, and proactivity. These are all traits that are beyond just general capabilities, but behaviors and characteristics of people at that level, of how do you get to the best solutions? How do you create the best experiences? So I think those types of traits are actually the tactile traits that I think will continue to be important.

29:35

I think there is also still a lot of capabilities and subject matter expertise as well. I think software engineering has been really transformed by AI. I think there are areas like biology, life sciences. These are all things that we're just kind of at the foot of the exponential on. Maybe software engineering, we're on the exponential. On some of these other areas, we're not quite there yet. And so I think you're seeing us ship things like Claude science, investing in these areas because those are areas that bring this technology to society and have a positive benefit for society. So I think there's a lot more to go there.

29:44

Another question I would ask is, as someone with kids, how do you think about what you are encouraging them to learn? Or how do you think you're going to nudge them to be successful in this wild new world that we're entering?

29:52

I actually think it's a lot of the same traits you and I probably grew up with, which is curiosity for learning, persistence, believing in your own inner voice, developing and then believing in your own inner voice. I have a four-year-old. I have an eight-year-old. It's on us to help. It's on me to help them develop their inner voice, and whether that's being opinionated and taking a stance to me, right, and developing that, encouraging that. I think those types of skill sets are things that are important in the future, and I like having their own individual voice.

30:04

That is so interesting. It's so related to the answer you had when asked about how to avoid brain rot, essentially an overreliance on AI, which is just keep focused on your own point of view and your own perspective before you overrely on AI. And just this idea you're describing of building that in kids is really important. That is so interesting. And I love how all this kind of connects: judgment, persistence, and a point of view of your own. Yeah. Both for kids and also adults. Yeah. Anything we think about for your.

30:19

Oh man. Well, the question I'm thinking about is just when to get them on some AI thing, when I have a three-year-old, so it's pretty hard for that. But how do you onboard them to this crazy thing? I was at an event recently, and a bunch of parents were talking about how they think about AI and their kids. And one person had a really interesting approach, which is keep them on the very early models so that they still have to struggle a bit and not get all the answers immediately. That was interesting. Like an open-source local model, not fable. Yeah. Oh yeah. And curiosity is something I keep mentioning Ben Mann, but his answer actually to this question has always stuck with me, which is curiosity. And also, he's a big fan of Montessori, which is what I'm encouraging for our kids. So there's something there.

30:26

Maybe a last question, just along, it's kind of along these lines, something Fiona Fung actually suggested I ask you, who's recently on the podcast. How do you stay recharged and not burn out being in the center of this crazy storm of AI as a mom working in, we're seeing the research work at Anthropic? We're living through the most unprecedented time working at, just being on the outside of Anthropic, it's crazy. I don't even know what it's like to be on the inside. What have you learned about avoiding burnout, staying recharged, staying sane during the middle of all this?

30:33

In 2024, we shipped four models for the whole year, or four series of models. And I think we did more than that volume in just Q2 of this year.

30:39

I think I've been really lucky with the team that we've grown and built, both the stakeholders on the research side and within our research product management team. I think that one of the magical parts about approaching all of this is that it's not an individual sport. There's a sense of radical ownership and team collaboration that I think sometimes does feel like a high-performance sport because you're in very critical decisions. There's new information about users, about training, and you have to make recommendations and judgments and decisions very quickly.

30:47

And nobody can do that sustainably by themselves. And so I think what's really helped is having a team that is incredible, who looks out for each other, who, the night before a launch, even if they're not the core DRI on that model, will stay up and help the DRI review the blog post and make edits and come up with better demos and know to be each other's extra hand.

30:52

I think it's very easy if you take all of this change on your own shoulders to feel like you're alone and to feel like you have to do everything. But I think one of the magical parts of Anthropic is this ability for us to figure out what are those opportunities to help each other and actually then taking the next mile of mind melding. We called it entering the hive mind. There was an article about this. And I think part of that is that it allows the team to replenish. It's not that you, I was just on PTO in June. It's not just that you can take PTO and you come back to 3x the amount of things to do. It's actually that you can take PTO and know the team can figure out the right things to do and that we individually can watch out for each other.

30:58

So I think that's a big part. I'm really lucky personally. Also, my partner is really supportive. This is year six of me working in AI, so Amazon and then Anthropic. And so he sees how much I just love the technology and what this can do. And that really helps, I think, also from a personal perspective as well. I love how many of these answers connect. So what I'm hearing here is just having other people, working with other people, relying on other people, helping each other out when things get crazy, which is a similar answer you had for just how to, how to find

31:12

come back to 3x the amount of things to do is actually that you can take PTO and know the team can figure out the right things to do and that we individually can watch out for each other. So I think that's a big part. I'm really lucky personally. Also, my partner is really supportive. This is year six of me working in AI, so Amazon and then Anthropic. He sees how much I just love the technology and what this can do, and that really helps. I think also, from a personal perspective as well, I love how many of these answers connect.

31:26

So what I'm hearing here is just having other people, working with other people, relying on other people, helping each other out when things get crazy, which is a similar answer you had for just how to find the joy and fun in this work. Just gain, be inspired by other people and see what they're doing.

31:34

Yeah. Yeah. And it's interesting, when Fiona was on the podcast recently, I was asking her what's changed in the world of software engineering, and she pointed out it's a lot lonelier now because now we're working with agents instead of other humans. Teams are smaller. People are having all these fleets they're talking to constantly. So this is just a reminder of the power of actual other humans around you.

31:36

We're asked to work and make decisions on really big things because you have more scale from the technology, right? And I think having individuals, having other folks who can have some level of mind meld with what you work on, how you approach maybe not exactly every detail, but what are the first principles, what are the assumptions you make, that helps them back up for you or push your decision and sharpen your thinking.

31:42

So I think we really try to, I really try to look for that when building the team, growing the team, hiring. Is this person going to care about their own ego and building out a big org, or are they going to care about contributing to Anthropic and contributing to that impact of the team and orienting towards folks who are low ego, team oriented? I think that's a big part of the sustainability.

31:47

Yeah. A lot of it always just comes back to culture and hiring. And I know I've heard a lot, just the reason Anthropic is able to move so fast. I remember that moment when something shipped every day of the month. It's like a calendar of launches, and people were talking about how is this possible. And what I heard a lot is just because everyone is so aligned around the mission and the values, it allows people to make decisions really quickly. Before we get to our very exciting lightning round, is there anything else, Dan, that you wanted to share, anything else you wanted to touch on, anything you want to maybe double down on of things we've talked about?

31:56

This was actually really fun because I feel like your questions actually sharpen some of my thinking around how the thoughts connect. I'm your real human Claude over here. One thing that I really want to convey or have people take away is, I think, one, in the ways of working, but also just this is a lot of growth and change, and having the joy in using this technology. If you're feeling like, in this moment, you don't have as much of that feeling of initial joy, how do you find people who do, if this is an area that you're excited and want to work on?

31:59

And I think developing skill sets, replenishing skill sets in many ways, of things like thinking from a first-principles manner about what you solve. I think fundamentally, you didn't ask me this, but there's this question in the community of, do we still need PMs when the models are so capable, when engineers are leaning in? I think the role of people who are user-centric, who go into the details of understanding what users are trying to accomplish, bubbling that up in an actionable manner and doing the relentless work to do that, that to me, it's a core of a product person. And I actually think we need more of that.

31:59

I think we are becoming very technology-layer-driven, and actually to make that impactful, you have to go deep, you have to be curious, you have to be super hands-on. And those are things that I think are also traits that have helped Anthropic from a product development and model development perspective and as part of the culture. And hopefully that's valuable for others as well.

32:03

Amazing. What an inspiring way to end it. Oh man. Yeah. And this is, I've been saying this too for a long time, just now that building is easy, the hard part becomes, as you said, what should we build, and is the thing we have built correct and good and worth leaning into? And to me, that's what PMs do and what PMs are good at. Yeah. Yeah. Yeah. And it's getting into the details of the user. Yeah. Empathy. Okay. Great. PMs are going to make it. Okay. PRD is not dead. All kinds of important lessons here. Diane, with that, we've reached our very exciting lightning round. I've got five questions for you. Are you ready? Yep.

32:21

First question. What are two or three books that you find yourself recommending most to other people? One personal one, I really like How to Raise an Adult. So, I'm a mom. I think a lot about what are the things that I want to instill in my kids. And that book is really helpful for describing, we're not trying to raise children, we're trying to raise adults. So just the framing of what does that mean, and what are the characteristics that we want to hone, harness, and foster in our kids. The other book that I was listening to on Audible recently is Incorruptible by Eric Ries. Incorruptible, Incorruptible. Incorruptible. Yes. Yes. Yeah. His recent podcast.

32:49

Yeah. And I just think the question of how to build great companies is important. I've personally just been most fascinated with how to keep great teams and great companies going further. And it was very interesting to just see his framing and reframing of the question.

32:58

I loved some of the examples around having metrics around culture. If you only measure revenue, then that's kind of how you're gauging against it, but if you have other better metrics, that's actually the way to sustain the values you care about. I've been trying to think about how to actually bring that to the team level of how do we better articulate, write our norms, a lot of the things we talked about on the team. So I think that's also a really good read. There you go. That'll be your next watch, everyone. As you're listening to this, the Eric Ries episode, such a good episode.

33:11

Yeah. And his book just came out, Incorruptible. Yes. And I think it was a New York Times bestseller. It's actually doing incredibly well, which I was really happy to see. Yeah, exactly. Next question. Favorite recent movie or TV show you've really enjoyed? Most people at Anthropic do not have time to watch things, but I'm curious if you have an answer. I would say during some time off last month, I did get to binge-watch Fallout on Amazon Prime. So that was actually, I kind of like, have you heard of it? Yeah. Yeah. It's based on the video game.

33:48

Yes. It's based on the video game. I think it was really witty. It's humorous. It's also super action-oriented. So highly recommend. Okay. Next question. Do you have a favorite product you recently discovered that you really love? I really do think Claude Tag is very interesting in terms of a product experience. We actually have different versions of this within Anthropic, and I think it's actually been a really, really powerful tool.

33:59

Yeah. It feels like, I think some people are like, what's the big deal? The fact that everyone at Anthropic is raving about it tells me something important is going on here. And I'm trying to actually get it working within my Slack community that I have for paid newsletter subscribers. How cool would that be? Yeah. Yeah. We're trying to figure out how it works when it's not a company, when it's just a bunch of people that don't know each other and how that might work, but we're trying it out. Okay. Two more questions. Your favorite life motto that you find yourself often coming back to in work or in life.

34:26

I really do think cloud tag is very interesting in terms of a product experience. We actually have different versions of this within Anthropic. And I think it's actually been a really, really, really powerful tool. Yeah. It feels like some people are like, what's the big deal? The fact that everyone at Anthropic is raving about it tells me something important is going on here. And I'm trying to actually get it working within my Slack community that I have for paid newsletter subscribers. How cool would that be? Yeah. Yeah. We're trying to figure out how it works when it's not a company, when it's just a bunch of people that don't know each other, and how that might work, but we're trying it out. Okay. Two more questions. Your favorite life motto that you find yourself often coming back to in work or in life.

34:32

So I was actually raised by my grandparents for the first 10 years of my life. And my parents were immigrant college and master students in the U.S. And my grandfather always says, no matter how far you go, there's always another level.

34:39

Which is, I think, a really good way, though a pretty intense way, of describing his life or philosophy. But I go back to that whenever there's something new or unprecedented that we experience. And I think the first half of this year, there was definitely a lot of that. There were a lot of new things that we were learning, I was learning. So just feeling like there's always another mountain, another opportunity to climb. Not good enough, Dan. We need to go better. We need to go bigger. Makes me think about actually another Ben Mann line from his podcast episode, that this is the most normal it's ever going to be. It's only going to get weirder and crazier. Yeah. Yeah. Oh my God. Okay. Final question. I was poking around at your LinkedIn. You were a high yield bond trader at JP Morgan Chase early in your career. You had this redacted hundred million dollar trading portfolio of some kind. What did you learn from that time in your life that has stuck with you? And, or is there a crazy story from that period? It was four years of your life.

34:47

I think I learned actually a lot that I apply here at, at, at, at Anthropic and other jobs thereafter. So when I was at JP Morgan, the trading floor, you could envision Wall Street. That's very different. Most traders, I think, are in front of a terminal. They're much more doing analyses on their computers. But it's still very, I would say, male dominated. And so I was the only woman, I was the only person with my background on the trading desk. And I learned that that was a very good environment to kind of build, one, my sense of authentic self, and two, that even if I was the most junior person, even if I may look different, the best ideas and having conviction in the best ideas, irregardless of all of those other factors, is the most important thing. And so I think just bringing that sense of how I show up more at work, I'm pretty vulnerable and authentic with my team. I try to really make sure that regardless of people's levels or tenures, if they have a great idea, how to help them pursue that, and to do also the same. So to put the idea out there, to actually have conviction in it, to do the follow through, to do the nitty gritty work to make it happen. So those were all things that I learned from trading. And yeah, I think it applies to any job in many ways.

34:54

That is beautiful. Where can people find you online if they want to follow you? And how can listeners be useful to you?

35:00

I don't have a large presence on social. I think the best way to find your work, my team's work, is really the Anthropic blog. And when we're publishing new models, new product experiences, I think in terms of useful for me, I think the best thing, number one, is your feedback. We actually, if you thumbs up or thumbs down on any of our product surfaces, if you contact your salesperson with feedback about the model, it will make its way to me. We actually, with every research model, I actually get pretty close to understanding favorability and feedback. So giving us that feedback, pushing Claude, telling us where it's falling down, those help us make Claude better. The other thing is, if you have folks in your network who seem like this type of profile person that I just talked about, I'm hiring, the team is growing. We really would love people who love this technology, who are deeply curious, first principles thinkers, who are fearless in questioning assumptions, and who have a tinkering, hackery spirit.

35:05

Wow. What a dream job. So open PM roles at Anthropic on the research team. Yes. And they apply, I assume, on the website, the careers page. Yes. Holy moly. All right, here we go. Enjoy the flood of resumes you're about to receive. Thank you. Dan, thank you so much for being here. Thank you so much for having me. Thank you for really helpful, thought provoking questions, helping me even connect the dots on how, how we work, how, how this whole technology is coming together and being product people in it. I really appreciate that. But thank you, Dan, for real.

35:45

Okay. Well, bye everyone. Thank you so much for listening. If you found this valuable, you can subscribe to the show on Apple Podcasts, Spotify, or your favorite podcast app. Also, please consider giving us a rating or leaving a review, as that really helps other listeners find the podcast. You can find all past episodes or learn more about the show at LennysPodcast.com. See you in the next episode. to get better will now have a lot more scrutiny and there will be more restrictions on who can use them, which feels like a big deal. How do you think about that? How does that change the way you operate?

36:01

I'm going to maybe leave the policy and the expert control side to, to focus on that and work on that. I think the product question and how we interact with these internally is, I think, as you mentioned, as frontier models become more capable, the safeguards and the ways of red teaming and testing and the pre-release process also needs to evolve and adapt quickly to, to address that. And so one example is, you know, before Fable models, we didn't have a strong, uh, let's say fallback UXs and systems because our, our, our goal is to make sure that like there is asymmetrical benefit for this technology

36:48

and to minimize like the downside or like a severe risk of, of it. And so we ended up building like fallback systems so that users will still get a great response from Opus 4.8 immediately. And so I think there's a piece around, uh, as we evolve and like improve safety systems, how do we continue to develop and deliver great user experiences? I think there's more that we can do on both sides. And so you'll see us innovating, improving on what we call now the model safeguards package, uh, more and more in the coming, coming weeks and months. What's really interesting and just like unexpected here is, creates this really interesting advantage

37:33

for Anthropic where you have access to the latest stuff. And this is going to happen at every lab. Everyone's going to keep improving and it's, it creates this unfair advantage within the labs to have access to the best stuff that other people can't get outside of your control. You'd prefer everyone use it. So it's a really interesting, this new feedback loop that's going to start where models that are so advanced are only accessible to certain companies. And that's going to be a whole new, unexpected, it's like a second order effect of all these restrictions. Our goal is to be, uh, to develop these systems and the models to be as inclusive as possible. Um,

38:05

I think our goal is to not have that happen, uh, for the general purpose, general use like technologies and to make it more accessible. I think, you know, this is like one of our top priorities right now to kind of reduce what we're seeing there. Yeah, that makes sense. I would imagine you'd want as many customers and people using this thing as possible. This episode is brought to you by Mercury, radically different banking loved by over 300,000 entrepreneurs. And now with command, I've been a customer of Mercury's for over six years. I have never once thought about leaving. Mercury is basically what happens when banking is built by

38:42

product people, not by bankers. They make it so easy. Dare I say fun to send invoices, move money around, set up virtual cards for folks on my team. Does your bank have an API, a terminal native CLI or an AI ready MCP server? I don't think so. And just recently they launched command, a conversational interface built directly into Mercury, which acts as your financial operator. I've been using command to transfer money around to figure out what categories I've been spending the most money in analyze my cash flows. And just today I used it to find out how much I've made from a

39:17

specific sponsor over the past year. I just asked how much have I made from X over the past year, 10 seconds later I have an answer. It is so freaking cool. Visit mercury.com to learn more and apply online in minutes. Mercury is a fintech company, not an FDIC insured bank. Banking services provided through Choice Financial Group and Column NA members FDIC. I want to talk a little bit about how the product role is changing and who is doing well in this new world. Now that AI is such a core part of of our life. When you're hiring PMs, product people, when you're looking at people that do well in

39:55

today's world, what are some things that you notice? What are you looking for more most? What are you looking for more? What's kind of like trending up and what you find is important and what's kind of trending down? We actually on my team have not changed our hiring loop for three years now. So what we actually look for and the traits and how we evaluate generalists like PINs, generalists like research product managers have actually been the same. So I think some of those traits, number one is first principles thinking. And this is really rather than pattern matching what you used to do

40:37

in let's say consumer product or B2B SaaS. But actually figuring out in this moment for this user group with this technology, what is the user value? Is there an example that a lot of people hear first principles thinking? They're like, yes, I got it. I'm good at this. What's an example of someone having really demonstrated really good first principles thinking? I think one example is, I think you think of a product manager as I own product strategy and delivering user value as, but I demonstrate day to day by writing a PRD or writing a product vision doc. And for my team as

41:19

like research product managers, the way to drive user value is to figure out the right user feedback, the evals, right, that then can be a personification of that user need. So like we do write some product documents and PRDs, but we actually have a saying on the team of evals are the new PRDs, right? Because in order to deliver that user value, it's not that exact artifact that people used to write in the last like one to two decades. It's a new way of working. And so the first think principles thinking would be, let me figure out what is the thing I should do to achieve my goals rather than here is a set of

42:06

activities that I've done and therefore I will continue to do. So the idea here is used to be have kind of an idea, create a PRD, talk to people about it, align on the plan, design it, build it, ship it, see how it goes, iterate. What I'm hearing here is it's like, okay, here's some feedback about something that's wrong or an opportunity. Step one is the eval is now how you define what the work is versus a PRD. Maybe step one would be understanding the user pain point. And so the way to even access that user pain point is different, right? In the past, we might do a user interview.

42:45

I think if you go like deep enough, you might have the user walk you through their user flow, the pixels. Here, you have to sweat the tokens as much as you sweat the pixels. And so one activity we have on the team is reading the transcripts and understanding what was the trajectories that failed very deeply to then say, was this like a hallucination? Was this cloud being overconfident? So like the theme of the failure actually has a lot of nuance. And then that allows you to build a description, a like sustained description of that pain point. So that could be essentially in a new eval. And is

43:32

an eval on distribution, right? Is it capturing both the positive situations where this is failing and also areas when it should actually not fail and then bring that back to, let's say research, so that we can make the improvements and actually measure the quality of, okay, when we have Opus 5.5, is this area improving or not? Is Claude now able to identify the right places in the document and pull the right synthesis out? So it's just the actionability, like, and shortening the distance to actionability for our stakeholders and partner teams like researchers to take action on.

44:16

Is there an example of something like this where you found an issue or opportunity and then wrote the eval? And what is, what is the eval looking like in, in most cases? What, when people want to picture an eval, what is that? What is, what do they picture? We actually pioneered this concept within Anthropic. So one of the early examples is the early cloud models were not very good at following specific schemas. So like things like outputs and JSON. And now that is fundamental to Claude being able to be a good agent, right? If you can't output a certain format, you don't know how to like access APIs, you can't call tools, et cetera. And so the initial

45:01

end to end was, I was hearing feedback around Claude 2 days. Claude was not very good at following instructions. So then digging in with users, what do you mean by Claude is not good at following instructions? Give me what situations this was happening. Like what's the exact like paragraph? What did you ask? What was Claude's response going to like that level of detail? And what I saw was something like 80% of what people meant in the early days for this failure was Claude would not write the right JSON. And so then, okay, let's generate maybe to start just 30 to 40 examples

45:42

of when Claude was not doing this thing correctly. And then that actually is your eval set. And you can have essentially a prompt and a response. And if that is not working in the right golden answer that you might have, then that means that the eval essentially is beneficial because it's identifying a pain point consistently. And so then we added that to our repositories for evals. And when we have versions of Claude, we actually run that eval and just check. I think at this point, it's always 100% or like 99.9. And so it's no longer a pain point. But in the early days, it was taking the user feedback, figuring out

46:29

actually what they mean. Can we reproduce it? Is it consistent? Is it a big issue? And then figuring out how to standardize it in a way that can be consumable for researchers. It's basically test-driven development for PMs is the world we're living now, where you write the test first. So is this just a core part of the product management job now at Anthropic Writing Evals? I think so. I also think it's something I've talked to other PMs at other companies about. And I think it's also more and more of the skill set more broadly. Because a lot of the products that we're

47:06

building is at the intersection of models with harnesses, with a set of contacts for a set of users. And so having things like evals actually is a way not just for folks working on models, but generally within product to get to better user experiences. Because you can't improve what you can't measure. And a lot of this is very still tactile-based. It's still very judgment-based. And so you have to stay close to the details. And also very non-deterministic, which is a big part of this. Just like it's not going to give you the same answer every time. So you got to describe it kind of more broadly. It's not going to be

47:47

an exact match. So this is a really interesting change in the way product happens and will happen is evals. Writing evals versus PRDs is a big part of this. Do you guys still do PRDs? Is there still like a one-pager describing a problem or is it replay? Okay. Now you're shaking your head yes. We are. We do. I think when there's a very defined problem, I think things like evals might be almost a shorthand. I think there's other cases where PRDs are really valuable. PRDs are great vehicles for getting a very large group of people aligned on a set of sources of truth about experience and set of

48:24

goals. So when we do have a model, we actually, for every model, we do have a PRD. Less necessarily for our researchers, but more for our growing product surfaces, for our engineering teams, for our stakeholders like legal and safety and others as just a source of truth of putting together what we're aiming to achieve so that a big group of people can row in the same direction. The other place where I do think PRDs are valuable are on the more ambiguous problems and opportunities. Right? So we, if we haven't shipped a thing like computer use, we don't necessarily have a set of like user

49:07

specific pain points always. And I think there's value in the product vision portions of a PRD to explore what could, even if a technology is not yet ready to work for everyone, how do you get it to work well for some group so you can explore the value, you can actually bring something that is coherent to a user group. So we do have PRDs. I think the application is a little different now. Okay, this is great. There's, I just had a, uh, Andrew for, he's the head of the Codex app at OpenAI and he's, you guys are aligned. Uh, PRD is not dead. Still very useful for specific projects and ideas. Uh,

49:51

great. Okay. We've closed, closed the book on. PRD is still kicking. Okay. So we've been talking a bit about just what kind of skills are kind of emerging for product people. Um, is there anything else that you find is shifted in what patterns, uh, are common across people that are doing well in this new AI world in terms of product managers and folks on the product teams? Is there anything else that you're like, okay, there's something you've got to shift or something you look for more people? I think maybe specifically, uh, for folks who might be mid career or folks who have been

50:27

more in a managerial like product, like leadership seat. Um, one thing that I think I feel pretty strongly about is in order to be good managers of teams and PMs working with this technology, you have to be really hands on yourself and have spent not just time tinkering, but actually shipping the technology and, and, and again, being in the details and sweating the tokens along with your PMs and your engineers and your teams. And so even for folks that I hire who have more tenured PM experience, the onboarding plans are exactly the same as somebody who is like more, uh, early career. And it's around

51:18

understanding users, reading like consented user feedback, talking to customers. I think there's something around, uh, being able to like understand what to do with this, what, what good looks like and having developed that in a very hands-on manner, that's important. Um, it's not necessarily easy for someone to, uh, agree or be able to see what a, what a good or great AI product or AI feature could look like if they haven't kind of experienced building themselves. Um, so I think, I think there is a, I, I, I do feel pretty strongly that like, you know, if you're a manager, you have to be hands-on, you have to

52:06

spend a portion of your time actually shipping, you, you have to kind of walk in the shoes of your teams. Uh, and, and that's, I, I always try to carve out a portion of time, uh, to, to actually like own one to two work streams when we have models in order to keep, like, keep my theory of mind, keep my sense of how the models are moving, how quickly it's improving, uh, so I can help the team make, make decisions and, and make better decisions. So what I'm hearing here is if you're not, no matter where you are in the ladder of hierarchy at a company, if you're not building yourself, if you're not actually talking

52:43

to Claude, talking to Codex, building stuff, you're not going to make it. And you should have fun working with his technology. I think that's the other piece. I think the folks that would be most successful regardless of their level are people who love working with AI and, and are exploring and experimenting and carving out the time, not just for the experimentation, but actually hands-on shipping end to end, getting the user feedback, I think has to be fundamental for everyone. I a hundred percent know what you mean there. Just like me sitting on my newsletter and this podcast, just talking about stuff and like, yeah, yeah, that sounds great. Like every time I actually

53:20

built something and I tinker with all kinds of little projects, you just like, okay, I see what's happening here. And you just get so much more. It's like hard to exactly describe what you, what you, experience actually working with the models and building stuff, but it's like a whole different world of like, okay, I see. Here's what they look. Here's what they're talking about computer use. Here's what they're talking about with this limitation of this UX situation. Yeah. Yeah. So it's just like, and you made this really interesting point that you have to have fun with it, which is not easy for a lot of people because they're pushed to use AI or they just don't know

53:50

exactly what to do with it. For people that are just like, I don't know, it's just so annoying. I just have to do this. I don't know what's just like, I hate this frigging thing. Why do I have to work with this? Things are changing so much. I'm tired. Advice for helping people find that, find that joy in this work. I think maybe I'll reemphasize something I said earlier around just that experimentation is not an individual sport. Like some of the moments where I think I've touched practically every version of research models across 20 plus versions of production claw at this point.

54:24

And I think part of the joy comes from seeing other people discover use cases too. And so maybe one idea here would be pairing with somebody who is excited and seeing what on a use case that you care about and working together versus identifying or trying to figure out the perfect use case yourself. Because that might feel like work. Working with others feels like joy a lot of the time. And is there more that we could do to bring that, bring other people along? That's something like a lot of times internally we have somebody who is like very curious and them sharing an idea of a new prototype actually

55:08

brings a ton more people who are like, oh, I didn't know this could work now with Claude. And so there's some virtuous cycles here and ways of continuing to have joy with this technology. That's such a good point. I think that's also why Twitter is so useful for a lot of this is you see other people sharing what they've done and it inspires you to come up with your own little ideas. And also it's just like fun to share your own thing that you've done. So that's a really good point. Just like find other people to kind of play around with and look for use cases. The thing I've also heard

55:41

a lot is just find like a problem you want to solve in your life or work and just open up Claude, Claude, tell it here's what I want to do. And it's incredible how far you can get just with like a vague idea of a problem you want to solve. Yeah. I think it gets hard in that there's so many different things that you could try. Yeah. And so you just like narrowing in on either pairing with someone, working with somebody who, who is, who have a lot of joy about this technology or figuring out something that you could immediately find value. Like either of those things allow you to go deeper

56:14

rather than like more high level about too many things. I, I find it hard to keep pace with the number of prototypes or products that are out there. And so my lens has been, how do I go deep in one to two of them myself? That's a, that's so interesting. You say that because that's exactly it. We just had the survey that I ran with my colleague, Noam, asking my readers just how they're feeling about all the things going on in the tech right now and AI. And one of the most interesting takeaways we had was to find that happiness is exactly what you said is go deep in a couple things versus trying to just

56:53

ton of little things, find a couple of things to really solve well, and then go deep. And that is a source because a lot of the happiness people feel is when they finally unlocked a way for AI to actually make their lives better versus just like a couple of messed up, broken half working things. Yeah. It's, it's, um, how do you go from this being a check the box? Right. And so like us as product people, it's then a exercise of product prioritization of your time and your energy. And if the goal is to experiment with joy, then how do you, what are the inputs that you need for that?

57:27

Yeah. Um, but yeah, I, I think a lot of the, um, I think the secret sauce of anthropic is the culture and the bottoms of nature of how people work and this like experimenting in public. Um, and by doing that, it's very much about how to bring other people along. Um, that ends up being, I think, really valuable. Yeah. I've heard this so many times from all the labs, just like, no, no, one's exactly sure how some of this is going to be used. And a lot of it is just putting stuff out early, seeing how people use it, seeing what it's, what's possible. And then using that information to build that product to lean

58:10

in. Yeah. Yeah. I'm curious how kind of on this thread of finding ways AI for AI to help you in your work and life, are there any interesting ways you've been using Claude lately in your work as a, as a PM? I think there's a lot of things with, um, you know, fable and things like tag. So there, there, I think tag is, um, in, in the very like early days, I think there's something around how you work in a different paradigm of allowing this, an agent to go off and work and then bring back, uh, product experiences to you. I think one area that it's not more recent, but one that, um, I bring up a lot with the team

58:51

and I think we could do more on using AI is just like how to use it to also be more, uh, to have better conversations with each other, to be better managers. I don't think it's necessarily, uh, just about raising the IQ of like experiences we build, but also I use it a lot and actually like prepping for how to have better conversations, um, in the moment during like crucial conversations. So I love that book. And so I actually have a skill that helps me figure out, am I having, am I going in the right level of detail given the situation at hand and actually helping me be a better manager and better supporter for the team? Um, so for, for like managers on the team, that's

59:40

actually a thing that I've been sharing more with, with, uh, with our managers. So, okay, how, how do you actually use, use Claude to, to, to make you a better coach? Because it's hard sometimes to find the right perfect words and the models have a lot of perfect and bright words. And, uh, I think there is something about how, how he can actually augment us from like an EQ perspective in addition to you. Oh man, there's so much interesting stuff there. So just to understand what you're doing there. So you, you built a skill, you're just like, help Claude build a skill, pulling in lessons from crucial

1:00:13

conversations, the book, which it knows enough about. You don't have to even give it the content. And then you use that skill to talk to Claude. Hey, I have this very difficult conversation coming up with a colleague. Give me some tips on how to approach it. Yeah. And it's, it's a great, uh, it's almost like, uh, coaching, like individualized, personalized coaching of just how to make you. And, and there's so much context switching that we do all day and having like Claude help me pair and help me. And maybe there are times where I end up not using suggestions from Claude, uh, but it

1:00:49

actually is, uh, ends up being very helpful for, for just coming up and brainstorming. Am I thinking about reactions in the right way? How do I actually, uh, go a bit deeper, faster, build trust faster, uh, be more direct? Yeah, man. I have so many questions here. This is so interesting. Uh, one is just like, there's concern. People are going to start talking the way AI writes because they're talking to AI so much. And it's going to be like, Diane, it's not this, but it's that, uh, I know that you're not doing that, but that's all, you know, a concern people have. Let me just ask about that. I guess, do you fear this? There's this, you know, brain rot,

1:01:25

uh, atrophy stuff. People talk about it. We're just so reliant on AI now, and we stop learning and thinking and, you know, overall AI thoughts on that being so close to it and being so integrated with, with AI constantly. A lot of actually thinking process and writing process are tied together for me personally. And so I think there are ways where I use Claude to augment my thinking, but what I want to make sure, and maybe this is what you're describing is Claude doesn't take over all of my thinking for me. And so I think depending on the situation, depending on how much more personal judgment I want to have in a situation, I might, uh, um, come up with my own

1:02:08

POV first and then work with Claude through that. Um, and making sure that like I maintain my sense and tone throughout, I think there are then other things like updates, right? We have like monthly business reviews. And then in those cases, it's much more, I watch, you want it to be standard and I want it to be much more like it gets a crisp crystallized information in the right way. And maybe, and I have a skill and like we're augmenting and improving our skill for that. But I want to get to a place where like the monthly business review, the writing of that is potentially asymmetrically less

1:02:48

valuable than the thinking. And so how do I get that piece delegated to Claude fully? And I'm more of a reviewer and a verifier of that information. So I think it depends on like what you're using Claude for and what you're trying to convey. And like, is there, is there a asymmetrical value in, in delegating more to Claude? What I'm also hearing the first tip is really great, which was think first, have a point of view, and then kind of use Claude as a, uh, sparring partner almost to evolve the idea, push back on the idea. Yeah. Yeah. And I think this is where things like actually our alignment research and safety

1:03:27

research is helpful because it, what you don't want is like a AI that just agrees with you, right? What you want is this technology to actually augment and grow and like get to a better outcome. And so sometimes it's having Claude push back makes me better. And so that's great. Like a coworker, I want somebody to push back when my ideas are not fully formed. I want to hear more about that. I've heard that when Ben Mann was on the podcast, he talked about the constitution that is built into Claude and how, unintuitively, the work and the focus on safety and alignment, as you said, and this constitution that describes how Claude should think and operate that

1:04:11

actually, you would think that would limit the abilities of Claude and make it less fun and interesting. It's exactly the opposite. Claude is the most interesting personality. I hear that constantly. It's just like, I much prefer talking to like open claw famously was built on Claude. And then people were forced to switch. We won't get into it. We're forced to switch to Chappie T and they're like, this is so bad. This is not who I'm used to talking to. So that is, I think a really interesting point. I just want to make sure we spend a little time on why is it, why is that the case? Just this

1:04:41

focus on alignment, safety, having this clear constitution. Why does that make Claude better and more interesting to talk to you also? In order to make Claude as like intelligent and as capable as possible, being able to have Claude actually push back in the right points and then add, it's like a yes or no and actually helps you come to a better conclusion. So I've used Claude to help with things like, are we making the right pricing decision on the next version of Claude? It's a little bit meta, but using a research version of Opus, asking it to figure out how it should price and being able to

1:05:23

come out with better outcomes is a goal at the end of the day. And so having AI not just be an assistant, not just be a doer and being delegated tasks by figuring out, is it doing the right thing? That's actually very integrated with knowing when to push back, right? That's part of knowing when you should be proactive. Proactivity is not necessarily always doing a thing that you are scheduled to do. It is knowing when to come up with a new idea. And so in order for Claude to be more useful, the general approach has to be that it knows when to push back. It's a core part of the characteristics together

1:06:08

of the models. That is so interesting. It's so interesting that that is what a big part of it, like it being less compliant is almost what makes it better and more useful because we need that. Like I've had so many people where they're like, hey, like AI told me I was right. And like, no, I wish I wish you could listen to other people. Yeah. And it comes back to our earlier point around thinking, right? How do you protect your thinking? If you have a AI that can be a thinking partner, a thinking partner doesn't just agree with you. It should add to you and you should come away at the

1:06:42

end of the day, having better ideas because you've worked with Claude. That should be the hero goal. Not just making your ideas 10% better. Yeah. I love this. And like, it used to be think 10x. I used to be the, the way, you know, founders push people. Like what if we 10x of this? And I love what I keep hearing is like, it's like, how do we go 1000x from this idea? What is the most ambitious version of this? I want to come back to something that I was thinking about as we were talking about, talking to Claude constantly. It's very clear when AI has written something still. It's funny that it's

1:07:15

a large language model. You would think of all things, it would be very good at writing. And interestingly, just no AI is very good at writing. It's always very clear. This was AI written. Do you think we'll get to a place where we will not know this was AI? I think it depends on what's the goal that you're looking to achieve by knowing or not? What's the eval? Yeah. What's the eval? I actually do think there's more that we could be doing on making Claude write better. There's actually very active efforts on my team and on the research side about making Claude write better, just generally.

1:07:55

I think it should be clear where an idea is being led by you or by Yuleni or me, Diane. I think it really depends on what's the goal of that writing. Like for something like a monthly business review, I would actually love to have that end to end be written by Claude. And obviously, and not to make it feel like it was written by a human. It's such an interesting point you're making. Like, is it actually better for us to know that it's AI versus not? Yeah. But it's also for maybe the lens is more around like verifiability or who's verifying the output, right? Like who's signing off? Maybe less around who's writing, but who's verifying,

1:08:45

who's signing off. That becomes like more what matters than who's writing it. Why do you think AI is not great at writing? Like my guess is it has studied all of the best writing in all of humanity. It's figured out here's the best way to write. And now that we, and it's just, there's only so many ways to write. And so we've just recognized, okay, this is what AI does. It has these tropes. Is that the core of it? Is there something else that's keeping it from being a great writer? Ironically, being a large language model of all things, you think it'd be really great at language. I think part of it is also, we need to invest more in training improvements to

1:09:27

make AI continuously strong on areas like writing. I think it's also, you know, like the technology's jagged edged, like we mentioned. So sometimes when the models were good at writing, but not agentic, our, our thesis is how do we make the models more agentic or call the right tools. Now that that's improved a bit, then it's, well, now these other areas actually become more of the rough edges. And so I think we're in one of those moments worth writing where we need to actually just focus and prioritize on training the models to be like great at this area. And like, that is an active, a very active area for us. So fun, thank you mentioned.

1:10:11

Okay. I'm glad. I'm glad. And also, it was going to be interesting once AI is so good. We're like, I don't know who wrote that, but to your point, sometimes we actually want to know that it's AI. That's really interesting. I never thought of it that way. The other interesting part of this is that there's that comedian who was joking that we're like on a plane and the wifi is down and we're just like, what the hell? The wifi is not working on this plane. The socks, how dare you? Like when you're like in a, in a tube in the sky flying like a bird and how dare you complain that

1:10:38

the wifi doesn't work? Like your point is there's so much advancement and so much power. We can't fix it all. We can't make it all work the best possible. And so basically AI writing has been not the priority and it feels like there's more investment happening there. Yeah. I think like tone and character is a priority. I think it's this advancement of the technology is a work in progress. And so we made, we, we see a leap or emergence of like a jump in agentic behaviors. And so that is a new normal. And then these other capabilities needs to continue like improving. Yeah.

1:11:16

And I think once we improve, let's say writing and like tone and character, uh, we probably will say like, how do we have Claude be even more proactive, like proactivity is an opportunity and that's human nature. Like we want to make ourselves better. We want to make this technology better. Um, so yeah, it, I, and I think we're applying it to, to AI, which is the right thing. We should be making it better. I want to ask you a couple of questions. I like to ask folks working at the very center of the future of that is coming. Um, one is where do you think human brains will continue to be most valuable over

1:11:53

the years? I know anthropics mission and, and vision is we'll reach a GI, a, uh, super intelligence. So in the future, maybe nowhere, but before we get there, where do you think human brains will continue to be most valuable as we approach that, that timeline? We started to talk about making Claude and models better at judgment. Um, especially in the last, um, year or so. I think judgment is one, and is an area where it's accumulation of so much nuance and so much experience and these systems haven't experienced as much as humans have. And so I think that hard earned, like, judgment

1:12:36

judgment is a, a, a area for, for product leaders and just generally, um, will continue to be really critical. There are so many things AIs can build. Which one are the things that, you know, an org like lab should build, right? A lot of that requires like human judgment, persistence. So proactivity, these are all traits are, are beyond just general capabilities, but just behaviors and characteristics of like people at that level of like, how do you get to the best solutions? How do you create the best experiences? So I think those types of traits are actually the tactile, uh, traits that I think will be, uh, continue to be important.

1:13:20

Um, I think there is also, uh, still a lot of like capabilities and subject matter expertise, as well. I think, you know, software engineering has been really transformed by AI. I think there's areas like, uh, biology, life sciences. These are all things that, um, we're just kind of at like the foot of the exponential on, like maybe software engineering, we're on the exponential on some of these area, other areas, we're not quite there yet. And so, um, I think you're seeing us ship things like cloud science, investing in these areas because those are areas that, um, I think it's just

1:14:02

bring the, this technology to society and having a positive benefit for society. So I think there's a lot more to go there. Another question I would ask is, um, as someone with kids, how do you think about what you are encouraging them to learn? Or do you think you're gonna nudge them to be successful in this wild new world that we're entering? I actually think it's a lot of the same traits like you and I probably grew up with, which is curiosity for learning, persistence, believing in your own inner voice, developing and then believing in your own inner voice. Like I have a four-year-old, I have a eight-year-old.

1:14:46

It's on us to help. Uh, it's on me to help them develop their inner voice and whether that's being opinionated and taking a stance to me, right. And developing that, encouraging that. Uh, I think that those types of skillsets are things that, um, is important in the future and I like having their own individual voice. That is so interesting. It's so related to the answer you had when asked about how to avoid a brain rot, essentially an overrelying on AI, which is just keep focused on your own point of view and your own perspective before you overrely on AI. And just this idea you're describing of building that

1:15:26

in kids is, is really important. Uh, that is so interesting. And I love how all this, all this kind of connects judgment, persistence in a point of view of your own. Yeah. Both for kids and also adults. Yeah. Anything we think about, um, for your. Oh man. Well, like the question I'm thinking about is just when to get them on like some AI thing, you know, when I, I have a three-year-old, so it's pretty hurt for that, but you know, how do you get, how do you onboard them to this crazy thing? I had, I was at an event recently and a bunch of parents were talking about how they think about AI and their kids. And one person had a really interesting

1:15:57

approach, which is, uh, keep them on the very early models so that they still have to struggle a bit and not get all the answers immediately. That was interesting. Like an open source local model, not fable. Yeah. Oh yeah. And curiosity is something, uh, I keep mentioning Ben Mann, but his answer actually to this question has always stuck with me, which is, um, curiosity. And also just like, he's a big fan of Montessori, which is what I'm encouraging for our kids. So there's something there. Maybe a last question, just along, it's kind of along these lines, something Fiona Fung actually

1:16:31

suggested I ask you, uh, who's recently on the podcast. How do you stay just recharged and not burn out being in the center of this crazy storm of AI as a mom, uh, working in, you know, uh, we're seeing the research work at Anthropic. Uh, I just like we're living through the most unprecedented time working at just like being, you know, being on the outside of Anthropic, it's crazy. I don't even know what it's like to be on the inside. Um, what have you learned about avoiding burnout, staying recharged, staying sane during the middle of all this? In 2024, we shipped four models for the, in the whole year

1:17:05

or four series of models. And I think we did more than that volume in just Q2 of this year. I think I've been really lucky with, uh, the team that we've grown and built both the stakeholders on the research side and within our research product management team. Um, I think that one of the magical parts about approaching all of this is that it's not an individual sport. Um, there's like a sense of radical ownership and team collaboration that I think sometimes it does feel like a high performance sport because you're in very critical decisions. There's new information about users, about training, and you have to make recommendations and judgments and decisions very quickly.

1:17:59

And nobody can do that sustainably by themselves. Um, and so I think what's really helped is having a team that is incredible, who looks out for each other, who, you know, the night before a launch, even if they're not the core DRI on that model will stay up and help the DRI who, uh, to review the blog post and make edits and come up with better demos and knowing to be each other's sort of extra hand. I think it's very easy if you take all of this change on your own shoulders to feel like you're alone and to feel like you have to do everything. Um, but I think one of the like magical parts of

1:18:42

Anthropic is this ability for us to, uh, figure out what are those opportunities to help each other and actually then taking the next mile of like mind melding. We called it like entering the hive mind. There was an article about this. And I think like part of that is just that allows like the team to replenish. It's not that you, I was just on PTO in June. It's not just that you can take PTO and you come back to like 3x the amount of things to do is actually that you can take PTO and know the team can figure out the right things to do and that we individually can like watch out for each other.

1:19:21

Um, so I think that's a big part. I'm really lucky just personally. Um, also my partner is really supportive. Um, this is year six of me working in AI. So Amazon and then Anthropic. And so he sees how much I just love the technology and what this can do. And that really helps. I think also, um, from like a personal perspective as well. I love, I love how many of these answers connect. So what I'm hearing here is just the having other people, working with other people, relying on other people, helping each other out when things get crazy. Uh, it, which is a similar answer you had for just how to, how to find

1:19:59

the joy and, and, and fun in this work just gets and gain, be inspired by other people and see what they're doing. Yeah. Yeah. And it's interesting when Fiona was on the podcast recently, she, I was asking her just like what's changed in the world of software engineering. And she pointed out, it's a lot lonelier now because now we're working with agents instead of other humans. Teams are smaller. People are having all these fleets they're talking to constantly. And so this is just a reminder of just the power of just actual other humans around you. We're asked to work and make decisions on really

1:20:29

big things because you have more scale from the technology. Right. And I think having individuals, having other folks more who can have some level of like mind meld with what you work on, how you approach maybe not exactly every detail, but what are the first principles? What are the assumptions you make that helps them, you know, back up for you or push your decision and sharpen your thinking? So I think, you know, we really try to like, I really try to look for that when like building the team, growing the team, hiring, like, is this person going to care about their own ego and building out

1:21:10

a big org or are they going to care about contributing to Anthropic and contributing to that like impact of the team and orienting towards folks who are like low ego, team oriented. I think that's, yeah, it's a big part of, I think the sustainability. Yeah. Just always, a lot of it's always just comes down back to culture and hiring. And I know I've heard a lot, just the reason Anthropic is able to move so fast. I remember that moment when like something shipped every day of the month, it's like a calendar of launches and people were talking about how is this possible. And what I heard a lot is just because everyone is so aligned around the mission

1:21:49

and the values, it allows people to make decisions really quickly. Before we get to our very exciting lightning round. Is there anything else, Dan, that you wanted to share anything else you wanted to touch on anything you want to maybe double down on of things we've talked about? This was actually really fun because I feel like your questions actually sharpen some of my thinking around how the thoughts kind of connect. I'm your real human clod over here. One thing that I really want to like convey or have people take away is I think one in the ways of working, but also just to like,

1:22:26

this is a, this is a lot of like growth and change and having the joy in using this technology. And like, if you're feeling like in this moment, you don't have as much of that feeling of initial joy, how do you find people who do? If this is an area that you're excited and like want to work on? And I think developing skill sets, replenishing skill sets in many ways of things like thinking from a first principles manner about what you solve. I think fundamentally, you didn't ask me this, but there's this question in the community of, do we still need PMs when the models are so capable, when engineers are leaning in?

1:23:10

I think the role of people who are user-centric, who go into the details of understanding what users are trying to accomplish, bubbling that up in an actionable manner and doing the relentless work to do that. Like that to me, it's a core of a product person. And I actually think we need more of that. I think we are becoming very technology layered driven and actually to make that impactful, it's, you have to go deep, you have to be curious, you have to be super hands-on. And those are things that I think are also traits that have, I think, helped anthropic from a product

1:23:51

development and model development perspective and as part of the culture. And hopefully that's valuable for others as well. Amazing. What an inspiring way to end it. Oh man. Yeah. And this is, I've been saying this too for a long time, just now that building is easy, the hard part becomes, as you said, what should we build and is the thing we have built correct and good and worth leaning into? And to me, that's what PMs do and what PMs are good at. Yeah. Yeah. Yeah. And it's getting into the details of the user. Yeah. Empathy. Okay. Great. PMs are going to make it. Okay. PRD is not dead. All kinds of,

1:24:29

all kinds of important lessons here. Diane, with that, we've reached our very exciting lightning round. I've got five questions for you. Are you ready? Yep. First question. What are two or three books that you find yourself recommending most to other people? One personal one, I really like how to raise an adult. So, uh, I'm a mom. I think a lot about what is the things that I want to instill in, in, in my kids. And that book is really helpful for describing, we're not trying to raise children or trying to raise adults. So just the framing of what does that mean and what does it mean? What are the characteristics

1:25:09

that we want to hone and like harness and foster in our kids? Um, the other book that I, uh, was listening to on audible recently is incorrigible by Eric Ries. So the, the, the, the, Incorruptible, incorruptible. Incorruptible. Yes. Yes. Yeah. He's recent podcast. Um, yeah. And I, I, I just, I think the question of how to build great companies is important. I personally just been most fascinated with how to keep great teams and great companies going further. And it was very interesting to just kind of see his framing and reframing of the question. Um, I loved some of the examples around having metrics around culture. You, if you can, if you only

1:25:52

measure revenue and then that's kind of how you're going against, but if you have other better metrics, that's actually the way, uh, to, to, to sustain the, the values you care about. I've been kind of trying to think about how to actually bring that to the team level of like, how do we better articulate, write our norms, a lot of the things we talked about on the team. So I think that's also a really good read. There you go. Uh, that'll be your next watch everyone. As you're listening to this, the Eric Ries episode, uh, such a good episode. Yeah. Uh, and his book just came out in corruptible. Yes. And I think it was like a New York times bestseller. Like it's

1:26:27

actually doing incredibly well, which I was really happy to see. Yeah, exactly. Next question. Favorite recent movie or TV show you've really enjoyed most people at Anthropic do not have time to do what, to watch things, but I'm curious if you have an answer. I would say, um, drink, uh, some time off last month, I did get to like binge watch fallout on Amazon Prime. So that was actually, uh, I kind of like, um, it's kind of, uh, have you heard of it? Yeah. Yeah. It's based on the video game. Yes. It's based on the video game. Uh, I think it's a, it was really, um, it's witty. It's humorous. It's also like super action oriented. So highly recommend.

1:27:08

Okay. Next question. Do you have a favorite product you recently discovered that you really love? I really do think like cloud tag is very interesting in terms of a product experience. Um, we actually have like different versions of this, uh, within Anthropic. And I think it's actually been really, uh, really, really, uh, powerful tool. Yeah. It feels like, I think some people are like, what's the big deal. The fact that everyone at Anthropic is like raving about it tells me something important is going on here. And I'm trying to actually get it working within my Slack community

1:27:40

that I have for paid newsletter subscribers. How cool would that be? Yeah. Yeah. We're trying to figure out how it works when it's not a company, when it's just a bunch of people that don't know each other and how that might work, but we're trying it out. Okay. Uh, two more questions. Your favorite life motto that you find yourself often coming back to in work or in life. So I was actually raised by my grandparents, uh, for the first 10 years of my life. And my parents were immigrant, uh, college and master students in the U S and, um, my grandfather always says, no matter how far you go, there's always another level.

1:28:18

Which, uh, um, is, I think, um, a really good way though, like a pretty, uh, intense way of describing, uh, his, his life or philosophy. But I go back to that whenever there's something new or unprecedented that we experienced. And I think, you know, first half of this year, there was definitely a lot of that. Like there was a lot of new things that we were learning. I was learning. Um, so just feeling like there's always like another mountain, another, uh, opportunity to climb. Not good enough, Dan. We need to go better. We need to go bigger. Uh, makes me think about actually another Ben Mann

1:28:56

line from his podcast episode that this is the most normal it's ever going to be. It's only going to get weirder and crazier. Yeah. Yeah. Oh my God. Okay. Final question. Uh, I was poking out at your LinkedIn. You were a high yield bond trader, JP Morgan chase early in your career. Uh, you had like, uh, you have this like redacted, uh, uh, a hundred million dollar trading portfolio of some kind. Uh, what did you learn from that time in your life that has stuck with you? And, or is there a crazy story from that period? It was four years of your life. I think I learned actually a lot that I, uh, apply here,

1:29:35

uh, at, at, at, at anthropic and other, uh, jobs thereafter. Um, so when I was at JP Morgan, um, the trading floor, uh, you could kind of envision like sort of wall full wall street. That's very different. Uh, most traders I think are in front of a terminal. They're much more doing analyses, uh, on their computers. Um, but it's still very, I would say like male dominated. And so, uh, I was the only woman, I was the only, uh, um, person with like my background, uh, on the trading desk. And I learned that the, that was a very good environment to kind of building one, my sense of authentic self.

1:30:20

And two, uh, that even if I was the most junior person, even if I may look different, uh, that the best ideas and having conviction and the best ideas, uh, irregardless of all of those other factors, like, is the most important thing. And so I think just bringing that sense of, um, how I show up more at work, um, I'm pretty vulnerable and authentic with my team. Uh, I try to really make sure that regardless of people's levels or tenures, if they have a great idea, how to help them pursue that. And to do also the same. Um, so to like put the idea out there to actually, um, have conviction in it,

1:31:09

to do the follow through, to do the like nitty gritty work to make it happen. Um, so those were all things that I learned from trading. Um, and yeah, I think applies to any, any job in many ways. That is beautiful. Where can people find you online if they want to follow you? And how can listeners be useful to you? I don't have a large presence on like, uh, social. Uh, I think the best way to, uh, find your work, uh, my team's work is really, uh, the anthropic blog. And when we're publishing new models, new product experiences, I think in terms of, uh, useful, uh, for me, I think the best thing,

1:31:52

number one is your feedback. Like we actually, if, if you thumbs up or thumbs down on any of our product surfaces, if you contact your salesperson with feedback about the model, it will make its way to me. Uh, we actually, with every like research model, I actually get pretty close into understanding favorability and feedback. Um, so giving us that feedback, pushing Claude, telling us where it's falling down, um, those help us make Claude better. Uh, the other, the other thing is like, if you have folks in your network who seem like this type of profile person that I just talked about, I'm hiring,

1:32:28

the team, uh, is growing. We really will love just people who love this technology, who are deeply curious, first principles thinkers, who are fearless in questioning assumptions. Um, and who have like a tinkering hackery spirit. Wow. What a dream job. So basically open, open PM roles, ad anthropic on the research team. Yes. And they apply, I assume on the website, the careers page. Yes. Holy moly. All right, here we go. Enjoy the flood of resumes you're about to receive. Thank you. Uh, Dan, thank you so much for being here. Thank you so much for having me. Thank you for, um, really helpful, thought provoking

1:33:11

questions. Um, helping me even connect the dots on how, how we work, how, how this whole technology is coming together and being product people in it. I really appreciate that. But thank you, Dan, for real. Okay. Well, bye everyone. Thank you so much for listening. If you found this valuable, you can subscribe to the show on Apple podcasts, Spotify, or your favorite podcast app. Also, please consider giving us a rating or leaving a review as that really helps other listeners find the podcast. You can find all past episodes or learn more about the show at Lenny's podcast.com. See you in the next episode.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note