Simon Willison in conversation with Cat Wu & Thariq Shihipar, Anthropic
Description
A long form Q&A with Cat Wu (Head of Product, Claude Code) and Thariq Shihipar (Engineer, Claude Code) from Anthropic, moderated by Simon Willison. The discussion focuses on the evolution of coding agents and how they have fundamentally shifted software development practices. Key Takeaways: Changing Developer Workflow (1:22 - 3:51): Coding agents like Claude Code have moved developers from manual, low-level implementation toward higher-level product strategy. The focus has shifted from writing every line of code to managing and refining outputs from increasingly capable models. The Rise of Proactive Agents (6:37 - 9:00): The introduction of Claude Tag marks a shift toward multiplayer, proactive agentic workflows in tools like Slack. It can monitor bugs, draft PRs, and retain team memory, currently landing over 65% of product PRs for the Anthropic team. Rethinking Software Engineering Norms (3:51 - 6:37): The speakers argue that traditional practices, such as avoiding rewrites or long-term waterfall planning, are becoming outdated. They emphasize the value of product sense, prototyping, and rapid iteration using high-quality test suites. Safety and Trust (30:57 - 37:53): A major portion of the discussion covers Auto Mode and safety. Anthropic has heavily invested in evals and red-teaming to mitigate risks like prompt injection and data exfiltration, making it the recommended way to handle long-running agentic tasks safely. Cultural Impact (38:13 - 41:50): The team encourages developers to be more ambitious and stop "negotiating against themselves." They suggest that because implementation is now cheaper, teams should focus on building the "bigger things" they previously thought were impossible or too resource-intensive. Future Outlook: Looking ahead, the speakers are excited about models becoming better interaction design partners (44:28) and continuing to bridge the gap between abstract ideas and production-ready software. ## Speakers ### Cat Wu Head of Product,
Summary
Generated by gpt-5.6-solAt-a-Glance
- Verdict: Watch fully
- Core thesis: As frontier models become capable of sustained autonomous work, the engineering bottleneck shifts from implementation to judgment, orchestration, verification, permissions, and choosing higher-value problems.
- Why it matters: Anthropic discloses concrete operating patterns for team agents, automated code review, model-specific prompting, contextual authorization, credential isolation, memory, and eval-driven trust that are directly applicable to production agent systems.
- Best use: Use the conversation as an architecture and operating-model reference for designing collaborative agents, progressively reducing human review, and securing long-running autonomous workflows.
Executive Summary
Anthropic describes a step change from closely supervising every Claude Code action to delegating entire features and long-running tasks. The consequence is not that engineering becomes trivial: implementation gets cheaper while product taste, business judgment, UX quality, prioritization, and verification become the scarce capabilities. Practices once considered dangerous, such as rewrites, become more viable when agents can treat the existing codebase as a specification and a strong test suite protects behavior.
Claude Tag is presented as the team and proactive evolution of Claude Code. It operates inside shared collaboration channels, can retain channel-level preferences, search organizational context, and continuously react to events without being manually invoked. Anthropic says its internal version now lands 65% of product-engineering PRs, while interactive Claude Code remains preferable for the hardest tasks requiring close iteration.
The strongest operational lesson is that autonomy must be earned through infrastructure. Anthropic moved selected outer-layer changes toward agent-only review over more than six months, using code ownership for critical components, automated review on every PR, incident-derived regression cases, robust CI, internal and external evals, and red-team testing. The same philosophy underlies Auto mode, which evaluates requested actions in conversational context, mediates sandbox and network escapes, and supports credentials that are usable by an agent without exposing the underlying secrets.
The discussion also challenges common prompting and tool-design assumptions. Anthropic reduced the system prompt for its frontier models by roughly 80%, removing examples and hard prohibitions that over-constrained model judgment, while retaining more explicit prompts for older models. It is also trending toward fewer, more orthogonal tools and letting models use general capabilities such as native Bash. The caveat throughout is that Anthropic does not claim complete eval coverage, published evidence for Auto mode was still forthcoming, and frontier models still lack consistently excellent product and interaction-design taste.
Key Takeaways
- Claim: As coding agents compress implementation from months to days or weeks, engineering value shifts toward deciding what deserves to be built and defining a sufficiently ambitious quality bar. | Evidence: Cat Wu contrasts the former six-to-twelve-month process of customer research, PRDs, cross-functional alignment, and implementation planning with cases that can now be built in about a week. The speakers say engineers increasingly need product and business sense, while PMs at Anthropic routinely prototype, design, automate launch operations, and fill engineering gaps themselves. | Implication: Ken should assess agent-system teams less by raw implementation throughput and more by problem selection, product taste, system-level judgment, and their ability to turn faster execution into larger business outcomes. | Caveat: Infrastructure and other critical systems still require close attention to implementation details, and the models' design and UX taste remains weaker than their ability to follow a detailed functional specification.
- Claim: The next useful abstraction beyond a personal coding agent is a persistent, multiplayer agent embedded in the team's collaboration layer. | Evidence: Claude Tag can live in a Slack channel, accept steering from multiple teammates, remember channel preferences, search public organizational context, and proactively monitor bug reports, create fixes, and tag the engineer who last touched the relevant code. Anthropic says its internal version lands 65% of product-engineering PRs. | Implication: For OpenClaw or an AI control plane, the collaboration surface should be treated as a shared execution environment with persistent policy, memory, visibility, and handoffs—not merely as a chat interface to a private agent. | Caveat: The 65% figure applies to Anthropic's product-engineering team rather than all company code, and the team says it is still learning the social dynamics of multiple people steering one agent session.
- Claim: Removing humans from portions of code review is feasible only through gradual, evidence-based expansion of an agent's authority. | Evidence: Anthropic retained mandatory human code owners for core areas such as the Claude Code system prompt, while allowing automated review to handle selected outer-layer changes after a six-plus-month trust-building process. When an incident occurs, the causative PR is used to improve the reviewer and is added to a permanent eval set to prevent regression. | Implication: Ken should structure autonomous deployment and review as a risk-tiered progression: establish measurable catch rates, preserve human ownership of high-impact surfaces, and convert every escaped defect into a regression test before expanding authority. | Caveat: Anthropic still requires human approval for critical core changes and does not claim that automated review is universally ready for all repositories or risk levels.
- Claim: Production-agent evals must test both task completion and undesirable interaction behavior, not just benchmark-level capability. | Evidence: Anthropic combines external evals with a larger internal suite that checks whether Claude fixes a fully specified bug and passes tests. It separately builds behavioral evals for user-visible failures such as telling users it is time to sleep or stopping after two of five requested steps to ask whether it should continue. New models are run against team-specific and company-wide evals before being substituted. | Implication: Ken should treat eval-set construction as an ongoing product discipline: capture completion quality, behavioral friction, security failures, and real production incidents rather than relying on a single aggregate benchmark. | Caveat: The speakers explicitly say they do not have complete confidence or full behavioral coverage; feedback is prioritized and converted into evals incrementally. They also argue that creating high-quality eval data, rather than eval tooling itself, is usually the limiting factor.
- Claim: More capable models may perform better with substantially less instruction because examples and absolute rules can suppress useful judgment. | Evidence: Anthropic reduced the Claude Code system prompt for its frontier models by about 80%. It removed many examples, supplied more context, used fewer 'do not' rules, and softened absolutes such as 'always verify' after identifying legitimate exceptions. Older models retain fuller prompts, so Anthropic now uses different system prompts by model. | Implication: Model routing should include prompt routing: Ken should not assume one universal system prompt across providers or capability tiers, and should test whether frontier models benefit from principles and context while smaller models need examples and explicit procedures. | Caveat: This finding is model-dependent and eval-driven; it should not be generalized to smaller or older models. Anthropic also lacks hard data showing that a frontier model can always generate optimal detailed prompts for weaker models.
- Claim: Agent tool sets should remain small and orthogonal, with dedicated tools retained only when they provide a clear capability, control, or user-interface advantage. | Evidence: Claude Code removed dedicated grep, glob, and other search tools in favor of native Bash, while keeping a file-edit tool partly because it lets the product deterministically detect edits and render an approval UI. The team says every added tool should have a distinct function so the model can easily choose among them. | Implication: Ken should resist adding narrow tools for operations the model can perform reliably through general primitives, but preserve explicit tools where observability, authorization, auditability, or deterministic UI treatment requires them. | Caveat: The speakers characterize tool design as partly empirical and 'more biology than physics'; some choices are driven by onboarding and interface needs rather than raw agent capability.
- Claim: Safe long-running autonomy requires contextual authorization layered with sandboxing, red-team evals, agent identity, and secret isolation. | Evidence: Auto mode uses a classifier to judge each tool or Bash action against the conversation and user request, allowing an action such as git push when requested and blocking it when the user said not to push. It also mediates sandbox and network escapes. Anthropic reports thousands of evals, multiple external red teams, internal use since January, separate Claude identities, trusted devices, and credential injection that lets an agent call Datadog without reading the actual token. | Implication: Ken should view natural-language permission classifiers as one layer in a defense-in-depth control plane, not as a substitute for least privilege, sandbox boundaries, scoped agent identities, secret brokering, audit logs, and adversarial testing. | Caveat: Anthropic says Auto mode does not catch 100% of attacks, and the detailed eval results promised in the conversation had not yet been published. Claims that its risk is lower than average human review therefore require independent validation.
Detailed Brief
Dogfooding, retention gates, and unexpected workflows
- Claims: Anthropic uses internal dogfooding and early-customer feedback as a primary feature-selection mechanism rather than shipping everything that cheaper implementation makes possible.; Features must clear internal active-use and retention thresholds before public launch, giving engineers a concrete adoption target and discouraging the release of unpolished experiments.; Observed user behavior can overturn the product team's assumptions about the correct workflow.
- Evidence: Remote Control initially seemed unnecessary to Cat Wu because she preferred cloud development sessions, but users adopted a distinct routine: leave local Claude Code sessions running on a plugged-in laptop and control them from a phone while away from the desk.; For collaborative feature work, a PM can ask Claude Tag for a first implementation and a recording, then bring design into the same session for refinement before engineering takes it to production.
- Caveats: Internal adoption at Anthropic may not predict behavior in organizations with weaker AI fluency, less open information sharing, or stricter compliance boundaries.
- Implications: Retention and repeated autonomous use are more informative than prototype novelty when evaluating agent features.; Mobile control and asynchronous resumability may become important control-plane surfaces even when the underlying work executes on local or remote infrastructure.
Emerging orchestration, memory, and multimodal use patterns
- Claims: Frontier models are becoming effective prompt writers and workflow orchestrators for other models, including subagents and specialized generation APIs.; Claude Tag's current team memory is deliberately simple and scoped to the collaboration context rather than implemented as a broad enterprise knowledge graph.; Coding-agent workflows extend beyond software development into media production, research, travel planning, and customized internal applications.
- Evidence: The speakers describe Claude generating detailed prompts for subagents, orchestrating multi-step Workflows, and prompting Gemini or image/video models before inspecting generated frames.; Claude Tag currently stores shared channel memory in one Markdown file per channel; individual sessions can contribute information back to that shared memory.; In a one-shot video-editing example, Fable transcribed a talk, rejected a corrupted slide recording, matched the transcript to HTML slide sources, dynamically tracked and cropped the moving speaker, and assembled the result using FFmpeg and Remotion.; A personal fighting-game project used generated 2D sprites and model-derived JSON hitboxes, while a climbing app researched direct flights, suitable routes on Mountain Project, accommodation, and short walking approaches.
- Caveats: Anthropic says memory design remains unintuitive and experimental, with no disclosed answer for how the file-based approach will scale.; The video-editing result is an anecdotal one-shot success rather than a measured reliability claim.
- Implications: A lightweight, scoped, inspectable memory artifact may be preferable to premature investment in a complex memory database.; The orchestrator's ability to formulate prompts, inspect outputs, and retry across specialized tools is becoming as important as direct code generation.
Notable Concepts & Terms
- Claude Tag: Anthropic's collaborative, proactive agent inside team communication channels; it adds multiplayer steering, persistent channel preferences, organizational search, and event-driven work to the Claude Code model.
- Auto mode: A contextual permission system that evaluates proposed agent actions against the user's instructions and conversation, while also mediating sandbox and network permission requests.
- YOLO mode: The informal label for broadly bypassing permission checks; the speakers discourage it in favor of the hardened Auto mode.
- Incident-derived evals: The practice of turning the PR behind a production incident into a permanent test case for the automated reviewer, creating a ratchet against repeating known failures.
- Lean, model-specific system prompts: Anthropic's practice of giving frontier models shorter, less prescriptive prompts while retaining more examples and constraints for older models.
- Contextual authorization: Permission decisions based not only on the requested tool call but also on what the user asked for or prohibited, such as conditionally allowing git push.
- Credential injection: A secret-brokering pattern in which the execution layer inserts the real credential into an approved outbound request, allowing the agent to use a service without exposing the secret in its context.
- Codebase as specification: The idea that an existing implementation captures branching behavior that may not exist in any written spec, making agent-assisted rewrites plausible when paired with a strong test suite.
Operator Notes / Why Ken Should Care
- Create explicit autonomy tiers for agent changes, with file-path or service-level risk classifications, named human owners for critical surfaces, and measurable promotion criteria for automated review.
- Build a production-incident ingestion loop that automatically converts escaped agent failures into reviewer, policy, and regression eval cases.
- Benchmark separate system prompts by model tier instead of inheriting one universal prompt across frontier, mid-tier, and low-cost routes.
- Audit the current tool catalog for overlapping capabilities; remove redundant narrow tools unless they materially improve permissions, telemetry, deterministic rendering, or user comprehension.
- Broker third-party credentials through scoped proxies or workload identities so agents can invoke approved services without reading reusable secrets.
- Require independent review of Anthropic's promised Auto mode security results before granting it authority over sensitive repositories, production deployment, or unrestricted outbound networking.
- Test channel-scoped Markdown memory as a transparent baseline before introducing a vector store or complex long-term memory layer; measure retrieval accuracy, stale-memory errors, and cross-channel leakage first.
- Decide deliberately whether broader public-channel access is acceptable: organizational search quality improves with visibility, but confidentiality boundaries and data minimization may outweigh the context benefit.
Source/Metadata
- Title: Simon Willison in conversation with Cat Wu & Thariq Shihipar, Anthropic
- Transcript words: 11447
- Duration seconds: 3090
- Timestamp note: No timestamps or chapter markers were present. The transcript contains several duplicated passages and some inconsistent product or model-name transcription.
Transcript
Welcome to this Fireside Chat. I have with me Tariq Shihipa and Kat Wu from Anthropic. We are going to be diving deep into Claude Code, and we'll probably talk a little about this Fable thing that's been out there in the news. Actually, on the subject of Fable, literally a minute and a half ago, Fable came back. Fable is now available to me, so if you all want to run out of the room and start using up your Fable credits, I wouldn't hold that against you. But we're going to have a great conversation, so please stick around. But please welcome Tariq and Kat for me. Thanks for having us. Yeah, we timed it for the chat, for sure. Yeah, this is why it's all happening. This year has been somewhat absurd. It's amazing. Claude Code came out in February of last year. It's under a year and a half old, and it was a bullet point on the Claude Sonnet 3.7 launch. I'd love to hear from you, how has your day-to-day basis changed in the past year, now that we have these coding agents that actually work for us? I remember when we first came out with Claude Code in Sonnet 3.7, you would give it this task, and you would have to closely monitor every single little thing that it tried to do. I remember I would read every permission prompt extremely carefully. I would frequently say no. I would always say no, no, no. Did you check this file? Did you check that file? And now it's been incredible. With every model generation, I feel we've all gotten a chance to just take a step back, delegate a lot more of the menial implementation to Claude, and it just freed up a lot of our time to think about more creative work. What is the right experience that we should be providing to our users now that we know Claude Code can implement a lot of it? And now with Fable, it's just a totally different step change improvement. We see for a lot of our use cases that you can actually one-shot a ton of features with Fable now. So it's been amazing to see the transition and to go through this with all of you in the community. [SPEAKER_02] Yeah, I remember the first text I got about Claude Code. One of my best friends was like, oh, you need to go try Claude Code. And it was just when Opus 4 came out. And I tried it, and I was like, oh. I need to work at Anthropic now. And that was Opus 4. I mean, great model, but yeah, permission prompts. And yeah, I think it's crazy how much amnesia we have, where I'm like, oh, auto mode has always been here, right? I don't even remember pressing yes and allow. And yeah, I think for me, the big thing that I'm trying to push myself is, oh, we have to do higher quality work than we've ever done before. The outputs are incredibly high quality. I've been using it to edit videos a bunch. And I'm like, okay, it has to meet the very exacting demands of our brand team, and in a couple hours, or we just can't do it. And so, yeah, I think that's how I'm sort of trying to shift with Fable, where it's like, okay, the best work we've ever done, faster than we've ever done it before. I've certainly been finding that myself. Software engineering is getting harder, because the level of ambition of the stuff we can take on has gone up. I have such higher expectations of myself now that I have these tools to back me up, which is fun, but it's a lot of work. Yeah, it's all the thinking is, you know, it's tiring. [SPEAKER_02] And so, what's a piece of conventional software engineering that was true a year ago that you don't think holds anymore in this new world? I think one of the biggest shifts that we're seeing in the eng skill set is, I think, two years ago, it was pretty typical for a product manager to go talk to a bunch of customers, and over the course of six months, align with cross-functional teams on some PRD, and then write this thorough spec and eng doc on how exactly we'll implement this before the first line of code gets written. And now, things are completely turned the opposite way. I think for a lot of engineers, the push I would give to a lot of folks in the room is to develop more of your business sense and product sense on what is it that we should build? Because now that the timeline between having this idea and building it is so much shorter, it's down from six to twelve months to maybe even a week, that means all of us need to have better taste on what is it that is worth building? What is it that will actually inflect the businesses that we're working on? So I think it's an increase in value on product taste and business sense, and a bit lower on execution in most product eng domains. Of course, for infra, there's still a very heavy emphasis on making sure all the details are right. Yeah, I think for me, it's rewrites are now good. I think that...the worst thing you could do is now actually fine. Yeah, exactly. All the, especially, Mythical Man-Month stuff, never rewrite. I'm a pro rewriting now. If you have a good test suite, I think actually the rewrite forces you to make sure you have a good test suite. But I think what people undercount is a code base is a spec. And maybe it's the only copy of the spec that you have, because no one knows every branching part of the code base. And yeah, you can take this as an artifact and distill it or create other versions of it. Obviously, yeah, we rewrote Bun in Rust. And you know, it works great. It's live for me right now. [SPEAKER_04] I've been a prototyper. And now I prototype things on my phone during the conference just so I've got something that I can pick up later on. And that's working now, which is kind of extraordinary. [SPEAKER_04] So the other big launch recently was Claude Tags, which is a week old now, I think, or at least for the rest of us. Claude Tags, I understand that's being used in Anthropic by non-engineers a great deal. What kind of things are non-engineers doing with Claude Tags? [SPEAKER_04] So Claude Tags is a Claude that lives in your team's collaboration tools. We launched it last week within Slack. The thing that's different about Claude Tags is it's multiplayer by default. And so once you add Claude Tags into a Slack channel, you can chime in, your teammates can chime in, and you can collaborate together on the work. [SPEAKER_04] So the other big launch recently was Claude Tag, which that's what a week old now, I think, or at least for the rest of us. Claude Tag, I understand that's being used in Anthropic by non-engineers a great deal. [SPEAKER_04] What kind of things are non-engineers doing with Claude Tag? So Claude Tag is a Claude that lives in your team's collaboration tools. We launched it last week within Slack. The thing that's different about Claude Tag is it's multiplayer by default. And so once you add Claude Tag into a Slack channel, you can chime in, your teammates can chime in, and you can collaborate together on the PR. [SPEAKER_04] The other big difference is that it's proactive instead of reactive. So you can tell Claude Tag, hey, monitor every bug report in this channel, put up a PR to fix it, and tag the engineer who last touched this part of the code base, [SPEAKER_04] and it'll do it for the lifetime of the channel without you having to manually tag it in. [SPEAKER_04] And then the third big shift that we've seen is we've added team memory into this. So if you tell Claude Tag your preferences in the channel, [SPEAKER_04] it'll remember this for every future post. So if you wanted to always debug outages, but you don't want it to debug warnings, [SPEAKER_04] just tell it that in natural language in the channel, and it'll remember it for you and everyone else on your team. Internally, we see Claude Tag as the evolution of Claude Code. So we see this as a large shift in how we work internally. Claude Tag currently lands 65% of our product ENG PRs. For all of Anthropic and for Claude Code or just for Claude Code? This is just for our product engineering team. So our internal version of Claude Tag lands 65% of our product ENG PRs right now. [SPEAKER_00] This is a huge shift. This is more than 50% of our PRs. [SPEAKER_00] And the way that we actually see people split work between Claude Code and Claude Tag is Claude Code is still the best place for your most complex tasks when you're interactively iterating with the agents. [SPEAKER_04] But Claude Tag is great for having it work proactively on your behalf so that you no longer need to manually kick off Claude Code for all of the bug reports that might come up for features that you're working on. [SPEAKER_00] Yeah, and for non-coding cases, I think we've seen people use Claude Tag, for example, before this talk, we asked Claude Tag, hey, when is Fable releasing? [SPEAKER_00] We wanted to make sure that we'd line it up with the announcement. And so Claude Tag would search our Slack and look at who's been saying what. So as a search engine for your company is really valuable. It has all the context for your product. So you can ask it metrics related questions. And oftentimes when you're making decisions, you want it to be informed by what do the metrics say? And then you hook it up to your event store. I've seen our marketing team do things like, hey, tell me about this feature. And they're not programmers, but Claude is a programmer. It can clone the code base and be like, yeah, this is the feature. This is what it looks like. This is a recording of me using the feature. So it just enables a whole wide variety of things. And I think we're still early on in figuring that out. Well, I feel like this is one of the fascinating things about the Claude Code story is you use Claude Code to build Claude Code. And you've been doing this since presumably before the public launch of Claude Code a year and a half ago. And yeah, one of the problems I've had with coding agents is I get how to use them as an individual, but I'm not really clear on how I use that in a team environment. It sounds like Claude Tag is your current answer to that sort of team collaborative layer for this stuff. Exactly. And a large percentage of our sessions are actually multiplayer right now. So that means maybe I say, hey, I think we should implement this new feature in CoWork. And I'll tag in Claude Tag to do a first pass at it. And then I'll tell Claude Tag, hey, just share a recording of your final implementation. And then I'll tag in design to take a look and they'll nudge it. And then they'll pass it on to Eng to take it to the finish line and get it out to prod. And so it's been this very fluid experience. We're still trying to iron out what the social dynamics are for steering the same session. But we found that people just observe how others use it and then follow those social norms. [SPEAKER_02] And it's actually been pretty easy for us, pretty intuitive for us to integrate Claude Tag into our teams. Yeah, I think it's great for teaching people and also reducing slop because if you see someone just be like, hey, at Claude, fix this. You're like, I think the fact that everyone is seeing you use Claude together sort of levels up how you use Claude as well. Right. You want to do work that you're proud to do in public where the quality doesn't fall off the cliff. So since we're talking about using Claude to build, how do you deal with the hardest problem in all of engineering? It's prioritization, right? How do you decide which features are worth building and shipping when building a feature is so much more inexpensive now? This is the hard thing. So there's a few ways we approach it. One is we dog food our products every single day. Whenever there's something that we want to be able to do in our products that we're not able to, instead of finding a different solution, we fix our products so that it can support this case. We have a very heavy dog food culture internally. So before we are able to share products with everyone in the world, we share it with everyone within Anthropic and we share it with some early customers who give us very honest feedback about it. The more brutal, the better. And we iterate until people love it. So we have an internal bar for the number of active users and the amount of retention a feature has to have before we share it with the world. And because this bar is very clear, every engineer knows what they're trying to hit. And I think this also levels up our polish because if the feature isn't polished, people will turn and then we shouldn't ship that feature. Do you have an example of a feature which surprised you when you rolled it out and the engagement was off the charts and it was something that was unlikely to be shipped that actually turned into a real product thing? I do have one. So a lot of folks on our team love remote control. So remote control lets you use your mobile device or Claude in the web browser to connect a local Claude code session running in your CLI. I never have this need because I just kick off the task directly on mobile and it runs in a cloud session and doesn't use my local environment. I think this is because I'm doing very easy coding tasks. But this is something where I didn't totally understand it. I was like, hey, people should just set up their remote dev environments. But in practice, once we rolled out remote control, so many people I talked to are like, OK, now what I do every night is I ask a lot of people tell me that they just plug their laptop into a power charger, close the screen or open a bunch of remote control sessions, lock the screen and then use their mobile phone from their couch to control Claude code. I never have this need because I just kick off the task directly on mobile and it runs in a cloud session and doesn't use my local environment. I think this is because I'm doing very easy coding tasks. But this is something where I didn't totally understand it. I was like, hey, people should just set up their remote dev environments. But in practice, once we rolled out remote control, everyone, so many people I talked to are like, OK, now what I do every night is a lot of people tell me that they just plug their laptop into a power charger, close the screen or open a bunch of remote control sessions, lock the screen and then use their mobile phone from their couch to control cloud code. And so this has been this flow that we're now leaning into that I didn't originally get. But now I do. [SPEAKER_00] I do exactly that. I get so much work on my laptop done for more comfortable environments because I can remote control it now. That's really fun. [SPEAKER_00] How does code review work? Are you reviewing? Does a human being review every line of production code that makes it into cloud code? And if not, what are you doing? How do you keep the quality up? [SPEAKER_00] Sure. Yeah, it varies on the task a lot. So for important areas, we have code owners. Right. And so the system prompt is an example where we have a code owner, you really need to submit your approval. And then [SPEAKER_00] So I guess the code owner is directly responsible for the quality of that area of the code. That's right. Yeah. Okay. [SPEAKER_00] They need to approve the PR that touches it. Right. That's right. We have code review. Our code review runs on every PR and oftentimes that's doing the bulk of the review. I think something I've seen on the team is for more complex PRs you might make an artifact to explain the PR so that other people can then review. And yeah, we just invest a lot into verification, CI/CD, things like that to make sure that anytime anything fails, we have a test. We have a really robust environment called code and test it. [SPEAKER_02] So yeah, there's just a multi-pronged approach to code review, I think. Do you have anything to add? [SPEAKER_02] In general, we are trying to move to a world where humans don't need to be in the loop. And so for the most critical core changes to the core of code and other cores of other products, there is always a code owner and they do manually review all the changes. [SPEAKER_02] But increasingly for the changes that are at the outer layers, we actually have code review fully review those. [SPEAKER_02] That sounds pretty scary, but there we've had this six plus month long process to get here. And I think there are baby steps that you take to build up trust with code review. [SPEAKER_02] So in the beginning, we would have human review for everything. And then increasingly we would say, okay, for code changes that touch these files, code review is catching 100% of the issues there. [SPEAKER_02] So we actually don't need a human to be manually reviewing those. And then also when we have incident review, we look at the PRs that cause the incident and we say, okay, how do we update code review to catch that? [SPEAKER_02] And then we also take those PRs and add it to an eval set to make sure that our future changes to code review never regress that metric. So it is a big step. Removing humans from the code review loop is a big step forward. I think it can sound scary and it's not something that you can do overnight, but it is something that you can do through many months of investment in the infrastructure to give you the confidence that code review is catching everything that you care about. So it's interesting. You mentioned building trust in the models. That's something I found is I know that Opus 4.8, if I ask it to build me a JSON endpoint that runs a SQL query and outputs JSON, it's just going to get it right. That's not something I have to review closely, but then a new model comes along and I still don't know how do I build trust in Fable quickly that it's not going to mess things up that Opus didn't. Is that something that you have to think about much? How does the new model affect your intuition for what it can do and what it can't do? So the main reason that we're building up this eval base over time is so that new models can be a drop in replacement. Because what we do when we have a new model is we run the whole eval set and we make sure that, for example, Fable is strictly better than Opus 4.8. And that gives us the confidence to drop it in. And those model evals for Anthropic as a whole, or are these Claude code team specific evals that you're using? [SPEAKER_00] We have both. So we have evals on our team and we run code review across every repo within Anthropic. And so we have evals for that. And for things like auto mode, we not only have evals across every user within Anthropic, but we've also commissioned multiple external testers to red team this, to create environments with prompt injections and malicious inputs and make sure that auto mode doesn't let any of those pass. So for Claude code itself, and this is a challenge I've had with stuff I'm building. I want to know if the system prompt improvement I made actually improved the product, right? That's the sort of most basic form of product specific eval. [SPEAKER_04] And I still don't have a great feel for how to do that. [SPEAKER_04] Is that something that you're doing such that you have complete confidence that this tweak that you've made to the system prompt does result in better output? [SPEAKER_04] We don't have complete confidence, but we do a lot to make sure that we don't regress. [SPEAKER_04] So the starting point that we have is we have a suite of external evals that we trust, and we complement that with an even larger suite of internal evals that we trust. To start, we mainly optimize for capability. So given a complete definition of a task and the full code base, does Claude make the right decisions and fully fix the bugs and pass all the tests? So that's the starting point. And that's the thing that we optimize for because it is most directly what users want. But there's a lot of behaviors that impact how users feel when they work with Claude code. For example, people really don't like it when Claude code says it's time to go to sleep. Or people really don't like it when it says, hey, I finished two out of five parts. Do you want me to continue? Yes, please continue. And so we're building up a set of behavioral evals to catch these. And as we get user feedback, please be loud with us about your user feedback. [SPEAKER_00] As we get user feedback, we just rank OK, these are the priority issues. [SPEAKER_00] And we go down one by one and build evals for each of them. [SPEAKER_00] So it's not 100% coverage, but we try to. It is a priority for us to increase the coverage. [SPEAKER_00] And how much overlap is there between the, how much interaction is there between the Claude code team and the teams at Anthropic who are training the models in the first place? [SPEAKER_00] Is that quite a close collaboration now? [SPEAKER_02] Across Anthropic, we all work quite closely together. And so we're building up a set of behavioral evals to catch these. [SPEAKER_04] And as we get user feedback, please be loud with us about your user feedback. As we get user feedback, we just rank. OK, these are the priority issues. And we go down one by one and build evals for each of them. So it's not 100% coverage, but we try to. It is a priority for us to increase the coverage. And how much overlap is there between the Claude code team and the teams at Anthropic who are training the models in the first place? [SPEAKER_02] Is that quite a close collaboration now? [SPEAKER_02] Across Anthropic, we all work quite closely together. So we meet often to talk about what do we expect the next generation of models to be able to do? I think our research team has also been amazing about showing this publicly. So we often talk in our blog posts about how we're targeting ever increasing longer horizon work, how we train Claude itself to be honest, harmless, and helpful. We also put a lot of effort into making sure that it's aligned with your intent. Even if your intent is expressed in a fuzzy way, of course, try your best to be specific about what you want so Claude has all the context. [SPEAKER_04] But even when you're not specific, we teach Claude to make good assumptions. [SPEAKER_04] And yeah, I think it's been a productive partnership. [SPEAKER_04] And so, Tharik, this morning you mentioned that the system prompt for Claude codes has been reduced by 80% because of Claude Fable. [SPEAKER_00] Can you go into a little bit more detail about what that looks like? [SPEAKER_00] What kind of things have you been able to drop? [SPEAKER_00] Yeah, so it wasn't just Fable. It was Opus 4.8 as well. And yeah, going forward, the future models. But we do have different system prompts for different models now. I think that some of the patterns we saw is that we were over-constraining Claude, right? [SPEAKER_02] So I think the initial maybe Opus 4-ish kind of models wanted a lot of examples. And removing examples was extremely helpful because it was just more creative than the examples we gave it. That's really interesting. Because one of the top prompting tips I give people is give it examples. Yeah. Examples are the easiest way. If that's no longer true, that breaks my prompting model a little bit. [SPEAKER_00] Yeah, same here. [SPEAKER_02] I think I was surprised to hear that. [SPEAKER_02] I think that now it's more about the shape of what you give the tools to Claude. [SPEAKER_02] And your system prompt and things like that. [SPEAKER_02] The other thing we did is we try to give it more context and fewer "do not do this" instructions. [SPEAKER_02] Because I think that it's just a very strong impulse to Claude. [SPEAKER_02] And especially if that conflicts with user instructions later on, that can be extremely confusing to Claude. [SPEAKER_02] Because you're like, oh, I've got this skill that says this and the system prompt says this. [SPEAKER_02] And so we try and have fewer hard constraints and more context and fewer instructions overall. [SPEAKER_02] Yeah, I think it's definitely a science that took a bunch of evals to build. [SPEAKER_02] I'm not sure if you had anything else on the lean system prompt. [SPEAKER_00] I think in general when you're prompting these models, you should always think about are there edge cases to the instruction that I'm giving it? [SPEAKER_00] And when we went back and we reviewed all the instructions on the Claude Code system prompt, we found a few cases where yes, this statement is 90% true. [SPEAKER_00] But there's a real 10% of cases where this is not true. [SPEAKER_00] And we didn't want to constrain the model or confuse it into thinking, hey, it should always do this. One good example is verification. Everyone here wants Claude to verify its work. And we had some instructions in the prompt that just said, if you make a front-end change, always verify. [SPEAKER_04] But there is a limit to it. [SPEAKER_04] Like, for example, if you're changing copy from one string to another string, and the user says, just make a quick fix and update the test, maybe you don't want to verify. [SPEAKER_04] And so we've also adjusted our wording from saying always verify, verify, verify, verify, to hey, most of the time when you're doing front-end work, you can't always understand the full experience by hitting the back-end endpoints. [SPEAKER_04] So when you make changes to the user experience, please run the app locally. [SPEAKER_00] And actually, in fact, that instruction probably isn't even good because what is a large change? [SPEAKER_00] Maybe it wants to change it, test it for small changes too. [SPEAKER_04] In general, whenever you give a prompt to the model, you should always think about the ways in which it could be misinterpreted by a well-intentioned other user or human, in order to better understand how the model might interpret it, and in order to make sure that you can soften the prompt such that it is actually 100% accurate, because you are giving this prompt to the model 100% of the time. But what's fascinating about that is you're relying on the model's judgment. [SPEAKER_04] And that's got to be an Opus Fable level thing. [SPEAKER_04] Models a year ago did not have the levels of judgment necessary to decide if they were going to test a change or not. [SPEAKER_04] That's absolutely fascinating. But that does break down if you're building for a wide range of models and trying to earn the cheaper models for cheaper tasks. [SPEAKER_04] We actually have a different system prompt for each model now, because of this very reason. So it's only our most frontier models that have this 80% token decrease, and the older models actually still have the full system prompt. Do you think Fable and Opus are smart enough to be able to prompt Haiku with more details, because they understand that Haiku has less judgment, has less taste? We haven't been able to eval this very precisely. I think it should be able to, but we don't have any hard data to show it. [SPEAKER_04] I think there's a tough thing with smaller models sometimes, because sometimes the larger models can be more token efficient on a hard problem than the smaller models. [SPEAKER_02] And so there's a little bit of that intuition to build about sometimes you really just want frontier intelligence almost all the time. But the Pareto curve shifts, and so it's hard to find. [SPEAKER_02] I mean, that's something I found fascinating. A year ago, I did not trust a model to write a prompt. [SPEAKER_02] Today, the good models are very good at prompting. A lot of my prompts are written by models, which feels absurd, but it actually works really well. And something that helped me come to terms with that was thinking about sub agents, which is entirely about a Claude model setting up a prompt for another Claude model so that it knows what to go and do. And so there's a little bit of that intuition to build about how sometimes you really just want frontier intelligence almost all the time, but the Pareto curve shifts, and so it's hard to find. I mean, that's something I found fascinating. A year ago, I did not trust a model to write a prompt. Today, the good models are very good at prompting. A lot of my prompts are written by models, which feels absurd, but it actually works really well. And something that helped me come to terms with that was thinking about sub agents, which is entirely about a Claude model setting up a prompt for another Claude model so that it knows what to go and do. Yeah, I think workflows are actually a really good example of this, because it's Claude not just prompting a single sub agent, but prompting the orchestration of many things. So yeah, it's a lot of these sub agents and each one of them gets a very detailed prompt. So it's almost like a level above just spawning a sub agent. So yeah, it's quite good at that. I've also been using on my personal machine, giving it the Gemini API and being like, oh, here, generate images. And it's so good. It's way less lazy than I am at prompting an image model, so yeah, it's just Claude prompting Claude all the way down. I think Claude also wrote the prompt for the workflow tool. For the workflow tool. I've read that prompt. It's a good prompt. [SPEAKER_02] I mean, that's actually a frustration I have with Anthropic generally. You publish the prompts for Claude. There's a web page with them on, but you don't include the tool prompts and the Claude code prompts. [SPEAKER_00] I still have to run a proxy to intercept them. I would love it if the Claude code prompts were deliberately published, because they're the documentation. They're how you know what the tool can do and how it works. [SPEAKER_00] I'll write down that feature. Please do. [SPEAKER_00] I'll have Claude tag do it. And also the diffs. Every now and then I'll diff the older and the newer prompt. And that's how I learn the capabilities of the new model. I'm really looking forward to seeing what this 80% reduction actually looks like. [SPEAKER_00] Yeah, this is on me. I have to make a post about this in detail. So what's your bar? Let's talk about tools a little bit. Claude code is basically a big bag of tools. What's your bar for introducing a new tool? How do you decide when it's worth doing that additional engineering at that level? Do you want to take it? Because you introduced one of the best tools we have. Yeah, my career peaked when I introduced the ask user question tool, I think. It's really hard, especially for some tools like ask user question is Claude's tool to ask you. And so it's hard to eval that and sometimes more of a user preference thing. So especially back then we had fewer evals. It was very dog fooding based. But yeah, I mean, I think overall we've been trying to trend towards fewer tools. [SPEAKER_02] I think the last set of tools we introduced were the task tool, and try to give Claude more general versions to do this. [SPEAKER_02] Right, because one of the most interesting tools is the file editing tool. [SPEAKER_02] Which, but you can have file editing as a tool or you can teach it to use sed and grep and do things that way. [SPEAKER_02] What's the latest evolution of your file editing tool? [SPEAKER_02] I think we still have one. But for example, we removed our grep and other search tools and glob tools for just native bash. [SPEAKER_02] And so yeah, we still have one. I think this is kind of what I said in my talk earlier, that the models are more of a biology than a physics. And so it's hard to, especially tool design, I think is quite hard. [SPEAKER_02] And I'm not sure if actually Kat disagrees and is like, oh, there's a science to the evals of it. But yeah, tool design is more of an art or a biology. [SPEAKER_02] I think I largely agree. But in general, as we introduce more tools, we try to keep the cardinality pretty low and make sure that every tool we add has a distinct function from every other tool. So that Claude can very easily distinguish when to call each. For file edit, actually, the reason that we have file edit is because we can render it because we used to show, or I guess we still do. We show people when Claude makes a file change and there's this nice dedicated UI that just says, do you approve this edit to this file? And the reason that we had a dedicated file edit tool was so that we could deterministically know that Claude was making a file so we could show people this nice UI. And for a lot of the new users who are onboarding, I think they still really like this experience. So we've kept it around. But for a lot of us who are on auto mode right now, or hopefully you're not on YOLO mode, but anyway. [SPEAKER_00] For a lot of us right now, I don't think it matters and we probably could just remove file edit and we'll be totally fine. [SPEAKER_00] So let's talk about auto mode. Or let's talk about safety and security in general. I am deeply aware of the risks of prompt injection and there are so many bad things that can happen if somebody else tells my Claude code what to do. [SPEAKER_02] I still mostly run Claude code in YOLO mode and feel incredibly guilty about it. [SPEAKER_02] What's the advice within Anthropic for safely running Claude code? Like what do you tell people to do? [SPEAKER_02] Why not auto mode? [SPEAKER_02] I am starting to use auto mode and I don't understand it enough to get how safe it is. But yeah, as of maybe three weeks ago, I'm defaulting to auto mode. [SPEAKER_02] Okay, so broadly within Anthropic, almost every single person uses auto mode. It is the best way to do long running work in Claude code while being safe. We've done extensive testing. We have thousands of evals. We've commissioned many red teamers to create adversarial environments in order to trick Claude code into doing bad actions. And we've mitigated every single issue that they've found. And so we're going to publish some evals in the coming weeks. [SPEAKER_00] But we've pretty much mitigated every attack. [SPEAKER_00] That is a big claim. That's very exciting if that holds up. [SPEAKER_00] We will share the evals for it so folks can assess. But we've been extremely diligent about identifying all of the ways in which Claude might mess up. And then updating auto mode to counter it. It doesn't catch a hundred percent of things. I don't think any would be way too strong of a claim. But for the main categories of risks that we're concerned about, like prompt injection, data exfiltration, [SPEAKER_02] And so we're going to publish some evals in the coming weeks. [SPEAKER_00] But we've pretty much mitigated every attack. [SPEAKER_00] That is a big claim. That's very exciting if that holds up. [SPEAKER_00] We will share the evals for it so folks can assess. [SPEAKER_00] But we've been extremely diligent about identifying all of the ways in which Claude might mess up. [SPEAKER_00] And then updating auto mode to counter it. [SPEAKER_04] It doesn't catch a hundred percent of things. [SPEAKER_04] I don't think any... Yeah, that would be way too strong of a claim. [SPEAKER_04] But for the main categories of risks that we're concerned about, like prompt injection, data exfiltration, the risks are far lower than the average human reviewer. [SPEAKER_04] So... [SPEAKER_04] A little bit on how auto mode works. I think it's useful to build this mental model. [SPEAKER_00] So whenever Claude is doing a turn, or a bash call, there's a sonic classifier that is judging the tool and also the context of the conversation, your instruction. [SPEAKER_00] Right? And so there are some things around like permissions which are dependent on your request. [SPEAKER_00] Right? So you don't want to give git push permissions all the time. [SPEAKER_00] But if you say, hey, push this to GitHub, you want it to do it. [SPEAKER_04] Right? And so auto mode will... Or if you say don't push, you want it to deny it. [SPEAKER_04] Right? And so auto mode will do... That particular thing happens to me a lot. [SPEAKER_04] Where it's auto mode... Claude tried to do this because it's very helpful and proactive. [SPEAKER_04] And auto mode saw, don't do this and it surfaced it. [SPEAKER_04] So it's good at the dynamic permissions that you yourself give inside of the prompt, which I think is really important. [SPEAKER_04] It also works well with our sandboxing infrastructure because sandboxing is one of those things where there are so many different edge cases. [SPEAKER_04] And it's hard for us to deterministically follow them. [SPEAKER_04] But if you have a sandbox and something needs to escape the sandbox, a network request, auto mode can then look at that request and be, does this make sense? Right? And just allow that in. [SPEAKER_04] So I hadn't realized auto mode is interacting with the networking sandbox as well. [SPEAKER_04] Yeah, exactly. So it's also part of sandbox. Yeah. It interacts with any permission prompt that the user would otherwise see. And how old is auto mode? Like, I feel like as a feature that I had access to, it's only a couple of months old, right? We've been using it within Anthropic since January. Okay. [SPEAKER_02] So we've been hardening it for quite a while. [SPEAKER_02] And it's obviously Anthropic is extremely focused on safety and security. [SPEAKER_02] And so we've been working broadly across our alignment and safeguards teams in order to enable the rollout internally, build out these evals, make auto mode even more robust before sharing it out with the world. [SPEAKER_02] I think my only main problem with auto mode is I don't understand it deeply enough. [SPEAKER_02] Like for anything that's looking after my security, I want to know as much as I can about how it works and what it protects me against and what it doesn't. So I can decide how much I can trust it. I think, yeah, Del was working on a post about this. So a little bit more about auto mode. This is also the reason Claude Tag is so good, right? Because Claude Tag uses auto mode. And like you can imagine that, I've heard a lot of build versus buy questions on a Slack bot. I'm like, please, you probably shouldn't build your own AI Slack bot. There's so many attack vectors. [SPEAKER_00] You know? [SPEAKER_00] And like you have a feedback channel that users can post feedback into it. [SPEAKER_00] Now your bot is reading it, right? [SPEAKER_00] And so I think that this, the work we put in with auto mode and we have a general Swiss cheese defense of security, right? [SPEAKER_00] We also, yeah, RL against this stuff and things like that. I think this is really what makes Claude Tag work. It just works seamlessly with your permissions. [SPEAKER_00] And yeah, we don't want to be prompt injected in your Slack. [SPEAKER_02] Do you have any, are there any more security things in the pipeline beyond auto mode? [SPEAKER_04] I think we're very secure. [SPEAKER_04] So with Claude Tag, you can permit, can provision your own credentials for Claude. [SPEAKER_00] So it doesn't need to act on your behalf. [SPEAKER_00] You can have Claude as an identity. [SPEAKER_00] And that also makes it easier to audit and inspect what Claude is doing. [SPEAKER_00] Well, I guess because Claude Tag is influenced by anyone who can talk to it. [SPEAKER_00] So it's got a much wider pool of people who are telling you what to do. [SPEAKER_04] That's right. Yeah. [SPEAKER_04] And of course we have probes as well, like with Mythos and, sorry, with Fable. [SPEAKER_04] And that's also a downstream effect of our safety and research work. [SPEAKER_04] And I think this is the moment where you sort of see AI, like Anthropic being an AI safety company really paying off when you, we really want Claude to be able to run in an aligned way over long periods of time. [SPEAKER_04] And yeah, Automode has to be basically flawless for this to work. [SPEAKER_04] Right. [SPEAKER_04] And it's sort of all downstream of our being an AI safety company. [SPEAKER_04] We also launched trusted devices for the remote control users out here who want to be safer. [SPEAKER_02] And for all of our remote environments, we support credential injection. So if you want Claude Code to be able to access Datadog, but you don't want Claude Code itself to hold the Datadog credential, you can set up our identity credential management system. So that the Datadog credentials are only usable by the agent, but not accessible by the agent. [SPEAKER_00] So we insert it on the fly when the agent tries to make a Datadog request. [SPEAKER_00] This is that token, that proxying trick, right? [SPEAKER_00] Where the proxy knows anytime somebody calls api.datadog.com address with a token, replace the token with the real thing. [SPEAKER_00] I love that pattern. [SPEAKER_00] I'm seeing that in a whole bunch of places. [SPEAKER_00] It feels so obviously right to me. [SPEAKER_02] Let's talk a little bit about the human element and you touched on this in the keynote this morning, but a lot of people are feeling a sense of loss. Now that so much of what they considered to be their role in building software is being subsumed by the models. [SPEAKER_02] How do you think about that? This is that token, that the proxying trick, right? Where the proxy knows anytime somebody calls this an API.datadog.com address with a token, replace the token with the real thing. [SPEAKER_00] I love that pattern. [SPEAKER_00] I'm seeing that in a whole bunch of places. [SPEAKER_00] It feels so obviously right to me. [SPEAKER_02] Let's talk a little bit about the human element and you touched on this in the keynote this morning, but a lot of people are feeling a sense of loss. [SPEAKER_00] Now that so much of what they considered to be their role in building software is being subsumed by the models. [SPEAKER_02] How do you think about that? [SPEAKER_00] Firstly, how has the past year and a half changed the way you think about your own craft and the value that you add? [SPEAKER_00] Yeah, I think for me and I think Cat is always such a good writer. [SPEAKER_00] Cat and Boris are such good reminders of having to be more ambitious. [SPEAKER_00] They're always saying you have to be ambitious, we're growing so fast. [SPEAKER_00] We have to be on the edge. [SPEAKER_00] We have to do the best work we can. [SPEAKER_00] I think that's a constant reminder for me where anytime I'm slow on something, I'm like, okay, can I do it faster? [SPEAKER_00] Can I be more ambitious here? [SPEAKER_00] I think the answer oftentimes is Claude because Claude is getting better as you go. [SPEAKER_00] So I'm like, oh, the last time I tried this, I was with the previous model or something. [SPEAKER_00] I think with their point on loss, I think this is real. And I do feel that if you're only trying to do the same work you were doing before LLMs and now it's a prompt, it is a sad feeling. And I think the way you offset that is by being more ambitious, right? [SPEAKER_04] I love how Jared is such a good example where he hand wrote all of the Zig code in his Oakland apartment, barely left his house. [SPEAKER_04] And then he had so much fun doing that. [SPEAKER_04] And now he's rewriting all of Bon into Rust and he's having so much fun doing that, right? [SPEAKER_04] And it's so much more ambitious. [SPEAKER_04] And that's how he offsets that. [SPEAKER_04] And I think just generally being like, okay, how do I do the bigger thing? [SPEAKER_04] And do more. [SPEAKER_04] And I think success is fun. [SPEAKER_04] And that's how I think about it. [SPEAKER_00] Right. Like it's changing your ambition. It's changing what you do because what you did before is a lot easier in quotes, but now we can take on these bigger challenges. Yeah. [SPEAKER_04] I think there's just on average, everyone has things they wish they did, and they were better at. [SPEAKER_04] And I think now it's like, let's do it. [SPEAKER_04] Kat, what does that look like from a product management perspective? [SPEAKER_04] I feel like the product role just changes every single month. [SPEAKER_04] And it's very much just identifying, okay, what are the priorities? All the PMs on our team are a mix of engineer, designer, PM. [SPEAKER_04] Most of the engineers on our team actually used to be full-time engineers in the past. [SPEAKER_04] And so for us, it really means plugging in whenever there's any kind of gap. [SPEAKER_02] So if it's like, okay, we have this idea and we didn't inspire any engineer to go build it, then we should just build it and then put it into a notebook and inspire people to take this to production. [SPEAKER_02] Or if the designs look a little off, let's take a page that's similar and do a first pass design and tag in someone who's very detail oriented to fill in the gaps. [SPEAKER_02] Or if we notice that our team and our product adoption is a bit bigger within the company and more people need to know what's coming down the pipe for quad code, quad tag and co-work. [SPEAKER_02] What we do then is okay, let us automate figuring out our whole launch calendar. [SPEAKER_02] Let's automate getting those status updates asynchronously. [SPEAKER_02] So we're not bugging people. [SPEAKER_02] And then let's figure out, okay, these are our three internal announced channels and make sure that our updates there are fully detailed and to the point. [SPEAKER_02] And so for us, it's very much just understanding what is the gap right now between a great idea and getting something to our customers. [SPEAKER_02] And then how do we automate it as much as possible? [SPEAKER_00] It sounds to me like with product management, there's always more to do, right? [SPEAKER_00] I feel like one of the things that makes me feel good is I've never worked at a company that didn't have a backlog of a thousand things they wanted to do and didn't have the resources to take on. [SPEAKER_00] So what's a moment when Claude has surprised you? [SPEAKER_00] When the model has done something that genuinely surprised you, you didn't think it would be able to do? [SPEAKER_00] Yeah. [SPEAKER_00] I mean, I posted a lot about Claude video editing, but most recently I gave a talk at the ACM Agentec conference and I was like, hey guys, do you have the edited video? I'd love to post it and share with my comps team. And they're like, oh, it's taking so long. And I'm like, okay, could you send me the raw files? [SPEAKER_02] So they send me the video of me talking on stage, the video of the deck and the audio file. [SPEAKER_02] And they're like, good luck. [SPEAKER_02] And so I give this to Claude, I give it my HTML deck as well. [SPEAKER_02] And I'm like, hey, can you just edit this together? [SPEAKER_02] And what it does is honestly incredible. [SPEAKER_02] I'm ready to ship it. It transcribes the entire video. It notices that sometimes the video of my deck is a little bit weird. [SPEAKER_00] There's a popup of an auto update in the middle. [SPEAKER_02] And it's like, oh, I probably shouldn't use the video of your deck. [SPEAKER_02] Actually, what I'm going to do is I'm going to slice up and figure out which slide you're on. [SPEAKER_02] And instead use your HTML source. [SPEAKER_02] Right. [SPEAKER_02] And so it's displaying the HTML source. [SPEAKER_02] Then it's got a video of me, but I'm only taking up a small part of the stage. [SPEAKER_02] And so it's cropping dynamically where I am on the stage. [SPEAKER_04] I'm pacing. [SPEAKER_04] So it's tracking me as I'm pacing and I've got a crop of me and the deck. [SPEAKER_02] And I probably shouldn't use the video of your deck. [SPEAKER_02] Actually, what I'm going to do is I'm going to slice up and figure out which slide you're on. And instead use your HTML source. Right. [SPEAKER_02] And so it's displaying the HTML source. [SPEAKER_02] Then it's got a video of me, but I'm only taking up a small part of the stage. [SPEAKER_02] And so it's cropping dynamically where I am on the stage. And I'm pacing. So it's tracking me as I'm pacing and I've got this crop of me, the deck. And then it's transcribing what I'm saying. This was Fable, right? This is Fable. Yeah, yeah, yeah. [SPEAKER_04] Absolutely. [SPEAKER_04] Yeah. [SPEAKER_04] And it was a good prompt, but it was a one shot prompt. [SPEAKER_04] And then I asked it to add some interesting animations and graphics. And I was blown away, and it just does all this stuff. It does FFmpeg, it does Remotion. It does. Okay. Yeah. I have to ask the followup. What can't it do? What are the things where you're still disappointed and you're waiting for Claude Fable 6 to figure out for you? I want it to have better design and UX taste. [SPEAKER_04] Uh huh. [SPEAKER_04] Like, I feel like it's at the point where if I give it a prompt with a detailed spec of how I want a feature to behave, it will usually behave that way. [SPEAKER_04] But the paddings might be off or the interface is just not delightful yet. [SPEAKER_04] I think it leans on existing best practices for apps, for how apps are designed. [SPEAKER_04] But I feel like for frontier AI products, there's so many new interaction experiences that we still have to design. [SPEAKER_00] I feel like there's an Opus aesthetic. [SPEAKER_00] You can look at something and go, yeah, that was created by Opus. [SPEAKER_00] It would be good if we could move beyond that. Yeah, yeah. [SPEAKER_00] Like, I'm very excited for future models to hopefully be interaction design thought partners. Hmm. [SPEAKER_02] What can't it do? I think I would love to see it interact more with the real world. [SPEAKER_02] Like, can it do this? [SPEAKER_02] Can it solve science, right? [SPEAKER_02] Can it orchestrate the experiments? And there's some amount of coding that goes into that. But there's also this other aspect of, the broader world that it needs. Claude Science is a new product that just came out a few days ago, right? [SPEAKER_04] Yeah, but I have no context on it. [SPEAKER_04] I was going to ask, is that part of Claude Code or is that a separate action? [SPEAKER_04] It's our partner team. [SPEAKER_04] Gotcha. [SPEAKER_04] Try it out though. [SPEAKER_04] So we've got a couple of closing questions. [SPEAKER_04] Which parts of Anthropic company culture do you think uniquely help Anthropic be productive with these tools that other companies should steal? [SPEAKER_04] What are the cultural hacks that people should be adopting from you? [SPEAKER_04] I'll share one and then you go. [SPEAKER_04] I'll share one for Claude Tag. [SPEAKER_04] So Claude Tag works best when you have it in a public channel and when most of your channels are public. [SPEAKER_04] Claude Tag is able to search across all public channels to get as much context as possible to give you the highest accuracy answer. [SPEAKER_04] And it's only able to do this if it has access to everything. [SPEAKER_04] Yeah, I mentioned this in my keynote, but I think it's so important to me. I want to reemphasize like, I think co-founders say we don't negotiate against ourselves, right? [SPEAKER_00] And I think this is really important where you can imagine trade-offs in your head and talk yourself out of doing something ambitious. [SPEAKER_04] Or you can just try and do the ambitious thing. And I think we're just so often being like, okay, what if we just did it? What if we just did it? Is this a real trade-off or not? Or if so, why? Where's the proof that it's a real trade-off and not just something that sounds reasonable? [SPEAKER_00] Right? And so I think, yeah, make the trade-offs show themselves to you, be as ambitious as you can. [SPEAKER_00] That's tough, cause that goes against 25 years of software experience that says the default answer should be no. Everything is a trade-off, everything has a cost. And now we're having to reimagine all of those intuitions. It's kind of fascinating. Okay. Final question for both of you. What is something, what's one of your favorite absurd things that you've built with Claude just because you could build it? I can go while you think. I'm working on a 2D Street Fighter fighting game with me as a character and my friends as well. And it uses Claude Code to prompt Gemini and honestly the Stable Diffusion model is pretty good to make video animations. [SPEAKER_02] And it works great. It's so good at prompting. It can verify the frames to check if this is a good animation. [SPEAKER_02] Is this Street Fighter 2 level 2D sprites that you're generating? [SPEAKER_02] Yeah, exactly. Yeah, yeah, yeah. 2D sprites. The animation looks amazing. And it can also figure out hitboxes. It can be like, oh, your fist is here, I'll draw the JSON hitbox. [SPEAKER_02] Yeah, yeah, yeah. It's incredible. Yeah, yeah. So I don't know if I'll put this out, but it's... [SPEAKER_02] I feel like we need a screenshot at least. This sounds amazing. [SPEAKER_00] Sure, yeah, yeah. I can make a screenshot happen. [SPEAKER_00] Mine is much more simple. I'm a big rock climber and a lot of my friends climb. And so we have this little app that we built with Claude Code where we just log all the projects that we're working on. [SPEAKER_00] And we also go outdoors together a lot. So we have Claude do all this research with Workflows. Workflows is amazing. We brand it as a coding tool, but it's amazing for doing deep research for travel. [SPEAKER_02] Yeah, yeah, yeah. It's incredible. Yeah, yeah. So I don't know if I'll put this out, but it's... [SPEAKER_02] I feel like we need a screenshot at least. This sounds amazing. Sure, yeah. I can make a screenshot happen. Yeah. Mine is much more simple. I'm a big rock climber and a lot of my friends climb. And so we have this little app that we built with Claude code where we just log all the projects that we're working on. And we also go outdoors together a lot. So we have Claude do all this research with Workflows. Workflows is amazing. We brand it as a coding tool, but it's amazing for doing deep research for travel. I also plan our team off sites. And it's good at finding venues that can fit all of us. Yeah, it has a lot of side benefits. But anyway, I also use Workflows to research all the climbing destinations that we might want to go to, what has direct flights from where all of us are located. [SPEAKER_00] It goes to Mountain Project and finds all the climbs that are in our grade level. It finds the Airbnb and it actually maps out... I don't hiking. And so I care a lot about it having a very short approach. So very short walking distance from where the car parks to where the rock actually is. And so it filters for this. So with existing apps, I have to manually click through Mountain Project. But with this, I just put in all of our preferences and it's just a custom app for us. [SPEAKER_00] So you're basically vibe coding JIRA for mountain climbing. Exactly. That's pretty fantastic. You know, we have time for a couple of audience questions. If you want to come forward and say this to me and I will repeat them so everyone can hear them. [SPEAKER_04] But yeah, I'll tell you what, anyone who gets here first gets to ask a question. Sorry for people at the back. [SPEAKER_04] Hey. [SPEAKER_04] Yeah. My question is, do you have a near plan to build more eval tools for us to build eval data sets or anything like that? [SPEAKER_04] And more ability tools to monitor the performance of agents and workflows? [SPEAKER_04] We've considered building eval tools, but I think the limiting factor actually tends to be that it takes a long time for customers to build really high quality evals. [SPEAKER_04] And so I think the tooling is less of the constraint and more of the skill set of how do you build a great eval? [SPEAKER_04] And that's an area where we're excited to both invest internally and also hopefully we can share some of the best practices externally. Hey, Kat. Hi there. My name is Sai. So my question was because I'm more interested in the memory and the multiplayer. How does memory being designed? So two questions, right? So how is memory being designed today? I assume it's around files. And second part of it is, have you thought about thinking in an orthogonal direction where you would actually need a data store to store these memories in the files to scale it better? So that's my question. Yeah, right now for Cloud Tag, the memory is channel specific. So every Cloud in that channel has a shared memory and then the instances have a session, but the session can contribute back to main memory. We do a lot of memory research and it's unintuitive. What is the right way to do memory? But yeah, we're always working on this. Yeah. Yeah. I mean, we're always running memory experiments. I don't have anything to share. How it works right now in Cloud Tag is a markdown file per channel. Yeah. Okay. Thank you. So I'm afraid we are out of time. Please join me in thanking Kat and Tariq and we will be around for more questions in the hallway. Thanks guys. Thank you. Thank you. I'll share one and then you go. Um, I'll share one for Claude Tag. So Claude Tag works best when you have it in a public channel and when most of your channels are public. Claude Tag is able to search across all public channels to get as much context as possible to give you the highest accuracy answer. And it's only able to do this if it has access to everything. Yeah, I mentioned this in my keynote, but I think I, it's so important to me. I want to reemphasize like, I think that co-founders say like we don't negotiate against ourselves, you know? And I think this is really important where you're like, you can imagine trade-offs in your head and talk yourself out of doing something ambitious, you know? Um, or you can just try and do the ambitious thing. And I think that like, we're just so often being like, okay, what if we just did it? Like, what if like, you know, like, is this a real trade-off or not? Right? Or like, and if so, like why? Like where's the proof that it's a real trade-off and not just like, it sounds reasonable. Right? And so I think just, yeah, like, you know, make the trade-offs show themselves to you, be as ambitious as you can. That's so, cause that goes against, I've got 25 years of software experience that says the default answer should be no. Everything is a trade-off, everything has a cost. And now we're having to reimagine all of those intuitions. It's kind of fascinating. Okay. And so final question for both of you. What is something, what's one of your favorite absurd things that you've built with Claude just because you could build it? I can go well, you think. I'm working on a 2D Street Fighter fighting game with me as a character and like my friends as well. And it uses Claude code to prompt, you know, Gemini and honestly the Sea Dance model is pretty good, like, to make like video animations. And it works great. Like, like, it's so good at prompting. It's like, you know, it can verify like the frames to check if this is a good animation. Is this Street Fighter 2 level 2D sprites that you're generating? Yeah, exactly. Yeah, yeah, yeah. Like, like 2D sprites. The animation looks amazing. And it can also figure out hitboxes. It can be like, oh, you know, your fist is like, here, I'll draw the JSON hitbox. Yeah, yeah, yeah. It's like, incredible. Yeah, yeah. So, I don't know if I'll put this out, but it's, uh... I feel like we need a screenshot at least. This sounds amazing. Sure, yeah, yeah. I can make a screenshot happen. Yeah, yeah. Mine's is much more simple. I'm a big rock climber and a lot of my friends climb. And so we have this little app that we built with Claude code that where we just log all the projects that we're working on. And we also go outdoors together a lot. So we have Claude do all this research with Workflows. Workflows is amazing. Like, we brand it as a coding tool, but it's amazing for doing deep research for travel. I also plan our team off sites. And it's good at finding venues that can fit all of us. Yeah, it has a lot of side benefits. But anyway, I also use Workflows to just research all the climbing destinations that we might want to go to, what has direct flights from, where all of us are located. It goes to Mountain Project and finds all the climbs that are in our grade level. It finds the Airbnb and it actually maps out... Like, I don't like hiking. And so I care a lot about it having a very short approach. So very short walking distance from where the car parks to where the rock actually is. And so it filters for this. And so it's like with existing apps, I have to like manually click through Mountain Project. But with this, I just put in all of our preferences and it's just a custom app for us. So you're basically vibe coding JIRA for mountain climbing. Exactly. That's pretty fantastic. You know, we have time for a couple of audience questions. If you want to come forward and say this to me and I will repeat them so everyone can hear them. But yeah, I'll tell you what, anyone who gets here first gets to ask a question. Sorry for people at the back. Hey. Yeah. My question is that, do you have a near plan to build more eval tools for us to build eval data set or anything like that? And more ability tools to monitor the performance of agents and workflows? We've considered building eval tools, but I think the limiting factor actually tends to be that it takes a long time for customers to build really high quality evals. And so I think the tooling is less of the constraints and more of the skill set of how do you build a great eval? And that's an area where we're excited to both invest internally and also hopefully we can share some of the best practices externally. Hey, Kat. Hi there. My name is Sai. So my question was because I'm more interested in the memory and the multiplayer. How does, how is memory being designed? So two questions, right? So how is memory being designed today? I assume it's around files. So, and second part of it is, have you thought about thinking in an orthogonal direction where you would actually need a data store to store these memories in the files to scale it better? So I think that's my question. Yeah, right now for Cloud Tag, the memory is channel specific. So every Cloud in that channel has a shared memory and then, you know, the instances have a session, but like the session can contribute back to main memory. We do a lot of memory research and it's, you know, can be kind of unintuitive. Like what, what is the right way to do memory? But yeah, we're always working on this. So, yeah. Yeah. I mean, we're always running a merit memory experiments. I don't have, you know, anything to, yeah. Like how it works right now in Cloud Tag is a markdown file per channel. Yeah. Okay. Thank you. So I'm afraid we are out of time. Please join me in thanking Kat and Tariq and we will be around for more questions in the hallway. Thanks guys. Thank you. Thank you.