How Anthropic Uses Claude Fable 5 With Mike Krieger
Description
Mike Krieger built one of the most consequential consumer apps of the last two decades as the cofounder of Instagram. He is now at the frontier of AI-native product development as head of Anthropic Labs, the team responsible for figuring out what the most capable AI models can do in the hands of real builders. When Krieger first got access to Fable 5 months before its public release, it was exciting and disorienting. “I feel like a total newbie again,” he remembers telling his team. The way he’d been thinking about productivity, strategy, and time management was out of date. The model had outpaced his workflows. Dan Shipper talked with Krieger for AI & I about what it looks like to build with a model as capable as Fable 5, including the new rhythms, challenges, and possibilities it reveals. If you found this episode interesting, please like, subscribe, comment, and share! To hear more from Dan Shipper: Subscribe to Every: https://every.to/subscribe Follow him on X: https://twitter.com/danshipper Get started with Braintrust at https://www.braintrust.dev/ Timestamps: 0:03 Introduction 1:48 How Fable completely reshaped Mike's workflow 4:48 When to use Sonnet versus Fable 10:06 What the media tracker Mike built over a weekend reveals about agent-native architecture 15:00 The cost to build has collapsed 19:03 Is software engineering over? 21:48 How Anthropic's engineering teams work today 38:39 The mechanics of verification 44:39 What people should use the model to build 47:24 Dynamic workflows Links to resources mentioned in the episode: Mike Krieger on X: https://x.com/mikeyk Anthropic Labs: https://www.anthropic.com Claude Code: https://claude.ai/code Every: https://every.to
Summary
Generated by claude-sonnet-4-5At-a-Glance
- Verdict: Watch fully
- Core thesis: Claude Fable 5 represents a fundamental shift from "code assistant" to "trusted teammate," enabling multi-day autonomous work with unprecedented follow-through, judgment, and system-level thinking.
- Why it matters: This isn't an incremental improvement—it changes what non-engineers can build, how software teams operate, and the economics of software development. Krieger (Instagram co-founder, now Anthropic Labs head) shows real production usage patterns, not launch-day hype.
- Best use: Watch for concrete workflow examples, verification strategies, cost/benefit analysis for production deployment, and the emerging "agent-native architecture" pattern where software modifies itself.
Executive Summary
Mike Krieger provides a rare insider view of Claude Fable 5 after months of heavy production use at Anthropic. The headline: this model fundamentally changes the unit of work from "write this function" to "complete this multi-day project." Krieger regularly assigns complex tasks overnight—like porting entire codebases from Python to TypeScript—and wakes to completed work including documentation of trade-offs made and problems encountered.
The key capability isn't speed or raw intelligence, but system-level judgment. Fable understands project architecture, pushes back on code review feedback when appropriate, tracks feature flags across sessions, and knows when to scaffold a mock backend versus demanding production infrastructure. This isn't better autocomplete—it's a model that acts like a senior engineer who thinks about the whole system.
Krieger's most striking example: building a personal media tracker where the app itself can be modified through conversational requests. Long-press the chat button, ask for changes, get a live preview with diffs. This "agent-native architecture" pattern—where AI can both use and modify the software—points to a future where the boundary between user and developer collapses.
The cost/workflow trade-offs are real: Fable is expensive and slow for quick queries (Krieger switched his iOS app back to Sonnet for NBA questions). But for substantial work, the economics flip: it's often cheaper to let Fable think deeply once than iterate 10 times with faster models. Verification becomes critical—Krieger emphasizes comprehensive screenshot galleries, video analysis, and architectural alignment conversations before execution.
Key Takeaways
- Claim: Fable enables true multi-day autonomous work with remarkable follow-through. Evidence: Krieger routinely delegates complex overnight tasks; model documents decisions, scaffolds workarounds when services fail, tracks incomplete work like unset feature flags across days-long sessions. Caveat: Requires upfront architectural planning and clear verification strategies. Implication: Software development shifts from "writing code" to "directing work" and "verifying outcomes."
- Claim: The model demonstrates unprecedented system-level judgment and taste. Evidence: Pushes back on code review feedback when appropriate, distinguishes between "restart the server now" versus "fix the memory leak architecturally," understands deployment complexity appropriate to project stage. Caveat: Still needs human accountability—engineers must understand trade-offs even if they didn't write the code. Implication: The skill of "AI-assisted engineering" becomes about expressing intent, making architectural decisions, and verification—not syntax.
- Claim: Cost per task drops dramatically despite higher per-token pricing. Evidence: Krieger's personal media tracker (production-quality full-stack app) built over a weekend fits within reasonable personal budget; Fable's one-shot completion often cheaper than 10 rounds with faster models. Caveat: Quick queries feel wasteful ("rocket launcher to kill a mosquito"); need model switching strategy. Implication: Economic model shifts from "minimize tokens" to "maximize successful task completion."
- Claim: Non-engineers can now build and maintain genuinely complex software over months. Evidence: Anthropic recruiting team member built comprehensive internal tool integrated with multiple systems—"first time in my life the thing in my head and the thing in the world are next to each other"; GTM employee maintains evolving tool for months without engineering help. Caveat: Still requires clear thinking about requirements and verification. Implication: Software development capacity no longer bottlenecked by engineer headcount; PM/engineer boundary dissolving.
- Claim: Dynamic Workflows + Fable = production-grade autonomous agent capability. Evidence: Complete Python-to-TypeScript codebase port executed as multi-step workflow over weekend; included deep understanding phase, module-by-module translation, incremental testing, adversarial verification, documentation of un-portable components. Caveat: Workflow design itself requires thoughtful planning. Implication: Long-horizon agentic work graduates from research demo to production tool; proper scaffolding and verification critical.
- Claim: Verification must evolve beyond code review to comprehensive behavioral testing. Evidence: Anthropic invests heavily in real staging environments with actual data, comprehensive screenshot/video capture for every PR, mock backends that stay in sync automatically, adversarial testing layers. Caveat: Requires infrastructure investment; video-based verification underexplored. Implication: Testing infrastructure and verification methodology become the critical bottleneck, not coding capacity.
- Claim: "Agent-native architecture" where software can modify itself is now practical. Evidence: Krieger's demo app with long-press-to-edit workflow; changes previewed live via Vercel, deployed to phone in real-time; app improves through normal usage. Caveat: Verification and rollback strategies essential; not appropriate for all software. Implication: New design pattern emerges: every app should expose full agent API and agent-driven modification capability.
Detailed Brief
Production Usage Patterns & Workflow Evolution
Claims:
- Fable changes the fundamental unit of work from turn-by-turn interaction to multi-day project delegation
- Engineers now maintain "dashboards" of multiple concurrent Claude sessions rather than single threads
- The model maintains context and follows through across days, tracking incomplete work like feature flags
- Architectural planning conversations happen upfront, then execution is delegated with high trust
Evidence:
Krieger describes wishing Claude "good night" and delegating complex tasks, waking to completed work—usually done by 2 AM. Model writes documentation of trade-offs, scaffolds temporary backends when services are down, tracks dependencies. He maintains 5-6 concurrent tabs for parallel workstreams versus one long-running high-context session. Instagram co-founder perspective: building Instagram v1 took 5 days of all-nighters; comparable complexity now fits in a weekend with intermittent attention.
Caveats:
Not pain-free transition—engineers who loved "dreaming about code" feel loss. Requires new skills: expressing intent, architectural thinking, verification discipline. Remote work on planes depends on stable connections and good context instructions. Need to explicitly manage model effort levels (medium vs. high) to avoid over-engineering simple tasks.
Implications:
Software engineering role shifts to: system design, context-holding across projects, verification/testing, incident response, team coordination. "Directly Responsible Individual" (DRI) ownership model persists because humans hold strategic context beyond what models access. The craft of implementation largely automated; craft of "what to build and why" becomes central.
Cost Economics & Model Selection Strategy
Claims:
- Per-token cost higher than Opus but per-completed-task cost often lower
- Quick queries feel wasteful; need intelligent model routing
- Professional software development economics clearly justify Fable-class models
- Personal/hobbyist usage requires more thoughtful budgeting
Evidence:
Krieger switched iOS app back to Sonnet for quick questions ("this is not Fable-worthy"), noticed "order of magnitude" speed/cost difference. But for substantial work, Fable's one-shot completion beats 9-10 iteration rounds with cheaper models. Personal media tracker project (full-stack production app) built over weekend fit within reasonable personal account budget. Anthropic's internal evolution: from "encouraging any usage" to "leaderboards" (bad incentives) to "demonstrate results, spend freely" (current).
Caveats:
Pricing process complex; economics vary by use case. Need to balance cost-consciousness with not over-rotating toward false economy. Still requires upfront architectural conversation (not free). Model selection overhead (which model for which task) becomes new cognitive burden.
Implications:
Business model shift: companies should measure "cost to satisfactory completion" not "cost per token." Tool/interface design must handle model routing intelligently—possibly per-surface sticky defaults (mobile = Sonnet, overnight batch = Fable). Productivity gains likely justify premium pricing for professional use; accessibility question remains for students/hobbyists.
Verification & Code Review in the Fable Era
Claims:
- Verification becomes the critical bottleneck, not coding capacity
- Comprehensive visual testing (screenshots, video) essential for every PR
- Model can respond to code review with judgment, not just compliance
- "Tight dev loop" principle still applies but looks different
Evidence:
Every PR generates screenshot galleries; video analysis catches animation jank invisible in stills. Model can be given FFmpeg and will scrub through video identifying issues. Code review interactions show discernment: model pushes back on feedback when appropriate ("I thought about it and I still disagree"), distinguishes immediate fixes from long-term refactoring. Anthropic invested in staging environments with real data but special affordances (skip onboarding, feature flags for testing). Fable automatically maintains mock backends that stay in sync with upstream changes.
Caveats:
Engineers must still "stand behind the work"—need follow-up conversations understanding trade-offs even if they didn't write code. Verification infrastructure requires investment. Video-based verification underexplored. Some engineers have moments of "not entirely sure, will find out before we merge this PR."
Implications:
Testing infrastructure investment critical for teams adopting Fable. Verification skills become premium. Code review culture must evolve: "did the model understand the trade-offs?" matters more than "is the syntax good?" New design patterns: comprehensive test suites, visual regression testing, behavioral verification over code inspection.
The "Agent-Native Architecture" Pattern
Claims:
- Phase one: every feature accessible via tool calls (table stakes)
- Phase two: agent can modify the software from within itself
- This enables software that improves through normal usage
- Pattern generalizes beyond personal projects to product design
Evidence:
Krieger's media tracker demo: long-press chat button triggers managed agent session that edits codebase, generates Vercel live preview, shows diffs. "Too low floating action button on iOS" → agent fixes it, live reloads on phone. Every software should have "ask Claude to add things by URL" capability—never navigate menus again. User becomes collaborator in ongoing software evolution.
Caveats:
Not appropriate for all software (safety, security concerns). Verification/rollback essential. "Don't particularly care about code quality or long-term maintainability" for personal project—wouldn't apply to production systems serving millions. Requires proper permissioning, audit trails.
Implications:
New product design pattern: every app exposes agent API + self-modification capability. Distinction between "user" and "developer" blurs. Software becomes continuously malleable to user needs. Opens question: what does version control, deployment, rollback look like in self-modifying systems?
Dynamic Workflows: Orchestrating Multi-Day Agent Work
Claims:
- Dynamic Workflows provide scaffolding for complex, multi-phase autonomous work
- Fable's capabilities make workflows far more powerful than with previous models
- Workflows can be composed conversationally, expressed in code, executed with clean UI
- Enable tasks previously impossible: full codebase ports, comprehensive testing suites
Evidence:
Krieger's weekend Python-to-TypeScript port via workflow: deep understanding phase → spec creation → module-by-module translation → incremental testing → adversarial verification → documentation of non-portable components. Came back to fully functional port with better architecture in some areas, clear notes on what couldn't translate. GTM team member built comprehensive internal tool using workflows + MCPs, works on it for months, now deploying to whole org. Engineering prototype-driven debates: PM builds janky proof-of-concept to open conversation.
Caveats:
Workflow design requires thought—iterative process with Claude Code to get right. Future work: tune subtasks to appropriate model size/effort (not everything needs high-effort Fable thinking). Product question: too many choices, need graspable buckets or surface-specific defaults.
Implications:
Agentic AI graduates from research demo to production tool with proper scaffolding. Long-horizon work (days/weeks) now practical with right verification. Future: automatic optimization of workflow steps to appropriate model capability/cost. Chat interface becomes workflow composition interface.
What Changes (and Doesn't) for Software Teams
Claims:
- Alignment conversations, ownership, incident response remain human-critical
- Parallelism explodes: many Claudes per human, but humans still hold DRI roles
- PM/engineer boundary dissolving; prototypes win arguments in new ways
- Understanding production behavior still requires human expertise
Evidence:
Teams still organize around areas of ownership, have alignment meetings, make trade-off decisions. But each person now orchestrates multiple concurrent agent sessions. Engineers need "understanding how things work in production" skills—incident response, knowing when to restart server vs. fix memory leak architecturally. PMs now build working prototypes to settle debates ("code wins arguments" redeemed—non-coder can produce code). Meta-work: maintaining dashboards of agent progress, triaging which PRs need attention.
Caveats:
Not pain-free transition. Skills evolving: less syntax debugging, more architectural thinking and intent expression. Meetings still happen (try to minimize). Complexity ceiling for non-engineers raised but not infinite. Human-to-human coordination remains essential.
Implications:
Engineering role: DRI ownership + architectural decisions + verification + incident response + context-holding. Engineer headcount may not be bottleneck anymore—product thinking and verification become constraints. New meta-skills: orchestrating multiple agents, expressing clear intent, comprehensive verification strategies. Culture shift: accountability for work you directed but didn't implement.
Notable Concepts & Terms
- DRI (Directly Responsible Individual): Anthropic's term for ownership model that persists in agent era—someone must hold context and accountability even when Claude does implementation
- Agent-Native Architecture: Design pattern where software exposes both agent API for usage and agent-driven modification capability; enables self-improving systems
- Dynamic Workflows: Anthropic feature for multi-phase autonomous work; scaffolds complex tasks as code-expressed DAGs; particularly powerful with Fable's judgment
- Verification-First Workflow: Krieger's evolved practice—every PR must include comprehensive screenshots/video, behavioral testing, not just code review; becomes the critical bottleneck
- MCP (Model Context Protocol): Anthropic's system for giving Claude access to internal tools with proper permissioning; enables personalized software composition
- "Tight Dev Loop" (evolved): Classic principle of fast iteration now means "comprehensive verification infrastructure" rather than "fast compile times"
- Managed Agents: Anthropic's term for agent sessions that run autonomously with UI for monitoring progress; used in Dynamic Workflows and self-modifying software demo
- Effort Levels (Extended Thinking Controls): Feature Krieger mentions using more with Fable—explicit setting for how much the model should "think" (low/medium/high); expensive model makes this more salient
Operator Notes / Why Ken Should Care
Direct Relevance to Ken's Work:
- Agent Ops / Orchestration: The "many Claudes per human" pattern and dashboard maintenance Krieger describes is exactly the agent orchestration problem Ken's thinking about. Anthropic is living this internally—concrete learnings on managing parallel agent sessions, DRI models, verification workflows.
- Content Production: If non-engineers at Anthropic are building production tools they maintain for months, Ken's content team could be doing far more sophisticated automation/tooling than currently assumed. The GTM example (recruiting team building custom integrated systems) maps directly to Every's ops.
- Product Strategy / Every Products: The "agent-native architecture" pattern (software that modifies itself through usage) is a potential competitive moat for AI-native products. Could Spiral, Lex, or other Every products adopt this? Would unlock continuous user-driven customization.
- Cost/Benefit for Every's Operations: Krieger's economics are clarifying: yes Fable is expensive per-token, but cheap per completed task for substantial work. If Every team is iterating 10 times on prompts with Sonnet, might be cheaper/faster to use Fable once. Needs analysis of actual usage patterns.
- Investing Lens: This interview is a rare view of production usage at the company building the model. Key signals: non-engineers maintaining complex tools for months (capability plateau?), verification becoming the bottleneck (opportunity for tooling), "agent-native architecture" as emerging pattern (investable category?).
- GTM / Positioning: Krieger's framing of "software engineering isn't over, it's radically different" is useful positioning. The skill is now intent expression, architectural thinking, verification—not syntax. This reframes AI coding tools from "replacement" to "capability expansion."
Cautions:
- Krieger has Anthropic-internal tooling (MCPs, staging environments, verification infra) that external users don't. Gap between his experience and typical user larger than it appears.
- Instagram founder perspective means his baseline for "how long should building take" is skewed—most people didn't do 5-day all-nighters for v1.
- Personal use case (media tracker) explicitly deprioritizes maintainability. Different calculus for production systems serving customers.
Action Items to Consider:
- Experiment with Fable for substantial Every projects where iteration overhead is high
- Map current Every workflows to Krieger's "many concurrent agents" pattern—where would parallel sessions help?
- Investigate verification infrastructure investment—screenshot automation, video testing tools
- Explore "agent-native architecture" for one Every product as experiment (Spiral most obvious candidate?)
- Analyze Every's actual AI spend patterns: are we iterating too much with cheaper models when Fable one-shot would be more efficient?
Bottom Line: This isn't hype—it's a credible technical leader with production experience showing real workflow changes. The shift from "code assistant" to "project teammate" is legitimate. For Ken's world: (1) content/ops teams are underutilizing capability, (2) agent orchestration patterns are investable, (3) verification tooling is the next bottleneck, (4) "agent-native" is a product design pattern worth exploring.
Watch Map
- Timestamp note: Timestamps/chapters not available in transcript
- 0:00-5:00 (estimated): Introduction, context on Krieger's role and preview of what makes Fable different beyond day-one impressions
- 5:00-15:00: Evolution of usage over months—"wish Claude good night" workflows, architectural planning conversations, managing concurrent sessions, comparison to Instagram building experience
- 15:00-25:00: Cost/economics discussion, model selection strategy (Fable vs. Sonnet for different use cases), professional vs. personal usage considerations
- 25:00-40:00: Demo of personal media tracker app with "agent-native architecture"—long-press to modify software from within itself, discussion of closing gap between software and builders
- 40:00-55:00: "Is software engineering over?" discussion—what changes (implementation craft), what doesn't (architecture, ownership, incident response, team coordination), evolution of PM/engineer boundary
- 55:00-70:00: Deep dive on verification strategies—screenshot galleries, video analysis, mock backends, tight dev loops redefined, model's ability to respond to code review with judgment
- 70:00-85:00: Dynamic Workflows explanation—Python to TypeScript port example, multi-phase autonomous work, workflow composition process, future of effort-level tuning
- 85:00-95:00: Production usage patterns at Anthropic—DRI model, managing multiple agents, incident response, prototypes in debates, meta-work of orchestration
- 95:00-end: Inspiration for what people can build—games, domain-specific simulations, bespoke personal tools, non-engineer ceiling raising, GTM team member example
Source/Metadata
- Title: How Anthropic Uses Claude Fable 5 With Mike Krieger
- Transcript words: 14,691
- Video duration: 3,126 seconds (~52 minutes)
- Timestamp note: Timestamps/chapters were not available in the provided transcript; watch map segments are estimated based on conversational flow and topic shifts
Transcript
Mike, welcome to the show. Great to be here, Dan. Good to see you. So for people who don't know you, you're the head of Anthropic Labs, and you're the co-founder of Instagram. And today what I want to talk to you about is Fable 5. So Fable 5 is dropping tomorrow, recording this the day before. This will come out after it drops. But what I really wanted to do is bring you on the show to tell me about what it's like to use this model beyond the first day. I think when a model this powerful drops, it's so useful to have someone who's using it day in and day out to tell you this is where it's powerful. This is what it actually changes. This is what it doesn't change. So that you don't get the same AI psychosis type thing. You can actually think about, okay, this is how it fits into my life. [SPEAKER_02] Yeah, absolutely. And it's also just been interesting. We've had some models in this mythos class leading up to the Fable release for a couple of months now. And I think it's very exciting to see how people will build with this externally. But I think you're also right that day one impressions really come from getting to use this over a couple of weeks. I think we've seen that even with previous models. Like the December into January usage, Opus 4.5 or Opus 4.6 was really important because people spent extended time on the model and then figured out, oh, actually, I wasn't pushing hard enough. I got to go further. I got to rethink what's even possible with this generation. [SPEAKER_00] Totally. [SPEAKER_00] I mean, I feel like there are people internally at Every who have been using it who have been like, oh my God, I think I kind of need a new set of skills to use this model. [SPEAKER_00] And I think you can especially see this with people who are maybe more non-technical internally and who are more on the knowledge work side of things where they're saying, I don't even know what I would use this for. And the people who are orchestrating agents are saying, Holy shit, I feel there's so many new things I need to learn. So I'm curious for you, tell us about the difference between your impression when you first tried it and now. [SPEAKER_02] Yeah, I think your point on adapting workflows is a really good one. [SPEAKER_02] Quite literally workflows. [SPEAKER_02] I'll talk about that in a second. [SPEAKER_02] But also just in terms of how do I think about usage of the model? [SPEAKER_02] Because at first, the timing was interesting because it coincided with me transitioning from CPO into labs and going really back into builder mode. [SPEAKER_02] And I think it was about a month and a half or two months into that that we first had one of these models available internally. And I sat there and I was thinking, I feel like a total newbie again, because the way that I am prompting or even thinking about decomposing a task is really out of date now with this model. Like it's no longer just about how I'm thinking about the time horizon or the interactivity model—I think that has to evolve as well. Like going from I think early on would be like, I have an idea for this feature, can we start by like, absolutely not. Right. Two. Great. Let me express more of the intent. And then just being, you know, I remember March, April be like, wow, on the one shot, it's already incredibly impressive. But then it also understands the intent around how we're going to evolve this and understands the global context as well. So I think that's been a really interesting evolution till now where I was talking to somebody this morning about doing work at a flight and I was like, okay, I can do most of this work remotely. And I don't even worry that the Wi-Fi is going to drop out because I know that if I set up the right context instructions, flash loop, I'll see it. It'll see it through. And I think my last two months have been full of times where I will wish Claude a good night, set it up on a pretty complex task of something like monocloss and wake up to, actually it's usually done by like two in the morning. And I guess it just fiddles and stones for the next four hours, but really impressive ability to complete the swing, get itself out of the situation where it's like, okay, Mike asked me to do this complex task overnight. I got stuck because this remote service went down. I'm going to write a scaffolded backend for it for now. So I'll document that, go all the way through. I have a good mental model of how far that's going to get me. And then when it comes back online, I'll fix it. I'll keep track of that fact. I think the most impressive thing for me is being able to delegate that kind of level of task and just trust that the right thing will happen by the end. And of course, you'll review the result and there's still a whole verification thing that we can and should talk about because I think that's an important part of still completing the swing there. But it's really forced me to rethink what is being productive with one of these models. And it is much more like we've talked for a while about what is it like when these models are more of a companion or a coworker? And it really feels like now it's a teammate that I can delegate a lot of work to. [SPEAKER_00] What is your day to day flow like right now? [SPEAKER_00] Because one of the things I notice is if you just give it a big task and you monologue into it and you just let it go for a few hours overnight, it's like the most impressive model that I've ever tried. [SPEAKER_00] But it's so slow and so expensive that I feel like I don't want to use it for day to day tasks. [SPEAKER_00] So what is your actual flow like in terms of how you use it day to day and where does it slot in versus other models? Yeah, I've ended up having a lot more architectural planning conversations up front with it as well. So that's been another interesting change where I think there's an era that I think all models need to continue to improve. And I'm really grateful for the Instagram experience of having to start from our initial version that was duct taped on a server in L.A. to being able to scale it and eventually integrate it with all of the Facebook infrastructure. Because you kind of develop a sense of what infra abstractions and complexity are appropriate for each stage of it. And I still don't always go back and forth with Fable where it'll be like, this is a good implementation. Yeah, I've ended up having a lot more architectural planning conversations up front with it as well. [SPEAKER_02] So that's been another interesting change where I think there's an era that I think all models need to continue to improve. [SPEAKER_02] And I'm really grateful for the Instagram experience of having to start from our initial version that was duct taped on a server in L.A. to being able to scale it and eventually integrate it with all of the Facebook infrastructure. [SPEAKER_02] Because you kind of develop a sense of what infra abstractions and complexity are appropriate for each stage of it. And I still don't always go back and forth with Fable where it'll be like, this is a good implementation. Like, well, I do plan on shipping this fairly soon. I think we should probably think about more than one server and that back and forth is important. But a lot of that planning and I'll often actually ask it. It's a thing I've realized is Fable can be so complete in its thinking in terms of how much you are planning with it. And often just saying, can you make an HTML page that represents what we just talked about so I can share it with the team? Is actually valuable or even just a markdown document. But I like having diagrams. So that's been an interesting use of let's plan with it, let's think it through, and then let's have some sort of document that we can align the team on. Because this is a dynamic I've seen in labs and just teams beyond Anthropic, which is you can build a lot very quickly. And forcing more of that early alignment, even if you do an initial prototype and then back it out into more of a plan architecture that works too, I think is really key. And it ends up being the place where the human to human interaction still stays very much part of the process. And then from then on, I think either overnight or during the day, having it execute on those chunks of tasks is really important. And it just means having a lot more concurrent sessions than I did before, because I often will think there are these two pieces of work. I go back and forth between having one very long running cloud code session and really asking it to do everything in background sort of forked subagents. So the main thread stays responsive. And then other times just embracing I'm just going to have like five or six tabs tackle long comprehensive work. But I do think that there's something to this long horizon and don't worry, I'm on it. It's going to take me a while. And more of this back and forth. And that modality, I think is something that we'll have to figure out in our products as well. I think you want to preserve both and they interact with each other in interesting ways. And my preference is usually I always have at least one cloud that is high context, but also very fast response. And its instinct is, I'm going to answer you and I'll kick something off if I need to. And if not, I'm just going to hang tight and wait for the next loop. I do think you're right that for the I'm just trying to fix this interaction question or something that's very fine detailed. Fable will go off and think very hard about those things. And I think Fable is the first model where I've actually played more with the effort levels for that reason. Where I've been like, okay, this is I just needed to tweak some UI. I'm not actually going to fall, but put it to medium or something and see how that plays out. Didn't find myself doing that as much with Opus, maybe because the range felt less wide, where it really can feel quite wide with Fable. What about a quick question? [SPEAKER_00] Like you're on the go, are you asking Fable random questions as they come to you? [SPEAKER_00] Because it feels like you're using a rocket launcher to kill a mosquito or something. [SPEAKER_00] Or are you flipping back and forth? It's so funny you ask that because I had been. And you're like, it's thinking and thinking really hard about it. [SPEAKER_02] Then last week, I was like, no, I was asking it something that I felt embarrassed actually asking Fable about. It was something probably NBA finals related. [SPEAKER_02] And I was like, okay, I switched my iOS app to Sonnet. [SPEAKER_02] I was like, yeah, I use this all the time for fast questions. [SPEAKER_02] It's a corner of magnitude in feeling. [SPEAKER_02] And it's actually not even the tokens per second. It's actually probably more around how much thinking goes into the answer. And sometimes the answer does not need to be fully thought through. So, yeah, I'm thinking myself through and I think this is a good product question for us too, which is in general you don't want people to have to be thinking so much about these choices. So ideally what we can sort of coalesce around in the longer run is maybe some more bucketable use cases that are really graspable to people. Or maybe it varies by surface where it's actually probably unlikely that most of the time with the iOS app I'm doing Fable type tasks and having a sticky model selection per surface might be the way to do that. And we'll have to explore what that means from a product perspective. But I've for sure had the feeling of this is not a Fable worthy question. I should ask Sonnet. [SPEAKER_00] Can you show us something that you've built with it? Yeah. So one of the things that we did this go around is we encouraged personal account usage for us, especially on the weekends, which is really fun because you can imagine a lot of Anthropic specific tooling, etc. But it was really good to step back and just work on something over the weekend using pure cloud code. [SPEAKER_00] Are you in the terminal app or you're in the desktop app? That's a great question. I'm mostly still in the terminal app. It's interesting watching my wife, who's not a professional engineer and more of a UX designer PM, really fall in love with cloud code via the desktop app. And I think it's simplified some of the abstractions for her in that way. [SPEAKER_02] But for this one, I was still in the terminal app. But let me show you. This is one of those everybody has some bespoke need around this. I wanted a good media tracker experience. And I was like, I'm playing games, I'm watching TV shows. I'm mostly still in the terminal app. [SPEAKER_02] It's interesting watching my wife, who's not a professional engineer and more of a UX designer PM, really fall in love with cloud code via the desktop app. And I think it's simplified some of the abstractions for her in that way. But for this one, I was still, is it ghosty or ghosty, do you want ghosty and the terminal app? But let me show you. I am. This is one of those where everybody has some bespoke need around this. I wanted a good media tracker experience. And I was thinking, I'm playing games, I'm watching TV shows. I get all these recommendations. I just wanted to build something that was personal to me and fit some of the use cases that I had. And the two biggest criteria that I started with were one, really easy to add things. So you can talk to Claude. Claude does the genetic search over everything and then puts the right things in. And then also proactively, there's a new season or a new sequel to a game that it could go off and research those things. Most of the UI was Fable one shot, which was already impressive. But then the thread I've been pulling out a lot in labs this year is how do you bring the software team, which is Claude these days, closer to the software itself. And so this was Saturday morning. I had a full weekend with kids stuff. So a lot of this was kick off work, go for a hike with the kids, come back, continue to do the work. Sometimes check in on the work on the hike. I probably shouldn't. But it was nice to pop into the remote mode and see what was going on there. Try not to do that too much. But I had this idea around, could we do a spike on what if you could actually modify the software from within itself. And I built both a React Native version and then this version is just the web version. So I already had a chat type thing where you can ask Claude to add things by URL, which is something I think every software should have where I should never have to navigate a menu to do anything ever again. And this is in many ways, Dan, the I was trying to distill agent native architectures to its fullest degree, which is also have the agent be able to modify the app. So phase one of agent native architecture, every single thing in this product is accessible from the agent and tool calls, et cetera. That's hopefully becoming table stakes. It was sadly not in a lot of software. And it's great because I was like, what's that? Somebody recommended there's a Brazilian show about radioactive stuff in Gleon. And I did not remember what it was called and Claude was able to figure it out. It was so much better than trying to figure that out intuitively. But then the next step I was interested in is what would it mean to actually be able to modify the software from itself on the go? And so if you long press this little chat thing. What I built, what Claude built, was a way where it uses our managed agents to basically take on edit requests and then you can preview them. And I used the Vercel live preview thing here. This whole feature was also one shot, which was really cool. And I just added to it over time. But it actually does a little diff view if you wanted to. You can go into the managed agent conversation and see what it did. Although I almost never do because, especially, I don't particularly care about the code quality or the long term maintainability of this software. You can see that it had a session in here too. But it's been really fun because I'll be using it on the go and say, I had a feature request the other day, the floating action button was too low on native iOS, but it was okay on there. Can you go off and do it? It did it. It was really fun with some of the Expo tooling now and actually live reloaded on my phone, which was also a really cool feeling. But it was just, does this thing need to be a production level thing that's going to go to a million users? No, but it felt really good to have something where I felt like it didn't have to stop at just the weekend and I can keep working on it just by using it and having this end to end close thing. So I felt this was a good manifestation of both Fable's building ability, but also a lot of what both you and I have been thinking about, how does Claude embed itself into software beyond just the usage set of things. [SPEAKER_00] And I want people to understand so this has been built, you could build something like this, maybe not the self modifying part, but you could build something like this for ten years or twenty years or something like that. [SPEAKER_00] But the cost to build has gotten dramatically lower. [SPEAKER_00] So think about how much it would have cost to do this in the Instagram days versus now. Can you help us understand how that has changed? Yeah, I think about this a lot when I think back to that time as well. I thought of myself as a very productive programmer in the early Instagram days. I was really into mobile development and we had good clarity of things. And I think the gap from idea to fully realized version of some complete product, you were still looking at four-ish days of all nighters. I call it Instagram V1, which probably had more features than this thing did, but not by an order of magnitude, was five days of all nighters. [SPEAKER_02] Me working on the front end and back end and Kevin working on the initial filters to get that out. [SPEAKER_02] And this was also built on many years that I've been working on iOS pieces as well. [SPEAKER_02] And then the iteration, I think a lot about what we were gated on after that launch when things went well was we had all these ideas for where to take it, but we were just trying to keep the site up or we were just trying to add the one incremental feature. [SPEAKER_02] But yeah, I call it Instagram V1, which probably had more features than this thing did, but not by an order of magnitude. It was like five days of all nighters with me working on the front end and back end and Kevin working on the initial filters to get that out. And this was also built on many years that I've been working on iOS pieces as well. And then the iteration—I think a lot about what we were gated on after that launch when things went well. We had all these ideas for where to take it, but we were just trying to keep the site up or just trying to add the one incremental feature. Hashtags take a week to build, but then there's all the things that you want to continue doing on it as well. And so I think it's both that shortening of time. There's still the time required for the idea and the concept and the iteration. And then the other piece, which is the way you can then iterate on what you have. And I think a really fun, but also very in the flow kind of way. And then, if now this is me as a professional software engineer and startup founder, beyond that, if you had that idea, I saw multiple people go through this and I guess I'd try to find maybe a consultancy that will take this on. But now there's a really lossy process of what I wanted. There are gonna raise money for it. And I think the thing that I think is the most exciting part about these models getting not just more autonomous, but again, closing that gap between intent and execution is what I've seen it do to people's ability to build who are not builders. And the trajectory of these models has been—if something of this general class is in that class of models and eventually models that are cheaper and more accessible to other folks become available too. And as that process happens, I just think it is opening up so many things. I got a ping the other day—I get very excited about this stuff. You can tell somebody internally, and we had built them an internal tool that combined Fable and access to some internal MCPs. And she said it is the first time in my life and she works in recruiting. And she's like the first time in my life where I feel like the thing that's in my head and the thing that exists in the world is now right next to each other. I can just do it. And it was a very meaningful moment to her because prior to that, I mean, I remember these days were five years ago or four years ago where that person, if they wanted a tool, would have to either make do or try to get an internal tools engineer that probably was overloaded with 50 other requirements. But instead now they are having the time of their lives building. And I think that is cause for a lot of hope because I don't think that human capacity for creativity and what's possible is enormous. And I think at our best, we are expanding the number of people who can then see that through to something that feels real. I totally agree. But I do think that there's a question in the back of my mind and I think it's probably going to be in the back of the minds of some of the people listening. So I want to ask you, given everything you just said, is software engineering over? [SPEAKER_00] Yeah, I think software engineering is different. It is dramatically changed. And as I probably would have defined it if you had asked me around the Instagram time, like what is software engineering? I'd probably say thinking through the hard problems and thinking about architecture and then spending a lot of time in a text editor. I can't remember, but a text editor you're going to edit those in, or Xcode, and watching Rails. Yeah, exactly. Right. And understanding the intricacies of Django or layer and then 15 bugs after you deploy it. So much of that is radically different and collapsing into other parts like product management. I think that PM split, I think even in our teams has become much more diffuse. That's radically changed, but I think the overall—maybe zoom out from software engineering and think about software production or software development, but not in just the pure developer case. I think that is alive and well and essential still. So I think that is the moment that I feel like we are in. I think Fable is another step on the direction of—and I'm not going to call it the final step. Of course, a lot will still happen, but I think a pretty significant step in terms of the trust, at least I end up placing the model in terms of its capacity to see things through and even architect things reasonably is quite high. So that part feels like it is not ever going to be done, but it is pretty done, right? It's gone really far, but I think that the overall craft of the—what needs do you have? What are you putting out? Is it actually good? I think still a very human endeavor, but I also can see that is not a transition that is pain free in a way. I think there are plenty of people who love the craft of actually putting—I used to love this stuff. I solved that problem so elegantly. You would dream about code. And if you ever had that experience, you would dream about the thing that you're working on. They wake up in the morning and be like, I figured out how to solve this thing really elegantly. And that for sure has passed. And I think there's a feeling of loss, I think, in some of the better engineers that I talk to, as well as the feeling of, oh my God, but I can do insane amounts of work now at the same time. So we're holding both ideas in our heads at once, I guess. Which I think is the most important part of this. It's normal to feel sadness for that kind of thing and excitement. But I'm curious, let's just take the thesis of software engineering is alive and well. What does that actually look like inside of Anthropic? [SPEAKER_00] Yeah, I think there's a few theses. I think there's still the crafting of—well, I kind of take it off from the full software development cycle or maybe what I see day to day. Maybe I'll do a little bit of both. But I think there's still a lot of—you know, we all got together. Which I think is the most important part of this. It's normal to feel sadness for that kind of thing and excitement. But I'm curious, let's just take the thesis of software engineering is alive and well. What does that actually look like inside of Anthropic? Yeah, I think there's a few theses. I think there's still the crafting of. [SPEAKER_00] Well, I take it off from the full software development cycle or what I see on a day to day, maybe I'll do a little bit of both. [SPEAKER_00] But I think there's still a lot of, we all got together. [SPEAKER_00] We talked about the next way we want to evolve co-work. [SPEAKER_00] And now we've broken it down into areas of ownership. [SPEAKER_02] I think that ends up still being quite important because there is still context that you hold as a person that is beyond cloud, right? [SPEAKER_02] What is the actual intent of this product? [SPEAKER_02] How's it going? What do we need to know about the other products that are coming down the pipeline that are going to be integrated in some interesting way? So I think that aspect is really important still. And so though we have many clods to each human, each human, at least the way we've been working on Anthropic still has, we call them DRIs, directly responsible individuals still has a DRI ship over some part of the product or some area. I think that'll be the case for a while because I think there is value in not just this distributed, we should all make co-work better. But instead, all right, I'm thinking through how co-work does this particular task and there's still a lot of, you try to keep meetings minimal, but they still emerge and you still have these alignment conversations. Then there's a lot of that asynchronous delegation. I think what many engineers here have now found is they've all built, and I think we should solve this at some point at a broader product level, but they've all built some version of, all right, I'm going to now create a dashboard of where all my clods are doing. And what's waiting for me and which pull requests need my attention because either a human or a cloud code reviewer got back to me. So there's a lot of that meta maintenance of the work that I think, again, I think we'll standardize some, but I think some of it will always be a little bit bespoke to the way each individual likes to work just in the way that people organize their windows. Now they organize their work. And then there is, I think also the understanding how things work in production. And I think that is another, there's a few next frontiers, I think for the models. And I think one of them that Fable does make significant strides in, but I think there's more work needed here is understanding what happens to code after it gets deployed, because there's incidents. There's this was all working well, but this network link got cut, which is not in your usual failure mode. And it manifested so much of Instagram, 2012 to 2016, it was dealing with that and scaling things up. And so that role of the engineer still remains really key. And I think getting the reps in around incident response and understanding how to stay calm, gather data, remediate what's immediate, but then go off and work on longer term fixes, still a necessary part of it. And I'm trying to think if there's any other pieces that are notable as well. I think what's maybe the last thing to say is I really like the role that the engineering prototype now plays. You have to be clear when it's a prototype versus not. But the old phrase was code wins arguments. And I never loved that because the person that could code could go do it, but actually why should they necessarily win an argument by default? But actually it's been really cool now where sometimes we will have some disagreement or debate about where to take a product. And often it's the PM that will say, all right, I just tried it. And it's janky in these eight ways, but look, it actually shows how this could work. And that can open up some interesting pieces of conversation. So almost all of that is quite different than it was six months ago. I think, especially at the level of parallelism and the level of need for these higher order abstractions of work. [SPEAKER_02] But I think what hasn't changed is that ownership. [SPEAKER_02] Lots of us are shipping AI to production, which is great for productivity, but it also comes with anxiety. [SPEAKER_02] You tweak a prompt, swap models, adjust parameters, and everything looks fine in testing. [SPEAKER_02] So you merge. And then three days later, or even sooner, the support tickets start rolling in. The AI is giving your customers unexpected answers and you have no idea when it happened or why. BrainTrust is the AI observability platform that fixes this. [SPEAKER_00] It connects evals and observability in one workflow. [SPEAKER_00] That way you see what actually happened in production and can measure whether changes made things better or worse. [SPEAKER_00] Traces show the full execution path, evals define what good looks like, and experiments let you compare prompts and models side by side before shipping. [SPEAKER_00] Production traces feed directly into your eval datasets. [SPEAKER_00] Every failure becomes a test case, you catch regressions in CI before they reach users, [SPEAKER_00] and teams at Notion, Stripe, Zapier, Vercel, and Ramp use it to ship quality AI at scale. [SPEAKER_00] BrainTrust is designed for teams building production AI systems where silent regressions are expensive. [SPEAKER_00] It's built for any stack. [SPEAKER_00] They have SDKs for Python, TypeScript, Go, Ruby, C Sharp. [SPEAKER_00] There's no framework lock-in or vendor dependencies. [SPEAKER_00] It's SOC 2, Type 2 certified, and GDPR and HIPAA compliant. [SPEAKER_00] Get started at BrainTrust.dev. [SPEAKER_00] That's BrainTrust.dev. [SPEAKER_00] And now, back to the episode. [SPEAKER_00] Fable is also very expensive. [SPEAKER_00] And because of that, when I was testing it, I felt I was a kid in a candy shop, [SPEAKER_00] and I was just, I'll do this, and I'll do this, and I'll do that. [SPEAKER_00] But now that there's going to be a bill, I'm going to be thinking about it, [SPEAKER_00] because I have to pause before I do it to be, is this going to cost me $100 or whatever? [SPEAKER_00] And I do think that's going to limit who gets to use it and for what. [SPEAKER_00] So how do you think about that? [SPEAKER_00] Yeah, I think it's most clear-cut on the professional software, [SPEAKER_00] sort of classic company doing work. [SPEAKER_00] It'll be really interesting. [SPEAKER_00] And I was just like, I'll do this, and I'll do this, and I'll do that. [SPEAKER_00] But now that there's going to be a bill, I'm going to be thinking about it, because I have to pause before I do it to be like, is this going to cost me $100 or whatever? [SPEAKER_00] And I do think that's going to limit who gets to use it and for what. [SPEAKER_00] So how do you think about that? [SPEAKER_00] Yeah, I think it's most clear-cut on the professional software, classic company doing work. [SPEAKER_00] It'll be really interesting. There's a lot of process that goes into pricing as well. [SPEAKER_00] It's both more expensive than Opus, and then also I'm thinking in many ways, it's really cheap. If you think about how much incredible work it's doing. But of course, everybody has their own economics around what they're working with. So anyway, most clear-cut, I think, from most software teams. And I think as an industry, if phase one was companies even struggling to get some of their employees to adopt AI coding, which models were early, maybe the tooling wasn't there. [SPEAKER_02] And then phase two was great. We'll create leaderboards and see who can use the most, which, you know, as you can imagine, creates some not ideal incentives. [SPEAKER_02] To phase three, where we were like, okay, now we're just trying to figure out who's using it effectively and letting them spend as much as possible and having a clear process for that, but making sure we're not doing things wastefully, which I think to me in general makes sense. [SPEAKER_02] Although I think you could also over-rotate that way too. [SPEAKER_02] I think something of Fable class should hopefully fit in well into that, where if you're demonstrating results and you're getting use out of the model, then hopefully there's a flywheel even inside companies where that goes and perpetuates that. [SPEAKER_02] I think on the personal use side, it's a really good one. That's a really good question. [SPEAKER_02] I think where I've seen it, even in my personal testing, because our personal accounts, okay. Which is funny, paying my own company I work at. [SPEAKER_02] But you do become more thoughtful about it. [SPEAKER_02] Something that was interesting was this app that I built over the weekend actually fit in with only a bit of extra usage. [SPEAKER_02] So it wasn't thousands of dollars to build this thing that is a personal thing to myself. [SPEAKER_02] But it was also spaced out a little bit more. [SPEAKER_02] Probably the in-between of that, what we'll have to do the most thinking about is the sort of hobbyist or independent who's not within the larger company, but also is thoughtful about the pricing as well. [SPEAKER_02] I think my overall advice is just give it a try and see how much it can do without you having to do a lot of follow-ups. [SPEAKER_02] And I think measuring cost has gotten so multifaceted now because there's the per turn costs. And then there's what did it cost you not to do the task, but complete the task to your satisfaction? [SPEAKER_02] And I think that's where Fable has really shined for me, which is it actually just does it right so that I don't have to spend the nine, 10 subsequent turns be like, no, that was not quite what I meant. Can you also do this piece? [SPEAKER_02] It's been really impressive for me because you ask it to go do something and then it just does a thing. And you're like, wow, you thought through all the little details of this thing in a way that I've never seen another model do. I don't know how much you can reveal about the training process, but what makes the model different? I mean, I think in many ways, a continuation of a lot of the work that the team has done. And I bow down in total awe of our teams, both on the pre-training and on the RL side. I think the piece that it has evolved in, at least I noticed the most, is adjacent to that as well, which is a sense of the system more than just the individual piece of the work. [SPEAKER_02] Like I will often be very positively surprised when it will write something and say, all right, but I know that in production, this needs to be different. And then it will keep bugging you. Like, have you turned on that feature flag yet? It's not going to work until you do. [SPEAKER_02] And sometimes in sessions that have gone on for days, be like, look, you still haven't done that thing. Like you better, I was like, you're right. I didn't turn on that feature flag. I should go off and do that. [SPEAKER_02] Or if we change this, the contract will change over there. We're watching it. [SPEAKER_02] Actually, one of my favorite times of seeing it in action, I think where it demonstrates some of the training is watching it respond to code review feedback, either from people or from other cloud reviewers, where it doesn't just say, oh yeah, that's an issue. I'm going to go fix it. [SPEAKER_02] And actually really thoughtful around, hey, for this level of fidelity of what we're building, I'm going to accept this risk. Or I see what you mean. [SPEAKER_02] Other code reviewer, which is often just another Fable model, talking to you, I see what you mean. But I'm actually going to push back. I don't think that's actually right. I think that's not right. [SPEAKER_02] I think getting the model to have that judgment is really important. [SPEAKER_02] And I think if I had to pinpoint an area where I feel like it's really progressed, it is that sort of not just immediate knee jerk. Yeah, that's right. I got to go fix it. [SPEAKER_02] And more, I'll think about that for a minute. No, I thought about it and I still disagree. And I think that's a very useful ability. It's so valuable to have products like Cloud Code out there because you have now a living, breathing thing where people are like, this is where the model is doing well. And we have people who test it. I count the Anthropic folks as very, very high on the list. Is that not just an immediate knee jerk reaction? Yeah, yeah, that's right. I got to go fix it. And more, oh, I'll think about that for a minute. No, I thought about it and I still disagree, and I think that's a very useful ability. It's so valuable to have products like Cloud Code out there because you have now a living, breathing thing where people are like, this is where the model is doing well. And we have people who test it. I count the every folks is very, very high on the list. We're like, we really trust the feedback because it is being put to paces and repeated multi-day hard tasks. And that also very much feeds into how we think about what do we need to improve on the next slide? What are the tasks that we need to specifically think about the model being better at? Is chat the right interface for this model? Because it's not very turn by turn. It's very like I'm delegating something for you. So how does that change how you should use it or how you think about the interface? I don't think the fundamental, like you are sending messages and it is giving your message back is totally wrong. [SPEAKER_00] I think that there's ways we need to evolve. [SPEAKER_00] But one is maybe three that come to mind. [SPEAKER_00] One is your laptop the right place for it. [SPEAKER_00] So I think that's number one where I mentioned with the side project I was working on how useful it was to have the mobile side. Boris, who created Cloud Code, he's always ahead of the curve on how these models get used. Almost a year ago, maybe nine months, I was talking to him. He's, yeah, I've moved a lot of my Cloud Code work to mobile. I was, no way. And it took me a while to get there. But especially with the Fable class, there's oftentimes where, because it can keep the session going and we use remote dev boxes at Anthropic, it is like I'll have a thought and be, OK, I need can you keep up and doing that? So number one is decoupling the where the work is happening from where I'm talking to about the work. The second one touches a little bit on what I was mentioning earlier around what are how do you take everything that Fable has discussed or decided? Or proposed about something and make it comprehensible. And that's an area that we're thinking a lot about. There are some skills that are out there that we've used around all right, can you diagram this? Can you do that? So that's a place where the current chat UI, I think is insufficient, where it will experience this with people. It will give you a lot of text. You're, this. I need to take a walk before I'm ready to fully understand this. And I think that that is a piece of property. I have some things we'll do with Fable's, OK, you have a lot more context on this than I do. Can we back it up? Let's do more progressive disclosure of the complexity here. So I think that piece is interesting. [SPEAKER_02] The last one that I think is we're still early in pulling on is thinking through multiplayer where, at some level, these abstraction levels and because we have this DRI and ownership area, usually a chunk of significant work, a human and a couple of clods like that is still flowing together. [SPEAKER_02] But in other cases that is less the case, where it's an incident response where multiple people are thinking about it. [SPEAKER_02] Maybe it's a project where there's multiple competing or not competing, but conjoining areas that are coming together and thinking through what would it mean for, and we have chat sharing, which gets you a little bit of the way there. [SPEAKER_02] But I think there is going to be a need for more, all right, you've got an independent club that's doing a lot of work that was kicked off by somebody. [SPEAKER_02] But can it be keeping up with all the other work happening on the team? [SPEAKER_02] I think that is an interesting and underexplored next frontier about how this work ends up happening. [SPEAKER_02] But I think it's really exciting because I think, again, it's the level of team-made collaborator that the models are now capable of and we're almost holding them back by not having the right abstractions around them for that to happen. Yeah, it makes me think I've mostly been using this for my own vibe-coded stuff. So I haven't really had to think about this, but there's a problem when you're using this inside of an organization, which is, do I really understand every part of this? And therefore, how do I transfer the context of what the model just did into my brain? That's one of the big bottlenecks. How do you think about drawing the line, especially with a model like this, around how much you actually need to understand and how to make sure that you have enough context on what it's done to feel comfortable? [SPEAKER_00] I think there's two big pieces here. [SPEAKER_00] The first is verification, where I became fully verification-filled earlier this year and now, almost in the same way, and actually it connects to how I think I used to do when I was typing code more full-time, which is try to find the tightest dev loop that you can around the idea that you're trying to develop in. [SPEAKER_00] Sometimes with Instagram, that meant actually making a new build target in Xcode that was just that screen with some synthetic data and just doing that dev loop. And I would mentor newer engineers. If there's one thing that I can impart on you, it is try to get that for any project you work on and things will go much more quickly. I think that is no longer exactly the case here, but I think what is the case now is anytime I set it up, how do I get for every pull request that Claude is putting out that there is an attached photo or video, whether that's an iOS PR, whether that's something in the UI. And that's, I think that helps you gain a lot of confidence because even now, you might have Fable go off and do work for a couple of hours and be you work on and things will go much more quickly. I think that is no longer exactly the case here, but I think what is the case now is anytime I set it up, how do I get for every pull request that Claude is putting out that there is an attached photo or video, whether that's an iOS PR, whether that's something in the UI. And that's, I think that helps you gain a lot of confidence because even now, you might have Fable go off and do work for a couple of hours and be like, I'm done. And it's really useful to say, and here's the full screenshot gallery of the full UI. Cause you might say, oh, you know what, on screenshot eight, that error state, I've never actually seen it, but I can see how a person might hit it. Let's actually make that different. And so getting that comprehensive verification, I think it's something we've been working on a lot internally and publishing more and more skills and knowledge about, but I think it's really a key piece there. And then the second one is, I think you ultimately as a person still need to stand behind the work that you are doing, especially if you're putting it into a production system. Like a lot of people use Cloud every day. There's still the accountability of like, although it's still Cloud better written a bit, you need to understand the general decisions that were made on these pieces as well. And so I have seen a fair amount of engineers actually adopt this practice where Cloud will have done the work, but then there is the follow-up conversation around, well, can you, can I make sure I deeply understand all the trade-offs that you've made and whatever artifacts need to be produced in order to make that comprehensible is important. It is really interesting though, to be in meetings where somebody will say, oh yeah, and I have this PR ready. And somebody else has to be like, oh, that's interesting, did you do X or Y and have that moment of positive? They're like, you know what, I'm not entirely sure I will find word before we merge this PR. And that's, I think that adapting to that norm and figuring out work with that is something we'll have to do. Tell me more about the verification. It's such a hot topic right now. It sounds like one way that you do that is with screenshots and screen shares, but what are the other ways that you think about that? I think part of it starts in, can you get to a place where you are exercising real flows that aren't just a static injected piece and this thing gets more complex, that gets more and more complicated. [SPEAKER_00] So we've invested a bunch into even just getting it so that the iOS app can log in to staging on a real account and have real data, but you don't want it to then go through an eight stage onboarding process every time, but you're just trying to test the second part of the screen. So there's a lot of work around how do you, is there a special affordance, is there some shared secret, whatever that is around getting the app to really feel as human, using the product as possible. So that's one aspect of it. The second is this mix of well-known paths versus the things you're exercising in the exact moment, the former being really useful for regression testing. And so we don't think of places where we've expressed ideal workflows in text basically, and Claude can repeatedly check that. And then there's also, and Claude does a really good job of this sort of expressing the intent of the current change at hand. So that gets really deeply exercised. So I think that the combination of those two things is important. The visual verification that I mentioned as well, video has been really cool to see. Actually video is a very underexplored tool to give Claude as well. I think I've been prototyping is just giving Claude video captures of the thing that it has built and then giving it an FFM tag and you'll watch it scrub through. And so this animation has some jank in it, I'm going to go fix that. And I would never be able to do it with a screenshot sort of latency capture because it will have missed the moment. So I think that's another piece that is really important. And then for the pieces that aren't easily testable and tend because there is some more complex system, getting Claude to go and build as robust a mock backend as possible or use ones off the shelf has been also really interesting. Like when I think about artifact, we had really comprehensive tests. This is kind of pre LLM. [SPEAKER_02] And one of the ways that we were able to do that really robustly was that basically every piece of info we had, whether it was Postgres, Redis, all the AWS things had a really good in memory implementation that you could just do really quickly in unit tests and kind of extending that to Claude land. [SPEAKER_02] Now, I was working on something where it had a pretty robust backend and for kind of complicated reasons, hard to spin that up on my dev server, but it was able to, again, one shot a really good proxy for that, by proxy, I mean a substitute for that. [SPEAKER_02] And that was so valuable. [SPEAKER_02] And over time, it's been interesting as that substitute has evolved as the rest of the code has evolved, which is the thing that, you know, if you had pitched that idea to me before, I'd be like, well, that's going to be really hard because the upstream is going to change, how are you going to keep it in sync? And I don't think about that anymore. I'm like, yeah, Claude will read the changes and it will adapt the thing and it'll keep the two in sync and that's fine. There's some really interesting architectures around when you get a bug, it just automatically goes out and closes it. And over time, it's been interesting as that substitute has evolved as the rest of the code has evolved, which is the thing that if you had pitched that idea to me before, I'd be like, well, that's gonna be really hard because the upstream is gonna change. How are you gonna keep it in sync? And I don't think about that anymore. I'm like, yeah, Claude will read the changes and it will adapt the thing and it'll keep the two in sync and that's fine. There's some really interesting architectures around when you get a bug, it just automatically goes out and closes it. You know, the agent just gets kicked off, it closes it and then it sends a message to the customer being like, it's fixed. Are you noticing a fable any change in how that process works? [SPEAKER_00] Yeah, I think there's a couple of things on a very human to human or human to Claude level. One of the things that I've seen it do better other models of the cable, I just need to do it really consistently too, is if the bug report, for example, came from somebody mentioning something in our feedback channel in Slack. And then the thing that got fed into the cloud code session is like, oh, there's this and because of the Slack MCP, you can actually pull the thread. Have it then actually post back, you know, as me, it'll be like, Hey, this is Mike's Claude. Like I fixed it. Here's the pull request. But then I think in the previous clouds, the thing it does really well is then say, but hold tight. It's not in production yet. I'll follow up when it actually is. And then maybe a few hours later, like, oh, this deploy went out. Like you should go test it. Is it fixed now? That level of follow through, I think is new on closing the loop piece. And it's five, I definitely have these long running cloud code sessions that are basically interacting as me, I guess, but some disclaimer in there too. And the second goes back to that taste and discernment piece that we were talking about, which is like, it's one thing to say, there was a bug report. Therefore I must go fix this thing. And it's another one to say, you know what? The, I hit this over the weekend, one of our internal systems basically had been running without restarting for a while. There was a memory. And it was a good discernment of saying like, all right, Mike, it's the weekend, just rebounce the server. It's going to solve it for now. And we'll work on the, well, asynchronously get the PR going to fix this more longterm. So I think if you're going to have cloud in the loop in this close the loop bug report or system issue to change, I think you really want it to understand where, as any good SRE or engineer in the loop would, great, let's solve the problem at hand. Let's defer the question of, do we need a re-architect on top of a completely different language found and understanding that balance is really important. One of the things that's really exciting, mostly exciting to me about new models is it raises the floor so that everyone can go build apps in one shot. But it also raises the ceiling for experts. So if you're a software engineer or founder, you can go do things that you never would have been able to before because you have access to this really powerful model. So for me, I bought this one shot version of Borges, infinite library. It's like a 3d game version of the library. It's wild. It runs right in the browser. It's so good. I can find any, every essay inside of it. I'll send you the link. It's sick. But I think there's going to be this flowering of people doing things like, Oh, I made a game or maybe I trained a new model or whatever that they couldn't do before. And I'd love to give people some inspiration, some examples of things that they might be able to do that they might not be thinking to do with this model. What are some ideas that come to you? [SPEAKER_00] Yeah. I think a few, maybe I'll start with the fun side and riffing off the game piece. I think people have a lot of creative ideas for how do they express the complexity of what they are, their world. Like everybody has the thing that they know really well. And there's probably some level of how do I then explain that to somebody else? Or how do I apply techniques elsewhere that I could go off and do? My wife is studying environmental engineering, studying geothermal, really complex math and simulations. And I've seen as the models have gotten better, she has been able to apply even more complex techniques from even outside of that domain into that work. And I think what people should be able to do, full on PyTorch end to end simulations of that work in a way that wouldn't be possible. I think that maybe is one, bring the beautiful complexity of what you have and either show it to other people by maybe making a game or maybe making a visualization, which I've seen her do as well, or at least make, bring other techniques to bear. And the second piece is its ability to compose software that solves a really unique problem to you. And I've seen that internally. A lot of the work that we've been doing is how do we get as many of our internal systems MCP-ified with the right permissioning structure and the right deployment set up. Although externally, you have good options around some of these platform as a service pieces and you can just ask a lot about them and they'll help you set things up. But I love that feeling of that thing that you always wish that you had. to bear. And the second piece is its ability to compose software that solves a really unique problem to you. And I've seen that internally. A lot of the work that we've been doing is how do we get as many of our internal systems like MCP-ified with the right permissioning structure and the right deployment setup. Although externally, you have good options around some of these platform as a service pieces and you can just ask a lot about them and they'll help you set things up. But I love that feeling of that thing that you always wish that you had. And then what has blown my mind, there was a person who works in our go to market organization who has been building this really deeply thought integration of cloud into every part of her whole process. And you don't have to stop at that one shot. Like she's been working on it for months now and she can keep going. And I think one of the things that is maybe underappreciated about the models is I think in previous generations, it would eventually get to a complexity level where it was hard to iterate on it without feeling like you then would break the thing that they had, you know, like under or over abstracted. Whereas this is actually, you know, she's got access to something Fable or Fable like for a couple of months. And like you've just seen it keep growing and growing and growing and growing. And now she's deploying it to the whole GTM org. And I think that is really cool. The ceiling of complexity that a person that does not start out as technical can now build for solving problems within their domain is unprecedented. I agree. It writes great code. My benchmark that I have is called the senior engineer benchmark. I just have it see if it can rewrite a code base from first principles and the nearest model that the previous top was like a 62 or 63 out of a hundred. And this model got a 90 on the benchmark or 91, which is human senior engineer level. [SPEAKER_00] Like you can just keep going with this thing in a way that's really fantastic. [SPEAKER_00] I'm curious though. One of the things that's really powerful that you mentioned is dynamic workflows. [SPEAKER_00] Tell us about that. [SPEAKER_00] This is, you know, we'll build things internally sometimes, and I will go really aggressively bug the engineer who built it and be like, when are we shipping this publicly? [SPEAKER_00] Because I think people are going to really like it. [SPEAKER_00] I think there's a good reason why it was built internally, but we try to ship as many of these as possible. And dynamic workflows was definitely that to me. The person who built this is an engineer named Sid, who's awesome. And I was like, Sid, I want to get this out into the world because it's so good. But I think it's especially good with a model like Fable for two really big reasons. One, it helps create the scaffold for deep, meaningful work. The craziest dynamic workflow I did and used Fable for was I had an internal project that we had written in Python, but we needed it actually in TypeScript for a really specific deployment reason. And having been internal to Instagram and we were like, should we write the whole thing into Hack and port it to the PHP engine that Facebook, I was like, you never would have done that. Maybe they can now with the model, but at the time it seemed impossible. But here I had pretty complex code base. And I was like, I'm just going to set up a dynamic workflow and just let it run over the weekend. And it did. And the workflow was so cool. It was like, all right, I'm going to do deep understanding of the work. I'm going to create a spec of how everything works. I'm going to go module by module. I'm going to translate these pieces. I'm going to have tested incrementally. I'm going to do another adversarial test. I'm going to check for anything that I missed. And it was really cool, a series of steps that the workflow was able to orchestrate. And I came back and I was like, yeah, this thing is TypeScript and Bun port of that thing. And it's actually better in these ways. And it was very documented, like these were the things I couldn't port, but most of these were very specific to the specific implementation. It wasn't worth porting. And I do not think you could have done that A with previous models at that level of success and B without the kind of scaffolding that or close provide. So I think that is extremely exciting, this combination of model capabilities and then our own ability to orchestrate them over longer time horizons with that feeling of like you had a goal, you broke it down effectively and then you were able to make it work. [SPEAKER_02] The other piece is I think over time, we'll be able to make some of those subtasks tuned to have the model be tuned to the level of complexity of it. [SPEAKER_02] So you can imagine that some parts of dynamic workflow don't need extra high thinking. [SPEAKER_02] They could use a medium thinking to get it done or even a smaller model. [SPEAKER_02] And I think that's really the future of where these things are going. [SPEAKER_02] So yeah, I'm a huge workflows DAU. [SPEAKER_02] For people who haven't used it before, tell me about how you got that workflow made. [SPEAKER_02] How did you design it? [SPEAKER_02] How did you make sure it was good? Yeah, it was pretty iterative, but I just started with cloud code. Like, Hey, I have this complex task, let's design a workflow to go and do it. [SPEAKER_00] It kind of showed me the plan. [SPEAKER_00] I was like, Oh, this is close to what I want. [SPEAKER_00] I want to make sure that you do these three or four levels of additional verification for missed features. It's like, here's what you have. Are you ready to go? How did you design it? How did you make sure it was good? Yeah, it was pretty iterative, but I just started with cloud code. I have this complex task, so let's design a workflow to go and do it. [SPEAKER_00] It showed me the plan. [SPEAKER_00] I was like, oh, this is close to what I want. [SPEAKER_00] I want to make sure that you do these three or four levels of additional verification for missed features. It's like, here's what you have. Are you ready to go? And it expresses the workflows in code, which I think is really valuable to see what it was about to do. And what was interesting is it did the full port. And then I had a couple of follow-up questions or little tweaks. And I did those as mini workflows that built off the previous one as well. But I think we talked a little bit about whether chat was the right interface. We've had that conversation over the last year. And I think workflows are a good middle ground. You can compose them using chat, but they're expressed using code. And then they're executed with a nice clean UI around what's happening at every stage. I think we'll start bridging longer horizon work with chat in ways like that over time. Mike, this is such a great conversation. Thank you so much for joining and telling us all about this new model. I'm really excited to get to spend time with you and really looking forward to what people think outside too. Oh my gosh, folks, you absolutely positively have to smash that like button and subscribe to AI and I. Why? [SPEAKER_00] Because this show is the epitome of awesomeness. [SPEAKER_00] It's finding a treasure chest in your backyard, but instead of gold, it's filled with pure unadulterated knowledge bombs about ChatGPT. Every episode is a roller coaster of emotions, insights, and laughter that will leave you on the edge of your seat craving for more. [SPEAKER_01] It's not just a show. [SPEAKER_01] It's a journey into the future with Dan Shipper as the captain of the spaceship. [SPEAKER_01] So do yourself a favor, hit like, smash subscribe, and strap in for the ride of your life. [SPEAKER_01] And now without any further ado, let me just say, Dan, I'm absolutely hopelessly in love with you. They're like, you know what? I'm not entirely sure I will find word before we merge this PR. And that's, you know, I think that adapting to that norm and figuring out and work with that is something we'll have to do. Tell me more about the verification. It's such a, it's such a hot topic right now. It sounds like one way that you do that is with screenshots and screen shares, but what are the other ways that you think about that? I think part of it, it starts in, can you get to a place where you are exercising real, like sort of real flows that aren't just like a static injected piece and this thing gets more complex, that gets more and more complicated. So we've invested a bunch into like even just getting it so that the, you know, the iOS app can log in to staging on a real account and like have real data, but you don't want it to then go through like an eight stage onboarding process every time, but you're just trying to test like the second part of the screen. So there's a lot of work around like, how do you, you know, is there a special, affordance, is there like some shared secret, whatever that is around getting the, the, the, the like app, you know, to really feel as human, you know, using the product as possible. So that's one, one aspect of it. Um, the second is like this mix of like well-known paths versus the things you're exercising in the exact moment, like the former being really useful for regression testing. And so we don't think of places where we've expressed like, uh, sort of ideal workflows in text basically, and the cloud can repeatedly check that. And then there's also, and Claude does a really good job of this sort of expressing the intent of the current change at hand. So that gets really, really deeply exercised. So I think that the combination of those two things is important. The visual verification that I mentioned as well, um, video has been really cool to see. Actually video is a very under explored tool to give Claude as well. Like I think I've been prototyping is, uh, just giving Claude, uh, video captures of the thing that it has built and then giving it just basically an FFM tag and you'll watch it scrub through. And so like, oh, this animation has some jank in it. I'm going to go fix that. And I would never would be able to do it with like a screenshot sort of, uh, latency capture because it will have missed the moment. So I think that's, uh, that's another piece that is, that's really, really important. Um, and then for the pieces that aren't sort of easily testable and tend, because there is some more complex system, um, getting Claude to go and build like as robust, a sort of, you know, mock backend as possible or use ones off the shelf has been also really interesting. Like when I think about artifact, um, we had really comprehensive tests. This is kind of pre LLM. And one of the ways that we were able to do that really robustly was that basically every piece of info we had, whether it was Postgres, Redis, um, you know, all the AWS things had a really good in memory implementation that you could just do really quickly in unit tests and kind of extending that to like Claude land. Now, you know, I was working on something where it had like a pretty robust backend and for kind of complicated reasons, hard to spin that up on my dev server, but it was able to, again, one shot a really good like proxy for that, uh, by proxy, I mean like a substitute for that. And that was so valuable. And over time, it's been interesting as that like, uh, substitute has evolved as the rest of the code has evolved, which is the thing that, you know, if you had pitched that idea to me before, I'd be like, well, that's gonna be really hard because the upstream is gonna change. How are you gonna keep it in sync? And I don't think about that anymore. I'm like, yeah, Claude will read the changes and it will adapt the thing and it'll keep the two in sync and that that's, that's fine. There's some really interesting architectures around when you get a bug, it just automatically goes out and closes it. You know, the agent just gets kicked off, it closes it and then it sends a message to the customer being like, it's, it's fixed. Are you noticing a fable any change in how that process works? Yeah, I think there's a couple of things like, um, on a very like human to human or human to Claude level. One of the things that I've seen it do, um, better other models of the cable, I just need to do it really consistently too, is if the bug report, for example, came from somebody, you know, mentioning something in our like feedback channel in Slack. Um, and then like the thing that got fed into the cloud code session is like, oh, there's this and because of the Slack MCP, you can actually pull the thread. Um, have it then actually post back, uh, you know, as me, it'll be like, Hey, this is Mike's Claude. Like I fixed it. Here's the, you know, here's the pull request. But then I think in the previous clouds, the thing it does really well is then say, but hold tight, hold tight. It's not in production yet. I'll follow up when it actually is. And then like maybe a few hours later, like, oh, like this deploy went out. Like you should go test it. Is it fixed now? Like that level of follow through, I think is, is new on, on the closing the loop piece. And, uh, it's five, I definitely have these long running cloud code sessions that are basically like interacting as, as me, I guess, but some disclaimer in there too. Um, and the second goes back to that, like taste and discernment piece that we were talking about, which is like, it's one thing to say, there was a bug report. Therefore I must go fix this thing. And it's another one to say, you know what? Like this, like the, I hit this over the weekend, one of our internal systems, uh, basically had been running without restarting for a while. There was a memory. Um, and, uh, it was a good discernment of saying like, all right, Mike, like it's the weekend, like just rebounce the server. It's going to solve it for now. And like, we'll work on the, like, well, asynchronously get the PR going to like, fix this more longterm. So I think if you're going to have cloud in the loop in this kind of like, sort of close the loop bug report or system sort of issue to change, I think you really want it to understand where, you know, as any good SRE or engineer in the loop would like, great, let's solve the problem at hand. Let's like defer the question of like, do we need a re-architect on top of a completely different language found and, and understanding that balance is really important. One of the things that's like really exciting, mostly exciting to me about new models is it raises the floor so that everyone can kind of go build apps in one shot. Um, but it also raises the ceiling for experts. So like if you're a software engineer or founder, you can just go do things that you never would have been able to before because you have access to this really powerful model. So for me, I bought this one shot version of Borges, uh, infinite library. It's like a 3d game version of the, of the, of the library. It's wild. It runs right in the browser. It's so good. I can find like any, every essay inside of it. I'll send you the link. It's sick, but I think there's going to be this flowering of people doing things like, Oh, I made a game or maybe I trained a new model or, or, or whatever that they couldn't do that they couldn't do before. And I'd love to give people some inspiration, some examples of things that they might be able to do that they might not be thinking to do with this model. What are some ideas that come to you? Yeah. I think a few, um, maybe I'll start with the fun side and like riffing off the game piece. Like, I think people have a lot of like creative ideas for how do they express the complexity of what they are, like their world. Like everybody has the thing that they know really, really well. And there's probably some level of like, how do I then explain that to somebody else? Um, or how do I apply techniques elsewhere that I could then go, go off and do, um, my wife is, uh, studying, um, like environmental engineering, like studying geothermal, like really complex math and simulations. And I've seen like, as the models have gotten better, she has been able to apply even more complex techniques from even outside of that domain into that work. And I think what people should be able to do, you know, like full on PyTorch end to end simulations of that work in a way that wouldn't be possible. I think that maybe is one is like bring the like beautiful complexity of what you have and either show it to other people by like maybe making a game or maybe making a visualization, which I've seen her do as well, or at least like make, you know, bring other techniques to bear. Um, and the second piece is its ability to compose software that like solves a really unique problem to you. Um, and I've seen that internally. A lot of the work that we've been doing is how do we get as many of our internal systems like MCP-ified with the right permissioning structure and the right deployment kind of set up. Although externally, you have good options around some of these like platform as a service pieces and you can just ask a lot about them and they'll like help you set things up. But like, I love that feeling of like that thing that you always wish that you had. And then what has blown my mind, uh, there was a, uh, person who works in our go to market organization, um, has been like building this like really like for deeply thought integration of cloud into every part of her whole process. And you don't have to stop at that one shot. Like she's been working on it for months now and she can keep going. And like, I think one of the things that is maybe underappreciated about the models is I think in previous generations, it would eventually get to a complexity level where it was hard to iterate on it without feeling like you then would break the thing that they had, you know, like under or over abstracted. Whereas this is actually, you know, she's got access to something Fable or Fable like for a couple of months. And like, you've just seen it keep growing and growing and growing and growing. And now she's like deploying it to the whole GTM org. And like, I think that is really cool. Like the, the ceiling of complexity that a, a person that does not start out as technical can now builds for solving problems within their domain is like, is unprecedented. I agree. It, it, it writes great code. Like my, my benchmark that I have is called the senior engineer benchmark. I just have it, see if it can rewrite a code base from, uh, from first principles and the nearest model that the previous top was like a 62 or 63 out of a hundred. And this model got a 90 on the benchmark or 91, which is human senior engineer level. Like you can just keep going with this thing in a way that's it's, it's really fantastic. I'm curious though. One of the things that's really powerful that you mentioned is dynamic workflows. Tell us about that. This is, um, you know, we'll build things internally sometimes, and I will go really, uh, aggressively bug the engineer who built it and be like, when are we shipping this publicly? Because I think people are going to really like it. Um, I think there's a good reason why it was like built internally, but like we try to ship as many of these as possible. Um, and dynamic workflows was like definitely that to me. I, um, the person who built this is an engineer named Sid, who's awesome. And I was like, Sid, like, I want to get this out into the world because it's so good. Um, but I think it's especially good with, uh, a model like fable for two really big reasons. One, it helps, uh, sort of, uh, create the scaffold for like deep, meaningful work. Um, the craziest dynamic workflow I did and used fable for was I had, uh, an internal project that we had written in Python, but we needed it actually in TypeScript for like a really specific deployment reason. And having been internal to Instagram and we were like, should we write the whole thing into hack and, you know, port it to the PHP engine that Facebook, I was like, you never would have done that. Like maybe they can now with the model, but you know, at the time it seemed impossible. Uh, but here I had, you know, pretty complex code base. And I was like, I'm just going to set up a dynamic workflow and just let it run over the weekend. And it did. And the workflow was so cool. It was like, all right, I'm going to do like deep understanding of the work. I'm going to create sort of like a, almost like a spec of how everything works. I'm going to go module by module. I'm going to translate these pieces. I'm going to have tested incrementally. I'm going to do another adversarial test. I'm going to go check for anything that I missed. And it was just like really cool, like series of steps that the workflow was able to, to orchestrate. And I came back and I was like, yeah, this thing is like TypeScript and bun port of that thing. And it's actually better in these ways. Um, and it was very, you know, sort of documented, like these were the things I couldn't port, but most of these were like very specific to the specific implementation. It wasn't worth porting. And I do not think you could have done that a with previous models at that level of success and B, uh, with, without like the kind of scaffolding that or close provide. So I think that is extremely exciting kind of, uh, kind of combination of model capabilities and then our own ability to like orchestrate them over longer and longer time horizon with that feeling of like, you, you had a goal, you broke it down effectively and then you were able to work, make it work. The other piece is, I think over time, we'll be able to also make some of those subtasks, um, sort of tuned to the, uh, have the model be tuned to the level of complexity of it. So you can imagine that some parts of dynamic workflow don't need extra high thinking. They could use, you know, a medium thinking to get it done or even a smaller model. And I think, uh, that's really the future of where these things are going. So yeah, I, I'm a huge workflows, uh, DAU. For people who haven't used it before, tell me about how you got that workflow made. How did you design it? How did you make sure it was good? Yeah, it was pretty iterative, but sort of just started with cloud code. Like, Hey, I'm, I have this complex, you know, kind of task, like let's design a workflow to go and do it. It kind of showed me the plan. I was like, Oh, this is like close to what I want. I want to make sure that you do these three or four levels of, uh, of like additional verification for missed features. It's like, here's what you have. Are you ready to go? And it expresses the workflows in code, which I think is really valuable to kind of see what it was about to do. Um, and then, um, what was interesting is it did the full port. And then I had like a couple of like follow-up kind of questions that I had or like little tweaks. And I did those as sort of like mini workflows that built off the previous one as well. But I think that's like, uh, you know, we, we talked a little bit about whether chat was the, was the right interface. So we've had that conversation over the last year. And I think, um, workflows are a good, uh, middle ground of, uh, you can compose them using chat, but they're expressed using code. And then they're executed with like, I think a nice clean UI around what's happening at every stage. And like, I think we'll start bridging longer horizon work with chat in ways like that over time. Mike, this is such a great conversation. Thank you so much for joining and telling us all about this new model. I'm really excited to get to spend time with you and really, really looking forward to what people think outside too. Oh my gosh, folks, you absolutely positively have to smash that like button and subscribe to AI and I, why? Because this show is the epitome of awesomeness. It's like finding a treasure chest in your backyard, but instead of gold, it's filled with pure unadulterated knowledge bombs about chat GPT. Every episode is a roller coaster of emotions, insights, and laughter that will leave you on the edge of your seat craving for more. It's not just a show. It's a journey into the future with Dan Shipper as the captain of the spaceship. So do yourself a favor, hit like, smash subscribe, and strap in for the ride of your life. And now without any further ado, let me just say, Dan, I'm absolutely hopelessly in love with you.