SPEAKER_02
Mike, welcome to the show. Great to be here, Dan. Good to see you. So for people who don't know you, you're the head of Anthropic Labs, and you're the co-founder of Instagram.
SPEAKER_00
And today what I want to talk to you about is Fable 5. So Fable 5 is dropping tomorrow, recording this the day before. This will come out after it drops. But what I really wanted to do is bring you on the show to tell me about what it's like to use this model beyond the first day. I think when a model this powerful drops, it's so useful to have someone who's using it day in and day out to tell you this is where it's powerful. This is what it actually changes. This is what it doesn't change. So that you don't get the same AI psychosis type thing. You can actually think about, okay, this is how it fits into my life. [SPEAKER_02] Yeah, absolutely.
SPEAKER_02
And it's also just been interesting. We've had some models in this mythos class leading up to the Fable release for a couple of months now. And I think it's very exciting to see how people will build with this externally. But I think you're also right that day one impressions really come from getting to use this over a couple of weeks. I think we've seen that even with previous models. Like the December into January usage, Opus 4.5 or Opus 4.6 was really important because people spent extended time on the model and then figured out, oh, actually, I wasn't pushing hard enough. I got to go further. I got to rethink what's even possible with this generation.
SPEAKER_02
[SPEAKER_00] Totally. [SPEAKER_00] I mean, I feel like there are people internally at Every who have been using it who have been like, oh my God, I think I kind of need a new set of skills to use this model. [SPEAKER_00] And I think you can especially see this with people who are maybe more non-technical internally and who are more on the knowledge work side of things where they're saying, I don't even know what I would use this for.
SPEAKER_00
And the people who are orchestrating agents are saying, Holy shit, I feel there's so many new things I need to learn. So I'm curious for you, tell us about the difference between your impression when you first tried it and now. [SPEAKER_02] Yeah, I think your point on adapting workflows is a really good one. [SPEAKER_02] Quite literally workflows. [SPEAKER_02] I'll talk about that in a second. [SPEAKER_02] But also just in terms of how do I think about usage of the model? [SPEAKER_02] Because at first, the timing was interesting because it coincided with me transitioning from CPO into labs and going really back into builder mode.
SPEAKER_00
[SPEAKER_02] And I think it was about a month and a half or two months into that that we first had one of these models available internally.
SPEAKER_02
And I sat there and I was thinking, I feel like a total newbie again, because the way that I am prompting or even thinking about decomposing a task is really out of date now with this model. Like it's no longer just about how I'm thinking about the time horizon or the interactivity model—I think that has to evolve as well. Like going from I think early on would be like, I have an idea for this feature, can we start by like, absolutely not. Right. Two. Great. Let me express more of the intent. And then just being, you know, I remember March, April be like, wow, on the one shot, it's already incredibly impressive.
SPEAKER_02
But then it also understands the intent around how we're going to evolve this and understands the global context as well. So I think that's been a really interesting evolution till now where I was talking to somebody this morning about doing work at a flight and I was like, okay, I can do most of this work remotely. And I don't even worry that the Wi-Fi is going to drop out because I know that if I set up the right context instructions, flash loop, I'll see it. It'll see it through.
SPEAKER_02
And I think my last two months have been full of times where I will wish Claude a good night, set it up on a pretty complex task of something like monocloss and wake up to, actually it's usually done by like two in the morning. And I guess it just fiddles and stones for the next four hours, but really impressive ability to complete the swing, get itself out of the situation where it's like, okay, Mike asked me to do this complex task overnight. I got stuck because this remote service went down. I'm going to write a scaffolded backend for it for now. So I'll document that, go all the way through. I have a good mental model of how far that's going to get me.
SPEAKER_02
And then when it comes back online, I'll fix it. I'll keep track of that fact. I think the most impressive thing for me is being able to delegate that kind of level of task and just trust that the right thing will happen by the end. And of course, you'll review the result and there's still a whole verification thing that we can and should talk about because I think that's an important part of still completing the swing there. But it's really forced me to rethink what is being productive with one of these models. And it is much more like we've talked for a while about what is it like when these models are more of a companion or a coworker?
SPEAKER_02
And it really feels like now it's a teammate that I can delegate a lot of work to. [SPEAKER_00] What is your day to day flow like right now? [SPEAKER_00] Because one of the things I notice is if you just give it a big task and you monologue into it and you just let it go for a few hours overnight, it's like the most impressive model that I've ever tried. [SPEAKER_00] But it's so slow and so expensive that I feel like I don't want to use it for day to day tasks. [SPEAKER_00] So what is your actual flow like in terms of how you use it day to day and where does it slot in versus other models?
SPEAKER_02
Yeah, I've ended up having a lot more architectural planning conversations up front with it as well. So that's been another interesting change where I think there's an era that I think all models need to continue to improve. And I'm really grateful for the Instagram experience of having to start from our initial version that was duct taped on a server in L.A. to being able to scale it and eventually integrate it with all of the Facebook infrastructure. Because you kind of develop a sense of what infra abstractions and complexity are appropriate for each stage of it. And I still don't always go back and forth with Fable where it'll be like, this is a good implementation.
SPEAKER_00
Yeah, I've ended up having a lot more architectural planning conversations up front with it as well. [SPEAKER_02] So that's been another interesting change where I think there's an era that I think all models need to continue to improve. [SPEAKER_02] And I'm really grateful for the Instagram experience of having to start from our initial version that was duct taped on a server in L.A. to being able to scale it and eventually integrate it with all of the Facebook infrastructure. [SPEAKER_02] Because you kind of develop a sense of what infra abstractions and complexity are appropriate for each stage of it.
SPEAKER_02
And I still don't always go back and forth with Fable where it'll be like, this is a good implementation. Like, well, I do plan on shipping this fairly soon. I think we should probably think about more than one server and that back and forth is important. But a lot of that planning and I'll often actually ask it. It's a thing I've realized is Fable can be so complete in its thinking in terms of how much you are planning with it. And often just saying, can you make an HTML page that represents what we just talked about so I can share it with the team? Is actually valuable or even just a markdown document. But I like having diagrams.
SPEAKER_02
So that's been an interesting use of let's plan with it, let's think it through, and then let's have some sort of document that we can align the team on. Because this is a dynamic I've seen in labs and just teams beyond Anthropic, which is you can build a lot very quickly. And forcing more of that early alignment, even if you do an initial prototype and then back it out into more of a plan architecture that works too, I think is really key. And it ends up being the place where the human to human interaction still stays very much part of the process.
SPEAKER_02
And then from then on, I think either overnight or during the day, having it execute on those chunks of tasks is really important. And it just means having a lot more concurrent sessions than I did before, because I often will think there are these two pieces of work. I go back and forth between having one very long running cloud code session and really asking it to do everything in background sort of forked subagents. So the main thread stays responsive. And then other times just embracing I'm just going to have like five or six tabs tackle long comprehensive work. But I do think that there's something to this long horizon and don't worry, I'm on it.
SPEAKER_02
It's going to take me a while. And more of this back and forth. And that modality, I think is something that we'll have to figure out in our products as well. I think you want to preserve both and they interact with each other in interesting ways. And my preference is usually I always have at least one cloud that is high context, but also very fast response. And its instinct is, I'm going to answer you and I'll kick something off if I need to. And if not, I'm just going to hang tight and wait for the next loop. I do think you're right that for the I'm just trying to fix this interaction question or something that's very fine detailed.
SPEAKER_02
Fable will go off and think very hard about those things. And I think Fable is the first model where I've actually played more with the effort levels for that reason. Where I've been like, okay, this is I just needed to tweak some UI. I'm not actually going to fall, but put it to medium or something and see how that plays out. Didn't find myself doing that as much with Opus, maybe because the range felt less wide, where it really can feel quite wide with Fable. What about a quick question? [SPEAKER_00] Like you're on the go, are you asking Fable random questions as they come to you?
SPEAKER_02
[SPEAKER_00] Because it feels like you're using a rocket launcher to kill a mosquito or something. [SPEAKER_00] Or are you flipping back and forth? It's so funny you ask that because I had been. And you're like, it's thinking and thinking really hard about it.
SPEAKER_00
[SPEAKER_02] Then last week, I was like, no, I was asking it something that I felt embarrassed actually asking Fable about. It was something probably NBA finals related. [SPEAKER_02] And I was like, okay, I switched my iOS app to Sonnet. [SPEAKER_02] I was like, yeah, I use this all the time for fast questions. [SPEAKER_02] It's a corner of magnitude in feeling. [SPEAKER_02] And it's actually not even the tokens per second.
SPEAKER_02
It's actually probably more around how much thinking goes into the answer. And sometimes the answer does not need to be fully thought through. So, yeah, I'm thinking myself through and I think this is a good product question for us too, which is in general you don't want people to have to be thinking so much about these choices. So ideally what we can sort of coalesce around in the longer run is maybe some more bucketable use cases that are really graspable to people.
SPEAKER_02
Or maybe it varies by surface where it's actually probably unlikely that most of the time with the iOS app I'm doing Fable type tasks and having a sticky model selection per surface might be the way to do that. And we'll have to explore what that means from a product perspective. But I've for sure had the feeling of this is not a Fable worthy question. I should ask Sonnet. [SPEAKER_00] Can you show us something that you've built with it? Yeah. So one of the things that we did this go around is we encouraged personal account usage for us, especially on the weekends, which is really fun because you can imagine a lot of Anthropic specific tooling, etc.
SPEAKER_02
But it was really good to step back and just work on something over the weekend using pure cloud code. [SPEAKER_00] Are you in the terminal app or you're in the desktop app? That's a great question. I'm mostly still in the terminal app. It's interesting watching my wife, who's not a professional engineer and more of a UX designer PM, really fall in love with cloud code via the desktop app. And I think it's simplified some of the abstractions for her in that way.
SPEAKER_00
[SPEAKER_02] But for this one, I was still in the terminal app.
SPEAKER_02
But let me show you. This is one of those everybody has some bespoke need around this. I wanted a good media tracker experience. And I was like, I'm playing games, I'm watching TV shows. I'm mostly still in the terminal app.
SPEAKER_00
[SPEAKER_02] It's interesting watching my wife, who's not a professional engineer and more of a UX designer PM, really fall in love with cloud code via the desktop app.
SPEAKER_02
And I think it's simplified some of the abstractions for her in that way. But for this one, I was still, is it ghosty or ghosty, do you want ghosty and the terminal app? But let me show you. I am. This is one of those where everybody has some bespoke need around this. I wanted a good media tracker experience. And I was thinking, I'm playing games, I'm watching TV shows. I get all these recommendations. I just wanted to build something that was personal to me and fit some of the use cases that I had. And the two biggest criteria that I started with were one, really easy to add things. So you can talk to Claude.
SPEAKER_02
Claude does the genetic search over everything and then puts the right things in. And then also proactively, there's a new season or a new sequel to a game that it could go off and research those things. Most of the UI was Fable one shot, which was already impressive. But then the thread I've been pulling out a lot in labs this year is how do you bring the software team, which is Claude these days, closer to the software itself. And so this was Saturday morning. I had a full weekend with kids stuff. So a lot of this was kick off work, go for a hike with the kids, come back, continue to do the work. Sometimes check in on the work on the hike. I probably shouldn't.
SPEAKER_02
But it was nice to pop into the remote mode and see what was going on there. Try not to do that too much. But I had this idea around, could we do a spike on what if you could actually modify the software from within itself. And I built both a React Native version and then this version is just the web version. So I already had a chat type thing where you can ask Claude to add things by URL, which is something I think every software should have where I should never have to navigate a menu to do anything ever again.
SPEAKER_02
And this is in many ways, Dan, the I was trying to distill agent native architectures to its fullest degree, which is also have the agent be able to modify the app. So phase one of agent native architecture, every single thing in this product is accessible from the agent and tool calls, et cetera. That's hopefully becoming table stakes. It was sadly not in a lot of software. And it's great because I was like, what's that? Somebody recommended there's a Brazilian show about radioactive stuff in Gleon. And I did not remember what it was called and Claude was able to figure it out. It was so much better than trying to figure that out intuitively.
SPEAKER_02
But then the next step I was interested in is what would it mean to actually be able to modify the software from itself on the go? And so if you long press this little chat thing. What I built, what Claude built, was a way where it uses our managed agents to basically take on edit requests and then you can preview them. And I used the Vercel live preview thing here. This whole feature was also one shot, which was really cool. And I just added to it over time. But it actually does a little diff view if you wanted to. You can go into the managed agent conversation and see what it did.
SPEAKER_02
Although I almost never do because, especially, I don't particularly care about the code quality or the long term maintainability of this software. You can see that it had a session in here too. But it's been really fun because I'll be using it on the go and say, I had a feature request the other day, the floating action button was too low on native iOS, but it was okay on there. Can you go off and do it? It did it. It was really fun with some of the Expo tooling now and actually live reloaded on my phone, which was also a really cool feeling. But it was just, does this thing need to be a production level thing that's going to go to a million users?
SPEAKER_02
No, but it felt really good to have something where I felt like it didn't have to stop at just the weekend and I can keep working on it just by using it and having this end to end close thing. So I felt this was a good manifestation of both Fable's building ability, but also a lot of what both you and I have been thinking about, how does Claude embed itself into software beyond just the usage set of things. [SPEAKER_00] And I want people to understand so this has been built, you could build something like this, maybe not the self modifying part, but you could build something like this for ten years or twenty years or something like that.
SPEAKER_02
[SPEAKER_00] But the cost to build has gotten dramatically lower. [SPEAKER_00] So think about how much it would have cost to do this in the Instagram days versus now. Can you help us understand how that has changed? Yeah, I think about this a lot when I think back to that time as well. I thought of myself as a very productive programmer in the early Instagram days. I was really into mobile development and we had good clarity of things. And I think the gap from idea to fully realized version of some complete product, you were still looking at four-ish days of all nighters.
SPEAKER_02
I call it Instagram V1, which probably had more features than this thing did, but not by an order of magnitude, was five days of all nighters.
SPEAKER_00
[SPEAKER_02] Me working on the front end and back end and Kevin working on the initial filters to get that out. [SPEAKER_02] And this was also built on many years that I've been working on iOS pieces as well. [SPEAKER_02] And then the iteration, I think a lot about what we were gated on after that launch when things went well was we had all these ideas for where to take it, but we were just trying to keep the site up or we were just trying to add the one incremental feature.
SPEAKER_00
[SPEAKER_02] But yeah, I call it Instagram V1, which probably had more features than this thing did, but not by an order of magnitude. It was like five days of all nighters with me working on the front end and back end and Kevin working on the initial filters to get that out.
SPEAKER_02
And this was also built on many years that I've been working on iOS pieces as well. And then the iteration—I think a lot about what we were gated on after that launch when things went well. We had all these ideas for where to take it, but we were just trying to keep the site up or just trying to add the one incremental feature. Hashtags take a week to build, but then there's all the things that you want to continue doing on it as well.
SPEAKER_02
And so I think it's both that shortening of time. There's still the time required for the idea and the concept and the iteration. And then the other piece, which is the way you can then iterate on what you have. And I think a really fun, but also very in the flow kind of way.
SPEAKER_02
And then, if now this is me as a professional software engineer and startup founder, beyond that, if you had that idea, I saw multiple people go through this and I guess I'd try to find maybe a consultancy that will take this on. But now there's a really lossy process of what I wanted. There are gonna raise money for it. And I think the thing that I think is the most exciting part about these models getting not just more autonomous, but again, closing that gap between intent and execution is what I've seen it do to people's ability to build who are not builders.
SPEAKER_02
And the trajectory of these models has been—if something of this general class is in that class of models and eventually models that are cheaper and more accessible to other folks become available too. And as that process happens, I just think it is opening up so many things. I got a ping the other day—I get very excited about this stuff. You can tell somebody internally, and we had built them an internal tool that combined Fable and access to some internal MCPs. And she said it is the first time in my life and she works in recruiting. And she's like the first time in my life where I feel like the thing that's in my head and the thing that exists in the world is now right next to each other. I can just do it.
SPEAKER_02
And it was a very meaningful moment to her because prior to that, I mean, I remember these days were five years ago or four years ago where that person, if they wanted a tool, would have to either make do or try to get an internal tools engineer that probably was overloaded with 50 other requirements. But instead now they are having the time of their lives building. And I think that is cause for a lot of hope because I don't think that human capacity for creativity and what's possible is enormous. And I think at our best, we are expanding the number of people who can then see that through to something that feels real.
SPEAKER_02
I totally agree. But I do think that there's a question in the back of my mind and I think it's probably going to be in the back of the minds of some of the people listening. So I want to ask you, given everything you just said, is software engineering over? [SPEAKER_00] Yeah, I think software engineering is different. It is dramatically changed. And as I probably would have defined it if you had asked me around the Instagram time, like what is software engineering? I'd probably say thinking through the hard problems and thinking about architecture and then spending a lot of time in a text editor.
SPEAKER_02
I can't remember, but a text editor you're going to edit those in, or Xcode, and watching Rails. Yeah, exactly. Right. And understanding the intricacies of Django or layer and then 15 bugs after you deploy it. So much of that is radically different and collapsing into other parts like product management.
SPEAKER_02
I think that PM split, I think even in our teams has become much more diffuse. That's radically changed, but I think the overall—maybe zoom out from software engineering and think about software production or software development, but not in just the pure developer case. I think that is alive and well and essential still. So I think that is the moment that I feel like we are in.
SPEAKER_02
I think Fable is another step on the direction of—and I'm not going to call it the final step. Of course, a lot will still happen, but I think a pretty significant step in terms of the trust, at least I end up placing the model in terms of its capacity to see things through and even architect things reasonably is quite high. So that part feels like it is not ever going to be done, but it is pretty done, right? It's gone really far, but I think that the overall craft of the—what needs do you have? What are you putting out? Is it actually good? I think still a very human endeavor, but I also can see that is not a transition that is pain free in a way.
SPEAKER_02
I think there are plenty of people who love the craft of actually putting—I used to love this stuff. I solved that problem so elegantly. You would dream about code. And if you ever had that experience, you would dream about the thing that you're working on. They wake up in the morning and be like, I figured out how to solve this thing really elegantly. And that for sure has passed. And I think there's a feeling of loss, I think, in some of the better engineers that I talk to, as well as the feeling of, oh my God, but I can do insane amounts of work now at the same time.
SPEAKER_02
So we're holding both ideas in our heads at once, I guess. Which I think is the most important part of this. It's normal to feel sadness for that kind of thing and excitement. But I'm curious, let's just take the thesis of software engineering is alive and well. What does that actually look like inside of Anthropic? [SPEAKER_00] Yeah, I think there's a few theses. I think there's still the crafting of—well, I kind of take it off from the full software development cycle or maybe what I see day to day. Maybe I'll do a little bit of both. But I think there's still a lot of—you know, we all got together. Which I think is the most important part of this.
SPEAKER_02
It's normal to feel sadness for that kind of thing and excitement. But I'm curious, let's just take the thesis of software engineering is alive and well. What does that actually look like inside of Anthropic? Yeah, I think there's a few theses. I think there's still the crafting of. [SPEAKER_00] Well, I take it off from the full software development cycle or what I see on a day to day, maybe I'll do a little bit of both. [SPEAKER_00] But I think there's still a lot of, we all got together. [SPEAKER_00] We talked about the next way we want to evolve co-work. [SPEAKER_00] And now we've broken it down into areas of ownership.
SPEAKER_00
[SPEAKER_02] I think that ends up still being quite important because there is still context that you hold as a person that is beyond cloud, right? [SPEAKER_02] What is the actual intent of this product? [SPEAKER_02] How's it going?
SPEAKER_02
What do we need to know about the other products that are coming down the pipeline that are going to be integrated in some interesting way? So I think that aspect is really important still. And so though we have many clods to each human, each human, at least the way we've been working on Anthropic still has, we call them DRIs, directly responsible individuals still has a DRI ship over some part of the product or some area. I think that'll be the case for a while because I think there is value in not just this distributed, we should all make co-work better.
SPEAKER_02
But instead, all right, I'm thinking through how co-work does this particular task and there's still a lot of, you try to keep meetings minimal, but they still emerge and you still have these alignment conversations. Then there's a lot of that asynchronous delegation. I think what many engineers here have now found is they've all built, and I think we should solve this at some point at a broader product level, but they've all built some version of, all right, I'm going to now create a dashboard of where all my clods are doing. And what's waiting for me and which pull requests need my attention because either a human or a cloud code reviewer got back to me.
SPEAKER_02
So there's a lot of that meta maintenance of the work that I think, again, I think we'll standardize some, but I think some of it will always be a little bit bespoke to the way each individual likes to work just in the way that people organize their windows. Now they organize their work. And then there is, I think also the understanding how things work in production. And I think that is another, there's a few next frontiers, I think for the models. And I think one of them that Fable does make significant strides in, but I think there's more work needed here is understanding what happens to code after it gets deployed, because there's incidents.
SPEAKER_02
There's this was all working well, but this network link got cut, which is not in your usual failure mode. And it manifested so much of Instagram, 2012 to 2016, it was dealing with that and scaling things up. And so that role of the engineer still remains really key. And I think getting the reps in around incident response and understanding how to stay calm, gather data, remediate what's immediate, but then go off and work on longer term fixes, still a necessary part of it. And I'm trying to think if there's any other pieces that are notable as well. I think what's maybe the last thing to say is I really like the role that the engineering prototype now plays.
SPEAKER_02
You have to be clear when it's a prototype versus not. But the old phrase was code wins arguments. And I never loved that because the person that could code could go do it, but actually why should they necessarily win an argument by default? But actually it's been really cool now where sometimes we will have some disagreement or debate about where to take a product. And often it's the PM that will say, all right, I just tried it. And it's janky in these eight ways, but look, it actually shows how this could work. And that can open up some interesting pieces of conversation. So almost all of that is quite different than it was six months ago.
SPEAKER_02
I think, especially at the level of parallelism and the level of need for these higher order abstractions of work.
SPEAKER_00
[SPEAKER_02] But I think what hasn't changed is that ownership. [SPEAKER_02] Lots of us are shipping AI to production, which is great for productivity, but it also comes with anxiety. [SPEAKER_02] You tweak a prompt, swap models, adjust parameters, and everything looks fine in testing. [SPEAKER_02] So you merge.
SPEAKER_02
And then three days later, or even sooner, the support tickets start rolling in. The AI is giving your customers unexpected answers and you have no idea when it happened or why. BrainTrust is the AI observability platform that fixes this. [SPEAKER_00] It connects evals and observability in one workflow. [SPEAKER_00] That way you see what actually happened in production and can measure whether changes made things better or worse. [SPEAKER_00] Traces show the full execution path, evals define what good looks like, and experiments let you compare prompts and models side by side before shipping. [SPEAKER_00] Production traces feed directly into your eval datasets.
SPEAKER_02
[SPEAKER_00] Every failure becomes a test case, you catch regressions in CI before they reach users, [SPEAKER_00] and teams at Notion, Stripe, Zapier, Vercel, and Ramp use it to ship quality AI at scale. [SPEAKER_00] BrainTrust is designed for teams building production AI systems where silent regressions are expensive. [SPEAKER_00] It's built for any stack. [SPEAKER_00] They have SDKs for Python, TypeScript, Go, Ruby, C Sharp. [SPEAKER_00] There's no framework lock-in or vendor dependencies. [SPEAKER_00] It's SOC 2, Type 2 certified, and GDPR and HIPAA compliant. [SPEAKER_00] Get started at BrainTrust.dev. [SPEAKER_00] That's BrainTrust.dev.
SPEAKER_02
[SPEAKER_00] And now, back to the episode. [SPEAKER_00] Fable is also very expensive. [SPEAKER_00] And because of that, when I was testing it, I felt I was a kid in a candy shop, [SPEAKER_00] and I was just, I'll do this, and I'll do this, and I'll do that. [SPEAKER_00] But now that there's going to be a bill, I'm going to be thinking about it, [SPEAKER_00] because I have to pause before I do it to be, is this going to cost me $100 or whatever? [SPEAKER_00] And I do think that's going to limit who gets to use it and for what. [SPEAKER_00] So how do you think about that? [SPEAKER_00] Yeah, I think it's most clear-cut on the professional software,
SPEAKER_02
[SPEAKER_00] sort of classic company doing work. [SPEAKER_00] It'll be really interesting. [SPEAKER_00] And I was just like, I'll do this, and I'll do this, and I'll do that. [SPEAKER_00] But now that there's going to be a bill, I'm going to be thinking about it, because I have to pause before I do it to be like, is this going to cost me $100 or whatever? [SPEAKER_00] And I do think that's going to limit who gets to use it and for what. [SPEAKER_00] So how do you think about that? [SPEAKER_00] Yeah, I think it's most clear-cut on the professional software, classic company doing work.
SPEAKER_02
[SPEAKER_00] It'll be really interesting. There's a lot of process that goes into pricing as well. [SPEAKER_00] It's both more expensive than Opus, and then also I'm thinking in many ways, it's really cheap. If you think about how much incredible work it's doing. But of course, everybody has their own economics around what they're working with. So anyway, most clear-cut, I think, from most software teams. And I think as an industry, if phase one was companies even struggling to get some of their employees to adopt AI coding, which models were early, maybe the tooling wasn't there.
SPEAKER_00
[SPEAKER_02] And then phase two was great. We'll create leaderboards and see who can use the most, which, you know, as you can imagine, creates some not ideal incentives. [SPEAKER_02] To phase three, where we were like, okay, now we're just trying to figure out who's using it effectively and letting them spend as much as possible and having a clear process for that, but making sure we're not doing things wastefully, which I think to me in general makes sense. [SPEAKER_02] Although I think you could also over-rotate that way too.
SPEAKER_00
[SPEAKER_02] I think something of Fable class should hopefully fit in well into that, where if you're demonstrating results and you're getting use out of the model, then hopefully there's a flywheel even inside companies where that goes and perpetuates that. [SPEAKER_02] I think on the personal use side, it's a really good one. That's a really good question. [SPEAKER_02] I think where I've seen it, even in my personal testing, because our personal accounts, okay. Which is funny, paying my own company I work at. [SPEAKER_02] But you do become more thoughtful about it.
SPEAKER_00
[SPEAKER_02] Something that was interesting was this app that I built over the weekend actually fit in with only a bit of extra usage. [SPEAKER_02] So it wasn't thousands of dollars to build this thing that is a personal thing to myself. [SPEAKER_02] But it was also spaced out a little bit more. [SPEAKER_02] Probably the in-between of that, what we'll have to do the most thinking about is the sort of hobbyist or independent who's not within the larger company, but also is thoughtful about the pricing as well. [SPEAKER_02] I think my overall advice is just give it a try and see how much it can do without you having to do a lot of follow-ups.
SPEAKER_00
[SPEAKER_02] And I think measuring cost has gotten so multifaceted now because there's the per turn costs. And then there's what did it cost you not to do the task, but complete the task to your satisfaction? [SPEAKER_02] And I think that's where Fable has really shined for me, which is it actually just does it right so that I don't have to spend the nine, 10 subsequent turns be like, no, that was not quite what I meant. Can you also do this piece?
SPEAKER_00
[SPEAKER_02] It's been really impressive for me because you ask it to go do something and then it just does a thing. And you're like, wow, you thought through all the little details of this thing in a way that I've never seen another model do. I don't know how much you can reveal about the training process, but what makes the model different? I mean, I think in many ways, a continuation of a lot of the work that the team has done. And I bow down in total awe of our teams, both on the pre-training and on the RL side.
SPEAKER_00
I think the piece that it has evolved in, at least I noticed the most, is adjacent to that as well, which is a sense of the system more than just the individual piece of the work. [SPEAKER_02] Like I will often be very positively surprised when it will write something and say, all right, but I know that in production, this needs to be different. And then it will keep bugging you. Like, have you turned on that feature flag yet? It's not going to work until you do.
SPEAKER_00
[SPEAKER_02] And sometimes in sessions that have gone on for days, be like, look, you still haven't done that thing. Like you better, I was like, you're right. I didn't turn on that feature flag. I should go off and do that. [SPEAKER_02] Or if we change this, the contract will change over there. We're watching it. [SPEAKER_02] Actually, one of my favorite times of seeing it in action, I think where it demonstrates some of the training is watching it respond to code review feedback, either from people or from other cloud reviewers, where it doesn't just say, oh yeah, that's an issue. I'm going to go fix it.
SPEAKER_00
[SPEAKER_02] And actually really thoughtful around, hey, for this level of fidelity of what we're building, I'm going to accept this risk. Or I see what you mean. [SPEAKER_02] Other code reviewer, which is often just another Fable model, talking to you, I see what you mean. But I'm actually going to push back. I don't think that's actually right. I think that's not right. [SPEAKER_02] I think getting the model to have that judgment is really important. [SPEAKER_02] And I think if I had to pinpoint an area where I feel like it's really progressed, it is that sort of not just immediate knee jerk. Yeah, that's right. I got to go fix it.
SPEAKER_00
[SPEAKER_02] And more, I'll think about that for a minute. No, I thought about it and I still disagree.
SPEAKER_02
And I think that's a very useful ability. It's so valuable to have products like Cloud Code out there because you have now a living, breathing thing where people are like, this is where the model is doing well. And we have people who test it. I count the Anthropic folks as very, very high on the list. Is that not just an immediate knee jerk reaction? Yeah, yeah, that's right. I got to go fix it. And more, oh, I'll think about that for a minute. No, I thought about it and I still disagree, and I think that's a very useful ability.
SPEAKER_02
It's so valuable to have products like Cloud Code out there because you have now a living, breathing thing where people are like, this is where the model is doing well. And we have people who test it. I count the every folks is very, very high on the list. We're like, we really trust the feedback because it is being put to paces and repeated multi-day hard tasks. And that also very much feeds into how we think about what do we need to improve on the next slide? What are the tasks that we need to specifically think about the model being better at? Is chat the right interface for this model? Because it's not very turn by turn.
SPEAKER_02
It's very like I'm delegating something for you. So how does that change how you should use it or how you think about the interface? I don't think the fundamental, like you are sending messages and it is giving your message back is totally wrong. [SPEAKER_00] I think that there's ways we need to evolve. [SPEAKER_00] But one is maybe three that come to mind. [SPEAKER_00] One is your laptop the right place for it. [SPEAKER_00] So I think that's number one where I mentioned with the side project I was working on how useful it was to have the mobile side. Boris, who created Cloud Code, he's always ahead of the curve on how these models get used.
SPEAKER_02
Almost a year ago, maybe nine months, I was talking to him. He's, yeah, I've moved a lot of my Cloud Code work to mobile. I was, no way. And it took me a while to get there. But especially with the Fable class, there's oftentimes where, because it can keep the session going and we use remote dev boxes at Anthropic, it is like I'll have a thought and be, OK, I need can you keep up and doing that? So number one is decoupling the where the work is happening from where I'm talking to about the work. The second one touches a little bit on what I was mentioning earlier around what are how do you take everything that Fable has discussed or decided?
SPEAKER_02
Or proposed about something and make it comprehensible. And that's an area that we're thinking a lot about. There are some skills that are out there that we've used around all right, can you diagram this? Can you do that? So that's a place where the current chat UI, I think is insufficient, where it will experience this with people. It will give you a lot of text. You're, this. I need to take a walk before I'm ready to fully understand this. And I think that that is a piece of property. I have some things we'll do with Fable's, OK, you have a lot more context on this than I do. Can we back it up? Let's do more progressive disclosure of the complexity here.
SPEAKER_02
So I think that piece is interesting.
SPEAKER_00
[SPEAKER_02] The last one that I think is we're still early in pulling on is thinking through multiplayer where, at some level, these abstraction levels and because we have this DRI and ownership area, usually a chunk of significant work, a human and a couple of clods like that is still flowing together. [SPEAKER_02] But in other cases that is less the case, where it's an incident response where multiple people are thinking about it.
SPEAKER_00
[SPEAKER_02] Maybe it's a project where there's multiple competing or not competing, but conjoining areas that are coming together and thinking through what would it mean for, and we have chat sharing, which gets you a little bit of the way there. [SPEAKER_02] But I think there is going to be a need for more, all right, you've got an independent club that's doing a lot of work that was kicked off by somebody. [SPEAKER_02] But can it be keeping up with all the other work happening on the team? [SPEAKER_02] I think that is an interesting and underexplored next frontier about how this work ends up happening.
SPEAKER_00
[SPEAKER_02] But I think it's really exciting because I think, again, it's the level of team-made collaborator that the models are now capable of and we're almost holding them back by not having the right abstractions around them for that to happen.
SPEAKER_02
Yeah, it makes me think I've mostly been using this for my own vibe-coded stuff. So I haven't really had to think about this, but there's a problem when you're using this inside of an organization, which is, do I really understand every part of this? And therefore, how do I transfer the context of what the model just did into my brain? That's one of the big bottlenecks. How do you think about drawing the line, especially with a model like this, around how much you actually need to understand and how to make sure that you have enough context on what it's done to feel comfortable? [SPEAKER_00] I think there's two big pieces here.
SPEAKER_02
[SPEAKER_00] The first is verification, where I became fully verification-filled earlier this year and now, almost in the same way, and actually it connects to how I think I used to do when I was typing code more full-time, which is try to find the tightest dev loop that you can around the idea that you're trying to develop in. [SPEAKER_00] Sometimes with Instagram, that meant actually making a new build target in Xcode that was just that screen with some synthetic data and just doing that dev loop. And I would mentor newer engineers. If there's one thing that I can impart on you, it is try to get that for any project you work on and things will go much more quickly.
SPEAKER_02
I think that is no longer exactly the case here, but I think what is the case now is anytime I set it up, how do I get for every pull request that Claude is putting out that there is an attached photo or video, whether that's an iOS PR, whether that's something in the UI. And that's, I think that helps you gain a lot of confidence because even now, you might have Fable go off and do work for a couple of hours and be you work on and things will go much more quickly.
SPEAKER_02
I think that is no longer exactly the case here, but I think what is the case now is anytime I set it up, how do I get for every pull request that Claude is putting out that there is an attached photo or video, whether that's an iOS PR, whether that's something in the UI. And that's, I think that helps you gain a lot of confidence because even now, you might have Fable go off and do work for a couple of hours and be like, I'm done. And it's really useful to say, and here's the full screenshot gallery of the full UI. Cause you might say, oh, you know what, on screenshot eight, that error state, I've never actually seen it, but I can see how a person might hit it.
SPEAKER_02
Let's actually make that different. And so getting that comprehensive verification, I think it's something we've been working on a lot internally and publishing more and more skills and knowledge about, but I think it's really a key piece there. And then the second one is, I think you ultimately as a person still need to stand behind the work that you are doing, especially if you're putting it into a production system. Like a lot of people use Cloud every day. There's still the accountability of like, although it's still Cloud better written a bit, you need to understand the general decisions that were made on these pieces as well.
SPEAKER_02
And so I have seen a fair amount of engineers actually adopt this practice where Cloud will have done the work, but then there is the follow-up conversation around, well, can you, can I make sure I deeply understand all the trade-offs that you've made and whatever artifacts need to be produced in order to make that comprehensible is important. It is really interesting though, to be in meetings where somebody will say, oh yeah, and I have this PR ready. And somebody else has to be like, oh, that's interesting, did you do X or Y and have that moment of positive? They're like, you know what, I'm not entirely sure I will find word before we merge this PR.
SPEAKER_02
And that's, I think that adapting to that norm and figuring out work with that is something we'll have to do. Tell me more about the verification. It's such a hot topic right now. It sounds like one way that you do that is with screenshots and screen shares, but what are the other ways that you think about that? I think part of it starts in, can you get to a place where you are exercising real flows that aren't just a static injected piece and this thing gets more complex, that gets more and more complicated.
SPEAKER_02
[SPEAKER_00] So we've invested a bunch into even just getting it so that the iOS app can log in to staging on a real account and have real data, but you don't want it to then go through an eight stage onboarding process every time, but you're just trying to test the second part of the screen. So there's a lot of work around how do you, is there a special affordance, is there some shared secret, whatever that is around getting the app to really feel as human, using the product as possible. So that's one aspect of it. The second is this mix of well-known paths versus the things you're exercising in the exact moment, the former being really useful for regression testing.
SPEAKER_02
And so we don't think of places where we've expressed ideal workflows in text basically, and Claude can repeatedly check that. And then there's also, and Claude does a really good job of this sort of expressing the intent of the current change at hand. So that gets really deeply exercised. So I think that the combination of those two things is important. The visual verification that I mentioned as well, video has been really cool to see. Actually video is a very underexplored tool to give Claude as well. I think I've been prototyping is just giving Claude video captures of the thing that it has built and then giving it an FFM tag and you'll watch it scrub through.
SPEAKER_02
And so this animation has some jank in it, I'm going to go fix that. And I would never be able to do it with a screenshot sort of latency capture because it will have missed the moment. So I think that's another piece that is really important. And then for the pieces that aren't easily testable and tend because there is some more complex system, getting Claude to go and build as robust a mock backend as possible or use ones off the shelf has been also really interesting. Like when I think about artifact, we had really comprehensive tests. This is kind of pre LLM.
SPEAKER_00
[SPEAKER_02] And one of the ways that we were able to do that really robustly was that basically every piece of info we had, whether it was Postgres, Redis, all the AWS things had a really good in memory implementation that you could just do really quickly in unit tests and kind of extending that to Claude land. [SPEAKER_02] Now, I was working on something where it had a pretty robust backend and for kind of complicated reasons, hard to spin that up on my dev server, but it was able to, again, one shot a really good proxy for that, by proxy, I mean a substitute for that. [SPEAKER_02] And that was so valuable.
SPEAKER_00
[SPEAKER_02] And over time, it's been interesting as that substitute has evolved as the rest of the code has evolved, which is the thing that, you know, if you had pitched that idea to me before, I'd be like, well, that's going to be really hard because the upstream is going to change, how are you going to keep it in sync?
SPEAKER_02
And I don't think about that anymore. I'm like, yeah, Claude will read the changes and it will adapt the thing and it'll keep the two in sync and that's fine. There's some really interesting architectures around when you get a bug, it just automatically goes out and closes it.
SPEAKER_02
And over time, it's been interesting as that substitute has evolved as the rest of the code has evolved, which is the thing that if you had pitched that idea to me before, I'd be like, well, that's gonna be really hard because the upstream is gonna change. How are you gonna keep it in sync? And I don't think about that anymore. I'm like, yeah, Claude will read the changes and it will adapt the thing and it'll keep the two in sync and that's fine. There's some really interesting architectures around when you get a bug, it just automatically goes out and closes it. You know, the agent just gets kicked off, it closes it and then it sends a message to the customer being like, it's fixed. Are you noticing a fable any change in how that process works?
SPEAKER_02
[SPEAKER_00] Yeah, I think there's a couple of things on a very human to human or human to Claude level. One of the things that I've seen it do better other models of the cable, I just need to do it really consistently too, is if the bug report, for example, came from somebody mentioning something in our feedback channel in Slack.
SPEAKER_02
And then the thing that got fed into the cloud code session is like, oh, there's this and because of the Slack MCP, you can actually pull the thread. Have it then actually post back, you know, as me, it'll be like, Hey, this is Mike's Claude. Like I fixed it. Here's the pull request. But then I think in the previous clouds, the thing it does really well is then say, but hold tight. It's not in production yet. I'll follow up when it actually is. And then maybe a few hours later, like, oh, this deploy went out. Like you should go test it. Is it fixed now? That level of follow through, I think is new on closing the loop piece. And it's five, I definitely have these long running cloud code sessions that are basically interacting as me, I guess, but some disclaimer in there too. And the second goes back to that taste and discernment piece that we were talking about, which is like, it's one thing to say, there was a bug report. Therefore I must go fix this thing. And it's another one to say, you know what? The, I hit this over the weekend, one of our internal systems basically had been running without restarting for a while. There was a memory. And it was a good discernment of saying like, all right, Mike, it's the weekend, just rebounce the server. It's going to solve it for now. And we'll work on the, well, asynchronously get the PR going to fix this more longterm. So I think if you're going to have cloud in the loop in this close the loop bug report or system issue to change, I think you really want it to understand where, as any good SRE or engineer in the loop would, great, let's solve the problem at hand. Let's defer the question of, do we need a re-architect on top of a completely different language found and understanding that balance is really important. One of the things that's really exciting, mostly exciting to me about new models is it raises the floor so that everyone can go build apps in one shot. But it also raises the ceiling for experts. So if you're a software engineer or founder, you can go do things that you never would have been able to before because you have access to this really powerful model. So for me, I bought this one shot version of Borges, infinite library. It's like a 3d game version of the library. It's wild. It runs right in the browser. It's so good. I can find any, every essay inside of it. I'll send you the link. It's sick. But I think there's going to be this flowering of people doing things like, Oh, I made a game or maybe I trained a new model or whatever that they couldn't do before. And I'd love to give people some inspiration, some examples of things that they might be able to do that they might not be thinking to do with this model. What are some ideas that come to you?
SPEAKER_02
[SPEAKER_00] Yeah. I think a few, maybe I'll start with the fun side and riffing off the game piece. I think people have a lot of creative ideas for how do they express the complexity of what they are, their world. Like everybody has the thing that they know really well. And there's probably some level of how do I then explain that to somebody else?
SPEAKER_02
Or how do I apply techniques elsewhere that I could go off and do? My wife is studying environmental engineering, studying geothermal, really complex math and simulations. And I've seen as the models have gotten better, she has been able to apply even more complex techniques from even outside of that domain into that work. And I think what people should be able to do, full on PyTorch end to end simulations of that work in a way that wouldn't be possible. I think that maybe is one, bring the beautiful complexity of what you have and either show it to other people by maybe making a game or maybe making a visualization, which I've seen her do as well, or at least make, bring other techniques to bear. And the second piece is its ability to compose software that solves a really unique problem to you. And I've seen that internally. A lot of the work that we've been doing is how do we get as many of our internal systems MCP-ified with the right permissioning structure and the right deployment set up. Although externally, you have good options around some of these platform as a service pieces and you can just ask a lot about them and they'll help you set things up. But I love that feeling of that thing that you always wish that you had.
SPEAKER_02
to bear. And the second piece is its ability to compose software that solves a really unique problem to you. And I've seen that internally. A lot of the work that we've been doing is how do we get as many of our internal systems like MCP-ified with the right permissioning structure and the right deployment setup. Although externally, you have good options around some of these platform as a service pieces and you can just ask a lot about them and they'll help you set things up. But I love that feeling of that thing that you always wish that you had.
SPEAKER_02
And then what has blown my mind, there was a person who works in our go to market organization who has been building this really deeply thought integration of cloud into every part of her whole process. And you don't have to stop at that one shot. Like she's been working on it for months now and she can keep going. And I think one of the things that is maybe underappreciated about the models is I think in previous generations, it would eventually get to a complexity level where it was hard to iterate on it without feeling like you then would break the thing that they had, you know, like under or over abstracted.
SPEAKER_02
Whereas this is actually, you know, she's got access to something Fable or Fable like for a couple of months. And like you've just seen it keep growing and growing and growing and growing. And now she's deploying it to the whole GTM org. And I think that is really cool. The ceiling of complexity that a person that does not start out as technical can now build for solving problems within their domain is unprecedented. I agree. It writes great code. My benchmark that I have is called the senior engineer benchmark. I just have it see if it can rewrite a code base from first principles and the nearest model that the previous top was like a 62 or 63 out of a hundred.
SPEAKER_02
And this model got a 90 on the benchmark or 91, which is human senior engineer level. [SPEAKER_00] Like you can just keep going with this thing in a way that's really fantastic. [SPEAKER_00] I'm curious though. One of the things that's really powerful that you mentioned is dynamic workflows. [SPEAKER_00] Tell us about that. [SPEAKER_00] This is, you know, we'll build things internally sometimes, and I will go really aggressively bug the engineer who built it and be like, when are we shipping this publicly? [SPEAKER_00] Because I think people are going to really like it.
SPEAKER_02
[SPEAKER_00] I think there's a good reason why it was built internally, but we try to ship as many of these as possible. And dynamic workflows was definitely that to me. The person who built this is an engineer named Sid, who's awesome. And I was like, Sid, I want to get this out into the world because it's so good. But I think it's especially good with a model like Fable for two really big reasons. One, it helps create the scaffold for deep, meaningful work. The craziest dynamic workflow I did and used Fable for was I had an internal project that we had written in Python, but we needed it actually in TypeScript for a really specific deployment reason.
SPEAKER_02
And having been internal to Instagram and we were like, should we write the whole thing into Hack and port it to the PHP engine that Facebook, I was like, you never would have done that. Maybe they can now with the model, but at the time it seemed impossible. But here I had pretty complex code base. And I was like, I'm just going to set up a dynamic workflow and just let it run over the weekend. And it did. And the workflow was so cool. It was like, all right, I'm going to do deep understanding of the work. I'm going to create a spec of how everything works. I'm going to go module by module. I'm going to translate these pieces. I'm going to have tested incrementally.
SPEAKER_02
I'm going to do another adversarial test. I'm going to check for anything that I missed. And it was really cool, a series of steps that the workflow was able to orchestrate. And I came back and I was like, yeah, this thing is TypeScript and Bun port of that thing. And it's actually better in these ways. And it was very documented, like these were the things I couldn't port, but most of these were very specific to the specific implementation. It wasn't worth porting. And I do not think you could have done that A with previous models at that level of success and B without the kind of scaffolding that or close provide.
SPEAKER_02
So I think that is extremely exciting, this combination of model capabilities and then our own ability to orchestrate them over longer time horizons with that feeling of like you had a goal, you broke it down effectively and then you were able to make it work.
SPEAKER_00
[SPEAKER_02] The other piece is I think over time, we'll be able to make some of those subtasks tuned to have the model be tuned to the level of complexity of it. [SPEAKER_02] So you can imagine that some parts of dynamic workflow don't need extra high thinking. [SPEAKER_02] They could use a medium thinking to get it done or even a smaller model. [SPEAKER_02] And I think that's really the future of where these things are going. [SPEAKER_02] So yeah, I'm a huge workflows DAU. [SPEAKER_02] For people who haven't used it before, tell me about how you got that workflow made. [SPEAKER_02] How did you design it? [SPEAKER_02] How did you make sure it was good?
SPEAKER_02
Yeah, it was pretty iterative, but I just started with cloud code. Like, Hey, I have this complex task, let's design a workflow to go and do it. [SPEAKER_00] It kind of showed me the plan. [SPEAKER_00] I was like, Oh, this is close to what I want. [SPEAKER_00] I want to make sure that you do these three or four levels of additional verification for missed features. It's like, here's what you have. Are you ready to go? How did you design it? How did you make sure it was good? Yeah, it was pretty iterative, but I just started with cloud code. I have this complex task, so let's design a workflow to go and do it. [SPEAKER_00] It showed me the plan.
SPEAKER_02
[SPEAKER_00] I was like, oh, this is close to what I want. [SPEAKER_00] I want to make sure that you do these three or four levels of additional verification for missed features. It's like, here's what you have. Are you ready to go? And it expresses the workflows in code, which I think is really valuable to see what it was about to do. And what was interesting is it did the full port. And then I had a couple of follow-up questions or little tweaks. And I did those as mini workflows that built off the previous one as well. But I think we talked a little bit about whether chat was the right interface. We've had that conversation over the last year.
SPEAKER_02
And I think workflows are a good middle ground. You can compose them using chat, but they're expressed using code. And then they're executed with a nice clean UI around what's happening at every stage. I think we'll start bridging longer horizon work with chat in ways like that over time. Mike, this is such a great conversation. Thank you so much for joining and telling us all about this new model. I'm really excited to get to spend time with you and really looking forward to what people think outside too. Oh my gosh, folks, you absolutely positively have to smash that like button and subscribe to AI and I. Why?
SPEAKER_02
[SPEAKER_00] Because this show is the epitome of awesomeness. [SPEAKER_00] It's finding a treasure chest in your backyard, but instead of gold, it's filled with pure unadulterated knowledge bombs about ChatGPT. Every episode is a roller coaster of emotions, insights, and laughter that will leave you on the edge of your seat craving for more. [SPEAKER_01] It's not just a show. [SPEAKER_01] It's a journey into the future with Dan Shipper as the captain of the spaceship. [SPEAKER_01] So do yourself a favor, hit like, smash subscribe, and strap in for the ride of your life.
SPEAKER_02
[SPEAKER_01] And now without any further ado, let me just say, Dan, I'm absolutely hopelessly in love with you. They're like, you know what? I'm not entirely sure I will find word before we merge this PR. And that's, you know, I think that adapting to that norm and figuring out and work with that is something we'll have to do. Tell me more about the verification. It's such a, it's such a hot topic right now. It sounds like one way that you do that is with screenshots and screen shares, but what are the other ways that you think about that? I think part of it, it starts in, can you get to a place where you are exercising real,
SPEAKER_02
like sort of real flows that aren't just like a static injected piece and this thing gets
SPEAKER_00
more complex, that gets more and more complicated. So we've invested a bunch into like even just getting it so that the, you know, the iOS app can log in to staging on a real account and like have real data, but you don't want it to then go through like an eight stage onboarding process every time, but you're just trying
SPEAKER_02
to test like the second part of the screen. So there's a lot of work around like, how do you, you know, is there a special, affordance, is there like some shared secret, whatever that is around getting the, the, the, the like app, you know, to really feel as human, you know, using the product as possible. So that's one, one aspect of it. Um, the second is like this mix of like well-known paths versus the things you're exercising in the exact moment, like the former being really useful for regression testing. And so we don't think of places where we've expressed like, uh, sort of ideal workflows in text basically, and the cloud can repeatedly check that.
SPEAKER_02
And then there's also, and Claude does a really good job of this sort of expressing the intent of the current change at hand. So that gets really, really deeply exercised. So I think that the combination of those two things is important. The visual verification that I mentioned as well, um, video has been really cool to see. Actually video is a very under explored tool to give Claude as well. Like I think I've been prototyping is, uh, just giving Claude, uh, video captures of the thing that it has built and then giving it just basically an FFM tag and you'll watch it scrub through. And so like, oh, this animation has some jank in it. I'm going to go fix that.
SPEAKER_02
And I would never would be able to do it with like a screenshot sort of, uh, latency capture because it will have missed the moment. So I think that's, uh, that's another piece that is, that's really, really important. Um, and then for the pieces that aren't sort of easily testable and tend, because there is some more complex system, um, getting Claude to go and build like as robust, a sort of, you know, mock backend as possible or use ones off the shelf has been also really interesting. Like when I think about artifact, um, we had really comprehensive tests. This is kind of pre LLM. And one of the ways that we were able to do that really robustly was that basically every
SPEAKER_02
piece of info we had, whether it was Postgres, Redis, um, you know, all the AWS things had a really good in memory implementation that you could just do really quickly in unit tests and kind of extending that to like Claude land. Now, you know, I was working on something where it had like a pretty robust backend and for kind of complicated reasons, hard to spin that up on my dev server, but it was able to, again, one shot a really good like proxy for that, uh, by proxy, I mean like a substitute for that. And that was so valuable. And over time, it's been interesting as that like, uh, substitute has evolved as the rest
SPEAKER_02
of the code has evolved, which is the thing that, you know, if you had pitched that idea to me before, I'd be like, well, that's gonna be really hard because the upstream is gonna change. How are you gonna keep it in sync? And I don't think about that anymore. I'm like, yeah, Claude will read the changes and it will adapt the thing and it'll keep the two in sync and that that's, that's fine. There's some really interesting architectures around when you get a bug, it just automatically goes out and closes it. You know, the agent just gets kicked off, it closes it and then it sends a message to the customer being like, it's, it's fixed.
SPEAKER_02
Are you noticing a fable any change in how that process works? Yeah, I think there's a couple of things like, um, on a very like human to human or human to
SPEAKER_00
Claude level. One of the things that I've seen it do, um, better other models of the cable, I just need to do it really consistently too, is if the bug report, for example, came from somebody, you know, mentioning something in our like feedback channel in Slack. Um, and then like the thing that got fed into the cloud code session is like, oh, there's
SPEAKER_02
this and because of the Slack MCP, you can actually pull the thread. Um, have it then actually post back, uh, you know, as me, it'll be like, Hey, this is Mike's Claude. Like I fixed it. Here's the, you know, here's the pull request. But then I think in the previous clouds, the thing it does really well is then say, but hold tight, hold tight. It's not in production yet. I'll follow up when it actually is. And then like maybe a few hours later, like, oh, like this deploy went out. Like you should go test it. Is it fixed now? Like that level of follow through, I think is, is new on, on the closing the loop piece.
SPEAKER_02
And, uh, it's five, I definitely have these long running cloud code sessions that are basically like interacting as, as me, I guess, but some disclaimer in there too. Um, and the second goes back to that, like taste and discernment piece that we were talking about, which is like, it's one thing to say, there was a bug report. Therefore I must go fix this thing. And it's another one to say, you know what? Like this, like the, I hit this over the weekend, one of our internal systems, uh, basically had been running without restarting for a while. There was a memory. Um, and, uh, it was a good discernment of saying like, all right, Mike, like it's the weekend,
SPEAKER_02
like just rebounce the server. It's going to solve it for now. And like, we'll work on the, like, well, asynchronously get the PR going to like, fix this more longterm. So I think if you're going to have cloud in the loop in this kind of like, sort of close the loop bug report or system sort of issue to change, I think you really want it to understand where, you know, as any good SRE or engineer in the loop would like, great, let's solve the problem at hand. Let's like defer the question of like, do we need a re-architect on top of a completely different language found and, and understanding that balance is really important.
SPEAKER_02
One of the things that's like really exciting, mostly exciting to me about new models is it raises the floor so that everyone can kind of go build apps in one shot. Um, but it also raises the ceiling for experts. So like if you're a software engineer or founder, you can just go do things that you never would have been able to before because you have access to this really powerful model. So for me, I bought this one shot version of Borges, uh, infinite library.
SPEAKER_00
It's like a 3d game version of the, of the, of the library. It's wild. It runs right in the browser. It's so good. I can find like any, every essay inside of it. I'll send you the link. It's sick, but I think there's going to be this flowering of people doing things like, Oh, I made a game or maybe I trained a new model or, or, or whatever that they couldn't do that they couldn't do before. And I'd love to give people some inspiration, some examples of things that they might be able to do that they might not be thinking to do with this model. What are some ideas that come to you? Yeah.
SPEAKER_00
I think a few, um, maybe I'll start with the fun side and like riffing off the game piece. Like, I think people have a lot of like creative ideas for how do they express the complexity of what they are, like their world. Like everybody has the thing that they know really, really well. And there's probably some level of like, how do I then explain that to somebody else?
SPEAKER_02
Um, or how do I apply techniques elsewhere that I could then go, go off and do, um, my wife is, uh, studying, um, like environmental engineering, like studying geothermal, like really complex math and simulations. And I've seen like, as the models have gotten better, she has been able to apply even more complex techniques from even outside of that domain into that work. And I think what people should be able to do, you know, like full on PyTorch end to end simulations of that work in a way that wouldn't be possible. I think that maybe is one is like bring the like beautiful complexity of what you have and
SPEAKER_02
either show it to other people by like maybe making a game or maybe making a visualization, which I've seen her do as well, or at least like make, you know, bring other techniques to bear. Um, and the second piece is its ability to compose software that like solves a really unique problem to you. Um, and I've seen that internally. A lot of the work that we've been doing is how do we get as many of our internal systems like MCP-ified with the right permissioning structure and the right deployment kind of set up. Although externally, you have good options around some of these like platform as a service
SPEAKER_02
pieces and you can just ask a lot about them and they'll like help you set things up. But like, I love that feeling of like that thing that you always wish that you had. And then what has blown my mind, uh, there was a, uh, person who works in our go to market organization, um, has been like building this like really like for deeply thought integration of cloud into every part of her whole process. And you don't have to stop at that one shot. Like she's been working on it for months now and she can keep going. And like, I think one of the things that is maybe underappreciated about the models is
SPEAKER_02
I think in previous generations, it would eventually get to a complexity level where it was hard to iterate on it without feeling like you then would break the thing that they had, you know, like under or over abstracted. Whereas this is actually, you know, she's got access to something Fable or Fable like for a couple of months. And like, you've just seen it keep growing and growing and growing and growing. And now she's like deploying it to the whole GTM org. And like, I think that is really cool. Like the, the ceiling of complexity that a, a person that does not start out as technical can now builds for solving problems within their domain is like, is unprecedented.
SPEAKER_02
I agree. It, it, it writes great code. Like my, my benchmark that I have is called the senior engineer benchmark. I just have it, see if it can rewrite a code base from, uh, from first principles and the nearest model that the previous top was like a 62 or 63 out of a hundred. And this model got a 90 on the benchmark or 91, which is human senior engineer level.
SPEAKER_00
Like you can just keep going with this thing in a way that's it's, it's really fantastic. I'm curious though. One of the things that's really powerful that you mentioned is dynamic workflows. Tell us about that. This is, um, you know, we'll build things internally sometimes, and I will go really, uh, aggressively bug the engineer who built it and be like, when are we shipping this publicly? Because I think people are going to really like it. Um, I think there's a good reason why it was like built internally, but like we try to ship as many of these as possible.
SPEAKER_02
Um, and dynamic workflows was like definitely that to me. I, um, the person who built this is an engineer named Sid, who's awesome. And I was like, Sid, like, I want to get this out into the world because it's so good. Um, but I think it's especially good with, uh, a model like fable for two really big reasons. One, it helps, uh, sort of, uh, create the scaffold for like deep, meaningful work. Um, the craziest dynamic workflow I did and used fable for was I had, uh, an internal project that we had written in Python, but we needed it actually in TypeScript for like a really specific deployment reason.
SPEAKER_02
And having been internal to Instagram and we were like, should we write the whole thing into hack and, you know, port it to the PHP engine that Facebook, I was like, you never would have done that. Like maybe they can now with the model, but you know, at the time it seemed impossible. Uh, but here I had, you know, pretty complex code base. And I was like, I'm just going to set up a dynamic workflow and just let it run over the weekend. And it did. And the workflow was so cool. It was like, all right, I'm going to do like deep understanding of the work. I'm going to create sort of like a, almost like a spec of how everything works. I'm going to go module by module.
SPEAKER_02
I'm going to translate these pieces. I'm going to have tested incrementally. I'm going to do another adversarial test. I'm going to go check for anything that I missed. And it was just like really cool, like series of steps that the workflow was able to, to orchestrate. And I came back and I was like, yeah, this thing is like TypeScript and bun port of that thing. And it's actually better in these ways. Um, and it was very, you know, sort of documented, like these were the things I couldn't port, but most of these were like very specific to the specific implementation. It wasn't worth porting.
SPEAKER_02
And I do not think you could have done that a with previous models at that level of success and B, uh, with, without like the kind of scaffolding that or close provide. So I think that is extremely exciting kind of, uh, kind of combination of model capabilities and then our own ability to like orchestrate them over longer and longer time horizon with that feeling of like, you, you had a goal, you broke it down effectively and then you were able to work, make it work. The other piece is, I think over time, we'll be able to also make some of those subtasks, um, sort of tuned to the, uh, have the model be tuned to the level of complexity of it.
SPEAKER_02
So you can imagine that some parts of dynamic workflow don't need extra high thinking. They could use, you know, a medium thinking to get it done or even a smaller model. And I think, uh, that's really the future of where these things are going. So yeah, I, I'm a huge workflows, uh, DAU. For people who haven't used it before, tell me about how you got that workflow made. How did you design it? How did you make sure it was good? Yeah, it was pretty iterative, but sort of just started with cloud code. Like, Hey, I'm, I have this complex, you know, kind of task, like let's design a workflow to go and do it.
SPEAKER_00
It kind of showed me the plan. I was like, Oh, this is like close to what I want. I want to make sure that you do these three or four levels of, uh, of like additional
SPEAKER_02
verification for missed features. It's like, here's what you have. Are you ready to go? And it expresses the workflows in code, which I think is really valuable to kind of see what it was about to do. Um, and then, um, what was interesting is it did the full port. And then I had like a couple of like follow-up kind of questions that I had or like little tweaks. And I did those as sort of like mini workflows that built off the previous one as well. But I think that's like, uh, you know, we, we talked a little bit about whether chat was the, was the right interface. So we've had that conversation over the last year.
SPEAKER_02
And I think, um, workflows are a good, uh, middle ground of, uh, you can compose them using chat, but they're expressed using code. And then they're executed with like, I think a nice clean UI around what's happening at every stage. And like, I think we'll start bridging longer horizon work with chat in ways like that over time. Mike, this is such a great conversation. Thank you so much for joining and telling us all about this new model. I'm really excited to get to spend time with you and really, really looking forward to what people think outside too.
SPEAKER_02
Oh my gosh, folks, you absolutely positively have to smash that like button and subscribe to AI and I, why?
SPEAKER_00
Because this show is the epitome of awesomeness. It's like finding a treasure chest in your backyard, but instead of gold, it's filled with pure
SPEAKER_02
unadulterated knowledge bombs about chat GPT. Every episode is a roller coaster of emotions, insights, and laughter that will leave you
SPEAKER_01
on the edge of your seat craving for more. It's not just a show. It's a journey into the future with Dan Shipper as the captain of the spaceship. So do yourself a favor, hit like, smash subscribe, and strap in for the ride of your life. And now without any further ado, let me just say, Dan, I'm absolutely hopelessly in love with you.