SPEAKER_01
Amazing. I think, right, yeah, we're definitely on. What's up, folks? I'm here with Dex. You prefer Dex?
SPEAKER_00
Or Dexter. I go with either. They say good engineers are lazy. Dex is half the
SPEAKER_01
syllables, so I go with Dex often. Very good. Yeah, Jip the Tur, I suppose. Yeah, I'm Matt Pocock. We're going to talk about loads of stuff. I watched Dex do a really great talk at AI Engineer. I don't know when you gave the talk, but it came out a month ago on YouTube, let's say. Yeah, it was late November or something, I think. Yeah, all about RALPHing, I suppose, finding ways to make AI coding actually work inside organizations, and it just sparked a huge fire in my brain, and I needed to, I guess, douse the fire with gasoline. So, Dex, you're here, and I just want to talk about all this stuff, really. I don't know if I've literally just signed up to a
SPEAKER_01
chat service so that I can talk to Dex, so I just want to get the chat open locally. Let's do it. Let me have a look, but why don't you introduce yourself for the folks, just while I'm doing this
SPEAKER_00
little bit of admin? Absolutely. What's up, y'all? I'm Dex. I do a lot of ranting and raving about coding agents and how to use them, and I spend a lot of time going to battle with the slop AI hype machine and trying to really focus on what actually works and solving hard problems. I think the most exciting parts are around a lot of what we do is still software engineering, and there's a lot of engineering to be done. Just because the API, just because the AI writes the code for you, doesn't mean that you don't have to think and that you don't have to be thoughtful about systems. How do we optimize the engineer-AI
SPEAKER_00
collaboration so that we're all spending time on the highest-leverage, ideally most exciting and fun part of building, which is shipping and designing systems and solving hard problems?
SPEAKER_01
Yeah, I am so, my YouTube channel was primarily a TypeScript YouTube channel. I'm still fascinated by TypeScript. I still freaking love TypeScript. I built my career. I can't imagine why. Yeah, I just like it, man, and I've been posting more about AI recently for the last year or so. I've come up with a course on AI and stuff, and so much of what I see on YouTube, and weirdly not on X at all, is AI slop, right? AI just produces crap. There's content slop, or
SPEAKER_00
the video itself is just AI-generated, and someone made a video about HumanLayer, and they're like, "HumanLayer is insane," and then you watch the video, and it's just a guy scrolling, or it's an AI browsing agent scrolling around completely irrelevant GitHub issues that have nothing to do with what he's talking about and nothing to do with anything that we've ever done. I'm just like, this is, if I turn this on and wasn't paying attention, I'd be like, wow, someone made a video about our stuff, and then you look a little bit more closely and you're like, nothing here is well thought out.
SPEAKER_00
Nothing here makes sense. It's just feeding the AI content machine. But so people take that
SPEAKER_01
though, they take that opinion because there is real slop out there, and then they put it on all code, right? So any code that is created by AI must be slop, right? And I kind of want you to say up front, that's not true, is it? From what you're seeing out there, AI can actually produce decent-quality outputs.
SPEAKER_00
Yes, AI can make very good code. You have to wield it right. As Beyang from Sourcegraph would say, it's like you have to know how to hold it, and you have to do some. If you've never written code before, it's going to be hard to get AI to write good code, and it's going to be hard to know whether the code is good or not. I don't want to, I think vibe coding is super interesting, and it's incredible. It's opened doors for lots of people to be able to do things. But what we're really, really focused on is how can we give tools to the staff and principal engineers at companies with millions of lines of code and thousands of engineers and enable them to,
SPEAKER_00
one, actually do their job faster with AI, two to three times faster on most things is what I think is best in class right now for brownfield codebases, and also, how can you build practices and platforms and systems that enable some kind of standardization across a team and across an org? Because even if you have a hundred engineers and they're all really good at shipping really high- quality code with AI, if they're all doing things in different ways, then you're going to end up with chaos anyway, even if all the code is peak, the exact same code you would have written by
SPEAKER_01
hand, three times slower. Yeah. I wish to give you a chance to plug as well. What are you working on right now? What are you building? Let's give you a chance to plug at the beginning.
SPEAKER_00
And we'll do a little plug. Yeah, so we're working on, it's very early, so we shipped, we launched in September, an open-source IDE for managing lots of parallel cloud code sessions. It's called CodeLayer. I've said this publicly once or twice now, so I guess we can pop it off again. There is a waitlist for it, but it's also open source, so you can go build it and play with it yourself. We like to joke the waitlist was a little bit of a psyop of, if you can't go to the GitHub repo and figure it out, we're probably not ready. It's an early product, and so I'm very excited for all the thousands of people
SPEAKER_00
who have gone and figured it out and showed up in our Discord and sent us PRs and things like that. We've basically, in the last six weeks, taken everything we've learned trying to roll that out to customers and rebuilt the entire product from scratch, redesigned the. It's funny, six months ago I was giving this talk about research, plan, implement. I was like, look, the magic is not in the prompts. There is no perfect prompt. The thing you should understand is context engineering and intentional compaction and really being intentional about how you manage your context windows and stuff. I said these exact words: the words won't be research, plan, implement. The
SPEAKER_00
prompts will be different. There might not even be three steps. There might be six. There might be two. I don't know what's going to happen in six months. This is me paying my respects to the bitter lesson or whatever it is. Yeah, and then we woke up in the middle of December like, you know what, RPI is not enough. We actually need six steps. Just based on, okay, how do we help people? People pick up these prompts and they try to use them, and if you use them for a thousand hours, you can get incredible results. We saw lots of people would try them, get really good results, give them to their team, and their team would have a vast spectrum of quality of results. So
SPEAKER_00
we're rebuilding the workflows, and it's more steps now. So it's like, okay, I don't want I don't know what's going to happen in six months. This is me giving my respects to the bitter lesson or whatever it is, yeah. And then we woke up in the middle of December: you know what? RPI is not enough. We actually need six steps, and just based on, okay, how do we help people? People pick up these prompts and they try to use them, and if you use them for a thousand hours, you can get incredible results. We saw lots of people would try them, get really good results, give them to their team, and their team would have a vast spectrum of quality of results. So
SPEAKER_00
we're rebuilding the workflows, and then it's more steps now. And so it's like, okay, I don't want to ask someone to learn how to wield six prompts. Three was already a lot, so how do we rebuild the product around more guided workflows and splitting up these very long, if you read the create plan prompt that's in the open source repo. Have you tried it? No? Okay, I'm not trying. So there's a prompt you can use, a slash create plan, and it has about 50 instructions in it. And it's a thing that I went up in June and talked about, 12 factor agents, and was like, don't use prompts for control flow. If you know, excuse me, if you know what the workflow is,
SPEAKER_01
[SPEAKER_00] use control flow for control flow because it's going to be much more reliable. It's guaranteed that [SPEAKER_00] things are going to happen in the order of this. So we split up the planning process into multiple steps, [SPEAKER_00] and so the deterministic code that wraps that and guides the user through those steps [SPEAKER_00] is, I think, what we're really obsessed with getting really, really tight and really, really well, so that [SPEAKER_00] your chance of getting really good results, and the parts of the conversation that you're [SPEAKER_00] involved in, the things you have to do, are the interesting, most high-leverage things you can
SPEAKER_01
[SPEAKER_00] do versus, well, if you don't sprinkle in these magic words at this part of the process, it might not work as well. Yeah, so it sounds like you are close, almost as close as humanly possible, atomically, to people actually doing this in the wild and trying to solve their problems.
SPEAKER_00
[SPEAKER_01] Is getting AI coding agents to behave properly, and teaching people how to hold it
SPEAKER_01
right, and building a product that helps people hold it right, and that is, I think, where I want to go today, which is, how do you hold this thing right? That's the question
SPEAKER_00
[SPEAKER_01] right now. Okay, why don't we do a little bit of scene setting, which is one thing I really love [SPEAKER_01] from your talk, is that you talked a lot about the constraints that LLMs have and that you have to work
SPEAKER_01
within those constraints. Specifically, you say that there's a dumb zone and a smart zone, right, in terms of context. Yeah. And what I might do is I might just do a bit of live diagramming, and [SPEAKER_00] let's do it. You want to do it? What's your diagram? I think I've been meaning to switch over. I think it's Scala Draw. It just looks, as I've said this publicly before, I think it looks like the scrawlings of a serial killer, and I think that TLDraw, hang on, let me, yeah, you can edit that, I think.
SPEAKER_00
Oh yeah, you dropped me a link. Amazing, let's do it. And because I know the hotkeys in TLDraw, I'm going to have to relearn. I'm going to be a little slow, but that's all right. So I think, hello, is [SPEAKER_01] this going to work? Oh no, I'm not actually sharing my screen. That's what I'm not doing. Hang on, here we go. [SPEAKER_01] Entire screen, there we go. Hello, we're here, we're in. Amazing. Oh hi. Oh, you're literally here. Awesome. [SPEAKER_01] All right, so LLMs have a context window, right? Let's say this is the context window, [SPEAKER_01] and somewhere here is a smart zone and a dumb zone, yep, right? Let's say the top bit is the smart
SPEAKER_00
[SPEAKER_01] zone and the bit here is the dumb zone. What does this mean concretely to people who are doing this stuff? Yeah, I mean, so it's like before coding agents, the answer was, you will always get better results the less context you use, right? And the way we think about a tool-calling agent is you're always just in a loop. You are taking this, so whatever's in your context window in the first place, so let's put this here. So you have your
SPEAKER_01
[SPEAKER_00] see, how much did they change the hotkeys? You have your system message, and then you have your
SPEAKER_00
built-in Claude tools, like agents and tasks and read, write, edit, all that stuff. You have whatever custom MCPs you have, or the tool search thing. But custom MCPs are no longer as much of a problem. I haven't messed with the tool search, but the promise of it seems like it'll probably work pretty well, yeah. And then you have your custom instructions, right? Your Claude MD or your Agents MD or whatever. I'm going to use a lot of Claude words, but
SPEAKER_01
[SPEAKER_00] this is true for no matter what LLM you use. They all have quadratic attention. They all work in this way. [SPEAKER_00] And then you're going to put in, and just to pause you, I'm going to pause you a lot, I think, which is, could you just explain quadratic attention to me? What does that mean?
SPEAKER_00
Yeah, so that's basically the idea that the longer your context window, the amount of compute and the quality of the responses that you get from the LLM, the amount of, I don't want to say compute intelligence, required increases quadratically with the number of tokens. And so if you have five tokens and you go to ten tokens, the ten tokens is going to be, you double the number of tokens, you four times the amount of computation that's needed to ingest all that context. [SPEAKER_01] and actually act on it. That sound right? Yeah, that's right. And that's per layer and per attention
SPEAKER_00
[SPEAKER_01] head too, right? So that's just going crazy. And you can have like 50, 80, these numbers [SPEAKER_01] aren't public, but it just goes nuts. Every single token you add quadratically scales into oblivion, and it makes it really dumb, yes. And a lot of the benchmarks for long context are about this. They've run on this thing called needle in a haystack, which I think is not actually useful. I mean, Jeff Hunn was talking about this, the Ralph guy was talking about this in March or April, like needle in a haystack is not useful because most of the time you don't need to read
SPEAKER_00
a hundred thousand words and pick the one sentence that matters. You can read a hundred thousand words and act on all of the information that's in there, or tease out the fifty thousand words that actually matter, and that's a much harder problem that we're not as good at benchmarking. I mean, people are working on benchmarks for long-context contacts, agents, and stuff like this, but yeah, back to the context window thing, though. Are you ready to jump back in? Does that sufficiently answer
SPEAKER_01
[SPEAKER_00] the quadratic attention question? You did it wonderfully. Well done, amazing. So we'll have [SPEAKER_00] your user message comes in here, right? Is this okay? That's great. Now, and then the models go
SPEAKER_00
over here with the TLDraw choice. It's okay. I'm ready to rock with it. It's about time I learned it. The agent's going to call some tools. You're going to get some, let's see, you're going to actually matter, and that's a much harder problem that we're not as good at benchmarking. People are working on benchmarks for long contacts, agents, and stuff like this, but yeah, back to the context window thing, though. Are you ready to jump back in? Does that sufficiently answer the quadratic attention question? You did it wonderfully. Well done. Amazing. So we'll have your user message comes in here, right? Is this okay? That's great. Now, and then the models go
SPEAKER_00
over here with the teal draw choice. Am I? It's okay. I'm ready to rock with it. It's about time I learned it. The agent's going to call some tools. You're going to get some, let's see, you're going to
SPEAKER_01
[SPEAKER_00] get some tool responses. This happens many, many times. If you're, what's in the readme of [SPEAKER_00] this project? That one's pretty easy. And then you eventually are going to get back some kind of
SPEAKER_00
assistant response, right? And your dumb zone, smart zone has, is, is, is smaller, but this is not to scale. Let's put it that way. Yeah, this context window is much longer. Yeah, at some [SPEAKER_01] point you're going to just run out of room in the smart zone, right? You will continue going [SPEAKER_01] down, continue long, long, long as context, and you'll hit some sort of barrier. Yes, exactly. And so the idea, yeah, the idea here is there's lots of different ways to use a coding agent. The most naive way is just to pop it open, ask it for some stuff. When it finishes that stuff, ask it for some
SPEAKER_00
more stuff. And the, there's the, I also like the smart zone, dumb zone. I think is
SPEAKER_01
[SPEAKER_00] is a rule of thumb, like 40 context usage, and that's based on how I count [SPEAKER_00] tokens. Actually, the way I count tokens and the way we count tokens at Riptide is slightly different [SPEAKER_00] from the way the one you get default in Cloud Code, because they include the end buffer in their
SPEAKER_00
percentage calculations, as far as they don't actually count that as available. The way to really understand this for different types of work and different types of tasks, the number of changes, it's flexible. Yeah, but the more context you use, the worse results [SPEAKER_01] you'll get, and I think the important thing is the paranoia, right? The feeling of [SPEAKER_01] I should be worried about this, or I should be thinking about this and trying to optimize for it, right?
SPEAKER_00
Yes. And the idea there is every time you're about to send a new message, you should ask yourself the question, could this be a new context? Is the information, and sometimes I'll keep it, right? Sometimes, well, there's a lot of good information in here, and I don't really have the patience to wait for the model to go read all those files over again in another context window or to have it, like, hey, I don't have the confidence that it will be able to accurately summarize everything we have so far so that I can onboard the new context window. But you should be asking yourself every time, should this be a new user message or should this be a new
SPEAKER_00
[SPEAKER_01] context window? And the hope, I suppose, is that you're not needing to ask yourself that, but you're [SPEAKER_01] designing systems and harnesses in a way that you optimize for the smart zone and avoid the dumb zone, [SPEAKER_01] right? Right, and that kind of leads us into Ralph, right? I love it. Yeah, right, so Ralph is a system that optimizes for always working in the early part of the context window. And doing so, I like to think of Ralph, this is, I've talked about this a little bit before, I like to think of Ralph as a control loop. Have you ever spent time in the Kubernetes world? Me? No, I've not
SPEAKER_00
touched it once. So the way it works is it's built on control loops, which is really simple. Your thermostat is a control loop, right? You read the current state of the world. Oh my God. All right, this is not a TL draw thing. I just can't type this morning. And then you read the desired state of the world, and then you take some action, and you just do this forever. You just take action to advance the current state of the world to the desired state of the world. Yeah, okay. That hockey is the same. Nice. And you just do this in a loop forever, right? And so when I think of Ralph, I think of Ralph the same as a control loop, where you have your system prompt, your built-
SPEAKER_00
in tools, your cloud MCPs, and your user message is telling it to read a couple files. And then what
SPEAKER_01
[SPEAKER_00] gets pulled in is basically the specs, which is your desired state of the world. You have it look at source, [SPEAKER_00] which is the current state of the world, and then you say implement one thing, right? [SPEAKER_00] Yeah. And the goal, when I think about the smart zone, the most practical advice I can give you is [SPEAKER_00] figure out a task that is, and this one thing, sizing this thing is important because you [SPEAKER_00] basically want to be able to do your edit, edit, edit tool calls, and then you're going to verify
SPEAKER_00
as run the tests or run the linter or whatever it is. Maybe it was broken, and you run a couple more edits, and then you run the test and it passes. And when I think about task sizing, or if you're doing a more one-off planning and you're not doing Ralph forever, you want to be able to do your tasks should be sized that you can make the changes, run the tests, fix any issues, run the tests, and then do the commit and push or whatever it is all before, in as little context as possible, right? There's trade-offs, because if you make the task too small, there's, you change one file and you change the signature of a function,
SPEAKER_00
and then the test won't pass until you change the other file. So some of these changes, you can't just do one edit per loop. I think someone, were you the one talking about the Cursor one where it was just like the loop was, someone was on X talking about the loop was too tight in some implementation [SPEAKER_01] of Ralph. So this, I kind of think of it like you've got some amount of water, right? And you [SPEAKER_01] need to, you can't fit all of this water in a single cup, right? And so you need to put it in multiple
SPEAKER_01
cups. And I think what people don't realize is that the cup is smaller than you think it is, right? The agent has less available to it than you think it is. And that's a really interesting question. I mean, there's so many interesting questions with Ralph. [SPEAKER_00] You have a lot of levers you can pull. Can we pull up the whiteboard again? Yeah, there you go. Yeah, so [SPEAKER_00] like here you have your specs. You can make the specs smaller and tighter. You can find a way to [SPEAKER_00] give it better ways to traverse the source so that's not taking up too much. You could rip out some of
SPEAKER_01
[SPEAKER_00] your custom MCPs. You can disable some of the built-in tools. You have lots of, and then it's [SPEAKER_00] like, okay, how big is this task? And this is probably the highest leverage thing you can do. But [SPEAKER_00] again, it's like you have lots of leverage you can pull to control what goes in here and how big [SPEAKER_00] your tasks are and how far, what is your average context length? That would actually [SPEAKER_00] be really interesting. If I were going to build a Ralph tool, part of what I would build is [SPEAKER_00] give people visibility and metrics into how much of the context is being used by each thing so that
SPEAKER_00
give it better ways to traverse the source so that's not taking up too much. You could rip out some of your custom MCPs. You can disable some of the built-in tools. You have lots of and and then it's okay, how big is this task? This is probably the highest leverage thing you can do, but again, you have lots of leverage you can pull to control what goes in here and how big your tasks are and how how how far, what is your average context length? That would actually be really interesting. If I were going to build a Ralph tool, part of what I would build is give people visibility and metrics into how much of the context is being used by each thing so that
SPEAKER_00
you know which place is your bottleneck, and you get to the end of the loop, how much context was used at each loop as it's exiting, right, and you can chart that per iteration and then understand, oh, this is always making it to 60 context. I need to make my tasks smaller. Yep, that's one [SPEAKER_01] thing I find. We can get into this later, but that's one thing I find really tricky with [SPEAKER_01] implementing this myself locally, is that the observability is so bad, right? I don't get any [SPEAKER_01] real stats here, but let's keep it at a high level for now, which is you are trying
SPEAKER_00
[SPEAKER_01] to build a system with Ralph. Whoops, I just completely knocked you off. That's okay. Where you're just [SPEAKER_01] trying to fill each cup a little bit, right, so that it has room in the cup to do tests, to do check the [SPEAKER_01] state of the world with a Playwright MCP or something. Yeah, and sizing the amount of water that [SPEAKER_01] you put in is mostly the whole game. How, though, if someone's going, okay, I like the sound of [SPEAKER_01] that. I want to run an agent in a loop to do some tasks. How do they integrate it into [SPEAKER_01] their organization? What's the one small thing they can do to, because because all
SPEAKER_00
[SPEAKER_01] we're doing here, it sounds stupid, Ralph Wiggum, right, but all we're doing here is we're just running [SPEAKER_01] the tools that we have already, like Claude Code in a loop, right? That's all we do. Yep. How do you get [SPEAKER_01] that into an organization, just for yourself to try things out? How do you set it up? Yeah, so there's a bunch of ways. The literal exact steps that I end up doing is, I don't know, I had this. I tell the story at some point, but we had, I was sitting with one of our front-end engineers who was going through a bunch of React code, and we were trying to debug something together, and
SPEAKER_01
[SPEAKER_00] he said, this code needs to be refactored. We need to refactor. I'm like, okay, great, you [SPEAKER_00] go fix the thing, and while we're working, I'm on the side, I'm chatting back and [SPEAKER_00] forth. I spent 30 minutes going back and forth with Claude of, build me the world's best React style [SPEAKER_00] guide, and it came, it read the code and came back with a bunch of questions, like do you want to allow
SPEAKER_00
barrel exports? Do you want to do something else? Do you want to use useEffect, or do you want to use Zustand stores? I see all these different patterns. Which ones are the right patterns? So 30 minutes going back answering this, it came back with 25 or 27 React rules, and I can share the PR, and you can put it on the video in the show notes or whatever, because the PR, I'll get to why the PR never got merged, but it came up with these rules. I spent another 30 minutes with the front-end engineer, and we went back and forth, iterated a couple of them, and then we dropped that in a repo, and we did a Ralph loop that was basically, I set up a GCP VM, I turned on a
SPEAKER_01
[SPEAKER_00] tmux session, I grabbed this style guide, and I made a slight variation on Jeff's prompt, right, [SPEAKER_00] because Jeff's prompt is read the specs and implement them. In this case, the desired state of [SPEAKER_00] the world was all the code adheres to the style guide, yeah, and instead of an implementation plan it
SPEAKER_00
was a refactoring plan, yeah. That was our artifact that we're iterating over and has all the tasks, and this thing went and churned for about six hours. I checked on it about six hours later, and I started to see the messages, like hey, we're done. I guess I'll do this too, and I'm like, that's when you know. Jeff always talks about Ralph being underbaked versus overbaked, or you think about art, when you train a machine learning model. Usually, overfit it, then you roll back to the checkpoint that is actually the right amount of model fit, with the same thing as if you leave it going too long, it will come up with more stuff to
SPEAKER_00
do. These models are trained to, if user asks you to do something, find a way to be [SPEAKER_01] useful, right? And to go to this last 90 percent, which is the way this works. In [SPEAKER_01] the implementation, at least the way I've seen it, I think it was in the original article maybe, [SPEAKER_01] and the way I've implemented it, is you run it in a bash loop, and then you say, you tell the LLM when [SPEAKER_01] you're done, emit some sort of sigil that I can then read from your output and then stop the loop. Right, and you specify something. I've never done that part. That's what the Anthropic plugin does.
SPEAKER_00
I haven't explored that. I think part of the joy of Ralph is just let it go, and it's on you to check on it, and you can always roll it back and experiment
SPEAKER_01
[SPEAKER_00] with the output. This is one of the exciting things, is I don't know, if you look [SPEAKER_00] in the Cursed Lang repo, which is the programming language that Jeff made with Ralph, but [SPEAKER_00] you see all kinds of weird emergent behavior. It dumps hundreds of markdown files everywhere [SPEAKER_00] as part of its work. You never told it to do that, but here, is it gonna let me share my screen? Let me see if I can. It may not. It may actually. It may. Trees, why are you doing that? I'd never thought of that. I'd literally never thought of not stopping it. You know what I mean? This is
SPEAKER_01
nervous me not wanting to spend too many tokens or something, but I never thought [SPEAKER_00] of just letting it run and run and run. Let's see here. Maybe I can. Yeah, let me see if I can share [SPEAKER_00] this, yeah, because let's see. Share screen. Let's share the window. That's fine. So [SPEAKER_00] here's the Cursed Lang repo, and so this is the thing that Ralph ran in for a very long time, for weeks [SPEAKER_00] and weeks and weeks and weeks and weeks. Yeah. Oh, interesting. It looks like it's been cleaned up a [SPEAKER_00] little bit. If you go to earlier commits here, you will see probably, let's just go back to
SPEAKER_00
here. You bump the size of the font a little bit? Yeah, yeah, yeah, yep. All right, now maybe we can go to the Rust branch. I don't think this one ended up getting cleaned up as much. Interesting. Okay. If you looked at earlier versions of this, there were just a hundred markdown files, all caps, of work in progress of Claude just spinning out thoughts and ideas and documentation and things like this, but there's this, you get these emergent and weeks and weeks and weeks and weeks, yeah. Oh, interesting. It looks like it's been cleaned up a little bit. If you go to earlier commits here, you will see probably, let's just go back to
SPEAKER_00
here. You bump the size of the font a little bit. Yeah, yeah, yeah, yep. All right, now maybe we can go to the rust branch. I don't think this one ended up getting cleaned up as much.
SPEAKER_01
[SPEAKER_00] Interesting, okay. If you looked at earlier versions of this, there were just a hundred
SPEAKER_00
markdown files, all caps, of just work in progress of claude just spinning out thoughts and ideas and documentation and things like this. But there's this, you get these emergent behaviors, and jeff had this thing too. It's like, oh, I left it running too long and it decided it needed it just came up with more stuff to build. It was like, okay, I think we should probably also have this programming language should have support for post quantum cryptography or something like this.
SPEAKER_01
[SPEAKER_00] And so I kind of think part of the fun of it is under-specify a little bit and see [SPEAKER_00] what happens. Obviously, if you want really good working production software, you should specify as much as possible. But yeah, yeah, that's kind of where I am. That's because I'm seeing this as just a way that I can do my work, and with this idea that I have. And I suppose I don't know, I don't know even where I picked that up, whether that is from, did I pick that
SPEAKER_00
[SPEAKER_01] from the ralph plugin? I probably did, didn't I? They have the completion promise and the max [SPEAKER_01] iterations. That's exactly what I do. That's exactly what I do. So that felt natural to me because what [SPEAKER_01] I wanted to do was choose a scope of work up front that was super well defined, and then just let [SPEAKER_01] ralph find its way to the end. Whereas what you're talking about there is choose an [SPEAKER_01] under-specified amount of work and just let ralph play in this zone, really. And what's [SPEAKER_01] to me, it feels like my version, you get to, that's a lot of upfront work for me, but it's kind
SPEAKER_00
[SPEAKER_01] of work that I want to do because I want to sharpen my ideas. I want to describe the desired state of the [SPEAKER_01] world really clearly. But yeah, I'm just interested in your thoughts there. What use cases do you see [SPEAKER_01] for an under-specified ralph versus an over-specified maybe ralph? I mean, I will also say, I've run ralph on a pile of spec. I've had ralph write the specs. I've been like, hey, go read all this documentation for this product and then extract out the clean room specifications that define actually how this would work from scratch, no implementation details. And then you run ralph on
SPEAKER_00
the specs. So you have one ralph building the specs, and another ralph is reading the specs, and then I think it was add ai features, and so it was having another directory full of specs that are just literally copy. And it's really almost lispy or pipeliney, where it's like you have a ralph that's taking the internet and turning it into specs. You have another ralph that's taking specs and turning them to what would this product look like if it had more ai? And then you have another one downstream that is reading those facts and turning them into working [SPEAKER_01] code. Yeah, wild. Okay, I guess it didn't go great. I didn't read the specs and I got a thing
SPEAKER_00
that I didn't love. And that's part of it too, is I think ralph is probably not the right final answer for how we build production software. I think it's probably, if anything, it's an incredible lesson in how context windows work. And that's kind of, I think, what jeff runs around doing, is like, don't use the entropic plug and learn the theory, because the theory is actually what makes you a better ai and coding engineer, and you should still use it, but you should [SPEAKER_01] understand why it works so well. Well, so let's define ralph as that kind of wild, freeform
SPEAKER_00
[SPEAKER_01] play around idea. Sure, there is an idea here, though, right? Running something in a loop that is really [SPEAKER_01] specified and tight, right? And that sounds like, if that's not ralph, then that's, I don't know, something
SPEAKER_01
else, right? But how do you, when you're talking to people who are doing this and they want long- running coding agents to just run afk, what's the structure that you recommend if it's not ralph? [SPEAKER_00] so that part, if you're doing full afk, this refactor plan, I can [SPEAKER_00] finish talking through how we did this and what I would do. The pr didn't get merged. Actually, [SPEAKER_00] here, I can pull it up and show it to you. I think that would be [SPEAKER_00] probably an interesting thing to look at. So we'll go to pull requests, close. You're gonna like the name of this
SPEAKER_01
[SPEAKER_00] one too. This is very low effort. Here it is, ralph is back, ralph is back. So this is 20 commits over [SPEAKER_00] six hours. Here is the react coding standards doc that I made with the engineer. Here's the, you [SPEAKER_00] can view actually the refactoring plan that it used to track all its progress. Let's see the rich
SPEAKER_00
so it's like, cool, we consolidated all the custom hooks, we added error boundaries everywhere, we made
SPEAKER_01
[SPEAKER_00] sure that we use proper forms, dates. It just did all this stuff that was in our spec. [SPEAKER_00] This didn't get merged. The reason why is because it's, thousands, how many lines is this?
SPEAKER_00
It's, yeah, 20,000 lines. And some of that is the plan files and stuff, but I was like, this was a cool experiment. I said to the engineer, he looked at it two days later and there was 100 merge conflicts, and I was like, okay, that's fine. We'll go do this later. You could always just rebase route, right? The nice thing about ralph too is you can just control c it, throw away all the code, and just same specs, update the code, and just run it again. And it was so low, it took 10
SPEAKER_01
[SPEAKER_00] minutes to set up the gcp instance and then took 10 minutes to turn it off later, yeah. I think if I [SPEAKER_00] was going to do something like this in a company for real, and a thing we've experimented with [SPEAKER_00] internally was a lot more towards do less. Ralph is a cool current state of the world, [SPEAKER_00] desired state of the world, make one change. I deployed something like this for an internal repo [SPEAKER_00] that we have, and I set it to only run. I run it on cron every night and only run three iterations of the [SPEAKER_00] loop. I want every morning to wake up to the code base, just for free. I just get the code base a
SPEAKER_01
[SPEAKER_00] little bit better. And we didn't merge all those prs, but every morning we wake up and just be like, [SPEAKER_00] oh yeah, this is great. And the specification is just here's how I want the code base to look,
SPEAKER_00
here's how I want it to be architected in the end world, and ralph just kind of does these little increments. So my advice is do not send your co-workers a 20,000-line pr that refactors the entire code base. That will not work in the real world. But you can use the concepts, and it's a building block that you can use to kind of hands off, don't think about it, just running in a github action [SPEAKER_01] every night and and and see what you get. Yeah, one thing that I'm thinking about actually, I was going A little bit better, and we didn't merge all those PRs, but every morning we wake up and just be like,
SPEAKER_00
"Oh yeah, this is great," and the specification is just, "Here's how I want the code base to look. Here's how I want it to be architected in the end world," and Ralph just does these little increments. So my advice is: do not send your co-workers a 20,000-line PR that refactors the entire code base. That will not work in the real world, but you can use the concepts, and it's a building block that you can use to hands off. Don't think about it, just running in a GitHub Action [SPEAKER_01] every night and see what you get. Yeah, one thing that I'm thinking about, actually, I was going
SPEAKER_00
[SPEAKER_01] to give this a go before we started, but we started early, so I didn't get a chance, but [SPEAKER_01] I have a couple of open source repos, and I just don't get time to triage all of the [SPEAKER_01] issues that come in. So there's an extremely simple Ralph loop here, right, which is you [SPEAKER_01] just feed all the issues into a loop, and you get it to work out whether it's, "I can categorize it," right, [SPEAKER_01] work out whether it's a feature request, which there's no real action to be taken, or you get [SPEAKER_01] it to make a reproduction of the bug or something, and then you maybe feed that to something else,
SPEAKER_00
[SPEAKER_01] another Ralph that's looking for a GitHub label or something to actually fix the bug, to make a PR. [SPEAKER_01] There's just a workflow pattern here, right, like you managing little queues and then Ralphs. Different Ralphs, different prompts are appropriate for different queues, yeah, and you tune those prompts [SPEAKER_01] with a bit of sitting on the loop and watching what it's doing, and then you let it go AFK overnight, [SPEAKER_01] as you're describing, and I suppose not too much at once is the way of thinking about it. I will also say you should be careful with taking GitHub issues from the community and
SPEAKER_00
feeding them to a clod that is in dangerously skip permissions, because it's technically untrusted input. We have a Ralph that runs through our Linear queue, but it is not allowed to see anything until we look at it and look for hidden prompts in HTML, markdown comments, and all of this stuff, and make sure, okay, this is actually okay for a model in DSP to process. Go, yeah, yeah, that [SPEAKER_01] makes total sense. Okay, so assuming then that you've got this loop set up, what are we talking about in [SPEAKER_01] terms of task size? Let's go there first. When you're prompting Ralph, yeah, the
SPEAKER_00
[SPEAKER_01] really nice thing I love about Ralph is that it frees you from the burden of choosing the next task. [SPEAKER_01] If you've just got a huge wadge of issues or something, as long as they're
SPEAKER_01
correctly described and you have the dependencies all mapped out, then you can just let Ralph do its thing and choose the next issue and then just go from there, which is just so nice. And when you're telling Ralph how big a change to make, what are you saying to it? What are you trying to get it as small as possible, or what do you think? So yeah, and
SPEAKER_00
in production, the tools we use are a little more human in the loop, again, because it's like the hardest software problems in the world are not going to be done completely
SPEAKER_01
[SPEAKER_00] autonomously. They're going to be done very much in collaboration with humans, but I think this [SPEAKER_00] advice is true in both worlds, which is when you're designing a plan, whether it's a plan you're [SPEAKER_00] going to give to Ralph or you're going to plan, we have a separate prompt that basically runs [SPEAKER_00] a parent sub-agent, runs each task in a sub-aid, sorry, parent agent, and then runs each implementation in a sub-agent, and the [SPEAKER_00] parent agent checks the work and then moves to the next phase, commits, and moves to the next phase.
SPEAKER_01
[SPEAKER_00] This is a plan that's been very vetted. The advice I keep finding myself, the models can't [SPEAKER_00] do, and it's the thing that a human needs to be in the loop right now, maybe we just need to [SPEAKER_00] make the prompts better, but models are not good at planning work in the way that I, a human, would
SPEAKER_00
do the work. I think there's something to be said for trying to steer them to, I don't know, I was doing something the other day with the buddy, and it was like, we have these 40 things on the front end, we're going to move them to the back end, serve them from an API, and then have the front end query the JSON and render it that way. And the model wanted to be like, "Cool, first we're going to move the 40 things, and then we're going to wire up the API endpoint, and then we're going to refactor the whole front end," three-phase plan, right? It's like, okay, if you were building this, what would you do? And it's like, well, I would move one thing,
SPEAKER_00
get the endpoint working, make sure one thing was working, probably learn some stuff in that process, and then hit some surprises. You want to minimize the size of the change in the same way that if you were an engineer, you wouldn't just copy-paste 40 files over and then go. Maybe some people would. I wouldn't, because I want tighter feedback loops, and I know there's going to be unknowns, and there's going to be blast radius to these changes that we didn't think about. I don't particularly want to sit, either you have someone with 15 years of experience in the code base, so they're just like, "Oh yeah, you're going to hit that issue," or you can sit there for three
SPEAKER_00
hours and research every single thing in the code base and try to get the plan perfect before you start, but there's a sweet spot here that is optimize for learning early on in the implementation plan, and then segment the tasks in the plan the same: how much code would you write if there was no AI? How much code would you write before you pulled up the web app and looked at it, or how much code would you write before you pause to run the tests? And that's a good task size for Ralph, or for a very human-on-the-loop AI where you're doing a smaller five- or six-step plan or something like this. That's my rule of thumb, always:
SPEAKER_00
your instincts as an engineer are still really, really good, and you should listen to them. Just because you're using AI, a lot of things change, but a lot of things don't. The programmatic [SPEAKER_01] programmer has this concept called tracer bullets, which is you should write code as it, you know this, [SPEAKER_01] right? Yeah, yeah, which is you should write code that goes through, that tells you where you're going, [SPEAKER_01] and that goes through all of the integration layers first, and you shouldn't do [SPEAKER_01] one huge change. It should be one change that goes through all the layers so you see that all the layers
SPEAKER_01
[SPEAKER_00] work, right? This is crazy. I read this book 12 years ago, and I loved this idea, and I've been [SPEAKER_00] explaining exactly what you said, this concept, to people for the last six months or whatever, and I forgot that it had a name. I love it. I literally read the book. I've had it on my shelf for three years, still in the wrap, but I got it out and I read it in one sitting the other day. It's just the most remarkable, well put together thing. Yeah, but this wisdom from 20 years ago is still super and that goes through all of the integration layers first, and you shouldn't do
SPEAKER_01
one huge change. It should be one change that goes through all the layers, so you see that all the layers [SPEAKER_00] work, right? This is crazy. I read this book 12 years ago, and I loved this idea, and I've been [SPEAKER_00] explaining exactly what you said, this concept, to people for the last six months or whatever, and I forgot that it had a name. I love it. I literally read the book. I've had it on my shelf for three
SPEAKER_00
[SPEAKER_01] years, still in the wrap, right? But I got it out and I read it in one sitting the other day. It's [SPEAKER_01] just the most remarkable, well-put-together thing. Yeah, but this wisdom from 20 years ago is still super [SPEAKER_01] important. It's, I would say, more important than ever, right? Because, yeah, and actually let's take a [SPEAKER_01] little sidebar on that, which is that if you're not reading those old books, we are programming in English [SPEAKER_01] now, and those books are written in the most clear, perfect way of describing good code that you're ever
SPEAKER_00
[SPEAKER_01] going to see, right? They are so, so good, and so I've been going back to loads of them, and I'm [SPEAKER_01] trying to buy four more now. It's just an amazing way to learn to prompt for coding, is reading those 20 real books. It really is incredible. I love this. I have another one that I've been obsessed with lately, which is, do you remember, I think it was Martin Fowler, it might be an Uncle Bob thing, but this idea of learning tests. Have you heard of this? No, no, no, I haven't. Okay, so a learning test is, I've been loving, especially if you're integrating with a system that you don't
SPEAKER_00
understand. We do a lot of integrations with the cloud code SDK, which the docs are getting much better, but the docs are docs, which means you can't fully trust them all the time, and it's closed source. So when we're going to build a feature, we'll get halfway
SPEAKER_01
[SPEAKER_00] through and be like, oh, we thought the behavior was this, but it wasn't, and the surprise would
SPEAKER_00
happen halfway through implementation, and then we have to throw everything out and roll back. And so a learning test is, you would build it in your unit test framework. You would put it in a place
SPEAKER_01
[SPEAKER_00] where it doesn't get run on every build because it's not for that, but it's a unit
SPEAKER_00
describe, expect, all this stuff, but it's for verifying how an external library behaves,
SPEAKER_01
[SPEAKER_00] whether it's bun.color or whether it's some giant battleship piece of software like [SPEAKER_00] the cloud agent SDK. You say, I think the docs say it works like this, I think it works like [SPEAKER_00] this. Go write a test that actually has assertions and console logs that explains the behavior of [SPEAKER_00] this thing, and often the assumptions were wrong. But what Claude is really, coding agents are really [SPEAKER_00] good at, is, okay, let me iterate on this until I have an understanding of what is the [SPEAKER_00] contract with this external system, and add a couple. The asserts are really nice because we don't
SPEAKER_01
[SPEAKER_00] run these all the time, but three months ago we wrote some learning tests about how session IDs work, [SPEAKER_00] and then they changed the session ID behavior, and I was like, I think this is wrong. Run the learning [SPEAKER_00] tests again, and it's like, yep, they broke the contract, or the contract changed, and here's [SPEAKER_00] how it behaves now. It's such a powerful thing to have to do up front as part of your research [SPEAKER_00] and design for building something new, and it's, again, a thing that, I mean, they were for this. [SPEAKER_00] They have not changed. The concept hasn't changed, and it's things that people have been doing for
SPEAKER_01
[SPEAKER_00] obviously you shouldn't write unit, because the wisdom is you shouldn't write unit tests on external [SPEAKER_00] software. It's their job to test it. You trust the contract and you trust the main container, and it's like, I don't know. We're on a total tangent here, but I had almost, when you quit a job and you think back, things you would have done differently on that job, we had a back-end team that was super unreliable and was shipping just crap. What I wanted to do was put, maybe their internal code was good, but they just kept breaking the tiny little JSON contracts that we'd have between front end and back end. They would
SPEAKER_01
put a null where things weren't supposed to be null and blah blah blah blah, misspell things, and so I just wanted to put a little Zod thing in between it just on the dev server just to