Matt Pocock

LIVE: Chat with AI Coding Wizard Dex Horthy

2277 summary words 10 min summary Watch video

Start with the signal

10 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Reliable AI coding in real organizations comes less from a magic prompt or fully autonomous agent and more from deliberately engineered workflows that constrain context, decompose work into testable increments, and retain human judgment at planning and review boundaries.
  • Why it matters: Dex offers reusable operating patterns for turning coding agents from one-off personal tools into controlled systems for large, brownfield codebases—while identifying the practical failure modes around context overload, oversized changes, weak observability, and untrusted inputs.
  • Best use: Use this as a design discussion for an AI coding-agent control plane: extract its context-budgeting, workflow orchestration, task-sizing, nightly-maintenance, and security patterns rather than treating Ralph as a turnkey autonomous-development recipe.

Executive Summary

Dex Horthy argues against both extremes of AI-coding discourse: AI-generated code is not inherently "slop," but capable output depends on skilled engineers designing the surrounding system. His stated best-in-class expectation for mature brownfield codebases is roughly 2–3x engineer productivity, not wholesale autonomous replacement. The organizational problem is consistency: even strong engineers will create operational chaos if each uses agents through incompatible practices.

The central technical argument is context engineering. Tool-using coding agents accumulate system prompts, tool definitions, repository instructions, user requests, file reads, tool outputs, edits, and test results in a finite context window. As that window fills, performance degrades; Dex's practical rule is to regard roughly 40% context use as a warning threshold, though the exact boundary varies. The response is to minimize irrelevant context, start fresh contexts when appropriate, and structure tasks so an agent can implement, test, repair, and commit before it enters the degraded zone.

He frames Ralph-style loops as a useful control-loop pattern: repeatedly compare a desired-state specification with the present codebase, make one bounded improvement, and verify it. But he explicitly says Ralph is not likely the final production-software architecture. Its deeper value is teaching how context windows, feedback loops, specifications, and task boundaries affect agent reliability. Riptide's product evolution reflects this: it moved from a three-prompt research-plan-implement workflow toward a more guided six-step process, with deterministic software handling control flow rather than embedding workflow logic in a giant prompt.

For real teams, Dex favors small, reviewable, reversible increments over massive autonomous rewrites. A six-hour Ralph experiment produced a 20,000-line React refactor PR that was technically impressive but operationally unusable after merge conflicts accumulated. A more viable pattern is a scheduled agent that makes only a few bounded iterations each night against an architectural specification, producing small changes for human review. He also emphasizes classic engineering practices—tracer bullets, early feedback through the full stack, and learning tests for uncertain external contracts—as especially valuable when agents accelerate implementation.

Key Takeaways

  • Claim: AI can produce high-quality code, but its quality is conditional on experienced engineering judgment and a deliberately designed human-agent workflow. | Evidence: Dex distinguishes exploratory "vibe coding" from enabling staff/principal engineers working in million-line codebases and says best-in-class AI assistance is currently about 2–3x faster on most brownfield work; he warns that 100 engineers using effective agents in 100 incompatible ways will still create organizational chaos. | Implication: Treat AI coding adoption as an engineering-platform and standardization problem, not merely a license or individual prompt-skills rollout. | Caveat: This is an experience-based estimate rather than a benchmarked, generalizable productivity measurement, and Dex does not claim that agents can autonomously solve the hardest software problems.
  • Claim: Context-window management is the primary reliability constraint in coding agents: more accumulated context generally produces worse agent behavior. | Evidence: Dex describes the context as containing system instructions, built-in tools, MCP tools, repo instructions such as CLAUDE.md/AGENTS.md, user messages, tool calls, file contents, and results. He calls the early portion the "smart zone" and the fuller portion the "dumb zone," using about 40% context consumption as a practical warning rule of thumb. | Implication: Build agent harnesses that expose context consumption by source and iteration, prune unnecessary tools/instructions, improve repository traversal, and deliberately reset or compact contexts rather than allowing long sessions to drift. | Caveat: The 40% figure is explicitly approximate and depends on the model, task, token-accounting method, and reserved output buffer.
  • Claim: Agent workflows should use deterministic code for known control flow, rather than asking a model to reliably execute a long sequence of prompt-encoded instructions. | Evidence: Riptide found that a roughly 50-instruction /create plan prompt produced highly variable outcomes across users. Dex cites his earlier "12 factor agents" advice—do not use prompts for control flow—and says the company split planning into more steps, with deterministic wrapper code guiding users through them. | Implication: For agent systems, make state transitions, gates, ordering, permissions, retries, and artifact handoffs explicit in orchestration code; reserve model reasoning for research, design alternatives, implementation, and interpretation. | Caveat: The exact workflow is not fixed: Dex notes that the previous research-plan-implement pattern became insufficient and was being rebuilt into a six-step guided workflow.
  • Claim: Ralph is best understood as a desired-state control loop, not as a single prompt or a promise of autonomous software delivery. | Evidence: Dex compares Ralph to Kubernetes controllers and thermostats: inspect current state, compare it with desired state, take an action that closes the gap, and repeat. In a coding loop, the desired state is a specification, the current state is the repository, and each iteration should implement and verify one bounded change. | Implication: Use the loop for scoped maintenance, refactoring, artifact generation, and supervised exploration, while putting explicit boundaries around production scope, permissions, budgets, stopping conditions, and review. | Caveat: Dex says Ralph is probably not the final answer for production development; leaving it under-specified or running indefinitely can create emergent but unwanted work, such as a programming-language agent deciding it should add post-quantum cryptography.
  • Claim: The highest-leverage design choice is task sizing: each agent task should fit an implement-test-fix-retest-commit loop without exhausting useful context. | Evidence: Dex advises sizing work so the agent can make edits, run tests or linters, repair failures, rerun verification, and commit/push in as little context as possible. He rejects both one-edit loops that cannot cross necessary dependencies and broad plans such as moving 40 frontend items first, then wiring an API, then refactoring the whole UI. | Implication: Plan work as vertical slices through relevant integration layers—equivalent to tracer bullets—so each increment validates assumptions early and limits blast radius. | Caveat: Tasks cannot always be atomically small because a valid change may require coordinated edits across files; the goal is a minimal complete feedback loop, not arbitrary fragmentation.
  • Claim: The production-friendly form of autonomous maintenance is small, periodic, reviewable change—not giant agent-generated rewrites. | Evidence: A six-hour experiment used an AI-generated React style guide and refactoring plan to create 20 commits and roughly 20,000 changed lines. It was not merged because it accumulated around 100 merge conflicts. Dex's internal alternative runs on cron nightly for only three iterations against an intended architectural state, yielding incremental PRs. | Implication: Deploy agents first as a continuous-improvement lane with narrow scope and disposable branches; avoid presenting coworkers with a massive refactor whose correctness and mergeability cannot be evaluated economically. | Caveat: Even small scheduled PRs still require review, and Dex notes that not every internal agent PR was merged.
  • Claim: Untrusted issue content must not be fed directly to a high-permission coding agent because public text can carry prompt-injection payloads. | Evidence: When Matt proposes an issue-triage/fix loop using public GitHub issues, Dex flags it as untrusted input. Riptide's Linear-queue agent does not process material until humans inspect it for hidden prompts in HTML and Markdown comments and determine that it is safe for an agent running with dangerous skip-permissions. | Implication: Separate intake/triage agents from agents with repository write, shell, or deployment permissions; sanitize and classify external content before it crosses into privileged execution contexts. | Caveat: The transcript gives a manual screening practice, not a complete automated trust-boundary or sandboxing design.

Detailed Brief

CodeLayer and the shift from prompt craft to guided workflow products

  • Claims: Dex says Riptide launched CodeLayer, an open-source IDE for managing many parallel cloud coding sessions, and then substantially rebuilt it after learning from early customer rollout.; The product direction is toward reducing the number of agent-operation skills an individual must internalize while preserving the high-leverage decisions for the human.; The stated design goal is to prevent quality from depending on users knowing when to inject particular wording into a long conversation.
  • Evidence: CodeLayer launched in September and is open source, despite also having a waitlist.; The original research-plan-implement approach was taught as a three-step process; the revised workflow has six steps because users who copied the previous prompts achieved a wide spectrum of results.; Dex says the former planning prompt alone contained approximately 50 instructions.
  • Caveats: CodeLayer is described as early and under active redesign, so its workflow should be read as a moving product hypothesis rather than a settled standard.; Dex explicitly expects prompt names, step counts, and mechanics to continue changing as models and operational learning evolve.
  • Implications: A useful coding-agent platform should productize repeatable workflow constraints and artifacts, rather than distribute an expert's prompt library and expect organizational consistency.; Evaluate agent tools on their ability to standardize planning, execution, verification, handoff, and observability across a team.

Specification pipelines and learning tests

  • Claims: Dex experiments with chained agents that transform source material into increasingly executable artifacts: documentation into clean-room specifications, specifications into AI feature ideas, and those specifications into code.; He advocates "learning tests" when integrating with external or poorly understood systems: tests designed to discover and record actual behavior rather than validate the team's own implementation.; Classic software-engineering literature is presented as a strong source for agent prompting because it expresses precise design and implementation concepts in natural language.
  • Evidence: One experiment had an agent read product documentation and extract implementation-independent clean-room specifications, then a downstream agent suggest what the product would look like with more AI, then another produce code.; For the Claude Code SDK, Dex creates unit-test-shaped experiments with assertions and console logs, excludes them from normal build runs, and uses them to establish the real contract of closed-source or imperfectly documented APIs.; A learning test about session IDs later revealed that the external system's session-ID behavior had changed.; The conversation explicitly invokes tracer bullets from The Pragmatic Programmer as a useful framing for end-to-end incremental changes.
  • Caveats: Dex reports that at least one specification-pipeline experiment produced an output he did not like, reinforcing that intermediate artifacts require review.; Learning tests do not eliminate the need for contract ownership or integration monitoring; they provide evidence about observed behavior at a point in time.
  • Implications: Make specifications, plans, and behavioral probes first-class versioned artifacts in the agent pipeline, with quality gates before downstream implementation.; For critical model, SDK, or service integrations, maintain a focused suite of contract probes that can detect behavior drift independently of ordinary application tests.

Notable Concepts & Terms

  • Ralph / Ralph Wiggum loop: A repeated coding-agent control loop: inspect the current repository state, compare it to a desired specification, make one bounded change, verify it, and repeat.
  • Smart zone / dumb zone: Dex's practical model of context degradation: an agent is more capable early in its context window and less reliable once tool results, files, and conversation history accumulate.
  • Context engineering: The operational discipline of deciding what enters an agent context, how it is compacted or reset, how tools are exposed, and how tasks fit within an effective context budget.
  • Control loop: The Kubernetes/thermostat analogy for agent orchestration: continuously move observed state toward declared desired state through constrained actions.
  • Tracer bullets: An incremental-development practice from The Pragmatic Programmer: implement a narrow end-to-end slice through integration layers first to learn quickly and reduce unknowns.
  • Learning tests: Non-routine test experiments that establish the actual behavior of an external dependency or SDK when documentation is incomplete, inaccurate, or subject to change.
  • Dangerously skip permissions: A high-permission agent execution mode that materially raises prompt-injection and unintended-action risk when the agent can consume untrusted external content.
  • Intentional compaction: Deliberately summarizing, pruning, or reinitializing agent context to preserve the information needed for the next task without carrying a degraded full history forward.

Operator Notes / Why Ken Should Care

  • Instrument agent runs with a context ledger: tokens consumed by system instructions, tools, repository guidance, file reads, tool outputs, tests, and each iteration; alert on sustained runs entering a predefined degradation threshold.
  • Define an orchestrated workflow schema in code with explicit artifacts and gates—research output, reviewed plan, bounded implementation task, verification result, commit/PR—instead of relying on one long prompt to manage sequencing.
  • Pilot a nightly maintenance agent on a low-risk repository branch with a hard cap of a few iterations, narrow architectural goals, mandatory tests, and human PR review; measure accepted change rate, merge conflicts, and rollback rate.
  • Create separate trust tiers for agent inputs. Do not pass GitHub issues, support tickets, webpages, or pasted documents into a repository-writing agent without sanitization, injection screening, and a least-privilege execution environment.
  • Add learning-test suites for critical external dependencies, especially agent SDK semantics, session behavior, model/tool contracts, and APIs whose documentation cannot be treated as authoritative.
  • Require plans to be expressed as vertical, end-to-end tracer-bullet slices; reject agent plans that batch broad migrations before validating a single working path.

Source/Metadata

  • Title: LIVE: Chat with AI Coding Wizard Dex Horthy
  • Transcript words: 8881
  • Duration seconds: 2780
  • Timestamp note: No usable timestamps or chapter markers were present in the supplied transcript.
Full transcript 9187 words · 39 min read
0:00

SPEAKER_01

Amazing. I think, right, yeah, we're definitely on. What's up, folks? I'm here with Dex. You prefer Dex?

0:06

SPEAKER_00

Or Dexter. I go with either. They say good engineers are lazy. Dex is half the

0:12

SPEAKER_01

syllables, so I go with Dex often. Very good. Yeah, Jip the Tur, I suppose. Yeah, I'm Matt Pocock. We're going to talk about loads of stuff. I watched Dex do a really great talk at AI Engineer. I don't know when you gave the talk, but it came out a month ago on YouTube, let's say. Yeah, it was late November or something, I think. Yeah, all about RALPHing, I suppose, finding ways to make AI coding actually work inside organizations, and it just sparked a huge fire in my brain, and I needed to, I guess, douse the fire with gasoline. So, Dex, you're here, and I just want to talk about all this stuff, really. I don't know if I've literally just signed up to a

1:03

SPEAKER_01

chat service so that I can talk to Dex, so I just want to get the chat open locally. Let's do it. Let me have a look, but why don't you introduce yourself for the folks, just while I'm doing this

1:14

SPEAKER_00

little bit of admin? Absolutely. What's up, y'all? I'm Dex. I do a lot of ranting and raving about coding agents and how to use them, and I spend a lot of time going to battle with the slop AI hype machine and trying to really focus on what actually works and solving hard problems. I think the most exciting parts are around a lot of what we do is still software engineering, and there's a lot of engineering to be done. Just because the API, just because the AI writes the code for you, doesn't mean that you don't have to think and that you don't have to be thoughtful about systems. How do we optimize the engineer-AI

1:59

SPEAKER_00

collaboration so that we're all spending time on the highest-leverage, ideally most exciting and fun part of building, which is shipping and designing systems and solving hard problems?

2:10

SPEAKER_01

Yeah, I am so, my YouTube channel was primarily a TypeScript YouTube channel. I'm still fascinated by TypeScript. I still freaking love TypeScript. I built my career. I can't imagine why. Yeah, I just like it, man, and I've been posting more about AI recently for the last year or so. I've come up with a course on AI and stuff, and so much of what I see on YouTube, and weirdly not on X at all, is AI slop, right? AI just produces crap. There's content slop, or

2:47

SPEAKER_00

the video itself is just AI-generated, and someone made a video about HumanLayer, and they're like, "HumanLayer is insane," and then you watch the video, and it's just a guy scrolling, or it's an AI browsing agent scrolling around completely irrelevant GitHub issues that have nothing to do with what he's talking about and nothing to do with anything that we've ever done. I'm just like, this is, if I turn this on and wasn't paying attention, I'd be like, wow, someone made a video about our stuff, and then you look a little bit more closely and you're like, nothing here is well thought out.

3:19

SPEAKER_00

Nothing here makes sense. It's just feeding the AI content machine. But so people take that

3:26

SPEAKER_01

though, they take that opinion because there is real slop out there, and then they put it on all code, right? So any code that is created by AI must be slop, right? And I kind of want you to say up front, that's not true, is it? From what you're seeing out there, AI can actually produce decent-quality outputs.

3:43

SPEAKER_00

Yes, AI can make very good code. You have to wield it right. As Beyang from Sourcegraph would say, it's like you have to know how to hold it, and you have to do some. If you've never written code before, it's going to be hard to get AI to write good code, and it's going to be hard to know whether the code is good or not. I don't want to, I think vibe coding is super interesting, and it's incredible. It's opened doors for lots of people to be able to do things. But what we're really, really focused on is how can we give tools to the staff and principal engineers at companies with millions of lines of code and thousands of engineers and enable them to,

4:22

SPEAKER_00

one, actually do their job faster with AI, two to three times faster on most things is what I think is best in class right now for brownfield codebases, and also, how can you build practices and platforms and systems that enable some kind of standardization across a team and across an org? Because even if you have a hundred engineers and they're all really good at shipping really high- quality code with AI, if they're all doing things in different ways, then you're going to end up with chaos anyway, even if all the code is peak, the exact same code you would have written by

4:57

SPEAKER_01

hand, three times slower. Yeah. I wish to give you a chance to plug as well. What are you working on right now? What are you building? Let's give you a chance to plug at the beginning.

5:05

SPEAKER_00

And we'll do a little plug. Yeah, so we're working on, it's very early, so we shipped, we launched in September, an open-source IDE for managing lots of parallel cloud code sessions. It's called CodeLayer. I've said this publicly once or twice now, so I guess we can pop it off again. There is a waitlist for it, but it's also open source, so you can go build it and play with it yourself. We like to joke the waitlist was a little bit of a psyop of, if you can't go to the GitHub repo and figure it out, we're probably not ready. It's an early product, and so I'm very excited for all the thousands of people

5:43

SPEAKER_00

who have gone and figured it out and showed up in our Discord and sent us PRs and things like that. We've basically, in the last six weeks, taken everything we've learned trying to roll that out to customers and rebuilt the entire product from scratch, redesigned the. It's funny, six months ago I was giving this talk about research, plan, implement. I was like, look, the magic is not in the prompts. There is no perfect prompt. The thing you should understand is context engineering and intentional compaction and really being intentional about how you manage your context windows and stuff. I said these exact words: the words won't be research, plan, implement. The

6:18

SPEAKER_00

prompts will be different. There might not even be three steps. There might be six. There might be two. I don't know what's going to happen in six months. This is me paying my respects to the bitter lesson or whatever it is. Yeah, and then we woke up in the middle of December like, you know what, RPI is not enough. We actually need six steps. Just based on, okay, how do we help people? People pick up these prompts and they try to use them, and if you use them for a thousand hours, you can get incredible results. We saw lots of people would try them, get really good results, give them to their team, and their team would have a vast spectrum of quality of results. So

6:54

SPEAKER_00

we're rebuilding the workflows, and it's more steps now. So it's like, okay, I don't want I don't know what's going to happen in six months. This is me giving my respects to the bitter lesson or whatever it is, yeah. And then we woke up in the middle of December: you know what? RPI is not enough. We actually need six steps, and just based on, okay, how do we help people? People pick up these prompts and they try to use them, and if you use them for a thousand hours, you can get incredible results. We saw lots of people would try them, get really good results, give them to their team, and their team would have a vast spectrum of quality of results. So

7:34

SPEAKER_00

we're rebuilding the workflows, and then it's more steps now. And so it's like, okay, I don't want to ask someone to learn how to wield six prompts. Three was already a lot, so how do we rebuild the product around more guided workflows and splitting up these very long, if you read the create plan prompt that's in the open source repo. Have you tried it? No? Okay, I'm not trying. So there's a prompt you can use, a slash create plan, and it has about 50 instructions in it. And it's a thing that I went up in June and talked about, 12 factor agents, and was like, don't use prompts for control flow. If you know, excuse me, if you know what the workflow is,

8:06

SPEAKER_01

[SPEAKER_00] use control flow for control flow because it's going to be much more reliable. It's guaranteed that [SPEAKER_00] things are going to happen in the order of this. So we split up the planning process into multiple steps, [SPEAKER_00] and so the deterministic code that wraps that and guides the user through those steps [SPEAKER_00] is, I think, what we're really obsessed with getting really, really tight and really, really well, so that [SPEAKER_00] your chance of getting really good results, and the parts of the conversation that you're [SPEAKER_00] involved in, the things you have to do, are the interesting, most high-leverage things you can

8:43

SPEAKER_01

[SPEAKER_00] do versus, well, if you don't sprinkle in these magic words at this part of the process, it might not work as well. Yeah, so it sounds like you are close, almost as close as humanly possible, atomically, to people actually doing this in the wild and trying to solve their problems.

9:00

SPEAKER_00

[SPEAKER_01] Is getting AI coding agents to behave properly, and teaching people how to hold it

9:09

SPEAKER_01

right, and building a product that helps people hold it right, and that is, I think, where I want to go today, which is, how do you hold this thing right? That's the question

9:24

SPEAKER_00

[SPEAKER_01] right now. Okay, why don't we do a little bit of scene setting, which is one thing I really love [SPEAKER_01] from your talk, is that you talked a lot about the constraints that LLMs have and that you have to work

9:36

SPEAKER_01

within those constraints. Specifically, you say that there's a dumb zone and a smart zone, right, in terms of context. Yeah. And what I might do is I might just do a bit of live diagramming, and [SPEAKER_00] let's do it. You want to do it? What's your diagram? I think I've been meaning to switch over. I think it's Scala Draw. It just looks, as I've said this publicly before, I think it looks like the scrawlings of a serial killer, and I think that TLDraw, hang on, let me, yeah, you can edit that, I think.

10:11

SPEAKER_00

Oh yeah, you dropped me a link. Amazing, let's do it. And because I know the hotkeys in TLDraw, I'm going to have to relearn. I'm going to be a little slow, but that's all right. So I think, hello, is [SPEAKER_01] this going to work? Oh no, I'm not actually sharing my screen. That's what I'm not doing. Hang on, here we go. [SPEAKER_01] Entire screen, there we go. Hello, we're here, we're in. Amazing. Oh hi. Oh, you're literally here. Awesome. [SPEAKER_01] All right, so LLMs have a context window, right? Let's say this is the context window, [SPEAKER_01] and somewhere here is a smart zone and a dumb zone, yep, right? Let's say the top bit is the smart

11:02

SPEAKER_00

[SPEAKER_01] zone and the bit here is the dumb zone. What does this mean concretely to people who are doing this stuff? Yeah, I mean, so it's like before coding agents, the answer was, you will always get better results the less context you use, right? And the way we think about a tool-calling agent is you're always just in a loop. You are taking this, so whatever's in your context window in the first place, so let's put this here. So you have your

11:23

SPEAKER_01

[SPEAKER_00] see, how much did they change the hotkeys? You have your system message, and then you have your

11:30

SPEAKER_00

built-in Claude tools, like agents and tasks and read, write, edit, all that stuff. You have whatever custom MCPs you have, or the tool search thing. But custom MCPs are no longer as much of a problem. I haven't messed with the tool search, but the promise of it seems like it'll probably work pretty well, yeah. And then you have your custom instructions, right? Your Claude MD or your Agents MD or whatever. I'm going to use a lot of Claude words, but

12:03

SPEAKER_01

[SPEAKER_00] this is true for no matter what LLM you use. They all have quadratic attention. They all work in this way. [SPEAKER_00] And then you're going to put in, and just to pause you, I'm going to pause you a lot, I think, which is, could you just explain quadratic attention to me? What does that mean?

12:20

SPEAKER_00

Yeah, so that's basically the idea that the longer your context window, the amount of compute and the quality of the responses that you get from the LLM, the amount of, I don't want to say compute intelligence, required increases quadratically with the number of tokens. And so if you have five tokens and you go to ten tokens, the ten tokens is going to be, you double the number of tokens, you four times the amount of computation that's needed to ingest all that context. [SPEAKER_01] and actually act on it. That sound right? Yeah, that's right. And that's per layer and per attention

12:52

SPEAKER_00

[SPEAKER_01] head too, right? So that's just going crazy. And you can have like 50, 80, these numbers [SPEAKER_01] aren't public, but it just goes nuts. Every single token you add quadratically scales into oblivion, and it makes it really dumb, yes. And a lot of the benchmarks for long context are about this. They've run on this thing called needle in a haystack, which I think is not actually useful. I mean, Jeff Hunn was talking about this, the Ralph guy was talking about this in March or April, like needle in a haystack is not useful because most of the time you don't need to read

13:29

SPEAKER_00

a hundred thousand words and pick the one sentence that matters. You can read a hundred thousand words and act on all of the information that's in there, or tease out the fifty thousand words that actually matter, and that's a much harder problem that we're not as good at benchmarking. I mean, people are working on benchmarks for long-context contacts, agents, and stuff like this, but yeah, back to the context window thing, though. Are you ready to jump back in? Does that sufficiently answer

14:01

SPEAKER_01

[SPEAKER_00] the quadratic attention question? You did it wonderfully. Well done, amazing. So we'll have [SPEAKER_00] your user message comes in here, right? Is this okay? That's great. Now, and then the models go

14:14

SPEAKER_00

over here with the TLDraw choice. It's okay. I'm ready to rock with it. It's about time I learned it. The agent's going to call some tools. You're going to get some, let's see, you're going to actually matter, and that's a much harder problem that we're not as good at benchmarking. People are working on benchmarks for long contacts, agents, and stuff like this, but yeah, back to the context window thing, though. Are you ready to jump back in? Does that sufficiently answer the quadratic attention question? You did it wonderfully. Well done. Amazing. So we'll have your user message comes in here, right? Is this okay? That's great. Now, and then the models go

14:55

SPEAKER_00

over here with the teal draw choice. Am I? It's okay. I'm ready to rock with it. It's about time I learned it. The agent's going to call some tools. You're going to get some, let's see, you're going to

15:07

SPEAKER_01

[SPEAKER_00] get some tool responses. This happens many, many times. If you're, what's in the readme of [SPEAKER_00] this project? That one's pretty easy. And then you eventually are going to get back some kind of

15:17

SPEAKER_00

assistant response, right? And your dumb zone, smart zone has, is, is, is smaller, but this is not to scale. Let's put it that way. Yeah, this context window is much longer. Yeah, at some [SPEAKER_01] point you're going to just run out of room in the smart zone, right? You will continue going [SPEAKER_01] down, continue long, long, long as context, and you'll hit some sort of barrier. Yes, exactly. And so the idea, yeah, the idea here is there's lots of different ways to use a coding agent. The most naive way is just to pop it open, ask it for some stuff. When it finishes that stuff, ask it for some

15:46

SPEAKER_00

more stuff. And the, there's the, I also like the smart zone, dumb zone. I think is

15:50

SPEAKER_01

[SPEAKER_00] is a rule of thumb, like 40 context usage, and that's based on how I count [SPEAKER_00] tokens. Actually, the way I count tokens and the way we count tokens at Riptide is slightly different [SPEAKER_00] from the way the one you get default in Cloud Code, because they include the end buffer in their

16:11

SPEAKER_00

percentage calculations, as far as they don't actually count that as available. The way to really understand this for different types of work and different types of tasks, the number of changes, it's flexible. Yeah, but the more context you use, the worse results [SPEAKER_01] you'll get, and I think the important thing is the paranoia, right? The feeling of [SPEAKER_01] I should be worried about this, or I should be thinking about this and trying to optimize for it, right?

16:42

SPEAKER_00

Yes. And the idea there is every time you're about to send a new message, you should ask yourself the question, could this be a new context? Is the information, and sometimes I'll keep it, right? Sometimes, well, there's a lot of good information in here, and I don't really have the patience to wait for the model to go read all those files over again in another context window or to have it, like, hey, I don't have the confidence that it will be able to accurately summarize everything we have so far so that I can onboard the new context window. But you should be asking yourself every time, should this be a new user message or should this be a new

17:29

SPEAKER_00

[SPEAKER_01] context window? And the hope, I suppose, is that you're not needing to ask yourself that, but you're [SPEAKER_01] designing systems and harnesses in a way that you optimize for the smart zone and avoid the dumb zone, [SPEAKER_01] right? Right, and that kind of leads us into Ralph, right? I love it. Yeah, right, so Ralph is a system that optimizes for always working in the early part of the context window. And doing so, I like to think of Ralph, this is, I've talked about this a little bit before, I like to think of Ralph as a control loop. Have you ever spent time in the Kubernetes world? Me? No, I've not

18:05

SPEAKER_00

touched it once. So the way it works is it's built on control loops, which is really simple. Your thermostat is a control loop, right? You read the current state of the world. Oh my God. All right, this is not a TL draw thing. I just can't type this morning. And then you read the desired state of the world, and then you take some action, and you just do this forever. You just take action to advance the current state of the world to the desired state of the world. Yeah, okay. That hockey is the same. Nice. And you just do this in a loop forever, right? And so when I think of Ralph, I think of Ralph the same as a control loop, where you have your system prompt, your built-

18:44

SPEAKER_00

in tools, your cloud MCPs, and your user message is telling it to read a couple files. And then what

18:50

SPEAKER_01

[SPEAKER_00] gets pulled in is basically the specs, which is your desired state of the world. You have it look at source, [SPEAKER_00] which is the current state of the world, and then you say implement one thing, right? [SPEAKER_00] Yeah. And the goal, when I think about the smart zone, the most practical advice I can give you is [SPEAKER_00] figure out a task that is, and this one thing, sizing this thing is important because you [SPEAKER_00] basically want to be able to do your edit, edit, edit tool calls, and then you're going to verify

19:22

SPEAKER_00

as run the tests or run the linter or whatever it is. Maybe it was broken, and you run a couple more edits, and then you run the test and it passes. And when I think about task sizing, or if you're doing a more one-off planning and you're not doing Ralph forever, you want to be able to do your tasks should be sized that you can make the changes, run the tests, fix any issues, run the tests, and then do the commit and push or whatever it is all before, in as little context as possible, right? There's trade-offs, because if you make the task too small, there's, you change one file and you change the signature of a function,

20:02

SPEAKER_00

and then the test won't pass until you change the other file. So some of these changes, you can't just do one edit per loop. I think someone, were you the one talking about the Cursor one where it was just like the loop was, someone was on X talking about the loop was too tight in some implementation [SPEAKER_01] of Ralph. So this, I kind of think of it like you've got some amount of water, right? And you [SPEAKER_01] need to, you can't fit all of this water in a single cup, right? And so you need to put it in multiple

20:32

SPEAKER_01

cups. And I think what people don't realize is that the cup is smaller than you think it is, right? The agent has less available to it than you think it is. And that's a really interesting question. I mean, there's so many interesting questions with Ralph. [SPEAKER_00] You have a lot of levers you can pull. Can we pull up the whiteboard again? Yeah, there you go. Yeah, so [SPEAKER_00] like here you have your specs. You can make the specs smaller and tighter. You can find a way to [SPEAKER_00] give it better ways to traverse the source so that's not taking up too much. You could rip out some of

21:06

SPEAKER_01

[SPEAKER_00] your custom MCPs. You can disable some of the built-in tools. You have lots of, and then it's [SPEAKER_00] like, okay, how big is this task? And this is probably the highest leverage thing you can do. But [SPEAKER_00] again, it's like you have lots of leverage you can pull to control what goes in here and how big [SPEAKER_00] your tasks are and how far, what is your average context length? That would actually [SPEAKER_00] be really interesting. If I were going to build a Ralph tool, part of what I would build is [SPEAKER_00] give people visibility and metrics into how much of the context is being used by each thing so that

21:45

SPEAKER_00

give it better ways to traverse the source so that's not taking up too much. You could rip out some of your custom MCPs. You can disable some of the built-in tools. You have lots of and and then it's okay, how big is this task? This is probably the highest leverage thing you can do, but again, you have lots of leverage you can pull to control what goes in here and how big your tasks are and how how how far, what is your average context length? That would actually be really interesting. If I were going to build a Ralph tool, part of what I would build is give people visibility and metrics into how much of the context is being used by each thing so that

22:16

SPEAKER_00

you know which place is your bottleneck, and you get to the end of the loop, how much context was used at each loop as it's exiting, right, and you can chart that per iteration and then understand, oh, this is always making it to 60 context. I need to make my tasks smaller. Yep, that's one [SPEAKER_01] thing I find. We can get into this later, but that's one thing I find really tricky with [SPEAKER_01] implementing this myself locally, is that the observability is so bad, right? I don't get any [SPEAKER_01] real stats here, but let's keep it at a high level for now, which is you are trying

22:45

SPEAKER_00

[SPEAKER_01] to build a system with Ralph. Whoops, I just completely knocked you off. That's okay. Where you're just [SPEAKER_01] trying to fill each cup a little bit, right, so that it has room in the cup to do tests, to do check the [SPEAKER_01] state of the world with a Playwright MCP or something. Yeah, and sizing the amount of water that [SPEAKER_01] you put in is mostly the whole game. How, though, if someone's going, okay, I like the sound of [SPEAKER_01] that. I want to run an agent in a loop to do some tasks. How do they integrate it into [SPEAKER_01] their organization? What's the one small thing they can do to, because because all

23:19

SPEAKER_00

[SPEAKER_01] we're doing here, it sounds stupid, Ralph Wiggum, right, but all we're doing here is we're just running [SPEAKER_01] the tools that we have already, like Claude Code in a loop, right? That's all we do. Yep. How do you get [SPEAKER_01] that into an organization, just for yourself to try things out? How do you set it up? Yeah, so there's a bunch of ways. The literal exact steps that I end up doing is, I don't know, I had this. I tell the story at some point, but we had, I was sitting with one of our front-end engineers who was going through a bunch of React code, and we were trying to debug something together, and

23:47

SPEAKER_01

[SPEAKER_00] he said, this code needs to be refactored. We need to refactor. I'm like, okay, great, you [SPEAKER_00] go fix the thing, and while we're working, I'm on the side, I'm chatting back and [SPEAKER_00] forth. I spent 30 minutes going back and forth with Claude of, build me the world's best React style [SPEAKER_00] guide, and it came, it read the code and came back with a bunch of questions, like do you want to allow

24:10

SPEAKER_00

barrel exports? Do you want to do something else? Do you want to use useEffect, or do you want to use Zustand stores? I see all these different patterns. Which ones are the right patterns? So 30 minutes going back answering this, it came back with 25 or 27 React rules, and I can share the PR, and you can put it on the video in the show notes or whatever, because the PR, I'll get to why the PR never got merged, but it came up with these rules. I spent another 30 minutes with the front-end engineer, and we went back and forth, iterated a couple of them, and then we dropped that in a repo, and we did a Ralph loop that was basically, I set up a GCP VM, I turned on a

24:47

SPEAKER_01

[SPEAKER_00] tmux session, I grabbed this style guide, and I made a slight variation on Jeff's prompt, right, [SPEAKER_00] because Jeff's prompt is read the specs and implement them. In this case, the desired state of [SPEAKER_00] the world was all the code adheres to the style guide, yeah, and instead of an implementation plan it

25:05

SPEAKER_00

was a refactoring plan, yeah. That was our artifact that we're iterating over and has all the tasks, and this thing went and churned for about six hours. I checked on it about six hours later, and I started to see the messages, like hey, we're done. I guess I'll do this too, and I'm like, that's when you know. Jeff always talks about Ralph being underbaked versus overbaked, or you think about art, when you train a machine learning model. Usually, overfit it, then you roll back to the checkpoint that is actually the right amount of model fit, with the same thing as if you leave it going too long, it will come up with more stuff to

25:59

SPEAKER_00

do. These models are trained to, if user asks you to do something, find a way to be [SPEAKER_01] useful, right? And to go to this last 90 percent, which is the way this works. In [SPEAKER_01] the implementation, at least the way I've seen it, I think it was in the original article maybe, [SPEAKER_01] and the way I've implemented it, is you run it in a bash loop, and then you say, you tell the LLM when [SPEAKER_01] you're done, emit some sort of sigil that I can then read from your output and then stop the loop. Right, and you specify something. I've never done that part. That's what the Anthropic plugin does.

26:30

SPEAKER_00

I haven't explored that. I think part of the joy of Ralph is just let it go, and it's on you to check on it, and you can always roll it back and experiment

26:39

SPEAKER_01

[SPEAKER_00] with the output. This is one of the exciting things, is I don't know, if you look [SPEAKER_00] in the Cursed Lang repo, which is the programming language that Jeff made with Ralph, but [SPEAKER_00] you see all kinds of weird emergent behavior. It dumps hundreds of markdown files everywhere [SPEAKER_00] as part of its work. You never told it to do that, but here, is it gonna let me share my screen? Let me see if I can. It may not. It may actually. It may. Trees, why are you doing that? I'd never thought of that. I'd literally never thought of not stopping it. You know what I mean? This is

27:17

SPEAKER_01

nervous me not wanting to spend too many tokens or something, but I never thought [SPEAKER_00] of just letting it run and run and run. Let's see here. Maybe I can. Yeah, let me see if I can share [SPEAKER_00] this, yeah, because let's see. Share screen. Let's share the window. That's fine. So [SPEAKER_00] here's the Cursed Lang repo, and so this is the thing that Ralph ran in for a very long time, for weeks [SPEAKER_00] and weeks and weeks and weeks and weeks. Yeah. Oh, interesting. It looks like it's been cleaned up a [SPEAKER_00] little bit. If you go to earlier commits here, you will see probably, let's just go back to

27:58

SPEAKER_00

here. You bump the size of the font a little bit? Yeah, yeah, yeah, yep. All right, now maybe we can go to the Rust branch. I don't think this one ended up getting cleaned up as much. Interesting. Okay. If you looked at earlier versions of this, there were just a hundred markdown files, all caps, of work in progress of Claude just spinning out thoughts and ideas and documentation and things like this, but there's this, you get these emergent and weeks and weeks and weeks and weeks, yeah. Oh, interesting. It looks like it's been cleaned up a little bit. If you go to earlier commits here, you will see probably, let's just go back to

28:32

SPEAKER_00

here. You bump the size of the font a little bit. Yeah, yeah, yeah, yep. All right, now maybe we can go to the rust branch. I don't think this one ended up getting cleaned up as much.

28:40

SPEAKER_01

[SPEAKER_00] Interesting, okay. If you looked at earlier versions of this, there were just a hundred

28:51

SPEAKER_00

markdown files, all caps, of just work in progress of claude just spinning out thoughts and ideas and documentation and things like this. But there's this, you get these emergent behaviors, and jeff had this thing too. It's like, oh, I left it running too long and it decided it needed it just came up with more stuff to build. It was like, okay, I think we should probably also have this programming language should have support for post quantum cryptography or something like this.

29:13

SPEAKER_01

[SPEAKER_00] And so I kind of think part of the fun of it is under-specify a little bit and see [SPEAKER_00] what happens. Obviously, if you want really good working production software, you should specify as much as possible. But yeah, yeah, that's kind of where I am. That's because I'm seeing this as just a way that I can do my work, and with this idea that I have. And I suppose I don't know, I don't know even where I picked that up, whether that is from, did I pick that

29:50

SPEAKER_00

[SPEAKER_01] from the ralph plugin? I probably did, didn't I? They have the completion promise and the max [SPEAKER_01] iterations. That's exactly what I do. That's exactly what I do. So that felt natural to me because what [SPEAKER_01] I wanted to do was choose a scope of work up front that was super well defined, and then just let [SPEAKER_01] ralph find its way to the end. Whereas what you're talking about there is choose an [SPEAKER_01] under-specified amount of work and just let ralph play in this zone, really. And what's [SPEAKER_01] to me, it feels like my version, you get to, that's a lot of upfront work for me, but it's kind

30:30

SPEAKER_00

[SPEAKER_01] of work that I want to do because I want to sharpen my ideas. I want to describe the desired state of the [SPEAKER_01] world really clearly. But yeah, I'm just interested in your thoughts there. What use cases do you see [SPEAKER_01] for an under-specified ralph versus an over-specified maybe ralph? I mean, I will also say, I've run ralph on a pile of spec. I've had ralph write the specs. I've been like, hey, go read all this documentation for this product and then extract out the clean room specifications that define actually how this would work from scratch, no implementation details. And then you run ralph on

31:08

SPEAKER_00

the specs. So you have one ralph building the specs, and another ralph is reading the specs, and then I think it was add ai features, and so it was having another directory full of specs that are just literally copy. And it's really almost lispy or pipeliney, where it's like you have a ralph that's taking the internet and turning it into specs. You have another ralph that's taking specs and turning them to what would this product look like if it had more ai? And then you have another one downstream that is reading those facts and turning them into working [SPEAKER_01] code. Yeah, wild. Okay, I guess it didn't go great. I didn't read the specs and I got a thing

31:48

SPEAKER_00

that I didn't love. And that's part of it too, is I think ralph is probably not the right final answer for how we build production software. I think it's probably, if anything, it's an incredible lesson in how context windows work. And that's kind of, I think, what jeff runs around doing, is like, don't use the entropic plug and learn the theory, because the theory is actually what makes you a better ai and coding engineer, and you should still use it, but you should [SPEAKER_01] understand why it works so well. Well, so let's define ralph as that kind of wild, freeform

32:20

SPEAKER_00

[SPEAKER_01] play around idea. Sure, there is an idea here, though, right? Running something in a loop that is really [SPEAKER_01] specified and tight, right? And that sounds like, if that's not ralph, then that's, I don't know, something

32:29

SPEAKER_01

else, right? But how do you, when you're talking to people who are doing this and they want long- running coding agents to just run afk, what's the structure that you recommend if it's not ralph? [SPEAKER_00] so that part, if you're doing full afk, this refactor plan, I can [SPEAKER_00] finish talking through how we did this and what I would do. The pr didn't get merged. Actually, [SPEAKER_00] here, I can pull it up and show it to you. I think that would be [SPEAKER_00] probably an interesting thing to look at. So we'll go to pull requests, close. You're gonna like the name of this

33:02

SPEAKER_01

[SPEAKER_00] one too. This is very low effort. Here it is, ralph is back, ralph is back. So this is 20 commits over [SPEAKER_00] six hours. Here is the react coding standards doc that I made with the engineer. Here's the, you [SPEAKER_00] can view actually the refactoring plan that it used to track all its progress. Let's see the rich

33:19

SPEAKER_00

so it's like, cool, we consolidated all the custom hooks, we added error boundaries everywhere, we made

33:25

SPEAKER_01

[SPEAKER_00] sure that we use proper forms, dates. It just did all this stuff that was in our spec. [SPEAKER_00] This didn't get merged. The reason why is because it's, thousands, how many lines is this?

33:35

SPEAKER_00

It's, yeah, 20,000 lines. And some of that is the plan files and stuff, but I was like, this was a cool experiment. I said to the engineer, he looked at it two days later and there was 100 merge conflicts, and I was like, okay, that's fine. We'll go do this later. You could always just rebase route, right? The nice thing about ralph too is you can just control c it, throw away all the code, and just same specs, update the code, and just run it again. And it was so low, it took 10

34:05

SPEAKER_01

[SPEAKER_00] minutes to set up the gcp instance and then took 10 minutes to turn it off later, yeah. I think if I [SPEAKER_00] was going to do something like this in a company for real, and a thing we've experimented with [SPEAKER_00] internally was a lot more towards do less. Ralph is a cool current state of the world, [SPEAKER_00] desired state of the world, make one change. I deployed something like this for an internal repo [SPEAKER_00] that we have, and I set it to only run. I run it on cron every night and only run three iterations of the [SPEAKER_00] loop. I want every morning to wake up to the code base, just for free. I just get the code base a

34:40

SPEAKER_01

[SPEAKER_00] little bit better. And we didn't merge all those prs, but every morning we wake up and just be like, [SPEAKER_00] oh yeah, this is great. And the specification is just here's how I want the code base to look,

34:56

SPEAKER_00

here's how I want it to be architected in the end world, and ralph just kind of does these little increments. So my advice is do not send your co-workers a 20,000-line pr that refactors the entire code base. That will not work in the real world. But you can use the concepts, and it's a building block that you can use to kind of hands off, don't think about it, just running in a github action [SPEAKER_01] every night and and and see what you get. Yeah, one thing that I'm thinking about actually, I was going A little bit better, and we didn't merge all those PRs, but every morning we wake up and just be like,

35:26

SPEAKER_00

"Oh yeah, this is great," and the specification is just, "Here's how I want the code base to look. Here's how I want it to be architected in the end world," and Ralph just does these little increments. So my advice is: do not send your co-workers a 20,000-line PR that refactors the entire code base. That will not work in the real world, but you can use the concepts, and it's a building block that you can use to hands off. Don't think about it, just running in a GitHub Action [SPEAKER_01] every night and see what you get. Yeah, one thing that I'm thinking about, actually, I was going

36:00

SPEAKER_00

[SPEAKER_01] to give this a go before we started, but we started early, so I didn't get a chance, but [SPEAKER_01] I have a couple of open source repos, and I just don't get time to triage all of the [SPEAKER_01] issues that come in. So there's an extremely simple Ralph loop here, right, which is you [SPEAKER_01] just feed all the issues into a loop, and you get it to work out whether it's, "I can categorize it," right, [SPEAKER_01] work out whether it's a feature request, which there's no real action to be taken, or you get [SPEAKER_01] it to make a reproduction of the bug or something, and then you maybe feed that to something else,

36:30

SPEAKER_00

[SPEAKER_01] another Ralph that's looking for a GitHub label or something to actually fix the bug, to make a PR. [SPEAKER_01] There's just a workflow pattern here, right, like you managing little queues and then Ralphs. Different Ralphs, different prompts are appropriate for different queues, yeah, and you tune those prompts [SPEAKER_01] with a bit of sitting on the loop and watching what it's doing, and then you let it go AFK overnight, [SPEAKER_01] as you're describing, and I suppose not too much at once is the way of thinking about it. I will also say you should be careful with taking GitHub issues from the community and

36:58

SPEAKER_00

feeding them to a clod that is in dangerously skip permissions, because it's technically untrusted input. We have a Ralph that runs through our Linear queue, but it is not allowed to see anything until we look at it and look for hidden prompts in HTML, markdown comments, and all of this stuff, and make sure, okay, this is actually okay for a model in DSP to process. Go, yeah, yeah, that [SPEAKER_01] makes total sense. Okay, so assuming then that you've got this loop set up, what are we talking about in [SPEAKER_01] terms of task size? Let's go there first. When you're prompting Ralph, yeah, the

37:33

SPEAKER_00

[SPEAKER_01] really nice thing I love about Ralph is that it frees you from the burden of choosing the next task. [SPEAKER_01] If you've just got a huge wadge of issues or something, as long as they're

37:44

SPEAKER_01

correctly described and you have the dependencies all mapped out, then you can just let Ralph do its thing and choose the next issue and then just go from there, which is just so nice. And when you're telling Ralph how big a change to make, what are you saying to it? What are you trying to get it as small as possible, or what do you think? So yeah, and

38:04

SPEAKER_00

in production, the tools we use are a little more human in the loop, again, because it's like the hardest software problems in the world are not going to be done completely

38:16

SPEAKER_01

[SPEAKER_00] autonomously. They're going to be done very much in collaboration with humans, but I think this [SPEAKER_00] advice is true in both worlds, which is when you're designing a plan, whether it's a plan you're [SPEAKER_00] going to give to Ralph or you're going to plan, we have a separate prompt that basically runs [SPEAKER_00] a parent sub-agent, runs each task in a sub-aid, sorry, parent agent, and then runs each implementation in a sub-agent, and the [SPEAKER_00] parent agent checks the work and then moves to the next phase, commits, and moves to the next phase.

38:46

SPEAKER_01

[SPEAKER_00] This is a plan that's been very vetted. The advice I keep finding myself, the models can't [SPEAKER_00] do, and it's the thing that a human needs to be in the loop right now, maybe we just need to [SPEAKER_00] make the prompts better, but models are not good at planning work in the way that I, a human, would

39:03

SPEAKER_00

do the work. I think there's something to be said for trying to steer them to, I don't know, I was doing something the other day with the buddy, and it was like, we have these 40 things on the front end, we're going to move them to the back end, serve them from an API, and then have the front end query the JSON and render it that way. And the model wanted to be like, "Cool, first we're going to move the 40 things, and then we're going to wire up the API endpoint, and then we're going to refactor the whole front end," three-phase plan, right? It's like, okay, if you were building this, what would you do? And it's like, well, I would move one thing,

39:32

SPEAKER_00

get the endpoint working, make sure one thing was working, probably learn some stuff in that process, and then hit some surprises. You want to minimize the size of the change in the same way that if you were an engineer, you wouldn't just copy-paste 40 files over and then go. Maybe some people would. I wouldn't, because I want tighter feedback loops, and I know there's going to be unknowns, and there's going to be blast radius to these changes that we didn't think about. I don't particularly want to sit, either you have someone with 15 years of experience in the code base, so they're just like, "Oh yeah, you're going to hit that issue," or you can sit there for three

40:07

SPEAKER_00

hours and research every single thing in the code base and try to get the plan perfect before you start, but there's a sweet spot here that is optimize for learning early on in the implementation plan, and then segment the tasks in the plan the same: how much code would you write if there was no AI? How much code would you write before you pulled up the web app and looked at it, or how much code would you write before you pause to run the tests? And that's a good task size for Ralph, or for a very human-on-the-loop AI where you're doing a smaller five- or six-step plan or something like this. That's my rule of thumb, always:

40:46

SPEAKER_00

your instincts as an engineer are still really, really good, and you should listen to them. Just because you're using AI, a lot of things change, but a lot of things don't. The programmatic [SPEAKER_01] programmer has this concept called tracer bullets, which is you should write code as it, you know this, [SPEAKER_01] right? Yeah, yeah, which is you should write code that goes through, that tells you where you're going, [SPEAKER_01] and that goes through all of the integration layers first, and you shouldn't do [SPEAKER_01] one huge change. It should be one change that goes through all the layers so you see that all the layers

41:16

SPEAKER_01

[SPEAKER_00] work, right? This is crazy. I read this book 12 years ago, and I loved this idea, and I've been [SPEAKER_00] explaining exactly what you said, this concept, to people for the last six months or whatever, and I forgot that it had a name. I love it. I literally read the book. I've had it on my shelf for three years, still in the wrap, but I got it out and I read it in one sitting the other day. It's just the most remarkable, well put together thing. Yeah, but this wisdom from 20 years ago is still super and that goes through all of the integration layers first, and you shouldn't do

41:50

SPEAKER_01

one huge change. It should be one change that goes through all the layers, so you see that all the layers [SPEAKER_00] work, right? This is crazy. I read this book 12 years ago, and I loved this idea, and I've been [SPEAKER_00] explaining exactly what you said, this concept, to people for the last six months or whatever, and I forgot that it had a name. I love it. I literally read the book. I've had it on my shelf for three

42:15

SPEAKER_00

[SPEAKER_01] years, still in the wrap, right? But I got it out and I read it in one sitting the other day. It's [SPEAKER_01] just the most remarkable, well-put-together thing. Yeah, but this wisdom from 20 years ago is still super [SPEAKER_01] important. It's, I would say, more important than ever, right? Because, yeah, and actually let's take a [SPEAKER_01] little sidebar on that, which is that if you're not reading those old books, we are programming in English [SPEAKER_01] now, and those books are written in the most clear, perfect way of describing good code that you're ever

42:41

SPEAKER_00

[SPEAKER_01] going to see, right? They are so, so good, and so I've been going back to loads of them, and I'm [SPEAKER_01] trying to buy four more now. It's just an amazing way to learn to prompt for coding, is reading those 20 real books. It really is incredible. I love this. I have another one that I've been obsessed with lately, which is, do you remember, I think it was Martin Fowler, it might be an Uncle Bob thing, but this idea of learning tests. Have you heard of this? No, no, no, I haven't. Okay, so a learning test is, I've been loving, especially if you're integrating with a system that you don't

43:12

SPEAKER_00

understand. We do a lot of integrations with the cloud code SDK, which the docs are getting much better, but the docs are docs, which means you can't fully trust them all the time, and it's closed source. So when we're going to build a feature, we'll get halfway

43:33

SPEAKER_01

[SPEAKER_00] through and be like, oh, we thought the behavior was this, but it wasn't, and the surprise would

43:41

SPEAKER_00

happen halfway through implementation, and then we have to throw everything out and roll back. And so a learning test is, you would build it in your unit test framework. You would put it in a place

43:52

SPEAKER_01

[SPEAKER_00] where it doesn't get run on every build because it's not for that, but it's a unit

43:58

SPEAKER_00

describe, expect, all this stuff, but it's for verifying how an external library behaves,

44:06

SPEAKER_01

[SPEAKER_00] whether it's bun.color or whether it's some giant battleship piece of software like [SPEAKER_00] the cloud agent SDK. You say, I think the docs say it works like this, I think it works like [SPEAKER_00] this. Go write a test that actually has assertions and console logs that explains the behavior of [SPEAKER_00] this thing, and often the assumptions were wrong. But what Claude is really, coding agents are really [SPEAKER_00] good at, is, okay, let me iterate on this until I have an understanding of what is the [SPEAKER_00] contract with this external system, and add a couple. The asserts are really nice because we don't

44:45

SPEAKER_01

[SPEAKER_00] run these all the time, but three months ago we wrote some learning tests about how session IDs work, [SPEAKER_00] and then they changed the session ID behavior, and I was like, I think this is wrong. Run the learning [SPEAKER_00] tests again, and it's like, yep, they broke the contract, or the contract changed, and here's [SPEAKER_00] how it behaves now. It's such a powerful thing to have to do up front as part of your research [SPEAKER_00] and design for building something new, and it's, again, a thing that, I mean, they were for this. [SPEAKER_00] They have not changed. The concept hasn't changed, and it's things that people have been doing for

45:25

SPEAKER_01

[SPEAKER_00] obviously you shouldn't write unit, because the wisdom is you shouldn't write unit tests on external [SPEAKER_00] software. It's their job to test it. You trust the contract and you trust the main container, and it's like, I don't know. We're on a total tangent here, but I had almost, when you quit a job and you think back, things you would have done differently on that job, we had a back-end team that was super unreliable and was shipping just crap. What I wanted to do was put, maybe their internal code was good, but they just kept breaking the tiny little JSON contracts that we'd have between front end and back end. They would

46:12

SPEAKER_01

put a null where things weren't supposed to be null and blah blah blah blah, misspell things, and so I just wanted to put a little Zod thing in between it just on the dev server just to

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note