Open Reader

Ralph Loops: Build Dumb AI Loops That Ship — Chris Parsons, Cherrypick

completed 1:48:25 May 04, 2026 Watch on YouTube

Current Status

completed

Video ID

2TLXsxkz0zI

RAG / Chat

Enabled
Ralph Loops: Build Dumb AI Loops That Ship — Chris Parsons, Cherrypick
Description

Dumb loops beat clever workflows. Most teams building with AI agents reach for multi-agent orchestration, planning graphs, and elaborate tool chains. Then they spend months debugging them. A single loop that processes one ticket at a time, evaluates its own output, and improves on the next run will outperform all of it. In this hands-on workshop you will build three things. First, a working Ralph Loop that processes real tickets end-to-end. Second, a synthetic feedback loop so you can test and iterate locally without waiting on production data. Third, a self-improving cycle where the loop's output quality gets better with every run without you touching the prompt. Speaker info: - https://x.com/chrismdp - https://www.linkedin.com/in/chrisparsons/ - https://github.com/chrismdp

Summary

Generated by claude-haiku-4-5-20251001

Ralph Loops: Build Dumb AI Loops That Ship

Main Topics

  • Ralph Loops Overview: Simple, repeatable loops where AI agents perform tasks iteratively, getting progressively better results
  • Evolution from Complex Workflows: Moving away from complicated orchestration tools (like N8N) to simpler AI-powered loops
  • Practical Implementation: Building a Pomodoro timer as a hands-on workshop example
  • Scaling AI Work: Using loops for everything from code development to business operations
  • Team Coordination and Constraints: Managing multiple agents and understanding system bottlenecks

Key Points

What are Ralph Loops?

  • Named after Ralph Wiggum from The Simpsons (who tries the same thing repeatedly until it works)
  • Core concept: Give AI a task, let it complete, then ask it to do the same task again
  • Modern AI models (Claude Opus 3.5+, GPT-4 Advanced) are much better at noticing incomplete work and self-correcting
  • No magic involved—just simple iteration that exposes what was missed

Why They Work Better Than Traditional Workflows

  • Old approach: Complex orchestration with all dependencies specified upfront (like N8N)
  • Brittle and fragile
  • Requires constant maintenance
  • High failure rates
  • New approach: Simple loops with good feedback
  • More coherent results
  • Self-correcting through iteration
  • Minimal maintenance overhead

The Loop Concept Is Universal

  • All work is essentially loops: pick next task → do task → iterate
  • Software developers: backlog → pick → implement → review → merge → release → repeat
  • Can be applied to: code, email, content, project management, operations, everything

Implementation Patterns

Basic Ralph Loop:

`

while true:

claude implement ticket_001

`

Advanced Loop with Cron:

`

loop every minute, implement next ticket from doc/tickets

`

Skills-Based Approach:

  • Package context and instructions into reusable "skills"
  • Version control skills like code
  • Pull skills into context when needed
  • Allows AI to know what tools/approaches are available

Key Success Factors

  • Clear Stopping Criteria: Tell the AI explicitly when to stop (context limit, irreversible actions needed, etc.)
  • Good Feedback Mechanisms: Build ways for AI to know if work is done well
  • Proper Context Management:
  • Fresh context per task avoids pollution
  • But larger context windows make this less critical
  • Reversibility Rule: Only let AI do things that can be undone without embarrassment
  • Version Control Everything: Tickets, skills, outputs—all need to be tracked

Practical Validation Approaches

  • Use sub-agents for validation (prevents confirmation bias)
  • Run adversarial reviews where a separate agent critiques the work
  • Simulate different audiences/personas to test from multiple angles
  • Implement comprehensive testing (unit, integration, end-to-end)
  • Use screenshot validation for UI work

Notable Quotes

> "The dumbest Ralph loop is literally that, just a while loop. And it just goes through and implements stuff."

> "All that a Ralph Loop is, is build this thing... and then it finishes and says, OK, I've done the thing. And then it says, OK, great, go and build this thing... and it goes, OK, I'll do it again."

> "Ultimately, because Claude is non-deterministic anyway, I think there's a high level of variability with any of those kinds of tests. So it's really difficult to think about how to construct a useful test in that way."

> "Is this reversible without embarrassment to me? And if the answer is no, don't do it."

> "If you don't work on that one bottleneck, all of the other work that you might do to optimize and improve the system is pointless and actually probably counterproductive."

> "Everything in fact is a Loop... Maybe as an engineer, I'm definitely on a Loop... Maybe as a CEO, I'm on a Loop... Maybe a lot of the cadences that I work on run in Loops too."

Takeaways

For Individual Developers

  • Start simple: Use basic Ralph loops before adding complexity
  • Don't optimize prematurely: Token costs are cheap; time is expensive
  • Build up your skills over time: Iterate on your prompts like you would code
  • Be intentional about what you keep: Only do work that plays to your unique strengths
  • Use loops for everything: Email drafts, content, code, project management

For Teams

  • Make it sequential first: Parallelization is harder and often unnecessary
  • Fix bottlenecks systematically: Use Theory of Constraints—identify and fix the main constraint first
  • Avoid over-specification: Waterfall approaches with AI fail; use just-in-time planning
  • Enable experimentation: Leaders must give teams air cover to try new approaches
  • Keep teams small: Coordination overhead often exceeds the benefits of larger teams
  • Use good coordination mechanisms: Ensure tickets are properly claimed and updated

For Security & Operations

  • Implement sandboxing: Use VPS, Docker sandbox, or separate environments
  • Understand the "lethal trifecta": Untrusted code + internet access + secret data = data loss
  • Use fine-grained permissions: Separate API keys, read-only access where possible
  • Never let AI send emails: Only draft them for human review
  • Review security-critical code: Don't trust AI with database migrations or customer data handling

For Knowledge Management

  • Use a vault approach: Markdown files + embeddings (like Obsidian + embeddings tools)
  • One note per thought: Zettelkasten method for discoverability
  • Include decision trails: Document why decisions were made for future context
  • Let Claude help organize: Have AI sessions clean up and structure your knowledge base

Broader Implications

  • Existential question: What work do you actually want to do? Automate the boring stuff, keep the strategic thinking
  • Future of work: AI handles "commodity" tasks; humans focus on unique value creation
  • Tools still evolving: Don't lock into current structures (like spec-driven tools); be flexible for future models
  • Skills marketplace needed: Better solutions needed for sharing and managing AI skills across teams
  • Everything is becoming a loop: From daily standups to startup operations—think in loops

Recommended Reading

  • The Goal by Eliyahu Goldratt (Theory of Constraints)
  • Simon Willison's articles on "lethal trifecta"
  • Andrew Kaparthy's article on "LLMs as a Wiki"
  • MCP (Model Context Protocol) documentation

Transcript

20489 words en Processed in 872.0s

Welcome. So this workshop is on Ralph loops. Hands up here who knows what a Ralph loop is. That's almost everyone. I'm guessing that the other folks who came in were just here because they thought that sounded weird or maybe looking for a quiet place to work. I don't know. But you're very welcome. So what we're going to do today, this is a two-hour workshop. We're going to, if you could just make a little bit of space if you need to as people are coming in, that would be really helpful. Thank you. This is a two-hour workshop. We're actually going to build Ralph loops together. We're going to do this together on our own laptops in order to make some stuff happen and get some things done. So it's not just about theory. This is a very practical thing. So if you've got a laptop, you're welcome to get it out in a second. We're actually going to try and do this ourselves. So I have a few slides, but not many. Most of this is going to be live demos and interaction points as well. And the idea is that at the end of this, you should be able to leave with something that works that will apply it to a toy code base just for fun to create a Pomodoro timer. But hopefully the idea is that you'll be able to use this on your real work when we get done. So another show of hands, that's okay. Who is using Claude code or Codex specifically to write code? Hands up. Quite a lot of people specifically to write code. Who is using it to write all their code? Who is no longer writing any code? That's quite a lot of people. Look around for a minute. That is a huge change. If I'd asked a bunch of programmers six months ago who was not writing any more code, you get a very different answer. So next question. Who is using either Claude code or Codex, and I'll include Cursor here as well, in their non-coding work? Okay, quite a lot of you. What about for all your normal non-coding work? Okay, interesting. Interesting. So you can just see the future in the room, right? There's a few people who are starting, but we're still on that journey for sure. And who has built Ralph loops before? Last show of hands. One or two people. Okay, great. I'm going to be looking to you for all the answers. So that's great. So just a little bit about me. My name is Chris Parsons. These days, I spend most of my time trying to help teams like the team I used to run figure out what on Earth to do with AI mostly. So I'm a CTO by background. I've done a couple of VC-backed startups and scale-ups. And this has taken me and my friends by storm rather, and we are trying to still all figure it out professionally together in terms of how to help our teams adopt and use AI. So that's what I do for a living these days. I have about 30 years or so of building software professionally. I've been the CEO of an agency. I've done a lot of agile consulting, remember that? And that kind of training as well back in the day. And funnily enough, all of those, it's a whole other talk, all of those principles and practices that we taught for years and no one really listened were still very much applicable to AI. So there we go. So these days, I'm actually running Ralph loops all the time, 24 hours a day to get my work done. So I'm using them to write my emails. I'm using them to check my calendar. I'm using them to write content and newsletters. I'm using them to help me do my client work. So I'm using them in absolutely everything. I also use them for code, which is what we're focusing on today, but they are very much applicable to every part of our lives. So by the end of the day, the idea is that you will be in the position where you can do that too. So this is how I used to work with AI until quite recently. You probably can't see that very well. This is an N8N workflow that I used in order to create my weekly newsletter. It took me probably a week to write, let alone actually test and debug. It's got a huge number of different things in here. This is like a featured article flow, which would read various different articles from my blog and figure out whether I posted it before, summarize it using AI and put it there. And then there's another one for grabbing links that I'd posted into a particular list. And it did a bit of commentary on that. It was really quite complicated and difficult to run and maintain. And it kind of worked okay, except that 2pm on a Monday, pretty much every Monday, I would get the dreaded notification from N8N that my workflow had failed. And I was just like, oh no. And then I'd have to go in and figure out in here what whatever had broken and try and run it and fix it. Now this is nothing against N8N. N8N is a really cool tool and it can do some really cool things. And I'd never have been able to orchestrate AI in the way that I was doing before without a tool like N8N. Despite being a coder, it's just so much easier to manage in here and you can manage all the API keys really easily. And it's a nice tool to stick things together. But it was so brittle to use at this kind of level of complexity. And I didn't get a huge amount of value out of it. And then every time I fixed it, I would do something else. And so honestly, it was probably easier for me to just write the newsletter than it would have been to maintain the thing that wrote the newsletter. And I probably had a slightly better newsletter. So this wasn't great. But this was, to my mind, a few months ago, really the only way to use AI. You had to kind of orchestrate it and manage it, give it the right data and handle all the context. And I thought that this was the future of automation, but it isn't really the future of automation. The future of automation is a lot more like something like this, running in Claude code. So this is obviously not the actual skill. But I have now in Claude code, a skill that writes newsletters for me. And it has all of those instructions. In fact, I copied and pasted the N8N JSON code from there into Claude code and said, write a skill based on this flow. And it did a great job. And then what it does is it goes through and does all of the things. But what's interesting about that is how does Claude code work? Well, it reads the first thing, it decides on the next step, and then it reads the next bit, and then it decides on the next step. And it kind of works through over some minutes to actually write and produce the newsletter that I was writing. And it's the same for code. It's the same for anything that we want to build using Claude code. Claude kind of just takes care of it. You describe the kinds of things you want, and it does it. Now, what's interesting is that Claude code fundamentally is running on a loop, isn't it? It just reads the skill, calls a tool, goes back to Is how does Claude code work? Well, it reads the first thing, it decides on the next step, and then it reads the next bit, and then it decides on the next step. And it works through over some minutes to actually write and produce the newsletter that I was writing. And it's the same for code. It's the same for anything that we want to build using Claude code. Claude just takes care of it. You describe the kinds of things you want, and it does it. Now, what's interesting is that Claude code fundamentally is running on a loop, isn't it? It just reads the skill, calls a tool, goes back to the beginning, reads the skill again, calls a tool, calls a tool. And then at some point, it figures out that it's done, and it stops, and it gives you your newsletter in whatever form you want it in. So what's interesting is that this ships much better, more coherent newsletters than the previous workflow. I still have to change them, write them, screw around with them, but they are a much better first draft than they ever were. And I haven't really touched this skill. All I really do with this skill is I say at the end of a newsletter writing process, please just update the skill with anything you can figure out from the session you should have done differently. And it makes the odd tweak here and there. So that's a loop. That is this form of working in loops with AI. So agents that originally start in workflows where you have quite complicated orchestration that looks a bit something hellish end up in quite a simple loop, perhaps with a bit better context. Now, this didn't work for the longest time, but it's now beginning to work with the latest models. And by the latest model, I really mean GPT 5.x, really GPT 5.12 onwards, and Claude Opus 4.6 or Sonnet 4.6 upwards. So those models started emerging around about the end of November. I've no idea about Mythos, by the way. I've spoken to people who've used it and they say it's good, but it's mostly marketing, but we'll see. But yes, we'll see where that takes us. Maybe we won't even need skills. Maybe we'll just say write a newsletter and I'll do it. Who knows? But what I'm trying to... My point is that rather than using complicated workflows, we're actually using skills and loops much more in context and loops. And any agent that we run is in some way a loop already. And this powerful looping construct is something that you could more generally apply. So what happens when we take loops a little bit further? So the first stage is this idea or the first idea came from Geoffrey Huntley a little while ago, ancient times in AI, which means probably about last June. And he said, basically, what we should do is whenever we finish using an AI to do anything, we should just try the thing again in some way. We should just give it exactly the same prompt and see what happens, see what it does again. And it sounds a bit stupid and it's based on... Does anyone know where this story comes from? Who knows why it's called a Ralph Loop? Like two people. It's called a Ralph Loop because of Ralph Wiggum, which is a Simpsons character who basically says, he just tries the same thing over and over and over again. And eventually it works. And it's all it is, really. All that a Ralph Loop is, is build this thing or do this thing inside a prompt. Then the AI goes away and does the thing. And then it finishes and says, OK, I've done the thing. And then it says, OK, great, go and build this thing and do all the things I said to build this thing. And it goes, OK, I'll do it again. And the groundbreaking nature of what that meant was that the AI would often review its code and realise it had missed something in some way, right? So it figured out that it wasn't quite finished. And this is quite a common problem with AI coding tools last year. It wasn't quite done. It didn't quite get to the end. And therefore, it would say, oh, yeah, I should have fixed that bit and then does it again. And then when it stops, it says, right, I've definitely finished now. 100%. It's done. It's finished. And then what you do, you give it another prompt again. Say, go away and build the feature. It's like, I've built the feature and then tries it. It looks like, oh, yeah, there was actually this tiny thing that I should have done. I really am now finished and so on and so on. So you can see the utility of going through that loop where you just build the feature and just then ask it to build the feature and then you ask it to build the feature. So that's the first stage of our loops. And what I'd like us to do is I'd like us to try that. So firstly, I'm going to do a bit of live coding. Hold on to your hats. We'll see how that goes. And we're going to try and do that process using Claude code to see where that takes us. So let me start change what I share. I'm going to start changing what I share. Sorry, it just takes a moment. Okay. Great. This one. Can everybody see that? Okay. Can people see that at the back? Okay. Do you want me to increase the size? Got thumbs up? Great. Okay. So this is a piece of code that I vibe coded in about three minutes last night. So it's not good. But that's the whole point. We're going to fix it. So it is literally a Pomodoro timer. And you can see how it works. If I go to Python and type Pomodoro start, woohoo, we've got a Pomodoro timer. That's all it does. It literally just does start. There is no way of finding out whether it's finished or complete or anything like that. But that's what we're going to change. The other cool thing, which is very important for any self-respecting vibe coded AI project is that it has tests. So look, it's got a test. There's one test. And the check to see whether it starts. So that's great. So if we just have a quick look, and you have to forgive me if you're not a Vim fan, because I am. Although I hardly use it now, it's quite sad. You know, 20 years of muscle memory just gone. But yes, so all it does is it literally just runs a start command, and then it saves in your .pomodoro in your home directory. It saves the time in which you started. It's really, really simple. So this is a very simple, quite straightforward project. The difference is that it has a new folder with different things in it. And these are tickets. So there are a whole bunch of ways in which we could improve this Pomodoro timer. And the first ticket is it would be really nice to know how long is left on our Pomodoro rather than just starting it. So what I've done is I've created a very simple ticket system to allow us to just capture some changes. These are not, and then it saves in your .pomodoro in your home directory. It saves the time in which you started. It's really, really simple. So this is a very simple, quite straightforward project. The difference is that it has a new folder with different things in it. And these are tickets. So there are a whole bunch of ways in which we could improve this Pomodoro timer. And the first ticket is it would be really nice to know how long is left on our Pomodoro rather than just starting it. So what I've done is I've created a very simple ticket system to allow us to just capture some changes. These are not this is one-shotted. So I have no idea if these tickets are actually good. In fact, I haven't looked at some of them. So we'll see how that goes. But the idea is that we can use these in order to start building a loop of work in order to get something done. So what I'm going to do to start with is I'm going to start Claude. And I'm literally going to say, write the first ticket. So bear with me while Claude fires up. In some ways, I'm quite glad they didn't actually release Mythos yesterday, because I think I don't think it would be working today if they did. That is really not working, is it? That's frustrating. Wow. Let's try again. I think there's a problem with Wi-Fi. Ah, okay. I mean, that could cause some problems to my talk, but we'll have to see how that goes. I don't have one of those fancy new Macs that allow you to run. No, this is actually locked up my computer. Can you believe that? It was working literally 10 minutes ago. Let me just have a look. Bear with me while I debug my machine. There is, I've got, I think I'm on a different Wi-Fi, so it should. No, yeah, the Wi-Fi has gone down. Fun. Tethering time. Might have to be. Okay, hang on a second while I tether to my phone, which I think has decent 5G, so we should be good. Okay. Let's see if that's any better. Hooray! Okay, let's try again. Not that one. Cool. So, code Pomodoro Workshop. Okay, is that big enough? Okay, good. Let's try this. Claude via the power of 5G. Look at that. It works. Fantastic. Okay. So, what I'm going to do is, we have, as I said, a very, very simple, stupid Pomodoro timer, and we're going to implement a ticket. So, what I'm going to do is, I'm going to say, implement this ticket. So, it's in doc tickets 001. Great. And I'm just going to say that, see what happens. So, what I'm going to do is, I'm going to read the ticket, which I showed you briefly earlier. It's very straightforward, and all it does is it implements a status to see how far we've got. And then, what I would like it to do, what I'm going to do after this is, I'm then going to say when it's done, because it will literally be two files. It's not going to be difficult for it to do. And then, I'm going to say, implement it again and see what happens. So, there's a few different ways that you can do this. And there is no one set way of doing a Ralph loop. It's really about the concept, not about anything else. Great. So, it's done the ticket. If I just quickly do a quick get diff, you can see that what it did is it added a status command. I think when I had a show of hands earlier, most of you are coders. So, hopefully, this is not tricky to follow. And then, we've got a new test. Look at that. It added a test. It didn't even ask it to. It added a test. Oh, my gosh. What is the world coming to? So, now, what I'm going to do is, I'm literally going to say the same thing. Now, a year ago, this would have been a really important step, because it would have definitely missed something. Whereas now, it's, we've already done it. It's fine. Right? So, Opus is now much better at noticing when things are done. Now, a traditional Ralph would just keep doing this. Right? And it would keep going with this implement this ticket, implement this ticket, implement this ticket. And this is boring. And it's not really going to do very much else. And at some point, actually, when I tried this earlier, it's interesting. It's done something different. It actually noticed that what it should have done is it actually should have updated the status to done. So, the process worked. That didn't work earlier. So, that's great. So, it's actually noticed something that it didn't do. So, there you can see the fundamental early principle of early Ralph loops. Right? The idea that you can just zoom through and do something. And it will find things eventually that it missed. Because it missed that, I'm just going to try once more. But I don't think it will come up with anything else. As I said, latest models really don't need this step in quite the same way. They tend to just get it done. And in fact, this time, it's just, oh, if you're running a Ralph loop that picks up the next ticket. Oh, that's hilarious. It's literally giving away my presentation. That's fantastic. Okay. What I'd like us to do is, I think as a starting point, the other thing you can do is you can just kill the context. And then you can do the same thing again. And you can say implement doc tickets 001. And I can't bother to spell it. It'll find it. And then, so now, what I'm doing is basically doing the same thing, but without it knowing about the previous context. So, it'd be quite interesting to see what it does with this. I'm assuming it'll find, assuming it found the ticket. Yeah, it did find the ticket. That was easy enough for it to do. It's just running the test to make sure that they work. And it all passes. Okay. So, it's happy. So, some people, when we first started using Ralph loops, is that they weren't doing it within the same status. And there was an early Claw plugin that just, on the stop hook, which is what runs right now, when it stops running, it would just do the same command again. So, rather like me just typing the same thing in each time. But that didn't really work very well, because it didn't get very far. Whereas now, what's more useful, or what people started doing, was just running Claw code in a loop. So, they would do something like while true, and then do Clawd implement ticket 001. Right. And then done. And then that will just go through, not quite actually, because I didn't do Clawd P, but effectively, that's what they were doing. Oh, no. Now I've really screwed it up. I really shouldn't have hit enter on that. Should I? So, rather like me just typing the same thing in each time. But that didn't really work very well, because it didn't get very far. Whereas now, what's more useful, or what people started doing, was just running Claude code in a kind of loop. So, they would do something like while true, and then do Claude implement ticket 001. Right. And then done. And then that will just go through, not quite actually, because I didn't do Claude P, but effectively, that's what they were doing. Oh, no. Now I've really screwed it up. I really shouldn't have hit enter on that. Should I? Okay. There we go. But that's effectively what people were doing. The dumbest Ralph loop is literally that, just a while loop. And it just goes through and implements stuff. Super, super easy. So, what is the next step in Ralphloot? Well, in fact, what we're going to do now is I'm going to get you to get to that point. And then I'm going to take questions from other folks. So, let me just switch back to here. I'm hoping. There we go. Yeah, great. So, what I'd like you to do, if you could crack open your laptops and grab the code from here. So, it's just on my GitHub as Pomodoro Workshop. You should be able to find that quite easily. And you saw how I ran it. It's very simple. You might need to set up Python in your machine. So, hopefully, that won't be too hard. Or you just run bare Python. Pomodoro.py will give you the command you can type. And then it's just a unit test thing to run the test Pomodoro. Super easy. And then in step four, I want you to fire up Claude code or codex. And I want you to try and build that ticket and make sure that it's working. And that would be a great starting point. But don't build any more tickets yet. Don't let it take you too far. And then if you are really used to that, and that is just a literal no-brainer for you, try it in codex. Try it in something else. Maybe try setting up something similar in one of your own projects. So, something different. So, I'm going to take questions now while people are typing away on that. I'm going to give you maybe a few minutes just to get that set up. And then we'll move on to the next step. Does anyone have any questions or comments or thoughts? I have a microphone here. If people would like to ask anything. Yeah. There we go. It should come on in a second. Hopefully. The guys at the back are... It may not be on. Is it on? Do you have to shout and repeat something? Yeah. That could work. Can I just check it's on first? Yeah. It looks on. That's weird. Shout and repeat anyway. Oh, there we go. There we go. Great. I've played a bit around with the BMAD method, which I don't know whether you've seen that. No. He's got a guy who's basically written a whole load of skills and commands for following a full agile process from... Oh, yes. I think I have seen that. He's got an agent for build it, test it, everything. And I guess... So have you done anything where using this kind of railflute process, you go through that cycle, you go through the full software development life cycle of each stage and then? Yes. I might get as far as that at the end. But yes, I have tried some of that stuff. It's really interesting. And it asks some really very good questions both around context and actually the value of the work. It's really interesting. So, we'll talk about that a bit more at the end. So ask again. If I haven't got to it, just ask the same question again and we'll get there in a raffling. Okay. Thank you. Anybody else got any questions? How are people getting on with setting that up? Has anyone managed to set it up? Wave at me if you've managed to get it running. Great. Great start. Has anyone managed to implement the first ticket? Hooray. A few people. You got a question? No, no, no. Oh, go, go. You've done it. Great. Fab. Yeah, I probably should have given different directions for asking a question versus having finished. That's great. So, a few people have got started. Fantastic. Great. So, you can probably tell where this is going. And if you were paying attention to the live demo, you'll already know the answer. I'll grab the mic from you so you don't have to keep holding that. Thank you. But yes, you don't have to just stop at one ticket. Now, is Matt Pocock happened to be in the room? I know that he's doing the workshop after this one. He is the person I got this from. So, if he's watching the video, thank you, Matt. This was a revelation to me back in September last year. He posted a brilliant YouTube video just about exactly how to take graph loops to the next level. Because when they first came out, I spotted it on the internet. I played with it. I was like, yeah, this is fine. It's cool. It kind of spots things that AI can do where it's missed things, and it can maybe do a slightly better job of things. But it's not going to change everything that I do. And then the answer is actually it does change the entire way that I work and approach code now. So, I guess the really interesting thing is not how do I make sure Claude has finished this one thing. It's what happens if I point this kind of loop at a whole pile of things to do, right? What happens when we point at a whole list of things? Now, I tried this. I wrote a blog post about this, which was a bit depressing because it just showed abject failure in the entire post, to be honest. But what I tried to do is I tried to get Claude, I think it was Claude at the time, to break up a big project into a lot of different tickets. And then I got it to break down all of those tickets into smaller tickets. And then I got it to figure out what all the dependencies were between those tickets and write them all down really carefully. And then I got it to figure out how it could use a ton of different agents. Sorry, did you have a question? That's fine. You've got someone there with a mic if you wanted to say it. One, two, one, two. Okay. I have a problem with, I had a problem with Wi-Fi and I didn't do the clone. Oh, I'm sorry. You wanted to go to this one. Yeah, a couple of minutes. Yeah, thank you. That's fine. I'll leave it on there. That's fine. Has everyone appreciated the slide? Okay, good. Okay, there we go. So, these slides, by the way, are created using a slide skill. agents. Sorry, did you have a question? That's fine. You've got someone there with a mic if you wanted to say it. One, two, one, two. Okay. I have a problem with, I had a problem with Wi-Fi and I didn't do the clone. Oh, I'm sorry. You wanted to go to this one. Yeah, a couple of minutes. Yeah, thank you. That's fine. I'll leave it on there. That's fine. Has everyone appreciated the slide? Okay, good. Okay, there we go. So these slides, by the way, are created using a slide skill. Using Nano Banana Pro is absolutely incredible at making slides. I didn't, I haven't, apart from the tiniest thing, adding a QR code, I haven't added any text to these slides. They're just flat images in Google Slides. They're absolutely incredible at making slides. What was I saying? Yeah. So the idea of just creating a Ralph loop to just do one thing seemed a bit pointless and it was just working around a few limitations. If you point this loop on a whole pile of work, then it becomes incredibly powerful. And as I was saying before, I created this huge complex dependency graph with a whole ton of different tickets about how it was going to build this really complex system for me. And then I got, I fired up like six or seven parallel agents and I was like, right, you do one stream and you do that stream and you pick up this ticket. And it just failed horribly because the system just couldn't figure out what had been done and what hadn't been done. And then picked up, there was lots of contention between tickets. Like, well, I can't do anything until you, until I get that shared ticket done. So I'm going to do it. And another Claude was like, well, I can't do anything until that shared ticket's done. So I'm going to do that too. And then they both implemented the same thing and it was a huge mess. So that was really very depressing. And I wrote a whole thing about it. And I was basically like, it's impossible to orchestrate large numbers of agents. You just can't do it, which was obviously nonsense. But that's how I felt at the time. And what was interesting was that what I'd done effectively was recreate the waterfall processes that were seen in some of the worst companies back when I was starting to code for myself in the 90s, where people would write requirements documents that you had to stagger to carry into the requirements meetings, where the entire project was specified up front with all the intricate dependencies, handed to the development team and then given two years to build. I can see some, perhaps, slightly more seasoned people in the room sort of nodding and smiling at me when they hear me talk about this. But yeah, I thankfully managed to avoid working on any of those teams. But some of my friends did and it was absolutely awful. And what I had done was basically I had given that to Claude to do. I'd given that waterfall process to Claude to organize and figure out as it went, which was really bad. So no wonder it didn't work. If humans can't do that, how was AI supposed to do any better? However, if instead of saying with all of your tickets, right, the first one is the most important, then this one, and then you should do this one. But don't think about this one until you've done that one. Instead of doing that, I'm going to go back for just a minute. Are we all good with this slide, by the way? Does anyone still need the slide? Okay. Okay, we're good. Right. Instead of doing that, you can just run a Ralph Fluke where you say something like, hey, just pick the next most important ticket. It's as simple as that. Just figure out. Here are all the tickets. Just figure out what is the most important next one to run. Okay. You don't have to worry about the dependencies. You don't have to figure it out yourself. The AI is quite capable of looking at all of them, figuring out the dependencies on the fly based on what's just been done and figuring out what the next most important thing to do is. That's actually quite easy for an AI to do. The one thing it cannot do so easily is manage that process in parallel. But to be honest, when we're running these kind of loops, if you're running them continually, the bottleneck is usually not the number of agents. It's usually you just keeping up with the AI, just doing things over and over again. So let's forget parallelism just for a minute and just start with a loop. See, if you can keep up with an AI, just one AI that's running continuously, you're fine. Don't worry about parallelism just yet. Don't worry about Gastown, any of that stuff just yet. As impressive as those projects are, you can just start with a simple loop. It is okay. So what I'd like you to do again at this point is I'm going to quickly show how this works for those of you who don't have your laptops, but then I'd like you to just try it on your computer. So again, let me find my mouse and then move to sharing my screen again. Okay. Great. So if I go back to Claude, in fact, I'll go back to Vim first and look at the tickets folder. So I've got a whole bunch of tickets here. I've got a status command. I've got a stop command. I've got custom durations. Never use that anyway. I use other things like labels and all of that. I could try and figure out the dependencies myself, but I really just don't need to. I can just simply go into Claude and say, implement the next most important ticket using TDD principles from doc tickets. Commit when done. Something like that. Okay. So let's see what it does. So it's now reading a whole bunch of tickets. As you can see, it's read number one, two, and three, and it's decided that the next one is number two. It's just going to do it. That's great. So it's going to work on it. Now, the interesting thing now is that once this finishes, now it's using TDD, so it read the test first. When it finishes, hopefully... Yep. It's marked it as done. Very good. Remember that time. And then it should commit. Let's see if it does. Sorry, is this in the same session as the previous? No, this is a brand new session, although I think it probably had the working directory from the previous one still. So... In fact, I think it has. So what it hasn't done is committed those atomically, which is definitely something I could improve in my prompt, but we can cover that in a minute. So then hopefully it's just going to do that. And then it's going to finish. [SPEAKER_10] Yep. Great. It's done it. Fantastic. Now what I can do is I can either do that again as a RALF loop, or I can just restart a new session, just do the same thing. And this time, it'll pick something else, and then it will keep working. Now, you can imagine that if I put this the working directory from the previous one still. So in fact, I think it has. So what it hasn't done is committed those atomically, which is definitely something I could improve in my prompt, but we can cover that in a minute. So then hopefully it's just going to do that. And then it's going to finish. [SPEAKER_10] Yep. Great. It's done it. Fantastic. Now what I can do is I can either do that again as a RALF loop, or I can just restart a new session, just do the same thing. And this time, it'll pick something else, and then it will keep working. Now, you can imagine that if I put this inside a while loop, then it should, in theory, work through all of the tickets in some way. Now, whether what I get at the end is actually what I want is a whole different question. But it will definitely get a lot of work done in a row. So I'd love you to try it. So if you've got the app working on your computers, see if you can get it to work through just as many tickets as you want to within that amount of time. So it should be able to just carry on. See if you can get it to actually maybe write a little bash script, just like I've done, where you do a while true and then clawed. [SPEAKER_13] I'll show you how to do that briefly. In fact, if I just quickly get reset, so it can start just from the beginning. And in fact, I'm actually going to go up one more. Hard, head, there we go. Great. Yeah, that's the right place to start. So if you can get it to do that, then that's great. The other thing that you can do is instead of using clawed like this, you can do clawed-p. Can you all see that? Okay, by the way, I'm not sure how I can make that. Yeah, clear is a good one. Yeah. Clear. Clawed. There we go. So what we can do is actually use clawed.p like this. And you can get it to output by just doing stream.json or something like that. What's it called? Something like stream.json. Hang on a second. Let's see. I think they've removed it. That's annoying. Never mind. We just won't see any output. So there's nothing to stop you setting it up like this. And then you can just do that. But you've got to set up clawed to have full permission so it doesn't properly. Yes. So the only way that this works is if you want to run this properly and for it to not stop, you have to be quite selective about the permissions that you give it. So the question was, presumably, you have to run clawed with full permissions for this to work. Yes. Yes, you do. It depends on what you're doing. Yeah. If you're working in a little sandbox project like this, the chances of it going elsewhere to find stuff out is very small. I have a project called Lockbox, and the sole purpose of that is to try and stop it doing stupid stuff. By when it reads untrusted tokens, which could potentially send it off track, it basically just prevents any kind of file system access or anything after that. So there are ways of managing it. So what this is doing in the background is you can't actually see it doing anything because I don't have that output mode. But you can figure out basically to run this in some kind of script. In fact, if I quit that, hopefully it will stop. There we go. No. Let's just keep going. Sorry. Clearly, a more production-ready RALF loop would not look like this. But you can see it's done a bunch of work. So if I go to here, you can see it's already started on the status command. And it's just working through that at the moment. So it started at the beginning again. It was just working. So there's a few things to be aware of here. One is that feedback is really, really important with RALF loops. You need to be able to have it run in a way that you can tell what it's doing and how it's doing it. So this kind of super basic one that I've given you there isn't very good. That's not one that I would recommend running in production. Equally, you need to figure out exactly what the prompt is for RALF. And that's a really, really important point. And I think what I'd like you to do when you're trying this is, yes, it's going to be building a bunch of tickets in a row, but equally it's going to be doing them in a way that you don't like. So for example, if I'm running this test, which is literally just implement the next most important thing, let's start with this one. I would probably do something like run simplify, which is a really useful skill from Claude, from the Anthropic team, when finished, and ensure you refactor to reduce duplication. You can imagine that you can create quite a complicated skill for this. And I'll show you my kind of actual skill for this at the end. But as you're working through this, do try and figure out if there's ways that you can improve what it's doing. So give it a go, let it make a decision, and then let it write some code, and then read the code and think, okay, what could I have actually improved about the process, and then reset everything and then improve the prompt after that. Have a go at that and see how that works. Whilst people are working through that on their machines, I'm happy to take questions. Yeah. Have you used the skills like superpowers? Which one, sorry? Superpowers one. The superpowers one, what I did is I pointed Claude at the entire repository and said, figure out anything that isn't currently in my skill set and implement them for me with my own context. And that worked quite well. So I haven't used those ones particularly, but I basically ripped them off. That's great. Then, because I use superpowers a lot, and then I just give tasks like this, and then ask it to run multiple agents in the background. Yeah. And how you've done that is what I've... Yeah, yeah. So what you can do is there is an agent teams version, which I think I've got turned off in this particular instance of Claude code. But what can happen is you can get Claude to use sub-agents within Teamux. So in fact, I think I might be able to turn that on if I can find the agent teams. There it is. Claude code experimental agent teams. So if I grab that and then run Claude with that on, then you should be able to say, use an agent team to implement the dot tickets in this repo, or something like this. And I don't... This isn't actually running within Teamux. So actually, maybe this won't work. So I might just try this again. Bear with me a second. So if I grab that, and then paste that there, and then grab that, and then run Teamux, and then run that, in theory, in theory, this should start pulling up other agents. And because, like I said, trying to orchestrate Claude code experimental agent teams. So if I grab that and then run Claude with that on, then you should be able to say, use an agent team to implement the dot tickets in this repo, or something like this. And I don't... This isn't actually running within Teamux. So actually, maybe this won't work. So I might just try this again. Bear with me a second. So if I grab that, and then paste that there, and then grab that, and then run Teamux, and then run that, in theory, in theory, this should start pulling up other agents. And because, as I said, trying to orchestrate Ralph Loops, orchestrate agents myself to try and organize all of the dependencies and complexities is actually really, really difficult to do. But what you can do is just give the job to Claude to do, and it does a much better job of managing that for it. So as you can see, it's already decided to print out the entire thing, the file name and the ticket for each. And it's got a whole bunch there. So it's actually decided that they're all sequential. So therefore, it should run none of them in parallel, which is kind of interesting. And then it should, in theory, start an implementation agent. Let's see if it's going to... I think it's just running as a sub-agent. Never mind. I'm not sure that's going to work. If you can get it to do it, then let me know. But basically, it's an experimental feature that only came out a few weeks ago that allow it to start sub-agents in Claude as well in order to do things. Any other questions while people are working through that? Yeah? You said you built a raft with the NA10 automation you had, right? Mm-hmm. So what was the feedback with criteria in that? So, what decides if it's a good newspaper article? That's great. I wasn't going to say if a website comes as engineering framework, but what good looks like? If that's dynamic, if that's static, do you put that in the Claude MD file? Yeah, great question. How does that work? So the question is, just to repeat the first half of that, is when I used the NA10 workflow in order to build a route for a newsletter creator, how did the agent know what good was? How did you define good as well? How did I define good? Okay, great question. So in terms of newsletters, I had already been writing my newsletter manually, so I knew roughly what I wanted it to read like and sound like. I also did a bunch of research using a research skill, which was something I built, which is something like great newsletters. And I also did things like that. I also said things like, this is a fantastic written news... In fact, I'm just going to... This is not what I do. I actually do. This is a fantastic newsletter that I've written or that I've read somewhere. Could you please figure out why this is so good? And what are the kind of editorial principles that went into this newsletter for it to work really, really well? And then I would just paste that into... Paste the newsletter in, get it to figure out what was good about the newsletter, and then I would check it, and then I would say yes. There still is an element of human taste here. You can't entirely get away with that. Having said that, I do also have a simulate audience skill, which basically uses a whole bunch of different personas for different clients or prospective clients that I work with. And then I would run the finished newsletter through that and say, run all of these in parallel. And then once you have finished that, figure out ways that I can improve this newsletter or newsletter skill in order to do that. So there's a number of different ways you can do that. The audience simulation is super experimental, but it's actually really effective and often will surface insights I just hadn't thought of. My personality, as I'm a bit slightly all over the place, slightly the way that I talk and communicate, often my clients are not like that. So I tend to barrage people with information, and sometimes my skill will say, okay, there's a lot of ideas in this, Chris. You just need to focus on one main point that makes sense. And I'm like, that's so helpful. So yes, what's quite helpful and interesting is that you can use AI to give feedback on AI like that. The great thing about this particular project that we're writing, this little Pomodoro thing that people are writing, is that it's a command line tool and it's really simple to know whether it works. So it's perfect for a RALF loop. And in fact, these little tools that we build for ourselves, for example, the newsletter is a skill perfect for this kind of loop. I will often say, I want to improve this skill. Could you please back and forth and write the content, then use another agent to read the content, decide if it's any good, come up with things to improve, then send that back in and just run that as a loop. There's a really cool skill. I wasn't going to tell you about this until the end, but I'll tell you now. There's a really cool feature inside Claude code called loop, where instead of creating, in fact, I'll start this in a new session, instead of just doing this thing where you have to create your own while loop, you can say loop every minute, build the next ticket from doc tickets, basically. And then what will happen is the loop will set up a kind of almost like a repeat timer. And as you see, it's got a cron create tool, which for the uninitiated just means do something every minute. This is what those five stars mean. And what it's going to do is it will literally just build the next ticket. When it finishes, it will then check the cron again, build the next ticket. When it finishes, it will check the cron again and keep going. So that's great for working through a bunch of tickets, but it isn't just applied to a set of things that you've got from before. If you think about it, I'll just leave that running up there. You could have a loop that does something like this. So I'll just loop every one hour, check linear for new bug reports from test. And then... I just leave that running. Oh yeah. Can't spell. I'm just going to... Just leave that running. And you're going to get... You're going to annoy your testing team. But anyway, the point is that you can run these kinds of loops in order to get work done in a quite an interesting, I guess, dynamic way, even though it's quite a simple loop. Just find the next thing, do the next thing. If you think about it, a lot of our work is just loops. If we're software developers, what do we do? We look at the backlog. We pick the top thing from the backlog. We pull it over to in progress. We assign it to ourselves. We check on the architecture. We figure out whether there's other contexts we need. I'm just going to leave that running. And you're going to get... You're going to annoy your testing team. But anyway, the point is that you can run these kinds of loops in order to get work done in a quite an interesting, dynamic way, even though it's quite a simple loop. Just find the next thing, do the next thing. If you think about it, a lot of our work is just loops. If we're software developers, what do we do? We look at the backlog. We pick the top thing from the backlog. We pull it over to In Progress. We assign it to ourselves. We check on the architecture. We figure out whether there's other contexts we need. We look at the change. We make the change. We submit a PR. We wait for reviews. We comment on the reviews. We reject the reviews. We implement the changes occasionally. We submit the PR. We merge the PR. We then go through the release process. Then we start again. Pick up the next ticket. And so on. That is a loop. It's quite a complicated one, as we talked about just a minute ago. But it is still a loop. It is possible to get an AI to run that entire loop. There's no reason not to. And that's effectively what's happening here. When you can set up... In fact, you would never actually write this. You would much more likely write something like this, where you'd say, every one hour linear bug finding. And you'd have a skill that encoded all of those chunks of information that I just gave it in a way that would work for you and your particular team. Does anyone not know what skills are? Before I go any further? No. I think almost everyone knows what skills are. If you haven't figured out what a skill is yet, then this is your homework. Go and understand how skills work. They are the best way that we have at the moment of packaging up useful little parcels of context and scripts and moving them to different places or creating different things. So for example, I have about 50 of them that I've written. And they just do lots of different things. The great thing about skills is that you can pull them into your context whenever you need them. So for example, I could say, do you know how to create images using nano banana? And I can ask the AI the question. And the answer is, I could look this up, but it actually knows that I have an images skill for this. But if you hadn't got one, it wouldn't know. But if I then do images and say, how do you create images? Give me the step-by-step. Then what it's going to do is pull in that images skill. And then it tells me exactly how it does it. And I've actually written, in fact, I will make that bigger so you can see, I've actually written a script within that skill that actually does the generation for me. So it's codified the process of doing that. And it just picks whichever model it wants to, and it gives it content. And I have these specific templates that I use in order to create specific nano banana skills. Nano banana is brilliant. This is how I created the presentation that you're looking at. I have a slide skill and an images skill that work in tandem in order to create these presentations. Cool. So let's see what the other thing has done. As you can see, it's already on ticket six. The great thing about Ralph Loops is you just keep working, keep talking about something else. And it's done a whole ton of stuff here. And it's just stopped at this point. But in a second, hopefully, if we just wait, it will start the whole process again. There we go. It's got the scheduled task to run and it's going again. So you can just leave Claude code sessions running with these kind of loops in them. They last about three days, so you do have to keep refreshing them. But you can just do that and keep it running even before you get to a more complicated writer script that wraps Claude to do a thing and all of those kinds of things. Any other questions? Anyone got anything interesting or surprising out of there, Ralph Loop? Has anyone tried this on their real work yet? This would be the interesting thing. Yeah. What was your experience? Have you still got the mic? Yeah, yeah. I just made a screenshot Ralph Loop for a website context engineering framework. So Claude just takes the screenshots and then looks at the layout because it has problems with the geometric spacing. It works well. Nice. Cool. So you're actually using Claude screenshotting to get feedback. Yeah. That's pretty advanced. Not many people are doing that. People are trying to use Playwright and things like that as well to take screenshots and the Claude in Chrome plugin that comes with Claude as well. You can use that in order to get it to drive Chrome and then take screenshots of what's going on. I've had mixed success with that because it's quite a complex thing for it to manage. But for basic screenshots, it works really well. For my images and content that I write, when it runs those images skills, it will always look at the images first to see whether there's any weird AI garbled text or whatever. And it will reject them without even showing me if there's a problem with an image. Was there another question or comment? Yeah, there's a question just back here. Can you just pass the mic? Is that OK? Thank you so much. Yeah, I think it's close to a question that has been already asked because I'm not quite familiar with rough loops. If I ask the agent to implement task one that has already been implemented, would it actually check the quality of what was implemented or only check if it was done or not? Great question. Yeah, it very much depends on what you set it up for. So there's no magic to a Ralph loop. It's just a loop. So this loop that I'm running at the moment, in fact, I probably should just say loop stop. Otherwise, it's going to keep going and use my quota. I think you can just stop like that. I've got quite a fully featured Pomodoro set up now. Come on. Time to stop. So it depends on... It entirely depends on what you write. So if you go through to... What was the loop that I set up? I think it was this one. I just said build the next ticket. That's very ambiguous and not very helpful. So it might not actually finish it. It may just decide to build it and not ship it. It may not actually be very helpful. So what's more interesting is if you go to... It's probably the easiest thing to do. If I load my Ralph skill... This is my actual skill that I use for Ralph loops. And what I'm doing... Actually, this is slightly out of date. But the one that I've got here is actually using a doc changes folder. You can see that there. But I'm using doc tickets in this example. But I've changed it on the latest one. But ultimately, you don't have to use a ticketing system like a flat file in the GitHub repository. So it might not actually finish it. It may just decide to build it and not ship it. It may not actually be very helpful. So what's more interesting is if you go to... It's probably the easiest thing to do. If I load my Ralph skill... This is my actual skill that I use for Ralph loops. And what I'm doing... Actually, this is slightly out of date. But the one that I've got here is actually using a doc changes folder. You can see that there. But I'm using a doc tickets in this example. But I've changed it on the latest one. But ultimately, you don't have to use a ticketing system like a flat file in the GitHub repository. You could use beads, which is Steve Yaghi's version of this kind of approach, which is quite cool. I've used it. You could use linear. You could use JIRA. You could... As long as you can get access to it from the AI, you can use whatever ticketing system you want. For this... For the purposes of this exercise that you're working through, I tend just to use flat files because they just work. You don't really need anything sophisticated. In the same way, the Ralph loop is entirely what you make it in terms of its effectiveness. So, for example, in this particular one, I've given it a proper kind of role in the sense that you are one engineer in a relay team, do exactly one change, then drop the context and start again. That's the idea. So, for this one, it's designed to be run in a shell script where it has an entirely fresh context each time because I didn't want the context to pollute each time. These days, I care much less about that because context is so much larger than it used to be. But when I wrote this, that was very important. As you can see, it's specifically for code that doesn't need human review before shipping. And then it tells you about when work should go in there, what the right time for the tool is, read the claw.md, change the format. This is the format of a ticket. This is all of the rationale. These are the different status values. I ask it to check git state to make sure that it hasn't got a working directory. It's also got recovery states. So, if it crashed, it knows that if there's a dirty working tree, but the tests are passing, you're probably done, but you might not be. So, just double check. If the tests are failing, then it's probably just mid-flight, but broken. So, you should probably just throw it away or just treat it differently. So, you can imagine that this was built up over time of trying to get this working, trying to understand what the user wants, make sure test passing is not enough, verify the actual behavior works, run things in parallel, mark it done, and more. There's an awful lot going on. So, with a real Ralph loop, you want to be building up over time for your specific project, exactly how to check something is working, exactly how to run the test in your particular framework and dialect, how you submit things to the test team, how you want to comment on particular changes, and what style you want to use, whether you want to pull off the thing that feels most obvious to you or the thing that's highest priority, or a mixture of both, depending on how you're feeling that day. Whatever you want needs to be coded in it. [SPEAKER_12] So, when you're writing Ralph loop, I'll show you a link to grab this one at the end, but there's no need to use just mine or something else. Just start with mine and then say, fix this for my project and allow it to change and morph and evolve. A couple of questions, so just come forward. Thank you. Hi. Thanks for the talk. Could you expand a bit more on the topic of sandboxing? Because that would be the thing stopping me from running this work. [SPEAKER_12] Yeah, it makes sense. Yeah, absolutely. So, there's a number of different ways to sandbox this. For this particular small project, I'm not doing that. Most of my work happens on a VPS, which is away from my main machine. It has a few keys on it that are specific to what I want it to do. And it can access developer tools. A lot of them, it can only access them read-only. It can also access my email. But again, it has quite strict, fine-grained clawed permissions for not sending emails, because that's quite important. I don't let it ever send an email. I only ever let it draft them. So, I use a combination of positioning the code physically, but away from the machine, on a different machine, on a VPS. I use clawed permissions for that as well. The permission system is a bit broken, but it mostly works. I'm trying to build lockbox to make it even better. What else do I do? So, the keys that I use are separate keys. So, the AI has access to its own keys, which I don't use for my other stuff. So, I can see the kind of audit trail of what it's done. So, there's a number of different ways of doing it. If you want to just run things simply on your own machine, there's Docker Sandbox, which is quite cool. It's a new feature in Docker, which just allows you to do Docker Sandbox clawed, and a run clawed within that sandbox. So, you can kind of isolate it within a specific container. That's quite powerful because it allows you to only change things within that specific place of the file system. The challenge with that is that it can still leak data from one of your systems to another of your systems. There's a thing called the lethal trifecta. I don't know if you've heard of that. Simon Willison coined it. It's an idea that if you have untrusted tokens, internet access, and access to secret important data you don't want to lose, you're going to lose that data, basically. That's the bottom line of it. So, you have to kind of minimize the amount of times that those things collide in the same context. So, yeah, lots to say about security and sandboxing specifically. I tend to run... I don't run with dangerously skip permissions, but I do run with a number of things turned on by default, but not everything. And you kind of have to go through and figure out what your risk profile is and how much you care about those things. [SPEAKER_12] And certainly, as you're giving... The main things to read up if you're interested is to read up about the lethal trifecta, if you weren't already aware of it, and kind of be thoughtful about how much power and permission you're giving to your agents, especially if you're using something like OpenClaw, which is unfortunately insecure by default. I know that they've been doing a huge amount of work on OpenClaw to make it more secure, but it is still a challenge for those kinds of agents. They do have access to a lot of things. Any other questions? Yeah. Yeah, you had a validation step in the loop. Mm-hmm. This might be anecdotal evidence, but as soon as I changed mine to use sub-agents here now for the validation step, it started finding things. Ah, interesting. Whereas, as long as you're doing the validation in the same step with the same context, it just pats itself on the back. I know that they've been doing a huge amount of work on OpenClaw to make it more secure, but it is still a challenge for those kinds of agents. They do have access to a lot of things. Any other questions? Yeah. Yeah, you had a validation step in the loop. Mm-hmm. This might be anecdotal evidence, but as soon as I changed mine to use sub-agents here now for the validation step, it started finding things. Ah, interesting. Whereas, as long as you're doing the validation in the same step with the same context, it just pats itself on the back. Yeah, yeah, yeah. That's a really good point. There's definitely confirmation bias going on with agents where they're like, oh, yeah, of course I wrote it. Fine. It was fine. I checked it a minute ago. Yeah, using sub-agents is really powerful because a sub-agent starts with only a small chunk of context. It doesn't start with a full context, right? So you can get much more power from it. So as a good example from this particular project, a really useful skill, which I mentioned earlier, is Simplify. Simplify is a clawed coding bundled skill. And what it does is it will look at the most recent changes, and it will run three sub-agents to try and figure out whether your code should improve. So you can see here what it's doing. So hopefully this will run. These will load. And it will probably find a bunch of problems. Yeah, great point. Great presentation. Thank you. Did you try OpenSpec or combine with OpenSpec or any other spec-driven? No. If I'm honest, I'm not a huge fan of spec-driven development. I know that's controversial and I'll qualify that. I worry that spec-driven development is taking us, at the extreme, is taking us back to the bad old days of Waterfall, where we would specify the entire or try and over-specify a project. Even these little set of tickets, I'm not that comfortable with. I feel like spec should be much more iterative than we can see. It's already fixing a bunch of things. That's quite cool. So it found, just to finish off that point, it found a bunch of issues there. And it's got some fixes. Yeah, so specs. I like just-in-time specs. I like the idea of building or thinking through what you're trying to build, creating some kind of plan and claw code and then executing it. That's fine. I'm happy about that. And I think that's a useful step. What I worry about is, A, I worry about things like Kiri where they've codified that into the tool. I worry that that will almost fossilize that one approach with AI that works today but may not work again when Mythos eventually comes out. So I worry that the tools are being too quick to jump to a specific structure of work that may not be the right thing in the future. So I'm cautious about that. I think it is obvious. It's a truism that AI needs more context in order to do well. So we should try and give it more context. But I think the idea of overdoing that and over-speccing a project is one to be careful of, as well as over-structuring our process based on what we know about agents today. Because then we'll end up with working with a new kind of AI-driven process that worked best with agents that came out in 2025 or 2026. You know, we'll still be using that in 2030 and that'll be a pointless waste of time. So those are the kind of concerns I have with it. Any other questions? Yeah, great talk. So you mentioned that you don't like spec-driven and you use RALF. So there is no human in the loop. So the question arises, does Claude actually need you there? Where is your input there? Great question. So I've been thinking about this quite a lot recently and having a bit of an existential crisis. I don't know about anyone else. But yes, what value am I adding here to this thing? Certainly not with writing a Pomodoro timer. I'm not sure I'm adding much value at all. I mean, I literally said I one-shotted those specs and there was no point there at all. I'm not saying that I don't like planning out a system. What I'm interested in at this point, and I don't have the answers, is the fact that AI often will pick better specs and write better specs that I can write and will often have a better idea of the kinds of direction my software should go in than I necessarily will have. So I like the idea of actually having RALF loops that create other RALF loops, potentially, or having RALF loops that track whole customer engagements or even whole startups. So I have a skill that I'm working on. Should I show this? I'm going to show it. It'll be fine. What could go wrong? Which is called startup. It's pretty ambitious. But the idea is that it should basically guide a product through an entire startup framework. So it is meant to be run as a loop. The idea is that it... Oh, I see you all taking pictures. Now I'm owning this thing. Damn it. But with great thanks to Ash Morrow, who wrote some brilliant stuff on this. I should say that for the tape. So really, really helpful to me. So I built this out of basically all of the cool books I've read about startups. So I'm a startup founder, co-founder, CTO. So this is near and dear to my heart. And what I'm trying to do here is I'm trying to give the AI enough context such that it could run my startup for me and potentially figure out what the next most important thing to work on is and then do that in a loop. And then there's a big outer loop that runs that says, OK, well, what's the next most important thing to do? Let's do that. So it doesn't work. But it's interesting and it's getting somewhere. And it will often... The first thing it does... I don't think I've got it to show. But... Oh, yeah. No, I will show it because it is hilarious. Hang on just a second. There's a... I asked it how it was doing on one of its loops. And it produced a startup update deck as an investor memo, which was... I didn't even ask it to do this. I'll show you the demo. Hang on a second. Because it's absolutely brilliant. Let's see if I can just show this window. There we go. Air skills. Startup update. And so, yeah, it said, yeah, I need to give him an update. So what I did is it said, basically, this is how far we've got. These are the problems nobody has solved. This is what we know that's real. These are the number of... [SPEAKER_09] This is a skills management tool that I'm working on in the background. These are all the kind of issues. And it came up with all of this cool stuff that could go into an investor deck. [SPEAKER_09] To be honest, it's not bad. It's not a bad... I think it's actually the GitHub for AI skills. But there we go. [SPEAKER_09] And, you know, who's going to pay for this thing? How much will they pay? Those numbers are definitely not right. There we go. Air skills. Startup update. So, yeah, I need to give him an update. So what I did is it said, this is how far we've got. These are the problems nobody has solved. This is what we know that's real. These are the number of... [SPEAKER_09] This is a skills management tool that I'm working on in the background. These are all the issues. And it came up with all of this cool stuff that could go into an investor deck. [SPEAKER_09] To be honest, it's not bad. I think it's actually the GitHub for AI skills. But there we go. [SPEAKER_09] Who's going to pay for this thing? How much will they pay? Those numbers are definitely not right. But what's interesting is that it decided that it wanted to do this and figure out all of these numbers based on this, which I think was hilarious. And it was quite proud of this deck, to be honest. And I have to be like, hang on a minute. We haven't... [SPEAKER_09] There's some serious thinking you need to do before you go to that. [SPEAKER_09] Will orgs pay for skills government? Would your orgs pay for skills governance? Great question. Not sure yet. So the reason for showing that is to point out that AI can do a lot. And it doesn't do startups well yet. But that's probably down to my skill file, not down to the agent itself. I have a feeling that there are an awful lot of things that potentially will be loops in the future. I only got that far on my slides. Oh, my gosh. Hang on a second. So we've done that. We've done that. If you are still working on this demo, I've got a couple of challenges for you if you'd like to do this. One is you could try upgrading your ticket format. If you like the raw markdown file, the doc tickets is fine. If you wanted to type BD install or install beads, it's super easy to do that. And you wanted to get Ralph Loop to work with your beads, try that out. See if that works. There's no pressure on you to achieve anything in this little folder. Beads is great because it only works within your folder. And it just installs a little tool. So it's quite a useful thing to try it on. So if you wanted to try a different ticket format, or you wanted to move this into your main project and connect your Ralph Loop to your ticketing system to see how that feels. Maybe not submit tickets yet, but you potentially could try that and see where that takes you. So that's an option. The other is the skill. You're going to need to keep upgrading and working on your Loop. The Loop basically contains all of the know-how about how you as a person will go through that. You can take it all the way through from do the next ticket, and you can take it all the way up to do the next step in the world dominating startup you're trying to build or whatever it is. It works for all of those things. And what's super interesting about this is that I'm more and more convinced that everything in fact is a Loop. Maybe as an engineer, I'm definitely on a Loop on a lot of the work that I do. Maybe as a project manager, I'm on a Loop. Maybe as a CEO, I'm on a Loop. Maybe a lot of the cadences that I work on run in Loops too. So I have a skill. And if you're running an open claw bot, you're doing a similar thing that just runs a heartbeat every 15 minutes. It just on my VPS, it just fires up clawed, checks a few things, checks my calendar, see if I've got anything happening, and sends me telegram messages. Maybe that's on a Loop. That is definitely on a Loop. It's 15 minutes. I have a Worker Loop, which I'll show you in a minute. And I have a Morning Loop where every morning at 6am, it comes up with a full briefing of my day, figures out exactly what I should be doing, and just gives me all the information that I need that's happened overnight. All the emails that come in, all of that stuff. The Worker Loop is particularly interesting because it basically... I'm not sure I can actually show this. Let's see if I can find something that I can show. No. The reason I can't is because it's got a bunch of client information in it, so I can't show you that. But what I can show you is, for example, this screen I can show you. So if I quickly switch to this... So this is an app. I'll make that slightly bigger. This is now how I run my Worker Loop. So this is an app that I wrote to manage projects. So I don't have tickets inside my work vault. I have project files. And each of the projects is a set of work that I need to do. And then every so often, I basically vibe coded a Kanban system. And a worker will pick up and do the next step on the project. So if the next step on the project is writing an email, because it has an overview, or it's checking things, or it's producing the slides for my project, it will do the next step. So this, for example, is the workshop prep spec project that it's working on. And it's got a bunch of front matter that is just looking like that. And it's got some questions for me. I haven't updated this. It needs to be updated. It's got the context. It's got a decision trail of things that it's done and why it's done them. And so for every different thing that's happening, it just is figuring out the next one. It's also got notes on other talks that I might be giving that didn't happen in the end. And it's got feedback from a previous workshop I did on a similar topic, which you can click on. Actually, I can't click on that. There's a bug. But it will basically show the notes from a feedback session. So this project pulls everything together from all of the contexts you can find and then does the next step in a loop. So you can run everything in a loop. You can run all of your work in a loop. When I wake up in the morning, normally, I have about 15 or 16 draft emails where people have got back to me and it's had a go at replying to them. I always have to edit them. They're always okay. But it definitely has a go at getting on with trying to schedule some of my work. I have very specific rules about what it can and can't do. My basic rule is, is this reversible without embarrassment to me? And if the answer is no, don't do it. But just make a little note in the project and hand it back to me. So sending emails is not allowed to do. Creating a slide deck, for example, this one, that's reversible. It doesn't cause me any embarrassment. So it just got on and did it and gave it to me. It doesn't post on LinkedIn for me. It doesn't send emails. It doesn't send messages. But it does get everything ready for me to review. To your point earlier, definitely has a go at getting on with trying to schedule some of my work. I have very specific rules about what it can and can't do. My basic rule is: is this reversible without embarrassment to me? And if the answer is no, don't do it. But just make a little note in the project and hand it back to me. So sending emails is not allowed to do. Creating a slide deck, for example, this one, that's reversible. It doesn't cause me any embarrassment. So it just got on and did it and gave it to me. It doesn't post on LinkedIn for me. It doesn't send emails. It doesn't send messages. But it does get everything ready for me to review. To your point earlier, which is a very long answer to your question, it has caused me to genuinely question what I'm good at and what I'm here for. Quite a lot of the time, I've got to a point where I'm just the email person who just checks emails and sends. Check emails, send. Check. That doesn't sound like a proper job. That doesn't feel good. So therefore, what does that mean for my work? And I've had to make a conscious decision. Which bits of my work do I want to do and which bits of my work don't I want to do? I don't want to be the email reviewer, but I do want to be the strategist. I do want to be helping organisations think through what on earth is going on with AI and how to fix it for their organisation. Now, I could get AI to do a bad first draft, but I don't want to be reviewing AI's draft. I actually want to be doing that thinking myself. So therefore, I basically said, don't do any of that work. I want to do that work. Just give me all the information I need and I'll do the work because I enjoy that work and I'm good at it. So AI can do all of the rubbish work, but it can't and it shouldn't do the work that I'm uniquely good at. But because everything is a loop and Ralph, this is getting so existential, because Ralph loops are really everything and can be used for everything. We have to start asking hard questions about which bits of work we actually want to do. What do we want to do out of this work? It's not just about what AI can do or can't do anymore. Yes, there's a question. There are loads of questions, but let's go at the back. I think your hand was up first. I think the chap's coming with the mic. Well, just for the recording, it's really helpful. Thank you. So with the open-ended tasks, how would you think about when to stop? Do you set KPIs up at the beginning or how do you know when it's done? Yeah, great question. I ask it to... So again, this comes down to if I just go back to... Sorry, different window. This one, it comes down to upgrading your loop and you have to basically tell it when to stop. So I don't just have one Ralph loop file that works for all of those different loops that I showed you earlier. I have different ones for each. And for example, the worker says, when you get to a point where you're either running out of context or you've got to a point where there's an irreversible thing to do, then I want you to stop and report. And what report means is, in this case, update the project file, which is just a file in a repository, with where you've got to and in a way that where you present it to me for review. So I'm working quite hard at the moment about that kind of presentation step, because I definitely don't want to be reviewing a diff. And just reading text, I find really difficult, just because there's often a huge wall of it and it's hard to parse. So I'm getting it to start giving me things step by step in slide format. I'm trying to get it to... That interface I showed you before, you can see a little bit how I'm trying to get it to show things in different places, in different ways, for it to get to what I need to uniquely do next, I suppose. So to answer your question, it depends on what you're doing. I think the most important bit is that you figure out what the edges are for yourself and have that real moment of, what do I actually want to be doing? How do I want to be involved in this work here? Not just AI is helping me do my work and a companion to me. It's much more now, which bits do I not even need to know about? Next question. There's one here. There's a mic just at the front somewhere, I think. Yeah, great. Thank you. Hi. So since you brought up this topic of our involvement, right? So at what stage do we really get involved? I mean, you mentioned you don't really review the diff. I guess the most important part of our work now is creating these tickets, right? I mean, first identifying what the most useful feature to implement is and then describing in a way that you foresee all the different edge cases or just explain it the best way possible so that the outcome is what you desired in the first place. Yeah. What's your process of creating these tickets? Yeah, great question. And so my concern is sometimes you don't really know yourself until you start implementing them in the way we used to do it, right? So during the development process, you encounter certain different cases where you need custom logic, you need to... It's just difficult to foresee these things from the get-go. And do you iteratively improve the tickets and you re-implement them or what's your process? Yeah, great question. So in terms of how I work, there are two modes of work that is done on my behalf. One is the fully automatic work that we're talking about here where the AI will just get something done. I don't need... When I've got a decent spec that I trust and there's a way of feeding back so that the AI knows that it's good, I don't need to be involved in that work. That can just happen. For every other piece of work I do, I have a... I work on it with and in Claude code. So I... With that system I showed you earlier... In fact, I'll go back to it so I can show you. Let me just change back to this window. Where is it? There it is. With this one, at the very bottom of any of these projects that it's running for me, there's a little thing here which I... This is just a Vscode app that I've written for me. Nobody else has access to this. It has a little VCP command which if I take that and I type this into a terminal window. So if I go back to this one for example, I think this is okay. Yeah, I can't easily show that just because my internet is not going to be able to connect to my VPS. But the point is, I'm able to go back to... Yeah, I'll go back to that. The point is, I'm able to type that in and just paste that into my VPS. What that does is, it starts a new Claude code session. That session grabs, knows where to find that project and it has access to this. It has a little VCP command which if I take that and I type this [SPEAKER_11] into a terminal window. So if I go back to this one for example, I think this is okay. Yeah, I can't easily show that just because my internet is not going to be able to connect to my VPS. But the point is that I'm able to go back to... Yeah, I'll go back to that. The point is that I'm able to type that in and just paste that into my VPS. What that does is it starts a new Claude code session. That session grabs, knows where to find that project and it pulls in all of the project context. So rather like loading a skill, it loads the project. It reads the entire project file and knows where everything else is. At that point, it's loaded in everything it needs in order to supercharge that session with me. And then we work together on it. So if I'm, for example, to your point about speccing tickets, I'd have a ticket, probably a project for that particular feature. And then I would say, okay, loading everything you know about that. And it would load them all in. And then we would work back and forth on speccing out those tickets. And then the output would be whatever I needed to get done in order to get that done. So if I'm working on something where I want to usefully and uniquely do that myself, that's when I would jump into a project with Claude code. And when I say do it myself, I don't mean typing it myself or I don't mean doing all the thinking. I normally mean I get Claude to interview me to ask questions so that I can give it the information it needs to formulate, to do the writing, because I don't like doing the typing. But I get it to pull the information out of me in order to get that work done. So those are the two modes. It's the back and forth iterating and then it's just the automatic stuff. And I should also point out that I don't like reading diffs. But ultimately, that's the only way that you can review code. When I'm reading a newsletter item, I don't want to read the diff, I want to read the newsletter. Whereas if I'm reviewing code that's important, that is for other people, then yes, I read the diffs. I don't like doing it. Nobody likes reading diffs. But I check to make sure it's working. And I will, I can't see myself not doing that for a while, especially not with security, like conscious code. Maybe with Mythos, I'll just delegate it. That'd be nice. Any other questions? There's one here. How do you deal with context rot? So for example, your example where you have a loop and it takes one task after the other, is it the same Claude code session that takes all those tasks? You'll have to experiment with that. With the slash loop command, yes, it is. It's the same session. When you run it as a kind of while loop outside Claude, then it's a different session. You have different trade-offs with that. With the same session, you have all the context of the previous tickets and the previous changes. That might be useful. In practice, I've not found that so useful because it can just pull the files as it goes. If you're not typing anything into the session, you're not really adding anything to that. So there's nothing really in there that's useful. So I've tended in the past to prefer starting a fresh context for each new session. But the loop is very, the slash loop is very easy to run and it just works. And especially Opus is very good at long context retrieval. So it's less, much less of an issue. Okay. Sorry. Yeah, there's one at the back as well. Is there a microphone as well? Okay, great. We'll come back to that. Are you reviewing sessions that are done by your loops or are you just reviewing diffs on the GitHub? Great question. Because I don't allow any of my workers to close a project. So I would always say, if you think you're done, tell me what's finished and I will close that off. So it could be that there's a big list of completed things that I need to check off myself. But I want to be that kind of final step of verification. The reason that I've added that is because I worry that I'll miss something. There's a thing that someone coined recently called cognitive debt, which is the idea of just not being up to speed with everything that your code base can do or all of the code in your code base. And that worries me. So I tend to want to at least understand how the code fits together and how the piece of work that I'm working on fits together. So I don't let AI get away with just putting something out of my sight without me having a chance to look at it. Otherwise, I feel like I'd lose track of what's happening. Yeah, because I mean, for example, I'm using sessions to track tickets. So instead of reviewing the code or this in the code, I'm just reviewing what the particular session was doing. And I even have a marking system with which session is on which status. Is there any way how you do it similarly? Similar? Yeah, I think the sessions and the status, I think that can work. I haven't tended to use sessions like that. What I've tended to do with sessions is I get Claude to every night go through all of the previous sessions that I've run that day across all the machines I run Claude. It saves them all into a JSON file for me. And then I get it to both figure out how my system could improve and also just what I did so that I haven't, I don't forget what happened. So it writes a little paragraph for how much I did. And I use that in order to track work. But it's not quite the same as one ticket per session. I quite like the idea of having one context per session. I think that's quite a nice idea. And one sorry, one like by per unit of work. I just haven't made that work. I found it really useful because then I can go back to the particular session when the particular thinking was happening. Yes, and I do do that for sometimes when I've got a project that's running over multiple sessions, I can go back to the previous session. Instead of the VCP command I showed you, I could type VCR and it'll do the same thing. But in practice though, I like the discipline of it having to pick up again. Because it means if it has to pick up again from a fresh context, it means that all of the information that was in that session has actually been codified into other places that any Claude code session or human could find. Which means that you end up with a much more, I guess, richer kind of repository of knowledge that you're working in. So there's a question mark around if sessions are truly not ephemeral and you've got them as a store, are they accessible as future context? If you treat them as ephemeral and make sure you capture everything within them into your repository anyway or into documentation files or whatever, I think that could be more powerful. So worth thinking through for sure. has to pick up again from a fresh context, it means that all of the information that was in that session has actually been codified into other places that any Claude code session or human could find. Which means that you end up with a much more richer kind of repository of knowledge that you're working in. So there's a question mark around if sessions are truly not ephemeral and you've got them as a store, are they accessible as future context? If you treat them as ephemeral and make sure you capture everything within them into your repository anyway or into documentation files or whatever, I think that could be more powerful. So worth thinking through for sure. Any other questions? Feels like we've come a long way from just write this ticket, but there we go. It's good. It's all good. Yeah. Thank you so much for the talk. I have a question. It seems like in the loop, some of the stats might not be necessary. Like you might go to the code and then find nothing there. Yeah. And would you consider optimizing it somehow or you just let the token burn? [SPEAKER_11] No, just burn the tokens. They're not that expensive. Depends what you're doing. I think we're at the point I should. This is a whole other thing. We're basically in the era of free tokens right now. I have a max 20 subscription. And I definitely use more than the average person probably who is paying for one of those. So I think at this point I would optimize for freeing your own time up as opposed to optimizing for burning a few more tokens. I don't think tokens will ever get that expensive. I think that the frontier models potentially will be very expensive, but we have really good, cheaper or freer alternatives just around the corner. Not quite as good for the latest kind of work that we're trying to do, but they're really, really good. So, you know, I think there was at least one that just kept the GLM one that just came out. It looks really promising. That's the ZAI one, I think. It just came out this week. [SPEAKER_00] Really, really interesting. I'm still running Claude, but that won't necessarily always be the case. So I think I would just burn them. I would, at the very beginning, you know, I spent a long time doing the whole optimization thing where I was doing this, if you weren't here at the beginning, this thing. You know, I spent a lot of time trying to screw around with all of this, but ultimately I just now let it run. It's much simpler. I do get quite close to the end of my max subscription sometimes though. I'm slightly nervous about what that means. I have to figure out how to get another account. Yeah, max to the $200 a month one. Yeah. Yeah. Yeah. I get pretty close to that every week. I'm quite about 80% now getting the jitters. Yeah. You had a question. Do you want to bring the mic back down? Is that okay? Thank you. I'm not looking at that anymore. Hello. Is this, I wanted to ask you about, thank you so much for the presentation. Yeah, sure. But fine tuning for the prompts, do you version it? Do you have datasets that you use to fine tune your entire loop? So in terms of versioning the Ralph loop specifically, like the prompt for the Ralph loop. Yeah. Yeah. So I use skills for that. So as I pointed out before, everything like that goes into the skill and I get Claude to write the skill for me. And that saves in either your .claud skills folder within your project, or it goes into your home directory under .claud skills. I use GitHub to version all of those for myself. I don't think Git is the right skills format for this long term. I think we need a new thing. Hence trying to build S skills, in fact, which is this idea of trying to make skills much more portable and shareable within teams, which I'm trying to figure out. So yes, I do version control them and I treat them as quite important code. And I don't actually, I do share some of them, but I don't share all of them routinely because they have a lot of my own IP in there. And actually a lot of my customers IP is in there too. The question is more with regards to the performance of the prompts. So you were saying that in the beginning you improve the prompts as you go along and are you versioning that going along and are you versioning the performance of the prompt overall? So when you say prompt, do you mean the skill itself that I'm using? Yes. Yes. Yes. Yes. Yes. So yes. So the skill, the prompt lives within the skill. So when I type slash bug tracking or slash Ralph, that is the prompt that gets written by Claude and managed by Claude, which means that the whole file is the prompt and therefore that is version controlled. So I always have a git running within that setup and then every time I change it, I update. [SPEAKER_11] But you remain subjective. How much you make sure you have improved. Let's say that you have, your data set will be an issue. I say that. And the expected output would be the new feature added to the wrapper. Yeah. So you could more, how do I evaluate whether it's any good or how? Yeah, exactly. How do you know you're actually improving? I see. So how do you know if you're improving? That's a really good question. I do stress test my skills. So with other skills, and I say, you know, is this skill any good? Could you improve it? Could you write it? I do spend quite a lot of my time tinkering with my system and my skills probably more than I should. I think it's a bit subjective at the moment. What I haven't done, and this would be a really good exercise, is to try running blind testing where you would run a set of tickets with one skill and a set of tickets with another. Ultimately, because Claude is non-deterministic anyway, I think there's a high level of variability with any of those kinds of tests. So it's really difficult to think about how to construct a useful test in that way, to know whether you're actually improving or not. In general, the more context you give into your prompt, the better it will do, up until a point which isn't very easy and obvious to figure out where it becomes worse. So it's about balancing that ultimately, but yeah, I haven't done a kind of objective improvement process. A great question though. It's a question just behind you. How are you version controlling the skills? I'm using GitHub at the moment. I do have a product that I'm trying to build. to construct a useful test in that way, to know whether you're actually improving or not. In general, the more context you give into your prompt, the better it will do, up until a point which isn't very easy and obvious to figure out where it becomes worse. So it's about balancing that ultimately, but yeah, I haven't done an objective improvement process. A great question though. It's a question just behind you. How are you version controlling the skills? I'm using GitHub at the moment. I do have a product that I'm trying to build, which is this thing here. So if you want my skill, by the way, that's how you get it. It's a project called AirSkills, which you saw a brief preview of earlier from my slide deck, that my agent put together for me. But the idea is that you can package and manage those skills as a unit. So you can create skills for your organization, you can create skill bundles, you can create a skill set for your org that works for different teams within your org. And then that all gets versioned and updated for you without everyone having to learn how to use Git and GitHub. That's the idea. It is a real pain at the moment. I found it really, really difficult to manage just for myself, even just putting a skill on GitHub. I can't imagine anyone from—there's quite a lot of friction for a coder like me. I can't imagine non-coders using that. So yeah, trying to build this. So yeah, run that command on your machine. You'll have my skill. Sorry, there's a question just here first and then go next. Yeah. How do you touch on this a little bit, but around the edges, how do you do knowledge management? So I guess I use Claude for "Why is my VPN not working?" And then I learned something and I want to record that. And then I'm like, I'm meeting with somebody with the transcription and I have that somewhere else. And I've got a bit of code that I'm writing and all, I've got all of these different contexts, but they're sort of very disorganized. Do you have a way of thinking about how you organize all of that? [SPEAKER_16] Yeah. So I have a code directory and I have a vault directory and those are the two directories I work in. So the code directory contains a few different projects that I work in a more classic way. The vault directory is where I do all of my other work. And frankly, I mostly start working in there, even if I'm working on code and just tell it where the code is, because the vault contains several thousand files with all of the different stuff that I have picked up, learned, worked on with Claude over the last several years. Well, not with Claude for that long, but you know what I mean? I started with Obsidian a long time ago and I've been working on that vault for a long time. And with Claude now, it just works on that for me. So when I do some research on how to fix my VPN or whatever it is, it just saves a file in there. I have some specific rules for how to structure and manage that. If you're interested in more, I haven't written a lot about this, but I know Andrew Kapathi has just written about it using LLMs as a wiki. That's a great article if you haven't seen it already. I know there's a Mina Jojovic actually funnily enough has done a thing on Mem Palace yesterday. That's another version of this. You can use that too. There's lots of different systems for that out there. The best way to get started is markdown files in a file system and use it like a wiki. So you run Obsidian in one window and Claude in the other and just work with it and save things as you go. So do you have an agent that then structures and puts those into folders or something like that? Yeah. So it depends on your method. I use the Zettelkasten approach, which is the one where you have one note per thought. So any thought just goes into a flat folder. Then I have a /projects thing, which has all of the projects that you saw, including one for this presentation, which is my kind of unit of work for an agent that we work on together. It does a lot and then I do some and then it does some. I have transcripts in there. All of the calls that I've ever recorded go in there. And I use a tool called Lian, which is a command line embeddings tool. So it basically just runs embeddings across the entire all of the texts in the repository, all of the transcripts, all of the links I've ever saved, including all of the content. It's huge. And then it can find things usefully and easily in there. So the best time to start that is today because it just takes years to put together. I need to write more about that. Any other questions? Yeah, there's one here. You said you had friction while versioning your skills. I've been using skills only for the last month, so I'm not aware of this friction. Can you explain what the friction is? I can. I'm sure there are—has anyone here had any friction with managing and using skills yet? Has anybody else? Yeah, quite a few different people. So yeah, it's an emerging thing. It's not surprising you haven't experienced it yet if you're not using it for very long. What I've found is that if you're just using them on your own, creating a file of skills and managing them is quite straightforward. Putting them in GitHub is quite straightforward. It's symlinks and GitHub repository. It's fine. Where it becomes difficult is how you share that. So how would you share a skill? Okay, well, if you want to use MCP skills, you have to then put it in its own GitHub repository. That feels quite heavyweight just for one skill. I'd have to have 50 of them in order to share all my skills. So that doesn't really work. Then it's more like, okay, if I don't want to do that, do I just send them the skill file? Do I send them a zip file? I mean, I can't think of a better way of doing it. Do I have to have a submodule in my skills folder for every single Git repository I share a skill with? It just doesn't make any sense. So I think Claude has got some stuff in there around plugin marketplaces where you can have a plugin which has a bunch of skills. That's the best way, but then you're versioning the plugin, not the skills. all my skills. So that doesn't really work. Then it's more like, okay, if I don't want to do that, do I just send them the skill file? Do I send them a zip file? I can't think of a better way of doing it. Do I have to have a sub module in my skills folder for every single Git repository I share a skill with? It just doesn't make any sense. So I think Claude has got some stuff in there around plugin marketplaces where you can have a plugin which has a bunch of skills. That's the best way, but then you're versioning the plugin, not the skills. So that's probably the most seamless way. It just doesn't work that well. Also, there's a challenge around if somebody contributes to your skill, do you want their changes or not? It will depend on what the contribution is. Are they local to just them or are they changes that could be generally incorporated? And that depends on the skill and depends on them. So you have to manage that. Do you run a backlog for each skill where you have tickets to improve the skills? I think these are all unsolved problems. I'm trying to solve some of those. But these are big problems we haven't figured out yet. Other questions? There's one right at the back. If there's a mic, that would be amazing. Thank you. One, two, one, two. Okay, it's working. My question is about how we can scale up this approach with Ralph Loop. But for the actual production team, I don't know, three engineers, how to coordinate, how to cooperate. Do you have any idea how we can organize it? Do you have any experience? That's a big question. How, so just to make sure I've understood it, how do you scale this up so that you can coordinate whole teams using this kind of looping approach? Is that yeah. [SPEAKER_17] With all the tickets and the skills, yeah. 100%, yeah. It's difficult. I think the teams that are where I've seen this work well is where they are proactive about updating tickets. The great thing is if you connect your ticketing system to the AI, it's really good at updating it. So you should definitely do that. Make sure that you claim the ticket and move it into the doing column before it starts work. And make sure that somebody else hasn't just done that before you start. Do you see what I'm saying? That's really important to avoid contention. Those have always been issues with bigger teams. Just in the same way, a couple of controversial things, just in the same way that Ralph loops work really well by just doing one thing in a loop and quite sequentially, it could well be that the coordination overhead in our teams is caused by the fact we've got too many people in our teams and maybe we should have smaller teams and just more of them, right? So maybe if you're trying to get 10 people to coordinate and using AI and Ralph loops and all of that, that's just not going to work. Maybe you need three and maybe that's the way to run that project and then you split it down and then you have the other seven people doing something else or whatever. Does that make sense? So making the problem go away is the first step and making sure that you're already using your coordination mechanisms as the second step. And then just try it and figure out what the bottlenecks are. Be really good at retrospectives with this stuff. I think retrospectives and teams are often pretty anemic. It's like, what should we do less of? What should we do more of? That's just a recipe for the same, or more of the same, and just changing tiny increments, which can be good. But ultimately, this requires a radical rethink. So be really conscious in making sure that your retrospectives are changing actual things about how you actually work or have space to try, let's just try using a route flip on all of our work for a week and see what happens. And if it doesn't work after two days, that's fine. And then if you are someone who's in a leadership capacity and is able to sponsor that kind of work, this is what it means to try and move to AI. If you want to transform your team, you are going to have to sponsor these kinds of experiments and be okay with failure because I'm speaking to the leaders in here for a minute, because it's going to be messy and it's going to fail a lot. But if you want real transformation, that's the only way to get it. You've got to give your team space to try a whole bunch of different things. So give them air cover. Yeah, that's a big and complicated question. I think if you're able to and have the agency to, just try it and see where it gets to. There's a whole separate thing called theory of constraints, which I haven't talked about at all, which is the idea that within any team, in any system, there is always a bottleneck. There's always one bottleneck that's the big bottleneck. If you don't work on that one bottleneck, all of the other work that you might do to optimize and improve the system is pointless and actually probably counterproductive. So this is why some teams when using AI tools and using advanced AI tools like Ralph Loots, which is just AI or what we're doing now, but just on steroids, some teams when they implement it actually go slower. Some teams go amazingly fast. Some teams go slow. Why is that? It's because they're not working on the constraint. The constraint in those teams might be the review process or the release process. If you release your code once a month and you're shipping 200 PRs, not 20 in that release, how do you think that's going to go? It's not going to go well. So that's why teams go slower, because what they need to do is fix their release process, not their coding speed. So always fix the thing that is the biggest bottleneck first, then figure out where the bottleneck moves in the system. And that's not predictable. It's random, so you have to figure that out. Then move and fix the next thing in the system. For more on that, read The Goal by Elio Goldratt from 1984, no less. It's an amazing book. One more. Is there another question down here? Is there another mic? Where's the mic? I have the mic. Oh, you've got the mic. Great. Keep going. Since you were asking about talking about the constraints part, this reminded me like, I'm part of the AI team and we have an AI team and they write a lot of microservices. It's in different repositories. Okay. How do you deal with coding now since, is it like one big monorepo or is it small repos? You have to try it different ways and see. I don't think that the Git architecture, whether it's many repos or one big one really matters. You can One more. Is there another question down here? Is there another mic? Where's the mic? I have the mic. Oh, you've got the mic. Great. Keep going. Since you were asking about talking about the constraints part, this reminded me. I'm part of the AI team and we have an AI team and they write a lot of microservices. It's in different repositories. Okay. How do you deal with coding now since, is it one big monorepo or is it small repos? You have to try it different ways and see. I don't think that the GitHub or sorry, Git architecture, whether it's many repos or one big one really matters. You can always start your AI in a folder above all of your other repos and just get it to work. It does a great job of that. So that's okay. I think the bigger question is what are the coordination patterns within your teams and your services? Who's responsible for what and how does that change with AI? I think that's a more interesting challenge. The main reason I'm asking this was because some of the microservices depend on others and then you have to release one of them. You need to release a tag and then update another. That's just... Yeah. I think what AI will do is it will expose all of the places in which that process is inefficient because it will do everything faster, which means that if you're seeing those bottlenecks where you are getting dependencies between your microservices, guess what? That's your biggest bottleneck. Therefore, you fix it. So how do you fix that bottleneck? Well, you might try atomic release system or you might build something that using a route loop that figures out a way of coordinating releases across multiple repos more successfully. I don't know. But that's what you do. That's where you work. Don't work on anything else until that's fixed, if that's the bottleneck. Yeah, there's a question here. Do you want to pass the mic? Yes. I should say I'm at the end of the content. There's... If you... Which probably was clear half an hour ago. The only other thing... I had a Q&A slide. If you are leaving, you're welcome to leave, but you're welcome to stay for more questions. I would really appreciate some feedback though. So this QR code is the only thing that I manually added to these slides. If you could just fill that in, that would be lovely and amazing. Thank you. It's literally only three minutes, four questions. It just helps me to improve and make sure that I do a good job of these workshops going forward. That's also my LinkedIn. I post a lot of content on there. Do connect with me. Do engage with me in some of this stuff. Disagree with me. I love disagreement. I love it when people say, surely, Chris, that's nuts. You shouldn't be doing that. Love those comments because it really helps me to think and improve, which is what I love to do. And because the Ralph Loops is doing all my other work. So got nothing else to do? Great. Thank you. So I just wanted to put that up there. Very happy to continue answering questions though. But if people wanted to drift away, then that might be a good time. Go for it. My question is regarding multi-agent orchestration tools. I'm curious if you've tried things like Steve Yeggie's Gastown or there's another guy who does MCP agent mail. Yeah, there's some really cool and interesting stuff. I still think we're in the wild west, literally, with Gastown and things like that. But we don't really know how that's going to go. I have tried Gastown. I couldn't really get it to work. But it was pretty early on. For me, I feel like the agent orchestration side of things is I think that we overcomplicate things by assuming that they need to be in parallel. I quite like the idea of just starting with a loop to start with. I don't feel the need to have my AI do lots of things at once before I can get just get it to do one thing well. It goes back to the theory of constraints thing again. I don't think speed of number of tokens per second is the bottleneck. I think that it's our ability to specify what we want and review what the AI has done. If that's my bottleneck, I don't want to introduce more agents. I haven't spent lots of time with those tools for that reason. I feel like they're solving a problem that not many people have yet. So, hello? Not just for speed, but for example, I don't know if you've experimented with MCP agent mail. So agents can lock files and speak to each other so they don't step on each other's toes. And you can use different models like Claude, Opus, and Codex work on the same project. So you get different brains working on the same project. Nice. Yeah, no, I haven't. It's not just for speed. I think I've heard of it, but I haven't tried it. Sounds like a super interesting idea. Rather like someone mentioned earlier about sub-agents trying to look at things from a different perspective. I've had a lot of value with doing that. I mentioned earlier my simulate audience approach, which takes ultimately the way that it works, by the way, is it takes transcripts and also survey responses on my website and it creates personas and has those personas think differently in parallel sub-agents to take fresh looks at content from different perspectives. So that whole idea of having two different things and two different models as well in that instance is a super interesting one. I think we'll see a lot more of that. I can see a lot of value in it. We definitely know that they're pretty good at agreeing with themselves and you often get a better contrarian take if you throw away the context and look at it again. So great principle. I think the tooling is still super early, which we all know, but they're interesting ideas for sure. There's a question just behind you. Yeah, just keep going. They'll turn on. So great talk, by the way. Thank you. Thank you. How much importance do you put on... So in terms of phases of how you develop, you're spending a lot of time building out the context, creating the tickets, and you have a system to run them in sequence. So in terms of phases, it gets pushed up. How much do you focus or emphasize on CICD, running automated tests, linting? Does that give you the confidence to reduce the amount of code you're reviewing? Absolutely. Well, yes and no. It depends what the code is. I think... Firstly, I think CICD good testing is absolutely essential. Linting and all of those things. If you want an AI to do a good job for you, why wouldn't you give it those tools to help it do a good job for you? Just the same way that humans do much better when they have linting and CICD and good tests. It's exactly the same. It's the same with clean code bases. So in terms of phases, it gets pushed up. How much do you focus or emphasize on CICD, running automated tests, linting? Does that give you the confidence to reduce the amount of code you're reviewing? Absolutely. Yes and no. It depends what the code is. I think firstly, CICD and good testing is absolutely essential. Linting and all of those things. If you want an AI to do a good job for you, why wouldn't you give it those tools to help it do a good job for you? Just the same way that humans do much better when they have linting and CICD and good tests. It's exactly the same. It's the same with clean code bases. It's worth doing all of that work to make an AI do well. So that does give me more confidence in what I'm doing. The challenge is if the AI writes the tests and also writes the code, then there's a good chance that it's got something wrong about what you're trying to build. So it doesn't make obvious mistakes. The things that it gets wrong is it just completely misunderstands a feature, builds and says, "Yep, that's fine," and then ships it. And then I'm like, "Oh my goodness." I don't quite let it ship all of the things. Only with pre-release projects do I do that. But yes, it does give me confidence in knowing that the thing is functionally acceptable for release or releasable. I still want to read the diffs because I still don't trust an AI with security. So I don't know if I've lost the screen. Oh, there we go, thank you. I think my computer went to sleep. I just don't quite trust the AI not to lose my customers' data, and I just won't compromise on that. So I will read the diffs because I don't want to be responsible for that. It doesn't feel like there are some problems for which you can trust AI fully, for example, linting, testing. There are some problems which you can't really trust AI just because it's not responsible to do so. So maybe specific changes around security. If you're running production database migration, you should probably check that they worked before running them in production. And there are some that are a bit more fuzzy and hazy. So UI testing is quite interesting. The idea of having a great feedback mechanism. If you're able to get an AI to click through your project to check that it works, that's really powerful. It works 50% of the time in my experience, but it can be quite useful, at least to have a first go at it. Having good end-to-end tests is actually really helpful if you're using Playwright or something like that. For the AirSkills project I showed you, I have some very comprehensive end-to-end tests that set up full file systems of skills and get the AI to create two different personas with running each, a publisher and a creator, and it just checks all the files are in the right places. And that's really, really useful for those kind of full end-to-end tests. Equally, if you're able to build those kind of feedback mechanisms and give the AI a chance to really know whether it's done well, that's a brilliant place to be. And so I'm always looking to figure out how if an AI could tell whether something was good or not, rather than me. And when I'm able to take myself out of that loop, it just massively improves the whole process. It's not always possible or desirable, but as much as I can, I do. You are designing the feedback process though. You're deciding on the criteria and then letting AI execute on top of it. Yes. If I'm working in a team of one, yes. If I am working with a product owner or product manager or a designer, I'm really interested in utilizing their skills and testers as well to figure out ways to design those. This is the value that they bring to these processes, right? It's how, just in the same way that coders are thinking, how could we avoid doing the typing ourselves now? What about if you're a product manager? How do you get the AI to do the easy stuff so that you don't have to do it? How do you go through that process? Same with testers. Super interesting area of research. Two things that come to mind on this. Have you, what works well for in a small team? What worked well for me is setting up adversarial reviews. You have spec, the dev agent goes, develops. There's a reviewer that does an adversarial review. You pass that context back. The dev agent iterates. Generally catches a lot of things, increases the amount of confidence I have to ship it. But even with that, the rate at which you can create specs and how much code you actually have the review, I end up being the bottleneck in the review process still. Have you? Yes. I'm always the bottleneck. I have 30 different things that I need to now review that my AI has done overnight or something. And I'm just like, "Oh my gosh." And the challenge for me is that a lot of that is not work that I should be doing. The only reason that it's given it to me is because I can't trust the AI with it. But any human could do that kind of work. So I'm now like, do I hire humans to just do the boring work? Is that ethical? This is an interesting, really interesting question for us to think through. But you're right. If we're able to design this, and I love your adversarial point to build on something else somebody else was saying, if we're able to do that and design these systems such that we don't have to be in the loop, I think that is better for all of us. Because I don't just want to give a human a terrible job. Till then, we're employed. So yeah, I guess. Thank you. No worries. Any other questions? Should we call it there? Folks, it's been a pleasure hanging out with you. Really, really fun. I appreciate you sharing this, but this transcript appears to contain only speaker labels with no actual dialogue or content to clean. There are no words, filler words, grammar issues, or text to process. Could you please provide the transcript with the actual spoken content included? So I haven't used those ones particularly, but I basically ripped them off. That's great. Then, because I use superpowers a lot, and then I just give tasks like this, and then ask it to like run multiple agents in the background. Yeah. And how you've done that is what I've... Yeah, yeah. So what you can do is there is an agent teams version, which I think I've got turned off in this particular instance of Claude code. But what can happen is you can get Claude to use sub-agents within Teamux. So in fact, I think I might be able to turn that on if I can find the agent teams. There it is. Claude code experimental agent teams. So if I grab that and then run Claude with that on, then you should be able to say, use an agent team to implement the dot tickets in this repo, or something like this. And I don't... This isn't actually running within Teamux. So actually, maybe this won't work. So I might just try this again. Bear with me a second. So if I grab that, and then paste that there, and then grab that, and then run Teamux, and then run that, in theory, in theory, this should start pulling up other agents. And because, like I said, trying to orchestrate Ralph Loops, orchestrate agents myself to try and organize all of the dependencies and complexities is actually really, really difficult to do. But what you can do is just give the job to Claude to do, and it does a much better job of managing that for it. So as you can see, it's already decided to print out the entire thing, the file name and the ticket for each. And it's got a whole bunch there. So it's actually decided that they're all sequential. So therefore, it should run none of them in parallel, which is kind of interesting. And then it should, in theory, start an implementation agent. Let's see if it's going to... I think it's just running as a sub-agent. Never mind. I'm not sure that's going to work. If you can get it to do it, then let me know. But basically, it's an experimental feature that only came out a few weeks ago that allow it to kind of start sub-agents in Claude as well in order to do things. Any other questions while people are working through that? Yeah? You said you built a raft with the NA10 automation you had, right? Mm-hmm. So what was the feedback with criteria in that? So like, what decides if it's a good newspaper article? That's great. Like, I wasn't going to say if a website comes as engineering framework, but what good looks like? If that's dynamic, if that's static, do you put that in the Claude MD file? Yeah, great question. How does that work? So the question is, just to repeat the first half of that, is when I used the NA10 workflow in order to build a route for a newsletter creator, how did the agent know what good was? How did you define good as well? How did I define good? Okay, great question. So in terms of newsletters, I had already been writing my newsletter manually, so I knew roughly what I wanted it to read like and sound like. I also did a bunch of research using a research skill, which was something like which I built, which is something like great newsletters. And I've also did things like that. I also said things like, this is a fantastic written news... In fact, I'm just going to... This is not what I do. I actually do. This is a fantastic newsletter that I've written or that I've read somewhere. Could you please figure out why this is so good? And what are the kind of editorial principles that went into this newsletter for it to work really, really well? And then I would just paste that into... Paste the newsletter in, get it to figure out what was good about the newsletter, and then I would check it, and then I would say yes. There still is an element of human taste here. You can't entirely get away with that. Having said that, I do also have a simulate audience skill, which basically uses a whole bunch of different personas for different clients or prospective clients that I work with. And then I would run the finished newsletter through that and say, run all of these in parallel. And then once you have finished that, figure out ways that I can improve this newsletter or newsletter skill in order to do that. So there's a number of different ways you can do that. The audience simulation is super experimental, but it's actually really effective and often will surface insights I just hadn't thought of. My personality, as I'm a bit slightly all over the place, slightly kind of the way that I talk and communicate, often my clients are not like that. So I tend to barrage people with information, and sometimes my skill will say, okay, there's a lot of ideas in this, Chris. You just need to focus on one main point that makes sense. And I'm like, that's so helpful. So yes, what's quite helpful and interesting is that you can use AI to give feedback on AI like that. The great thing about this particular project that we're writing, this little Pomodoro thing that people are writing, is that it's a command line tool and it's really simple to know whether it works. So it's perfect for a RALF loop. And in fact, these little tools that we build for ourselves, like for example, the newsletter is a skill perfect for this kind of loop. I will often say, I want to improve this skill. Could you please back and forth and write the content, then use another agent to read the content, decide if it's any good, come up with things to improve, then send that back in and just run that as a loop. There's a really cool skill. I wasn't going to tell you about this until the end, but I'll tell you now. There's a really cool feature inside Claw code called loop, where it is instead of creating, in fact, I'll start this in a new session, instead of just doing this thing where you have to create your own while loop, you can say loop every minute, build the next ticket from doc tickets, basically. And then what will happen is the loop will set up a kind of almost like a repeat timer. And as you see, it's got a cron create tool, which for the uninitiated just means do something every minute. This is what those five stars mean. And what it's going to do is it will literally just build the next ticket. When it finishes, it will then check the cron again, build the next ticket. When it finishes, it will check the cron again and keep going. So that's great for working through a bunch of tickets, but it isn't just applied to a set of things that you've got from before. If you think about it, I'll just leave that running up there. You could have a loop that does something like this. So I'll just loop every one hour, check linear for new bug reports from test. And then... I just leave that running. Oh yeah. Can't spell. I'm just going to... Just leave that running. And you're going to get... You're going to annoy your testing team. But anyway, the point is, is that you can run these kinds of loops in order to get work done in a quite an interesting, I guess, dynamic way, even though it's quite a simple loop. Just find the next thing, do the next thing. If you think about it, heck of a lot of our work is just loops. If we're software developers, what do we do? We look at the backlog. We pick the top thing from the backlog. We pull it over to... In progress. We assign it to ourselves. We check on the architecture. We figure out whether there's other contexts we need. We look at the change. We make the change. We submit a PR. We wait for reviews. We comment on the reviews. We reject the reviews. We implement the changes occasionally. We submit the PR. We merge the PR. We then go through the release process. Then we start again. Pick up the next ticket. And so on. That is a loop. It's quite a complicated one, like we talked about just a minute ago. But it is still a loop. It is possible to get an AI to run that entire loop. There's no reason not to. And that's effectively what's happening here. When you can set up... In fact, you would never actually write this. You were much more likely to write something like this, where you'd say, every one hour linear bug finding. And you'd have a skill that encoded all of those chunks of information that I just gave it in a way that would work for you and your particular team. Does anyone not know what skills are? Before I go any further? No. I think pretty much almost everyone knows what skills are. If you haven't figured out what a skill is yet, then this is your homework. Go and understand how skills work. They are the best way that we have at the moment of packaging up useful little parcels of context and scripts and moving them to different places or creating different things. So for example, I have about 50 of them that I've written. And then they just do lots of different things. The great thing about skills is that you can pull them into your context whenever you need them. So for example, and I'll just do this, I could say, do you know how to create images using nano banana? And I can ask the AI the question. And the answer is, well, I could kind of look this up, but it actually knows that I have an images skill for this, funnily enough. But if you hadn't got one, it wouldn't know. But if I then do images and say, how do you create images? Give me the step step-by-step. Then what it's going to do is it's going to pull in that images skill. And then it tells me exactly how it does it. And I've actually written, in fact, I will make that bigger so you can see, I've actually written a script within that skill that actually does the generation for me. So it's codified the process of doing that. And it just picks whichever model it wants to, and it gives it content. And I have these specific templates that I use in order to create specific nano banana skills. Nano banana is brilliant. This is how I created the presentation that you're looking at. I have a slide skill and an images skill that work in tandem in order to create these presentations. Cool. So let's see what the other thing has done. As you can see, it's already on ticket six. The great thing about Ralph Leaves is you just keep working, keep talking about something else. And it's done a whole ton of stuff here. And it's just stopped at this point. But in a second, hopefully, if we just wait, it will start the whole process again. There we go. It's got the scheduled task to run and it's going again. So you can just leave Claw code sessions running with these kind of loops in them. They last about three days, so you do have to keep refreshing them. But you can just do that and keep it running even before you get to a more complicated writer script that wraps Claw to do a thing and all of those kinds of things. Any other questions? Anyone got anything interesting or surprising out of there, Ralph Loop? Has anyone tried this on their real work yet? This would be the interesting thing. Yeah. What was your experience? Have you still got the mic? Yeah. Yeah, yeah. I just made a screenshot Ralph Loop for a website context engineering framework. So Claude just takes the screenshots and then looks at the layout because it has problems with the geometric like spacing. It works well. Nice. Cool. So you're actually using Claude screenshotting to get feedback. Yeah. Yeah, that's pretty advanced. Not many people are doing that. The people are trying to use like Playwright and things like that as well to take screenshots and the Claude in Chrome plugin that comes with Claude as well. You can use that in order to get it to drive Chrome and then take screenshots of what's going on. I've had mixed success with that because it's quite a complex thing for it to manage. But but for just basic screenshots, it works really well. For my images and content that I write, when it runs those images skills, it will always look at the images first to see whether there's any kind of a weird AI garbled text or whatever. And it will reject them without even showing me if there's a problem with an image. Was there another question or comment? Yeah, there's a question just back here. Can you just pass the mic? Is that OK? Thank you so much. Yeah, I think it's close to a question that has been already asked because I'm not quite familiar with rough loops. If I ask the agent to implement task one that has already been implemented, would it actually check the quality of what was implemented or only check if it was done or not? Great question. Yeah, it very much depends on what you set it up for. So there's no kind of magic to a Ralph loop. It's just a loop. So this loop that I'm running at the moment, in fact, I probably should just say loop stop. Otherwise, it's going to keep going and use my quota. I think you can just stop like that. I've got quite a fully featured Pomodoro set up now. Come on. Time to stop. So it depends on... It entirely depends on what you write. So if you go through to... What was the loop that I set up? I think it was this one. I just said build the next ticket. That's very ambiguous and not very helpful. So it might not actually finish it. It may just decide to build it and not ship it. It may not actually be very helpful. So what's more interesting is if you go to... It's probably the easiest thing to do. If I load my Ralph skill... This is my actual skill that I use for Ralph loops. And what I'm doing... Actually, this is slightly out of date. But the one that I've got here is actually using a doc changes folder. You can see that there. But I'm using a doc tickets in this example. But I've changed it on the latest one. But ultimately, you don't have to use a ticketing system like a flat file in the GitHub repository. You could use beads, which is Steve Yaghi's version of this kind of approach, which is quite cool. I've used it. You could use linear. You could use JIRA. You could... You know, as long as you can get access to it from the AI, you can use whatever ticketing system you want. For this... For the purposes of this exercise that you're working through, I tend just to use flat files because they just work. You know, it's not... You don't really need anything sophisticated. In the same way, the Ralph loop is entirely what you make it in terms of its effectiveness. So, for example, in this particular one, I've given it a proper kind of role in the sense that you are one engineer in a relay team, do exactly one change, then drop the context and start again. That's the idea. So, for this one, it's designed to be run in a shell script where it has an entirely fresh context each time because I didn't want the context to pollute each time. These days, I care much less about that because context is so much larger than it used to be. But when I wrote this, that was very important. As you can see, it's specifically for code that doesn't need human review before shipping. And then basically, it tells you about when work should go in there, what the right time for the tool is, read the claw.md, change the format. This is the format of a ticket. This is all of the rationale. These are the different status values. I ask it to check git state to make sure that it hasn't got a working directory. It's also got recovery states. So, if it crashed, it knows that if there's a dirty working tree, but the tests are passing, but you're probably done, but you might not be. So, just double check. If the tests are failing, then it's probably just mid-flight, but broken. So, you should probably just throw it away or just treat it differently. So, you can imagine that this was built up over time of trying to get this working, trying to understand what the user wants, make sure test passing is not enough, verify the actual behavior works, run things in parallel, mark it done, blah, blah, blah, blah. There's an awful lot going on. So, with a real Ralph loop, you want to be building up over time for your specific project, exactly how to check something is working, exactly how to run the test in your particular framework and dialect, how you submit things to the test team, how you want to comment on particular changes, and what style you want to use, whether you want to pull off the thing that feels most obvious to you or the thing that's highest priority, or a mixture of both, depending on how you're feeling that day. Whatever you want needs to be coded in it. So, when you're writing Ralph loop, I'll show you a link to grab this one at the end, but there's no need to use just mine or something else. Just start with mine and then say, fix this for my project and allow it to change and morph and evolve. A couple of questions, so just come forward. Thank you. Hi. Thanks for the talk. Could you expand a bit more on the topic of sandboxing? Because that would be the thing stopping me from running this work. Yeah, it makes sense. Yeah, absolutely. So, there's a number of different ways to sandbox this. For this particular small project, I'm not doing that. Most of my work happens on a VPS, which is away from my main machine. It has a few keys on it that are specific to what I want it to do. And it can access developer tools. A lot of them, it can only access them read-only. It can also access my email. But again, it has quite strict, fine-grained clawed permissions for not sending emails, because that's quite important. I don't let it ever send an email. I only ever let it draft them. So, I use a combination of positioning the code physically, not physically, but away from the machine, on a different machine, on a VPS. I use clawed permissions for that as well. The permission system is a bit broken, but it mostly works. I'm trying to build lockbox to make it even better. I use... What else do I do? So, the keys that I use are separate keys. So, the AI has access to its own keys, which I don't use for my other stuff. So, I can see the kind of audit trail of what it's done. So, there's a number of different ways of doing it. If you want to just run things simply on your own machine, there's Docker Sandbox, which is quite cool. It's a new feature in Docker, which just allows you to do Docker Sandbox clawed, and a run clawed within that sandbox. So, you can kind of isolate it within a specific container. That's quite powerful because it allows you to only change things within that specific place of the file system. The challenge with that is that it can still leak data from one of your systems to another of your systems. There's a thing called the lethal trifecta. I don't know if you've heard of that. Simon Willison coined it. It's an idea that if you have untrusted tokens, internet access, and access to secret important data you don't want to lose, you're going to lose that data, basically. That's the bottom line of it. So, you have to kind of minimize the amount of times that those things collide in the same context. So, yeah, lots to say about security and sandboxing specifically. I tend to run... I don't run with dangerously skip permissions, but I do run with a number of things turned on by default, but not everything, basically. And you kind of have to go through and figure out what your risk profile is and how much you care about those things. And certainly, as you're giving... The main things to read up if you're interested is to read up about the lethal trifecta, if you weren't already aware of it, and kind of be thoughtful about how much power and permission you're giving to your agents, especially if you're using something like OpenClaw, which is unfortunately insecure by default. I know that they've been doing a huge amount of work on OpenClaw to make it more secure, but it is still a challenge for those kinds of agents. They do have access to a lot of things. Any other questions? Yeah. Yeah, you had a validation step in the loop. Mm-hmm. This might be anecdotal evidence, but as soon as I changed mine to use sub-agents here now for the validation step, it started finding things. Ah, interesting. Whereas, as long as you're doing the validation in the same step with the same context, it just pats itself on the back. Yeah, yeah, yeah. That's a really good point. There's definitely confirmation bias going on with agents where they're like, oh, yeah, of course I wrote it. Fine. It was fine. I checked it a minute ago. Yeah, using sub-agents is really powerful because a sub-agent starts with only a small chunk of context. It doesn't start with a full context, right? So you can get much more power from it. So as a good example from this particular project, a really useful skill, which I mentioned earlier, is Simplify. Simplify is a clawed coding bundled skill. And what it does is it will look at the most recent changes, and it will run three sub-agents to try and figure out whether your code should improve. So you can see here what it's doing. So hopefully this will run. These will load. And it will probably find a bunch of problems. Yeah, great point. Great presentation. Thank you. Did you try OpenSpec or combine with OpenSpec or any other spec-driven? No. If I'm honest, I'm not a huge fan of spec-driven development. I know that's controversial and I'll qualify that. I worry that spec-driven development is taking us, at the extreme, is taking us back to the bad old days of Waterfall, where we would specify the entire or try and over-specify a project. Even these little set of tickets, I'm not that comfortable with. I feel like spec should be much more iterative than we can see. It's already fixing a bunch of things. That's quite cool. So it found, just to finish off that point, it found a bunch of issues there. And it's got some fixes. Yeah, so specs. I like just-in-time specs. I like the idea of building or thinking through what you're trying to build, creating some kind of plan and claw code and then executing it. That's fine. I'm happy about that. And I think that's a useful step. What I worry about is, A, I worry about things like Kiri where they've codified that into the tool. I worry that that will almost fossilize that one approach with AI that works today but may not work again when Mythos eventually comes out. So I worry that the tools are being too quick to jump to a specific structure of work that may not be the right thing in the future. So I'm cautious about that. I think it is obvious. It's a truism that AI needs more context in order to do well. So we should try and give it more context. But I think the idea of overdoing that and over-spec-ing a project is one to be careful of, as well as over-structuring our process based on what we know about agents today. Because then we'll end up with working with a new kind of AI-driven process that worked best with agents that came out in 2025 or 2026. You know, we'll still be using that in 2030 and that'll be a pointless waste of time. So those are the kind of concerns I have with it. Any other questions? Yeah, great talk. So you mentioned that you don't like spec-driven and you use RALF. So basically there is no human in the loop. So the question arises, does Claude actually need you there? Where is your input there? Great question. So I've been thinking about this quite a lot recently and having a bit of an existential crisis. I don't know about anyone else. But yes, what value am I adding here to this thing? Certainly not with writing a Pomodoro timer. I'm not sure I'm adding much value at all. I mean, I literally said I one-shotted those specs and there was no point there at all. I'm not saying that I don't like planning out a system. What I'm interested in at this point, and I don't have the answers, is the fact that AI often will pick better specs and write better specs that I can write and will often have a better idea of the kinds of direction my software should go in than I necessarily will have. So I like the idea of actually having RALF loops that create other RALF loops, potentially, or having RALF loops that track whole customer engagements or even whole startups. So I have a skill that I'm working on. Should I show this? I'm going to show it. It'll be fine. What could go wrong? Which is called startup. It's pretty ambitious. But the idea is that it should basically guide a product through an entire startup framework. So it is meant to be run as a loop. The idea is that it... Oh, I see you all taking pictures. Now I'm owning this thing. Damn it. But with great thanks to Ash Morrow, who wrote some brilliant stuff on this. I should say that for the tape. So really, really helpful to me. So I built this out of basically all of the cool books I've read about startups. So I'm a startup founder, co-founder, CTO. So this is near and dear to my heart. And what I'm trying to do here is I'm trying to give the AI enough context such that it could run my startup for me and potentially figure out what the next most important thing to work on is and then do that in a loop. And then there's a big outer loop that runs that says, OK, well, what's the next most important thing to do? Let's do that. So it doesn't work. But it's interesting and it's getting somewhere. And it will often... The first thing it does... I don't think I've got it to show. But... Oh, yeah. No, I will show it because it is hilarious. Hang on just a second. There's a... I asked it how it was doing on one of its loops. And it produced a startup update deck as an investor memo, which was... I didn't even ask it to do this. I'll show you the demo. Hang on a second. Because it's absolutely brilliant. Let's see if I can just show this window. There we go. Air skills. Startup update. And so, yeah, it said, yeah, I need to give him an update. So what I did is it said, basically, this is how far we've got. These are the problems nobody has solved. This is what we know that's real. These are the number of... This is a skills management tool that I'm working on in the background. These are all the kind of issues. And it came up with all of this cool stuff that could go into an investor deck. To be honest, it's not bad. It's not a bad... I think it's actually the GitHub for AI skills. But there we go. And, you know, who's going to pay for this thing? How much will they pay? Those numbers are definitely not right. But what's interesting is that it decided that it wanted to do this and figure out all of these numbers based on this, which I think was hilarious. And it was quite proud of this deck, to be honest. And I have to kind of be like, hang on a minute. We haven't... There's some serious thinking you need to do before you kind of go to that. Will orgs pay for skills government? Would your orgs pay for skills governance? Great question. Not sure yet. So anyway, the reason for showing that is more to kind of point out that AI can do a heck of a lot. And it doesn't do startups well yet. But that's probably down to my skill file, not down to the agent itself. I have a feeling that there are an awful lot of things that potentially will be loops in the future. I only got that far on my slides. Oh, my gosh. Hang on a second. So we've done that. We've done that. If you are still... If you're not just listening to me and are still working on this demo, I've got a couple of challenges for you if you'd like to do this. One is you could try upgrading your ticket format. If you like the raw markdown file, the doc tickets is fine. If you wanted to just type BD install or install beads, it's super easy to do that. And you wanted to kind of get Ralph Loop to work with your beads, try that out. See if that works. There's no pressure on you to achieve anything in this little folder. Beads is great because it only works within your folder. And it just installs a little tool. So it's quite a useful thing to try it on. So if you wanted to try a different ticket format, or you wanted to kind of move this into your main project and connect your Ralph Loop to your ticketing system to see how that feels. Maybe not submit tickets yet, but you potentially could try that and see where that takes you. So that's an option. The other is the skill. You're going to need to keep upgrading and working on your Loop. The Loop basically contains all of the know-how about how you as a person will go through that. You can take it all the way through from do the next ticket, and you can take it all the way up to do the next step in the world dominating startup you're trying to build or whatever it is. You know, it works for all of those kind of things. And what's super interesting about this is that I'm I'm more and more convinced that everything in fact is a Loop. Maybe as an engineer, I'm definitely on a Loop on a lot of the work that I do. Maybe as a project manager or project manager, I'm on a Loop. Maybe as a CEO, I'm on a Loop. Who knows? Maybe certainly a lot of the kind of cadences that I work on run in Loops too. So I have a skill. And if you're running an open claw bot, you're doing a similar thing that just runs a heartbeat every 15 minutes. It just on my VPS, it just fires up clawed, checks a few things, checks my calendar, see if I've got anything happening, and sends me telegram messages. Maybe that's on a Loop. I mean, that is definitely on a Loop. It's 15 minutes. I have a Worker Loop, which I'll show you in a minute. And I have a Morning Loop where every morning at 6am, it comes up with a full briefing of my day, figures out exactly what I should be doing, and just gives me all the information that I need that's happened overnight. All the emails that come in, all of that stuff. The Worker Loop is particularly interesting because it basically... I'm not sure I can actually show this. Let's see if I can find something that I can show. Let's see. No. The reason I can't is because it's got a bunch of client information in it, so I can't show you that. But what I can show you is, for example, this screen I can show you. So if I quickly switch to this... So this is an app. I'll make that slightly bigger. This is now how I run my Worker Loop. So this is an app that I wrote to manage projects. So I don't have tickets inside my work vault. I have project files. And each of the projects is a set of work that I need to do. And then every so often, I basically Vibe Coded a Kanban system. And a worker will pick up and do the next step on the project. So if the next step on the project is writing an email, because it has an overview, or it's checking things, or it's producing the slides for my project, it will do the next step. So this, for example, is the workshop prep spec project that it's working on. And it's got a bunch of front matter that is just looking like that. And it's got some questions for me. I haven't updated this. It needs to be updated. It's got the context. It's got a decision trail of things that it's done and why it's done them. And so basically, for every different thing that's happening, it just is figuring out the next one. It's also got notes on other talks that I might be giving that didn't happen in the end. And it's got feedback from a previous workshop I did on a similar topic, which you can click on. Actually, I can't click on that. There's a bug. But it will basically show the notes from a feedback session. So this project pulls everything together from all of the contexts you can find and then does the next step in a loop. So you can run everything in a loop. You can run all of your work in a loop. When I wake up in the morning, normally, I have about 15 or 16 draft emails where people have got back to me and it's had a go at replying to them. I always have to edit them. They're always okay. But it definitely has a go at getting on with trying to schedule some of my work. I have very specific rules about what it can and can't do. My basic rule is, is this reversible without embarrassment to me? And if the answer is no, don't do it. But just make a little note in the project and hand it back to me. So sending emails is not allowed to do. Creating a slide deck like, for example, this one, that's reversible. It doesn't cause me any embarrassment. So it just got on and did it and gave it to me, for example. It doesn't post on LinkedIn for me. It doesn't send emails. It doesn't send messages. But it does get everything ready for me to review. To your point earlier, which is a very long answer to your question, it has caused me to genuinely question what I'm good at and what I'm here for. Quite a lot of the time, I've got to a point where I'm just the email person who just checks emails and send. Check emails, send. Check. That doesn't sound like a proper job. That doesn't feel good. So therefore, what does that mean for my work? And I've had to make a conscious decision. Which bits of my work do I want to do and which bits of my work don't I want to do? I don't want to be the email reviewer, but I do want to be the strategist. I do want to be helping organisations think through what on earth is going on with AI and how to kind of fix it for their organisation. Now, I could get AI to do a bad first draft, but I don't want to be reviewing AI's draft. I actually want to be doing that thinking myself. So therefore, I basically said, don't do any of that work. I want to do that work. Just give me all the information I need and I'll do the work because I enjoy that work and I'm good at it. So AI can do all of the rubbish work, but it can't and it shouldn't do the work that I'm uniquely good at. But because everything is a loop and Ralph, this is getting so existential, because Ralph loops are so really everything and can be used for everything. We have to start asking hard questions about which of the bits of work we actually want to do. What do we want to do out of this work? It's not just about what AI can do or can't do anymore. Yes, there's a question. There's like loads of questions, but let's go at the back. I think your hand was up first. I think the chap's coming with the mic. Well, just for the recording, it's really helpful. Thank you. So with the open-ended tasks, how would you think about when to stop? So do you set KPIs up at the beginning or how do you know when it's done? Yeah, great question. I ask it to... So again, this comes down to if I just go back to... Sorry, different window. This one, it comes down to upgrading your loop and you have to basically tell it when to stop. So I don't just have one Ralph loop file that works for all of those different loops that I showed you earlier. I have different ones for each. And for example, the worker says, when you get to a point where you're either running out of context or you've got to a point where there's an irreversible thing to do, then I want you to stop and report. And what report means is, in this case, update the project file, which is just a file in a repository, with where you've got to and in a way that where you present it to me for review. So I'm working quite hard at the moment about that kind of presentation step, because I definitely don't want to be reviewing a diff. And just reading text, I find really difficult, just because there's often a huge wall of it and it's hard to parse. So I'm getting it to start giving me things step by step in slide format. I'm trying to get it to... That interface I showed you before, you can kind of see a little bit how I'm trying to get it to show things in different places, in different ways, for it to get to what I need to uniquely do next, I suppose. So to answer your question, it depends on what you're doing. I think the most important bit is that you figure out what the edges are for yourself and have that real moment of, what do I actually want to be doing? How do I want to be involved in this work here? Not just AI is helping me do my work and a companion to me. It's much more now, which bits do I not even need to know about? Next question. There's one here. There's a mic just at the front somewhere, I think. Yeah, great. Thank you. Hi. So since you brought up this topic of our involvement, right? So at what stage we really get involved? I mean, you mentioned you don't really review the diff. I guess the most important part of our work now is creating these tickets, right? I mean, first identifying what the most useful feature to implement is and then describing in a way that you foresee all the different edge cases or just explain it the best way possible so that the outcome is what you desired in the first place. Yeah. What's your process of creating these tickets? Yeah, great question. And so like my concern is sometimes you don't really know yourself until you start implementing them in the way we used to do it, right? So during the development process, you encounter certain different cases where you need like a custom logic, you need to... It's just difficult to foresee these things from the get-go. And do you like iteratively improve the tickets and you re-implement them or what's your process? Yeah, great question. So in terms of how I work, there are two modes of work that is done on my behalf. One is the fully automatic work that we're talking about here where the AI will just get something done. I don't need... When I've got a decent spec that I trust and there's a way of feeding back so that the AI knows that it's good, I don't need to be involved in that work. That can just happen. For every other piece of work I do, I have a... I work on it with and in Claude code. So I... With that system I showed you earlier... In fact, I'll go back to it so I can show you. Let me just change back to this window. Where is it? There it is. With this one, at the very bottom of any of these kind of projects that it's running for me, there's a little thing here which I... This is just a Vicoded app that I've written for me. Nobody else has access to this. It has a little VCP command which if I take that and I type this into a terminal window. So if I go back to this one for example, I think this is okay. Yeah, I can't... I can't easily show that just because my internet is not going to be able to connect to my VPS. But the point is, is that I'm able to go back to... Yeah, I'll go back to that. The point is, is that I'm able to type that in and just paste that into my VPS. What that does is, it starts a new Claude code session. That session grabs, knows where to find that project and it pulls in all of the project context. So rather like loading a skill, it loads the project. It reads the entire project file and knows where everything else is. At that point, it's loaded in everything it needs in order to supercharge that session with me. And then we work together on it. So if I'm, for example, to your point about speccing tickets, I'd have a ticket, I don't know, probably a project for that particular feature. And then I would say, okay, loading everything you know about that. And it would load them all in. And then we would work back and forth on speccing out those tickets. And then the output would be whatever I needed to get done in order to get that done. So if I'm working on something that where I want to usefully and uniquely do that myself, that's when I would jump into a project with Claude code. And when I say do it myself, I don't mean typing it myself, or I don't mean doing all the thinking. I normally mean I get Claude to interview me to ask questions so that I can give it the information it needs to formulate, to do the writing, because I don't like doing the typing. But I get it to pull the information out of me in order to get that work done. So those are the two modes. It's the back and forth iterating. And then it's just the automatic stuff. And I should also point out that I don't like reading diffs. But ultimately, that's the only way that you can review code. When I'm reading a newsletter item, I don't want to read the diff, I want to read the newsletter. Whereas if I'm reviewing code that's important, that is for other people, then yes, I read the diffs. I don't like doing it. Nobody likes reading diffs. But I check to make sure it's working. And I will, I can't see myself not doing that for a while, especially not with security, like conscious code. Maybe with Mythos, I'll just delegate it. That'd be nice. Any other questions? There's one here. How do you deal with context rot? So for example, your example where you have a loop, and it takes one task after the other, is it the same Claude code session that takes all those tasks? You'll have to experiment with that. With the slash loop command, yes, it is. It's the same session. When you run it as a kind of while loop outside Claude, then it's a different session. You have different trade-offs with that. With the same session, you have all the context of the previous tickets and the previous changes. That might be useful. In practice, I've not found that so useful because it can just pull the files as it goes. If you're not typing anything into the session, you're not really adding anything to that. So there's nothing really in there that's useful. So I've tended in the past to prefer starting a fresh context for each new session. But the loop is very, the slash loop is very easy to run and it just works. And especially Opus is very, very good at long context retrieval. So it's less, much, much less of an issue. Okay. Sorry. Yeah, there's one at the back as well. Is there a microphone as well? Okay, great. We'll come back to that. Are you reviewing sessions that are done by your loops or are you just reviewing diffs on the GitHub? Great question. Because I don't allow any of my workers to close a project. So I would always say, if you think you're done, tell me what's finished and I will close. I will close that off. So it could be that there's a big list of completed things that I need to check off myself. But I want to be that kind of final step of verification. The reason that I've added that is because I worry that I'll miss something. There's a thing that someone coined recently called cognitive debt, which is the idea of just not being up to speed with everything that your code base can do, or all of the code in your code base. And that worries me. So I tend to want to at least understand how the code fits together and how the piece of work that I'm working on fits together. So I don't let AI get away with just putting something out of my sight without me having a chance to look at it. Otherwise, I feel like I'd lose track of what's happening. Yeah, because I mean, for example, I I'm using sessions to track tickets. So instead of reviewing the code or this in the code, I'm just reviewing what the particular session was doing. And I even have like a marking system with which session is on which status. Is there any way how you do it similarly? Similar? Yeah, I think the sessions and the status, I think that can work. I haven't tended to use sessions like that. What I've tended to do with sessions is I get Claude to every night go through all of the previous sessions that I've run that day across all the machines I run Claude. It saves them all into a JSON file for me. And then I get it to both figure out how my system could improve and also just what I did so that I haven't I don't forget what happened. So it writes a little paragraph for how much I did. And I use that in order to kind of track work. But it's not quite the same as one ticket per session. I quite like the idea of having like one context per session. I think that's quite a nice idea. And one sorry, one like by per unit of work. I just haven't made that work. I found it really useful because then then I can go back to the particular session when the particular thinking was happening. Yes, and I do do that for sometimes when I've got a project that's running over multiple sessions, I can go back to the previous session. Instead of the VCP command I showed you, I could type VCR and it'll do the same thing. But in practice though, I like the discipline of it having to pick up again. Because it means if it has to pick up again from a fresh context, it means that all of the information that was in that session has actually been codified into other places that any Claude code session or human could find. Which means that you end up with a much more, I guess, richer kind of repository of knowledge that you're working in. So there's a question mark around if sessions are truly not ephemeral and you've got them as a store, are they accessible as future context? If you treat them as ephemeral and make sure you capture everything within them into your repository anyway or into documentation files or whatever, I think that could be more powerful. So worth thinking through for sure. Any other questions? Feels like we've come a long way from just write this ticket, but there we go. It's good. It's all good. Yeah. Thank you so much for the talk. I have a question. It seems like in the loop, some of the stats might not be necessary. Like you might go to the code and then find nothing there. Yeah. And would you consider to optimize it somehow or you just let the token burn? No, just burn the tokens. They're not that expensive. Depends what you're doing. I think we're at the point I should. This is a whole other thing. We're basically in the era of free tokens right now. You know, I have a max 20 subscription. And I definitely use more than the average person probably who is paying for one of those. So I think at this point I would optimize for freeing your own time up as opposed to optimizing for for burning a few more tokens. I don't think tokens will ever get that expensive. I think that the frontier models potentially will be very expensive, but we have really good, cheaper or freer alternatives just around the corner. Not quite as good for the latest kind of work that we're trying to do, but they're really, really good. So, you know, I think there was at least one that just kept the GLM one that just came out. It looks really promising. That's the ZAI one, I think. It just came out this week. Really, really interesting. I'm still running Claude, but that won't necessarily always be the case. So I think I would just burn them. I would, like I said at the very beginning, you know, I spent a long time doing the whole optimization thing where I was doing this, if you weren't here at the beginning, this thing. You know, I spent a lot of time trying to screw around with all of this, but ultimately I just now let it run. It's much simpler. I do get quite close to the end of my max subscription sometimes though. I'm slightly nervous about what that means. I have to figure out how to get another account. Yeah, max to the $200 a month one. Yeah. Yeah. Yeah. I get pretty close to that every week. I'm quite about 80% now getting the jitters. Yeah. You had a question. Do you want to bring the mic back down? Is that okay? Thank you. I'm not looking at that anymore. Hello. Is this, I wanted to ask you about, thank you so much for the presentation. Yeah, sure. But fine tuning for the prompts, do you version it? Do you have datasets that you use to fine tune your entire loop? So in terms of versioning the Ralph loop specifically, like the prompt for the Ralph loop. Yeah. Yeah. So I use skills for that. So as I pointed out before, everything like that goes into the skill and I get Claude to write the skill for me. And that saves in either your .claud skills folder within your project, or it goes into your home directory under .claud skills. I use GitHub to version all of those for myself. I don't think Git is the right skills format for this long term. I think we need a new thing. Hence trying to build S skills, in fact, which is this idea of trying to make skills much more portable and shareable within teams, which I'm trying to figure out. So yes, I do I do version control them and I treat them as quite important code. And I don't actually, I do share some of them, but I don't share all of them routinely because they, there's a lot of my own IP in there. And actually a lot of my customers IP is in there too. The question is more with regards to the performance of the prompts. So, um, you were saying that in the beginning you, as you go along, uh, you improve the prompts as you go along and, but are you versioning that going along and are you versioning the performance of the prompt overall? So when you say prompt, do you mean the, um, the skill itself that I'm using? Yes. Yes. Yes. Yes. Yes. So yes. So the skill, so the prompt lives within the skill. So when I type slash bug tracking or slash Ralph, uh, that, that is the prompt that, that, um, that gets written by Claude, um, and managed by Claude, which means that the, um, uh, that whole file is, is, is, is the prompt and therefore that is version controlled. So I, I always, I have a git running within that setup and then every time I change it, I update, um, update. But you, but you, but you remain subjective. How much you make sure you have improved. Let's say that you have, your data set will be, uh, an issue. I say that. And the expected, uh, output would be, uh, the new feature added to the wrapper. Yeah. So you could more, how do I evaluate whether it's any good or how? Yeah, exactly. How do you know you're actually improving? I see. So how do you know if you're improving? That's a really good question. Um, I do, um, stress test my skills. So with other skills, uh, and I say, you know, is this skill any good? Could you improve it? Could you write it? Um, I do, I, I, I spent quite a lot of my time tinkering with my system and my skills probably more than I should. Um, I think, um, it's a bit subjective at the moment. What I haven't done, and this would be a really good exercise, is to try running, um, blind testing where you would run a set of tickets with one skill and a set of tickets with another. Ultimately, because Claude is non-deterministic anyway, I think there's a high level of variability with any of those kinds of tests. So it's, it's really difficult to think about how to, to construct a useful test in that way, to know whether you're actually improving or not. Um, in general, the more context you give, um, into your prompt, the better it will do, up until a point which isn't very easy and obvious to think, figure out where it becomes worse. So, um, it's about kind of balancing that ultimately, but yeah, I haven't done a kind of objective improvement process. A great question though. It's a question just behind you. How are you version controlling the skills? I'm using GitHub at the moment. Um, I do have a product that I'm trying to build, which is, I mean, this thing here. So if you want my skill, by the way, that's how you get it. Um, it's a project called AirSkills, which you saw a brief preview of earlier from my, my slide deck, uh, that my agent put together for me. But the idea is that, um, you, uh, can package and manage those skills as a unit. So you can create skills for your organization, you can create skill bundles, you can, um, create a skill set for your org that works for different teams within your org. Um, and then that all gets versions and updated for you without everyone having to learn how to use Git and GitHub. That's the idea. Um, it is a real pain at the moment. I found it really, really difficult to manage just for myself, even just putting a skill on GitHub. Uh, it's, you know, I can't imagine anyone from, there's quite a lot of friction for, for a coder like me. It's, I can't imagine non coders using that. So, so yeah, trying to, trying to build this. Um, so yeah, run that command on your machine. You'll have my skill. Um, sorry, there's a question just here first and then go next. Yeah. Um, how do you sort of touch on this a little bit, but sort of around the edges, how do you do knowledge management? So I guess, you know, I use Claude for why is my VPN not working? And then I learned something and I want to record that. And then I'm like, I've got some, um, I'm meeting with somebody with the transcription and I have that somewhere else. And I've got a bit of code that I'm writing and all, I've got all of these different contexts, but they're sort of very disorganized. Do you have a way of thinking about how you organize all of that? Yeah. So I have a code directory and I have a vault directory and those are the two directories I work in. So the code directory contains a few different projects, um, that I work in a more classic way. The vault directory is where I do all of my other work. And, and frankly, I mostly start working in there, even if I'm working on code and just tell it where the code is, uh, because the vault contains, uh, several thousand files with all of the different stuff that I have picked up, learn, um, worked on with Claude over the last several years. Well, not with Claude for that long, but you know what I mean? Um, I started with, uh, Obsidian a long time ago and I've been working on that vault for, for a long time. Um, and, and with Claude now, it just works on that for me. So when I do some research on how to fix my VPN or whatever it is, it just saves a file in there. I have some specific rules for how to kind of structure and manage that. Um, if you're interested in more, I haven't written a lot about this, but I know Andrew Kapathi has just written about it using LLMs as a wiki. That's a great article if you haven't seen it already. Um, I know there's a, uh, Mina Jojovic actually funnily enough has done a thing on Mem Palace yesterday. That's another version of this. Um, uh, you can, you can use that too. There's lots of different systems for that out there. The best way to get started is it's markdown files in a file system and use it like a wiki. So, uh, you run Obsidian in one window and, um, Claude in the other and just kind of work with it and save things as you go. So do you have an agent that then structures and puts those into folders or something like that? Yeah. So it depends on your method. I use the Zettelkasten approach, which is the one where you have one note per thought. So any thought of all just goes into a flat folder. Then I have a slash projects, uh, thing, which has all of the projects that you saw, um, including one for this presentation, um, which, which is my kind of unit of work for an agent that we work on together. It does a lot and then I do some and then it does some. Uh, I have, um, transcripts in there. All of the calls that I've ever recorded go in there. Um, and I use a tool called Lian, uh, which is a command line embeddings tool. So it basically just runs embeddings across the entire, all of the texts in the repository, all of the transcripts, all of the links I've ever saved, including all of the content. It's huge. Um, and, um, then it can find things usefully and easily in there. Um, so you, you, the best time to start that is today because it just takes years to put together. I need to write more about that. Any other questions? Yeah, there's one here. You said you had friction while versioning your skills. I've been using skills only for last month, so I'm not aware of this friction. Can you explain what the friction is? Um, I can. I'm sure there are, has anyone here had, had any kind of friction with my kind of managing and using skills yet? Has anybody else? Yeah, quite a few different people. So yeah, it's, it's a, it's an emerging thing. It's not surprising you haven't experienced it yet if you're not using it for very long. What I've found is that if you're just using them on your own, creating a file of skills and managing them is quite straightforward. Putting them in GitHub is quite straightforward. It's symlinks and GitHub repository. It's fine. What, where it becomes difficult is how you share that. So how would you share a skill? Okay, well, if you want to use MPX skills, you have to then put it in its own GitHub repository. That feels quite heavyweight just for one skill. I'd have to have 50 of them in order to share all my skills. So that doesn't really work. Then it's more like, okay, if I don't want to do that, I just, do I just send them the skill file? Do I send them a zip file? I mean, I can't think of a better way of doing it. Do I have to have a sub module in my skills folder for every single Git repository I share a skill with? It just doesn't make any sense. So I think, I think the idea of Claude has got some stuff in there around plugin marketplaces where you can have a plugin which has a bunch of skills. That's the best way, but then you're versioning the plugin, not the skills. So that's probably the most seamless way. It just, it just doesn't work that well. Also, there's a challenge around if you, if somebody contributes to your skill, do you want their changes or not? It will depend on what the contribution is. Are they local to just them or are they changes that could be generally incorporated? And that depends on the skill and depends on them. So you have to then manage that. So do you run a backlog for each skill where you have tickets to improve the skills? Do you see, do you see the kind of, I think these are all unsolved problems. I'm trying to, my contribution is trying to solve some of those. But these are big problems we haven't figured out yet. Other questions? There's one right at the back. If there's a mic, that would be amazing. Thank you. One, two, one, two. Okay, it's working. My question is about how we can scale up this approach with Ralph Loop. But like for the actual production team, like I, I don't know, three engineers, how to coordinate, how to cooperate. Do you have any idea how we can organize it? Do you have any experience? That's a big question. How, so just to make sure I've understood it, how do you scale this up so that you can coordinate whole teams using this kind of looping approach? Is that, yeah. With all the tickets and the skills, yeah. 100%, yeah. It's difficult. I think the teams that are, where I've seen this work well is where they are proactive about updating tickets. The great thing is if you connect your ticketing system to the AI, it's really good at updating it. So you should definitely do that. Make sure that you claim the ticket and move it into the doing column before it starts work. And make sure that somebody else hasn't just done that before you start. Do you see what I'm saying? That's really important to avoid contention. Those have always been issues with, with bigger teams. Just in the same way, a couple of controversial things, just in the same way that Ralph loops work really well by just doing one thing in a loop and quite sequentially, you know, it could well be that the coordination overhead in our teams is caused by the fact we've got too many people in our teams and maybe we should have smaller teams and just more of them, right? So maybe, maybe if you're trying to get 10 people to coordinate and using AI and Ralph loops and all of that, that's just not going to work. Maybe you need three and maybe that's the way to kind of run that project and then you split it down and then you have another, the other seven people doing something else or whatever. Does that make sense? So, so making the problem go away is the first step and making sure that you're already using your coordination mechanisms as the second step. And then just try it and, and figure out what, what the bottlenecks are. Be really, you know, be really good at retrospectives with this stuff. I think retrospectives and teams are often pretty anemic. It's like, what should we do less of? What should we do more of? That's just a recipe for the same, or more of the same, and just changing tiny increments, which can be good. But ultimately, this requires a radical rethink. So be really conscious in making sure that your retrospectives are changing actual things about how you actually work or, or have space to try, let's just try using a route flip on all of our work for a week and see what happens. You know, and if it doesn't work after two days, that's fine, you know. And then if you are someone who's in a leadership capacity and is able to sponsor that kind of work, this is what it means to try and move to AI. If you want to transform your team, you are going to have to sponsor these kinds of experiments and be okay with failure because, so I'm speaking to the leaders in here for a minute, because it's going to be messy and it's going to, it's going to fail a lot. But if you want real transformation, that's the only way to get it. You've got to give your team space to try a whole bunch of different things. So give them air cover. So yeah, that's a big and complicated question. I think if you're able to and have the agency to, just try it and see where it gets to. There's a whole separate thing called theory of constraints, which I haven't talked about at all, which is the idea that within any team, in any system, there is always a bottleneck. There's always one bottleneck that's the big bottleneck. If you don't work on that one bottleneck, all of the other work that you might do to optimize and improve the system is pointless and actually probably counterproductive. So this is why some teams when using AI tools and using advanced AI tools like Ralph Loots, which is just AI or what we're doing now, but just on steroids, some teams when they implement it actually go slower. Some teams go amazingly fast. Some teams go slow. Why is that? It's because they're not working on the constraint. The constraint in those teams might be the review process or the release process. If you release your code once a month and you're shipping 200 PRs, not 20 in that release, how do you think that's going to go? It's not going to go well. So that's why teams go slower, because what they need to do is fix their release process, not their coding speed. So always fix the thing that is the biggest bottleneck first, then figure out where the bottleneck moves in the system. And that's not predictable. It's random, so you have to figure that out. Then move and fix the next thing in the system. For more on that, read The Goal by Elio Goldratt from 1984, no less. It's an amazing book. One more. Is there another question down here? Is there another mic? Where's the mic? I have the mic. Oh, you've got the mic. Great. Keep going. Since you were asking about, like talking about the constraints part, this reminded me like, I'm part of the AI team and we have an AI team and they write a lot of microservices. It's in different repositories. Okay. How do you deal with like coding now since, is it like one big monorepo or is it like small, small repos? You have to try it different ways and see. I don't think that the the GitHub or sorry, Git architecture, whether it's many repos or one big one really matters. You can always start your AI in a folder above all of your other repos and just get it to work. It does a great job of that. So that's okay. I think the bigger question is what are the coordination patterns within your teams and your services? Who's responsible for what and how does that change with with AI? I think that's a more interesting challenge. The main reason I'm asking this was because some of the microservices depend on others and then you have to release one of them. You need to release a tag and then update another. That's just... Yeah. I think what AI will do is it will expose all of the places in which that process is inefficient because it will do everything faster, which means that if you're seeing those bottlenecks where you are getting dependencies between your microservatives, guess what? That's your biggest bottleneck. Therefore, you fix it. So how do you fix that bottleneck? Well, you might try atomic release system or you might build something that using a route loop that figures out a way of coordinating releases across multiple repos more successfully. I don't know. But that's what you do. That's where you work. Don't work on anything else until that's fixed, if that's the bottleneck. Yeah, there's a question here. Do you want to pass the mic? Yes. I should say, I'm kind of at the end of the content. There's... If you... Which probably was clear half an hour ago. The only other thing... I mean, I had a Q&A slide. If you are leaving, you're welcome to leave, but you're welcome to stay for more questions. I would really appreciate some feedback though. So this this QR code is the only thing that I manually added to these slides. If you could just fill that in, that would be lovely and amazing. Thank you. It's literally only three minutes, four questions. It just helps me to improve and make sure that I do a good job of these workshops going forward. That's also my LinkedIn. I post a lot of content on there. Do connect with me. Do engage with me in some of this stuff. Disagree with me. I love disagreement. I love it when people say, surely, Chris, that's nuts. You shouldn't be doing that. Love those kind of comments because it really helps me to think and improve, which is what I love to do. And because the Ralph Loops is doing all my other work. So got nothing else to do? Great. Thank you. So I just wanted to put that up there. Very happy to continue answering questions though. But if people wanted to drift away, then that might be a good time. Go for it. My question is regarding multi-agent orchestration tools. I'm curious if you've tried things like Steve Yeggie's Gastown or there's another guy who does like MCP agent mail. Yeah, there's some really cool and interesting stuff. I still think we're in the wild west, literally, with Gastown and things like that. But we don't really know how that's going to go. I have tried Gastown. I couldn't really get it to work. But it was pretty early on. For me, I feel like the agent orchestration side of things is I think that we overcomplicate things by assuming that they need to be in parallel. I quite like the idea of just starting with a loop to start with. I don't feel the need to have my AI do lots of things at once before I can get just get it to do one thing well. It kind of goes back to the theory of constraints thing again. I don't think speed of number of tokens per second is the bottleneck. I think that it's our ability to specify what we want and review what the AI has done. If that's my bottleneck, I don't want to introduce more agents. I haven't spent lots of time with those tools for that reason. I feel like they're solving a problem that not many people have yet. So, hello? Not just for speed, but for example, I don't know if you've experimented with MCP agent mail. So agents can lock files and speak to each other so they don't step on each other's toes. And you can use different like cloud, opus, and codecs work on the same project. So you get different brains working on the same project. Nice. Yeah, no, I haven't. It's not just speed. I think I've heard of it, but I haven't tried it. Sounds like a super interesting idea. Rather like someone mentioned earlier about sub-agents trying to look at things from a different perspective. I've had a lot of value with doing that. I mentioned earlier my simulate audience approach, which takes ultimately the way the way that it works, by the way, is it takes like transcripts and also survey responses on my website and it creates personas and has those personas think differently in parallel sub-agents to take fresh looks at content from different perspectives. So that whole idea of having it, having two different things and two different models as well in that instance, is a super interesting one. I think we'll see a lot more of that. I can see a lot of value in it. We definitely know that they're pretty good at agreeing with themselves and you often get a better contrarian take if you throw away the context and look at it again. So great principle. I think the tooling is still super early, which we all know, but they're interesting ideas for sure. There's a question just behind you. Yeah, just keep going. They'll turn on. So great talk, by the way. Thank you. Thank you. How much importance do you put on... So in terms of phases of how you develop, you're spending a lot of time building out the context, creating the tickets, and you have a system to run them in sequence. So in terms of phases, it gets pushed up. How much do you focus or emphasize on CICD, running automated tests, linting? Does that give you the confidence to reduce the amount of code you're reviewing? Absolutely. Well, yes and no. It depends what the code is. I think... Firstly, I think CICD good testing is absolutely essential. Linting and all of those things. If you want an AI to do a good job for you, why wouldn't you give it those tools to help it do a good job for you? Just the same way that humans do much better when they have linting and CICD and good tests. It's exactly the same. It's the same with clean code bases. It's worth doing all of that work to make an AI do well. So that does give me more confidence in what I'm doing. The challenge is if the AI writes the tests and also writes the code, then there's a good chance that it's got something wrong about what you're trying to build. So often it doesn't make kind of obvious mistakes. The things that it gets wrong is it just completely misunderstands a feature, builds and says, yep, that's fine, and then ships it. And then I'm like, oh my goodness. I don't quite let it ship all of the things. Only with pre-release projects do I do that. But yes, it does give me confidence in knowing that the thing is, I guess, functionally acceptable for release or releasable. I still want to read the diffs because I still don't trust an AI with security. So I don't know if I've lost the screen. Oh, there we go, thank you. I think my computer went to sleep. I just don't quite trust the AI not to lose my customers' data, and I just won't compromise on that. So I will read the diffs because I don't want to be responsible for that. It doesn't feel like there are some problems for which you can trust AI fully, like, for example, linting, testing. There are some problems which you can't really trust AI just because it's not responsible to do so. So maybe specific changes around security. If you're running production database migration, you should probably check that they worked before running them production. And there are some that are a bit more fuzzy and hazy. So UI testing is quite interesting early, you know, the idea of having a great feedback mechanism. If you're able to get an AI to click through your project to check that it works, that's really powerful. It works 50% of the time, in my experience, but it can be quite useful, at least to have a first go at it. Having good end-to-end tests is actually really helpful if you're using Playwright or something like that, which it maintains, for the AirSkills project I showed you, I have some very comprehensive end-to-end tests that set up full file systems of skills and get the AI to create two different personas with running each, you know, a publisher and a creator, and it just checks all the files are in the right places. And that's really, really useful for those kind of full end-to-end tests. Equally, if you're able to build those kind of feedback mechanisms and give the AI a chance to really know whether it's done well, that's a brilliant place to be. And so I'm always looking to figure out how, if an AI could tell whether something was good or not, rather than me. And when I'm able to take myself out of that loop, it just massively improves the whole process. It's not always possible or desirable, but as much as I can, I do. You are designing the feedback process though. You're deciding on the criteria and then letting AI execute on top of it. Yes. If I'm working in a team of one, yes. If I am working with a product owner or product manager or a designer, I'm really interested in utilizing their skills and testers as well to figure out ways to design those. This is the value that they bring to these processes, right? It's how, just in the same way that coders are thinking, how could we avoid doing the typing ourselves now? What about, if you're a product manager, how do you get the AI to do the easy stuff so that you don't have to do it? You know, how do you go through that process? Same with testers. Super interesting area of research. Two things that come to mind on this. Have you, what works well for, in a small team, what worked well for me is setting up adversarial reviews. Yep. You have spec, the dev agent goes, develops. Nice, yes. There's a reviewer that does an adversarial review. You pass that context back. The dev agent iterates. Generally catches a lot of things, increases the amount of confidence I have to ship it. But even with that, the rate at which you can create specs and how much code you actually have the review, I end up being the bottleneck in the review process still. Have you? Yes. I'm always the bottleneck. I have 30 different things that I need to now review, that my AI has done overnight or something. And I'm just like, oh my gosh. And the challenge for me is that a lot of that is not work that I should be doing. The only reason that it's given it to me is because I can't trust the AI with it. But any human could do that kind of work. So I'm now like, do I hire humans to just do the boring work? Is that ethical? You know, this is kind of an interesting, really interesting questions for us to think through. But you're right. If we're able to design this, and I love your adversarial point to build on something else somebody else was saying, if we're able to do that and design these systems such that we don't have to be in the loop, I think that is better for all of us. Because I don't just want to give a human a terrible job. Till then, we're employed. So. Yeah, I guess. Thank you. No worries. Any other questions? Should we call it there? Folks, it's been a pleasure hanging out with you. Really, really fun. I am just aologian, just aologian, just aologian, just aologian, just aologian, just aologian, just aologian, just aologian, just aologian, just aologian, just aologian, just aologian, just aologian,