SPEAKER_00
Hello there! My name is Raymond Weidekamp and today I'm going to talk about recursive coding agents which is this idea of applying the lessons of recursive language models, RLMs, to coding agents. This is some work that I have done both in my independent research, raw works, and also more recently in my role at OpenPros. So to motivate this a little bit, we all want outcomes. We all want agents that are working on our behalf. We want reliable co-workers that are getting things done while we're doing something fun, while we're out on a hike, while we're chilling, while we're doing the do. And my argument and my experience is that the bottleneck to this is not intelligence. The models are intelligent enough. They know all kinds of things. They know the entire internet, but they can't reliably deliver outcomes. And so I can't trust them. So as a very simple example, one day I get almost a fully working SaaS app from a single prompt, granted a long prompt. The next day, and I swear this actually happened, cloud code empties the entire contents of my Solana wallet. Oops! Okay. So that doesn't really instill trust. So at the bottom here, we've got this progression. Okay. And we all want to move towards the one on the right where we're sitting there and meditating and things are manifesting. And so where does that come from? This is from the AI engineer code. It's actually from the back of the t-shirt engineer code, November 2025, man. I hope you're there. If you weren't watching on YouTube, it was amazing. So here's the thesis. The thesis is today's agents are mismanaged geniuses. The intelligence is there and the missing layer is how do we specify and manage and reuse and verify the work? So this framing, the phrase the mismanaged genius, comes from Alex Zhang, Zed Li and Omar Khatab at MIT. And Alex and Omar are part of the authors of the original recursive language models paper. I've also talked a little bit about this recently on turning post. I forgot to mention that these slides are actually a website, recursive coding agents.com. So you can click on them by going to this website. So everything I'm going to show in here is interactive. Okay. What are recursive language models? So I like to say that in an RLM, the context itself is the object of computation. And this is essentially a marriage of tool calling and reasoning. We're going to talk more about that in the next slide, but the idea is that the full prompt is not a simple user query. The full prompt is a variable. The full prompt could be a file or many files. And we have this read evaluate print loop repl that the agent is interacting with in the original paper. That's Python. And the RLM is instructed to operate symbolically on that prompt. So don't just read the whole thing into your context window, explore it symbolically and even more, you don't even directly export symbolically, or maybe you do a little bit of poking around, but have other LLMs. And I guess other RLMs, if you allow the recursion depth to be greater than one, have these other recursive sub agents. And again, we'll get a little bit into the weeds of the lingo, sub RLMs, sub LLMs, do this symbolic manipulation to pick apart the answer and then work our way back up to a final answer. So it looks something like this in this tree below. So my take here is that RLMs are the new reasoning models. And I see this as the next paradigm of test time compute, inference time compute, whatever you want to call it. And why does it seem obvious or maybe like, hey, why is this even a thing? I think it's very elegant because it's a very elegant marriage of two things, reasoning and code execution. So the code execution is reasoning. And so instead of we had long chain of thought as a prompting strategy that evolved into reasoning models that explicitly expressed the chain of thought as their reasoning tokens. We already had function calling, tool calling, parallel tool calling. And RLMs really puts that together in a way that gets amazing results. So three very simple examples, one from the original paper, Oolong, the RLMs can process information that is many orders of magnitude larger than their context window. Tens of millions, millions of tokens. What I showed in my own independent work was that the default RLM harness is itself a really powerful memory system. So RLM with no modifications is essentially like a top 10 memory system and up there with all the people custom making memory systems. And there's probably billions of dollars going into that. And with a little bit of modification, you can get really amazing results using it as memory. I was also able to show state-of-the-art results where the RLM framework and specifically the DSPi implementation of it was able to get state-of-the-art results on long reasoning tasks. Now there's this new benchmark, long COT. I won't go into the details in depth, but the idea of this benchmark was that the problems are hard specifically because they require so many steps of reasoning, in the analysis sequence or the chain of thought, that most reasoning models, including the top ones can't hold the thread for long enough. If you allow the RLM to solve the problem using a combination of code and recursive calls to sub-agents, then a very small model, QUEN 3.5 9b, you could run this on a laptop, can actually beat, so QUEN 3.5 9b as an RLM can beat OPUS and GPT 5.4, all the top frontier models as LLMs on these long reasoning tasks. So they're extremely, extremely powerful. So powerful that they are arguably too hot to benchmark. So two examples here on the left, a very high profile case where the Symbolica team has this RLM agent harness called Agentica. Within hours of ARC AGI 3 being released where the top scores of all the frontier models were around two or three percent, the Symbolica team showed thirty something percent. This is crazy. They blew it out of the water within hours using RLMs as a framework. So much so that it's very much upset the ARC Prize team. And so they gave them what I'm interpreting as a consolation tweet, which as far as I'm reading the situation was essentially saying, congratulations, but you didn't solve the problem the right way. And we don't like RLM harnesses. And so you can have this nice tweet, but we refuse to actually do the full private part of the ARC AGI evaluation,
SPEAKER_00
Two or three percent. The Symbolica team showed thirty something percent. This is crazy. They blew it out of the water within hours using RLMs as a framework. So much so that it's very much upset the ARC Prize team. And so they gave them what I'm interpreting as a consolation tweet, which as far as I'm reading the situation was essentially saying congratulations, but you didn't solve the problem the right way. And we don't like RLM harnesses. And so you can have this nice tweet, but we refuse to actually do the full private part of the ARC AGI evaluation, which to me is just insane. In my own work and on the long COT benchmark, my results as well as Alex from MIT, the RLM first author, encouraged the leaderboard maintainers to actually make a separate open harness leaderboard, so that the results of the RLMs could be showcased without contaminating the original intent of the leaderboard, which was basically no tool calling is allowed. So my take on this is I don't care. I don't care whether it's latent space or reasoning tokens or code execution. I want results and I want AI programs that get those results. Okay. So this can feel close to a lot of other things. And I built a little rubric. There's a companion GitHub repo for this that you can go through if you want to see. And so what do we need to be an RLM? We have an executable environment. The prompt is externalized. There's code. That's actually the thing calling the model. The model is able to pick the decomposition of the problem into the sub calls or sub agents. And the state itself is staying symbolic, right? So obviously plain LLMs and RAG and things like that don't meet those coding agents and sub agents and loops. They get close, but they're not quite there. And again, the rubric here is not to start fights or nitpick. It's just trying to explain what's the essence of RLM and recursive coding agents. Another example that's close, but in a cigar would be hard coded map reduce. And I would put this project called Lambda RLM in that category, which is essentially a way of saying okay, I'm decomposed the problem using Lambda calculus into a map reduce. And then there are like LLM calls in that executing the map reduce, but the LLM is not deciding or the RLM is not deciding how to decompose the problem. And that I see as a key element of this that makes it very agent native, you might say. Okay. So now RLMs, what about recursive coding agents? Okay. It looks the same to me. We just swap RLMs and LLMs for agents and sub agents. And don't we have the same thing? And yeah, you do. And I think that you could take this perspective of trick question. Like RLM is a coding agent and it's already recursive. So that's fine. And I don't think that argument is wrong, but it doesn't really move anything forward. And so what I'm interested in is this question of how can we apply the principles of RLMs to coding agents and make them actually useful for coding agents. And I've been very obsessed with this problem since the October RLM blog post came out in 2025. Okay. So I show some of my experiments on recursive coding agents. The first one was simply wrapping Alex's RLM package as a CLI. So the idea was I just want to give my coding agent an RLM as a tool call. So let's say we need to go sift through a hundred million token corpus. Well now it just used this tool. And then the RLM does the RLM thing. So that's interesting and it can be very useful. And a few other people have built things like that. But then I thought, well, what would it really mean? What would be possible or how could it be possible to make the coding agent fully recursive? So the coding agent harness calls itself, like the exact version of itself. And how might you implement that? And that is what I called YPI. Y stands for the Lambda calculus Y combinator and PI in case you haven't heard of it is a really awesome coding agent. It's incredibly minimal and it's specifically designed to be extensible. So Mario wants you to write extensions for PI for new features that you would like rather than trying to stuff your ideas into the main agent. So it's meant to me a very minimal core that you extend however you like. And when I originally had the recursive coding agent idea and wanted to do it with PI, I was not able to use PI extensions to achieve this goal. And so I had to fork it instead. I'm very excited to report that in anticipation of this talk, I revisited this and now PI has evolved, and the PI extensions have evolved such that you can make it fully recursive with a pure extension. So I have both the pure recursive extension, PI recursive, package as well as the YPI wrapper. That is a convenience wrapper for this. So this is very quite literally a recursive coding agent in the sense that PI calls PI calls PI calls PI. You can set the depth however you want. And now I want to show a few other notable projects in the space. So obviously there's the original implementation from Alex, our alum, dspi.rlm is my go-to, especially when I'm doing benchmarking. And that's how I got all these amazing results on some of these benchmarks. Axe, I think is a very interesting one because it's incredibly agent native. So the ax started out as a TypeScript variation on dspi when our limbs came out, they obviously implemented it. But they did it in this way that enables the ax agent to write a whole TypeScript interface to another ax agent and go all the way down the recursive rabbit hole, which I think is really cool and very interesting just to showcase that this REPL could be anything. There's an example from Dan at Open Pros who made the UNIX RLM. This is pure bash and the environment is just the Linux file system. So that's a whole other angle of thinking about what's possible with RLM. And then lastly, and we'll talk more about Open Pros, but Open Pros as a language actually enables you to convert any coding agent into an RLM. And I'll talk a little bit more about how to do that towards the end of the talk. And then the Open Pros repo also contains a harness that executes. It's a coding agent harness that will let you use codex SDK or cloud code underneath and do an RLM style execution of these pros programs.
SPEAKER_00
This is pure bash and the environment is just the Linux file system. So that's a whole other angle of thinking about what's possible with RLM. And then lastly, and we'll talk more about open pros, but open pros as a language actually enables you to convert any coding agent into an RLM. And I'll talk a little bit more about how to do that towards the end of the talk. And then the open pros repo also contains a harness that executes. It's a coding agent harness that will let you use codex SDK or cloud code underneath and do an RLM style execution of these pros programs. So is cloud code and RLM. This is the question that keeps getting asked and the original answer on the day of release day of the blog post was no, no, no, it's not. But it was literally the first question that was asked on the very same day to the launch tweet. It's saying, Hey, this is cloud code sub agents, right? Go back to my rubric. If you want to dig into some of the nitty gritty details. But arguably now it is. So over here on the right, we can see Omar saying, Hey, congratulations. Anthropic, cloud code is finally an RLM now that you have dynamic workflows. So what changed? And I think this is an interesting way of explaining what's powerful about recursive coding agents and what RLMs even are, by using this example. So dynamic workflows were released just a few weeks ago. And they make cloud code recursive or capable of doing these recursive workflows. And I would highly encourage you to read this blog post called a harness for every task. It shows six different workflow patterns that are very powerful. Obviously there's many more that you can achieve. And just to show it, I wrote two workflows for cloud code, one that is explicitly not an RLM. So that's a hard coded map reduce workflow. And one that I'm arguing is, you could think of it as deep research over a file system. So pick a handle, assign that to an agent, have it go do some analysis, bring back what it finds, et cetera. Again, these are in the companion repo. And then now to open pros. So dynamic workflows are cool. They're only in cloud code. They're also not the only way to do this. So what if you don't like cloud code or what if you do like cloud code, but you don't want to use dynamic workflows, you want something else. So this is what open pros is all about. So open pros is technically a programming language, but it is not compiled by your computer. It's compiled by your coding agent. It's a markdown spec. It's logical English. You don't need to learn any crazy syntax. And there's actually a command pros, right? That will get cloud code or codex or your favorite or PI or your favorite coding agent to write a dot pros.md file for you. So in that way, it's similar to the cloud code ultra code command where decide and write the workflows for you. And the pros has the ability to turn any agent that's got a file system and sub agents into an RLM. So this is an open source repo. You can check it out. And I've also written a little bit more in depth about this for Turing post in this article. And the key thing that I want to bring up with regards to the RLMs is that open pros can explicitly declare the sub agent work. So again, I've made two demo pros.md files that are in this companion repo. If you want to dig into the code, I'm not going to do that here in the slides, where you can break a problem up into smaller pieces that are assigned to sub agents, verify the work of those sub agents in the parent agent session. And what's even cooler is you can actually in open pros the features that I added to the language are that you can add skills and tools as explicit dependencies. So you can imagine a workflow where a certain sub agent needs a very specific skill to do its role in the workflow or must have access to a certain CLI tool, for example, or it can't run and do its job. And so there's a way in pros to actually wire those in as dependencies to ensure that not only is the way the work is done what you want, but actually that the sub agents are specifically configured with the tools and skills that they need to successfully do the work that you are declaring in the pros contract. Okay. Super cool. What can you actually do? So I've got two examples from cloud dynamic workflows, two examples from open pros repo. Scale migrations. This was the launch post example. Refactor, a huge thing, all with a big swarm and parallel and then merge the whole thing together. That's super cool. This idea of going after a directory and then deep research or deep analyze or deep process in some way recursively inside that is another example I have here. You can do audits, bug sweeps. You can do adversarial things such as having a skeptical agent or a red team set of agents that are going to try to improve the system adversarially or in parallel. And then one really cool thing that you can do with open pros that I just added recently is it goes back to my very first slide, right? So one day I get this amazing result the next day they trade away all my crypto currency, which is very small, thankfully. But how do we get these things to be more reliable and how do we get them to be, you know, we have a golden session and we have a great day and now we want to capture that and reuse it over and over again. So I built a system where you can take a golden session for cloud code code, codex, PI, whatever you want. And it's a pros program that will actually have the agent deconstruct that session and turn it into a reusable pros workflow. That again can involve this idea of recursive coding agents to get you to a reliable way of getting to that golden state of performance over and over again. Recursive coding agents for the win. I really think this is very powerful. RLMs just blew my mind when they first came out. And again, as you can probably see through this talk, I've been absolutely obsessed with applying the ideas of RLMs to coding agents. The three things that I hope you'll take away from this is one, trust is reliability. How can we trust something that isn't reliable? And again, my argument and this idea of the mismanaged genius is that the next step is not more raw intelligence. It's actually behavioral textual orchestration. I personally believe that our
SPEAKER_00
of performance over and over and over again. Recursing, recursive coding agents for the win.
SPEAKER_00
I really think this is very powerful. Our LMS just blew my mind when they first came out. And again, as you can probably see through this talk, I've been absolutely obsessed with applying the ideas of our LMS to coding agents. The three things that I hope you'll take away from this is one, trust is reliability. How can we trust something that isn't reliable? And again, my argument and this idea of the mismanaged genius is that the next step is not more raw intelligence. It's actually behavioral textually orchestration. I personally believe that our LMS represent this new paradigm of test time, compute inference time, compute, where tool calling and reasoning are unified and we reason through tool calling and we can recursively iterate. And one of those tools is to call another agent to go do it on some other specific task or subset of the problem. And then also, I hope we settle a little bit of this drama around, wait, our LMS, aren't they just coding agents? Yes. Coding agents can be our LMS. They aren't automatically our LMS. And so I've showed a couple of different cloud code dynamic workflows that can turn cloud code into an RLM, as well as some ways of doing this with open pros that you can use with any coding agent. So I see this as an incredibly powerful way of working with coding agents. I hope you will dig in more to recursive coding agents.com, but with great power comes great responsibility. So until next time, please recurse responsibly. Thank you very much.
SPEAKER_00
that are getting things done while we're doing something fun, while we're out on a hike, while we're cold chilling, while we're doing the do. And my argument and my experience is that the bottleneck to this is not intelligence. The models are intelligent enough. They know all kinds of things. They know the entire internet, but they can't reliably deliver outcomes. And so I can't trust them. So as a very simple example, you know, one day I get almost a fully working SaaS app from a single prompt, granted a long prompt. The next day, and I swear this actually happened,
SPEAKER_00
cloud code empties the entire contents of my Solana wallet. Oops! Okay. So that doesn't really instill trust. So, uh, at the bottom here, we've got this pro this progression. Okay. And we all want to move towards the one on the right where we're just sort of sitting there and meditating and, and things are manifesting. And so where does that come from? This is from the AI engineer code. It's actually from the back of the t-shirt engineer code, November, 2025, man. I hope, I hope you're there. If you weren't watching on YouTube, it was, it was amazing. So here's the thesis.
SPEAKER_00
The thesis is today's agents are mismanaged geniuses. The intelligence is there and the missing layer is how do we specify and manage and reuse and verify the work? So this, uh, framing this phrase, the mismanaged genius, uh, comes from Alex Zhang, Zed Li and Omar Khatab at MIT. Um, and Alex and Omar are, uh, part of the authors of the original recursive language models paper. Uh, I've also talked a little bit about this recently on turning post. Um, I forgot to mention that these slides are actually a website, recursive coding agents.com. So you can click on them, uh, by going to this website. So everything
SPEAKER_00
I'm going to show in here is, is interactive. Okay. What are recursive language models? So I like to say that in an RLM, the context itself is the object of computation. Um, and this is essentially a marriage of tool calling and reasoning. We're going to talk a lot more, more about that in the next slide, but the idea is that the full prompt is not a simple user query. The full prompt is a variable. The full prompt could be a file or many files. Um, and we have this read evaluate print loop repl, um, that the agent is interacting with in the original paper. That's Python. And the RLM is instructed to operate
SPEAKER_00
symbolically on that prompt. So don't just read the whole thing into your context window, um, explore it symbolically and, uh, even more, you don't even directly export symbolically, or maybe you do a little bit of poking around, but have, uh, other LLMs. Uh, and I guess other RLMs, if you allow the recursion depth to be, to be greater than one, uh, have these, uh, other recursive, um, sub agents. And again, we'll get a little bit, uh, a little bit into the weeds of the lingo, um, sub RLMs, sub LLMs, uh, do this symbolic manipulation to pick apart the answer and then work our way back up to a final answer. So it looks
SPEAKER_00
something like, like this in this, in this tree below. So my take here is that RLMs are the new reasoning models. And I see this as the next paradigm of test time, compute inference time, compute, whatever you want to call it. And why does it seem obvious or maybe like, Hey, why is this even a thing? Um, I think it's very elegant because it's a very elegant marriage of two things, reasoning and code execution. So the code execution is reasoning. Um, and so instead of we, we had long, um, we had chain of thought as a prompting strategy that evolved into reasoning models that explicitly
SPEAKER_00
expressed the chain of thought as their reasoning tokens. We already had function calling, tool calling, parallel tool calling. Um, and RLMs really puts that together in a way that gets amazing results. So three very simple examples, one from the original paper, Oolong, the RLMs can process information that is many orders of magnitude larger than their context window. Tens of millions, millions of tokens. Uh, what I showed in my own independent work was that the default RLM harness is itself a really powerful memory system. Um, so RLM with no modifications is essentially like a top 10
SPEAKER_00
memory system and like, you know, up there with all the people custom making memory systems. And there's probably billions of dollars going into that. Um, and, uh, with a little bit of modification, you can get really amazing results, uh, using it as memory. I was also able to show state-of-the-art results where the RLM framework and specifically the DSPi implementation of it was able to get state-of-the-art results on long reasoning tasks. Now there's this new benchmark, long COT. I won't go into the details in depth, but the idea of this benchmark was that the problems are hard specifically because they
SPEAKER_00
require so many, um, steps of reasoning, uh, in, in the like analysis sequence or the chain of thought, um, that most, uh, reasoning models, including the top ones can't hold the thread for long enough. Um, if you allow the, uh, RLM to solve the problem, uh, using a combination of code and recursive calls to sub-agents, then a very small model, QUEN 3.5 9b, you could run this on the laptop, uh, can actually beat, so QUEN 3.5 9b as an RLM can beat OPUS and, um, and GPT 5.4, all the top frontier models as LLMs on these long reasoning tasks. So they're extremely, extremely powerful.
SPEAKER_00
So powerful that they are arguably too hot to benchmark. So two examples here on the left, a very high profile, uh, case where, um, the Symbolica team has this RLM agent harness called Agentica. Within hours of ARC AGI 3 being released where the top scores of all the frontier models were around two or 3%, the, uh, Symbolica team showed 30 something percent. This is crazy. They blew it out of the water within hours using RLMs as a framework. So much so that it's very much upset the ARC Prize team. And so, uh, they gave them what I'm interpreting as a consolation tweet, which as far as I'm reading the situation was essentially saying, you know, la-di-da, congratulations,
SPEAKER_00
but you didn't solve the problem the right way. And we don't like RLM harnesses. And so, uh, you can have this nice tweet, but we refuse to actually do the full private part of the ARC AGI evaluation, uh, which to me is just insane. Uh, in my own, uh, work and, and on the long COT benchmark, um, my results as well as Alex, uh, from MIT, the RLM first author, uh, encouraged, let's say the, uh, the, uh, leaderboard maintainers to actually make like a separate open harness leaderboard, so that the results of the RLMs could be showcased without contaminating the original intent of the leaderboard, which was basically no tool calling is allowed. So my take on this is I
SPEAKER_00
don't care. I don't care whether it's latent space or reasoning tokens or code execution. I want results and I want AI programs that get those results. Okay. So this can feel close to a lot of other things. And I, I built a little rubric. There's a companion GitHub repo for this that you can go through if you want to see. Um, and so what do we need to be an RLM? We have an executable environment. The prompt is externalized. There's code. That's actually the thing calling the model. The model is able to pick the decomposition of the problem into the sub calls or sub agents. And the state itself is staying symbolic, right? So obviously plain LLMs and rag and things like
SPEAKER_00
that don't, don't meet those coding agents and sub agents and loops. They get close, but they're not quite there. And again, the rubric here is not to like start fights or nitpick. It's just trying to explain like what, what's the essence of, of RLM, uh, and recursive coding agents. Another example that's close, but in a cigar would be hard coded map reduce. And I would put, um, uh, this, this project called Lambda RLM in that category, which is essentially a way of saying, okay, I'm decomposed the problem, uh, using Lambda calculus into a map reduce. And then there are like LLM calls in that executing the map reduce, but the, but the LLM is not deciding or the RLM is not
SPEAKER_00
deciding how to decompose the problem. And that I see as like a key element of this that makes it very agent native, you might say. Okay. So now RLMs, what about recursive coding agents? Okay. It looks the same to me. We just swap RLMs and LLMs for agents and sub agents. And don't we have the same thing? Uh, and yeah, you do. And I think that you could take this perspective of ha ha, trick question. Like RLM is a coding agent and it's already recursive. So wadida. And that's fine. And I don't think that argument is wrong, but it doesn't really move anything forward. And so
SPEAKER_00
what I'm interested in is this question of how can we apply the principles of RLMs to coding agents and make them actually useful for coding agents. And I've been very obsessed with this problem since the October, uh, RLM blog post came out in 2025. Okay. So I show some of my experiments on recursive coding agents. Um, the first one, it was simply wrapping Alex's RLM package as a CLI. So the idea was, I just want to give my coding agent an RLM as a tool call. So let's say we go, we need to go sift through a hundred million token corpus. Uh, well now it just used this, that tool. And then
SPEAKER_00
the RLM does the RLM thing. So that's interesting and it can be very useful. Uh, and a few other people have built, uh, things like that. Uh, but then I thought, well, what would it really mean? What would be possible or how could it be possible to make the coding agent fully recursive? So the coding agent harness, like calls itself, like the exact version of itself. Um, and how might you implement that? And that is what I called YPI. Uh, Y stands for the Lambda calculus, uh, Y combinator and PI in case you haven't heard of it is really awesome coding agent. It's incredibly minimal and it's
SPEAKER_00
specifically designed to be extensible. So Mario, uh, wants you to write extensions for PI for new features that you would like rather than trying to stuff your ideas into the, the main, um, agent. So it's meant to me a very minimal core that you extend however you like. And when I originally had the, uh, recursive coding agent idea and wanted to do it with PI, I was not able to use PI extensions to achieve this goal. Um, and so I had to fork it instead. I'm very excited to report that in anticipation of this talk, I revisited this and now PI has evolved, um, and the PI extensions have evolved such that you can, uh,
SPEAKER_00
make it fully recursive with a pure extension. So I have both the pure recursive, um, extension, PI recursive, uh, package as well as the Y PI wrapper. That is like a convenience wrapper for this. So this is very quite literally a recursive coding agent in the sense that PI calls PI calls PI calls PI. You can set the depth however you want. Um, and now I want to show a few other notable projects in the space. So obviously there's the original, uh, implementation from Alex, our alum, dspi.rlm is, my go-to, uh, when I'm especially doing benchmarking. Um, that's how I got all these amazing results,
SPEAKER_00
uh, on some of these benchmarks. Axe, I think is a very interesting one because it's incredibly agent native. So the ax started out as a typescript variation on dspi when our limbs came out, they obviously implemented it. Uh, but they did it in this way. That's, um, enables then ax agent to like write a whole typescript interface to another ax agent and go all the way down, uh, the, the recursive rabbit hole. Um, which I think is really cool and very interesting just to showcase that this like REPL could be anything. There's a, an example from Dan at open pros who made the UNIX RLM.
SPEAKER_00
This is pure bash and the environment is just the Linux file system. So that's a whole nother angle of thinking about, uh, what's possible with RLM. And then lastly, and we'll talk more about open pros, um, but, uh, open pros as a language, uh, actually enables you to convert any coding agent into an RLM. And I'll talk a little bit more about how to do that towards the end of the talk. Uh, and then the open pros repo also contains a harness that executes. It's a coding agent harness that, uh, will let you use codex SDK or cloud code underneath and do an RLM style execution of these pros programs.
SPEAKER_00
So is cloud code and RLM. This is like the, the question that keeps getting asked and the original answer on the day of release day of the blog post was no, no, no, it's not. Uh, but it was literally the first question that was asked or like asked on the very same day, uh, to the launch tweet. Um, it's basically saying, Hey, this is cloud code sub agents, right? Uh, go back to my rubric. If you want to like dig into some of the, the nitty gritty details. Um, but arguably now it is. So, uh, over here on the right, we can see Omar saying, Hey, congratulations. Anthropic, uh, cloud code is
SPEAKER_00
finally an RLM now that you have dynamic workflows. So, so what changed? And I think this is an interesting way of explaining, um, what's powerful about recursive coding agents and what RLMs even are, by using this example. So dynamic workflows were released just a few weeks ago. Um, and they make cloud code recursive or capable of doing these recursive workflows. And I would highly encourage you to read this blog post called a harness for every task. It shows six different workflow patterns that are very powerful. Obviously there's many more that you can achieve. And just to, to show it, I,
SPEAKER_00
I wrote two workflows for cloud code, uh, one that is explicitly not an RLM. So that's like a hard coded map reduce workflow. And one that I'm arguing is that is, uh, you could kind of think of it as like deep research over a file system. So, you know, pick, pick a handle, um, assign that to an agent, have it go do some analysis, bring back what it finds, et cetera. Uh, again, these are in the companion repo. Uh, and then now to open pros. So, so dynamic workflows are cool. They're only in cloud code. They're also not the only way to do this. So what if you don't like cloud code or what if you do like
SPEAKER_00
cloud code, but you don't want to use dynamic workflows, you want something else. So this is what open pros is all about. So open pros is technically a programming language, but it is not compiled by your computer. It's compiled by your coding agent. Um, it's a markdown spec. It's logical English. You don't need to learn any kind of crazy syntax. And there's actually a command pros, right? That will get cloud code or codex or your favorite or PI or your favorite coding agent to write a dot pros.md file for you. Um, so in that way, it's similar to the, the, um, cloud code, um, ultra code command where like decide and write the workflows for you.
SPEAKER_00
And the, the pros has the ability to turn any agent that's got a file system and sub agents into an RLM. So this is an open source repo. You can check it out. Um, and I've also written a little bit more in depth about this for Turing post in this article. And the key thing that I want to bring up with regards to the RLMs is that open pros can explicitly declare the sub agent work. So again, I've made two kind of demo pros.md files that are in this companion repo. If you want to dig into the code, I'm not going to do that here in the slides, uh, where you can break a problem up into smaller
SPEAKER_00
pieces that are assigned to sub agents, verify the work of those sub agents, um, in the, in the parent agent session. Uh, and what's even cooler is you can actually in open pros to the features that I added to the language are that you can add skills and tools as explicit dependencies. So you can imagine a workflow where a certain sub agent needs a very specific skill to do its role in the workflow or, uh, must have access to a certain CLI tool, for example, or it can't run and do its job. And so there's a way in pros to actually wire those in as dependencies to ensure that, um, not only that the
SPEAKER_00
way the work is done, um, is what you want, but actually that the sub agents are specifically configured with the tools and skills that they need to successfully do the work that you are declaring in the pros contract. Okay. Super cool. What can you actually do? So I've got two examples from cloud dynamic workflows, two examples from open pros, uh, repo scale migrations. This was kind of the launch post example, refactor, like a huge thing, uh, all with a big swarm and parallel and then merge the whole thing together. That's super cool. Um, this idea of like go after a directory and then deep
SPEAKER_00
research or deep analyze, uh, or deep process in some way recursively, uh, inside that is, is another example I have here. You can do audits, uh, bug sweeps. You can do adversarial things such as having a skeptical agent or, um, you know, a red team, uh, set of agents, uh, that are going to, uh, try to, uh, improve the system adversarially or in parallel. And then, uh, one really cool thing that you can do with open pros that I just added recently is kind of goes back to my very first slide, right? So like one day I get this amazing result the next day they trade away all my, all my, uh, crypto currency,
SPEAKER_00
which is very small, thankfully. Um, but like, how do we get these things to be more reliable and how do we get them? You know, we, we have like a golden session and we have a great day and now we want to capture that and reuse it over and over again. So I built a system where you can take a golden session for cloud code code, code X, PI, whatever you want. And it's a pros program that will actually have the agent deconstruct that session and turn it into a reusable pros workflow. Um, that again can involve this idea of recursive coding agents, um, to get you to a reliable way of getting to that golden state
SPEAKER_00
of performance over and over and over again. Recursing, recursive coding agents for the win. I really think this is very powerful. Our LMS just kind of blew my mind when they first came out. And again, as you can probably see through this talk, I've been absolutely obsessed with applying the ideas of our LMS to coding agents. The three things that I hope you'll take away from this is one, like trust is reliability. Like how can we, how can we trust something that isn't reliable? And again, my argument and this idea of the mismanaged genius is, uh, that the next step is not, um, more raw
SPEAKER_00
intelligence. It's actually, uh, behavioral textually orchestration. I personally believe that, um, our LMS represent this new paradigm of test time, compute inference time, compute, uh, where tool calling and reasoning are unified and we reason through tool calling and we can, uh, recursively iterate. And one of those tools is to call another agent to go do it on some other specific task or subset of the problem. Uh, and then also, I hope we settle a little bit of this drama around like, wait, like our, our LMS, um, actually new, aren't they just coding agents? Yes. Coding agents can be our LMS. They aren't automatically our LMS. Um,
SPEAKER_00
and so I've showed a couple of different cloud code dynamic workflows that can turn cloud code into an RLM, uh, as well as some ways of doing this with open pros that you can use with any coding agent. So I see this as an incredibly powerful way of working with coding agents. I hope you will dig in more to recursive coding agents.com, but with great power comes great responsibility. So until next time, please recurse responsibly. Thank you very much.