SPEAKER_00
Hello and welcome to this online track talk for the AI Engineer World's Fair 2026.
SPEAKER_00
Today we are going to explore the concept of RLM, also known as recursive language models, and how we can use those concepts for larger code bases. My name is Shashi, I am a founder of Superagentic AI. First of all, let's be clear that the RLM paper has been published by MIT and Friends. As you can see, there's a full paper you can read about it. But the purpose of this talk is how you can use the concepts of RLM in your own workflow to implement your own harnesses. So first of all, what's the problem? If you're using coding agents for smaller repos or mono repos, they work exceptionally well. But if you have ever tried it with large mono repos with large context, there's a context problem. As the context grows, the performance degrades. And if you're working with mono repos, this problem gets worse. In this talk, we will see how the selected code base and the concept of RLMs are relevant for larger code bases. If you use coding agents, then you probably saw that there are different approaches that other coding agent harnesses have taken to solve this problem. The most common approach is searching using tools like grep. So there's a file system and the coding agent harnesses search using these tools. The second approach, you've probably seen, is semantic search or local search. So the idea here is you can search through the code and curate the context. Another approach is long context gets compressed and you can use the summarized version of the context. And there are some memory solutions available in the market as well that you can use to persist the memory for the coding agent. First of all, let's explore the RLM idea. The core thesis of RLM is you need to externalize the context management into a programmable execution environment. Meaning you should have a separate dedicated environment so that the model can operate on that. In this case, for example, your whole repository is treated as data that the model can operate on. Then the model can write code to inspect, slice, and compute the relevant chunks. The value you can then feed into the main context window. So rather than putting everything into the model's context, create a separate dedicated environment, give the model a coding agent or REPL, and then the model writes code to curate the context that can be used in the main. So it's another context management technique proved to be very effective. Could also be used as a memory layer for your coding agents. Let me summarize this by giving you a simple analogy. Imagine you are a lead software engineer assigned to a new project with a huge code base. Imagine that's a monorepo. How does that lead engineer deal with the code? So rather than reading each line of code line by line, the engineer probably inspects the code base, makes some notes, sees what the project's dependencies are, and how it is structured. Maybe something in the repository is not understood by the engineer. The engineer probably asks another engineer or expert to get some ideas. And the same concept is applied in RLM. So large projects like files, docs, texts, and configs, because the repository has a lot of things. And the programmable REPL is a notebook that the engineer makes notes about the code base that can be used. While researching, the engineer may be using other techniques or writing some script to search something from the repo. And then if stuck, the engineer asks another engineer or specialist. That's where the LLM query comes in. And the LLM query is basically asking another model environment to get an answer from. And once they get an answer, the loop continues. And at the end, it returns the clean note synthesis. So the recursion part here is the engineer asks another specialist using LLM query. That can be one question or that can be a number of questions. So this is where the recursion comes into the picture. The loop is basically your repo as your context. And then the model writes the REPL code to get some relevant context. That returns the bounded observation. And if the loop needs more information, it passes through the LLM query where it asks another language model or another system to get the response, returns the value, and continues the loop. And the loop gets terminated until we get our final results. So the code base is different. It has directories. It has tests. It has some imports. It has dependencies. It has pictures. It has configuration files. So the code base is not only just the text. It is structured data. And the model needs to understand and reason over the text. That's why I chose this scenario to use code basis to prove these concepts of RLM. Now let's switch gears and talk about our own library that we created at SuperAgentic AI called RLM code. You can see RLM code's landing page here. This is a research playground where you can implement the concepts of RLM. We have documentation that you can take a look at, and there's the GitHub repository. It is a completely open source project that you can use and play with. RLM itself is a concept and a pattern. And you can implement that concept and pattern in your own way. The official authors also wrote some implementations in their GitHub repo. It's called RLM and RLM minimal. You can refer to that implementation of RLM in dspy.rlm. Omar is the author of RLM and he is also the author of another popular framework called dspy. So dspy has an RLM implementation inside it. However, you should treat them as completely different. RLM is a pattern and you can implement it in your own ways. You can find that various other people have implemented RLM in their own way. And similarly, we implemented RLM code as our own independent harness that we will be using in this live demo. RLM code is a reference implementation to demonstrate how the RLM concepts work under the hood. So we have implemented something called RLM mode. We are using RLM as it is. We are not adding anything on top of RLM's ideas and RLM's paper. We are using the same concept of recursive calls and REPL execution. However, you can run it with a local model. You can run it with a cloud-based model. You can plug it into any observability framework of your choice. And that gives you a lot of flexibility around RLM. You can also plug it into the framework of your choice. For example, you can use Pydantic AI or Google ADK or something similar and implement ideas of RLM over there. In order to demonstrate this, we have created a source code repository where you can try this concept by yourself using MIT's RLM paper and RLM code. And we will see how these things work in practice. So basically, we will show you the loop. You will understand this once we see this live demo and what all these files are doing. Where the context has been created, where the Python REPL has written code, and where it's passed to the LLM query, and how we get the final results. So that will be covered as part of the live demo. So we can cover everything here. So in a nutshell, how it looks like is basically it creates the REPL and then observation and the final recursive language output. So let me jump into the live demo now. Okay, let's do the live demo of these concepts of RLM and RLM code and how it works in larger code bases. So I have a code repository here I have checked out and let's open it into the editor so that we can see what's inside it. So as you can see, there's a demo target which is using RLM code source as a demo here, and then we have some instructions that you can follow along yourself. So basically you have a readme file that you can use with your local model or we are going to use with Gemini. So we will try this script and see what happens. So right now you can see we're using Docker as a sandbox. If you see the Docker container has just been started for this RLM. And now coming back to our execution, you can see that the execution just finished. And in this execution, what you have seen, basically, in the first step the model has written the REPL code that you can see here, and then it built the evidence. And after that it also made calls to the LLM query with some prompt and got the result back. And after that it gives the final answer. And as you can see here, we have all these steps coming back to the final answer. And here you can see it made two tool calls and how many tokens are used for this model, that you can see here. And the good thing is that you can see all these traces in the RLM code repositories. So for example, you can see all the runs. This is the run that we just did. You can see all the sessions and all the observability that you can plug into any of your favorite observability platforms. So this is the CLI path we just demonstrated, but we also have this kind of coding agent style experimental harness where you can try the same thing. So first of all, let's connect with the Gemini model. So you can connect with the Gemini model using the command connect and you can have a provider and the model name. So you can also run the doctor command and see if everything is okay. Seems like the doctor command found some warnings, but this is related to deep agent ADK and other frameworks, which is not relevant to this demo. And the interesting part where we will be seeing is basically you are sending the prompt. So for what we did now, we ran the command and we asked a question. We specify the budget so that we don't spend too much on this run. But once we do that, as you can see, you have the maximum steps recursion depth and it completed this run and coming back with the results. We can also see this in the research lab where we can see this spin has been completed. We can see some rewards. We can also see the trajectory, which is the important part where we can see all the RLM loop. For example, the REPL code and the final output. So and also we can see the events when it started and when it ended. So you can play around with this RLM code terminal user interface, which is a harness, and you can experiment with your RLM ideas in here. So I'm going to quit this for now and let's switch back to the slides. In a nutshell, what we just saw is basically our context has been loaded. We have some REPL code written to extract some snippets. We also saw the LLM query has been called to get some more context from another model. This is where the recursion comes into the picture. And we got the final result and we got the traces in JSONL format that you can import into any of the observability platforms of your choice. We also saw these results coming from different files. You can take a look at the source code that will be available for you. Let's talk about the real thing: how an AI engineer could use these concepts in real life. And there are several things. For example, if you're dealing with large source code and you want to, for example, root cause analysis or onboarding of repositories or some unfamiliar repos, so there are several use cases you can try from here and probably try to use RLM concepts over there. Basically, you can design your own harness based on your needs so that it should capture the whole trajectory, all these things like the planning, coding, observation, sub-call budget, and the final output. Now coming back to the final point about RLM concepts and where it's being used, I have recently come across a lot of posts on X saying the RLM concepts have been used in some proprietary things like managed agent dynamic workloads using the RLM concepts under the hood. So they have implemented one or more forms of RLM inside their agent harnesses. Recently I saw that the Codex harness is writing Python code in the REPL that you can see to curate the context, that is one form of RLM I have seen myself. And obviously the cloud managed agents or Gemini managed agents, they're all kind of concepts of RLM. So basically you can get the harness in the sandbox and then you can do the stub. And the recent things about dynamic workflows where one agent given a task can spawn multiple agents that have their separate sandboxes, they can work together, and give back the final results. And the idea is basically generally coming from RLMs. A lot of software factories concepts are probably using RLMs, but we are not sure yet. However, some cloud code engineers from Anthropic acknowledged on X that they have used concepts of RLM. You can use this RLM concept on your large context repository. And if you have any questions, feel free to reach out to me. And finally, thank you so much for listening to my talk.
SPEAKER_00
we can use those concepts for larger code bases. My name is Shashi, I am a founder of Superagentic R. First of all, let's be clear that RLM paper has been published by MIT and Friends. As you can see there's a full paper, you can read about it. But the purpose of this talk is how you can use the concepts of RLM and you can use into your own workflow to implement your own harnesses. So, first of all, what's the problem? If you're using the coding agents for smaller repos or mono repos, they work exceptionally well. But if you have ever tried it with the mono repos, with the large context, you know there's a context problem. As the context grows,
SPEAKER_00
the performance degrades. And if you're working with the mono repos, this problem gets worse. In this talk, we will see we selected the code base and the concept of RLMs are relevant for the larger code bases. If you use the coding agents, then you probably saw that there are different approaches that other coding agent harnesses have been taken to solve this problem. The most common approach is searching using the tools like grep. So basically there's a file system and the coding agent harnesses search using these tools. The second approach, you've probably seen that the semantic search
SPEAKER_00
search or the local search. So idea here is basically you can search through the code and curate the context. Another approach is the long context get compressed and you can use the summarized version of the context. And there are some memory solutions available in the market as well that you can use to persist the memory for the coding agent. First of all, let's explore the RLM idea. The core thesis of the RLM is you need to externalize the context management into programmable execution environment. Meaning you should have a separate dedicated environment so that model can operate on that.
SPEAKER_00
In this case, for example, your whole repository is treated as a data that model can operate on. Then model can write the code to inspect, slice and compute the relevant chunks. The value you can then feed into the main context window. So basically rather than putting everything into the model's context, create a separate dedicated environment, give them a coding agent or REPL, model and then model write the code to curate the context that can be used into the main. So it's another context management technique proved to be very effective. Could be also be used as a memory layer for your coding agents.
SPEAKER_00
Let me summarize this giving you a simple analogy. Imagine you are a lead software engineer and assigned to the new project with a huge code base. Imagine that's a monorepo. How does that lead engineer deals with the code? So rather than reading each line of code line by line, engineer probably inspect the code base, make some notes, see what are the project's dependencies, how it is structured. Maybe something else is not understood by the repository. Engineer probably asks to another engineer or expert to get some ideas. And the same concept is applied in the RLM. So large project like the files and docs and
SPEAKER_00
texts and configs because the repository has a lot of things. And the programmable REPL is kind of a notebook that engineer makes a note about the code base that can be used. Researching, he may be using other techniques, or maybe writing some script to search something from the repo. And then if he stugs, then he asks another engineer or specialist where it comes to the LLM query. And LLM query is basically asking another model environment to get an answer from. And once they get answer, then the loop continues. And at the end, it returns the clean note synthesis. So the recursion part here is engineer asks another
SPEAKER_00
specialist using LLM query. That can be one question or that can be number of questions. So this is where the recursion comes in picture. The loop is basically your repo as your context. And then the model writes the REPL code to get some relevant context. That returns the bounded observation. And if loop needs more information, it passes through the LLM query where it asks another language model or another system to get the response, return the value, and continue the loop. And the loop gets terminated until we get our final results.
SPEAKER_00
So the code base is different. It has directories. It has tests. It has some imports. It has dependencies. It has tests. It has pictures. It has configuration files. So the code base is not only just the text. It is a structured data. It has a structured data. And the model needs to understand and reason over the text. That's why I chose this scenario to use the code basis to prove these concepts of RLM. Now let's switch the gear and talk about our own library that we created at SuperAgentic AI called RLM code. You can see RLM code's landing page here. This is just a research playground where you can implement the concepts of RLM.
SPEAKER_00
We have documentation that you can take a look and there's the GitHub repository. It is completely open source project that you can use it and play with it. RLM itself is a concept and a pattern. And you can implement that concept and pattern in your own way. There are official authors also wrote some implementation in their GitHub repo. It's called RLM and RLM minimal. You can refer that implementation of RLM in dspy.rlm. So Omar is author of RLM and he is also author of another popular framework called dspy. So dspy got RLM implementation inside it. However, you should treat they are completely different.
SPEAKER_00
So RLM is a pattern and you can implement in your own ways. You can find there are various other people implemented RLM in their own way. And in the similar way, we implemented RLM code as our own independent harness that we will be using in this live demo. RLM code is just a reference implementation to demonstrate how the RLM concepts works under the hood. So we have implemented something called RLM mode. We are using RLM as it is. We are not adding anything on top of RLM's ideas and RLM's paper. We are using the same concept of recursive calls, REPL execution. However, you can run it with a local model.
SPEAKER_00
You can run it with a cloud-based model. You can plug into any observability framework of your choice. And that gives you like a lot of flexibility around RLM. You can also plug it into the framework of your choice. For example, you can use PyDentic AI or Google ADK or something similar framework and implement ideas of RLM over there. In order to demonstrate this, we have created a source code repository where you can try this concept by yourself using MIT's RLM paper and RLM code. And we will see how these things work in a practice. So basically, we will show you the loop. This, you will understand this once we see this live demo and what all these files are doing.
SPEAKER_00
Where, where's the context has been created, where the Python REPL has written a code, and where it's passed to the LLM query, and how we get the final results. So that will be covered as part of the live demo. So we can cover this, everything here. So in nutshell, how it looks like is basically, it creates the REPL and then observation and the final recursive language output. so let me jump into the live demo now okay let's do the live demo of these concepts of rlm and rlm code and how it works in a the larger code basis so i have a code repositories here i have checked out and let's open it into the editor so that we can see what's inside it so as you can see
SPEAKER_00
there's a demo target which is um we are using rlm code source as a as a demo here and then we have some instructions that you can follow along um yourself so basically you have a readme file that you can use to use with your local model or so we are going to use with the gemini so we will try this script and see what happened so right now you can see we're using the docker as a sandbox if you see the docker container has been just started for this rlm and now coming back to our execution you can see that execution i just finished and in this execution what you have seen basically in the first step model has written the
SPEAKER_00
repl code that you can see here and then it's built the evidence and after that it also made the calls to the llm query with some prompt and got the result back and after that it gives the final answer and as you can see here we can have all this all the step coming back to the the final answer and here you can see the it made the two tool calls and how many the tokens is used for this model that you can see it here and the good thing is that you can see all these traces in the rlm code repositories so for example you can see all the runs this is the run that we just did you can see all the sessions and all the observability that you can plug it into any of your
SPEAKER_00
observed favorite observability platform so this is the cli path we just demonstrated but we also have this kind of coding agent style experimental harness where you can try the same thing so first of all let's connect with the gemini model so you can connect with the gemini model using the command connect and you can have a provider and the model name so you can also run the doctor command and see if everything is okay seems like doctor command found some warning but this is related to deep agent adk and other frameworks which is not relevant to this demo and the interesting part where we will be saying is basically you are
SPEAKER_00
sending the prompt so for what we did now we ran the command and we ask the question we specify the budget so that we don't know spending too much uh on this run but once we do that as you can see you have the maximum steps recursion depth and it completed its this run and coming back with the results we can also see that this thing into the research lab where we can see this spin has been completed we can see some rewards we can also see the trajectory which is important part where we can see the all the rlm loop for example the the ripple and the code and the final output so and also we can see the the events
SPEAKER_00
when it started and when it ended so you can play around with this rlm code terminal user interface which is kind of harness and you can experiment your rlm ideas in here so i'm going to quit this for now and let's switch back to the slides in a nutshell what we just saw basically our context has been loaded we have some repl code written to extract some snippets we also saw the llm query has been called to get some more context from another model this is where the recursion comes in picture and we got the final result and we got the the traces in json l format that you can import it into the any of the observability platform of your choice
SPEAKER_00
we also saw these results coming from different piles you can take a look at the source code that will be available uh for you let's talk about the real thing how ai engineer could use this concepts in the real life and there are few things for example if you're dealing with this large source code and you want to for example root cause analysis or onboarding of the repositories or some unfamiliar repos so there are few use cases you can from here and probably try to use rlm concepts over there basically you can design your own harness um based on your needs so that should capture the whole trajectory all these things like
SPEAKER_00
the planning coding observation sub call budget and the final output now coming back to the final point about rlm concepts and where it's been used i have recently came across a lot of the post on x saying the rlm concepts have been being used into the some of the proprietary things like the managed agent dynamic workloads using the rlm concepts under the hood so they have implemented one or more forms of rlm inside their agent harnesses recently i saw that the codex harness is writing the python python code in the ripple that you can see to curate the context that is the one form of rlm i have seen myself and
SPEAKER_00
obviously the clouds manage agents or gemini managed agents they're all kind of concepts of rlm so basically you can get the harness in the sandbox and then you can do the stub and the recent things about the dynamic workflows where one agent given that given the task you can spawn multiple agents that have their separate sandboxes they can work together and give back the final results and the idea is basically generally coming from the um rlms a lot of software factories concepts are probably using the rlms but we are not sure yet however some of the cloud code engineers from anthropicas accepted on x that
SPEAKER_00
they have used concepts of rlm you can use this rlm concept on your large context repository and if you have any questions then feel free to reach out to me and finally thank you so much for listening to my talk