RLM: Recursive Language Models for Large Codebases - Shashi, Superagentic AI
Description
Large codebases break coding agents: they lose the architecture and drown in tool output as context grows. This talk introduces Recursive Language Models (RLM) from a MIT paper a pattern that loads the repo into a programmable REPL where the model writes code to inspect it and recursively delegates focused sub-questions via llm_query. With a live demo on RLM Code (independent, unofficial), you'll see the loop run end to end on local and cloud models, with a fully inspectable trajectory. Speakers: - Shashi (Superagentic AI): Building tools and frameworks for AI Agents X/Twitter: https://x.com/Shashikant86 LinkedIn: https://www.linkedin.com/in/shashikantjagtap/ GitHub: https://github.com/Shashikant86
Summary
Generated by claude-sonnet-4-5At-a-Glance
- Verdict: Watch fully
- Core thesis: Recursive Language Models (RLM) solve large codebase context problems by externalizing context management into programmable REPL environments where models write code to curate their own context, enabling recursive sub-queries to specialized models instead of cramming everything into main context windows.
- Why it matters: If you build or operate coding agents at scale (especially mono repos), current approaches (grep, semantic search, compression) degrade with size—RLM demonstrates a fundamentally different pattern where the agent programs its own context extraction loop, potentially unlocking better performance and observability for large-context AI workflows.
- Best use: Watch fully to understand the RLM loop mechanics and open-source RLM-code implementation; reference the demo sections for practical integration patterns into agent harnesses; use the conceptual framework (engineer analogy, REPL + recursion) when designing your own context-management systems for agentic workflows.
Executive Summary
Shashi (founder of Superagentic AI) presents Recursive Language Models (RLM)—a pattern published by MIT—and demonstrates RLM-code, an open-source reference implementation. The core problem: coding agents work well on small repos but degrade on large mono repos due to context limits. Traditional solutions (grep search, semantic search, long-context compression, memory layers) help but don't fully solve the structural data and reasoning challenges of huge codebases with directories, configs, dependencies, and mixed file types.
The RLM solution: externalize context management into a dedicated programmable execution environment (like a Jupyter REPL). Instead of feeding the entire repo into the LLM's context window, the model writes Python code to inspect, slice, and extract relevant chunks. When the model needs deeper understanding, it recursively calls another LLM (an 'LLM query') with a focused sub-question, gets an answer, and continues. This mirrors how a lead engineer onboards to a new codebase: make notes in a notebook, write scripts to explore structure, and ask specialists when stuck. The loop terminates when the curated context and observations are sufficient.
Shashi demos RLM-code against its own source as the target repository. The model writes REPL code, builds evidence, makes recursive LLM query calls, and returns the final answer with full trajectory logs (JSONL traces). The tool supports local or cloud models (Gemini demo shown), Docker sandboxing, and pluggable observability. He emphasizes RLM is a pattern—not tied to one framework—and can be integrated into PyDantic AI, Google ADK, or custom harnesses. Real-world adoption signals: Anthropic's Claude Code (acknowledged using RLM concepts), Gemini managed agents, and 'dynamic workloads' where one agent spawns multiple sub-agents in separate sandboxes.
Practical use cases include root-cause analysis, onboarding to unfamiliar repos, and any scenario requiring structured reasoning over large codebases. The talk encourages AI engineers to design custom harnesses capturing the full trajectory (planning, coding, observation, sub-call budget, final output) and to experiment with the open-source RLM-code repo for hands-on learning.
Key Takeaways
- Claim: RLM externalizes context management: instead of cramming the entire repo into the LLM's context window, create a dedicated programmable REPL environment where the model writes code to curate context. | Evidence: Demo shows the model writing Python REPL code to inspect RLM-code's own source, extracting snippets and building evidence before feeding curated chunks into the main context. The official RLM paper (MIT) and implementations (dspy.rlm, RLM-minimal) formalize this pattern. | Caveat: RLM is a pattern/concept, not a turnkey product; requires integration work, sandbox setup (Docker), and careful trajectory/observability design. Token budgets and recursion depth need tuning to avoid runaway costs. | Implication: For Ken's agent systems operating on large codebases or structured data: adopt RLM-style REPL loops to prevent context-window degradation, improve observability (full trajectory logs), and enable recursive sub-agent calls for specialized tasks. RLM-code is a working open-source reference to fork/customize. | Timestamp: 00:00 to 03:30
- Claim: Recursion comes from 'LLM query' calls: when the model's REPL-based exploration hits a knowledge gap, it invokes another LLM environment with a focused sub-question, gets an answer, and continues the loop. | Evidence: Demo trace shows two LLM query tool calls mid-execution, each with a prompt and response. Analogy: a lead engineer asks a specialist colleague when stuck. The loop terminates once sufficient observations are gathered. | Caveat: No explicit discussion of how to prevent infinite recursion or budget blow-out; speaker mentions 'budget' flag in CLI but doesn't detail failsafes. Recursion depth tuning is critical. | Implication: Ken can design hierarchical agent systems where parent agents spawn specialist sub-agents (similar to dynamic workloads) and budget token usage per sub-call. Full observability (JSONL traces) becomes essential for cost control and debugging. | Timestamp: 06:20 to 07:40
- Claim: Codebases are structured data (directories, configs, dependencies, images), not flat text—models must reason over structure, making RLM's programmable approach superior to naive grep or semantic search. | Evidence: Speaker lists code base components: directories, tests, imports, dependencies, configs, pictures. RLM lets the model write scripts to navigate and slice this structure programmatically, whereas grep/semantic search treat it as unstructured text. | Caveat: No quantitative benchmark comparing RLM vs. semantic search or compression on large repos. Claims are conceptual and demo-based, not empirical performance data. | Implication: For Ken's content/workflow systems: if operating on mixed-media repos (docs, code, configs, assets), RLM's structured inspection via code beats keyword search. Consider RLM patterns when building RAG pipelines for non-text-heavy corpora. | Timestamp: 08:10 to 09:00
- Claim: RLM-code is open-source, model-agnostic, and pluggable: works with local or cloud models (Gemini, etc.), integrates with any observability platform, and can be embedded in frameworks like PyDantic AI or Google ADK. | Evidence: Demo connects to Gemini via CLI, runs in Docker sandbox, exports JSONL traces. GitHub repo and docs provided. Speaker emphasizes it's a 'reference implementation' to demonstrate RLM mechanics, not a proprietary lock-in. | Caveat: Experimental harness; CLI and TUI shown are research-grade, not production-hardened. Warnings appear in 'doctor' command (related to deep-agent ADK dependencies), though speaker dismisses them as irrelevant to core demo. | Implication: Ken can fork RLM-code, swap in preferred LLMs (local or cloud), and integrate with existing observability/cost-tracking stacks. Use it as a learning sandbox to understand RLM loop mechanics before building custom harnesses for production workflows. | Timestamp: 10:00 to 14:30
- Claim: Real-world adoption signals: Anthropic's Claude Code acknowledged using RLM concepts, Gemini managed agents use RLM patterns, and 'dynamic workloads' (one agent spawning multiple sub-agents in separate sandboxes) are inspired by RLM. | Evidence: Speaker cites X posts and acknowledgments from Anthropic engineers. Mentions Codex harness writing Python in REPL to curate context as a visible RLM form. Cloud managed agents and software factories likely use RLM under the hood. | Caveat: No official public statements or technical deep-dives from Anthropic/Google cited; claims are based on social media posts and speaker's observation. 'Probably using RLMs' is speculative for some vendors. | Implication: Ken should track RLM-inspired architectures in leading agentic products (Claude Code, Gemini agents) as signals of emerging best practices. If competitors adopt RLM patterns, consider it a validated approach worth integrating into proprietary agent systems. | Timestamp: 16:00 to 17:20
- Claim: Use cases beyond coding: root-cause analysis, onboarding to unfamiliar repos, and any large-context structured data scenario where you need the model to 'explore and take notes' programmatically. | Evidence: Speaker lists specific use cases: large source code repos, mono repo onboarding, unfamiliar codebases. The engineer analogy (inspect, note, script, ask) generalizes to other domains with structured assets. | Caveat: No non-code examples shown; demo is purely codebase-focused. Applicability to other domains (e.g., multi-file datasets, config forests, doc repos) is plausible but unproven in this talk. | Implication: For Ken: test RLM patterns on non-code structured data (e.g., multi-file campaign assets, API schema repos, internal doc wikis). If the REPL+recursion pattern holds, it could generalize beyond engineering workflows into content ops and knowledge management. | Timestamp: 15:30 to 16:00
Detailed Brief
Problem: Coding Agents Degrade on Large Mono Repos
- Claims: Coding agents work well on small/single repos but fail on large mono repos due to context-window limits; Existing solutions (grep search, semantic search, long-context compression, memory layers) help but don't fully solve the problem; Context growth causes performance degradation; mono repos amplify this issue
- Evidence: Speaker states 'if you have ever tried it with large mono repos with large context, there's a context problem. As the context grows, the performance degrades.'; Lists common approaches: grep tools, semantic/local search, compression, memory solutions; No quantitative benchmarks provided; claims are experiential and qualitative
- Caveats: No direct comparison of RLM vs. existing approaches on standardized benchmarks (e.g., SWE-bench); Grep/semantic search may still be faster for narrow queries; RLM trades latency for comprehensive context curation; Speaker doesn't discuss cost implications (token usage) of RLM vs. simpler methods
- Implications: Ken's agent systems should adopt RLM-style patterns if operating on large, heterogeneous codebases where grep/semantic search falls short; Budget RLM recursion carefully; full trajectory logs are critical for cost/performance tuning; Consider hybrid approach: use semantic search for initial candidate retrieval, then RLM REPL for deep inspection of top candidates
RLM Core Mechanics: REPL + Recursion Loop
- Claims: Core thesis: externalize context management into a programmable execution environment (REPL); Model writes code (Python) to inspect, slice, and compute relevant chunks from the repo; When more info is needed, model makes 'LLM query' calls (recursive invocations to another model environment); Loop terminates when final synthesized note/answer is ready; bounded observation prevents runaway
- Evidence: Demo trace shows: (1) REPL code written, (2) evidence built, (3) two LLM query calls with prompts/responses, (4) final answer returned; Analogy: lead engineer inspects codebase, makes notes in a notebook, writes scripts, asks specialist colleagues when stuck; Docker sandbox spins up for isolated execution; JSONL traces capture full trajectory (planning, coding, observation, sub-calls, output)
- Caveats: Recursion depth and token budget must be tuned; speaker shows budget flag in CLI but doesn't detail failsafes or auto-cutoff logic; No discussion of how to handle REPL errors or partial observations; error recovery not shown; Single-threaded loop in demo; unclear if parallelization (multiple REPL sessions) is supported or advisable
- Implications: Ken can architect agent systems with hierarchical task decomposition: parent agent spawns REPL sub-agents for specialized slicing, then aggregates results; Observability is paramount: full trajectory logs enable post-hoc analysis, cost attribution, and iterative prompt/harness tuning; Consider RLM as a memory layer: curated context persists across sessions, avoiding re-inspection of unchanged code
RLM-code: Open-Source Reference Implementation
- Claims: RLM-code is a research playground/reference implementation, not a proprietary product; Model-agnostic: works with local models or cloud APIs (Gemini demo, but pluggable); Integrates with any observability framework; exports JSONL traces for custom tooling; Can be embedded in frameworks like PyDantic AI, Google ADK, or custom harnesses
- Evidence: GitHub repo and documentation available; speaker emphasizes 'completely open source'; Demo shows CLI (
rlm-code run) and experimental TUI harness (rlm-code connect gemini); Docker sandbox for REPL isolation; supportsdoctorcommand for health checks (warnings about deep-agent ADK dependencies, but not critical); Trajectory/session data stored locally; can import into favorite observability platforms - Caveats: Experimental/research-grade quality; CLI warnings suggest incomplete integration with some dependencies; No production SLA, error handling, or enterprise features mentioned; Speaker treats it as a learning tool, not a drop-in replacement for production coding agents
- Implications: Ken should fork/clone RLM-code as a sandbox to prototype RLM patterns in existing agent workflows; Swap in preferred LLMs (OpenAI, Anthropic, local Llama models) to test cost/performance trade-offs; Use JSONL traces to feed Ken's observability stack (Langfuse, LangSmith, etc.) and track token/cost attribution per sub-call; If building custom harness, reference RLM-code's architecture (REPL manager, LLM query router, trajectory logger) as design blueprint
Real-World Adoption and Industry Signals
- Claims: Anthropic's Claude Code acknowledged using RLM concepts (via X posts); Gemini managed agents and cloud code engineers use RLM patterns; Codex harness writes Python in REPL to curate context—visible RLM form; Dynamic workloads (agent spawning sub-agents in separate sandboxes) inspired by RLM; Software factories 'probably' using RLM, but not confirmed
- Evidence: Speaker cites X posts and engineer acknowledgments; no direct links or technical papers provided; Observation of Codex harness behavior; speaker claims to have 'seen myself'; Managed agent products (Google, Anthropic) described as using 'concepts of RLM under the hood'
- Caveats: No official public documentation or deep-dives from vendors cited; claims based on social media and speaker's analysis; Speculative for some vendors ('probably using RLMs' for software factories); Lack of direct confirmation or architectural details limits verifiability
- Implications: Ken should monitor leading agentic products (Claude Code, Gemini Code Assist, Devin-style tools) for RLM-inspired architectures; If top players adopt RLM patterns, it signals a validated approach worth integrating into proprietary systems; Track academic/industry discourse on RLM (MIT paper follow-ups, conference talks) for emerging best practices and benchmarks; Consider RLM as a competitive differentiator if peers have not yet adopted programmable context management
Practical Use Cases and Next Steps for AI Engineers
- Claims: Use cases: large source code analysis, mono repo onboarding, unfamiliar repo exploration, root-cause analysis; Design custom harnesses capturing full trajectory: planning, coding, observation, sub-call budget, final output; Experiment with RLM-code repo to understand loop mechanics hands-on
- Evidence: Speaker lists specific scenarios where RLM shines: 'dealing with large source code,' 'onboarding of repositories,' 'unfamiliar repos'; Demo shows end-to-end loop from context load → REPL code → LLM query → final answer → JSONL traces; Source code and demo target (RLM-code's own repo) available for self-experimentation
- Caveats: No non-code examples; applicability to other structured data domains (e.g., multi-file datasets, config repos) is plausible but unproven; Speaker doesn't provide step-by-step integration guide for existing agent frameworks; Cost and latency implications for production use cases not quantified
- Implications: Ken should prototype RLM patterns on high-value use cases: onboarding engineers to complex repos, automating root-cause analysis for incidents, or curating context for code-review agents; Test RLM on non-code structured data (e.g., multi-file campaign assets, API schema repos) to validate generalizability; Use RLM-code's observability outputs to build dashboards for agent performance, cost per recursion depth, and sub-call efficiency; Consider RLM as a building block for multi-agent orchestration: parent agent delegates sub-tasks to REPL-equipped child agents, then aggregates results
Notable Concepts & Terms
- RLM (Recursive Language Models): Pattern published by MIT for externalizing context management into programmable REPL environments where models write code to curate their own context and recursively invoke sub-models for deeper reasoning; fundamentally different from cramming everything into a single context window.
- LLM Query: Recursive call within RLM loop where the model invokes another LLM environment with a focused sub-question to fill knowledge gaps; analogous to asking a specialist colleague; enables hierarchical reasoning and prevents context overload.
- REPL (Read-Eval-Print Loop): Interactive programming environment (here, Python) where the model writes and executes code to inspect, slice, and extract relevant context from structured data (e.g., codebases); core execution substrate for RLM.
- RLM-code: Open-source reference implementation by Superagentic AI demonstrating RLM concepts; model-agnostic, Docker-sandboxed, exports JSONL traces, and can be plugged into frameworks like PyDantic AI or Google ADK.
- Bounded Observation: Output of each REPL execution step in the RLM loop; curated context chunks returned to the main model before the next recursion or final synthesis; prevents runaway context growth.
- Dynamic Workloads: Agentic pattern where one agent spawns multiple sub-agents in separate sandboxes to work in parallel, then aggregates results; inspired by RLM's recursive and programmable execution model.
- Trajectory Logging (JSONL traces): Full observability capture in RLM-code: planning, REPL code, observations, LLM query calls, token usage, and final output; enables cost attribution, debugging, and iterative harness tuning.
- dspy.rlm: RLM implementation within the DSPy framework (authored by Omar Khattab, also author of RLM paper); treats RLM as a module in DSPy's prompt optimization pipeline; separate from RLM-code but shares core pattern.
Operator Notes / Why Ken Should Care
- For Ken's agent systems: if operating on large mono repos or multi-file structured data, RLM's programmable context curation via REPL can prevent context-window degradation that kills grep/semantic search approaches at scale.
- RLM enables hierarchical multi-agent orchestration: parent agent delegates sub-tasks to REPL-equipped child agents (each with bounded context), then aggregates results—critical for dynamic workflow automation and software factories.
- Full trajectory observability (JSONL traces) is a forcing function for cost control and performance tuning; integrate RLM-code traces into Ken's observability stack (Langfuse, LangSmith) to track token/cost per recursion depth and sub-call.
- RLM as memory layer: curated context can persist across sessions, avoiding re-inspection of unchanged code—useful for long-running coding agents or incremental codebase updates.
- Industry signals (Claude Code, Gemini agents, Codex) suggest RLM patterns are being adopted by leading agentic products; early adoption could be a competitive differentiator in Ken's product/GTM strategy.
- RLM-code is open-source and model-agnostic; fork it to prototype RLM patterns in existing workflows, swap in preferred LLMs, and customize trajectory capture for domain-specific needs (e.g., content ops, config management).
- Test RLM on non-code use cases (multi-file campaign assets, doc wikis, API schema repos) to validate generalizability beyond engineering workflows; the REPL+recursion pattern may unlock structured reasoning in other domains.
- Budget RLM recursion carefully: token usage scales with depth and sub-call count; use budget flags and observability to prevent runaway costs in production.
- For AI investing/GTM: RLM is an emerging pattern with academic backing (MIT) and early adoption signals from top labs (Anthropic, Google)—track follow-on research and product integrations as indicators of long-term viability.
Watch Map
- 00:00: Introduction: Shashi from Superagentic AI introduces RLM (MIT paper) and talk scope—using RLM concepts for large codebases.
- 01:00: Problem statement: Coding agents degrade on large mono repos due to context limits; existing solutions (grep, semantic search, compression) are insufficient.
- 03:00: RLM core thesis: externalize context management into programmable REPL environment; model writes code to curate context instead of cramming everything into context window.
- 05:00: Engineer analogy: lead engineer inspects codebase, makes notes, writes scripts, asks specialists when stuck—same pattern as RLM loop (REPL + LLM query).
- 07:00: RLM loop mechanics: repo as context → model writes REPL code → bounded observation → LLM query for recursion → loop terminates with final synthesis.
- 08:00: Why codebases? Structured data (directories, configs, dependencies, images) requires reasoning beyond flat text; RLM's programmable approach is superior.
- 10:00: RLM-code introduction: open-source reference implementation by Superagentic AI; model-agnostic, Docker-sandboxed, pluggable observability.
- 11:00: RLM as pattern, not product: can be implemented in dspy.rlm, custom harnesses, or frameworks like PyDantic AI/Google ADK.
- 12:00: Live demo setup: RLM-code repo, Docker sandbox, Gemini model connection, demo target is RLM-code's own source.
- 13:00: Demo execution: model writes REPL code → builds evidence → two LLM query calls → final answer. JSONL traces shown.
- 14:00: Observability: runs/sessions viewable in research lab; traces can be imported into any observability platform.
- 15:00: Experimental TUI harness: connect to model, run doctor, set budget, execute prompt, view trajectory/rewards/events in research lab.
- 16:00: Use cases: large source code, mono repo onboarding, root-cause analysis, unfamiliar repos. Design custom harnesses capturing full trajectory.
- 17:00: Industry adoption: Anthropic Claude Code, Gemini managed agents, Codex harness, dynamic workloads—all inspired by or using RLM concepts.
- 17:30: Closing: reach out with questions, thank you.
Source/Metadata
- Title: RLM: Recursive Language Models for Large Codebases - Shashi, Superagentic AI
- Transcript words: 4706
- Duration seconds: 1047
- Timestamp note: Timestamps estimated from transcript cues and demo flow; no hardcoded chapter markers provided.
Transcript
Hello and welcome to this online track talk for the AI Engineer World's Fair 2026. Today we are going to explore the concept of RLM, also known as recursive language models, and how we can use those concepts for larger code bases. My name is Shashi, I am a founder of Superagentic AI. First of all, let's be clear that the RLM paper has been published by MIT and Friends. As you can see, there's a full paper you can read about it. But the purpose of this talk is how you can use the concepts of RLM in your own workflow to implement your own harnesses. So first of all, what's the problem? If you're using coding agents for smaller repos or mono repos, they work exceptionally well. But if you have ever tried it with large mono repos with large context, there's a context problem. As the context grows, the performance degrades. And if you're working with mono repos, this problem gets worse. In this talk, we will see how the selected code base and the concept of RLMs are relevant for larger code bases. If you use coding agents, then you probably saw that there are different approaches that other coding agent harnesses have taken to solve this problem. The most common approach is searching using tools like grep. So there's a file system and the coding agent harnesses search using these tools. The second approach, you've probably seen, is semantic search or local search. So the idea here is you can search through the code and curate the context. Another approach is long context gets compressed and you can use the summarized version of the context. And there are some memory solutions available in the market as well that you can use to persist the memory for the coding agent. First of all, let's explore the RLM idea. The core thesis of RLM is you need to externalize the context management into a programmable execution environment. Meaning you should have a separate dedicated environment so that the model can operate on that. In this case, for example, your whole repository is treated as data that the model can operate on. Then the model can write code to inspect, slice, and compute the relevant chunks. The value you can then feed into the main context window. So rather than putting everything into the model's context, create a separate dedicated environment, give the model a coding agent or REPL, and then the model writes code to curate the context that can be used in the main. So it's another context management technique proved to be very effective. Could also be used as a memory layer for your coding agents. Let me summarize this by giving you a simple analogy. Imagine you are a lead software engineer assigned to a new project with a huge code base. Imagine that's a monorepo. How does that lead engineer deal with the code? So rather than reading each line of code line by line, the engineer probably inspects the code base, makes some notes, sees what the project's dependencies are, and how it is structured. Maybe something in the repository is not understood by the engineer. The engineer probably asks another engineer or expert to get some ideas. And the same concept is applied in RLM. So large projects like files, docs, texts, and configs, because the repository has a lot of things. And the programmable REPL is a notebook that the engineer makes notes about the code base that can be used. While researching, the engineer may be using other techniques or writing some script to search something from the repo. And then if stuck, the engineer asks another engineer or specialist. That's where the LLM query comes in. And the LLM query is basically asking another model environment to get an answer from. And once they get an answer, the loop continues. And at the end, it returns the clean note synthesis. So the recursion part here is the engineer asks another specialist using LLM query. That can be one question or that can be a number of questions. So this is where the recursion comes into the picture. The loop is basically your repo as your context. And then the model writes the REPL code to get some relevant context. That returns the bounded observation. And if the loop needs more information, it passes through the LLM query where it asks another language model or another system to get the response, returns the value, and continues the loop. And the loop gets terminated until we get our final results. So the code base is different. It has directories. It has tests. It has some imports. It has dependencies. It has pictures. It has configuration files. So the code base is not only just the text. It is structured data. And the model needs to understand and reason over the text. That's why I chose this scenario to use code basis to prove these concepts of RLM. Now let's switch gears and talk about our own library that we created at SuperAgentic AI called RLM code. You can see RLM code's landing page here. This is a research playground where you can implement the concepts of RLM. We have documentation that you can take a look at, and there's the GitHub repository. It is a completely open source project that you can use and play with. RLM itself is a concept and a pattern. And you can implement that concept and pattern in your own way. The official authors also wrote some implementations in their GitHub repo. It's called RLM and RLM minimal. You can refer to that implementation of RLM in dspy.rlm. Omar is the author of RLM and he is also the author of another popular framework called dspy. So dspy has an RLM implementation inside it. However, you should treat them as completely different. RLM is a pattern and you can implement it in your own ways. You can find that various other people have implemented RLM in their own way. And similarly, we implemented RLM code as our own independent harness that we will be using in this live demo. RLM code is a reference implementation to demonstrate how the RLM concepts work under the hood. So we have implemented something called RLM mode. We are using RLM as it is. We are not adding anything on top of RLM's ideas and RLM's paper. We are using the same concept of recursive calls and REPL execution. However, you can run it with a local model. You can run it with a cloud-based model. You can plug it into any observability framework of your choice. And that gives you a lot of flexibility around RLM. You can also plug it into the framework of your choice. For example, you can use Pydantic AI or Google ADK or something similar and implement ideas of RLM over there. In order to demonstrate this, we have created a source code repository where you can try this concept by yourself using MIT's RLM paper and RLM code. And we will see how these things work in practice. So basically, we will show you the loop. You will understand this once we see this live demo and what all these files are doing. Where the context has been created, where the Python REPL has written code, and where it's passed to the LLM query, and how we get the final results. So that will be covered as part of the live demo. So we can cover everything here. So in a nutshell, how it looks like is basically it creates the REPL and then observation and the final recursive language output. So let me jump into the live demo now. Okay, let's do the live demo of these concepts of RLM and RLM code and how it works in larger code bases. So I have a code repository here I have checked out and let's open it into the editor so that we can see what's inside it. So as you can see, there's a demo target which is using RLM code source as a demo here, and then we have some instructions that you can follow along yourself. So basically you have a readme file that you can use with your local model or we are going to use with Gemini. So we will try this script and see what happens. So right now you can see we're using Docker as a sandbox. If you see the Docker container has just been started for this RLM. And now coming back to our execution, you can see that the execution just finished. And in this execution, what you have seen, basically, in the first step the model has written the REPL code that you can see here, and then it built the evidence. And after that it also made calls to the LLM query with some prompt and got the result back. And after that it gives the final answer. And as you can see here, we have all these steps coming back to the final answer. And here you can see it made two tool calls and how many tokens are used for this model, that you can see here. And the good thing is that you can see all these traces in the RLM code repositories. So for example, you can see all the runs. This is the run that we just did. You can see all the sessions and all the observability that you can plug into any of your favorite observability platforms. So this is the CLI path we just demonstrated, but we also have this kind of coding agent style experimental harness where you can try the same thing. So first of all, let's connect with the Gemini model. So you can connect with the Gemini model using the command connect and you can have a provider and the model name. So you can also run the doctor command and see if everything is okay. Seems like the doctor command found some warnings, but this is related to deep agent ADK and other frameworks, which is not relevant to this demo. And the interesting part where we will be seeing is basically you are sending the prompt. So for what we did now, we ran the command and we asked a question. We specify the budget so that we don't spend too much on this run. But once we do that, as you can see, you have the maximum steps recursion depth and it completed this run and coming back with the results. We can also see this in the research lab where we can see this spin has been completed. We can see some rewards. We can also see the trajectory, which is the important part where we can see all the RLM loop. For example, the REPL code and the final output. So and also we can see the events when it started and when it ended. So you can play around with this RLM code terminal user interface, which is a harness, and you can experiment with your RLM ideas in here. So I'm going to quit this for now and let's switch back to the slides. In a nutshell, what we just saw is basically our context has been loaded. We have some REPL code written to extract some snippets. We also saw the LLM query has been called to get some more context from another model. This is where the recursion comes into the picture. And we got the final result and we got the traces in JSONL format that you can import into any of the observability platforms of your choice. We also saw these results coming from different files. You can take a look at the source code that will be available for you. Let's talk about the real thing: how an AI engineer could use these concepts in real life. And there are several things. For example, if you're dealing with large source code and you want to, for example, root cause analysis or onboarding of repositories or some unfamiliar repos, so there are several use cases you can try from here and probably try to use RLM concepts over there. Basically, you can design your own harness based on your needs so that it should capture the whole trajectory, all these things like the planning, coding, observation, sub-call budget, and the final output. Now coming back to the final point about RLM concepts and where it's being used, I have recently come across a lot of posts on X saying the RLM concepts have been used in some proprietary things like managed agent dynamic workloads using the RLM concepts under the hood. So they have implemented one or more forms of RLM inside their agent harnesses. Recently I saw that the Codex harness is writing Python code in the REPL that you can see to curate the context, that is one form of RLM I have seen myself. And obviously the cloud managed agents or Gemini managed agents, they're all kind of concepts of RLM. So basically you can get the harness in the sandbox and then you can do the stub. And the recent things about dynamic workflows where one agent given a task can spawn multiple agents that have their separate sandboxes, they can work together, and give back the final results. And the idea is basically generally coming from RLMs. A lot of software factories concepts are probably using RLMs, but we are not sure yet. However, some cloud code engineers from Anthropic acknowledged on X that they have used concepts of RLM. You can use this RLM concept on your large context repository. And if you have any questions, feel free to reach out to me. And finally, thank you so much for listening to my talk. we can use those concepts for larger code bases. My name is Shashi, I am a founder of Superagentic R. First of all, let's be clear that RLM paper has been published by MIT and Friends. As you can see there's a full paper, you can read about it. But the purpose of this talk is how you can use the concepts of RLM and you can use into your own workflow to implement your own harnesses. So, first of all, what's the problem? If you're using the coding agents for smaller repos or mono repos, they work exceptionally well. But if you have ever tried it with the mono repos, with the large context, you know there's a context problem. As the context grows, the performance degrades. And if you're working with the mono repos, this problem gets worse. In this talk, we will see we selected the code base and the concept of RLMs are relevant for the larger code bases. If you use the coding agents, then you probably saw that there are different approaches that other coding agent harnesses have been taken to solve this problem. The most common approach is searching using the tools like grep. So basically there's a file system and the coding agent harnesses search using these tools. The second approach, you've probably seen that the semantic search search or the local search. So idea here is basically you can search through the code and curate the context. Another approach is the long context get compressed and you can use the summarized version of the context. And there are some memory solutions available in the market as well that you can use to persist the memory for the coding agent. First of all, let's explore the RLM idea. The core thesis of the RLM is you need to externalize the context management into programmable execution environment. Meaning you should have a separate dedicated environment so that model can operate on that. In this case, for example, your whole repository is treated as a data that model can operate on. Then model can write the code to inspect, slice and compute the relevant chunks. The value you can then feed into the main context window. So basically rather than putting everything into the model's context, create a separate dedicated environment, give them a coding agent or REPL, model and then model write the code to curate the context that can be used into the main. So it's another context management technique proved to be very effective. Could be also be used as a memory layer for your coding agents. Let me summarize this giving you a simple analogy. Imagine you are a lead software engineer and assigned to the new project with a huge code base. Imagine that's a monorepo. How does that lead engineer deals with the code? So rather than reading each line of code line by line, engineer probably inspect the code base, make some notes, see what are the project's dependencies, how it is structured. Maybe something else is not understood by the repository. Engineer probably asks to another engineer or expert to get some ideas. And the same concept is applied in the RLM. So large project like the files and docs and texts and configs because the repository has a lot of things. And the programmable REPL is kind of a notebook that engineer makes a note about the code base that can be used. Researching, he may be using other techniques, or maybe writing some script to search something from the repo. And then if he stugs, then he asks another engineer or specialist where it comes to the LLM query. And LLM query is basically asking another model environment to get an answer from. And once they get answer, then the loop continues. And at the end, it returns the clean note synthesis. So the recursion part here is engineer asks another specialist using LLM query. That can be one question or that can be number of questions. So this is where the recursion comes in picture. The loop is basically your repo as your context. And then the model writes the REPL code to get some relevant context. That returns the bounded observation. And if loop needs more information, it passes through the LLM query where it asks another language model or another system to get the response, return the value, and continue the loop. And the loop gets terminated until we get our final results. So the code base is different. It has directories. It has tests. It has some imports. It has dependencies. It has tests. It has pictures. It has configuration files. So the code base is not only just the text. It is a structured data. It has a structured data. And the model needs to understand and reason over the text. That's why I chose this scenario to use the code basis to prove these concepts of RLM. Now let's switch the gear and talk about our own library that we created at SuperAgentic AI called RLM code. You can see RLM code's landing page here. This is just a research playground where you can implement the concepts of RLM. We have documentation that you can take a look and there's the GitHub repository. It is completely open source project that you can use it and play with it. RLM itself is a concept and a pattern. And you can implement that concept and pattern in your own way. There are official authors also wrote some implementation in their GitHub repo. It's called RLM and RLM minimal. You can refer that implementation of RLM in dspy.rlm. So Omar is author of RLM and he is also author of another popular framework called dspy. So dspy got RLM implementation inside it. However, you should treat they are completely different. So RLM is a pattern and you can implement in your own ways. You can find there are various other people implemented RLM in their own way. And in the similar way, we implemented RLM code as our own independent harness that we will be using in this live demo. RLM code is just a reference implementation to demonstrate how the RLM concepts works under the hood. So we have implemented something called RLM mode. We are using RLM as it is. We are not adding anything on top of RLM's ideas and RLM's paper. We are using the same concept of recursive calls, REPL execution. However, you can run it with a local model. You can run it with a cloud-based model. You can plug into any observability framework of your choice. And that gives you like a lot of flexibility around RLM. You can also plug it into the framework of your choice. For example, you can use PyDentic AI or Google ADK or something similar framework and implement ideas of RLM over there. In order to demonstrate this, we have created a source code repository where you can try this concept by yourself using MIT's RLM paper and RLM code. And we will see how these things work in a practice. So basically, we will show you the loop. This, you will understand this once we see this live demo and what all these files are doing. Where, where's the context has been created, where the Python REPL has written a code, and where it's passed to the LLM query, and how we get the final results. So that will be covered as part of the live demo. So we can cover this, everything here. So in nutshell, how it looks like is basically, it creates the REPL and then observation and the final recursive language output. so let me jump into the live demo now okay let's do the live demo of these concepts of rlm and rlm code and how it works in a the larger code basis so i have a code repositories here i have checked out and let's open it into the editor so that we can see what's inside it so as you can see there's a demo target which is um we are using rlm code source as a as a demo here and then we have some instructions that you can follow along um yourself so basically you have a readme file that you can use to use with your local model or so we are going to use with the gemini so we will try this script and see what happened so right now you can see we're using the docker as a sandbox if you see the docker container has been just started for this rlm and now coming back to our execution you can see that execution i just finished and in this execution what you have seen basically in the first step model has written the repl code that you can see here and then it's built the evidence and after that it also made the calls to the llm query with some prompt and got the result back and after that it gives the final answer and as you can see here we can have all this all the step coming back to the the final answer and here you can see the it made the two tool calls and how many the tokens is used for this model that you can see it here and the good thing is that you can see all these traces in the rlm code repositories so for example you can see all the runs this is the run that we just did you can see all the sessions and all the observability that you can plug it into any of your observed favorite observability platform so this is the cli path we just demonstrated but we also have this kind of coding agent style experimental harness where you can try the same thing so first of all let's connect with the gemini model so you can connect with the gemini model using the command connect and you can have a provider and the model name so you can also run the doctor command and see if everything is okay seems like doctor command found some warning but this is related to deep agent adk and other frameworks which is not relevant to this demo and the interesting part where we will be saying is basically you are sending the prompt so for what we did now we ran the command and we ask the question we specify the budget so that we don't know spending too much uh on this run but once we do that as you can see you have the maximum steps recursion depth and it completed its this run and coming back with the results we can also see that this thing into the research lab where we can see this spin has been completed we can see some rewards we can also see the trajectory which is important part where we can see the all the rlm loop for example the the ripple and the code and the final output so and also we can see the the events when it started and when it ended so you can play around with this rlm code terminal user interface which is kind of harness and you can experiment your rlm ideas in here so i'm going to quit this for now and let's switch back to the slides in a nutshell what we just saw basically our context has been loaded we have some repl code written to extract some snippets we also saw the llm query has been called to get some more context from another model this is where the recursion comes in picture and we got the final result and we got the the traces in json l format that you can import it into the any of the observability platform of your choice we also saw these results coming from different piles you can take a look at the source code that will be available uh for you let's talk about the real thing how ai engineer could use this concepts in the real life and there are few things for example if you're dealing with this large source code and you want to for example root cause analysis or onboarding of the repositories or some unfamiliar repos so there are few use cases you can from here and probably try to use rlm concepts over there basically you can design your own harness um based on your needs so that should capture the whole trajectory all these things like the planning coding observation sub call budget and the final output now coming back to the final point about rlm concepts and where it's been used i have recently came across a lot of the post on x saying the rlm concepts have been being used into the some of the proprietary things like the managed agent dynamic workloads using the rlm concepts under the hood so they have implemented one or more forms of rlm inside their agent harnesses recently i saw that the codex harness is writing the python python code in the ripple that you can see to curate the context that is the one form of rlm i have seen myself and obviously the clouds manage agents or gemini managed agents they're all kind of concepts of rlm so basically you can get the harness in the sandbox and then you can do the stub and the recent things about the dynamic workflows where one agent given that given the task you can spawn multiple agents that have their separate sandboxes they can work together and give back the final results and the idea is basically generally coming from the um rlms a lot of software factories concepts are probably using the rlms but we are not sure yet however some of the cloud code engineers from anthropicas accepted on x that they have used concepts of rlm you can use this rlm concept on your large context repository and if you have any questions then feel free to reach out to me and finally thank you so much for listening to my talk