It’s Tokens All The Way Down: How RLMs are Different — Kevin Madura, AlixPartners
Description
Ask a model to sum twelve numbers scattered across thirty thousand tokens and it will often get it wrong. Give it a Python REPL and it writes a few lines of regex instead. That is the shift Kevin Madura, director of advanced technology at AlixPartners, describes in recursive language models. An RLM treats its context as an object living in a REPL rather than as tokens it must attend to, so it can slice, compute over, and iterate on the input as a variable. The second property matters as much: it can delegate to another model, including itself, with its own parameters, so a hard problem decomposes recursively and only the results that matter return to the main context. On a long chain of thought benchmark that moved accuracy from 2.6 percent to 45.4 percent, with the largest gains on tasks that reduce cleanly to code. Madura frames it against what most teams do now. RAG stuffs the window until quality rots. Agents and tool calls shuttle JSON strings back and forth, leaving logic, execution, and results loosely coupled. An RLM keeps all three in one environment. He walks a cohort retention analysis where three data frames go in, the model reasons in its own REPL as if typing in a notebook, and decides itself when to stop and submit a typed answer. Then case studies: consolidating long invoices with no chunking or embedding, surfacing patterns in raw logs, optimizing an agent harness from its own traces, and generating a security report across five hundred thousand lines of code. His closing bet is that models post trained to be RLM aware will make this much stranger. Speaker info: - https://x.com/kmad - https://www.linkedin.com/in/kevinmadura/ - https://kmad.ai Timestamps: 0:00 - What a recursive language model is 1:36 - Recursive decomposition and delegating to sub models 3:25 - Benchmarks and the cost curve 4:31 - A deterministic shell, with the model filling the middle 5:24 - Context rot, and why an RLM avoids it 6:21 - How this differs from RAG, agents, and too
Summary
Generated by gpt-5.6-terraAt-a-Glance
- Verdict: Watch fully
- Core thesis: Recursive language models (RLMs) replace token-heavy context management with an LLM-controlled computational environment—typically a Python REPL—where models can inspect data symbolically, write code, and recursively delegate subtasks to other models.
- Why it matters: This is a practical architecture pattern for making agent systems more reliable on massive documents, dataframes, codebases, logs, and traces without repeatedly serializing intermediate state into prompts or tool-call strings.
- Best use: Use the talk to evaluate RLMs as a control-plane primitive for long-context analysis and agent-harness optimization, then prototype the pattern on one high-value, decomposable workload rather than treating it as a universal replacement for RAG or standard tools.
Executive Summary
Kevin Madura presents RLMs as a distinct agent architecture, not merely a prompting or tool-calling technique. An RLM keeps large input context as an object in a symbolic execution environment, usually a Python REPL. The primary model can inspect or compute over that object, write code against it, and call subordinate LLMs to handle bounded pieces of the work. The main model receives only the salient results rather than being forced to attend over every source token.
The claimed benefit is relief from context-window degradation and the operational burden of chunking, embedding, retrieval, and serializing intermediate outputs as JSON or strings. Madura contrasts RLMs with RAG, conventional agents, tool calls, and coding agents: those approaches often move text back into model context, whereas the RLM can retain state as variables and perform computation directly where the data lives.
The best workloads are large, dense, decomposable, and relatively latency-tolerant: document sets, invoices and contracts, dataframes, codebases, logs, agent traces, and code-friendly reasoning tasks. He cites benchmark results and small demonstrations suggesting large gains where code can precisely locate, filter, or calculate over information. He also gives examples of RLM-driven cohort analysis, invoice consolidation, security review of a 500,000-line vulnerable application, and meta-optimization of agent harnesses from traces.
The talk is directionally compelling but should be read as an architecture briefing rather than independent validation. Several comparisons are speaker-reported, some examples are acknowledged as favorable to RLMs, and the coding-agent comparison is explicitly described as potentially unfair. The actionable insight is to preserve deterministic contracts at the outer boundary—typed inputs, outputs, iteration budgets, and security controls—while allowing the model broad autonomy inside a sandboxed computational workspace.
Key Takeaways
- Claim: An RLM differs from ordinary tool-using agents because it operates on context as a symbolic object in an execution environment rather than repeatedly exchanging serialized strings. | Evidence: Madura describes the standard environment as a Python REPL: a large document, dataframe, or other input remains a variable that the model can query, transform, and compute over. Conventional tool calls typically pass JSON or strings to another program and receive strings back. | Implication: For agent systems handling structured or very large state, prioritize designs where the model can operate on governed in-memory objects rather than stuffing transformed state back into prompts after every step. | Caveat: The distinction is architectural rather than absolute: some newer workflow products increasingly preserve intermediate results in script variables, blurring the boundary with advanced coding/workflow agents.
- Claim: Recursive delegation is the core scaling mechanism: the primary model decides how to decompose a task, writes logic as needed, and can delegate subproblems to the same or a different LLM. | Evidence: The main model can invoke sub-LLMs from the REPL with selected parameters; those submodels can recursively interpret and solve their own subtasks. In the dataframe example, a main model could send a large subset to a sub-LM for analysis and harvest its result. | Implication: Treat recursion, model routing, stop conditions, and cost limits as first-class control-plane concerns, not incidental implementation details. | Caveat: Delegation must be bounded operationally; the example exposes a configurable maximum-iteration limit, while the model itself decides when it is ready to submit a final result.
- Claim: RLMs can mitigate context rot because the full corpus does not need to occupy the main model's token context; only selected slices and computed results are surfaced. | Evidence: Madura frames the issue as a 'dumb zone' after a context window becomes overly full. With an RLM, the corpus lives in the REPL and the model chooses how to access it or offload analysis, rather than attending to all tokens simultaneously. | Implication: RLMs are most worth testing where retrieval and manually engineered chunking are consuming disproportionate engineering effort or producing incomplete cross-document reasoning. | Caveat: This does not eliminate the need for architecture or evaluation; it shifts context management into model-directed exploration, code execution, and submodel calls.
- Claim: The approach appears especially effective for problems whose solution can be expressed through code, filtering, extraction, or deterministic computation over long inputs. | Evidence: Madura cites a reported long-chain-of-thought result of 45.4% accuracy for an RLM versus 2.6% overall on compared tasks, with strong performance in logic puzzles, chess, and chemistry. His simple example is summing 12 values hidden across 30,000 tokens, where regex/code is more dependable than asking a model to retain and add all values unaided. | Implication: Select pilot tasks where correctness can be checked against a programmatic ground truth—reconciliation, tabular analysis, code inspection, log analysis, or structured extraction—before using RLMs for ambiguous judgment-heavy workflows. | Caveat: He calls the base-model comparisons somewhat unfair and says the coding-agent comparison needs more rigorous experimentation.
- Claim: A robust RLM implementation should combine a deterministic outer contract with model autonomy inside the execution loop. | Evidence: Madura's mental model is to define task intent, expected inputs, desired typed outputs, and high-level guidance, while allowing the model to choose the implementation. His cohort-retention example takes three dataframes, specifies what to analyze and the required output types, then lets the RLM generate and run code before issuing a final typed submit. | Implication: For production adoption, preserve schemas and acceptance tests at interfaces while avoiding premature hand-authored orchestration for every internal reasoning step. | Caveat: The speaker's 'defer everything to the model' framing is aspirational; production systems still need defined boundaries, iteration caps, observability, and output validation.
- Claim: RLMs are not the default for all workloads; they trade latency and execution complexity for better handling of scale and decomposition. | Evidence: Madura recommends them for large or dense context, long-horizon sessions, decomposable tasks, and potentially extremely large outputs. He says to skip them when the task already fits in context, requires low latency, or is better served by a model that is already a strong coder. | Implication: Route requests by workload shape: retain straightforward single-context or low-latency paths, and activate an RLM path only after a size, complexity, or verifiability threshold is crossed. | Caveat: The transcript does not supply detailed latency, reliability, sandboxing, or total-cost measurements for real production deployments.
- Claim: RLMs may become a meta-optimization layer for agent systems, using long execution traces to improve the harness rather than merely optimizing individual workflow prompts. | Evidence: Madura cites Halo, a project that applies an RLM to agent traces in order to recommend a better harness, and mentions Sam Hogan of Inference.net using RLM-style analysis of production workload traces to identify work that can be deferred to another model such as GLM 5.2. | Implication: Long traces from orchestration systems should be treated as analyzable design data: an RLM could identify repeated failure patterns, needless model escalation, opportunities for routing, and brittle workflow steps. | Caveat: These are emerging project examples rather than validated, general production outcomes in the transcript.
Detailed Brief
Implementation ecosystem and emerging design convergence
- Claims: The RLM pattern is already appearing both as dedicated libraries and as a capability folded into broader frameworks.; Typed communication between parent and submodels may make recursive decomposition more maintainable and potentially allow cheaper models to perform bounded subtasks effectively.; Madura sees model post-training for RLM-aware behavior as a possible inflection point, because models could learn to exploit recursion and symbolic environments natively.
- Evidence: Named implementations include Predict RLM, DSPy, Axe, and FastRLM.; Predict RLM is described as targeting knowledge work over spreadsheets and PDFs; Trampoline AI is cited as using it for tasks such as turning a directory of complex invoices into a consolidated inventory.; Predict RLM reportedly uses DSPy to establish schemas for handoffs between the main model and submodels.; Madura notes that Anthropic Workflows stores intermediate results in script variables and says an Anthropic conference speaker referenced the RLM paper as an influence.
- Caveats: The claimed performance benefit from DSPy-enforced schemas for cheaper models such as Qwen is Madura's hypothesis; he says it needs experiment-based confirmation.; Product and framework references demonstrate activity in the ecosystem, not interoperability, security maturity, or equivalent capability across implementations.
- Implications: A practical evaluation should separate the core pattern from any individual framework: symbolic state, constrained execution, recursive calls, typed handoffs, and final validation are the reusable elements.; Schema-enforced parent/submodel contracts are a promising lever for observability and model-routing discipline, especially if lower-cost models handle narrow extraction or computation tasks.
Concrete deployment candidates
- Claims: Knowledge-work document processing is a natural near-term RLM use case because source files are lengthy, heterogeneous, and often require cross-file consolidation.; RLMs can analyze arbitrary structured operational data without first turning all of it into prompt text.; Large codebases are a candidate for broad inspection tasks such as generating a security report.
- Evidence: An AWS engineer is cited as using an RLM experimentally on log data to surface useful findings.; Madura describes an OWASP intentionally vulnerable web application experiment in which a small amount of RLM code analyzed roughly 500,000 lines of code and produced a security report.; The presented cohort-retention workflow has the model inspect three dataframes directly in its REPL, generate analysis code, and return formatted findings and recommendations.
- Caveats: The security example uses an intentionally vulnerable application, so it does not establish efficacy or safety on a real enterprise codebase.; The transcript does not address permissions, data isolation, network egress, code-execution sandboxing, or human review requirements—critical omissions for use with sensitive documents, production logs, and source code.
- Implications: The highest-value initial applications are internal, read-only analyses with clear evaluation datasets and limited blast radius.; Any RLM deployment over proprietary data should be designed as an execution-security problem as much as an LLM-quality problem.
Notable Concepts & Terms
- Recursive Language Model (RLM): An LLM architecture in which the model works within a symbolic environment, can write/run code against context objects, and can recursively delegate subtasks to other LLMs.
- Symbolic environment / Python REPL: The execution workspace where data remains as variables and can be directly inspected or computed over instead of serialized back into prompt tokens.
- Context rot / dumb zone: The degradation in model performance as a context window becomes too full; RLMs aim to avoid exposing all source tokens to the primary model at once.
- Bitter Lesson: The speaker's guiding belief that increasing model capability and compute should let systems defer more implementation choices to learned model behavior rather than hand-crafted rules.
- DSPy: A framework central to Madura's design approach; here it is valued for declaring input/output contracts and enforcing schemas between main and submodel calls.
- Typed outputs / final submit: A production boundary in which the system declares the expected output structure upfront and only accepts a model result when it submits that structure.
- Halo: A cited project that applies RLMs to lengthy agent traces to improve the agent harness itself, rather than only tuning a single workflow.
- Predict RLM: A cited RLM-oriented implementation for knowledge work over artifacts such as PDFs and spreadsheets, including schema-aware delegation through DSPy.
Operator Notes / Why Ken Should Care
- Choose one read-only, high-volume workload with programmatic evaluation—such as multi-file invoice extraction, operational-log investigation, dataframe analysis, or repository inventory—and benchmark an RLM path against the current RAG/tool-call path on accuracy, latency, total token cost, and human-review burden.
- Define the outer contract before prototyping: allowed data objects, tool permissions, typed final schema, maximum recursion depth, maximum iterations, per-run spend cap, timeout, and audit trace requirements.
- Run RLM code only in a restricted sandbox with no implicit network access, scoped filesystem/data permissions, resource quotas, secret isolation, and explicit approval gates before any side-effecting action.
- Instrument traces so a later RLM-based meta-analysis can identify repeated context bloat, bad routing decisions, unnecessary escalation, and workflow steps suitable for deterministic replacement.
- Avoid replacing simple low-latency or single-context tasks with an RLM; use workload routing so added recursion and execution overhead are incurred only when data size and decomposition justify it.
Source/Metadata
- Title: It’s Tokens All The Way Down: How RLMs are Different — Kevin Madura, AlixPartners
- Transcript words: 6189
- Duration seconds: 1247
- Timestamp note: No usable timestamps or chapters were provided in the transcript; substantial portions of the latter half appear duplicated.
Transcript
Kevin Madura Reviewer Awesome. Thanks everyone for being here. My name is Kevin Madura. I'm from a company called Alex Partners. We're a consulting firm. I'm here to talk to you today about RLMs. Just curious, show of hands, who here is familiar with RLMs? So we know how much time to spend on it. Okay, so not many. All right, well that's good. So we'll start with what an RLM is and why it's different. So RLM is recursive language model and really the key difference here is that it treats the context as an object that it can interact with symbolically in its environment. So it differs from a tool call in the sense that typically when you do a tool call, it's JSON or some type of string that's being sent, being interpreted elsewhere, maybe by some other program and that's returning effectively as a string. The key difference here is that it's interacting with a symbolic environment. So typically that's a REPL, Python REPL. So that's key difference number one. Key difference number two is that it has the ability to delegate to another LLM, often to itself. You can specify whether it's the same model or a different model, but fundamentally because it lives in this environment, you can offload or make a sub call to another LLM with particular parameters that also lives in that REPL environment. And so you get this ability to recursively decompose problems and apply and have the LLM basically decide how to apply certain logic or certain interpretations or write its own code to solve those problems. And then that recurses down. So the sub LLMs can do the same sort of thing in terms of understanding and interpreting what it thinks it needs to do. And I added this last one here. It's largely bitter lesson-pilled in my opinion, right, and shared by Alex and the rest of the creators of it. But as models get better, you should be able to defer more and more to the model for it to figure out on its own what it needs to do. So I don't know if this is the actual kind of starting point for our LLMs. This is one that I consider to be one of the first kind of inklings of it. This is a tweet from Omar, who is Alex's advisor for RLMs. And this was a concept that he had come up with where it was basically an ability to use DSPy and some other techniques to take in arbitrary length inputs. And basically the use case here would be summarizing an arbitrarily long document and coming up with a table of contents and some summary of that content. But at least to me this is the first inkling of, okay, context windows might not be something you need to deliberately manage, although there's of course benefits to doing so. There could be ways to exceed the context windows using some of these clever techniques. And so if you read the paper and some of the blog posts that are out there from Alex and Omar, it has demonstrably better performance on some of these long context tasks. So Oolong is one benchmark where the intent of the benchmark is to measure model performance on answering questions about excessively long context. This is another one, BrowseComp, where it needs to iterate through a large body and corpus of text and answer particular questions about it. You can see the blue line at the top there is the RLM. It's very good performance as compared to some of these other models. And even on the price curve, the purple is actually just using tool calling with GPT-5 calling a BM25 tool. And that's actually even more expensive for worse performance than an RLM. So it's worth reading into if you're interested in some of the benchmarks and how RLMs perform. But fundamentally, an RLM, again, takes in your input and you're deferring to the model about how to decompose the process, what code it needs to write. And it is very tightly integrated with the REPL itself. So it by itself defines what it needs to do. And so I kind of had this mental model in terms of, and I'm very DSPy-pilled, if you can tell by now, basically a student of Omar and the rest of the group there, where you have this relatively deterministic shell of what you want to do. Like, what is your intent? What is your actual task that you're trying to accomplish? You define that in terms of your inputs and your outputs, and some type of guidance or prompt or what have you to the model to say, this is generally what I want to achieve, go off and do it. Here's the things that you can expect as your input. Here's what I want out of it. Go figure out the rest. And so this applies for using something like DSPy, but I think it applies to RLMs as well, because you don't have to worry as much now about how the actual implementation works in the middle. You can just have some guarantees about the inputs and the outputs, and you can let the model figure out the rest of that part of it. So a lot of this comes down to if you were at, I think it was Code in November in New York City, Dex had this great talk about broader context engineering, and he coined something like the dumb zone, which is great out at the bottom there. But the point is that we all know that there's context rot, right? Once you fill up the context window to a certain degree, performance starts to degrade. And so RLMs somewhat get around this problem because the context itself doesn't fill up as quickly because you're deferring a lot of the subtasks to the submodels, and it's the full kind of context and the inputs aren't exposed to the context window itself. It lives as a variable in the REPL. And so the main LM can choose how to access that. It can offload some of these subtasks to sub-LMs. And really the only context that it gets back are the things that actually matter. So in terms of how it's meaningfully different, RAG, of course, you kind of just stuff the context window. You want it to limits there. Agents are largely just bringing strings back, and you don't have this tight coupling between the logic, the execution, and the results. And so you still run into the same sort of problem there. Same thing with tool calling and code act. And then RLMs, as I mentioned, the LLM is actually just interacting with the context, the results as variables in the REPL so that it can do additional computation on versus it trying to attend to all these different tokens in its context window. It's a meaningfully different way of the LLM interacting with the actual content itself. And so people always say, okay, what's the difference between that and coding agents? In my mind, the largest difference is that the way that tool calls are done is passing strings back and forth. But you can see with the release recently of Workflows that Anthropic is doing something fairly similar. And they, at the CAIS conference, I think it was Tarek or someone similar, mentioned the Arlen paper as a key driver of Workflows and how they've implemented it. And you can see here the intermediate results for Workflows live in script variables, i.e. a variable in the context. So it's driving some of these breakthroughs and some of these techniques from the labs as well. I'll skip through this a bit just because I have about 10 minutes left. But generally speaking, when you want to use it, it's obviously for large or dense input context. An underexplored area is outputs as well. So if you have some type of task where you need to generate hundreds of thousands of lines or whatever it might be, RLMs, I think, would be a good candidate for that as well. Obviously tasks that are amenable to some type of decomposition. So if you want to look through the entire, I don't know, the whole tax code as an example and try and find loopholes or something, you can't obviously put all of that into context at once. You could use an LLM to crunch through all of that and iteratively explore and use sub-agents to explore interesting areas of the tax code, bring back those sections, and then reason over that. And then just generally for longer horizon sessions. And when you want to skip it, of course, it would be something that fits in context. You want something that's low latency or the model itself is as strong of a coder. And our friend Raymond here did some great performance testing on the long chain of thought benchmark. I'll leave this link as a leave behind after, but just to give you a sense of how well it performs in some of these tasks. It's a meaningful jump overall from 2.6 to 45.4% accuracy on many of these tasks. And you can see it performs really well on things that are amenable to code. So logic puzzles and chess and chemistry and things like that where it can dynamically write code, bring in only the relevant part of the context, compute that, and then return the result, where the main model is really just harvesting the results from the sub-LMs instead of trying to do that by itself. I put together a few super simple examples. I mean, these are somewhat unfair, I suppose, to the base model, but it makes the point that there are certain tasks that base models just aren't really fit to do themselves. I'll leave this link as a leave behind after, but just to give you a sense of how well it performs in some of these tasks. It's a meaningful jump overall from 2.6 to 45.4% accuracy on many of these tasks. And you can see it performs really well on things that are amenable to code. So logic puzzles and chess and chemistry and things like that where it can dynamically write code, bring in only the relevant part of the context, compute that, and then return the result, where the main model is really just harvesting the results from the sub-LMs instead of trying to do that by itself. I put together a few just super simple examples. These are somewhat unfair to the base model, but it makes the point that there are certain tasks that base models just aren't really fit to do themselves because they have to attend all these different tokens at once in the context window, where you need to use some type of coding approach. So in this random example, summing 12 numbers that are buried across 30,000 tokens, the LLM trying to figure all that out by itself and give you the answer isn't always going to work as well as something that you can write regex for or something similar. And then the same sort of thing, particularly for data frames, where because the LLM can interact with the data frame within the REPL, it just has a much better understanding of the content and can iterate through that much more quickly than having to pass tool calls back and forth in terms of JSON strings and that sort of thing. And then I threw this in there in terms of running the same experiments with a coding agent. I didn't look into this too deeply. There's probably some unfair math going on here, but you can see that it was totally bloated in terms of the way that Cloud Code tried to solve these tasks. So there's more work to be done there in terms of running experiments to compare base models versus RLMs versus something like a coding agent. But for certain tasks, for production workloads, my sense is you probably don't want to just do Claude-P, your prompt, and hope for a good result. You want more of a structured approach to your inputs, your outputs, and you want a defined pipeline for doing so, which reduces your cost, it reduces your complexity, it reduces your bloat, where RLMs can shine. So in the real world, there are a bunch of different open source libraries that implement RLMs at some level. Some of them are more RLM focused, like Predict RLM would be a good example of that, versus others are just integrating it into the broader approach or the broader framework. DSPy, obviously, there's Axe, which is really interesting work that's being done there. Predict RLM is more focused on knowledge work, so it works with spreadsheets and PDFs and that sort of thing, and then FastRLM. And then there's a tweet yesterday from this guy Sam Hogan, who runs Inference.net. He's using an RLM to basically run and extract insights from your particular production workload traces, so that they can see what makes sense to defer off to something like a GLM 5.2, and do that iteratively and automatically as your traffic goes through. So the point is, you don't have to worry about context engineering. You can just throw the RLM at it and have it figure it out. I only have five minutes left, so we won't go through this whole example, and I'll skip to some of the traces, because that's probably the most interesting. But this is all you would really need to do in terms of a simple cohort retention analysis, something that you might give to a data scientist. But this concept of applying an RLM to a complex data structure like a data frame becomes very easy to do. This is all the code you need to do it. I'm feeding in three different data frames, I'm saying these are the sorts of things you need to look for, these are the output types that I want. And then just let the RLM go on it, and I'll show you some of the traces. So it has its own REPL where it can interact with those data frames, and you can see it reasoning through. Okay, first I need to do this. It's writing the code. And because it's living in the REPL with the data frame, you don't have this additional bloat of the tool calls back and forth. It's actually interacting directly with the data frame as if it was typing in its own Jupyter notebook. And there's significant advantages for doing so. And so you can see the sorts of outputs that it gets as a result, and it by itself will iterate. And in this case, it didn't, but it has the option to defer to sub-LMs to do, okay, now I have this big subset of the data. Sub-LM, go off and do this analysis, give me the result. And it can do that iteratively over time. But the point is that the LM is directly interacting with the data frame in its REPL and iterating through the results. So this platform compound is RLM and DSPy native. So it gives you this really nice breakdown of the reasoning. It separates out the code that's being generated. And ultimately, you can see the final output, which is here where it's formatting. Okay, here are the key findings that I have. Here are the recommendations. And then you have this final submit, which is the final answer that gives you the typed outputs that you defined up front. And the key thing here is that the LLM itself is deciding when to stop. So you have a variable of max iterations. So you can decide whether you want it to have a maximum of 10 or 100 or whatever it is. But it will by itself explore the data, understand what needs to happen. And then when it itself is comfortable, it can run submit and give you the final output. Again, being better lesson-pilled, this will get better over time. You can just defer everything and it will figure out what to do. And so the hope would be you don't have to, I mean, we're already at whatever this is, 20 lines of code or something. But you can see a world where you can continue to go up levels of abstraction. As long as you can define what your objective is and what you want it to do, the model will figure out the rest. So we just walked through a bunch of this. But these are the different steps that it took in this example and the code that it wrote. And then I'll just breeze through a few real world case studies and where it's actually being used. So I mentioned Predict RLM before. So the company Trampoline AI. They're doing really interesting work in applying RLMs for different pieces of knowledge work. So natively interacting with PDFs and spreadsheets and that sort of thing. So in this relatively simple example, okay, I have a bunch of invoices that I need to create one consolidated inventory out of. As we all know, invoices can be complicated. They can be very long. They can be all over the place. To do that today without RLMs or this sort of framework gets very complicated very quickly. I have a lot of battle scars to prove it. But with something like an RLM, you don't need to worry as much about, okay, if I have a 200 page invoice or contract or whatever it is, you can let the RLM just churn through all of that and give you the result instead of having to worry about chunking and embedding and doing all these different strategies to try and get around the content. And then you can also add to the context window management that we've all had to do previously. So it allows you to focus on the abstractions and what you actually want to do instead of the context engineering itself, which I think is a really helpful output of all this. And an interesting tidbit for all the DSPy fans in the room, Predict RLM uses DSPy to determine the schemas between the main LM and the sub LM calls, which I personally think is a nice feature because you have a lot more readability and maintainability. So you understand exactly what the model is trying to achieve. And the model can be much more precise and prescriptive about the types of data that it's looking for from the sub LM. And I would want to do some experiments to test this out, but I would think that this would improve performance for cheaper models like Quinn or some of the other ones because you're specifying the inputs and outputs and you're enforcing those types coming back. And so you get all the benefits of the RLM being able to turn through all this information, but you have a lot more of the structure in between where when it's handing off to a sub LM, it enforces some of those schemas. This is an example from an AWS engineer from a couple of days ago, where you're just playing around with it. So you understand exactly what the model is trying to achieve. And the model can be much more precise and prescriptive about the types of data that it's looking for from the sub LM. And I would want to do some experiments to test this out, but I would think that this would improve performance for cheaper models like a Qwen or some of the other ones because you're specifying the inputs and outputs and you're enforcing those types coming back. And so you get all the benefits of the RLM being able to turn through all this information, but you have a lot more of the structure in between where when it's handing off to a sub LM, it enforces some of those schemas. This is an example from an AWS engineer from a couple of days ago, where you're just playing around with it. But I just thought it was a nice example of how you can throw arbitrary data at RLMs. In this case, it was a bunch of log data to surface some interesting results and he found it useful. There's a project called Halo, which uses an RLM to look at traces of different agent tasks. And basically the promise of Halo is that instead of optimizing a particular workflow or DSPy or other framework structure itself, it's actually iterating on the harness. So it's like a meta abstraction or meta optimization of the harness itself. And it uses an RLM because as we all know, tracing can get very long and complicated. So the RLM can not only take in all of that context, but also leverage the structure of those traces to recommend a better harness. And then this last one, this is all the code you need. I ran this little experiment. There's an intentionally vulnerable application called from OWASP, but basically it's a web app with a bunch of vulnerabilities in it. This is all the code you need on the right hand side to run basically an agent to run through 500,000 lines of code to generate some type of security report. That's just an arbitrary example. But the point is you don't need a lot of context engineering. You don't need a lot of structure around it to achieve what you want to do. And so you can feed in an arbitrary size code base into this and get some type of insights out. So you can imagine that being applied to other areas as well. So I know I rushed through everything a little bit, but I'm happy to answer questions afterwards. I'll leave you with this. The biggest promise I see here is just imagine a world where the models are actually post-trained and actually RLM aware. I think things will get pretty crazy pretty quick when they actually know how to use and take advantage of the RLM methodology natively. So thank you so much for your time. the results as variables in the REPL so that it can do additional computation on versus it trying to attend to all these different tokens in its context window. It's a meaningfully different way of the LLM interacting with the actual content itself. And so people always say, okay, what's the difference between that and coding agents? In my mind, the largest difference is that the way that tool calls are done is passing strings back and forth. But you can see with the release recently of Workflows that Anthropic is doing something fairly similar. And they, at the CAIS conference, I think it was Tarek or someone similar, mentioned the Arlen paper as a key driver of Workflows and how they've implemented it. And you can see here the intermediate results for Workflows live in script variables, i.e. a variable in the context. So it's driving some of these breakthroughs and some of these techniques from the labs as well. I'll skip through this a bit just because I have about 10 minutes left. But generally speaking, when you want to use it, it's obviously for large or dense input context. An underexplored area is outputs as well. So if you have some type of task where you need to generate hundreds of thousands of lines or whatever it might be, RLMs, I think, would be a good candidate for that as well. Obviously tasks that are amenable to some type of decomposition. So if you want to look through the entire, I don't know, the whole tax code as an example and try and find loopholes or something, you can't obviously put all of that into context at once. You could use an LLM to crunch through all of that and iteratively explore and use sub-agents to explore interesting areas of the tax code, bring back those sections, and then reason over that. And then just generally for longer horizon sessions. And when you want to skip it, of course, it would be something that fits in context. You want something that's low latency or the model itself is as strong of a coder. And our friend Raymond here did some great performance testing on the long chain of thought benchmark. I'll leave this link as a leave behind after, but just to give you a sense of how well it performs in some of these tasks. It's a meaningful jump overall from 2.6 to 45.4% accuracy on many of these tasks. And you can see it performs really well on things that are amenable to code. So logic puzzles and chess and chemistry and things like that where it can dynamically write code, bring in only the relevant part of the context, compute that, and then return the result, where the main model is really just harvesting the results from the sub-LMs instead of trying to do that by itself. I put together a few just super simple examples. I mean, these are kind of, they're somewhat unfair, I suppose, to the base model, but it makes the point that there are certain tasks that base models just aren't really fit to do themselves because they have to attend all these different tokens at once in the context window, where you need or want to use some type of coding approach to that. So in this random example, summing 12 numbers that are buried across 30,000 tokens, the LLM trying to figure all that out by itself and give you the answer isn't always going to work as well as something that you can write regex for or something similar. And then the same sort of thing, particularly for data frames, and we'll walk through a brief example here, where because the LLM can interact with the data frame within the REPL, it just has a much better understanding of the content and can iterate through that much more quickly than having to pass tool calls back and forth in terms of like JSON strings and that sort of thing. And then I threw this in there in terms of running the same experiments with a coding agent. Now, I didn't look into this too deeply. There's probably some unfair math going on here, but you can see that it was totally bloated in terms of the way that Cloud Code tried to solve these tasks. So there's more work to be done there, of course, in terms of like running experiments to compare base models versus RLMs versus something like a coding agent. But there's a, for certain tasks, for like production workloads, my sense is you probably don't want to just do Claude-P, your prompt, and like hope for a good result. Like you want more of a structured approach to your inputs, your outputs, and you want a defined pipeline for doing so, which reduces your cost, it reduces your complexity, it reduces your bloat, all that sort of thing, where RLMs can shine. So in the real world, there are a bunch of different open source libraries that implement RLMs at some level. Some of them are more RLM focused, like a predict RLM would be a good example of that, versus others are kind of just integrating it into the broader approach or the broader framework. DSPy, obviously, there's Axe, which is really interesting work that's being done there. Predict RLM is more focused on like knowledge work, so it works with spreadsheets and PDFs and that sort of thing, and then Fast RLM. And then there's a tweet yesterday from this guy Sam Hogan, who runs inference.net. He's using an RLM to basically run and extract insights from your particular production workload traces, so that they can see what makes sense to defer off to something like a GLM 5.2, and do that iteratively and automatically as your traffic goes through. So point being, you don't have to worry about context engineering. You can kind of just throw the RLM at it and have it figure it out. I only have five minutes left, so we won't go through this whole example, and I'll skip to some of the traces, because that's probably the most interesting. But this is all you would really need to do in terms of a simple, in this case, it's like a cohort retention analysis, something that you might give to a data scientist. But this concept of applying an RLM to a complex data structure like a data frame becomes very easy to do. This is all the code you need to do it. Where I'm feeding in three different data frames, I'm saying, these are the sorts of things you need to look for, these are the output types that I want. And then just let the RLM go on it, and I'll show you some of the traces. And so it has its own REPL where it can interact with those data frames, and you can see it reasoning through. Okay, first I need to do this. It's writing the code. And because it's living in the REPL with the data frame, you don't have this additional bloat of the tool calls back and forth. It's actually interacting directly with the data frame as if it was typing in its own Jupyter notebook. And there's significant advantages for doing so. And so you can see the sorts of outputs that it gets as a result, and it by itself will iterate. And in this case, it didn't, but it has the option to defer to sub-LMs to do, okay, now I have this big, whatever, this big subset of the data. Sub-LM, go off and do this analysis, give me the result. And it can do that iteratively over time. But the point is that the LM is directly interacting with the data frame in its REPL and kind of iterating through the results. And so this platform compound is RLM and DSPy native. So it gives you this really nice breakdown of the reasoning. It separates out the code that's being generated. And ultimately, you can see the final output, which is here where it's formatting. Okay, here are the key findings that I have. Here are the recommendations. And then you have this final submit, which is the final answer that gives you the typed outputs that you had to find up front. And the key thing here is that the LLM itself is deciding when to stop. So you have this, you have a variable of max iteration. So you can decide whether you want it to have a maximum of 10 or 100 or whatever it is. But it will by itself explore the data, understand what needs to happen. And then when it itself is comfortable, it can run submit and give you the final output. Again, being better lesson-pilled, this will get better over time. You can kind of just defer everything and it will figure out what to do. And so the hope would be you don't have to, I mean, we're already, you know, whatever this is, 20 lines of code or something. But you can see a world where you can continue to go up levels of abstraction. As long as you can define what your objective is and what you want it to do, the model will kind of figure out the rest. So we just walked through a bunch of this. But these are the different steps that it took in this example. And the code that it wrote. And then I'll just breeze through a few real world case studies and where it's actually being used. So I mentioned predict RLM before. So the company Trampoline AI, I think it is. They're doing really interesting work in applying RLMs, like I mentioned before, for different pieces of knowledge work. So natively interacting with PDFs and spreadsheets and that sort of thing. So in this relatively simple example, okay, I have a bunch of, I have a directory of invoices that I need to create one consolidated inventory out of. As we all know, invoices can be complicated. They can be very long. They can kind of be all over the place. To do that today without RLMs or this sort of like framework gets very complicated very quickly. I have a lot of battle scars to prove it. But with something like an RLM, you don't need to worry as much about, okay, if I have a 200 page invoice or contract or whatever it is, you can let the RLM just churn through all of that and give you the result instead of having to worry about chunking and embedding maybe and doing all these different strategies to try and get around the content. And then you can also add to the context window management that we've all had to do previously. So it allows you, the point there is that you can focus on the abstractions and what you actually want to do instead of the context engineering itself, which I think is a really helpful output of all this. And an interesting tidbit for all the DSPy fans in the room, predict RLM uses DSPy to determine the schemas between the main LM and the sub LM calls, which I personally think is a nice feature because you have a lot more readability and maintainability. So you understand exactly what the model is trying to achieve. And the model can be much more precise and prescriptive about the types of data that it's looking for from the sub LM. And I would want to do some experiments to test this out, but I would think that this would improve performance for cheaper models like a Quinn or some of the other ones because you're specifying the inputs and outputs and you're enforcing those types coming back. And so you get all the benefits of the RLM being able to turn through all this information, but you have a lot more of the structure in between where when it's handing off to a sub LM, it enforces some of those schemas. This is an example from an AWS engineer from a couple of days ago, where you're just kind of playing around with it. But I just thought it was a nice example of you can kind of just throw arbitrary data at RLMs. In this case, it was a bunch of log data to surface some interesting results and he found it useful. There's a project called Halo, which uses an RLM to look at traces of different agent tasks. And basically the promise of Halo is that instead of optimizing a particular workflow or DSPy or other framework structure itself, it's actually iterating on the harness. So it's like a meta abstraction almost or meta optimization of the harness itself. And it uses an RLM because as we all know, tracing can get very long and complicated. So the RLM can not only take in all of that context, but also leverage the structure of those traces to recommend a better harness. And then this last one, this is all the code you need. I ran this little experiment. There's an intentionally vulnerable application called, it's from OWASP, but basically it's a web app with a bunch of vulnerabilities in it. This is all the code you need on the right hand side to run basically an agent to run through whatever it is, 500,000 lines of code to generate some type of security report. That's just an arbitrary example. But the point is you don't need a lot of context engineering. You don't need a lot of structure around it to achieve what you want to do. And so you can feed in an arbitrary size code base into this and get some type of insights out. So you can imagine that being applied to other areas as well. So I know I rushed through everything a little bit, but I'm happy to answer questions afterwards. I'll leave you with this. The biggest promise I see here is just imagine a world where the models are actually post-trained and actually like RLM aware. I think things will get pretty crazy pretty quick when they actually know how to use and kind of take advantage of the RLM methodology natively. So thank you so much for your time.