AI Engineer

Recursive Coding Agents - Raymond Weitekamp, OpenProse

4022 summary words 18 min summary Watch video

Start with the signal

18 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: RLMs (Recursive Language Models) represent a paradigm shift in AI agents by unifying tool calling and reasoning through symbolic manipulation of context, and their principles can dramatically improve coding agent reliability when agents recursively call themselves to decompose and verify work.
  • Why it matters: Addresses the core reliability problem blocking AI agent deployment: models are intelligent enough but can't consistently deliver outcomes; RLMs achieved 30%+ on ARC-AGI where frontier models scored 2-3%, and small models beat GPT-4/Opus on long reasoning tasks.
  • Best use: Study for building production-grade agent systems, especially if Ken needs reliable task decomposition, verification workflows, or wants to convert golden coding sessions into reusable workflows; contains specific implementation paths via Claude Code, OpenProse, and PY recursive extensions.

Executive Summary

Raymond Weitekamp (OpenProse) argues that the bottleneck for AI agents isn't intelligence but reliability and trust. His thesis: today's agents are 'mismanaged geniuses' that need better specification, decomposition, and verification layers. He experienced this viscerally—one day Claude Code built a working SaaS app from a prompt, the next day it emptied his Solana wallet. The solution: apply Recursive Language Model (RLM) principles to coding agents so they symbolically manipulate context, recursively decompose problems into sub-agent tasks, and verify work hierarchically rather than stuffing everything into a context window.

RLMs are a marriage of reasoning and tool calling where the context itself becomes the object of computation. Instead of reading millions of tokens into context, RLMs explore symbolically via a REPL (Python in the original paper) and spawn sub-agents to handle decomposed tasks. This approach achieved shocking results: Symbolica's Agentica harness scored 30%+ on ARC-AGI within hours of release (vs. 2-3% for frontier models), causing the ARC Prize team to refuse full evaluation. On the LongCoT benchmark, a 9B parameter model (Qwen 3.5) using RLM beat Opus and GPT-4 on long reasoning tasks. Weitekamp's own work showed RLM harnesses are top-10 memory systems by default and achieve state-of-the-art on long reasoning benchmarks.

Weitekamp implemented several recursive coding agent systems: YPI (a Y-combinator-style recursive wrapper for PY agent that now works as a pure extension), CLI wrappers for Alex Zhang's RLM package, and two OpenProse workflows demonstrating RLM patterns. He argues Claude Code became an RLM when Anthropic released dynamic workflows (February 2025), which allow the agent to decide problem decomposition rather than executing hardcoded map-reduce patterns. OpenProse is a markdown-based 'programming language' compiled by coding agents (not computers) that can turn any agent with file system access and sub-agent capability into an RLM.

Practical applications include: scale migrations/refactors via parallel swarms, deep recursive analysis of directories, adversarial verification via skeptical red-team agents, bug sweeps, and—critically—a system to deconstruct 'golden sessions' (successful agent runs) into reusable OpenProse workflows. Weitekamp's takeaway: trust requires reliability, RLMs unify tool calling and reasoning into a new test-time compute paradigm, and coding agents can be RLMs but aren't automatically—you need explicit decomposition control, symbolic state management, and recursive sub-agent patterns.

Key Takeaways

  • Claim: The bottleneck for AI agent deployment is reliability and trust, not raw intelligence—agents can't consistently deliver outcomes despite knowing the entire internet. | Evidence: Weitekamp's personal example: Claude Code built a working SaaS app from one prompt, then emptied his Solana wallet the next day. Symbolica's Agentica scored 30%+ on ARC-AGI vs. 2-3% for frontier models, but ARC Prize refused to validate it because they dislike RLM harnesses. | Caveat: No quantification of how often 'golden sessions' vs. failures occur across different agent types; Solana wallet incident is anecdotal and context is missing (was it a bug, prompt injection, or user error?). ARC Prize's objection suggests some believe RLMs solve problems 'the wrong way.' | Implication: For Ken's agent systems: prioritize decomposition, verification, and reusability architecture over chasing frontier model intelligence. If deploying production agents, build in explicit verification loops and adversarial checks. Consider that benchmarks may not capture RLM value if maintainers reject tool-calling approaches. | Timestamp: 00:45
  • Claim: RLMs treat context itself as the object of computation via symbolic manipulation and recursive sub-agent calls, enabling processing of millions of tokens beyond context window limits. | Evidence: Original Oolong paper showed RLMs handle tens of millions of tokens. Weitekamp showed default RLM harness is a top-10 memory system without modifications. On LongCoT benchmark, Qwen 3.5 9B as an RLM beat Opus and GPT-4 as LLMs on tasks requiring extremely long reasoning chains that frontier models couldn't hold. | Caveat: No mention of latency costs or computational overhead of recursive decomposition vs. single long-context calls. Benchmark maintainers created separate 'open harness' leaderboards to avoid contamination, suggesting RLM results aren't directly comparable to pure LLM approaches. | Implication: For Ken: RLMs offer a path to handle massive context (codebases, document corpuses) and long reasoning without hitting context limits or losing thread. Consider dspi.rlm implementation for benchmarking/prototyping. Trade-off may be increased latency and orchestration complexity vs. raw throughput. | Timestamp: 05:20
  • Claim: Claude Code became an RLM when Anthropic released dynamic workflows (Feb 2025), which let the agent decide problem decomposition rather than execute hardcoded patterns. | Evidence: Omar Khattab (RLM co-author) congratulated Anthropic for making Claude Code 'finally an RLM' with dynamic workflows. Weitekamp wrote two Claude workflows: one explicitly not an RLM (hardcoded map-reduce) and one that is (deep research over file system where agent picks decomposition). Anthropic's blog post 'A Harness for Every Task' shows six workflow patterns. | Caveat: Dynamic workflows are Claude-only; no equivalent in Cursor, Codex, or other agents yet. Weitekamp's rubric for 'is it an RLM' includes agent picking decomposition—hardcoded map-reduce doesn't count even if it uses sub-agents. Lambda RLM project falls short because it forces decomposition via lambda calculus. | Implication: For Ken: If using Claude Code, dynamic workflows unlock RLM capabilities now. If using other agents (Codex, PY, etc.), OpenProse provides a cross-platform path to RLM patterns. Design workflows where the agent decides 'how to split the problem' rather than forcing a predetermined structure. Key differentiator is adaptive decomposition. | Timestamp: 17:15
  • Claim: OpenProse is a markdown-based 'programming language' compiled by coding agents that can turn any agent with file system and sub-agent support into an RLM by explicitly declaring sub-agent work, skills, and tool dependencies. | Evidence: OpenProse repo includes harness supporting Codex SDK and Claude Code underneath. Weitekamp added features to declare skills and tools as explicit dependencies so sub-agents are configured with required capabilities. Example: 'prose write' command has agent write a .prose.md file (similar to Claude's ultra code command). Two demo .prose.md files in companion repo show problem decomposition and parent-session verification of sub-agent work. | Caveat: OpenProse is early-stage and open source but adoption/maturity unclear. Requires agent to 'compile' markdown specs, which adds interpretation layer and potential failure mode. No comparison of OpenProse vs. native dynamic workflows performance or reliability. | Implication: For Ken: If building cross-platform agent workflows or need portability across Claude/Codex/PY, OpenProse offers a unified abstraction. The dependency declaration feature (skills/tools for sub-agents) is powerful for ensuring sub-agents have required capabilities before execution. Consider for standardizing agent workflows across team or product portfolio. May be overkill for single-agent use cases. | Timestamp: 19:30
  • Claim: YPI (Y-combinator recursive PY extension) demonstrates fully recursive coding agents where the agent harness calls itself, and as of this talk PY evolved to support this as a pure extension rather than requiring a fork. | Evidence: Weitekamp originally forked PY to implement recursion but revisited before this talk and found PY extensions now support it natively. YPI wraps PY so 'PY calls PY calls PY' with configurable depth. PY is designed by Mario to be minimal and extensible rather than feature-stuffed, encouraging users to write extensions for new capabilities. | Caveat: No performance metrics, depth limits, or failure modes discussed for recursive PY. Unclear how recursion depth is determined or what happens if sub-agents diverge or loop. PY's minimalism may mean more setup/extension work vs. batteries-included agents like Claude Code. | Implication: For Ken: If building custom agent systems or need lightweight/extensible base, PY + recursive extension offers a path to RLM patterns without vendor lock-in. The pure-extension approach means clean separation of concerns. Consider for research/prototyping or when you need full control over agent behavior and don't want opinionated defaults. | Timestamp: 14:00
  • Claim: A 'golden session' workflow in OpenProse can deconstruct successful agent runs and convert them into reusable .prose.md workflows, addressing the reliability problem where agents perform inconsistently. | Evidence: Weitekamp built an OpenProse program that takes a golden Claude/Codex/PY session and has the agent deconstruct it into a reusable workflow involving recursive coding agents for reliable reproduction. This addresses the 'one day SaaS app, next day empty wallet' problem by capturing what worked and making it repeatable. | Caveat: No examples of deconstructed workflows or success rate of reproduction. Unclear how the system handles sessions that partially succeeded or required human intervention. No mention of how to identify 'golden' vs. 'good enough' sessions or whether this works for novel tasks vs. only similar repeats. | Implication: For Ken: This is the killer feature for production deployment—convert one-off wins into systematic capabilities. Prioritize building session recording/replay infrastructure if deploying agents at scale. Consider pairing with adversarial verification (red-team agents) to stress-test deconstructed workflows. Potential to build a library of validated, reusable agent workflows for common tasks (migrations, refactors, audits). | Timestamp: 22:10
  • Claim: RLMs are 'too hot to benchmark' because they achieve results so far beyond frontier models that benchmark maintainers refuse to validate them or create separate leaderboards to avoid contamination. | Evidence: Symbolica's Agentica scored 30%+ on ARC-AGI vs. 2-3% for frontier models within hours of release, but ARC Prize gave a 'consolation tweet' and refused private evaluation because they 'don't like RLM harnesses.' Weitekamp and Alex Zhang (RLM co-author) convinced LongCoT maintainers to create a separate 'open harness' leaderboard so RLM results wouldn't contaminate the no-tool-calling intent. | Caveat: Benchmark rejection suggests philosophical disagreement over what counts as 'solving' vs. 'brute forcing' problems. RLMs may be optimizing for benchmark performance via tool scaffolding rather than demonstrating general reasoning. Separation of leaderboards makes it hard to compare RLM vs. pure LLM approaches apples-to-apples. | Implication: For Ken: RLMs may be overfitting to benchmarkable tasks or the benchmarks themselves may not capture the value prop. Focus on real-world reliability and outcome delivery rather than benchmark scores. If evaluating agents, design custom benchmarks that reflect actual use cases (e.g., 'successfully refactor this codebase without breaking tests') rather than relying on academic leaderboards. The political/philosophical resistance suggests RLMs challenge existing evaluation paradigms. | Timestamp: 08:50

Detailed Brief

The Reliability Problem and Mismanaged Genius Thesis

  • Claims: Agents are intelligent enough (know the entire internet) but can't reliably deliver outcomes, blocking trust and deployment; The bottleneck is not more intelligence but better specification, management, reuse, and verification of work; Today's agents are 'mismanaged geniuses' per Alex Zhang, Zed Li, and Omar Khattab (MIT, RLM co-authors)
  • Evidence: Personal example: Claude Code built working SaaS app from one prompt, then emptied Solana wallet the next day; Vision from AI Engineer conference (Nov 2025): moving from manual agent wrangling to 'meditating while things manifest'; Weitekamp's experience at Raw Works (independent research) and OpenProse (current role)
  • Caveats: Solana wallet incident lacks context (bug vs. prompt injection vs. user error unclear); No quantified failure rates or systematic analysis of when/why agents fail; Anecdotal evidence from one practitioner rather than industry-wide data
  • Implications: Production AI systems need architectural focus on decomposition, verification loops, and reusability rather than chasing frontier model capabilities; The 'behavioral textual orchestration' layer is where investment should focus for reliability gains; Trust is a function of consistency, not peak performance—agents need to be boringly reliable

RLM Core Mechanics and Paradigm Shift

  • Claims: In RLMs, context itself is the object of computation via symbolic manipulation; RLMs are a marriage of tool calling and reasoning where code execution is reasoning; The full prompt is a variable (file or many files), not a simple user query; agent operates symbolically on it via REPL; RLMs represent a new paradigm of test-time/inference-time compute, unified tool calling + reasoning; RLMs evolved from chain-of-thought prompting → reasoning models → function/tool calling → unified recursive approach
  • Evidence: Original Oolong paper: RLMs process tens of millions of tokens, orders of magnitude beyond context windows; Weitekamp's results: default RLM harness is top-10 memory system without modifications; LongCoT benchmark: Qwen 3.5 9B as RLM beat Opus and GPT-4 as LLMs on long reasoning tasks requiring sustained thread; dspi.rlm implementation achieved state-of-the-art on long reasoning tasks
  • Caveats: No latency/cost analysis for recursive decomposition vs. single long-context calls; Benchmark maintainers created separate leaderboards, suggesting RLM results aren't directly comparable to pure LLM approaches; Recursion depth, failure modes, and convergence guarantees not discussed
  • Implications: For massive context tasks (codebases, document analysis), RLMs offer alternative to hitting context limits; Trade-off may be latency/orchestration complexity for reliability and ability to handle unbounded context; Consider dspi.rlm for prototyping/benchmarking; it's Weitekamp's go-to for achieving results; Test-time compute shifts from 'think longer in latent space' to 'decompose, execute, verify symbolically'

RLM Rubric and What Counts (Philosophical Clarity)

  • Claims: To be an RLM: executable environment, externalized prompt, code calling the model, model picks decomposition, state stays symbolic; Plain LLMs, RAG, coding agents with loops get close but aren't RLMs without agent-controlled decomposition; Hardcoded map-reduce (e.g., Lambda RLM project) doesn't count—agent must decide decomposition; Claude Code became an RLM only with dynamic workflows (Feb 2025) because agent can now choose decomposition
  • Evidence: Weitekamp built rubric in companion GitHub repo with criteria breakdown; Omar Khattab congratulated Anthropic for making Claude Code 'finally an RLM' with dynamic workflows; Lambda RLM decomposes via lambda calculus into map-reduce, but LLM doesn't decide the decomposition structure; Weitekamp wrote two Claude workflows: one explicitly not RLM (hardcoded map-reduce), one that is (agent picks decomposition for deep research)
  • Caveats: Rubric isn't meant to 'start fights or nitpick' but clarify essence—there's philosophical debate over definitions; Dynamic workflows are Claude-only currently; other agents lack equivalent; The line between 'agent picks decomposition' and 'follows hardcoded pattern' can be fuzzy in practice
  • Implications: For Ken: When evaluating or building agent systems, the key differentiator is whether the agent adaptively decomposes problems vs. follows predetermined patterns; Hardcoded orchestration (map-reduce, DAGs) isn't enough—need agent to reason about problem structure and choose approach; This explains why some impressive-looking agent systems aren't RLMs: they execute complex workflows but don't control the workflow structure

Implementation Paths: Claude Dynamic Workflows, OpenProse, YPI

  • Claims: Claude Code dynamic workflows (Feb 2025) enable RLM patterns with 'A Harness for Every Task' showing six patterns; OpenProse is markdown-based 'programming language' compiled by coding agents, works with any agent (Codex, Claude, PY); OpenProse can declare sub-agent work, skills, and tool dependencies explicitly, ensuring sub-agents have required capabilities; YPI is Y-combinator recursive PY wrapper, now achievable as pure PY extension without forking; Other notable implementations: dspi.rlm (Weitekamp's benchmarking go-to), Ax (TypeScript RLM with agent-native interfaces), UNIX RLM (pure bash)
  • Evidence: Weitekamp wrote two Claude workflows demonstrating RLM vs. non-RLM patterns in companion repo; OpenProse repo includes harness supporting Codex SDK and Claude Code; 'prose write' command generates .prose.md files; Two demo .prose.md files in companion repo show problem decomposition and parent verification; PY evolved so recursive extension works without forking; PY designed by Mario to be minimal/extensible; Ax (TypeScript) enables ax agent to write TypeScript interface to another ax agent, full recursive descent; Dan at OpenProse made UNIX RLM with pure bash and Linux file system as environment
  • Caveats: OpenProse is early-stage open source, adoption/maturity unclear; No performance comparisons between OpenProse and native dynamic workflows; YPI recursion depth limits, failure modes, convergence not discussed; PY minimalism means more setup/extension work vs. batteries-included agents; Ax TypeScript approach may be niche for TypeScript-first teams
  • Implications: For Claude users: dynamic workflows are the fastest path to RLM capabilities now; For cross-platform or custom systems: OpenProse offers unified abstraction with dependency management; For research/prototyping or full control needs: PY + recursive extension provides lightweight, extensible base; For benchmarking/achieving SOTA: dspi.rlm is proven implementation; The REPL can be anything (Python, TypeScript, bash)—choose based on problem domain and existing tooling

Practical Applications and Use Cases

  • Claims: Scale migrations/refactors: parallel swarms merge work together; Deep recursive analysis: go after directory, recursively process/analyze/research inside; Adversarial verification: skeptical agents or red-team swarms improve system in parallel; Audits and bug sweeps: systematic recursive checking; Golden session deconstruction: convert successful one-off runs into reusable .prose.md workflows
  • Evidence: Scale migration was Anthropic's launch post example for dynamic workflows; Weitekamp built two examples each for Claude dynamic workflows and OpenProse repo; Golden session deconstruction system in OpenProse can take Claude/Codex/PY session and agent deconstructs it into reusable workflow; Examples include refactoring huge codebases, directory-level deep research, adversarial improvement
  • Caveats: No success rates, time-to-completion, or cost metrics for practical applications; Golden session deconstruction examples not shown; unclear how it handles partial success or human intervention; No guidance on identifying 'golden' vs. 'good enough' sessions or applicability to novel vs. repeated tasks
  • Implications: For Ken's agent systems: prioritize session recording/replay infrastructure to capture and systematize wins; Pair golden session capture with adversarial verification (red-team agents) to stress-test before deployment; Build library of validated, reusable agent workflows for common tasks (refactors, migrations, audits, bug sweeps); Deep recursive directory analysis useful for codebase understanding, documentation generation, or compliance checks; Adversarial patterns (skeptical agents) address reliability concern by building verification into workflow

Benchmark Drama and Evaluation Challenges

  • Claims: RLMs are 'too hot to benchmark'—results so far exceed frontier models that maintainers reject or segregate them; Symbolica's Agentica scored 30%+ on ARC-AGI vs. 2-3% for frontier models within hours, ARC Prize refused private evaluation; Weitekamp and Alex Zhang convinced LongCoT maintainers to create separate 'open harness' leaderboard to avoid contamination; Weitekamp's take: 'I don't care' about latent space vs. reasoning tokens vs. code execution—wants results and reliable AI programs
  • Evidence: Symbolica's Agentica result on ARC-AGI 3 within hours of release using RLM harness; ARC Prize 'consolation tweet' saying congrats but 'didn't solve the right way,' refusing full private evaluation; LongCoT benchmark maintainers created separate leaderboard with no-tool-calling constraint; Omar Khattab (RLM co-author) and Weitekamp both expressed frustration with benchmark gatekeeping
  • Caveats: Benchmark rejection suggests philosophical disagreement over what counts as solving vs. scaffolding/brute-forcing; RLMs may be optimizing for benchmark performance via tool access rather than demonstrating general reasoning; Separate leaderboards make apples-to-apples comparison difficult; Some view RLM approach as 'cheating' by using tools, others see it as pragmatic progress
  • Implications: For Ken: Don't rely on academic benchmarks to evaluate agent systems—design custom benchmarks reflecting real use cases; RLMs may be overfitting to benchmarkable tasks or benchmarks may not capture value prop; Focus on outcome delivery and reliability in production rather than leaderboard position; Political/philosophical resistance to RLMs suggests they challenge existing evaluation paradigms—may be a feature, not a bug; Consider building internal evaluation harnesses for tasks like 'refactor codebase without breaking tests' or 'implement feature from spec with <X errors'

Notable Concepts & Terms

  • RLM (Recursive Language Model): Paradigm where context itself is object of computation; agent symbolically manipulates context via REPL, recursively spawns sub-agents to decompose problems, state stays symbolic. Unifies tool calling and reasoning.
  • Mismanaged Genius: Weitekamp's thesis (from Alex Zhang, Zed Li, Omar Khattab at MIT): agents have sufficient intelligence but lack specification, management, reuse, and verification layers needed for reliable outcomes.
  • dspi.rlm: DSPy implementation of RLMs; Weitekamp's go-to for benchmarking, achieved state-of-the-art on long reasoning tasks. Python-based, proven for SOTA results.
  • Dynamic Workflows (Claude): Anthropic feature (Feb 2025) that made Claude Code an RLM by letting the agent decide problem decomposition. 'A Harness for Every Task' blog post shows six patterns.
  • OpenProse: Markdown-based 'programming language' compiled by coding agents (not computers). Logical English spec that can turn any agent with file system + sub-agents into an RLM. Supports declaring sub-agent work, skills, and tool dependencies.
  • YPI (Y-combinator PY): Recursive wrapper for PY coding agent where PY calls PY calls PY. Named after lambda calculus Y combinator. Now achievable as pure PY extension without forking. Demonstrates fully recursive coding agent.
  • Golden Session: Successful agent run that produced desired outcome. Weitekamp built OpenProse system to deconstruct golden sessions into reusable .prose.md workflows for reliable reproduction.
  • LongCoT Benchmark: Benchmark where problems are hard because they require extremely long reasoning chains. Frontier models can't hold thread; Qwen 3.5 9B as RLM beat Opus/GPT-4 as LLMs.
  • Agentica (Symbolica): RLM agent harness from Symbolica team that scored 30%+ on ARC-AGI 3 within hours (vs. 2-3% for frontier models), causing benchmark controversy and ARC Prize refusal to fully evaluate.
  • PY Agent: Minimal, extensible coding agent by Mario designed for users to write extensions rather than feature-stuffing core. Weitekamp used it to demonstrate pure-extension recursive agent.
  • UNIX RLM: Implementation by Dan at OpenProse where REPL is pure bash and environment is Linux file system, demonstrating RLMs can use any REPL, not just Python.
  • Ax (TypeScript RLM): TypeScript variation on DSPy with RLM implementation where ax agent can write full TypeScript interface to another ax agent, enabling recursive descent with typed interfaces.

Operator Notes / Why Ken Should Care

  • This is a strategic talk for production agent deployment: the 'golden session to workflow' feature directly solves the 'one day great, next day disaster' problem Ken likely faces with AI coding agents.
  • If Ken's team uses Claude Code, dynamic workflows are available now and worth immediate exploration for RLM capabilities.
  • For cross-platform or custom agent systems, OpenProse offers a vendor-neutral path with explicit dependency management (skills/tools for sub-agents)—useful if Ken has heterogeneous agent portfolio.
  • The benchmark drama (ARC Prize refusing to evaluate, separate leaderboards) signals that RLMs are powerful enough to disrupt evaluation paradigms—Ken should design internal benchmarks for real use cases rather than chase academic leaderboards.
  • The 'mismanaged genius' framing is excellent for explaining to stakeholders why throwing GPT-5 at the problem won't fix reliability—it's an architecture and orchestration problem, not an intelligence problem.
  • Practical applications (scale migrations, adversarial verification, bug sweeps, golden session capture) map directly to software development workflows Ken's team likely needs.
  • dspi.rlm is the proven implementation for SOTA results—worth prototyping with if Ken's exploring RLM patterns in research or internal tools.
  • The recursive depth and verification patterns (parent agent verifies sub-agent work) are directly applicable to Ken's agent systems for improving reliability and trust.
  • OpenProse's explicit dependency declaration (sub-agents need specific skills/tools) addresses configuration management problem in multi-agent systems.
  • Weitekamp's emphasis on 'I don't care about latent space vs. code execution, I want results' aligns with pragmatic operator perspective Ken likely shares.

Watch Map

  • 00:00: Introduction and motivation: reliability problem, Solana wallet incident, 'mismanaged geniuses' thesis
  • 04:30: What are RLMs: context as computation object, marriage of tool calling and reasoning, symbolic manipulation via REPL
  • 06:15: RLM results: Oolong (millions of tokens), top-10 memory system, LongCoT benchmark (Qwen 9B beats GPT-4/Opus)
  • 08:50: Benchmark drama: Symbolica's 30%+ on ARC-AGI, ARC Prize refusal, LongCoT separate leaderboard
  • 11:20: RLM rubric: what counts as an RLM vs. coding agents, hardcoded map-reduce doesn't qualify
  • 13:00: Recursive coding agents: applying RLM principles to coding agents, YPI (recursive PY wrapper) demo
  • 15:45: Notable implementations: dspi.rlm, Ax (TypeScript), UNIX RLM (bash), OpenProse harness
  • 17:15: Is Claude Code an RLM: dynamic workflows (Feb 2025) made it one, Omar Khattab's congratulations
  • 19:30: OpenProse deep dive: markdown language compiled by agents, skills/tools as dependencies, 'prose write' command
  • 22:10: Practical applications: scale migrations, deep recursive analysis, adversarial verification, golden session deconstruction
  • 23:15: Takeaways: trust=reliability, RLMs as new test-time compute paradigm, coding agents can be RLMs with right patterns
  • 23:45: Conclusion: recurse responsibly, check out recursivecodingagents.com

Source/Metadata

  • Title: Recursive Coding Agents - Raymond Weitekamp, OpenProse
  • Transcript words: 7442
  • Duration seconds: 1428
  • Timestamp note: Timestamps provided based on transcript cues and standard talk pacing; not explicitly marked in transcript, inferred from content flow.
Full transcript 3530 words · 33 min read
0:00

SPEAKER_00

Hello there! My name is Raymond Weidekamp and today I'm going to talk about recursive coding agents which is this idea of applying the lessons of recursive language models, RLMs, to coding agents. This is some work that I have done both in my independent research, raw works, and also more recently in my role at OpenPros. So to motivate this a little bit, we all want outcomes. We all want agents that are working on our behalf. We want reliable co-workers that are getting things done while we're doing something fun, while we're out on a hike, while we're chilling, while we're doing the do. And my argument and my experience is that the bottleneck to this is not intelligence. The models are intelligent enough. They know all kinds of things. They know the entire internet, but they can't reliably deliver outcomes. And so I can't trust them. So as a very simple example, one day I get almost a fully working SaaS app from a single prompt, granted a long prompt. The next day, and I swear this actually happened, cloud code empties the entire contents of my Solana wallet. Oops! Okay. So that doesn't really instill trust. So at the bottom here, we've got this progression. Okay. And we all want to move towards the one on the right where we're sitting there and meditating and things are manifesting. And so where does that come from? This is from the AI engineer code. It's actually from the back of the t-shirt engineer code, November 2025, man. I hope you're there. If you weren't watching on YouTube, it was amazing. So here's the thesis. The thesis is today's agents are mismanaged geniuses. The intelligence is there and the missing layer is how do we specify and manage and reuse and verify the work? So this framing, the phrase the mismanaged genius, comes from Alex Zhang, Zed Li and Omar Khatab at MIT. And Alex and Omar are part of the authors of the original recursive language models paper. I've also talked a little bit about this recently on turning post. I forgot to mention that these slides are actually a website, recursive coding agents.com. So you can click on them by going to this website. So everything I'm going to show in here is interactive. Okay. What are recursive language models? So I like to say that in an RLM, the context itself is the object of computation. And this is essentially a marriage of tool calling and reasoning. We're going to talk more about that in the next slide, but the idea is that the full prompt is not a simple user query. The full prompt is a variable. The full prompt could be a file or many files. And we have this read evaluate print loop repl that the agent is interacting with in the original paper. That's Python. And the RLM is instructed to operate symbolically on that prompt. So don't just read the whole thing into your context window, explore it symbolically and even more, you don't even directly export symbolically, or maybe you do a little bit of poking around, but have other LLMs. And I guess other RLMs, if you allow the recursion depth to be greater than one, have these other recursive sub agents. And again, we'll get a little bit into the weeds of the lingo, sub RLMs, sub LLMs, do this symbolic manipulation to pick apart the answer and then work our way back up to a final answer. So it looks something like this in this tree below. So my take here is that RLMs are the new reasoning models. And I see this as the next paradigm of test time compute, inference time compute, whatever you want to call it. And why does it seem obvious or maybe like, hey, why is this even a thing? I think it's very elegant because it's a very elegant marriage of two things, reasoning and code execution. So the code execution is reasoning. And so instead of we had long chain of thought as a prompting strategy that evolved into reasoning models that explicitly expressed the chain of thought as their reasoning tokens. We already had function calling, tool calling, parallel tool calling. And RLMs really puts that together in a way that gets amazing results. So three very simple examples, one from the original paper, Oolong, the RLMs can process information that is many orders of magnitude larger than their context window. Tens of millions, millions of tokens. What I showed in my own independent work was that the default RLM harness is itself a really powerful memory system. So RLM with no modifications is essentially like a top 10 memory system and up there with all the people custom making memory systems. And there's probably billions of dollars going into that. And with a little bit of modification, you can get really amazing results using it as memory. I was also able to show state-of-the-art results where the RLM framework and specifically the DSPi implementation of it was able to get state-of-the-art results on long reasoning tasks. Now there's this new benchmark, long COT. I won't go into the details in depth, but the idea of this benchmark was that the problems are hard specifically because they require so many steps of reasoning, in the analysis sequence or the chain of thought, that most reasoning models, including the top ones can't hold the thread for long enough. If you allow the RLM to solve the problem using a combination of code and recursive calls to sub-agents, then a very small model, QUEN 3.5 9b, you could run this on a laptop, can actually beat, so QUEN 3.5 9b as an RLM can beat OPUS and GPT 5.4, all the top frontier models as LLMs on these long reasoning tasks. So they're extremely, extremely powerful. So powerful that they are arguably too hot to benchmark. So two examples here on the left, a very high profile case where the Symbolica team has this RLM agent harness called Agentica. Within hours of ARC AGI 3 being released where the top scores of all the frontier models were around two or three percent, the Symbolica team showed thirty something percent. This is crazy. They blew it out of the water within hours using RLMs as a framework. So much so that it's very much upset the ARC Prize team. And so they gave them what I'm interpreting as a consolation tweet, which as far as I'm reading the situation was essentially saying, congratulations, but you didn't solve the problem the right way. And we don't like RLM harnesses. And so you can have this nice tweet, but we refuse to actually do the full private part of the ARC AGI evaluation,

0:05

SPEAKER_00

Two or three percent. The Symbolica team showed thirty something percent. This is crazy. They blew it out of the water within hours using RLMs as a framework. So much so that it's very much upset the ARC Prize team. And so they gave them what I'm interpreting as a consolation tweet, which as far as I'm reading the situation was essentially saying congratulations, but you didn't solve the problem the right way. And we don't like RLM harnesses. And so you can have this nice tweet, but we refuse to actually do the full private part of the ARC AGI evaluation, which to me is just insane. In my own work and on the long COT benchmark, my results as well as Alex from MIT, the RLM first author, encouraged the leaderboard maintainers to actually make a separate open harness leaderboard, so that the results of the RLMs could be showcased without contaminating the original intent of the leaderboard, which was basically no tool calling is allowed. So my take on this is I don't care. I don't care whether it's latent space or reasoning tokens or code execution. I want results and I want AI programs that get those results. Okay. So this can feel close to a lot of other things. And I built a little rubric. There's a companion GitHub repo for this that you can go through if you want to see. And so what do we need to be an RLM? We have an executable environment. The prompt is externalized. There's code. That's actually the thing calling the model. The model is able to pick the decomposition of the problem into the sub calls or sub agents. And the state itself is staying symbolic, right? So obviously plain LLMs and RAG and things like that don't meet those coding agents and sub agents and loops. They get close, but they're not quite there. And again, the rubric here is not to start fights or nitpick. It's just trying to explain what's the essence of RLM and recursive coding agents. Another example that's close, but in a cigar would be hard coded map reduce. And I would put this project called Lambda RLM in that category, which is essentially a way of saying okay, I'm decomposed the problem using Lambda calculus into a map reduce. And then there are like LLM calls in that executing the map reduce, but the LLM is not deciding or the RLM is not deciding how to decompose the problem. And that I see as a key element of this that makes it very agent native, you might say. Okay. So now RLMs, what about recursive coding agents? Okay. It looks the same to me. We just swap RLMs and LLMs for agents and sub agents. And don't we have the same thing? And yeah, you do. And I think that you could take this perspective of trick question. Like RLM is a coding agent and it's already recursive. So that's fine. And I don't think that argument is wrong, but it doesn't really move anything forward. And so what I'm interested in is this question of how can we apply the principles of RLMs to coding agents and make them actually useful for coding agents. And I've been very obsessed with this problem since the October RLM blog post came out in 2025. Okay. So I show some of my experiments on recursive coding agents. The first one was simply wrapping Alex's RLM package as a CLI. So the idea was I just want to give my coding agent an RLM as a tool call. So let's say we need to go sift through a hundred million token corpus. Well now it just used this tool. And then the RLM does the RLM thing. So that's interesting and it can be very useful. And a few other people have built things like that. But then I thought, well, what would it really mean? What would be possible or how could it be possible to make the coding agent fully recursive? So the coding agent harness calls itself, like the exact version of itself. And how might you implement that? And that is what I called YPI. Y stands for the Lambda calculus Y combinator and PI in case you haven't heard of it is a really awesome coding agent. It's incredibly minimal and it's specifically designed to be extensible. So Mario wants you to write extensions for PI for new features that you would like rather than trying to stuff your ideas into the main agent. So it's meant to me a very minimal core that you extend however you like. And when I originally had the recursive coding agent idea and wanted to do it with PI, I was not able to use PI extensions to achieve this goal. And so I had to fork it instead. I'm very excited to report that in anticipation of this talk, I revisited this and now PI has evolved, and the PI extensions have evolved such that you can make it fully recursive with a pure extension. So I have both the pure recursive extension, PI recursive, package as well as the YPI wrapper. That is a convenience wrapper for this. So this is very quite literally a recursive coding agent in the sense that PI calls PI calls PI calls PI. You can set the depth however you want. And now I want to show a few other notable projects in the space. So obviously there's the original implementation from Alex, our alum, dspi.rlm is my go-to, especially when I'm doing benchmarking. And that's how I got all these amazing results on some of these benchmarks. Axe, I think is a very interesting one because it's incredibly agent native. So the ax started out as a TypeScript variation on dspi when our limbs came out, they obviously implemented it. But they did it in this way that enables the ax agent to write a whole TypeScript interface to another ax agent and go all the way down the recursive rabbit hole, which I think is really cool and very interesting just to showcase that this REPL could be anything. There's an example from Dan at Open Pros who made the UNIX RLM. This is pure bash and the environment is just the Linux file system. So that's a whole other angle of thinking about what's possible with RLM. And then lastly, and we'll talk more about Open Pros, but Open Pros as a language actually enables you to convert any coding agent into an RLM. And I'll talk a little bit more about how to do that towards the end of the talk. And then the Open Pros repo also contains a harness that executes. It's a coding agent harness that will let you use codex SDK or cloud code underneath and do an RLM style execution of these pros programs.

0:13

SPEAKER_00

This is pure bash and the environment is just the Linux file system. So that's a whole other angle of thinking about what's possible with RLM. And then lastly, and we'll talk more about open pros, but open pros as a language actually enables you to convert any coding agent into an RLM. And I'll talk a little bit more about how to do that towards the end of the talk. And then the open pros repo also contains a harness that executes. It's a coding agent harness that will let you use codex SDK or cloud code underneath and do an RLM style execution of these pros programs. So is cloud code and RLM. This is the question that keeps getting asked and the original answer on the day of release day of the blog post was no, no, no, it's not. But it was literally the first question that was asked on the very same day to the launch tweet. It's saying, Hey, this is cloud code sub agents, right? Go back to my rubric. If you want to dig into some of the nitty gritty details. But arguably now it is. So over here on the right, we can see Omar saying, Hey, congratulations. Anthropic, cloud code is finally an RLM now that you have dynamic workflows. So what changed? And I think this is an interesting way of explaining what's powerful about recursive coding agents and what RLMs even are, by using this example. So dynamic workflows were released just a few weeks ago. And they make cloud code recursive or capable of doing these recursive workflows. And I would highly encourage you to read this blog post called a harness for every task. It shows six different workflow patterns that are very powerful. Obviously there's many more that you can achieve. And just to show it, I wrote two workflows for cloud code, one that is explicitly not an RLM. So that's a hard coded map reduce workflow. And one that I'm arguing is, you could think of it as deep research over a file system. So pick a handle, assign that to an agent, have it go do some analysis, bring back what it finds, et cetera. Again, these are in the companion repo. And then now to open pros. So dynamic workflows are cool. They're only in cloud code. They're also not the only way to do this. So what if you don't like cloud code or what if you do like cloud code, but you don't want to use dynamic workflows, you want something else. So this is what open pros is all about. So open pros is technically a programming language, but it is not compiled by your computer. It's compiled by your coding agent. It's a markdown spec. It's logical English. You don't need to learn any crazy syntax. And there's actually a command pros, right? That will get cloud code or codex or your favorite or PI or your favorite coding agent to write a dot pros.md file for you. So in that way, it's similar to the cloud code ultra code command where decide and write the workflows for you. And the pros has the ability to turn any agent that's got a file system and sub agents into an RLM. So this is an open source repo. You can check it out. And I've also written a little bit more in depth about this for Turing post in this article. And the key thing that I want to bring up with regards to the RLMs is that open pros can explicitly declare the sub agent work. So again, I've made two demo pros.md files that are in this companion repo. If you want to dig into the code, I'm not going to do that here in the slides, where you can break a problem up into smaller pieces that are assigned to sub agents, verify the work of those sub agents in the parent agent session. And what's even cooler is you can actually in open pros the features that I added to the language are that you can add skills and tools as explicit dependencies. So you can imagine a workflow where a certain sub agent needs a very specific skill to do its role in the workflow or must have access to a certain CLI tool, for example, or it can't run and do its job. And so there's a way in pros to actually wire those in as dependencies to ensure that not only is the way the work is done what you want, but actually that the sub agents are specifically configured with the tools and skills that they need to successfully do the work that you are declaring in the pros contract. Okay. Super cool. What can you actually do? So I've got two examples from cloud dynamic workflows, two examples from open pros repo. Scale migrations. This was the launch post example. Refactor, a huge thing, all with a big swarm and parallel and then merge the whole thing together. That's super cool. This idea of going after a directory and then deep research or deep analyze or deep process in some way recursively inside that is another example I have here. You can do audits, bug sweeps. You can do adversarial things such as having a skeptical agent or a red team set of agents that are going to try to improve the system adversarially or in parallel. And then one really cool thing that you can do with open pros that I just added recently is it goes back to my very first slide, right? So one day I get this amazing result the next day they trade away all my crypto currency, which is very small, thankfully. But how do we get these things to be more reliable and how do we get them to be, you know, we have a golden session and we have a great day and now we want to capture that and reuse it over and over again. So I built a system where you can take a golden session for cloud code code, codex, PI, whatever you want. And it's a pros program that will actually have the agent deconstruct that session and turn it into a reusable pros workflow. That again can involve this idea of recursive coding agents to get you to a reliable way of getting to that golden state of performance over and over again. Recursive coding agents for the win. I really think this is very powerful. RLMs just blew my mind when they first came out. And again, as you can probably see through this talk, I've been absolutely obsessed with applying the ideas of RLMs to coding agents. The three things that I hope you'll take away from this is one, trust is reliability. How can we trust something that isn't reliable? And again, my argument and this idea of the mismanaged genius is that the next step is not more raw intelligence. It's actually behavioral textual orchestration. I personally believe that our

0:19

SPEAKER_00

of performance over and over and over again. Recursing, recursive coding agents for the win.

0:30

SPEAKER_00

I really think this is very powerful. Our LMS just blew my mind when they first came out. And again, as you can probably see through this talk, I've been absolutely obsessed with applying the ideas of our LMS to coding agents. The three things that I hope you'll take away from this is one, trust is reliability. How can we trust something that isn't reliable? And again, my argument and this idea of the mismanaged genius is that the next step is not more raw intelligence. It's actually behavioral textually orchestration. I personally believe that our LMS represent this new paradigm of test time, compute inference time, compute, where tool calling and reasoning are unified and we reason through tool calling and we can recursively iterate. And one of those tools is to call another agent to go do it on some other specific task or subset of the problem. And then also, I hope we settle a little bit of this drama around, wait, our LMS, aren't they just coding agents? Yes. Coding agents can be our LMS. They aren't automatically our LMS. And so I've showed a couple of different cloud code dynamic workflows that can turn cloud code into an RLM, as well as some ways of doing this with open pros that you can use with any coding agent. So I see this as an incredibly powerful way of working with coding agents. I hope you will dig in more to recursive coding agents.com, but with great power comes great responsibility. So until next time, please recurse responsibly. Thank you very much.

0:36

SPEAKER_00

that are getting things done while we're doing something fun, while we're out on a hike, while we're cold chilling, while we're doing the do. And my argument and my experience is that the bottleneck to this is not intelligence. The models are intelligent enough. They know all kinds of things. They know the entire internet, but they can't reliably deliver outcomes. And so I can't trust them. So as a very simple example, you know, one day I get almost a fully working SaaS app from a single prompt, granted a long prompt. The next day, and I swear this actually happened,

1:20

SPEAKER_00

cloud code empties the entire contents of my Solana wallet. Oops! Okay. So that doesn't really instill trust. So, uh, at the bottom here, we've got this pro this progression. Okay. And we all want to move towards the one on the right where we're just sort of sitting there and meditating and, and things are manifesting. And so where does that come from? This is from the AI engineer code. It's actually from the back of the t-shirt engineer code, November, 2025, man. I hope, I hope you're there. If you weren't watching on YouTube, it was, it was amazing. So here's the thesis.

1:58

SPEAKER_00

The thesis is today's agents are mismanaged geniuses. The intelligence is there and the missing layer is how do we specify and manage and reuse and verify the work? So this, uh, framing this phrase, the mismanaged genius, uh, comes from Alex Zhang, Zed Li and Omar Khatab at MIT. Um, and Alex and Omar are, uh, part of the authors of the original recursive language models paper. Uh, I've also talked a little bit about this recently on turning post. Um, I forgot to mention that these slides are actually a website, recursive coding agents.com. So you can click on them, uh, by going to this website. So everything

2:40

SPEAKER_00

I'm going to show in here is, is interactive. Okay. What are recursive language models? So I like to say that in an RLM, the context itself is the object of computation. Um, and this is essentially a marriage of tool calling and reasoning. We're going to talk a lot more, more about that in the next slide, but the idea is that the full prompt is not a simple user query. The full prompt is a variable. The full prompt could be a file or many files. Um, and we have this read evaluate print loop repl, um, that the agent is interacting with in the original paper. That's Python. And the RLM is instructed to operate

3:27

SPEAKER_00

symbolically on that prompt. So don't just read the whole thing into your context window, um, explore it symbolically and, uh, even more, you don't even directly export symbolically, or maybe you do a little bit of poking around, but have, uh, other LLMs. Uh, and I guess other RLMs, if you allow the recursion depth to be, to be greater than one, uh, have these, uh, other recursive, um, sub agents. And again, we'll get a little bit, uh, a little bit into the weeds of the lingo, um, sub RLMs, sub LLMs, uh, do this symbolic manipulation to pick apart the answer and then work our way back up to a final answer. So it looks

4:15

SPEAKER_00

something like, like this in this, in this tree below. So my take here is that RLMs are the new reasoning models. And I see this as the next paradigm of test time, compute inference time, compute, whatever you want to call it. And why does it seem obvious or maybe like, Hey, why is this even a thing? Um, I think it's very elegant because it's a very elegant marriage of two things, reasoning and code execution. So the code execution is reasoning. Um, and so instead of we, we had long, um, we had chain of thought as a prompting strategy that evolved into reasoning models that explicitly

4:58

SPEAKER_00

expressed the chain of thought as their reasoning tokens. We already had function calling, tool calling, parallel tool calling. Um, and RLMs really puts that together in a way that gets amazing results. So three very simple examples, one from the original paper, Oolong, the RLMs can process information that is many orders of magnitude larger than their context window. Tens of millions, millions of tokens. Uh, what I showed in my own independent work was that the default RLM harness is itself a really powerful memory system. Um, so RLM with no modifications is essentially like a top 10

5:39

SPEAKER_00

memory system and like, you know, up there with all the people custom making memory systems. And there's probably billions of dollars going into that. Um, and, uh, with a little bit of modification, you can get really amazing results, uh, using it as memory. I was also able to show state-of-the-art results where the RLM framework and specifically the DSPi implementation of it was able to get state-of-the-art results on long reasoning tasks. Now there's this new benchmark, long COT. I won't go into the details in depth, but the idea of this benchmark was that the problems are hard specifically because they

6:21

SPEAKER_00

require so many, um, steps of reasoning, uh, in, in the like analysis sequence or the chain of thought, um, that most, uh, reasoning models, including the top ones can't hold the thread for long enough. Um, if you allow the, uh, RLM to solve the problem, uh, using a combination of code and recursive calls to sub-agents, then a very small model, QUEN 3.5 9b, you could run this on the laptop, uh, can actually beat, so QUEN 3.5 9b as an RLM can beat OPUS and, um, and GPT 5.4, all the top frontier models as LLMs on these long reasoning tasks. So they're extremely, extremely powerful.

7:15

SPEAKER_00

So powerful that they are arguably too hot to benchmark. So two examples here on the left, a very high profile, uh, case where, um, the Symbolica team has this RLM agent harness called Agentica. Within hours of ARC AGI 3 being released where the top scores of all the frontier models were around two or 3%, the, uh, Symbolica team showed 30 something percent. This is crazy. They blew it out of the water within hours using RLMs as a framework. So much so that it's very much upset the ARC Prize team. And so, uh, they gave them what I'm interpreting as a consolation tweet, which as far as I'm reading the situation was essentially saying, you know, la-di-da, congratulations,

8:08

SPEAKER_00

but you didn't solve the problem the right way. And we don't like RLM harnesses. And so, uh, you can have this nice tweet, but we refuse to actually do the full private part of the ARC AGI evaluation, uh, which to me is just insane. Uh, in my own, uh, work and, and on the long COT benchmark, um, my results as well as Alex, uh, from MIT, the RLM first author, uh, encouraged, let's say the, uh, the, uh, leaderboard maintainers to actually make like a separate open harness leaderboard, so that the results of the RLMs could be showcased without contaminating the original intent of the leaderboard, which was basically no tool calling is allowed. So my take on this is I

8:57

SPEAKER_00

don't care. I don't care whether it's latent space or reasoning tokens or code execution. I want results and I want AI programs that get those results. Okay. So this can feel close to a lot of other things. And I, I built a little rubric. There's a companion GitHub repo for this that you can go through if you want to see. Um, and so what do we need to be an RLM? We have an executable environment. The prompt is externalized. There's code. That's actually the thing calling the model. The model is able to pick the decomposition of the problem into the sub calls or sub agents. And the state itself is staying symbolic, right? So obviously plain LLMs and rag and things like

9:39

SPEAKER_00

that don't, don't meet those coding agents and sub agents and loops. They get close, but they're not quite there. And again, the rubric here is not to like start fights or nitpick. It's just trying to explain like what, what's the essence of, of RLM, uh, and recursive coding agents. Another example that's close, but in a cigar would be hard coded map reduce. And I would put, um, uh, this, this project called Lambda RLM in that category, which is essentially a way of saying, okay, I'm decomposed the problem, uh, using Lambda calculus into a map reduce. And then there are like LLM calls in that executing the map reduce, but the, but the LLM is not deciding or the RLM is not

10:21

SPEAKER_00

deciding how to decompose the problem. And that I see as like a key element of this that makes it very agent native, you might say. Okay. So now RLMs, what about recursive coding agents? Okay. It looks the same to me. We just swap RLMs and LLMs for agents and sub agents. And don't we have the same thing? Uh, and yeah, you do. And I think that you could take this perspective of ha ha, trick question. Like RLM is a coding agent and it's already recursive. So wadida. And that's fine. And I don't think that argument is wrong, but it doesn't really move anything forward. And so

11:04

SPEAKER_00

what I'm interested in is this question of how can we apply the principles of RLMs to coding agents and make them actually useful for coding agents. And I've been very obsessed with this problem since the October, uh, RLM blog post came out in 2025. Okay. So I show some of my experiments on recursive coding agents. Um, the first one, it was simply wrapping Alex's RLM package as a CLI. So the idea was, I just want to give my coding agent an RLM as a tool call. So let's say we go, we need to go sift through a hundred million token corpus. Uh, well now it just used this, that tool. And then

11:51

SPEAKER_00

the RLM does the RLM thing. So that's interesting and it can be very useful. Uh, and a few other people have built, uh, things like that. Uh, but then I thought, well, what would it really mean? What would be possible or how could it be possible to make the coding agent fully recursive? So the coding agent harness, like calls itself, like the exact version of itself. Um, and how might you implement that? And that is what I called YPI. Uh, Y stands for the Lambda calculus, uh, Y combinator and PI in case you haven't heard of it is really awesome coding agent. It's incredibly minimal and it's

12:38

SPEAKER_00

specifically designed to be extensible. So Mario, uh, wants you to write extensions for PI for new features that you would like rather than trying to stuff your ideas into the, the main, um, agent. So it's meant to me a very minimal core that you extend however you like. And when I originally had the, uh, recursive coding agent idea and wanted to do it with PI, I was not able to use PI extensions to achieve this goal. Um, and so I had to fork it instead. I'm very excited to report that in anticipation of this talk, I revisited this and now PI has evolved, um, and the PI extensions have evolved such that you can, uh,

13:23

SPEAKER_00

make it fully recursive with a pure extension. So I have both the pure recursive, um, extension, PI recursive, uh, package as well as the Y PI wrapper. That is like a convenience wrapper for this. So this is very quite literally a recursive coding agent in the sense that PI calls PI calls PI calls PI. You can set the depth however you want. Um, and now I want to show a few other notable projects in the space. So obviously there's the original, uh, implementation from Alex, our alum, dspi.rlm is, my go-to, uh, when I'm especially doing benchmarking. Um, that's how I got all these amazing results,

14:03

SPEAKER_00

uh, on some of these benchmarks. Axe, I think is a very interesting one because it's incredibly agent native. So the ax started out as a typescript variation on dspi when our limbs came out, they obviously implemented it. Uh, but they did it in this way. That's, um, enables then ax agent to like write a whole typescript interface to another ax agent and go all the way down, uh, the, the recursive rabbit hole. Um, which I think is really cool and very interesting just to showcase that this like REPL could be anything. There's a, an example from Dan at open pros who made the UNIX RLM.

14:46

SPEAKER_00

This is pure bash and the environment is just the Linux file system. So that's a whole nother angle of thinking about, uh, what's possible with RLM. And then lastly, and we'll talk more about open pros, um, but, uh, open pros as a language, uh, actually enables you to convert any coding agent into an RLM. And I'll talk a little bit more about how to do that towards the end of the talk. Uh, and then the open pros repo also contains a harness that executes. It's a coding agent harness that, uh, will let you use codex SDK or cloud code underneath and do an RLM style execution of these pros programs.

15:28

SPEAKER_00

So is cloud code and RLM. This is like the, the question that keeps getting asked and the original answer on the day of release day of the blog post was no, no, no, it's not. Uh, but it was literally the first question that was asked or like asked on the very same day, uh, to the launch tweet. Um, it's basically saying, Hey, this is cloud code sub agents, right? Uh, go back to my rubric. If you want to like dig into some of the, the nitty gritty details. Um, but arguably now it is. So, uh, over here on the right, we can see Omar saying, Hey, congratulations. Anthropic, uh, cloud code is

16:07

SPEAKER_00

finally an RLM now that you have dynamic workflows. So, so what changed? And I think this is an interesting way of explaining, um, what's powerful about recursive coding agents and what RLMs even are, by using this example. So dynamic workflows were released just a few weeks ago. Um, and they make cloud code recursive or capable of doing these recursive workflows. And I would highly encourage you to read this blog post called a harness for every task. It shows six different workflow patterns that are very powerful. Obviously there's many more that you can achieve. And just to, to show it, I,

16:46

SPEAKER_00

I wrote two workflows for cloud code, uh, one that is explicitly not an RLM. So that's like a hard coded map reduce workflow. And one that I'm arguing is that is, uh, you could kind of think of it as like deep research over a file system. So, you know, pick, pick a handle, um, assign that to an agent, have it go do some analysis, bring back what it finds, et cetera. Uh, again, these are in the companion repo. Uh, and then now to open pros. So, so dynamic workflows are cool. They're only in cloud code. They're also not the only way to do this. So what if you don't like cloud code or what if you do like

17:27

SPEAKER_00

cloud code, but you don't want to use dynamic workflows, you want something else. So this is what open pros is all about. So open pros is technically a programming language, but it is not compiled by your computer. It's compiled by your coding agent. Um, it's a markdown spec. It's logical English. You don't need to learn any kind of crazy syntax. And there's actually a command pros, right? That will get cloud code or codex or your favorite or PI or your favorite coding agent to write a dot pros.md file for you. Um, so in that way, it's similar to the, the, um, cloud code, um, ultra code command where like decide and write the workflows for you.

18:13

SPEAKER_00

And the, the pros has the ability to turn any agent that's got a file system and sub agents into an RLM. So this is an open source repo. You can check it out. Um, and I've also written a little bit more in depth about this for Turing post in this article. And the key thing that I want to bring up with regards to the RLMs is that open pros can explicitly declare the sub agent work. So again, I've made two kind of demo pros.md files that are in this companion repo. If you want to dig into the code, I'm not going to do that here in the slides, uh, where you can break a problem up into smaller

18:54

SPEAKER_00

pieces that are assigned to sub agents, verify the work of those sub agents, um, in the, in the parent agent session. Uh, and what's even cooler is you can actually in open pros to the features that I added to the language are that you can add skills and tools as explicit dependencies. So you can imagine a workflow where a certain sub agent needs a very specific skill to do its role in the workflow or, uh, must have access to a certain CLI tool, for example, or it can't run and do its job. And so there's a way in pros to actually wire those in as dependencies to ensure that, um, not only that the

19:39

SPEAKER_00

way the work is done, um, is what you want, but actually that the sub agents are specifically configured with the tools and skills that they need to successfully do the work that you are declaring in the pros contract. Okay. Super cool. What can you actually do? So I've got two examples from cloud dynamic workflows, two examples from open pros, uh, repo scale migrations. This was kind of the launch post example, refactor, like a huge thing, uh, all with a big swarm and parallel and then merge the whole thing together. That's super cool. Um, this idea of like go after a directory and then deep

20:20

SPEAKER_00

research or deep analyze, uh, or deep process in some way recursively, uh, inside that is, is another example I have here. You can do audits, uh, bug sweeps. You can do adversarial things such as having a skeptical agent or, um, you know, a red team, uh, set of agents, uh, that are going to, uh, try to, uh, improve the system adversarially or in parallel. And then, uh, one really cool thing that you can do with open pros that I just added recently is kind of goes back to my very first slide, right? So like one day I get this amazing result the next day they trade away all my, all my, uh, crypto currency,

21:07

SPEAKER_00

which is very small, thankfully. Um, but like, how do we get these things to be more reliable and how do we get them? You know, we, we have like a golden session and we have a great day and now we want to capture that and reuse it over and over again. So I built a system where you can take a golden session for cloud code code, code X, PI, whatever you want. And it's a pros program that will actually have the agent deconstruct that session and turn it into a reusable pros workflow. Um, that again can involve this idea of recursive coding agents, um, to get you to a reliable way of getting to that golden state

21:46

SPEAKER_00

of performance over and over and over again. Recursing, recursive coding agents for the win. I really think this is very powerful. Our LMS just kind of blew my mind when they first came out. And again, as you can probably see through this talk, I've been absolutely obsessed with applying the ideas of our LMS to coding agents. The three things that I hope you'll take away from this is one, like trust is reliability. Like how can we, how can we trust something that isn't reliable? And again, my argument and this idea of the mismanaged genius is, uh, that the next step is not, um, more raw

22:26

SPEAKER_00

intelligence. It's actually, uh, behavioral textually orchestration. I personally believe that, um, our LMS represent this new paradigm of test time, compute inference time, compute, uh, where tool calling and reasoning are unified and we reason through tool calling and we can, uh, recursively iterate. And one of those tools is to call another agent to go do it on some other specific task or subset of the problem. Uh, and then also, I hope we settle a little bit of this drama around like, wait, like our, our LMS, um, actually new, aren't they just coding agents? Yes. Coding agents can be our LMS. They aren't automatically our LMS. Um,

23:12

SPEAKER_00

and so I've showed a couple of different cloud code dynamic workflows that can turn cloud code into an RLM, uh, as well as some ways of doing this with open pros that you can use with any coding agent. So I see this as an incredibly powerful way of working with coding agents. I hope you will dig in more to recursive coding agents.com, but with great power comes great responsibility. So until next time, please recurse responsibly. Thank you very much.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note