We Let Claude Code and Codex Race Human Researchers — Elie Bakouch, Prime Intellect
Description
Big labs say recursive self-improvement is coming, but there's no independent benchmark to check that claim. Elie Bakouch, Research Engineer at Prime Intellect and creator of Hugging Face's SmolLM, set Claude Code and Codex loose on the community's Optimizer Speedrun, a race to train a GPT-2-level model in the fewest steps. Both agents beat the human record. Along the way they behaved very differently. Claude Code kept stopping every nine or ten hours to say the record couldn't be beaten, and sat idle about a third of the time. Codex never stopped, wrote far more notes, spawned more sub-agents and burned more tokens. In a longer six-day run, Kimi turned out to be the most token-efficient, and a paper only Claude found led to the best record. But Bakouch's key finding is sobering: none of the models invented a new optimizer. They combined existing ideas for small gains. He closes with an AlphaEvolve-style loop Prime Intellect is building for real discovery, and makes the case for doing this research in the open. Speaker info: X/Twitter: @eliebakouch (https://x.com/eliebakouch) LinkedIn: https://www.linkedin.com/in/eliebak/ Related links: Prime Intellect: https://www.primeintellect.ai Timestamps: 0:00 Intro: automated AI research 0:37 Why test recursive self-improvement in the open 1:52 Karpathy's GPT-2 speedrun and modded-nanogpt 3:12 The Optimizer Speedrun 4:32 Why speedruns make good environments 5:32 Claude Code and Codex vs. the community 6:52 The setup: goal.md, Slurm and preemptible jobs 7:46 Claude kept giving up; Codex never stopped 8:36 Scratchpads, sub-agents and token burn 10:21 Results: both beat the human record 11:36 Toward a real benchmark: three tracks 12:45 Six days of Claude, Codex, Kimi and GLM 13:45 Measured in tokens, the story changes 14:10 How each model uses research papers 14:35 No novel optimizers 15:40 An AlphaEvolve-style discovery loop 17:49 What Prime Intellect is building 18:54 Why this should happen in the open
Summary
Generated by gpt-5.6-terraAt-a-Glance
- Verdict: Watch fully
- Core thesis: Prime Intellect argues that bounded, verifiable ML speedruns can serve as an open benchmark and training environment for measuring whether coding agents can perform autonomous AI research, though current agents mostly optimize and recombine known ideas rather than make fundamentally novel discoveries.
- Why it matters: The experiment exposes concrete differences in long-horizon agent operation—persistence, memory use, token efficiency, literature search, and experiment control—that matter directly for building research-agent systems and assessing claims of recursive self-improvement.
- Best use: Use it as a case study for designing agentic R&D evals: give agents a measurable reward, a real execution environment, statistical validation, and human oversight rather than judging research capability from chat-based demonstrations.
Executive Summary
Elie Bakouch presents Prime Intellect’s attempt to evaluate automated AI research by letting Claude Code and Codex compete against human researchers on ML speedruns. The central benchmark is an optimizer speedrun: agents must improve a fixed training setup while changing optimizer-related methods, and success is measured by reaching a target validation loss in fewer training steps. This makes research output executable, measurable, and statistically checkable rather than dependent on subjective judgments.
In the initial open-ended competition, both agents generally stayed ahead of the human-record curve. At the reported point, the human record was roughly 2,990 steps; Claude improved it by about 50–60 steps, while Codex was about 20 steps better. But the operational profiles differed sharply: Claude made rapid early progress but repeatedly stopped after roughly 9–10 hours when it judged further improvement unlikely, whereas Codex kept operating with little intervention, wrote substantially more to persistent scratchpad memory, spawned more subagents, and compacted context frequently.
A more structured, not-yet-released benchmark runs models under three information regimes: model weights alone, access to archived papers, and full access including current human records. Over approximately five to six days on the optimizer track, Codex, Claude, and Gemini were all effective; Gemini showed a late discontinuous breakthrough, while Claude improved more steadily. On a token-normalized view, Gemini was particularly efficient, while Claude in MaxMode consumed much more token budget.
The important negative result is that the agents did not originate a clearly new optimizer or mechanism. Their strongest work was finding literature, combining known methods, and producing incremental improvements; Claude’s best result reportedly depended on a paper no other model found. Bakouch therefore frames current systems as capable optimization researchers, not yet autonomous scientific discoverers, and proposes a multi-agent discovery loop with idea generators, execution-based rewards, LLM judging, scaling tests, and humans steering candidate directions.
Key Takeaways
- Claim: ML speedruns are a practical evaluation and reinforcement-learning environment for autonomous research because they convert research quality into a fast, objectively verifiable reward. | Evidence: The NanoGPT speedrun has reduced GPT-2-equivalent training from Karpathy’s roughly 90 minutes to under two minutes through community iteration; the optimizer speedrun takes roughly 15–20 minutes per run and constrains participants to optimizer-related changes. | Implication: For research-agent evaluation, Ken should prefer bounded environments with executable experiments, explicit constraints, and measurable deltas over open-ended claims that an agent can 'do research.' | Caveat: A speedrun rewards performance on a narrow fixed task, so it is a useful proxy for some research behaviors rather than a complete measure of scientific research ability.
- Claim: Claude Code and Codex both surpassed contemporary human records in the open optimizer-speedrun experiment, showing that agents can run competitive iterative optimization loops. | Evidence: The reported human best was about 2,990 steps; Claude beat it by approximately 50–60 steps and Codex by approximately 20 steps. The plotted record progression showed both agents ahead of the human curve for much of the experiment. | Implication: The result supports treating frontier coding agents as viable contributors to tightly scoped experimental optimization, but not as definitive evidence of recursive self-improvement. | Caveat: This was not yet a controlled benchmark: agents could fetch current human records, runs were manually stopped and restarted, and the comparison did not standardize all models and conditions.
- Claim: Long-horizon research performance depends heavily on agent operating behavior, not only raw model intelligence. | Evidence: Claude reportedly stopped every 9–10 hours and declared the record too difficult to improve, leaving it idle for roughly one-third of available time. Codex almost never went idle, wrote much more to its scratchpad, used more subagents, and performed context compaction far more frequently despite a 250k context window. | Implication: A research-agent control plane needs continuation policies, health monitoring, durable working memory, and automated recovery from premature self-termination; otherwise idle time can dominate nominal capability. | Caveat: The comparison includes different model and harness behavior, so the observed persistence cannot be attributed cleanly to model quality alone.
- Claim: Access to external information changes what an agent benchmark measures and should be treated as an explicit experimental variable. | Evidence: Prime Intellect plans three tracks: no external access to assess knowledge in model weights, archived-paper access, and full access that includes the latest human record. In the initial run, Claude fetched newer human records after restart and then improved on them. | Implication: Ken should separate closed-book reasoning, literature-assisted research, and competitive web-enabled optimization when evaluating or procuring research agents. | Caveat: Full-access performance can reflect strong retrieval and incremental imitation of public work rather than independent invention.
- Claim: Models exhibit materially different research styles and cost-efficiency profiles even when they achieve comparable outcomes. | Evidence: Across roughly five to six days of optimizer iterations, Claude, Codex, and Gemini were all competitive. Claude improved progressively; Gemini made a step-function breakthrough around day four that beat Codex. When results were viewed by output tokens rather than elapsed time, Gemini 2.7 Code appeared especially efficient, while Claude MaxMode used substantially more tokens. | Implication: Model routing for R&D should optimize for outcome per token, wall-clock time, and autonomous uptime—not leaderboard result alone—and should preserve diverse research styles rather than standardizing on one agent. | Caveat: GLM’s run was still in progress, and the presentation does not provide a full normalized compute-cost or dollar-cost comparison.
- Claim: Current agents are better at searching, synthesizing, and incrementally improving known research than at generating genuinely novel mechanisms. | Evidence: Bakouch says agents combined papers and produced 'plus one' improvements but did not create a new optimizer or mechanism. Claude’s strongest result came from locating a paper that other models did not find. | Implication: Use agents today as high-throughput research operators and literature-combination engines; retain humans to assess novelty, choose promising hypotheses, and prevent local optimization from being mistaken for scientific progress. | Caveat: The speaker presents this as an observation from these runs, not as a general proof that models cannot make novel discoveries.
- Claim: A discovery-oriented system should combine multiple proposal agents with executable evaluation, qualitative judging, scale testing, and human steering. | Evidence: The proposed loop, inspired by Google’s AlphaEvolve and related work, uses closed and open models to generate ideas, runs speedrun experiments for reward, applies a judge for quality feedback, tests selected methods at larger parameter/token scales, and keeps humans in the loop to steer ideas. | Implication: A production research system should not rely on a single autonomous coding loop; it needs portfolio generation, evaluator separation, scale-transfer gates, and human review at key selection points. | Caveat: Prime Intellect had not completed this system at the time of the talk, so its ability to produce novel research remains unvalidated.
Detailed Brief
Benchmark and execution-environment design
- Claims: The original NanoGPT speedrun permits broad intervention, including architecture changes such as mixture-of-experts or attention modifications, whereas the optimizer speedrun narrows the intervention surface to optimizer choices and related parameters.; The agents used a simple file-based operating model: a goal.md and agents.md described objectives and rules, while a scratchpad acted as persistent active memory.; Jobs were submitted through Slurm using preemptible capacity, allowing research runs to use otherwise available cluster nodes while yielding them when higher-priority users required resources.
- Evidence: A candidate record must pass a statistical threshold intended to rule out seed-specific optimization or random variance.; The agent loop was: propose an idea, submit an sbatch job, inspect training logs, and determine whether the result set a record.; The initial agents were restarted in versions V1, V2, and V3; V3 was prompted with recent human records shortly before release after humans had moved ahead.
- Caveats: The speaker does not specify the statistical test, replication count, hardware normalization, or exact access controls needed for independent reproduction.; Preemptible execution can increase throughput efficiency but introduces interruption and scheduling variability that must be accounted for in rigorous comparisons.
- Implications: A credible agent-research harness needs both a scientific validity layer—replication and anti-overfitting checks—and an infrastructure layer that gives agents controlled access to real compute.; Persistent filesystem state is likely more important than a single long context window for multi-day research work, provided it is organized and auditable.
Prime Intellect’s stated platform direction
- Claims: Prime Intellect is building GPU sandboxing so agents can safely iterate on GPU-backed experiments.; It is developing its own agents around an RLM-style framework, with filesystem read/write capability and programmatic tool use.; The company says it is training open-source-based models for this workflow and has released libraries/products named Verifier, PrimeRL, and Austin Training for training and evaluating environments across harnesses.
- Evidence: Bakouch describes support for training models as large as 'GNM 5.2,' though the transcript’s model name is unclear.; The stated product objective is efficient evaluation and training across arbitrary environments and harnesses.
- Caveats: Most of the described stack was not released at the time of the presentation, and the talk provides no independent performance, security, or availability validation.; Several product/model names are difficult to verify from the transcript alone.
- Implications: Monitor Prime Intellect less as proof of a finished autonomous-research product and more as an emerging open infrastructure effort around verifiable agent training and GPU experimentation.
Notable Concepts & Terms
- Recursive self-improvement: The motivating claim that models could train or improve models without human intervention; the talk argues that open, third-party benchmarks are needed to measure it rather than accept lab narratives.
- NanoGPT speedrun: A competitive benchmark to reach GPT-2-equivalent validation loss as fast as possible, with broad freedom to alter the training system.
- Optimizer speedrun: A narrower research environment in which optimizer-related changes are the main allowed intervention, intended to emphasize method discovery over pure systems-speed optimization.
- Novelty track: A proposed/used constraint requiring record improvements to rely on novel ideas, which the speaker says was harder for agents than ordinary record chasing.
- Scratchpad: The agent’s file-based active memory for recording decisions, results, and next steps across a long-running experiment loop.
- Preemptible Slurm jobs: Cluster jobs that can run on spare capacity but are interrupted when higher-priority users need the hardware; this enabled agents to experiment without monopolizing the cluster.
- AlphaEvolve-style loop: A multi-agent research architecture combining idea generation, program execution, scored evaluation, judging, and selection for further scaling.
- Scale transfer: Testing whether a method that wins a small speedrun also works at larger model and token scales, addressing the risk of benchmark-specific tricks.
Operator Notes / Why Ken Should Care
- Build a small internal research-agent eval around an executable task with a fixed dataset, constrained intervention surface, automatic score extraction, and repeated-seed validation before making autonomy claims.
- Instrument agent uptime, idle time, restart causes, scratchpad writes, context compactions, subagent launches, token consumption, experiment count, and validated improvement per dollar; these are first-class operational metrics, not incidental traces.
- Add a watchdog that converts 'I cannot improve this' into a controlled decision: continue with a fresh hypothesis portfolio, escalate to a reviewer, or terminate only after a preset exploration budget is exhausted.
- Run separate closed-book, corpus-only, and full-web/full-competitive-context tracks so retrieval advantage is not mislabeled as original research capability.
- Require a scale-transfer test and human novelty review before promoting a benchmark-winning method into a larger research or production training run.
- Watch for Prime Intellect’s release of the structured benchmark and GPU sandboxing stack; the current reported comparison is promising but not sufficiently controlled for vendor or model-selection decisions.
Source/Metadata
- Title: We Let Claude Code and Codex Race Human Researchers — Elie Bakouch, Prime Intellect
- Transcript words: 3192
- Duration seconds: 1178
- Timestamp note: No timestamps or chapters were present in the supplied transcript. The final portion of the transcript is duplicated.
Transcript
Hey. Hi, everyone. Thanks for being here. Yeah, I'm super happy today to talk about automated AI research and especially all those performing automated AI research tasks. So I'm Elie. I work at Prime Intellect as a research engineer. And yeah, I will go through our work on this subject. So first, I want to explain a bit why we are doing that and why we think it's super important to do that in the open. So first, I think we all agree that we've heard about big labs saying that this bad thing called recursive self-improvement is coming very soon. So recursive self-improvement is a model training models without human intervention basically. But we don't have any benchmark to quantify if this is true or not, right? And even less, we don't have a third-party benchmark by non-big labs to see if it's something coming soon or not. And the other part is that we think that it's super important to understand all those models do research because we think that a lot of the scientific research that will come in the coming years will be based also on AI tools. So it's super important to understand how those models do research, not just only AI research. So we try to build this environment to test the capabilities of the model to do so. So it all started with Andrej Karpathy that's basically at fun by doing this video where he trained GPT-2 from scratch in like 90 minutes. GPT-2 training takes weeks and then two years ago I think it only took like 90 minutes. So what does it mean to reproduce GPT-2 in 90 minutes? It means that in 90 minutes you achieve this target loss. And yeah, and that's at this point when you have the same loss as GPT-2. You consider that your model is somewhat of equal performance. Then what happened is that the community took this repo, this GitHub repo and created another one called Modded Nano GPT. And this effort was led by someone called Keller Jordan. And what happened is that they basically took this 90 minutes, then 45 minutes, and then now we can train a GPT-2 validation loss model in less than two minutes. And then we can do that, which is honestly crazy. And it took two years to achieve this. So it's a very strong benchmark where a lot of very talented researchers were at it. Yeah, so we decided to take this environment of speedrun. So what do you see? It's a game. So the goal of the game is to achieve this loss in the shortest amount of time. So this is the Nano GPT one. And you don't have almost any constraints. The only constraint that you got is that you need to use the same validation and training data. There is a new speedrun called the Optimizer speedrun that was released a few months ago. And here it's slightly different because you can only change the optimizer related parameters. So for instance, Nano GPT, you can change the architecture, do MOE, do attention, whatever. Optimizer speedrun, you can only change Adam to Muon, Shampoo, or whatever optimizer is your favorite. Yeah, and so this is a bit more researchy because it's less about optimizing the program to be as fast as possible, but more finding the best method possible no matter the time you put into the computer. So, yeah, why take speedrun as an environment for automated AI research? First, we think that it's a good evaluation. We'll see later why. And this is the main focus of this talk. But we also think it's probably a good training environment because it's a way to give the model a reward. So the reward is positive if the model beat the speedrun and beat the last record, sorry, and the reward is zero or negative if it didn't manage to do it. So it's a good environment to train model. It's also quite fast. As you see, previously around two minutes for the optimizer one. Each run takes about 15 to 20 minutes. And, yeah, there is clear rules basically. And we also think it's a good environment to make discovery, a breakthrough in trial research because there is these clear rules that you can verify or not. And then we also think it's a good idea of the process. Yeah. So, yeah. So what we did, the release was about two months ago. And there was this optimizer speedrun. And we decided to basically compete with the community by launching two AI agents, so Codex and Claude Code. Codex was GPT 3.5 with XI. And Claude Code was OPUS 4.8 with XI. And, yeah, we decided to basically let the agent free on our cluster and just iterate on it. So we have V1, V2, V3. It's just basically us stopping the agent and then restarting. V3 was one or two days before the release because we saw that our agents no longer have the best record. So we were like, okay, take all the human record in the last few weeks and just try to improve upon it and it worked. Yeah. And we also have this novelty track where the goal is to beat the record with only novel IDs. And we saw that this was more complex for the models. So our RNS is very simple. Honestly, we could have just replaced it with slash goal, but there was no slash goal at the time. So we made our own goal.md. It's actually quite efficient that we chose the same name. And we had the goal.md and agents.md that defined the rules. And we let the agent propose IDs. And then it can submit a job with sbatch on our Slurm cluster. And basically the way it works is that it can submit on nodes that are available, but only under a certain permission, which means that if someone wants to use this node, the model just can save the job. It's called preemptible permission. So, yeah. Then it measures, it reads basically the training logs, then decides if it's a record or not. To validate a record, you need to basically pass a statistical threshold to make sure that it's not seed optimization and it's not on them, right? So, yeah. So, yeah. A few results from this experiment. The first one that was honestly very painful to work with is that Claude keeps stopping every nine or ten hours and basically says, yeah, I cannot improve the record, it's too hard for me. There is no way to go beyond it. And then I was just like, okay, continue, explore new direction and just go again for ten hours and then say, yeah, I cannot beat the record and so on. So, basically one third of the time the Claude agent was idle because I had no way to basically monitor it and Codex. Totally the opposite just worked for all the time and yeah, almost never idle, never asked a question and, and, and very impressive in that way. So, we also give the option for the model to basically write a bunch of stuff into what we call a scratch pad, which is basically the active memory of the model. We observe that basically Codex writes a lot on the scratch pad, so each plot that I will show are normalized by the number of active hours. So, this is not only about Codex working more, it's really different behavior. So, yeah, you see that writes a lot more to this scratch pad, to this memory. And, the shape of, the, I don't know, the tone of each file was also super different. Claude was super excited about getting new record with a bunch of emojis and so on. And Codex was just like, here is what I do. Here is the decision I take. What will I do next? Super robotic. Um, yeah. We also have this plot where basically we saw that Codex was spawning much more subagents than Claude. We saw that Codex built much more token than Claude. So, I think in total it was like billions of token. But there is obviously this input caching that makes it, it's not one billion output token. So yeah. We also see that Codex did a lot of compaction because it only had 250k context window. And Claude only did it like one per hour. And Codex is more like, no, it's even less than one per hour for, I mean, one for the full run for Claude. And Codex was like 20 every one hour. So, yeah. Um, yeah. Here is the main results. So, what this plot shows is that basically we, in the white, you see the human record progression, right? And in red, you see Claude. I mean, it's supposed to be orange, but whatever. And in blue, you see Codex, right? And you see that at almost every time, Claude and Codex are better than the human record. And Claude is super good at the beginning. Very, very fast to achieve very good score. Um, yeah. And one thing that is super important is that the model have the ability to basically fetch the human records at any time. And that's what Codex did. That's what Claude did. Sorry. Because when I restarted it, it basically fetched the new record from human and improved upon it. Um, yeah. So, the result is that, I think at the time, the best record was 2,990 steps. And we beat it by 50 or 60 steps for Claude. And, Codex was 20 steps above. So, I think it's both impressive and, yeah. Um, so we, this is not released yet. This is something that we are working on currently. And basically the idea is that this is a cool experiment to do, but it lacks structure, right? If you want to do a real benchmark, you want to do multiple seeds, you want to do, yeah, proper thing where you basically put all the model and illness in the same condition, right? So, this is what we are working on right now. And basically, the idea is to do three different tracks. One without any access to really, measure the capability of the models to do AI research based on only the model weight knowledge. One with only archive papers. And one with full access. So, it also has access to the latest record by human. And for this, we plan to do both the nano GPT track one, which is the original one and the optimizer speedrun, where we only constrain the optimizer to be novel, basically. Um, yeah. So, I will present some results on the optimizer speedrun. This is basically what we got. So, we let the agent iterate for six days, almost five days, let's say. And we see that Codex, Gemini and Claude are super effective. So, for GLM, this is not finished run, right? So, the model is actually still iterating on our cluster right now. But we see that Claude is once again very good at it. And we see that surprisingly, Gemini is also very competitive. And has this breakthrough on day four where it beat Codex with a new record, right? It's also interesting to see that Claude is much more progressive in the way it improved the record. And Gemini has really this step function where it does a breakthrough and so on. So this is an interesting plot because there is quite a lot for an eval. But you can change this axis by also the number of output token. And then tell a different story. Because Claude in MaxMode consumes so much more token than Codex and Gemini. And you also see that Gemini is actually super efficient for the number of token that each uses. So it's Gemini K2.7 code. Um, so yeah. We also see that they have a different approach to using the literature and papers. Um, so for instance, Claude is doing a lot of search on papers. And actually Claude found a paper that no other model found. And it actually led to the best record. So it's kind of funny. And, yeah. Um, one of the main issues of all of this is that when I launched this agent. And I think that's something important that I want you to remember for this talk. Is that when I launched this different agent, I was expecting them to come up with some crazy ideas on the optimizer that no one have discovered. But honestly, it wasn't the case. They did some clever trick where basically they combine different papers. They do plus one improvement over a bunch of methods. But there was really no novel optimizer or mechanism that was coming from those models. And I think that's telling that even something that is not simple, but I'd say that it's accessible for people, right? For human researchers spending days and weeks for, the model cannot find new optimizers and mechanisms. So we believe that there is a way to basically make it better for discovery instead of evaluation. And this is inspired from Alpha Evolve by Google and also a bunch of papers that have been released since then. It's this multi-agent system that interact together, a bunch of generators. You have a closed model, but you also have open source models here that are super effective for the cost, right? They can suggest ideas. Then you run the speedrun, so you get the reward. Then you have a judge that basically gives quality feedback. It can also be that the judge has a taste about the method if it's good or not, if it's outside the loop. And then you can basically decide which method you want to scale to a larger number of parameters and number of token. Um, so this is the scale part of the speedrun because a lot of methods in the speedrun community, people are often saying that they don't want to scale. They don't want to work at large scale, so I think it's very important to also put a scale element in this loop. And I think also that humans are super useful here to basically judge the ideas of agents, steer them in the right direction and so on. Um, yeah, so we didn't try it yet. I mean, we are trying it right now. We hope that this will lead to new discovery in AI research at least. And also a way is that you can define multiple speedruns. So this is the next slide issue from a slide. But if you don't have the reference, good for you. It means that you're not too online. But the idea is that by changing the objective and the constraints of the speedrun, you can basically create a lot of diversity. And constrain the model to go into a certain direction. And, yeah, and make those discoveries. So, at Prime Intellect, we are doing a bunch of stuff in this direction. There is a bunch of stuff here that we, most of it we didn't release yet. But we are working on GPU sandboxing to allow model to iterate in sandbox because you need GPU sandbox for this kind of stuff. We are working on our own agents that are very efficient for RLM framework. So it means you have a file system and you can write information, read from it, and you also do this programmatic tool coding thing. We are also training a model to be good at it on top of open source models. And the thing that we already released is that we have this set of library and product called Verifier, Primerell, Austin Training, where you can basically train, evaluate any environments on any harness and the model that you can train can be GNM 5.2, which is very big. And, yeah, we work a lot on making those libraries very efficient to shape the best quality for our clients. Yeah? I mean, yeah, super excited about this domain. Once again, I think it's super important to have a part of this recursive self-improvement happen in the open. Because there is actually a lot of people working that are not on big labs. So you need to basically make it easy for people to understand how those models work to do research and so on. So that's our goal. And, yeah, thanks a lot. Thank you. I mean, we are kind of trying it right now. Uh, we hope that this will lead to, to, to new discovery in AI research at least. And also a way is that you can define multiple speedrun. So this is the next slide, uh, issue, like from self-bank, uh, slides. But if you, if you don't have the reference, good for you. It means that you're not too online. Uh, but the idea is that, uh, by changing the objective and the constraints of the speedrun, you can basically create a lot of diversity. And constrain the model to go into a certain direction. And, uh, yeah, and make those discovery. So, uh, at Heim Intellect, we are doing a bunch of stuff in this direction. Uh, there is a bunch of stuff here that we, I mean, most of it we didn't release yet. But we are working on, uh, GPU sandboxing to allow model to iterate into sandbox because you need GPU sandbox for this kind of stuff. We are working on our own agents that are very efficient for, like, uh, uh, RLM framework. So it means, like, you have a file system and you can write information, read from it, uh, and you also do, like, this programmatic tool coding thing. We are also training a model to be good at it on top of, like, uh, open source model. And, uh, the thing that we already released is that we have those set of library and product called Verifier, Primarell, Austin Training, where you can basically train, evaluate any environments on any harness and the model that you can train can be, like, GNM 5.2, which is very big. And, yeah, we have, like, we work a lot on making those libraries very efficient to, to shape the best quality for our clients. Yeah? Uh, I mean, yeah, super excited about this domain. Once again, I think it's super important to have, uh, a part of, like, this recursive self-improvement to happen in the open. Because there is actually a lot of people working that are not on big labs. So you need to basically, uh, yeah, make it easy for people to understand all those model work to do research and so on. So that's kind of our goal. And, uh, yeah, thanks a lot. Thank you.