Hey. Hi, everyone.
Thanks for being here. Yeah, I'm super happy today to talk about automated AI research and especially all those performing automated AI research tasks. So I'm Elie. I work at Prime Intellect as a research engineer. And yeah, I will go through our work on this subject. So first, I want to explain a bit why we are doing that and why we think it's super important to do that in the open. So first, I think we all agree that we've heard about big labs saying that this bad thing called recursive self-improvement is coming very soon. So recursive self-improvement is a model training models without human intervention basically.
But we don't have any benchmark to quantify if this is true or not, right? And even less, we don't have a third-party benchmark by non-big labs to see if it's something coming soon or not. And the other part is that we think that it's super important to understand all those models do research because we think that a lot of the scientific research that will come in the coming years will be based also on AI tools. So it's super important to understand how those models do research, not just only AI research. So we try to build this environment to test the capabilities of the model to do so.
So it all started with Andrej Karpathy that's basically at fun by doing this video where he trained GPT-2 from scratch in like 90 minutes. GPT-2 training takes weeks and then two years ago I think it only took like 90 minutes. So what does it mean to reproduce GPT-2 in 90 minutes? It means that in 90 minutes you achieve this target loss. And yeah, and that's at this point when you have the same loss as GPT-2. You consider that your model is somewhat of equal performance. Then what happened is that the community took this repo, this GitHub repo and created another one called Modded Nano GPT. And this effort was led by someone called Keller Jordan.
And what happened is that they basically took this 90 minutes, then 45 minutes, and then now we can train a GPT-2 validation loss model in less than two minutes. And then we can do that, which is honestly crazy. And it took two years to achieve this. So it's a very strong benchmark where a lot of very talented researchers were at it. Yeah, so we decided to take this environment of speedrun. So what do you see? It's a game. So the goal of the game is to achieve this loss in the shortest amount of time. So this is the Nano GPT one. And you don't have almost any constraints. The only constraint that you got is that you need to use the same validation and training data.
There is a new speedrun called the Optimizer speedrun that was released a few months ago. And here it's slightly different because you can only change the optimizer related parameters. So for instance, Nano GPT, you can change the architecture, do MOE, do attention, whatever. Optimizer speedrun, you can only change Adam to Muon, Shampoo, or whatever optimizer is your favorite. Yeah, and so this is a bit more researchy because it's less about optimizing the program to be as fast as possible, but more finding the best method possible no matter the time you put into the computer. So, yeah, why take speedrun as an environment for automated AI research?
First, we think that it's a good evaluation. We'll see later why. And this is the main focus of this talk. But we also think it's probably a good training environment because it's a way to give the model a reward. So the reward is positive if the model beat the speedrun and beat the last record, sorry, and the reward is zero or negative if it didn't manage to do it. So it's a good environment to train model. It's also quite fast. As you see, previously around two minutes for the optimizer one. Each run takes about 15 to 20 minutes. And, yeah, there is clear rules basically.
And we also think it's a good environment to make discovery, a breakthrough in trial research because there is these clear rules that you can verify or not. And then we also think it's a good idea of the process. Yeah. So, yeah. So what we did, the release was about two months ago. And there was this optimizer speedrun. And we decided to basically compete with the community by launching two AI agents, so Codex and Claude Code. Codex was GPT 3.5 with XI. And Claude Code was OPUS 4.8 with XI. And, yeah, we decided to basically let the agent free on our cluster and just iterate on it. So we have V1, V2, V3. It's just basically us stopping the agent and then restarting.
V3 was one or two days before the release because we saw that our agents no longer have the best record. So we were like, okay, take all the human record in the last few weeks and just try to improve upon it and it worked. Yeah. And we also have this novelty track where the goal is to beat the record with only novel IDs. And we saw that this was more complex for the models. So our RNS is very simple. Honestly, we could have just replaced it with slash goal, but there was no slash goal at the time. So we made our own goal.md. It's actually quite efficient that we chose the same name. And we had the goal.md and agents.md that defined the rules.
And we let the agent propose IDs. And then it can submit a job with sbatch on our Slurm cluster. And basically the way it works is that it can submit on nodes that are available, but only under a certain permission, which means that if someone wants to use this node, the model just can save the job. It's called preemptible permission. So, yeah. Then it measures, it reads basically the training logs, then decides if it's a record or not. To validate a record, you need to basically pass a statistical threshold to make sure that it's not seed optimization and it's not on them, right? So, yeah. So, yeah. A few results from this experiment.
The first one that was honestly very painful to work with is that Claude keeps stopping every nine or ten hours and basically says, yeah, I cannot improve the record, it's too hard for me. There is no way to go beyond it. And then I was just like, okay, continue, explore new direction and just go again for ten hours and then say, yeah, I cannot beat the record and so on. So, basically one third of the time the Claude agent was idle because I had no way to basically monitor it and Codex. Totally the opposite just worked for all the time and yeah, almost never idle, never asked a question and, and, and very impressive in that way.
So, we also give the option for the model to basically write a bunch of stuff into what we call a scratch pad, which is basically the active memory of the model. We observe that basically Codex writes a lot on the scratch pad, so each plot that I will show are normalized by the number of active hours. So, this is not only about Codex working more, it's really different behavior. So, yeah, you see that writes a lot more to this scratch pad, to this memory. And, the shape of, the, I don't know, the tone of each file was also super different. Claude was super excited about getting new record with a bunch of emojis and so on. And Codex was just like, here is what I do.
Here is the decision I take. What will I do next? Super robotic. Um, yeah. We also have this plot where basically we saw that Codex was spawning much more subagents than Claude. We saw that Codex built much more token than Claude. So, I think in total it was like billions of token. But there is obviously this input caching that makes it, it's not one billion output token. So yeah. We also see that Codex did a lot of compaction because it only had 250k context window. And Claude only did it like one per hour. And Codex is more like, no, it's even less than one per hour for, I mean, one for the full run for Claude. And Codex was like 20 every one hour. So, yeah.
Um, yeah. Here is the main results. So, what this plot shows is that basically we, in the white, you see the human record progression, right? And in red, you see Claude. I mean, it's supposed to be orange, but whatever. And in blue, you see Codex, right? And you see that at almost every time, Claude and Codex are better than the human record. And Claude is super good at the beginning. Very, very fast to achieve very good score. Um, yeah. And one thing that is super important is that the model have the ability to basically fetch the human records at any time. And that's what Codex did. That's what Claude did. Sorry.
Because when I restarted it, it basically fetched the new record from human and improved upon it. Um, yeah. So, the result is that, I think at the time, the best record was 2,990 steps. And we beat it by 50 or 60 steps for Claude. And, Codex was 20 steps above. So, I think it's both impressive and, yeah. Um, so we, this is not released yet. This is something that we are working on currently. And basically the idea is that this is a cool experiment to do, but it lacks structure, right? If you want to do a real benchmark, you want to do multiple seeds, you want to do, yeah, proper thing where you basically put all the model and illness in the same condition, right?
So, this is what we are working on right now. And basically, the idea is to do three different tracks. One without any access to really, measure the capability of the models to do AI research based on only the model weight knowledge. One with only archive papers. And one with full access. So, it also has access to the latest record by human. And for this, we plan to do both the nano GPT track one, which is the original one and the optimizer speedrun, where we only constrain the optimizer to be novel, basically. Um, yeah. So, I will present some results on the optimizer speedrun. This is basically what we got.
So, we let the agent iterate for six days, almost five days, let's say. And we see that Codex, Gemini and Claude are super effective. So, for GLM, this is not finished run, right? So, the model is actually still iterating on our cluster right now. But we see that Claude is once again very good at it. And we see that surprisingly, Gemini is also very competitive. And has this breakthrough on day four where it beat Codex with a new record, right? It's also interesting to see that Claude is much more progressive in the way it improved the record. And Gemini has really this step function where it does a breakthrough and so on.
So this is an interesting plot because there is quite a lot for an eval. But you can change this axis by also the number of output token. And then tell a different story. Because Claude in MaxMode consumes so much more token than Codex and Gemini. And you also see that Gemini is actually super efficient for the number of token that each uses. So it's Gemini K2.7 code. Um, so yeah. We also see that they have a different approach to using the literature and papers. Um, so for instance, Claude is doing a lot of search on papers. And actually Claude found a paper that no other model found. And it actually led to the best record. So it's kind of funny.
And, yeah. Um, one of the main issues of all of this is that when I launched this agent. And I think that's something important that I want you to remember for this talk. Is that when I launched this different agent, I was expecting them to come up with some crazy ideas on the optimizer that no one have discovered. But honestly, it wasn't the case. They did some clever trick where basically they combine different papers. They do plus one improvement over a bunch of methods. But there was really no novel optimizer or mechanism that was coming from those models.
And I think that's telling that even something that is not simple, but I'd say that it's accessible for people, right? For human researchers spending days and weeks for, the model cannot find new optimizers and mechanisms. So we believe that there is a way to basically make it better for discovery instead of evaluation. And this is inspired from Alpha Evolve by Google and also a bunch of papers that have been released since then. It's this multi-agent system that interact together, a bunch of generators. You have a closed model, but you also have open source models here that are super effective for the cost, right? They can suggest ideas.
Then you run the speedrun, so you get the reward. Then you have a judge that basically gives quality feedback. It can also be that the judge has a taste about the method if it's good or not, if it's outside the loop. And then you can basically decide which method you want to scale to a larger number of parameters and number of token. Um, so this is the scale part of the speedrun because a lot of methods in the speedrun community, people are often saying that they don't want to scale. They don't want to work at large scale, so I think it's very important to also put a scale element in this loop.
And I think also that humans are super useful here to basically judge the ideas of agents, steer them in the right direction and so on. Um, yeah, so we didn't try it yet. I mean, we are trying it right now. We hope that this will lead to new discovery in AI research at least. And also a way is that you can define multiple speedruns. So this is the next slide issue from a slide. But if you don't have the reference, good for you. It means that you're not too online. But the idea is that by changing the objective and the constraints of the speedrun, you can basically create a lot of diversity. And constrain the model to go into a certain direction.
And, yeah, and make those discoveries. So, at Prime Intellect, we are doing a bunch of stuff in this direction. There is a bunch of stuff here that we, most of it we didn't release yet. But we are working on GPU sandboxing to allow model to iterate in sandbox because you need GPU sandbox for this kind of stuff. We are working on our own agents that are very efficient for RLM framework. So it means you have a file system and you can write information, read from it, and you also do this programmatic tool coding thing. We are also training a model to be good at it on top of open source models.
And the thing that we already released is that we have this set of library and product called Verifier, Primerell, Austin Training, where you can basically train, evaluate any environments on any harness and the model that you can train can be GNM 5.2, which is very big. And, yeah, we work a lot on making those libraries very efficient to shape the best quality for our clients. Yeah? I mean, yeah, super excited about this domain. Once again, I think it's super important to have a part of this recursive self-improvement happen in the open. Because there is actually a lot of people working that are not on big labs.
So you need to basically make it easy for people to understand how those models work to do research and so on. So that's our goal. And, yeah, thanks a lot. Thank you. I mean, we are kind of trying it right now. Uh, we hope that this will lead to, to, to new discovery in AI research at least. And also a way is that you can define multiple speedrun. So this is the next slide, uh, issue, like from self-bank, uh, slides. But if you, if you don't have the reference, good for you. It means that you're not too online. Uh, but the idea is that, uh, by changing the objective and the constraints of the speedrun, you can basically create a lot of diversity.
And constrain the model to go into a certain direction. And, uh, yeah, and make those discovery. So, uh, at Heim Intellect, we are doing a bunch of stuff in this direction. Uh, there is a bunch of stuff here that we, I mean, most of it we didn't release yet. But we are working on, uh, GPU sandboxing to allow model to iterate into sandbox because you need GPU sandbox for this kind of stuff. We are working on our own agents that are very efficient for, like, uh, uh, RLM framework. So it means, like, you have a file system and you can write information, read from it, uh, and you also do, like, this programmatic tool coding thing.
We are also training a model to be good at it on top of, like, uh, open source model. And, uh, the thing that we already released is that we have those set of library and product called Verifier, Primarell, Austin Training, where you can basically train, evaluate any environments on any harness and the model that you can train can be, like, GNM 5.2, which is very big. And, yeah, we have, like, we work a lot on making those libraries very efficient to, to shape the best quality for our clients. Yeah? Uh, I mean, yeah, super excited about this domain.
Once again, I think it's super important to have, uh, a part of, like, this recursive self-improvement to happen in the open. Because there is actually a lot of people working that are not on big labs. So you need to basically, uh, yeah, make it easy for people to understand all those model work to do research and so on. So that's kind of our goal. And, uh, yeah, thanks a lot.
Thank you.