It is so lovely to be here. So I wanted to share today some thoughts that I have around who gets to be at the frontier of discovery. So modern computer science as a field has only existed for the last 77 years. It's bizarre when you think about it. So World War II, all the transistor technology that was developed for radio, we finally had our first versions of the computer. But when you think about it, that's only two generations of people working on these tools. However, within that time, even for computer science, who and what and what topics we work on has dramatically changed. And I think it's an interesting setting because actually, if you look back across science as a whole, how we do discovery has been markedly different at different points in time. So when we started, the whole idea of a researcher was what we call a gentleman scientist, typically someone with adequate wealth to dabble in discovery. And these were all individuals, independent researchers. When we see the first associations emerge with the rural society in the 1600s, the idea of being a scientist as a full-time job was very special. And these became the predominant spaces for discovery. Why am I talking about this? Because, to be honest, the professionalization of science led to what I recall, and my good friend Roseanne Lu calls, the unreasonably narrow path. So I'm an AI researcher. This is Jan's actual career. And it's interesting, if you were to be an AI researcher at the forefront, you had to follow this exact narrow path. You had to get into the right PhD program. You had to then go to the right industry lab. You had to do sufficiently interesting work. And then, finally, you got to contribute to the frontier. This was my story too. So a lot of my work has been at efficiency at scale. I did my PhD and worked at DeepMind and a lot of different frontier labs. But this was, in many ways, a very aggressively filtered system. If you did not make it or you were not curious about the right problem at the right time, you didn't have a place to play at a frontier lab. And this is the standard successful scientist. You have a famous advisor, hopefully. You hopefully get one or two important internships. And what's interesting about this is computer science was really about representing the world. And it was all the tools that do that. But the reason why most people go into computer science was the question at the end of it. If you can represent the world, what questions can you answer? So today I'm going to talk about what I think is one of the most profound and interesting topics, which is why this is so important for computer science, that this is changing. So I'll also speak to this. It was double compounded in computer science because of the need for compute. So this resulted in jokes. It's fun that Merve was here. This is her tweet about GPU poor versus GPU rich. It led to barriers of entry on who can contribute frontier AI. So I put here company A, company B, company C. But to be honest, if we polled, I think there would be significant majority votes about who those companies are. But basically, a handful of frontier labs have been able to build the technology we use. It's also determined who gets to participate in breakthroughs and who doesn't. So this is a map of what Stanford calls where statistically significant breakthroughs have come through. And you can see whole sections of the world are completely left out. And so, for me, this is a very important question worth answering. Who gets to shape the frontier? Who gets to answer the questions at the end of the pursuit? We've seen that the shift has dramatically changed from academia to industry. And it's also meant that we ship the same model to everyone. Why that's particularly interesting is that most people intuitively understand that you shouldn't ship the same model to billions of people. And they also understand that it's not a particularly good use of compute, right? You're spending the same amount of compute on everything. And some problems are hard and some are very easy. So where does that leave us? What's my talk for today? I would like to say that we are ripe for a revolution. And we are ripe for a revolution in who gets to participate at the frontier of AI. I'll tell you two reasons why I'm bullish on this. I'll definitely cover one, and then I'll actually see interest and timing because I want to leave plenty of time for questions. I think that's, unless you have a few questions and a bit of banter, these things can be boring. So we'll see. But I'll cover one definitely that I'm actively thinking about. And this is: what if we could allow anyone to build the same frontier intelligence as that in labs? And I've been in a few labs. I've done my tour of duty. And this is the core question I care about now: how do you build intelligence that continuously adapts and that builders everywhere can have more control? So instead of taking years of training to learn how to build the tools, scientists just skip to the questions. So a few weeks ago, we released Autoscientist. And Autoscientist is really about how do you automate the training of models itself. I'll share a few things that are really interesting about this. One, it's co-optimized the entire loop. So it's from data to alignment, and it chooses and self-evolves based upon the domain and the type of data. What's interesting as well is that it actually outperforms research staff, mainly because a lot of our research staff has experience with certain model types. And we're testing it across many different model architectures, different size models, as well as dense and mixture of experts. And that search base is a lot broader. And so exploiting it using how do you self-improve from experience and scale is very effective. I think this is very interesting. This only worked when we co-optimized the data. So there's a lot of auto research projects right now, which basically treat data as the agent. It decides whether to create data or not or what to do. Frankly, we did not get the returns for how much you can squeeze out of performance until you control for data quality. So we actually co-optimized, based on all the adaptation we did with the data, exactly what we would do with the model. And that was super interesting. It speaks to the need to control the entire flow. What's fun about this is really what Autoscientist is doing is it combines all the knowledge it gained from the adaptive data component with also the knowledge of the domain and also the ability to self-improve for a domain and to learn from other components of that domain. What is a cheeky fact, and this is quite fun, you'll notice all these percentages for win rates are 60 plus. And that's because we put the budget stopping it above 60. So once it was above 60, our agentic flow could exit. But we've since removed that barrier, and you can see it just go up over time, which is super fascinating. And then I think what's interesting is it changes a lot of the hyperparameters. Typically, humans are much more wary about changing all at once, and so you get massive exploitation of the search space. I see this as crucial for how do you reduce the amount of compute you use for customization because you train with much more predictability, but also how do you leverage your domain knowledge to really unlock how you build frontier AI. And this was fun. We announced a beta four weeks ago. The excitement is most acute for medical and science. And that's largely, I think, legal and code as well. But these are domains where typically current models fall short and also domains where, in many ways, the degree of last-mile customization is really acute. And this is the core point. I think this is fun. I think I might even have time to cover the other point I want to make about why now is very important for changing who shapes. But really, the main factor that this does is increase your innovation cycle and also increase the likelihood that when you train and spend compute, you'll succeed. And those combined factors are super interesting. One thing that we're doing next is extending that so even your test-time compute should be adaptive based on your task. So this kind of brings me back to where I started and the grumpy statement I said, which is we have this verified super narrow compounding issue of barriers to entry. One is that you need to do this very narrow funnel of who gets to build frontier AI, and the other is that typically compute and cost really dominate. We want to change that. We decided, okay, we're going to cover languages from day one, 242 languages. And also a big interest for us is actually non-verifiable tasks. I think this is super interesting because this is really the bulk of everyday tasks that people do. And it's really where the meat of what is interesting for progress is going to be over the next year. And this leads me into our mandate. We care deeply about how do you accelerate learning in a way that models should be able to learn from their environment. So right now we've moved from an era of the model is monolithic. When I was at different parts of my research career, basically your whole team would be around building a model.
One is that you need to do this very narrow funnel of who gets to build Frontier AI, and the other is that typically compute and cost really dominate. We want to change that. We decided, okay, we're going to cover languages from day one: 242 languages. And also a big interest for us is actually non-verifiable tasks. I think this is super interesting because this is really the bulk of everyday tasks that people do. And it's really where the meat of what is interesting for progress is going to be over the next year. And this leads me into our mandate. We care deeply about how do you accelerate learning in a way that models should be able to learn from their environment.
So right now we've moved from an era of the model being monolithic. When I was at different parts of my research career, your whole team would be around building a model. You give it to someone else to serve, and you have someone else do the front end. And actually now, the most important intelligence is a model that interacts. And so this idea of how efficiently are you going to interact, how will you continuously learn from the environment, is pretty core. And I think about it a lot. So let's see. I think I do have time, right? How are we doing for time? Oh, I do. I have plenty. This is lovely.
So we'll have time for questions, and I'll share a little bit about what I think the next component is. I think core to this: if we just did auto scientists, but it still took enormous compute to do Frontier AI trainings, I think we'd be in a bit of a pickle, right? I'd be saying, oh, great, you can use this agent, but don't worry, just bring your 10,000 GPUs with you. But I think there's another trend which makes this very important timing, and rooms like this probably much more optimistic than they have been a few years ago about who can build Frontier AI. And one of that is the rules of where you get rate of return for compute are totally changing.
So I wrote a paper about this called Slow Death of Scaling. But empirically, we do now know that pre-training size in particular is not your most lucrative axis of scale. And what does this mean? If pre-training scale isn't going to dominate performance, it actually really greatly changes who can create the best recipes for innovation. Because pre-training compute typically has to be co-located. It has to be, in many ways, large volume to accommodate for redundancy. Inference compute and other places where you actually apply compute, typically you can have much more distributed.
It's also much higher return given the amount of flops. And so it's interesting when we talk about what is the state of pre-training compute. We know it's not giving the same returns, largely because our architecture is saturated. So we see much smaller models outperforming much larger ones. This is the OpenLLM leaderboard. And this is the daily submission of the best small model under 13B versus all the larger models. And you can see over time that ratio totally flips. And also there's the grumpy assessment that most recent models that have severely played with just increasing model size haven't provided the same stepwise change as their predecessors.
And a lot of that is because where the most returns for performance are now are on a broader action space. And this is really what I was getting at when we move from an algorithm to expanding optimization space in new places. And what's fun about that is that these are new places where the barriers to entry are much more nimble and where recipe and algorithm and research matter again. And things like how do you automate that discovery. And so this is what I'll state, and I think then we should open up for questions.
And I would encourage good grumpy questions or fun positions. Let's make use of the time. I know I was told earlier that almost no talks have time for questions. I find that so disappointing. So we'll need some brave people to start the conversation. But I will say this means all bets are off. And I would say it's a very good time to be working on intelligence because instead of just a handful of people getting to create it, it's much more now about the question you want to answer. At the end of the day, the reason why people did a computer science PhD was to learn the tools to get to the question. And now you can just get to the question, which is super meaningful.
Okay, let me open up. Where should we start? We have an abundance. I hear there's no microphone. So if you want to ask a question, you want to make a statement, I will indulge a statement if it's interesting. Yeah, go for it. Just raise your hand and I'll repeat it afterwards. Let me just get to the end of this in case people want to reach me afterwards. Nice.
Yes, go ahead. Gentleman in the fourth row, go for it. You mentioned, yeah, I'm here looking for a market price. Can you point to how? So I think how is twofold. One is there's very few people who know how to train frontier models. I would say realistically probably less than 5,000 in the world at scale. I think that type of knowledge, that's a very exploitable search space. And actually as humans, all those configurations, we're not particularly good at.
It's kind of like secret knowledge we pass as if we're apprentices. So that's one. Once you automate a lot of that knowledge, you just accelerate innovation cycles, which means that you can explore and do more questions. Typically what people often miss is that the cost of asking something informs what is asked. And if you make it cheaper to ask something, you change the volume of things that are asked, which is super interesting. The other reason, though, I do think it's very much a facet of the changing nature of compute. So agentic compute, post-training compute, matters a significant amount for performance.
That does not require the same type of, dare I say, hoarding of GPUs. But I think it's very different compute purchasing dynamics. And again, it means that the person with the best idea has a higher chance of winning, which is fun. Nice. What else? Who wants to go? I see, yeah, we can go up here. And then I saw a hand back there. Okay. Yes, I do see you. The glare is high, but you go first and then we'll come up here. So once you're allowed to figure out a lot of that thinking, one of the challenges is to adapt the whole model.
Somebody will take your base model and adapt those to be something. How do you see that? Yeah. So the question, I'll just repeat it, because I think there's probably people in the room who want to know. So the question was, one of the, I guess, counterpoints from some frontier labs about not enabling frontier AI outside is a safety question. So I think it would be, I definitely am not one of those people who says that open source doesn't carry any risk. So when you make a tool more readily available, there's a profile of risk associated with it. Autoscientists is, to be fair, about enabling people to customize their models.
You can think of that as a slightly different question from whether those are open source. It's giving people way more control, whether that's local or private or within their company. It's about how do they own their own intelligence? What do I think broadly about the impact of open source on safety? The dynamic has often conflated that real risk of wider access with a slight sense that it restrains who can actually participate. And I think that's a delicate balance. And I think you have to acknowledge risk while also navigating that and acknowledging that it limits who can participate. Yeah, so nuanced answer. So I guess I should be more bombastic on that one.
But I guess I have been in this discussion a few times, and I find the binary views on other sides miss a lot. But anyways, okay, go ahead. Are there any specific research ideas or technologies that predict the paradigm of the type of learning? Oh, I think for automating and speeding up learning, one of the core questions is how do you balance what you store in the parametric space and the non-parametric space? And actually, one of the most interesting things, I mentioned that this only worked because we co-optimize data and model. It will only work to do an auto scientist for harnesses if you also co-optimize it with a model.
And so it's interesting. It's actually a long-horizon problem. And that's super fascinating to think about, where you're optimizing the choices for each and co-training, which is cool. Nice. I think we have time for maybe two more, and then we can pass on to the next speaker. Nice. Go ahead. So you talked a little bit about this, actually working on the full-training side of the model is cheaper than the first thing, obviously. But still, especially talking about really large models, reinforcement learning, even
Oh, I think for automating and speeding up learning, one of the core questions is: how do you balance what you store in the parametric space and the non-parametric space? And actually, one of the most interesting things, I mentioned that this only worked because we co-optimize data and model. It will only work to do an auto scientist for harnesses if you also co-optimize it with a model. And so it's interesting. It's actually a long horizon problem. And that's super fascinating to think about, where you're optimizing the choices for each and co-training, which is cool. Nice. I think we have time for maybe two more, and then we can pass on to the next speaker. Nice.
Go ahead. So you talked a little bit about this, actually working on the full-training side of the model, it's cheaper than the first thing, obviously. But still, especially talking really large models, reinforcement learning, even fine-tunes, it's pretty difficult, so you talked a little bit about the smaller models. But I still think most frontier smaller models still rely on the bigger knowledge, like the smaller models, like the smaller models. Yeah, actually that's an excellent point. I think the question amounts to two points. One, are larger models necessary for distillation benefits? And then second, frontier models are still pretty large.
So I think for the second one, frontier models are still pretty large. Yes. I don't think I'm arguing that. My argument is slightly different. And my argument is that no frontier AI lab is going to 4X the size of that model again for pre-training. So it's almost like we know we're at an upper ceiling for this architecture. If someone comes out with a new architecture, that's totally different. The architecture determines your ceiling. And I'm saying we are probably at the ceiling of size, which means that that's fun, because it means, okay, it's what you innovate within that. So size does matter. I think that's a very good point to bring up.
Meaning I'm not advocating everyone uses 0.8B, but I am saying that we now have a more equal playing field at the top. Second point is interesting, distillation, the impact, certainly. So data quality in general means you use capacity a lot more. So what you will see in pre-training is instead of size, people are just moving post-training further back, which is very fascinating and a bigger lever. So I agree distillation is helpful. It's just that, again, we've hit the ceiling. And so it's almost like no one is going to supersize their model. Or if they do, it's not clear it's beneficial except for a small size of the distribution, which is very much the long tail.
And that's interesting, where that tradeoff is worth that much pre-training compute. So very good question. One more, and then I think we are done. Yes, go ahead. Can you add the old installers? Can you add the GPUs to the data? Yeah, it's actually in beta. So you can, I shared here, you can try it in beta. So we actually are offering the GPUs for free. Okay, oh, that's a nice question. I promise I don't know this gentleman. But yes, I think actually we're trying to remove the compute hurdle, and I think it's quite cool to see. So feel free to take a look at the beta. Nice, lovely, thank you so much. Really nice, thank you. Thank you. Nice.
I think we have time for maybe two more and then we can pass on to the next speaker. Nice. Go ahead. So you talked a little bit about this, actually working on the full-training side of the model, it's like cheaper than the first thing, obviously. But like, still, especially talking like really large models, like reinforcement learning, even 5-tunes, it's pretty, like, difficult, so like you talked a little bit about the smaller models. But I still think like most, like, frontier smaller models still rely on, like, the bigger knowledge, like, the smaller models, like, the smaller models. Yeah, actually that's an excellent point. I think the question amounts to two points.
One, are larger models necessary for distillation benefits? And then second, so frontier models are still pretty large. So I think for the second one, frontier models are still pretty large. Yes. I don't think I'm arguing that you, my argument is slightly different. And my argument is that no frontier AI lab is going to 4X the size of that model again for pre-training. So it's almost like we know we're at an upper ceiling, for this architecture. If someone comes out with a new architecture, that's totally different. You can, the architecture determines your ceiling. And I'm saying we are probably at the ceiling of size, which means that that's fun,
because it means, okay, it's what you innovate within that. So size does matter. I think that's a very good point to bring up. Meaning I'm not advocating everyone uses 0.8B, but I am saying that we now have a more equal playing field at the top. Second point is interesting, distillation, like, the impact, certainly. So data quality in general means you use capacity a lot more. So what you will see in pre-training is instead of size, people are just moving post-training further back, which is very fascinating and a bigger lever. So I agree distillation is helpful. It's just that, again, we've hit the ceiling.
And so it's almost like no one is going to supersize their model. Or if they do, it's not clear it's beneficial except for a small size of the distribution, which is very much the long tail. And that's kind of interesting, like, where that tradeoff is worth that much pre-training compute. So very good question. One more and then I think we are done. Yes, go ahead. Can you add the old installers? Can you add the GPUs to the data? Yeah, it's actually in beta. So you can, I shared here, you can try it in beta. So we actually are offering the GPUs for free. Okay, oh, that's a nice question. I promise I don't know this gentleman.
But yes, I think actually we're trying to remove the compute hurdle and I think it's quite cool to see. So feel free to take a look at the beta. Nice, lovely, thank you so much. Really nice, thank you.
Thank you.