Hi everyone, my name is Varun.
I'm the pre-training lead at RCAI. And the talk I'm going to be giving today is called The Base Model is Dead, but not really. The idea of the base model that we have is built on this idea of training on super large-scale web text and the base model being a reflection of the whole knowledge of the human internet. You can see, I've taken these from a bunch of different papers on the entire LLM training process. Our own model, RCA, Trinity Lodge Thinking. The process looked like the simplified diagram on the left. I've taken the top one from GLM 4.5, the bottom one from GLM 5. All these have a pre-training phase. And pre-training is the stage where the model
accumulates world knowledge, builds useful representations, all through next-token prediction on web text. I've got a simplified transformer diagram, decoder-only transformer, and a screenshot from the GPT-3 paper that talks about how language models can learn how to do in-context learning through unsupervised or self-supervised, or some even just call it supervised learning, through next-token prediction. The way that older base models were trained was, like I said, mostly on things that reflected the entirety of human knowledge. So Common Crawl, which is a commonly available web scrape, made up most of the training dataset for GPT-3.
WebText2, another web scrape dataset. Some sources from books as well. And Wikipedia is a high-quality representation of human knowledge. You can see that web text alone here, including Wikipedia, makes up roughly 85% of the whole training mix. Looking at the bottom with LLaMA-3, web text still makes up a majority of the model's training data, with 50% of the tokens corresponding to general knowledge. Back then, post-training was mostly shaping the model to use the parts, to surface the knowledge that it accumulated in pre-training in a chat interface. So mostly allowing the model to adapt to a chat template, to the question-answer format,
and be useful in an interaction that way. RL was mostly just a cherry on top, shaping the flavor of the interactions more than conferring extra knowledge or quality onto the base model itself. Now, in this realm of how language models used to be, pre-training and the base model defined how good you were able to get a model. It was the bulk of the compute budget, and it was the core of the training process. However, this changed a lot last year, when OpenAI, I guess 2024 actually, OpenAI released 01, pioneering reasoning models, and DeepSeq also released R1 in January 2025, allowing the whole world to know how to build these types of language models.
And now we have this new use for reinforcement learning, which is no longer a cherry on top, but it can dramatically improve the performance of the model on various different tasks. The famous graphs in 01 there, talking about AIME performance, competitive math contest. And then even later in the year, we saw cloud code come into being as a way for developers to easily use language models in a terminal to build out applications as models got stronger and stronger on things like function calling. And then people realized you could RL this end to end. And now models could learn how to interact with software environments and build software and perform really useful work.
And so the question then becomes: is your standard base model still the best, what would be the best prior for this large-scale reinforcement learning phase that reasoners and agentic models now use? And we can see in a few open research papers what the trend is, where the trend is going. And, interestingly enough, it seems not super clear yet. I have my opinions on synthetic data being the way forward, but I've got two contrasting perspectives here on the slide. The top image is from the MEI thinking one paper where they make a really large point to not use any synthetic data or any data from any other language model. And they really try to
filter that web scrape for this as well in order to adhere to the previous paradigm of using human knowledge as a way to bootstrap model representations and capabilities. But I would say that this is also, even though they stuck with no synthetic data, the data mix that they've chosen here is still totally different from what you'd expect in a classical language model. And the main reason for that is that web text, which used to make up to 85% of the training data in GPT-3, is now all the way down at 15%. And that just shows that the value of web text contributing to the downstream performance of the models on RL and stuff is still important,
but taking a backseat to things like code and STEM abilities as the models gain more real-world use cases related to those. The other approach is to bring post-training data and large-scale synthetic data back through, pull it back through the process into the pre-training phase. The bottom chart I've taken from NemoTron 3 Ultra, where they reveal their data recipe, and I'm not sure how readable it is, but these top three on the left pie chart, the top three on the right side of it, they're all labeled SFT, with SFT as a prefix. And that's the type of question-and-answer chat dataset that you'd expect to see only in post-training. But by pulling it back into the process,
they're able to get the model to learn the shape of these conversations and what kind of tasks they might be expected to do downstream from the very beginning of the pre-training process. And this follows a similar trend in diminishing the amount of web text used in the model. Yeah.
It's really interesting to see the NemoTron series lean so heavily into synthetic data, but MAI thinking one lean in the opposite direction. I've just got this slide here as an easy contrast that people can see on the amount of web text and the amount of books and stuff being less of a percentage here. And GPT-3 didn't even used to have any specific code datasets, but now code is the dominating data subset that we have in pre-training recipes. So I mentioned synthetic data, but what is actually, how is synthetic data used?
There's a lot of talk around synthetic data that blindly tossing it into a model can cause the model to collapse and performance to tank. But there's been a lot of work and even a large-scale example of this turning out really well. So in our own model, Trinity Lodge, we had a large amount of web-scale synthetic data, mostly through rephrasing, where you take a seed data item and you upsample it in the mix by generating synthetic rephrases. So the model sees the same information in multiple ways. The bottom two screenshots are from Kimi K2, an even larger-scale model that broadly used this across the whole pre-training dataset.
The top-right screenshot is from a paper that resulted in the datasets Swallow Code and Swallow Math, which are early examples of this. But the trend seems to be that synthetic data not only allows you to get more and more tokens, but also clean up tokens, get higher-quality tokens, and have tokens that are shaped more like instruct or agentic tasks all the way back in pre-training, allowing the model to lower those task representations from the very beginning. Another reason that it's beneficial to add post-training data early in pre-training is now with MoEs.
One of the biggest pain points in training an MoE is dealing with load balancing, where experts can specialize over the course of training, and load-balancing objectives aim to achieve broadly equal utilization of the experts in a given batch, or sequence, depending on the objective. So, without post-training data early in pre-training, and with an MoE, one really easy pitfall that we can fall into is this, as illustrated in the MAI thinking one report, which is that the data distribution that the model sees in post-training is really, really different compared to what it sees in pre-training.
And this can cause massive imbalances, and MAI overcame it by really cranking up the load-balancing coefficient during the SFT stages. But ideally you don't want to mess with the balance that far into training, and the model should learn stable representations from really early on. Another interesting thing that is changing in base models now is that there's this whole advent of mid-training, which is exposing the model to the distribution that it would see during post-training in RL, and added longer context, so for things like agentic traces to be allowed into the mix, and to help prepare the model that way.
A lot of models, though, are training with much longer context in pre-training, and there's no reason that these datasets can't be pulled back into the mix to allow for more stable representations from the very beginning. I think a better way to understand the current phase of LLM training isn't so much pre-training, mid-training, post-training, RL, it all gets a bit muddy that way, but there's two broad paradigms that really help build LLMs today, and that's supervised learning through next-token prediction, and RL. And RL is becoming more and more important. The bottom thing is a screenshot from an interview with the head of Xiaomi's Mimo Labs,
where she talks about how they allocate compute between research, pre-training, and post-training, and pre-training and post-training in the final model have a roughly equal compute allocation. Composer 2.5 takes us to the extreme where COSO really sank much, much more RL compute into the model than the model had ever seen in supervised learning. But with RL dominating such a massive amount of the compute budget, it makes sense to view supervised learning as a way specifically to prepare the model for, to build useful representations for RL instead of it being the bulk of what the model would be used for previously.
There's been some work on how supervised learning affects RL.
I really like this one paper where the main takeaways are basically that the base model needs to have some exposure to the atomic skills that it would need to compose during RL, and the model can learn to extrapolate from there during RL, given the environment has a sufficient level of difficulty. I had to put in the classic AlphaGo graph there where RL eventually overtakes supervised learning. It's unclear if we'll see something like that for language models, because, of course, human language is such an insane distribution to have to learn through reinforcement learning alone.
But it's definitely possible that we might see diminished supervised learning and more and more RL, which makes this kind of thinking of a base model as atomic skills for RL more and more valuable. Another thing that some labs are doing is introducing novel data during supervised learning. And by novel, I mean something that the model really wouldn't have seen the shape of before. An easy example is reasoning traces. They don't really look like a ton of what humans output. And another interesting thing is training for test-time compute schemes by warming the model up to them during SFT or even pre-training itself. These screenshots were taken from Xiphar's XIA1 paper.
And I think that they're very interesting ways of thinking about how data can affect the skills needed to explore well in RL. In conclusion, base models have moved from general human knowledge and world priors to reasoning and agentic behavior priors. Of course, that's reductive in a way, that reasoners and agents are the main way we see bots, the main way we see these chatbots used now. But if a new paradigm were to take off, a new way of interacting with the models, it makes sense to think of a base model as building a prior for that instead of just building off a massive scrape of web text. And yeah, thanks for listening. Thanks for your time. And yeah. And yeah.