The Base Model Is Dead — Varun Singh, Arcee AI
Description
The old story is that a base model is a mirror of the internet, a good model of human web text that everything else gets bolted onto. Varun Singh, who leads pre-training at Arcee AI, argues that story is dead: no modern base model reflects the web the way GPT-3 once did. Instruction data and synthetic reasoning traces have moved earlier and earlier into training, and a distinct mid-training stage has emerged for longer datapoints that look much more like the downstream capabilities you actually want. Reading recent open recipes, from Nemotron to Kimi K2, the pattern is clear: raw web text is taking a backseat. The rest of the talk is what that shift does to how you build. Once reinforcement learning became the thing that got models to reason, the base model stopped being a cherry on top and started needing to carry the prior that RL builds on, which changes the data mix and pulls post-training-flavored data forward. Singh walks through the practical pitfalls his team hit training the Trinity series, like getting the balancing coefficients right and establishing stable representations early so the model is prepared for what it must compose during RL. The message is that as capabilities advance, the base model's job keeps redefining itself, and pretending it still just mirrors the internet will cost you. Speaker info: - https://x.com/stochasticchasm - https://www.linkedin.com/in/varun-singh-cs Timestamps: 0:00 - The base model as a mirror of the web 1:26 - How knowledge accumulates in training 2:49 - When instruction data moves earlier 4:11 - After o1: RL and reasoning 5:41 - What prior the base model must carry 6:18 - Filtering web text, adding synthetic 8:01 - Reading the open data recipes 9:41 - Lessons from training Trinity 12:02 - Balancing coefficients and early stability 13:30 - Why RL keeps raising the stakes 15:55 - The base model's shifting job
Summary
Generated by gpt-5.6-terraAt-a-Glance
- Verdict: Watch fully
- Core thesis: The traditional web-trained base model is no longer the main determinant of LLM capability; supervised training should increasingly be designed as a prior that prepares models for large-scale reinforcement learning, reasoning, and agentic work.
- Why it matters: For teams building agent systems, the implication is that model quality will depend less on generic web knowledge and more on whether pre-training data, architecture, and training distributions anticipate tool use, long-horizon traces, and RL environments.
- Best use: Use this as a concise technical framework for evaluating model-training claims, especially when comparing general-purpose base models against models optimized for coding, reasoning, and agents.
Executive Summary
Varun Singh argues that the old concept of a base model—a broad representation of human internet knowledge acquired primarily through next-token prediction over web text—is becoming inadequate. Historically, pre-training consumed most compute and determined downstream quality, while post-training mostly made a capable model usable in chat. In the reasoning-and-agent era, RL is no longer cosmetic alignment: it can materially improve performance on mathematics, software engineering, tool use, and other verifiable tasks.
This changes the role of supervised learning. Rather than treating pre-training as the product and RL as a final polish, Singh proposes viewing supervised learning as preparation: it should establish the atomic skills, data distributions, task formats, context lengths, and stable model representations that RL can later compose and optimize. The relevant prior is increasingly reasoning and agentic behavior rather than a maximal general-web prior.
He points to a data-mix shift across modern training recipes. GPT-3 used roughly 85% web text including Wikipedia, while newer models sharply reduce general web data and increase code, STEM, instructional, and synthetic data. Microsoft MAI Thinking 1 reportedly limits web text to 15% while avoiding synthetic model-generated data, whereas NVIDIA's Nemotron 3 Ultra explicitly brings SFT-style question-answer data into pre-training and leans heavily into synthetic data.
The operational message is not that base models literally no longer matter. They matter differently: a base model must expose the capabilities that RL will need, including task-specific formats and possibly long agent traces. This is particularly consequential for MoE training, where an abrupt shift from generic pre-training data to post-training data can destabilize expert routing and require aggressive late-stage load-balancing intervention.
Key Takeaways
- Claim: RL has shifted from a minor post-training alignment layer to a major source of capability gains, making the original base-model objective insufficient as the central measure of model quality. | Evidence: Singh contrasts earlier chat-oriented post-training with OpenAI o1's reasoning results, DeepSeek R1's public demonstration of reasoning-model construction, and coding agents that can be trained end-to-end to interact with software environments. He also cites Xiaomi Mimo Labs as allocating roughly equal compute to pre-training and post-training, and says Kimi's K2.5 reportedly used more RL compute than supervised-learning compute. | Implication: When evaluating frontier models for agentic work, assess their RL environment, reward/verifier quality, and training distribution—not just base-model size, benchmark pretraining quality, or general knowledge coverage. | Caveat: The speaker does not claim supervised learning can disappear: human language remains an exceptionally broad distribution, and he frames a full AlphaGo-like replacement of supervised learning by RL as uncertain for language models.
- Claim: Modern pre-training data is shifting away from broad web text toward code, STEM, instruction-shaped, and agentically useful data. | Evidence: GPT-3's mix was described as roughly 85% web text including Wikipedia, while LLaMA 3 still used 50% general-knowledge tokens. Singh says MAI Thinking 1 reduces web text to 15%, and notes that code—absent as a dedicated GPT-3 subset—has become a dominant component of current recipes. | Implication: Generic internet-corpus breadth is becoming a weaker proxy for usefulness in coding and agent deployments; data composition should be judged against the target interaction and environment. | Caveat: The transcript presents selected published recipes rather than a universal optimal mix; MAI Thinking 1 and Nemotron represent contrasting approaches.
- Claim: Synthetic data can improve pre-training when it is used to clean, reshape, and diversify useful seed information rather than indiscriminately inflate token counts. | Evidence: Arcee's Trinity Large Thinking used web-scale synthetic rephrasing to present the same seed information in multiple forms. Singh also cites Kimi K2's broad use of this approach and Swallow Code/Swallow Math as earlier examples. He argues synthetic generation can yield higher-quality tokens and make pre-training examples resemble instruct or agentic tasks. | Implication: For a proprietary agent or domain model, synthetic traces, reformulations, and task-shaped examples are potentially valuable, but should be treated as curated training assets with validation—not cheap substitutes for quality source data. | Caveat: Singh acknowledges the concern that blindly adding synthetic data can cause model collapse or degrade performance; data generation and filtering quality are therefore central.
- Claim: Introducing post-training-like distributions earlier can be especially important for MoE models because late distribution shifts can create expert-load imbalance. | Evidence: Singh describes MoE load balancing as a core difficulty: experts specialize during training, and objectives try to keep utilization broadly balanced. He says MAI Thinking 1 encountered a large pre-training-to-SFT distribution shift and addressed it by substantially increasing the load-balancing coefficient during SFT. | Implication: MoE model builders should test routing and utilization under the eventual instruction, reasoning, and agent-trace distribution early, rather than assuming generic-corpus balance will persist through post-training. | Caveat: The proposed remedy is architectural/training-process specific to MoEs; it is not presented as a general reason every dense model must mix SFT data into pre-training.
- Claim: The useful role of a base model is to supply atomic skills that RL can compose, not merely to store a broad world model. | Evidence: Singh cites work suggesting that a model needs exposure to the atomic skills required for a downstream task, after which RL can extrapolate and compose them if the environment is sufficiently difficult. He compares the direction conceptually with AlphaGo, where RL ultimately surpassed supervised learning. | Implication: Before investing in agent RL, identify the irreducible prerequisite skills—tool syntax, code patterns, planning formats, domain primitives, and long-context behavior—and ensure they are present in supervised training. | Caveat: This depends on having an RL environment with adequate difficulty and feedback; RL cannot reliably create capabilities from entirely absent prerequisite representations.
- Claim: Training data should include novel interaction forms and test-time-compute patterns before RL so the model can explore them effectively later. | Evidence: Singh identifies reasoning traces as a novel format that does not closely resemble ordinary human output, and points to xAI's XIA1 paper as an example of warming models during SFT or pre-training for test-time-compute schemes. | Implication: If Ken's systems depend on deliberation loops, structured tool trajectories, or extended contexts, model selection and fine-tuning should prioritize prior exposure to those exact interaction shapes over generic chat fluency. | Caveat: The transcript does not specify which test-time-compute formats generalize best or quantify their incremental value.
Detailed Brief
Two competing recipes for the new base-model prior
- Claims: The field has not converged on whether synthetic data is necessary for the next generation of base models.; Both the human-data-heavy and synthetic-data-heavy approaches still depart from the older web-dominant paradigm.
- Evidence: MAI Thinking 1 reportedly emphasizes avoiding synthetic data and filtering web scrapes to exclude data from other language models.; Nemotron 3 Ultra reportedly incorporates SFT-prefixed, question-answer-style data directly into pre-training, rather than reserving it for post-training.; Singh frames these as opposing examples despite their shared reduction in the proportional role of general web text.
- Caveats: Published data recipes offer incomplete visibility into filtering, quality controls, generation methods, and experimental controls, so their reported mixes should not be read as direct causal proof.
- Implications: The strategic decision is not simply human versus synthetic data; it is whether the resulting corpus creates robust representations for the downstream training regime and deployment interface.
Replacing phase labels with a two-paradigm view
- Claims: Singh considers the labels pre-training, mid-training, post-training, and RL increasingly blurry and less useful than distinguishing supervised next-token learning from RL.; Mid-training reflects the practical need to expose models to distributions closer to RL and agent deployment, including longer contexts.
- Evidence: He describes mid-training as bringing in the distribution likely to appear in post-training RL and adding longer context for agentic traces.; He argues that datasets used at mid-training could often be introduced earlier to establish more stable representations from the start.
- Caveats: This is a conceptual reframing, not a claim that pipeline stages have disappeared operationally.
- Implications: Training-pipeline discussions should focus on distribution continuity and capability prerequisites across stages, rather than treating phase names as hard boundaries.
Notable Concepts & Terms
- Base model prior: Singh recasts the base model as the representation and skill foundation needed for later RL, rather than a general-purpose compression of all web knowledge.
- Atomic skills: Prerequisite capabilities that supervised learning must establish before RL can combine and optimize them into more complex behavior.
- Synthetic rephrasing: Generating alternate expressions of seed data to increase coverage, improve task shape, and reinforce information without relying solely on new raw-source tokens.
- SFT data in pre-training: Moving instruction-style question-answer examples upstream so models learn expected downstream task formats before conventional post-training.
- MoE load balancing: The process of preventing a mixture-of-experts model from overusing or underusing particular experts; sharp distribution shifts can destabilize it.
- Mid-training: An intermediate stage that acclimates a model to longer contexts and distributions closer to reasoning, agent traces, or RL workloads.
- Test-time compute: Inference-time methods that let a model spend additional computation on reasoning or search; Singh argues models can be warmed up for these schemes in training.
Operator Notes / Why Ken Should Care
- Add a model-evaluation criterion for distribution match: compare candidate models' exposure to tool calls, coding tasks, structured reasoning, long contexts, and agent trajectories against the target workflow.
- For any proposed agent-RL program, write an atomic-skill inventory first and use it to identify missing supervised-training or fine-tuning data before building rewards.
- If considering MoE-based models or custom MoE training, require routing-utilization testing on the intended production distribution, not only generic pre-training or chat benchmarks.
- Treat synthetic agent traces and rephrased domain data as an experiment with quality gates: retain source provenance, use held-out evaluations, and test for degradation before scaling generation.
- When assessing vendor claims about a strong 'base model,' ask how much of the observed capability comes from supervised data, RL compute, reward environments, and test-time compute rather than parameter count alone.
Source/Metadata
- Title: The Base Model Is Dead — Varun Singh, Arcee AI
- Transcript words: 2102
- Duration seconds: 1064
- Timestamp note: No timestamps or chapter markers were present in the supplied transcript.
Transcript
Hi everyone, my name is Varun. I'm the pre-training lead at RCAI. And the talk I'm going to be giving today is called The Base Model is Dead, but not really. The idea of the base model that we have is built on this idea of training on super large-scale web text and the base model being a reflection of the whole knowledge of the human internet. You can see, I've taken these from a bunch of different papers on the entire LLM training process. Our own model, RCA, Trinity Lodge Thinking. The process looked like the simplified diagram on the left. I've taken the top one from GLM 4.5, the bottom one from GLM 5. All these have a pre-training phase. And pre-training is the stage where the model accumulates world knowledge, builds useful representations, all through next-token prediction on web text. I've got a simplified transformer diagram, decoder-only transformer, and a screenshot from the GPT-3 paper that talks about how language models can learn how to do in-context learning through unsupervised or self-supervised, or some even just call it supervised learning, through next-token prediction. The way that older base models were trained was, like I said, mostly on things that reflected the entirety of human knowledge. So Common Crawl, which is a commonly available web scrape, made up most of the training dataset for GPT-3. WebText2, another web scrape dataset. Some sources from books as well. And Wikipedia is a high-quality representation of human knowledge. You can see that web text alone here, including Wikipedia, makes up roughly 85% of the whole training mix. Looking at the bottom with LLaMA-3, web text still makes up a majority of the model's training data, with 50% of the tokens corresponding to general knowledge. Back then, post-training was mostly shaping the model to use the parts, to surface the knowledge that it accumulated in pre-training in a chat interface. So mostly allowing the model to adapt to a chat template, to the question-answer format, and be useful in an interaction that way. RL was mostly just a cherry on top, shaping the flavor of the interactions more than conferring extra knowledge or quality onto the base model itself. Now, in this realm of how language models used to be, pre-training and the base model defined how good you were able to get a model. It was the bulk of the compute budget, and it was the core of the training process. However, this changed a lot last year, when OpenAI, I guess 2024 actually, OpenAI released 01, pioneering reasoning models, and DeepSeq also released R1 in January 2025, allowing the whole world to know how to build these types of language models. And now we have this new use for reinforcement learning, which is no longer a cherry on top, but it can dramatically improve the performance of the model on various different tasks. The famous graphs in 01 there, talking about AIME performance, competitive math contest. And then even later in the year, we saw cloud code come into being as a way for developers to easily use language models in a terminal to build out applications as models got stronger and stronger on things like function calling. And then people realized you could RL this end to end. And now models could learn how to interact with software environments and build software and perform really useful work. And so the question then becomes: is your standard base model still the best, what would be the best prior for this large-scale reinforcement learning phase that reasoners and agentic models now use? And we can see in a few open research papers what the trend is, where the trend is going. And, interestingly enough, it seems not super clear yet. I have my opinions on synthetic data being the way forward, but I've got two contrasting perspectives here on the slide. The top image is from the MEI thinking one paper where they make a really large point to not use any synthetic data or any data from any other language model. And they really try to filter that web scrape for this as well in order to adhere to the previous paradigm of using human knowledge as a way to bootstrap model representations and capabilities. But I would say that this is also, even though they stuck with no synthetic data, the data mix that they've chosen here is still totally different from what you'd expect in a classical language model. And the main reason for that is that web text, which used to make up to 85% of the training data in GPT-3, is now all the way down at 15%. And that just shows that the value of web text contributing to the downstream performance of the models on RL and stuff is still important, but taking a backseat to things like code and STEM abilities as the models gain more real-world use cases related to those. The other approach is to bring post-training data and large-scale synthetic data back through, pull it back through the process into the pre-training phase. The bottom chart I've taken from NemoTron 3 Ultra, where they reveal their data recipe, and I'm not sure how readable it is, but these top three on the left pie chart, the top three on the right side of it, they're all labeled SFT, with SFT as a prefix. And that's the type of question-and-answer chat dataset that you'd expect to see only in post-training. But by pulling it back into the process, they're able to get the model to learn the shape of these conversations and what kind of tasks they might be expected to do downstream from the very beginning of the pre-training process. And this follows a similar trend in diminishing the amount of web text used in the model. Yeah. It's really interesting to see the NemoTron series lean so heavily into synthetic data, but MAI thinking one lean in the opposite direction. I've just got this slide here as an easy contrast that people can see on the amount of web text and the amount of books and stuff being less of a percentage here. And GPT-3 didn't even used to have any specific code datasets, but now code is the dominating data subset that we have in pre-training recipes. So I mentioned synthetic data, but what is actually, how is synthetic data used? There's a lot of talk around synthetic data that blindly tossing it into a model can cause the model to collapse and performance to tank. But there's been a lot of work and even a large-scale example of this turning out really well. So in our own model, Trinity Lodge, we had a large amount of web-scale synthetic data, mostly through rephrasing, where you take a seed data item and you upsample it in the mix by generating synthetic rephrases. So the model sees the same information in multiple ways. The bottom two screenshots are from Kimi K2, an even larger-scale model that broadly used this across the whole pre-training dataset. The top-right screenshot is from a paper that resulted in the datasets Swallow Code and Swallow Math, which are early examples of this. But the trend seems to be that synthetic data not only allows you to get more and more tokens, but also clean up tokens, get higher-quality tokens, and have tokens that are shaped more like instruct or agentic tasks all the way back in pre-training, allowing the model to lower those task representations from the very beginning. Another reason that it's beneficial to add post-training data early in pre-training is now with MoEs. One of the biggest pain points in training an MoE is dealing with load balancing, where experts can specialize over the course of training, and load-balancing objectives aim to achieve broadly equal utilization of the experts in a given batch, or sequence, depending on the objective. So, without post-training data early in pre-training, and with an MoE, one really easy pitfall that we can fall into is this, as illustrated in the MAI thinking one report, which is that the data distribution that the model sees in post-training is really, really different compared to what it sees in pre-training. And this can cause massive imbalances, and MAI overcame it by really cranking up the load-balancing coefficient during the SFT stages. But ideally you don't want to mess with the balance that far into training, and the model should learn stable representations from really early on. Another interesting thing that is changing in base models now is that there's this whole advent of mid-training, which is exposing the model to the distribution that it would see during post-training in RL, and added longer context, so for things like agentic traces to be allowed into the mix, and to help prepare the model that way. A lot of models, though, are training with much longer context in pre-training, and there's no reason that these datasets can't be pulled back into the mix to allow for more stable representations from the very beginning. I think a better way to understand the current phase of LLM training isn't so much pre-training, mid-training, post-training, RL, it all gets a bit muddy that way, but there's two broad paradigms that really help build LLMs today, and that's supervised learning through next-token prediction, and RL. And RL is becoming more and more important. The bottom thing is a screenshot from an interview with the head of Xiaomi's Mimo Labs, where she talks about how they allocate compute between research, pre-training, and post-training, and pre-training and post-training in the final model have a roughly equal compute allocation. Composer 2.5 takes us to the extreme where COSO really sank much, much more RL compute into the model than the model had ever seen in supervised learning. But with RL dominating such a massive amount of the compute budget, it makes sense to view supervised learning as a way specifically to prepare the model for, to build useful representations for RL instead of it being the bulk of what the model would be used for previously. There's been some work on how supervised learning affects RL. I really like this one paper where the main takeaways are basically that the base model needs to have some exposure to the atomic skills that it would need to compose during RL, and the model can learn to extrapolate from there during RL, given the environment has a sufficient level of difficulty. I had to put in the classic AlphaGo graph there where RL eventually overtakes supervised learning. It's unclear if we'll see something like that for language models, because, of course, human language is such an insane distribution to have to learn through reinforcement learning alone. But it's definitely possible that we might see diminished supervised learning and more and more RL, which makes this kind of thinking of a base model as atomic skills for RL more and more valuable. Another thing that some labs are doing is introducing novel data during supervised learning. And by novel, I mean something that the model really wouldn't have seen the shape of before. An easy example is reasoning traces. They don't really look like a ton of what humans output. And another interesting thing is training for test-time compute schemes by warming the model up to them during SFT or even pre-training itself. These screenshots were taken from Xiphar's XIA1 paper. And I think that they're very interesting ways of thinking about how data can affect the skills needed to explore well in RL. In conclusion, base models have moved from general human knowledge and world priors to reasoning and agentic behavior priors. Of course, that's reductive in a way, that reasoners and agents are the main way we see bots, the main way we see these chatbots used now. But if a new paradigm were to take off, a new way of interacting with the models, it makes sense to think of a base model as building a prior for that instead of just building off a massive scrape of web text. And yeah, thanks for listening. Thanks for your time. And yeah. And yeah.