Open Reader

Data Quality Is the Compute Multiplier — Ari Morcos, DatologyAI

completed 19:05 Jul 31, 2026 Watch on YouTube

Current Status

completed

Video ID

_PdK6x7PQNM

RAG / Chat

Enabled
Data Quality Is the Compute Multiplier — Ari Morcos, DatologyAI
Description

Swap compute for data on the scaling curve and the same money buys a better model, which is why Ari Morcos calls data quality the compute multiplier and the most underinvested part of training. His frame is an oil refinery for data rather than a firehose: clean, curate, create, and compose, with quality classifiers, deduplication, and synthetic generation each earning their place, and the sequencing across stages mattering as much as any single step. The scarce resource now is not tokens but signal per token, and finding data that is optimal for a given target is where the leverage hides. The proof points are concrete. Better curated data lets a small multilingual model beat far larger ones trained on many more tokens, and it buys real inference efficiency because a model reaches the same quality with less. Morcos points to DatologyAI's customer results, from Thomson Reuters gaining on proprietary legal data in mid-training to Arcee's Trinity reaching the open frontier on public data alone. The closing argument is blunt: it is cheaper to manufacture high quality data than to buy more compute, so data curation is quietly shaping the future of model training. Speaker info: - https://x.com/arimorcos - https://www.linkedin.com/in/arimorcos/ - http://www.arimorcos.com/ Timestamps: 0:00 - Data is all we think about 0:52 - Why good data became scarce 2:19 - Swapping compute for data on the curve 3:48 - An oil refinery for data 5:52 - Curation work at DatologyAI 6:54 - Proof: small models beating bigger ones 8:58 - Inference efficiency from better data 9:24 - Multilingual gains 12:16 - Synthetic data done right 14:10 - Thomson Reuters and Arcee results 17:43 - Cheaper than buying compute

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: High-quality, task-specific data curation steepens model scaling curves, allowing teams to achieve comparable or better model performance with far less training and inference compute.
  • Why it matters: As frontier inference capacity tightens and reasoning models consume substantially more tokens, data quality is presented as the most practical lever for reducing the cost and risk of building, adapting, and operating capable models.
  • Best use: Use this as a model-development strategy brief: apply its data-quality framework to any custom-model, mid-training, synthetic-data, or domain-adaptation roadmap, while treating the reported gains as vendor-presented results to validate independently.

Executive Summary

Ari Morcos argues that data quality should be treated as a compute multiplier rather than a hygiene exercise. His core mechanism is straightforward: better data increases marginal information gained per token, steepening the performance-versus-data/compute curve. In an environment where H100 costs have risen, reasoning models consume roughly 8x the tokens of non-reasoning models, and API capacity may be constrained, this can be more valuable than simply acquiring more compute.

The proposed operating model has four stages: clean data, curate it for quality/relevance/diversity, create synthetic variants from the best source documents, and compose data mixtures and curricula across training phases. Morcos emphasizes that there is no universally optimal dataset: the right corpus depends on the intended tasks. Effective curation therefore includes task-distribution matching, semantic redundancy reduction, domain balancing, benchmark decontamination, and deliberate up/downsampling.

The talk's strongest practical claim is that data decisions compound across pre-training, domain mid-training, and post-training. In a Thomson Reuters legal-model example, 100 billion curated mid-training tokens reportedly raised LegalBench performance by about five points without degrading general capabilities; the subsequent post-training gains nearly tripled versus beginning from a default instruction-tuned model. The proposed explanation is that a better initial policy makes downstream optimization more productive.

Morcos also frames synthetic data as controlled re-expression rather than knowledge generation: transform vetted source documents into many answerable formats, such as true/false questions, to expand diversity without relying on a closed model to introduce new facts. His conclusion is not merely to collect more tokens, but to repeatedly expose models to high-signal data and to treat scalable data scoring, selection, and curriculum design as a frontier engineering capability.

Key Takeaways

  • Claim: Data quality can substitute for substantial amounts of training compute by making each token more informative and steepening the model scaling curve. | Evidence: Morcos cites Datology's vision-language-model results: curation of a roughly 25-billion-token Mammoth input dataset produced about a 14 absolute-point improvement with other conditions held constant; he also claims performance near Qwen 3.5 4B with 145x less training compute. | Implication: Before funding a larger training run, Ken should treat data selection, deduplication, task alignment, and mixture design as first-order optimization variables with potentially higher ROI than additional GPU budget. | Caveat: These are speaker/vendor-presented benchmark results rather than independently reproduced comparisons, and the transcript does not provide the full training configurations or benchmark methodology.
  • Claim: The optimal training corpus is use-case-specific; quality means relevant, diverse, information-dense data mixed to match the target task distribution. | Evidence: Morcos contrasts legal versus healthcare models, argues that robustness failures often originate in insufficiently diverse training examples, and lists quality classifiers, topical taxonomies, semantic redundancy reduction, quality/relevance-based sampling, and task-distribution matching as curation tools. | Implication: A custom-model program should start from a defined task and evaluation suite, then construct a data policy around those tasks rather than treating a larger generic corpus as automatically superior. | Caveat: The talk does not provide a universal scoring function or a concrete recipe for setting mixture weights; these require evaluation against the target workload.
  • Claim: Data curation can improve inference efficiency as well as training efficiency, including by producing more concise responses. | Evidence: For the cited VLM comparisons, Morcos says curated-data models were markedly more concise and achieved similar performance to Qwen 3.5 using approximately 35x fewer FLOPs per correct answer. | Implication: For high-volume agent systems, evaluate curated or adapted models on quality-adjusted cost per successful task, not just raw benchmark accuracy or tokens per answer. | Caveat: Conciseness is only valuable where it does not remove required reasoning, citations, safety checks, or user-needed detail; the transcript does not establish behavior across agentic workflows.
  • Claim: Small curated training runs can de-risk large-scale training because curated-data scaling behavior may predict outcomes at much higher compute. | Evidence: In multilingual experiments, two dense Llama-style models trained on curated data for one trillion tokens defined a trend line that Morcos says passed near RCI's Trinity Large, a hyper-sparse MoE trained with about 50x more compute. The multilingual mix was only 8% of tokens, with most individual non-English languages receiving at most six billion tokens. | Implication: Use deliberately designed pilot runs and scaling experiments as explicit investment gates before committing to a hero run, especially where compute capacity is scarce. | Caveat: Extrapolating from small dense models to a much larger sparse MoE can fail when architecture, optimization, data coverage, or downstream objectives change materially.
  • Claim: Domain mid-training can improve specialized capability without catastrophic forgetting when most of the mixture preserves the original pre-training distribution, and it can make downstream post-training much more effective. | Evidence: With Thomson Reuters, continued pre-training on 100 billion tokens reportedly increased LegalBench by about five points despite being less than 1% of the original pre-training budget, while general performance also rose. Applying the same post-training harness afterward reportedly produced nearly 3x the gain versus post-training a default instruction-tuned model. | Implication: Do not treat pre-training, mid-training, and post-training as separate handoffs. Design the data and evaluation plan jointly, with domain mid-training positioned to improve the starting policy for later fine-tuning or reinforcement stages. | Caveat: The preservation of general capability depends on maintaining an appropriate general-data majority and on the target domain; overspecializing the mixture can still degrade broad performance.
  • Claim: Synthetic data is most useful when it rephrases high-quality source material into diverse formats, rather than when it indiscriminately expands a random corpus. | Evidence: Morcos's example transforms a corporate-takeover document into formats such as true/false questions. He argues this can expand a vetted document into hundreds of templates, increasing diversity while keeping source information grounded; he explicitly warns that rephrasing random documents does not yield good results. | Implication: For proprietary knowledge bases, build synthetic-data pipelines around selected canonical documents plus automated and human validation, not bulk prompt-based generation over unfiltered repositories. | Caveat: Grounding in a source document reduces the need for the generator to supply knowledge, but transformation quality still has to be checked for faithfulness, ambiguity, and task relevance.
  • Claim: A competitive custom or general-purpose model need not require hundreds of millions of dollars if the data pipeline is strong, though this is a claim based on a specific customer case. | Evidence: Morcos says RCI trained Trinity Large on 17 trillion curated public-data tokens, without proprietary data or closed-model usage, and reached open-frontier competitiveness for under $20 million total across compute, salaries, R&D, repetitions, and several models. He argues narrower-domain models can be viable in the high-six-figure to low-million-dollar range. | Implication: The relevant strategic question is not whether Ken can match a frontier lab broadly, but whether a bounded domain, proprietary evaluation advantage, and disciplined data program justify owning a specialized model layer. | Caveat: The total-cost claim lacks a detailed cost breakdown, hardware access assumptions, model specifications, and independently verifiable benchmark results; it should not be used as a generic budget guarantee.

Detailed Brief

The four-C data pipeline

  • Claims: Cleaning alone is necessary but insufficient: it makes data ingestible, not necessarily high-value for learning.; Composition includes both selecting the final blend of sources and sequencing those blends across training stages; Morcos describes multi-phase training as standard and suggests continuous curricula as a further optimization.; Repeated exposure to high-quality data can be preferable to adding low-quality data, up to an unspecified threshold.
  • Evidence: Cleaning examples include heuristic filters for malformed or near-empty documents and aggressive benchmark decontamination using a low n-gram threshold.; Curation tools include classifiers, topic taxonomies, semantic—not merely exact-match—redundancy reduction, and targeted upsampling/downsampling.; The company metaphor is an oil refinery: it refines public, proprietary, and licensed tokens rather than primarily sourcing net-new tokens.
  • Caveats: Aggressive benchmark decontamination supports more credible benchmark interpretation, but benchmark cleanliness alone does not prove real-world task performance.; The transcript gives no operational thresholds for deduplication, reweighting, repetition, or curriculum transitions.
  • Implications: Data infrastructure should preserve provenance and allow corpus-level decisions to be revised as target-task evaluations change.; Benchmark-contamination controls belong in the data pipeline before model-comparison claims are trusted.

Multilingual curation and transfer effects

  • Claims: Curation can improve non-Western-language performance even when multilingual tokens are a minority of total training data.; Improving English data reportedly helps non-English performance through cross-lingual transfer, with a stronger effect for languages more similar to English.; The transfer works in both directions, although Morcos says non-English-to-English gains are smaller.
  • Evidence: The cited multilingual experiment used an 8% multilingual share; most languages had no more than six billion tokens.; Morcos compares curated models favorably with a frontier defined by Qwen and Liquid models and references Cohere's tiny AYA as a multilingual baseline.
  • Caveats: Language similarity to English affects transfer magnitude, so English-centric curation is not a substitute for adequate direct coverage of more distant languages.; The talk does not identify individual languages, datasets, or performance levels needed to assess equity or low-resource-language coverage.
  • Implications: Multilingual quality plans should combine shared high-quality general data with direct language-specific curation rather than relying solely on translation or English improvements.

Notable Concepts & Terms

  • Data quality as a compute multiplier: The central idea that higher-signal examples can yield the performance of a substantially larger compute budget by improving the learning curve.
  • Marginal information gain per data point: Morcos's technical framing for selecting examples based on how much they teach the model, rather than simply maximizing corpus size.
  • Clean, curate, create, compose: Datology's four-stage framework: remove unusable data, select and balance valuable data, synthetically re-express it, then combine and sequence it across training.
  • Benchmark decontamination: Removing training examples that overlap downstream benchmark material so reported results are less likely to be inflated by memorization.
  • Task-distribution matching: Aligning the training-data mixture with the type and frequency of tasks the deployed model is expected to solve.
  • Rephrasing: A synthetic-data approach that converts a high-quality source document into many formats while grounding content in that document rather than asking a generator to invent knowledge.
  • Continuous curriculum: Changing the data mixture or ordering dynamically through training instead of using fixed, discrete training phases.
  • Catastrophic forgetting: Loss of a model's general capabilities during domain adaptation; Morcos argues it can be mitigated through a mixture retaining mostly general-distribution data.

Operator Notes / Why Ken Should Care

  • Require every model-training or adaptation proposal to include a task-specific data map: source provenance, quality criteria, redundancy policy, language/domain coverage, contamination controls, and intended mixture weights.
  • Add a pre-hero-run stage-gate: run small curated-data scaling experiments, measure task quality and quality-adjusted inference cost, then use the results to decide whether larger compute spend is warranted.
  • For proprietary corpora, prioritize a grounded synthetic-data pilot using a small set of vetted canonical documents; audit generated examples for source faithfulness, ambiguity, and coverage before scaling.
  • If pursuing domain adaptation, test a general-plus-domain mid-training mixture before changing the post-training stack; compare post-training lift from the adapted base against the current instruction-tuned base.
  • Track token length and FLOPs per successful task in agent evaluations, since a data intervention that preserves accuracy while reducing answer length may relieve inference-capacity pressure.
  • Treat Datology's cost and benchmark claims as diligence leads: request exact model architectures, public evaluation protocols, compute accounting, corpus composition, and independent reproducibility evidence before using them in an investment or build-vs-buy decision.

Source/Metadata

  • Title: Data Quality Is the Compute Multiplier — Ari Morcos, DatologyAI
  • Transcript words: 4995
  • Duration seconds: 1145
  • Timestamp note: No timestamps or chapter markers were present in the supplied transcript; the transcript also contains repeated closing sections.

Transcript

3818 words en Processed in 131.9s

Good morning, everybody. My name's Ari Marcos. I'm the CEO and co-founder of Datology AI, and really excited to kick off the data quality track today. Data quality is what we live and breathe at Datology. It's all we think about. In fact, the company's name literally means the science or study of data. So, very excited to see the increasing excitement and interest in this area and an amazing lineup of talks today. So today, I'm going to tell you about why data quality is a compute multiplier that we're all overlooking and where we can make massive gains by just working on better data. We've seen that compute availability over the last six months has become extremely scarce and is only getting worse. We saw H100 prices reverse their several-year-long drop, which is normal for hardware, and all of a sudden come up, where now they're about 40% up from their lows at the end of last year. As test-time compute has become a critical part of models, and as we have put more and more thinking tokens in, we're seeing that token usage is absolutely skyrocketing. Reasoning models use eight times as many tokens as non-reasoning models, and that's projected to 5x again in the next year or so. So the number of tokens we're pushing through goes higher and higher, and that constrains compute even further. And this has led to a world where it's not implausible today that we might see access to some of the frontier APIs actually get limited or go away. As one example, Google just capped Meta's Gemini usage because of inference constraints. OpenAI has effectively started selling token futures, where you can guarantee token capacity some amount of time into the future. This is only necessary because they are legitimately wondering if there might be a world where access to frontier API tokens is limited, not as a business decision, but because there's just simply not enough inference and first-party products will be prioritized. So in a world where compute is increasingly scarce and you need to make models better, what do you do? Well, we work on data. You're going to hear me say this over and over again. Data quality is a compute multiplier because what it does is it makes the learning curve steeper. So this is just a very simple schematic of performance on the y-axis as a function of data on the x-axis. Note that this x-axis, you could swap it out for data, for compute, for time, for dollars. They're all the same x-axis fundamentally. And if you can make data quality better, you can turn this gray curve into this blue curve. And that means that now you can get dramatically better performance for the same compute budget as if you had trained with far more compute. And similarly, you can get the same performance for a much smaller compute budget, which is exactly showing you how you can get performance as if you had spent 10 times as much on compute and really shows this compute multiplier point. So how do you actually do this? Fundamentally, the idea is we want to make it so that we get the maximum signal per token and per batch. For those that are a little bit more technical, what we want to do here is maximize the marginal information gain per data point when we show it to the model. What data is going to teach the model the most? And that's all about finding data that's relevant to the use cases that you want. One thing that is very important is that there's no one golden data set to rule them all that's good for everything, no matter what you want to do. A data set's only going to be optimal with respect to a particular set of output tasks that you want the model to do. So if you want a great legal model, you're going to want legal data more than healthcare data, and vice versa. It needs to be diverse. A lot of the issues we see with model robustness and brittleness come from training on data that's not diverse enough. So the model can answer a question correctly if it's presented just so, but if it's presented a little bit differently, now everything breaks. It needs to be information-dense, and you have to mix the data correctly. This is a hugely difficult part. You now have many different sources. How do you combine them to actually drive the largest improvement in performance? So this is a high level of what we do at Datology here. You can think of us as the oil refinery for data. We don't source new tokens like many data providers. Rather, we take existing tokens coming from public data sets, proprietary data sets, and licensed data sets and make them way better. And how do we do that? We do that through these four Cs: clean, curate, create, and compose. So cleaning is fairly straightforward. This is doing things like heuristic filters, all in Gopher and things like that, removing documents that have only 10 characters in them or are all wingdings. That's basic table stakes. Benchmark decontamination is incredibly important. As I'm sure you all know, bench-maxing has become a real problem and makes it very difficult to interpret model results. So we rigorously decontaminate all of our training data with respect to all downstream benchmarks with a pretty low n-gram threshold to ensure that that's the case. That gets you to a point where now you can feed the data into the model, but it's still not very good. So then how do you make it better? It's a combination of many things, ranging from quality classifiers and taxonomy across different topics and balancing that, redundancy reduction, so removing data points that are not the same, that are semantically similar, but convey very similar information, even if they're not the same pixels themselves, say. Upsampling and downsampling data points based on the quality and the relevance, and then task distribution matching, identifying what data do you actually need in order to solve this given task. That now gives you a data set that is very high quality, but is typically still too small. And that's where synthetic data comes in. Now we can go and rephrase that data as effectively a very fancy form of data augmentation to produce dramatically more data in many different formats. And this helps a lot, both with data size and with diversity, because we can really inject a lot of diversity into this. And then finally, how do you combine these data sets and how do you sequence them across different training stages? It's now become table stakes that any large model is generally trained for at least three phases of data. How do you do that? And can you actually even do continuous curricula and things like that, which is a lot of what we work on at Datology? And that ultimately gets you a much better data set out. All right, so that's a high level of what we need to do. What can you actually get out of this? Can this actually really make a massive difference? About half of our team at Datology are just researchers, and we do all of our own research on how we do data curation effectively. Because this is such a critical part of the model-building pipeline, there's very little published here because there's a very strong disincentive to share how you do this. The foundational paper for Datology is the one I wrote when I was at Meta called Beyond Neural Scaling Laws, which was fortunate to get a best paper at NeurIPS a couple years ago, which showed that if you choose your data correctly, you can actually bend the scaling laws itself. You can change the exponent. And that's because you're now not wasting your time looking at redundant or unnecessary data. That was very much the proof of principle for all of Datology. And we've since expanded this into many public research releases we've shared of various ways to improve models just through data curation. I'm going to go through a couple of those results now and show you what we've been able to achieve. So first, let's talk about vision-language models. How can we improve VLMs just through data curation alone? So in this case, what we did is we took the Mammoth dataset. This is a fairly small dataset, about 25 billion tokens, that we use for the purposes of training the fusion adapter layer between your text model and your vision model. And what you can see, this is a scaling plot where we have error on the y-axis as a function of log FLOPs on the x-axis. So there's about a scale of a thousand from the leftmost part of this plot to the rightmost part. And what you can see is that if you look at the Pareto frontier defined by many of the best public VLMs like the Qwen 3 series and 3.5, InternVL, etc., you can see that models trained on Datology's data are able to go well beyond that frontier. And I'll note, this is actually without any post-training as well. So you can get very strong performance across many different benchmarks. So just looking at taking the input dataset, that gray diamond there, that's the input dataset we use to do our curation. You can see that just through curation, you're able to get around a 14 absolute percentage point improvement, holding everything else constant just through better data alone. And not only that, you can also see that we can roughly match the performance of Qwen 3.5 4b, of a thousand from the leftmost part of this plot to the rightmost part. And what you can see is that if you look at the Pareto frontier defined by many of the best public VLMs like the Quen 3 series and 3.5, intern VL, etc., you can see that a model trained on Datology's data are able to go well beyond that frontier. And I'll note, this is actually without any post-training as well. So you can get very strong performance across many different benchmarks. So just looking at taking the input dataset, that gray diamond there, that's the input dataset we use to do our curation. You can see that just through curation, you're able to get around a 14 absolute percentage point improvement, holding everything else constant just through better data alone. And not only that, you can also see that we can roughly match the performance of Quen 3.5, 4b, come with about a percentage of it while using 145x less training compute. In a world with less compute, how do you do more? You make data better, and now it's as if you had 100 times the compute. Interestingly, I mentioned that reasoning models are using tokens at a very high rate as well. Well, another thing that we found is that data curation can also lead to more concise answers depending on how you represent the data. So what's plotted here is the mean number of tokens per response across all the same set of models for the largest part that we just showed. And you can see that models trained on datology, those three blue lines right at the top, are all extremely concise. And if we do the same sort of plot, but now on the x-axis, instead of log training flops, this is now log flops per response. So this is inference efficiency. You can still see that we go well by beating that out of frontier and roughly get similar performance to Quen 3.5 with 35 times fewer flops per correct answer. So data curation can make a huge impact in VLMs. What about text models? One of the most challenging things about many models is that they work very well on English data, but they don't work well for non-Western use cases in general. The internet is an extremely biased view of the world that does not represent the world uniformly at all, and this has major implications for fairness and for the usability of these models across the world. I don't want to live in a future where only developed countries can access this very effectively. So how do you do this? Well, curation again can be a massive lever here. So what I'm plotting here now is a similar plot, error on the y-axis as a function of log flops, about 100x going from left to right here. This is highlighting multilingual MMLU performance. You can see we have a pareto frontier here defined by many models, the Quen models, some of the liquid models. The green square is tiny AYA, coheres best multilingual model. You can see again that we're well off the pareto frontier with a couple things I really want to highlight. First off, we only use 8% of the data here as multilingual tokens. So most languages actually only had at max 6 billion tokens here. So these are not massive amounts of data in the non-English languages that are going in here. You can again see we get the same sort of compute multiplier effect. We're a little better than Quen 3 while having roughly 8x less compute budget here. So you can make a huge improvement. One last thing I want to show here is that if you look at the two blue points on the upper left here, those are both dense llama style models trained for a trillion tokens on curated data. The point on the lower right here I'll come back to, but is a model trained by one of our customers, RCAI, Trinity Large, that was trained on 17 trillion tokens and is a hyper sparse MOE. And what you can see is that if you take the line defined by the two smaller models, it goes mostly right through that blue star, which is Trinity Large, despite it being trained with 50x more training compute. So if you use your data correctly and you simulate token scarcity appropriately, you can also get very predictable scaling to much larger models, and you can de-risk a run with 50 or 100 times less compute effectively before you actually go and scale up the hero run and find that maybe it doesn't end up where you want it to be. Another interesting scientific result I want to share here is that we also see very strong cross-lingual benefits from curation. So what's plotted here is the non-English accuracy, wherein the left bar is showing not curating anything at all, and then the right bar just curating the English. We also curate all the non-English data, and that leads to much better performance, but I just wanted to show this because I think it's quite interesting that curating English data benefits non-English performance. And that's because we see this cross-lingual transfer, where the model understands how English relates to, say, Spanish, and so therefore making it better at English would also make it better at Spanish to some extent. And interestingly, we see the magnitude of that transfer is strongly correlated with the similarity between English and that language. And we also see it go the other way, although the effect's a little bit smaller, where curating the non-English data also helps to benefit English data performance. Okay, let me talk a little bit about synthetic data. We take an approach to synthetic data that we call rephrasing. This is something that our team pioneered several years ago and has now become table stakes for building a very strong model. In any way, I think we'll hear a lot about synthetic data in various forms throughout the day. But fundamentally, with Beyond Web, our goal is: how can we define a synthetic data platform that works extremely well and can be applied to anyone's proprietary data and documents? Fundamentally, we want to help folks build models that wouldn't be able to do so otherwise, and that's where data quality can make an absolute difference. So to give an example, we might take a document like this about a corporate takeover, and we might convert that into one of hundreds of templates, one of which might be a series of true-false questions. By doing this, there's a couple things that are really great. Number one, because all the information is coming from the document on the left, you don't have any issue with model collapse, and you can actually train models that are much better than the rephrasing model because the rephrasing model doesn't actually have to teach and understand all the concepts. All it needs to do is transform the left document into true-false questions accurately, which is a much easier task. And then you do this into many, many different formats throughout data. This effectively increases diversity, and it makes it so you learn a lot more from the highest quality data points. One thing that's really critical here: what do you rephrase? All documents are not created equal for rephrasing. If you just pick random sets of documents to rephrase, you will not get a great result. But if you find the high quality documents and rephrase them, it can make a big difference. And we've seen if you compare this to lots of other public synthetic corpora, we can get much better performance much faster. And critically, this can be applied to any proprietary data in one of our customers' own environments. All right, let me spend the last few minutes just quickly talking about a couple of what we've seen, the things we've seen with our customers where this can actually drive early gains. So first, one of our customers, Thompson Reuters, has really focused on post-training quite a bit. And they have a very sophisticated post-training infrastructure with the goal of building better legal models on their proprietary high quality legal data that they have. So we partnered with them to mid-train a model first on a combination of their data and public data to then make much better legal reasoning models. So what do we see? Well, first off, look at the left here. In this case, we took an open source model and then just did continued pre-training or mid-training on 100 billion tokens. What you'll see here is that we see that legal capabilities go up at about five percentage points, as measured by LegalBench, after you do this 100 billion mid-training, which was less than 1% of the pre-training budget. But you don't get catastrophic forgetting. You also see the general capabilities go up as well. I'm sure that many of you have seen or experienced when you try to adapt the model to a particular domain, you lose general performance. Not if you use the data correctly. The key here is actually the majority of the data we showed the model was actually data that was representative of the pre-training distribution. There was only some of the domain specific data, and that's necessary to prevent the model from losing the capabilities that it had before. And you can solve this entirely through better data. But this actually isn't the most exciting part here. As I mentioned, the TR team had done a lot of work on post-training. It had a very sophisticated post-training harness. Well, they then applied that to the mid-trained model versus just the default instruction tuned model. And what they found was that the gain deriving from post-training, so the y-axis here is a delta as a result of post-training, almost tripled when you applied it to the mid-trained model versus to the default instruction tuned model. And that's because its policy, when it starts, is now much more accurate general performance. Not if you use the data correctly. The key here is actually the majority of the data we showed the model was actually data that was representative of the pre-training distribution. There was only some of the domain-specific data, and that's necessary to prevent the model from losing the capabilities that it had before. And you can solve this entirely through better data. But this actually isn't the most exciting part here. As I mentioned, the TR team had done a lot of work on post-training. It had a very sophisticated post-training harness. Well, they then applied that to the mid-trained model versus just the default instruction-tuned model. And what they found was that the gain deriving from post-training, so the y-axis here is a delta as a result of post-training, almost tripled when you applied it to the mid-trained model versus to the default instruction-tuned model. And that's because its policy, when it starts, is now much more accurate, and it can make much better inference. So even if you don't change the post-training data at all, showing your model better domain-specific data can actually make post-training two to three times more effective out of the box. Which I think really goes to show not only how important data can be in these factors, but it also actually goes to show how we really should be thinking about all these stages synergistically rather than as three completely independent stages of pre-training, and then I hand it off to somebody else who mid-trains, and then I hand it off to somebody else who post-trains. Okay. In the last minute or so, I just want to quickly talk about one other one, which is RCI, who I mentioned, that large model, Trinity Large. You'll actually hear from Varun, who is the pre-training lead for this model later today. So look forward to that talk. But in this case, they trained a model. This is a fully open, US-made model on 17 trillion tokens that we curated from public data sets. No proprietary data involved here and no closed model usage. So no asking Claude to do this for you. And with that, RC was able to train a model that is competitive with the open frontier, matches GLM-5 and Kimmy on many tasks, and even outperforms Claude on a couple tasks. But I think what's most exciting about this is that the RC team had not trained a model prior to the middle of last year when they started working with us. And critically, in total, across salaries, across compute, across R&D, across everything for this and several other models, they were able to get to a model that's competitive with the open frontier for less than $20 million total. That includes all the repetitions, that includes compute, that includes everything. So if you hear this story over and over again, oh, if I want to customize a model, it's going to cost hundreds of millions of dollars. That's just not true. You can train an immensely powerful model, especially in a narrow domain, for high six figures, million dollars. It's very doable to get a model that's extremely performant. This is for a general purpose, so this is the upper bound of that. And data quality is how you can do that. All right, so last slide here, just summarizing. Focus on what's going to give you the most signal per token. That's the thing that matters a lot more than more tokens. It is almost always better to repeat high-quality data than it is to show low-quality data at a certain point, up to a threshold. But focus on how can you get that. Data quality remains the single most under-leveraged compute multiplier. If you're sitting in a world where you want to build a model or customize a model and you're limited on compute, how do you get past that? Invest in data. And that's something that can do a tremendous amount of effort. And then finally, this is a frontier research engineering problem. You need to be able to score and understand data across many different axes. That's a frontier research problem. And then have that scale up to petabytes of data, massive scale, and can be very critical, and it can lead to tremendous leverage. And with that, I'll say thank you. I will note that we're hiring for a bunch of different roles on the left here. If you're interested in building or customizing your own model and would like to get much better data out of the box or apply it to your own data, we'd love to chat with you. And thank you very much. And thank you very much. Thank you. was actually data that was representative of the pre-training distribution. There was only some of the domain specific data and that's necessary to prevent the model from losing the capabilities that it had before. And you can solve this entirely through better data. But this actually isn't the most exciting part here. As I mentioned, the TR team had done a lot of work on post-training. It had a very sophisticated post-training harness. Well, they then applied that to the mid-trained model versus just the default instruction tuned model. And what they found was that the gain deriving from post-training, so the y-axis here is a delta as a result of post-training, almost tripled when you applied it to the mid-trained model versus to the just default instruction tuned model. And that's because its policy, when it starts, is now much more accurate and it can make much better inference. So even if you don't change the post-training data at all, showing your model better domain specific data can actually make post-training two to three times more effective out of the box. Which I think really goes to show not only how important data can be in these factors, but it also actually goes to show how we really should be thinking about all these stages synergistically rather than as three completely independent stages of pre-training and then I hand it off to somebody else who mid-trains and then I hand it off to somebody else who post-trains. Okay. In the last minute or so, I just want to quickly talk about one other one, which is RCI, who I mentioned that large model, Trinity Large. You'll actually hear from Varun, who is the pre-training lead for this model later today. So look forward to that talk. But in this case, they trained a model, this is a fully open, US-made model on 17 trillion tokens that we curated from public data sets. No proprietary data involved here and no closed model usage. So no asking Claude to do this for you. And with that, RC was able to train a model that is competitive with the open frontier, matches GLM-5 and Kimmy on many tasks and even outperforms Claude on a couple tasks. But I think what's most exciting about this is that the RC team had not trained a model prior to the middle of last year when they started working with us. And critically, in total, across salaries, across compute, across R&D, across everything for this and several other models, they were able to get to a model that's competitive with the open frontier for less than $20 million total. That includes all the repetitions, that includes compute, that includes everything. So if you hear this story over and over again, oh, if I want to customize a model, it's going to cost hundreds of millions of dollars. That's just not true. You can train an immensely powerful model, especially in a narrow domain, for high six figures, million dollars. It's very doable to get a model that's extremely performant. This is for a general purpose, so this is kind of the upper bound of that. And data quality is how you can do that. All right, so last slide here, just kind of summarizing. Focus on kind of what's going to give you the most signal per token. That's the thing that matters a lot more than more tokens. It is almost always better to repeat high-quality data than it is to show low-quality data at a certain point up to a threshold. But focus on kind of how can you get that. Data quality remains the single most under-leveraged compute multiplier. If you're sitting in a world where you want to build a model or customize a model and you're limited on compute, how do you get past that? Invest in data. And that's something that can do a tremendous amount of effort. And then finally, this is a frontier research engineering problem. You need to be able to score and understand data across many different axes. That's a frontier research problem. And then have that scale up to petabytes of data, massive scale, and can be very critical and it can lead to tremendous leverage. And with that, I'll say thank you. I will note that we're hiring for a bunch of different roles on the left here. If you're interested in building or customizing your own model and would like to get much better data out of the box or apply it to your own data, we'd love to chat with you. And thank you very much. and thank you very much. Thank you.