AI Engineer

The Messy Reality of Scale: Synthetic Data and Pre-Training — Marah Abdin & Robert McHardy, poolside

1946 summary words 9 min summary Watch video

Start with the signal

9 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Scaling a capable open-weight base model requires treating data diversity, synthetic-data generation, distributed-training correctness, numerical precision, and observability as one integrated recipe rather than independent optimization projects.
  • Why it matters: The talk offers reusable patterns for building synthetic-data control planes and for detecting silent failures that can waste large GPU runs or quietly degrade model quality.
  • Best use: Use it as a practical architecture and reliability reference for agent-data pipelines, model-training observability, and deciding where synthetic data should augment rather than replace organic corpora.

Executive Summary

Marah Abdin and Robert McHardy describe their transition from earlier Laguna models to newer open-weight releases, arguing that scale exposed failures in both their data mix and training stack. Their central lesson is that high-quality seed data alone does not scale: repeated exposure to a narrow high-quality corpus can shape a model too early, so larger training budgets require more token uniqueness, recall-oriented web sampling, and targeted synthetic augmentation.

Their synthetic-data strategy is explicitly complementary to organic data, not a substitute for it. Synthetic examples are used to surface latent structure—reasoning, planning, organization, and cross-domain representations—that exists implicitly or inconsistently in web data. Their XS.2 pre-training mix used 13% synthetic data, while their overall corpus has grown to six trillion tokens. They organize generation as modular pipelines with inputs, metadata, generators, validation, and filtering; complexity should rise only where the output is valuable enough to justify it.

McHardy’s half is chiefly a reliability argument: model quality depends as much on correct distributed implementation and numerical behavior as on data. Their team hashes model replicas to enforce invariants and aborts training when weights diverge. This caught silent data corruption from a faulty GPU; separately, a tensor-parallel BF16 accumulation near the language-model head lost sufficient precision as activations grew, causing training to stop converging until the operation was moved to FP32.

The new 118B-total-parameter, 8B-active-parameter Laguna S model is presented as a test of whether this combined recipe transfers to larger scale: 30T tokens on 4,000 GPUs. Early base-model results are strongest in the coding and agentic proxy evaluations the team prioritizes, while knowledge-heavy MMLU Pro results lag some competitors. The useful takeaway is not the benchmark claim alone, but the operational posture: specialize data and evaluations around the intended agent behavior, instrument every plausible invariant, and assume scale will reveal both obvious and silent failure modes.

Key Takeaways

  • Claim: Synthetic data should complement organic data by making implicit useful structure explicit, rather than attempting to replace the web corpus. | Evidence: The team uses synthetic generation to expose rationale, planning, and structure, fill coverage gaps, and regularize both token presentation and the teaching signal; synthetic data made up 13% of XS.2's pre-training mix. | Implication: For agent or foundation-model data systems, use synthetic data selectively to repair specific representational and coverage failures in organic data rather than treating generated volume as inherently valuable. | Caveat: The speaker explicitly says she does not view synthetic data as a replacement for organic data in the current state of the field.
  • Claim: At larger training budgets, maximizing nominal data quality without enough uniqueness can be harmful because repeated high-quality documents over-shape the model early. | Evidence: When scaling from Laguna M toward later models, the team hit non-optimal repetition in its higher-quality data; replacing repeated seed tokens with multi-mode rewrites produced consistently better ablation results than repeating seeds. | Implication: Track repetition and effective token diversity as first-class training variables; do not assume a small curated corpus remains optimal once the token budget increases. | Caveat: The presented results are ablations and the speaker asks listeners to take the exact numbers with a grain of salt.
  • Claim: Synthetic-data generation should be built as composable pipelines, with task complexity decomposed until the generator can execute each part reliably. | Evidence: The proposed six-part pipeline consists of seeds/primary inputs, metadata, secondary inputs, a generator function, and supplementary filters and validators. Examples include cheap seed-heavy rephrasing, staged novel generation, converting math problems to code, and iterative judge/evolver workflows. | Implication: Design agentic data workflows as small inspectable stages with explicit validators and escalation paths, instead of asking a single model call to produce complex artifacts end to end. | Caveat: Cheap generation works best when the seed carries substantial structure; harder tasks require more orchestration because a teacher model otherwise falls into bias, loses correctness, and loses diversity.
  • Claim: A configurable multi-agent generation control plane enables both scalable synthetic-data production and stronger quality control. | Evidence: Their Hive system represents generation as a queue of agents with configurable prompts, parameters, models, inputs, outputs, entry/exit conditions, and help frequencies; orchestrators can change downstream instructions, choose or skip agents, and a supervisor oversees the orchestration layer. | Implication: For Ken's agent systems, separate worker execution from orchestration and supervisory policy so that routing, dynamic instructions, retries, and quality policing can evolve without rewriting each generation workflow. | Caveat: The talk explains the architecture conceptually but does not provide implementation details, comparative costs, or measured quality gains for Hive.
  • Claim: Distributed training requires active invariant checking because hardware faults can silently corrupt a run without any difference in code, configuration, or data. | Evidence: The team periodically hashes weights across data-parallel replicas, which should remain identical, and crashes training on divergence. One run with markedly spikier loss and much larger gradient norms was traced to a broken GPU causing silent data corruption despite otherwise identical conditions. | Implication: Treat correctness monitors as part of the training product: encode expected invariants, continuously test them, and fail fast rather than interpreting anomalous loss curves as ordinary optimization noise. | Caveat: Replica hashing only works where redundant replicas process equivalent state; it does not cover every source of computation corruption in a real training run.
  • Claim: Numerical precision decisions that appear safe at smaller scale can halt learning as activation magnitudes change during training. | Evidence: Around 50,000 steps in Laguna M.1, activations before the unembedding grew until a tensor-parallel accumulation in BF16 no longer had adequate precision. Loss flattened and gradient norms trended upward; resuming from the checkpoint with FP32 accumulation restored convergence and reduced gradient norms. | Implication: Audit precision-sensitive reductions and accumulations independently of the broader mixed-precision policy, and monitor activation scales and gradient-norm trends throughout long runs. | Caveat: This was a specific tensor-parallel accumulation failure, not evidence that BF16 is generally unsuitable for model training.
  • Claim: The team's scaled recipe appears to improve coding and agentic performance, but its specialization creates measurable tradeoffs against broad knowledge benchmarks. | Evidence: Laguna S is a 118B-total, 8B-active model trained on 30T tokens across 4,000 GPUs. In early base-model evaluations it exceeds XS.2, Laguna M.1, GLM 4.5 Air, NemoTron 3 Super, and DeepSeek V4 FlashMax on several coding evaluations, and leads the compared models on SWE-bench Agentless Multilingual; it trails NemoTron and DeepSeek on MMLU Pro. | Implication: Evaluate model and data investments against the target operating workload—such as coding agents—not generic leaderboard position, while making the resulting capability tradeoffs explicit. | Caveat: These are pre-post-training base-model evaluations, so they are only partly predictive of final released-model performance; the team intentionally deprioritized MMLU Pro relative to coding and agentic use.

Detailed Brief

Synthetic-data portfolio: match pipeline cost to information structure

  • Claims: The team spans a continuum from low-cost, highly scalable synthetic generation to expensive, orchestrated workflows.; Rephrasing was expanded beyond generic multi-mode rewriting into specialized raw-code-to-code-plus-text and STEM-focused pipelines.; Cross-domain porting can increase representational coverage by translating a problem into a different modality, such as turning math problems into code.
  • Evidence: The speaker characterizes rephrasing as cheap and scalable because it can lean heavily on the original seed.; For complex educational or otherwise high-value outputs, the team uses multi-stage workflows rather than relying on a single teacher-model completion.; A staged novel-generation example begins with setting, characters, character styles, plot, and twists before generating individual chapters.
  • Caveats: The talk provides no quality thresholds, cost model, or acceptance-rate data for deciding exactly when a pipeline should move from cheap rewriting to multi-stage orchestration.
  • Implications: Create an explicit production taxonomy for synthetic tasks: seed-heavy transforms, structured multi-stage creation, cross-domain conversions, and iterative adversarial/refinement loops.; Reserve high-cost agent orchestration for datasets where increased correctness, pedagogical structure, or diversity has a clear downstream value.

Observability gaps remain even after replica checking

  • Claims: Adding FP8 training via DeepGem FP8 kernels introduced a race condition that produced illegal memory accesses, NaNs in gradients, and an additional partially silent corruption mode.; The team observed roughly 0.5% of gradients becoming corrupted and replaced with random values.; Standard replica hash checks have a blind spot because ordinary training does not provide redundant replicas executing identical forward and backward passes on identical data.
  • Evidence: The issue was traced to the FP8 kernels after debugging; the team says it has a fix in a public pull request.; They are developing a hash checker intended to validate forward and backward behavior at least in a dry-run setting.
  • Caveats: The transcript does not establish whether the kernel issue is broadly reproducible, whether the proposed fix has been merged, or what performance cost the proposed dry-run checking will impose.
  • Implications: Separate runtime-health monitoring from deterministic correctness testing: a healthy job without crashes or NaNs can still be wrong.; Before deploying new low-precision kernels or compiler paths at full scale, run controlled redundancy-based validation that can expose nondeterministic corruption.

Notable Concepts & Terms

  • Token uniqueness / non-optimal repetition: The data-scaling problem where repeatedly consuming a limited high-quality corpus causes premature over-shaping of model behavior; rewriting is used to replace repeated token exposure.
  • Multi-mode rephrasing: A scalable synthetic-data approach that generates varied rewrites of seed content to increase diversity while preserving the underlying information.
  • Cross-domain porting: Converting content between modalities or domains—such as math to code—to create new training representations and capabilities.
  • Hive: The team's configurable synthetic-data orchestration system, built around queued agents, orchestrators, and a supervisor layer.
  • Model replica hashes: Periodic hashes of data-parallel replica weights used to enforce the invariant that equivalent replicas remain identical and to detect silent corruption.
  • Tensor parallel unembedding accumulation: A precision-sensitive distributed reduction near the language-model output head; BF16 accumulation failed when activation scales increased and was corrected with FP32.
  • Laguna S: The unreleased 118B-total-parameter, 8B-active-parameter model used to test whether the team's data, architecture, numerics, and observability recipe holds at larger scale.
  • SWE-bench Agentless Multilingual: An evaluation the team uses during pre-training as a proxy for agentic performance, reflecting its stated focus on coding agents rather than broad knowledge benchmarks.

Operator Notes / Why Ken Should Care

  • Add configurable orchestration, validator, and supervisor layers to synthetic-data workflows rather than embedding routing and quality-control logic inside individual agents.
  • Define a data-diversity dashboard for any large-scale corpus: repeated-token exposure, source recurrence, rewrite share, acceptance rates, and downstream capability movement.
  • For new model-serving or training kernels, require a pre-production correctness gate that includes deterministic dry runs, anomaly thresholds for NaNs and gradients, and redundancy-based checks where feasible.
  • Audit every distributed reduction and accumulation for precision risk; monitor activation distributions and gradient norms over time rather than only loss.
  • When comparing coding-agent models, prioritize task-specific agent evaluations alongside broad benchmarks and document any intentional capability tradeoffs.

Source/Metadata

  • Title: The Messy Reality of Scale: Synthetic Data and Pre-Training — Marah Abdin & Robert McHardy, poolside
  • Transcript words: 3240
  • Duration seconds: 1051
  • Timestamp note: No timestamps or chapter markers were present in the supplied transcript.
Full transcript 3139 words · 14 min read
0:00

Hi, everybody.

0:12

Thanks for coming to our talk. My name is Mar Abdin. I'm from PULSA. I'm the synthetic lead for our data team. And today my colleague Robert and I will be talking a little bit about some of the challenges that we feel as we've scaled our models over here. Particularly if you haven't heard, we've switched recently from releasing our models toward enterprise to also releasing toward everybody. We've actually put out two open-weight models. They're on Hugging Face a few weeks ago. We have LagunaM and LagunaXS. We also put out a tech report which has a ton of detail if you're interested.

0:48

As you can see here by the LagunaM and XS.2, this actually is because we have switched out quite a few things between those two models. And so a big flavor of this talk is going to be about how to transition from one to two. And in fact, we've continued to do so. And now we actually have a newer version, and Robert will give a sneak peek about a soon-to-be-released model. Okay. So I will be particularly talking about the synthetic data part of things. So there's three things that we did on the data side to resolve some of the issues that we've seen with scale.

1:21

One is that we implemented an automixer that just gives us the chance to do a cheaper sweep on clusters of our sets before moving on to more expensive experiments. And then we improved, just rethought, our sampling of web data for higher recall. And then the third one is that we relied a lot more on synthetic data in a few forms, which I will go into. Okay. So before we go into what does that mean and what have we done, et cetera, why would we use synthetic data? And the thing is that at least at Pulseye we don't see it as a way to replace organic data. I don't see it so in the current state of the world at least, but it is a way to complement it.

2:06

And the thing is that organic data has a lot in it that is implicitly hidden. A lot of things that could teach the model are not presented in the most optimal way sometimes. And so synthetic data gives us a track to extract some of these features and then project them on some new planes. And this is how we get to expose implicit rationale, implicit planning, implicit structure, and a way for us to fill gaps and regularize not only how we present the tokens, but also how we are teaching the model. For XS.2 in particular, we settled on 13% of the mix. This is only pre-training stages before post-training.

2:47

And since then we've just been continuously generating more data in a bunch of directions. Now we have a six trillion token corpus that's continuously growing. Okay. So what I just said is that we saw some limitations switching from LogoNAM to 0.1 or 0.2 models. And so one of those things is that we started on data. This is really not a crazy problem. We very intuitively started from a place on a smaller scale where we were focusing on quality versus quantity, maybe a little too much, because eventually when we started scaling our models, we had to scale our training budget.

3:16

And with that came some limitations because we started hitting repetition, non-optimal repetition, on some of our higher quality data, which started shaping the model a little too early. So one of the ways that we've, particularly for this token uniqueness problem, relied on, which is a very common form of synthetic data, is rephrasing, which you've just heard in Beyond Web, for example. It's become pretty trendy these days. And you can see here that this is an ablation result, so take the numbers with a grain of salt.

3:38

But what presents for you consistently is the diff between using the orange, which would be just the seeds with repetition, and then the green, which would be replacing some of those repeated tokens with the multimode rewrites, or at least all of them are at least reducing the repetition. And so, for rephrasing in particular, we did the, what everyone's doing with the generic multimode, very scalable pipeline, but we also took it a step farther. And we did two other specialized pipelines, one to go from raw code to code and text and one to go specifically for STEM data, just because this is a very cheap, scalable pipeline.

4:07

So you have to rely very heavily on the seed, and we pushed a little further on that for the STEM documents. Okay. So if you think of everything in a modular way, you can think of every synthetic data pipeline as composed of the same six components. And so, you have your seeds, your primary inputs, your metadata, your secondary inputs, your generator function, which can be an agent with tools or an LLM with some prompt, with some prompt input, and then some supplementary functions like filters and validators and so on. And really, you can compose just about all pipelines, from very simple to very expensive pipelines, like this.

4:42

And on that note, we have covered quite a bit of a wide scope on the axis of complexity, and you can think of it as, at one end, you have the cheap, scalable pipelines that use smaller models and can get away with it because they're seed-heavy examples of phrasing. And then on the other end, you have more complex pipelines with a little more orchestration in the workflows. This is reserved when we're building on something that's worth it, educational data, but really this is how we're not blocked or limited by whatever the teacher model can do.

5:01

And this is how we can be ambitious in our synthetic data, because the rule of thumb is if a task is too hard for your model, then your model will start to fall in its bias, lose correctness, lose diversity. So break down the task, make it simpler. And I will give some examples of shapes rather than something more concrete about how did we use this modularity. And one shape is the form rewriting, it's just rephrasing. We already talked about this. Multi-stage pipelines and multi-stage workflows: basically, this is what I also just said. You take a step and break it down into multiple steps.

6:05

You can aggregate the process and slowly build up the generation. An example of this would be if you wanted to generate a novel, for example. You could generate one chapter at a time, but you could also take it a little slowly. First generate the setting, the character names, the character styles, the plot, some twists, and then from there go into generating chapters one by one. You will absolutely get a better novel. Okay. Third is cross-domain porting, which really is just moving from one mode to another. An example would be translating code. Another example would be something we did, which is take our math problems and convert them to code.

6:51

The last one is multi-term role. And so by that, all I mean is that instead of having a very singular or linear, or even non-linear, view of things, you have more of an iteration. This encapsulates pretty much everything. And an example of that would be multi-term chats when you have two agents talking to each other, or a task evolution pipeline where you have a judge and an evolver going back and forth for some key amount of time. And so on. Okay. So lastly, before going off to Robert, I do want to mention that because of the modularity of the way we think about this, we can implement infrastructure that's pretty configurable. So this is how we present Hive.

7:38

Hive is basically a way for us to easily construct generations where now you have a queue of agents that you define. Each one has its prompt, its parameters, its model, its inputs, outputs. But it also has, and you can configure, when it enters the queue, when it exits, how many help frequencies come in. And then we have orchestration in the middle between agents. And these orchestrators are really useful because they give you more flexibility in presenting a hierarchy between the LLMs that are generated. This is how you police them.

8:06

But also give them some form of creativity and dynamically change the instructions for the next agent or choose which agent goes next, which agent is skipped, and so on. And lastly, you have the supervisor, which basically is someone who polices the orchestrator and has more of a global view. Cool. Okay. That's it for me and Zendegade. I hope you learned something interesting. Handing over to Robert for pre-trainings. Yeah. All right. Thank you, Maura. Yeah. Because I will talk a little bit more on the actual pre-training side rather than just data.

8:58

I liked in the previous talk, the speaker made a point that we should treat different data mixes holistically, different training stages. I want to make the same point that we should treat data and implementation of your training code base, correctness of it, and so on, also holistically. If you've got data that sucks, you can't train a good model. If you've got a training code base that sucks, you also can't. So I specifically focus on architecture work and distributed training and so on. And the way we look at things in my team is we don't trust anything.

9:12

There are so many things that can go wrong when you scale models to billions of parameters, to hundreds of billions of parameters, training on thousands of GPUs and so on. And I want to show you some of the learnings that we got from training Laguna M.1 and some of the surprising things that happen at scale. So one thing we do is we've got these model replica hashtags. So essentially when we train a model, we've got multiple replicas of the same model, right? Distributed data parallel. And we know there's invariants. The weights should always be the same across all of these replicas. That's something you can verify, right?

9:51

You can calculate the hash over the weights and you know that should always be the same across all replicas. So we do that in training and periodically compare them. If all of these hashes are identical, then we know we can continue training. If they're not identical, we know something has gone seriously wrong because that should never happen. And we crash the training. And I will give you some examples now of things that we've not shared before publicly like this. So I hope they're interesting. So first example here of shit that happens at scale is broken GPUs.

10:19

On the left-hand side, we've got two loss curves and on the right-hand side, the corresponding gradient norms that we observed during training. And you can see that these loss curves look quite different, right? The purple one has got quite some bumps, looks a bit spiky. The gradient norms are huge for that run. And there's actually no difference in terms of model configuration, training data, training implementation between these runs. They're exactly the same run. Just in one of them, we got unlucky and we had a broken GPU included. That broken GPU caused silent data corruption and therefore made the training behave the way it did.

10:45

And that is one of those cases that you can catch with these hashtags because you know this computation should be the same across all replicas. But it wasn't. Which brings me to the next instance of shit that happens at scale. In this case, exploding gradients. Again, we're looking at two different loss curves and the corresponding gradient norm curves. The purple run is our initial training run for Laguna M.1. We're a bit further into training here, around 50,000 steps or so. And you can see it stops converging, right? It just flattens out. And the reason here was that during training, the activations grew and grew right before the LM head, the un-embedding.

11:25

And we have to perform some sort of accumulation here because we used tensor parallel for the un-embedding. And that accumulation was performed in BF16 by default. And because of the growing scale that we observed in the activations, there wasn't enough numerical precision available anymore to do this accurately. And hence the model just couldn't learn anymore. And this is also a very dramatic point for this to happen because from there on it really backpropagates into the full model trunk. The orange curve is essentially just adding a fix on that. So we took the checkpoint from the purple curve. We moved that accumulation into FP32.

11:52

And then from there on the model started converging again. The gradient norm, as you can see, actually started decreasing. Before then we had an increasing trend. And this is also something you can only observe at scale. And that will break your model if you're not careful about it. So as Mara said, we took all of these insights on data, on numerics, and so on, and we turned them from M.1 into Access.2. That's why we say it's a newer-generation model.

12:27

This included increasing diversity, reducing data repetitions, all of these numerical things I just mentioned, and just adding more observability and checks on that side, as well as generally optimizing the training and the architecture. And Access.2, if you look at it, it's open weights, right, so you can download and use it for free. It's one of the most competitive models for its size and for coding specifically. That's why we focus on agentic coding. So we're pretty happy with it. However, you can say that the model with 33 billion parameters is pretty small. You probably wouldn't observe any issues anyway training it.

12:51

So what was important to us was to scale this, right? And this is where Laguna S comes in. This model is not public yet, so this is a preview. As I said, we treated Access as a test bed in a sense, and with Laguna S, we scaled this to a model that's 118 billion total parameters and 8B active parameters. Again, we trained it on 30 trillion tokens on 4,000 GPUs. So the scale was sufficient to not only test whether all the improvements we made on the data and architecture side hold, but also if any of these numerical issues come up again. And of course, something happened. In this case, it doesn't actually have anything to do with scale, so it was just unfortunate.

13:24

In this case, we had a race condition because we added FP8 training based on DeepGem FP8 kernels that are also open source. We noticed these because we hit illegal memory accesses as well as NaNs in the gradients, which, after a while of debugging, we traced back to those kernels. There's also an unobservable effect that you wouldn't know about if you don't know that there's an issue. In our case, we noticed about 0.5% of the gradient gets suddenly corrupted, essentially replaced by random values. We do have a fix available that's in a PR right now. It's not been mentioned to DeepGem yet, but it's public on that QR code if anyone is interested.

13:47

And it's also an interesting point because there's a blind spot in the hashtags. In real training runs, you don't have any redundancy where you have the same model weights and the same data, so you can never check if forward and backward actually behave the same across different model replicas. So you can also never check if there's a race condition in that. That's something that we're working on right now, to essentially have a hash checker that can also do that, at least as a dry run. And I want to end on some early results from this new model and demonstrating how it performs against some open-weight models and also against our previous models.

14:21

So first I want to caveat this with these are base model evals, right? They are partly indicative of how the final model will look, but also not perfectly, right? There's still post-training happening. Not all of these will translate one-to-one to the final model. But if you look at them specifically on the coding part of the evals, so for instance multiple E, live code bench, big code bench, Laguna S is not only stronger than XS.2, which is our previous smaller model that performed very well, but also than a much larger M.1.

14:45

And it's also much better than GLM 4.5 Air, which is admittedly a bit older, NemoTron 3 Super, which is quite recent, and then DeepSig V4 FlashMax, which is quite recent and a fair bit larger. We can see it's competitive on Big Bench Hard, for instance. It doesn't achieve the top eval results compared to these models, but it's quite close. We also see it's quite close on eval plus, and quite importantly for us, it does very well on speed bench agentless multilingual, which we use to sort of proxy agentic performance during pre-training. And in that case, it performs much better than all the other models we tested here.

15:17

I also want to point out that, of course, it's not the strongest model in the world, right? For instance, MMLU Pro Knowledge Benchmark is something we don't care about that much compared to coding, because we want to build the strongest agentless model. So here, compared to NemoTron and DeepSig, we have to say that they perform much better. And this mainly comes down to data, right? It's a data gap that we could plug if we wanted to. But I think the point is all of the things we've found before were included in a recipe. The recipe held, it scaled, and we will continue scaling it from here. So this model will also be available sometime in the future relatively soon.

16:16

Again, open weights, so all of you can download it and use it for free. And with that, I want to thank everyone for attending our talk. I added also a link to our careers page and our Twitter page if you want to check it out. And yeah, thank you very much. Thank you very much. But I think the point is all of the things we've found before were included in a recipe. The recipe held, it scaled, and we will continue scaling it from here. So this model will also be available sometime in the future relatively soon. Again, open weights, so all of you can download it and use it for free. And with that, I want to thank everyone for attending our talk.

17:06

I added also a link to our careers page and our Twitter page if you want to check it out. And yeah, thank you very much. Thank you very much.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note