.
All right. Hi, everybody. My name's Jack. I'm here to talk about scaling compute on context, and also our startup Ngram, which launched last week. This isn't going to be a super detail-oriented talk where I go through a lot of experiments we've been running or talk too much about what our models do. I just want to frame the high-level problem of what we call scaling compute on context. People have many names for this. It's maybe a sub-problem of continual learning, or maybe even just the answer that we see to the problem. A little bit about myself at first. I'm Jack. I'm a researcher. I'm part of the startup Ngram. You can see me on the left in this picture.
We just launched last week with this picture on Twitter, so if you want to look us up and see more about what we do afterwards, please feel free. I'm also happy to talk after the talk concludes. So I'm going to start with this question about breadth versus depth. This is Terence Tao. He's likely the world's most famous mathematician. And he's also a really heavy user of AI, even an advocate for AI in math. And one of the things that you'll hear him talk about is how AI knows every single public mathematical topic.
And it can make connections between things that you wouldn't expect and help bridge gaps in the literature in a way that no human can know because it's read so much. But it maybe lacks the depth that you would look for from, for example, a graduate student who spent five years practicing in one area, who gets this almost subconscious intuition for the problem space. And so I think about what we're doing, which is training your data into models at a high level, as learning facts and skills like traditional AI. But at a low level, it's about depth. And the thing that we're after, this idea of scaling compute on context, is the pursuit of depth in AI.
I can give you some examples of why I think our approach is important. It's related to continual learning in general, as other people propose it. So one thing that models don't have is knowledge of what happens after they're trained. So even Fable whatever five that's coming out today was probably pre-trained with a cutoff, I would guess, at least a month ago, and has no idea if Mexico won their game last night. Actually, I don't know. I meant to check before the talk, but it's nothing to do with my pre-training cutoff because I'm a human. I'm capable of acquiring new information.
Another thing that models are bad at is hard, difficult, long-tail skill and knowledge acquisition that doesn't appear a lot in the training data. So models still are quite bad at writing AMD kernels. There are not that many good kernels written on AMD GPUs that are public, and they're intended to acquire this knowledge through pre-training, but they don't because it doesn't occur very often. And then I think maybe the most pressing case for you is, why can't I have a model like ChatGPT that knows all of my emails or knows the way that I like to write things or knows where my family and I like to vacation every year and all these little details of your life?
And there's a pretty basic reason for this, which is that ChatGPT and models like it are trained on public data. So they don't know things about your partnership with another company like SunTrust Bank. They don't know really anything about you unless you happen to be famous enough to appear in the pre-training data. So I feel this is not just an intellectual or academic problem. It's the core problem with the current paradigm in AI, that models cannot acquire new knowledge after training in a personalized way. So by definition, models have to be trained on data that's open to the public, and they can't learn the depth of the things that you know.
So just going back to Terence Tao again, how do we change this? How do we teach new things to models in a way that lets them acquire this kind of expertise or really deep skill set that we're looking for? And I'll say there's a lot of names for this. People call it sleep-time compute, continual learning, neural memory, write-time compute, note-taking, dreaming, studying, machine studying. In classical AI, maybe it's called amortized inference, and I think I'm calling it scaling compute on context. And it's almost all describing the same thing, which is something people really want.
But I think the reason why it doesn't have even a set, agreed-upon name is because the paradigm is very early and hasn't been solidified the way, for example, pre-training or post-training have. So maybe I'll take a second and talk about scaling. There's basically three axes that we use to scale AI models. We can train them on more data, we can train them for longer or add compute, or we can make the models themselves bigger, give them more capacity to acquire new information. And this is the main driver of progress from the last, really the entirety of the deep learning revolution comes from these three axes of scaling.
And the results are extremely compelling, but they're still limited to these public data sources. Models are really good at Wikipedia, they know everything about Reddit, papers on Archive, code on GitHub. And now they have this new layer of post-training data that's experts who are hired through data acquisition companies like Scale AI, Surge AI, and Mercore. But they're still by definition creating publicly available data because it's something that the model could tell to a user. So scaling is basically only used on public data, and yet it's so powerful.
So I think I'm trying to go a bit faster, so I'm not going to dwell on this, but this is the plot from Meter about how models get better every month and can complete tasks that take a longer time. And this is purely an artifact of scaling. I think the core question that I want to talk to you about today is how do we apply scale to your data? I think scale is clearly the thing that drives progress. It's not necessarily new algorithms or great new ideas.
I think maybe there's an element of data that's important, but really the thing that makes the new generation of models like Fable and GPT whatever that's coming out next month so good is that they basically scale along all three axes. I'm sure they have new data. They're certainly training for longer, and they make the models bigger. And this is how models keep getting better and will continue to get better. But I think the missing element is that this is always on public data. So models are getting better at coding in the way that is public on GitHub. They're getting better at doing math in ways that are written in public textbooks.
But they're not getting more knowledge of you or your life or your work. And that's what we're trying to change here. So if we approach the problem from first principles, I think the core limitation is that you have a fixed data budget. So say I want to scale in some fashion to train a model that knows the data from Ngram better, our company. We can't create new data. So the data-scaling axis is out the window. I think we also probably agree that we can't train a model from scratch on our data. So we probably want to start from a pre-trained model.
Or maybe another way to look at it is there's a ton of information about the outside world that is useful for understanding what happens within our company or your own context of choice. So you very likely want to start from a pre-trained model. This leaves us with essentially one axis of scaling, which is compute. And this brings us to the title of the talk today, which is Scaling Compute on Context. Just a small tangent while I have you is that I think one thing that's been beneficial for us to realize is that the amount of data isn't really fixed. There are a lot of ways that you can get more data afterwards.
Maybe in the pure math problem that I'll propose, you have this fixed data set and you want to train it into a model. But really if you're studying a textbook or trying to learn a new language, it's not really that you're limited to the words of the textbook itself. There's a lot of stuff you can do. You can find other textbooks. You can go on the internet and search for related things. You can even be proactive and talk to speakers of the language or people who know the thing that you're trying to learn. So in practice, I think the data access is very interesting and not actually fixed.
But from a core idealistic standpoint, the way we think about things is more or less, how do you scale more compute given the same data? And for the math heads in the room, I'm not going to write any equations.
But I think you can think of this as a box that you're dropped into. And all you have is this one pre-trained model. Maybe it's, I don't know, GLM 5.2. Maybe you somehow hacked into Anthropic and stole the weights of Claude. There's a lot of stuff you can do. You can find other textbooks. You can go on the internet and search for related things. You can even be proactive and talk to speakers of the language or people who know the thing that you're trying to learn. So in practice, I think the data access is very interesting and not actually fixed.
But from a core idealistic standpoint, the way we think about things is more or less, how do you scale more compute given the same data? And for the math heads in the room, I'm not going to write any equations. But I think you can think of this as a box that you're dropped into. And all you have is this one pre-trained model. Maybe it's, I don't know, GLM 5.2. Maybe you somehow hacked into Anthropic and stole the weights of Claude. And now you're trying to do it that way. But you have the pre-trained model. And then you have this unstructured data set D. So maybe this is all the emails you've ever written.
It's all the transcripts from every meeting your company's ever had. It's some very large unstructured corpus. And the question is, how do we create a better theta that knows D? And I'm going to walk through a few ideas that you could try or that people have tried and point to some links. And you can also ask me questions at the end. So the core question is something like how to produce a new model. I could call it data star that knows D. And I think the definition of know, this is a very load-bearing term in this question. And maybe that's where people get the leeway to propose new ideas. But this is essentially what you want to do.
This is what every continual learning startup is trying to do. This is more or less what we're doing at Ngram. So I'll start with a very simple idea, which is, okay, maybe you can just train the model on the data. You can use next token prediction and train it like an LLM. And I think you'll find unless you have a D that's so wide it can simulate the effect of pre-training, which no one has, then this doesn't work very well. I'll walk through an example real quick. Say we have this set of 10K financial reports. You want the model to know these. You want it to be in the weights. You want the model to answer questions about them.
You want the model to be able to create new ones. You want all these behaviors to be encoded into theta. And then you just train theta on the context that you have. You can get to a loss of 0.0001. And you can end up with a model that knows the data perfectly well. And then when you generate from it, it basically collapses. So this strategy, this naive idea of, oh, take the context that you have and train on it indefinitely, one, it's clearly bounded because there's some information that just gets perfectly transferred into the model and then you no longer learn. So this is not an indefinite axis of scaling. But two, it just frankly doesn't work.
Just doing this next token prediction on the data you have doesn't produce a model that has interesting generalization properties like normal models. It can't answer any question unless the question is perfectly encoded in the data with its answer, which is never the case in practice. So let's think about another idea since this is not quite as easy as we thought. What if we try to trick the model to think the data is in context? Because we know models are really good when you paste stuff into context. In-context learning is magical. One idea is you can do compaction, which is similar to the way that Cloud Code or Codex or Open Code, what have you, does compaction.
You take this really long context, which is D, and then you try to compress it into some set of KVs that can represent the data to the model in a very succinct way. And there's some interesting approaches to do this. You can do it in a learned way. This is a very cute paper that has a greedy algorithm for approximating KV compaction. So basically, if your data is small enough to fit into context, there are some interesting ways to compress it to something very small and pretend your model knows this.
I think there's multiple problems with this, the main one being it only applies to things that are in context, but it also misses, I think, some of the magic that you can get from taking gradients. So there's an alternate way of doing it, which is you can train the model to think the data is in context. And there's some interesting approaches here. I think Roanock was talking about on-policy distillation. This is a powerful tool for doing knowledge transfer where you have text and you show it to the model and then you make the model think that the text is in context. That's more or less the trick of on-policy distillation.
The on-policy part just means you update the model throughout training. It works. It's a pretty good algorithm. I think there are also some core problems with it, maybe the main one being what data do you actually do this with? You can't really distill the raw documents. So techniques like self-study from the cartridges paper on the left here try to generate question-and-answer pairs conditioned on D and then train the model to behave as if it is seeing D in context when it's answering questions. I think this is close to the behavior you want but also has some properties that are not necessarily appealing. And I'll get to that in a few slides.
I think there's one more idea that I think is interesting, or maybe two. I think a lot of the magic in deep learning or in LLMs, the reason why GPT-5 is so amazing, is basically because of pre-training. I think there's a lot of caveats to this statement, but pre-training is amazing for knowledge acquisition. I can ask Claude what I don't know result I got in a paper that I've written, and it actually knows this, which is incredible. And you can argue maybe they do one of these synthetic data tricks, but it more or less is knowledge that's acquired through pre-training. And so one way to teach data to a model is to simulate pre-training in some way.
And these are three pretty interesting approaches to do that, to craft synthetic data conditioned on D and then train data for longer on the synthetic data as if you're continuing pre-training. I think there are caveats to this approach, like you overwrite some of the pre-training. I think it's difficult to scale. But I think this is pretty promising. Maybe one blocker is you then have to post-train the model after doing this. So a lot of people don't actually start with good pre-trained base models. They have post-trained models, which makes this hard. But I do like this line of work, and these papers are interesting resources if you're interested in learning more.
I need to go faster. There's one more. Let's skip Andre. There's one more interesting idea, which is you can craft unsupervised reinforcement learning environments and do RL. It's pretty similar to the previous suggestion, except instead of doing some type of distillation you're just using RL loss like GRPO or whatever. I think all of these are promising but also missing maybe some core component. The thing that we're really after is to give the model more knowledge of D or to get better depth in your domain. We want to be able to add compute arbitrarily in a way that makes the model better.
So I think none of the approaches I propose do this, basically for classical machine learning reasons, which is that whatever you do, you have to define the data set and then you train on the data set and eventually things saturate. So even if it's really hard, unless your model is under-parameterized, eventually it will learn all the data. And this doesn't give the property, the beautiful scaling properties, that we see out of pre-training. It's kind of like a data wall in the synthetic sense where, when you create synthetic data from D and train on it, you eventually hit this upper bound where you've learned all of the synthetic data. And then you have to do it again.
And so I think a lot of the missing components here are how do you do it again? What's this second stage of training look like? So you can do almost any of the techniques I just mentioned. You could do the attention matching or some type of self-study thing or some continued pre-training. But eventually you will fit the data and you'll know some about D, but you won't know everything and you'll no longer have this property where you can add compute and give the model more depth. So I think a lot of the exciting work here comes from this idea of self-improvement.
It's a bit overworked as well, but I think this is actually the magic behind a lot of successful RL systems like AlphaGo, which is that AlphaGo makes its own training questions harder by getting better through training. And so I think one thing that everyone is looking for is a technique that can make models better, which makes them train themselves better.
Or this is maybe a long way of saying self-improvement. What's this second stage of training look like? You can do almost any of the techniques I just mentioned. You could do the attention matching, or some type of self-study thing, or some continued pre-training. But eventually you will fit the data, and you'll know some about D, but you won't know everything, and you'll no longer have this property where you can add compute and give the model more depth. So I think a lot of the exciting work here comes from this idea of self-improvement.
It's a bit overworked as well, but I think this is actually the magic behind a lot of successful RL systems like AlphaGo, is that AlphaGo makes its own training questions harder by getting better through training. And so I think one thing that everyone is looking for is a technique that can make models better, which makes them train themselves better. This is maybe a long way of saying self-improvement.
You generate data, and then the model gets a bit better, and then you generate better data recursively. And I think this is something that we're working on a lot at Ngram, is how do you make, when we started the company, we generated curves that look just like this blue curve. Where no matter how much data we generate or how much we train, we do plateau because there's this almost natural upper bound to how much you can learn in one go from D. But I think it turns out that there are more sophisticated things you can do that make the training gradually harder, that make the model better over time.