AI Engineer

Scaling Compute on Context — Jack Morris, Engram

1927 summary words 9 min summary Watch video

Start with the signal

9 min read

Summary

At-a-Glance

  • Verdict: Skim
  • Core thesis: AI’s next personalization and domain-expertise frontier is to spend scalable training compute on a fixed private context so a pretrained model develops durable depth, rather than merely retrieving or reading that context at inference time.
  • Why it matters: This frames a core limitation of current public-data-trained models: they can improve broadly through frontier scaling but do not naturally accumulate knowledge of a company, user, or rare technical domain.
  • Best use: Use this as a conceptual map of the emerging continual-learning/neural-memory space and its unresolved scaling problem, not as an implementation guide; the talk names promising approaches but withholds Engram’s technical method and results.

Executive Summary

Jack Morris argues that present LLM scaling creates extraordinary breadth but not personalized or domain-specific depth. A model may know public mathematics, GitHub code, and web knowledge, yet it does not know a company’s internal meetings, a user’s email history, current events after its cutoff, or scarce expertise such as AMD kernel programming. His proposed problem is to turn a pretrained model plus an unstructured private corpus D into a new model that genuinely “knows” D.

The central constraint is that private-context learning starts with a limited corpus and normally must retain the broad competence of a pretrained foundation model. That eliminates training from scratch and makes simply adding more raw private data difficult. Morris therefore calls the remaining principal lever “scaling compute on context”: repeatedly applying training compute to the same contextual material, while potentially using search, external sources, or interaction to expand the effective learning material.

He reviews several approaches: naïve continued next-token training, context/KV compaction, distillation through generated question-answer pairs, synthetic continued pretraining, and unsupervised RL environments. His assessment is that each can transfer some information but eventually saturates: the model fits a finite training set or synthetic derivative of it, may lose prior knowledge, and lacks the open-ended scaling behavior seen in large-scale pretraining.

The claimed missing ingredient is recursive self-improvement: training procedures that make the model capable of producing increasingly difficult and useful training material about the original context. Morris compares the desired dynamic to AlphaGo, where improvement creates harder future learning problems. He says Engram observed early plateau curves and is working on more sophisticated curricula that keep training progressively harder, but he supplies no experiments, benchmarks, or mechanism for how it does so.

Key Takeaways

  • Claim: The strategic gap in current AI is not broad public knowledge but durable, personalized depth over private and long-tail context. | Evidence: Morris contrasts Terence Tao’s observation that AI can connect vast public mathematical literature with the intuition of a graduate student who has spent years in one specialty. He cites private emails, company meeting transcripts, vacation preferences, and a company partnership with SunTrust Bank as information absent from standard public-data pretraining. | Implication: For enterprise agents, a long context window or RAG layer is not necessarily equivalent to a model that has internalized organizational conventions, history, and expertise. | Caveat: “Knowing” a corpus is left intentionally broad; the talk does not define a benchmark that distinguishes memorization, reliable retrieval, generalization, and useful domain judgment.
  • Claim: Scaling frontier models on public data does not automatically make them more useful on an organization’s private context. | Evidence: Morris identifies the standard scaling axes as more data, more training compute, and larger model capacity. He argues these have driven progress on Wikipedia, Reddit, arXiv, GitHub, and expert post-training datasets, but those sources remain shareable knowledge rather than proprietary context. | Implication: Organizations seeking compounding internal AI capability need a separate strategy for private-data adaptation rather than assuming successive foundation-model releases will acquire it. | Caveat: He presents this as a high-level framing rather than evidence that all post-training or memory systems are incapable of private personalization.
  • Claim: Naïvely continuing next-token training on a small private corpus is not sufficient for useful context learning. | Evidence: Using 10,000 financial reports as an example, Morris says a model can be driven to a loss of 0.0001 and effectively memorize the reports, yet generation collapses and the model cannot answer questions unless both question and answer are directly encoded in the corpus. | Implication: Fine-tuning internal documents for low training loss should not be treated as proof that a model can reason over, compose from, or generalize beyond those documents. | Caveat: The example is illustrative; no architecture, evaluation setup, or comparison data is provided.
  • Claim: Current approaches to teaching a model a context can transfer knowledge but each has structural limits. | Evidence: Morris describes KV/context compaction for data that fits in context; on-policy distillation that trains a model to behave as if documents were present; self-study or “cartridges” approaches that generate corpus-conditioned Q&A; synthetic-data continued pretraining; and unsupervised RL using losses such as GRPO. | Implication: A context-learning stack should be evaluated for retention, coverage, generalization, and its ability to improve past an initial adaptation pass—not only for short-term task performance. | Caveat: He does not claim these methods are ineffective; rather, he argues they are bounded by the data or synthetic data one defines and can introduce issues such as overwritten pretrained knowledge or the need to post-train again.
  • Claim: The hard research problem is achieving an ongoing compute-to-depth scaling curve after the first pass over a fixed context. | Evidence: Morris says that whether training uses attention matching, self-study, continued pretraining, or other methods, it eventually fits the selected data and hits a synthetic-data-style wall. He characterizes Engram’s early results as blue curves that plateaued regardless of how much data they generated or how long they trained. | Implication: Claims that a system learns continuously from enterprise data should be tested for marginal gains under additional compute and repeated learning cycles, rather than inferred from a one-time adaptation demo. | Caveat: The talk offers no measured scaling curves, definition of “depth,” or evidence that the proposed approach has overcome the plateau.
  • Claim: Recursive self-improvement and progressively harder generated training material may be the route to non-saturating context learning. | Evidence: Morris invokes AlphaGo as a successful system in which improving play produces harder training questions. He describes the target loop as generating data, improving the model, then generating better data recursively, with training made gradually harder over time. | Implication: The highest-value architecture in this category would not just index or distill a corpus once; it would create validated, adaptive learning tasks that expose increasingly subtle gaps in a model’s understanding of that corpus. | Caveat: This is an aspiration and research direction, not a demonstrated product capability in the talk; self-generated curricula can also reinforce errors if their quality is not externally grounded.

Detailed Brief

How Morris distinguishes context learning from inference-time context access

  • Claims: Compaction methods attempt to make a model act as though a long corpus remains in its active context by compressing it into a smaller representation, such as KVs.; The speaker believes this is useful for context that is already small enough to enter the model’s window, but it does not capture all of the benefits available from updating model parameters with gradients.; The desired output is a new parameter set, theta-star, derived from a pretrained model and corpus D, rather than a model that merely receives D again at query time.
  • Evidence: He relates compaction to the context-compaction behavior used in tools such as Claude Code, Codex, and OpenCode.; He references a learned and a greedy-algorithm line of work for approximating KV compaction.; He frames the formal input as a pretrained model—illustratively “GLM 5.2” or stolen Claude weights—and a large unstructured corpus such as every company meeting transcript or every email a person has written.
  • Caveats: No operational criterion is supplied for when parameter learning should replace RAG, long-context prompting, caching, or a hybrid approach.; Private-data training introduces practical questions not addressed here, including data governance, deletion, permissioning, provenance, model rollback, and cross-tenant contamination.
  • Implications: The relevant product distinction is between transient access to context and durable adaptation, which have different costs, privacy properties, failure modes, and evaluation requirements.; Any platform pursuing learned organizational memory should preserve an auditable source layer even if some knowledge is distilled into weights.

Why the fixed-data assumption is useful but incomplete

  • Claims: Morris uses fixed corpus D as the clean theoretical setup for asking how additional compute can improve a model’s understanding.; In practice, he argues the effective data budget can grow: a learner can retrieve adjacent documents, search the web, or proactively interact with people who possess relevant knowledge.
  • Evidence: He compares the process to studying a textbook or learning a language, where the learner can consult other books, search online, and speak with knowledgeable people rather than only reread the original material.
  • Caveats: Expanding the corpus may add low-quality, contradictory, or permission-incompatible information; the talk does not explain how new material is selected or verified.
  • Implications: A mature context-learning system may need an active data-acquisition loop, not just an offline training job over an existing document store.; The quality of task generation and evidence selection is likely as important as raw compute in avoiding self-reinforcing synthetic-data loops.

Notable Concepts & Terms

  • Scaling compute on context: Morris’s label for applying additional training compute to a private or domain-specific corpus so a pretrained model develops increasingly deep knowledge of it.
  • Breadth versus depth: The distinction between an LLM’s broad coverage of public knowledge and the specialized intuition accumulated through prolonged work in a narrow domain.
  • Continual learning / neural memory / sleep-time compute / write-time compute: Overlapping labels for the emerging effort to let models acquire knowledge after their original training rather than only use it in an inference prompt.
  • Theta-star: The talk’s conceptual target: an adapted parameter set produced from a pretrained model and contextual corpus D that better knows that corpus.
  • On-policy distillation: A knowledge-transfer approach in which the model is trained, with continual updates, to respond as if source information were in its context.
  • KV compaction: Compressing a long context into a smaller key-value-style representation so a model can condition on it within a limited context window.
  • Synthetic-data wall: The claimed saturation point where a model has learned the generated derivative of a corpus, so further training no longer produces deeper knowledge.
  • GRPO: A reinforcement-learning loss mentioned as one possible tool for training against unsupervised environments constructed from the context.

Operator Notes / Why Ken Should Care

  • When assessing vendor claims of “continual learning” or “AI memory,” require a test showing incremental quality gains after multiple training cycles on the same private corpus, with a predefined retention and generalization benchmark.
  • Keep retrieval-grounded citations and source permissions as a system of record; do not rely solely on knowledge embedded in weights for regulated, changeable, or deletion-sensitive organizational information.
  • Separate three architecture decisions in internal AI design: inference-time context access, parameter adaptation for stable domain behavior, and active acquisition of new evidence or expert feedback.
  • Treat self-generated training data as a quality-control risk: add external validation, held-out tasks, and regression tests before allowing recursive generated curricula to influence production models.
  • Track whether adaptation harms baseline capability or instruction following, since Morris identifies catastrophic overwriting and re-post-training requirements as meaningful tradeoffs.

Source/Metadata

  • Title: Scaling Compute on Context — Jack Morris, Engram
  • Transcript words: 3578
  • Duration seconds: 1182
  • Timestamp note: No timestamps or chapters were present in the supplied transcript.
Full transcript 3477 words · 16 min read
0:00

.

0:12

All right. Hi, everybody. My name's Jack. I'm here to talk about scaling compute on context, and also our startup Ngram, which launched last week. This isn't going to be a super detail-oriented talk where I go through a lot of experiments we've been running or talk too much about what our models do. I just want to frame the high-level problem of what we call scaling compute on context. People have many names for this. It's maybe a sub-problem of continual learning, or maybe even just the answer that we see to the problem. A little bit about myself at first. I'm Jack. I'm a researcher. I'm part of the startup Ngram. You can see me on the left in this picture.

0:59

We just launched last week with this picture on Twitter, so if you want to look us up and see more about what we do afterwards, please feel free. I'm also happy to talk after the talk concludes. So I'm going to start with this question about breadth versus depth. This is Terence Tao. He's likely the world's most famous mathematician. And he's also a really heavy user of AI, even an advocate for AI in math. And one of the things that you'll hear him talk about is how AI knows every single public mathematical topic.

1:37

And it can make connections between things that you wouldn't expect and help bridge gaps in the literature in a way that no human can know because it's read so much. But it maybe lacks the depth that you would look for from, for example, a graduate student who spent five years practicing in one area, who gets this almost subconscious intuition for the problem space. And so I think about what we're doing, which is training your data into models at a high level, as learning facts and skills like traditional AI. But at a low level, it's about depth. And the thing that we're after, this idea of scaling compute on context, is the pursuit of depth in AI.

2:27

I can give you some examples of why I think our approach is important. It's related to continual learning in general, as other people propose it. So one thing that models don't have is knowledge of what happens after they're trained. So even Fable whatever five that's coming out today was probably pre-trained with a cutoff, I would guess, at least a month ago, and has no idea if Mexico won their game last night. Actually, I don't know. I meant to check before the talk, but it's nothing to do with my pre-training cutoff because I'm a human. I'm capable of acquiring new information.

3:01

Another thing that models are bad at is hard, difficult, long-tail skill and knowledge acquisition that doesn't appear a lot in the training data. So models still are quite bad at writing AMD kernels. There are not that many good kernels written on AMD GPUs that are public, and they're intended to acquire this knowledge through pre-training, but they don't because it doesn't occur very often. And then I think maybe the most pressing case for you is, why can't I have a model like ChatGPT that knows all of my emails or knows the way that I like to write things or knows where my family and I like to vacation every year and all these little details of your life?

3:38

And there's a pretty basic reason for this, which is that ChatGPT and models like it are trained on public data. So they don't know things about your partnership with another company like SunTrust Bank. They don't know really anything about you unless you happen to be famous enough to appear in the pre-training data. So I feel this is not just an intellectual or academic problem. It's the core problem with the current paradigm in AI, that models cannot acquire new knowledge after training in a personalized way. So by definition, models have to be trained on data that's open to the public, and they can't learn the depth of the things that you know.

4:21

So just going back to Terence Tao again, how do we change this? How do we teach new things to models in a way that lets them acquire this kind of expertise or really deep skill set that we're looking for? And I'll say there's a lot of names for this. People call it sleep-time compute, continual learning, neural memory, write-time compute, note-taking, dreaming, studying, machine studying. In classical AI, maybe it's called amortized inference, and I think I'm calling it scaling compute on context. And it's almost all describing the same thing, which is something people really want.

4:54

But I think the reason why it doesn't have even a set, agreed-upon name is because the paradigm is very early and hasn't been solidified the way, for example, pre-training or post-training have. So maybe I'll take a second and talk about scaling. There's basically three axes that we use to scale AI models. We can train them on more data, we can train them for longer or add compute, or we can make the models themselves bigger, give them more capacity to acquire new information. And this is the main driver of progress from the last, really the entirety of the deep learning revolution comes from these three axes of scaling.

5:33

And the results are extremely compelling, but they're still limited to these public data sources. Models are really good at Wikipedia, they know everything about Reddit, papers on Archive, code on GitHub. And now they have this new layer of post-training data that's experts who are hired through data acquisition companies like Scale AI, Surge AI, and Mercore. But they're still by definition creating publicly available data because it's something that the model could tell to a user. So scaling is basically only used on public data, and yet it's so powerful.

6:07

So I think I'm trying to go a bit faster, so I'm not going to dwell on this, but this is the plot from Meter about how models get better every month and can complete tasks that take a longer time. And this is purely an artifact of scaling. I think the core question that I want to talk to you about today is how do we apply scale to your data? I think scale is clearly the thing that drives progress. It's not necessarily new algorithms or great new ideas.

6:39

I think maybe there's an element of data that's important, but really the thing that makes the new generation of models like Fable and GPT whatever that's coming out next month so good is that they basically scale along all three axes. I'm sure they have new data. They're certainly training for longer, and they make the models bigger. And this is how models keep getting better and will continue to get better. But I think the missing element is that this is always on public data. So models are getting better at coding in the way that is public on GitHub. They're getting better at doing math in ways that are written in public textbooks.

7:06

But they're not getting more knowledge of you or your life or your work. And that's what we're trying to change here. So if we approach the problem from first principles, I think the core limitation is that you have a fixed data budget. So say I want to scale in some fashion to train a model that knows the data from Ngram better, our company. We can't create new data. So the data-scaling axis is out the window. I think we also probably agree that we can't train a model from scratch on our data. So we probably want to start from a pre-trained model.

7:45

Or maybe another way to look at it is there's a ton of information about the outside world that is useful for understanding what happens within our company or your own context of choice. So you very likely want to start from a pre-trained model. This leaves us with essentially one axis of scaling, which is compute. And this brings us to the title of the talk today, which is Scaling Compute on Context. Just a small tangent while I have you is that I think one thing that's been beneficial for us to realize is that the amount of data isn't really fixed. There are a lot of ways that you can get more data afterwards.

8:21

Maybe in the pure math problem that I'll propose, you have this fixed data set and you want to train it into a model. But really if you're studying a textbook or trying to learn a new language, it's not really that you're limited to the words of the textbook itself. There's a lot of stuff you can do. You can find other textbooks. You can go on the internet and search for related things. You can even be proactive and talk to speakers of the language or people who know the thing that you're trying to learn. So in practice, I think the data access is very interesting and not actually fixed.

8:54

But from a core idealistic standpoint, the way we think about things is more or less, how do you scale more compute given the same data? And for the math heads in the room, I'm not going to write any equations.

9:10

But I think you can think of this as a box that you're dropped into. And all you have is this one pre-trained model. Maybe it's, I don't know, GLM 5.2. Maybe you somehow hacked into Anthropic and stole the weights of Claude. There's a lot of stuff you can do. You can find other textbooks. You can go on the internet and search for related things. You can even be proactive and talk to speakers of the language or people who know the thing that you're trying to learn. So in practice, I think the data access is very interesting and not actually fixed.

9:38

But from a core idealistic standpoint, the way we think about things is more or less, how do you scale more compute given the same data? And for the math heads in the room, I'm not going to write any equations. But I think you can think of this as a box that you're dropped into. And all you have is this one pre-trained model. Maybe it's, I don't know, GLM 5.2. Maybe you somehow hacked into Anthropic and stole the weights of Claude. And now you're trying to do it that way. But you have the pre-trained model. And then you have this unstructured data set D. So maybe this is all the emails you've ever written.

10:21

It's all the transcripts from every meeting your company's ever had. It's some very large unstructured corpus. And the question is, how do we create a better theta that knows D? And I'm going to walk through a few ideas that you could try or that people have tried and point to some links. And you can also ask me questions at the end. So the core question is something like how to produce a new model. I could call it data star that knows D. And I think the definition of know, this is a very load-bearing term in this question. And maybe that's where people get the leeway to propose new ideas. But this is essentially what you want to do.

11:00

This is what every continual learning startup is trying to do. This is more or less what we're doing at Ngram. So I'll start with a very simple idea, which is, okay, maybe you can just train the model on the data. You can use next token prediction and train it like an LLM. And I think you'll find unless you have a D that's so wide it can simulate the effect of pre-training, which no one has, then this doesn't work very well. I'll walk through an example real quick. Say we have this set of 10K financial reports. You want the model to know these. You want it to be in the weights. You want the model to answer questions about them.

11:42

You want the model to be able to create new ones. You want all these behaviors to be encoded into theta. And then you just train theta on the context that you have. You can get to a loss of 0.0001. And you can end up with a model that knows the data perfectly well. And then when you generate from it, it basically collapses. So this strategy, this naive idea of, oh, take the context that you have and train on it indefinitely, one, it's clearly bounded because there's some information that just gets perfectly transferred into the model and then you no longer learn. So this is not an indefinite axis of scaling. But two, it just frankly doesn't work.

12:40

Just doing this next token prediction on the data you have doesn't produce a model that has interesting generalization properties like normal models. It can't answer any question unless the question is perfectly encoded in the data with its answer, which is never the case in practice. So let's think about another idea since this is not quite as easy as we thought. What if we try to trick the model to think the data is in context? Because we know models are really good when you paste stuff into context. In-context learning is magical. One idea is you can do compaction, which is similar to the way that Cloud Code or Codex or Open Code, what have you, does compaction.

13:19

You take this really long context, which is D, and then you try to compress it into some set of KVs that can represent the data to the model in a very succinct way. And there's some interesting approaches to do this. You can do it in a learned way. This is a very cute paper that has a greedy algorithm for approximating KV compaction. So basically, if your data is small enough to fit into context, there are some interesting ways to compress it to something very small and pretend your model knows this.

13:41

I think there's multiple problems with this, the main one being it only applies to things that are in context, but it also misses, I think, some of the magic that you can get from taking gradients. So there's an alternate way of doing it, which is you can train the model to think the data is in context. And there's some interesting approaches here. I think Roanock was talking about on-policy distillation. This is a powerful tool for doing knowledge transfer where you have text and you show it to the model and then you make the model think that the text is in context. That's more or less the trick of on-policy distillation.

14:10

The on-policy part just means you update the model throughout training. It works. It's a pretty good algorithm. I think there are also some core problems with it, maybe the main one being what data do you actually do this with? You can't really distill the raw documents. So techniques like self-study from the cartridges paper on the left here try to generate question-and-answer pairs conditioned on D and then train the model to behave as if it is seeing D in context when it's answering questions. I think this is close to the behavior you want but also has some properties that are not necessarily appealing. And I'll get to that in a few slides.

15:00

I think there's one more idea that I think is interesting, or maybe two. I think a lot of the magic in deep learning or in LLMs, the reason why GPT-5 is so amazing, is basically because of pre-training. I think there's a lot of caveats to this statement, but pre-training is amazing for knowledge acquisition. I can ask Claude what I don't know result I got in a paper that I've written, and it actually knows this, which is incredible. And you can argue maybe they do one of these synthetic data tricks, but it more or less is knowledge that's acquired through pre-training. And so one way to teach data to a model is to simulate pre-training in some way.

15:31

And these are three pretty interesting approaches to do that, to craft synthetic data conditioned on D and then train data for longer on the synthetic data as if you're continuing pre-training. I think there are caveats to this approach, like you overwrite some of the pre-training. I think it's difficult to scale. But I think this is pretty promising. Maybe one blocker is you then have to post-train the model after doing this. So a lot of people don't actually start with good pre-trained base models. They have post-trained models, which makes this hard. But I do like this line of work, and these papers are interesting resources if you're interested in learning more.

16:12

I need to go faster. There's one more. Let's skip Andre. There's one more interesting idea, which is you can craft unsupervised reinforcement learning environments and do RL. It's pretty similar to the previous suggestion, except instead of doing some type of distillation you're just using RL loss like GRPO or whatever. I think all of these are promising but also missing maybe some core component. The thing that we're really after is to give the model more knowledge of D or to get better depth in your domain. We want to be able to add compute arbitrarily in a way that makes the model better.

17:05

So I think none of the approaches I propose do this, basically for classical machine learning reasons, which is that whatever you do, you have to define the data set and then you train on the data set and eventually things saturate. So even if it's really hard, unless your model is under-parameterized, eventually it will learn all the data. And this doesn't give the property, the beautiful scaling properties, that we see out of pre-training. It's kind of like a data wall in the synthetic sense where, when you create synthetic data from D and train on it, you eventually hit this upper bound where you've learned all of the synthetic data. And then you have to do it again.

17:55

And so I think a lot of the missing components here are how do you do it again? What's this second stage of training look like? So you can do almost any of the techniques I just mentioned. You could do the attention matching or some type of self-study thing or some continued pre-training. But eventually you will fit the data and you'll know some about D, but you won't know everything and you'll no longer have this property where you can add compute and give the model more depth. So I think a lot of the exciting work here comes from this idea of self-improvement.

18:38

It's a bit overworked as well, but I think this is actually the magic behind a lot of successful RL systems like AlphaGo, which is that AlphaGo makes its own training questions harder by getting better through training. And so I think one thing that everyone is looking for is a technique that can make models better, which makes them train themselves better.

18:54

Or this is maybe a long way of saying self-improvement. What's this second stage of training look like? You can do almost any of the techniques I just mentioned. You could do the attention matching, or some type of self-study thing, or some continued pre-training. But eventually you will fit the data, and you'll know some about D, but you won't know everything, and you'll no longer have this property where you can add compute and give the model more depth. So I think a lot of the exciting work here comes from this idea of self-improvement.

19:11

It's a bit overworked as well, but I think this is actually the magic behind a lot of successful RL systems like AlphaGo, is that AlphaGo makes its own training questions harder by getting better through training. And so I think one thing that everyone is looking for is a technique that can make models better, which makes them train themselves better. This is maybe a long way of saying self-improvement.

19:38

You generate data, and then the model gets a bit better, and then you generate better data recursively. And I think this is something that we're working on a lot at Ngram, is how do you make, when we started the company, we generated curves that look just like this blue curve. Where no matter how much data we generate or how much we train, we do plateau because there's this almost natural upper bound to how much you can learn in one go from D. But I think it turns out that there are more sophisticated things you can do that make the training gradually harder, that make the model better over time.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note