All right. Hey, everyone. Thanks for coming by over here. Hope you enjoyed the chat by Parth on how to measure continual learning. Now I'm here to talk about how we scale it up.
So a bit of background about me. I went to this company called Windsurf where I was growing the research team over there. We trained this model called SWE1 that ended up leading to the $2 billion acquisition at DeepMind, and then I ended up giving up all the acquisition money to start Trajectory, where we're building the platform for continual learning.
So, to dive in, I think it's useful to talk a little bit about what AI progress has looked like for the past couple of years. We've been in this mode where we've been rapidly scaling up benchmarks that have saturated within first years and then months. And that has continued on for the past couple of years and will continue to grow as well. The problem is it's becoming more time-consuming and more expensive. We're seeing domains where it takes four hours, six hours, 24 hours, or even several days in order to scale up benchmarks. And obviously you guys have seen the massive amount of money that the labs are pouring into scaling up RL environments and data.
Now, the thing is, we are left with a bunch of benchmarks that are getting more time-consuming, more and more expensive, and perhaps more concerningly, they're not tied to real-world use cases where people are using AI. Now, the trillion token problem here is that we are actually spending hundreds of trillions of tokens every single day on inference. And we're generating great amounts of data on how models in the real world are failing, how they're doing well, and that should be a signal that we should be capturing and training on. And this is actually how humans learn, right? We're continuously updating in the real world and getting smarter every single day.
And we're seeing now the entire field talking about continual learning, from AI experts like Ilya, Karpathy, Sholto, and Demis, to also many industry leaders as well. Satya talking about how it helps companies. And there was actually a great video by Dwarkesh a couple days ago talking a lot about the intricacies of scaling up continual learning. And obviously the goal is, we started with more data training on Internet-scale pre-training, scaling up to benchmarks, but the real unlock moving forward will be continual learning. So what's the problem? Why aren't we here yet? The main thing is there's several problems right now with our current algorithms.
So first is we have a task distribution mismatch. We're scaling up these benchmarks, which might not even be tied to reality of what people are training on. And second, oftentimes some of these methods aren't even sampling truly on policy with what's happening online. Third is we have a paradigm right now where we're sending out multiple rollouts. And this requires huge amounts of infrastructure to make sure our environments are one-to-one copies of the real world. And it ends up being a bias that we're adding to our training paradigm. And then finally, third is we are shoving every single reward into one scaler in order to train on. When the real world is messy, it's noisy, and it has rich amounts of data that should be per-token signal.
So we have all of these broken criteria, and we're training on it. And I want to take a little bit of a step through memory lane to see how we've been dealing with algorithms in the past. So the first part, we started with SFT. These were the days of instruction fine-tuning, GPT-3.5, ChatGPT. And if you take a look at these four criteria, we had parallelism solved back then. It was just one use case or example that we needed to train on. So that was a solved problem. And then the reward for SFT is per token, which is great. But we are not sampling on policy, and the task distribution is just some sort of benchmark that we've curated or some sort of data set.
Then we moved on to DPO, RLHF. This is when ChatGPT really started taking off. And we finally got online task distributions that we were able to train on. But sampling, it was a little bit better because we were actually sampling from a model, but it was still off policy. But then we ended up losing some of the key infrastructure, easy infrastructure, that we had with SFT. Now we suddenly had pairs that we were training on. And then the reward went back to sequence level.
Then we moved on to GRPO, which is the mode that we're in right now. We took a Faustian bargain and wanted to max on-policy rollouts, which is extremely powerful. We now have models that are capable of amazing things because they are on policy and able to grow and not have all of the catastrophic forgetting problems that we've had with previous methods. But on the other hand, we're working with off-policy task distributions. Our parallelism has now exploded, meaning that we need really robust environment infrastructure. And then finally, for rewards, we're back to this paradigm of training on the entire sequence.
So can we get to a world where all of these are true, where we have an online task distribution, where sampling on policy, we have one parallelism, so we don't need any of this crazy infrastructure, and then finally, our reward is actually token level? Well, this is what I'm excited to share and see how we can scale this up.
So first, let's just go through what our best algorithm of post-training is today. Hopefully, you guys are all familiar with how GRPO works at a high level. You start off with some sort of task as an input. You do a bunch of parallel rollouts. So we'll call these 01 through 04. And then you have some sort of end-state reward. So you classify these, let's say, as a couple numbers. And then GRPO works on advantage, the idea that you have some sort of mean that you're calculating over. And you're trying to shift the distribution of the model to be better on the ones that you're better on, and then obviously move away from the ones that were worse than the mean. So this is how GRPO works.
Now, let's take a look at a new continual learning algorithm called self-distillation policy optimization. Sorry, first, we'll talk through why GRPO actually is still not enough. So the first part is, if you remember, the tasks. So these benchmarks are obviously very different. It requires a bunch of parallelism. And then it's almost like you're drinking through a straw in order to get the reward. The way to think about it is, imagine you were trying to write an essay, and your teacher just gave you a score of 87 out of 100. You'd have to run through so many different examples to get to the idea of what a good essay is. And it's very simple and efficient. And so these are the fundamental problems with RL.
Now, we're going to go through a new algorithm called on-policy self-distillation. So you guys are probably familiar with distillation in general. There's been obviously a lot of talk about mythos and blocking that from happening for a lot of open-source models. The high level here is that you start with some sort of data set. It's usually off-policy or not actually tied to reality, but some sort of fixed data set. And what you have is a student model and then a smarter teacher model, so usually a bigger model. And you're trying to basically fit the log probs of the student to the teacher. And both of these are passed in with the same data.
All right. So this is normal distillation. Now, there's a new innovation called on-policy distillation. The only thing we're doing is actually swapping out the fixed data for instead a rollout of what the student would have rolled out. All right. So we take a rollout of the student. That is our trajectory. And then we're trying to fit the log probs again of a smarter teacher model to the student. Great.
Now, the problem is when we're trying to push the frontier, we don't magically have some smarter model, right, like a teacher. And so now what do we do if we're already at the smartest model? Well, here's the final algorithm. And this is where the self-distillation part comes in. If you basically take the student model and give it some sort of what we call privileged information, a hint about the world, and put that into the prompt, suddenly that student is a little bit smarter. And that's essentially the key idea of on-policy self-distillation. You take what's called this hint, put it into the beginning of the prompt, and now you match the log probs of the student without that hint to the teacher with that hint.
And this is an extremely powerful algorithm. To visualize it a little bit more, let's say you have this student prompt, right? Find the derivative of this function. What a hint would be is if you had some sort of environment information or some sort of guidance on how you should actually solve the problem. The simplest form of this is literally just an example of a golden solution. You put that into the teacher prompt and then say, here is guidance on how you would solve the problem. If you take the student model and give it what we call privileged information, a hint about the world, and put that into the prompt, suddenly that student is a little bit smarter.
And that's essentially the key idea of on-policy self-distillation. You take what's called this hint, put it into the beginning of the prompt, and now you match the log probs of the student without that hint to the teacher with that hint. And this is an extremely powerful algorithm. To visualize it a little bit more, let's say you have this student prompt, right? Find the derivative of this function. What a hint would be is if you had some environment information or some guidance on how you should actually solve the problem.
The simplest form of this is literally just an example of a golden solution. You put that into the teacher prompt and then say, here is guidance on how you would solve the problem. And you can imagine now that solving that problem from the teacher's perspective becomes a little bit easier, and we're trying to shift to those log probs had it known the answer in the first place. And so by doing this, we've solved several key problems that existed with RL. So if you remember the task distribution, now we can suddenly take something that is truly from online without needing to have benchmarks that are created.
The second part is we're still on-policy sampling, which is great. But now there's no parallel rollouts. We don't need a group of eight in order to roll out, but just from a single example, we're able to get information. So that takes away the environment bottleneck and all of these other infrastructures. And then finally, the most exciting part is we're matching every single log prob of every token.
So there's massively rich feedback about what this algorithm is doing. To see it a little more in action, and this is a really exciting part of self-distillation, it's actually not just the top token that you sampled that you're making better, but instead the entire vocabulary. So for every single token, there's a vocabulary of, let's say, 65K tokens that you're optimizing over. And let's say the model generated some sort of for loop or a range in Python.
And essentially what we're doing is saying, hey, the teacher, for some of these areas where it's blue, it wasn't the top sampled token of the student, but instead we're pushing that distribution to sample a brand new token. And this is really exciting because we're not just taking now a distribution like RL and slightly sharpening it, but we're instead actually shifting entire distributions. So this is a really exciting part about OPSD. Now, we might be saying OPSD is awesome and we've solved this problem.
For a short horizon task, it works incredibly well. If you take a look at LiveCodeBench, actually we found that GRPO saturates around Sonnet-level performance and doesn't really push the frontier. But because we're actually shifting distributions, we're able to get to brand new territory of results with a lot of data sets. And then the really cool part is with RL, a fundamental limitation, and if you guys have ever trained RL models, is the models like to think a lot, right? The more tokens that you expend, it's just going to do better. But with OPSD, you don't have that problem.
And so the actual tokens to solve some of these really difficult challenges actually collapses, which is really exciting for token efficiency. And for some of these short horizon tasks, you can actually plug and play OPSD right now with several open source projects. So OpenClaw is a really good example of this, OpenClaw RL. And you can actually use it to do simple chatbots, learn from your behaviors with unstructured data, which is super, super awesome. All right. So we might be thinking, okay, continual learning is solved. We can all go home and be super happy. The thing is, this works really well for small models, short horizon tasks, like something like a chatbot.
But this is where academic papers kind of end and where you really need to scale things up to start to see the limitations. So at Trajectory, we've been scaling up this algorithm to 120Bs, to 500Bs, to 1 trillion parameter models. And as soon as you get to the 120B range with not just one or two tool calls, but 50 or 100, things start to break apart a little bit. So first of all, eval accuracy is all over the place, right? The range is going really high. Run-to-run variance is extremely high. And then also, we start to see a lot of tool call errors.
The model is not behaving according to the format that it was trained on in the first place with the instruction fine-tuning. All right, so what are the problems that we started to see with this algorithm as we were scaling it up? So the first part is actually really funny and what I call the butt weight problem. So on shorter tasks, you're fitting to a distribution, everything's great. But when you move on to longer tasks, what you get is this student model. It's going off and doing whatever it thinks is on-policy. And then at some point, because you're so divergent in a long task, the teacher is going to try to course-correct every single time it gets a chance.
And so what you end up with is the teacher model just continuously trying to improve this token of weight or maybe or some of these hedging words. And there's this really interesting word on the bottom left that you can see, or word cloud, that as steps go on with OPSD and you really scale it up, you start to see some of these words like weight and then but start to appear. And then you actually end up in this really interesting local suboptimal position where everything just turns into maybe. And on the right here, you can kind of see some visualization of this. As the task goes on, you get two different distributions that are really divergent.
And what you end up with is the model trying to be in the middle of both of those, which is obviously really suboptimal. All right, so there's a couple of ways to solve this. One potential solution is to define step-level divergence. So in a tool-calling trajectory, let's say you have 100 tool calls going on. What we can do is start to look at the KL divergence of the student model and the teacher model as time goes on. And now we can use this as a weighting factor. So not just like normal KL where we use that as a KL penalty, but instead we're actually multiplying the token weight of every single step based on this divergence property.
And now the cool part of this is that, on the left side, we have a normal trajectory that is pretty in distribution and everything's fine, right? Everything is a weight of one. On the second part, we have a trajectory that diverges pretty heavily. And so in here, we're only going to modify W1 as the first step so that we can train that, get that right, and then move on, which is awesome. And then finally on the third part, and this is really cool by having independent weighting for every step, we can actually have a scenario where the model might go off track. We don't want to heavily weight that in.
But then later on in the trajectory, it might get back on track again, and we're fine with that and we'll mildly shift the distribution. So this is one way that we've been able to overcome this for long-horizon tool calling. The second part, and this is actually really nefarious with OPSD. So with RL, the number one problem that people face is reward hacking, as you guys are all aware of. And that's a game that continuously RL researchers have had to play. Well, there is an equivalent for OPSD as well, and that is hint leakage. The way you can think about it is we're taking this hint, right, putting it in the beginning of a prompt.
Well, if the student had no way of knowing what that hint would be, then you're going to end up with some weird scenarios and skipping some steps along the way. So the way this manifests is, let's take this example here where you're trying to find the last three digits of this formula. And a normal hint might be actually giving the correct steps and then maybe a final answer and saying, hey, actually, hint, the last three digits are all zeros. Well, you can see on the right here that when you actually roll out the model, you end up with something that says, oh, actually, I know what the solution is. It's zero, zero, zero.
So let me go back and put that into my reasoning trace and then figure out what's going on. Well, you can imagine that this is not going to occur whatsoever in the real world, and so we've ended up in this really strange position. And so there's a lot of care that you have to put into how you design these hints and making sure there's not leakage of information. So there's one kind of trivial solution to this that you can imagine, and this is just literally using an LLM to filter out these hints. So an example of this is, let's say a user can't log in. You have some sort of problem like this.
Well, you can see on the right here that when you actually roll out the model, you end up with something that says, "Oh, actually, I know what the solution is." It's zero, zero, zero. So let me go back and put that into my reasoning trace and then figure out what's going on. Well, you can imagine that this is not going to occur whatsoever in the real world, and so we've ended up in this really strange position. And so there's a lot of care that you have to put into how you design these hints and making sure there's not leakage of information. So there's one trivial solution to this that you can imagine, and this is just literally using an LLM to filter out these hints.
So an example of this is, let's say a user can't log in. You have some sort of problem like this. And there might be some sort of information, like exactly the solution, right? You'll find their SSO token that is expired. And you can have an LLM translate that into what is something reasonable that they should have known. And that's the process of looking through the logs, but not actually giving it the solution that would shortcut some of its vital reasoning. So this is one trivial solution, and it works decently well. But there are some more satisfying algorithmic approaches, too. One of these is called residual guidance.
The general idea is, in a hint, most of the time you need to actually get through the entire hint in order to get the full information. So what if you, let's say, cut it in half? Then you have a partial hint. And then you say, okay, this is the partial teacher. We're able to get the log probs of a slightly smart teacher. And then we have the full teacher as well, and that's with the full hint. Now what we can do is actually take the linear combination of both of these. And this gives us a good idea of how strong the hint is and how out of distribution it is for the original model.
And what you end up with here is, on the left you have normal OPD, where you might be entirely shifting distributions and there's almost no overlap between your model and the original solution. To this cool world where, let's say, half the hint is actually quite close to your distribution, but the full hint is very off. And you take a linear combination, and so then you're not shifting the model into unknown territory. So these are just some of the solutions to a lot of the challenges that we've had with scaling up OPSD. But it is a very, very powerful algorithm.
And with a lot of these combinations together, we're actually able to scale this up to a 120b model on Mercore Apex agents, which often requires 100-plus tool calls in order to achieve. And it's a really powerful algorithm that has even surpassed RL as well. So now we've finally arrived at an algorithm. It's not necessarily the algorithm to solve continual learning, but definitely one that is a huge step forward, that keeps the on-policy nature that makes RL so powerful. But then it also has, finally, the task distribution that is online. It has parallelism that is singular, so we don't need all of this infrastructure. And then finally, it is per-token dense reward.
So the really exciting part for us, and what we're really focused on in NetTrajectory, is this just gives you one taste of the entire continual learning loop. And there is a really exciting world that is about to come where software, in general, just gets smarter every single time it's used. And that is the most exciting unlock that's going to happen in 2026, 2027, and as we scale up. So a little bit about Trajectory. We are building the platform that turns every interaction into model improvement, harness improvement, and the entire agentic loop just getting smarter over time.
And we're building this platform where we take in agent traces data from production, we're able to optimize that as a self-serve loop, and then deploy that as a continually learning system. And we have a control plane that goes over all of that. So very quickly, our team is super awesome. We're from DeepMind, Meta, Superintelligence, OpenAI, and a lot of great product builders as well. And we've given early access to a lot of companies, so Harvey, Dekagon, Rogo, and they're super excited about what we're doing. If you're interested in any of the research that we're doing, or any of the product that we're building, definitely let me know.
Keep in touch, and happy to answer any questions. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Is it continual in the sense of very little latency so that it can come into your time, or is this something that happens? Yeah, that's a great question. So what I would say is, as a research community right now, we're in this zone of what I call pseudo-continual learning, where there's still some level of batch updates offline, and then re-uploading the model. I think it's partly an infrastructure question.
It's partly still an algorithmic question as well, of how do you truly get, when you have 10,000 rollouts going out in a product, merging those together, the infrastructure to pull all of those together. So those are some of the problems that we are solving on Trajectory. But I wouldn't say we're anywhere close to the end-to-end solution. I have a question about harness. I have a study of different harnesses. I think it's not just a little bit of a challenge. Totally. I think that's actually one of the most underexplored and most exciting questions, not only just harness improvement, right? And I think there's some literature out there now starting to explore that.
But the really exciting part is, how does the model and the harness interplay with each other as you're both updating them? That's some of the stuff that we're now exploring with our current customers and really doing those things online. But it's completely underexplored territory, and there's some really exciting innovations to be made there as well. Take all these questions outside. Cool. All right. Thanks so much, guys. We're able to get the log probs of a slightly smart teacher. And then we have the full teacher as well, and that's with the full hint. Now what we can do is actually take the linear combination of both of these.
And this gives us a good idea of how strong the hint is and how out of distribution it is for the original model. And what you end up with here is, so on the left you have normal OPD, where you might be entirely shifting distributions and there's almost no overlap between your model and the original solution. To this kind of cool world where, let's say, half the hint is actually quite close to your distribution, but the full hint is very off. And you take a linear combination, and so then you're not shifting the model into unknown territory. So these are just some of the solutions to a lot of the challenges that we've had with scaling up OPSD.
But it is a very, very powerful algorithm. And with a lot of these combinations together, we're actually able to scale this up to a 120b model on Mercore Apex agents, which often requires 100 or plus tool calls in order to achieve. And it's a really powerful algorithm that has even surpassed RL as well. So now we've finally arrived at a algorithm. It's not necessarily the algorithm to solve continual learning, but definitely one that is a huge step forward, that keeps the on policy nature that makes RL so powerful. But then it also has, finally, the task distribution that is online. It has parallelism that is singular, so we don't need all of this infrastructure.
And then finally, it is per token dense reward. So the really exciting part for us, and what we're really focused on in NetTrajectory, is this just gives you one taste of the entire continual learning loop. And there is a really exciting world that is about to come where software, in general, just gets smarter every single time it's used. And that is the most exciting unlock that's going to happen in 2026, 2027, and as we scale up. So a little bit about trajectory. We are building the platform that turns every interaction into model improvement, harness improvement, and the entire agentic loop just getting smarter over time.
And we're building this platform where we take in agent traces data from production, we're able to optimize that as a self-serve loop, and then deploy that as a continually learning system. And we have a control plane that goes over all of that. So very quickly, our team is super awesome. We're from DeepMind, Meta, Superintelligence, OpenAI, and a lot of great product builders as well. And we've given early access to a lot of companies, so Harvey, Dekagon, Rogo, and they're super excited about what we're doing. If you're interested in any of the research that we're doing, or any of the product that we're building, definitely let me know.
Keep in touch, and happy to answer any questions. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Is it continual in the sense of very little latency so that it can come into your time, or is this something that happens? Yeah, that's a great question. So what I would say is, as a research community right now, we're in this zone of what I call pseudo-continual learning, where there's some still level of batch updates offline, and then re-uploading the model. I think it's partly an infrastructure question.
It's partly still an algorithmic question as well, of how do you truly get, when you have 10,000 rollouts going out in a product, merging those together, the infrastructure to pull all of those together. So those are some of the problems that we are solving on trajectory. But I wouldn't say we're anywhere close to the end-to-end solution. I have a question about harness. I have a study of different harnesses. I think it's not just a little bit of a challenge. Totally. I think that's actually one of the most underexplored and most exciting questions is not only just harness improvement, right? And I think there's some literature out there now starting to explore that.
But the really exciting part is, how does the model and the harness interplay with each other as you're both updating them? That's some of the stuff that we're now exploring with our current customers and really doing those things online. But it's completely underexplored territory, and there's some really exciting innovations to be made there as well. Take all these questions outside.
Cool. All right. Thanks so much, guys.