Yeah, thank you Jack. I'm really grateful for the opportunity to speak here. Today I'm going to be sharing some of our frontier work on post-training and how we envision a future where agents can learn new skills on the job.
Over the last year or so, we've seen agents develop really strong reasoning skills, and they've learned to use agentic harnesses to solve longer and longer horizon tasks, which involve many turns and tool calls on complicated environment states. We're seeing an increasing demand for agents that can be deployed in a plug-and-play way into how enterprises use the agents. For instance, if they already have some method of calling the agent to do a task, they would want to be able to train a custom model to do that task instead, and that requires new ways of looking at post-training that allow you to adapt to any harness, including ones that you don't necessarily have access to the source code of.
I wanted to talk about a few different levels of post-training where each one builds on top of the last. One way that we think of this is a framework comparing it to how humans do learning, where you learn simple tasks first, and you can compound your understanding to more and more complicated tasks. Over the last year, we've, I would say, mastered or gotten a lot of reps with these simple single-turn Q&A tasks and some longer-horizon synthetic environment tasks. What we're increasingly seeing is we want to be able to adapt to custom harnesses and be able to train directly on those instead, and we think of those like internships, where you want the model to do a specific task, but you don't necessarily know exactly how the task will play out because you don't own the harness. Finally, I want to share some visions we have for the future of the custom model training space, where we think that there will be these agentic citizens, which you can deploy once, and they'll be able to adapt to many different types of out-of-distribution tasks and learn from their interactions.
First, I want to talk about the training setup for these simple Q&A tasks. We have something that looks like this, where you have an orchestrator, and the orchestrator is in charge of driving the rollouts. The orchestrator holds a task spec, which you can think of for now as just the simple prompt and answer, so something like a math question and a corresponding numerical answer. The orchestrator will send this prompt to a model and then get an answer back. Then it will send the answer to a grader and have it be graded. Once all of this is done, we want to improve our model based on that interaction, or maybe a batch of interactions, and the way we do that is through a training engine, which takes in the graded chats and produces a weight update. That weight update is then synced to some inference engines, and once those inference engines are updated, then we can start this entire process over again, where the orchestrator will have new problems to send to the model completion endpoint, and then you'll be able to get more chats, grade them, and train again. The key thing to note here is that the only thing you need for improving your model is the graded chats in some format, and once you have those, the training engine can compute weight updates to improve your model.
What's important here is that the chats are in a very specific format because we're constraining everything to be inside of our training stack. In this simple setup for Q&A, you don't have anything living outside of the training stack. You have the code of how to run the rollout and how everything is formatted, so it's a very controlled environment. However, this is limited because we can only do single-turn tasks in this way. If we want to do longer-horizon tasks and we want to build higher-order skills into our models, we need to also increase the complexity of our environment.
With synthetic environments, we have a very similar setup, but we offload a lot of the environment state outside of the training stack. You still have the same orchestrator from before, but the task spec is maybe a little bit more complicated, and the environment state is living outside of the training stack. The task spec might now include things like tool cost specs or maybe an initial state for your environment, like a file system. The orchestrator is now in charge of running many turns in series where maybe first it asks the model for how it wants to respond, and then if the model wants to call some tools, it will then call the sandbox to actually modify the environment state or read the environment state and then return those results back to the model. After all that is said and done, you get a full task trace out of this, and that task trace is then sent to a grader for grading. Very similar to what we had before, you'll be able to take the graded chats and use them to do a weight update.
The main thing to highlight here is that this orchestrator and sandbox setup is replayable, which is just saying that for any specific prompt, you can always roll back to the initial state and rerun it. You can do that in parallel or in series. The reason that's important is because the main method that we use for reinforcement learning today is GRP. We have a lot of work that involves comparing many rollouts for the same prompt and then comparing, relatively, which one is better than the other. The training engine will then make an edit to the model to upweight the trajectories that were more successful and then downweight the ones that were less successful.
Some challenges that we face in this setup are that the environment is something that you want to use to replicate reality so that after you're done training, the improvements that you've seen actually translate to when you deploy these models into production. The main problem has two names, which are both the same problem: environment fidelity and reward hacking. Essentially, the agent is exposed to an environment, and any quirks of your environment will end up being something that your agent may learn a model of.
We have some examples that we've seen. In a training run in the past, we had some networking issues causing our environment to have tool calls that failed maybe around 10% of the time. If that is the case, then we actually saw that the model would then start outputting shorter and shorter responses. This was really surprising to us because in our reward function we actually didn't have any length penalty. We couldn't really tell why this was happening, but what's going on here is, if you think about the model as a human walking along a sidewalk and the tool call failures as potholes in the sidewalk, it makes a lot of sense that because there are so many potholes, the model doesn't want to run for that long because it might fall in a pothole and then get a zero reward for the rollout.
Conversely, it's also possible that your model just learns to output more and more gibberish over time, depending on what your environment looks like. In a different case, we had a training run where we have sandbox timeouts so that they don't run forever. We usually filter out the rollouts that timed out from being trained on. One thing we saw was that if your tool calls take a long time, then if the model feels like the problem is really hard, it will actually just be incentivized to abuse the tool calls and call a lot of them in quick succession and try to time out the sandbox so it avoids getting a reward of zero. It just gets the rollout dropped.
As we scale to more and more complicated tasks, the task of replicating these environments becomes increasingly difficult because it's very difficult to perfectly simulate reality, and any mistake that you make, even if it's not intentional, will end up inducing these subtle undesirable behaviors in your model. That brings us to our next topic of bring your own harness, where we're asking, if the agent learns the exact environment distribution, why don't we just use that for our training directly, the real environment? You will no longer need to replicate anything. You can just use exactly how it's going to be used in production.
This solves a lot of problems and sounds really good. The architecture looks something like this, where we now have almost everything outside of our training stack. The only thing we have left is the model completion endpoint and some way to record the requests and responses that go in and out of the model. Everything else lives outside of the training stack and can be run in whatever fashion an existing enterprise or customer might be using. These would be existing enterprise harnesses, and essentially the reason this is nice is because we can meet customers where they're at. If they're already using the model in a certain way, we can just take our training methodology and plug it right in, and then we can help them improve the model for exactly the way that they're using it. All the orchestration loops and logic will now live outside of the training stack.
Now, this sounds really good, but the challenge here is in the data, where as you're deploying this into production and you have less and less control over how the rollouts play out, you also have less signal to learn from because you have less control over how the rollouts play out. Then you have to do that because the data is not in a familiar format.
This topic is touched on in a related work by NVIDIA. This is a paper from around a month ago where they introduce Polar, which is essentially a way to think about transitioning from a harness where you are in charge of micromanaging every aspect of the rollouts, like what I was previously talking about, and transitioning to some method of just listening in on a black-box harness, and you would no longer know exactly what the logic in here is.
Some challenges we face in this setting are non-replayability and offline or off-policy data. I think both of these are describing the same issue, which is that because we've moved so much of the logic outside of our training stack, we just don't have any way of enforcing invariance or data structures that we like. We have to be more flexible about the way we do training, and because of that, it just becomes harder to train your model and make gradient updates.
An example would be for GRPO, which is the traditional method. You would want to have many rollouts in parallel for your task, and that may not be possible anymore. Suppose a customer support chat and you have a record of how one of your chats went. There's not really a way that you could then go back and think, if I responded in this other way, would the user have been happier? There's no way to then get the user's response again.
But we're optimistic because we think that humans can do this kind of learning, and so it should be possible to formulate some kind of method that would work for models as well. If a human was in a customer support chat, they could understand somehow that, based on the customer's reaction, what they said was wrong or what they said was good, and then be able to internalize improvements for the future. I want to talk about some of the frontier research directions we have toward solving this problem. There's three main topics, which are self-distillation, automated data pipelines, and qualitative feedback ingestion.
Self-distillation is a pretty new technique, which is still, I would say, relatively narrowly scoped. We've seen successes in inducing specific new behaviors with models, but it's still an open research question of how general we can push it.
Automated data pipelines is an idea where maybe if you take a big batch of traces, you would be able to automatically flag undesirable behaviors or failure modes and then be able to put together a nice batch of training data automatically, and then send that to the model and help it improve. Currently, this is pretty manual, or human in the loop, where we go through traces ourselves and we're looking for these failure modes manually and then describing how we can improve the model and then curating those data sets ourselves.
Finally, I think an interesting direction is qualitative feedback ingestion. As you move to these production settings, sometimes you don't have access to a clear-cut binary grade or a numerical grade. Oftentimes, what you receive back is, hey, for this chat, the customer had this piece of feedback. If we can find a way to update our models based on that information, that would also prove extremely helpful. In fact, self-distillation is one way in which we're exploring how we can do that, but it's still a pretty open question.
Finally, I wanted to share a little bit about a vision for what the future of post-training might look like if we extrapolate out and take this progression to its end. I think eventually we might reach a setting where, instead of just limiting ourselves to thinking about a specific task that we can improve the model on, we can actually just think of the model as one deployment that can interact in many, many different settings. The task that you think about might just be the task of improving yourself on everything. This model may be used for all sorts of different tasks, maybe across different users as well, and be able to do some kind of reflection or introspection on, okay, for this type of interaction, here's how I self-evaluate and think that I'm doing, and then for this other type of interaction, here's how I think I'm doing, and then being able to automatically take these interactions and compute weight updates from them and improve.
Going back to the point that the agent learns every nook and cranny in your environment, the exact environmental distribution, what if the environment was just every interaction that the agent ever has? Then, in addition, we had some way that the model could evaluate itself.
One challenge that we've seen is, if you're only focusing on improving on one task at a time or flagging one failure mode at a time, you're playing a game of whack-a-mole where as soon as a new thing pops up, you need to scramble and create new data or new environments and improve the model in that way. With this kind of self-improving system that understands interactions from the environment, understands every interaction from the environment, you would get around this problem and you wouldn't need to worry about it anymore.
I want to leave people on this quote from a paper around a year ago, which I find extremely relevant now, which is that AI is at the cusp of a new period in which experience will become the dominant medium of improvement and ultimately dwarf the scale of human data used in today's systems. Yeah, if you all have any questions, I'm happy to take them after, and you can also email me at that. Thank you. Thank you. interactions and the way we do that is through a training engine which takes in the graded chats and produces a weight update. That weight update is then synced to some inference engines and once those
inference engines are updated then we can start this entire process over again where the orchestrator will have like new problems to send to the model completion endpoint and then you'll be able to get more chats, grade them and train again. The key thing to note here is that the only thing you need for improving your model is the graded chats in some format and once you have those the training engine can compute weight updates to improve your model. What's important here is that the chats are in a very specific format because we're sort of constraining everything to be inside of our training stack. So in
this simple setup for Q&A you don't have anything living outside of the training stack. You basically have the code of how to run the rollout and how everything is formatted so it's like in a very controlled environment. However, this is kind of limited because we can only kind of do single turn tasks in this way. If we want to do longer and long horizon tasks and we want to build like higher order skills into our models, we need to also increase the complexity of our environment. So with synthetic environments we have a very similar setup but we offload a lot of the environment state outside of the
training stack. So you still have the same orchestrator from before but the task spec is maybe a little bit more complicated and the environment state is living outside of the training stack. So the task spec might now include things like tool cost specs or like maybe an initial state for your environment like a file system. And the orchestrator is now in charge of running many turns in series where maybe first it asks the model for how it wants to respond and then if the model wants to call some tools it will then call the sandbox to actually like modify the environment state or read the environment state and then return
those results back to the model. After all that is said and done you get a full task trace out of this and that task trace is then sent to a grader for grading. And very similar to what we had before you'll be able to take the graded chats,
you'll be able to then use them to do a wait update. The main thing to highlight here is that this orchestrator and sandbox setup is replayable which is basically just saying that for any specific prompt you can always like roll back to the initial state and like rerun it. You can do that in parallel or you can do that in series. But the reason that's important is because the main sort of method that we use for reinforcement learning today is GRP.
So, we have a lot of work that involves comparing many rollouts for the same prompt and then comparing relatively which one is better than the other. And the training engine will then up like make an edit to the model to upweight the trajectories that were more successful and then downweight the ones that were less successful.
So, some challenges that we face in this setup is that the environment is something that you want to basically use to replicate reality so that after you're done training like the improvements that you've seen actually translate to when you deploy these models into production. So, some of the main sort of thing is that you're not going to be able to do that. And the main sort of problem is like has kind of two names which are both the same problem, environment fidelity and reward hacking.
Essentially the agent is exposed to an environment and sort of any quirks of your environment will end up being something that your agent may like learn a model of. So, we have some examples that we've seen where in a training run in the past we had some like networking issues causing our environment to have tool calls that failed maybe around 10% of the time. If that is the case then we actually saw that the model would then start outputting shorter and shorter responses. Now this was really surprising to us because in our reward function we actually didn't have any length penalty.
So, like we couldn't really tell why this was happening but really what's going on here is if you think about maybe the model is like a human like walking along a sidewalk and like the tool call failures are like potholes in the sidewalk. Like it makes a lot of sense that because there's so many potholes the model doesn't want to run for that long because it might fall in a pothole and then get a zero reward for the rollout.
And then conversely it's also possible that your model just learns to like output more and more gibberish over time depending on like what your environment looks like. So, in a different case we had a training run where we have sandbox timeouts just so that they don't run forever.
And we usually like filter out the rollouts that timed out from being trained on. One thing we saw was that if your tool calls take a long time then if the model feels like the problem is really hard it will actually just be incentivized to like abuse the tool calls and just like call a lot of them in quick succession and try to time out the sandbox so it avoids getting a reward of zero. It just gets the rollout dropped.
So, as we scale to like more and more complicated tasks the task of like replicating these environments becomes increasingly difficult because it's very very difficult to like perfectly simulate reality and sort of any mistake that you make even if it's not intentional will end up inducing these like subtle undesirable behaviors in your model.
So, that brings us to our next topic of bring your own harness where we're basically asking like if the agent learns the exact environment distribution why don't we just use that for our training like just directly the real environment you will no longer need to replicate anything you can just like use exactly how it's going to be used in production. This solves a lot of problems and sounds really good. The architecture looks something like this where we now have almost everything outside of our training stack. The only thing we have left is the model completion endpoint and some way to record the requests and responses that go in and out of the model.
Everything else kind of lives outside of the training stack and can be run in whatever fashion that like an existing enterprise or like customer might be using. So, these would be like existing enterprise harnesses and essentially the reason this is nice is because we can meet customers where they're at. Like if they're already using the model like if they're already using the model in a certain way we can just take our like training methodology and just like plug it right in and then we can help them improve the model for like exactly the way that they're using it. So, all the orchestration loops and logic will now live outside of the training stack.
Now, this sounds really good but the challenge here is in the like data where as you're deploying this into production and you have less and less control over how the rollouts play out, you also have like less signal to learn from because you have less control over how the rollouts play out. And then you have to do that just because the data is not in like a familiar format. And this topic is touched on in a related work by NVIDIA. This is like a paper from around a month ago where they introduce Polar which is essentially a way to think about transitioning from a harness where you are kind of in charge of micromanaging every aspect of the rollouts.
Kind of like what I was previously talking about and transitioning to some method of just listening in on a black box harness and you would no longer know exactly what the logic in here is.
So, some challenges is that some challenges we face in this setting are non replayability and offline or off policy data. I think both of these are describing the same issue which is just that because we've moved so much of the logic outside of our training stack, we just don't have any way of like enforcing sort of invariance or like data structures that we like. We have to be more flexible about the way we do training and because of that it just becomes harder to train your model and make gradient updates. So, an example would be for GRPO which is like the traditional method.
You would want to have many rollouts in parallel for your task and that may not be possible anymore. If you think about suppose like a customer support chat and you have a record of how one of your chats went, there's not really a way that you could then go back and think, oh, if I like said or if I responded in this other way, like would the user have been happier? Like there's no way to then get the user's response again. But we're optimistic because like we think that humans can do this kind of learning and so it should be possible to like formulate some kind of method that would work for models as well.
Like if a human was in a customer support chat, they could understand somehow that like based on the customer's reaction, like what they said was wrong or what they said was good and then be able to like internalize improvements for the future. So, I want to talk about some of the frontier research directions we have towards like solving this problem. There's kind of three main topics which are self distillation, automated data pipelines and qualitative feedback ingestion. Self distillation is a pretty new technique which is still, I would say, relatively like narrowly scoped.
So, we've seen successes in inducing like specific new behaviors with models but it's still an open research question of like how general can we push it. Automated data pipelines is an idea which maybe if you take like a big batch of traces, you would be able to like automatically like flag undesirable behaviors or failure modes and then be able to like put together like nice batch of training data automatically. And then send that to the model and help it improve.
Currently, this is like pretty manual or like human in the loop where like we go through traces ourselves and we're like looking for these failure modes manually and then like describing how we can improve the model and then curating those data sets ourselves. And then finally, I think an interesting direction is qualitative feedback ingestion. So, as you move to these like production settings, sometimes you don't have access to like a clear cut binary grade or like a numerical grade. Oftentimes, what you receive back is like, hey, for this chat, the customer had this like piece of feedback.
If we can find a way to update our models based on that information, that would also prove extremely helpful. And in fact, like self distillation is one way in which we're exploring how we can do that but it's like a pretty, it's still a pretty open question.
Finally, I wanted to share a little bit about a vision for what the future of post training might look like if we sort of extrapolate out and take this sort of progression to its end. I think eventually we might reach a setting where instead of just limiting ourselves to thinking about a specific task that we can improve the model on, we can actually just think of the model as just one deployment that can interact in many, many different settings. And like the task that you think about might just be the task of improving yourself on everything. And this model may be used for all sorts of different tasks, maybe across different users as well.
And be able to sort of do some kind of reflection or introspection on like, okay, for this sort of type of interaction, here's how I like self evaluate and think that I'm doing. And then for this other type of interaction, here's how I think I'm doing. And then being able to automatically take these interactions and compute weight updates from them and improve.
And so, yeah, going back to the point that the agent learns every nook and cranny in your environment, the exact environmental distribution. What if like the environment was just like every interaction that the agent ever has? And then in addition we had some way that the model could evaluate itself.
Basically, one like question that we've, or sorry, one challenge that we've seen is like, if you're only focusing on like, improving on one task at a time or like flagging one failure mode at a time, you're kind of playing a game of whack-a-mole where as soon as a new thing pops up, you need to scramble and like create new data or new environments and improve the model in that way. With this kind of like self improving system that understands interactions from like, understands every interaction from the environment, you would like kind of get around this problem and you wouldn't need to worry about it anymore.
So I want to leave people on this quote from a paper around a year ago, which I find extremely relevant now, which is that AI is at the cusp of a new period in which experience will become the dominant medium of improvement and ultimately dwarf the scale of human data used in today's systems.
And yeah, if you all have any questions, I'm happy to take them after and yeah, you can also email me at that. Thank you. Thank you.