Bringing Continual Learning into Enterprises — Samuel Denton, Applied Compute
Description
A Qwen thinking model was taking up to 80 turns to submit on SWE bench. Applied Compute wanted it wrapping up by turn 40 and got the submit tool call rate from 22% to 60% with test pass rate flat. The interesting part is the mechanism: because the rollout was conditioned on an old production trace that never called the tool, the teacher never touched the tool call tokens at all. It moved the reasoning path toward the call instead, and the call followed. Sam Denton's frame is a grid. One axis is how online the traces are, from a single dump of production traces to a unified engine where serving and training are the same loop. The other is where the hint comes from, either static priors, such as knowing a support agent is too quick to refund, or a hint built dynamically from what the on policy model just did. Applied Compute works two corners of that grid. Offline hints on offline traces need no replayable environment and can improve an enterprise agent from a data dump on day one. Online hints on online traces have the far higher ceiling, and that is what fixed a customer whose harness required unusual hyperlink formatting: rewarding the format directly and finetuning on correct examples both degraded coding ability, while a hint written against each rollout took correct formatting from 15% to 80%. Two things he says make it work in practice. Let a judge pick where in the rollout the hint goes and distill only the next few steps, since the learning signal decays with distance from the hint. And mask which tokens you learn from, because the teacher has strong opinions about connector words that have nothing to do with the lesson. Throughout, the constraint he keeps is doing all of this without a golden answer to distill toward. Speaker info: - https://x.com/samueldenton - https://www.linkedin.com/in/sam-denton-161b50126/ Timestamps: 0:00 - The distillation spectrum, offline to online 2:46 - The holy grail: serving and training as one loop 4:00 - Where the hint come
Summary
Generated by gpt-5.6-terraAt-a-Glance
- Verdict: Watch fully
- Core thesis: Enterprises can begin improving agents immediately from static production traces using offline hints, then progress to higher-ceiling continual learning by generating rollout-specific online hints and updating models during production use.
- Why it matters: The talk offers a practical taxonomy and several implementation patterns for turning agent traces into post-training signal without requiring golden labels for every task.
- Best use: Use it to design an agent-improvement pipeline: start with targeted behavior changes from logged traces, then add online, judge-driven hinting for adaptive continual learning.
Executive Summary
Samuel Denton frames continual learning as two interacting spectra: whether traces are offline or generated by the current policy online, and whether the instructional “hint” is static or dynamically derived from a specific rollout. This creates a four-quadrant distillation taxonomy. Applied Compute concentrates commercially on offline-trace/offline-hint training for immediate enterprise value and online-trace/online-hint training for the most scalable long-term improvement loop.
The central operating premise is that production traces are useful even when they lack gold answers or fully replayable environments. An enterprise can provide a batch of agent traces plus a desired behavior prior—such as avoiding unnecessary refunds, completing work before a turn limit, or following customer-specific output conventions—and use self-distillation to steer future behavior. The company positions this as a way to improve agents from day one rather than waiting for a fully online training stack.
The strongest empirical examples are narrow but operationally relevant. On SWE-bench, an offline hint about nearing a 40-turn limit increased an agent’s task-completion tool-call rate from roughly 22% to 60% while its test-pass rate stayed approximately flat. In a production coding-agent case, dynamically generated online hints increased correct, out-of-distribution hyperlink formatting from about 15% to about 80%, outperforming a static instruction and avoiding the coding-performance degradation reportedly seen with reward shaping or SFT.
The practical lesson is to make learning local and selective. Rather than prepend a generic hint to an entire trajectory and train on every subsequent token, use an LLM judge to identify the relevant decision point, inject a rollout-specific hint there, and distill only the next step or a short window. A second filter, “relevance mass self-distillation,” selects only teacher tokens that carry the desired behavioral signal, reducing imitation of stylistic filler and the risk of broad capability degradation.
Key Takeaways
- Claim: Continual-learning systems should be designed across two independent dimensions: trace freshness and hint freshness. | Evidence: Denton defines offline versus online production traces, and offline versus online hints; their intersection yields four distillation modes, from static traces plus static behavioral priors to fully on-policy rollouts plus rollout-specific hints. | Implication: Ken can choose the least operationally demanding training setup that fits current data and environment constraints, instead of treating real-time online learning as a prerequisite. | Caveat: The categories are presented as conceptual spectra rather than rigid system boundaries.
- Claim: Static production traces can produce targeted agent improvements without gold answers, environment replay, or direct edits to the desired action tokens. | Evidence: Applied Compute used an offline trace/offline hint setup on SWE-bench: a hint warned the agent that it was near a 40-turn limit and should finalize, verify, and submit. Task-completion tool-call rate rose from about 22% to 60%, while test-pass rate was roughly unchanged or slightly higher. | Implication: For existing agent fleets, logged trajectories plus explicit behavior priors can be a viable first post-training dataset, especially for correcting repeated operational failure modes. | Caveat: The reported result targets a specific completion behavior; the transcript does not provide sample size, confidence intervals, model-training details, or broader task-generalization evidence.
- Claim: Adding even a single on-policy rollout step to offline trajectories can improve distillation because the teacher can reinforce the actual action token that the improved policy should take. | Evidence: In the SWE-bench example, the fully offline trace never contained the task-completion tool call. Denton says one on-policy step conditioned on the offline trace let the teacher encourage that tool token and yielded a better combined SWE-bench pass rate than the fully offline setup. | Implication: A practical middle path is to preserve historical traces but periodically regenerate selected decision points with the current policy, avoiding the infrastructure burden of fully online training. | Caveat: This hybrid approach assumes the relevant agent state can be reconstructed sufficiently to generate an on-policy step, even if the full environment is not replayed.
- Claim: Online, rollout-specific hints are materially more effective than a static instruction for teaching rare or out-of-distribution output conventions. | Evidence: For a customer-specific coding-agent hyperlink format, a dynamic hint referring to the agent’s prior rollout increased correct formatting from roughly 15% to roughly 80%. Applying the same static formatting reminder to every rollout improved the behavior much less. | Implication: When an agent fails in context-dependent ways, feedback should reference what it actually did and where it diverged, rather than repeat generic policy text. | Caveat: The transcript reports one customer-harness formatting case, so the magnitude should not be assumed to transfer to every behavioral objective.
- Claim: Rewarding or supervised-fine-tuning a narrow behavioral requirement can degrade broader agent capability; distillation can be a safer steering mechanism. | Evidence: Denton says that adding a reward for the required hyperlink format, and SFT on correctly formatted traces, caused degradation in overall coding-agent performance, whereas their online-hint distillation achieved the target behavior without stated regression. | Implication: For high-value but narrow enterprise constraints, evaluate capability preservation explicitly and consider targeted distillation before applying broad reward or SFT pressure. | Caveat: No quantitative base-capability results or training controls are shown, so this is a directional implementation lesson rather than proof that distillation universally dominates reward modeling or SFT.
- Claim: The learning signal should be applied at the relevant trajectory step and over a short horizon, not indiscriminately over the entire rollout. | Evidence: Applied Compute uses a judge to select where to inject hints, then distills the immediately following step or a few subsequent steps. Denton reports that the KL learning signal declines as tokens get farther from the hint. | Implication: Agent-training pipelines should log and score step-level decision contexts, not merely label entire conversations or task outcomes. | Caveat: This introduces dependence on judge quality: a poor judge can identify the wrong intervention point and generate misleading learning signal.
- Claim: Selective token learning can make self-distillation more robust for unusual behaviors and reduce destructive imitation. | Evidence: Their “relevance mass self-distillation” uses an LLM judge to sample and select teacher tokens worth learning from, excluding irrelevant preferences such as connector-word choices. Denton says this improved learning of very out-of-distribution behavior while better avoiding catastrophic degradation. | Implication: Do not train uniformly on teacher outputs: separate behavior-causal tokens from incidental style tokens to improve controllability and preserve general competence. | Caveat: The transcript references a company blog post rather than providing the algorithm, selection criteria, or quantitative graph values.
Detailed Brief
The four-quadrant deployment taxonomy
- Claims: Offline hint plus offline trace is positioned as the enterprise entry point: analyze an existing trace dump and apply a general behavior prior across it.; Offline hint plus online trace applies a fixed desired behavior during current-policy rollouts.; Offline trace plus an on-policy step is a hybrid mode: use historical context, generate one current-policy action, and construct feedback against that action.; Online hint plus online trace closes the continual-learning loop by generating feedback from the model’s complete production rollout and using it to update future behavior.
- Evidence: The recurring example of an offline behavioral prior is a customer-support agent that gives refunds too readily.; Denton describes the online/online quadrant as usable either in a replayable environment or while serving actual production traffic.; Applied Compute characterizes the offline/offline mode as immediate value and the online/online mode as the highest-ceiling route for raising overall evaluations.
- Caveats: The talk does not address deployment gates, rollback mechanisms, data retention, privacy controls, or how quickly online model updates should be promoted to production.; The claim that infrastructure will increasingly unify inference and training is a strategic forecast, not an implementation plan.
- Implications: Continual learning should be treated as a maturity path, with data collection, judging, replay, and promotion controls added progressively rather than built all at once.; The economically useful distinction is not simply offline versus online training; it is whether the feedback can adapt to the policy’s actual failure in context.
What the examples imply about evaluation
- Claims: Targeted behavior improvements require separate metrics for the new behavior and the underlying task quality.; Combined success should measure whether the agent both takes the desired operational action and solves the original task.
- Evidence: The SWE-bench setup separates task-complete rate, test-pass rate, and a combined pass rate defined as their intersection.; The stated goal was to induce submission before turn 40 without degrading the base task’s test performance.
- Caveats: The transcript does not specify whether the model was evaluated on held-out tasks, how contamination was controlled, or whether tool-call frequency could be gamed independently of useful completion.
- Implications: For agent operations, every behavioral fine-tuning objective should have a paired non-regression suite and an intersection metric; a compliance metric alone is insufficient.; Turn limits, formatting requirements, and tool-use policies should be evaluated as part of end-to-end task success rather than as isolated classifier scores.
Notable Concepts & Terms
- Distillation spectrum: A continuum from one-time learning from static production traces to an integrated inference-and-training loop that updates a serving model from current traffic.
- Offline hinting: Applying static priors or task rubrics to traces, such as discouraging excessive refunds or prompting timely task completion.
- Online hinting: Constructing feedback dynamically from the policy’s own rollout, enabling context-specific correction rather than a universal instruction.
- On-policy step: A current-model rollout action generated against a historical trace context; it can bridge offline data and policy-specific learning.
- Per-step hinting: Using a judge to locate the decision point that needs correction, then distilling only the immediate next action or short future window.
- Relevance mass self-distillation: A judge-guided token-selection method intended to train on behaviorally relevant teacher output while ignoring incidental phrasing and style.
- KL learning signal: The divergence-based distillation signal; Denton uses its decay with distance from a hint to justify localized training windows.
- SWE-bench pass rate: In this talk, the intersection of task submission/completion and test passing, used to prevent a behavioral metric from obscuring task-quality regression.
Operator Notes / Why Ken Should Care
- Instrument agent traces at the decision-step level, including tool calls, state/context, outcomes, and policy version, so targeted intervention points can be reconstructed later.
- Create a behavioral-improvement backlog from observed production failures, with each item defined by a target metric, a base-capability non-regression metric, and an end-to-end intersection metric.
- Prototype offline-trace/offline-hint distillation first for one measurable operational behavior before committing to a live continual-learning architecture.
- For policies involving unusual customer-specific formats or workflow rules, test rollout-specific corrective hints against static system-prompt reminders and against SFT/reward baselines.
- Add an LLM-judge validation layer for both hint placement and token selection, then audit judge errors before allowing any learned policy update to reach production.
- Require staged evaluation, canary promotion, and rollback controls for online-learning candidates; the presentation demonstrates learning gains but does not cover production safety governance.
Source/Metadata
- Title: Bringing Continual Learning into Enterprises — Samuel Denton, Applied Compute
- Transcript words: 5124
- Duration seconds: 1142
- Timestamp note: No usable timestamps or chapter markers were present in the supplied transcript; the latter portion substantially repeats earlier material.
Transcript
Sam Denton Reviewer All right. Can you hear me? Yeah, I'll take that as a yes. Cool. So we'll go ahead and get started here. Today, what we're going to be talking about is bringing continual learning into enterprises and how we're doing it at Applied Compute. A bit of an intro: my name is Sam Denton. I lead the platform research team at Applied Compute. So here's our loose agenda for the day. We're going to start by talking about the distillation spectrum and how we think about different areas on the spectrum of distillation. Then we're going to talk about where value accrues along this distillation spectrum. We'll show a bunch of data on how distillation is working in enterprises. If I have time, I'll try and get to some tips and tricks we found at Applied Compute to make self-distillation work and continual learning work in the enterprise. And then, finally, we'll wrap up and talk about what we've learned. Also, I'm going under the assumption that most people have some kind of context on distillation and self-distillation. I figured by 3 p.m. on a continual learning day, you'd had enough of it. So I'm just going to dive right into it. So first, I want to talk about the distillation spectrum and how we think about the spectrum at Applied Compute. So I want to define this offline and online distillation spectrum. So we'll start on one end of the spectrum over here, which is offline distillation. What this looks like is you get a single batch of traces from a production agent, and you're meant to do something with it. How do you learn from a bunch of production traces of some agent out in the wild? You want to learn via hindsight. You want to learn via the mistakes the agent made, whatever you can learn from this one-time batch of production traces. And this is the offline end of the spectrum. In the middle, we have something that might look like a daily batch of production traces. So maybe you deploy a model, and then at the end of every day, you collect a bunch of traces, and you figure out, what can I do with these traces? How do I make use of these traces? This is not as offline as a single lump of production traces, but it's not fully online in the sense of all the way on the right here. We have this unified engine of putting inference and training together, right? And this is the holy grail of continual learning, where I have a model that's serving production traffic. It does a rollout. It creates a trace. We figure out how to learn from that trace. We update the model, and then we serve the next production request. So there are a bunch of different points along this spectrum where we think distillation and continual learning can be useful, but this is how we think about the two ends of the spectrum and how we try to do continual learning across the whole spectrum. On the left side, on the offline side of traces, this is where a lot of enterprises are today. They say, okay, we have a bunch of production traffic, we have a bunch of traces. Figure out a way to make our agent better. It's clearly doing something, but it clearly can be better. And so, how do you make our agent better today, given this one-time batch of production traces? And on the right side, this is the full, complete flywheel, right? This is, we have some enterprises that are willing to start getting something into production, which looks like this fully online continual learning, where we essentially deploy a model, and then we're able to make updates as the model is serving production traffic. So that's the way. Our goal at Applied Compute is to meet enterprises where they are, right? So they're across this spectrum, and we want to provide value across both ends of the spectrum. So this is online and offline distillation, but there's a whole other axis here we think about, which is hinting, right? So the whole point of distillation is we have some kind of information that allows us to create a teacher model, which is smarter than the student model or the on-policy model. In order to create a teacher that's smarter than this on-policy model, we need to create some kind of hint or have some kind of privileged information. And so the question is, where does this hint come from? So, in the offline hinting world, we're deriving hints from some static or offline data. So this can be potentially known rubrics for a single task. It can be general priors about behavior that needs to get better, such as a customer support agent that is too willing to give refunds, for example, right? This is some known behavior that you're trying to improve. Or it could just be general things that we've seen in production about loss reports and saying, oh, the model tends to miss on questions like this. But it's independent of the online model's rollout. So there's a whole world of creating offline hints from static data. There's also online hints, right? And so online hinting are hints that are dynamically constructed from the online rollout. The idea here is that we can also inject other pieces of information, like behavior priors and things like that. But the goal is to create a hint that is completely dependent on the online rollout and the online policy that's doing the work. So we have these two online and offline spectrums. We have hinting and we have distillation, which leaves us with a very nice two-by-two grid, right? And so here are the four types of distillation that we see. And again, these are all spectrums, right? So I've drawn lines and put things in boxes where boxes sometimes don't make sense as boxes, but more as spectrums. But in general, these are the four quadrants of the continual learning distillation spectrum. So, in the first box, we have an offline hint paired with an offline production trace. So what this looks like is we take a trace from production, and we have some prior about how production traces generally aren't good enough. Again, maybe customer support is too quick to give refunds, things like that. And we construct a hint that we apply to all of these offline production traces. And then we do this distillation task, and eventually we create some smarter model from offline hints paired with offline production traces. In the second quadrant, we have offline hints paired with on-policy and online production traces. So again, what this looks like is we have a hint of a behavior we're trying to improve. We let an online agent do a rollout. We inject some hint that we're injecting into all of our rollouts, try to update the policy model, and then continue serving production traffic. In the third quadrant, we have off-policy traces with some on-policy step and hints that are constructed against that one on-policy step. I'll talk a little bit more about what an on-policy step means later on. But basically, the idea is that the trace that led to the point where I inject a hint was fully off-policy. It was some production trace that came from a few days ago. And we are using our on-policy model to just roll out one step without actually having to interact with the environment. And we construct the hint based on what that on-policy model did in that one step. So again, we have an offline production trace. We pick a moment in time to do some online step. And then we construct a hint based on what the on-policy model did in that one online step. And then, finally, we have this fourth quadrant, which is online hints with online production traces. So what this means is we have an on-policy model do a bunch of things in production. It finishes. We stop. We look at what the online model did. We then create a hint dynamically from the full rollout that the online model did. And then we construct a hint. And then we do some distillation against this online production trace with an online hint. So, in general, at Applied Compute, we're really, really focused on quadrant one and quadrant four, though we do research across all four quadrants of this table. And so quadrant one, basically, is how we meet enterprises who are ready to have their production agents improve today. And so what this looks like is we don't actually have to have replayability of a production environment, right? We can take a bunch of production traces and we just look at what happened. And then we can essentially construct these offline hints for behavior changes that we're trying to target and improve. Again, not giving refunds quite as often. I'll talk a little bit about formatting things or reasoning, the amount of reasoning we're trying to encourage. But basically, these are off-policy traces where we inject some offline hint. And this allows us to really target specific behaviors. On quadrant four, this is our most scalable solution to climbing overall evals. So this can be done with either a replayable environment or actually serving production traffic. And because we're constructing these hints online in a dynamic way, we can actually cater to a bunch of different behaviors via distillation, right? And so, what this looks like is we don't actually have to have replayability of a production environment. Right? We can take a bunch of production traces, and we just look at what happened. And then we can essentially construct, we have these offline hints for behavior changes that we're trying to target and improve. Again, not giving refunds quite as often. I'll talk a little bit about formatting things or reasoning, the amount of reasoning we're trying to encourage. But these are off-policy traces where we inject some offline hint. And this allows us to really target specific behaviors. On quadrant four, this is our most scalable solution to climbing overall evals. So, this can be done with either a replayable environment or actually serving production traffic. And because we're constructing these hints online in a dynamic way, we can actually cater to a bunch of different behaviors via distillation. Right? And this is how we complete this flywheel, where we have an online model serving production traffic, constructing hints online dynamically based on what it did, and then updating our model accordingly. So, again, this is our full training taxonomy. We've done work across all four. And today, I'm really going to focus on quadrant one and quadrant four, which is where we spend most of our time. In terms of how that grid maps to where value accrues, this is how we think about it. We can improve for free today, and we can raise all ceilings tomorrow. We can improve for free today by using offline production traces. Give us a dump of your production data. We'll find a way to make it valuable. And then, as we deploy an online policy model, we can then raise the ceilings continuously by updating the model as it's serving production traffic. And I think the most important thing I want to call out here is that when we think about how we do this, we want to do this without having access to some golden answer. I think this is something that generally frustrates me a lot in the distillation space, is a lot of distillation work is done assuming you have some kind of golden answer that you can distill into the model. And this is often not the case. And so, in general, we want to think about how we can do continual learning and distillation without having some beautifully golden rubric to accompany every task. As infra collapses between serving and training, we're automatically going to raise the ceiling continuously via online distillation. So, again, we have this spectrum: one-time batch of offline production data, and then online traces. And our goal is to provide value across the full spectrum. Cool. So, let's spend a little bit of time talking about the data and some of the results across these four quadrants. So, in the world of offline traces, offline hint, we have this setup, right, where our goal is to essentially take a QN 3.5 thinking model against Sweebench and get it to submit its reasoning faster than it normally does. So, on Sweebench, we found that this model was essentially taking up to 80 turns to submit its answer. And what we wanted to do was encourage it to call a tool to submit its task before turn 40. And the reason for that was to prove to ourselves that we could get it to wrap up its reasoning quickly by turn 40 without letting it do its normal full reasoning chain. And so, we'll talk about three metrics here. The first is the task complete rate, which is the percentage of the time that the agent calls this tool to quote-unquote finish a submission. The second is test pass rate. So, this is how we measure the regression in performance on the base task. And this is the percentage of the time that the environment passes all tests accompanying the Sweebench task, irrespective of whether the agent submitted the task via this tool call. And then finally, we have our Sweebench task rate, or pass rate, which is how we combine these two metrics. It's the intersection of those two behaviors. And so, the goal is we want to raise the Sweebench pass rate performance while not degrading the test pass rate. And I've included a hint here of what this looks like in practice. It says something like, you are near your 40-turn limit. There's only about three turns left. You often keep exploring and forget to wrap up investigating. So, finalize and verify your fix, and then call this tool before you run out of time. So, this is what the results look like. So, we were able to target the specific behavioral change, which was to call this tool when it wanted to submit a task, and without any degradation to the overall performance. So, you can see that maroon color is the test pass rate. It's relatively constant. In fact, it goes up a little bit. But the task complete call rate increases dramatically from about 22% to 60%. So, we're able to add this behavior. And I think the really interesting, surprising result here is, again, this is in the fully offline world. And so, we're taking a production trace, or a trace that was created ahead of time, that never called this task complete tool. And we're nudging with the student and teacher models, the student model towards calling this task complete tool call, without ever specifically changing the tokens for the tool call. Because, again, the rollout is conditioned on the quote-unquote production trace, right? And so, it never had the reasoning path to think to call the tool call. And so, the teacher doesn't force the tool call. It just starts to force the reasoning path towards the tool call, without ever actually changing the tool call. Which I think was really cool and surprising to us. Now, that being said, there actually is a little bit of a cheat here that we can use. Which is that you can, and as I mentioned earlier, roll out just one step from the on-policy model, given an offline production trace. And when we do this, we're obviously seeing that the student model learns to wrap up its reasoning. And eventually, the teacher starts encouraging it to actually call this tool token. And so, you can see by having something that's a little bit more on-policy, we're able to increase the Sweebench pass rate more than in the fully offline world. So, again, this is offline trace, offline hint, with just one step on-policy. Cool. So, then what does it look like in the fully online, online trace, online hint world, where it's serving production traffic? So, for a certain production use case we had, we needed to teach a coding agent to use very specific formatting for hyperlinks due to a certain harness nuance of one of our customers. And obviously, this coding agent needed to not regress on any of the base coding agent capabilities. Now, the problem here is that these hyperlink formats were very, very out of distribution for previously post-trained models. And so, when we tried things like adding a reward for specific hyperlink formatting, or even doing SFT on traces where we knew the hyperlink was correctly formatted, we saw that there was this degradation in overall coding agent performance. And so, what we did here is set this up as an online trace with an online hint. So, what this looked like is that we would do a rollout, then we would inject a hint specific to the rollout that occurred from the on-policy model, and then say, in your prior rollout, you'd formatted hyperlinks like this. Next time, make sure to format hyperlinks in this way instead. And so, what we were able to see is that the percentage of correct hyperlink formatting jumped drastically from about, I guess, 15% all the way up to around 80%. And the other graph here shows what happens if we try to do offline hinting. So, this is for every single rollout, apply the same hint, which says, remember that when you do hyperlinks, you have to format it this way. And you can see that we do climb the behavior a little bit, but far less than in this online hinting world. So, here we've seen two different results: one where we can use offline hinting and offline traces to climb from production traces, and then another one where we're able to actually use the on-policy-ness of the model and online hints to improve a behavior when we're serving production traffic. Okay, cool. I think I have enough time here to talk about a little bit of tips and tricks here. So, the first is that we found that per-step hinting is very, very important to making distillation work. So, rather than injecting a hint to the beginning of a rollout, we use a judge to essentially decide where in the rollout we should be injecting hints, and then actually have found that it's best to just do distillation on that next step that occurs, or maybe a few steps forward, rather than the entire rollout. Because that's really the turn and the moment in time that you want to have the teacher teach something to the student. You can also see in this graph here that this KL learning signal goes down as you get further and further away from the hint, which makes sense. Another trick that we've used is something called relevance mass self-distillation. And there's a blog post on our website about how we've done this. Okay, cool. I think I have enough time here to talk about a little bit of tips and tricks here. So, the first is that we found that per-step hinting is very, very important to making distillation work. So, rather than injecting a hint to the beginning of a rollout, we use a judge to essentially decide where in the rollout we should be injecting hints, and then actually have found that it's best to just do distillation on that next step that occurs, or maybe a few steps forward, rather than the entire rollout. Because that's really the turn and the moment in time that you want to have the teacher teach something to the student. You can also see in this graph here that this KL learning signal basically goes down as you get further and further away from the hint, which makes sense. Another trick that we've used is something called relevance mass self-distillation. And there's a blog post on our website about how we've done this. But essentially, the idea is that we use an LLM judge to sample and choose which tokens we actually learn from, from our teacher model. Because often we'll see that the teacher model has preferences of certain connector words that are not really relevant to actually what we're trying to teach the student. And we can see in the graph at the bottom that we're able to increase our ability to learn a very, very out-of-distribution behavior, while also being better about avoiding catastrophic degradation. Cool. So, overall, where does that leave us? So, obviously, I assume everyone here is on the distillation train, but it's a very, very valuable tool towards continual learning. We introduced a spectrum of offline and online rollouts, as well as offline and online hinting, and how we use them towards distillation. So, we use offline hinting with offline production traces to provide value on day one to enterprise clients. Give us production traces, and we can teach it a certain behavior. We then use online hinting and online production traces to do this highest-ceiling continuous learning improvement across multiple improvement areas, because that judge is able to adapt to whatever the online model does in production. And finally, I just want to say thank you to the team that worked on this. A lot of the work was done by others. I just got to present it. And we're hiring and having a lot of fun working on research problems around continual learning. So, if you're interested, reach out to hiring at applycompute.com or, yeah, just email me as well. So, thank you, everyone. So, congratulations. And what we wanted to do was encourage it to call a tool to submit its task before turn 40. And the reason for that was basically to prove to ourselves that we could get it to sort of wrap up its reasoning quickly by turn 40 without letting it do its normal sort of full reasoning chain. And so, we'll talk about three metrics here. The first is the task complete rate, which is the percentage of the time that the agent calls this tool to quote unquote finish a submission. The second is test pass rate. So, this is how we measure the regression in performance on sort of the base task. And this is the percentage of the time that the environment passes all tests accompanying the Sweebench task, irrespective of whether the agent submitted the task via this tool call. And then finally, we have our Sweebench task rate, or pass rate, which is how we basically combine these two metrics. It's the intersection of those two behaviors. And so, the goal is we want to raise the Sweebench pass rate performance while not degrading the test pass rate. And I've included a hint here of what this looks like in practice. It says something like you are near your 40 turn limit. There's only about three turns left. You often keep exploring and forget to wrap up investigating. So, finalize and verify your fix, and then call this tool before you run out of time. So, this is what the results look like. So, we were able to target the specific behavioral change, which was to call this tool when it wanted to submit a task, and without any degradation to the overall performance. So, you can see that sort of maroon color is the test pass rate. It's relatively constant. In fact, it goes up a little bit. But the task complete call rate increases dramatically from about 22% to 60%. So, we're able to add this behavior. And I think the really interesting, surprising result here is, again, this is in the fully sort of like offline world. And so, we're taking a production trace, or a trace that was created ahead of time, that basically never called this task complete tool. And we're nudging with the student and teacher models, the student model towards calling this task complete tool call, without ever specifically changing the tokens for the tool call. Because, again, the rollout is conditioned on the quote unquote production trace, right? And so, it never had the reasoning path to think to call the tool call. And so, the teacher doesn't force the tool call. It just starts to force the reasoning path towards the tool call, without ever actually changing the tool call. Which I think was really cool and surprising to us. Now, that being said, there actually is like a little bit of a cheat here that we can use. Which is that you can, and as I mentioned earlier, you can roll out just one step from the on policy model, given an offline production trace. And when we do this, we're obviously see that the student model sort of learns to wrap up its reasoning. And eventually, the teacher starts encouraging it to actually call this tool token. And so, you can see by having something that's a little bit more on policy, that we're able to increase sort of the SWE bench pass rate more than in the fully offline world. So, again, this is sort of offline trace, offline hint, with just one step on policy. Cool. So, then what does it look like in sort of the fully online, online trace, online hint world, where it's serving production traffic? So, for a certain production use case we had, we needed to teach a coding agent to use very specific formatting for hyperlinks due to a certain harness, certain harness nuance of one of our customers. And obviously, this coding agent needed to not regress on any of the base coding agent capabilities. Now, the problem here is that these hyperlink formats were very, very out of distribution for previously post-trained models. And so, when we tried things like adding a reward for specific hyperlink formatting, or even doing SFT on traces where we knew the hyperlink was correctly formatted, we saw that there was this sort of degradation in overall coding agent performance. And so, what we did here is set this up as an online trace with an online hint. So, basically, what this looked like is that we would do a rollout, then we would inject a hint specific to the rollout that occurred from the on policy model, and then say, in your prior rollout, you'd formatted hyperlinks like this. Next time, make sure to format hyperlinks in this way instead. And so, what we were able to see is that the percentage of correct hyperlink formatting jumped drastically from about, I guess, 15% all the way up to around 80%. And the other line, the other graph here shows what happens if we try to do offline hinting. So, this is basically for every single rollout, apply the same hint, which says, remember that when you do hyperlinks, you have to format it this way. And you can see that we do climb the behavior a little bit, but far less than in this online hinting world. So, here we've seen sort of like two different results, one where we can use offline hinting and offline traces to climb from production traces, and then another one where we're able to actually use sort of the on-policy-ness of the model and online hints to improve a behavior when we're serving production traffic. Okay, cool. I think I have enough time here to talk about a little bit of tips and tricks here. So, the first is that we found that per-step hinting is very, very important to making distillation work. So, rather than injecting a hint to the beginning of a rollout, we use a judge to essentially decide where in the rollout we should be injecting hints, and then actually have found that it's best to just do distillation on that next step that occurs, or maybe a few steps forward, rather than the entire rollout. Because that's really the turn and the moment in time that you want to have the teacher teach something to the student. You can also see in this graph here that this KL learning signal basically goes down as you get further and further away from the hint, which makes sense. Another trick that we've used is something called relevance mass self-distillation. And there's a blog post on our website about how we've done this. But essentially the idea is that we use an LLM judge to sample and choose which tokens we actually learn from, from our teacher model. Because often we'll see that the teacher model has preferences of certain connector words that are not really relevant to actually what we're trying to teach the student. And we can see in sort of the graph at the bottom that we're able to increase our ability to learn a very, very out of distribution behavior, while also being better about avoiding catastrophic degradation. Cool. So, overall, where does that leave us? So, obviously, I assume everyone here is sort of on the distillation train, but it's a very, very valuable tool towards continual learning. We introduced a spectrum of offline and online rollouts, as well as offline and online hinting, and how we use them towards distillation. So, we use offline hinting with offline production traces to provide value on day one to enterprise clients. Give us production traces and we can teach it a certain behavior. We then use online hinting and online production traces to do this highest ceiling sort of continuous learning improvement across multiple improvement areas, because that judge is able to adapt to whatever the online model does in production. And finally, I just want to say thank you to the team that worked on this. A lot of the work was done by others. I just kind of got to present it. And we're hiring and having a lot of fun working on research problems around continual learning. So, if you're interested, reach out to hiring at applycompute.com or, yeah, just email me as well. So, thank you, everyone. So, congratulations.