Open Reader

Reinforcement Learning without Verifiable Rewards — Will Brown, Prime Intellect

completed 19:26 Jul 31, 2026 Watch on YouTube

Current Status

completed

Video ID

AQv3qRCG6Gw

RAG / Chat

Enabled
Reinforcement Learning without Verifiable Rewards — Will Brown, Prime Intellect
Description

Reinforcement learning has been easy to sell where the answer is checkable, like math or code, and Will Brown's talk is about everything else. Most valuable tasks have no clean verifier, so Prime Intellect's work is on how you build reward signal when there is no ground truth waiting. He frames RL simply first, a model acting in a harness with tools and skills, getting a reward, and nudging its weights, then asks how you keep climbing once you leave the verifiable island behind. His answer leans on environments as the anchor. You can set up judges, generate question and answer pairs grounded in real documents and repos, and use a reverse direction trick where you hide something, like a bug or a backdoor, so the model can learn to find it again, which conveniently gives you a difficulty dial to keep tasks not too easy and not too hard. He is direct about the dangers: reward hacking will find you if you are not careful, so you inspect traces, run small experiments, and bring in expert understanding. The goal he keeps returning to is making this a real science, with open models and shared benchmarks, where environments turn into new tasks and higher levels of ability. Speaker info: - https://x.com/willccbb - https://www.linkedin.com/in/willcb/ - https://willcb.com Timestamps: 0:00 - RL without verifiable rewards 1:17 - How RL works, simply 2:43 - The tooling that powers it 4:13 - Where verifiable rewards run out 6:24 - Being careful about reward design 8:04 - Making RL a science 9:20 - Judges and grounded question answer pairs 10:46 - The reverse direction trick 14:19 - Calibrating difficulty 15:08 - Hunting for reward hacks 18:27 - Environments as the anchor

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: For real-world agent tasks without deterministic success criteria, the practical path to continual improvement is to turn production traces and existing artifacts into grounded tasks, controllable simulations, judge-derived rubrics, and iteratively refined rewards.
  • Why it matters: This is a concrete architecture for moving agent improvement beyond prompt edits and static benchmarks toward an observable production-to-evaluation-to-training feedback loop.
  • Best use: Use it as a design framework for building an agent-learning control plane: capture traces, derive task distributions and failure modes, evaluate with scalable judges, train in replayable simulations, and reserve human review for high-leverage judgments.

Executive Summary

Will Brown argues that reinforcement learning with verifiable rewards (RLVR) covers unusually clean domains such as math, code tests, and deterministic database state changes, but does not directly solve the messier work agents must do in production: research, report writing, user interaction, refunds, booking, and open-ended tool use. In these settings, the central problem is not merely selecting an RL algorithm; it is manufacturing reliable learning signal without hand-authoring an expensive benchmark and reward function for every task.

His proposed answer is to make environments the central object. An environment combines task, world, harness, and reward/evaluation logic, and can support more than RL: synthetic-data generation, supervised fine-tuning, on-policy distillation, prompt optimization, evaluation, and scientific iteration. Production traces, document corpora, code repositories, pull requests, and completed work artifacts become raw material for discovering the real task distribution rather than relying on assumed benchmark distributions.

The operational pattern is to work backward from known or reachable end states. For documents, generate grounded Q&A from source material and remove the initial retrieval path before asking a model to solve the task. For code, use real PR descriptions, diffs, and tests, then remove parts of the completed solution to create reconstructable training problems. For less controllable tools or web apps, build simulators from production data so the system can plant known answers, expose backend state, and create verifiable tasks even when production itself is not directly verifiable.

Brown also frames test-time compute as a core ingredient in reward and environment design. Multiple models can inspect completed traces, identify mistakes retrospectively, derive reusable evaluation rubrics, calibrate task difficulty, red-team for reward hacks, and improve simulators. Small RL runs are themselves diagnostic experiments: they reveal behavioral changes and reward exploits that static evaluation design misses. The intended end state is a continual-learning loop in which production failures become new tasks and humans review only the most consequential judgments.

Key Takeaways

  • Claim: The main bottleneck for applying RL to production agents is creating trustworthy signal for open-ended tasks, not choosing among policy-gradient variants. | Evidence: Brown characterizes GRPO, REINFORCE, and related methods as variants within the same policy-gradient goal of increasing reward, while noting that real tasks such as writing reports, handling refunds, booking flights, and buying products rarely have one clean, deterministic answer. | Implication: For agent systems, invest first in environment, trace, evaluator, and reward-design infrastructure; algorithm selection is secondary if the system cannot establish what good behavior looks like. | Caveat: RLVR remains highly effective in domains with deterministic checks, including numerical math answers, code tests or linters, and known end-state database changes.
  • Claim: Production traces should define the evolving task distribution when the real-world distribution is unknown upfront. | Evidence: He recommends collecting deployed-agent traces such as user prompts and orchestrator-to-subagent calls, then treating them as source material even before labels exist; document corpora and code repositories serve the same grounding role. | Implication: Instrument agents to retain replayable inputs, tool calls, state transitions, outputs, and outcomes; these traces become a strategic dataset for evaluation, failure analysis, simulation, and future training. | Caveat: Traces identify what occurs in production but do not by themselves provide correct labels or prove that a sampled task is solvable.
  • Claim: Working backward from completed artifacts creates supervision where direct verification of the original task is difficult. | Evidence: For document tasks, models can generate and cross-check questions grounded in sampled documents, then solve after the original search path is removed. For code, real PR diffs, descriptions, and test cases can be partially removed so the model must recover a known reachable end state. | Implication: Convert successful tickets, resolved incidents, completed workflows, approved documents, and code changes into counterfactual replay tasks rather than relying exclusively on manually authored benchmarks. | Caveat: The constructed task must still be validated for answerability and should not inadvertently leak the solution through retained context.
  • Claim: High-fidelity simulators can convert non-verifiable production tool-use tasks into controllable, verifiable RL environments. | Evidence: Brown describes simulating MCP tools, CLIs, web applications, and other systems whose production backend state cannot yet be controlled. By grounding simulators in production traces and iterating between simulator and real behavior, the team can control backend state, plant answers, and reverse-engineer tasks from known outcomes. | Implication: For valuable but hard-to-evaluate workflows, build a sandbox or digital twin around the critical state transitions; use production replay to validate it before allowing training improvements to influence live behavior. | Caveat: Simulator fidelity is a material risk: training can optimize behavior for a simulated world that diverges from production, so the simulator must be continuously checked and refined against real traces.
  • Claim: LLM judges are most useful as retrospective, compute-intensive analysts that turn trace review into reusable rubrics and targeted tasks. | Evidence: Brown says errors are often easier to identify after seeing the full chain of events. Multiple models can inspect traces, agree on whether behavior was wrong, search for failure patterns, and distill that analysis into cheaper rubric questions for later auditing and task generation. | Implication: Separate expensive offline adjudication from inexpensive online monitoring: use broad, multi-model analysis to discover failure modes, then operationalize only calibrated rubrics for routine gating. | Caveat: A judge should not be presumed correct simply because it is an LLM or because several models agree; high-impact rubrics still require human calibration and periodic review.
  • Claim: Reward hacking and poor task difficulty only become visible through iterative training experiments, so environment design must include train-and-observe loops. | Evidence: He warns that loose reward proxies create boundary exploits and recommends red teaming, adversarial prompt optimization, trace mining, and small RL runs that log tool-call patterns and judge behavioral changes. He also notes RL needs tasks with an advantage gap: neither too easy nor too hard, but with a meaningful gap between one rollout and a set of rollouts. | Implication: Treat every reward and environment as a hypothesis under experiment: run limited-scope training, compare trajectory distributions and exploit rates, then revise before scaling training or deployment. | Caveat: A behavior that appears to improve a scalar reward may still violate the intended spirit of the task; reward metrics must be inspected alongside trajectory-level evidence.
  • Claim: Continual learning for agents requires more than RL because RL refines skills but is weaker at incorporating dense new world knowledge. | Evidence: Referencing Echo and Prime Intellect's follow-up work, Brown argues for combining RL with supervised signal from the environment so a model develops a likelihood-based understanding of expected environment outputs—a more native world model—rather than only learning rewarded action patterns. | Implication: When agents repeatedly fail due to missing domain facts, system semantics, or tool behavior, add curated/replayed environment observations and supervised updates rather than expecting reward-only training to discover the knowledge reliably. | Caveat: The talk presents this as a direction and references external work, not as a fully specified production recipe with comparative performance data.

Detailed Brief

Environment-centered post-training stack

  • Claims: Brown treats environments and evaluations as the same reusable object rather than separate assets for training versus testing.; A single environment can serve synthetic-data generation, SFT, RL, on-policy distillation, prompt optimization, agent/harness iteration, and formal evaluation.; The broader objective is to let domain experts own optimized model weights for their specific workflows instead of depending exclusively on the training distribution of frontier base models.
  • Evidence: His environment definition is task plus world, where the world may include a Docker image, codebase, task-specific tools, skills, applications, browser tabs, scoring rules, verifiers, or rewards.; Prime Intellect's described stack spans GPU orchestration, its PrimeRL framework, environment/task/harness/verifier components, and a Lab platform for hosted training, evaluation, inference, experiment monitoring, and deployment.
  • Caveats: The talk is a conceptual and product-adjacent synthesis; it does not provide implementation cost estimates, benchmark numbers, or a detailed governance model for updating production model weights.; The transcript contains substantial repeated segments near the end, so the useful unique content is shorter than the reported 5,289-word count.
  • Implications: Architecture should make environments portable and versioned assets, with the same task specification reusable across offline evaluation, simulation, data generation, and training.; The durable internal asset is not just a model prompt or a benchmark score, but a versioned corpus of environments, traces, evaluators, rubrics, and replay data.

Human role and operating model

  • Claims: Automation should raise the abstraction level of reward and environment design rather than eliminate expert judgment.; The human should decide the highest-value normative questions—what behavior is acceptable, unacceptable, or strategically desired—while models and search handle broad trace analysis and refinement.
  • Evidence: Brown repeatedly describes using inference compute to mine real-world data, refine signals, and surface the most important cases upward to experts.; He compares this shift to coding agents moving humans to a higher level of abstraction.
  • Caveats: The talk does not specify escalation thresholds, reviewer sampling rates, approval requirements, or rollback mechanisms for an online learning system.
  • Implications: The operating model should explicitly define which judgments can be delegated to judges and which must remain human-approved, especially where failure has customer, financial, compliance, or security consequences.

Notable Concepts & Terms

  • RLVR (Reinforcement Learning with Verifiable Rewards): The clean case where success can be deterministically checked, such as unit tests, numerical answers, or known database state; Brown uses it as the contrast to real-world agent learning.
  • Environment: The central reusable object comprising a task, a world, an agent harness, and scoring or reward logic; it can power training, evaluation, simulation, and data generation.
  • Advantage gap: The performance separation between a single rollout and better outcomes found across multiple rollouts; tasks need an appropriate gap to generate useful RL learning signal.
  • Grounding: Anchoring generated tasks and evaluations in source material such as documents, repositories, or production traces so supervision is tied to real evidence rather than unconstrained model generation.
  • Reverse engineering from end states: Starting with a completed, known-reachable artifact or state, removing pieces of the solution, and training an agent to recover it; this creates verifiability for originally open-ended tasks.
  • World simulator: A controllable approximation of a tool, web application, or backend system that permits state control and planted answers, enabling verifiable training tasks when live systems are opaque.
  • Scaling judges / test-time compute: Using additional inference, multiple models, and search to inspect traces, identify errors, derive rubrics, calibrate difficulty, red-team reward functions, and improve simulations.
  • Echo: Referenced research direction combining reinforcement learning with supervised signal from environment outputs so models learn environment knowledge and not only rewarded behaviors.

Operator Notes / Why Ken Should Care

  • Create a trace-retention schema for every agent workflow: user intent, context, model version, prompts, tool calls, tool responses, state diffs, final output, human intervention, and post hoc outcome.
  • Select one high-value workflow with messy success criteria and build a replay environment from completed successful and failed cases before attempting online RL.
  • Require a simulator-fidelity evaluation against held-out production traces before using simulated performance as a deployment or training decision metric.
  • Establish a rubric lifecycle: multi-model offline discovery, expert calibration, versioned deployment, trajectory-level audit sampling, and retirement or revision when exploited.
  • Run small, isolated training experiments with behavioral telemetry and rollback gates; do not scale an updated policy based on aggregate reward alone.
  • Use supervised updates from environment observations for persistent knowledge gaps, while reserving RL for action sequencing, tool-use policy, and skill refinement.

Source/Metadata

  • Title: Reinforcement Learning without Verifiable Rewards — Will Brown, Prime Intellect
  • Transcript words: 5289
  • Duration seconds: 1166
  • Timestamp note: No timestamps or chapters were present in the supplied transcript; the transcript also includes repeated passages and a repeated closing artifact.

Transcript

3969 words en Processed in 151.4s

Thanks all for coming to AI Engineer and checking out the post-training session. Hopefully lots of fun stuff today and throughout the conference. I'm Will Brown, I lead applied research at Prime Intellect, and today I want to talk about reinforcement learning without verifiable rewards. Many people may have been learning about RLVR over the past year or so, year and a half, as this stuff has really taken off and become the main way that we think about scaling reinforcement learning. But often we don't actually have verifiable rewards. Messy real-world tasks, often we're figuring out as we go, we're having our agents run around and we, in hindsight, maybe can look at what they did and say, okay, this was good, this was bad. But sometimes sitting down and specifying, okay, this is the rule, this is the goal, is not always so straightforward. This is going to be synthesizing a lot of work we've been doing, as well as from the broader research literature and some of the things we're building to support extending RL into more messy real-world tasks. So, recap quickly of how reinforcement learning works. I would imagine if you're in the post-training session here, you've probably heard a little bit about RL, but for those of you at home and for those who are still looking for the crash course, generally we have an agent, which we're going to call a model plus a harness, which we place into an environment. An environment, we're going to call a task, plus a world. A world you could think of as, maybe it's a Docker image, maybe it's a code base, maybe it is a collection of task-specific tools, maybe it's some skills, maybe it is a bunch of applications or browser tabs or things like this, as well as a scoring rule, verifiers, or rewards, whatever you want to call them. The agent and the environment are going to interact in a loop. At the end, we'll have some reward of how well the agent did in the environment for this task. Then reinforcement learning is all about creating an advantage. The advantage is really about taking the reward, minusing some baseline, maybe doing some scaling. Now you have a set of rollouts from the agent in the environment that you can use to then update the policy. The policy here is just the model weights themselves. The goal here is to take a gradient, which nudges the model towards getting higher reward. So all of the RL stuff people talk about, whether it's GRPO or reinforce or Cisco or any of the other new algorithms people come up with, they're all in this policy gradient framework, which is just about saying, okay, how do I make the model do things that have higher reward? At Prime and Elect, we build a lot of tooling to power all of this at every layer. We start at the compute layer and do lots of large-scale GPU orchestration. We build the PrimeRL training framework, which powers all of our large-scale reinforcement learning and other algorithms running. We build environments. We have task sets and harnesses and verifiers as tools that you can mix and match to assemble complex worlds for agents to learn from real-world feedback. We have a training platform called Lab, which is anchored around environments, where we do both hosted training and evaluations, as well as inference. This is to allow people to monitor their experiments and manage the training runs and iterate on their evals and deploy these models. Ultimately, the models are starting generally from some open-source base model, and you're optimizing it for your task. The goal that we're really trying to enable is for more people to be able to become their own research lab, take ownership over the intelligence of their own model weights, and optimize for the tasks that they care about with themselves as the experts steering the model, which means we need to make it way easier for people to do this. Currently, for a lot of people, it's still really hard. I think you can go, we're all here learning more about how it works because it's hard. We don't know how it all works, and we're figuring it out as we go in many cases, but we've spent a lot of effort and a lot of time building stuff that hopefully makes this a bit easier for people. So what's an environment? An environment is tasks, harness, and rewards, but it's not just for RL. I think a lot of people think RL when they think environment, but environments and evals are really the same thing. You can use these same objects for generating synthetic data, which then you could use for SFT. You can do RL, or you could do algorithms like on-policy distillation. You could do prompt optimization like JEPA. You can use it as a scientific test bed to iterate on your agents and your harnesses. Verifiable rewards are the easy case, where we can check exactly whether something was done correctly or not. For math, often if you have a numerical answer, you can just parse this out of a box in the answer from the model and check. For code, maybe you want to use test cases or a linter or something like this. For tool use, often you have some database state, which you know what you're expecting at the end, and you can check this deterministically. These are the easy cases where the reward design problem is not so difficult. But most real-world tasks are not this verifiable. For a lot of real agent work tasks, we're having agents do things like write reports that maybe are analyzing a bunch of documents or research. Or maybe ask it to do things like book flights or buy things, but there isn't always a clean best answer here. And there are also notions that are fuzzier, of interacting with users, like handling a refund. What does it mean to handle this well? Here the signal is less clear. And there's a lot of different tricks we might want to explore and techniques we want to develop to ensure that this can be done reliably and scalably. Making evals is hard because oftentimes the benchmarks out there that we might look at in the new model releases, it's a set of a few hundred tasks that a bunch of researchers spent months handcrafting and talking to experts. Maybe they worked with data vendors and spent lots and lots of money getting these to be very precisely refined. This isn't very scalable out of the box, especially for things that are more open-ended, where there's no clean check for what's good or not. Often the real-world situations can be unbounded. You don't always know what things are going to be in the distribution. A lot of these cases, we are figuring out the distribution as we go, and classical machine learning will tell you, you can train for the distribution, but generalizing outside of the distribution is an undefined problem. Then, especially with RL, we have to be very careful about reward hacking. Reward hacking is when you have a loose proxy for your objective that is undefined at the boundaries, and then models, if you train with RL, they can learn to exploit this and find weaknesses where there's some path towards climbing the reward that doesn't actually give you what you want. So really the goal of what we would hope all of this builds into is continual learning, which is a big buzzword that I think a lot of people talk about in many different ways, but I'm going to use it to mean a very particular thing, which is that we want models to be deployed in relatively realistic, complex, messy settings, and to be able to learn as they go, where they are doing things, they are making mistakes, they are then able to observe and catch these mistakes after they happen and use this to not do the same thing again. In some cases, people want to try to do this at the harness layer or the prompt layer, but ultimately you want a system that can evolve autonomously to be able to get better over time with humans in the loop at the right level of abstraction. I think currently the level of abstraction for doing this is far too low for it to be practical for most people. This means we need new methods to be able to automatize as much of the difficult processes as possible. Many of these actually are automatizable, they just are difficult problems to solve. There's a few techniques you can use to start making progress here. But one of the goals here is to do online reinforcement learning so that you can iterate on this process as you go. As well as beyond just RL, there are other things where you might want to incorporate world knowledge into the model itself, not just in terms of skill refinement. RL's great for refining skills, but less so for incorporating dense new knowledge. And so blending these two together is also an important goal. Ultimately what we want from this is to be able to deploy agents into production and have them improve as they go, and have these experiments be monitorable and traceable and replayable so that we can treat model optimization very much as a science and make this science accessible to people who have a very wide variety of use cases they want to deploy agents for, which don't all live in the training distributions of the big models. And so we want to be able to do this on top of the best to start making progress here. But one of the goals here is to do online reinforcement learning so that you can iterate on this process as you go. As well as beyond just RL, there are other things where you might want to incorporate world knowledge into the model itself, not just in terms of skill refinement. RL's great for refining skills, but less so for incorporating dense new knowledge. And so blending these two together is also an important goal. And ultimately, what we want from this is to be able to deploy agents into production and have them improve as they go, and have these experiments be monitorable and traceable and replayable so that we can treat model optimization very much as a science and make this science accessible to people who have a very wide variety of use cases they want to deploy agents for, which don't all live in the training distributions of the big models. And so we want to be able to do this on top of the best and biggest open models in the world and make this accessible. And so how do you manufacture signal? There's a bunch of techniques that we found very useful. One is grounding. And so grounding roughly means that you have some source material, and in machine learning generally, you want to have some notion of supervision. There's something you're learning from. And in messy situations, we don't necessarily always have clean supervision, but we can get pretty reliable supervision if we are careful about the techniques we use. And so grounding is one where you have some source material, and the ability to do an A-B test of with and without is a very useful way of creating this capability gap where a model will do better if it has something in context. And this gap is something we can exploit to create signal that we can then learn from. Judges are also really useful. We're relying on the fact that LLMs are already really powerful general reasoners for many things. And if we set these judges up in the right way, then they can spend compute to make decisions about whether an action was good or bad. We also want to be scaling search. So search, in many ways, is something we can apply at many different layers of the pipeline, both in terms of creating tasks, as well as the worlds, as well as the criteria for which we want to be giving judges for answering questions about the quality of a rollout. And so for source material, one very useful version of this, especially for this continual learning goal, is production traces themselves. And so what we found is super helpful is taking existing traces from a deployed agent and treating these as the source material where we don't necessarily know upfront what the distribution of tasks is, but as an agent is deployed, you start collecting more and more examples of, let's say, user prompts or calls from an orchestrator agent down into a sub-agent. And this starts becoming the distribution. We don't have labels yet, but it tells us at least what we want to look for. And so this is one very useful category, especially for other things like search or for code. You have corpora documents, you have repos that are also very useful for anchoring your learning as well. And so taking these sources, these raw materials, as places to search for tasks from is a very useful way of starting to create this environment out of nothing. Well, it's not nothing. It's something from the real world. And because you have the real world, you want to use that, the real world, your production environment, your agent traces, as the source from which you want to learn, even if you don't have supervision yet. And so one thing you have to do here is get tasks. And so documents are actually a pretty easy version of this where you can just sample documents, you can have models generate question-answer pairs grounded in the documents. You can verify that those questions are answerable with other models that are still grounded in the documents. And then the actual task at hand involves throwing away the initial search. And so you get to work backwards. And so this general principle of working backwards is starting from the solution or something that's close to the solution and where your real task is further upstream. This is a very useful way of having a, you can verify the easy problem and then learn on the hard problem. And so anything where you can move backwards like this is super useful for getting supervision for free. In code, you can use real-world PRs, the diffs, the descriptions, the test cases, removing different pieces of different files to be able to have models start learning over code bases because you can take something that is a completed artifact and start breaking it down into smaller pieces. And then have replaying these pieces of getting to an end state that you know is reachable be a task that you train on. And so this idea of wanting to know that an end state is reachable, and that you can then take steps back, throw away the solution, and then learn to find it again, can be applied more generally beyond code as well. And so we talk about world simulators broadly as the sort of thing we might want to do in messier environments, which are not just production, which are not just doc search or code. And so a lot of the ones that we've been working on at Prime Intellect are related to things like tool use and web applications where we don't actually have full controllability of the back-end state. There are some MCP tools or CLI tools or websites or applications where we can't actually program them yet. And so what we want to do is learn to simulate them. And so we found that using combinations of universal back-end infrastructure and test-time scaling and search and iterating between the simulator and the real, and the real we can ground in these production traces, this then allows us to create really high-fidelity simulators. And so we found that these simulators are actually really great for RL because you can, one, if you have production data, make your simulator better and better over time, but also you have full controllability over the back end. And so you can actually do this reverse engineering where you get to plant the answer. You can start from the end and work backwards. And so that you have this verifiability baked into the simulator, even if you don't have it in the real-world production deployment, because you don't know in advance if a task was solvable at the time that you are being asked it. And in terms of doing this, a very useful thing is scaling judges. And so a lot of times we will have a model that does something, and it will make mistakes along the way. And it's easier to tell what went wrong in hindsight. And so the fact that you've already seen the chain of events after, and you can look backwards and say, okay, the model made a mistake here. This thing doesn't feel quite right. Or we asked seven different models and they all agree this thing is wrong. This is a very useful way of spending compute to do search, to then extract rubrics. These rubric questions are very effective at distilling down the search into something that we can then use to more cheaply audit and also ground once we have these rubrics as a way of saying, okay, we need tasks that target these kinds of failure modes as well. All of this is under the umbrella of scaling search with test-time compute. And so we can scale search for mining traces by, if we have offline production traces, we can just look at them more and think about them more and have more models play with them. We can calibrate difficulty. So RL, to have the advantage gap that we mentioned, needs to have a separation between what one model will do once and what a collection of rollouts will do. And so you want tasks that are not too easy, not too hard, and you want to be searching for these and iterating on generating more of them. And so this is another area where you can spend compute to refine the difficulty of your task distributions, your task sets. For simulators, when you're building web applications or tools that need to simulate complex behavior, you can spend compute on searching these and then you can have agents refine the implementations of them. And this allows for increasingly high-fidelity environments. You can do this on verification, both at train time as well as offline when you're creating these rubrics. You can do things like red teaming with adversarial prompt optimization to explore for backdoors. Then you can look for traces and spend compute mining these traces for understanding: was this a reward hack, or was this actually in the spirit of the task? And I think these things can feel like reward hacking can sneak up on you if you're not careful, but in many cases, the basic simple things actually work quite well, where if the reward hacks are the sorts of things where a human can look at them and be like, oh yeah, that's a reward hack, judges are often quite good at doing this as well. If you tell the model not to do this, it won't necessarily do it in the rollout, but in hindsight, you can reflect on this and spend compute, especially if you're collecting these over time and you are building up your corpus of examples of reward hacks. You can understand the sorts of things that go wrong and address this by, again, spending inference compute on refining your implementation. a reward hack, or was this actually in the spirit of the task? And I think these things can feel like reward hacking can sneak up on you if you're not careful for it, but in many cases, the basic simple things actually work quite well, where if the reward hacks are the sorts of things where a human can look at them and be like, oh yeah, that's a reward hack. Judges are often quite good at doing this as well. They just don't necessarily, if you tell the model not to do this, it won't necessarily do it in the rollout, but in hindsight, you can reflect on this and spend compute to, especially if you're collecting these over time and you are building up your corpus of examples of reward hacks, you can understand the sorts of things that go wrong and address this by, again, spending inference compute on refining your implementation, refining your rewards, as well as validating these by training. And so we find that in many cases, you can do a lot up front, but also there are things that don't show up until you actually start doing RL. And so part of this is folding training experiments themselves into the process of environment design, where you can do small runs with individual models on one environment and see what happens. And you can understand the behavior changes. You can have metrics that log the types of tool calls that are being done, that are judges asking questions about the traces to understand how behavioral patterns are changing. And all of these are very useful ways of getting something from nothing and using compute as the thing that allows you to refine your understanding. And ultimately, what you want is to surface the most important pieces up to the human, the highest level of the human being, the highest level of the human being, the highest level of the human being. So that all of this is deferring to the human for the most important pieces of actually employing expert judgment to say, this is good, this is bad, this is what I want, this is not what I want. And these are all the ways that we gain confidence in the environments. I know we're running a little short on time, so I wanted to recap a couple blogs that we've put out recently that are demonstrating pieces of this. We have a blog called General Agent, which is demonstrating this for tool use, this online loop of generating, solving, and synthesizing new tasks and gating based on this pass rate, which then we train on, and we see a great uplift on popular benchmarks for tool use. Additionally, beyond just RL, we found that it's quite important to think about cases where there's information in the world that RL alone will not explore. And so there's this great work, Echo, from some researchers that we are friends with and have been collaborating with. And then we did our own deep dive into this as well to look at what happens when you have an agent that is not just training with reinforcement learning, but is also getting supervised learning signal from the environment itself. Which then allows the model to understand things like having a native world model of the environment, understanding what to expect because it has a likelihood model of the tokens that the environment itself will generate. And these are the sorts of things that, in many cases, allow the model itself to more adeptly navigate the world, and not just refine its skill, but get new information into its weights over time as well. And so all of this is in spirit of making post-training easier, making continual learning easier, giving people the ability to create agents and not worry too much about having to fuss with the research pieces and fine-tune all the small details. Today we still do, but we're seeing paths forward of how we start automating this more and more by having environments as the anchor, which we can then spend compute on refining. We can use compute to mine the data we have from the real world to refine the signals, where the humans are just, in the same way that with coding agents, we're going to higher levels of abstraction. We can do this with environment and reward design as well. And all of this is what allows us to ultimately close the loop where models are then able to stay within the guardrails we give them, they go find the issues in production, and then they turn these back into new tasks that can then be trained on for getting better in the real world. We work hands-on with startup enterprises to help them train their models. We are also hiring quite a lot. If you want to get in touch for either of these, find me after the talk. Thanks a bunch. If you can find me, find me. If you can find me, find me. If you can find me, find me. If you can find me, find me, find me. and find me, find me, find me. between what one model will do once and what a collection of rollouts will do. And so you want tasks that are not too easy, not too hard, and you want to be searching for these and iterating on generating more of them. And so this is another area where you can spend compute to refine the difficulty of your task distributions, your task sets. For simulators, when you're building web applications or tools that need to simulate complex behavior, you can spend compute on searching these and then you can have agents refine the implementations of them. And this allows for increasingly high fidelity environments. You can do this on verification, both at train time, as well as offline when you're kind of creating these rubrics. You can do things like red teaming with adversarial prompt optimization to kind of explore for backdoors. Then you can look for traces and spend compute mining these traces for understanding was this a reward hack or was this actually kind of in the spirit of the task. And I think these things can kind of feel like reward hacking can kind of sneak up on you if you're not careful for it, but in many cases, the basic simple things actually work quite well, where if the reward hacks are the sorts of things where a human can look at them and be like, oh yeah, that's a reward hack. Judges are often quite good at doing this as well. They just don't necessarily, if you tell the model not to do this, it won't necessarily do it in the rollout, but in hindsight, you can reflect on this and spend compute to kind of, especially if you're collecting these over time and you are building up your corpus of examples of reward hacks, you can understand the sorts of things that go wrong and address this by kind of, again, spending inference compute on refining your implementation, refining your rewards, as well as validating these by training. And so we find that in many cases, you can do a lot up front, but also there are things that don't show up until you actually start doing RL. And so part of this is folding in training experiments themselves into the process of environment design, where you can do small runs with individual models on one environment and see what happens. And you can understand the behavior changes. You can have metrics that log the types of tool calls that are being done that are like judges asking questions about the traces to understand how behavioral patterns are changing. And all of these are very useful ways of kind of getting something from nothing and using compute as the thing that allows you to refine your understanding. And ultimately what you want is to surface the most important pieces up to the human, the highest level of the human being, the highest level of the human being, the highest level of the human being. So that all of this is deferring to the human for the most important pieces of like actually employing expert judgment to say, this is good, this is bad, this is what I want, this is not what I want. And these are all the ways that we kind of gain confidence in the environments. I know we're running a little short on time, so I wanted to recap a couple blogs that we've put out recently that are kind of demonstrating pieces of this. We have a blog called General Agent, which is demonstrating this for tool use, this online loop of generating, solving, and synthesizing new tasks and gating based on this pass rate, which then we train on and we see a great uplift on popular benchmarks for tool use. Additionally, beyond just RL, we found that it's quite important to think about cases where there's information in the world that RL alone will not explore. And so there's this great work echo from some researchers that we are friends with and have been collaborating with. And then we did our own kind of deep dive into this as well to look at what happens when you have an agent that is not just training with reinforcement learning, but is also getting supervised learning signal from the environment itself. Which then allows the model to understand things like having a native world model of the environment, understanding what to expect because it has a likelihood model of the tokens that the environment itself will generate. And these are the sorts of things that in many cases allow the model itself to kind of more adeptly navigate the world, and not just like refine its skill, but like get new information into its weights over time as well. And so all of this is in spirit of making post training easier, making continual learning easier, giving people the ability to create agents and not worry too much about having to fuss with the research pieces and fine tune all the small details. Today we still do, but we're kind of seeing paths forward of how we start automating this more and more by like having environments as the anchor, which we can then spend compute on refining. We can kind of use compute to mine the data we have from the real world to refine the signals where the humans are kind of just, in the same way that with coding agents, we're kind of going to higher levels of abstraction. We can do this with environment and reward design as well. And all of this is what allows us to ultimately close the loop where models are then able to stay within the guardrails we give them, they go find the issues in production, and then they turn these back into new tasks that can then be trained on for getting better in the real world. We work hands-on with startup enterprises to help them train their models. We are also hiring quite a lot. If you want to get in touch for either of these, find me after the talk. Thanks a bunch. If you can find me, find me. If you can find me, find me. If you can find me, find me. If you can find me, find me, find me. and find me, find me, find me.