Rethinking Environments for Long-Horizon Work — Rayan Garg, Theta Software
Description
Everyone wants agents that handle long horizon work, but Rayan Garg starts with the awkward question of what long horizon even means. One popular answer measures the time horizon as the task length at which an agent crosses a success threshold, like the sixteen hour mark, which is a useful endpoint but a noisy one, since human time estimates vary and the same wall clock hides very different amounts of real difficulty. How you choose to measure this has an outsized effect on what you conclude about a model. From there Theta Software's work is about designing the environments and verifiers that make those measurements honest. A task can be artificially stretched by forcing serial dependencies, or made genuinely hard when a bad early query cascades through everything after it, and as environments grow more complex, standardized evaluation gets harder and correctness is best verified from the final state rather than a judge's guess. Garg walks through collapsing a huge state space with sample trajectories, being careful that judges do not see information they should not, and reusing agents to sift artifacts like CI logs. The recurring principle is that long horizon progress lives or dies on environment and verifier design, not on the headline benchmark number. Speaker info: - https://x.com/RayanGarg - https://www.linkedin.com/in/rayan-garg/ Timestamps: 0:00 - What does long horizon mean? 1:13 - Time horizon and the threshold metric 3:17 - Why the metric is noisy 4:20 - Measuring what actually matters 6:38 - Creating tasks and environments 7:42 - When a bad early step cascades 10:01 - Why standardized evaluation is hard 11:17 - Verifying from the final state 13:46 - Judges, tools, and reused agents 17:45 - Rubrics, QA, and careful grading
Summary
Generated by gpt-5.6-terraAt-a-Glance
- Verdict: Watch fully
- Core thesis: Long-horizon agent capability should be evaluated in stateful, multi-tool, ambiguous environments, with agentic judges that inspect both final world state and trajectory rather than relying on simplistic task duration or deterministic pass/fail checks.
- Why it matters: This is directly relevant to designing reliable agent control planes and evaluation loops: weak environments and weak verifiers can make agents appear capable while rewarding brittle shortcuts, benchmark gaming, or work that does not actually produce the intended operational outcome.
- Best use: Use it as a design review framework for building or buying long-running agent environments, especially for defining stateful tasks, instrumenting trajectories, and implementing read-only judge agents with robust rubric QA.
Executive Summary
Theta argues that “long horizon” is not a binary property and cannot be represented cleanly by one metric. Human-time estimates, such as METR-style thresholds where a model succeeds 50% of the time on tasks estimated to take humans 16 hours, are useful but noisy because agents and humans solve work differently. Token count, steps, and tool calls are also informative but model- and harness-dependent. The practical recommendation is to use multiple measurements rather than treat any single duration statistic as capability proof.
The more meaningful measure of a long-horizon environment is not simply how many steps it contains, but whether it requires interdependent decisions across tools and changing external state. A task made long by concatenating independent subtasks is less diagnostic than one in which an early log query, diagnosis, or action changes the conditions and consequences of later work. Theta emphasizes multi-tool coordination, state transitions, and ambiguous starting information as the ingredients that better approximate real operational work.
The presentation’s central implementation argument is that verification becomes the hard problem once agents move beyond math and testable coding tasks. For open-ended software and business workflows, deterministic checkers are often insufficient. Theta recommends an agentic judge that can inspect the environment’s final state, query a structured record of the agent’s trajectory, use rubrics to assign granular credit, and detect reward hacking or invalid paths such as sandbox escape and hidden-test access.
The speakers also criticize current finance-agent benchmarks as often too short, narrow, saturated, and weakly supervised to establish real long-horizon competence. Their alternative is deeper task environments with detailed reward models, broad domain coverage, and judge QA. The useful takeaway for Ken is less the vendor comparison and more the operating principle: evaluation quality, permissions, state observability, and reward design will determine whether long-running agents are genuinely reliable.
Key Takeaways
- Claim: Long-horizon capability should be treated as a relative, multidimensional measure rather than a binary label or a single task-duration number. | Evidence: The speakers contrast METR’s human-time framing—e.g., a 50% success threshold on tasks estimated to take a human 16 hours—with model-centric measures such as tokens, steps, and tool calls. They note that a task taking 500,000 tokens for one model or harness may not have comparable difficulty for another. | Implication: Ken should avoid using one benchmark score, autonomous-runtime claim, or token count as a go/no-go measure for agent readiness; track outcome success alongside trajectory length, tool/state complexity, and operational reliability. | Caveat: Human-time estimates become especially noisy for expert work, where task duration depends heavily on whether the reference worker is in the top 10%, 1%, or 0.1% of the relevant field.
- Claim: The important distinction in long-horizon work is sequential, stateful complexity—not merely a large number of steps or subtasks. | Evidence: Theta contrasts a codebase-analysis task that can be parallelized across subagents with an incident-style workflow where an incorrect early Grafana, GitHub, or CloudWatch interpretation cascades into later actions. They describe meaningful environments as requiring agents to move information across tools and respond to changing state. | Implication: When designing agent workflows or evals, make earlier choices affect later available information, system state, constraints, or remediation paths; otherwise apparent long-horizon success may only reflect decomposition and parallel execution. | Caveat: Artificially chaining unrelated independent tasks can create a long trajectory without testing the decision dependence that makes real long-horizon work difficult.
- Claim: Ambiguity is necessary for realistic capability measurement, but it makes standardized evaluation substantially harder. | Evidence: The speakers define ambiguity as incomplete initial instructions and artifacts, arguing that real workers must explore rather than receive complete specifications. They note that ambiguous tasks create many valid solution paths and therefore many ways an agent can be correct. | Implication: Ken should distinguish between production constraints that need explicit policy boundaries and task information that can safely remain incomplete; for the latter, evaluate outcomes and principles rather than exact action sequences. | Caveat: A verifier that compares agents against a single reference answer or reference trajectory will wrongly reject valid alternatives and can collapse exploration into one prescribed path.
- Claim: For economically valuable, open-ended workflows, the verifier must usually be an agentic judge rather than a single deterministic checker or basic LLM review call. | Evidence: Theta argues that deterministic verification worked well for domains such as math and data-structure coding, but cannot reliably evaluate whether complex software or operational work changed a real environment correctly. Their judge evaluates both final environment state and the path used to reach it. | Implication: For OpenClaw-style operational agents, build evaluation as a first-class service with environment access, state checks, trajectory analysis, and calibrated scoring—not as a post-hoc prompt asking whether an agent did well. | Caveat: Judges themselves can be inconsistent, particularly when rubrics are overly dense or tasks are beyond current model capability; rubric and judge QA are therefore part of the system, not an afterthought.
- Claim: A judge needs access to the same operational evidence as the worker agent, but with protections that prevent the judge from altering the environment. | Evidence: In Theta’s deployment-failure example, the worker reads GitHub CI/CD and CloudWatch logs, changes code, opens a PR, and triggers redeployment. The judge must inspect post-deployment GitHub or AWS evidence to verify the fix, since tool-call logs alone do not establish correctness. | Implication: Separate worker and evaluator credentials. Give evaluators broad observability, immutable audit access, and state-query capabilities, while explicitly denying write, deploy, approval, and remediation privileges. | Caveat: The judge should have read-only permissions or equivalent safeguards so that it cannot accidentally trigger deployments or mutate the environment it is evaluating.
- Claim: Long trajectories must be made queryable; passing an entire trajectory into one judge context window is not a scalable verification design. | Evidence: Theta proposes storing trajectories in a database, enriching them with metadata, using subagents, and segmenting phases such as log investigation, code modification, and post-change validation. This allows a judge to locate critical failure points instead of reviewing an undifferentiated transcript. | Implication: Instrument agent runs as structured event streams with tool calls, state snapshots, artifacts, phase labels, and decision provenance, so evaluators and operators can retrieve the small set of events needed to establish correctness.
- Claim: Many current long-horizon finance benchmarks may overstate readiness because their tasks are short, narrow, already partly saturated, and insufficiently granular in reward design. | Evidence: Theta cites GDP Val, Banker Toolbench, and Apex Agents, arguing that their average human-task durations fall below the long-horizon thresholds being discussed. It says Apex Agents’ investment-banking section has a 57% pass@1 rate for fully solved cases, while GDP Val focuses narrowly on Excel tasks and Apex largely on investment banking. Theta reports its own finance sample averages 15 human hours per task across 50 tasks and says models still struggle. | Implication: Before relying on an agent benchmark for procurement, model routing, or investment conclusions, inspect task breadth, statefulness, saturation, evaluator design, and the reward granularity—not just headline pass rates. | Caveat: These are Theta’s comparative claims about competing benchmarks and should be treated as a vendor’s assessment; the transcript does not provide the underlying benchmark methodology or model-score table.
Detailed Brief
Verifier and rubric design patterns
- Claims: Deterministic verifiers remain useful as components of a hybrid evaluation system rather than being replaced wholesale by model judges.; The density of a reward signal is a key learnability variable, but increasing rubric detail indefinitely can reduce consistency if the judge cannot apply the rubric reliably.; Partial-credit methods can evaluate downstream reasoning even when an earlier assumption or sub-answer was wrong.
- Evidence: Theta describes deterministic checks as possible artifact or metric generators for a judge to review.; It describes dynamic evaluation-time rubrics that effectively grade later work conditional on treating an earlier answer as correct, analogous to awarding follow-on credit on a test.; For rubric QA, the speakers name gold-case, no-op, and variant tests, and add coverage and expert-agreement testing as tasks become longer and AI increasingly helps generate or assess rubrics.
- Caveats: A highly granular rubric is not automatically better: if its judge cannot apply criteria consistently, additional reward density can create noisy training signals and waste compute.; The talk identifies the QA categories but does not provide implementation thresholds, inter-rater agreement targets, or a concrete calibration procedure.
- Implications: Treat deterministic checks as evidence-producing sensors within a broader evaluator, not as the only source of truth.; Establish a rubric test suite before using evaluator scores for model training, production promotion, or automated remediation authority.
What environment complexity should capture
- Claims: Tool count alone is insufficient: complexity comes from coordination across external dependencies and from the degree to which actions mutate or reveal consequential state.; A task environment should model the exploratory nature of real work, where instructions and artifacts are incomplete and the agent must decide what evidence to seek.
- Evidence: The tool examples named are Grafana for observability, GitHub for CI/CD, AWS CloudWatch, databases, and code repositories.; The presentation repeatedly uses a production deployment failure as its representative stateful task: investigate logs, identify cause, edit code, open a PR, redeploy, and inspect whether the system actually works.
- Caveats: Ambiguous environments increase the set of valid trajectories, raising the cost and difficulty of evaluation.; The speaker does not distinguish which kinds of production state changes should be safely simulated versus permitted in a live environment.
- Implications: Prioritize sandboxed but realistic environments that expose the actual cross-system evidence an operator would use, rather than toy tasks that only mimic tool syntax.; Design escalation boundaries separately from evaluation realism: an agent can be evaluated on a full remediation chain without granting unrestricted production mutation rights.
Notable Concepts & Terms
- Human horizon / METR-style task horizon: A capability framing based on the estimated time a human needs to complete a task at a specified model success threshold; useful as one comparative metric but not a complete definition of difficulty.
- Sequential complexity: Difficulty created when early actions and interpretations change the validity or consequences of later decisions; presented as more meaningful than simply accumulating many independent steps.
- Stateful environment: An environment in which agent actions affect external systems or the information available later, such as repositories, logs, deployments, and databases.
- Judge model / critic model: An evaluator agent used when deterministic verification cannot establish whether an open-ended task was completed correctly or safely.
- Trajectory verification: Reviewing the process an agent used, not just its final output, to detect invalid shortcuts, unsafe actions, hidden-information access, and reward hacking.
- Queryable trajectory: A structured, stored, enriched record of an agent run that a judge can search by phase, event, metadata, or failure point instead of receiving one massive raw transcript.
- Dynamic evaluation-time rubric: A partial-credit approach that conditionally accepts an earlier assumption to assess whether subsequent reasoning or execution was internally correct.
- Reward density: How much granular feedback a rubric provides during evaluation or training; valuable for learnability but risky if the judge applies dense criteria unreliably.
Operator Notes / Why Ken Should Care
- Create a standard agent-run schema that logs tool inputs and outputs, artifact versions, state snapshots, permissions used, phase boundaries, approvals, and final outcome evidence.
- Implement a separate evaluator identity with read-only access to the same logs, repositories, deployment status, and databases visible to worker agents; prohibit evaluator-side writes and deployment triggers.
- Add adversarial verifier tests covering no-op behavior, superficially plausible but ineffective fixes, hidden-information attempts, sandbox escape, unauthorized tool use, and alternate valid solution paths.
- For any long-running workflow scorecard, report stateful dependency count, irreversible or consequential state changes, outcome success, and verifier confidence alongside runtime and token consumption.
- Do not use a single golden trajectory as the acceptance criterion for open-ended workflows; define invariant-based outcome checks and safety constraints instead.
- Audit external benchmark claims before using them for routing or investment decisions: require evidence on task duration methodology, domain breadth, saturation, reward rubric granularity, and evaluator access to final state.
Source/Metadata
- Title: Rethinking Environments for Long-Horizon Work — Rayan Garg, Theta Software
- Transcript words: 7281
- Duration seconds: 1274
- Timestamp note: No timestamps or chapters were present in the supplied transcript. The latter portion substantially repeats earlier material, so the effective unique content is shorter than the stated word count.
Transcript
It's great to see all of you here today. We're super excited to talk about one of our favorite topics here at Theta. Before we get started, we just want to introduce ourselves. So, hi, I'm a co-founder and CTO at Theta Software. Hi, I'm Ryan. I'm a co-founder and CEO at Theta Software. Prior to this, I was previously a founding engineer at Deep Silicon, where we did research into ternary models. Awesome. So, I can get us started with the topic today. We're going to be talking about aural environments within the context of long horizon tasks. And I think the most important thing for us to start with at the beginning is to talk about the trends and what long horizon actually means. We all know that the horizon at which AI agents can work autonomously is accelerating really fast. These are just some of the metrics that you can look at to see how this progress is really accelerating. But I think it's really important to actually define what the time horizon here actually means. We've gotten this data from one of the most common benchmarks out there that you've probably heard of for time horizons, if you've seen it in your Twitter feed all the time. It comes from Meter. And Meter has one response or answer to this really important question about how we actually define long horizon. It's really important to understand because what we considered long horizon a year ago probably isn't long horizon in our definition today. And what's long horizon today probably won't be long horizon in a year or two. And I think that gets to our first point, which is that long horizon is really a scalar metric. It's useful for measuring relative tasks. One task might be more long horizon than another. But it's really hard to define it into a binary category of this task is long horizon, this task is not, especially as the scope changes over time. So I think the first way we can talk about defining this is how Meter looks at it, which is human horizon, meaning can we use humans as a benchmark of, oh, this task takes humans a certain amount of time. So if AI agents can do that, then they've reached this certain critical level of a time horizon. And the way Meter does it is they have thresholds for the tasks they care about. So they have a 50% threshold, meaning if a certain model reaches a 16-hour threshold on this benchmark, that means it can achieve tasks with a 50% success rate that take a human 16 hours. And there's a really rigorous methodology of how they actually measure how it took a human 16 hours, but we will avoid some of those details. The other way we usually think about what long horizon actually means is not with reference to humans, but instead with reference to models. So some of the relevant model units we usually care about are things like tokens, how many tokens are consumed in a trajectory, how many steps to take, how many tool calls it takes. And these can be really noisy, right? Because I'm sure you guys have used different models. A lot of the Codex models are seen as more token-efficient than some of the Claude models. And it's a pretty noisy estimate for a couple of reasons. One is which model you're using, like I just said, and different harnesses you care about have a pretty big impact on how many tokens are actually consumed on a task. So this can be pretty hard to interpret when you're not holding variables constant. If a task takes a GPT model 500,000 tokens, that doesn't really tell you a lot about what that task would look like for Claude models until you actually run it on those Claude models. But despite it being a pretty noisy metric, it's actually really useful and important for us to understand because the amount of tokens that are consumed tells us a lot about how difficult a task actually is for an AI agent to tackle autonomously. You have to deal with things like compaction over long horizons. They don't really stay coherent over enough steps or trajectory length that you achieve. So even though it's a noisy metric, it can be really useful when, if we look at what a GPT 5.5 model can do now, and then you use the same model generation and see, oh, now it can actually achieve a million-trajectory based on an increased context window or improved compaction endpoint. That tells us a lot about how autonomous AI agents can actually go for long periods of time in that sense. And it really defines for us what the technical frontier actually means for models right now. Maybe not really human-adjacent. It's really hard to say how many tokens a task takes for a human because we don't really think in tokens, but still very useful in that sense. So these are two different approaches we can think about, but what's actually the right way to think about this? The answer is that we probably want to think about all of these. And if we just look at one of these metrics in isolation, it's probably not a great way of measuring things. So I went through some of the weaknesses with measuring with model-specific metrics like tokens and steps. But there are also a lot of weaknesses in the other approach of relying on humans. What's long horizon for a human isn't necessarily that difficult for a model, depending on what the actual task you care about is. There are a lot of tasks that are really tedious and time-intensive. Maybe some financial analyst has to go into an Excel file and fix a bunch of formatting issues throughout the task. Maybe they're changing the theming of the colors in the actual file. That might be really tedious for a human. It might take them days to do that if it's a really big Excel file. But for a model, it can maybe write a Python script or find some other cool trick to do that really quickly. And that's not really hard for it to do. But you never really expect a financial expert to do that because most of them don't really know how to write these Python scripts. So I think that's one thing to note. And the other that I briefly touched upon before is that the methodology of how we actually measure this has a really big impact. And if someone is out there saying, hey, we have some tasks or environments that are 16 hours long on average, and someone else has 20 hours, that's really hard to compare across people because there are so many different things in the methodology that really impact what that actually means. It can mean the quality of the experts they're using. Some more experienced experts might actually be way more efficient at doing a certain type of financial or coding task, whatever it is. And I think this becomes really, really important as we start shifting toward the frontier of even human capabilities. Meter talks about this, but as you shift toward more long horizon tasks and tasks that only the top 10%, the top 1%, the top 0.1% of humans can really do, these estimates start to get really, really noisy. And it's something that we really have to consider. And I think the way agents work is developing on its own separate path. And there are a lot of different bottlenecks and different things that AI agents are better at than even the way humans work. And with that in mind, as these paths diverge, of how humans do work and what their limitations are and what agents do and what their limitations are, it's really important to keep both these metrics in mind because they tell and paint different pictures of what's actually relevant. And you don't really get the whole picture by just looking at one in that sense. Yep. So now the question becomes, how do you measure model capabilities? And this is a really important question because fundamentally, long horizon tasks aren't the only thing we care about. This is the larger question that we want to think about every time we're trying to create tasks, create environments to train our models. And so the first way we can think about this is environment complexity, and specifically environment complexity related to tool coordination. So how many tools or external dependencies does the agent have to coordinate? How many tools or external dependencies does the agent have to move information across? So if we start off thinking about what the world looked like before a long horizon task world, we'll notice that there was a low-complexity world where the agent maybe had to read one file or one set of files in a code base. And that's what a task entailed. But now we can see increasingly, as these tasks become more long horizon, what is important to define for measuring model capabilities is, okay, the agent should be using a ton of different tools like Grafana for observability to parse logs, or GitHub for CI CD, or AWS CloudWatch, or reading and writing to a database. And we're going to notice that as we start to have these agents in these environments use many tools, that we also start to think about environment complexity in regard to state changes, which is external dependency does the agent have to coordinate? How many tools or external dependencies does the agent have to move information across? So if we start off thinking about what the world looked like before a long horizon task world, we'll notice that there was a low-complexity world where the agent maybe had to read one file or one set of files in a code base. And that's what a task entailed. But now we can see increasingly, as these tasks become more long horizon, what is important to define for measuring model capabilities is, okay, the agent should be using a ton of different tools like Grafana for observability to parse logs, or GitHub for CI CD, or AWS CloudWatch, or reading and writing to a database. And we're going to notice that as we start to have these agents in these environments use many tools, we also start to think about environment complexity in regards to state changes, which is effectively the degree to which the environment changes throughout the task. And so fundamentally, the way we want to think about this is, all long horizon tasks aren't equal. So for example, one task can maybe be made artificially long horizon by chaining together unrelated independent tasks. However, that doesn't actually tell us or meaningfully measure the model capabilities. Instead, a key component of this is actually being able to have the earlier decisions in the environment influence the later decisions. And this comes back to how the agents are asked to interact with the tools, how these tools change the state of the environment, etc. So we can look at a concrete example for this. One example where you'll see parallelizable complexity, which is effectively not involving a lot of state changes, is when you can maybe have an agent analyzing a large code base. And then the agent needs to spawn off multiple subagents, and it can very easily parallelize this. You can look at a lot of the different files in parallel, come back to the master agent, and then wrap this all up. But meanwhile, if we look at sequential complexity, we'll see if you have to use a dashboard or logs, a bad early query or a misread can cascade into these downstream steps that really start to have major consequences later on. It's all dependent on how you use those tools and how the state of the environment changes. So the third area that we also need to consider for measuring model capabilities is ambiguity. Ambiguity is defined as the information you give the agent and the environment when starting the task. So this could be the instructions, this could be the artifacts, etc. And increasingly, as these agents work with more artifacts at the start, we want to have them mirror the work that humans really do. And the work that humans really do has a lot to deal with ambiguity. They don't always have the most complete information, and they want to let exploration happen. And so we believe that to measure model capabilities, we need to test the model's ability to explore throughout the environment as well and explore these artifacts similarly to how a human would. Now, the trade-off with this is that if you are going to have ambiguity in the materials you give, there's a lot more possible paths that the agent could take. There's a lot more ways the agent could be right. And that means that standardized evaluation gets much, much harder. Awesome. So I'm going to talk about one of the hardest things there are to build in environments, and one of the most complex things to really think about, because there's a lot of nuance, which is the verifier in the environment. How do we actually know that the work the agent did was correct and give it some reward signal during the training process? So I think there's a few challenges here. To give a high-level overview, tasks are getting more complex, the environment's getting more complex, the trajectories are getting longer, and we've shifted a lot from a lot of the early RL that we were doing. In recent times, it was really in hard verifiable domains, and that's why we saw these gains in math and data structure-style coding problems. But what's happened over time is now we really care about a bunch of economically valuable work in software, across all the domains, as a way to put it, where we can't just run a Python script or run test cases or write a proof to really see whether or not the output was correct or whether or not the environment was changed correctly. We have to start using other techniques, and the main way we're really going to use that is to introduce a judge model or a critic model, as some people put it, and they can add a lot of nuance to how we actually look at a few things here. Very critical for how we actually determine correctness and assign reward, they'll look at two things mainly. One is usually either the final state of the environment and how it's impacted, and the other is looking at the trajectory of how the model that you're actually training made changes to the state of the environment as well and what correctness looked like there. So why do we actually use judges and usually rubrics as a technique? I think it's really important to understand before we can even understand how to use domains. There's an entire class of problems that are really important, and a lot of the problems we care about, that you can't really write a deterministic verifier for. They would be really impractical, brittle, or just downright impossible depending on what the problem setup really is. And I think the other thing also is that, like I said, we're going to look at the trajectory. And not all solutions are really created equal, and not all paths of those solutions are equal either. The worst case of a bad solution we can get is some reward hacking that happens. Lots of different types of reward hacking can happen depending on the setup or the task you care about. An agent can escape a sandbox when you see privileged information you shouldn't be seeing about maybe a hidden test suite for a coding task. This is all behavior that we want to prevent, obviously, because those are not actually valid solutions we care about. And a lot of mitigating this is going to require strengthening your verifier and your environment setup, but the judge is really, really important in actually catching this behavior. And that's an important reason why we actually look at the trajectory that the agent took to get there. Yeah, and I think there's a lot of careful things you want to be doing here. One is that there's nuance in how much guidance or explicit rigidness you want to add to the trajectories that the model can actually take. If we enforce this too tightly, we collapse the state space of how many actual paths the agent explores. And that can be really bad, especially because I think some of the simpler approaches we've seen with judges early on is, hey, we'll just give it a reference answer or solution or maybe a sample trajectory of what a good solution looks like and then just compare against what the model did and say, hey, does it match up with that? And that really does not work for these more ambiguous or open-ended tasks because there's so many possible correct solutions. It's basically impossible to account for every single one. And we want to check for more robust methods that allow for these different solutions. So now that we've established why we use judges, we want to go through some of the general heuristics and principles we think about when we're designing good judges, some of the things that we think about at Theta. So I think the first important consideration to make is that judges are agents too. As environments get really complex, oftentimes a consideration we have is, hey, we have to make sure the harness can scale and match up with whatever environment you have. Maybe that means introducing a bunch of new tools and making sure your harness can support those tools really well, and the agent has clear observability of what's happening in the environment. But I think, like we said, the way the judge determines correctness is that it oftentimes has to look at the state of the environment itself as well. So a lot of the harness that you've designed for the agent might also be reused for the judge as well. I think the best way to illustrate this is the example we have here. Let's say you define a task where there's some deployment failure with the software engineering tasks of some platform you're deploying, and the agent's task is to sift through the CI CD logs on GitHub, look through the CloudWatch logs, figure out whatever happened, apply the changes you care about to the code base, and then open a PR and kick off a redeploy there once the PR is merged. For a lot of that, if the judge actually wants to verify whether or not this is correct, besides just looking at the tool calls the agent made, which are usually not very reliable, it actually has to also check the GitHub logs. It might check the AWS logs or the GitHub logs after the deployment happened to make sure things are actually working properly. So it's really important that the judge has access to the environment in the same way, with some important safeguards, of course. One is that we don't want the judge to make an accidental failure with the software engineering tasks of some platform you're deploying, and the agent's task is to sift through the CI CD logs on GitHub, look through the CloudWatch logs, figure out whatever happened, apply the changes you care about to the code base, and then open a PR and kick off a redeploy there once the PR is merged. For a lot of that, if the judge actually wants to verify whether or not this is correct, besides just looking at the tool calls the agent made, which are usually not very reliable, it actually has to also check the GitHub logs. It might check the AWS logs or the GitHub logs after the deployment happened to make sure, oh, are things actually working properly. So it's really important that the judge has access to the environment in the same way, with some important safeguards, of course. One is that we don't want the judge to make an accidental mutation in some way to the environment after the agent is done. So you want to be very careful about that. Maybe that means enforcing read-only permissions for a lot of this information. It can't actually kick off a deployment or anything like that. So those are things to be careful about. But I think this is really, really important, especially where there's a lot of open-ended approaches. And the only way we can really verify correctness is to actually look at the state itself. The answer isn't obvious whether or not the agent completed the task just from looking at the trajectory. So I think that's one example where this approach is really, really important. I think the other thing to be notable of is, as the environments get more complex, the agent structures get longer and longer. And part of the reason we also need the judge to be an agent is that you can't just use this really basic approach of taking the trajectory and stuffing it in the context window of the judge and have it be a basic LLM call. These structures can get really, really long and really, really complex. So we need to do a lot more thoughtful processing of the trajectory in some meaningful way. So that might mean we put it into some database, we use sub-agents to actually enrich certain information. Maybe we parse out specific phases that the agent was actually in. Maybe the beginning part was it going through logs. The second part was actually writing code. The third part was actually it checking what happened after that. These are all different things we want to do. And in that sense, we need to make the trajectory itself queryable. So that might mean enriching information, like I just said, or some other metadata we can look at at certain steps. It's really important so the agent can find critical steps, like failure points, and verify whether or not those are actually failures. Making that usable for the agent is really, really important. I think another important thing to consider is learnability of our environments. The most important thing here is just the density of the reward signal, and a lot of that comes from your rubric and how the judge is defining that. So I think you'd be very careful with overloading with density in your rubric. A lot of times, especially for frontier problems that models aren't really capable of yet, judges will really struggle to apply that rubric consistently. So there's a lot of QA we need to do to make sure judges are able to apply that information correctly. There's other learnability factors that we care about and we measure in environments, like the distribution of tasks and the actual underlying data there is. And these are all things we think about for learnability, and it's really important. Otherwise, you're just wasting a bunch of compute on problems where the model can't actually effectively learn. These are some emerging rubric judge patterns we've seen. I'll quickly skim over this. Oftentimes, deterministic verifiers aren't completely dead. Oftentimes, we use them in tandem with judges, maybe generating artifact for the judge to actually look over, where you're maybe collecting metrics, or an interesting thing that we could also use is dynamic evaluation time rubrics, where we're actually generating, we're maybe giving partial credit where we've baked in some assumptions that the models made and assume they're correct. It's like grading a test assuming if you got the first part wrong, let's just assume it's correct. Did they get the rest of the part right? That can be really important as well for assigning credit there. I'll skip over this part. I will let Rain just close things off with some things about QA for rubrics. Yep. So for each rubric we produce, we run a couple different tests. We'll go into all of them. Some of them are pretty basic, right? Gold, no op, variants. These are tests you want to be considering regardless for your verifiers. But I think increasingly, as you involve AI in the process of even creating rubrics or verifying rubrics or aiding experts, you need to have more and more tests, especially as the tasks become more long horizon. And so that really touches on the coverage and the expert agreement. But I think what we wanted to close off with today is why a lot of this stuff matters, right? We spent a lot of time earlier in this presentation defining what long horizon means. And a huge reason we did that is because we feel like a lot of the literature and data, a lot of the literature shows that a lot of the data being produced right now and being used to train and evaluate models is actually flawed. So we present three major benchmarks in the area of finance predominantly. And so this is GDP Val, Banker Toolbench, and Apex agents. There's a couple of notable issues here. First, if you look at the average human hours per task, based on what Meter has defined for a lot of the leading frontier models, a lot of these different average human hours per task fall far below that. And so they wouldn't actually be considered long horizon tasks. The second notable issue here is that we see that these benchmarks are already reasonably saturated. And we think this is a downstream effect of the average human hours per task. So it's really important to look at the metrics that are being used here. If you look at the Apex agents IB section of this benchmark that they put out, pass at one effectively means that for 57% of cases, the tasks are 100% solved. That is effectively telling us that there's a large part of these tasks that models are solving, similar to what we've seen already. But I think a third key important part here is the breadth. For each of these different benchmarks, particularly GDP Val, they have a very narrow set of Excel tasks that they consider for finance. And for Apex agents, they're largely focused on IB. What this means is that a lot of these more important areas for learnability, like credit, debt, risk in the domain of finance, don't really get covered. And then I think lastly, I'll note there, the reward signal, as you've already mentioned, is really important. And in regards to the reward signal here, we'll notice that there's really, if you look at what you need for a rubric, you need very granular, detailed reward signal. You need, we have 20 different sub criteria and 10 different sub criteria per criteria. So I think there's a lot of room that's left when you read these benchmarks into how granular a reward signal they're giving, which is really important for being able to go ahead and train your models. With that, I think I wanted to round off with a couple of stats about the data we produce. Here, you can look at some statistics for our finance data. We can see that the human time to complete one task on average is 15 hours over a 50 task sample set. Furthermore, it takes models a pretty long time to work through these tasks. And after all of that, across all the domains we care about within finance, for example, they still struggle significantly. And so here we provide MENET5, notably different than all of these previous scores we see here. So thanks for taking the time to talk with us today. Thank you. to to to They always don't have the most complete information, and they want to let exploration happen. And so we believe that to measure model capabilities, we need to test the model's ability to explore and explore throughout the environment as well and explore these artifacts slimmer to how a human would. Now, the trade-off with this, right, is that if you are going to have ambiguity in the materials you give, there's a lot more possible paths that the agent could take. There's a lot more ways the agent could be right. And that means that standardized evaluation gets much, much harder. Awesome. So I'm going to talk about one of the hardest things there are to build environments, and one of the most complex things to really think about, there's a lot of nuance, which is the verifier in the environment. How do we actually know that the work the agent did was correct and give it some reward signal during the training process? So I think there's a few challenges here. I think just to give a high-level overview, tasks are getting more complex, the environment's getting more complex, the trajectories are getting longer, and we've shifted a lot from a lot of the early RL that we were doing in recent times was really in hard verifiable domains, and that's why we saw these gains in math and kind of data structure-style coding problems. But what's happened over time is now we really care about a bunch of economically valuable work in software, if all the domains is a way to put it, where we can't just run a Python script or run test cases or write a proof to really see whether or not the output was correct or whether or not the environment was changed correctly. We have to start using other techniques, and the main way we're really going to use that is kind of introduce a judge model or a critic model, as some people put it, and they kind of can add a lot of nuance to how we actually look at a few things here. Very critical for how we actually determine correctness and assign reward. They'll look at two things mainly. One is usually either the state, final state of the environment and kind of how it's impacted, and the other is looking at the trajectory of how the model that you're actually training made kind of changes to the the state of the environment as well and what kind of correctness look like there. So, you know, why do we actually use judges and usually rubrics as a technique? I think it's really important to understand before we can even understand how to use domains. There's like an entire class of problems that are really important and a lot of the problems we care about that really you can't really write a deterministic verifier for. They would be really impractical, brittle, or just downright impossible depending on what the problem setup really is. And, you know, I think the other thing also is that, like I said, we're going to look at the trajectory. And, you know, not all solutions are really created equal and not all paths of those solutions are equal either. You know, the worst case of a bad solution we can get is some reward hacking that happens. Lots of different types of reward hacking can happen depending on the setup or the task you kind of care about. You know, agent can escape a sandbox when you see privileged information you shouldn't be seeing about maybe a hidden test suite for like a coding task. This is all behavior that we want to prevent obviously because those are not actually really valid solutions we really care about. And, you know, a lot of this mitigating this is going to require strengthening your verifier and your environment setup but the judge is really, really important in actually catching this behavior. And that's an important reason of why we actually look at the trajectory that the agent actually took to get there. Yeah, and I think there's a lot of careful things you want to be doing here. One is that there's nuance in how much guidance or like explicit rigidness you want to add to the trajectories that the model can actually take. If we kind of enforce this too tightly, we collapse the state space of how many actual paths the agent actually explores. And that can be really bad, especially because I think some of the more simple approaches we've seen with judges early on is, hey, we'll just give it a reference answer or solution or maybe a sample trajectory of what a good solution looks like and then just compare against what the model did and say, hey, does it match up with that? And that really does not work for these more ambiguous or open-ended tasks because there's so many possible correct solutions. It's basically impossible to account for every single one. And we want to check for more robust methods that allow for these different solutions. So now that we've kind of established why we use judges, we want to go through some of the general heuristics and kind of principles we think about when we're designing good judges, some of the things that we think about at Theta. So, you know, I think the first important consideration to make is that judges are agents too. You know, so as environments get really complex, oftentimes a consideration we kind of have is like, hey, we have to make sure the harness can kind of scale and kind of match up with whatever environment you kind of have. Maybe that means introducing a bunch of new tools and making sure your harness can support those tools really well, the agent has clear observability of what's happening in the environment. But I think, like we said, the way the judge determines correctness is that it oftentimes has to look at the state of the environment itself as well. So a lot of the harness that you've designed for the agent might also be reused for the judge as well. I think the best way to illustrate this is the example we have here. Let's say you define a task where, you know, there's some deployment failure with the software engineering tasks of some platform you're deploying and the agent's task is to like sift through the CI CD logs on GitHub, look through the CloudWatch logs, figure out whatever happened, kind of apply the changes you care about to the code base and then open a PR and kind of kick off a redeploy there once the PR is merged. For a lot of that, if the judge actually wants to verify whether or not this is correct, besides just like looking at the tool calls the agent made, which are usually not very reliable, it actually has to also check the GitHub logs. It might check the AWS logs or the GitHub logs after the deployment happened to make sure, oh, are things actually working properly. So it's really important that the judge has access to the environment in the same way with some important safeguards, of course. One is that we don't want the judge to make an accidental mutation in some way to the environment after the agent is done. So you want to be very careful about that. Maybe that means enforcing read-only permissions for a lot of this information. It can't actually kick off a deployment or anything like that. So those are things to be careful about. But I think this is really, really important, especially where there's a lot of open-ended approaches. And the only way we can really verify correctness is to actually look at the state itself. The answer isn't obvious of whether or not the agent completed the task just from looking at the trajectory. So I think that's one example where this approach is really, really important. I think the other thing to be notable of is, you know, as the environments get more complex, the agent structures get longer and longer. And part of the reason we also need the judge to be an agent is that you can't just use this really basic approach of taking the trajectory and stuffing it in the context window of the judge and kind of have it be a basic LLM call. These structures can get really, really long and really, really complex. So we need to do a lot more thoughtful processing of the trajectory in some meaningful way. So, you know, that might mean we put into some database, we use sub-agents to actually enrich certain information. Maybe we parse out specific phases that the agent was actually in. Maybe the beginning part was it going through logs. The second part was actually writing code. The third part was actually it checking what happened after that. These are all different things we want to do. And in that sense, we need to make the trajectory itself queryable. So that might mean enriching of information like I just said, or some other metadata we can kind of look at at certain steps. It's really important so the agent can find critical steps, you know, like failure points and verify whether or not those are actually failures. Making that kind of usable for the agent is really, really important. I think another important thing to consider is learnability of our environments. You know, the most important thing here is just the density of the reward signal and a lot of that comes from your rubric and kind of how the judge is defining that. So I think you'd be very careful with just, you know, overloading with density in your rubric. A lot of times, especially for frontier problems that models aren't really capable of yet, judges will really struggle to apply that rubric consistently. So there's a lot of QA we kind of need to do to make sure judges are able to apply that information correctly. You know, there's other learnability factors that we care about and we measure in environments like the distribution of tasks and the actual underlying data there is. And these are all kind of things we think about for learnability and it's really important otherwise you're just wasting a bunch of compute on problems where the model can't actually effectively learn. You know, these are some emerging rubric judge patterns we've seen. I'll quickly skim over this. You know, oftentimes you, deterministic verifiers aren't completely dead. Oftentimes we use them in tandem with judges, maybe generating artifact for the judge to actually look over where you're maybe collecting metrics or an interesting thing that we could also use is dynamic evaluation time rubrics where we're actually generating, you know, we're maybe giving partial credit where we've baked in some assumptions that the models made and assume they're correct. It's like grading a test assuming like if you got the first part wrong, let's just assume it's correct. Did they get the rest of the part right? That can be really important as well as well for kind of assigning credit there. I'll skip over this part. I will let Rain just close things off with some things about QA for rubrics. Yep. So for each rubric we produce, we run a couple different tests. We'll go into all of them. Some of them are pretty basic, right? Gold, no op, variants. These are tests you want to be considering regardless for your verifiers. But I think increasingly, you know, as you involve AI in the process of even creating rubrics or verifying rubrics or aiding experts, you need to have more and more tests, especially as the tasks become more long horizon. And so that really touches on the coverage and the expert agreement. But I think what we wanted to close off with today is why a lot of this stuff matters, right? We spent a lot of time earlier in this presentation defining what long horizon means. And a huge reason we did that is because we feel like a lot of the literature and data, a lot of the literature shows that a lot of the data being produced right now and being used to train and evaluate models is actually flawed. So we present three major benchmarks in the area of finance predominantly. And so this is GDP Val, Banker Toolbench and Apex agents. There's a couple of notable issues here. First, if you look at the average human hours per task, based on what Meter has defined for a lot of the leading frontier models, a lot of these different average human hours per task fall far below that. And so they wouldn't actually be considered long horizon tasks. The second notable issue here, right, is that we see that these benchmarks are already reasonably saturated. And we think this is a downstream effect of the average human hours per task. So it's really important to look at the metrics that are being used here. If you look at, you know, the Apex agents IB section of this benchmark that they put out, pass at one effectively means that for like 57% of cases, the tasks are 100% solved. That is effectively telling us that like, there's a large part of these tasks that models are solving similar to what we've seen already. But I think a third key important part here is the breadth. For each of these different benchmarks, particularly GDP Val, they have a very narrow set of Excel tasks that they consider for finance. And for Apex agents, they're largely focused on IB. What this means is that a lot of these more important areas for learnability, like, you know, credit, debt, risk in the domain of finance don't really get covered. And then I think lastly, I'll note there, the reward signal, as you've already mentioned, is really important. And in regards to the reward signal here, we'll notice that there's like really, you know, if you look at what you need for a rubric, you need very granular, detailed reward signal. You need, you know, we have 20 different sub criteria and 10 different sub criteria per criteria. So I think there's a lot of room that's left when you read these benchmarks into how granular reward signal they're giving, which is really important for being able to go ahead and train your models. With that, I think I wanted to round off with a couple of stats about the data we produce. Here, we, you know, you can look at some statistics for our finance data, we can see that the human time to complete one task on average is 15 hours over a 50 task sample set. Furthermore, it takes models a pretty long time to work through these tasks. And after all of that, across all the domains we care about within finance, for example, they still struggle significantly. And so here we provide MENET5, notably different than, you know, all of these previous scores we see here. So thanks for taking the time to talk with us today. Thank you. to to to