Open Reader

Everything Is a Rollout — Alex Shaw + Ryan Marten, Terminal-Bench, Harbor, Laude Institute

completed 21:11 Jul 24, 2026 Watch on YouTube

Current Status

completed

Video ID

jRCpXUjz4CI

RAG / Chat

Enabled
Everything Is a Rollout — Alex Shaw + Ryan Marten, Terminal-Bench, Harbor, Laude Institute
Description

Alex Shaw and Ryan Marten present a rollout-centered view of evaluating and improving AI agents. Drawing on their work on Harbor, Terminal-Bench, and OpenThoughts-Agent, they connect sandboxed environments, agent evaluations, and optimization workflows into a practical framework for generating and learning from rollouts. Speakers: Alex Shaw — Member of Technical Staff, Laude Institute Alex is the creator of Harbor, a framework for evaluating and optimizing agents and language models in sandboxed environments. Ryan Marten — Member of Technical Staff, Laude Institute Ryan builds Harbor and works on research-to-production efforts including Terminal-Bench and OpenThoughts-Agent. Harbor: https://www.harborframework.com/ GitHub: https://github.com/harbor-framework/harbor

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Agent systems should be developed like machine-learning systems: treat their behavior as empirically measured rather than deterministic, and make sandboxed rollouts, task environments, verifiers, and feedback loops the core development infrastructure.
  • Why it matters: This is directly applicable to building reliable agent operations: a proprietary, outcome-gradeable eval suite lets Ken select models, prompts, tools, and agent architectures on actual business work rather than vendor claims or public benchmarks.
  • Best use: Use it as a design framing for an internal agent-evaluation control plane, especially for converting production failures and repeated human corrections into reusable test environments and optimization data.

Executive Summary

Alex Shaw argues that conventional software engineering assumed developers could predict what code would do before execution, whereas agentic systems are probabilistic black-box systems whose performance must be managed empirically. His central reframing is that agents—not merely AI-generated code—should be treated as a form of machine learning, so agent development needs ML-like data, evaluation, optimization, and anti-overfitting practices.

The proposed unit of work is a rollout: give an agent an instruction and a sandboxed computer environment, let it act until a stopping condition, capture its trajectory, and run a verifier that produces one or more rewards. Aggregate many rollouts across tasks to obtain an eval result. Harbor is presented as an open-source common format and execution layer for these environments, intended to run arbitrary agents, models, sandboxes, and tasks in parallel.

Shaw's practical recommendation is to start with the outcome that matters and the ability to grade it, then compare all candidate models and agent configurations against that internal eval. He identifies four near-universal evaluation categories: agents building the company's product, agents using the company's product, agents powering product features, and agents automating internal processes.

The talk goes beyond benchmarking: rollout infrastructure can power distributed agentic map-reduce jobs, generate trajectories for supervised fine-tuning or reinforcement learning, and mine prior agent sessions for recurring mistakes. The important operating loop is therefore not a one-off benchmark score, but continuously converting task outcomes and failure evidence into better environments, skills, prompts, routing choices, and agent harnesses.

Key Takeaways

  • Claim: Agent development is closer to machine learning than deterministic software engineering because agent behavior cannot be assumed from the program text alone. | Evidence: Shaw contrasts a regex phone-number extractor, whose output he says is predictable across one million executions, with a GPT-based extractor that handles more varied formatting but cannot be known to produce exactly the same output every time; he extends François Chollet's claim about agentic coding to agents generally. | Implication: Ken should require empirical reliability evidence for agent workflows and avoid treating prompts, tool definitions, or generated code as sufficient proof of production behavior. | Caveat: The comparison is conceptual rather than evidence that every agent call is meaningfully nondeterministic; simple tasks may still be highly reliable.
  • Claim: The core agent-development analogs to ML are environments as data, evals as validation, skills/prompts/tools/models as weights, environment feedback as the loss signal, and reward hacking as an overfitting risk. | Evidence: Shaw maps ML training data to environments; model weights to skills, prompts, tools, and model selection; optimization to text-based methods such as JEPA or coding-agent loops; and gradient steps to pull requests into an agent repository. | Implication: Version and evaluate agent configuration changes as rigorously as model changes, with explicit regression coverage and checks that agents are satisfying the intended outcome rather than gaming the reward. | Caveat: The mappings are useful operating metaphors, not literal equivalences: agent optimization often changes prompts, tools, policies, and workflows rather than differentiable parameters.
  • Claim: A useful agent eval is an executable environment composed of an instruction, a sandbox in which the work occurs, and a verifier that grades whether the required outcome was achieved. | Evidence: The Harbor rollout starts a sandbox, passes it to an agent, records the resulting trajectory until a stopping condition, gives the sandbox to a verifier, and aggregates the verifier's reward across many task rollouts. | Implication: For business-critical workflows, define success conditions and evidence of completion before selecting an agent or model; verifier quality is a first-class reliability and governance requirement. | Caveat: A weak or incomplete verifier produces a misleading score, and the talk explicitly allows verifiers to range from programmatic tests to rubrics or another agent.
  • Claim: Internal evals create model independence by allowing teams to choose on their own cost-performance frontier rather than trusting provider brands or public benchmarks. | Evidence: Shaw cites Satya Nadella's advice to start with the eval that matters and the ability to grade the outcome, then “welcome all models.” He cites Ramp's internal-codebase benchmark, Ramp SWE-bench, as a way to choose coding agents for Ramp's own use cases without simply maximizing token spend. | Implication: Ken can use a representative internal suite to make model-routing and vendor decisions based on measured unit economics and task success, not generalized leaderboard performance. | Caveat: An internal benchmark only transfers to production to the extent that its task mix, constraints, and grading resemble real work.
  • Claim: Every company that works through computers has multiple high-value surfaces for agent evaluation and automation. | Evidence: Shaw identifies four observed Harbor use cases: evaluating how agents build a company's product; how agents use the product; how agents power product features; and how agents automate internal processes. He notes that agent usability may increasingly require a headless or developer-accessible product mode. | Implication: Prioritize processes by automability and gradeability, then decide whether the immediate objective is internal efficiency, an agent-native product interface, or a customer-facing agent feature. | Caveat: Not every computer-based process is suitable for full automation; the value depends on task repetitiveness, available access, safety constraints, and whether outcomes can be graded.
  • Claim: Parallel rollout execution is critical because long-running agent tasks otherwise make iteration loops too slow. | Evidence: In a Terminal-Bench 2.1 example, Harbor runs 64 Codex/GPT-5.5 rollouts in parallel; Shaw describes a planned launch workflow meant to run up to 10,000 rollouts asynchronously and notes that individual rollouts can take substantial time. | Implication: Build capacity controls around batch evaluation—budget caps, task sampling, run trace retention, and verifier audits—while parallelizing enough to keep the agent-improvement loop fast. | Caveat: Large parallel experiments increase spend and can hide systematic evaluation defects if tasks, verifiers, or environments are not validated first.
  • Claim: Rollout traces are not only for scoring; they can become operational data for failure analysis, synthetic data generation, and subsequent agent optimization. | Evidence: Shaw demonstrates an agentic map-reduce workflow over Codex sessions: a cheap Cursor CLI map step writes each correction as mistake/reason/correction, then a stronger model produces recurring failure categories. He says trajectories are also used for SFT, rewards plus token trajectories for RL, and text-feedback-driven evolutionary or hill-climbing optimization. | Implication: Treat production traces and human interventions as a feedback dataset: periodically cluster failure modes, turn recurring ones into regression tasks, and use the resulting suite to assess changes before deployment. | Caveat: Using agent-generated analyses to improve agents can compound mistaken judgments unless sampled outputs and failure labels receive human review.

Detailed Brief

Harbor's interoperability and ecosystem position

  • Claims: Harbor is positioned as three things at once: a format for agentic environments, an open-source rollout framework, and a registry of training and evaluation environment sets.; The format's purpose is to make environments portable among teams and systems, increasing what Shaw calls data velocity.; Harbor supports more than the basic single-step rollout, including multi-step interactions, separate verification sandboxes, artifact collection, and simulated users.
  • Evidence: Shaw says Harbor had roughly 300 to 400 eval sets at the time of the talk.; Named projects built on or integrated with Harbor include Terminal-Bench/Terminal-Bench 2.1, Frontier Suite, Banker Toolbench by Handshake, RuneBench, Scale's Atlas suite, Poolside's model-training evals, Cognition's Frontier Code, LangChain Deep Agents and sandboxes, and Snorkel's Senior Suitebench.; Senior Suitebench is described as measuring agent performance under ambiguity with behavioral feedback; RuneBench measures play in RuneScape.
  • Caveats: The speaker presents an ecosystem snapshot and adoption claims from a product talk; these references establish usage examples but do not independently validate Harbor's relative performance, maturity, or suitability for a particular deployment.
  • Implications: A common environment format can reduce lock-in between agent frameworks, sandbox providers, and model vendors, provided task definitions and verifier interfaces are kept portable.; Benchmark diversity is expanding beyond code patches toward long-horizon, tool-using, ambiguous, and domain-specific work, so a narrow coding benchmark should not serve as a general agent-readiness proxy.

Rollouts as a distributed production primitive

  • Claims: The same orchestration layer used for evals can run high-volume, non-evaluation agent work.; Shaw calls this agentic map-reduce: distribute many isolated agent jobs, then intelligently aggregate their outputs.; A mixed-model architecture can use a cheaper, faster model for broad map work and a more capable model for a narrower aggregation or judgment step.
  • Evidence: Examples include processing reimbursement receipts, semantically searching Obsidian files, and analyzing many pull requests or prior agent sessions.; In the demo, Shaw uses Cursor CLI for map tasks and a model he calls “fable five” for the quality-sensitive reduce summary, with a configured concurrency limit of 32 on Modal.
  • Caveats: The talk does not specify the privacy, access-control, retention, or approval controls required when production documents, code, receipts, or conversation traces are sent into distributed sandboxes.
  • Implications: For high-volume document or trace analysis, separate execution from aggregation: use low-cost workers for bounded extraction and reserve expensive reasoning models for final synthesis, escalation, or review.

Notable Concepts & Terms

  • Rollout: The fundamental agent execution unit: an agent attempts a task in an environment, produces a trajectory, and receives a verified outcome or reward.
  • Environment: An executable task package containing instructions, a controlled workspace or sandbox, and a mechanism to determine success.
  • Verifier: The grading layer that determines task completion, potentially using programmatic tests, rubrics, or another agent; it defines the trustworthiness of the score.
  • Trajectory: The record of an agent's actions and outputs during a rollout; it can support debugging, failure analysis, SFT, and RL.
  • Reward hacking: The agent satisfies the measured reward without fulfilling the intended task, the agent-system analogue of overfitting and a reason to audit verifiers.
  • Agentic map-reduce: A distributed pattern in which many agents process items independently and a reducer model synthesizes the results.
  • Terminal-Bench: A computer-use/terminal-agent benchmark used in the talk's parallel rollout example.
  • Harbor: The open-source environment format, rollout execution framework, and environment registry promoted in the talk.

Operator Notes / Why Ken Should Care

  • Select one recurring, computer-mediated workflow with measurable completion criteria and build a small internal regression suite before expanding agent deployment.
  • Require each agent workflow to retain task inputs, action traces, terminal artifacts, verifier outputs, model/version identifiers, and cost/latency metrics so failures are reproducible.
  • Create a review process for verifiers and reward functions; explicitly test whether an agent can obtain a passing score while violating the real business intent.
  • Use a two-tier experimentation policy: low-cost models for broad task exploration and stronger models only where measured evaluation results justify them.
  • Turn repeated operator overrides and corrections into labeled failure categories, then add representative cases to the pre-deployment evaluation suite.
  • Before adopting a distributed rollout platform for production data, define sandbox isolation, credential scoping, data-retention, and human-approval controls.

Source/Metadata

  • Title: Everything Is a Rollout — Alex Shaw + Ryan Marten, Terminal-Bench, Harbor, Laude Institute
  • Transcript words: 4408
  • Duration seconds: 1271
  • Timestamp note: No timestamps or chapter markers were present in the supplied transcript. The transcript contains repeated passages near the latter half and closing repetitions.

Transcript

3538 words en Processed in 154.5s

Alex Shaw Reviewer Awesome. Thank you so much. So, as you said, my name is Alex Shaw. I work at Lot Institute. And I'll be speaking today about Harbor, which is an agent evaluation and RL environment framework. And the title of my talk is Everything is a Rollout. And I think you'll see, as I get into it, why we titled the talk that way. But first, I want everybody to come travel back in time with me to the year 2018, so eight years ago. And we're going to talk about some of the things that were going on in 2018. So, probably you're going to see Avengers: Infinity War later today. The second one is just getting released. GPT-1 was just released. And you probably didn't even notice, although maybe some people did. Musically, this random startup was about to rebrand to a product called TikTok. And you might have just learned about AirPods when you saw somebody walking around with headphones that had no cord. So, what about software engineering? What did software engineering look like in 2018? We're going to read this tweet from Stay Sassy, Sassy with two A's, about what it was like to write code in 2018. So, it says: It's 2018 and your co-worker just sent you a 400-line pull request. You get a cup of coffee and sit down to review it. It's beautiful, elegant micro-refactors, crispy method names. You catch a few things, but that's okay. It's part of the dance. They didn't consider extensibility on part of their API. Here's a comment, buddy. And this is actually just part of the tweet, so it keeps going. You should look it up if you want. And it's obviously written humorously. But the thing is, it does feel a little bit nostalgic, just like some of these other things that were going on in 2018. Now, the thing is, this only changed maybe 6 or 12 or 18, if you're a very early adopter, months ago. But it already feels like the distant past in some ways. And I think it's time to start talking about, well, what will the history books say about software engineering? And when I say software engineering, I mean that style of 2018 software engineering. It will probably say, the history books will probably say, a lot of things. But I think one thing for sure that they'll say is: software engineering was when you knew what the code would do before you ran it. So that brings me, then, to a different tweet from Francois Chalet, where he says, agentic coding is a form of machine learning. Generated code is best treated as a black-box artifact whose behavior and generalization should be managed via empirical evaluation, like with any ML model. And that brings us to the next part in this talk, which is to compare and contrast what agent development looks like versus what more traditional software engineering development looks like, and why it demands a new set of tools to really understand what's going on and have confidence and trust. So here's a 2018 program right here. So you can tell already the purpose of the program is to extract phone numbers from text. And we have a regex right here that looks for the phone number. And I can say with 100% confidence what will happen if I run this program one million times in a row. So now let's update it to the 2026 version. So I swap out my regex and instead, obviously, I throw in my model call instead, and I say, extract this phone number. So in some ways, this is actually a more powerful program because the regex was actually a little bit brittle. It would have missed any phone number that wasn't formatted exactly like how it was specified. Whereas I'm pretty confident that this program with GPT 5.5 will catch a lot of the phone numbers that are formatted weirdly. However, if I ran this exact program one million times, I'm not 100% confident that it will print the same thing every single time or that I know exactly what it will print. It probably gets it right almost every time. This is a pretty simple task. But this is obviously far simpler than the things we're asking these models to do. And the uncertainty only increases as the complexity of the task increases. So now let's come back to Francois' tweet and let's update it a little bit. We'll generalize it. So he says agent decoding. And my claim is, well, just agents in general are a form of machine learning. And then he says generated code is best treated as a black-box artifact. And I say agent performance itself is best treated as a black-box artifact. And that brings us to our new paradigm. So we know now agent development is more similar to machine learning than it is to software engineering. So what are the things that we should keep an eye out for? The tools that you use for machine learning and also the pitfalls in machine learning? And what are the analogs for those with agent development? So we have machine learning on the left, agent building on the right. So training data, that now looks like environments. And then your test and your validation set, that looks like evals, which, if you take off the mask, is also actually just environments, at least the style that we'll be talking about today. Weights, your model weights, are now skills, prompts, tools, the model, whether you're picking between models or actually updating the model yourself. Your loss function now looks more like environment rewards and feedback. Your backprop or optimizer is now some context text-based optimization algorithm like JEPA, or even just running a coding agent in a loop. Your gradient descent step looks like a pull request into your repo. And overfitting looks like reward hacking, or also just overfitting. That's also possible for agent development. So we have these things on the left. There are products and libraries and platforms that were built for these purposes. And then for everything on the right, we're just getting started. So we built Harbor to answer some of these questions. We're building it to answer more of these questions, but other people are also building interesting products and interesting frameworks to help people tighten this loop of agent development. And then something else to consider is that, as popular as machine learning was, and it spawns some of the largest companies in the world, you have Databricks, which is worth hundreds of billions of dollars. Already, the number of people using and building agents is probably a magnitude larger than the number of people that ever were doing machine learning. And that trend will only continue. And now let's look at the last piece of Francois' tweet that I want to call out, which is these agent performance, da, da, da, should be managed via empirical evaluation. And that brings us to our next question, which is: how do I actually evaluate an agent? And this is when we start to get into what Harbor does. So the answer, as we already gave away earlier, is environments. That's your evals and your training data. And what is an environment in this case? Well, we need an instruction. We need some way to tell the agent what it's supposed to do. And then we need some place for the agent to try to do this thing. And right now, all of the agents run on computers and they do stuff on computers. So we'll put it into a computer, but we'll put it into a virtual computer, so a sandbox. And then we need some way of telling whether or not the agent actually did the thing that we told it to do in the sandbox, within some amount of time or other stopping condition. And that's your verifier, which is like some programmatic tests or rubrics or agent that comes in to see what happened. So in Harbor, you specify this as a file directory. And this specific directory layout has become relatively standard in a lot of the environment space. So lots of people have adopted it and use it as a way to specify environments, which is useful because then it can easily pass between hands and becomes interoperable. And then, okay, cool, we have a bunch of data now. We've implemented a bunch of these environments. Now, how do I actually use it to start understanding what my agent can and can't do and how to make it better? So that's a Harbor rollout. And so what you do is you start with your tasks, and then you take it, you start up your sandbox, and then step one, you pass that sandbox to your agent. So you're either running your agent outside the sandbox and executing commands into it, or you're running it inside the sandbox and it's calling whatever its commands are as part of its program. And then it runs for some amount of time until it hits a stopping condition, produces some trajectory, and we'll come back to that later because that's important. And then you pass that sandbox to the verifier, which runs some verification process. And then finally, you stop the sandbox. The verifier produces a reward or a set of rewards. And then you take those, you aggregate them across a bunch of rollouts in a data set with some agent, and that becomes your avow result. So this process here, which looks relatively simple, is actually extremely universal. And it's also a little bit overly simplified. So Harbor by now allows for a lot of different flavors of rollouts. So you can do multi-step, you can run your verification in a separate sandbox, you can collect artifacts, you can simulate a user. and executing commands into it, or you're running it inside the sandbox, and it's calling whatever its commands are as part of its program. And then it runs for some amount of time until it hits a stopping condition, produces some trajectory, and we'll come back to that later because that's important. And then you pass that sandbox to the verifier, which runs some verification process. And then finally, you stop the sandbox. The verifier produces a reward or a set of rewards. And then you take those, you aggregate them across a bunch of rollouts in a data set with some agent, and that becomes your avow result. So this process here, which looks relatively simple, is actually extremely universal. And it's also a little bit overly simplified. So Harbor by now allows for a lot of different flavors of rollouts. So you can do multi-step, you can run your verification in a separate sandbox, you can collect artifacts, you can simulate a user. But this is the bare-bones approach that they're all built off of. So what is Harbor more specifically? One, it's a format for specifying agentic environments. Two, it's an open-source framework for performing rollouts in parallel using any agent with any model in any sandbox on any task. And it's a registry of popular training and avow environment sets. I think we have three or 400 avow sets by now. And in fact, I think two or three benchmarks even came out today that run with Harbor. So we're excited about that. And in general, it is trying to be a common language for environments. So a way for people to specify things that are extremely interoperable, and it allows you to maximize data velocity and just increases progress in the industry. So who needs avowals? Now that we understand how to make them, we understand how to use them. Well, now who should actually be doing this? And the answer is every single company that uses computers. And I think that's probably close to all of the companies in the world. And the reason is because if you're doing something on a computer, then you should be seeing if you can use AI to automate part or all of that process that you're currently performing on a computer, because that will increase the productivity of your company and therefore increase the value that your company generates. So I like this quote from Satya Nadella from the Applied Compute podcast he did last week. He says, if you want to build an agentic system, start with the avowal that matters and your ability to grade the outcome. And then say, I welcome all models. So I like that last line because what he goes to say is that as soon as you have an avowal, the power is now in your hands. You can consider every single model. You don't have to trust brand. You don't have to trust somebody else's avowal. You don't have to trust a public avowal. And you can skate the Pareto however you desire to balance that cost-performance trade-off. So step one, build the avowal. Step two, optimize against it. So what will people evaluate? And I'm going to list four things. And these four things are based off of what we actually see people using Harbor to build evals for right now. So the first type of avowal is people evaluating how well agents build their products. So everybody that's building a software product has some internal code base or set of internal code bases. I think Ramp recently announced Ramp Sweebench, which is a sweep bench built off of their internal code base. And what that allows them to do is pick the coding agent or the model that performs best on their internal use cases. And they don't have to maximize token spend. Instead, they can make informed and educated decisions about how they build their products with agents. The second type of avowal that we see people build is how to evaluate how well agents use your product. So anybody that builds a software product is probably moving towards a world where they offer some sort of headless mode. So you see some companies have always been this way. They've been developer-first, like HubSpot and Stripe and things like that. And the idea is, if you can make an avowal to see how well agents use your product and then iterate on your product to make it more usable for agents, you're going to get more usage, and therefore your product will become more valuable to agents. And then three, evaluate how well agents power product features. And then four, evaluate how well agents automate internal processes. So depending on what type of company you are, one or more of these might apply to you, but everybody should be considering right now how they can do one of these things. Okay, so what are the different use cases of the Harbor rollout? So we've talked about evals. That's the most popular right now. Oh, here we go. I actually have a video even of us doing an eval in Harbor. So this is how you can run it from the command line or have your agent run it from the command line. So you can see here it says Harbor run. We're running, in this case, Terminal Bench 2.1 with Codex and GPT 5.5. And we're saying let's run 64 in parallel. And then we actually type in the launch command right here, which is something we're rolling out soon, which allows you to launch the rollouts on Harbor servers, which means you can just fire and forget and go to sleep and then come back the next day, and 10,000 rollouts are done for you. And you can see in here what it looks like to look at those rollouts happening. But they don't complete in the time span of this video because rollouts can take a while, which is another case for parallelizing as much as you possibly can to tighten that loop and maximize your throughput. So there are people who create tasks and sell those as data. So there's actually a multi-billion-dollar market right now that exists probably around Harbor data and also other types of data. But it's very, I guess, lucrative right now, where experts can specify certain types of tasks that they would like automated and then sell those to one of the labs that's training models to improve on that capability. You can also do what we call prod rollouts. So remember, Harbor is evaluate any agent with any model. I guess not even evaluate. We'll just say run any agent with any model in any sandbox on any task. And that actually doesn't mean you have to do it for evaluation. It also doesn't mean you have to do it for training. It could literally be that you want to do what we've been calling agentic map reduce, which is you just want to run a ton of agents on distributed computes of sandboxes in parallel and then somehow aggregate those results, probably also intelligently. So things like looking over a bunch of trajectories to detect reward hacking or processing a bunch of receipts for reimbursements or searching over your Obsidian files to figure out semantically where you wrote some note or processing a bunch of PRs to ask a question about it. So we see a bunch of different use cases. This is actually an emergent use case. So we didn't build Harbor for this, but we see people doing it a lot. So we actually built this feature and launched it just for this. So it's Harbor exec. It's going to go away soon, but you can see in this scenario, actually, let me see if I can pause this. So I'm actually going to go back a little bit before we kicked it off. So I'm saying Harbor exec input is all of my codec sessions from the June 20s. And then I say prompt, if I corrected the agent, write in analysis.json with mistake, reason, and correction. And then my reduce prompt is summarize recurring mistakes and failure categories into concise feedback.md file. And then you can see in this case, I'm running on modal. I do the map step with cursor CLI because it's cheap and fast. And then I do the reduce step with fable five because I wanted to have an accurate summary, and I care about intelligence in that scenario. And then I'm limiting it to 32 because I didn't actually want to process all of my sessions for this demo, but there's no reason you couldn't do 10,000 sessions or a thousand sessions. And then maybe you have to go and understand the results more deeply. But yeah, you can see here, I think, so we kicked this off. A bunch of these rollouts are in parallel. This one is running locally on my computer with cloud sandboxing but orchestrated locally. And then you can see here all of the cursor rollouts running right now. But those will take a second to finish. And then you can see the reduce step here where fable recommended me recurring mistakes. And these recurring mistakes, for example, could be used to inform the next batch of Harbor tasks that I create to evaluate my agent and pick a better agent or train up a skill or something like that. So another use case of Harbor. And then we see people taking the trajectories and doing SFT. And then we also see people taking the reward or rewards the trajectory in the form of tokens and doing actual reinforcement learning. For example, Tinker launched an integration with Harbor. And then people also do other types of optimizations. So we mentioned JEPA. But any of these I think, so we kicked this off. A bunch of these rollouts are in parallel. This one is running locally on my computer with cloud sandboxing but orchestrated locally. And then you can see here all of the cursor rollouts running right now. But those will take a second to finish. And then you can see the reduce step here where fable recommended me recurring mistakes. And these recurring mistakes, for example, could be used to inform the next batch of Harbor tasks that I create to evaluate my agent and pick a better agent or train up a skill or something like that. So another use case of Harbor. And then we see people taking the trajectories and doing SFT. And then we also see people taking the reward or rewards the trajectory in the form of tokens and doing actual reinforcement learning. For example, Tinker launched an integration with Harbor. And then people also do other types of optimizations. So we mentioned JEPA. But any of these evolutionary methods, you take the text feedback in the form of a trajectory and the eval, and you can do some sort of auto hill climbing with a harness or a skill. So in the last couple minutes, I just want to talk about some cool things people have built with Harbor. I'll try to breeze through this. So Frontier Suite, an ultra long horizon software engineering benchmark, built on Harbor. Banker Toolbench, an investment banking benchmark built by Handshake on Harbor. Swix saying that his team at Cognition migrated all their evals to Harbor. This Runebench is one of my favorites. It's a benchmark to measure how well agents can play RuneScape. Scale launched their whole suite, Atlas suite, on Harbor. Poolside uses Harbor to do all of their evaluations for model training. Kevin Gu created Auto Agent, which is a self-optimizing agent that all you have to plug in is a Harbor eval. After query post trained a model using Harbor. Cognition released Frontier Code recently, which is a Harbor benchmark. And Langchain just integrated Deep Agents and their sandboxes into Harbor. And then actually just today, Snorkel released Senior Suitebench, which is a benchmark for measuring how well agents can function under ambiguity with behavioral feedback. So that's a brief overview of Harbor and the different things you can do with it. I would encourage everybody here to check it out. And then this is a link to my Twitter. And the reason I'm linking my Twitter, usually I link Harbor Docs, but we're actually actively trying to grow the Harbor team right now. So if you want to get involved, then shoot me a DM on Twitter. And if you just want to see the docs, then click the docs link on my Twitter, and that's the best way to get there. So thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much. and and and and and and and on my computer with like cloud sandboxing but orchestrated locally. And then you can see here all of the cursor rollouts running right now. But those will take a second to finish. And then you can see the reduce step here where fable recommended me like recurring mistakes. And like these recurring mistakes, for example, could be used to inform the next batch of Harbor tasks that I create to evaluate my agent and pick like a better agent or train up a skill or something like that. So another use case of Harbor. And then we see people taking the trajectories and doing SFT. And then we also see people taking the reward or rewards the trajectory in the form of tokens and doing actual reinforcement learning. For example, Tinker launched an integration with Harbor. And then people also do other types of optimizations. So we mentioned JEPA. But any of these like evolutionary methods, you take the text feedback in the form of a trajectory and the eval, and you can do some sort of like auto hill climbing with a harness or a skill. So in the last couple minutes, I just want to talk about some cool things people have built with Harbor. I'll try to breeze through this. So Frontier Suite and Ultra Long Horizon software engineering benchmark built on Harbor. Banker Toolbench and investment banking benchmark built by Handshake on Harbor. Swix saying that his team at Cognition migrated all their evals to Harbor. This Runebench is one of my favorites. It's like a benchmark to measure how well agents can play RuneScape. Scale launched their whole suite, Atlas suite on Harbor. Poolside uses Harbor to do all of their evaluations for model training. Kevin Gu created Auto Agent, which is a self-optimizing agent that all you have to plug in is a Harbor eval. After query post trained a model using Harbor. Cognition released Frontier Code recently, which is a Harbor benchmark. And Langchain just integrated Deep Agents and their sandboxes into Harbor. And then actually just today, Snorkel released Senior Suitebench, which is a benchmark for measuring how well agents can function under ambiguity with behavioral feedback. So that's kind of a brief overview of Harbor and the different things you can do with it. I would encourage everybody here to check it out. And then this is a link to my Twitter. And the reason I'm linking my Twitter, usually I link Harbor Docs, but we're actually actively trying to grow the Harbor team right now. So if you want to get involved, then shoot me a DM on Twitter. And if you just want to see the docs, then click the docs link on my Twitter, and that's the best way to get there. So thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much. and and and and and and and