Harness Engineering: How to Build Software When Humans Steer, Agents Execute — Ryan Lopopolo, OpenAI
Description
https://openai.com/index/harness-engineering/ Speaker info: - https://x.com/_lopopolo - https://www.linkedin.com/in/ryanlopopolo/ - https://github.com/lopopolo With a special post keynote Q&A with Vibhu Sapra (https://x.com/vibhuuuus), cohost for https://latent.space/p/harness-eng
Summary
Generated by gpt-5.6-terraAt-a-Glance
- Verdict: Watch fully
- Core thesis: As coding models become capable of end-to-end implementation, engineering leadership should shift from writing and reviewing code toward building repository-native harnesses that supply context, enforce standards, and let agents execute autonomously.
- Why it matters: This is a concrete operating model for turning abundant model capacity into reliable software throughput without creating an unreviewable, high-churn codebase.
- Best use: Use it as a design reference for an agent-native SDLC: repository structure, skills, CI review agents, durable guardrails, and the organizational shift from manual review to systematic failure elimination.
Executive Summary
Ryan Lopopolo argues that implementation is no longer the bottleneck in software engineering: sufficiently capable coding agents can produce, refactor, migrate, test, and delete code at effectively abundant scale. The scarce resources are human judgment and attention, plus model context. The engineer's job therefore becomes staff-engineer-like orchestration: deciding what matters, specifying what good work means, and deploying many agent trajectories against that work.
His proposed answer is “harness engineering”: make the repository, local tools, documentation, tests, CI, and pull-request process legible to agents. Rather than relying on a large initial prompt or manual code review, encode non-functional requirements—security, reliability, architecture, QA, package boundaries, and coding conventions—as durable instructions that are surfaced at the right point in an agent's workflow.
The most operationally useful pattern is to convert recurring review feedback into automated remediation loops. Documentation defines standards; custom lints, source-structure tests, and persona-specific reviewer agents detect violations; failure messages explain exactly how to repair them; and the implementation agent self-corrects before merge. This removes repetitive human review while improving consistency across every agent-produced patch.
Lopopolo recommends keeping humans at the control plane rather than inside each execution loop. Work begins as well-defined tickets, agents receive a small number of high-quality skills to operate the environment, GitHub PRs become the shared collaboration surface, and repository architecture is standardized so agents can reason locally. His long-term target is to allocate a token budget and prioritized roadmap, then have agents continuously advance the product with minimal synchronous human intervention.
Key Takeaways
- Claim: Engineering should treat code generation as abundant capacity and move human effort to prioritization, systems design, delegation, and specification of quality. | Evidence: Lopopolo says his team has built software exclusively with agents for nine months, describes each engineer as able to drive 5, 50, or 5,000 agent-equivalents, and frames the scarce resources as human time, human/model attention, and context window rather than implementation labor. | Implication: Ken should evaluate engineering leverage through control-plane quality—backlog definition, acceptance criteria, guardrails, and routing—not seat count or lines of hand-written code. | Caveat: The claim depends on models being capable enough for the team's codebase and on having enough token budget; it does not imply that raw output is inherently safe or acceptable.
- Claim: A harness should primarily deliver the right instructions and context at the right time, not become a sprawling custom agent platform. | Evidence: He characterizes prompts, rules files, skills, lint failures, test failures, reviewer comments, and embedded agent SDK checks as variations of prompt injection. He advises against front-loading all standards; for example, let an agent first prototype a React UI, then enforce decomposition and statelessness at lint/test time. | Implication: Build progressive disclosure into agent workflows: broad task intent first, then targeted constraints at planning, implementation, validation, and merge gates. | Caveat: Requirements still need to be explicitly captured somewhere; deferring them too late can cause unnecessary rework if they alter fundamental product or architecture choices.
- Claim: The highest-leverage reliability strategy is to turn repeated human review comments into durable, automated failure prevention. | Evidence: His team uses security and reliability review agents on every CI push, custom ESLint rules, and source-code tests. A concrete example is enforcing retries and timeouts around every fetch call; another is limiting files to 350 lines to preserve context efficiency. He calls Friday “garbage collection day,” when engineers eliminate recurring categories of PR slop rather than repeatedly commenting on them. | Implication: Create a feedback-to-control loop: classify review comments, identify the missing context or invariant, encode it in repository documentation and checks, and measure whether the failure class declines. | Caveat: Automation should target durable, recurring failure classes; trying to encode every stylistic preference or edge case can create noisy gates and reduce delivery velocity.
- Claim: Repository architecture is part of the prompt: standardized, local, bounded code structure makes agent output more predictable and scalable. | Evidence: After beginning with a single-package Electron app that became messy, his team moved to a 750-package PNPM workspace organized by business domain and stack layer. He advocates package privacy, explicit dependency edges, canonical shared utilities, and “one way” to implement common patterns such as concurrency, ORM access, CI scripts, and lint rules. | Implication: For agentic development, favor clear ownership boundaries, canonical primitives, and migration-driven standardization over locally optimized but inconsistent implementations. | Caveat: A 750-package structure reflects this team's needs and should not be copied mechanically; the underlying principle is scoping most changes to a comprehensible repository subtree.
- Claim: Use first-party coding-agent harnesses as the execution substrate, and customize the repository's correctness layer rather than rebuilding model mechanics. | Evidence: Lopopolo argues that labs post-train models in the context of their native harnesses—including patch and shell-tool semantics—so using Codex or comparable first-party harnesses through their SDK/app server lets teams inherit ongoing post-training gains. His team makes Codex the entry point and gives it skills to launch the app, start observability, and operate browser tooling. | Implication: Keep the execution layer replaceable where practical, but concentrate internal engineering on stable artifacts: task contracts, documentation, CI validation, environment skills, and acceptance controls. | Caveat: This creates some dependency on vendor interfaces and behavior; teams need regression checks for model/harness changes even if they avoid maintaining their own execution harness.
- Claim: PRs and repository markdown can serve as the shared collaboration plane for humans and agents, provided throughput is protected from mandatory-review bottlenecks. | Evidence: The team uses GitHub PRs as a hub-and-spoke environment akin to a collaborative document. Contributors—including reviewer agents—can offer feedback, while the implementation agent can acknowledge, defer, or reject it. Lopopolo warns that forcing every comment to be addressed can cause agents to be “bullied” into low-value changes. | Implication: Separate must-fix policy failures from advisory review, and allow agent execution to use judgment within explicit merge criteria. | Caveat: Non-blocking feedback needs clear severity thresholds and automated protections for genuinely critical issues; otherwise throughput can become an excuse for accepting regressions.
- Claim: Autonomy requires investing in the whole SDLC, not just coding: planning, documentation, QA, artifact validation, and user-feedback triage all need agent-operable tools and standards. | Evidence: He estimates token consumption is roughly split across planning/ticket curation, documentation/implementation, and CI. When the team reached deployed software, agents were weak at smoke-testing built artifacts because documentation and tools for downloading, launching, and validating critical user journeys had not been built. | Implication: Do not declare an agentic engineering system mature based on code generation alone; explicitly map and instrument every handoff from roadmap through production verification and operational feedback. | Caveat: The transcript ends while he begins discussing user-feedback triage, so it does not provide a complete operational model for post-deployment support and product learning.
Detailed Brief
Practical harness design: skills, environment access, and source-level controls
- Claims: The development environment should be designed outside-in for the coding agent rather than as a human IDE with an agent bolted on.; A small, improving set of skills is preferable to thousands of narrowly scoped skills, because local tooling changes rapidly and humans cannot reliably track its complexity.; Source-code properties can be tested directly, beyond conventional behavior tests and syntax-level linting.
- Evidence: Skills teach Codex how to launch the application, start local observability for logs and telemetry, and attach Chrome DevTools through a local CLI and daemon.; The team centralizes leverage around approximately five to ten skills.; Their higher-level repository checks enforce package privacy, dependency direction between stack layers, schema deduplication, and use of a canonical async-helper implementation.
- Caveats: Skills need stable interfaces and documentation, otherwise they merely relocate operational complexity from humans to agents.; Structural checks should encode meaningful architecture constraints rather than create arbitrary bureaucracy.
- Implications: Treat agent-operable developer tooling as a first-class product surface, with discoverable commands, observable outcomes, and actionable failures.; Use source-level verification to protect shared abstractions and architectural boundaries that models may violate when optimizing for a local patch.
Trust, QA, and code review economics
- Claims: Trust in agent output rises when QA requirements are documented and verifiable, rather than dependent on a reviewer informally inspecting a patch.; Review bottlenecks are also merge-conflict generators: high PR throughput plus long-lived reviews caused significant conflict even on a three-person team.; Plans should not be rubber-stamped; an unread approved plan can lock in bad instructions for downstream execution.
- Evidence: A product-minded engineer documented what a good QA plan entails, including feature inventory, critical user journeys, and expected evidence attached to PRs.; With each engineer generating three to five PRs per day, the team reduced conflicts by restructuring the code tree and shortening PR open time.; Lopopolo recommends that, if plans are used, they be reviewed as standalone PRs with actual human line-by-line approval before implementation begins.
- Caveats: Tests and QA evidence improve confidence but do not prove all product, security, or operational requirements are satisfied.; Standalone plan approval is useful for consequential work but may impose unnecessary overhead on simple, reversible tasks.
- Implications: Establish a lightweight decision rule for when planning artifacts require human approval versus when a ticket can proceed directly to autonomous execution.; Track PR age and merge-conflict rate as harness-quality metrics, not merely developer-process metrics.
Notable Concepts & Terms
- Harness engineering: The discipline of shaping repositories, tools, context, validation, and workflow so agents can perform software engineering reliably over long horizons.
- Token billionaire: Lopopolo's label for operating at extremely high model-usage scale; he says he spends more than one billion output tokens per day.
- Code is free: A strategic premise that generation, refactoring, deletion, and broad migration no longer consume scarce human implementation capacity; correctness and operational acceptance become the constraints.
- Progressive disclosure: Surfacing only the context and requirements needed at a particular stage of the agent trajectory instead of overwhelming the model with all standards at task start.
- Persona-oriented documentation: Durable definitions of quality from roles such as frontend architect, reliability engineer, security reviewer, and product/QA owner, which can drive specialized review agents.
- Garbage collection day: A recurring team practice of converting the week's recurring code-review complaints into documentation, tests, lints, or agent checks that prevent recurrence.
- LLM as fuzzy compiler: The view that repository context and validation act like constraints and optimization passes, while different models are interchangeable code-generation backends that should still yield acceptable software.
- Codebase as prompt: The idea that directory structure, naming, package boundaries, canonical utilities, documentation, and failure messages all condition what an agent can infer and produce.
Operator Notes / Why Ken Should Care
- Audit the last 30 days of human PR comments and bucket them by recurring failure class; choose the top three and convert each into a repository rule, actionable error message, and CI check.
- Define a minimum agent task contract for tickets: objective, success metrics, relevant scope, acceptance evidence, and explicit escalation conditions.
- Consolidate agent tooling into a small set of versioned repository skills that can start services, access observability, run tests, and validate shipped artifacts without manual setup.
- Add model/harness regression evaluation before changing coding-agent versions, focused on the team's actual architectural and reliability invariants rather than generic coding benchmarks.
- Instrument agentic delivery with PR open time, merge-conflict frequency, review-comment recurrence, autonomous fix rate, CI failure categories, and escaped defects.
- Avoid mandatory response to every automated or human PR comment; define blocking severities and preserve implementation-agent discretion for non-critical feedback.
- Map the portions of the SDLC still requiring human clicks or judgment—especially release smoke tests and post-launch feedback triage—and prioritize building agent-operable tools for the highest-leverage gaps.
Source/Metadata
- Title: Harness Engineering: How to Build Software When Humans Steer, Agents Execute — Ryan Lopopolo, OpenAI
- Transcript words: 8089
- Duration seconds: 2780
- Timestamp note: No timestamps or chapters were provided. The transcript includes duplicated interview passages and ends mid-discussion of post-deployment work.
Transcript
Our next speaker is here to speak about harness engineering. How to build software when humans steer and agents execute. Please join me in welcoming to the stage member of technical staff at OpenAI, Ryan LaPopolo. Good morning, London. I'm super excited to be here today. I'm Ryan LaPopolo. And for the last nine months, I have had the privilege of building software exclusively with agents. I am a token billionaire. And I believe that in order for us to get into our AGI future, we want everybody to be token billionaires, to use the models to do the full job. And what that means is to lean into the idea that the models are capable of being a full software engineer. And I've lived that experience by banning my team from even touching their editors, to have to work through the models in order to get the job done. And today I'm going to talk to you a little bit about what it means to lean into that and operationalize the way you work, the code spaces you live in, and the processes on your teams in order to get the agents to do the full job. I believe I'm preaching to the choir here when I say that the way we build software has changed. In the last six months, we have seen coding agents take over the world, and capability has continually advanced at a super fast pace to have these models and the harnesses within which they live take more complex actions, do more complicated work with higher reliability over longer time horizons. And the place we've gotten to here is that implementation is no longer the scarce resource of what it means to do the job of software engineering. Code is free. We have an abundance of code to solve the problems that we come across in our day-to-day as we run our teams, build software, and solve user problems. Hiring the hands on the keyboards as part of our teams is only constrained by GPU capacity and token budgets. And each engineer today in this room has access to five, 50, or 5,000 engineers worth of capacity, 24-7, every day of the year. The only thing that needs to happen in our roles is to figure out how to productively deploy these resources into our code and into our teams to make use of this new capacity. And in this world, skill sets are shifting more towards systems thinking, system design, and delegation in order to make use of this abundant capacity to produce code to solve problems. And there are three reasons that this happened, all of which happened in late 2025. For me, the magic moment was GPT 5.2, which, when it came out, was able to do the full job of a software engineer. The models at this point are good enough where they're isomorphic to you and me in terms of the ability to produce code at high quality that solves real user problems in real code bases. Code is free. And I know this is maybe a scary thing to hear because code carries maintenance burden, but it's free to produce, free to refactor, and it is not a thing to get hung up on anymore. We think of code as burden because it's a synchronous attention drain on the human engineers on our team. But the models are incredibly patient. They are infinitely parallel. So the ability to produce, maintain, refactor, and delete code is no longer a forcing function on figuring out how to allocate resources on your engineering teams. So be AGI-pilled here is to believe that the models are capable of producing every line of code we could ever possibly need, figuring out when to delete them, figuring out when to refactor them, or make them more reliable. And it's your role as software engineers to figure out how to unblock your team of agents and humans driving those agents from being able to drive them over long-horizon work to do the full job. The idea here is that every one of you is a staff engineer. You have as many team members as you can possibly drive concurrently and have tokens to support, and you need to look one day, one week, six months into the future to figure out what structures you need to put in place to productively harness this infinite capacity to produce code. The scarce resources in this world that we see today are three things: human time, human and model attention, and model context window. And in the world where human time and attention is scarce, the role is to think about where that time is going, figure out ways to productively automate it, and move that synchronous human time into higher-leverage activities. In a world where human time is scarce, and human time is required to produce code, we have a stack rank. Things are either P0s or P2s. Those P3s will never get done. However, in a world where code is free and infinitely abundant, all those P3s get kicked off immediately, maybe 4x in parallel, we pick one that solves the problem, and in it goes. I've had the privilege of building a ton of agents internally at OpenAI to improve the productivity of my coworkers. And when code is free, all these internal tools can have good localization and internationalization from day one. I can make tools that my colleagues in London, Dublin, Paris, Brussels, Zurich, and Munich are able to experience in their native languages without really having to trade against any of my other team's capacity in order to make high-quality tools. We should be working with the assumption that the best parts of software engineering that we all know, live, and breathe are available in any product that we could ever build all the time. Humans no longer need to concern themselves with implementation. The important thing is not the code, but the prompt and the guardrails that got you there. This is why leaving breadcrumbs, documentation, ADRs, persona-oriented documentation around what a good job looks like, all the historical logs of tickets and code reviews, this is the process that got you and your teams to the code and products that you have today, and this is what needs to happen in order to get your agents there as well. Your job is to build systems, software, and structures that enable your team to be successful. And to do that, we need to make them legible to those agents that are driving the implementation. That means structuring them in a way that's native to the agents, writing them in a way that is respectful of scarce context, which is this other scarce resource here, and figuring out ways to make the tokens that are required to do the job easy to predict. That means making things the same as much as possible so we can limit the amount of attention the model needs to activate in order to do the job. Large-scale refactoring in this world is free, so making things the same is something that you are all able to do. There's never going to be a migration that hangs open for six months now that you can't get the last parts of the code base to do, because you can just fire off 15 agents to drive that work to completion. This is what it means to have a migration, right? We can finish them now. Come on. That's good. That's good. Clap. There's this meta-epistemiological question here about what it means to do a good job, and doing a good job as a software engineer is hard. It requires years of being in the industry to fully internalize what it means to write high-quality, maintainable, reliable code that our teammates are able to build on top of, that is going to accrue leverage to the code base. To do a single patch well probably requires 500 little decisions along the way around the under-specified non-functional requirements that go into producing good code. The agents, the models during their training, have seen trillions of lines of code that make every possible choice of those non-functional requirements that you could ever imagine. So it's our job to specify those non-functional requirements, to write them down in a way that the agents can see this is what it is to do a good, acceptable job that's going to produce a merge patch. And if the agents aren't doing that, it's our job to figure out ways to refine and restrict their output such that the code they write is acceptable. You can simply say do not produce slop. Don't accept slop, and you won't get slop in your code base. But to do that requires taking short-term velocity hits in order to back up or double-click into a task to figure out what it is the agents are struggling with in your environment, put the guardrails in place so they stop making those mistakes, and then figure out ways to step back and spend your time on higher-leverage activities once you solve some of the blockers in the short term. When I think about empowering my team in this way, everyone is an expert in what it is they bring. I have a diverse full-stack team that is expert in front-end architecture, back-end scalability, and being product-minded. And each one of those different personas fleshes out the skill set of my team by bringing a different understanding, a different set of solves for those non-functional requirements. Getting teammates to write those down actually means that every engineer driving agents gets the best of every single person on my team. I don't need to block on low-signal code review in order to learn what it means to write a good QA plan. To have one engineer on my team document that in a durable way means every agent trajectory is going to get a good QA plan, and we can do this once in a high-leverage way that we're able to stack on top of. So, how can we get the agents to do a good job? When I think about empowering my team in this way, everyone is an expert in what it is they bring. I have a diverse full-stack team that is experts in front-end architecture, back-end scalability, being product-minded. And each one of those different personas fleshes out the skill set of my team by bringing a different understanding, a different set of solves for those non-functional requirements. Getting teammates to write those down actually means that every engineer driving agents gets the best of every single person on my team. I don't need to block on low-signal code review in order to learn what it means to write a good QA plan. To have one engineer on my team document that in a durable way means every agent trajectory is going to get a good QA plan, and we can do this once in a high-leverage way that we're able to stack on top of. So, how can we get the agents to do a good job? What are some of the tools and techniques we have in order to essentially prompt inject our agents and continually remind them of what it means to make those specific choices that we expect around those non-functional requirements? And there's a bunch of ways we can do this. We can write good agents.md files. However, with auto-compaction, which is a thing that has continued to improve, GPT 5.4 and Codex is fantastic at auto-compaction. I essentially never have to write slash new anymore. I've got some pictures on my Twitter of me strapping my laptop into the back of my car so I can continue running inference while I'm commuting to and from work. And in this world, you have to build for that expectation that context will get paged out over time. We need to be continually refreshing context as the agent goes about doing a task. And the ways we can do that are by having reviewer agents look at the code along the way through the lens of what it means to be successful. Right? We have security and reliability review agents in our code base that are continually running as part of every push in CI that look at those documentations and the proposed patch and do simple things like say, are there timeouts and retries on this bit of network code? Has the code that has been introduced have a secure interface that is impossible to misuse? I'm sure everyone here has been paged at some point for network code that failed in production causing an outage that could have been remediated by a retry and a timeout. And I know I'm guilty of putting that retry and timeout in, merging the bug fix, and otherwise ignoring that. I am not a reliable reviewer or author of code with respect to this nonfunctional requirement. However, taking the time to write some docs, write a lint that is bespoke to my code base that is going to look at every time I call fetch to make sure that there's a retry and a timeout wrapped around it means I've durably solved this problem. And I'm able to do it because I lean on this axiom that code is free, that the agents are able to do a good job, that I can completely migrate the code base to solve this problem durably once and for all. And in order to operate in this way, we need to step back and look at the durable classes of failures that the agents and the humans in the code base are making time after time, figure out why we're spending time on it, devise a solution to systematically eliminate this class of misbehavior, and then continue to observe, refine, and make additional choices on those nonfunctional requirements. One really neat trick I use here is that you can write tests about the source code as well that are separate from lints. If we know that context is limited, we can write a test that limits the fact that files are no longer than 350 lines. We're adapting our code base to the harness, to the models, to do a little bit of engineering to be context efficient and squeeze more juice out of the model capability that we have today. The other things we can think about are providing good error messages that give actual remediation steps to the model and to humans for how to proceed next. It's not enough to say we've got a lint failure because we're waiting in a loop, or that we have an unknown at this deep part of the code base and why is the model writing a function called isRecord. So, if we're not familiar, what we need to do is provide a prompt via a lint or a test failure that says, no, no, no, you shouldn't have an unknown here at all because we parse don't validate at the edge and you certainly have a type here which was derived from Zot, load-bearing infrastructure for our AI future. You can just prompt things. Everything I've talked about here today is a prompt. You can do this without touching the model weights at all. A funny digression here is it seems like each advancement we've had in the complexity of the way we write code to interact with these models comes from both increasing capability in the models and increasingly niche ways for injecting prompts into those models. Prompts, I'm sure you're aware, are prompts. Powers, prompts. Rules files, prompts. Skills, prompts. These lint error messages that I am talking about, prompts. Review agents that inject comments onto the PR that we require the agent to address before it is able to propose it for merge, prompts. You're going to find lots of ways to insert prompts into your code. And one way you can do that is by embedding agent SDKs into your tests. They're going to review the code base for acceptability using prompts that get embedded into the code. And if I find myself spending a ton of time writing prompts, we can actually shell out to the agent for that as well. I've pointed codecs at all of the prompting cookbooks we have on the OpenAI developer guide and told them to synthesize a skill out of them for how to write prompts. Which means when I find a need to write prompts in order to improve my agent performance locally in the code, I use the skill to write prompts that I wrote with the agent looking at the prompts to write the prompts. All the leverage that you're encoding into your repository, your team, and the agents in this way stacks incredibly well. To pull back to this idea that a single product-minded engineer on my team was able to give us a big lift, they know what it means to write a good QA plan. To write a good QA plan, though, you have to document all the features that you have, the critical user journeys, and how users engage with your applications, web apps, APIs, and services. Once you write those down on how to write a good QA plan, with the expectation that all user-facing work has a QA plan, now a review agent is able to assert expectations around what it means to prove that you have effectively written the feature. A QA plan indicates what media should be attached to the PR for the humans and agents to know that you've got a good QA plan. So, if you're not sure if you've done a good job, which has the consequence of me trusting the output more, needing to shoulder surf the agent less, and removing myself from the loop even more, to delegate more and more of the work to agents. And all of this is just making sure the agents have the tools and tokens and context to do the full job, to remove myself from the need as a synchronous driver. The models crave tokens. We can operationalize our code base to give them tokens to drive them forward, using sub-agents and all these other techniques to refine the agent output. I'm excited to let you all know today, in the way you all do, that you can just go build things. Do not hesitate to remove yourselves from the loop by getting the agents to do the full job, because they can. Thank you. Thank you. Very excited to bring on our guest. We've got Ryan LaPopolo today. He just gave the keynote. Very exciting speaker. The man is Full Send Hyper Engineering at OpenAI. So, a little bit of background. We did a Latent Space episode with him. We shipped it the other day. The story, he wrote this great article called Harness Engineering. And we're like, wow, this is pure gold. We have him on the podcast. He's a token billionaire, spending over a billion output tokens a day. That's like over a thousand dollars. So, man is really living it. We want to keep this exciting, ask good questions, ask interesting stuff, ask things that people can learn from. But let's welcome Ryan onto the stage. Hi folks, how's it going? Excited to be here. London has been fantastic and excited to walk through what it is that we do and how we work here. I think you've got to come on this. Oh, I do. The camera is just here. I got blinded by the QR code. Okay, so background, we have about an hour. Scan this QR code, you should get Slido. Slido will let you ask questions. If you see interesting stuff, you can thumbs them up and we'll try to get through them. Unfortunately, the first one, I can't super do, but let's just kick it off. Ryan, can you show us your actual working setup with the laptop? Yeah, here, beach, margarita, linear, right? Oh, wow. I'll say, what's the podcast we put out? We go through some of the work. But let's welcome Ryan onto the stage. Hi folks, how's it going? Excited to be here. London has been fantastic, and excited to walk through what it is that we do and how we work here. I think you've got to come on this. Oh, I do. The camera is just here. I got blinded by the QR code. Okay, so background, we have about an hour. Scan this QR code, you should get Slido. Slido will let you ask questions. If you see interesting stuff, you can thumbs them up and we'll try to get through them. Unfortunately, the first one, I can't super do, but let's just kick it off. Ryan, can you show us your actual working setup with the laptop? Yeah, here, beach, margarita, linear, right? Oh, wow. I'll say what's the podcast we put out. We go through some of the work. But if you want to talk about it, I guess without actually showing us, what's your workflow like, what's your setup, how do you approach a task? Sure. So the way me and my team work is to start with tickets, right? We have chunks of work that we want to do, features we want to add to our apps, reliability work that we want to do. We give that ticket to an agent along with a couple of skills that enable it to manipulate our app. We want the entry point to the development process to be codex, not an environment which we build around it. So we do things outside in, right? Codex is the entry point the same way you would be. And we give it tools. We give it instructions on how to cook. So rather than creating a shell that our app and codex get spawned into, we have a skill that teaches codex how to launch the app, that teaches codex how to spin up that local observability stack to give it logging and telemetry. We give it a skill that enables it to boot up Chrome DevTools and attach to the application with a local CLI that will connect via some daemon that we have. So the whole way we have set up the repository and all of the local dev tools is for codex to invoke them first. That means we have a bunch of little mini harnesses within the code base that make it really easy for us to slot in additional guardrails. A big package of custom ESLint rules which get wired into every PNPM package in the workspace. We have another local dev harness that allows us to add higher-level wholesome tests that assert the structure of the code itself rather than either the syntax or the behavior of the code. Things like package privacy, dependency edges between different layers of our stack, these sorts of things. Making sure that across multiple files, odd schemas are deduplicated, that there's a single canonical implementation of our async helpers, these sorts of things. Because the way we have seen the agents work is to sometimes optimize for local coherence of a package rather than using our shared utilities and things like that. So having observed that behavior, we have built a bunch of little pseudo-linter source code verification things that shake out some of that bad behavior so the humans don't get distracted paying attention to that in reviews. Stuff like that. But the setup optimizes for the agent to do the job and for the humans to not have to keep track of the high churn in the code base. We centralize our leverage around five to ten skills. We don't go super wide on skills, preferring to make the existing skills better because at least I find that the infrastructure within the repository, all the local developer tools, change super frequently. And I don't really have the bandwidth to keep track of this. So we hide all that complexity beneath the skills that the human has to invoke and let the agent just figure it out. One neat thing here is when we moved from using Chrome DevTools protocol directly to having this daemon thing, I didn't know that had happened for three weeks. It was totally fine because Codex was able to do the thing with the documentation and things that we had in place. And part of this you can get more detail in your article. So some background, you wrote a great piece called Harness Engineering. There's a whole section in there on how you thought about skills, thousands of skills versus simplifying it to just quite a few. But okay, continuing on, how do you stop yourself from overengineering harnesses? And a little bit of a similar follow-up is, do you often build small tools for yourself, if ever? Do you build custom tools? Yeah, so I think this is gesturing in the direction of the bitter lesson here, right? Which is, how do I make sure the work that I do isn't completely obsoleted by an increase in model capability? And the way I have thought about that is doing the bare minimum amount of context management to pull in requirements for the agent to do an acceptable job over the course of its work. And context is a thing that I don't think will ever be obsoleted, right? The models must be told the requirements of the task, which guardrails to pay attention to, these sorts of things. So a good harness is really operationalized around giving the model text at the right time so it can look at the work it has done and the information around what a good job looks like. And fundamentally the models are trained to follow instructions. All the harness should do is surface instructions to the model at the right time. So we do want to minimize that too, right? You don't want to front-load all those instructions because then you overwhelm the agent. But all of these requirements around what a good job does need to be paid attention to over the entire course of the PR, right? So figuring out ways to either defer or just-in-time surface those instructions is what a good harness should do, right? If you know that you want your React components, right, to be decomposed so that they make good snapshot tests for individual more stateless pieces, right, you don't need to load that up front. Instead, you should let the agent cook and prototype and experiment with the UI you want to build. And then at lint or test time say, okay, you've done the work. In order to finish it, you have to break this apart so that your components are small and as stateless as possible and have local dependencies on hooks instead of prompt drilling or whatever it is that you want the code to look like. And then the agent will say, oh, this is a new instruction for me. Let me take the patch as written, modify it to make sure that it adheres to the instructions, and then up it goes to GitHub. And this sort of thing is not going to be obsoleted by increases in model capability. It's really just about getting that right text, that right context, to the agent at the right time. Can we talk about an example of a good harness? So a lot of people are asking about the codex model, the codex harness. How does that compare to other harnesses? So cloud code, open code, how do you guys take these decisions into play? You don't work directly on codex, but if there's stuff you can speak about about the codex harness, what you guys see as you architect it out. Yeah, so one thing that I think is super powerful is this notion that the labs are not just post-training the models, but post-training the models in the context of the harness in which they are primarily deployed in, right? The apply patch tool or the specific quoting semantics of how to invoke the bash tool are in the loop for the post-training process for the harnesses from the labs, which means there is leverage to be had by depending on these first-party harnesses directly, at least this is what I believe. And as such, being able to direct through them via things like the SDK or manipulating the codex app server directly means you get to ride the wave of all that leverage in post-training. Instead, focus on the parts that you care about, which is what correct code looks like. I have high confidence that things like clog code and codex will continue to get better. That is the responsibility of the teams working on these coding agents. So in my role, where I don't really want to focus on the coding harness at all, figuring out ways to plug into them in ways that steer the agent means my job can move up to thinking about differences in model behavior between releases rather than deeply understanding the nuts and bolts of the harness. Instead, I can think about what it means to drive the behavior that I want based on the observed behavior rather than the inner mechanics of the thing. It's a perfect follow-up to the next question, which is, do you have any recommendations for collaboration platforms? So when you're in the software development lifecycle, is there any platform that you use for agents, engineers, developers all to collaborate on working on anything? Any tips, any tools? Yeah, so in this world, it has largely been just markdown files in the repository and GitHub that have been the primary hub-and-spoke sort of thing. So in my role, where I don't really want to focus on the coding harness at all, is figuring out ways to plug into them in ways that steer the agent. That means my job can move up to thinking about differences in model behavior between releases rather than deeply understanding the nuts and bolts of the harness. Instead, I can think about what it means to drive the behavior that I want based on the observed behavior rather than the inner mechanics of the thing. It's a perfect follow-up to the next question, which is: do you have any recommendations for collaboration platforms? So when you're in the software development lifecycle, is there any platform that you use for agents, engineers, developers all to collaborate on working on anything? Any tips, any tools? Yeah, so in this world, it has largely been just markdown files in the repository and GitHub that have been the primary hub. Hub-and-spoke sort of thing. If you think about collaborating on a document, you open Google Docs, you write something, you ask for feedback, people comment, you apply suggestions, these sorts of things. This is a little clean-room environment just for this work artifact that you're producing. A PR has a similar purpose. So we treat that as a big hub-and-spoke broadcast domain where all of the agents and humans collaborate together. And because we optimize for throughput, we don't block on any contribution to that. Folks can either review or not. Agents can either review or not. The implementation agent can acknowledge, defer, or reject any feedback that it gets. Really allowing each participant in the production of diffs to make their own judgments around what it means to deliver, receive, and respond to feedback. And this has a nice property of not putting the model in a box in a bunch of places. We want them to use their good reasoning. So being super prescriptive around every bit of feedback that must be addressed can have this catastrophic failure mode of your coding agent being bullied by all of the reviewers. When really, we want to bias toward code being accepted: not perfect, not drowning in minutiae, and these sorts of things. How should people get started with using coding agents? People that have been doing a lot of manually written code, how do they start to transition? What should they offload? How do they come over that barrier? Okay, I'm still checking every PR. I'm copy-pasting from codex. How should the average engineer start to use these tools? I think there's two ways to approach this problem. One is to start using the coding agents to improve your confidence in the code itself as it is written today. I think we would all agree that more tests is probably a good thing, right? To assert that our programs are well specified and behave correctly as our users interact with them is a good thing. And the agents are super good at looking at the existing code, with some context around how it is meant to be used, and writing tests that assert that behavior. So using this to improve your confidence in the quality of the code will also increase the agent's ability to succeed. And then you can successfully navigate it, which means you don't have to worry as much about doing super detailed review of the agent output. The other way to think about this is to look at how you are spending your time. Is it staring at your editor, writing code? Is it waiting for tests to run? Is it waiting for human review feedback? Is CI slow, and you're waiting on that? Maybe you have a ton of flaky tests. Is it waiting for us to do tests and using the agents to incrementally automate the parts where you are spending your time? Because ultimately, the high-leverage parts of our jobs are to define the work that must be done, prioritize and schedule that work, and then effectively empower folks on our team to do that work. And the more we can delegate and move into this sequencing and orchestration role, even if you just think about managing your teams, the more parallel and the deeper individual executions of those delegations we're able to do. If I put primitives in place that make it super easy to spin up ways to respond to events on my Kafka queue, I don't really need to be in the weeds with every engineer making sure they implement a consumer correctly. And these same building-block-style techniques apply really well to the agents and stack really well too. A fun one: how do you work with agents in your car? So I have not used the new voice mode that launched in CarPlay recently. Not ready for that. But usually what I'll do is kick off a task right before I leave the office, tether my laptop to my phone, buckle it into the backseat, and let it cook in the 30 minutes it takes me to get home. Most of the time, with the skills we invoke that tell the agent you're operating on a task, you go until the tests are green. I don't have to reach back there and poke yes, continue onto the thing. And I'm basically able to more fully saturate my day with token consumption. The dream here is that I actually have 50 agents running 24/7, and I don't have to interact with them at all. And the way to do that is to define the work well, figure out ways for it to automatically be scheduled, and remove myself from having to click the button. Every time I have to type continue to the agent is a failure of the harness to provide enough context around what it means to continue to completion. Wow. Good statement at the end. Every time you have to interact with the agent is a failure. Okay. So the following question scales this out, right? As your org knowledge maps scale, what practical steps do you have to enable progressive disclosure? So as you have a larger and larger code base, as you have more people, how do you scale your agents to work better with this? Yeah. So when I initially started this project that I was working on, blank repository, create electron app, the single package, all this sort of stuff, and eventually ended up with a mess, right? Because there's no package privacy that allows me to enforce invariance around what APIs are public versus which ones are not. The agent didn't have concrete hooks in the file system to determine which domains were separate from the other ones. So we ended up going full 10,000-engineer-organization heavy on the architecture. 750 packages in the PNPM workspace, isolated by business logic domain or layer of the stack. Individual small util packages that encapsulate reusable functionality that we rely on being used, that we can encode leverage in. And I do think that in this world, even if you don't actually have microservices, structuring your repositories in ways that you can actually scope the directory subtree you are looking in to be able to do most of the change helps. And code in the file system is also text, which means it's effectively prompts that you're giving to your coding agent. So making the code as much the same as possible makes it so that regardless of where in the repository your agent is looking, it develops a ton of transferable context. You should have one way to do a bounded concurrency helper. You should have one way to construct observable and instrumented side-effectful command. You should have one ORM. You should have one programming language. You should have one way of writing CI scripts. You should have one way of adding additional lint rules. These sorts of things. Because it means that the tokens that you want the model to produce are easier to predict and more consistently predicted regardless of where it looks. So I would say figure out ways to structure the code so it is local to a subtree in the repository for most of the ways you would interact with that system. And then figure out a way to use these agents to completely migrate the code base to be the same. Empower someone on your team to be a dictator to say this is the way it must be done, or figure that out together. And write it down. Evolve the code so that it reflects that reality, these sorts of things. We've got a few questions on code review. Sure. How do you approach code review now that you have such high velocity? Do you just not read the code? Do you just trust the test coverage? How do you write good tests? How do you offload that sigma of, you have a mental blocker: I need to manually check everything before I merge PRs? So that same idea where you have to look at where you're spending your time and figure out ways to spend less of it. When we started, the first thing to do was figure out how to get the agent reliably producing code that we would accept. And a big challenge we ran into is with each engineer producing three to five PRs per day, even on a team of three, merge conflicts were super miserable. Because these PRs tended to be pretty big. We were working on the same parts of the code base. So that's where we moved in two directions. One was to tree out the code a bit more to minimize these merge conflicts, but also minimize the amount of time PRs were open. How do you write good tests? How do you offload that sigma of, you have a mental blocker? I need to manually check everything before I merge PRs. So that same idea where you have to look at where you're spending your time and figure out ways to spend less of it. When we started, the first thing to do was figure out how to get the agent reliably producing code that we would accept. And a big challenge we ran into is with each engineer producing three to five PRs per day, even on a team of three, merge conflicts were super miserable, right? Because these PRs tended to be pretty big. We were working on the same parts of the code base. So that's where we moved in two directions. One was to tree out the code a bit more to minimize these merge conflicts, but also minimize the amount of time PRs were open so that we were reducing the likelihood of a merge conflict actually occurring. And the reason PRs were staying open so long was because we needed code review, because humans were being the blocker in this scenario. So in order to do that piece automatically, I essentially asked every engineer on the team to take one day a week, Fridays, we called it garbage collection day, where our entire job was to take every bit of slop we had observed over the course of the week that was making a PR difficult to merge and figure out ways to categorically eliminate it from ever happening in the first place. Which is where we started closing this loop between the feedback that humans were giving on the PR, indicating some context failure on behalf of the agent, getting that into the repository, and then figuring out ways to automatically prompt inject the agent so that it would self-heal when it produced this bad behavior. And this is how you go from synchronous human time spent giving feedback as code review comments, to documentation in the repository, to automatically surfacing this documentation either via a failing test or a review, and then the user agent who is primed to review the code as written in the context of these docs. But all of that happens by putting those docs in a single place that all these processes are able to attach to. We asked folks to bucket the types of review feedback they were giving into the persona they were operating as. Like front end architect, reliability engineer, scalability sort of thing. And for each of those personas, we spun up a review agent that gets triggered on every push that says, is this code good? Surface any P2s or above that would block this PR from merging based on this documentation that says what good looks like. And with that, and just continuously appending to these files, we started to see slop reduce, reduce, reduce. People have questions about your billion tokens. Where do you think those are split up? So how much of it is on code review? Where is the majority of that usage coming from? And a follow-up for people that are just getting started. Say they've jumped and done a $200 Pro plan, right? If you had to cut your usage by a fifth, how should people maximize their usage? You run into usage limits. You don't want to just copy-paste a million lines of code every six hours, no prompt cache hit. But how should we think about that? Yeah, so I would say it's probably a third, a third, a third between planning, ticket curation, documentation, implementation, and stuff that runs in CI. Do you use plan mode? I've used exec plans, which was an early version of this that we published, which is sort of a proto-skill that says, this is how you should do this. You should structure a plan with milestones and acceptance criteria. I haven't really used plan mode as part of the harness at all. My expectation here is that I should be able to drop a ticket in and have it do the job anyway without diverting through a plan. Because most of the time I'm never going to read it anyway. So I find that if you do use a plan and you approve it without reading it at all, you're actually encoding a bunch of instructions that you don't necessarily want followed. So if you are going to use plans, my recommendation is to push those up as single PRs with just the plan where you actually have humans review every line of it and block on human approval before they get merged and then kicked off. Because you're effectively potentially wasting your time on a rollout with instructions that are bad. So you want to minimize the time that happens. But I do think that getting tokens to be spent in CI is a necessary part here because writing code no longer is the hard part. Getting code accepted and advancing the code and product forward is what it takes to make that written code be valuable. And we have all heard the aphorism that senior engineers give good code reviews. We expect our senior engineers as agents to do the same. Someone asked, is code a disposable build artifact? Yes. Yes. I think we touch on this with Symphony, which is this agent orchestrator that we released. This idea that we can publish a library that's actually a super well-defined spec that the code is a compiled artifact of. And I think using LLM as fuzzy compiler is an interesting mental model to have, right? All of the context that we're putting in the code base for harness engineering is effectively constraints and optimization passes on which code is acceptable to build in the first place. And this is pretty similar to the static analysis and optimization passes that something like LLVM would do in the process of compiling Rust code. And swapping out one model for another is like changing your code generation backend from LLVM to Cranelift in the Rust compiler. And you would expect that all of the rules around what acceptable Rust code looks like produce valid sound machine code out the back, even if the generation process is different and you end up with different x86 instructions. So same sort of mindset for LLMs swapping out different models sort of thing. We want the structure around the code to limit how it is written to things that would be acceptable to us. At a high level, can you give us a picture of what future you're building for? Does context still matter? How do people do engineering, harness engineering, context engineering? What does the future look like? The future that I want to build toward here is where I'm able to take a token budget and a quarter, a half, or a year's worth of work, take the human input to rank what is most important, success metrics, reliability metrics, give it to the machines, and have them continually work and advance my product forward without my hands explicitly on the wheels at all. As we have gone through very early prototyping to internal alpha, internal beta, external alpha, I have felt that new parts of the software engineering process have started from zero and we've had to build up capability, like these pentagonal personality charts, right? Where I spike in this direction, maybe I'm weak over here. And when we get to deployed software for the first time, the agent's ability to do QA smoke testing on our built artifacts before they're promoted to distribution was weak. We hadn't invested any time in this. There were no docs. There were no tools that the agents could use, like download the built artifact, launch it, poke around to make sure that our most critical user journeys were well validated and tested. So because I don't want to be touching the computer, we needed to figure out ways for the agents to build themselves tools to do that part. And there's a whole universe of software engineering outside of writing code, right? I am triaging user feedback.