Open Reader

Loop Engineering from First Principles — Kyle Mistele, HumanLayer

completed 17:57 Jul 25, 2026 Watch on YouTube

Current Status

completed

Video ID

xIt_mTQp6mY

RAG / Chat

Enabled
Loop Engineering from First Principles — Kyle Mistele, HumanLayer
Description

A coding agent will happily hand you a 40,000 line pull request that nobody can review and that quietly does the wrong thing. Kyle Mistele's argument is that the fix is not a better prompt but a better loop, borrowed from control theory: a thermostat senses the error between where a system is and where you want it, emits a control signal, and measures again, over and over. Infrastructure as code already approximates this. The point is to design agent loops the same way, so each iteration makes a small, readable change you can actually verify, instead of one giant diff you have to trust. The working example is migrating a codebase one procedure at a time. A sensor, often just Grep or a structural search, finds the smallest unmigrated piece, a controller picks what to work on next, and an actuator agent makes the change against golden patterns defined by hand, gated by deterministic CI like a single loop iteration in CircleCI. The loop tracks its own PRs in version control, refuses to stack a new change while an earlier one is still open, and keeps improving the code incrementally, even while the team is away. Speaker info: - https://x.com/0xBlacklight - https://www.linkedin.com/in/kyle-mistele - https://blacklight.sh Timestamps: 0:00 - Introduction: the 40,000 line PR problem 1:55 - Why more code is not the goal 3:37 - Is the generated code any good? 4:43 - Control loops from control theory 5:49 - Infrastructure as code and Ralph loops 7:06 - Applying control loops to coding 8:35 - Migrating a codebase one procedure at a time 10:40 - Tracking the loop in version control 12:35 - The actuator agent and golden patterns 13:51 - Wiring the loop into CI 15:44 - Avoiding stacked PRs and scaling the controller

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Agent coding loops should be engineered as closed-loop control systems that make small, measurable, reviewable changes—not blind autonomous loops that generate massive PRs.
  • Why it matters: This is a practical architecture for applying coding agents safely to production codebases: deterministic sensing, bounded agent actions, human feedback persistence, PR-level flow control, and controlled throughput scaling.
  • Best use: Use it as a design pattern and implementation checklist for GitHub Actions-based agent workflows that continuously migrate, remediate, or enforce standards in a real repository.

Executive Summary

Kyle Mistele argues that the industry’s enthusiasm for autonomous coding loops has drifted toward a dangerous pattern: repeatedly prompt an agent, use more agents to verify output, and accept code volumes no person can realistically review. He does not reject loops or RALF-style approaches; rather, he argues they are sharp tools that become unsuitable when teams operate complex, customer-facing, regulated, or reliability-sensitive systems.

His alternative borrows from control theory. Treat the repository as a dynamic system: measure the gap between its current state and a defined desired state, choose a bounded incremental intervention, apply it, inspect the result, and repeat. In this framing, deterministic tools should handle deterministic work such as finding violations and selecting targets, while agents are reserved for the non-deterministic implementation work.

The concrete HumanLayer example is a gradual migration of RPC procedures to Effect. An AST-based sensor identifies unmigrated procedures; a baseline scan committed to version control prevents new violations from entering main; an agent with a carefully evolved skill and handwritten 'golden patterns' migrates a small target and opens a PR. Human review is not removed—it becomes the feedback signal that improves the loop.

The most operationally useful details are the feedback-file mechanism and flow-control guardrails. Review feedback is stored in a versioned Markdown file and injected into future agent runs, while each loop allows only one open PR at a time to prevent unattended automation from creating duplicative, conflicting review queues. Once quality is proven, velocity can be increased through larger bounded batches, isolated implementation contexts, or parallel workflows assigned to different reviewers.

Key Takeaways

  • Claim: Do not use agent loops as a substitute for readable, reviewable engineering; design them to converge through small changes with feedback. | Evidence: Mistele contrasts the prevailing 'prompt plus loop' approach with 40,000-line PRs that nobody wants to read, and notes that even heavily agent-verified code becomes effectively unreadable at sufficient volume. | Implication: For production agent systems, optimize for incremental PRs and human comprehension rather than maximum code-generation throughput. | Caveat: He explicitly says RALF is innovative and useful for constrained problems, especially solo or non-critical work; the problem is applying blind loops to team-operated, mission-critical systems.
  • Claim: A reliable agentic loop should be built as a control system with a set point, sensor, controller, actuator, feedback, and protection against disturbances. | Evidence: The proposed model maps a codebase to a dynamic system: the sensor measures current state, the set point defines desired state, the controller chooses an incremental correction, and the actuator applies it before the system is remeasured. | Implication: Before automating a task, Ken should require three explicit answers: what can be measured, what bounded change can be made, and how will the quality of that change be evaluated.
  • Claim: Use deterministic tooling for deterministic detection and selection; do not spend agent context or introduce agent variance where a rule can do the job. | Evidence: For identifying TypeScript RPC procedures not migrated to Effect, HumanLayer uses ast-grep rather than an agent or plain grep. Mistele says he would not send an agent to perform deterministic code’s job, and suggests deterministic target selection such as choosing the first sorted violation or the smallest procedure. | Implication: Build an explicit deterministic policy layer around agent execution: rule-based discovery, stable ordering, scoped inputs, and only then an agentic implementation step. | Caveat: An agent can be used for selection when prioritization is genuinely contextual, such as choosing procedures based on error rates, instrumentation gaps, or APM data.
  • Claim: Prevent regression before running a remediation loop, otherwise teammates and agents can continuously undo the loop’s work. | Evidence: HumanLayer runs a full scan on main, deterministically sorts violations, commits the baseline to version control, and checks each new PR for added unmigrated procedures. Mistele calls this a 'disturbance dampener' because it blocks new violations while the migration loop removes existing ones. | Implication: Any autonomous cleanup or migration program should pair backlog reduction with a CI admission control that prevents new debt of the same class.
  • Claim: The effectiveness of an agent actuator depends more on an iteratively refined skill and local examples than on a one-time prompt. | Evidence: HumanLayer gives the coding agent a skill plus a control signal, includes a response template, and creates handwritten 'golden patterns' before letting the agent operate. The stated rationale is that agents replicate patterns; absent local examples, they fall back to documentation or internet-learned conventions. | Implication: Treat agent instructions, response contracts, and repository-specific examples as versioned operational artifacts with a continuous improvement process. | Caveat: The skill will require ongoing iteration based on actual failed or suboptimal loop outputs; upfront prompt design alone is insufficient.
  • Claim: Human oversight can be retained without making the loop operationally burdensome by converting review comments into persistent, versioned steering instructions. | Evidence: A reviewer posts a '/iterate' comment on a loop-created PR. The workflow loads the PR diff, description, comments, review comments, and the skill into the agent context; the agent fixes the branch and updates a Markdown feedback file that is committed to version control and injected into subsequent runs. | Implication: Implement a low-friction human-on-the-loop interface where review corrections update the system’s durable operating memory rather than disappearing in chat or individual PR threads.
  • Claim: Throughput should be gated by review capacity, then increased deliberately after the loop demonstrates reliability. | Evidence: Each workflow labels its PRs and exits before doing work if any PR carrying that loop’s label remains open. This caps the loop at one pending PR, avoiding stacked, duplicate, and conflicting outputs. For a backlog of 150 procedures, Mistele suggests increasing pace via batches of three to five, separate context windows per migration, or four parallel workflows assigned across four reviewers. | Implication: Make reviewer availability a first-class backpressure signal in agent orchestration, and scale concurrency only alongside an explicit review and merge capacity plan. | Caveat: Increasing batch size or parallelism raises output volume and should follow demonstrated confidence in the loop; separate implementation contexts are presented as cheaper and more reliable than combining many changes into one context.

Detailed Brief

Reusable control-loop applications beyond code migration

  • Claims: The framework applies wherever a repository has a measurable deviation, an incremental repair action, and a way to assess results.; The controller can become outcome-aware rather than merely mechanically compliant by incorporating operational data into its target-selection signal.
  • Evidence: Mistele lists API compliance against an external OpenAPI specification, MCP-server compliance with the current MCP specification, mirroring a Python project to TypeScript or vice versa, and maintaining a fork against upstream as candidate loops.; For the Effect migration, the controller could prioritize procedures with the highest error rates, weakest instrumentation, or APM coverage gaps, then provide that telemetry to the actuator agent so the migration improves operational quality rather than only changing syntax.; He identifies ast-grep as language-agnostic and suitable for multilingual monorepos, with granular include/exclude paths and rules layered over time.
  • Caveats: The talk provides an architecture pattern rather than measured evidence that the proposed loop improves defect rates, cost, or developer throughput across organizations.; Non-deterministic sensors and controllers may be appropriate, but they introduce variance and should not replace deterministic checks where rules are sufficient.
  • Implications: Prioritize loop candidates that have a crisp, machine-observable set point rather than beginning with broad requests such as 'improve the codebase.'; Connect controller prioritization to reliability, security, customer impact, or observability data when a compliance backlog is too large to process uniformly.

Reference implementation topology

  • Claims: Existing CI systems are sufficient orchestration infrastructure for this class of loop; a separate agent cluster is not required.; A single scheduled iteration should perform sense, control, actuation, and PR creation, with deterministic commit and publishing steps outside the agent.
  • Evidence: Mistele recommends GitHub Actions, GitLab CI, CircleCI, or equivalent because they already have repository access, secret handling, dispatch mechanisms, and schedulers.; The workflow runs daily, generates one small PR, and uses the agent’s final response as the PR description after deterministic commit, push, and PR-creation actions.
  • Caveats: The transcript does not specify authorization boundaries, secret-scoping practices, branch protections, test gates, or rollback policy; these remain necessary for deployments in sensitive repositories.
  • Implications: Keep the agent’s authority narrowly focused on a branch-level code change while preserving deterministic CI, protected-branch, test, and approval mechanisms around merge and release.

Notable Concepts & Terms

  • Agentic control loop: A software-agent workflow designed as feedback control: observe a measurable repository state, make a bounded correction, assess the result, and iterate.
  • Set point: The explicitly defined desired state of the codebase, such as all RPC procedures following an Effect implementation pattern.
  • Sensor: The mechanism that detects deviation from the desired state; it can be deterministic tooling, an agentic evaluator, or a hybrid pipeline.
  • Controller: The target-selection and change-sizing policy that converts detected violations into a safe incremental action.
  • Actuator: The coding agent plus its repository-specific skill that executes the selected code modification.
  • Disturbance dampener: A regression-prevention CI check that blocks new violations while an automation loop removes the existing backlog.
  • Golden patterns: Handwritten, idiomatic local examples supplied to agents so they replicate the team’s desired implementation rather than generic external patterns.
  • Flow control: A backpressure mechanism that suppresses new loop runs while a prior loop-created PR is still open, tying automation output to human review capacity.

Operator Notes / Why Ken Should Care

  • Select one high-volume engineering backlog with a deterministic detector—such as deprecated API usage, missing observability wrappers, insecure patterns, or spec drift—and define its desired end state before deploying an agent.
  • Add a baseline-and-diff gate to CI before enabling remediation automation, so the debt backlog cannot grow while the loop processes it.
  • Require every coding loop to label its PRs and enforce a configurable maximum number of open PRs per loop; start with one.
  • Create a version-controlled feedback artifact per loop and a reviewer command path that feeds PR-level corrections into both the current repair and future runs.
  • Separate deterministic discovery, prioritization, commit, PR creation, and merge controls from the agent’s implementation task; scope agent context to the smallest selected unit of work.
  • Before increasing concurrency, measure review latency, conflict rate, rework from '/iterate' cycles, test failures, and cost per accepted PR.

Source/Metadata

  • Title: Loop Engineering from First Principles — Kyle Mistele, HumanLayer
  • Transcript words: 4042
  • Duration seconds: 1077
  • Timestamp note: No usable timestamps or chapters were present in the supplied transcript.

Transcript

3385 words en Processed in 161.0s

Hey, everybody. Hi, my name's Kyle. I'm co-founder of a company called Human Layer, and I'm here to talk about loops. I think we've all been building loops lately, and I realized recently, I think we're all doing it wrong. Loops are really powerful. Don't get me wrong, but so much of the discourse around them is hype-driven and really not helpful, right? I think we have this idea, as an industry, that we can pipe a prompt and a loop to a coding agent and that we can build software this way, right? Maybe we're investing a lot of time in verifiers. Maybe you have six different code review agents, but at the end of the day, if we're doing this, we're still building 40,000-line PRs that nobody wants to read, right? And this isn't to throw shade at Jeff Huntley, right? This is, RALF is innovative. It's a sharp tool that works very well for certain types of problems. It works very well if you're not building on a team, and it works very well if you're not working on critical systems. But most of us are working on teams, and we don't fit in that box. So today, I want to talk about how to build loops that work in large, complex code bases for systems that have real customers, real users, real regulatory obligations, and service level agreements, and everything else that keeps us from shipping YOLO 40,000-line PRs straight to production. In other words, I want to talk about how to build loops for the real world. If you're not aware, this post actually dates back to July. It went viral this past January, which is when a lot of us, I think, started building loops. And of course, much more recently, I'm sure you all are going to see this slide a lot this week, but Peter Steinberger said that we shouldn't be prompting coding agents anymore, right? We should just be designing loops that prompt our agents. Of course, Open Clause notoriously built on loops. Loops build the code. Loops review the code. They merge and release the code. They find and fix the bugs. There's even loops for finding and fixing bugs in the loops that are merging the things, right? It's loops all the way down. Boris Cherny, the creator of Cloud Code, also recently said that this is his entire job as an engineer now: just writing loops to prompt Claude. And eventually, we might not even need loops, right? We're just going to have swarms of agents designing loops to prompt agents building swarms for loops. And I don't know, somewhere we're writing production code, I assume. And in fact, all of our loops that we're building are producing so much code that we can't possibly read all of it, right? So we might as well just not read any of it, right? We're investing in verification and in code review. But all this code is read-only. This is the thesis of a conference that was here in town last month. A lot of smart people at the Frontier Labs think that this is the future of software development. And if you're doing this, you're moving 10x faster, and everybody else is getting left behind. Now, it's not clear how well this works yet. It took six months to fix the Cloud Code terminal flicker. The OpenCode team wrote a renderer in a fraction of that time. And OpenClaw, of course, also notoriously has stability issues. What is abundantly clear, however, is that this shit is really expensive if you don't work at a Frontier Lab and have an unlimited token budget. And all this code that we're writing is actually really expensive, right? Matt Pocock talked about this recently. Bad code is much more expensive in the age of agents than it has ever been at any point in the past. So today I want to talk about what I think works in the real world and what we've started doing at HumanLayer, which, to be clear, is still building loops, right? I think loops are super powerful. But we can design loops and still read the code. In fact, we can design loops that make it easier to read the code because the loops are making the code better. We can solve hard problems in complex code bases with loops. And we can build our software factory incrementally. But to do this is going to take some real engineering, y'all. So let's talk about control theory. Control theory is all about how we drive a dynamic system, which would be your code base, toward some desired stable or optimal end state, right? You have a sensor that measures the current state of the world. You have your set point, the desired state of the world. And the difference between those two things is your measured error. You have a controller that reads that measured error and turns it into a control signal about an incremental change to apply to the system. We have an actuator that applies that change to the system, which is undergoing disturbances in the meantime. And then we re-measure, re-compute our measured error, and we're back where we started. Now, this sounds really complicated, and it can be. I have a twin brother, actually, who's an aerospace engineer. This is how they keep fighter jets from falling out of the sky. But it's probably a little bit simpler than most of y'all think. Does anyone have one of these? A thermostat uses a control loop, right? For any of our European friends in the audience, it's part of something we have here in the States. It's called air conditioning. And most of us probably actually use control loops on a daily basis, right? Kubernetes auto-scaling systems are built on control loops. Infrastructure as code uses a desired-state, current-state, iterative-change control loop pattern. Postgres' auto-vacuum and React's Virtual DOM both use or approximate control loops. Control loops are ideal when we have a system that we want to change, a problem we can measure, and a way to get feedback on the result of that change. Like good software engineers have always been taught to do, control loops change the system incrementally instead of just trying to get straight to the end state immediately, all at once, and risk blowing everything up, right? They help us avoid oversteering and destabilizing the system, and it minimizes risk. So control loops are the opposite of what I'm going to call our blind RALF loop. They're how we avoid PRs that look like this, because nobody wants to review this, right? Which is not to say that all RALF loops are blind loops. The best RALFs are actually applying control theory. I know Jeff Huntley is out in the hall somewhere wandering around. If you go talk to him, he's going to tell you the same thing, right? That RALF is a teaching device, and I think some of us read it a little too literally, but this is how we should have always been building loops. But the other issue with RALF loops is they're not incremental, right? It's just a bash loop. So we have to build agentic control loops. And to do that, we start by defining a set point, which is the desired end state of our code base with respect to some property of it, and we add a sensor. There's a lot of ways to build a sensor. It can be strictly deterministic: your ESLint rules, your AST grep, your Packwerk, or it can be non-deterministic. You can have an agent and a skill and a bunch of natural language rules, and you could also just have a pipeline, a combination of the two. So how do we build agentic compo- whoops. There we go. Now, this is all theory, right? Practically speaking, and because we're using agents, we can blur the lines a little bit between system components. AidenBuy's React Doctor, for example, is fantastic. It is a great way to catch all of the React slop that Claude snuck into your code base last week. But it's a hybrid sensor and controller. It tells you what all the problems are with your React code, and also, by the way, what the top three things are that you should fix, and how you fix them. Similarly, our controller and actuator might actually just be a single agent deciding on an incremental change to make and then applying it in the same context window. But I want to zoom in on the controller a little bit, because without one, or without a well-tuned one, we might make too large a change all at once, or we might make the wrong change entirely. And if you put that in a loop, you're in trouble pretty quickly. So we can use control loops to root out bad patterns and to clean up our code, but we can actually use them for all sorts of things, right? We could make sure that our API is compliant with someone else's OpenAPI spec. We can make sure that our MCP server is compliant with whatever version of the MCP specification that we're currently on. Haven't checked. You could mirror a project from Python into TypeScript, or vice versa. You could even maintain your Vite-based slop fork of Next.js against the upstream. The key questions are: can we find something we can measure? Can we apply changes incrementally? And can we get feedback on the quality of those changes? one, we might make too large of a change all at once, or we might make the wrong change entirely. And if you put that in a loop, you're in trouble pretty quickly. So we can use control loops to root out bad patterns and to clean up our code, but we can actually use them for all sorts of things, right? We could make sure that our API is compliant with someone else's OpenAPI spec. We can make sure that our MCP server is compliant with whatever version of the MCP specification that we're currently on. Haven't checked. You could mirror a project from Python into TypeScript, or vice versa. You could even maintain your Vite-based slop fork of Next.js against the upstream. The key questions are: can we find something we can measure? Can we apply changes incrementally? And can we get feedback on the quality of those changes? To illustrate that, I'm going to walk through a control loop that we use internally at HumanLayer. For our loop, we are incrementally migrating our RPC API to Effect. We adopted it for some of our race-prone code. We like it, so we're adopting it across the rest of our code base. If you've never seen Effect code before, the code on the right is just the trivial procedure on the left rewritten in Effect. The syntax is really weird. We're psychos. We really like it. It's not for everybody. That's okay. This isn't a talk about Effect, so we'll keep moving. Ooh, clicker's not working. Cool. So, step one. We have to build our sensor to find unmigrated procedures. We can have an agent do this, or we could use grep or ripgrep. But instead, we're going to use ast-grep, because it's really powerful. It's a great tool to have in your toolbox for building loops. It's language agnostic. It's out of band from your TypeScript config or ESLint rules, which, if you're a TypeScript developer, you've watched Claude disable with inline comments. So we can just write a simple rule that finds unmigrated procedures based on the pattern above. And over time, we can even layer on more rules that describe other patterns we want to get rid of, with granular include and exclude paths. If you have a multilingual monorepo like we do, it'll work for any language you could possibly imagine. And we can just scan our code base, and it'll produce a long list of violations. Way too long, in fact. It'll give you about 50 keys per violation. So we're just going to filter it down to four, and we're going to sort it deterministically. Why are we doing that? At the beginning, I said this was going to be practical. And so we're going to step outside of our control loop paradigm for a second. Because before we start incrementally migrating procedures one at a time, we need to enforce that all new procedures are using Effect, right? So we're going to run a full scan once on main, sort all the violations deterministically, and track it in our version control. And then on every new PR, we can see if the branch added any unmigrated procedures, right? So this is our control loop, and our system is undergoing disturbances. In this case, all of our teammates shipping clots lot. And this is how we make sure that they're not undoing our loop's work. This doesn't map directly to a part of the control loop, but we can squint a little bit and call it a disturbance dampener. So now that we've stopped the bleeding, we can actually design our controller. For a simple controller, we could just deterministically pick the first violation from the list. You can use bash and jq. Or we could get a little cleverer and use ast-grep to find the smallest unmigrated procedure, and always pick the smallest one to reduce the risk. We could have an agent make the decision if we really want to. I don't think you should ever send an agent to do deterministic code's job, but you certainly can. In fact, depending on the complexity, we could have the agent pick the procedure to migrate and just do it at the same time, like we just talked about. But we can make this even more powerful, right? Because we're not just migrating to Effect for the sake of it. We're doing it because it's helpful for handling errors and for helping us instrument our code better. And so what we could do, if we want to get really clever, is look at our telemetry and figure out which procedures have the most errors or the least instrumentation, or have a gap in our APM, right? And when we send a control signal to our actuator agent, we can include not just the procedure to migrate, but also all the data about the things that we're trying to fix with this migration so that the actuator agent can actually make the code better instead of just doing a one-to-one migration. Oh man. There we go. So next is building our actuator. Our actuator is just an agent plus a skill. Bring your CLI coding agent of choice. You should spend a lot of time on the skill. Not all of that should be up front. You'll want to iterate on it over time based on what works. At HumanLayer, we like to build out what we call golden patterns by hand before setting the agent loose. These are just idiomatic handwritten examples for the agent to follow, because they're just pattern replicators, and otherwise you're getting what's in the docs or what the agent knows from the internet. And so we pipe the skill plus our control signal into our actuator agent. And the skill, of course, should include a response template. And the agent's going to work and work and work. And it'll produce a final response. And then we're going to deterministically commit and push and create a PR, using the final message as our PR description. Now all we have to do is actually run the loop, right? My recommendation is to use GitHub Actions or your GitLab or your CircleCI or whatever else you're using, because it has access to your code, it has access to your secrets, and it has great dispatch and scheduling primitives, right? We don't need a new cluster for this. So we can write a workflow that runs a single iteration of the loop: sense, control, actuate, and create a PR. And then we can schedule this to run once a day. And every morning we walk into the office to a small incremental PR that's low risk. And when we first did this, it was actually really frustrating, and we turned the loop off. Because we had to constantly update the skill, we had to constantly check out the branch, change the skill, change the code, commit and push, and our loop was actually really high friction, right? But there's a better way to do this where we can put a human on the loop in a really low-friction way to re-steer it when it goes wrong. And the way to do this is to just create a feedback file that's tracked in version control, just as a markdown file, right? We can deterministically load it into our actuator agent's context every time that it runs after we run the controller. Then we can add a label to the PR, right? Each workflow needs to be able to identify PRs that it created, since there might be a bunch of different loops running, and we only want workflows to respond to feedback from comments on their PRs. And we're going to add a comment trigger to each loop workflow. So when a user leaves a /iterate comment on the PR, the loop workflow is going to pick that up. It's going to deterministically load all of the PR context, the diff, the comments, the review comments, and the description into the agent's context along with the skill. And it's going to instruct the agent to fix the code, but also to update that feedback file, right? It looks kind of like this. And the benefit of doing this this way is that now that feedback file with instructions is tracked in your version control. You can see how you've changed it over time. You can revert it if you need to. So the next thing we're going to do is add flow control. Because the other problem that we had when we did this was that if we were at a customer site for a week, or if we were traveling, or spent six days working on slides instead of writing code, the PRs from all of our loops would just stack up. They'd duplicate work. They'd conflict. And we wouldn't get around to doing it. And the loop's work is important, but it's not that important. And so now we just had all this junk we had to deal with that wasn't important. So this is actually a really easy problem to fix. Because each loop and its workflow has a label that gets attached to PRs, when the workflow first runs, before we check out the code and install the dependencies and run our sense, control, and actuate steps, we can just check and see if the last PR that we created, or any PR with the loop's label on it, is open. And if so, we just shut down, right? Because this means that the last time that a human reviewed the code from this loop was before the loop ran, right? No human reviewed the last output, so there's no reason to stack up even more work for humans to review. This way, we have exactly one PR at most open per loop at a time. And we wouldn't get around to doing it. And the loop's work is important, but it's not that important. And so now we just had all this junk we had to deal with that wasn't important. So this is actually a really easy problem to fix. Because each loop and its workflow has a label that gets attached to PRs, when the workflow first runs, before we check out the code and install the dependencies and run our sense control actuate steps, we can just check and see if the last PR that we created, or any PR with the loop's label on it, is open. And if so, we just shut down, right? Because this means that the last time that a human reviewed the code from this loop was before the loop ran, right? No human reviewed the last output, so there's no reason to stack up even more work for humans to review. This way, we have exactly one PR at most open per loop at a time. No stacking, no duplication, hopefully no conflicts. And of course, once you're feeling confident in the loop, we're going to want to speed it up, right? I have 150 RPC procedures to migrate. If I do one at a time, it's going to take six months, which is way longer than I want to wait. Fortunately, there's a lot of ways to pick up the velocity of our loop. We could have our controller pick three procedures to migrate instead of one at a time, or five. We could have our controller pick three or five and then do each of those in a separate implementation phase, which will be both cheaper and more reliable since each migration gets its own context window. Or we could just run the workflow four times and give one PR to each of four people on the team. So let's put it all together. We built a control loop that improves our code incrementally, and we're actually reading the code. It has adaptive flow control, so we're not creating a bunch of loop or a bunch of work that nobody wants to review. And we can re-steer it on the fly in a super low-friction way. If you want to try this yourself, we built a skill. Please try it out. My Twitter handle is down there on the bottom. Please share it. I would love to see what you build. And if you get excited by this, at HumanLayer, we're hiring here in San Francisco. And if you're working on mission-critical systems and want to figure out how to get more out of AI, we'd love to chat. Thank you so much. Thank you. to review. This way we have exactly one PR at most open per loop at a time. No stacking, no duplication, hopefully no conflicts. And of course, once you're feeling confident in the loop, we're going to want to speed it up, right? I have 150 RPC procedures to migrate. If I do one at a time, it's going to take six months, which is way longer than I want to wait. Fortunately, there's a lot of ways to pick up the velocity of our loop. We could have our controller pick three procedures to migrate instead of one at a time, or five. We could have our controller pick three or five and then do each of those in a separate implementation phase, which will be both cheaper and more reliable since each migration gets its own context window. Or we could just run the workflow four times and give one PR to each of four people on the team. So let's put it all together. We built a control loop that improves our code incrementally, and we're actually reading the code. It has adaptive flow control, so we're not creating a bunch of loop or a bunch of work that nobody wants to review. And we can re-steer it on the fly in a super low friction way. If you want to try this yourself, we built a skill. Please try it out. My Twitter handle is down there on the bottom. Please share it. I would love to see what you build. And if you get excited by this, at HumanLayer, we're hiring here in San Francisco. And if you're working on mission critical systems and want to figure out how to get more out of AI, we'd love to chat. Thank you so much. Thank you.