AI Engineer

Harness Engineering is not Enough: Why Software Factories Fail — Dex Horthy, HumanLayer

2063 summary words 9 min summary Watch video

Start with the signal

9 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: AI coding harnesses, agent loops, tests, and automated review can accelerate implementation, but they cannot currently prevent long-term codebase degradation because coding models are trained and rewarded mainly for short-horizon functional correctness rather than maintainability.
  • Why it matters: For teams operationalizing coding agents in production systems, treating code generation as effectively free and eliminating human review can create compounding architectural debt, rising incidents, and eventual recovery work that erases the apparent speed gains.
  • Best use: Use this as a practical operating-model argument for keeping humans accountable for design and code ownership while applying AI aggressively to planning, implementation, and review preparation.

Executive Summary

Dex Horthy argues against the "lights-off software factory" thesis: the idea that teams can let agents generate code, rely on automated tests and monitoring, and stop reading pull requests. His central claim is not that teams are using coding agents incorrectly or failing to spend enough tokens. It is that current models have a structural training limitation: they are optimized to solve bounded tasks and pass verifiable tests, not to preserve or improve a codebase's design quality over months and years.

He frames the problem through the evolution of the software factory. Agentic implementation has compressed coding from days to minutes or hours, but architecture, review, integration, and testing remain bottlenecks. The tempting response is to automate those controls too and route incidents and user requests directly back into an agent factory. Horthy says this fails in complex, rapidly evolving repositories because agents can make locally correct changes that increase coupling, add defensive patches, and make later changes harder.

The training argument is the substantive core. Benchmarks such as SWE-bench reward whether a patch fixes the stated issue and passes tests; they generally cannot score whether the solution introduced poor abstractions, unnecessary exception handling, type casts, or future maintenance hazards. Since architectural costs often emerge only after months, the reward signal is difficult to attribute back to a single coding episode. Emerging long-horizon and multi-PR benchmarks are progress, but Horthy does not believe judge models alone can reliably substitute for human quality judgment.

His alternative is not slower, traditional development. It is AI-assisted, design-first engineering: use AI to accelerate product clarification, system architecture, program-level design, and implementation sequencing before code is generated. For substantial work, teams should align on desired behavior, component contracts, types, method signatures, call graphs, and vertical implementation slices. This makes the resulting PR conform to an agreed plan, reduces rework, and keeps line-by-line review feasible. Small changes can still go directly to an agent; the key is to retain human ownership at the points where maintainability decisions are made.

Key Takeaways

  • Claim: The current push toward fully autonomous "lights-off" software factories is unsafe for complex production codebases because coding agents do not reliably maintain code quality over time. | Evidence: Horthy cites reported AI-agent-related outages, a Pharos AI report claiming lower pull-request review quality, more and longer review comments, more unreviewed merges, and increases in incidents and bugs per developer after broad adoption of AI coding tools. HumanLayer also tried going fully lights-off in July 2025 and eventually encountered problems that required humans to reopen and understand code the team had stopped reading. | Implication: Do not generalize success on greenfield demos or isolated tasks into a policy of removing human code ownership for business-critical systems. | Caveat: He explicitly excludes low-stakes vibe-coded projects from this critique; a side project used by a dozen people has very different constraints from a production system or complex brownfield repository.
  • Claim: More sophisticated prompting, harness engineering, agent loops, and review agents cannot fully solve the problem because the primary bottleneck is model training, not merely agent orchestration. | Evidence: He contrasts early CLI agents such as Aider and Codebuff with Claude Code's success, arguing that the decisive advantage came from a model lab training a model against the specific harness it distributed. He also references an OpenAI team position that harness builders without model weights and the ability to reinforce-learn within their harness remain disadvantaged relative to teams that own both. | Implication: Evaluate agent platforms separately on workflow orchestration and on whether their model training/evaluation regime is aligned with the long-term engineering outcomes you need. | Caveat: Harnesses and loops still improve execution and can raise the quality floor; the claim is that they do not eliminate the underlying limitation.
  • Claim: Existing coding-agent evaluation rewards local functional correctness rather than maintainability, encouraging patches that pass tests while worsening design. | Evidence: In the SWE-bench-style example, the agent receives a binary reward if old and newly introduced tests pass after its patch. The evaluator does not penalize unnecessary try/catch blocks, type-casting workarounds, excessive coupling, or other design compromises that can make future modifications riskier. | Implication: A high benchmark score or passing CI should not be treated as evidence that an agent-generated change is architecturally sound or safe to merge without meaningful review. | Caveat: Passing tests remains necessary and valuable; it is simply an incomplete proxy for quality.
  • Claim: Maintainability is intrinsically harder to train and verify than task completion because the cost of bad architecture is delayed and distributed across later changes. | Evidence: Horthy uses Martin Fowler's "shotgun surgery" code smell as the example: a codebase becomes difficult to modify because one change requires edits across many unrelated locations. The damage may only become visible months later, making it difficult to attribute a training reward or penalty to the original agent trace. | Implication: Treat maintainability as a long-horizon operational metric requiring direct governance, rather than assuming it will emerge from short-horizon test optimization. | Caveat: The speaker presents this as an informed assessment rather than a proven benchmark result, acknowledging that good benchmarks for an agent's ability to improve codebase quality do not yet exist.
  • Claim: The productive near-term model is not to abandon code review, but to move human judgment earlier so implementation and review are constrained by an agreed design. | Evidence: Horthy's proposed workflow is: product review of the problem and desired behavior; architecture review covering system design, component contracts, data models, and constraints; program design covering types, method signatures, layout, and call stacks; then vertical slices defining implementation order, multi-repo coordination, and checkpoints. | Implication: Build an explicit escalation threshold: autonomous handling for bounded low-risk tasks, but mandatory human-led design alignment before agents implement larger, cross-cutting, or production-sensitive changes. | Caveat: He says small changes can still go directly to an agent; the heavier planning process is intended for non-trivial work in complex systems.
  • Claim: Upfront AI-assisted alignment can make human code review faster rather than making it a throughput bottleneck. | Evidence: His operating heuristic is that roughly 30 minutes spent on pre-planning and alignment can save hours in review. A well-scoped PR is easy to inspect because the reviewer can compare it against decisions already made; by contrast, even a PR needing only 20% rework creates substantial intellectual and emotional load for both author and reviewer. | Implication: Measure review burden by rework rate, scope clarity, and deviation from approved design—not merely by PR count or elapsed review time. | Caveat: This depends on planning artifacts being concrete enough to constrain implementation, rather than being generic architecture documentation.

Detailed Brief

What a software factory optimizes—and where the lights-off model breaks

  • Claims: Traditional software delivery already separated slow implementation and review stages from earlier planning activities designed to reduce rework.; Agentic systems compress the implementation stage first, then attempt to automate review, regression testing, incident response, and intake of user requests.; Once every input is routed into an agent queue, the organization risks optimizing for volume of generated changes rather than quality and ownership of the system.
  • Evidence: The speaker traces a conventional loop from work tracking through implementation, automated and manual testing, pull-request review, production, user feedback, and monitoring-driven incident response.; The lights-off endpoint removes code reading while investing in automated tests, monitoring, rollout systems, and agent-generated remediation PRs.; He argues that agents often begin struggling after only three to six months of accelerated development, not solely in decade-old legacy systems.
  • Caveats: The cited industry observations are directional claims from the speaker rather than a full causal analysis of every reported incident or code-quality trend.; The transcript does not provide operational thresholds for when a repository has crossed into the high-risk condition.
  • Implications: The relevant control point is not simply deployment safety; it is whether the organization can still understand, safely change, and debug its own software after repeated agent-written modifications.; Rapidly changing young codebases may need maintainability controls earlier than legacy-code terminology suggests.

The evaluation frontier is improving, but does not yet close the maintainability gap

  • Claims: Longer-horizon code benchmarks are attempting to test beyond isolated bug fixes.; Better verifier design can discourage some obvious gaming behavior, but model-based quality judges remain limited as a full replacement for human engineering judgment.
  • Evidence: SWE Marathon from Abundant AI is described as using approximately 400-hour tasks, including reproducing all features of Microsoft Excel.; Deep SWE from Data Curve is described as using large tasks on open-source repositories that were not built in the real world and therefore are less likely to have been in training data.; Frontier Code from Cognition is described as evaluating multi-PR tasks and penalizing tests that do not fail on the pre-patch code, while also using a judge model for code-quality rules.
  • Caveats: Horthy acknowledges the distinction between benchmarks and verifiers, even while using benchmark structure to reason about likely training incentives.; His assertion that a model able to judge good code would already write it is a heuristic argument, not a formal impossibility claim.
  • Implications: Monitor progress in long-horizon, multi-change, and anti-gaming evaluation rather than relying on standard issue-resolution benchmark performance.; Even stronger automated evaluation should be introduced as a guardrail layered into human-led engineering governance, not as a justification to remove it.

Notable Concepts & Terms

  • Lights-off software factory: A development model in which teams stop reading generated code and rely on agents, automated testing, monitoring, and rollout systems to ship and remediate changes.
  • Harness engineering: The agent environment and orchestration layer around a model: tools, loops, sandboxes, prompts, computer use, and execution workflows.
  • Shotgun surgery: Martin Fowler's code smell in which a seemingly local change requires edits across many areas of a system; Horthy uses it as a concrete manifestation of lost maintainability.
  • SWE-bench: A coding-agent evaluation pattern based on resolving real repository issues and passing tests; the talk uses it to show why short-horizon correctness rewards omit design quality.
  • Program design: A planning layer between architecture and implementation that specifies types, method signatures, program layout, and call stacks before an agent writes code.
  • Vertical slices: An implementation plan that sequences end-to-end increments, including multi-repository coordination and validation checkpoints, instead of broad horizontal task plans.
  • Brownfield: Conventionally an old legacy system, but Horthy broadens it to include fast-moving repositories that become difficult for agents to navigate after only a few months of accumulated change.

Operator Notes / Why Ken Should Care

  • Define a change-risk policy that permits direct agent implementation only for bounded, reversible, low-blast-radius work; require explicit product, architecture, and program-design artifacts for cross-cutting changes.
  • Add codebase maintainability indicators to the AI-development scorecard: rework after review, number of files/components touched per feature, incident recurrence, time to diagnose failures, and rate of architectural exceptions or workaround patterns.
  • Do not use merged PR count, test pass rate, or agent benchmark performance as the primary approval criteria for autonomous production changes.
  • Require that design artifacts specify component contracts, data ownership, types/interfaces, call paths, and validation checkpoints sufficiently concretely for reviewers to compare generated code against intent.
  • Treat fully autonomous incident-to-PR pipelines as an experiment behind strong blast-radius controls, not as a default production operating model.
  • Evaluate coding-agent vendors on their support for collaborative planning, review traceability, and long-horizon quality verification—not only on code-generation speed.

Source/Metadata

  • Title: Harness Engineering is not Enough: Why Software Factories Fail — Dex Horthy, HumanLayer
  • Transcript words: 4999
  • Duration seconds: 1157
  • Timestamp note: No timestamps or chapters were present in the supplied transcript. The transcript also contains a repeated closing/workflow segment.
Full transcript 3855 words · 22 min read
0:00

Music Music Music Music Music Music Music Music Music Music Music Music Guys, give it up for all the great speakers today so far. All right. This is Harness Engineering Is Not Enough and Why Software Factories Fail. And we're going to click, maybe. Oh, that's way too many slides. Hold on, guys. Okay. So we're all racing to put AI coding into production. And there's been lots said about loop engineering. And we should probably write more loops. And, yeah, I don't know. I guess we're doing loops now. Strong DM built a Lightsoft software factory where nobody even reads the code. And the prevailing narrative is we should just spend more tokens. You are the bottleneck.

0:14

The models are good enough. Code is free. Just ship more stuff. But at the same time, we are starting to see the cracks. Our friend Mario at AI Engineer Europe begged us to slow down. Because companies that should not be having outages because of coding agents are having outages due to coding agent mishaps. Code bases are falling apart faster than they ever have before. And our friends at Pharos AI even did a report since we all picked up all these AI coding tools in January, maybe February. Pull request code review quality is way down. We're having more comments, longer comments, and tons of PRs being merged without any review at all. Incidents are way up.

0:14

Bugs per developer are way up. And many people will tell you that you're holding it wrong. That's the only reason. You're not. Well, maybe you are. But that's not the point. I've spoken a lot about how to hold it better when it comes to working with AI. Probably a million views on YouTube at this point across a bunch of different talks. And the basic thing is, as engineers, we've been told that if token maxing isn't working, then it's a skill issue. You just need to spend more tokens. Let go of reading the code. That with enough harness engineering, if we maybe sprinkle some magic words, adversarial review, on enough of our PR bots, we can get the best of both worlds.

0:16

10 to 100x faster, high quality, and nobody has to do that thing we all hate called code review. I'm here to convince you today that this is, in fact, not a skill issue. That no amount of harness engineering or loops maxing can solve what is fundamentally a model training issue. That's why we say the harness is not enough. And to understand this, we have to grapple with and dig into how coding models are trained. I'm going to talk about what I think the shortcomings are with some of the current benchmarks and what better ones might look like. And we'll talk about how to move faster safely in the meantime. It's going to sound like a rant, but there is hope here.

0:16

I'm going to talk about our journey and a bunch of the landmines we've hit building in this world. A bunch of exciting new techniques that we've been working with a lot of our users and customers to develop. And, I think, how we all, as a community, get to the next chapter of agentic engineering after whatever this thing that we're in. So we use a lot of words here.

0:25

I'm going to zoom out a little bit. I want to give you a brief history of the software factory. And actually, I just learned this last week. The term software factory was defined at a NATO conference in 1968. We're going to start around 2022, right before AI started coming around. And basically, in a typical 2022 software factory, you will have some people building stuff. You'll have engineers, you'll have PMs. Maybe you have some sort of leadership team that is driving the vision here. And they all decide that stuff needs to get done.

0:28

And so you put it in a tracker, a Linear, a JIRA, a beads, some sort of state machine that tracks what needs to be done. And then someone goes and grabs something off there and they build the thing. And there may be some automated testing in that process. Maybe some manual testing in that process.

0:35

At a certain point, we make this pull request thing. It says, okay, cool, we've got to run a bunch of checks, automated stuff. A human's going to review the change and review the code.

0:46

And perhaps we might even have a human pull it down and test it somehow. And if anything goes wrong here, we loop back to someone builds the thing. And eventually we're ready for prod. And so we ship it to production. And once it's in prod, it makes contact with our users. And users do a thing that we all love. Users love to complain. I love our users.

1:06

But, yeah, they're going to ask for things. They're going to find bugs. They're going to file feature requests. And that goes back to your team. You might also add monitoring. And so, what do we want more than anything else? We want to wake up engineers at 3 in the morning when something breaks. So they can get dragged out of bed to try to go fix it. And we go on and on in this loop. And we ship a bunch of code. And one thing that we noticed here is that teams figured this out decades ago, that this someone builds the thing step is usually going to take hours or days in most cases. And the review part will also take hours or days for large things.

1:51

And so teams started doing this upfront planning, architecture proposals, sprint planning, and would collaborate on these things as a team with the hopes that we might decrease the percent chance that something would need to be reworked, that we would be able to reduce the time spent reviewing every line of code because we aligned on everything ahead of time. This brings us to the agentic software factory. Every company and their mother is talking about how they built a coding agent factory that ships 75% of their code now. Literally everybody. And so if we look at the software factory from 2022, we just replace someone builds the thing with an agent builds the thing.

2:05

And we have an orchestration and a harness and a sandbox and a model and computer use. And I'm not going to get into the details of that. You can watch 100 talks about that this week, I'm sure. But now the building part takes minutes or hours, but this human part still takes hours or days if you're going to review the code and you're going to test the changes. And so we bring in agentic code review. And we bring in agentic regression testing. And it makes this part faster. But it's probably still the bottleneck. But we can do more loops here. Why not? Let's do some more loops. So we can route all incidents straight into the factory.

3:00

Why does someone need to get woken up and try to fix it when they can just wake up to a pull request? And maybe that fixes the issue for you. You can take all the user feedback and just stick it straight into the factory so that people ask for stuff and it gets built. And now your only job is how much things can you stuff into the queue of stuff to do and how fast can you review and test the changes? Which brings us, of course, to, I'm sure you know, the Lights Off software factory, where basically Dan Shapiro coined this as we no longer read the code. We say, you know what? This is going great. That code review thing? No thanks. We're just not going to do that anymore.

3:38

And we invest into all these other parts of the system. Your testing, your monitoring, your rollout, everything else. We just write more code and build those systems better. And now our job really is just how much stuff can we ask the agent to build? I am going to posit that this does not work. And this is why software factories fail. As an aside, what I'm going to say has nothing to do with vibe coding. So Addy had this great post. I'm just going to literally take his quote verbatim. A developer vibe coding a side project a dozen people will ever run and a team keeping a 10-year-old enterprise system alive for another quarter share almost no constraints worth naming.

4:17

And most of what you hear on the internet is one of these groups of people telling the other group of people how to live their lives. So if you love vibe coding, please go on. At Human Layer, what we care about is how do we help people solve hard problems in complex code bases. We use the word brownfield a lot, which historically has meant some 10-year-old Java thing. I actually think agents really start to struggle after maybe three to six months, especially with the pace at which we can ship now. You can ask me how I know this, and I will tell you that it is because in July 2025, we tried this. We went full lights off.

4:37

And if you have tried this seriously for a number of months, you probably found at least one issue that the agent couldn't solve. Even with your most advanced prompting, you do research, you do reproductions, you just have to go and dig into that code base that you stopped reading three months ago to try to figure out what's broken. So if you love vibe coding, please go on. At Human Layer, what we care about is how do we help people solve hard problems in complex code bases. We use the word brownfield a lot, which historically has meant some 10-year-old Java thing.

4:50

I actually think agents really start to struggle after maybe three to six months, especially with the pace at which we can ship now. You can ask me how I know this, and I will tell you that it is because in July 2025, we tried this. We went full lights off. And if you have tried this seriously for a number of months, you probably found at least one issue that the agent couldn't solve. Even with your most advanced prompting, you do research, you do reproductions, you just have to go and dig into that code base that you stopped reading three months ago to try to figure out what's broken.

4:57

And in the meantime, your site was down, your users were pissed, and if you were like me, you were probably miserable reading all this slop code that you let slip into your system. And what I want to get to is models have a shortcoming. They can't maintain an improved code base quality over time, not without a decent amount of human steering. And when I say maintainability, I'm talking about issues like it becomes really, really hard to make a change in one part of the code base without breaking other parts of the code base. This is Martin Fowler's shotgun surgery, textbook code smell. I'm not going to say much more about maintainability.

5:13

There's a bunch of books that you can go read about it. In fact, John Osterhood is actually here speaking this week, so you can go ask him in person about the philosophy of software design if you want to. But it brings us to this question of why can't models do software maintainability? And you may also be saying, but Dex, surely the models have gotten much better since then. They've gotten better in some ways, but they're still about the same in others. If you want to solve one-off problems or vibe code a new marketing site, yes, they got way better since 2025 and 2024. But as far as improving code base quality, I think they have not gotten much better.

6:07

Now, I cannot prove this because there are no good benchmarks for a model's ability to maintain code base quality. And I'll get into where we're going with that. But if you've worked with coding agents for a while, a lot of people are posting about this. You probably have this vibe that they generally make things worse over time and make the code base harder to work in. And to figure out why this happens, I want to zoom out to the first great coding agent. Why did Cloud Code go from nothing to 4 billion, and I think now they're at 9 billion in revenue, in under a year? Because there were great CLI agents before Cloud Code. You had Ader, you had Codebuff.

6:31

There was a bunch of tools in this category. They had all the same tools: read, write, edit, grab, bash. So what was the difference? The difference was that this was the first time that a model lab trained a model against the harness that they were going to distribute it to users in. And it got really, really good at, this is just some of the tools, but it got really, really good at calling these sorts of tools in an agentic loop.

6:41

In fact, the OpenAI team did a talk in November about, if you are a harness builder and you don't own the model weights and you can't RL the model in your harness, you will always be at a disadvantage compared to somebody who owns both the model and the harness. And I'm going to cite a couple slides from my buddy Calvin French Owen, who was a MTS on Codex during the initial launch. But LMs are just next token predictors. This is a slide from over a year ago where, as you're doing your agentic loop, context window goes in, next step comes out. And we're going to try to do this.

7:05

I haven't actually timed this, but we're going to see if we can do coding agent reinforcement learning in 60 seconds. So what we're going to do is we want to train a model to get better at tool calling, better at solving software problems. We're going to generate a bunch of, we're going to give it a problem and we're going to generate a bunch of traces. Try to solve the problem a bunch of different times. We're going to score them all on correctness and did the test pass and all this stuff. And then we're going to reinforce. We're going to make the bad behavior less likely, and we're going to update the weights to make the good behavior more likely.

7:19

One of the classic ones here is SWE-bench Multilingual. They're about 15-minute tasks. They're from open source repos like Redis, JQ, and Django and all this stuff. And they have binary one or zero rewards on: did you fix the problem you were trying to fix, and did you do it without breaking anything else? And if we look at actually a real problem from one of these benchmarks, this is Fastlane, which is a Ruby project. Basically there was some issue where we weren't checking for nil, and we have a stack trace blowup because you have a null pointer exception.

7:38

And in this benchmark, you have a base commit that we're going to check out before the issue was solved by a human in the past. We're going to give it a test patch that says here's what the behavior should be afterwards. We have a golden patch. Both of these are hidden from the model. And so we have the agent go try to solve the problem. We store its patch. We undo all the changes it made to any test files, because I'm sure you've seen models comment out tests just to get things working. And then we're going to apply our golden test patch. And then we're going to run the test. Did the old test and did the new test pass? And if they both pass, then we get the reward.

8:42

Otherwise, we don't. And so models are trying to get the test to pass. There's no way in this system that we can penalize it for poor program design or for eroding the maintainability of our systems. That's why we get things like this. Try-catches around things that probably don't need a try-catch. Or things like this. I think Vi Bob gave us this example earlier of casting things to other things just so the model can, just, just, just wants to get the test to pass. And so if you can't verify the maintainability of the code, it gets way harder to train on this stuff. So remember this picture.

9:28

Verifying code quality and maintainability is orders of magnitude harder than the code runs and the tests pass. Because the cost function of bad architecture is measured in months and years. If you have a coding episode and then you only find out months later that somebody vibed this a little bit too hard, it's really hard to propagate that reward signal back across the gap. And now the frontier is getting better, slowly. And since I know someone's going to be in the YouTube comments about this, yes, I know benchmarks and verifiers are different and they actually have to be separate data sets.

9:55

But they're shaped the same, and the structure of these benchmarks is directionally correct. So we're going to look at these as what is the future of evaluating code maintainability? There's a really cool one called SWE Marathon from Abundant AI where they do 400-hour tasks of clone all of Microsoft Excel, every single feature. And they have some sophisticated reward channel stuff. Deep SWE from Data Curve is also large tasks on OSS repos that are not actually in the training set because they were never actually built in the real world. And then you have Frontier Code from Cognition, which is multi-PR tasks.

10:23

They do interesting things like, hey, if the model writes tests that don't fail on the pre-patch code, then it gets penalized. And we have a judge model that says, okay, did this follow all of our code quality rules? So we're getting better. But I think models judging quality can only go so far. Because if the model knew what good code looks like, it would probably write it in the first place. And review agents and throwing more tokens at the problem can raise the floor. But we're still constrained by what we can teach during RL. And so I will posit that for now we're stuck reading the code. But we can still move pretty fast.

11:11

And of course, there's a world where this is solved in the future. And if you want to just keep YOLOing prompts until you get to GPT-70, you don't have to think about this. By all means, please. But bitter lesson be damned, we've got some problems to solve. So let's engineer our way out of this. So turning the lights back on, we're going to put the code review back. We're going to embrace this approach of how do we plan up front to reduce the chance that we have a long or difficult review process. We're going to find leverage, we're going to use AI to help with this. And review agents and throwing more tokens at the problem, it can raise the floor.

11:41

But we're still constrained by what we can teach during RL. And so I will posit that for now we're stuck reading the code. But we can still move pretty fast. And of course, there's a world where this is solved in the future. And if you want to just keep YOLOing prompts until you get to GPT-70, you don't have to think about this. By all means, please. But bitter lesson be damned, we've got some problems to solve. So let's engineer our way out of this. So turning the lights back on, we're going to put the code review back. We're going to embrace this approach of, how do we plan up front to reduce the chance that we have a long or difficult review process?

12:07

We're going to find leverage. We're going to use AI to help with this. The first thing we're going to do is we're going to do some sort of product review. Understanding what problem we're solving, what's the desired behavior, maybe looking at mockups. Here's a product review I was working on yesterday with a mockup of a new feature. Once we have our product review, we're going to, by the way, small stuff still just goes straight to the agent. But once we have the product review, we're going to also do architecture. System architecture, a lot of people have been doing this for a while. Component contracts, data models, constraints.

12:30

This is an example of a doc that we build to understand how these systems are going to fit together and what's the high-level picture of it. From there we do something that I think is really underemphasized in agent encoding these days, which is program design. I think people assume that once you get the architecture right, the model can just cook. But we often look into the types and the method signatures. The program layout and the call stacks. So here's some examples. I don't think you'll be able to read this one, but this is the level of abstraction we're at, is how are we actually going to lay this stuff out, and how are these systems going to interact?

12:44

Dylan Mulroy from Cloudflare talks a lot about how he's using these call graphs as part of his planning process. I think this is exactly right. And then once we've done the program design, we can do this thing called vertical slices, which is the order of implementation, multi-repo coordination, how are we going to build this across our entire system, and how are we going to check it along the way. I've talked a little bit about how models have horizontal plans. I won't go too deep into it. If you want to learn more about this, you can go watch our talk from AI Engineer Miami, a couple shots of a doc like this going through the tests and the steps in between each phase.

12:57

The main idea here is 30 minutes over here in pre-planning and alignment can save you hours in review. And so it's actually feasible to still read every line of code. Let me skip this part. The summary here is you don't have too many PRs. If you're drowning in PRs, you actually have too many bad PRs. Because a good PR is a joy to review. You're just reading through it, like, yep, this is great. This is what we discussed. This is what we talked about. But even if a PR needs 20% rework, which is generous for a lot of AI vibe-coded slop, it's an emotional and intellectual burden on both the reviewer and the submitter.

13:28

And so if you use model-assisted planning and alignment, your alignment is shorter because you used AI to get all the information at once. Your code review is faster because you aligned up front. And your coding is faster because AI did it. And so now you're actually really moving faster. But you're still reading everything and you're still owning the code. So, closing advice. It's easy to hear all this and be a little bummed out. I really like the world where we just YOLO everything and we can just not have to ever read code ever again. But we're engineers. And these are just constraints. And models are good at certain things and they're not good at other things.

14:04

And so go figure out how to solve problems given a set of constraints. Use loops. They're great. Go solve hard problems. Seek leverage. If you want to help with this, we're building a human layer. Human layer is an AI IDE and collaboration platform. It's building blocks for your software factory and soon-to-be better verifiers for software quality. We've got a big move for cloud code and codex-style collaborative workspace. It walks you through the workflows for doing this sort of work. And we are talking to design partners. We are hiring founding engineers here in San Francisco. And these slides are live. You can go get them right now.

14:48

You can try human layer at humanlayer.com. It's free for small teams. Go solve hard problems in complex code bases. Thank you all for your energy. think about this, by all means, please. But bitter lesson be damned, we've got some problems to solve. So let's engineer our way out of this. So turning the lights back on, we're going to put the code review back. We're going to embrace this approach of like, how do we plan up front to reduce the chance that we have a long or difficult review process. We're going to find leverage, we're going to use AI to help with this. The first thing we're going to do is we're going to do some sort of product review.

15:24

Understanding what problem we're solving, what's the desired behavior, maybe looking at mockups. Here's a product review I was working on yesterday with a mockup of a new feature. Once we have our product review, we're going to, by the way, we don't, small stuff still just goes straight to the agent. But once we have the product review, we're going to also do architecture. System architecture, a lot of people have been doing this for a while. Component contracts, data models, constraints. This is an example of a doc that we build to understand how these systems are going to fit together and what's like the high level picture of it.

15:52

From there we do something that I think is really under emphasized in agent encoding these days, which is program design. I think people assume that once you get the architecture right, the model can just cook. But we often look into the types and the method signatures. The program layout and the call stacks. So here's some examples, I don't think you'll be able to read this one, but this is like the level of abstraction we're at is how are we actually going to lay this stuff out and how are these systems going to interact. Dylan Mulroy from Cloudflare talks a lot about how he's using these call graphs as part of his planning process. I think this is exactly right.

16:25

And then once we've done the program design, we can do this thing called vertical slices, which is the order of implementation, multi-repo coordination, how are we going to build this across our entire system and how are we going to check it along the way. I've talked a little bit about how models have horizontal plans. I won't go too deep into it. If you want to learn more about this, you can go watch our talk from AI Engineer Miami, a couple shots of a doc like this going through the tests and the steps in between each phase. The main idea here is 30 minutes over here in pre-planning and alignment can save you hours in review.

16:56

And so it's actually feasible to still read every line of code. Let me skip this part. Basically the summary here is like you don't have too many PRs. If you're drowning in PRs, you actually have too many bad PRs. Because a good PR is a joy to review. You're just reading through it like, yep, this is great. This is what we discussed. This is what we talked about. But even if a PR needs 20% rework, which is generous for a lot of AI vibe coded slop, it's an emotional and intellectual burden on both the reviewer and the submitter. And so if you use model-assisted planning and alignment, your alignment is shorter because you used AI to get all the information at once.

17:38

Your code review is faster because you aligned up front. And your coding is faster because AI did it. And so now you're actually really moving faster. But you're still reading everything and you're still owning the code. So closing advice. It's easy to hear all this and be a little bummed out. I really like the world where we just YOLO everything and we can just, like, not have to ever read code ever again. But we're engineers. And these are just constraints. And models are good at certain things and they're not good at other things. And so go figure out how to solve problems given a set of constraints. Use loops. They're great. Go solve hard problems. Seek leverage.

18:15

If you want to help with this, we're building a human layer. Human layer is an AI IDE and collaboration platform. It's building blocks for your software factory and soon to be better verifiers for software quality. We've got sort of a big move for cloud code and codex style collaborative workspace. It walks you through the workflows for doing this sort of work. And we are talking to design partners. We are hiring founding engineers here in San Francisco. And these slides are live. You can go get them right now. You can try human layer at humanlayer.com. It's free for small teams. Go solve hard problems in complex code bases. Thank you all for your energy. . . . . .

19:10

.

19:15

. .

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note