AI Engineer

From Blind Spots to Merged PRs: Continuous Agentic Performance Optimization - May Walter, Hud

1841 summary words 8 min summary Watch video

Start with the signal

8 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Reliable autonomous performance optimization requires agents to investigate real production behavior at function level, apply domain-specific diagnostic skills and ROI guardrails, verify fixes against runtime evidence, and surface only a small number of human-reviewable opportunities.
  • Why it matters: This is a concrete control-plane pattern for moving agents beyond code generation into trusted, recurring production-improvement workflows without flooding engineers with speculative PRs.
  • Best use: Use it as a design reference for an agentic SRE/performance workflow: production-to-code context, constrained investigation playbooks, verification gates, prioritization, and a deliberately narrow human review interface.

Executive Summary

Mai Walter, co-founder and CTO of Hud, frames performance debt as an investigation-capacity problem rather than a fixing-capacity problem. Teams often know how to implement an optimization once identified, but cannot predict or justify the research time required to find it. The resulting cycle is reactive: issues are ignored until they become urgent, then fixed in crisis mode.

Hud's proposed answer is a weekly autonomous workflow that combines production traces, latency data, queries, and business criticality with code-level reasoning. A GitHub Actions job invokes an agent, gives it runtime intelligence through MCP, identifies and scores high-ROI bottlenecks, proposes a change, reruns relevant tests, and reports the verified opportunity in Slack for a human decision.

The central implementation lesson is that raw observability data is not sufficient agent context. Agents reason over functions and files, while conventional metrics are organized around services, endpoints, CPU, and memory. Hud therefore builds a function-level "prod to code" data model, captures deep forensic evidence only for anomalous requests, and supplies reusable investigation skills for recurring diagnostic questions.

Walter emphasizes that autonomous agent workflows need a far higher trust threshold than interactive coding assistance. A plausible fix is not enough: it must explain an observed production issue, be verified, be worth the review cost, and be filtered by impact, business importance, and change risk. The operational output should be one compelling, readable recommendation at a time—not dozens of auto-generated PRs.

Key Takeaways

  • Claim: Continuous optimization is valuable because it automates an engineering phase that usually never happens: recurring investigation for small but meaningful production improvements. | Evidence: Walter describes the common pattern in which teams defer slow-page or stability work because investigation may take anywhere from an hour to a week; only when degradation becomes severe does the issue receive emergency attention. | Implication: The best agent opportunity is not merely accelerating an existing task; it is creating a recurring discovery loop for neglected operational debt.
  • Claim: Production-to-code context is the prerequisite for trustworthy performance agents. | Evidence: Hud maps endpoint, event-consumer, and cron-job behavior down to individual functions and outbound calls, allowing an agent to trace a seven-second endpoint delay to a specific function, database call, LLM call, or microservice dependency. | Implication: For agent systems that modify production code, invest in a context layer that translates telemetry into the same function/file abstractions used by coding agents. | Caveat: The speaker argues that conventional service-level metrics alone leave a reasoning gap; this approach depends on having sufficient instrumentation and a reliable mapping from runtime behavior to source-level entities.
  • Claim: More telemetry does not automatically improve agent performance; forensic evidence must be selectively retrieved and structured. | Evidence: Walter says runtime context commonly has both failure modes at once: large volumes of low-signal data and missing logs, traces, or metrics needed to explain an incident. Hud captures deeper evidence only when requests exceed a configured P99 or other threshold. | Implication: Design retrieval around anomaly-triggered evidence and specific investigative questions rather than dumping broad observability datasets into an agent context window.
  • Claim: Agent reliability improves materially when diagnostic methodology is encoded as reusable skills rather than left to ad hoc query generation. | Evidence: Hud found high evaluation variance when agents repeatedly had to formulate complex ClickHouse queries. It added skills such as tracing an HTTP 500 to its origin and comparing processes on memory-spiking pods against a baseline. | Implication: Treat agent troubleshooting as a library of tested investigative procedures, not a single broad prompt asking the model to inspect production data. | Caveat: The skills are tied to the organization’s telemetry model and database semantics; they must be maintained as the stack and failure modes evolve.
  • Claim: The workflow must reject plausible-but-unverified and superficially safe fixes. | Evidence: Early agents suggested changes that theoretically could explain a slowdown but did not actually resolve the production behavior. Walter also flags the common "lazy fix" of catching an exception instead of determining why it was thrown. | Implication: Require causal grounding and post-change validation; do not treat syntactic plausibility, passing tests alone, or exception suppression as sufficient evidence for an autonomous remediation recommendation.
  • Claim: Prioritization should optimize for ROI and reviewer attention, not for the count of possible fixes or even raw technical impact. | Evidence: Rather than opening 80 PRs, Hud starts with one recommendation at a time and scores hot-path frequency, business criticality such as payment or signup flows, and implementation risk such as migrations. Its example contrasts an endpoint normally at 200 ms but occasionally at 45 seconds due to using MongoDB distinct instead of a search index. | Implication: Build a ranking policy that explicitly prices review burden and deployment risk, then present humans with concise evidence designed to earn prioritization. | Caveat: The scoring depends on business-specific definitions of critical flows and acceptable latency, which cannot be inferred solely from generic telemetry.
  • Claim: Autonomous engineering requires roughly 80–90% trust, a materially different standard from using an agent interactively in an IDE. | Evidence: Walter argues that an 80% useful coding assistant can work when an engineer is present to steer it, but an unattended recurring automation will lose trust if it generates low-value or incorrect work. Hud’s goal is to hand off only issues that are worth fixing and fixes that are runtime-verified. | Implication: Separate interactive-agent evaluation from autonomous-workflow evaluation; measure precision of surfaced work, verification quality, review acceptance, and ongoing trust rather than only task completion. | Caveat: The 80–90% figure is presented as the speaker’s operating threshold rather than a measured benchmark or universal standard.

Detailed Brief

Reference workflow architecture

  • Claims: The workflow was designed to be vendor-neutral across execution harnesses, compute environments, and models because agent tooling changes rapidly and customers may use different coding agents.; The system needs explicit security around tool calls, permissions, and authentication, plus triggers for both scheduled reviews and event-driven investigations such as SLO breaches.; A maintenance and feedback loop is a first-class requirement because agent logic, team expectations, and reliability standards change over time.
  • Evidence: Hud’s initial implementation uses GitHub Actions on a weekly schedule, Claude Code as the selected agent, MCP to access runtime intelligence, and Slack as the delivery channel.; The same conceptual setup could use Cursor or Copilot and could deliver to Teams or email; the implementation choices were presented as examples rather than requirements.; The runtime query layer is based on ClickHouse over functions, endpoints, and forensic events.
  • Caveats: The talk is a vendor presentation and does not provide quantitative benchmarks for fix acceptance rate, false-positive rate, latency reduction, or engineering time saved.; ClickHouse-specific query complexity was a real source of variability in this implementation; a different telemetry stack will need equivalent semantic abstractions and tested access patterns.
  • Implications: Keep orchestration, model selection, data access, and notification channels replaceable rather than embedding the optimization logic inside one agent vendor’s proprietary workflow.; Treat permissions and tool authentication as architecture requirements from the start, since this workflow reads production evidence and can generate code changes.

Candidate optimization classes and human-facing output

  • Claims: Some performance smells are especially suitable for recurring, evidence-led detection because they are both common in mature codebases and relatively straightforward to remediate.; The human-facing artifact should explain the production symptom, probable cause, proposed change, and business relevance in a compact form.
  • Evidence: Hud explicitly searches for artificial delays such as timeouts and sleeps, N+1 queries, missing indexes, and sequential asynchronous work.; Walter characterizes older codebases with decades of history and hundreds of contributors as likely to contain many such low-hanging opportunities.; The report lets a reviewer choose among creating a ticket, fixing the item independently, creating a PR, or simply inspecting the recommendation.
  • Caveats: A static code-analysis finding can identify a possible issue, but Walter argues it is inadequate as the sole basis for a recurring autonomous workflow because it may not reflect actual production impact.
  • Implications: Start with a narrowly defined catalog of remediations where detection signals and expected validation paths are well understood.; Make reports decision-oriented rather than agent-centric: the artifact should help an engineer or product owner decide whether the opportunity deserves attention.

Notable Concepts & Terms

  • Prod to code: Hud’s function-level mapping of production behavior back to the code abstractions that coding agents can reason over.
  • Runtime intelligence layer: The proposed context system combining function-level execution information with selectively collected forensic evidence for production investigation.
  • Plausible unverified: A failure mode where an agent proposes a technically credible change that does not actually explain or fix the observed production problem.
  • Forensic context: Detailed request-level evidence retrieved for abnormal executions, rather than continuously supplying all raw telemetry to the agent.
  • Agent skills: Reusable, structured diagnostic procedures that reduce variance in querying and interpreting production data.
  • Human-friendly report: A deliberately concise recommendation that makes the case for one prioritized optimization instead of opening a large batch of PRs.
  • Highest-ROI optimization: A ranking principle balancing invocation frequency, business-flow importance, expected impact, and implementation/review risk.
  • Autonomous trust threshold: The higher confidence bar required for unattended agent workflows compared with interactive IDE assistance, which can be corrected live by an engineer.

Operator Notes / Why Ken Should Care

  • Prototype a weekly production-to-code investigation loop on one bounded service or business-critical flow; do not begin with autonomous PR creation across the repository.
  • Define an explicit candidate score before implementing agent remediation: request frequency, customer/business criticality, estimated latency or error impact, change risk, and expected reviewer effort.
  • Create a small versioned skill library for recurring investigations—slow request attribution, error-origin tracing, database/query regression analysis, and resource-spike comparison—and evaluate each skill against known incidents.
  • Set hard promotion gates from finding to recommendation: observed production anomaly, source-level causal hypothesis, narrow proposed change, test result, and runtime or representative-flow verification.
  • Track recommendation precision operationally: percentage accepted for investigation, percentage merged, verified post-deploy improvement, reviewer time, false-positive categories, and trust degradation caused by noisy output.
  • Avoid bulk PR generation; initially cap the workflow to a single highest-ROI recommendation per reporting cycle and iterate on report readability and acceptance.

Source/Metadata

  • Title: From Blind Spots to Merged PRs: Continuous Agentic Performance Optimization - May Walter, Hud
  • Transcript words: 4017
  • Duration seconds: 1365
  • Timestamp note: No timestamps or chapter markers were present in the supplied transcript.
Full transcript 3697 words · 18 min read
0:00

Hi everyone. I don't know if you're familiar with what I'm about to show, but remember when someone from the product is saying that some page is too slow and maybe we can optimize it. And then someone from engineering would say, yeah, probably, but we'd have to dig in to find out. So we just leave it as is. And then a few weeks later, Jenny would say, no, no, no, but it's actually way too slow now. We have to prioritize it. How long is it going to take?

0:30

And we would say something like, well, somewhere between an hour and a week. We'd have to look at it to find out. And then when she asks who can dig into it, the answer is only Dave. He's the only one who knows that code, and everyone else who wrote it left a decade ago. And amazingly, this still happens in teams all the time because we would know how much time it takes to do a specific optimization, but we never know how long it's going to take to investigate it and to find it. And miraculously, every time we look, we actually find things that can be done around performance and stability, but we never stop to proactively look for them because it doesn't make sense.

1:15

So hi, I'm Mai, co-founder and CTO at HUD. We're building a runtime intelligence layer for coding agents that captures function-level context and deep forensic context on things that matter so that coding agents can help you fix what's going on with production. And as a part of that, we built an agent workflow that helps with continuously optimizing performance of production applications. And I'm going to share a bit more about what we built, what were the challenges along the way, and hopefully you can take something out of it and apply it in your day-to-day. So we're surrounded by a lot of agency. So we're surrounded by a lot of agent PRs that all look good to us.

1:57

And one agent wrote them, and the other one says, yeah, I went over it and it looks fine. And we still feel that urge to verify before we just merge it to production. And then when it doesn't work, we end up asking ourselves, how the hell did we get to this place? And if you feel like that, I want you to know that you're not alone. So Google just published their Dora metrics for 2026. And we can see that the biggest impact of AI adoption on engineering is individual effectiveness, or that feeling of, oh my God, I'm so fast. I can do everything in the world. And for me personally, lack of sleep is another symptom of that.

2:34

But the second one is software delivery instability. And the throughput is actually not impacted as much as we expected. So we feel more effective. We're more effective individually. But as a team, our throughput is the same, and our software breaks more often, which is not exactly what we were hoping for with this revolution just yet. But maybe there is an agent for that. And if we can build faster with AI and we can fix faster with AI, then we can get those gains that we were talking about.

3:03

So I'm going to start with why we even wanted to do this. Then we're going to go over the tech and the process, which I think is also really important, and then share some gotchas and takeaways along the way. So debt leaks faster than we can bail. We have these issues that we ignore because they're not important enough, then they degrade to the point where they are important enough. And then we reach this crisis mode where we are all hands on deck. We fix it in emergency mode, and then we go straight back to ignoring, which is a leaky bucket by definition.

3:37

And that mostly happens because the research phase is a black box. It could take an hour or weeks, and we have to pay that debt and make sure that we spend time and engineering time on it in order to even know what can be done about it. And that's hard. It's legitimately hard to prioritize something where you're not sure exactly what you're going to get out of it. But what if we can automate that investigation? So we can run on a weekly basis with real production context, analyze the sweet spot, and flag the high-ROI opportunities in a way that's scored.

4:13

So it runs automatically without us having to stop and do something about it. It has the production context in mind and can give high-ROI scored performance opportunities of the things that are easy and impactful. Kind of like that performance sprint that you run every few months, just automated. Now let's talk about how we can actually do it because the dream is very nice, but the devil's in the details. So first of all, we wanted an infrastructure for the agentic workflows that would be vendor-neutral in terms of compute and where it runs in terms of the harness and also in terms of the model.

4:54

Things are changing all the time. We wouldn't want to constrain ourselves to a specific vendor, specific model, or anything like that. And especially with HUD, we want our customers to be able to use whatever agents they wish. We also want it to be secure in terms of the tool calls, permissions, and authentication. We want some trigger system, whether those are scheduled runs like the weekly run or a set of webhooks. For example, if we see some SLO breach, we would want to investigate it. And we really wanted it to be easy to maintain and update over time.

5:29

I think one of the biggest learnings we've had with agentic workflows is that even if they work out of the box or we get to a point where we're happy about them, in time, we evolve and our expectations go up. So just being able to maintain and update those logics and build that feedback loop was very important for us so that it's reliable and that people actually trust the outcome. Specifically, we chose to work with GitHub agentic workflows for that. And we can choose whatever agent we want to work with and build the workflows on top of that. But there are many other great tools that could be used for that. So this is how it looks.

6:11

You can see that there's a description of the job and the goal and the analysis. And then we go over the GitHub repository and generate that weekly deep insight report analyzing production data and finding those low-hanging, high-ROI opportunities. So for this specific setup, GitHub Actions runs weekly, and it uses cloud code. That was our specific choice. And then captures the runtime intelligence overhead via MCP so that we can look at the different endpoints, connect to the function level of what happens there, and send a report to Slack. Of course, that could have been Teams or an email or anything like that. It could have been cursor or co-pilot.

6:57

It's just the setup that we started with to make sure that we have something that runs without us in the loop and sends that report to somewhere we actually live in, which is Slack. So we want to take the production context, the traces, the queries, the latencies, then analyze them with the agent, score and flag which opportunities matter, because if you can optimize something, but it runs every three weeks, or you can reduce 20 milliseconds, then it doesn't matter. And then the most important part here is the difference.

7:16

So the agent actually fixes it, reruns the test, and sees the impact that it had on that specific flow that was optimized so that the human gets something after we already detected it. We understood the root cause. We understood why it matters to the business. And we verified that the fix actually impacted that time. And then a human reviews that, and the loop is actually closed. So it's not, hey, I have this idea of something you can do. It's, here's something that works and we believe would make an impact on these specific business flows in production that are running 7000 times a week. And then the human gets in the loop as a review gate.

7:58

And of course, it didn't work out of the box, if you were wondering. So the first hurdle we had along the way is what we call plausible unverified. So the agent would suggest a fix. It sounds right. It would look fairly real. And then after we verified it, it just didn't work. And it's true that the agent suggested something that could theoretically cause that slowdown. But what we wanted is to ground it on what's actually happening in production. Second part was complex queries. We specifically use ClickHouse. It's just an amazing columnar database, but it's also slightly different than the classic SQL patterns. And I'll talk about that in a bit.

8:37

And third one is the lazy fix. I'm sure you guys also experienced that when something throws an exception and the agent says, well, maybe we can just catch that exception and then everything will be fine. But what we really want is to understand why that exception was even thrown in the first place or why the results are lagging. So those were things that we found that methodology could be very, very impactful with. So that's not just the understanding of the data and how to connect it. It's also being quite thorough on what we want to do in that process and building the playbook of how a senior engineer would do that. And of course, we need the context to be right.

8:59

So the problem with context, there are only two problems with context. You either have too much of it or you have too little of it. And I'll talk about that in a bit. And the third one is the lazy fix. I'm sure you guys also experienced that when something throws an exception and the agent says, well, maybe we can just catch that exception and then everything will be fine. But what we really want is to understand why that exception was even thrown in the first place or why the results are lagging. So those were things that we found that methodology could be very, very impactful with. So that's not just the understanding of the data and how to connect it.

9:43

It's also being quite thorough on what we want to do in that process and building the playbook of how a senior engineer would do that. And of course, we need the context to be right. So the problem with context, there are only two problems with context. You either have too much of it or you have too little of it. And what we found around runtime context, especially from production, is that more often than not, you actually have these two problems together because, on one hand, you have a lot of low-signal data that is hard for the agent to reason over.

10:12

And on the other, you might not have all the logs and traces and metrics that you need in order to investigate that issue, which would still leave some room for assumptions and theories that are not necessarily what our users are experiencing in production. And also, when we talk about metrics like service level, CPU and memory, or endpoints in the P90s, they are not connected to the function level. So our coding agents reason over code and they look at these metrics, and there are some relations between them, but they don't exactly speak the same language.

10:54

So when we ask questions about what's taking time and what can I do about it, we're often finding that there are some gaps there. And again, accuracy could drop from that. So what we did is what we call prod to code, which is to be able to explain what's going on in production at the same level that agents reason over because the agent context lies on a function and file level, not on a service and endpoint level. So our context is running on a function level, and it is also connected to the endpoint or event consumer or cron job that ended up starting this task.

11:15

So you can ask a question like, hey, this endpoint that sometimes takes seven seconds, where is the time spent and what can be done about it, whether that's a function or an outbound call to a database, an LLM, or another microservice? And with that, you basically have the complete function-level context for every single function, like this service map but on a function level, the different invocations, where they come from, how often they run, and the deep forensic context only when it's needed.

11:44

Only when we see requests that are taking longer than the P99 or some threshold that we can define, then we will capture that forensic evidence so that we can say, hey, here is an example of a request that took longer. Let's find out why. Let's find out why. And we also have the ability to see that on top of the code, which is the HUD, the heads-up display, but I guess it's a way to explain how that data set is actually structured in a way that is much more comprehensive for a coding agent to read and reason over. And then we talked about the complex queries and what to do with them.

12:30

So the basic layer gave us the HUD query language, which is basically ClickHouse queries over that structure of functions and endpoints and forensics. On top of that, we also added a set of skills. We found that sometimes just querying the data is enough, but being able to get to the right query and to ask it again and again really created a lot of variance in our evals. And the skills actually help work with that data so that agents can use it. So, for example, if we're talking about a 500, an HTTP 500, we would want to understand where that error came from.

13:10

If we're talking about a memory spike, we want to understand what was running on those specific pods at that specific time where memory was higher and compare it to a baseline so that we see the diff. And all of these things were extremely helpful to be able to be a bit more methodological around how that works and to create more consistent results. And on top of that, there are a set of automations. So if we have the data, the query language, and a set of skills, we can build automations on top of them, like auto-fixing issues as they arise or detecting code that isn't even running for the last 60 days and eliminating it.

13:49

So, for this automated performance improvement automation that we are talking about right now, for performance, one example of that is looking for artificial delays like timeouts and sleeps, n plus one queries, missing indexes, sequential asyncs, and so on and so forth. These are specific things that are much easier to find when you actually look for them. And in a code base that's 20 years old and has hundreds of contributors, it makes sense that you'll find quite a lot of those. And removing them is fairly easy and impactful.

14:30

And then when you're asking something like, why are my endpoints taking long, you can actually find the specific reasons and not just guess a bunch of static code analysis, which could get you some result. I'm not saying it's never going to work. But when you're talking about an automation, we have to think about how to build something that is robust enough for us to trust over time. And then we said, okay, now that our evals are looking good, we run weekly, we find real slow endpoints from production that are invoked, we score them, we test them, we verify them. Maybe we can just open tool requests and everyone will fix everything and the world will be amazing.

15:09

Well, that's not exactly how that worked. Because people are still people, and no one wants to wake up for a rain of 80 pull requests, as small as they can be. That's just not how people operate. And no one has time for that. We're too busy building other things, and it's fine. So what we actually do is we use that priority to make sure that we only flag the ones that matter. And we actually started with one at a time to create that appetite and that habit. So we look at whether this is a hot path in terms of how often it runs and how critical it is for the business. The business impact, as in if this is something that has to do with payments or sign-up flows.

16:06

Obviously, we are more sensitive to that. Like asking ourselves, would we be able to convince the product manager to prioritize it? And we also looked at the risk. If it's a risky change that requires a migration or anything like that, then obviously it would require more time from the human who's reviewing it. And in that case, we're not necessarily looking for the highest-impact ones, but for the highest-ROI ones. And because we automate the investigation, we can look at the impact and the risk together and only surface and require attention on the ones that we believe are the right ones that are worth the engineering time.

16:35

So instead of opening 80 PRs, we built this human-friendly report that basically tries to convince you that it's worth your while. Something like, hey, this endpoint, it's usually taking around 200 milliseconds, but every once in a while, it takes 45 seconds. And it happens because you're using distinct and not the search index of Mongo. There's a very short explanation of what's happening and what's the fix. And then you can either create a ticket and fix it on your own and create a PR or just look at it.

17:04

And we found that building these small gists that are humanly readable and easy to understand made a huge, huge difference because we're still living in this hybrid world where humans are in the loop. And we want to respect our place in people's lives and to make sure that we flag the things that really matter and that we have some confirmation, not only that the fix is good, but also that this issue is worth fixing. And then when you look at that endpoint and you deploy that change and all of a sudden it's flat again, it is pretty satisfying. And then next time you'll get that report, maybe you'll have a look, and it will be easier to convince you.

17:44

So four things that I learned that could be relevant for you. One is we need to define what matters, and the scoring and the guardrails are what makes this reliable. We can automate a bunch of things, and it becomes easier to just create some sloth. But when we start with what are the things that are worth it, even though it's much cheaper to fix these things than it was a few years ago, it's still not free. And therefore, we need to understand that it's worth the impact, the human reviewer, and the risk that it entails if it does. The second part is that a lot of what we're talking about today is accelerating developers and what they do.

18:46

And I think what's interesting about this experience is that we automated something that is not just doing it faster. We're automating a phase that just did not happen in the day-to-day life of engineers. No one actually stopped every week and had a look at whether there are low-hanging fruits that could be relevant. One is we need to define what matters, and the scoring and the guardrails are what makes this reliable. We can automate a bunch of things, and it becomes easier to just create some sloth. But when we start with what are the things that are worth it, even though it's much cheaper to fix these things than it was a few years ago, it's still not free.

19:28

And therefore, we need to understand that it's worth the impact, the human reviewer, and the risk that it entails if it does. Second part is that a lot of what we're talking about today is accelerating developers and what they do. And I think what's interesting about this experience is that we automated something that it's not just doing it faster. We're automating a phase that just did not happen in the day-to-day life of engineers. No one actually stopped every week and had a look on whether there are low-hanging fruits that could be relevant.

20:06

But now we can use the agents to not only do the things that we're doing faster, but also to help us surface opportunities that we would probably never do without them. And this sounds simple, but it's hard to actually apply it in the day-to-day. But context over cleverness works almost every time. If they have the right context, if they have the right skills, if they know exactly what they need, these agents get much, much more useful. And every time a new model comes out, things get better. And yet I do believe that, at least for most of the cases, the models are good enough already to be able to automate that.

20:37

But it is up to us to help steer and guide to the right directions to make sure that we get the results that we want. And they're not necessarily the absolute right thing because it doesn't necessarily exist. And when we know our domain and our business, we understand which issues matter more, what's worth fixing, what latencies actually impact our customers' experience in the most significant way. And that helps us make sure we focus on the right things. And in that aspect, agentic engineering is not like coding with an agent.

21:19

If something works 80% of the time and you're using it with your cursor and your IDE, that's fine because you're there, you're in context, and you can help fix and steer. If we're talking about an automation that runs autonomously, we have to have very high confidence that we're doing the right thing. And that it's not going to just stray off and hand us a bunch of things that we would either waste time on reviewing or just not be able to trust over time. And I think the hardest part of this automation was to get to a point where we feel confident enough that the issue is worth fixing and that the fix is verified in runtime.

22:03

So that it is handed to a person when we have fairly good confidence that it actually works. And yet they still review it, but we know it's worth their time. And getting to that agentic engineering automation level requires crossing towards the 80, 90% trust. And it's something that is dramatically different than using an agent directly as an engineer. So I hope that was helpful. And if there are any questions or anything like that, I'm available and always love to geek out on the AI in the SDLC. And looking forward to hear about cool automations that you build on your own. Thank you. Thank you.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note