Open Reader

How to Kill the Code Review — Ankit Jain, Aviator

completed 16:25 Aug 17, 2026 Watch on YouTube

Current Status

completed

Video ID

YgEv7IQzGdM

RAG / Chat

Enabled
How to Kill the Code Review — Ankit Jain, Aviator
Description

Over 30% of changes now merge with no review at all, and the wait on the ones that do get reviewed is four times what it used to be. Ankit Jain's read is that the debate about when we stop reading code line by line is already over, because we stopped. His sharper point is what replaced it: an AI writes the code, an AI reviews the code, the two go back and forth in a web UI, and a human skims the thread and merges. When AI reviews and nobody reads, he says, we have configured the wrong thing. He is also here to correct his own five layer trust model from a few months earlier, which missed that review was never only about correctness. It also carries knowledge sharing, mentorship, architectural feedback, and onboarding, and that half has to survive. Spec driven development does not rescue it, because a spec written up front with no feedback loop is the 1970 waterfall model, and the decisions that actually matter end up in the prompts, which teams throw away the moment the pull request opens. His proposal keeps them: capture the session, turn those decisions into acceptance criteria, pair that with a registry built from your own recurring review comments, and generate a test plan that a verification system runs against a live preview. The review surface becomes intent and evidence instead of the diff. Speaker info: - https://x.com/ankitxg - https://www.linkedin.com/in/ankitjaindce/ - https://www.latent.space/p/reviews-dead Timestamps: 0:00 - The five layer trust model, and what it got wrong 1:17 - We already stopped reviewing, and AI reviews nobody reads 3:05 - What code review was actually for 4:21 - Spec driven development is waterfall again 5:53 - Intent lives in the prompts, which we throw away 6:42 - The AI slop registry 7:58 - Session to acceptance criteria to test plan 11:51 - Deterministic where you can, and reviewing intent not the diff 14:12 - Homework: mine your last 1,000 review comments

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Skim
  • Core thesis: AI should not merely accelerate line-by-line code review; teams should shift review toward validating captured implementation intent and behavioral evidence, while codifying recurring defects as reusable automated guardrails.
  • Why it matters: As AI raises code volume and makes conventional PR review a larger bottleneck, the proposed model offers a useful control-plane pattern: preserve human governance and architectural alignment while automating repeatable semantic checks and verification.
  • Best use: Use this as a conceptual design input for AI-native engineering workflows, especially how agent-session context, acceptance criteria, test plans, and verification evidence could replace diff-centric approvals.

Executive Summary

Ankit Jain argues that conventional code review is already failing under increased code volume: review wait time has become the new delivery bottleneck, many changes merge without review, and AI-generated code reviewed by AI in a GitHub UI often leaves humans doing only a superficial final check. His central correction to a prior "five-layers trust model" is that code review is not principally about finding bugs; its enduring human function is alignment across teammates, including mentorship, shared architectural understanding, onboarding, and governance.

He rejects a simplistic version of spec-driven development in which a complete specification is handed to an agent and verified afterward. That reproduces waterfall assumptions: important decisions emerge interactively during implementation, specs are rarely updated, and LLM behavior is not deterministic. The key artifact to preserve is therefore not just the original ticket or formal spec, but the evolving intent expressed in the human-agent coding session.

The proposed replacement review loop converts decisions from the coding session into acceptance criteria, combines them with a maintained "AI Slop Registry" of recurring review findings and invariants, generates a test plan, and verifies it against a runnable preview. Reviewers then inspect intent, architectural choices, rejected alternatives, and behavioral evidence such as browser actions, screenshots, and database snapshots rather than scrutinizing every changed line.

The strongest practical idea is to mine recurring review comments and turn them into durable controls. The weakest part is that the talk remains product-led and conceptual: it asserts operational metrics and a verification architecture but provides no validation data, registry design details, ownership model, false-positive controls, or guidance for high-risk changes. It is worth skimming for the workflow pattern, not watching as an implementation playbook.

Key Takeaways

  • Claim: The bottleneck in AI-assisted software delivery is moving from code production to review, and conventional review is increasingly unable to provide reliable control. | Evidence: Jain cites 861% code churn, rising incident-to-PR ratios, review time increasing to roughly 4x prior levels, and more than 30% of changes merging without review. | Implication: Ken should assume that adding AI reviewers to existing PR workflows will not by itself solve governance; the review object and evidence model need redesign. | Caveat: The talk does not define the sources, populations, or measurement methodology behind these figures, so treat them as directional motivation rather than planning-grade benchmarks.
  • Claim: Human review should retain alignment work while tooling takes on more semantic-accuracy checking. | Evidence: The speaker frames review's human value as knowledge sharing, mentorship, architectural feedback, onboarding, and collaboration—not only bug catching—and says "alignment must survive" even as semantic accuracy becomes more tool-assisted. | Implication: Design human approval gates around intent, architecture, and policy tradeoffs, rather than expecting people to manually audit every AI-generated implementation detail. | Caveat: This model is aimed at collaborative engineering teams, not solo developers or fully autonomous "dark factory" workflows.
  • Claim: Static spec-driven development is insufficient for agentic coding because critical intent emerges during the interactive implementation session. | Evidence: Jain compares a spec-to-agent-to-verification workflow to waterfall: the original spec is written before unknowns emerge, teams continue interacting with coding agents to resolve ambiguity, and the newly made decisions typically never flow back into the specification. | Implication: Agent orchestration should persist session decisions, questions, constraints, and rejected alternatives as first-class artifacts that can feed verification and later review. | Caveat: He still calls spec-driven development valuable; his objection is to treating the initial spec as complete and deterministic.
  • Claim: Recurring code-review feedback should be captured as an "AI Slop Registry"—a growing set of organization-specific invariants and guardrails. | Evidence: His concrete homework is to mine the last 1,000 review comments for repeatable findings; each captured pattern is intended to prevent that same comment from needing to be made again on future PRs. | Implication: A team can convert reviewer judgment into compounding institutional controls, but should prioritize high-frequency, objectively verifiable comment classes rather than attempting to encode every stylistic preference. | Caveat: He explicitly warns of a J-curve: building the registry has real up-front cost before it reduces review effort.
  • Claim: The proposed review surface is a generated test plan plus verification evidence, not the raw code diff. | Evidence: The flow is: capture session responses as acceptance criteria; combine criteria with registry invariants into a test plan; run a preview; then present evidence that intended capability and acceptance criteria were met. For a payment-form change, he suggests agent browser interaction, screenshots, and database snapshots as evidence. | Implication: For AI-generated changes, require an auditable evidence bundle tied to acceptance criteria—especially for user-facing workflows—rather than treating a passing AI review comment as sufficient assurance. | Caveat: Screenshots and agent-driven browsing are useful evidence but do not prove all correctness properties; the speaker says the system should be deterministic where possible and use LLMs only where necessary.
  • Claim: Test-plan generation should draw from independent session-level intent rather than solely from code generated by the same agent. | Evidence: Jain argues that when the same agent produces both code and its test plan, it is unlikely to create tests that expose its own mistakes; he attributes a related observation to "Dex" from the prior day. | Implication: Avoid self-certifying agent loops: establish separation between implementation generation and the source of acceptance criteria, with human-authored or human-confirmed intent as the anchor. | Caveat: The transcript does not specify a rigorous independence mechanism, such as a separately prompted model, adversarial test generator, or human test-plan approval.

Detailed Brief

Proposed AI-native verification architecture

  • Claims: A coding session should become a durable source of requirements because it contains the decisions that distinguish useful human contribution from raw code generation.; LLMs can reduce the burden of drafting and maintaining test plans, while humans govern the criteria and interpret the resulting evidence.; The resulting workflow is presented as closer to behavior-driven development than classic test-driven development because the test plan is English-readable and can involve product managers and designers.
  • Evidence: The talk's pipeline is described as session decisions to acceptance criteria, acceptance criteria plus invariants to a test plan, then preview-based verification.; The product being piloted is Aviator Verify, positioned as combining the alignment side with semantic-accuracy detection through the AI Slop Registry.; The speaker distinguishes deterministic verification of measurable criteria from LLM fallback for tasks that cannot be fully deterministic.
  • Caveats: No details are given on how acceptance criteria are versioned, linked to tickets and code, resolved when requirements conflict, or protected from prompt/session contamination.; The claim that teams may no longer need to maintain tests for new features is aspirational and should not be interpreted as a safe replacement for durable regression, unit, integration, or security test suites.
  • Implications: The useful abstraction is an evidence-producing verification pipeline whose inputs are intent and organizational invariants, not a generic AI code-review bot.; Cross-functional review may become more practical if behavioral criteria are legible to non-engineers, but accountability for approving those criteria must remain explicit.

Notable Concepts & Terms

  • Alignment: The human-centered purpose of review: shared understanding, architectural feedback, mentorship, governance, and confirmation that a change reflects what the team intended.
  • Semantic accuracy: Whether the implementation actually behaves correctly and avoids known defect classes; Jain positions this as increasingly amenable to automated tooling.
  • AI Slop Registry: Jain's proposed repository of recurring review findings, policies, and invariants that becomes a compounding, organization-specific set of automated guardrails.
  • Session becomes criteria: The design principle of extracting decisions made during human-agent coding interaction and converting them into explicit acceptance criteria.
  • Review surface: The artifact reviewers inspect; Jain proposes shifting it from a line-level diff to intent, acceptance criteria, test plan, architectural decisions, and verification evidence.
  • Deterministic where you can, LLM where you must: A control principle: use deterministic checks for measurable requirements and reserve LLM judgment for tasks that cannot be specified or verified mechanically.
  • Self-certifying agent loop: The failure mode in which the same agent creates code and the test plan used to validate it, reducing the likelihood that the tests uncover the agent's own errors.

Operator Notes / Why Ken Should Care

  • Run a retrospective over a bounded sample of recent PR comments and classify each as: deterministic lint/static rule, invariant or policy check, behavioral acceptance criterion, architectural judgment, or non-actionable preference.
  • For one low-risk, user-facing workflow, pilot an evidence bundle containing human-confirmed acceptance criteria, a generated test plan, preview execution, browser evidence, and relevant state/database assertions.
  • Require implementation agents and verification agents to use separate contexts or roles; anchor verification in ticket, PRD, and session-derived criteria rather than implementation-generated rationale alone.
  • Define explicit ownership, versioning, exception handling, and deletion rules for any review-comment registry before it becomes a blocking control plane.
  • Do not remove existing regression or security tests based on real-time generated test plans until comparative escape-rate and false-assurance data exists.

Source/Metadata

  • Title: How to Kill the Code Review — Ankit Jain, Aviator
  • Transcript words: 4507
  • Duration seconds: 985
  • Timestamp note: No timestamps or chapters were present in the supplied transcript. The latter portion is substantially duplicated, so the unique substantive content is shorter than the stated word count.

Transcript

2520 words en Processed in 95.0s

OK. Ooh, hello. Hey, everyone. Thanks for joining in. Today, we will be talking about how to kill code reviews, everyone's favorite topic. I'm Ankit, co-founder of Aviator at Aviator VR, building an AI code verification platform. So we'll bring in some of the ideas and concepts that we build in our product. But first, let's dive in a little bit. A few months ago, I wrote a post on Laying Space about how to kill code review, creating a framework of a five-layers trust model. This model was focused around how do we actually, layer by layer, build trust into the code that can then be merged without needing line-by-line review. And I got some things right, and I got some things wrong. So this talk will be about diving a bit more into it. I'm not going to talk about specific layers, but we will talk about some of the concepts that emerged from this session. So let's talk about the problem. We are looking today at the volume of code increasing every day, and we are struggling to keep up. So when we think about how long it will take us to actually stop reading code line by line, the reality is we've already stopped reviewing it. There is 861% code churn. That means we are producing more code. The incidence-to-PR ratio is increasing. That means even if you are doing reviews today, they are not effective. So the median time of review is increasing. We have just increased the bottleneck from coding, now coding is solved, to now reviewing, where everything just gets stuck there. You're spending 4x the time that you were spending before just waiting for the reviews. And today, over 30% of changes are actually getting merged without a review at all. So let's think about it a little bit. Think about AI reviews. Everyone is probably using some form of AI reviews today. When AI writes the code and AI reviews the code, why are we doing it in a UI? Right? We open GitHub, there are maybe two or three AI coding agents who are doing the reviews. The review goes back past the code, and you're talking back to the user or the agent, and it gets resolved. And you're doing back and forth with the agent. Where is the human here in the loop? You're eventually just looking at, OK, if AI has reviewed it, most of the things probably will be found. Let's do a skimming of it and merge it. So when AI reviews and nobody reads, we have configured the wrong thing. So let's take a step back. Code reviews are not very old. They are maybe 15 to 20 years old. In 2006 was when Google launched Montrean internally, and they made formal code review a thing. If you think about it, Windows, back in the day, the first versions were actually built without reviews. But if we look carefully, code review is not just about code reviews. Obviously, we are looking at it as catching bugs, understanding conventions, identifying security issues. But code review is also about alignment. And this was one piece that was missing from my five-layers model that I talked about a few months ago. So a big part of code reviews is knowledge sharing, mentorship, architectural feedback, onboarding, being able to collaborate. Again, if you're doing white coding, you're working on a solo project, this is not a talk for you. If you are working in teams, which I believe most of you folks are, collaborating in teams, you're not likely using completely dark factories, orchestrators where nobody looks at the code. You're actually collaborating in teams. You need to do knowledge sharing, which is the alignment part. And that is the most important aspect of reviews. So for semantic accuracy, we can build better tooling, but alignment must survive. So let's dive into alignment. What does it mean? In today's world, can we actually think of a better model than aligning just based on reading line-by-line code? So most folks have probably heard about spec-driven development by now. Spec-driven development is, OK, we write a spec, it covers all the details, we pass it to an agent, it generates code, and then we verify. So what's wrong here? If you look back in 1970, this is what the waterfall model was. You have requirements, you have specification, you implement, and then you verify. But there's no feedback loop. The spec is written before we identify everything else. Right? That's why today everyone still wants to use coding sessions, whether it's cloud code, codex, cursor, whatever you're using, you want to interact with the agents. And the reason you're interacting with agents is because there were certain things that are not clear in the spec, and we still need to capture that. And second, as you implement, you identify more issues, and you never go back and update the spec. Because if you're doing spec-driven development, it's already done. Once the spec is done, you expect the code will come deterministically. But guess what? LLMs are not deterministic. They're going to make decisions themselves. So that's why spec-driven development is a great methodology, but it falls short in day-to-day software development. But there are some interesting aspects of this that we should carry forward. The most important part is the intent. And intent doesn't only live in the spec. Intent lives in your Jira ticket. That's the goal, right? It's where you express what we want to do. It lives in your PRDs. It's details. It's a plan. But most importantly, it lives in your PRDs today. This is where the real decisions are being made. You start with, OK, this is a Jira ticket I'm going to look at. But you're going back and forth with the agent, and this is where all the user decisions are being made. But what we do today is we create a change, we create a pull request, and then we throw away the PRDs. And this is one of the things that we need to change. Let's also talk about semantic accuracy, because you're saying, hey, OK, I understand the alignment part, but there are still bugs in the code. Who's going to look at that? LLMs are also not great at this. And we already talked about how AI agents, reviewers, may not also be perfect. So this is where I introduce you to the concept of AI Slop Registry. So think about this. If you're reviewing code today manually, and I expect everyone should be doing some degree of this, we are essentially possibly identifying the same issues over and over again. Can we actually capture these concepts and codify them so that we don't have to always create that review feedback one by one? You also have all of those things automatically identified. The beauty of this is if you do it a few times, you now build a system that actually learns over time. So think of this as you're doing more training on top of the standard LLM that you have actually extracted, built on top of. So AI Slop Registers now can create better results because they're trained, they're learning from the review experience that you as humans are providing. Every recurring comment is now a guardrail that you don't have to review again. OK. So let's try to put both of these, alignment and semantic accuracy, together. It is two halves of the same problem. We are trying to understand what are the core mechanics of review. How do we actually break it down into two parts, which is the alignment and the semantic accuracy, and bring them together into a single loop? So first thing is you take your session and you capture the user responses, and that essentially forms your acceptance criteria. The acceptance criteria then tied with your AI Slop Register that you are now constantly maintaining finally creates a test plan. And this is a test plan that then gets verified. This is part of the system that we are building, the verification system, where it spins up a preview, takes your test plan, and makes sure it actually works end to end. Even if the code looks right, does it actually work? So this verification part has become interesting. And now the kicker is this is now your review surface. You're not reviewing code line by line, but rather you're looking at the evidence of what was the intent, did the user actually implement the capability that was defined in their intent, and did the behavior actually meet the requirements that we had in the acceptance criteria. So you're still having the architectural decisions, you're still having these arguments, but the review surface changes. So just giving a walkthrough of how we have built our system, the session becomes the criteria. So all these decisions that you are doing here with the agent, you're asking, you're providing this feedback. Even if there is a simple task, many times you're going back and forth, the agent will stop to ask questions. These are the decisions that we need to capture. This is the intent. This is what makes your review, makes your collaboration, more valuable. This is how you teach your junior engineers how to improve over time. These are the decisions that make a software engineer valuable today. We convert those into acceptance criteria. This is where you can also leverage LLMs to do so. So you capture these user decisions and make sure you can actually create a test plan based on this. I know test plan creation is always painful, which is why I would always recommend people use LLMs for this purpose. And finally, the criteria plus invariant is what makes the test plan. And we build the verification systems to actually capture the test plan, run your previews, and be able to test based on this test plan. So even imagine if you're building a new feature, you don't have to maintain tests at all. This is creating tests in real time. And this is where you can leverage the power of LLMs because test plan maintenance and creation can be very painful. But the value of the human in the loop here is the governance and the review part, and the part where you're reviewing the test plan and not the code. Right? I know it's been over 20 years since we came up with test-driven development. This, in some ways, is closer to behavior-driven development, where the test plan is now something that even you can share with your product managers, your designers, everyone can participate in the test plan. And then the test plan is now in English. At the same time, we have deterministic verification that actually verifies whether the particular test criteria has been measured. Let's move on. So this is where I would say the system is not supposed to be perfect. It's deterministic where it can be, but LLM where you must. Not everything can be deterministic. Not every system can be built in a way that is 100% built on deterministic systems. This is where you use LLM as a fallback. Let me give an example. If you are making a change in your web application, what you can do is it creates a test plan of what the behavior changes. Let's say you introduce a new payment form. So it creates a new payment form. The verification here is does the payment system change. An AI agent can go and browse through your application to fill out a form and capture screenshots as evidence, and then take those screenshots, as well as your database snapshots, to identify whether the criteria was met. So the screenshot testing or sampling can then still be done by agents, but at the same time you're creating more solid evidence that now a reviewer can look at and build more confidence that this actually works. So now reviewers are reviewing the intent, not the diff. You're reviewing the intent decisions, what we set out to build, what we tried and rejected, and capture all of these things from the sessions. Remember, capturing it from the sessions is the key. If we try to build it from the code, you end up in the same situation that we were talking about. I think Dex was talking about yesterday, which is if your code is built by the same agent that is actually building a test plan, it's not going to build a test plan that will actually catch issues. So that's why it's important to actually use the session information to build out a test plan. You can discuss architectural decisions, how you're creating the data models, how these services interact with each other. So you've moved one level above. So instead of reviewing line by line, you're actually having discussions on architecture, which are very critical for any kind of collaboration. And then you look at the evidence, everything that was collected from the verification. So here's homework for everyone. Go home and mine your last 1,000 review comments and build out an AI Slop Register for the things that are repeatable. So a vast majority of the comments that we're providing in your code review are something that we repeat over and over again. This compounds with every merged PR. Every time you capture something as a register, you don't have to capture that comment again. And this is where you can actually codify some of the best practices of doing code, maintain semantic accuracy, and at the same time, do not lose the collaboration part of the review. It does follow a J-curve. So the pain is real. You will have to spend some time to actually make it pay off because initially creating a registry can take some time. And this is where I would recommend you folks come and try out our product. So code review is not just about code review. It is about really getting the alignment. And where we can build better tools is creating the semantic accuracy and defining your AI Slop Register. So if you remember one thing from today, remember code review is not just about code review. It is about getting the alignment. And yes, we are piloting our new product called Verify. Please join and be our early design partners. We are working with a few companies to pilot out a new verification system. This combines both the alignment side of things, as well as building tools and capabilities for detecting semantic accuracy using the AI Slop Register. Thank you, everyone. Thanks for joining. Thank you, everyone. Thanks for joining. that I talked about a few months ago. So a big part of code reviews is knowledge sharing, mentorship, architectural feedback, onboarding, being able to collaborate. This is, again, if you're doing white coding, you're working on a solo project, this is not a talk for you. If you are working in teams, which I believe most of you folks are, collaborating in teams, you're not likely using completely dark factories, orchestrators where nobody looks at the code. You're actually collaborating in teams. You need to do knowledge sharing, which is the alignment part. And that is the part which is the most important aspect of reviews. So for semantic accuracy, we can build better tooling, but alignment must survive. So let's just kind of like dive into alignment. What does it mean? Like in today's world, can we actually think of better model than aligning just based on reading line-by-line code? So most folks have probably heard about spec-driven development by now. So spec-driven development is like, okay, we write a spec, you know, it covers all the details, we pass it to an agent, it generates a code, and then we verify. So what's wrong here? If you look back in 1970, this is what waterfall model was. You know, you have requirements, you have specification, you implement and then you verify. But there's no feedback loop. You know, the spec is written before we identified everything else. Right? That's why today everyone still wants to use your coding sessions, whether it's cloud code, codex, cursor, whatever you're using, you want to interact with the agents. And the reason you're interacting with agents is because there were certain things which are not clear in the spec, and we still need to capture that. And second is, as you implement, you identify more issues, and you never go back and update the spec. Because like, if you're doing a spec-driven development, it's already done. Once the spec is done, you expect, like, you know, the code will come deterministically. But guess what? LLM is not deterministic. It's going to make decisions itself. So that's why spec-driven development is a great methodology, but it falls short in day-to-day software development. But there are some interesting aspects of this which we should carry forward. The most important part is the intent. And intent doesn't only live in the spec. Intent lives in your Jira ticket. That's the goal, right? Like, it's kind of like where you express what we want to do. It lives in your PRDs. It's kind of like details. Like, it's a plan. But most importantly, it lives in your PRDs today. This is where the real decisions are being made. You start with like, okay, this is a Jira ticket I'm going to look at. But you're going back and forth with the agent, and this is where all the user decisions are being made. But what we do today is we create a change, we create a pull request, and then we throw away the PRDs. And this is one of the things that we need to change. Let's just first talk also about semantic accuracy, because like, you know, you're saying, hey, okay, I understand the alignment part, but like there are still bugs in the code. Who's going to look at that? LLMs are also not great at this. And we already talked about how AI agents may, reviewers may not also be perfect. So this is where I introduced you to the concept of AI Slop Registry. So think about this. We are reviewing, if you're reviewing code today manually, and I expect everyone should be doing some degree of this. We are essentially possibly identifying the same issues over and over again. Can we actually capture these concepts and codify them so that we don't have to always create those reviewing feedback one by one? You actually also have all of those things automatically identified. The beauty of this is if you do it a few times, you now build sort of like a system which actually learns over time. So think of this as sort of like, you know, you're doing more training on top of the standard LLM that you have actually extracted, like built on top of. So AI Slop Registers now can create better results because it's trained, it's learning from the review experience that you as humans are providing. Every recurring comment is now a guardrail that you don't have to review again. Okay. So let's kind of like try to put both of these alignment and the semantic accuracy together. It is two halves of the same problem. We are trying to understand what are the core mechanics of review. How do we actually break it down into two parts, which is the alignment and the semantic accuracy, and bring them together into a single loop. So first thing is you go take your session and you capture the user responses and that essentially forms your acceptance criteria. The acceptance criteria then tied with your AI Slop Register that you are now constantly maintaining, finally creates a test plan. And this is a test plan which then gets verified. This is part of the system that we are building, is the verification system where it spins up a preview, takes your test plan, and makes sure it actually works end-to-end. Even if the code looks right, does it actually work? So this verification part has become interesting. And now the kicker is this is now your review surface. You're not reviewing code line by line, but rather you're looking at the evidence of what was the intent, did the user actually implement the capability that was defined in their intent, and did the behavior actually meet the requirements that we had in the as acceptance criteria. So you're still having the architectural decisions, you're still having these arguments, but the review surface changes. So just giving a walkthrough of like how we have built our system, the session becomes a criteria. So all these decisions that you are doing here with the agent, you're asking, you know, you're providing this feedback. Like even if there is a simple task, many times you're going back and forth, agent will stop to ask questions. These are the decisions that we need to capture. This is the intent. This is what makes your review, makes your collaboration more valuable. This is how you teach your junior engineers on how to improve over time. These are the decisions which make a software engineer valuable today. We convert those into an acceptance criteria. This is where you can also leverage LLM to do so. So you capture these user decisions and make sure you can actually create a test plan based on this. I know test plan creation is always painful, which is where I would always recommend people to use LLM for this purpose. And finally, the criteria plus invariant is what makes the test plan. And we build the verification systems to actually capture the test plan, run your previews, and be able to test based on this test plan. So even imagine if you're building a new feature, you don't have to maintain test at all. This is creating test in real time. And this is where you can leverage the power of LLM because the test plan maintenance and creation can be very painful. But the value of human in the loop here is the governance and the review part and the part where you're reviewing the test plan and not the code. Right? I know it's been like over 20 years we came up with test-driven development. This in some ways is closer to behavior-driven development where the test plan is now something which even you can share with your product managers, your designers, everyone can participate in the test plan. And then the test plan is now in English. At the same time, we have now deterministic verification which actually verifies whether the particular test criteria has been measured. Let's move on. So this is where I would say the system is not supposed to be perfect. It's deterministic where it can be, but LLM where you must. Not everything can be deterministic. Not every system can be built in a way which is 100% built on deterministic systems. This is where you use LLM as a fallback. Let me give an example. Like if you are making a change in your web application, what you can do is it creates a test plan of what the behavior changes. Let's say you introduce a new payment form. So it creates a new payment form. The verification here is does the payment system changes. An AI agent can go and browse through your application to fill out a form and capture screenshots as evidence and then take those screenshots as well as your database snapshots to identify whether the criteria was met. So the screenshot testing or sampling can then still be done by agents, but at the same time you're creating more solid evidence which now a reviewer can look at and build more confidence that this actually works. So now reviewers are reviewing the intent, not the diff. You're reviewing the intent decisions, what we set to build out, what we tried and rejected, and capture all of these things from the sessions. Remember, capturing it from the sessions is the key. If we try to build it from the code, you end up in the same situation that we were talking about. I think Dex was talking about yesterday, which is if your code is built by the same agent which is actually building a test plan, it's not going to build a test plan which will actually catch issues. So that's why it's important to actually use the session information to build out a test plan. You can discuss architectural decisions, so how you're creating the data models, how these services interact with each other. So you're moved one level above. So you're essentially, instead of reviewing line by line, you're actually having discussions on architecture, which are very critical for any kind of collaboration. And then you look at the evidence, everything that was collected from the verification. So here's a homework for everyone. Go home and mine your last 1,000 review comments and build out an AI slot register for the things which are repeatable. So a vast majority of the comments that we're providing in your code review are something that we repeat over and over again. This compounds with every merged PR. Every time you capture something as a register, you don't have to capture that comment again. And this is where you can actually codify some of the best practices of doing code, maintain semantic accuracy, at the same time, do not lose the collaboration part of the review. It does follow a J-curve. So pain is real. You will have to spend some time to actually make it pay off because initially creating a registry can take some time. And this is where I would recommend you folks can come and try out our product. So code review is not just about code review. It is about really getting the alignment. And where we can build better tools is creating the semantic accuracy and defining your AI slot register. So if you remember one thing from today, remember code review is not just about code review. It is about getting the alignment. And yes, we are piloting our new product called Verify. Please join and be our early design partners. We are working with a few companies to pilot out a new verification system. This combines both the alignment side of things, as well as building tools and capabilities for detecting semantic accuracy using the AI slot register. Thank you everyone. Thanks for joining. Thank you everyone. Thanks for joining.