Open Reader

ReviewDebt: a practical framework for scoring every pull request — Sachin Gupta, Ebay

completed 25:00 Jul 12, 2026 Watch on YouTube

Current Status

completed

Video ID

TJPInBjhE4Q

RAG / Chat

Enabled
ReviewDebt: a practical framework for scoring every pull request — Sachin Gupta, Ebay
Description

Coding agents ship PRs faster than humans can trust them. The gap is filling up with a debt nobody is measuring — and it's about to swallow your engineering velocity. Every team in 2026 measures coding agents the same way: PR count, lines of code, cycle time, developer NPS. None of those see the real cost — bloated diffs, weak tests, ambiguous rationale, ownership sprawl, and human reviewers spending more time verifying AI code than they used to spend writing their own. This talk introduces ReviewDebt: a practical framework for scoring every pull request on the hidden review burden it creates. The scoring is deterministic — diff size, test-coverage delta, ownership spread, generated-code smells, evidence and rationale gaps — so the number is defensible in a real engineering review. We'll walk three real PRs side-by-side (clean human PR, high-debt AI PR, refactored AI PR), watch the scoring play out signal by signal, and look at a 90-day dashboard from a production backend org where review debt climbs in lockstep with AI-PR share. Speakers: - Sachin Gupta: Sachin Gupta is a Staff Software Engineer with 15+ years building backend platforms at internet scale, currently focused on the runtime trust boundaries that LLM coding agents blur and the creator of HeapLens, a Java heap analyzer extension used in 50+ countries. LinkedIn: https://www.linkedin.com/in/guptasachin1/ GitHub: https://github.com/sachinkg12

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: AI coding agents increase review debt—the compounding gap between code produced and code humans have genuinely reviewed, understood, and trusted—so teams should measure review burden rather than celebrate throughput alone.
  • Why it matters: For teams deploying coding agents, faster code generation can shift the bottleneck to senior-review capacity, architectural coherence, and incident accountability unless PR shape and review evidence are actively governed.
  • Best use: Use it as a practical design reference for an agent-aware PR policy, deterministic review-burden scorer, and engineering-management dashboard rather than as a generic warning against AI coding.

Executive Summary

Sachin Gupta argues that AI-assisted development has created an unmeasured operational liability: review debt. His definition is the accumulating gap between code an agent produces and code that humans have actually reviewed, trusted, and understood. Unlike ordinary technical debt, he argues, it compounds through three loops: unreviewed code becomes future agent context; reviewers focus on local syntax rather than architectural choices; and increased output resets leadership expectations without adding proportional review capacity.

The proposed response is ReviewDebt, a deterministic 0–100 PR score derived from five signal families: diff size and coupling, test evidence, directory/ownership spread, AI-authorship indicators, and evidence/rationale quality. Gupta deliberately rejects LLM-as-judge scoring because a changing model would make scores unstable and difficult to defend in engineering or leadership reviews.

The central nuance is that this is not an anti-AI metric. AI authorship is intended as a small, information-only amplifier, not a penalty. The largest burden drivers are PR complexity, weak tests, broad ownership boundaries, and insufficient explanation. A well-shaped AI-authored PR scored 7/100 in his example, while another PR scored 60/100 largely because of diff/claim mismatch and missing tests, with AI indicators contributing only 5 points.

His adoption plan is deliberately lightweight: backfill roughly 200 merged PRs, calibrate thresholds against known team experience, surface non-blocking scores on new PRs, require author evidence above a threshold such as 50, and track each team's weekly debt slope. The presentation is useful because it translates vague concerns about agent-generated code into concrete review workflow controls and a management metric, although its public-repo sample and estimated-review-time claims should be validated locally before operational use.

Key Takeaways

  • Claim: The meaningful risk of coding agents is not code generation itself but an expanding review-capacity gap that conventional velocity metrics conceal. | Evidence: Gupta cites GitHub's 2025 Octoverse report as showing commits up 25% year over year while commit comments, used here as a proxy for review activity, fell 27%. He also cites a Faros AI benchmark in which median PR review time rose 441.5% and 31% more PRs were merged with no review. | Implication: Ken should avoid using PR count, merge speed, or AI adoption rate as standalone proof of engineering productivity; pair them with review-capacity, evidence-quality, and post-merge risk measures. | Caveat: The cited benchmarks are presented without methodology in the transcript, and commit comments are an imperfect proxy for review quality; treat the directional argument as stronger than the precise figures.
  • Claim: Review debt compounds because weakly reviewed code becomes input to future agent work, architectural review is displaced, and organizational velocity expectations remove the slack needed to repay the debt. | Evidence: The three feedback loops named are: agents use repository code through fine-tuning, RAG grounding, or in-context suggestions; reviewers contract attention to syntax and obvious bugs when PRs are largely generated; and leadership sees increased throughput without hiring reviewers proportionately. | Implication: Agent systems that retrieve from or learn from a repository need stronger provenance and quality gates: poor review is not merely a local defect risk but can contaminate the context for subsequent generated changes.
  • Claim: Review burden can be measured with deterministic repository and PR checks rather than an LLM judging code quality. | Evidence: ReviewDebt uses five signal families with 10 deterministic checks: diff size/coupling, test evidence gap, directory and ownership spread, AI-authorship indicators, and evidence/rationale gaps. Gupta argues that LLM scoring is a moving target when models change and cannot be readily defended in an engineering review. | Implication: A PR-control plane should make scoring explainable and reproducible, while reserving deeper semantic analysis or human judgment for the high-risk cases that the deterministic score surfaces. | Caveat: Deterministic signals measure reviewability and burden, not semantic correctness or test quality; for example, the test metric only establishes whether tests were added, not whether they assert the right behavior.
  • Claim: PR shape—not AI authorship alone—is the primary driver of review burden; AI detection should only direct additional attention, not stigmatize agent usage. | Evidence: In the 60/100 example, soft AI-authorship indicators accounted for only 5 points, while 55 points came from diff-size/claim mismatch and missing tests. Conversely, an AI-authored PR with tests, green CI, and an explicitly identified risky path scored 7/100; its AI indicator was described as information-only and contributed 2 points. | Implication: Set identical substantive review standards for human and agent-assisted PRs. Use agent provenance, where available, for routing or prioritization, never as a standalone merge blocker or performance signal. | Caveat: AI-authorship heuristics—co-author footers, branch prefixes such as Codex/Copilot/Cursor, and 'generated by' language—can be absent, suppressed, or inconsistently used, so they are not reliable attribution mechanisms.
  • Claim: Cross-team scope and missing author rationale are disproportionately expensive because they fragment reviewer context and eliminate the ability to assess intent. | Evidence: The ownership signal counts distinct code-owner teams touched by a diff; Gupta argues that a multi-team PR requires multiple approvals and no individual retains the whole mental model. His high-rationale-gap example had the title 'fix leaky test,' an 18-character PR body, and a commit message of 'updates,' versus a reviewable PR that included symptom, diagnosis, change, and benchmark link. | Implication: Require agent users to personally author the rationale: what failed, why the chosen change addresses the cause, what behavior is expected, and which risky paths were tested. Split cross-cutting work into coherent per-owner PRs where possible.
  • Claim: The score should drive an escalation workflow and trend monitoring, not become an automatic quality verdict or blocking gate from day one. | Evidence: Gupta's bands are 0–24 low burden, 25–49 normal, 50–74 'needs evidence before senior review,' and 75+ 'split or request more context.' His rollout sequence is to score the previous 200 merged PRs, calibrate locally, post a non-blocking PR comment, aggregate weekly by team, and discuss the weekly slope in retrospectives and roadmap reviews. | Implication: Start with visibility and author-justification requirements rather than hard merge blocks. The key management indicator is whether review debt is rising over time, not whether any isolated PR has a high score. | Caveat: The transcript's proposed weights, thresholds, and estimated-review-minute figures are defaults rather than validated universal standards; a score must be calibrated against a team's own historical PRs and incidents.
  • Claim: In the presented public-repository scan, volume and structural complexity created most review burden, while AI indicators remained relatively stable. | Evidence: Across 524 PRs in three unnamed public repositories, AI indicators reportedly appeared on 5–20% of weekly PRs, while one repository accumulated an estimated 186 senior-review hours over 27 days versus 43 hours for another. Only four PRs reached 'needs evidence' or high bands, and these were described as large migrations, SDK rewrites, or multi-team refactors. One PR was estimated at 5,036 review minutes (about 84 hours) and scored 73. | Implication: Capacity planning should monitor sustained merge volume and concentration of large cross-cutting PRs, not merely the share of PRs visibly attributed to an AI tool. | Caveat: The sample is small, the repositories are unnamed, and the estimated-review-effort model is not fully specified, so these numbers demonstrate the framework's intended use rather than establish industry-wide rates.

Detailed Brief

Signal-family mechanics and intended reviewer behavior

  • Claims: Diff size and coupling are not simply line-count measures: a change spread across unrelated files imposes a nonlinear mental-model cost on reviewers.; The test evidence gap is calculated as test lines added divided by production lines added; it is a coverage-presence signal, not a proof that the test suite captures intended behavior.; A healthy PR should yield a near-empty report rather than constant tooling noise; the scanner's clean example scored 0/100, estimated six review minutes, and fired no checks.; The practical value is not the composite number alone but structured output: a list of fired checks, reviewer focus areas, and author actions that lower the burden.
  • Evidence: Gupta contrasts an agent tendency to address symptoms at multiple call sites with a human tendency to route a correction toward a root cause, resulting in a more coherent diff.; He warns that generated tests may validate the implementation's current behavior, including bugs, rather than specify what the system should do.; The clean-PR output intentionally had empty reviewer-focus and author-next-action fields, illustrating that a scorer should be quiet when no specific intervention is needed.
  • Caveats: Line-based test ratios can be gamed and do not evaluate behavioral coverage, assertions, integration risk, or the correctness of expected outcomes.; Ownership spread depends on accurate CODEOWNERS or equivalent repository metadata; incomplete ownership maps will understate coordination costs.
  • Implications: Design the PR bot as an explanation and triage layer, not a generalized code-quality oracle.; Track which checks repeatedly fire by team or repository area; recurring patterns can reveal weak architecture boundaries, ownership ambiguity, or agent workflow problems.

Operating rules Gupta recommends for AI-assisted PRs

  • Claims: Each PR should contain one logical change, not merely a small arbitrary number of changed lines.; Tests should ship with the change, but the human author must confirm they express intended behavior rather than merely mirror generated implementation.; Cross-cutting work should be divided into per-team or per-owner PRs to preserve one approval context and one understandable mental model.; The agent should not author the PR rationale; requiring the human author to write the 'why' is the moment of accountability for what is being shipped.; Teams should reject superficial merge habits such as 'the AI did the boring part,' deferring detection to QA, treating smaller PRs as automatically safe, or approving with a casual LGTM.
  • Evidence: The proposed operating formula is 'one approval, one context, one mental model' for ownership-contained changes.; Gupta frames the review-debt metric as an audit bridge between increased AI throughput and accountability when an AI-assisted change later causes an incident.
  • Caveats: Some production changes inherently require coordinated multi-owner modifications; splitting them may reduce reviewability but can introduce sequencing, compatibility, and rollout complexity.
  • Implications: For agent-generated changes, enforce an explicit human acceptance boundary: the named author owns the intent, test expectations, rollout rationale, and any cross-team coordination.

Notable Concepts & Terms

  • Review debt: The compounding gap between generated code and code that has been meaningfully reviewed, understood, and trusted by humans.
  • ReviewDebt score: A proposed deterministic 0–100 PR review-burden score intended to trigger evidence requests, splitting, or additional context rather than judge code quality directly.
  • AI-authorship amplifier: A small, information-only scoring signal based on indicators such as co-author footers, AI-tool branch names, or generated/assisted wording; it should not independently penalize AI usage.
  • Test evidence gap: Test lines added divided by production lines added, used as a simple indication that test evidence accompanied a change, while explicitly not measuring test quality.
  • Directory and ownership spread: The number of distinct code-owner territories touched by a PR, used as a proxy for coordination overhead and fragmented reviewer context.
  • Evidence and rationale gap: The absence of an explanation of why a change is needed and how it addresses the problem, which Gupta considers especially destructive to reviewability.
  • Debt slope: The weekly trend in aggregate review burden by team; Gupta considers its direction more useful to management than the level of any single score.

Operator Notes / Why Ken Should Care

  • Run a retrospective scorer over at least 200 merged PRs and manually compare the highest-scored items with known painful reviews, incidents, and reversions before selecting weights or thresholds.
  • Implement a non-blocking PR check that requires a human-written rationale and test/validation evidence above a calibrated burden threshold; begin with author justification rather than automatic merge denial.
  • Ensure CODEOWNERS, service ownership, and repository boundaries are current before using ownership-spread signals for routing or reporting.
  • Add agent provenance to internal PR metadata where feasible, but treat it as routing context only; do not rely on co-author footers or branch naming as a complete source of truth.
  • Create a weekly dashboard that pairs AI-assisted merge volume with aggregate review-burden slope, reviewer load, review latency, and post-merge incidents to test whether the metric predicts operational risk.
  • Audit agent PR templates so the human submitter—not the agent—must explicitly state the symptom, root-cause hypothesis, intended behavior, validation performed, rollout risk, and rollback plan.

Source/Metadata

  • Title: ReviewDebt: a practical framework for scoring every pull request — Sachin Gupta, Ebay
  • Transcript words: 5460
  • Duration seconds: 1500
  • Timestamp note: No timestamps or chapters were present in the supplied transcript. The transcript also contains substantial duplicated passages near the end.

Transcript

3683 words en Processed in 232.6s

Hi everyone, I'm Sachin Gupta, and I'm a software engineer. The title is exactly what it sounds like. Your coding agent is creating review date. Before I start, sit with the title for a second. Notice what it is not saying. I'm not saying that coding agents are bad. I'm not saying they don't make us faster. I'm saying they are creating a kind of debt that nobody is measuring. And that debt is going to come to you. Over the next few minutes, here is what I'm going to do. I will define what review date is. I'll walk five signal families that are going to compose it. I'll score three real pull requests side by side. And I will show you a cross-repo scan of over 500 PRs, which is nothing but three public code bases. So let's go. Look here. Here's the gap nobody is measuring. GitHub's 2025 Octover's report that covers almost every public pull request on the planet shows that the commits climb 25% year over year. And now, over the same year, comments on the commits drop 27%. Now these comments are nothing but the proxy for review activity. Code production volume actually went up, but the review attention went down. They moved in opposite directions, in the same year. Now look at the teams furthest along the AI adoption curve. Faro's AI tracks this cohort in their 2026 benchmark. Median PR review time is up by 441.5%. And if you calculate, you'll figure out that the review PRs take 5.4 times longer than what they used to. And then 31% more PRs are now merged with no review at all. So AI is producing the code very fast. AI is producing the pull request very fast. But humans cannot responsibly review them at that pace. This gap is called reviewed at. It actually accrues quietly. It compounds. And right now, nobody has a number for it. But by the end of this talk, I will make sure that you will have a certain number. Now, if you look at this particular slide, here is the story every team is telling right now. PR per developer are up 16%. That's the Faro's AI acceleration whiplash benchmark, and that too from April 2026. There were 22,000 developers, 4,000 teams. And the median PR size is up 63%: 44 lines to 72 lines per pull request. That is 16 months of long-term data from DX 2026 study, which includes 400 organizations. Now, if you see the cycle time open to merge, that is modestly down. And it is coming from the same DX study. The framing is very generous. PR throughput grew actually about 8%. And the AI usage rose about 65%. The gain that we see today is real. But it is actually smaller than the hype that is present. Every one of these numbers are real. None of them is a lie. But every one is a vanity metric. PR count goes up when one PR spreads into seven PRs. Median PR size going up is not a benefit. It's actually bloating. Cycle time going down when reverse stop pushing back. These things tell you the speed of production. They do not tell you the speed of trust. Now, we will go into the next slide to see more. Now, if you go looking for the second story, here is what you're going to find. Reviewer fatigue. Engineers carry far more review load than a year ago. They're not happier about it. Late night merges. PRs are sitting unreviewed for three days, four days, and then they suddenly get a thumbs up at 11 p.m. before a Friday deadline. Test theater. Tests are getting added, but they are actually asserting what the code did, not what the code should do. They lack in behavior, including bugs. Architectural drift. Again, the same problem that is solved in three different ways in three different ways. Nobody is actually holding up the architectural thread. Finally, incident lag. When something breaks, the bug lands weeks or months after merge. Nobody connects the dot back to the AI author change. These costs don't show up in any dashboard until now. We'll look into the next slide to know more about it. So now, let me give you the definition. Review debt. It's the accumulating gap between code your agent has produced and the code humans have actually reviewed, trusted, and understood. It rhymes with technical debt, but it is more like a financial debt because it compounds. It actually accrues interest, but this interest is not money. This interest is paid in human attention. It compounds because of the three feedback loops. First, the agent learns from your code base: fine-tuning, rag-grounding, in-context suggestions. Code that was not deeply reviewed yesterday grounds tomorrow's PR. This debt becomes generative in nature. Second, reviewers seed the architectural call. When most of a PR was generated, attention contracts to syntax and obvious bugs. Big picture decision moved from review time to never. Third, velocity expectations reset. Once leadership sees the new throughput, you don't get to hire reviewers in proportion. There is no slack left to pay the debt back. Each loop on its own is survivable. Together, you will have a runway. But the next question is, how are you going to measure it? So, how do we measure review debt? There are five signal families that have 10 deterministic checks. Remember, the keyword here is deterministic. Every check is computable from a pull request and its repository. No, we don't need any language model. Why? Because LLM, when it acts as a judge, it possibly can break two things. Number one, the score becomes a moving target. The same PR, the score that you got for the same PR will score differently when your model will change. Two, the score stops being defensible in an engineering review. You cannot put it on a slide. You cannot put it in front of a leadership. So, what do you want? You want a number that is traceable to a deterministic computation, different size and coupling. Second, test evidence gap. Third, directory and ownership spread. Fourth, AI authorship indicators. Fifth, evidence and rational gaps. These are the five families that will have 10 checks under them. Let's walk each of them. The first one, the first signal that we have is def size and coupling. This is the simplest one to calculate. It's also the most often misread. What it measures is the net lines changed, the files touched, whether changes cluster in one module or are sprawling across many. Why is the agent struggling here? Agents are biased toward fix at the call site. A human engineer routes a fix to the root cause. That keeps the difference very small. But an agent will reach into many files looking for the same symptom. The reviewer cost of a sprawling difference is not proportional to the size. It is actually much steeper. So the cross-pile coupling explodes the mental model and reviewer has to hold. We'll look into the second signal, which is test evidence gap. So what is test evidence gap? Test evidence gap is nothing but test lines added divided by the production lines added per pull request. So AI-authored PRs, they ship with a far lower test-to-code ratio. Sometimes there is no test file in place, and if there is any, it's minimal. So this gap is actually a consistent gap. And the problem with this particular thing is the number. The reason this number is brutal is the agent generate test. They generate a lot of tests. But these tests basically assert what the code is doing. It does not generate the test cases what the code should actually do. They lock in the behavior, including the bugs. The ratio does not capture the quality gap. It just measures whether the tests showed up at all. So the deeper signal is the next step that is a layer down. We'll see the next one, which is directory and ownership spread. So what is directory and ownership spread? Count the distinct code owner teams whose file appear in the difference. Now what does it mean? Basically, a well-shaped PR concentrates in one team territory. But a stalling one reaches across many team files, and the reward cost is enormous in this scenario. No single human hold the whole mental model. You need multiple approvals from multiple engineers in multiple different contexts. The coordination overhead easily exceeds the time the agent save producing the code. This is where you start to feel the economics of review debt. The agent gave you three hours of typing. You spend those three hours by multi-party reviewer attention bank. And the next one we are going to see is AI authorship indicators. So what is AI authorship indicator? Before everyone gets defensive, this is not for the blame. We are not flagging that engineers use coding agents. We wait the scores to a PR shape like agent-assisted authorship get extra reviewer attention. So that way you know, this is the PR that is directly coming from the agent. And there are the very common three detection modes. I think everyone must have seen it by now. One is the co-author! Which is basically one of your strongest signal. You see that co-authored by copilot. Second one contexts. The coordination overhead easily exceeds the time the agent saves producing the code. This is where you start to feel the economics of review debt. The agent gave you three hours of typing. You spend those three hours by multi-party reviewer attention bank. And the next one we are going to see is AI authorship indicators. So what is an AI authorship indicator? Before everyone gets defensive, this is not for blame. We are not flagging that engineers use coding agents. We weight the scores to a PR shape like agent-assisted authorship get extra reviewer attention. So that way you know, okay, this is the PR that is directly coming from the agent. And there are the very common three detection modes. I think everyone must have seen it by now. One is the co-author, which is one of your strongest signals. You see that co-authored by Copilot. Second one is the branch name pattern. Generally, you must have seen it says Codex, Copilot, Cursor prefixes. Third one is generated by, assisted by, which is nothing but your PR body or commit message phrases. We did a real data check on three public repos, which comprised 524 PRs. I am not going to name the company, but for the sake of this, we are just calling it A, B, and C. Now if you see, the highest signal was the co-author footer, and the second one, the lowest signal, was 0% on repo C. Even though it was done by a coding agent, they might have blocked it, saying any co-authored or generated by, etc. The next thing that we are going to see is evidence and rationale gaps. So evidence and rationale gaps. I think this is one of the most deterministic that I personally feel. It's the one that destroys reviewability the fastest. What it measures, basically: does the PR explain the why, or just the what? On the left, if you see, we have a high gap. On the right, we have a low gap. The title says fix leaky test, and all the data that you are seeing here is actually from some public reports. Now the PR body length is 18 characters. The commit message says updates. Obviously, if you give it to me, I won't be able to review this. Okay, I cannot accept this. But on the right, if you see, when there is actually a low gap, the title actually is telling you what the change is all about. The body has a symptom, the diagnosis, the change, and the link to the benchmark. Now a reviewer can do the job. So in our regression fixtures and hand-coded teams, this signal destroys reviewability fastest. In the real open source data, PR bodies follow conventional commit format, so it fires most rarely there. But if you see, the regression fixtures are just the demos. Okay, let's see how those five signals combine. You will only have one number that is going from 0 to 100, and the exact weights are basically your defaults. Now if you are planning to adopt this, I would like you to do this first. Run it backwards over the last 200 PRs that have been merged in your company. Calibrate the weight against your team's actual experience. The score has to feel right against your actual gut. You are going to divide it into four bands. If you are 0 to 24, which means you have very low review burden. If you have 25 to 49, you are normal, proceed with the standard care. 50 to 74, you need evidence from the author before senior review. And if you are 75 and above, definitely it's high, split or request more context. Same shape as technical debt categories, but the unit is different. Now we are going to see what the scanner actually did. Now if you see, this is the clean PR. Unfortunately, I won't be able to give you a live demo today, but I will walk you through a CLI. I want to show you three scored pull requests side by side, the level where the framework becomes useful as a repeatable review conversation. This is the first one. If you see, this is what the scanner says when a PR is well shaped. Score 0 out of 100. The burden is basically nil. It's 0. We have a low review burden. The estimated minutes is 6. No checks were fired. And if you look at the structure here, that is why the list is empty. Reviewer focus is empty. Author next action is empty. The framework only generates specific advice when it has something specific to say. A healthy PR produces a tiny, almost ceremonial report. That's exactly what you want. And the takeaway is most healthy PRs produce zero noise. They get the work done without making any noise. This is not a tool that complains by default. Running it on every PR costs you one comment that says look good. That's it. Let's see the high debt PR now. Now what you see actually on the screen is from a public repo. This is a real scenario from the scanner regression suite. If you see, the score is here, 60 out of 100. It requires evidence. It's on the orange, it says needs evidence. 86 estimated minutes of review effort. Look at the structured output. It's actually not a score, a wire list explaining what fired, a reviewer focus list telling the reviewer what to do next, an author next action list telling the author how to bring the score down. This is the part most PR quality tools miss. The score alone is useful. The structured advice is what actually moves the team behavior. Now if you read the bullet, it says PR has soft indicators of AI-assisted authorship. This is information only, not a definitive claim, and not a penalty on its own. This particular sentence is in the scanner output. The AI indicator check contributes 5 of the 60 points, which is nearly about 8%. The other 55 came from a different size claim mismatch and missing tests. Those would be high-burden signals on any PR, agent-authored or not authored. This is not an anti-AI scorecard. Basically, this is more of a review burden scorecard. The agent did not cause the score, but the shape of the PR that is created by the agent did this. Let's see the next one. This is the PR, basically, which is AI-authored. It's very well shaped, and the score is 7, which is low review burden. Now if you see, the same agent, it has done the same kind of work, but the tests were added. So the CI is green. The risky path is called out in the PR description. That's why we got the score as 7 out of 100. It's very low burden, took only 14 minutes. The AI indicator check still fires, but it contributes two points. One, low severity. Second, the framework explicitly says, quoting the report itself, information only, not a definitive claim, and not a penalty on its own. So when your team is going to ask, will this tool penalize us for using coding agents? No. The answer is on your screen: low review burden. It's a very well shaped AI PR. We are going to see the next one, which is basically our three public repos over 524 PRs. I wanted to show you a 90-day AI sub-slope across three public repos and 524 pull requests. The review burden climbs even when the AI authorship doesn't. If you see this particular scan, we observed three things. One of them is the volume is the actual variable. AI authorship was flat, and it was 5-20% steady state, but the review burden was not. If you see, one repo accumulated 186 senior reviewer hours in 27 days, and another one took 43. The window length is same, but the volume was not the same. The burden was not the same. That is what PR volume is. Second is amplifier only holds up under real data, which means AI indicators fired on 5-20% of PRs every week across all three repos. None of those AI positioning contracts holds against real codebases. Third one is complexity drives burden, not authorship, which means across 524 PRs, four landed in needs evidence or high bands, which means all four were structural changes, which were large migrations, SDK rewrites, multi-team refactors. The framework scores complexity fairly. AI-driven volume creates the condition under which these accumulate. Where does the review debt actually show up in absolute terms? The scanner saw 524 real pull requests in three public repos. 12-28 senior review hours accumulated across three public repos that were scanned over 27- to 90-day windows. Second one is 9 PRs per day, sustained merge rate at high-velocity repos. So if you see, it is all about the volume that is basically driving the burden. Third one is 5036. These are the minutes that are basically spent on a single PR, which is nothing but 84 hours of review effort estimated for one single pull request. The score that we received there was 73, which said it requires an evidence spend. Now 5-20 percent of these PRs every week fire the AI indicator signal. Again, remember, this is volume which is causing the review debt, which is changing the game altogether. We'll move to the next slide. We have shown the cost. Now what do we do about it? So first thing is you should have one logical change per PR. I'm not saying the scanner saw 5-24 real pull request in three public repos. 12-28 senior review hours accumulated across three public repos that were scanned over 27 to 90 days windows. Second one is 9 PR per day sustained merge rate at high velocity repos, so if you see it, all about the volume that is driving the burden. Third one is 50-36. These are the minutes that are spent on a single PR, which is nothing but 84 hours of review effort estimated for one single pull request. The score that we received there was 73, which said it requires an evidence spend now. 5-20 percent of these PR every week, across every week, fire the AI indicator signal. Again, remember, this is volume which is causing the review debt, which is changing the game altogether. We'll move to the next slide. We have shown the cost. Now, what do we do about it? So, first thing is you should have one logical change per PR. I'm not saying not small PR in the abstract. One logical change, that is enough. Second, test shift with the change. Even if the agent wrote the code, even if the agent wrote the test, the human author confirms the test assert what the code should do, not what the code is doing, what the code is supposed to do. Okay, third, stay in one owner territory. Cross-cutting work, splitting into per-team PR, remain in your own territory. One approval, one context, one mental model. Author writes the why. The agent should not write the PR body. That's the moment the human author commits to understanding what they are actually shipping. Same review standard for the AI PRs as human PRs. The AI indicator amplifier only fires when other signals are weak. Keep the other four strong. The amplifier is invisible, so make sure there is no exception for AI. None of these is actually requiring you to have a new tool. These are the moves that you already are aware of. You just need to implement it. Now, how are you planning to adopt this? If you see, we again have five steps: back fill, threshold, surface, aggregate, and talk about it. Back fill: so you run the scorer over your last 200 merge PR. Look at the one that has the highest score. Second, threshold: set up justify line default. Let's say you want to give it 50, so any PR scored at 50 or above requires a comment from the author. As simple as that. Surface it: post the score as a PR comment on every PR. Don't block it, just for the visibility purpose, surface it so that everybody should know what's happening over there. Aggregate weekly per team, which means the slope of each team deadline is the leading indicator. That's what your engineering manager should watch. Talk about it: bring the number to a retrospective, every roadmap view, in every meeting. You should discuss that number, like what was your score when you did this particular PR? That's why this number matter, because of this conversation. Absolutely not. Without a number, what you are just converging is Y, but if you have a number, it's actually structured. You can tell, okay, with AI coding agents, we have a rollout which has added X percentage of the throughput, but you need to make sure you are also adding the Y points of the review debt that has. If not, then who is accountable when an AI author change causes an incident? Where is the audit trail? So review debt is the bridge between these two columns. It's the first number that lets you have the right-hand-side conversation without abandoning the left-hand-side gains. sinn sinn sinn sinn sinn sinn sinn sinn sinn sinn sinn sinn sinn sinn over time matters most than the level, so plot it weekly and then discuss it with your team, maybe in a meeting or in a sprint planning, whatever the way you feel like. Have this conversation. Bring the number to your next review, say that, okay, we do have extra put wide it and this is the slope, and then you are moving away from the vibe culture to the measurement culture. One quick callout on anti-patterns to avoid. First thing is approved-with-comment mergers. The AI did the boring part. No. Will you catch it in QA? No. PR are smaller. No, no. Stop writing LGTM on the PRs. Make sure whatever we have discussed, the habits that we have discussed, we are going to follow them. This is the entire point of the talk. Thank you very much for listening. any noise. This is not a tool that complains by default. Running it on every PR costs you one comment that says look good. That's it. Let's see the high debt PR now. Now you see what you see actually on the screen is from a public repo. This is a real scenario from the scanner regression suite. If you see the score is here 60 out of 100. It requires evidence. It's on the orange it says needs evidence. 86 estimated minutes of review effort. Look at the structured output. It's actually not a score. A wire list explaining what fired. A reviewer focus list telling the reviewer what to do next. An author next action list telling the author how to bring the score down. This is the part most PR quality tool miss. The score alone is useful. The structured advice is what actually moves the team behavior. Now if you read the bullet book, it says PR has soft indicators of AI assisted authorship. This is information only not a definitive claim and not a penalty on its own. This particular sentence is in the scanner output. The AI indicator check contributes 5 of the 60 points which is nearly about 8%. The other 55 came from a different size claim mismatch and missing test. Those would be high burden signals on NAPR. Agent authored or not authored. This is not an anti-AI scorecard. Basically, this is more of a review burden scorecard. The agent did not cause the score, but the shape of the PR that is created by the agent did this. Let's see the next one. This is the PR basically, which is AI authored. It's very well shaped and the score is 7, which is low review burden. Now if you see the same agent, it has done the same kind of work, but the tests were added. So the CI is green. The risky path is called out in the PR description. That's why we got the score as 7 out of 100. It's very low burden, took only 14 minutes. The AI indicator check still fires, but it contributes two points. One, low severity. Second, the framework explicitly says, quoting the report itself, information only, not a definitive claim, and not a penalty on its own. So when your team is going to ask, will this tool penalize us for using coding agents? No. The answer is on your screen, low review burden. It's a very well shaped AI PR. We are going to see the next one, which is basically our three public repos over 524 PRs. I wanted to show you a 90-day AI sub slope across three public repos and 524 pull request. The review burden climbs even when the AI authorship doesn't. If you see this particular scan, we observe three things. One of them is the volume is the actual variable. AI authorship was flat and it was 5-20% steady state, but the review burden was not. If you see one repo accumulated 186 senior reviewer hours in 27 days and another one took 43. The window length is same but the volume was not the same. The burden was not the same. That is what PR volume is. Second is amplifier only holds up under real data which means AI indicators fired on 5-20% of PRs every week across all three repos. None of those AI positioning contract holds against real code basis. Third one is complexity drives burden not authorship which means across 5-24 PRs four landed in needs evidence or high bands which means all four were structural changes which were large migration SDK rewrites multi team refactors the framework score complexity fairly AI driven volume creates the condition under which these accumulate ! where does the review date actually show up in absolute terms the scanner saw 5-24 real pull request in three public repos 12-28 senior review hours accumulated across three public repos that were scanned over 27 to 90 days windows second one is 9 PR per day sustained merge rate at high velocity repos so if you see it all about the volume that is basically driving the burden third one is 5036 these are the minutes that are basically spent on a single PR which is nothing but 84 hours of review effort estimated for one single pull request the score that we received there was 73 which said it requires an evidence spend now 5-20 percent of these PR every week across every week fire the AI indicator signal again remember this is volume which is causing the review debt which is changing the game all together we'll move to the next slide we have shown the cost now what do we do about it so first thing is you should have one logical change per PR I'm not saying like not small PR in the abstract one logical change that is enough second test shift with the change even if the agent wrote the code even if the agent wrote the test the human author confirms the test assert what the code should do not what the code is doing what the code is supposed to do okay third stay in one owner territory cross cutting work splitting into per team PR remain in your own territory one approval one context one mental model author writes the why the agent should not write the PR body that's the moment the human author commits to understanding what they are actually shipping same review standard for the AI PRs as human PRs the AI indicator amplifier only fires when other signals are weak keep the other four strong the amplifier is invisible so make sure there is no exception for AI none of these is actually requiring you to have a new tool these are the moves that you already are aware of you just need to implement it now how are you planning to adopt this if you see we again have five steps back fill ! threshold surface aggregate and talk about it back fill so you run the scorer over your last 200 merge PR look at the one basically that has the highest score second threshold set up basically justify line default let's say you want to give it 50 so any PR scored at 50 or above requires a comment from the author as simple as that surface it post the score as a PR comment on every PR don't block it just for the visibility purpose surface it so that everybody should know what's happening over there aggregate weekly per team which means the slope of each team deadline is the leading indicator that's what your engineering manager should watch talk about it bring the number to a retrospective every roadmap view in every meeting you should discuss about that number like what was your score when you did this particular PR that's why this number matter because of this conversation absolutely not without a number what you are just converging is Y but if you have a number it's actually structured you can tell like okay with AI coding agents we have a rollout which has added X percentage of the throughput but you need to make sure you are also adding the Y points of the review debt that has if no then who is accountable when an AI author change causes an incident where is the audit trail so review debt is basically the bridge between these two columns it's the first number that lets you have the right hand side conversation without abandoning the left hand side gains sinn sinn sinn sinn sinn sinn sinn sinn sinn ! sinn sinn ! sinn over time matters most than the level so plot it weekly and then discuss it with your team maybe in a meeting or in a sprint planning whatever the way you feel like have this conversation bring the number to your next review like say that okay we do have extra put wide it and this is the slope and then you are basically moving away from the vibe culture to the measurement culture one quick call out on anti-patterns to avoid first thing is approved with comment mergers the ai did the boring part no will you catch it in qa no pr are smaller no no stop writing lgtm on the prs make sure whatever we have discussed the habits that we have discussed we are going to follow them this is the entire point of the talk thank you very much for listening