AI-Generated Code Is Already Competing With Human Code — Daksh Gupta, Greptile
Description
More than a quarter of the pull requests Greptile reviewed in April showed signs of being written largely or entirely by AI agents, up from under 1% a year earlier. Daksh Gupta, co-founder of Greptile, analyzed more than a million pull requests a month from companies like NVIDIA, Coinbase and Scale to test a simple question: are fully vibe-coded PRs actually any good in enterprise codebases? He compares agent-written PRs (Codex, Claude Code, Devin, Cursor) with human-written ones on four measures: revert rates, revert rates by PR size, the severity of issues Greptile flags (P0/P1/P2), and how many review rounds it takes to get to a mergeable PR. On all four, agent code landed in the same range as human code, and humans were actually more likely to introduce P0 bugs. The differences show up in *how* each one fails: Claude is about 1.5x more likely than humans to introduce SQL injection, Devin is half as likely to cause auth bypasses, and N+1 queries are far more common from Cursor. Daksh also shares what this means for code review. The median Greptile user makes fewer than 50 commits a month, while the 99th percentile makes close to 1,000, which makes manual review impossible at that scale. He argues that code validation should answer three questions: Does this change violate the user contract? Does it make a future violation more likely? Does it do what the author intended? He then shows how Greptile answers them with blast-radius analysis and sandboxed browser agents that try to break the app. Speaker info: - Greptile: https://greptile.com - Daksh Gupta on X: https://x.com/dakshgup - Daksh Gupta on LinkedIn: https://www.linkedin.com/in/dakshg/ - Daksh's website: https://dakshgupta.com Timestamps: 00:00 Intro 00:25 Daksh Gupta & Greptile 00:57 From tab complete to fully autonomous agents 02:09 Are fully vibe-coded PRs any good? 02:29 How to detect vibe-coded PRs: author fields, co-author footers, branch prefixes 03:49 25%+ of PRs vibe coded, up from under 1% a y
Summary
Generated by gpt-5.6-terraAt-a-Glance
- Verdict: Watch fully
- Core thesis: Enterprise pull requests that are largely generated end-to-end by coding agents now appear comparable to human-authored PRs on observed revert, bug-severity, and review-cycle measures, but they require scalable, agent-native validation rather than conventional manual review.
- Why it matters: If agentic coding raises PR volume by 10-100x at the productive edge, software delivery is constrained less by code generation than by an automated control plane that can verify intent, user-contract safety, and deployability.
- Best use: Use this as a directional data point and a practical framing for designing AI-code governance: route every agent PR through sandboxed, context-aware validation and track model-specific failure patterns rather than applying a single generic review policy.
Executive Summary
Daksh Gupta of Greptile argues that autonomous coding agents have moved beyond demo use into real enterprise development. Drawing on Greptile's review stream of more than one million pull requests per month across thousands of companies, he estimates that roughly 25% of reviewed PRs are now completely or largely AI-generated, up from under 1% early in the prior year. His central observation is that adoption appears to be continuously diffusing as models improve, rather than arriving only in visible jumps around model releases.
On the question of quality, Gupta reports no material aggregate disadvantage for agent-authored code using three operational proxies: reverted PRs, severe issues detected by Greptile, and review iterations before merge. Codex PRs were reportedly reverted about 1 per 1,000, Devin PRs about 1 per 3,500, and human PRs about 2.5 per 1,000; agent and human PRs also showed little difference in review cycles. Three of four evaluated agents produced fewer P0 issues than human-authored PRs in Greptile's findings.
The more useful conclusion is not that all agents are equally safe, but that their errors differ. Gupta found model-specific failure profiles—for example, Claude-associated PRs were 1.5x as likely as human PRs to produce SQL-injection findings, while Devin was roughly half as likely to produce off-by-one errors. This makes model-aware evaluation, policy, and testing a better operating model than treating AI code as a single category.
He frames the bottleneck created by high-volume agents as validation. Greptile's first-principles test for a mergeable PR is whether it violates user contracts, increases the future propensity for such violations, or fulfills the stated author intent. Their proposed mechanism is a swarm of context-aware agents that inspect related code, run the application in a sandbox, resolve dependencies, mock inputs, and exercise browser behavior. Gupta says nearly one-fifth of Greptile-reviewed PRs already merge without human review or testing, subject to quality guardrails.
Key Takeaways
- Claim: Largely autonomous, end-to-end AI coding is already a substantial part of enterprise PR production rather than a niche "vibe coding" phenomenon. | Evidence: Greptile reviewed more than 1 million PRs per month across thousands of companies and inferred that about one-quarter of PRs in a given month were completely or largely AI-generated; fewer than 1% showed such evidence early in the prior year. | Implication: Planning assumptions for engineering organizations should shift from occasional AI assistance to a near-term environment in which a meaningful share of change volume is agent-created and needs systematic governance. | Caveat: AI authorship was inferred from signals such as co-author footer text and agent-style branch prefixes, not from a definitive provenance field; GitHub commit authors alone identified less than 1% of PRs as coming from Codex, Claude, or Cursor.
- Claim: On Greptile's observed quality proxies, agent-generated PRs were broadly comparable to human-authored PRs and sometimes showed lower severe-bug rates. | Evidence: Reported revert rates were approximately 1 per 1,000 for Codex PRs, 1 per 3,500 for Devin PRs, and 2.5 per 1,000 for human PRs; three of four tested agents produced fewer P0 findings than humans, with similar broad results for P1s and P2s. | Implication: A blanket presumption that agent-authored code is inherently lower quality is not supported by this dataset; admission criteria should be evidence-based and applied to both human and agent changes. | Caveat: These are observational, vendor-generated metrics rather than a randomized comparison. Reverts and automated bug findings are incomplete quality measures and may not capture maintainability, architectural fit, security coverage beyond detected patterns, or business correctness.
- Claim: The apparent quality parity is not readily explained by agents merely receiving smaller or simpler PRs. | Evidence: Gupta tested PR size against revert behavior and reported little, if any, correlation differentiating reverted human PRs from reverted agent PRs by size. | Implication: Organizations should test their own task-routing assumptions rather than restricting agents to trivial tickets by default; well-scoped work may be a useful starting point, but it is not the only explanation for agent performance. | Caveat: PR size is only one proxy for task complexity; it does not fully control for differences in requirements ambiguity, dependency criticality, or architectural novelty.
- Claim: Agent PRs require about the same amount of review-and-fix iteration as human PRs in Greptile's data. | Evidence: The speaker reports 2.1 review cycles to merge for Devin PRs and 2.45 for Codex PRs, with human PRs "right in the middle" and no meaningful statistical difference overall. | Implication: The scalable unit is not a one-shot coding agent; it is a closed-loop system in which an agent receives machine-generated findings, repairs the branch, and repeats until it meets merge criteria. | Caveat: This measure reflects an environment where agents are commonly used to address Greptile's review comments, so it demonstrates an effective agent-plus-validation loop rather than autonomous code generation in isolation.
- Claim: Different coding agents exhibit distinct error distributions, making model-specific controls necessary. | Evidence: Using several million Greptile review comments, Gupta found Claude-associated code was about 1.5x more likely than human code to receive SQL-injection findings, whereas Devin was about half as likely as humans to produce off-by-one issues. | Implication: Maintain per-model scorecards for failure classes—especially security, data access, correctness, and performance—and use those scorecards to inform routing, test augmentation, and human-escalation rules. | Caveat: The talk provides only selected examples and does not disclose the full chart, model versions, task mix, or statistical confidence intervals.
- Claim: AI-driven PR throughput makes manual review, conventional testing, and outsourced QA structurally unable to remain the sole safety gate. | Evidence: Among Greptile users, the median engineer produces about 50 PRs per month, the 90th percentile about 500, and the 99th percentile thousands; Gupta characterizes the upper tail as producing PRs at the rate of new ideas. | Implication: Engineering leadership should invest in automated validation capacity before encouraging unconstrained agent throughput, otherwise review queues and unverified change volume become the limiting factor.
- Claim: A practical automated merge gate should validate contract safety, future risk, and stated intent—not merely linting or test pass/fail. | Evidence: Greptile's three questions are whether a change violates a user contract, increases the propensity of a future contract violation, and fulfills the author's described intent. Its approach includes codebase-context analysis plus sandbox execution, dependency resolution, mocked inputs, and browser-agent interaction. | Implication: Define a policy-driven autonomous-merge lane only for changes that can be evaluated against explicit contracts and validated in an isolated runtime; retain escalations for ambiguous intent, sensitive systems, and insufficient observability. | Caveat: The speaker states that almost one-fifth of Greptile-reviewed PRs merge without human review or testing, but does not specify the risk tiering, service criticality, rollback design, or guardrails used for those merges.
Detailed Brief
Measurement design and what the data can establish
- Claims: The speaker deliberately uses production outcomes rather than code-style judgments to evaluate AI-generated PRs.; He treats a revert as a proxy for a bad PR, automated P0/P1/P2 findings as a proxy for defect severity, and iteration count to merge as a proxy for review burden.; He emphasizes that uptake has looked continuous over the measured period, with no obvious discontinuities tied to new model launches.
- Evidence: The analysis draws on Greptile's PR-review corpus and an average of about four Greptile comments per PR, yielding several million review comments over several months.; The dataset includes organizations described as enterprises or serious product companies, with named customers including NVIDIA, Coinbase, Scale, Datadog, and American Express.
- Caveats: The presentation does not provide sample sizes by agent, confidence intervals, precise date ranges, definitions for "largely AI-generated," or independent replication.; Because Greptile is both the data source and the evaluator that generates bug findings, its results should be treated as strong operational evidence but not a neutral benchmark of all coding-agent quality.
- Implications: For internal evaluation, record provenance explicitly at PR creation rather than relying on metadata heuristics; retain agent, model, version, prompt/task class, tool permissions, and validation results.; Measure post-merge incidents, rollbacks, security escapes, cycle time, and maintenance cost alongside review findings so that quality is assessed across the full software lifecycle.
The validation architecture implied by agent-scale development
- Claims: The speaker's product approach is to validate changes with a swarm of agents that inspect both modified files and code related to those files.; Runtime validation is central: the system is intended to install dependencies, start the application locally, mock inputs, and interact with the application to find breakage.; The target operating state is higher autonomous-merge coverage without relaxing correctness requirements.
- Evidence: Gupta distinguishes this approach from simply automating one existing function such as QA, testing, or manual code review.; He explicitly describes browser agents clicking through locally running applications in a sandbox to attempt to break changes.
- Caveats: Sandbox execution can provide strong confidence only to the extent that environments, dependencies, test data, user workflows, and external integrations are faithfully represented.; The talk does not address authorization boundaries, secret management, supply-chain protections, or containment requirements for agents that resolve dependencies and execute arbitrary PR code.
- Implications: Treat autonomous code validation as a control-plane problem: isolated execution, least-privilege credentials, network controls, policy checks, auditable evidence, and rollback readiness are core design requirements rather than implementation details.
Notable Concepts & Terms
- End-to-end agentic PR: A pull request generated largely or completely by an autonomous coding agent rather than a human using AI only for autocomplete or isolated edits.
- PR provenance signals: The metadata heuristics Gupta used to infer AI authorship: commit author, AI co-author text in PR footers, and coding-agent branch-name prefixes.
- P0 / P1 / P2 findings: Greptile's severity categories for review-detected defects, used in the talk as a proxy for relative code quality.
- Model-specific failure profile: The idea that each coding agent has a distinct distribution of errors, such as higher SQL-injection propensity for Claude-associated PRs in this dataset.
- User contract: The expected behavior or guarantees an application provides to users; Gupta proposes contract preservation as the first merge-validation question.
- Intent validation: Checking whether the implemented PR actually accomplishes what its author claimed it should do, beyond simply compiling or passing tests.
- Sandboxed runtime validation: Executing a PR in an isolated environment, resolving dependencies and exercising mocked inputs or browser flows to detect behavior-level failures.
- Autonomous merge: A PR merging without human review or human testing after automated validation; Greptile reports this for nearly one-fifth of PRs it reviews.
Operator Notes / Why Ken Should Care
- Instrument PR provenance now: require disclosure of agent/model/version, task source, tool permissions, and whether an agent made the final commit; avoid trying to infer this later from Git metadata.
- Build a model-by-failure-class dashboard using your own PR and incident data, then set routing and test policies by model and code-risk tier rather than using a blanket AI-code rule.
- Create an autonomous-merge policy with explicit eligibility criteria: isolated sandbox execution, contract tests, intent specification, sensitive-path exclusions, observable rollout, and fast rollback.
- Stress-test the validation environment before scaling agent throughput: dependency installation, mocked external systems, browser/UI flows, secrets isolation, network egress, and untrusted-code containment are likely control-plane bottlenecks.
- Do not use revert rate alone as the adoption gate; add production incidents, security escapes, architectural regression, ownership burden, and post-merge maintenance measures.
Source/Metadata
- Title: AI-Generated Code Is Already Competing With Human Code — Daksh Gupta, Greptile
- Transcript words: 4631
- Duration seconds: 760
- Timestamp note: No timestamps or chapter markers were present. The supplied transcript substantially repeats the talk's main body, so the stated word count overstates unique spoken content.
Transcript
DAKSH SHARMA How's everyone doing? Awesome. My name is Daksh. I'm one of the co-founders of a company called Grubtile. And at Grubtile, we're working on AI agents that validate pull requests with full context of the code base. The way Grubtile does this is every time there's a pull request, we have a swarm of agents. And they go and look at every file that's changed, every file that's related to the files that changed to figure out if there's bugs. And it also spins up your code in a sandbox, then solves the dependencies, spins up localhost, clicks around to try to break things, mocks inputs, and everything else to figure out if the code is broken. But I'm going to talk about something a little bit different, which is fully autonomous coding agents. So I first moved to San Francisco three years ago to work on AI coding. And this is because GPT 3.5, which came out in 2022, was the first model that seemed like it was really good at programming. At the time, code complete, tab complete, was the main paradigm for coding. The most exciting thing that was happening was your tab complete from cursor and you had co-pilots, code complete. And that was the way in which AI coding works. In 2024, for the first time, multi-file editing started to work. Cursor was the first to do this, and then a couple of other products followed. But now, for the first time, AI could simultaneously edit multiple files at once. But things got really interesting in 2025, because for the first time we had autonomous agents. They could be given a task and they could go and fulfill the task and create entire pull requests all at once. And this came to a precipice in December of last year, when coding agents became literally completely autonomous. And anyone that's working in AI coding knows that December, when the new models came out, was a watershed moment in the history of AI coding. Because these products for the first time were truly autonomous in how they programmed. And all over Twitter, there were these crazy stories of people that were having these agents spin off and open 100 pull requests a day. They were experimenting with polyphasic sleep so they could stay awake to re-trigger their agents from time to time. And as someone that started programming before agents, I was very skeptical that real companies could program this way. So I was really curious. Are these fully viable PRs really actually good? Is this something that's Twitter hype, and you can have independent developers, and maybe small startups with no real customers do this, but anyone with real customers with an actually commercially viable code base, could they still use these end-to-end coding agents? And at Grubtile, we're fortunate to work with some very large companies. We work with NVIDIA, with Coinbase, with Scale, with Datadog, with American Express. And we got really excited to see if these coding agents were being used really at these large companies, and if they were any good at doing real-world coding. So as an amateur data scientist, I decided to go and dive into our data. We review more than a million pull requests a month, and we tend to get a lot of interesting data on what's good and bad about these pull requests. We do this for thousands of companies, and they're usually enterprises or at least companies with serious products with real customers. And so we figured studying this data would yield some really interesting results. It was surprisingly hard to figure out from all the pull requests that Grubtile was reviewing which of the pull requests was actually generated by AI. It was surprisingly difficult to do this. The first thing that I tried was I started looking at the GitHub author field. GitHub, every commit has an author field, and so I went through the authors. And it turns out that less than 1% of pull requests had Codex or Claude or Cursor as the author. But it wasn't intuitive to me that only about 1% of all code was fully AI generated. It just seemed that that number would be much higher. So I started to look for other signals. Now, thankfully, Claude and Codex and some other products leave a PR description in the footer. In the PR description, they put something along the lines of co-authored by Claude or co-authored by Cursor and so on. And it gave me a little bit more signal, more data on which PRs were ostensibly largely, if not completely, AI generated. The third thing I looked at was branch name prefixes. If you use Codex, you know that Codex names branches. And it seemed reasonable that if someone was using these coding agents in such a way that they were writing the branch names or writing the PR descriptions, that the PRs were largely AI generated. That seemed like a reasonable assumption to me. So now I had a good set of signals to indicate whether a pull request was in fact completely or largely AI generated or not. And it came to the result that about a quarter of all the pull requests that Grubtile was reviewing in any given month were completely or at least largely generated by AI. Then I backtracked this data to the last 12 months. And it turns out that this number is growing really fast. In fact, early last year, fewer than 1% of pull requests had any evidence of being completely AI generated. So this number is growing very fast as model performance is getting better. What's also interesting is that you can't tell when new models came out on this chart. It seemed like the progress is very continuous. These things are just generally diffusing into the economy at a generally high pace. So then came the next question. Everyone's coding. Everyone's producing these PRs that are end-to-end agentic. Not human plus AI, but literally AI. The question is, are these PRs any good? First question, what does it mean for a PR to be good? That seemed like actually a very important question to answer before we got into judging these AI-generated PRs. And so I took a few methods to try this. The first one I tried was revert rates. If a pull request was reverted, it was probably pretty bad. And so that seemed like a sincere guess at what a bad PR could be. And so I started tracking, which PRs were in fact reverted. GitHub conveniently names its branches with a revert-PR number and PR name. And so I was able to track the rate at which these things were being reverted. And I found some interesting data. Codex PRs were reverted about one out of every thousand pull requests. Devon, one out of every three and a half thousand pull requests. Humans were right in the middle at about two and a half. So there didn't seem to be a very big difference between the rate at which pull requests were reverted from people versus agents in my study. Now, I was very skeptical of this. And I figured this is probably because humans are making agents do the easier work, the simpler, well-scoped pull requests. And the more complex work was being done by humans. And of course, complex work had a higher propensity to be reverted because the pull requests were more complicated and had larger surface areas, risks, and so on. So I started to measure the average size of PR and whether the revert rates were aligned for both of those. And I found really interesting data. Turns out there's actually very little correlation, if any, between the size of PRs that humans were getting reverted versus size of PRs from agents that were being reverted. So this does not seem like strong evidence that the human PRs are any better than the agent PRs. The second set of signals I started looking at was Grubtile comments. Now, Grubtile is reviewing all these pull requests. And Grubtile finds P0s, P1s, P2s across all these changes. And it seemed reasonable that if Grubtile was finding more bugs in the code, that the code was most likely worse. So then I started tracking the counts of P0s, P1s, and P2s in all these changes. And interestingly, once again, there was not that much of a difference. In fact, three out of the four agents that we tested performed better than humans in terms of the rate at which they were producing P0s. They were producing fewer P0s than humans were. Similarly, the case with P1s and P2s. Broadly speaking, human-generated PRs were about equal in quality to agent-generated PRs based on this data. Then I started looking at a third piece of data. The most common way to use Grubtile is to have it review a pull request, get some set of comments, and then have the agent go and look at those comments, address them, and make a new commit on the pull request branch. That is the most common way to use Grubtile. And so it seems reasonable that if the pull requests were higher quality, it would require fewer iterations before they would be merged. And so I started tracking the number of iterations, the number of review rounds, between when the pull request was opened and when it was merged. Sure enough, very little difference. Devon's PRs, 2.1. Codex PRs, 2.45. Number of review cycles to merge. And humans, right in the middle. Once again, very little, if any, statistical difference between human-generated and AI-generated pull requests in terms of how many iterations before they were ready to merge. So now that there wasn't any strong evidence that agent-generated PRs were worse than human-generated PRs, I got curious if there were qualitative differences. Maybe the types of ways in which agents failed were different from the types of ways that humans failed. And so then I looked at the corpus of Grubtile's comments. It makes, on average, about four comments per pull request. So we had this corpus of several million comments across the last several months. And so I started scanning them for specific phrases and words. For instance, SQL injection or N plus one query. And I plotted the frequency with which these terms occur in Grubtile comments for these various agents. And so I made this chart of patterns of failure. To interpret this chart, you can assume that 1x is the human propensity for producing that type of error across the entire chart. And you can see that there's actually quite a lot of variation in the types of failures that these agents seem to produce. For instance, Claude is one and a half times more likely to produce a SQL injection error than humans. Devon is about half as likely as humans to produce an off-by-one issue. I found it very interesting that there was this much variation in how these agents were performing, and how different their failure modes were from humans. So it turns out that in spite of my initial skepticism around the enterprise usability of end-to-end coding agents, the evidence seems to suggest that they're here and they probably can contribute in real meaningful ways to enterprise coding environments. And so we started to think a little bit more about what code review would look like in such a world. Here's an interesting stat. Today, Grubtile is used by several tens of thousands of engineers every single week to review all of their code. So we also know, generally speaking, the number of pull requests that each of these people are writing. The median Grubtile user writes 50 pull requests a month. So call it about two per workday. The 90th percentile writes 500 pull requests per month. That is a drastic difference between the median and the P90. The P99 is in the thousands of pull requests. That isn't very interesting in its own sense because that means that the people on the margins are actually producing pull requests at the rate at which they're coming up with new ideas. And beyond being interesting, it also makes you wonder, well, there are existing systems for validating this code, which is manual code review, of course, testing, maybe you have a QA firm that you work with. Naturally they can't scale to that same degree. And at Grubtile, we decided to take a first principles view of what really good validation could look like. Instead of saying that we wanted to automate QA or automate testing or automate code review, we took a step back and said, what would we need to do for anyone to be able to merge hundreds of pull requests a month in an enterprise environment where it really matters if the code is correct? What would need to happen between when the code was expressed into a pull request and when it was merged into a pull request and deployed safely? We figured we only actually had to answer three questions. The first one is, does this change violate the user contracts? The second one, does it increase the propensity of a future violation of the user contract, whatever the user contract might be for that application? And third, does it fulfill the intent that the author described? Does the pull request do the thing that the author wanted to do? And so we started approaching this problem from the ground level and said, OK, agents can probably figure out if something's going to violate the user contract and detect bugs. If you let it spin up the code in a sandbox, have it install the dependencies, mock the inputs, run the browser agents, you can probably start to discover most of the issues that could occur. You can get a pretty high degree of confidence on merge. Today, almost a fifth of all the pull requests that Grubtile reviews are merged without any human review or without any human testing. Because I find it very interesting. And that is a number that we care a lot about and want to bring higher and higher, of course, within the guardrails of producing really high quality code. Thank you so much. My name is Daksh, one of the co-founders of Grubtile. We have a booth here which you should come to. And if you're interested in trying Grubtile, you can find us at grubtile.com. You can try it today for free and we'd love to hear your feedback. Thank you so much. And all over Twitter, while these crazy stories of people that were having these agents spin off and open 100 pull requests a day, they were experimenting with polyphasic sleep so they could stay awake to re-trigger their agents from time to time. And as someone that started programming before agents, I was very skeptical that real companies could program this way. So I was really, really curious. Are these fully viable PRs really actually good? Is this something that's sort of a Twitter hype, and you can have independent developers, and maybe like small startups with no real customers do this, but anyone with real customers with an actually commercially viable code base, could they still use these end-to-end coding agents? And at Grubtile, we're fortunate to work with some very large companies. We work with NVIDIA, with Coinbase, with Scale, with Datadog, with American Express. And we got really excited to see if these coding agents were being used really at these large companies, and if they were any good at doing real-world coding. So as an amateur data scientist, I decided to go and dive into our data. We review more than a million pull requests a month, and we tend to get a lot of interesting data on what's good and bad about these pull requests. We do this for thousands and thousands of companies, and they're usually enterprises or at least companies with serious products with real customers. And so we figured actually studying this data would yield some really interesting results. It was surprisingly hard to figure out from all the pull requests that Grubtile was reviewing which of the pull requests was actually generated by AI. It was surprisingly difficult to do this. The first thing that I tried was I started looking at the GitHub author field. GitHub, every commit has an author field, and so went through the authors. And it turns out that less than 1% of pull requests had Codex or Claude or Cursor as the author. But it wasn't intuitive to me that only about 1% of all code was fully AI generated. It just seemed intuitive that that number would be much higher. So I started to look for other signals. Now, thankfully, Claude and Codex and some other products leave a PR description in the footer. In the PR description, they put something along the lines of co-authored by Claude or co-authored by Cursor and so on. And it gave me a little bit more signal, a little bit more data on which PRs were ostensibly largely, if not completely, AI generated. The third thing I looked at was branch name prefixes. If you use Codex, you know that Codex names as branches. And it seemed reasonable that if someone was using these coding agents in such a sense that they were writing the branch names or writing the PR descriptions, that the PRs were largely AI generated. That seemed like a reasonable assumption to me. So now I had a good set of signals to indicate whether a pull request was in fact completely or largely AI generated or not. And it came to the result that about a quarter of all the pull requests that Grubtel was reviewing in any given month, were completely or at least largely generated by AI. Then I backtracked this data to the last 12 months. And it turns out that this number is growing really fast. In fact, early last year, fewer than 1% of pull requests that any evidence of being completely AI generated. So this number is going very, very fast as model performance is getting better. What's also interesting is that you can't tell when new models came out on this chart. It seemed like the progress is very continuous. These things are just generally diffusing into the economy at a generally high pace. So then came the next question. Everyone's vibe coding. Everyone's producing these PRs that are end-to-end agentic. Not human plus AI, but literally AI. The question is, are these PRs any good? First question, what does it mean for PR to be good? That seemed like actually a very important question to answer before we got into judging these AI generated PRs. And so I took a few methods to try this. The first one I tried was revert rates. If a pull request was reverted, it was probably pretty bad. And so that seemed like a pretty sincere sort of guess at what a bad PR could be. And so I started tracking, okay, which PRs were in fact reverted. GitHub conveniently named its branches with a revert-PR number and PR name. And so I was able to track the rate at which these things were being reverted. And I found some interesting data. Codex PRs were reverted about one out of every thousand pull requests. Devon wants every three and a half times every thousand pull requests. Humans were right in the middle at about two and a half. So there didn't seem to be a very big difference between the rate at which pull requests were reverted from people versus agents in my study. Now, I was very skeptical of this. And I figured, okay, this is probably because humans are making agents do the easier work, the simpler, well-scoped pull requests. And the more complex work was being done by humans. And of course, complex work had a higher propensity to be reverted because the pull requests were more complicated and larger surface areas, risks, and so on. So I started to measure the average size of PR and whether the revert rates were aligned for both of those. And I found really interesting data. Turns out there's actually very little correlation, if any, between the size of PRs that humans were getting reverted versus size of PRs from agents that were being reverted. So this does not seem like actually very strong evidence that the human PRs are any better than the agent PRs. The second set of signals I started looking at was gruptile comments. Now, gruptile is reviewing all these pull requests. And gruptile finds p0s, p1s, p2s across all these changes. And it seemed reasonable that if gruptile was finding more bugs in the code, that the code was most likely worse. So then I started tracking the counts of p0s, p1s, and p2s and all these changes. And interestingly, once again, there was not that much of a difference. In fact, three out of the four agents that we tested performed better than humans in terms of the rate at which they were producing p0s. They were producing fewer p0s than humans were. Similarly, the case with p1s and p2s. Broadly speaking, human-generated PRs were about equal in quality to agent-generated PRs based on this data. Then I started looking at a third piece of data. The most common way to use gruptile is to have it review a pull request, get some set of comments, and then have the agent go and look at those comments, address them, and make a new commit on the pull request branch. That is the most common way to use gruptile. And so it seems reasonable that if the pull requests were higher quality, it would require fewer iterations before they would be merged. And so I started tracking the number of iterations, the number of review rounds, between when the pull request was opened and when it was merged. Sure enough, very little difference. Devon's PRs, 2.1. Codex PRs, 2.45. Number of review cycles to merge. And humans, right in the middle. Once again, very little, if any, statistical difference between human-generated and AI-generated pull requests in terms of how many iterations before they were ready to merge. So now that there wasn't any strong evidence that agent-generated PRs were worse than human-generated PRs, I got curious if there were qualitative differences. Maybe the types of ways in which agents failed were different from the types of ways that humans failed. And so then I looked at the corpus of gruptile's comments. It makes, on average, about four comments per pull request. So we had this corpus of several million comments across the last several months. And so I started scanning them for specific phrases and words. For instance, SQL injection or n plus one query. And I plotted the frequency with which these terms occur in gruptile comments for these various agents. And so I made this chart of patterns of failure. To interpret this chart, you can assume that 1x is the human propensity for producing that type of error across the entire chart. And you can kind of see that there's actually quite a lot of variation in the types of failures that these agents seem to produce. For instance, Claude is one and a half times more likely to produce a SQL injection error than humans. Devon is about half as likely as humans to produce an off-bypass issue. I found it very interesting that there was this much variation in how these agents were performing. And how different their failure modes were from humans. So it turns out that in spite of my initial skepticism around the enterprise usability of end-to-end coding agents, the evidence seems to suggest that they're here and they probably can contribute in real meaningful ways to enterprise coding environments. And so we started to think a little bit more about what code review would look like in such a world. Here's an interesting stat. Today, gruptile is used by several tens of thousands of engineers every single week to review all of their code. So we also know, generally speaking, the number of pull requests that each of these people are writing. The median gruptile user writes 50 pull requests a month. So call it about two per workday. The 90th percentile writes 500 pull requests per month. That is a drastic difference between the median and the P90. The P99 is in the thousands of pull requests. That isn't very interesting in its own sense because that means that the people on the margins are actually producing pull requests at the rate at which they're coming up with new ideas. And beyond being interesting, it also makes you wonder, well, there are existing systems for validating this code, which is, manual code review, of course, testing, maybe you have a QA firm that you work with, naturally can scale to that same degree. And at gruptile, we decided to take sort of a first principles view of what really good validation could look like. Instead of saying that we wanted to automate QA or automate testing or automate code review, we took a step back and said, what would we need to do for anyone to be able to merge hundreds of pull requests a month in an enterprise environment where it really matters if the code is correct? What would need to happen between when the code was expressed into a pull request and when it was merged into a pull request? And deployed safely. We figured we only actually had to answer three questions. The first one is, does this change violate the user contracts? The second one, does it increase the propensity of a future violation of the user contract, whatever the user contract might be for that application? And third, does it fulfill the intent that the author described? Does the pull request do the thing that the author wanted to do? And so we started approaching this problem from base ground level and said, OK, agents can probably figure out if something's going to violate the user contract and detect bugs. If you let it spin up the code in a sandbox, have it install the dependencies, mock the inputs, run the browser agents, you can probably start to discover most of the issues that could occur. You can get a pretty high degree of confidence on merge. Today, almost a fifth of all the pull requests of Grubtile reviews are merged without any human review or without any human testing. Because I find it very interesting. And that is a number that we care a lot about and want to bring higher and higher, of course, within the guardrails of producing really high quality code. Thank you so much. My name is Dax, one of the co-founders of Grubtile. We have a booth here which you should come to. And if you're interested in trying Grubtile, you can find us at grubtile.com. You can try it today for free and we'd love to hear your feedback. Thank you so much. Thank you.