DAKSH SHARMA How's everyone doing? Awesome. My name is Daksh. I'm one of the co-founders of a company called Grubtile. And at Grubtile, we're working on AI agents that validate pull requests with full context of the code base. The way Grubtile does this is every time there's a pull request, we have a swarm of agents. And they go and look at every file that's changed, every file that's related to the files that changed to figure out if there's bugs. And it also spins up your code in a sandbox, then solves the dependencies, spins up localhost, clicks around to try to break things, mocks inputs, and everything else to figure out if the code is broken.
But I'm going to talk about something a little bit different, which is fully autonomous coding agents. So I first moved to San Francisco three years ago to work on AI coding. And this is because GPT 3.5, which came out in 2022, was the first model that seemed like it was really good at programming. At the time, code complete, tab complete, was the main paradigm for coding. The most exciting thing that was happening was your tab complete from cursor and you had co-pilots, code complete. And that was the way in which AI coding works.
In 2024, for the first time, multi-file editing started to work. Cursor was the first to do this, and then a couple of other products followed. But now, for the first time, AI could simultaneously edit multiple files at once.
But things got really interesting in 2025, because for the first time we had autonomous agents. They could be given a task and they could go and fulfill the task and create entire pull requests all at once. And this came to a precipice in December of last year, when coding agents became literally completely autonomous. And anyone that's working in AI coding knows that December, when the new models came out, was a watershed moment in the history of AI coding. Because these products for the first time were truly autonomous in how they programmed.
And all over Twitter, there were these crazy stories of people that were having these agents spin off and open 100 pull requests a day. They were experimenting with polyphasic sleep so they could stay awake to re-trigger their agents from time to time.
And as someone that started programming before agents, I was very skeptical that real companies could program this way. So I was really curious. Are these fully viable PRs really actually good? Is this something that's Twitter hype, and you can have independent developers, and maybe small startups with no real customers do this, but anyone with real customers with an actually commercially viable code base, could they still use these end-to-end coding agents?
And at Grubtile, we're fortunate to work with some very large companies. We work with NVIDIA, with Coinbase, with Scale, with Datadog, with American Express. And we got really excited to see if these coding agents were being used really at these large companies, and if they were any good at doing real-world coding.
So as an amateur data scientist, I decided to go and dive into our data. We review more than a million pull requests a month, and we tend to get a lot of interesting data on what's good and bad about these pull requests. We do this for thousands of companies, and they're usually enterprises or at least companies with serious products with real customers. And so we figured studying this data would yield some really interesting results. It was surprisingly hard to figure out from all the pull requests that Grubtile was reviewing which of the pull requests was actually generated by AI. It was surprisingly difficult to do this.
The first thing that I tried was I started looking at the GitHub author field. GitHub, every commit has an author field, and so I went through the authors. And it turns out that less than 1% of pull requests had Codex or Claude or Cursor as the author. But it wasn't intuitive to me that only about 1% of all code was fully AI generated. It just seemed that that number would be much higher.
So I started to look for other signals. Now, thankfully, Claude and Codex and some other products leave a PR description in the footer. In the PR description, they put something along the lines of co-authored by Claude or co-authored by Cursor and so on. And it gave me a little bit more signal, more data on which PRs were ostensibly largely, if not completely, AI generated.
The third thing I looked at was branch name prefixes. If you use Codex, you know that Codex names branches. And it seemed reasonable that if someone was using these coding agents in such a way that they were writing the branch names or writing the PR descriptions, that the PRs were largely AI generated. That seemed like a reasonable assumption to me. So now I had a good set of signals to indicate whether a pull request was in fact completely or largely AI generated or not. And it came to the result that about a quarter of all the pull requests that Grubtile was reviewing in any given month were completely or at least largely generated by AI.
Then I backtracked this data to the last 12 months. And it turns out that this number is growing really fast. In fact, early last year, fewer than 1% of pull requests had any evidence of being completely AI generated. So this number is growing very fast as model performance is getting better. What's also interesting is that you can't tell when new models came out on this chart. It seemed like the progress is very continuous. These things are just generally diffusing into the economy at a generally high pace.
So then came the next question. Everyone's coding. Everyone's producing these PRs that are end-to-end agentic. Not human plus AI, but literally AI. The question is, are these PRs any good? First question, what does it mean for a PR to be good? That seemed like actually a very important question to answer before we got into judging these AI-generated PRs.
And so I took a few methods to try this. The first one I tried was revert rates. If a pull request was reverted, it was probably pretty bad. And so that seemed like a sincere guess at what a bad PR could be. And so I started tracking, which PRs were in fact reverted. GitHub conveniently names its branches with a revert-PR number and PR name. And so I was able to track the rate at which these things were being reverted.
And I found some interesting data. Codex PRs were reverted about one out of every thousand pull requests. Devon, one out of every three and a half thousand pull requests. Humans were right in the middle at about two and a half. So there didn't seem to be a very big difference between the rate at which pull requests were reverted from people versus agents in my study.
Now, I was very skeptical of this. And I figured this is probably because humans are making agents do the easier work, the simpler, well-scoped pull requests. And the more complex work was being done by humans. And of course, complex work had a higher propensity to be reverted because the pull requests were more complicated and had larger surface areas, risks, and so on.
So I started to measure the average size of PR and whether the revert rates were aligned for both of those. And I found really interesting data. Turns out there's actually very little correlation, if any, between the size of PRs that humans were getting reverted versus size of PRs from agents that were being reverted. So this does not seem like strong evidence that the human PRs are any better than the agent PRs.
The second set of signals I started looking at was Grubtile comments. Now, Grubtile is reviewing all these pull requests. And Grubtile finds P0s, P1s, P2s across all these changes. And it seemed reasonable that if Grubtile was finding more bugs in the code, that the code was most likely worse. So then I started tracking the counts of P0s, P1s, and P2s in all these changes. And interestingly, once again, there was not that much of a difference. In fact, three out of the four agents that we tested performed better than humans in terms of the rate at which they were producing P0s. They were producing fewer P0s than humans were. Similarly, the case with P1s and P2s.
Broadly speaking, human-generated PRs were about equal in quality to agent-generated PRs based on this data. Then I started looking at a third piece of data. The most common way to use Grubtile is to have it review a pull request, get some set of comments, and then have the agent go and look at those comments, address them, and make a new commit on the pull request branch. That is the most common way to use Grubtile. And so it seems reasonable that if the pull requests were higher quality, it would require fewer iterations before they would be merged.
And so I started tracking the number of iterations, the number of review rounds, between when the pull request was opened and when it was merged. Sure enough, very little difference. Devon's PRs, 2.1. Codex PRs, 2.45. Number of review cycles to merge. And humans, right in the middle. Once again, very little, if any, statistical difference between human-generated and AI-generated pull requests in terms of how many iterations before they were ready to merge.
So now that there wasn't any strong evidence that agent-generated PRs were worse than human-generated PRs, I got curious if there were qualitative differences. Maybe the types of ways in which agents failed were different from the types of ways that humans failed.
And so then I looked at the corpus of Grubtile's comments. It makes, on average, about four comments per pull request. So we had this corpus of several million comments across the last several months. And so I started scanning them for specific phrases and words. For instance, SQL injection or N plus one query. And I plotted the frequency with which these terms occur in Grubtile comments for these various agents.
And so I made this chart of patterns of failure. To interpret this chart, you can assume that 1x is the human propensity for producing that type of error across the entire chart. And you can see that there's actually quite a lot of variation in the types of failures that these agents seem to produce. For instance, Claude is one and a half times more likely to produce a SQL injection error than humans. Devon is about half as likely as humans to produce an off-by-one issue. I found it very interesting that there was this much variation in how these agents were performing, and how different their failure modes were from humans.
So it turns out that in spite of my initial skepticism around the enterprise usability of end-to-end coding agents, the evidence seems to suggest that they're here and they probably can contribute in real meaningful ways to enterprise coding environments. And so we started to think a little bit more about what code review would look like in such a world.
Here's an interesting stat. Today, Grubtile is used by several tens of thousands of engineers every single week to review all of their code. So we also know, generally speaking, the number of pull requests that each of these people are writing. The median Grubtile user writes 50 pull requests a month. So call it about two per workday. The 90th percentile writes 500 pull requests per month. That is a drastic difference between the median and the P90. The P99 is in the thousands of pull requests.
That isn't very interesting in its own sense because that means that the people on the margins are actually producing pull requests at the rate at which they're coming up with new ideas. And beyond being interesting, it also makes you wonder, well, there are existing systems for validating this code, which is manual code review, of course, testing, maybe you have a QA firm that you work with. Naturally they can't scale to that same degree.
And at Grubtile, we decided to take a first principles view of what really good validation could look like. Instead of saying that we wanted to automate QA or automate testing or automate code review, we took a step back and said, what would we need to do for anyone to be able to merge hundreds of pull requests a month in an enterprise environment where it really matters if the code is correct? What would need to happen between when the code was expressed into a pull request and when it was merged into a pull request and deployed safely?
We figured we only actually had to answer three questions. The first one is, does this change violate the user contracts? The second one, does it increase the propensity of a future violation of the user contract, whatever the user contract might be for that application? And third, does it fulfill the intent that the author described? Does the pull request do the thing that the author wanted to do?
And so we started approaching this problem from the ground level and said, OK, agents can probably figure out if something's going to violate the user contract and detect bugs. If you let it spin up the code in a sandbox, have it install the dependencies, mock the inputs, run the browser agents, you can probably start to discover most of the issues that could occur. You can get a pretty high degree of confidence on merge.
Today, almost a fifth of all the pull requests that Grubtile reviews are merged without any human review or without any human testing. Because I find it very interesting. And that is a number that we care a lot about and want to bring higher and higher, of course, within the guardrails of producing really high quality code. Thank you so much. My name is Daksh, one of the co-founders of Grubtile. We have a booth here which you should come to. And if you're interested in trying Grubtile, you can find us at grubtile.com. You can try it today for free and we'd love to hear your feedback. Thank you so much. And all over Twitter, while these crazy stories of people
that were having these agents spin off and open 100 pull requests a day, they were experimenting with polyphasic sleep so they could stay awake to re-trigger their agents from time to time. And as someone that started programming before agents, I was very skeptical that real companies could program this way. So I was really, really curious. Are these fully viable PRs really actually good? Is this something that's sort of a Twitter hype, and you can have independent developers, and maybe like small startups with no real customers do this, but anyone with real customers with an actually commercially viable code base, could they still use these end-to-end coding agents?
And at Grubtile, we're fortunate to work with some very large companies. We work with NVIDIA, with Coinbase, with Scale, with Datadog, with American Express. And we got really excited to see if these coding agents were being used really at these large companies, and if they were any good at doing real-world coding.
So as an amateur data scientist, I decided to go and dive into our data. We review more than a million pull requests a month, and we tend to get a lot of interesting data on what's good and bad about these pull requests. We do this for thousands and thousands of companies, and they're usually enterprises or at least companies with serious products with real customers. And so we figured actually studying this data would yield some really interesting results. It was surprisingly hard to figure out from all the pull requests that Grubtile was reviewing which of the pull requests was actually generated by AI. It was surprisingly difficult to do this.
The first thing that I tried was I started looking at the GitHub author field. GitHub, every commit has an author field, and so went through the authors. And it turns out that less than 1% of pull requests had Codex or Claude or Cursor as the author. But it wasn't intuitive to me that only about 1% of all code was fully AI generated. It just seemed intuitive that that number would be much higher. So I started to look for other signals. Now, thankfully, Claude and Codex and some other products leave a PR description in the footer. In the PR description, they put something along the lines of co-authored by Claude or co-authored by Cursor and so on.
And it gave me a little bit more signal, a little bit more data on which PRs were ostensibly largely, if not completely, AI generated. The third thing I looked at was branch name prefixes. If you use Codex, you know that Codex names as branches. And it seemed reasonable that if someone was using these coding agents in such a sense that they were writing the branch names or writing the PR descriptions, that the PRs were largely AI generated. That seemed like a reasonable assumption to me. So now I had a good set of signals to indicate whether a pull request was in fact completely or largely AI generated or not.
And it came to the result that about a quarter of all the pull requests that Grubtel was reviewing in any given month, were completely or at least largely generated by AI. Then I backtracked this data to the last 12 months. And it turns out that this number is growing really fast. In fact, early last year, fewer than 1% of pull requests that any evidence of being completely AI generated. So this number is going very, very fast as model performance is getting better. What's also interesting is that you can't tell when new models came out on this chart. It seemed like the progress is very continuous.
These things are just generally diffusing into the economy at a generally high pace. So then came the next question. Everyone's vibe coding. Everyone's producing these PRs that are end-to-end agentic. Not human plus AI, but literally AI. The question is, are these PRs any good? First question, what does it mean for PR to be good? That seemed like actually a very important question to answer before we got into judging these AI generated PRs. And so I took a few methods to try this. The first one I tried was revert rates. If a pull request was reverted, it was probably pretty bad. And so that seemed like a pretty sincere sort of guess at what a bad PR could be.
And so I started tracking, okay, which PRs were in fact reverted. GitHub conveniently named its branches with a revert-PR number and PR name. And so I was able to track the rate at which these things were being reverted. And I found some interesting data. Codex PRs were reverted about one out of every thousand pull requests. Devon wants every three and a half times every thousand pull requests. Humans were right in the middle at about two and a half. So there didn't seem to be a very big difference between the rate at which pull requests were reverted from people versus agents in my study. Now, I was very skeptical of this.
And I figured, okay, this is probably because humans are making agents do the easier work, the simpler, well-scoped pull requests. And the more complex work was being done by humans. And of course, complex work had a higher propensity to be reverted because the pull requests were more complicated and larger surface areas, risks, and so on. So I started to measure the average size of PR and whether the revert rates were aligned for both of those. And I found really interesting data. Turns out there's actually very little correlation, if any, between the size of PRs that humans were getting reverted versus size of PRs from agents that were being reverted.
So this does not seem like actually very strong evidence that the human PRs are any better than the agent PRs. The second set of signals I started looking at was gruptile comments. Now, gruptile is reviewing all these pull requests. And gruptile finds p0s, p1s, p2s across all these changes. And it seemed reasonable that if gruptile was finding more bugs in the code, that the code was most likely worse. So then I started tracking the counts of p0s, p1s, and p2s and all these changes. And interestingly, once again, there was not that much of a difference.
In fact, three out of the four agents that we tested performed better than humans in terms of the rate at which they were producing p0s. They were producing fewer p0s than humans were. Similarly, the case with p1s and p2s. Broadly speaking, human-generated PRs were about equal in quality to agent-generated PRs based on this data. Then I started looking at a third piece of data. The most common way to use gruptile is to have it review a pull request, get some set of comments, and then have the agent go and look at those comments, address them, and make a new commit on the pull request branch. That is the most common way to use gruptile.
And so it seems reasonable that if the pull requests were higher quality, it would require fewer iterations before they would be merged. And so I started tracking the number of iterations, the number of review rounds, between when the pull request was opened and when it was merged. Sure enough, very little difference. Devon's PRs, 2.1. Codex PRs, 2.45. Number of review cycles to merge. And humans, right in the middle. Once again, very little, if any, statistical difference between human-generated and AI-generated pull requests in terms of how many iterations before they were ready to merge.
So now that there wasn't any strong evidence that agent-generated PRs were worse than human-generated PRs, I got curious if there were qualitative differences. Maybe the types of ways in which agents failed were different from the types of ways that humans failed. And so then I looked at the corpus of gruptile's comments. It makes, on average, about four comments per pull request. So we had this corpus of several million comments across the last several months. And so I started scanning them for specific phrases and words. For instance, SQL injection or n plus one query.
And I plotted the frequency with which these terms occur in gruptile comments for these various agents. And so I made this chart of patterns of failure. To interpret this chart, you can assume that 1x is the human propensity for producing that type of error across the entire chart. And you can kind of see that there's actually quite a lot of variation in the types of failures that these agents seem to produce. For instance, Claude is one and a half times more likely to produce a SQL injection error than humans. Devon is about half as likely as humans to produce an off-bypass issue.
I found it very interesting that there was this much variation in how these agents were performing. And how different their failure modes were from humans. So it turns out that in spite of my initial skepticism around the enterprise usability of end-to-end coding agents, the evidence seems to suggest that they're here and they probably can contribute in real meaningful ways to enterprise coding environments. And so we started to think a little bit more about what code review would look like in such a world. Here's an interesting stat. Today, gruptile is used by several tens of thousands of engineers every single week to review all of their code.
So we also know, generally speaking, the number of pull requests that each of these people are writing. The median gruptile user writes 50 pull requests a month. So call it about two per workday. The 90th percentile writes 500 pull requests per month. That is a drastic difference between the median and the P90. The P99 is in the thousands of pull requests. That isn't very interesting in its own sense because that means that the people on the margins are actually producing pull requests at the rate at which they're coming up with new ideas. And beyond being interesting, it also makes you wonder, well, there are existing systems for validating this code, which is,
manual code review, of course, testing, maybe you have a QA firm that you work with, naturally can scale to that same degree. And at gruptile, we decided to take sort of a first principles view of what really good validation could look like. Instead of saying that we wanted to automate QA or automate testing or automate code review, we took a step back and said, what would we need to do for anyone to be able to merge hundreds of pull requests a month in an enterprise environment where it really matters if the code is correct? What would need to happen between when the code was expressed into a pull request and when it was merged into a pull request? And deployed safely.
We figured we only actually had to answer three questions. The first one is, does this change violate the user contracts? The second one, does it increase the propensity of a future violation of the user contract, whatever the user contract might be for that application? And third, does it fulfill the intent that the author described? Does the pull request do the thing that the author wanted to do? And so we started approaching this problem from base ground level and said, OK, agents can probably figure out if something's going to violate the user contract and detect bugs.
If you let it spin up the code in a sandbox, have it install the dependencies, mock the inputs, run the browser agents, you can probably start to discover most of the issues that could occur. You can get a pretty high degree of confidence on merge. Today, almost a fifth of all the pull requests of Grubtile reviews are merged without any human review or without any human testing. Because I find it very interesting. And that is a number that we care a lot about and want to bring higher and higher, of course, within the guardrails of producing really high quality code.
Thank you so much. My name is Dax, one of the co-founders of Grubtile. We have a booth here which you should come to. And if you're interested in trying Grubtile, you can find us at grubtile.com. You can try it today for free and we'd love to hear your feedback. Thank you so much.
Thank you.