Open Reader

How GitHub Deals With 17 Million Pull Requests a Month

completed 28:07 Jun 17, 2026 Watch on YouTube

Current Status

completed

Video ID

OCEVqy8kl7Q

RAG / Chat

Enabled
How GitHub Deals With 17 Million Pull Requests a Month
Description

Last year, there were 1 billion commits on GitHub. This year, Kyle Daigle expects that number to exceed 14 billion, a two-component explosion caused by more humans—and their agents—issuing pull requests. In March alone, 17 million pull requests on GitHub were created by agents. Daigle is the COO of GitHub and Microsoft’s chief marketing officer for developer products. He’s been at GitHub for 13 years, and is paying close attention to how AI is expanding the platform’s user base. Along with agents, legal, sales, and marketing professionals are building apps with the GitHub Copilot app. The line between developer and non-developer is disappearing. On this episode of AI & I, guest host Mike Taylor sat down with Daigle at Microsoft Build to discuss how GitHub is building infrastructure for an agent-native world: agentic code review, model routers that automatically select the right model for the task, and a philosophy that the most durable advantage in this market is developer choice. If you found this episode interesting, please like, subscribe, comment, and share! Want even more? To hear more from Mike Taylor: Subscribe to Every: https://every.to/subscribe Follow him on X: https://x.com/hammer_mt Timestamps for YouTube: 00:00:52: Introduction 00:03:27: The agentic PR flood 00:04:33: GitHub's approach to helping open-source maintainers manage the surge 00:06:15: What 14 billion commits means for code quality 00:08:03: Moving from per-seat licensing to usage-based pricing 00:09:45: Kyle's dual role as GitHub COO and Microsoft's chief marketing officer for developers 00:13:03: Developer choice as competitive moat 00:14:57: How to balance dogfooding your own tools with staying honest about the competition 00:19:45: Hill climbing, frontier tuning, and solving the model-routing problem 00:24:45: Kyle's agentic communication hack Links to resources mentioned in the episode: Kyle Daigle on X: https://x.com/kdaigle Mike Taylor on Every: https://every.to/@mike_2114 Mike’s pi

Summary

Generated by claude-sonnet-4-5

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: GitHub is experiencing exponential agent-driven code creation (17M agent PRs/month, 14B commits projected vs. 1B last year), forcing rethinking of developer tooling, pricing models, open source maintenance, and personalized AI workflows
  • Why it matters: GitHub COO provides concrete metrics on agent adoption at scale, reveals enterprise strategy for hill-climbing/frontier-tuning, addresses the $200-to-$2000 subscription cost crisis, and shows how personalization (not raw capability) is the long-term moat
  • Best use: Ken should watch to understand real agent economy metrics, GitHub's choice/interop positioning vs. competitors, how enterprises handle model routing/cost explosion, and practical self-improvement agent loops from a COO who builds his own feedback agents

Executive Summary

Kyle Daigle, GitHub COO and Microsoft's CMO for Developers, reports that GitHub saw 17 million agent-created pull requests in March 2024 alone—exploding from 1 billion commits all of last year to a projected 14 billion this year if growth were linear (which he says it won't be, implying faster). This is not 'slop': developers are using 1-to-N agents to multiply their context and skills, and all that code lands on GitHub regardless of which coding tool created it. GitHub is preparing infrastructure for the next wave, focused on supporting 'everyone's agent moment,' not just GitHub Copilot.

The business model is in flux. Freemium worked when humans slept, but agents run 24/7. GitHub hasn't locked a new pricing model yet, but Daigle sees two truths: (1) individual devs need free/affordable baseline access (like free private repos became standard), and (2) enterprises running 150 agents simultaneously need usage-based tiers. The real cost explosion—$200/month subscriptions ballooning to $2,000—will be solved by automatic model routing (choose expensive frontier models only when needed, drop to small models for find-and-replace) and frontier-tuning/hill-climbing that personalizes models to user context, reducing wasted tokens.

Open source maintainers are drowning in agent PRs. GitHub's answer: agentic code review (finds vulnerabilities, implements fixes via comments), agentic merge (handles CI/policies/manual steps), and modular controls (like Mitchell Hashimoto's vouch system) that each community can configure. GitHub won't impose a standard until one emerges organically. Daigle insists on developer choice as the differentiator: GitHub partners with Anthropic, OpenAI, Google, and any model/agent provider, refusing to become a walled garden even as competitors build mousetraps.

Daigle personally runs a self-improvement agent loop: he feeds all his emails, Slack, and scripts into Claude ('Baxter'), which reads the last seven days and gives him critical feedback on clarity, metaphor usage, and comms patterns. He finds humans accept robot critique better than peer critique. This personal practice reflects GitHub's long-term bet: personalization/context/memory (not raw model capability) is the durable advantage, because models will commoditize and local SLMs will handle much of the work soon. Hill-climbing (weekly evals, thumbs-up/down data, acceptance rates, user sentiment checks) is the company-wide discipline, and frontier-tuning lets enterprises do the same without manual fine-tuning effort.

Key Takeaways

  • Claim: GitHub processed 17 million agent-created pull requests in March 2024, and commits are on track for 14 billion this year (vs. 1 billion all of last year) | Evidence: Daigle: 'In March, there were 17 million pull requests that were created by agents… We're on track to be 14 billion if the growth is linear this year, which it will not be.' This excludes human PRs—agents alone drove that figure. | Caveat: Daigle notes growth won't be linear, implying acceleration. He doesn't break down how many are merged vs. abandoned, or quality distribution. | Implication: Ken: This is the clearest quantitative signal of agent coding adoption at scale. GitHub is infrastructure for the agent economy, and the exponential curve means developer tooling, CI/CD, and code review must be rearchitected for 10x-100x volume. Opportunity for tools that triage/filter/rank agent output. | Timestamp: timestamp unavailable
  • Claim: The $200-to-$2000 subscription cost explosion will be solved by automatic model routing and frontier-tuning, not by users manually choosing models per task | Evidence: Daigle: 'My train of thought is slipping in and out of a hard problem to a simple problem… that's a find and replace. But am I going to go off 4.8 or 5.5 and down to Haiku? Probably not. But the tools could.' He cites GitHub's model router and Microsoft Foundry's API-level router as solutions. | Caveat: No timeline given for when model routing becomes seamless. Daigle admits he personally doesn't bother switching models mid-task, so the UX is not yet solved. | Implication: Ken: For agent/AI SaaS operators, the winning move is invisible model routing that detects task intent (hard reasoning vs. find-and-replace) and drops to cheaper models automatically. Users won't tolerate manual switching. This is a product requirement, not a nice-to-have. | Timestamp: timestamp unavailable
  • Claim: Personalization/context/memory—not raw model capability—is the long-term moat, because models will commoditize and local SLMs will soon handle much of the work | Evidence: Daigle: 'The thing that seems to be true from the beginning to OpenClaw to now is this idea of personalization or context or fine-tuning with context or memory… there's experiments, but not a long-term vision for this across the industry.' He contrasts short-term focus (multi-agent sessions) with long-term bet (agents that intuit your style without manual instruction). | Caveat: He says 'we're not very far off' from local SLMs doing serious work, but gives no specific timeline or benchmarks. | Implication: Ken: If models commoditize (Daigle's assumption), differentiation comes from agents that know you—your codebase, communication style, preferences—without manual setup. Invest in memory/context infrastructure, not chasing the latest model. Local SLM inference could disrupt cloud token economics sooner than expected. | Timestamp: timestamp unavailable
  • Claim: Open source maintainers get agentic code review (finds vulnerabilities, implements fixes via comments) and agentic merge (handles CI, policies, manual steps), plus modular controls for community-specific workflows | Evidence: Daigle: 'Copilot code review… finds a lot more novel vulnerabilities and you can just comment and the agent will take that on and go implement the change… I can set exactly what I want to allow GitHub Copilot to do and say, OK, now go merge this PR and wait for CI and wait for policies.' He cites Mitchell Hashimoto's vouch system as one approach GitHub won't force on everyone. | Caveat: GitHub won't impose a standard until one emerges organically. Maintainers must configure controls themselves, which could be high-friction for smaller projects. | Implication: Ken: Open source tooling is shifting from 'accept/reject PR' to 'delegate review and merge to agents, but set guardrails.' For content/workflow agents, this means users will expect similar delegation + control patterns. Don't automate without configurable guardrails. | Timestamp: timestamp unavailable
  • Claim: Hill-climbing (weekly evals, acceptance rates, thumbs-up/down data, user sentiment) is GitHub's core improvement loop, and hard metrics alone can mislead if user sentiment crashes | Evidence: Daigle: 'Every week we're talking about the hill climbing results… sometimes the hard measures and evals and rubrics will show we've made an improvement, but user sentiment will crash, even with the same latency and performance. It's overfitting.' He says Satya, Mustafa, and Jacob (Copilot lead) emphasize this constantly. | Caveat: No specific examples given of when evals improved but sentiment crashed, so we don't know what metrics were overfitted. | Implication: Ken: For agent product development, evals are necessary but insufficient. Track both task success (hard) and user sentiment (soft) in tight weekly loops. Overfitting to benchmarks is a real risk—users notice when agents feel 'off' even if scores improve. | Timestamp: timestamp unavailable
  • Claim: Daigle runs a personal self-improvement agent loop: Claude ('Baxter') reads his last 7 days of emails/Slack/scripts and gives critical feedback on clarity, metaphors, and communication patterns | Evidence: Daigle: 'Every day I get a comms report… Kyle, you keep saying this, this isn't super clear based on how you speak. It'll give me examples of metaphors that are clearer… humans are way more willing to take critical feedback from robots than other humans.' He feeds this interview transcript into it afterward. | Caveat: He uses a separate Claude instance (can't access work stuff) for this, so there's manual data loading. No details on prompt structure or how feedback is applied. | Implication: Ken: This is a template for executive/operator agent use cases: continuous feedback loops on communication style, not just task completion. Humans accept robot critique more readily than peer feedback (Daigle's 'Hubot/ChatOps' observation). Build agents that review user output and suggest improvements, not just generate output. | Timestamp: timestamp unavailable

Detailed Brief

Agent Economy Metrics and Infrastructure Shift

  • Claims: 17 million agent PRs in March 2024, 14 billion commits projected (vs. 1 billion last year); Growth is not linear—will accelerate beyond 14B; Code creation is not 'slop'; developers use agents to multiply context and skills; All code ends up on GitHub regardless of which tool created it, so GitHub must support 'everyone's agent moment'
  • Evidence: Daigle shared March 2024 PR numbers publicly on Twitter; Comparison: 1 billion commits in October 2023 (GitHub Universe announcement); Daigle: 'We're no longer in the super early adoption… we're climbing that hill to see what we can build when it's Kyle and 1-to-N agents using my skills, my resources, my context'; Peter Steinberger example: 150 agents running simultaneously
  • Caveats: No breakdown of merge rate, quality distribution, or % of agent PRs that are accepted; Daigle doesn't specify what 'not linear' means—could be exponential or S-curve; Metrics conflate all agent tools (Cursor, Windsurf, etc.), not just GitHub Copilot
  • Implications: Ken: Developer tooling must scale for 10x-100x code volume. CI/CD, code review, testing infrastructure are bottlenecks.; Opportunity: Tools that triage/rank/filter agent output, or provide 'acceptance likelihood' scores for maintainers.; GitHub's neutrality (supporting all agent tools) positions it as infrastructure layer, but also means it can't force standards—creates fragmentation risk.

Pricing Model Evolution: From Freemium to Usage-Based Hybrid

  • Claims: Freemium worked when humans slept; agents work 24/7, breaking the model; GitHub hasn't locked a new pricing structure yet; Individual devs need affordable baseline (like free private repos became standard); Enterprises running 150 agents need usage-based tiers; $200-to-$2000 cost explosion will be solved by automatic model routing and frontier-tuning
  • Evidence: Daigle: 'I need to make sure you, the dev, have what you need to be successful. And then work with enterprises to make sure they have what they need to do at scale'; Historical parallel: GitHub moved from 'free public repos only' to free private repos when it became clear individuals needed privacy; Model router examples: GitHub Copilot's model router, Microsoft Foundry's API-level router; Daigle's personal token waste: 'I'll get an agent to do enormous work, then there's a smallish thing—change naming—that's find-and-replace. But am I going to drop to Haiku? Probably not. But the tools could.'
  • Caveats: No timeline for when model routing becomes seamless; No concrete pricing tiers announced; Daigle admits he doesn't manually switch models, so UX is unsolved; Frontier-tuning is enterprise-focused (M365 data); unclear how individuals get similar personalization
  • Implications: Ken: For SaaS/AI product pricing, the shift is from seat-based to hybrid (baseline + usage). Expect user backlash if pricing isn't transparent.; Automatic model routing is a product requirement, not optional. Users won't tolerate manual model selection mid-task.; Enterprises will demand cost controls (budget caps, model restrictions) as part of procurement.; Opportunity: Third-party cost optimization tools (model routers, token monitors) for companies not on GitHub/Microsoft stack.

Open Source Maintainer Tools and Community-Driven Standards

  • Claims: Maintainers are drowning in agent PRs; GitHub's answer: agentic code review, agentic merge, and modular controls; Agentic code review finds vulnerabilities and implements fixes via comments; Agentic merge handles CI, policies, and manual steps after review; GitHub won't impose standards—will wait for community consensus (e.g., Mitchell Hashimoto's vouch system)
  • Evidence: Daigle: 'Copilot code review… you can just comment and the agent will take that on and go implement the change'; Daigle: 'I can set exactly what I want to allow GitHub Copilot to do and say, OK, now go merge this PR and wait for CI and wait for policies'; Mitchell Hashimoto's vouch system cited as one approach, but 'there's just as many communities that don't want to use that system'; Daigle: 'We're focusing on building blocks of controls for maintainers. As we all learn together and maintainers send feedback, we'll cement a system if one emerges'
  • Caveats: No details on how agentic review/merge handle adversarial PRs or security risks; Modular controls require maintainers to configure—high friction for small projects; Community-driven approach means slow standardization, possible fragmentation
  • Implications: Ken: Open source governance is shifting from human gatekeeping to agent-assisted gatekeeping with configurable policies.; For content/workflow agents, users will expect similar patterns: delegation + guardrails, not full automation.; Opportunity: Tools that auto-generate control policies based on project type (e.g., security-critical vs. doc repos).; Risk: If GitHub waits too long for community consensus, a competitor could impose a standard and gain network effects.

Personalization and Hill-Climbing as Long-Term Moats

  • Claims: Personalization/context/memory is the durable advantage, not raw model capability; Models will commoditize; local SLMs will handle much work soon; Hill-climbing (weekly evals, acceptance rates, thumbs-up/down data, user sentiment) is the improvement loop; Hard metrics can mislead if user sentiment crashes (overfitting risk); Frontier-tuning lets enterprises personalize models without manual fine-tuning effort
  • Evidence: Daigle: 'The thing that seems to be true from the beginning to OpenClaw to now is this idea of personalization or context or fine-tuning with context or memory… there's experiments, but not a long-term vision'; Daigle: 'We're not very far off from having serious ability to use something above a small language model on a local device'; Daigle: 'Every week we're talking about hill climbing results… sometimes hard measures show improvement but user sentiment crashes, even with same latency. It's overfitting.'; Frontier-tuning example: 'Using MAI thinking one as base model, it shows real results without having to do all that extra work' (M365 data as corpus); Seven new Microsoft models launched via this process
  • Caveats: No timeline for local SLM capability or specific benchmarks; No details on what 'novel vulnerabilities' agentic review finds vs. traditional tools; Frontier-tuning is enterprise-only right now (requires M365 data corpus)
  • Implications: Ken: If models commoditize, the moat is memory/context infrastructure. Invest in systems that learn user preferences implicitly, not via manual setup.; Local SLM inference could disrupt cloud token economics faster than expected—watch for model size/capability crossover.; For product development, track both task success (hard) and user sentiment (soft) in tight weekly loops. Overfitting to benchmarks is a real trap.; Hill-climbing is a discipline, not a one-time optimization. Enterprises will expect similar continuous improvement from SaaS vendors.

Developer Choice and Anti-Walled-Garden Positioning

  • Claims: Developer choice is GitHub's core differentiator vs. competitors; GitHub partners with Anthropic, OpenAI, Google, and any model/agent provider; Microsoft Build 2024 was first to feature external contributors in keynote/primary sessions; Industry is in 'unintentional walled garden' moment; GitHub refuses to build a mousetrap; Daigle personally uses Mac, PC, Linux to avoid myopia; swaps between boxes every Saturday
  • Evidence: Daigle: 'We'll partner with everyone to make that as simple as possible… the ability to do that across the entirety of building software is a real superpower of ours'; Daigle: 'We'll both let you bring that to us or we'll offer it through GitHub and GitHub Copilot. That choice is core. If we back down, developers will still choose—they'll just be stuck in another mousetrap.'; Build 2024 featured Swix, Peter Steinberger, and other external speakers in keynote; Daigle: 'I have my Mac, my PC, and my Linux box… I code most Saturdays, swapping between the boxes because I want to understand what's that experience'; GitHub Copilot app tested on Windows because 'developers on Windows also deserve great apps'
  • Caveats: No details on how GitHub enforces interop standards or handles conflicting agent APIs; Build external speakers were still curated by Microsoft—not fully open call for papers; Personal multi-platform testing doesn't scale to all GitHub employees
  • Implications: Ken: In competitive markets, positioning as 'open platform' vs. 'best integrated experience' is a strategic choice. GitHub bets on openness; competitors bet on tight integration.; For GTM, GitHub's anti-mousetrap message resonates with devs who fear lock-in (especially post-AI hype cycle).; Risk: Openness can mean slower feature velocity if you must support every external tool. GitHub's scale lets them absorb that cost; smaller players can't.; Operator note: Daigle's personal multi-platform discipline is a leadership culture signal—dogfooding across ecosystems, not just your own stack.

Self-Improvement Agent Loops and Human-Robot Feedback Dynamics

  • Claims: Daigle runs a personal agent (Claude 'Baxter') that reviews his last 7 days of emails, Slack, scripts and gives critical feedback; Humans accept robot critique more readily than peer critique; Agent loop focuses on communication clarity, metaphor usage, and patterns; This practice reflects GitHub's long-term bet on personalization; Historical precedent: GitHub's Hubot/ChatOps showed humans prefer robot feedback
  • Evidence: Daigle: 'Every day I get a comms report… Kyle, you keep saying this, this isn't super clear. It'll give me examples of metaphors that are clearer.'; Daigle: 'Humans are way more willing to take critical feedback from robots than other humans… when Baxter tells me how terrible I did, I feel way better going, tell me why.'; Daigle feeds interview transcripts into Claude after the fact; Separate Claude instance (can't access work stuff) for personal review; Hubot/ChatOps at GitHub (historical): 'We used to say humans are way more willing to take critical feedback from robots than other humans'
  • Caveats: No details on prompt structure or how feedback is systematically applied; Manual data loading required (separate Claude instance); Not clear if this is a one-off experiment or a repeatable workflow for others
  • Implications: Ken: This is a high-value use case for executive/operator agents: continuous feedback on communication/writing, not just task generation.; Human-robot feedback dynamic is exploitable: people are less defensive with AI critique. Build agents that review user output (emails, docs, code) and suggest improvements.; For content/workflow agents, positioning as 'coach' rather than 'automator' may increase adoption.; Operator note: Daigle's loop is a template for self-improvement workflows—could be productized for execs, writers, developers.

Notable Concepts & Terms

  • Hill-climbing: GitHub's weekly improvement loop: run evals, measure acceptance rates/thumbs-up-down/user sentiment, adjust models/prompts, repeat. Daigle emphasizes hard metrics can mislead if sentiment crashes—overfitting risk. Term used repeatedly by Satya, Mustafa, Jacob (Copilot lead).
  • Frontier-tuning: Microsoft's approach to personalizing models (e.g., MAI thinking one) using enterprise data (M365 docs, chats) without manual fine-tuning effort. Shows 'real results' per Daigle. Distinct from traditional fine-tuning; leverages context/memory automatically.
  • Agentic code review: GitHub Copilot feature that finds 'novel vulnerabilities' and implements fixes via comment. Maintainer comments on PR, agent executes changes. Shifts review from 'approve/reject' to 'delegate and refine.'
  • Agentic merge: GitHub Copilot app feature where maintainer sets policies (CI, checks, etc.), then agent handles merge steps automatically. Daigle: 'Set exactly what I want to allow… now go merge this PR and wait for CI and policies.'
  • Model router: GitHub Copilot and Microsoft Foundry feature that auto-selects models based on task intent (hard reasoning vs. simple find-and-replace). Daigle's answer to $200-to-$2000 subscription cost explosion—invisible to users.
  • Vouch system: Mitchell Hashimoto's open source PR control system (cited by Daigle). Maintainer requires contributors to be 'vouched' before accepting PRs. GitHub won't force this as standard, waiting for community consensus.
  • Baxter (Daigle's Claude instance): Daigle's personal self-improvement agent: reads last 7 days of emails/Slack/scripts, gives critical feedback on clarity and metaphors. Named affectionately; separate from work Claude instance. Example of human-robot feedback dynamic.
  • ChatOps / Hubot: GitHub's historical chat-based automation bot. Daigle cites lesson: 'Humans are way more willing to take critical feedback from robots than other humans.' Informs current agent design philosophy.
  • Mousetrap / walled garden: Daigle's term for tools that lock users in via tight integration, then make it painful to switch. He contrasts GitHub's 'developer choice' positioning (partner with all model/agent providers) vs. competitors building mousetraps.

Operator Notes / Why Ken Should Care

  • Ken: GitHub's 17M agent PRs/month and 14B commit projection are the clearest quantitative signals of agent coding adoption at scale. Use these as benchmarks for estimating agent economy growth—if code generation is 10x-ing, downstream tooling (CI/CD, testing, review) must scale similarly.
  • Ken: The $200-to-$2000 subscription cost explosion is a universal SaaS problem for AI products. Daigle's answer—automatic model routing and frontier-tuning—should inform your pricing/product strategy. Don't expect users to manually switch models; build intent detection and routing into the product.
  • Ken: Personalization/context/memory is the long-term moat per Daigle, not raw model capability. If models commoditize (his assumption), differentiation comes from agents that know the user. Prioritize memory/context infrastructure over chasing latest model releases.
  • Ken: Hill-climbing (weekly evals + sentiment checks) is a discipline, not a one-time optimization. Track both hard metrics (task success) and soft metrics (user sentiment) in tight loops. Overfitting to benchmarks is a real trap—users notice when agents feel 'off' even if scores improve.
  • Ken: Daigle's personal self-improvement agent loop (Claude 'Baxter' reviewing his emails/Slack/scripts) is a high-value use case template. Humans accept robot critique more readily than peer feedback. Build agents that review user output and suggest improvements, not just generate output.
  • Ken: GitHub's anti-walled-garden positioning (partner with all model/agent providers) is a strategic GTM choice. In competitive markets, 'open platform' vs. 'best integrated experience' is a tradeoff. GitHub bets on openness; competitors bet on tight integration. Choose based on your scale and customer lock-in risk.
  • Ken: Open source governance is shifting from human gatekeeping to agent-assisted gatekeeping with configurable policies. For content/workflow agents, users will expect similar patterns: delegation + guardrails, not full automation. Don't automate without configurable controls.
  • Ken: Daigle's multi-platform discipline (Mac, PC, Linux; swaps every Saturday) is a leadership culture signal—dogfooding across ecosystems, not just your own stack. For product development, this prevents myopia and ensures you understand cross-platform pain points.

Watch Map

  • timestamp unavailable: Opening: Mike Taylor introduces Kyle Daigle (GitHub COO, Microsoft CMO for Developers); context-setting on changing developer demographics and agent economy
  • timestamp unavailable: Agent economy metrics: 17M agent PRs in March 2024, 14B commits projected (vs. 1B last year); 'not slop,' developers using 1-to-N agents
  • timestamp unavailable: Open source maintainer tools: agentic code review, agentic merge, modular controls; GitHub won't impose standards, waiting for community consensus (Mitchell Hashimoto vouch system example)
  • timestamp unavailable: Pricing model shift: freemium breaking down (agents work 24/7); no new model locked yet; $200-to-$2000 cost explosion solved by model routing and frontier-tuning
  • timestamp unavailable: COO + CMO dual role: GitHub's developer-first culture, Microsoft Build external speakers (first time), anti-walled-garden positioning, partnerships with Anthropic/OpenAI/Google
  • timestamp unavailable: Long-term moat: personalization/context/memory, not raw model capability; hill-climbing (weekly evals + sentiment checks); frontier-tuning for enterprises; local SLM future
  • timestamp unavailable: Dogfooding and experimentation: Daigle's multi-platform discipline (Mac, PC, Linux); GitHub Copilot app tested on Windows; culture of using competitors' tools to avoid myopia
  • timestamp unavailable: Self-improvement agent loop: Daigle's Claude 'Baxter' reviews last 7 days of emails/Slack/scripts, gives critical feedback; humans accept robot critique more readily; historical Hubot/ChatOps lesson
  • timestamp unavailable: Closing: Mike reveals he made an AI clone of Kyle to practice the interview; Kyle shares he does similar with Claude for self-feedback

Source/Metadata

  • Title: How GitHub Deals With 17 Million Pull Requests a Month
  • Transcript words: 5060
  • Duration seconds: 1687
  • Timestamp note: Timestamps were not present in the transcript; chapter notes are descriptive only

Transcript

4832 words en Processed in 162.1s

Hi, I'm Mike Taylor. I'm the head of tech consulting at Every, and I sat down with Kyle Daigle, the COO of GitHub, and talked to him about what is happening on the front lines of coding agents. We have 17 million pull requests coming in every month to GitHub now. It's growing exponentially, and that puts him at the forefront of what's happening in this new economy. We talked about how that affects our users, as well as how this affects open source maintainers, and we covered a topic which is dear to everyone's hearts. How do I stop my $200 a month coding agent subscription from ballooning into a $2,000 a month usage limit? In this interview, we did something a little bit different, which is I told Kyle I had made an AI clone of him to practice the interview, and he revealed something surprising in return. Here's the conversation. Every is the only subscription you need to stay at the edge of AI. If you care about being on top of the latest models and using the latest tools, you have to subscribe to Every to separate out the signal from the noise. Go to every.to.subscribe today. Hey, Kyle. Thanks for spending some time with me at the conference. Yeah, of course. Yeah, it was good to meet you as well the day before, and I feel like we already kind of covered a few of these questions, but I think it would be good for the wider audience so they can understand what's going on here. Absolutely. Yeah. So the first thing I think is we were talking about what's really interesting is that the demographics of the customer are changing, right? A lot of people who previously maybe never used GitHub or never used developer products before are now using them. So how has that changed the way that you decide the product roadmap? Yeah, I mean, I think for GitHub in particular, we've always really had this expansive view of what a developer is. I started as a developer before I would have ever called myself a dev, where I was just writing code, but it was just for me. And I went personally on a completely different career path. I didn't go to school for computer science. I was going to art school. I wrote code to pay for art school, which was a very silly decision as an adult now, I guess. But then, that journey of just creating tools with the team and delivering them to people who can have that same experience of wanting to build an app that's for me or for my family, maybe as a startup, maybe as a business. [SPEAKER_00] We very much have serious developer tools, and all the largest businesses are using GitHub. But when I look at something like the GitHub Copilot app, I see just as many developers that are using AI every day, running multiple projects, with all kinds of agent sessions at the same time. And I see our legal team at GitHub using the GitHub Copilot app or the finance team, or I was meeting with a customer today and they were saying the same thing. A lot of the folks that the industry would call knowledge workers or non-trade developers are using these tools to build little apps or assets for them. And so while our focus is very much on developers, I think we want to make it easier for people to choose to write some code and make sure there's always an on-ramp into writing some software with things like the GitHub Copilot app. Yeah. And then how do you deal with the burden of all of that extra activity? There's a flood of PRs now, and open source maintainers I talked to are drowning. How do you help them? Yeah. I mean, I think for all developers we're building tools like the Copilot code review. It's now agentic. So it finds a lot more novel vulnerabilities and you can just comment and the agent will take that on and go implement the change if you want. So I think that code review step is in some ways overlooked as a really great way to get PRs to a place that are much more easily reviewed. I think that the agentic merge in the app is another place where we see a lot of times internally and in the community, you may comment on something that might have a code review and you might go through and get it almost all the way there. But then there's all those manual steps just to finish processing the PR instead. I can go in and set exactly what I want to allow GitHub Copilot to do and say, OK, now go merge this PR and wait for CI and wait for policies and all of that. [SPEAKER_03] I think that's a big part on the open source side. It's a unique set of needs because you don't control who's sending everything in or you haven't really historically. [SPEAKER_03] And that's been really where we've been focusing is giving maintainers more tools to decide, well, do you want to accept all of these PRs? Who do you want to accept them from? How much work do you need to do to prove that you're going to contribute something that is going to be meaningful to this project? And that's something that we want to provide tools to open source maintainers, but really leave them in control. Every community is choosing a slightly different way to approach the problem. [SPEAKER_00] And for GitHub, we've always wanted to leave that in their hands, give them tools and enable them. [SPEAKER_00] But if a standard comes out of that or most are using a certain practice, we'll lock that in. But we don't really ever want to be the first to create a standard or an approach. I think Mitchell Hashimoto shared the vouch system that they use. And I was getting questions like, well, why aren't you roll this out to everybody? But there's just as many communities that don't want to use that system because they have their own ideas of how it should work. And so for now, we're focusing on the building blocks of controls for maintainers. And then as we all are learning together and as maintainers send feedback in, we'll cement an entire system if one emerges. Yeah. And I feel like you have a front row seat to this new agent economy where I think you said public on Twitter that you've had more pull requests submitted per month than you did all last year. How are those stats exploding? Yeah. I mean, we're seeing way more activity on GitHub. We've always been talking about our users for many, many years and that growth. But this year we're seeing the growth of developers having agents building with them. And so last year in October at GitHub Universe, we shared there's a billion commits on GitHub for the full year. We're on track to be 14 billion if the growth is linear this year, which it will not be. In March, there were 17 million pull requests that were created by agents. Yeah. That's just the agent pull requests. Okay, yeah. And so there's so much more code being created. [SPEAKER_03] And I think at times everyone goes, oh, this is all just slop. But this year we're seeing the growth of developers having agents building with them. And so last year in October at GitHub Universe, we shared there's a billion commits on GitHub for the full year. [SPEAKER_00] We're on track to be 14 billion if the growth is linear this year, which it will not be. [SPEAKER_00] In March, there were 17 million pull requests that were created by agents. [SPEAKER_00] Yeah. [SPEAKER_00] That's just the agent pull requests. Okay, yeah. And so there's so much more code being created. [SPEAKER_03] And I think at times everyone goes like, oh, this is all just slop. [SPEAKER_03] This is all just code that's getting pushed up. [SPEAKER_03] And no one cares. [SPEAKER_03] It's not really true. We're all just actually getting to the point where we're no longer in the super early adoption. We're definitely not at the peak, but we're climbing that hill to see what can we build when it's not just Kyle building, but it's Kyle and one, two to N agents that are using my skills, using my resources, using my context and so on and so forth. And so we're investing heavily in preparing for the next wave of growth because it doesn't seem to be growing and plateauing. It's just going to continue to grow because no matter where you're building or what tools you're using to build, all of that code ends up on GitHub or that's where you're sharing it with the world or that's where you're collaborating in a PR. And so we need to be able to support everyone's agent moment and not just GitHub Copilot. Yeah, yeah. Yeah. And how does the business model change? Because I think freemium makes sense in a human centered world where we go to bed, but the agents are still working while we're asleep now. So does that change to usage based? Does that, you see that where things are going? Yeah. I mean, I don't think we know yet ultimately. I think we very much right now Kyle's going to have a license or Kyle's using GitHub.com for free. And we've always had API rate limits and things like that. And that's usually where folks are seeing the agent back pressure, I think. I think the goal is that if you want to be able to do way more, if you want to be able to have Peter Steinberger says, 150 agents are doing everything all at once, that's great. We want to be able to enable that to you. But at the same time, I want you to have a great core GitHub experience and at the very least there's some amount of agent usage as part of that that is necessary. I need similar to how we way, way, way back, right? You'd have free public repos, but you didn't have free private repos. And then we said, okay, well, actually, it's fair for an individual to have some code that they don't want to put out into the world. And we'll give you free private repos to allow you to do that. So GitHub is always evolving as the industry and community does. But we're always focused on, I need to make sure you, the dev, have what you need to be successful. And then work with enterprises to make sure they have what they need to do at scale, which is usually a little bit different than what an individual dev is doing. Yeah, yeah. And I guess the business model of pricing, that all leads back into the wider Microsoft orbit. [SPEAKER_00] And because you have a dual role now, right? Partial responsibility for the wider marketing org. So do you want to talk me through how that's changed and how you prioritize between those two? Yeah, I mean, I've been at GitHub for a very long time, 13 years. And as a developer myself and leading engineering teams for a lot of that time. And I think what's always been unique about GitHub is we really focus on the dev. We're building tools for the developers. And the fact that people like enterprises are buying them is awesome. And that's definitely great. But we're not building for the buyers. We're building for the developers 100%. And so in this, and that's been my focus as the COO of GitHub, which I continue to do. And then now as the chief marketing officer of developer for Microsoft, my goal is to look across all of Microsoft's tooling, their developer tools, their technology that they're bringing to developers. And making sure that we're bringing holistic solutions that you can use that are authentic to developer experiences. And at events like this where we've taken a very different approach to build this year. We're in San Francisco, first off. The vibe is a bit different than the conference hall set up. Really focused on can I go to a session? [SPEAKER_03] Can I use the thing? I don't want to be pitched on a thing. I have to be able to use it. Expo hall and so on and so forth. It's really bringing that expertise and love and focus on the developer that GitHub's always had to have an even broader impact throughout all of Microsoft. [SPEAKER_00] Yeah. And do I hear you say that this is the first Build that you've had external contributors? It's the first Build that I think by intention we focused on having speakers from the community in these primary sessions. That includes in the keynote we had a bunch of folks like Peter. There's sessions from Swix and others as well. I think that it's important. Software development is a team sport. It seems silly to think that there's any one company, one group inclusive of GitHub and Microsoft and everyone that can just answer every single question. That's not how software gets made. We're all at least using open source and we're building on the backs of these giant open source projects. Let's invite people in that can help tell their part of the story together because I deeply believe that that's what developers want. I know that's what I want. I know that's what my friends that are developers want. And when we look at the events and we hear the feedback, they're excited to see people from Microsoft, from GitHub. And then I get to see this outside perspective at this event. It's really meaningful. Yeah, that makes sense. And it's a very competitive market, right? Sure. Let's invite people in that can help tell their part of the story together because I deeply believe that that's what developers want. I know that's what I want. [SPEAKER_00] I know that's what my friends that are developers want. [SPEAKER_00] And when we look at the events and we hear the feedback, they're excited to see people from Microsoft, from GitHub. And then, oh, I get to see this outside perspective at this event. It's really meaningful. Yeah, that makes sense. And it's a very competitive market, right? Sure. Yeah. [SPEAKER_03] The most competitive market probably. [SPEAKER_03] Maybe the last competitive market. I'm not sure. But how do you differentiate in all of that given the pace of change is so quick? Yeah. I think we continue to focus on our roots, which is we care a lot about developer choice. It's always been true. We care about building for builders and enabling builders. And so I think we're in a moment that's really interesting because we've gone from an era of having a ton of APIs, all this access, to a little bit of an unintentional walled garden setup, where you get a kind of affinity. [SPEAKER_00] Sometimes I'll say it's like a little bit of a mousetrap. [SPEAKER_00] And then you realize, oh, this thing's really interesting over here. [SPEAKER_00] And then I have to learn a new thing or a new tool or a new account. [SPEAKER_00] And I think for us, we always want to enable developers that are building with GitHub to go use these other tools. [SPEAKER_00] And we'll partner with everyone to make that as simple as possible. [SPEAKER_00] And while I think there are other folks that are doing similar things, I think the ability to do that across the entirety of building software and not just the cogen side or not just the collaboration review side, but across everything is a real superpower of ours. [SPEAKER_00] And so I think you'll see us invest in our own tech. We talked about the new Microsoft AI models, which we'll continue to bring to developers. [SPEAKER_03] We're also continuing to partner with Anthropic and OpenAI and Google and anyone who's bringing a model to market or a coding agent to market. [SPEAKER_03] We'll partner with you and we'll both let you bring that to us or we'll offer it through GitHub and GitHub Copilot. [SPEAKER_03] That choice is core. [SPEAKER_03] And that's something that I won't back down on. [SPEAKER_03] Because if we do, developers will still choose. [SPEAKER_03] They'll just be stuck in another mousetrap. [SPEAKER_03] And we don't want the world of software to be like that. Yeah, yeah, yeah. And how do you make decisions internally when there was a news cycle recently about how code licenses are being canceled? How do you make the tradeoff between dogfooding your own products, using the new models you made or using the GitHub Copilot desktop app, versus letting developers experiment with other tools? Yeah, I mean, we all use a variety of tools because otherwise you lose track or you're too invested in your own work. So for me, I've been a daily driver of a MacBook for many years. I use Windows PCs on the weekends when I play video games. And I got this role and I have my Mac, my PC, and my Linux box. So I can make sure that every weekend I code most Saturdays, I do my kids' sports activities in the morning. And then in the afternoon I'm coding and swapping between the boxes because I want to understand what's that experience? The GitHub Copilot app I only use on Windows because I want to make sure that developers who are on Windows also deserve great apps. It's not just the audience that's on a Mac. And that's true across our teams, especially when we're looking at coding agents, harnesses, desktop apps, memory management, everything. [SPEAKER_00] We have this really great culture of experimentation. [SPEAKER_00] Everyone is building and using these tools. [SPEAKER_00] Obviously we're putting most of our energy into our own tools. It's such a blind spot that happened to GitHub in the past where when you're doing something and you're doing it well, you really laser focus. [SPEAKER_03] And that's what every piece of startup energy says, right? [SPEAKER_03] It's look down and just keep moving and keep moving fast. [SPEAKER_03] And I think that's myopic. I think while I can't spend every day using every tool, when something comes out, I want to know why this is really great. Why are people having a great experience with this? Not only so I can understand, but so I can figure out for our goals, for our goal of developer choice, I don't need this. I want to focus over here, but I want to know why a dev would pick these tools. And the same thing goes for our teams. Yeah. And how do you filter? Because obviously a lot of these ideas are relatively short-lived. Yeah, yeah. Enterprise product development cycles are longer lived. Yes. How do you decide? Yeah, I think right now we're in a moment where we're really looking at the short term in capturing the ability to have a multitude of agent sessions. This idea of because that seems quite clear, everyone's doing it. How can we cement it? But it seems clear on the longer term path models are going to continue to get better. The prices of tokens, token economics, is going to be a bigger factor in what models everyone is using. And I do strongly believe that we're not very far off from having serious ability to use something above a small language model on a local device to do some of our work. And so if I assume that I have all this optionality when it comes to tokens, the thing that seems to be true from the beginning to open claw to now is this idea of personalization or context or fine tuning with context or memory. All of these ideas seem to be a truth that's been there since ChatGPT came out or GitHub Copilot came out. And there are experiments, but not a long term vision for this across the industry. So I think it's a good example of where I need to get you to use agents incredibly well. [SPEAKER_00] And so if I assume that I have all this optionality when it comes to tokens, effectively, the thing that I think seems to be true from the beginning to open claw to now is this idea of personalization or mine or context or fine tuning with context or memory. [SPEAKER_00] All of these ideas seem to be a truth that's been there since ChatGPT came out or GitHub Copilot came out. [SPEAKER_00] And there's experiments, but not a long term vision, I think, for this across the industry. [SPEAKER_00] So I think it's a good example of where I need to get you to use agents incredibly well. [SPEAKER_00] A lot of them, because if you're using agents, you're not just going to be staring at a single agent working. [SPEAKER_00] But that's not going to give you a long term great experience. [SPEAKER_00] Using an agent that you feel like is completing a thought for you will give you that great experience, especially if you did not have to personally codify that thought. [SPEAKER_03] Yeah. [SPEAKER_03] To your agent. [SPEAKER_03] Always remember that I insert thing. [SPEAKER_03] That's a lot of work. [SPEAKER_03] A hundred percent. It should be able to intuit that or potentially, again, post trainer fine tune or frontier tune a model that deeply understands me and how I'm using the work. [SPEAKER_03] That is how we're looking at it. [SPEAKER_03] It's like sometimes it's short term and sometimes we got to take a bunch of bites of the apple or a bunch of attempts at the long term to get to something really tangible to help us move forward. Cool. And I heard the term hill climbing a hundred times yesterday. Yes. And I'm a big proponent of that because I experimented with DSPY, auto research, a few others. Can you talk a little bit about how that's become a big focus? Yeah. I mean, I think Satya and Mustafa talk about it a fair bit and Jacob leading the copilot group. The biggest thing that we've learned is we need to use the tools as a core way to improve the underlying use of the models, our own models, et cetera. And just the evals that are necessary to ensure that we're actually improving from things like using the thumbs up, thumbs down data that comes in to using whether you're accepting it and how much you're accepting. All of that data is enormous to create that magical type of experience that's not just for you, but for everyone. And so every week we're talking about the hill climbing results. We're looking at the data, we're looking at the improvement, we're looking at both the hard measures and the soft measures, because sometimes the hard measures and evals and rubrics will show that we've made an improvement. But user sentiment will crash. [SPEAKER_03] Yeah. Even with the same latency and performance. [SPEAKER_03] It's overfitting. A hundred percent. [SPEAKER_00] And so being able to really do that loop incredibly quickly. [SPEAKER_00] And then I think the main goal is giving everyone one of these hill climbing machines and not have you have to do it the hard way that we've all been doing it. [SPEAKER_00] But particularly if you're in an enterprise and you are using M365, we know so much about that data, or we could know so much about that data because of all the assets, all the documents, the chats. [SPEAKER_00] And so being able to turn on something like frontier tuning and using MAI thinking one as the base model, it shows real results without having to do all that extra work. [SPEAKER_00] And it's been interesting because when I first heard about this, I'll be honest, I was thinking this is a magic parlor trick, that is not going to really work. [SPEAKER_03] Yeah. [SPEAKER_03] We all have all this data and what are we going to do with it? [SPEAKER_03] We have to do all this effort to make it work. [SPEAKER_03] But I think so much has come down the pipe to allow us to just use the data and improve, look at the workflow and improve and just keep doing the hill climbing. [SPEAKER_03] That's why I think we say it so much is that it's not these moonshots. It is just climb, climb, improve, new eval, improve, new data, improve, and just keep going to get to the point where we're able to launch these models, seven models for ourselves. And then allow customers to use the same or similar tooling to do it. Is that the answer to stopping the $200 subscription becoming a $2,000 subscription? [SPEAKER_00] I mean, I think the $200 subscription to $2,000 is really going to be not only making these models or frontier tuning these models so they know you better. But I also think it's really going to be about how can we, particularly for developers, help you automatically choose the models and potentially either have a model in that step. Like the model router in GitHub. A hundred percent. Exactly. [SPEAKER_01] Like auto model router with task intent in GitHub. Microsoft Foundry has a model router as well that can do this sort of at an API level. Yeah. The more that we can help you tell us a bit of where your bars are. This is an incredibly hard problem and I'm willing to go all the way to the top or I just want to sit here and let us help choose the models. Because there's a lot of times where a lot of the reasons my tokens are expensive is because we're all going and choosing our model of the day or week or hour, and those models are incredibly expensive. But my train of thought is slipping in and out of a hard problem to a simple problem. Yeah. I personally feel I'll get an agent to do an enormous amount of work. And then there's always that last step that is a smallish thing, oh, I don't actually like change all the naming of this to this. Yeah. [SPEAKER_00] And that's a find and replace. But am I going to actually go and oh, I want to save tokens right now. So I'm going to go off 4.8 or 5.5 and down to Haiku or something, probably not. But the tools could. [SPEAKER_03] Yeah. [SPEAKER_03] And I think that will really help us, particularly in the enterprise, but even for individual developers and folks that are building automations and using their co-pilot SDK to power that. It'll help them too. Yeah. I did something a little bit weird. I hope you don't find it creepy, but I made an AI version of you to practice this interview. No way. Yeah. And it's actually been pretty spot on so far. And hopefully you think the questions have been good. They've been great. But the tools could. Yeah. [SPEAKER_03] And I think that will really help us, particularly in the enterprise, but even for individual developers and folks that are building automations and using their co-pilot SDK to power that. It'll help them too. Yeah. I did something a little bit weird. I hope you don't find it creepy, but I made an AI version of you to practice this interview. No way. Yeah. And it's actually been pretty spot on so far. [SPEAKER_00] And hopefully you think the questions have been good. [SPEAKER_00] They've been great. [SPEAKER_00] Now I want to see what AI Kyle said. [SPEAKER_00] And yeah, it's just in the terminal. [SPEAKER_00] I didn't have, I didn't go the full whack and make a video thing. [SPEAKER_00] Sure, sure, sure. [SPEAKER_00] But I found it immensely useful. [SPEAKER_00] I just wanted to ask, what other weird things are you seeing people do internally or externally? [SPEAKER_00] Oh, man. [SPEAKER_00] So it's so funny that you say that because I do a very similar thing where I have both via the app and then I have a claw that can't talk to work stuff, you know? [SPEAKER_00] Just so I have separation of state where I spend a lot of time having it read everything I write and say. [SPEAKER_00] This interview will get fed into it ultimately. [SPEAKER_00] Yeah. [SPEAKER_00] Yeah. And every day I get a comms report that's not, what Kyle said. But Kyle, you keep saying this. Yeah. [SPEAKER_03] Okay. [SPEAKER_03] This isn't super clear. [SPEAKER_03] Based on how you speak. [SPEAKER_03] Because I find that I write and speak in a very particular way that I want to use a lot of metaphors. And so it'll just give me examples of metaphors that are clearer. Yeah. [SPEAKER_03] I find that the self-improvement loop as a human from these agents to be incredibly powerful. [SPEAKER_03] We used to talk about it way back with Hubot at GitHub, chat ops. [SPEAKER_03] And we used to say, humans are way more willing to take critical feedback from robots than other humans. [SPEAKER_03] Yeah, it's less threatening. [SPEAKER_03] A hundred percent. [SPEAKER_03] And so when my open claw that I affectionately named Baxter tells me how terrible I did in something, I feel way better going, tell me why. [SPEAKER_03] And then ensure that when I'm writing emails, when I'm writing a script or I'm reviewing details, that you're giving me that feedback. [SPEAKER_03] So a lot of my agent loop is really about me and less about the software side. I still have all those tools, too, you know. But it's always looking backwards. It's going, okay, the last seven days I'm going to read all Kyle's emails, Slack messages, you know, and then give me feedback.