[SPEAKER_00] All right. Can you all hear me? Great. There we go. All right. Well, we're only 35 to 40 minutes late, but thank you for sticking around. We're going to talk about why CI/CD is dead and we're going to propose that continuous compute is going to be the next thing. Maybe.
SPEAKER_00
All right. So just a quick introduction. We're going to have two speakers. One's getting mic'd up. My name is Madison and I'm a partner at NEA Investing in Technology. I do focus in infra and dev tools and I formerly used to be a Meta AI researcher. So I used to lead data and AI teams. I got really frustrated by the state of infrastructure and so I jumped into venture to do something about it from the top down. And then I'm also going to introduce on behalf of my partner here, Hugo Santos. So he's the CEO of Namespace, which is building high performance compute infrastructure.
SPEAKER_01
And at this point, what we believe is going to eclipse the new CI/CD wave. He also formerly led microservices at Google. Yeah. Great to be here with you folks. So we're going to talk about why agentic software is breaking traditional CI/CD. Obviously, we're not going to get through this today, but the point is on the left side, what started off in agentic software was really monolithic agents. We were really using the LLM as one engine, but now we're moving into the right side, which is microservices with agents. And that's how we really need to think about software development in an agentic world. So the lifecycle is very fragmented. This is quite a mess, right? We've really brought together all these traditional CI/CD systems: build, test, deploy, DevOps. But we also now have new IDEs. We have autonomous agentic engineering solutions. And then we have our traditional DevOps in the middle, which we believe is really going to innovate in the next year. So let's explain why we think it's dead. So first, how do CI/CD pipelines work today? Well, we all know human developers are currently submitting one, maybe a couple of diffs when they're just writing it themselves. And those PRs then take your colleagues a bunch of time to review. Then you have to go through GitHub Actions and run build, test, and deploy steps. And then finally, you're addressing those failed test cases and maybe you're iterating on the diff. So in that scenario, it was really just one or two a week. So now, how do we think about this at agent scale? You've got agents using the exact same systems, but they have N number of PRs, maybe N number of repos. Still takes a similar amount of time to verify unless you're using review bots, which gets a little crazy. And then we correct those failed cases just like we did in the past scenario. So what ends up happening? With a human, pretty predictable. And you've got local caches, which are often warm. With an agent, this starts to get really complicated. You have thousands of short-lived branches. It's all trying to pull the same code base in a few different directions. You start to get to a point where merging all these different versions together is really impossible. And that is where we start to have a huge problem.
SPEAKER_00
So let's look at, in real time, GitHub Activity has gotten absolutely crazy. The white line here is the actual number of commits in the last couple of months. And then the number of lines added versus deleted. This is just an unbelievable spike.
SPEAKER_00
So how do we start with replacing CI/CD? Well, the starting point should be at the acceleration. So obviously right now, I know a lot of you are struggling with very slow build, test, and deploy times for your CI/CD solutions. This is a very common problem. But where we're headed is being able to first speed that up by inserting over the existing GitHub Actions and other underlying infrastructure for CI/CD. So that cache is really going to become the orchestration layer in this scenario. And this is really critical to do through a hardware and software co-design.
SPEAKER_00
So what does this start to look like? And how does this start to eclipse previous CI/CD? So first, we have our intake, which requires ingress shaping and rate limiting. Then we move to our cache. And this is the next big step. How do we think about orchestrating and making sure we're routing to the right infrastructure? From there, we can even move into agentic identity for software and thinking about retries at scale.
SPEAKER_00
And then if you don't believe me, let's ask the experts. So Mitchell Hashimoto, one of the coolest dev rel at scale, he's also the former founder of HashiCorp, wrote exactly what he would do to fix GitHub today. And a lot of this has to do with even shutting down Copilot, thinking about how do you actually just evolve GitHub to be first in the cloud era, but second, actually really enabling inference at scale. And then we've got a number of other data points on the left hand side that we need to be able to serve AI and agentic users first, or we die. And thinking about friendly code storage solutions that may also help. So there's a lot of frustration around existing CI/CD, but this is really just the starting point. We've only just started to see agentic software takeover. So now I'm going to pass it to Hugo to talk more about what a real solution can look like. Yeah, so I'm fortunate, and me and my team, we spend a lot of time with companies today that are going from how traditional CI/CD look like into how we think it's going to look into the future. And giving a little bit of a hint, it's agents all the way down. So we work with companies like Fall and Zed and Ramp and many others that are really at the forefront of everything around development. And you probably recognize yourselves in between these two bits where up to six months, humans were writing all the code very slowly, and some of them actually fairly quickly, but in hindsight fairly slowly. We package all these changes in PRs. We do validation as part of those PRs. And behind the scenes, the machines are a little bit slow, but all of that is hidden behind the human latency. And many of you might already be seeing a bit of what's happening today where code generation is very cheap, work is much more continuous, and that forces the validation to go into the inner loop. So what you might not realize is up to this point, you as a human, you are the agent. You have a state of mind, here's what I'm trying to accomplish, and then I'm going through all of these phases. Okay, I start the pull request, and the pull request within your team says, well, you didn't quite follow the right format, so go back to the beginning. You're in the loop. Now your changes are in the PR, the tests are running, they fail, you need to go and change something in the code. You're back in the loop. A human reviewer comes back and says, well, it didn't quite use the right API, please go and change it. You're back in the loop.
SPEAKER_00
So what you might not realize is up to this point, you as a human, you are the agent. You have a state of mind, here's what I'm trying to accomplish, and then I'm going through all of these phases.
SPEAKER_00
Okay, I start the pull requests, and the pull requests within your team says, well, you didn't quite follow the right format, so go back to the beginning. You're in the loop. Now your changes are in the PR, the tests are running, they fail, you need to go and change something in the code. You're back in the loop. A human reviewer comes back and says, well, it didn't quite use the right API, please go and change it. You're back in the loop. And then when you go and get your code, you're finally done, and you go and merge it. The merge queue says, well, another colleague managed to get some code ahead of you, and you have to go back in the loop.
SPEAKER_00
And when you're human skill, this opportunity to merge, the time that you go from when you're working the code until the code goes into the repository can be large because there's only so many changes that you're doing at the same time. But as you accelerate, this opportunity to merge is really important because the rate of change increases dramatically. So we talked a little bit about this.
SPEAKER_01
[SPEAKER_00] The PR is used as the unit of work, and that's what's really designed for human review. It expects a bit of delayed feedback. It's expected to go into discrete handoffs where you send it over to the reviewer and then it comes back. [SPEAKER_00] CI matters because it's validating the work that you're doing. It's doing things like, well, are you introducing a regression? Are you compiling and building your code from well-known source? Are there other changes that are going on that would be conflicting with this change? Is this change allowed? So all of that is part of this validation process that is automated.
SPEAKER_00
Human reviewers are overwhelmed. You've heard this many times. I don't have to repeat it. And the interesting thing is that the act of merging is starting to look a lot like high performance database problems, where you have serialization and you have a single ledger where every single change needs to go in and you need to lock the database in order to be able to commit. And the time that you have to lock when there are humans is large, but when there's machines, it's short. So the time to merge really matters. We need a new architecture.
SPEAKER_00
This is already how our team is working today and how we see a lot of the companies working today already that are at the forefront. There are no PRs. We start with intent and plan. This is what we want to achieve and we codify it. That's the spec. Someone writes it down. It might be in a linear ticket. It might be on Slack. It's somewhere. Somewhere you have written down what is the goal. What are you trying to achieve?
SPEAKER_00
That goes into a loop and this loop is a typical agentic harness. So it might be your cloud code. It might be we're big AMP fans. So in our case, it's often AMP. It might be cursor. It might be factory. You go into a loop and here the agent will check out your code and we'll start moving towards implementing your plan.
SPEAKER_00
Very importantly, the agent makes use of some of these invariants. Well, it checks out a well-known commit. So it doesn't just start from anything, for example. Then internal, what is internal validation? Well, it goes and uses the assets that exist in the repository to actually validate that the change is correct. So it builds it. It tests it. Then it comes back and tells you as a human, hey, I've just finished. Does it look good? Should I change something else? And you say yes. So you say continue. Continue is probably the word that we use the most nowadays. And then it just goes back and continues through the plan.
SPEAKER_00
Eventually you're done and you go into the merge queue and then it goes into the ledger. So your repository, your Git repository is a ledger. This is fast, but it's not fast enough because in this external validation, you still have a human in the loop. So where do we think we're moving towards? And this is in the span of weeks to months, not years. It's a world where generating code becomes much faster. It's already fast, but inference will only get faster.
SPEAKER_00
Internal validation. So running your builds and tests need to be extremely fast as well. And that's where you cannot go and spend 15 minutes running your tests or 45 minutes or any sort of minutes because you are delaying the whole loop. And external validation no longer has humans. We have other agents that are evaluating the changes. So you may have a security focus LLM. You may have an API conformance base LLM that is providing feedback within the loop to the changes that your main harness is then incorporating back into the code. And it's doing this very quickly.
SPEAKER_00
When it's done and in order to do it very quickly, it actually needs to be running in a stateful environment. Memory is important. State is important because if you're starting things from scratch all the time, you're just going to delay things even further. So the statefulness of it is really important within this loop. [SPEAKER_01] You are getting world signals from time to time. Things like, well, the plan changed or someone else got the change in. So the harness is also adapting its intent and plan, which then creates a new loop.
SPEAKER_00
[SPEAKER_01] And then when you're done, because there are so many changes going on, and you haven't yet as a human, the team hasn't accepted this change. You don't go directly into the repository. You go into a pre-queue, which we're starting to call a pre-merge, where there's a queue of changes that are done. They would have been merged if the process of merging was fast enough. But the reality is that you will have so many of these running in parallel and operating on the same parts of the code base that you need a process that reconciles them so that you can have serializability. So that you actually can guarantee that all the changes go back into the ledger, into your repository.
SPEAKER_00
[SPEAKER_01] And that's the point where you get external approval. That's where the human comes in, where they look at not the code, but that the intent was this and the result was that. And the result might be, here's the video of the feature working. It might be here's the output of the security focus LLM on this particular change. And it's not on one commit or one PR. It might actually be on multiple of them. So you may even have multiple agents independently working on features that go into this pre-merge queue and semantically get grouped into something that you as a human can manage. [SPEAKER_01] And that's the point where you get external approval.
SPEAKER_01
That's where the human comes in, where they look at not the code, but the intent and the result. And the result might be here's the video of the feature working.
SPEAKER_00
[SPEAKER_01] It might be here's the output of the security focus LLM on this particular change. [SPEAKER_01] And it's not on one commit or one PR. [SPEAKER_01] It might actually be on multiple of them. [SPEAKER_01] So you may even have multiple agents independently working on features that go into this pre-merge queue and semantically get grouped into something that you as a human can manage. [SPEAKER_01] Because there's going to be way too many. [SPEAKER_01] We already see that today where within our team, our volume of what we would call PRs from the past is four times as big as before. [SPEAKER_01] It's impossible for a human reviewer to look at every single PR.
SPEAKER_00
[SPEAKER_01] And if we think a little bit more into the future after this, if this process is extremely quick, one thing that may end up happening is that you may have to step into the multiverse.
SPEAKER_00
[SPEAKER_01] Okay. [SPEAKER_01] Where the starting point where the intent and plan gets applied is not the tip of the ledger. [SPEAKER_01] It's not the latest commit of your repository, because that is moving. [SPEAKER_01] There's many candidates. [SPEAKER_01] So the agents may actually be working on multiple commits at the same time to address the same plan. [SPEAKER_01] And in order to get there, this inner loop needs to be extremely quick.
SPEAKER_00
[SPEAKER_01] And it adds up in terms of capacity. [SPEAKER_01] So resource usage will also blow up because of all of the candidates that you're going to be exploring at the same time. [SPEAKER_01] This is the world that we think that we're moving towards. [SPEAKER_01] We're obsessed about performance and efficiency. [SPEAKER_01] So we're spending a lot of energy finding ways to maintain efficiency within this loop.
SPEAKER_00
[SPEAKER_01] And part of it is well, don't do work that is not necessary. [SPEAKER_01] Don't start things from scratch all the time. [SPEAKER_01] Have agents work a lot more as we did as engineers on our own workstations that were much more incremental. [SPEAKER_01] And that's the world that we're moving towards. [SPEAKER_01] Did CI go away? [SPEAKER_01] Well, CI still matters, but it's just shifted because the principles of, well, for example, does the code actually work? [SPEAKER_01] No longer is a separate phase, but it's just part of this loop. [SPEAKER_01] Every single iteration is going through validation.
SPEAKER_00
[SPEAKER_01] Now it's still going through enforcing those invariants as well. [SPEAKER_01] Like you still want to have, for example, for compliance reasons, you still want to have guarantees that you're starting from a well known checkout that you don't have someone in your company that came in and added other code that was never vetted. [SPEAKER_01] And you're starting from there. [SPEAKER_01] So those invariants need to still be enforced, but they enforce on a continuous basis. [SPEAKER_01] Coordination moves away from CI. [SPEAKER_01] So CI no longer has to guide different changes and make sure that different tests are passing in order for changes to be committed.
SPEAKER_00
[SPEAKER_01] That needs to be part of the overall loop. [SPEAKER_01] And governance is still important. [SPEAKER_01] But it also gets much more lifted into the harness and how the harness is coercing the change towards following everything that your team has codified within these processes. [SPEAKER_01] And that's it. [SPEAKER_01] This is where we leave that the world is moving towards. [SPEAKER_01] If you're interested about this topic, us at namespace spend a lot of time thinking about it.
SPEAKER_00
[SPEAKER_01] There's other folks in the industry as well. [SPEAKER_01] It's a crazy world and we need to be ready for it.
SPEAKER_00
[SPEAKER_01] Thank you. [SPEAKER_01] Thank you. [SPEAKER_01] And yeah, let's go for lunch. comes back and says, well, you know, it didn't quite use the right API, please go and change it. You're back in the loop. And then when you go and get your code, you're finally done, and you go and merge it. The merge queue says, well, you know, another colleague managed to get some code ahead of you, and you have to go back in the loop. And when you're human skill, this opportunity to merge, the time that you go from when you're working the code until the code goes into the repository can be large because there's only so many changes that you're doing
SPEAKER_00
at the same time. But as you accelerate, this opportunity to merge is really, really important because the rate of change increases dramatically. So we talked a little bit about this. Like the PR is kind of used as the unit of work, and that's what really designed for human review. It's, it's, it's, it expects a bit of delayed feedback. It's, it's expected to, to go into kind of discrete handoffs where you send it over to the reviewer and then it comes back. CI matters because it's kind of validating the work that you're doing. It's doing things like, well, are you introducing a regression? Are you compiling and building your code from well-known source?
SPEAKER_00
Are there other changes that are going on that would be conflicting with this change? Is this change allowed? So all of that is kind of part of this validation process that is automated. Human reviewers are overwhelmed. You've heard this many times. I don't have to repeat it. And the interesting thing is that this, the, the act of merging is starting to look a lot like high performance database problems, where you have serialization and you have a single ledger where every single change needs to go in and you need to lock the database, you know,
SPEAKER_00
in order to be able to commit. And the time that you have to lock when there are humans is large, but when there's machines, it's short. So the time to merge really matters. We need a new architecture. This is already how our team is working today and how we see a lot of the companies working today already that are at the forefront. There are no PRs. We start with intent and plan. This is what we want to achieve and we codify it. That's the spec. Someone writes it down. It might be in a linear ticket. It might be on Slack. It's somewhere. Somewhere you have written down what is the goal. What are you trying to achieve?
SPEAKER_00
That goes into a loop and this loop is a typical Asian harness. So it might be your, might be your cloud code. It might be, we're, we're big AMP fans. So in our case, it's often AMP. It might be cursor. It might be factory. You go into a loop and here the agent will check out your code and we'll start kind of moving towards the, and implementing your plan. Very importantly, already makes use of some of these invariants. Well, it checks out a well known commit. So it doesn't just start from, from anything, for example. Then internal, what is internal validation? Well, it goes and uses the assets that exist in the repository to actually validate that the change is correct.
SPEAKER_00
So it builds it. It tests it. Then it comes back and tells you as a human, uh, hey, I've just finished. Does it look good? Should I change something else? And you say yes. So you say continue. Like continue is probably the word that we use the most nowadays. And then it just goes back and continues through the plan. Eventually you're done and you go into the merge queue and, and then it goes into the ledger. So your repository, your Git repository is, it's kind of like a ledger. Uh, this is fast, but it's not fast enough. Because in this external validation, you still have a human in the loop.
SPEAKER_00
Uh, so where do we think we're kind of moving towards? And this is in the span of weeks to months, not years. It's, it's a world where generating code becomes much faster. It's already fast, but inference will only get faster. Uh, internal validation. So running your builds and tests need to be extremely fast as well. And that's where you cannot go and spend 15 minutes running your tests or 45 minutes or any sort of minutes because you are delaying the whole loop. And external validation no longer has humans. We have other agents that are evaluating the changes. So you may have, um, a security, uh, focus LLM.
SPEAKER_00
You may have an, um, uh, API conformance, uh, base LLM that is providing feedback within the loop to the changes that your main harness is then, uh, incorporating back into the code. And he's doing this very quickly. Um, when it's done and in order to do it very quickly, it actually needs to be running in a stateful environment. Memory is important. Memory is important. State is important because if you're starting things from scratch all the time, you're just going to delay things even further. So the statefulness of it is really important within this loop.
SPEAKER_01
Um, you are getting world signals from time to time. Things like, well, the plan changed or someone else got the change in. So the harness is also adapting his intent and plan, which then creates a new loop. And then when you're done, because there are so many changes going on, uh, and you haven't yet really, as a human, the team hasn't accepted this change. You don't go directly into the repository. You go into a pre-queue, which we're starting to call a pre-merge, where there's a queue of changes that are done. They would have been merged, uh, if we, if the process of merging was fast enough.
SPEAKER_01
But the reality is that you will have so many of these running in parallel and operating on the same parts of the code base that you need a process that reconciles them so that you can have serialize, serializability. So that you actually can guarantee that all the changes go back to back into the, into, into your ledger, into your repository. And that's the point where you get external approval. That's where the human comes in, where looks at not the code, but that the, this was the intent and this was the result. And the result might be, here's the video of the feature working.
SPEAKER_01
It might be, uh, here's the, uh, the output of the security focus LLM on, on this particular change. And it's not on one commit or one PR. It might actually be on multiple of them. So you may even have multiple agents, uh, independently working on features that go into this pre-merge queue and semantically get grouped into something that you as a human can manage. Because there's going to be way too many. We already see that today where within our team, where our, our volume of what we would call PRs from, from the past is four times as big as before. It's impossible for a human reviewer to look at every single PR.
SPEAKER_01
Um, and if we think a little bit more into the future after this, if this process is extremely quick, one thing that may end up happening is that you may have to step into the multiverse. Okay. Where, uh, the starting point where the intent and plan gets applied is not the tip of the ledger. It's not the latest commit, uh, that of your repository, because that is moving. There's many candidates. So the agents may actually be working on multiple commits at the same time to address the same plan. And in order to get that, uh, to get there, this inner loop needs to be extremely quickly, uh, extremely quick. And, um, it adds up in terms of capacity.
SPEAKER_01
So resource usage will also blow up because of all of the candidates that you're going to be exploring at the same time. This is the world that we think that we're moving towards. Uh, we're, uh, obsessed about performance and efficiency. So we're, uh, spending a lot of energy finding ways to maintain efficiency within this loop. And part of it is, uh, well, don't do work that is not necessary. Don't start things from scratch all the time. Uh, have agents work a lot more as we did as engineers that in our own workstations that were much more incremental. And, and that's kind of the world that we're moving towards. Uh, did CI go away?
SPEAKER_01
Well, CI still matters, but it's just shifted because the principles of, uh, well, for example, does, does the code actually work? No longer is a separate phase, but it's just part of this loop. Every single iteration is going through validation. Now it's still going through enforcing those invariants as well. Like you still have want to have, for example, for compliance reasons, you still want to have guarantees that you're starting from a well known, uh, checkout that you don't have someone in the, in the, in your company that came in and added other code that was never vetted. And you're starting from there.
SPEAKER_01
So those invariants needs to still be enforced, but they enforce on a continuous basis. Coordination moves away from CI. So CI no longer has to, uh, kind of guide different changes and making sure that different tests are passing in order for changes to be committed. That needs to be part of the overall loop. And governance is still important. Uh, but it also gets much more lifted into the harness and how the harness is, uh, uh, coercing the change towards following everything that your team has codified, um, within these processes. And that's it. This is where we, we, we leave that the world is moving towards.
SPEAKER_01
Um, if you're interested about this topic, uh, us at namespace, um, spend a lot of time thinking about it. There's other folks in the industry as well. Uh, it's a crazy world and we need to be ready for it. Uh, thank you. Thank you. And, uh, yeah, let's go for lunch.