All right. Hello, everyone.
SPEAKER_00
[SPEAKER_00] We're getting to the back end of the conference. I don't know if that's a good thing or not for you. Start the weekend or maybe sad that the conference is over. So, going to be talking about the missing primitive for agent swarms. So, talking about sub-agents and swarms within the context of coding agents and the infrastructure underneath them. So, just a quick introduction. My name is Lou. I'm the field CTO at a company called Ona. Previous life, I was principal engineer and platform engineer. Joined Ona doing product management. And now, I work a little bit more with our customers on the field side.
SPEAKER_00
So, in my world at least, everyone is trying to build a form of a software factory. I've seen a few different talks at this conference of similar ideas, trying to take coding agents and then apply them across the software development lifecycle. I'm wondering actually if this statement is entirely correct because I've definitely chatted to a lot of people over the course of this week that are not yet at this point of thinking about this. But it's been very much occupying my head space for the last few months as well.
SPEAKER_00
So, I added a definition in this slide as well just to quickly define what I believe to be a software factory, which is the commitment to incrementally moving the human out of the loop within the SDLC. Such that the human is not proactively interacting with the computer. I say that because I've seen a few talks and some people talk about software factory with these parallel agents, like one individual IC running lots of coding agents at the same time. It's not my personal definition. My personal definition is that you're slowly bringing the human out of it and then work is flowing from development into production theoretically in an automated fashion.
SPEAKER_00
But we're extremely early in terms of where we're at with software factories. So, obviously, there's a bunch of funky little visualizations that I've been showing. And a lot of these are actually taken from this website, backgroundagents.com, that I created, which seemed to resonate quite a lot. So, please do have a look at it. It encapsulates some of these ideas as well. But we have different patterns really for running these agents at scale, coding agents at scale.
SPEAKER_00
So, one of them really is this swarm pattern, which for my definition is starting off with an individual intent, firing that out to a number of different agents, and then funneling that back in, let's say, to an individual PR or task. And that's almost your typical sub-agent process that we've seen quite a few times. Fleets is something you can't see this very well in the room, unfortunately. But on the top right, the fleets is where you're fanning out agents across, let's say, a number of different repositories inside of an organization, which is a capability we've had in owner for quite some time now.
SPEAKER_00
And you see it a little bit popping up here or there, but not so much. But I think it will be something that organizations will start to take advantage of a lot more in the future. And then at the bottom, you've got events. So, if you think about how do we take the human out of the loop and build this software factory, you need to know how and when are you going to trigger those agents, when do they come online. And a lot of this already exists with existing webhook infrastructure and things like that. That PR is raised, linear ticket is created, et cetera.
SPEAKER_00
One thing that's been useful for us over the last couple of months is a lot of these large companies have also come forward and shared some of their implementations of this infrastructure. So, I'll mention a few notable ones. But Stripe has one. They're what they call Minions, which is built on top of their existing infrastructure, where they've then plugged in these coding agents and they're able to then drive thousands of pull requests inside of Stripe. Another very noteworthy one, and Ramp has been so loud on social media these last couple of weeks. It's been quite insane.
SPEAKER_00
But they also built one that they internally call Inspect, which again is their infrastructure for running these background agents. And it's something that we've been doing now for quite some time. So, Owner as a platform has been infrastructure for development environments for about six years. But over the last year or two, obviously integrating further with the agent side of that as well. What Owner effectively does is allow you to spin up any number of different development environments. But one of the additional features that we have that I mentioned is this fleet feature.
SPEAKER_00
So, you can automate on schedules or triggers agents that spin up to resolve issues across a number of different repositories.
SPEAKER_00
Use cases for this is things like CVE remediation or bumping test coverage or enforcing something at scale. As it stands today with current LLMs, that seems to be often tasks that are somehow simple. But the hard part is that you're doing this across a number of different teams. Thousands of teams. Or thousands of repositories. Which is what I'm showing you here. So, it works as a workflow creation. You're adding prompts, scripts, and things like that. And then using that to drive this change across your organization. I did want to give a notable shout out also to the Harness Engineering blog that OpenAI and Ryan created.
SPEAKER_00
Because this encapsulates much of this mindset of trying to then take and encode as much of your process into your repository, into your context files, your agents.md, in order to build effectively that software factory. Over the last couple of days, I had a few questions with people talking about Harness Engineering and what is it. For me, it's really another extension on context engineering whereby everything in your repository, from skills to Agents.md to unit tests, everything that you could possibly use to give feedback to your agent is for me Harness Engineering.
SPEAKER_00
So, I include this little visualization in the top corner because as well, for me, Harness Engineering is very much about doing things, letting the agent run through, figuring out where the agent gets lost, and then encoding that knowledge back into your repository or context. Again, to try and get the agent flowing through the software factory as much as it can. So, if you take a step back and then think about from an infrastructure level what you need effectively to build this form of software factory, the first piece of that puzzle is a runtime. Or you need somewhere for the agent to run. And I believe this mostly is pretty much a solved problem now.
SPEAKER_00
You then need a way to orchestrate these. So, you need to run them at scale.
SPEAKER_00
So, you need a way to run these agents, scale up, scale down horizontally. You need some way to trigger them. Again, to try and get the agent flowing through the software factory as much as it can. So, if you take a step back and then think about from an infrastructure level what you need effectively to build this form of software factory, the first piece of that puzzle is a runtime. Or you need somewhere for the agent to run. And I believe this mostly is pretty much a solved problem now.
SPEAKER_00
You then need a way to orchestrate these. So, you need to run them at scale. So, you need a way to run these agents, scale up, scale down horizontally. You need some way to trigger them. But for me, one of the biggest difficulties if you try and build this today is effectively agent coordination. So, how do you get the agents to interact with each other, pick up tasks from each other, how do they collaborate?
SPEAKER_00
For runtimes, lots of different approaches here. But you can run agents as separate threads. You can isolate them more in work trees. You can then go one step of abstraction further and put them in containers and VMs or micro VMs. Or what ONA does basically, we really call these dev environments. So, the sandbox conversation has made this very blurry. But we at least believe that for running proper development tasks, it has to be inside of a virtual machine.
SPEAKER_00
The reason for that is for the isolation from a security standpoint, a container is not bulletproof isolation boundary. So, if you have an agent running in there and you want to secure it, there's challenges for a container. They're also bursty if you run them on Kubernetes or in pods. You have noisy neighbor problems. You're going to have compute contention across different containers. And only with having the full isolation of a VM will you be able to effectively do this properly. Let me run back through my presentation. Sorry. Cool.
SPEAKER_00
So, let me actually just quickly, it's a good point actually at this point. I will show you a quick demo of how this actually looks inside of the ONA interface because a lot of this is a little bit theoretical. I did a quick recording of this yesterday in my hotel room just in case the Wi-Fi was terrible in here. But let me just run you through this.
SPEAKER_00
So, in here you see the ONA interface. On the left-hand side you have all the different tasks that I have running. So, I kicked off two different tasks here. The first one I asked ONA to implement me Symfony. So, Symfony has a spec in the repository that talks about how to implement Symfony. It's a very detailed spec. So, I wanted to use it as an example. I gave one of the agents and asked it to spin this up using process-based agents. So, sub-agents running within the environment itself. So, take the VM, run the agent inside, and that agent will spin up sub-agents within that VM.
SPEAKER_00
The other one I asked it to is effectively to run me a fleet with a number of different VMs. So, the agent is actually then empowered to create other VMs inside of the platform. And it can spin up technically infinite of these. So, wherever you're running this, you're only really inhibited by as much as you're willing to pay and as much as your cloud provider can scale to. So, if I run this through, what we see, and I might have to skip through here a little bit, is we see... I jumped too fast.
SPEAKER_00
This bottom agent, the VM one, then it will spawn these three different sub-agents. So, it spins up the different VMs which we see coming in on the left-hand side. The parent is the controlling agent, and the sub-agents obviously then are given small bits of context, individual tasks to complete, and then we'll do message passing back to that parent agent to control and govern the overall task itself. One challenge for sure we have is how do we build the UX for this? Like, as you build more and more complicated tasks, how do you think about managing and controlling these sub-agents, and how do you think about the UX on top of them?
SPEAKER_00
So, as this progresses, you'll see also on the left-hand side, eventually you start to see... As the agents come online, they start their environments and they start to work through tasks. Both of these have, for some reason, seven sort of sub-items that they're working through, so you can see those. When that gets to the end, it's then obviously going to terminate those VMs. That's all going to collapse down and your task is complete.
SPEAKER_00
The second form of UX for this that we have with the sub-agents is this one here with the process level, which I'll pause so that it's not jumping around. When you launch a process level sub-agent, it happens all within the single agent window. So, you actually see at the bottom here a stack of a number of different sub-agents that are starting. And then when you click on those, you can then open up a new chat window, which is almost your new context, and use that. So, you've got this like two different forms of this. One entirely isolated VMs to scale out these swarms, and another one at the process level where you can run it within the individual VM itself.
SPEAKER_00
So, lots of stuff going on there, but I wanted to show you conceptually, people say, what does this swarm look like in reality? How does this actually look? And this is how it looks within ONA. So, give me one second to fly back through.
SPEAKER_00
So, coming back to the software factory idea. So, how do I know about some of the challenges about software factories? And that is because I've also tried to build one. Obviously, I implement many of these ideas in our own projects, but I also, similarly to what OpenAI did building out their symphony projects and their Harness Engineering blog, is, okay, what can we build a project without touching any lines of code, and can we automate as much of this process as possible, such that the agent can actually then develop everything as autonomously and self-driven as possible?
SPEAKER_00
And it turns out, yes, the technology is there today, but there are some different challenges that you have.
SPEAKER_00
What's missing? One thing that you find out is that the SDLC is not this. You know, this is the conceptual SDLC that we present and talk about with each other. The very coarse-grained five steps of the SDLC. Agents don't respect this and they don't understand. I mean, these boxes contain a ton of complexity. So, if we take something like plan or a plan stage, actually within that, our SDLC has a ton of different sort of microsteps almost to it. And if we're wanting then to train agents to then step through the SDLC, we need to find ways to actually break down the SDLC into some of these microsteps. What's missing? One thing that you find out is that the SDLC is not this.
SPEAKER_00
This is the conceptual SDLC that we present and talk about with each other. The very coarse-grained five steps of the SDLC. Agents don't respect this and they don't understand. These boxes contain a ton of complexity. So, if we take something like plan or a plan stage, actually within that, our SDLC has a ton of different microsteps to it. And if we're wanting then to train agents to step through the SDLC, we need to find ways to actually break down the SDLC into some of these microsteps. And that happens all the way through the SDLC. So, if you want to build some form of software factory, we then need to start figuring out how to solve these microsteps.
SPEAKER_00
How do we get agents to sufficiently follow those steps and do them in deterministic ways as well? And the hard part of this is context.
SPEAKER_00
Probably not a huge innovation to everyone in this room. Context windows are some of the hardest parts about working with LLMs because, as the last speaker said here as well, context rot. Context, once the context window becomes consumed, the agent starts to lose track of where it's going and things like that. It gets less effective. They also skip steps typically. They want to please us. They're quite sycophantic. So, they might ask them to write some tests and then they're going to skip some tests in order to complete the task. So, a lot of the tools that we have today for coordinating these agents as you start to spin them up.
SPEAKER_00
Once you've got this infrastructure, you have the runtimes, you have the orchestration. The missing piece then you also need is coordination. So, if you use something like GitHub, GitHub is not a coordination layer for agents. It gets incredibly overwhelming. You can have your agents raise a pull request. You can review it. You can then solve the merge conflicts. You can fix the CI build, et cetera. But this gets incredibly noisy for you as a human to make sense of where you should step in to intervene with that agent itself. And GitHub is a poor solution for this. Symphony is built on top of linear but suffers the same problem.
SPEAKER_00
We're reusing existing human tools in very weird ways for agents. But I think we are now at the cusp of effectively solving this. And I think it might be solved in a few different ways. One is through state machines. By building out workflows and effectively state machines that we've had for a very long time. And then building these as our versions of our SDLC. I do believe to a degree some of the ideas from durable executions comes into this as well. Lots of companies have pioneered this over time to be able to run a process in a durable fashion. But we do need to solve also the gates and compliance section of this.
SPEAKER_00
And I do believe actually there is definitely a gap right now also for packaging this in some form of CLI construct whereby you can run this locally in a development environment. But also then remotely in a CI or some other fashion. But that's effectively it. If anyone wants to chat and go a little bit deeper, I think we've got a couple of minutes. We can also take some questions if we need to. Yes, my definition of software factory is moving this human on the loop so they're not necessarily driving each of these individual changes. Context and context management is by far the hardest part of building this software factory.
SPEAKER_00
Out of these primitives, I do believe we've effectively solved the runtime. There are many options for this now. Sandboxes and containers and other solutions. The orchestration is effectively solved. The triggers are solved. But the thing that's missing for me is coordination. Also, one of our folks here is here from our security team. And security is another piece of the puzzle that we really need to solve as well to drive more automation. But it's this coordination layer that I think is largely what's missing. Obviously, a lot of this is a little high level.
SPEAKER_00
On the 6th of May, we'll run a virtual summit to go into these topics about software factories and background agents.
SPEAKER_00
So if anyone here would like to attend, you can do. It's on backgroundagents.com. And we have a CFP up for it as well. If anyone would like to speak, that would also be great. There's not many people in the world doing this. So if you're at least experimenting with it, it would be great to also talk to you as well. The other thing is a few people that I've been speaking to at the conference, especially about software factory, complained that obviously a lot of these talks are a little theoretical, which is true. So some of you might want to be really getting into the weeds of this.
SPEAKER_00
So actually next week, with Zach from Ono, we're actually going to build one of these in public. Starting from scratch, build a software factory. We'll do this over the course of two weeks in public to show what that looks like with today's technology. Just to really show all the ins and outs of the workflows if anyone wants to see what that actually looks like as well. But there you go. Thank you very much. We do have two and a half minutes if anyone has questions. But if not, I'll also be outside and downstairs. Yes. [SPEAKER_01] Thank you for the talk. [SPEAKER_01] Super helpful.
SPEAKER_00
[SPEAKER_01] Could you maybe go back to the slide where you were talking about the problems with the coordination layer? Yep. So I need to repeat the question actually for the stream. So the question was problems with the coordination layer and sort of more concreteness about the solution. What could that potentially look like? So I have, we have a number of different prototypes for this internally. And I've also been trying to think about how we do with this. So I think one form factor that works is the CLI. So I've seen some other solutions, graph based ones. There are some open source things right now where defining the workflow effectively as a graph.
SPEAKER_00
Like what you drive out of mermaid diagrams like NA10 type of workflows. Where you can define some of these prompts in that form. I think that ultimately needs to be packaged in some form of CLI though. What could that potentially look like? So I have a number of different prototypes for this internally. And I've also been trying to think about how we do with this. So I think one form factor that works is the CLI. So I've seen some other solutions, graph based ones. There are some open source things right now where defining the workflow effectively as a graph. Like what you drive out of mermaid diagrams like NA10 type of workflows.
SPEAKER_00
Where you can define some of these prompts in that form. I think that ultimately needs to be packaged in some form of CLI though. What I mean by that is I think that if you have a local running agent, code, whatever, any local running CLI, I think that now needs to have something that integrates with this tool so it can invoke this to say, hey, have I achieved this part of my SDLC and can I now proceed to the next part of it as a CLI? If you want as well, I can have a full spec for this. I can show you after the talk as well. We can go a bit deeper into it. But I'm seeing this solved in a number of different ways. People are solving this in a few different ways.
SPEAKER_00
There's the NA10 workflow type of diagram way. CLI gateway is the prototype that I have as well. And then there's a range of different hacky interim solutions. There's one even the open claw folks have as well, which the name is eluding me.
SPEAKER_00
But I think ACPX, which they built on top of ACP, which builds out some of this workflow. There's a GitHub one, Fabro, some folks are playing around with. But it's very nascent right now. Yeah. Any other questions? Are you using some protocol for the CLI? Because ACP, there's eight-way, there's a bunch of, yeah. I don't know if you will agree on the standards. [SPEAKER_01] Currently not. [SPEAKER_01] So the implementation I have, I'm trying to understand whether or not to release it as an implementation or as a standard, to be honest. Because I largely don't really care about the implementation.
SPEAKER_00
I almost care about the standard for it so we can collaborate on the standard. But it's not built on top of ACP, at least not yet. But it might be. Not yet either, no. Solving a slightly different problem space. Cool. All right. Thank you very much. Enjoy the conference, everyone. Thank you. Thank you. Thank you. So some of you might want to be really getting into the weeds of this. So actually next week, with Zach from Ono, we're actually just going to build one of these in public. Starting from scratch, build a software factory. We'll do this over the course of two weeks in public to show what that looks like with today's technology.
SPEAKER_00
Just to really show all the ins and outs of the workflows if anyone wants to see what that actually looks like as well. But there you go. Thank you very much. We do have two and a half minutes if anyone has questions. But if not, I'll also be outside and downstairs. Yes. Thank you for the talk. Super helpful. Could you maybe go back to the slide where you were talking about the problems with the coordination layer?
SPEAKER_01
Yep.
SPEAKER_01
So I have been, so I need to repeat the question actually for the stream. So the question was problems with the coordination layer and sort of more concreteness about the solution.
SPEAKER_00
What could that potentially look like? So I have, we have a number of different sort of prototypes for this internally. And I've also been trying to think about how we, what we do with this. So I think one form factor that works is the CLI. So I've seen some other solutions, graph based ones. There are some open source things right now where defining the workflow effectively as a graph. You know, like kind of like what you drive sort of mermaid diagrams out of like NA10 type of workflows. Where you can define some of these prompts in that form. I think that ultimately needs to be packaged in some form of CLI though.
SPEAKER_00
What I mean by that is I think that if you have a local running agent, code code, whatever, any local running CLI, I think that now needs to have something that integrates with this tool so it can invoke this to say, hey, have I achieved this part of my SDLC and can I now proceed to the next part of it as a CLI? If you want as well, I can, I have a full spec for this. I can show you after the talk as well. We can go a bit deeper into it. But I'm seeing this solved in a number of different ways. People are solving this in a few different ways. There's the NA10 sort of workflow type of diagram way. CLI sort of gateway is the prototype that I have as well.
SPEAKER_00
And then there's a range of different sort of hacky interim solutions. There's one even the open claw folks have as well, which the name is eluding me. But I will, I think, ACPX, which they built on top of ACP, which builds out some of this workflow. There's a GitHub one, Fabro, some folks are playing around with. But it's very nascent right now.
Yeah. Any other questions? Are you using some protocol for the CLI? Because, you know, ACP, there's eight-way, there's a bunch of, yeah. I don't know if you will agree on the standards. Currently not. So the implementation I have, I'm trying to understand whether or not to release it as an implementation or as a standard, to be honest.
SPEAKER_00
Because I largely kind of don't really care about the implementation. I almost care about the standard for it so we can collaborate on the standard. But it's not built on top of ACP, at least not yet. But it might be. Not yet either, no. Solving, I feel like, a slightly different problem space.
SPEAKER_00
Cool. All right. Thank you very much. Enjoy the conference, everyone. Thank you.
SPEAKER_00
Thank you. Thank you.