Open Reader

Scaling Agents on Kubernetes with acpx and ACP — Onur Solmaz, OpenClaw

completed 19:00 May 21, 2026 Watch on YouTube

Current Status

completed

Video ID

VaS2h-dY1-4

RAG / Chat

Enabled
Scaling Agents on Kubernetes with acpx and ACP — Onur Solmaz, OpenClaw
Description

OpenClaw receives 300 to 500 pull requests per day. Most arrive AI generated, most are not mergeable, and every one of them is signal about something broken in the codebase. Onur Solmaz built acpx to process them without him in the loop. acpx is a headless CLI for the Agent Client Protocol. It replaces PTY scraping with structured agent to client communication and drives sessions through a node based workflow graph: reproduce the bug, judge the implementation, check for conflicts, run a review loop, emit structured JSON. Onur runs parallel Codex sessions from Discord channels while traveling, one channel per task. The talk ends with disposable agent pods on Kubernetes, a Go operator that provisions a full compute environment per task, wires it into Slack, and tears it down when the work is done. Speaker info: - https://x.com/onusoz - https://www.linkedin.com/in/osolmaz/ - https://github.com/osolmaz

Summary

Generated by claude-haiku-4-5-20251001

Scaling Agents on Kubernetes with ACPX and ACP — Summary

Main Topics

  • Agent Integration & Interoperability: Building standardized protocols for agent-to-client communication (ACP) and agent-to-agent coordination
  • OpenClaw Community & Workflow: Managing large-scale open-source contributions through AI-assisted PR review and automated workflows
  • Enterprise Agent Orchestration: Deploying autonomous agents on Kubernetes for scalable, task-specific operations
  • Developer Tools & Automation: Leveraging AI agents for code review, refactoring, and development workflows across multiple platforms (Discord, Slack, Teams)

Key Points

Background & Experience

  • Speaker has been building AI coding harnesses since before ChatGPT, evolving from JupyterLab extensions to full autonomous agent systems
  • Long-time OpenClaw contributor; currently focused on agent interoperability and enterprise adoption

The Problem with Current Agent Deployment

  • Fragmentation: Different platforms (VS Code, Zed, Google) each build separate agent integrations—massive wasted effort
  • Single Instance Limitation: Current chatbot-style deployment (one agent per Slack/Teams workspace) doesn't scale for enterprise needs
  • Manual Management: Creating multiple agent instances requires manual app manifests and configuration

ACP (Agent Client Protocol)

  • Purpose: Standardizes human-to-agent interaction (distinct from MCP, which provides tools to models, and Agent Protocol, for agent-to-agent communication)
  • Benefit: Build once, deploy everywhere across different platforms and editors
  • Current Adoption: Only Zed had full ACP adapters when speaker needed them

ACPX - ACP CLI & Workflow Engine

  • "Swiss Army knife" for ACP—enables calling any ACP-compatible agent from the command line
  • Includes A10 workflow engine for automating repetitive development tasks
  • Automates OpenClaw PR triage workflow:
  • Intent discovery
  • Implementation judgment
  • Conflict detection
  • CI validation
  • Review loops with shallow bug fixes (not deep refactoring)

Enterprise Agent Vision

  • On-Demand Disposable Agents: Spin up task-specific agents rather than one permanent instance
  • Key Components Needed:
  • Kubernetes orchestration
  • Agent harness (OpenClaw, Codex, etc.)
  • Standardized protocols (ACP)
  • State/data synchronization (similar to Dropbox/rsync)
  • Proper multi-agent provisioning on chat platforms

TextCortex Spritz Project

  • Open-source Go operator for Kubernetes managing agent orchestration
  • Use Case: Error reporting/debugging on Slack—dispatches autonomous agents to investigate prod issues
  • Architecture: Full Kubernetes pods per agent (acknowledges resource overhead but values the autonomy and power)
  • Features:
  • React-based UI hosted in cluster
  • Not locked to specific agent framework (abstracted via ACP)
  • Can switch agents/models without architecture changes

Notable Quotes

> "You need to apply agents generously on problems... take yourself out of the loop and solve it with agents."

> "By the time the PR ends up in front of you, should be resolved ideally and can be resolved."

> "Because OpenClaw showed the power. When you give a full computer to an agent, it's a lot more powerful."

> "At OpenClow, you have fire hoses... The biggest challenge of the project currently is how do you absorb all the needs and wants of these people?"

> "You're automating the automator."

Takeaways

  • Standardization is Critical: ACP-style protocols reduce wasted engineering effort and enable true multi-platform agent deployment
  • Workflow Automation: Mechanical tasks (PR review, refactoring, conflict resolution) should be programmed as agent workflows, not done manually
  • Enterprise Scale Requires New Patterns: Move from single persistent agents to on-demand, disposable agents spawned per task, coordinated via Kubernetes
  • Architecture Matters: Giving agents full compute resources (pods vs. sandboxed environments) significantly increases capability
  • Interoperability Over Vendor Lock-in: Abstract agent implementations via ACP to avoid dependency on any single framework
  • Practical Path Forward: Textcortex Spritz demonstrates a working implementation of autonomous agents on Kubernetes for enterprise workflows

Resources:

  • ACPX: GitHub repo at TextCortex/Spritz
  • Available for consultation on internal deployment and large-scale issue processing

Transcript

2707 words en Processed in 234.8s

Welcome. So the talk is building an ACP at OpenClo. It's also about other things like how to put open source agents on open source agent frameworks on Kubernetes and stuff. So I hope today I'll share in a nice way what I've been working on in the last two months. A little bit about me, very brief. I've been building harnesses since a few months before ChatGPT came out. I built a JupyterLab extension over OG codex model, like DaVinci Code 2. Back in the days that eventually... So I'm currently working for a startup. That initial coding harness turned into its current harness over time, like a ship of theses. It got ripped apart and put back together so many times. And yeah, I'm a founding engineer there. I've been in the industry since three and a half, four years. I'm using OpenClo since CloudBot first dropped in Discord. I went in there. I've been following Peter since he wrote CloudCode is my computer. I was thinking, this guy's crazy because CloudFour is on it. I haven't given my machine to it. But he ventured forward and he's paid the way for us. And when I saw CloudBot in Discord, my mind was blown. Next day, I installed it at the company in our cluster. It was talking to people on Discord. And everybody else's minds were blown. And we were used to selling the enterprise for a few years now. So that's why I started by adding an MS Teams integration in case it might be useful at some point. And I became a maintainer along the way. I was there when it was renamed two times. And today, my focus is on agent interoperability and orchestration. One of my goals is to accelerate enterprise adoption of OpenClo and adjacent software and also address that OpenClo is not secure. Well, it will be secure or Peter talked a lot about that earlier. So I don't have anything else to add on top of it. It's work in progress. Sorry. I'll use this. So I started on developer workflows right away. So I created a PR and called it Discord-driven development. Well, Telegram-driven development suits better because it's TDD. And then right after I set up personally and I realized Opus is not so reliable for complex... It wasn't back in the days. It's now better. Agents are improving. But Codex was my main harness and I wanted to use Codex in Discord. But it was basically... I was playing a telephone game. I was telling Opus to tell Codex to do something. God knows what it's saying because wording matters when prompting. But it was working somehow. And I would go and look at the Codex session. And then it paraphrased what I was saying. But eventually, I got some stuff done. But I knew this could be done much easier. And today I'm running a full ID on Discord. I got parallel workloads. You have one to five channels. At any point, I'm working with one to five agents. So you see Codex 1 to 5. And then close for testing the OpenClose ACP feature. And basically that's how some of my channels look. And it's very good for coding on the go because I'm a guy who's addicted to side projects. And I like to just... AI is making things a lot easier to do these sorts of things. In parallel, you have an inspiration. You execute on a weekend. And you just get it ready. You ship it. ACPX, which I'm going to talk about, is similar. I built it through Discord. And what you do is you bind the Discord channel to Codex through ACP. You can also use Codex AppServer Protocol. Harold, another maintainer, developed it. And here is me using it before I'm flying to London to create a PDF about converting the docs of ACP into PDF. And then I have to say put it in temp because Codex doesn't know about the harness. And it cannot send me in Discord. So then I go to another channel and tell it to send it to me on that channel. So we are developers and our tools. We don't have time to polish them very well. But we know the advantages and disadvantages. What is ACP? So it's not... Most people say, is it MCP? MCP is for giving tools to the model. ACP is for standardizing agent to client interaction. It's a shout out to Zed. I actually forgot to put Zed logo. Zed is building a new editor in Rust, more efficient, lower memory usage, not electron. So I'm using Zed also since last fall. And if you use Codex on VS Code or Cloud Code, they are all building different plugins. It's not exactly... It's so much wasted work. If only you could just standardize them under one interface and you just build it once and then you ship it. And that's what the idea... That's the idea they had. It's also more... It's much less duplicated work. There are competing standards. So agent... Agent protocol, that's for agent talking to agent. ACP, agent client protocol, is for human talking to agent. But agent can use the human one to talk to other agents as well. In the long run, as these protocols get adopted, we will support all of them. We will weigh the advantages and then we will use them somehow. But when I needed the most, I needed adapters for Codex and Cloud Code and only Zed had built them. And Google didn't have them back at the time. So that's why I chose it. When you're adding functionality to OpenClo, you do it through a CLI. So I said, okay, let's create a CLI for ACP. Let's call any other agent over the command line. So that's how it started. And slowly turning into a Swiss Army knife for ACP, as I will show in a bit. So at OpenClo, you have fire hoses. We have over 60 PRs total. 300 to 500 per day on average are open. And basically overnight, people woke up and decided they want... They like OpenClo. So we have tens of thousands of stakeholders who want to add features to OpenClo. And the biggest challenge of the project currently is how do you absorb all the needs and wants of these people? And how do you balance them? You can't please everyone. How do you create an elegant system without creating AI slope that can cater everyone's needs? And Peter's workflow gave me an idea. So this is something I do. You go and you ask your clanker, what is this? So a PR comes your way. Most of the time, it's AI-generated description. If the human put thought into it, great. But you need to ask what it's doing. You ask, is this the best possible fix? Most of the time, it's no. You need to take this data point. It's crucial feedback from the user. So you need to put a category on the code. And then you either continue... He wrote it. You do some back and forth with the agent. The reason is people just run into an issue with their OpenClo. And it can use GitHub. And then there's a please fix. And then they just send some slope your way. This is... You can merge it, but you can also fully discard. You need to take this data point. It's crucial feedback from the user. So you need to categorize it, put it in a bin. It tells you when some part of the code is broken. And Vincent also talked about that a little bit. And then you go and then you either continue. He wrote it. You do some back and forth with the agent. The reason is people just run into an issue with their OpenClo. And it can use GitHub. And then there's a... Please fix. And then they just send some slope your way. This is... You can merge it, but you can also fully discard. You need to take this data point. It's crucial feedback from the user. So you need to put it, categorize it, put it in a bin. It tells you when some part of the code is broken. And Vincent also talked about that a little bit. And here is his codec session doing that. On one side, there's one that says it's a good fix. And on the other side, it's just saying it's not bad. And this is so mechanical. And once you do this over and over, you realize you're repeating something. If only you could program something to automate it. So you're automating the automator. So I created... I started to create workflows in an abstract way. So item comes PR. And then you find the intent. You're judging implementation. You're looking into... If that's conflicts. If reviews gives you issues that need to be addressed. So you need to make the CI pass. If it's not passing already, most people just don't care about that. All this mechanical work. By the time the PR ends up in front of you, should be resolved ideally and can be resolved. And that's what I'm working on. In the workflow, we have the shameful Ralph review refactor loops. I'm a believer that some people say give the... Just running an agent in a loop does not necessarily have to be something that will create slope. As long as you're not making it design something, but you're making it uncover shallow bugs that can be easily fixed. Should be fine. So in the abstract workflow that I created, which is actually called turn into a program. You can tell it to do superficial refactors. And then you can tell if it needs a fundamental refactor relate to the human. Resolving conflicts also doesn't need... No, it was hard back in the days. I don't think anyone is resolving conflicts by hand by now. So we are basically creating standard operating procedures for agents. That's a fancy word for workflow. So that's what I built into ACPX. It's an A10 workflow engine, but it's driving a codec session. You can see on the right. Let me show you in action. So this was 1PR. It's loading. Let's speed it up a bit. So there's some programmatic parts. So this is just replaying what it's doing. It's reproducing the bug. It's judging refactor. It's reviewing now. I'm doing a review loop. And then review didn't bring anything. It's what I do, but I make it output JSON structured data so I can put it in an A10 workflow. And I will talk... This is a general workflow engine. You can use it on other things as well. You need to apply agents generously on problems. So I see it as an ointment that you apply generously. On any problem that can be solved with agents, you need to take yourself out of the loop and solve it with agents. And personal agents, I think, are on a... There's a spectrum enterprise and personal agent. And normally you see the work you use at the computer and then the PC... The PC you use at work and the PC you use at home used to be relatively similar. But that will not be the case with agents because at work you will be using... Consuming a lot more inference. And that means in enterprise there will be a lot more money to be made. So that's why I'm also excited about enterprise agents and OpenClaw's potential. That's why I believe in on-demand disposable agents. If you use OpenClaw on Slack or Teams or Discord, it's just one instance. You create an app. The problem is you can't really talk to multiple instances of this. You create a connection. You create a Slack app. To connect another agent and another name and another profile picture, you have to create another app. And you have to create an app manifest. It's something that shouldn't be managed manually by clicking. And the platforms like chat apps don't have this standard yet where you can do multi-agent provisioning. Cosmetically create different agents. This is what... This is on the screen. You see I asked ChatGP to generate the idea. You have agents and then the name can be generated by an underlying app. And you can talk to them separately. This is not supported and this must be supported for this vision to work. Until it's supported, I'm using it on another UI. Because we are all gonna start one agent per task. And there will be... They will work on these tasks. And they will be creating files, editing files, and it will all be synchronized. It will be... It will be a tad bit different than what you're used to with your personal agent. And to have that, you need to have a few key components. You need to have Kubernetes. You need to have an agent harness. Open Cloud could be one of them. Codex Cloud Code could be one of them. It could be ACP. It could be GitHub. You give read-write access and you do state data synchronization. Maybe you do something that's used R-Sync. Something like whatever algorithm Dropbox is using. And there are some projects that are taking this on. I've been working on this. This is outside of Open Cloud. This is my day job. And it's an open source orchestrator. It's a Go operator that basically handles the complicated parts. You know there's a user experience. You want to create a concierge on Slack. A concierge agent. And you're talking to it. But you get bottlenecked because you're a hundred employees on Slack. So you need to... For some other task, you may need to create a new one. And then it creates and gives you a website link. I'm gonna skip the... Shout outs to Cognition and Devin. Because they invented the category. But I'm running low on time. So the repo is at TextCortex Spritz. I'm just gonna demo it for our use case. We use it on error reporting currently. So if you're on Slack, you can ask it to dispatch an agent to debug it. You know, you're asking any new bugs after prod release and it's saying something. And then you ask it to create an agent. If I could put that agent into Slack, I would. But I can't do that. So I have to put it in another UI. And this is an open source project. You can take this. UI is also a React app hosted in the cluster that you're deploying these Helm charts to. And it starts the conversation there. It starts working on the problem. This is Codex Web or Devin or anything. But it's actually using a full Kubernetes pod. Wasteful, but I think it's a better abstraction. Because OpenClaw showed the power. When you give a full computer to an agent, it's a lot more powerful. And I believe that as well. I think OpenHands uses Firecracker. So I'm also not so well-versed on all the different virtualization frameworks. So I'm also learning along the way. But I have a working product that's running on Kubernetes. And you can use this product. Yeah, it starts the conversation there. It starts working on the problem. This is like Codex Web or Devin or anything. But it's actually using a full Kubernetes pod. Wasteful, but I think it's a better abstraction. Because OpenClaw showed the power. When you give a full computer to an agent, it's a lot more powerful. And I believe that as well. I think OpenHands uses Firecracker. So I'm also not so well-versed on all the different virtualization frameworks. So I'm also learning along the way. But I have a working product that's running on Kubernetes. And you can use this product. If you're interested in deploying internally and using Codex on the web, and then just spin things off. If you're a back-end... If you have an open source project, first of all, I can help you set this up. If you have a system like Inlet of just hundreds of issues per day, I can help you process it. This is all the wiring around Slack. And keeping those agents on. And the user experience. And then... You're not locked in any agent. You can switch. It's all abstracted with ACP. Yeah. That was my talk. Thank you for listening. Some social links in case you want to get in contact with me. I just want to make clear the OpenClo side and TexCortex side. So the last part was about the work I do at TexCortex. Just to give a disclaimer. Thank you for listening. I guess...