[SPEAKER_00] Alex Hancock Hey everybody. My name is Alex Hancock. Today I'm going to talk about a universal remote control for AI. Before I start, I just want to say the previous speaker said that MCP client maintainers haven't implemented support for tasks because they're smart. I'm an MCP client maintainer. I can tell you it's just because I'm lazy. I haven't done it.
Okay, so a little bit about me before we start. I am a software engineer at Block, which is the parent company of Cash App and Square and Title. We have a few different things going on now. I've worked there for a long time. I worked on Square product stuff and Cash App stuff, but I've been doing open source AI for the last couple years. Specifically, I work on this open source harness project called Goose, which started as an internal project at Block. Yeah, some Goose fans out there. And then, yeah, we open sourced it and we donated it to the Linux Foundation. So now the IP is there, but lots of us from Block still work on it. I'm also a maintainer of MCP, the Model Context Protocol. I work on the Rust SDK for that project. And more recently, I've also started some work on ACP, the Agent Client Protocol, which is what I'm going to talk about today.
So I think we have an issue with harnesses that I want to propose to you all today as a problem, and then recommend a solution. What I've been noticing recently is that we've got lots of great harnesses out there. There are ones from the labs, ones from different companies. There's lots of open standards-based ones. But I noticed that the interface to them is often custom or bespoke, and in the worst case, you might have some harnesses where there's literally only one client application you can use to control that harness.
I think this has a couple issues with it, but the analogy I'll make with the web is it would be like if you had to use one browser or one given protocol to connect to every website. That just wouldn't work. You wouldn't have something like the open web if that were the reality with browsers. So I think we can do better. The thing about standards is that they create ecosystems and markets. I would argue that in the agentic AI space, we have a good standard for the agent going out and doing things—calling tools, taking actions, and other systems reading resources, reading data. We've all benefited as a community from having MCP. The most powerful thing about MCP is not anything about MCP itself, but it's that everyone uses MCP. That's why we have thousands or tens of thousands of servers around the world and all the agents can connect to them and do things in those other systems.
I would say that we don't yet have a good solution or a standard for client software to tell agents what to do, giving it tasks, telling it what to work on, and getting updates. So I'm going to put forward an option today that I think is a good option, that we on our team have been working on, and we think is a good solution in the open standard space. This is ACP. Agent Client Protocol is the name of this project. It came from editor companies. It came from the Zed text editor folks and JetBrains folks. The Zed folks and JetBrains folks teamed up and proposed a standard for clients to be able to control harnesses. It makes sense if you put yourself in their shoes. What they wanted to be able to do is write a single high-quality client implementation in an editor, maybe in Zed or in IntelliJ or something like that, and be able to control any harness with that single client implementation, sending tasks, getting results back, seeing what files are being edited, etc. It makes a ton of sense if you put yourself in their shoes.
But we saw this on the Goose team, and we think that there is a much broader utility than just editors. It's relatively neutral and it doesn't have many editor-specific features. So we think that this can be spread to a wider range of client software.
To go into a little bit more depth about ACP's design and what you can do with it: It lets you establish connections between clients and agent harnesses that have a given set of capabilities associated with the connection. And then you can make sessions. Within sessions, you can send user messages—the things that a user is typing into the app or that the client software wants to send. The agent can then respond to those with text, images, audio, etc., or updates about what's going on. For example, if a tool is called, it can send a tool call notification and explain what tool was called and what the metadata was. And it can also send things like permission requests, so if the client software needs to show the user "Should I do this tool call? Yes or no?" it can go over this protocol. It's pretty simple in its design. It uses JSON RPC messages. The thing we like about it most is that it's extensible. You're not limited to just what's in the vanilla protocol; you can add custom methods. The convention is you put an underscore and then you put your custom methods. The thing I like about this is that if enough harness projects or client projects adopt this, we can start to see what we're all doing that's the same. If the Codex team has some custom methods, the Goose team has some custom methods, the client team has some custom methods—whoever—we can see what emerges in the ecosystem and what makes sense to get on a standards track and bring into the protocol itself. So this is shaped by usage and shaped by the community.
I'm going to do a demo of a standard I/O version of this. So I'm going to open Zed. I just have a really simple project here. I'll say "tell me about this project." This is a single HTML file. You can see I was able to type my query into Zed. The agent in play here is Goose. So it's using Goose's ACP interface. It's sending text back. It's sending tool call information back about what it read and what it did. And then it found that it's a single HTML file and explained it. I'll do another one. This is one from a company called Poolside AI. I'll say "tell me about this project" in the same project. This is a terminal-based client getting exactly the same experience from the same agent. One implementation on the harness side, and you can now use any client. You can see it did the same thing. It showed me some text results back. It showed a tool call. And then it showed it streaming in a summary.
So that's a basic demo showing two clients talking to the same agent over standard I/O locally in this case. But local obviously isn't enough. If you want this to take off, you have to be able to do remote. Agents are going to be running in the cloud. So when we came to this project, we saw that it did not have remote support yet. So we specified an HTTP transport. There's an HTTP version and a WebSocket upgrade. The messages are the same. The protocol semantics are the same. But there's a new transport that is just landing now that enables remote.
The way we think about this on the Goose team, the agentic stack has these four important components: You have the client, which is the app that the user is using or a headless app running somewhere on the machine. There's the harness, which is the program that implements the tool calling loop. There are the tools themselves, often MCP. And then there's the model.
If you do a remote transport for the Agent Client Protocol, and MCP has remote transport for tool calling, and the models have all had remote endpoints like OpenAI APIs for a long time now, you have the flexibility to move all of these four components around. They could all be on the same machine. The harness could be on a different machine than the client. The model could be the only thing that's remote. The tools could be the only thing that's remote. Aligning on standards and making sure that they have good transport stories is what's going to let us move all the pieces of this agentic stack around.
I can show a quick demo of this as well. This is a client just to show how easy it is to create clients for this. I just vibe coded this last night. I'll say "write a poem." This is again connecting to that same process on my machine. In this case, I'm running it over the network, but it's on my machine. It's connecting and sending Goose instructions for what to do remotely. So this could be in a container. It could be up in the cloud. But the messages are the same and the library you use is the same, so you can just switch between local and remote very easily.
So if you want to get plugged into this ecosystem, start experimenting with support, either making your own clients or adding stuff to harnesses. This will link you to the Agent Client Protocol site for how to get started. There's a number of clients and agent servers already out there. This ranges from editors, desktop applications, mobile applications, terminal-based things. There's a proliferation.
I think the use cases are potentially huge if we get some interoperability going here. You can have people make personal clients that are exactly how you want them, orchestrating your agents. You could have clients created for certain business domains or an individual company. You could customize a white-label client and have it work with all the harnesses. I also think if we make a new category here, we're going to see the quality of the clients go up. Any time you get an ecosystem or a marketplace going with many options, users can vote with their feet. If clients aren't meeting their needs, people will start to compete on the quality of the user experience. Overall, I think this should drive up the user experience of using AI.
That's what I've got today. Thank you very much. If you want to chat with me, find me after or send me an email. Happy to get you plugged into this work. Thank you. So I think I think we have an issue with harnesses that I want to I want to try to put to you all today Propose to you all today as a problem and then and then recommend a solution so What I've been noticing recently? Is that we've got lots of great harnesses out there right there are ones from the labs there one from ones from different companies There's lots of open standards based ones But I noticed that the interface to them is often custom or bespoke and in the worst case
It's like you might have some harnesses where there's literally only one client application you can use to control that harness right and I think this has a couple issues with it, but the analogy that I'll make with the web is it would be like if you had to use one browser or a One given protocol to connect to a to every website, right? That just wouldn't work You wouldn't have something like the open web if if that were the reality With browsers and so I think we can do better and the thing about standards by finding a standard and the thing about standards is that they create ecosystems and markets and I would argue that in the agentic AI space we have a
Good standard for the agent going out and doing things right calling tools taking actions and other systems reading resources reading data We've all benefited as a community from having MCP Right and the most powerful thing about MCP is not anything about MCP itself, but it's that everyone uses MCP and That's why we have you know thousands or tens of thousands of servers around the world and all the agents can connect to them and Go and do things in those other systems. I would say that we don't yet have a good solution or a standard for for client software to tell agents what to do giving it tasks telling it what to work on and getting updates and
so I'm going to put forward an option today that I think is a good option that that we on our team Have been working on and we think is a good a good solution in the open standard space and This is ACP so agent client protocol is the name of this project and it came from the editor companies It came from like if you've used the Zed text editor or you've used any of JetBrains products the Zed folks in the JetBrains JetBrains folks teamed up and proposed the standard for clients to be able to control harnesses and it makes sense if you put yourself in their shoes, right? What they wanted to be able to do is write a single high quality client implementation in an editor
Maybe in Zed or in IntelliJ or something like that and be able to control any harness by with that single client implementation Sending tasks getting results back Seeing what files are being edited etc. It makes a ton of sense if you put yourself in their shoes, right? But we saw this on the goose team and we think that there is a much broader utility than just editors, right? So it's a it's it's relatively neutral and it doesn't have many editor specific features And so we think that this can can go to a be spread to a wider range of client software To go into a little bit more depth about ACP's design and what you can do with it
It lets you establish connections between clients and agent harnesses that have a given Set of capabilities associated with the connection and then you can make sessions Within sessions you can send user messages the things that a user is maybe typing into the app or that the client software wants to send the agent can then respond to those with text More images or audio text etc or updates about what's going on So like if a tool is called it can send a tool call notification and explain what tool was called and what the metadata was And it can also send things like permission requests so that if the client software needs to show the user
You know should I do this tool call yes or no it can go over this This protocol and it's it's pretty simple and its design It uses json RPC messages and the thing we like about it most is that it's extensible as well So you're not limited to just what's in the vanilla protocol you can add custom methods So the pre the convention is you put an underscore and then you start to put your custom methods And the thing I like about this is that if enough harness projects or client projects adopt this We can start to see what we're all doing. That's the same right Like if the codex team has some custom methods the goose team has some custom methods
The client team has some custom methods whoever we can see what emerges in the in the ecosystem And what makes sense to get on a standards track and bring into the protocol itself so that this is sort of shaped by usage and shaped by the community I'm going to do a demo of a standard I O version of this So I'm going to open zed and I just have a really simple project here Where I'll say tell me about this project and so this is a single html file so you can see I was able to type my query into zed and this is the agent in play here is goose So it's using goose's acp interface and it's you can see it's like sending text back. It's sending tool call information back
About what it read and what it did and then it found you know that it's a single html file and explained it and I'll do another I'll do another one. This is one from a company called poolside AI I'll say tell me about this project in the same project and so this is a terminal based client Getting exactly the same experience from the same agent one implementation on the harness side and you can now use any client Right and so you can see it did the same thing it showed me Some text results back it showed a tool call and then it showed it. It's streaming in a summary So that's a basic demo showing two clients talking to the same agent Over standard IO locally in this case
But local obviously isn't enough right if you want this to take off you have to be able to do remote as well Agents are going to be running in the cloud and so when we came to this project We saw that it did not have remote support yet. So we specified an HTTP transport There's an HTTP version and there's a web socket upgrade and so now the messages are the same the protocol semantics are the same But there's a new transport that is just landing now that enables remote and The way we think about this on the goose team the agentic stack is there's sort of these four important components, right? You have the client
Which is like the app that the user is using or a headless app running somewhere on the machine? There's the harness which is the program that implements the tool calling loop There are the tools themselves as often MCP and then there's the model, right? and if you do a remote transport for the agent client protocol and MCP has remote transport for tool calling and And the models have kind of all had remote endpoints like responses apis for a long time now you have the flexibility to move All of these four components around they could all be on the same machine the harness could be on a different machine than the client
The model could be the only thing that's remote the tools could be the only thing that's remote Aligning on standards and making sure that they have good transport stories is what's going to let us move all the pieces of this agentic stack around And I can show a quick demo of this as well So this is a client just to show how easy it is to create clients for this I just vibe coded this you know last night and I'll say write a poem So this is again connecting to that same process on my machine In this case, I'm running it over the network, but it's on my machine It's connecting and sending goose instructions for what to do
Remotely so this could be in a container it could be up in the cloud But the messages are the same and the library you use is the same so you can just switch between local and remote very very easily So if you want to get plugged into this ecosystem start experimenting with support either making your own clients or adding stuff to harnesses This is a this will link you to the agent client protocol site for How to get started there's a number of clients and and agent servers already out there This ranges from editors desktop applications mobile applications terminal based things like there's a proliferation and I think the use cases are
Are potentially huge right if we if we get some interoperability going here because you can have people can make personal clients That's exactly how you want it orchestrating your agents You could have sort of clients created for certain business domains or an individual company or A set of clients from a company you could customize like a white label Client and have it work with all the harnesses and I also think if we make a new category here We're going to see quality of the clients go up right because any any time you get an ecosystem or a marketplace going and there's many options Users can vote with their feet if clients aren't meeting their needs
And so people will start to compete on the quality of the user experience and and like overall I think this should drive up The user experience of using AI That's what I've got today. Thank you very much And if you want to chat with me find me after or send me an email Happy to get you plugged into this work Thank you Thank you