Open Reader

MCP Tasks (async): Why Aren't Any Agents Supporting Them? — Cornelia Davis, Temporal

completed 23:54 Aug 02, 2026 Watch on YouTube

Current Status

completed

Video ID

s4r6nk5WsZw

RAG / Chat

Enabled
MCP Tasks (async): Why Aren't Any Agents Supporting Them? — Cornelia Davis, Temporal
Description

You invoke a tool and expect an answer, but real work takes time, and over that time connections drop, networks blip, and processes crash. Cornelia Davis, a distributed systems veteran who wrote the book on cloud native patterns, argues that this is exactly the gap the MCP tasks specification exists to close, and walks through why almost no agents support it yet. A task lets a tool run long, report progress, and pause for human input without losing its place, which means the interaction has to be durable: it survives the client disconnecting and picks up right where it left off. She demonstrates it with an invoice processing flow, a dashboard tracking task state, and a step that waits for a human to submit input before the backend continues, then traces how the spec evolved from V1 to V2. The design she keeps returning to is a stateless core with the harder long running behavior layered on as an extension, RPC requests replaced by the server pushing updates, and life cycle state carefully mapped so clients know what to resume. Her honest takeaway is that just because you can open a long lived stateful connection does not mean you should, and that getting durable long running tasks right is what will finally let agents handle work that does not finish in a single call. Speaker info: - https://x.com/cdavisafc - https://www.linkedin.com/in/corneliadavis/ Timestamps: 0:00 - What MCP tasks are, and why they're hard 1:29 - A distributed systems point of view 2:34 - A first look at a task running 4:03 - What a task actually allows 4:43 - Why long running work breaks 6:02 - Durability across disconnections 7:04 - Demo: invoice processing dashboard 9:10 - Waiting for human input 11:18 - What changed in tasks V1 12:35 - The stateless core 16:37 - Extensions and server pushed updates 20:09 - V2 and what you need to implement

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: MCP Tasks are necessary for durable, human-interactive, long-running agent tools, but the original V1 protocol imposed enough client-side state and session complexity that broad agent support was rationally delayed; the forthcoming V2 design is materially cleaner but still leaves persistence and large-scale notification problems to solve.
  • Why it matters: Any agent system that moves beyond short request-response tool calls—especially into approvals, ERP operations, retries, or multistep back-office workflows—needs a durable async control plane rather than treating tool execution as a single model turn.
  • Best use: Use this as an implementation and architecture briefing for deciding whether to build MCP Tasks support now, what durability guarantees to require from clients and servers, and where to rely on a workflow engine such as Temporal.

Executive Summary

Cornelia Davis explains MCP Tasks through a purchase-order workflow in which an agent invokes a long-running invoice-processing MCP tool. The tool validates against an ERP, pauses for human approval, retries payment activity, and eventually returns a result while other purchase-order work proceeds in parallel. The central point is that an async tool cannot merely return a job handle: it must survive disconnected clients, crashed servers, network failures, and delayed human input.

The talk argues that low adoption of MCP Tasks V1 was not a failure of agent vendors but a sensible response to protocol complexity. V1 made recovery depend in part on a stateful task-list endpoint, lacked filtering for task lookup at scale, and tunneled interactive input through an awkward long-lived task-result connection. The reference behavior also created practical concurrency limitations for multiple simultaneous input-required tasks.

V2 moves MCP toward a stateless core with Tasks as an extension. It removes task listing and replaces the long-lived interactive result-session pattern with an explicit client-to-task update endpoint—effectively a durable signal into a workflow. The task lifecycle remains intact: working, input-required, back to working, then completed, canceled, or failed.

However, V2 shifts important responsibility to the client: task IDs must be persisted, because without them a client cannot recover a task after losing local state. It also does not solve fleet-scale observation by itself: polling each of a million tasks with a million clients is untenable. Davis sees the notifications portion of the specification as the likely path to change detection at scale and intends to make an easier implementation available through FastMCP.

Key Takeaways

  • Claim: MCP Tasks address a different class of tool than ordinary MCP request-response calls: durable, asynchronous work that can pause for external input and resume later. | Evidence: The demo invoice tool validates against an ERP, enters an "input required" state for a human approval, resumes after approval is submitted, retries ERP payment operations, and then completes while purchase-order back-office work runs in parallel. | Implication: For agent workflows involving approvals, external systems, or lengthy execution, model-facing tool calls need to create and track a durable work item rather than hold an HTTP/session connection open.
  • Claim: Durability is the defining requirement of MCP Tasks, not merely async execution. | Evidence: Davis states that once launched, a task must not disappear despite network interruptions, disconnected clients, unavailable humans, process crashes, or server outages; in the demo, a purchase order submitted before the client and server were started was retained and later processed. | Implication: Evaluate any Tasks implementation on crash recovery, task identity persistence, retry behavior, and reattachment—not just whether it returns a task handle. | Caveat: The demonstration establishes the desired behavior through her workflow-based implementation, not that every current MCP client/server provides this guarantee.
  • Claim: MCP Tasks V1 was difficult to support because its client protocol was stateful and its interactive semantics depended on a long-running connection. | Evidence: V1 included task get/cancel/list/result operations; task list was used for recovery after client or network loss, while input-required handling was tunneled through task result over a kept-open connection that had to be recovered if it died midstream. | Implication: Do not treat missing V1 support in agents as evidence that long-running MCP tools lack value; it reflects substantial protocol-handler and recovery complexity. | Caveat: V1's lifecycle model itself is not the problem; Davis considers the lifecycle semantics sound and retained in V2.
  • Claim: The V1 task-list recovery mechanism was structurally unsuitable for large agent fleets. | Evidence: The task-list endpoint had no filter, so a client trying to rediscover one task could need to scan a backend containing a million tasks. | Implication: Task discovery should be based on durable, client-owned task identities and targeted lookup, not global server-side enumeration.
  • Claim: MCP Tasks V2 materially improves implementability by making the core protocol stateless and replacing server-driven interactive sessions with explicit client updates. | Evidence: Davis describes V2 as a stateless MCP core with Tasks moved into an extension; task list is removed, task result no longer uses the long session-based pattern, and a client-side update endpoint acts like a Temporal signal into the long-running task. | Implication: If building new support, target the V2 interaction model rather than investing heavily in V1-specific session and task-list machinery. | Caveat: V2 remains "relatively involved" to implement and was described as forthcoming, so teams should verify the final specification and ecosystem support before standardizing on exact APIs.
  • Claim: V2's removal of task listing makes client persistence of task IDs operationally mandatory, even if the specification language only says clients "should" persist them. | Evidence: Davis notes that the specification itself acknowledges that if a client does not persist task IDs, there is no way to retrieve the task later after local state is lost. | Implication: Treat task ID storage as a durable client-control-plane concern: persist it with tenant, agent/run, authorization context, and task-state metadata so agents can resume, cancel, or receive results after restart.
  • Claim: Polling individual tasks remains a major unresolved scaling issue even under the cleaner V2 protocol. | Evidence: Davis gives the example of one million tasks producing one million clients issuing GETs; she identifies the MCP notifications protocol as a promising alternative in which clients ask whether anything changed and only then fetch the affected task. | Implication: For high concurrency, design around event/notification-driven task wakeups and targeted retrieval; do not launch a large agent fleet on per-task polling alone. | Caveat: She had not yet fully implemented or validated the notifications approach at the time of the talk.

Detailed Brief

Workflow-engine implementation pattern

  • Claims: The speaker's client-side task tracker is itself implemented as a workflow, rather than as ephemeral application logic.; A server-side Tasks implementation must map generic MCP task lifecycle states onto the application's domain state machine.
  • Evidence: Her Temporal dashboards show separate durable purchase-order and invoice-processing workflows, with a "task tracker workflow" representing her MCP client implementation.; The invoice domain process contains ERP validation, approval wait states, reconciliation/payment steps, and programmed retries; its domain transitions are mapped to MCP states such as working and input required.
  • Caveats: Temporal is the speaker's chosen implementation substrate, not a normative MCP requirement.; The transcript refers viewers to an earlier MCP Dev Summit presentation for fuller server-side implementation detail.
  • Implications: Separate domain workflow state from MCP protocol state, then explicitly map between them.; A durable workflow runtime is a natural fit when a task requires timers, retries, human signals, recovery after failure, and parallel branches.

Concurrency and ecosystem trajectory

  • Claims: Multiple long-running tasks should be expected in agentic back-office systems, rather than modeled as a single serial interaction.; The V1 reference behavior had a practical FIFO limitation for simultaneous input-required work.
  • Evidence: The second demo submits multiple purchase orders concurrently.; Davis says that, under the V1 reference implementation, when multiple tasks required input, the client could respond only to the first one; her implementation included work to address that gap.; She states an intention to provide a simpler implementation through FastMCP in the near future.
  • Caveats: The FastMCP work was described as planned future work rather than an already released capability.
  • Implications: Require independent task addressing and concurrent human-response handling in any production agent client.; Track FastMCP and the final V2 Tasks extension before deciding whether to build a proprietary client protocol handler.

Notable Concepts & Terms

  • MCP Tasks: An MCP extension for invoking a long-running tool asynchronously, receiving a task handle, tracking lifecycle, delivering results, and supplying later input.
  • Task lifecycle: The protocol-level states—working, input required, working again, then completed/canceled/failed—that let agents reason about durable asynchronous work independently of a specific business domain.
  • Task ID persistence: The client responsibility to durably retain task identifiers so it can resume interaction after restart or disconnection, especially once global task listing is removed.
  • Elicitation / input required: The mechanism by which a running task requests additional information or approval from a client or human; it is the key interaction that made V1 session handling difficult.
  • Temporal signal: A durable message sent into a running Temporal workflow; Davis uses it as the conceptual analogue for V2's explicit task update operation.
  • Stateful vs. stateless protocol: V1 depended on server-side state discovery and long-lived interaction sessions, whereas V2 aims for a stateless core to improve resilience and large-scale operation.
  • FastMCP: The MCP framework used in the demo and the intended target for a simpler Tasks implementation, potentially reducing custom protocol-handler work.
  • Notifications protocol: A prospective mechanism for discovering changed tasks without polling every task individually; it is presented as necessary for million-task-scale operation.

Operator Notes / Why Ken Should Care

  • Set a platform rule that any long-running MCP tool must expose durable task identity, explicit lifecycle state, cancellation, targeted result retrieval, and a restart/reattachment test plan.
  • Persist task records client-side before treating a task invocation as accepted; include task ID, agent/run ID, tenant, invoking principal, tool/server identity, and current state.
  • Avoid adopting MCP Tasks V1-specific task-list or long-lived result-session behavior for net-new systems; validate V2's final extension semantics and client compatibility first.
  • Add a load-design decision for task observation: use notifications or an event-driven wakeup layer when task volumes are high, and prohibit naive per-task polling at fleet scale.
  • Evaluate whether workflow orchestration should be the underlying runtime for approval-heavy or failure-prone agent tools, with domain-state-to-MCP-lifecycle mapping made explicit.
  • Monitor the planned FastMCP implementation and test whether it satisfies concurrent input-required handling, task recovery after client restart, and authorization boundaries around task updates.

Source/Metadata

  • Title: MCP Tasks (async): Why Aren't Any Agents Supporting Them? — Cornelia Davis, Temporal
  • Transcript words: 7086
  • Duration seconds: 1434
  • Timestamp note: No usable timestamps or chapter markers were present in the supplied transcript; the latter portion contains duplicated transcript content.

Transcript

3859 words en Processed in 189.4s

I know it's one minute ahead, but these 20-minute sessions are really short, so I'm going to get started. So the title of my talk you all have seen, because you're all here, which is why the heck aren't any agents supporting MCP tasks? If you don't know what tasks are, don't worry, you will know in just a moment. But the first answer to that question is, well, because they're smart. The people who are building those clients are smart. What I mean by that is that the MCP tasks specification that came out in November was marked as experimental. And so, well, you might shrug and say, well, gosh, those clients and servers, they're all supporting a whole bunch of experimental things. Why not MCP tasks? Well, again, you'll see the answer to that as we move forward. The next answer to that question is, well, they're pretty involved. There's a lot of complexity in here, and that's what I want to do over the next 20 minutes is teach you some of that complexity. Quick intro. My name is Cornelia Davis. I'm a technologist at Temporal. We're distributed systems stuff. I have a long history in distributed systems, did a whole bunch of stuff in the microservices era, including Cloud Foundry, Kubernetes, GitOps, Weaveworks, all of that stuff. And I even wrote a book about that. That's who I am. Today's agenda in the next 19 minutes is that rather than just talking about things in the abstract, I'm going to ground us in a very concrete example. So I'm going to give you the lay of the land of that concrete example. Then I'm going to give you an overview of MCP tasks. Quick question. Who here wants to do things with tasks? Async MCP tools. Okay. So I'm going to give you a little bit of an overview. Then we're going to talk about MCP tasks V1. That's the spec that came out in November. And spoiler alert, there's a new one coming out in July. So that comment that I made about them being smart about not implementing it yet, well, there's some pretty radical changes. So I'm going to show you what's happening with V2. And I actually have some live demos to show all this working. And then we'll have some takeaways at the end. So the use case that we're going to talk about here is a simple purchase order use case. So the use case is you're going to get in a purchase order, and then it's going to go through a number of steps. It's going to record the fact that the goods were received. And then it's going to do, in parallel, some back office stuff, updating inventory, sending out notifications. And then, in parallel to that, it's going to pay some invoices. Now the invoicing is going to happen via an MCP tool. Now that MCP tool has itself a number of steps. So it's going to validate against an ERP. Then it's going to have a little human in the loop to request approval, maybe. Then it's going to reconcile against the ERP again, do a little bit more human in the loop, and so on. So you can see that on the right-hand side, that MCP server, that's going to be a tool that's doing the invoice processing for us, is long-running. It's not going to work in a request-response style. And that's what MCP tasks are all about. And what we're going to do, and today's talk is not about Temporal, but really what I did here was just show you a couple of snippets of the code. And yes, I will be sharing all the code for what I'm showing today. A couple of snippets here. And the real point that I want you to look at is that reject or approve. That is showing you that there is a mechanism for signaling into a long-running process. And that's really the point. And that's what we need, is that this is all about asynchronous. So you understand what MCP tasks are now? MCP tasks are allowing you to have an MCP tool that you can invoke, and then it is long-running in the background, and then eventually you can get back some response. So let's talk about that MCP tasks overview. This is a very simple sequence diagram. It's exactly what you all would expect when I tell you that MCP tasks are long-running tasks. You're going to invoke a tool, and instead of getting back a response, you're going to get a handle. And you can interact with that handle, right? Obvious, right? This isn't rocket science. Looks easy enough, right? Well, it turns out that if you actually want this to work over long horizons, it gets a little bit more complicated than that. So what are some of those complications? Well, you can have all sorts of, the longer something runs, the more likely there's going to be some kind of infrastructure blip that's going to cause a problem in that long-running task. So you could have network blips, you could have network challenges, you could have humans that you're waiting for, they're in a loop part, and they go away on vacation like I'm about to, yay, day after tomorrow. Or processes can crash. So your agent can go down, the agent that's processing the purchase order can go down, or your MCP server can go down as well. So all of those problems you need to deal with, and those are the things that make it a little bit more difficult. Now, in addition to what I've told you about MCP tasks so far, that you're going to get back a handle that you can interact with, by the specification, those MCP tasks can't disappear. This is verbiage from the spec itself that says once you've launched a task, it has to be durable. What that means is all of these things that I just showed you on the previous screen, clients, humans going away on vacation, servers going down, clients going down, connections disconnecting, the task needs to survive that, and you need to be able to interact with that task when the infrastructure comes back. And I'm going to show you how all of that is done. Now, there are server-side elements that talk about how you make the server side durable. And I did a talk at the MCP Dev Summit in March, and this is the QR code that will take you to that YouTube video. And that's where I go into a lot of detail about the server side and what you need to do with the server side. Today, as you saw, is an extension of that work where I'm talking about the client side. So, without further ado, let me go into a demo. For those of you who know me, I'm always doing demos. So, what we have here is we have a dashboard. I am not doing this through a chat interface because, frankly, it's more efficient for me to click a couple of buttons here to show you this rather than trying to type things in. So, I have a user interface here that's showing you the number of purchase orders that have been submitted. I'm going to submit a simple purchase order, so that's just a button that is kicking things off. And in a moment, if the demo gods are with me, it says submitted, we should see the purchase order pop up here, and it should show some, ah, here's why it's not working, because I haven't started my servers. So, remember I said it has to work even when the servers aren't running? I forgot to show you here that what I'm doing in these two windows is, in the upper window, I'm starting the backend. This is the MCP server, and in the lower window, I am starting the MCP client, and you'll see what that client is in a moment. You can see in the splash screen there that I am using Fast MCP on the client side. So, let's go back here and notice that even though I submitted that, even though my servers weren't running, that submission did go through. So, it's captured that. So, what you can see here, and you didn't see it cycle through, but on the far right-hand side, the invoice task initially showed you that it was submitted, then it showed you that it was working, and now it's asking for input required. I can come over here, let me show you what's going on at the backend and at the frontend. What I have here are some dashboards that are showing those running processes. On the right-hand side, you have the backend, that's where the invoice processing is, and you can see the name here. Let me increase the font size there a little bit. So, you can see that this is running the invoice, and on the left-hand side, you can see that it's running the PO. I'll explain that task tracker thing in just a moment. So, if we go into the invoice, we can see that it has the process that we talked about earlier. It validated against the ERP, and now it's waiting for human input. It's waiting for that approval. Over on the PO side, we can also see the process that I showed you earlier, which is to say, let's go back here, it is, so, ah, yes, come over here, let me show you what's going on at the backend and at the frontend. What I have here are some dashboards that are showing those running processes. On the right-hand side, you have the backend, that's where the invoice processing is, and you can see the name here. Let me increase the font size there a little bit. So, you can see that this is running the invoice, and on the left-hand side, you can see that it's running the PO. I'll explain that task tracker thing in just a moment. So, if we go into the invoice, we can see that it has the process that we talked about earlier. It validated against the ERP, and now it's waiting for human input. It's waiting for that approval. Over on the PO side, we can also see the process that I showed you earlier, which is to say, let's go back here, it is, yes, it did that record, it recorded that the goods were received. Then, in parallel, it's invoking the invoice processor, MCP task, and notice that there's this line item here that says task tracker workflow. Yes, indeed, that is my MCP client implementation. Remember, I said nobody's implemented this on the client side? Well, I created my own implementation here. But in parallel with doing the invoice processing, we also had this back office stuff that was happening. So, if I come back over here, and I click on input required, I can approve this, and I'll hit submit, and we come over here, and you'll see in just a moment that the signal is going to come into the backend. I need to refresh. Oh, there it goes. So, the approval came into the backend, and now the backend is going ahead with its additional process, paying the invoice, and you'll see a number of line items there. There's some retries that have been programmed in here, but you can see here that it took a few tries before the ERP went through. We paid the line item, and now you can see that the task completed. So, everything's completed. If I go back to the dashboard that you saw at the top, you can see that all of those processes completed, okay? So, that's the basic stuff. And I can run that again, but I already gave you, inadvertently gave you, the example of the infrastructure was down. I could have killed that server halfway through, and it would have continued exactly as you saw here. Okay? So, you saw it at the very beginning. All right, let's go back to slides. So, that's the first demo. So, let's talk about tasks version one. So, in tasks version one, there were a number of tool semantics. And again, I go over these tool semantics in a lot more detail in that MCP Dev Summit talk. But there's one really interesting thing that I want to draw your attention to, which is that tasks come with, one of the things that the specification defines is a life cycle for tasks. And that's what you see here on the screen. It has working. It has working. It can go into an input required. From input required, it can go back to working. And then eventually, it'll complete or be canceled or fail. So, that's one of the things that's super interesting about the task specification, is that it's about the life cycle of the task. There's a whole bunch of other semantics there as well around obtaining inputs and delivering results. And I'm going to go through this fairly quickly because I already mentioned some of this is going away. So, this is what the tool semantics were before the task semantics. Notice that tools slash call is exactly the same. There's some metadata that you pass in when you want it to be async. And then there's task get cancel list as well as task result. And so, the top four are request-response style. The bottom one keeps a connection open. It keeps a connection alive. And the sequence diagram that you can see here is the basic stuff. Now, there's two major challenges with this particular version of the protocol. The first one is right here: task list. This is a stateful protocol. So, what that means is that, remember I said that the server was responsible for durability? Well, this particular endpoint allows me to go to the server and say, hey, what tasks do you have? So, if the client has gone away, if the user took too long to respond, if my network dropped out and I had to reconnect, I can use this task list to go back to the server and say, what have you got? And then you can continue on with that. That works fine if you have one task or two tasks, or maybe it even works if you have 10 tasks. But what happens if you've got a whole slew of agents out there and you've got a million tasks at the backend? Spoiler alert: there is no filter on that endpoint. So, you would have to go through a million tasks to find the one that you're looking for that you want to interact with. So, this is going away. This is going away. You'll see in just a moment. But that's one of the challenges. Just because you can doesn't mean you should. The other one is the task result because that is where we were tunneling the input required. So, in the case of task result, this sequence diagram is really simple. It doesn't have the interactivity. What we have, as soon as you have input required, is the top and bottom are just fine, but this middle section has this weird protocol where you open a long-running connection and then the server elicits a response from the client. That gets super tricky. And I'm running short on time, so I'm not actually going to show you this demo. Happy to show it to you. I'll be around all day tomorrow too. So, I'm happy to show it to you. But I want to show you instead. Here's the architecture of what you need to build on the server side. Notice that this is using Fast MCP. So, Fast MCP already has support for server side and some client side stuff as well. But the interesting thing is notice that little box on the left-hand side on the lower part where it says MCP client protocol handler. That protocol handler, with the ugliness that I just showed you around results, actually looks like this. And I can show this to you running, and it has all sorts of complexity in it. I got to have the long-running connection. Well, what happens if my connection dies in the middle of that? How do I pick up where I left off when I come back? You'll see that a big part of what the task specification does is it talks about durability. So, back to the question of why the heck aren't there any clients that are supporting this protocol? Yeah, that's why. Super involved. It's still involved with v2, but it gets better. So, let me tell you about that. So, in May, Angie Jones, who's responsible for developer experience at the Agentic AI Foundation, which is where MCP now lives, posted this blog. And one of the things that made me jump up and celebrate a little bit is that the protocol is going stateless. So, as somebody who's been working in the microservices world for a long time, stateful protocols are the absolute worst thing in large-scale distributed systems. So, the protocol is going stateless. It's also doing a number of other things. So, the first bullet is the stateless core. The second bullet is interesting because they also have structured MCP so that there's a core and there are extensions. If some of you were in the room for the previous two talks, they talked about MCPUI. Two talks ago, they mentioned extension. Well, that's what's happening here in the v2 MCP protocol, is that they have extensions. And tasks have become an extension. So, let me tell you a little bit about how tasks changed from v1 to v2, and I do want to give you one more demo. So, on the left-hand side, you can see what the protocol was before. These are the RPC requests that you were doing over the wire. On the right-hand side, you can see a couple of things. Task list has gone away. Good. Wasn't particularly useful anyway, especially at large scale. And instead of having this input required going over a long-running session, you now have an endpoint that allows you from the client side to say, here's an update. So, if you remember a while ago, I showed you that screenshot that said Temporal has this notion of a signal. That's effectively what this is. It's a way of signaling into this long-running task. The task result stays, but it changes because it no longer has this long session-based protocol. But I put the picture on the right-hand side here to emphasize the fact that the lifecycle management of these tasks is unchanged. That's actually sound. Now, I go into this a lot more detail in the talk that I keep referring to. On the server side, in invoice processing, I have my own state machine that the invoice is going through. Task list has gone away. Good. Wasn't particularly useful anyway, especially at large scale. And instead of having this input required going over a long-running session, you now have an endpoint that allows you from the client side to say, here's an update. So, if you remember a while ago, I showed you that screenshot that said Temporal has this notion of a signal. That's effectively what this is. It's a way of signaling into this long-running task. The task result stays, but it changes because it no longer has this long session-based protocol. But I put the picture on the right-hand side here to emphasize the fact that the lifecycle management of these tasks is unchanged. That's actually sound. Now, I go into this in more detail in the talk that I keep referring to. On the server side, in invoice processing, I have my own state machine that the invoice is going through. And so, part of what you're doing when you implement these server-side tasks is you're mapping from the lifecycle states of the task over to the domain state machine that's running the application, the MCP server in the backend, or the tool. So, list again goes away. Now, remember I said that the MCP task specification has durability all over it? With this change, given that lists are gone, you now are required on the client side, required, there's a little parenthetical remark here. The spec right now says that clients should persist task IDs. But it also points out that if you don't persist task IDs, there is no way to get it back. So, I'm not quite sure why this doesn't have an all-caps must. The other thing that I want to point out, and I already mentioned it, is that you're going to have potentially a lot of agents that are processing POs or a lot of agents that are doing a lot of things. And so, having multiple things running, I think, is really crucial as well. So, with that, I'm going to go to the second demo. And I'm going to go back to my purchase order here. So, what I'm going to do now is I'm going to submit a number of things. And I'm actually still demoing here because I have 13 seconds left. I'm not going to switch over to my V2. You'll see that from the high level, it actually looks exactly the same. I am going to show you what the client-server protocol looks like in the V1 case. It's really quite ugly. But you'll notice here that I've submitted a bunch of different ones. I can tell you with the V1 protocol, the reference implementation, if you had input required on multiple, even though you can see that there's many of them in flight, on the client side there were FIFO. So, you can only respond to the first one. And part of the protocol that I implemented was to get around that gap. So, let's come over here. We can refresh both of these, and you can see that there's going to be a bunch of POs in flight. And now I want to show you the task tracker. So, if we go into the task tracker, that's the MCP client. And now let me just expand this so we can see it in a little bit more detail. What you can see here is that, remember that protocol, I showed you that big long sequence diagram, there's a lot of steps involved in that. And what I've done here is I've implemented it as a workflow. And you can see here that there's some elicitation handling that's going from the server side back to the client. So, I won't go into any more details because I'm literally out of time now, but I want to share two more things. And that is, going from V1, remember this ugly picture, to V2 in the client-server protocol, much, much cleaner. Much easier to implement. So, speaking of implementing, here's a summary of all the things that you need to do if you want to implement tasks. Still relatively involved. Here's a picture. I'm going to make these slides available in the Git repo that I'm about to show you. And here's the Git repo that I'm about to show you. And while you're getting that screenshot, I'm going to tell you about two pieces of work that I'm continuing with. Number one, even though this is better, it still doesn't scale to the millions. Why? Because if I've got a million tasks running, I've got a million clients that are doing gets against each and every one of those tasks. That does not scale. There is a part of the MCP task specification that is a notifications protocol, which I haven't gotten far enough into yet, but it's showing promise. Which is going to allow you to, instead of having a million clients polling their tasks, have a single endpoint where they can say, has something changed? And if it has, tell me which one, and now I'll go pull that task. So it's definitely from a scale perspective. The other thing that we're doing is, in the very near future, in the next month or two, we're going to have an implementation of all of this where it's going to be much simpler for you. My goal is to actually implement it in Fast MCP so that you can use the same protocol, the same framework that you're using probably for your MCP servers today. So without further ado, that is it. Thank you to the next speaker for letting me go a few minutes long, and I'll be around, I'll step out. If you have any questions, find me in the hallway. Thank you. down, clients going down, connections disconnecting, the task needs to survive that, and you need to be able to interact with that task when the infrastructure comes back. And I'm going to show you how all of that is done. Now, there's elements, there's server side elements that talk about how you make the server side durable. And I did a talk at the MCP Dev Summit in March, and this is the QR code that will take you to that YouTube video. And that's where I go into a lot of detail about the server side and what you need to do with the server side. Today, as you saw, is an extension of that work where I'm talking about the client side. So, without further ado, let me go into a demo. For those of you who know me, I'm always doing demos. So, what we have here is we have a dashboard. I am not doing this through a chat interface because, frankly, it's more efficient for me to click a couple of buttons here to show you this rather than trying to type things in. So, I have a user interface here that's showing you the number of purchase orders that have been submitted. I'm going to submit a simple purchase order, so that's just a button that is kicking things off. And in a moment, if the demo gods are with me, it says submitted, we should see the purchase order pop up here, and it should show some, ah, here's why it's not working, because I haven't started my servers. So, remember I said it has to work even when the servers aren't running? I forgot to show you here that what I'm doing in these two windows is in the upper window, I'm starting the backend. This is the MCP server, and in the lower window, I am starting the MCP client, and you'll see what that client is in a moment. You can see in the splash screen there that I am using Fast MCP on the client side. So, let's go back here and notice that even though I submitted that, even though my servers weren't running, that submission did go through. So, it's captured that. So, what you can see here, and you didn't see it cycle through, but on the far right hand side, the invoice task initially showed you that it was submitted, then it showed you that it was working, and now it's asking for input required. I can come over here, let me show you what's going on at the backend and at the frontend. What I have here are some dashboards that are showing those running processes. On the right hand side, you have the backend, that's where the invoice processing is, and you can see the name here. Let me increase the font size there a little bit. So, you can see that this is running the invoice, and on the left hand side, you can see that it's running the PO. I'll explain that task tracker thing in just a moment. So, if we go into the invoice, we can see that it has the process that we talked about earlier. It validated against the ERP, and now it's waiting for human input. It's waiting for that approval. Over on the PO side, we can also see the process that I showed you earlier, which is to say, let's go back here, it is, so, ah, yes, so it did that record, it recorded that the goods were received. Then, in parallel, it's invoking the invoice processor, MCP task, and notice that there's this line item here that says task tracker workflow. Yes, indeed, that is my MCP client implementation. Remember, I said, nobody's implemented this on the client side? Well, I created my own implementation here. But in parallel with doing the invoice processing, we also had this back office stuff that was happening. So, if I come back over here, and I click on input required, I can approve this, and I'll hit submit, and we come over here, and you'll see in just a moment that the signal is going to come into the back end. Ah, I need to refresh. Oh, there it goes. So, the approval came into the back end, and now the back end is going ahead with its additional process paying the invoice, and you'll see a number of line items there. There's some retries that have been programmed in here, but you can see here that it took a few tries before the ERP went through. We paid the line item, and now you can see that the task completed. So, everything's completed. If I go back to the dashboard that you saw at the top, you can see that all of those processes completed, okay? So, that's the basic stuff. And I can run that again, but in the, I already gave you, inadvertently gave you the example of, the infrastructure was down. I could have killed that server halfway through, and it would have continued exactly as you saw here. Okay? So, you saw it at the very beginning. All right, let's go back to slides. So, that's the first demo. So, let's talk about tasks version one. So, in tasks version one, there were a number of tool semantics. And again, I go over these tool semantics in a lot more detail in that MCP Dev Summit talk. But there's one really interesting thing that I want to draw your attention to, which is that tasks come with, one of the things that the specification defines is a life cycle for tasks. And that's what you see here on the screen. It has working. It has working. It can go into an input required. From input required, it can go back to working. And then eventually, it'll complete or be canceled or fail. So, that's one of the things that's super interesting about the task specification is that it's about the life cycle of the task. There's a whole bunch of other semantics there as well around obtaining inputs and delivering results. And I'm going to go through this fairly quickly because I already mentioned, some of this is going away. So, this is what the tool semantics were before the task semantics. Notice that tools slash call is exactly the same. There's some metadata that you pass in when you want it to be async. And then there's task get cancel list as well as task result. And so, the top four are request response and style. The bottom one keeps a connection open. It keeps a connection alive. And the sequence diagram that you can see here is kind of the basic stuff. Now, there's two major challenges with this particular version of the protocol. The first one is right here. Task list. This is a stateful protocol. So, what that means is that the, remember I said that the server was responsible for durability? Well, this particular endpoint allows me to go to the server and say, hey, what tasks do you have? So, if I have had, if the client has gone away, if the user took too long to respond, if my network dropped out and I had to reconnect, I can use this task list to go back to the server and say, what have you got? And then you can continue on with that. That works fine if you have one task or two tasks or maybe it even works if you have 10 tasks. But what happens if you've got a whole slew of agents out there and you've got a million tasks at the backend? Spoiler alert, there is no filter on that endpoint. So, you would have to go through a million tasks to find the one that you're looking for that you want to interact with. So, this is going away. This is going away. You'll see in just a moment. But that's one of the challenges. Just because you can doesn't mean you should. The other one is the task result because that is where we were tunneling the input required. So, in the case of task result, this sequence diagram is really simple. It doesn't have the interactivity. What we have as soon as you have input required is the top and bottom are just fine, but this middle section has this weird protocol where you open a long running connection and then the server elicits a response from the client. That gets super tricky. And I'm running short on time, so I'm not actually going to show you this demo. Happy to show it to you. I'll be around all day tomorrow too. So, I'm happy to show it to you. But I want to show you instead. Here's basically the architecture of what you need to build on the server side. Notice that this is using Fast MCP. So, Fast MCP already has support for server side and some client side stuff as well. But the interesting thing is notice that little box on the left hand side on the lower part where it says MCP client protocol handler. That protocol handler with the ugliness that I just showed you around results actually looks like this. And I can show this to you running and it has all sorts of complexity in it. I got to have the long running connection. Well, what happens if my connection dies in the middle of that? How do I pick up where I left off when I come back? You'll see that a big part of what the task specification does is it talks about durability. So, back to the question of why the heck aren't there any clients that are supporting this protocol? Yeah, that's why. Super involved. It's still involved with v2, but it gets better. So, let me tell you about that. So, in May, Angie Jones, who's responsible for developer experience at the Agentic AI Foundation, which is where MCP now lives, posted this blog. And one of the things that made me jump up and celebrate a little bit is that the protocol is going stateless. So, as somebody who's been working in the microservices world for a long time, stateful protocols are the absolute worst thing in large scale distributed systems. So, the protocol is going stateless. So, the protocol is going stateless. It's also doing a number of other things. So, the first bullet is the stateless core. The second bullet is interesting because they also have structured MCP so that there's a core and there's extensions. If some of you were in the room for the previous two talks, they talked about MCPUI. Two talks ago, they mentioned extension. Well, that's what's happening here in the v2 MCP protocol, is that they have extensions. And tasks have become an extension. So, let me tell you a little bit about how tasks changed from v1 to v2 and I do want to give you one more demo. So, on the left hand side, you can see what the protocol was before. These are the RPC requests that you were doing over the wire. On the right hand side, you can see a couple of things. Task list has gone away. Good. Wasn't particularly useful anyway, especially at large scale. And instead of having this input required going over a long running session, you now have an endpoint that allows you from the client side to say, here's an update. So, if you remember a while ago, I showed you that screenshot that said, temporal has this notion of a signal. That's effectively what this is. It's a way of signaling into this long running task. The task result stays, but it changes because it no longer has this long session based protocol. But I put the picture on the right hand side here to emphasize the fact that the lifecycle management of these tasks is unchanged. That's actually sound. Now, I go into this a lot into more detail in the talk that I keep referring to. On the server side, in invoice processing, I have my own state machine that the invoice is going through. And so, part of what you're doing when you implement these server side, these tasks, is you're mapping from the lifecycle states of the task over to the domain state machine that's running the application, the MCP server in the backend or the tool. So, list again goes away. Now, remember I said that the MCP task specification has durability all over it? With this change, given that lists are gone, you now are required on the client side, well, kind of required, there's a little parenthetical remark here. The spec right now says that clients should persist task IDs. But it also points out that if you don't persist task IDs, there is no way to get it back. So, I'm not quite sure why this doesn't have an all caps must. The other thing that I want to point out is that I already mentioned it, is that you're going to have potentially a lot of agents that are processing POs or a lot of agents that are doing a lot of things. And so, having multiple things running, I think is really crucial as well. So, with that, I'm going to go to the second demo. And I'm going to go back to my purchase order here. So, what I'm going to do now is I'm going to submit a number of things. And I'm actually still demoing here because I have 13 seconds left. I'm not going to switch over to my V2. You'll see that from the high level, it actually looks exactly the same. I am going to show you what the client server protocol looks like in the V1 case. It's really quite ugly. But you'll notice here that we have, I've submitted a bunch of different ones. I can tell you with the V1 protocol, the reference implementation, if you had input required on multiple, even though you can see that there's many of them in flight, on the client side there were FIFO. So, you can only respond to the first one. And part of the protocol that I implemented was to get around that gap. So, let's come over here. We can refresh both of these and you can see that there's going to be a bunch of POs in flight. And now I want to show you the task tracker. So, if we go into the task tracker, that's the MCP client. And now let me just expand this so we can see it in a little bit more detail. What you can see here is that, remember that that protocol, I showed you that big long sequence diagram, there's a lot of steps involved in that. And what I've done here is I've implemented it as a workflow. And you can see here that there's some elicitation handling that's going from the server side back to the client. So, I won't go into any more details because I'm literally out of time now, but I want to share two more things. And that is, so going from V1, remember this ugly picture, to V2 in the client server protocol, much, much cleaner. Much easier to implement. So, speaking of implementing, here's a summary of all the things that you need to do if you want to implement tasks. Still relatively involved, here's a picture. I'm going to make these slides available in the Git repo that I'm about to show you. And here's the Git repo that I'm about to show you. And while you're getting that screenshot, I'm going to tell you about two pieces of work that I'm continuing with. Number one, even though this is better, it still doesn't scale to the millions. Why? Because if I've got a million tasks running, I've got a million clients that are doing gets against each and every one of those tasks. That does not scale. There is a part of the MCP task specification that is a notifications protocol, which I haven't gotten far enough yet, but it's showing promise. Which is going to allow you to, instead of having a million clients polling their tasks, it's going to have a single endpoint where they can say, has something changed? And if it has, tell me which one, and now I'll go pull that task. So it's definitely from a scale perspective. The other thing that we're doing is in the very near future, in the next month or two, we're going to have an implementation of all of this where it's going to be much simpler for you. My goal is to actually implement it in Fast MCP so that you can use the same protocol, the same framework that you're using probably for your MCP servers today. So without further ado, that is it. Thank you to the next speaker for letting me go a few minutes long, and I'll be around, I'll step out. If you have any questions, find me in the hallway. Thank you.