Every step you take, every call you make: the reliable agent stack — Giselle van Dongen, Restate
Description
A research agent asks a human to approve its plan, then waits. The wait might last a month, spanning restarts and redeploys, and while it waits the function consumes no serverless execution time at all. Giselle van Dongen builds Restate, and that suspended promise is the clearest illustration of what durable execution buys you. Her talk skips the agent loop itself and goes after the infrastructure underneath it, the part she argues teams keep rebuilding badly: retry logic, recovery logic, session isolation, and stopping something already in flight. Restate runs as a server in front of your agent service, proxying requests and holding an open connection that acts as a lifeline. As the agent works it emits events to a journal, and that journal replays the process back to its exact failure point rather than starting over. The design borrows from Apache Flink and Meta's event infrastructure, and it pushes invocations rather than polling, which is where the low latency comes from. The demo is a deep research agent living in Slack. A planner proposes subtopics, parallel sub agents go to work, and an injected tool failure shows a web search retrying and completing instead of sinking the run. Then it gets more interesting. Because a session is modeled as a virtual object with its own key and isolated state, van Dongen can talk to a run already in progress. She tells it midflight to focus on frontier models and a classifier decides that is relevant, so it signals the live loop. She then tells it to research something else entirely, and the cancellation rewinds down the call chain, killing sub agents before the controller. Speaker info: - https://x.com/vdgiselle - https://www.linkedin.com/in/giselle-van-dongen/ Timestamps: 0:00 - Three waves, and where agents are heading 1:53 - The infrastructure layer nobody wants to write 2:45 - Four ingredients of a durable foundation 4:29 - How the journal recovers a failed run 5:20 - Demo: a deep research agent in Slack 7:51 - Making
Summary
Generated by gpt-5.6-terraAt-a-Glance
- Verdict: Watch fully
- Core thesis: Production agents need a durable execution and stateful coordination layer—not merely an agent framework—so they can recover precisely from failures, manage concurrent sessions, accept mid-run input, and be cancelled or governed safely.
- Why it matters: As agents become persistent, asynchronous systems connected to organizational tools and other agents, reliability, session isolation, human approvals, cancellation, and centralized policy/rate control become core control-plane requirements.
- Best use: Use this as an architecture reference for evaluating or designing a durable agent runtime, particularly for long-running OpenClaw-style workflows, human-in-the-loop operations, and shared model/tool gateways.
Executive Summary
Giselle van Dongen argues that agent infrastructure must evolve beyond the SDK-centric approach that works for prototypes. Her premise, framed through Karpathy's progression from chat interfaces to tool-using agents to persistent asynchronous entities, is that production agents are distributed, stateful backend processes. Consequently, teams must solve failure recovery, retries, state consistency, inter-agent communication, and operational control rather than treat an agent as a single request-response function.
The talk presents Restate, an open-source durable execution platform, as the infrastructure layer for those concerns. Restate sits in front of an agent service, journals events exchanged during execution, and uses that journal to replay an agent from its last durable step after failure rather than restarting the entire run. The demo uses a Slack deep-research agent: a planner proposes a research plan, waits for user approval, launches parallel researchers, and ultimately produces a report. An injected web-search outage is retried and recovered without losing earlier progress.
The more distinctive design point is Restate's virtual object model for persistent agent sessions. A session has a unique ID, isolated key-value state, and serialized execution, allowing a user to add relevant context to an active run or cancel an obsolete run and propagate cancellation down to its sub-agents. This addresses two common agent-operating problems: concurrent messages corrupting shared session history, and a user needing to redirect work before a lengthy task completes.
Finally, the speaker shows how durable primitives can support a centralized LLM gateway rather than leaving calls inline in every agent. That gateway can enforce policy checks and per-department flow control—for example, capping one department at 300 simultaneous calls. The presentation is a strong architectural demonstration, although it is product-led: it asserts operational and latency benefits without comparing trade-offs, migration costs, or alternatives such as Temporal, cloud workflow services, or actor systems.
Key Takeaways
- Claim: Agent frameworks and memory layers are useful for prototypes, but production agent platforms require an additional distributed-systems layer for durability, recovery, coordination, and control. | Evidence: The speaker contrasts agent SDKs with the infrastructure teams otherwise must build themselves: deployment infrastructure, retry logic, recovery logic, and mechanisms to run long-lived stateful distributed processes. | Implication: Separate agent quality/evaluation from runtime reliability in the stack; do not assume an agent framework's tool and memory abstractions provide production-grade failure semantics. | Caveat: The talk deliberately excludes agent evaluation quality; Restate addresses execution reliability and orchestration, not whether an agent's reasoning or outputs are correct.
- Claim: Durable execution should recover an agent at its last completed durable step, even after a failure long after the workflow began, instead of rerunning the whole task. | Evidence: Restate journals events sent from the agent service; the speaker describes wrapping a Python LLM call in Restate.run so a failure 'two hours or two months later' can resume at that point. In the demo, a web-search API failure is retried and the sub-agent completes without restarting prior LLM and search work. | Implication: For expensive, long-running, or side-effecting agent work, make tool calls, model calls, waits, and state transitions explicit durable boundaries so recovery does not duplicate work or lose progress.
- Claim: Durability is also a practical human-in-the-loop primitive because an agent can suspend cheaply while awaiting an approval that may take weeks. | Evidence: The Slack research demo creates a durable promise, asks a human to approve the plan, suspends the process, and resumes when the Slack response arrives; the speaker notes that a suspended serverless function does not consume function execution time. | Implication: Model approvals as persisted, resumable state transitions rather than as long-held processes, and separately define identity, timeout, and escalation controls for those approvals. | Caveat: Durable waiting solves runtime continuity, but the transcript does not address approval authorization, escalation, expiration, or audit-policy design.
- Claim: Persistent agents are better modeled as session-scoped stateful actors than as one-shot workflows when users need to interact with work already in progress. | Evidence: Restate's 'virtual object' has a unique session ID, isolated key-value state such as chat history, and handlers that run durable functions. In the demo, 'focus on frontier models' is classified as relevant and signaled into the active research loop. | Implication: Give each durable agent session a stable identity and isolated state store; this enables follow-up context, memory, and control operations without waiting for a full run to finish.
- Claim: Concurrency isolation and hierarchical cancellation are required control-plane behaviors for multi-session and multi-agent systems. | Evidence: The speaker says Restate permits only one execution at a time for a given virtual object, queuing a second execution so simultaneous Slack messages do not overwrite session state. When the user changes the topic to AI policy, the coordinator cancels the active run and cancellation propagates down the call chain to spawned sub-agents. | Implication: Define concurrency scopes explicitly—typically per user, session, tenant, or resource—and require cancellation propagation across child agents and tools to stop obsolete spend and side effects. | Caveat: Serializing execution per session prevents state races but can introduce queuing and responsiveness trade-offs for high-frequency interactions within the same session.
- Claim: A distributed LLM gateway provides a clean place to add policy enforcement and capacity controls as agent deployments scale. | Evidence: The presenter moves an inline LLM call into a dedicated handler that performs a policy check before calling the model; other agents call it through Restate communication primitives. She gives an example limit of 300 concurrent LLM-gateway calls for one department. | Implication: Centralize shared LLM access where possible so authorization, routing, rate limits, spend controls, observability, and provider changes are managed once rather than embedded independently in every agent. | Caveat: The presentation demonstrates policy and concurrency controls conceptually but does not specify model-routing policies, identity propagation, cost accounting, or security enforcement mechanics.
- Claim: Restate's architecture aims to provide workflow-grade reliability with low latency and serverless compatibility through push-based invocation rather than orchestrator polling. | Evidence: The platform persists journal events in a distributed log, with an event loop that writes state, sets timers, or invokes other agents. The speaker claims a push model enables 45 ms p99 for a 10-step workflow, calls it suitable for waking serverless functions, and describes high availability through multiple instances that snapshot to object storage. | Implication: Benchmark durability platforms against the actual workload mix—short synchronous paths, long waits, fan-out, tool retries, and throughput—not only against their headline latency claims. | Caveat: The 45 ms p99 figure is a vendor-presented performance claim with no workload definition, deployment configuration, benchmark methodology, or comparison baseline in the transcript.
Detailed Brief
Restate programming and operational model
- Claims: The primary application interface is an HTTP handler augmented with a Restate SDK context.; Actions invoked through that context become journaled events, allowing ordinary application functions to gain long-running, durable, stateful behavior.; The platform is not agent-specific; the stated positioning is a general durable backend foundation that can host agent workloads.
- Evidence: The demo's deep-research handler receives a Restate object context, and an ordinary Python LLM call becomes durable by being wrapped in Restate.run.; The internal event loop interprets service events by persisting embedded state, setting a timer, or dispatching a request to another agent.; The speaker says the distributed-log design is inspired by Apache Flink and Meta's core event infrastructure, with former Meta event-infrastructure architects involved.
- Caveats: The presentation does not cover operational failure testing, replay/side-effect idempotency discipline, state-retention policy, or how code/version changes interact with long-lived executions.; Claims about architectural inspiration are provenance context, not independent proof of reliability or fitness for a specific environment.
- Implications: A proof of concept should test not only successful paths but forced process termination, tool timeouts, duplicate external side effects, redeploys during waits, and cancellation during fan-out.; Durable execution adoption should include standards for which calls are safe to replay and how external systems receive idempotency keys.
Deployment and ecosystem positioning
- Claims: Restate is available as open source, self-hosted software, a bring-your-own-cloud deployment, and a managed cloud offering.; The speaker says the system is packaged as a single binary and supports six SDKs plus integrations with popular agent frameworks.; The claimed BYOC benefit is that data need not leave the customer's cloud account.
- Evidence: For high availability, the described deployment approach is to run multiple instances and snapshot to object storage.; The speaker emphasizes that developers can also use any LLM SDK and create custom agents by wrapping selected steps with Restate constructs.
- Caveats: No transcript detail establishes supported languages, framework coverage, data residency guarantees, access-control architecture, pricing, or managed-service operational boundaries.; Keeping data in a customer's cloud account does not by itself establish security, governance, or compliance adequacy.
- Implications: Before platform selection, validate SDK fit, tenancy boundaries, observability export, secret handling, backup/recovery objectives, and the operating model for self-hosted versus BYOC versus managed service.
Notable Concepts & Terms
- Durable execution: Journaled execution semantics that let an agent resume from a completed step after crashes, restarts, or long delays rather than restarting the entire task.
- Virtual object: Restate's session-scoped stateful actor abstraction: a unique ID, isolated key-value state, and serialized durable handlers for a persistent agent session.
- Durable promise: A persisted suspension point used to pause an agent for an external event such as a human approval and resume later without holding compute.
- Execution ID: The identifier used to inspect an active run, retrieve output, signal new state into it, or cancel it.
- Signal: A mechanism for injecting new context or state into an already-running agent loop, enabling mid-flight user steering.
- LLM gateway: A separate durable service handler that centralizes model calls and becomes an enforcement point for policy checks and departmental flow control.
- Push-based invocation: Restate sends invocations to services rather than relying on workers to poll an orchestrator, which the speaker argues lowers latency and fits serverless wake-up behavior.
- Zombie failure: A class of distributed-systems failure referenced by the speaker alongside network partitions; it motivates using a resilient runtime rather than ad hoc retry logic.
Operator Notes / Why Ken Should Care
- Run a small durability bake-off for one long-running, tool-using agent: kill the worker during an LLM/tool call, fail an external API, redeploy during a human approval, and cancel during sub-agent fan-out; compare Restate with the current orchestration option.
- Define an execution-control contract for agent sessions: stable session IDs, concurrency key, signal schema, relevance/reroute decision policy, cancellation propagation, and user-visible cancellation status.
- Move shared model access behind a gateway before scaling internal agents; enforce tenant/department identity, model allowlists, concurrency quotas, budget limits, trace logging, and idempotency handling there.
- Require a design review for replayed side effects: determine which external calls are journal-safe, where idempotency keys are required, and how duplicate Slack messages, purchases, tickets, or writes are prevented.
- Validate vendor claims around 45 ms p99, HA snapshots, framework compatibility, and data residency using workload-specific tests and security review rather than relying on this presentation.
Source/Metadata
- Title: Every step you take, every call you make: the reliable agent stack — Giselle van Dongen, Restate
- Transcript words: 6315
- Duration seconds: 1249
- Timestamp note: No usable timestamps or chapters were present. The transcript contains substantial repeated sections and ends with extraction noise repeating 'November 10th.'
Transcript
Hi everyone. This talk will be about how to run agents reliably in production. It will not be about the eval part, but it will be about all the other things you need to get going in order to run agents resiliently. So the infrastructure layer basically. I want to set the scene with this quote of Andrej Karpathy of last week. It describes that the way we interact with agents and LLMs has been evolving in three waves. The first wave was an LLM being something like a website where we go to, we ask it a question, it thinks for a few seconds and then gives us a response. The second wave was going towards agents. It was an app that we download to our computer. It has some tools at its disposal and it can do some work with our interaction. Now the third wave will be going more and more towards persistent and asynchronous entities. So agents being long-running processes in our infrastructure with access to tools and other agents around the organization and context. And so as our use cases are evolving more and more from single agents to agentic platforms that connect parts around the organization, our infrastructure layer should also evolve with that. So when we look at the types of tools that are currently out there to implement agents, a lot of innovation has been done on sites such as agent-seks and memory. And agent-seks are really cool to implement POCs and get started quickly, but they don't necessarily help with connecting the distributed bits around an organization. And if you want to implement more complex agentic systems, you actually need all of those things. So that is the layer that you see below here where you have to deploy extra infrastructure, you need to write things like retry logic, recovery logic, and all of that is actually pretty complex to get right. But completely necessary to run long-running stateful and distributed processes in production. So today I want to talk about an open source framework called Restate. And you can see it a bit as a flexible, durable foundation that lets you build any backend. So it's not specific for agents, but as agents are also just a type of a backend, it also works well for them. The ideas behind Restate come from Apache Flink, which is a popular distributed stream processing engine, and also from some of the ex-architects behind Meta's core event infra. So what are the ingredients in Restate? Basically four parts. First of all, it makes sure that a single run of an agent is resilient. This is called durable execution in the industry. Think about things like when an agent runs for a week and then crashes, we want to be able to bring it back and let it continue exactly at the point where it failed. We don't want it to start over from the beginning. Another area here is running many concurrent sessions in parallel. Imagine running thousands of concurrent agent sessions at the same time and needing to make sure that state is always consistent and that different agents don't interfere with each other. And then going more towards things like communication between agents, between agents and MCP servers and other tools. And finally also control. Making sure that when an agent, for example, is doing something you don't want it to continue or when it's stuck, being able to actually cancel or kill the execution. So the way that you can think of it is as follows. Restate is basically a server which runs in front of your agent service. So as a separate component. It sits there a bit like a message broker or a proxy. And when there's a request for your agent, Restate proxies the request to the service and pushes it to the service basically. And from that moment there's an open connection between Restate and the agent. And that connection will basically be a bit like a lifeline for the agent. So as the agent is doing stuff, it sends events over to Restate. And Restate will use that journal of events to recover the process after a failure. So from a slightly higher level explanation, you could say that it's turning a normal function in your application into something that is long running, durable and stateful. Without having to do a lot of the complex things you otherwise need to do for this. So my talk today will be mainly a demo. So I'll be showing you a research agent that's connected to Slack. Imagine we are working at some company and we want to make a Slack agent available to all of our employees. So if I go here into Slack, I can hear in this channel for example ask what is new in AI. Now let's have a look at what it's doing under the hood. So if I go back here, I have here the Restate UI. This is a bit like a cockpit for your agents. So you can see a registry of all the agents that are currently registered. And you can also see for example which execution is currently happening. So here is the deep research agent that I spinned up a few seconds ago. We can see what it's currently doing now. It called first an LLM and then it sent me an answer via Slack. This first LLM call was a planner agent. So what it did is it planned the research and sent me a list of subtopics that it wants to research. Now if I press here approve, then this will unblock the workflow and will spin up a set of parallel research agents. So this is basically the classical deep research workflow, right? You have a planner, then a set of sub research agents, and then finally someone who writes a report on this. Like a writer agent. And so this journal you see here on the left, that is basically the events that get sent from the agent to the Restate server. And if this now crashes at some point, this journal is what will be used to recover the execution to the point where it failed. I don't know if there were some errors. I injected a bit of tool errors in here. Yeah, here you can for example see that the sub agent first did an LLM call, then started doing some web searches. And eventually one of the web searches didn't go through because the API was down. And then you see here on the right how it got retried and eventually completed successfully. So instead of starting over, it uses the journal to recover the progress. Let's now have a look at what this looks like in code. So the basic unit of how you implement applications in Restate is by writing HTTP handlers. And those handlers become durable by using the Restate SDK. So here in this case, we have here our deep research handler. And here as a first argument, we have a Restate object context. And the way you can imagine that is basically as that connection to that Restate server. Whenever I do an action on this Restate object, it will lead to an event being sent to Restate. So for example, when I did that planner LLM call, what actually happened under the hood was it executed here this Python function. This is just a simple light LLM call. And the way I made it durable is by wrapping it in Restate.run. So what happens is by doing these durable steps, if this fails somewhere here, two hours or two months later, it will recover to exactly that point. So that's the idea of durable execution. You're always able to recover a process to where it was. You can also use that for other things, not necessarily for failure recovery. For example, imagine we want to ask a human to approve something, and this approval might take weeks or a month. This process needs to be able to survive restarts and redeploys over those kind of long periods of time. And so with durable execution, you can actually also suspend a function and bring it back when it's able to make progress. So in the case of a human approval, what we do here is basically we create a durable promise, which lives in that journal, a bit like a suspension point. Then we ask a human to click that button in Slack, as I showed in the beginning. And while we are waiting, this process actually suspends. So if it's running on serverless, this is not using execution time on our functions. Once the response comes in, this then gets unblocked and can continue where it left off. So what we see here is a bit like a workflow. It's a set of steps that get executed durable. But when we think about agents and also the way that Karpathy described it in the tweet, it's more like a persistent stateful entity that lives for a longer period of time, that has some memory. So a workflow is not the nicest way to model this kind of thing. So the way that we can model this in Restate is by using something called a virtual object. So imagine in the use case that I'm showing this Slack research agent, imagine that I don't want to wait for 10 minutes to give it some follow-up context or maybe I think about something else that I should have told it. I want to actually be able to interact with it, not wait till that research is finished before I can send a follow-up. And so this is basically what a virtual object in Restate is. It's a bit like a stateful actor. It has a unique ID, for example, a session ID. It has some key value state that is isolated for that specific session that you can write to. Imagine, for example, your history of messages. And it also has a set of handlers that can execute durable functions for this session. So here, the way I implemented this use case that I mentioned of interacting with a running process is as follows. So imagine in the use case that I'm showing this Slack research agent, imagine that I don't want to wait for 10 minutes to give it some follow-up context or maybe I think about something else that I should have told it. I want to actually be able to interact with it, not wait till that research is finished before I can send a follow-up. And so this is basically what a virtual object in Restate is. It's a bit like a stateful actor. It has a unique ID, for example, a session ID. It has some key value state that is isolated for that specific session that you can write to. Imagine, for example, your history of messages. And it also has a set of handlers that can execute durable functions for this session. So here, the way I implemented this use case that I mentioned of interacting with a running process is as follows. This is a session controller. Again, it has this Restate object context at its disposal to do things in a recoverable way. It can write to the session store. Here I'm retrieving the chat history. And one thing that's interesting there is that in order to run these kind of sessions in very high parallelized ways with thousands of sessions at the same time, we need to make sure that agents do not interfere with each other. Imagine I'm sending two messages on Slack and now two agents are actually overwriting each other's session state. To prevent that, this will guarantee that only one execution is running at a time. So a second execution will be queued behind the current one. Then let's have a look at how we implement this, interacting with another execution. So an execution in Restate has a unique identifier. And you can use that identifier to connect to it from other processes. For example, to retrieve the output, but also to cancel it, or maybe to signal it, to inject state into an already running agent loop. And so this is a very flexible type of capabilities that you can do to implement things like, for example, signaling an already ongoing agent loop. So what we do here is if there is a current execution ongoing, then we will ask an LLM, is this something that is relevant for the current agent loop? If that is the case, inject this via a signal. If it's not really relevant for what we are currently doing, then cancel what you are currently doing and start over again with this new information. And so this goes further than workflows. It goes more towards writing persistent stateful entities that can interact with each other and have memory at their disposal. So let me show you how this works. So here if I now ask again, what is new in AI? And I wait a few seconds, then it should respond again with a plan. And then I can say, for example, some extra info, focus on frontier models, let's say. So once I have the plan, I will inject that bit of extra state. Now let's look at the UI of what this is now doing. So here I have that controller which I just showed. It started calling an LLM to classify this new input. Once this comes back, it will probably decide that it should signal it because it's still relevant to the research it's currently doing. So this injects that new message into the ongoing agent loop. So let me show you in the deep research agent again. So we first called an LLM, then asked us. Then we injected this new message of focus on frontier models. And then it took that into account and started over again. Here I can now, for example, also say something like forget about that. Research AI policy. And if I send this, then the coordinator will decide to cancel the ongoing run and start a new one that will research this new topic. And so this cancellation is basically a signal that gets sent down the stack or the call chain. So if my agent was already spinning up sub-agents first, those sub-agents would be canceled, then the controller itself. And that way, it would basically rewind the stack and give agents the ability to roll back. Okay, so this went a bit more into the direction of stateful persistent entities that we can interact with over longer periods of time. Now the last part of the demo that I want to show is going more towards being able to write highly customized applications. Imagine that we deploy this in production, but then a few months later, a new model provider brings out a new model, for example, Fabulous. And even though the model is very good, it's also very expensive. And we notice that this research agent is actually starting to cost a lot. These kinds of things that pop up halfway through a project require you to then deploy a lot of new extra infrastructure or find a good way to solve this. This is the kind of thing that Restate really excels at. It doesn't really peg you into a specific way of how you should write your application. It basically gives you a durable programming model that lets you implement an application in the way that fits for you and also extend it if necessary. So first I showed this LLM call in the first example as an inline step. It was just a Python function that got persisted. But imagine this use case that we want to actually have more control over those LLM calls. For example, what you can do is then pull this out into its own handler. And this handler can now do things like, for example, a policy check and then do the LLM call. And the other agents instead of doing this LLM call inline can now use Restate's distributed communication primitives to actually just call this LLM gateway instead of doing it as an inline step. And this service fabric that lets you communicate between agents also gives you some things like flow control. So we can, for example, say one department is only allowed to run 300 calls to this LLM gateway at the same time. So the reason why I showed this was just to show you that it's basically a resilient foundation. It makes sure that your process can recover from even more advanced types of infrastructure failures. Things like network partitions and zombie failures. And it gives you tooling to extend and customize as your use case grows. Let's go back to the slides to have a little more of an idea of how this thing is actually implemented on the inside. Because it's actually a pretty interesting design or architecture. So the way it's implemented is basically by having an event driven distributed log implementation. So inside the box you basically on one side have the clients, on the other side the services. And inside the box is a log which persists all those journal events and an event loop. And that event loop basically gets the events from the service based on what the event is. It either persists some state in the embedded state store or it sets a timer or it sends a request to another agent. And by doing that you basically have a durable foundation for whatever an application is doing. The design of this distributed log is heavily inspired by the way that the core event infra layer at Meta works. It's basically an iteration on top of that. And some of those architects now have designed that for Restate as a more generic solution that is available in open source. There are two important things related to this architecture that make it interesting. The first one is that it works as a push model. So whereas most workflow orchestrators actually pull for new tasks, pull from the workflow server, Restate actually pushes the invocations. And the benefit you get from that is that it has much lower latency. So you can use these kinds of workflow guarantees in functions around your application and have latency. So, for example, 45 milliseconds p99 for a 10 step workflow. Pushing invocations also works very well for serverless because they require you to basically send the request and wake up the function. So this design that I show here includes everything you need. It includes the state store where we were embedding the state as the UI. It's a single binary. So it's pretty easy to operate as well. To run it in a highly available way, you just spin it up multiple times and let it snapshot to object storage. So Restate has six different SDKs. We also have integrations for most of the popular agent frameworks out there. And of course, because it's just a flexible layer, you can also just use any LLM SDK and implement custom agents by just wrapping some steps into these SDK constructs. So it's open source. You can self-host it. We also have a BYOC offering where we deploy Restate in your cloud. You can also use a cloud account and that gives you the benefit that data doesn't leave your cloud account. Otherwise, there's also a managed cloud offering. This was mainly what I wanted to show. If you want to explore the code a bit further, there is here the GitHub repo. It's publicly available. If you like the project, then have a look at the Restate repo itself. We are hiring across the board for all sorts of roles going from engineering to marketing, especially also here in the Bay Area. So if you're interested in that, then definitely check out our careers page. And I will be outside in front of the conference hall here if you want to ask any questions or learn more about Restate. Thank you very much. So it's open source. You can self-host it. We also have a BYOC offering where we deploy Restate in your cloud. You can also use a cloud account and that gives you the benefit that data doesn't leave your cloud account. Otherwise, there's also a managed cloud offering. This was mainly what I wanted to show. If you want to explore the code a bit further, there is the GitHub repo. It's publicly available. If you like the project, then have a look at the Restate repo itself. We are hiring across the board for all sorts of roles going from engineering to marketing, especially also here in the Bay Area. So if you're interested in that, then definitely check out our careers page. And I will be outside in front of the conference hall here if you want to ask any questions or learn more about Restate. Thank you very much. So as the agent is doing stuff, it sends events over to Restate. And Restate will use that journal of events to recover the process after a failure. So from a slightly higher level explanation, you could say that it's turning a normal function in your application into something that is long running, durable and stateful. Without having to do a lot of the complex things you otherwise need to do for this. So my talk today will be mainly a demo. So I'll be showing you a research agent that's connected to Slack. Imagine we are like working at some company and we want to make a Slack agent available to all of our employees. So if I go here into Slack, I can hear in this channel for example ask what is new in AI. Now let's have a look at what it's doing under the hood. So if I go back here, I have here the Restate UI. This is a bit like a cockpit for your agents. So you can see a registry of all the agents that are currently registered. And you can also see for example which execution is currently happening. So here is the deep research agent that I spinned up a few seconds ago. We can see what it's currently doing now. It called first an LLM and then it sent me an answer via Slack. This first LLM call was a planner agent. So what it did is it planned the research and sent me a list of subtopics that it wants to research. Now if I press here approve, then this will unblock the workflow and will spin up a set of parallel research agents. So this is basically like the classical deep research workflow, right? You have a planner, then a set of sub research agents, and then finally someone who writes a report on this. Like a writer agent. And so this journal you see here on the left, that is basically the events that get sent from the agent to the Restate server. And if this now crashes at some point, this journal is what will be used to recover the execution to the point where it failed. I don't know if there were some errors. I injected a bit of like tool errors in here. Yeah, here you can for example see that the sub agent first did an LLM call, then started doing some web searches. And eventually one of the web searches didn't go through because the API was down. And then you see here on the right how it got retried and eventually completed successfully. So instead of starting over, it uses the journal to recover the progress. Let's now have a look at what this looks like in code. So the basic unit of how you implement applications in Restate is by writing HTTP handlers. And those handlers become durable by using the Restate SDK. So here in this case, we have here our deep research handler. And here as a first argument, we have a Restate object context. And the way you can imagine that is basically as that connection to that Restate server. Whenever I do an action on this Restate object, it will lead to an event being sent to Restate. So for example, when I did that planner LLM call, what actually happened under the hood was it executed here this Python function. This is just a simple light LLM like LLM call. And the way I made it durable is by wrapping it in Restate.run. So what happens is by doing these durable steps, if this fails somewhere here, two hours or two months later, it will recover to exactly that point. So that's the idea of durable execution. You're always able to recover a process to where it was. You can also use that for other things, not necessarily for failure recovery. For example, imagine we want to ask a human to approve something, and this approval might take weeks or a month. This process needs to be able to survive restarts and redeploys over those kind of long periods of time. And so with durable execution, you can actually also suspend a function and bring it back when it's able to make progress. So in the case of a human approval, what we do here is basically we create a durable promise, which lives in that journal, a bit like a suspension point. Then we ask a human to click that button in Slack, as I showed in the beginning. And while we are waiting, this process actually suspends. So if it's running on serverless, this is not using execution time on our functions. Once the response comes in, this then gets unblocked and can continue where it left off. So what we see here is a bit like a workflow. It's a set of steps that get executed durable. But when we think about agents and also the way that Carpathie described it in the tweet, it's more like a persistent stateful entity that lives for a longer period of time, that has some memory. So a workflow is not the nicest way to model this kind of thing. So the way that we can model this in Restate is by using something called a virtual object. So imagine in the use case that I'm showing this Slack research agent, imagine that I don't want to wait for 10 minutes to give it some follow-up context or maybe I think about something else that I should have told it. I want to actually be able to interact with it, not wait till that research is finished before I can send a follow-up. And so this is basically what a virtual object in Restate is. It's a bit like a stateful actor. It has a unique ID, for example, a session ID. It has some key value state that is isolated for that specific session that you can write to. Imagine, for example, your history of messages. And it also has like a set of handlers that can execute durable functions for this session. So here, the way I implemented this use case that I mentioned of interacting with a running process is as follows. This is a session controller. Again, it has like this Restate object context at its disposal to do things in a recoverable way. It can write to the session store. Here I'm retrieving the chat history. And one thing that's interesting there is that in order to run these kind of sessions in very high parallelized ways, with thousands of sessions at the same time, we need to make sure that agents do not interfere with each other. Imagine I'm sending two messages on Slack and now two agents are actually overwriting each other's session state. To prevent that, this will guarantee that only one execution is running at a time. So a second execution will be queued behind the current one. Then let's have a look at how we implement this, like interacting with another execution. So an execution in Restate has a unique identifier. And you can use that identifier to connect to it from other processes. For example, to retrieve the output, but also to cancel it, or maybe to signal it, to inject a bit of state into an already running agent loop. And so this is like a very flexible type of capabilities that you can do to implement things like, for example, signaling an already ongoing agent loop. So what we do here is if there is a current execution ongoing, then we will ask an LLM, is this like something that is relevant for the current agent loop? If that is the case, inject this via a signal. If it's not really relevant for what we are currently doing, then cancel what you are currently doing and start over again with this new information. And so this goes a little bit further than workflows. It goes a bit more towards like writing persistent stateful entities that can interact with each other and have memory at their disposal. So let me show you how this works. So here if I now ask again, what is new in AI? And I wait a few seconds, then it should respond again with a plan. And then I can say, for example, some extra info, focus on frontier models, let's say. So once I have the plan, I will inject that bit of extra state. Now let's look at the UI of what this is now doing. So here I have that controller which I just showed. It started calling an LLM to classify this new input. Once this comes back, it will probably decide that it should signal it because it's still relevant to the research it's currently doing. So this injects that new message into the ongoing agent loop. So let me show you in the deep research agent again. So we first called an LLM, then asked us. Then we injected this new message of focus on frontier models. And then it took that into account and started over again. Here I can now, for example, also say something like forget about that. Research AI policy. And if I send this, then the coordinator will decide to cancel the ongoing run and start a new one that will research this new topic. And so this cancellation is basically like a signal that gets sent down the stack of, or the call chain. So if my agent was already spinning up sub-agents first, those sub-agents would be canceled, then the controller itself. And like that, it would basically rewind the stack and give agents also the ability to roll back. Okay. So this went a bit more into the direction of like stateful persistent entities that we can interact with over longer periods of time. Now the last part of the demo that I want to show is going more towards like being able to write highly customized applications. Imagine that we deploy this in production, but then a few months later, a new model provider brings out a new model, for example, Fabulous. And even though the model is very good, it's also very expensive. And we notice that this research agent is actually starting to cost a lot. These kind of things that pop up halfway through a project require you to then deploy a lot of new extra infra or like find a good way to solve this. This is the kind of things that Restate really excels at. It doesn't really peg you into a specific way of how you should write your application. It basically gives you like a durable programming model that lets you implement an application in the way that fits for you and also extend it if necessary. So first I showed this LLM call in the first example as an inline step. It was just a Python function that got persisted. But imagine this use case that we want to actually have a bit more control over those LLM calls. For example, what you can do is then pull this out into its own handler. And this handler can now do things like for example a policy check and then do the LLM call. And the other agents instead of doing this LLM call inline can now use Restates like distributed communication primitives to actually just call this LLM gateway instead of doing it as an inline step. And this service fabric that lets you communicate between agents also gives you some things like flow control. So we can for example say one department is only allowed to run 300 calls to this LLM gateway at the same time. So the reason why I showed this was just to show you a bit like that it's basically just a resilient foundation. It makes sure that your process can recover from even more advanced types of infrastructure failures. Things like network partitions and zombie failures. And it gives you like tooling to extend and customize as your use case grows. Let's go back to the slides to have a little more of an idea of how this thing is actually implemented on the inside. Because it's actually a pretty interesting design or architecture. So the way it's implemented is basically by having an event driven distributed log implementation. So inside the box you basically on one side have the clients, on the other side the services. And inside the box is a log which persists all those journal events and an event loop. And that event loop basically gets the events from the service based on what the event is. It either persists some state in the embedded state store or it sets a timer or it sends a request to another agent. And by doing that you basically have a durable foundation for whatever an application is doing. The design of this distributed log is heavily inspired by the way that the core event infralayer at Meta works. It's basically like an iteration on top of that. And some of those architects now have designed that for Restate as a more generic solution that is available in open source. There are two important things related to this architecture that make it interesting. The first one is that it works as a push model. So whereas most workflow orchestrators actually pull for new tasks, pull from the workflow server, Restate actually pushes the invocations. And the benefit you get from that is that it has a much lower latency. So you can use these kind of workflow guarantees in functions around your application and have like a latency. So for example 45 milliseconds p99 for like a 10 step workflow. Pushing invocations also works very well for serverless because they require you to basically send the request and wake up the function. So this design that I show here includes everything you need. It includes as well that state store where we were embedding the state as the UI. It's a single binary. So it's pretty easy to operate as well. To run it in like a highly available way, you just spin it up multiple times and let it snapshot to object storage. So Restate has six different SDKs. We also have integrations for most of the popular agent frameworks out there. And of course, because it's just like a flexible layer, you can also just use any LLM SDK and implement custom agents by just wrapping some steps into these SDK constructs. So it's open source. You can self-host it. We also have a BYOC offering where we deploy Restate in your cloud. You can also use a cloud account and that gives you the benefit that data doesn't leave your cloud account. Otherwise, there's also a managed cloud offering. This was mainly what I wanted to show. If you want to explore the code a bit further, there is here the GitHub repo. It's publicly available. If you like the project, then have a look at the Restate repo itself. We are hiring across the board for all sorts of roles going from engineering to marketing, especially also here in the Bay Area. So if you're interested in that, then definitely check out our careers page. And I will be outside in front of the conference hall here if you want to ask any questions or learn more about Restate. Thank you very much. I would like to catch up on November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10th November 10 November 10th November 10th November 10 November 10th November 10 November 10 November 10th November 10 November 10 November