Open Reader

Two Roads to Durable Agents: Replay vs. Snapshot — Eric Allam, CEO, Trigger.dev

completed 16:36 May 10, 2026 Watch on YouTube

Current Status

completed

Video ID

svCnShDvgQg

RAG / Chat

Enabled
Two Roads to Durable Agents: Replay vs. Snapshot — Eric Allam, CEO, Trigger.dev
Description

Replay-based durability — wrapping every step in a journal, replaying on recovery, requiring deterministic code — is how everyone makes agents durable today. It works until it doesn't: the journal grows with every turn, the structure starts constraining how you write code, and an agent that needs to run for hours starts looking less like a transaction and more like a session. This talk separates the problem in two: context durability (the append-only log of everything the LLM saw, which already fits in a database) and execution durability (the files, memory, and subprocesses that live in the compute layer, which don't). The answer to the second half isn't a smarter log — it's OS-level snapshot and restore. Eric Allam walks through how Trigger.dev built this on Firecracker microVMs, getting snapshots down to 14 megabytes compressed with sub-second save and hundred-millisecond restore times, and why IBM mainframes in 1966 got there first. Speaker info: - https://x.com/maverickdotdev - https://www.linkedin.com/in/eric-allam/ - https://github.com/ericallam

Summary

Generated by claude-haiku-4-5-20251001

Two Roads to Durable Agents: Replay vs. Snapshot

Main Topics

  • Evolution of Backend Infrastructure: From stateless to stateful compute paradigms
  • Agent Durability Challenge: Making LLM agents resilient, long-running, and recoverable
  • Two Complementary Approaches: Replay model vs. snapshot/restore methodology
  • Technical Implementation: Using Firecracker Micro VMs for efficient execution snapshots

Key Points

Historical Context

  • 30 years of stateless architecture: CGI → LAMP stack → serverless all follow the "request + DB = response" model
  • Workflow engines (10-15 years ago): Introduced replay mechanisms by caching steps to handle failures in multi-step processes
  • Agent paradigm shift: LLMs orchestrate code execution (not vice versa), fundamentally changing infrastructure needs

The Replay Model Problem

  • Wraps every side effect as a cached step in a replay journal
  • Generates audit trails and enables failure recovery
  • Critical limitation for agents: Replay logs grow exponentially with agent interactions
  • After a few hours of agent work, logs become prohibitively large
  • Agent capability duration doubling every 4-7 months
  • Agents need sessions lasting days, not hours

The Two-Part Solution

Context Layer (Append-Only Log)

  • Stores all system messages, user messages, tool calls, responses
  • Provides durability across code versions and crashes
  • Extremely scalable using existing primitives (databases, object storage)

Execution Layer (Snapshot/Restore)

  • Captures machine state: files, memory, processes, sub-processes
  • Allows safe shutdown between user interactions
  • Enables cost-effective durability without keeping VMs running

Technical Implementation

  • Foundation: CRIU (Checkpoint/Restore In Userspace) from 2011
  • Current solution: Firecracker Micro VMs (2024)
  • Optimization challenges solved:
  • Memory compression (512MB → 14MB compressed)
  • Seekable compression for on-demand page restoration
  • Snapshot time: <1 second
  • Restore time: ~200 milliseconds
  • Throughput: 15,000 VM starts per minute

Notable Quotes

> "The LLM orchestrates the code, right?" — Highlighting the fundamental shift from traditional workflows

> "An agent isn't a transaction. It's a session, right? And it lasts as long as the user wants it to last."

> "For 30 years, we had stateless compute as the core of backend infrastructure. And I think agents are forcing this move to become stateful compute."

> "You can almost render a video. The FPS would be about 30 FPS." — Describing the speed of VM operations

Takeaways

  • Architectural Paradigm Shift: Agents require moving from stateless to stateful compute infrastructure
  • Replay Model Limitations: Traditional workflow durability solutions don't scale for long-running agent sessions
  • Hybrid Approach is Essential: Combine append-only context logs with VM snapshots for complete durability
  • Recovery Flexibility: Different failure scenarios require different recovery strategies (timeout retries vs. context restoration)
  • Practical Technology: Firecracker Micro VMs with compression enable cost-effective, sub-second snapshot/restore cycles
  • Open Source Coming: fcrun tool will provide Docker-like CLI for easy snapshot/restore management
  • Historical Validation: Checkpoint/restore concepts trace back to 1966 IBM mainframes—applying proven concepts to modern AI agents

Transcript

2547 words en Processed in 172.7s

Music How's everyone doing? It's a full room, look at this thing. Let's get started. Okay, so here is our agent. It's got the turn loop, it's got the LLM loop. This little example works well enough running on your own machine, but what if we want to deploy these to production backends and run them on our servers? So what do we want them to do, right, when they run on our servers? We want them to do long-running meaningful work. Should be durable across turns and versions of our code, and it should be able to recover from errors. So I'm Eric. I'm one of the founders of Trigger.dev, and we've been trying to make it easy to deploy these types of agents to production for the last few years. And what I like about this little meme here is which one is the agent and which one is the human? I like to think. Yeah. Yeah. So this talk is about the fundamental shift that agents are posing to back-end infrastructure and some of the ideas for how to achieve these durable agents. So before we go into that, I want to do a little history lesson here. Let's take a step back and see how we got here. So the very first dynamic web backend was CGI back in 1993. Anyone here ever done CGI stuff? Cool. Thanks. So the model was really simple. HTTP request comes in. The server forks a whole new process. Request data goes in. [SPEAKER_00] The process does some stuff. And then it writes the response to standard out. And then the process goes away. So it's completely stateless. Shortly after that, PHP came out, which turned into the LAMP stack. And the LAMP stack reused the PHP process, right? But it kept the principle that all you needed to do to create a response was the request, some state from the database, and then it would do the request. So the second request would come in and it would do all the same work again, and it would produce the response. So this became sort of request plus DB equals the responses, which became known as the shared nothing architecture, right? So looking at another way, shared nothing means that the compute layer is stateless, right? There's nothing, there's no meaningful state in the compute. The state is in the database, right? So this became the dominant backend infrastructure for the last 30 years, right? Everything that followed from this, like Ruby on Rails, Node.js, serverless, it all follows the same paradigm. As web applications became more complicated and sophisticated, they started performing these side effects outside of the request and DB lifecycle. These side effects are async tasks. So they started out simple, send an email, charge a credit card, resize an image. But soon they became these multi-step side effects, right? Like this process order example here, where you do things in sequence, right? You'd quickly run into the problem of how to handle failures in something like this, right? So send receipt fails. You can't just retry the whole process order thing again without charging the credit card twice, which is bad. So about 10 to 15 years ago, workflow and durable execution engines were adopted to solve this problem, right? So you'd write your code like this now, where you'd wrap every single side effect in a step that becomes cached as it's executed. So now that solves the problem nicely when you call process order for the second time. You skip the things you've already done and then you do the thing that you want to do originally, right? And you don't charge the credit card twice. So this is what I call the replay model. So it builds durable execution on top of existing stateless compute architecture, which I covered was how everything works, right? So you get this nice side effect of an execution history, an audit trail of everything that happened. And also by being able to resume to a specific point in time, you can recover from a failure. But you can also wait for something else to happen, right? So you can wait for a human to do something and then you can resume execution. Some of the downsides of this replay system is that now you have to wrap everything in these steps and everything outside of steps has to be deterministic. You get this rigid structure. You have to write your code in a certain way or things break. And replay journaling and replay journal versioning is tricky if you deploy a new version. So this is the very simple and truncated history of the state of the world in 2023 when LLMs came out. At first they really fit neatly into this paradigm, right? They would just become another step in a workflow, right? They would classify some text or something, but it was still in this old workflow era, right? So not long after that we got tool calling and tool calling got good and we were introduced to the agent loop, right? And the big difference there is code is no longer orchestrating the LLM. The LLM orchestrates the code, right? So we're back at our agent loop, right? And what happens if we, yeah, you can see that. If we want to make this agent loop durable, can we do it with this replay model, right? What does that look like? So what does that look like? Every LLM call, right, becomes a step in the replay journal. Every tool call becomes a step. On resume, the function re-executes on top and replays all that stuff, right? So after a single turn of the LLM not doing too much, this is what the replay log looks like, right? And as you keep interacting with the agent, the log grows and grows and grows. At a certain point, you might hit some fundamental limit of your replay system. That could be either too many entries or the entries grow too large. But this falls over once you hit that limit. And there's this measure of how long agents can actually do meaningful work. And apparently it's doubling every four to seven months. So right now we're on about a few hours. So after a single turn of the LLM not doing too much, this is what the replay log looks like, right? And as you keep interacting with the agent, the log grows and grows and grows. At a certain point, you might hit into some fundamental limit of your replay system. That could be either too many actual entries or it could be the entries grow too large. But this falls over once you hit that limit. And there's this measure of how long agents can actually do meaningful work. And apparently it's doubling every four to seven months. So right now we're at about a few hours. But not too long from now we'll be at multiple days of length as these agents build to actually do meaningful work. So replay gave us these durable transactions. But an agent isn't a transaction. It's a session, right? And it lasts as long as the user wants it to last. Multi-step workflows start and end and sessions keep going for as long as possible. So if we take a step back and think about what an agent needs to be durable from first principles, I think of it as an agent having these two halves, right? The first half is the context. So this is all your system messages, user messages, tool calls, tool results, assistant responses, right? This is all the actual context, everything that went in and out of the LLM. This is extremely valuable, obviously. You want to make that durable, right? But you also have this execution layer. And as agents are more complicated, doing more things, they want a machine, right? They want to be able to do stuff like they could do on your laptop, right? They want to be able to write files, use memory, create sub-processes. And I think of both of these as super valuable pieces of state, but they can be treated separately. So the context is first and the most important. And it's just an append-only log of everything that happened, right. And you can make this log durable using any primitive that already exists, like a database, object storage, distributed file system. There's a ton of technologies that are coming out that are specialized in making this durable, right? And when that context log is saved somewhere, now you can have durability across versions of your code, right? So you upgrade your harness and you can still use that same context, right? Maybe the machine crashes and you can save somewhere so you can pick up where you left off, right? And append-only logs scale really well. But what about making this execution side durable, right? For the types of agents right now that are doing meaningful work, there's a lot of state that happens in the compute layer that we might want to save. Maybe you've cloned a GitHub repo, you've installed some packages, you've got some data sets of memory, you're running a dev server, right? You sandbox and a sub-process, whatever it is, right? You can't really make that durable using a log. And how do we get this to work, right? So you have to wait for some amount of time for the next user message, right? And we can't just keep the machine running. It would be nice, but we can't. It would be too expensive. So instead of recreating the execution state from a log, we should use snapshot and restore. So this allows us to snapshot the machine, shut it down, save it to disk, and then when the user message comes in, we restore it, right? So this gives us durability across turns. So when the user goes to lunch, right, we don't have to run the machine the whole time. It allows us to preserve everything that the agent was doing. And effectively compared to running the machine live, it's pretty cheap. So I think if you combine these two things, then you get a durable agent, right? You've got the context, so context durability and execution durability, right? And this also allows you to recover from errors. So one of the whole points of having these durability guarantees is to recover, right? And it depends on what happened, what went wrong. You can recover in different ways. So say the LLM isn't working for some reason. That never happens, but you never know if it happens. And it takes a long time to retry. Maybe it says, wait, wait, 15 minutes so you retry your next message. But you don't want to wait in memory, so you snapshot, and then you restore when you can retry. But if there's something wrong with the machine, maybe you've shipped a bug, or maybe there's just an issue with the machine, right? It crashes. You have the context log, and you can recover that. So for 30 years, we had this stateless compute as the core of backend infrastructure. And I think agents are forcing this move to become stateful compute. And at the heart of that, I think, is going to have to be this snapshot and restore capability. But this isn't actually new. This is an IBM mainframe from 1966, and it actually has checkpoint and restore. Because they would run these super expensive jobs for hours. And something went wrong, and they couldn't afford to run it all again. So they would add these checkpoints into their code, right? Fast forward to 2011. This thing called Cree was developed. It was a way to suspend and restore a process from user space. So it would basically inject a process with this parasite. And then they would force the process to dump everything to memory. And then it would remove all the traces of the parasite. And it actually worked. In 2024, we actually shipped this. And we've done millions of snapshot restores since. It's transparent for the process. So the process doesn't have to participate in it. And it's compatible with container runtimes, which is good. So the downsides are you can only checkpoint a process. So if you're doing stuff with FFmpeg or you've got a Chrome instance running or anything else, right? It doesn't work. It only captures open files. And then it would remove all the traces of the parasite. And it actually worked. In 2024, we actually shipped this. And we've done millions of snapshot restores since. It's transparent for the process. So the process doesn't have to participate in it. And it's compatible with container runtimes, which is good. So the downsides are you can only checkpoint a process. So if you're doing stuff with FFmpeg or you've got a Chrome instance running or anything else, right, it doesn't work. It only captures open files. So if you're working with the file system, it has to be open at the time of snapshot or you won't get snapshot. And then also if you, it's nice that it's compatible with containers. But once you are compatible with containers, you have to work with registries and push and pull. And it gets very slow. So last year we moved to Firecracker Micro VMs. And this allows us to snapshot the entire machine, right? So everything that's on a machine, on a VM, we can snapshot it. And then we can restore it and it pick up right where it left off, no matter what was happening in the machine, right? But if you do that in a naive way, it can be quite expensive. So say you have a default machine size of 512 megabytes. If you do a snapshot, it's 512 megabytes on disk. So that's not great. So obviously you've got network transfer costs, you've got storage costs, and there's a lot of memory there that's not actually being used. So we actually solved this with compressing it. We actually use a seekable compression. So when we restore, we actually don't restore all the memory pages at once. We capture when it needs to be restored and decompress that little bit that needs to be restored at time. We also have a couple other techniques for layering the snapshot. And we can get the snapshot down to 14 megabytes compressed. And that's a knob you can tweak, and depending on what kind of performance you want, you can compress more or less. So that's pretty much all we had to do other than all of that. And once we did that, we got super fast snapshot and restore times. So this is a graph comparing Cree and Firecracker, but basically the moral of the story is that snapshots are slightly under a second and restores are a couple hundred milliseconds. We've actually bundled all of this into a tool that's going to be open sourced here soon. It's called fcrun, depending on who you ask. So this allows you, it's like a Docker-like CLI. So you can drop in replacement for the Docker command for running containers in Firecracker VMs and snapshotting and restoring them. So for example, you can run Alpine and it's super fast. And you can snapshot running VM and it's super fast. You can fork a VM, also very fast. This is a little benchmark for TTI, so basically how long it takes the VM to become interactable with the internet. So we're doing 15,000 VM starts per minute. You can almost render a video. The FPS would be about 30 FPS. So it's extremely, extremely fast. So this is going to be powering our future compute layer. But it's open source. Not yet, but very soon. So back to where we started with our little agent loop here. And we've made it durable now by doing two different things: context log and execution snapshots. So we got durability across versions, durability across turns, across failures.