Open Reader

Your Agents Need a Save Button - Hamza Tahir, ZenML

completed 17:07 Jul 18, 2026 Watch on YouTube

Current Status

completed

Video ID

bZISsg7H7DA

RAG / Chat

Enabled
Your Agents Need a Save Button - Hamza Tahir, ZenML
Description

Most of an agent's life is spent waiting - on a tool, a human, the next step - and the whole time you're holding a live process awake and billing for it. Multiply that across every agent your org wants to run overnight and the math stops working. A save button fixes the obvious stuff: freeze an agent to durable state, drop its compute to zero, bring it back in milliseconds when there's work, and if it crashes, resume from the last save instead of re-burning every token from the top. The interesting part is what a save button does after the run is over. Reload an agent to any point in its trajectory, change one thing, a prompt, a model, a tool, and watch whether it does better. A finished run stops being a log you read and becomes something you re-run and improve. I'll argue that this one primitive, a checkpoint, is the most underrated thing in agent infrastructure, and follow it down to the sandboxes and Kubernetes-shaped infra the industry is quietly racing to build so a million saved agents can sleep for free. You'll leave with a model for running and improving agents at scale without paying to keep them awake. Speakers: - Hamza Tahir (ZenML): Hamza Tahir is co-founder and CTO of ZenML and co-founder of Kitaru, the durable runtime for AI agents. He's spent a decade building production ML and AI infrastructure used by JetBrains, the German Bundeswehr, and Adeo, and writes and speaks on what actually breaks when agents hit production. X/Twitter: https://x.com/htahir111 LinkedIn: https://www.linkedin.com/in/hamzatahirofficial GitHub: http://github.com/htahir1

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Agent observability should evolve from read-only traces to a durable, checkpointed execution runtime that lets teams replay real production runs, modify one variable, diff outcomes, and make safer deployment decisions.
  • Why it matters: This provides a concrete control-plane pattern for agent systems: recover the full executable state behind a trace so model, tool, policy, and routing changes can be evaluated against real workloads rather than guessed from isolated tests.
  • Best use: Use it to assess whether Ken's agent stack captures enough code, artifacts, environment, and intermediate state to support counterfactual evaluation and cohort-level release gates.

Executive Summary

Hamza Tahir argues that conventional agent traces are insufficient because they record emitted telemetry after the fact but remain disconnected from the execution runtime. A trace may show tool inputs and outputs, but generally loses in-memory state, filesystem artifacts, exact code, configuration, and environment—the elements needed to answer the operational question: what would have happened if the agent had used a different model, tool response, or policy?

His proposed solution is an agent "save button": checkpoint the runtime state throughout an execution, including code, artifacts, configuration, and execution environment, then fork/replay from any checkpoint. Rather than rerunning an entire workflow, the system can reuse the known state before an intervention point, execute only the downstream branch, and compare the changed run with the original baseline.

The key operating loop is checkpoint, replay, diff, decide. Teams should start with meaningful cohorts of real production runs—such as high-cost, high-latency, or risky cases—apply a proposed intervention at scale, compare final outcomes and cost/latency distributions, and then ship, route, or hold the change with human approval. This makes production executions the basis for evaluation rather than relying exclusively on synthetic eval sets.

The presentation is also a product demonstration for ZenML's new open-source tool, Kitaru. Its central caution is important: a single replay that appears cheaper or equivalent is anecdotal. In the demo, replacing a model with GPT-5-nano looked favorable in one support case, but a cohort analysis concluded the cheaper model should not be shipped for that dataset.

Key Takeaways

  • Claim: Read-only tracing is not enough for reliable agent debugging or evaluation; teams need execution state that can be resumed and altered. | Evidence: Tahir distinguishes telemetry spans and tool I/O from the missing runtime context: variables in state, in-flight filesystem data, decision logic, exact code, and the execution environment such as a Docker image or sandbox. | Implication: Treat traces as an observability surface, not as a sufficient evaluation substrate; agent infrastructure should preserve the executable context needed to reproduce and fork behavior.
  • Claim: Checkpointed agent runs enable practical counterfactual tests of models, tools, policies, and intentional failures. | Evidence: The speaker gives examples of swapping to a cheaper open-source model, mocking a tool with a different return value, degrading a dependency deliberately, testing policy changes, and checking whether a support request should have escalated to a human. | Implication: Design agent workflows with explicit checkpoint boundaries around LLM calls and consequential tool calls, allowing targeted interventions without rerunning known-good earlier steps. | Caveat: Replay validity depends on capturing the relevant state and controlling dependencies; the transcript does not establish that every external side effect or nondeterministic source can be faithfully reproduced.
  • Claim: A replay system can reduce evaluation cost and increase diagnostic precision by reusing the baseline state before the changed checkpoint. | Evidence: In the Kitaru demo, a run is forked after a tool call and the first three checkpoints are skipped because their state has already been captured; only the downstream execution is run after changing the model to GPT-5-nano. | Implication: For expensive multi-step agents, partial re-execution should be a core optimization: preserve upstream context once, then test downstream alternatives cheaply and comparably.
  • Claim: Model-cost optimization must be evaluated as an outcome tradeoff, not a token-price decision. | Evidence: A single support-agent replay using GPT-5-nano produced a similar visible result at lower input/output token cost, yet the subsequent cohort analysis returned a "don't ship" recommendation. Tahir cites a Braintrust study as evidence that naive cheaper-model swaps can create a false economy when bots fail to resolve requests. | Implication: Set release criteria around task resolution, safety/escalation behavior, and business value alongside latency and cost; do not approve a routing change from token savings alone. | Caveat: The speaker does not provide the cohort size, quality metric, or economics behind the demo's "don't ship" verdict, so the conclusion demonstrates the methodology rather than proving a general rule about GPT-5-nano.
  • Claim: One or two replays cannot establish agent reliability; decisions should be based on cohort-level distributions from real production executions. | Evidence: Tahir cites Tau-bench for the point that a model passing 60% of the time is self-consistent only about a quarter of the time, framing a single replay as an anecdote. He recommends cohorts selected from expensive, slow, or risky runs. | Implication: Build segmented evaluation cohorts and compare distributions of success, escalation, failure mode, cost, and latency before changing defaults, policies, or model routing. | Caveat: Cohort replay itself can become expensive, so selection and prioritization matter rather than indiscriminately replaying all traffic.
  • Claim: Production-grounded simulation can materially accelerate improvement loops when historical executions are replayable. | Evidence: Tahir references a June 1 DoorDash blog post describing replay of customer bots in a simulated environment; he says the process fell from hours to five minutes for hundreds of simulations, with 90% fewer hallucinations and results within two points of production. | Implication: Production replay is a credible path to higher-fidelity evaluation, but Ken should independently validate reported benchmarks and define metrics that correspond to his own agent outcomes. | Caveat: These are secondhand claims about DoorDash's published results, and the transcript does not define the hallucination metric, baseline, or the meaning of "two points."
  • Claim: At large scale, an agent-accessible runtime and MCP interface can make cohort-analysis results actionable rather than manually inspected. | Evidence: Kitaru's replay many produces a JSON report; the demonstrated MCP server is asked to read it, analyze the cohort, fetch execution artifacts, and flag red flags. Tahir notes that manual comparison is manageable for roughly 10 cases but not thousands. | Implication: Expose replay results and artifacts through a constrained query interface, then automate triage and recommendation while preserving deterministic metrics and human approval for production changes. | Caveat: An LLM-generated recommendation should not replace release governance; Tahir explicitly retains a human in the loop at the final decision.

Detailed Brief

Kitaru demonstration: what a replayable runtime records and compares

  • Claims: Kitaru is positioned below the agent harness/framework as a durable runtime layer, rather than as a standalone tracing UI.; The product permits interventions both at the model layer and at the tool/policy layer, then presents original and forked executions side by side.; A final artifact comparison is intended to reveal not only that execution diverged, but whether the operational decision changed.
  • Evidence: At an individual tool-call checkpoint, the UI displays the configuration, executed code, and incoming/outgoing artifacts; it snapshots state between checkpoints along with the environment.; The demo mocks a lookup policy tool with another function from the codebase while holding the model constant, producing different logs and a different downstream artifact.; In the three-run comparison, all versions restrict the account charge, but the first two conclude the case needs review while the policy-modified replay concludes it is safe to answer.
  • Caveats: Different end decisions are not inherently improvements; a policy branch that changes "needs review" to "safe to answer" must be judged against actual safety and support-quality standards.; The transcript presents product behavior in a controlled demo and does not cover operational requirements such as secrets handling, state retention policies, replay isolation, or protection against re-triggering external effects.
  • Implications: A useful agent diff needs semantic outcome artifacts and decision states, not merely token counts and timing spans.; Checkpoint/replay infrastructure should distinguish pure/replayable steps from side-effecting steps and enforce safe simulation or mocking for the latter.

Operational playbook and governance model

  • Claims: The proposed workflow is explicitly production-led: begin with real runs and their checkpointed state rather than relying solely on synthetic tests.; The speaker's intended end state is a partially automated release loop in which systems identify meaningful cohorts, evaluate candidate changes, and surface a recommendation for approval.
  • Evidence: Tahir summarizes the sequence as: select production cohorts, replay a change, diff against the known baseline, decide, then ship, route, or hold.; Suggested cohort filters are high cost, long duration, and high-risk executions.
  • Caveats: Real production data may carry privacy, retention, authorization, and customer-data restrictions that must be addressed before making runtime state broadly queryable.; A human remains the final approver in the speaker's model, even if an agent performs the cohort analysis.
  • Implications: Release governance can evolve from static pre-deployment evals toward evidence-based routing policies that are continuously refreshed from production behavior.; Cohort construction becomes an important product and risk decision: prioritize slices where regressions would be expensive, safety-sensitive, or difficult to detect.

Notable Concepts & Terms

  • Agent save button: The metaphor for persistent checkpoints of agent execution state, enabling historical recovery and counterfactual replay.
  • Durable runtime: A runtime layer beneath the agent harness that preserves code, artifacts, configuration, environment, and state rather than only emitting telemetry.
  • Checkpoint, replay, diff, decide: The proposed closed-loop evaluation methodology: preserve a baseline, fork it with an intervention, compare outcomes, and make a deployment decision.
  • Harness layer: The framework/application layer used to build agents; Tahir argues it should sit on a runtime capable of durable state capture.
  • Cohort replay: Applying the same counterfactual change across a selected set of production runs to measure distributions rather than relying on a single example.
  • Kitaru: ZenML's newly launched, open-source tool demonstrated as a checkpointing and replay runtime with execution diffs, batch replay, and an MCP server.
  • MCP server: The interface used in the demo to let an LLM/agent query replay reports and runtime artifacts for large-scale cohort analysis.
  • False economy: The risk that a lower-cost or faster model appears favorable on infrastructure metrics while reducing task resolution or business value.

Operator Notes / Why Ken Should Care

  • Audit whether the current agent platform can reconstruct a run from a checkpoint with exact prompt/model configuration, code version, tool inputs/outputs, artifacts, environment, and controlled secrets references.
  • Define replay safety rules before implementation: classify side-effecting tools, require mocks or sandboxing during counterfactual runs, and prevent external writes, payments, messages, or account changes.
  • Create a release-gate cohort strategy for high-cost, long-running, and safety-sensitive tasks; require outcome-quality and escalation metrics alongside token cost and latency.
  • Run a small proof of concept comparing baseline versus candidate model/tool-policy routes on recorded production cases, then inspect where final decision artifacts—not just text outputs—diverge.
  • Evaluate Kitaru as a reference implementation, but separately diligence data retention, access controls, execution isolation, nondeterminism handling, and integration fit with the existing harness and observability stack.
  • Keep human approval on high-impact routing and policy changes even if an MCP-connected analysis agent generates the recommendation.

Source/Metadata

  • Title: Your Agents Need a Save Button - Hamza Tahir, ZenML
  • Transcript words: 5116
  • Duration seconds: 1027
  • Timestamp note: No timestamps or chapters were provided. The transcript contains a substantial duplicated segment beginning after the conclusion of the first pass.

Transcript

2799 words en Processed in 125.2s

Have you ever looked at your agent execution and asked yourself the question, why did it do that? What if it had done a different thing? Would it have been cheaper? Would it have been faster? Well, you can do all these things if your agents have a save button. We've had the save button for documents for decades now. Since the 1980s, people have been used to pressing Control-S, Command-S, or autosaving while you're working, to have a persistent state. But agents don't have that today. The only thing we have that is closest is a trace. A trace gives you the emitted telemetry data of how an agent calls tools and the input and output of that state. Now, while this is a good start, it is actually very disconnected from the runtime in which these agents actually execute. So all the variables that are in state, all the file system that is in flight, the decisions that it makes in the code, the actual code itself, all of that is lost, and it is only stamped as a read-only trace by the end, which is sitting in another tool far away from where the actual code is. And I think this is what's missing today in the industry: we don't have a clear connection between the observability spans that are emitted with Odell and the execution. And maybe at this point, you might be wondering, but why even bother? Why do I need to have a save button? Well, save allows you to replay. You can go back in history and ask the what-if question. What are the types of questions you might want to ask? You might want to swap the model. Maybe you use an open-source model that is cheaper. Maybe you mock a tool and you override what it returns. Maybe you degrade it intentionally to see what would happen if things are wrong. And these sorts of questions are only possible if you have that state. And there is a category of the stack emerging which actually allows that. This sits on top of the harness, sits on top of the frameworks that allow you to create agents, and puts a durable runtime below that and augments the traces that are emitted with the actual code execution and the things that are around it to actually complete the state of the system. And the good news is that once you have such a system, in production you already have the information you need to ask those questions that are relevant to making your system better, cheaper, and faster. Production already has the traces. It already has these state checkpoints, ideally from the runtime, that can allow you to go back in time and ask those questions. For example, let's take an agent example which does a customer resolution and refunds after a chargeback dispute. You can then see if the order status is changing languages or whether the request should have been escalated, or maybe a smaller model would have handled it, if the runtime is checkpointing each of the state as it goes along, almost like an autosave, a Command-S, a Control-S in your agent. And once you have that, you can even close the loop. This conference is all about loops, so this is nothing different. You have a cohort of runs that you think maybe matter because maybe they're too expensive, they took too long. You replay a change, you diff it, you see what would have happened when you have the baseline, which you know what happened in the first place. And then you decide and you route and you ship back. That's closing the loop on your evals. It's evaluating using your production traces. So it's evaluating using your production checkpoints. So checkpoint, replay, diff, decide. And this is really the methodology that I've seen, and I've seen others do, which has really scaled. For example, DoorDash. DoorDash has a blog post on the 1st of June where they talk about having a simulated environment where they replayed customer bots and they've done what-if scenarios and seen how they could have made it better. And where it used to take them hours and hours to do this, now they've reduced it to five minutes with hundreds of simulations, have 90% less hallucinations, and they're still two points within what they've seen in production. So the simulations are pretty good because they're grounded in what's already happened. So we're going to just walk through this in an example, and we're going to see how this works. For this demo, we're going to be using a tool called Kitaru. Kitaru is a very new tool that is launched by the team at ZenML. ZenML has been around for many years and is a player in the orchestration space. I'm one of the co-founders. And Kitaru is something we've launched recently, which allows you to have a runtime layer below your harness layer and also connect to your traces and do all the checkpointing that we were just talking about, and then run replay scenarios. So you can see here, I have all of my... I have a support agent which looks at my customer requests and escalates it when it needs to, to humans. You can see various things as you might expect. Every tool call, I can see a timeline view if I wanted. And here, the difference is that if I click on a particular checkpoint like a tool call, I can see the configuration where it ran, the code it took, and the artifacts which came in and out of it. And this combination of code and the artifacts that it created and the environment in which it ran, whether it was a Docker image or a sandbox, those are all snapshotted in state here between the checkpoints. And you can see that here in this particular example, there were a few tool calls every time it went to the LLM, and I can actually see how long it took. But imagine I wanted to do something different. Imagine I wanted to change the model. Maybe at this point, I wanted to use a cheaper model. Would it have done the same stuff afterwards? To do this in Kitaru, it's quite easy. All you have to do... Now, to do this in Kitaru is very easy. All you have to do is take your execution and replay it at a particular point. So I'm just going to copy this over, and I'm going to put it in my terminal and see what happens. So here what I'm doing is essentially saying, okay, after this particular tool call, change the model to GPT-5-nano, which is obviously a bit cheaper. And then what's going to happen here is there's going to be a new execution that's launched from 71, and you can see the first three checkpoints are skipped. So they're all skipped because Kitaru already has the state of all the checkpoints before that. All it needs to do is just change this particular checkpoint and then start executing from here. And this is really cool because now I can see what would have happened in this scenario if there was a cheaper model. So this looks pretty similar to me, but what if I wanted to do an even different change? What if I wanted to mock a tool or if I wanted to change a tool call? In order to do that, you can mock up, for example, the lookup policy tool. And here it's also very easy. I can just go back here, replay. And this time, rather than changing the model, I am changing the lookup policy. And I'm mocking it with another function in my code base, which returns a different lookup policy. And I'm just trying to understand if the policy had changed, what would have happened? This time I'm holding the model constant. Now, the interesting thing here is that because I have the code, it's very easy for me to do tool calls and to change these particular things and do more experiments than I would have had if I was completely disconnected from the code base. And here you can see that the code base is a little bit different. So I see my logs here. They look a bit different. And here you can see that I got a slightly different artifact, maybe from the tool call, and it published the thing. So now again, I have three runs. I have the original run and I have the two replays. Now, what if I wanted to see them side by side? So Kitaru actually has this very handy diff command that lets me give an original ID of an execution and then allows me on the other side to actually see them side by side so I can see what happened. So there's a few warnings here about some artifacts, but in a second it will go and give me a URL. I can copy this URL and I can actually put it directly here. And now I have a very nice comparison of the original and the two forks. I can see a bunch of things here, but I think what's really interesting is this view. So in this view, you can see that the first part of the baseline is the same. So these things were skipped. They were exactly as they were. The state is exactly where it was. But now right after that, in the third replay, it's a little bit different. The tool call happened a bit differently because we used a different policy, and then something changed after here. Here it took a little bit longer, and you can really start looking at the final result. And if I click on the final result, I can actually see the artifact side by side. I can see what decision it actually ended up making. So you can see that here it restricted the account charge. Here it also restricted it. Here it also restricted it. The first two, it actually needs review. And the third, because we changed the policy, is safe to answer. So this might or might not be good in your scenario, but is this what you expected to see? Maybe, because these two were cheaper at the end of the day. For the number of tokens consumed, they were cheaper, input and output. And you can see a bunch of detail here in the UI, which might be useful for these analyses. Okay, but this is just one point. What if I wanted to do this across a cohort? What if I wanted to have a bunch of runs which actually may be sorted by cost? So maybe I took all of my expensive ones, and I want to do one change across the entire cohort. Change all the tool calls that happened for this particular set of configurations across the cohort, or change the model itself and use a cheaper model across the cohort. So this gives me a bigger distribution of information and replays that I can use. And now I can just replay this. And the way you would replay it is you can use the replay many command, where you actually start from a particular point and you do the same as we did before, just across the cohort. This is going to take a while, so I'm not going to do it now. But once you execute it, in this particular case I just emitted it to JSON. So I have a very nice JSON here. And this JSON gives me a bunch of things. So this is a bit hard to do with the UI. There's a lot of things going on. You can't just do many, many, many comparisons here. But what you could do, and what I love to do personally, is use the Kitaru MCP server. And here what you do is you can just say, hey, read this JSON report and do an analysis on what you think I should be doing across the cohort. And I think this is what is really important, is using agents and LLMs to analyze these cohorts across a plethora of data. Because at some point, 10 is probably easy to do, but what if you have thousands? And doing thousands and thousands is hard. And this is where skills and MCP servers get really relevant. And having the runtime be queryable and go into your execution and fetch the artifacts is very important. So it's going to be doing a lot of things. It's trying to run an analysis around these decisions, and it's going to flag any red flags that happen. So while this is going on, maybe we can go back and continue our presentation. So here you can see, yes, if you change the model, it can get very cheap. And of course, you don't want to do it across just one sample, but across many, because then you can see a bigger variation. I think one thing which is important here is, and to be very honest, while we've been doing this with our users and customers, what I've personally seen is that having a naive model swap usually or oftentimes doesn't work. So just changing, for example, to a cheaper model and just looking at the cost one single-dimensionally, obviously it could be that you're spending a lot less money. But what happens if your support bot is not resolving the requests? So this is a study from Braintrust, an excellent study where they actually looked at that. And they saw that there could be a false economy if you do a naive model swap, because it might look on paper that you're faster and you're cheaper. But at the end of the day, you have to look at the value created. So it's a tradeoff between how much money you want to spend and the result. And also, if you look at the Tau bench, what we must understand is that a model that passes 60% of the time is only self-consistent about a quarter of the time. Which basically means that one replay is just an anecdote, and having a cohort analysis is way, way, way better. Because then you can really see across a population and across scale what would have happened, not just looking at one estimate. This can get very expensive, of course, and this is where you have to be really smart about what you replay and have tooling that really helps you. And this you can really bake into your production process as well. Again, if you take a step back and look at the playbook, you can start from real runs, not synthetic, but real runs, real production, checkpointed state, build cohorts that matter. Maybe take the expensive ones, maybe take the long ones, maybe take the risky ones. The first thing is, never ship anything by just replaying one or two things, and just do this at scale and ship, route, and hold, and try to automate that loop as much as possible. Maybe there's an agent that's doing that for you. That's even better. But you just have a human in the loop at the end. Or you just have a human in the loop at the end. Now, let's see where our cohort analysis has gotten. Okay, so we're done. So the verdict is don't ship. Even though it looked like from a single replay that it was cheaper to do and we reached the same result, across a bunch of those support cases, you actually saw that our agent concludes that you shouldn't be using a cheaper model in this particular case for your data. But this might be different for your data. In conclusion, if you want to be replaying your agent executions and answering the questions, what if I had done something different while designing this agent? Or what if the agent could be driven to do something different? You can do this if you model your agent with your harness in a runtime that can checkpoint state and is able to replay that state from code with different scenarios. If you want to use Kitaru, the tool that I showed that allows you to do it, you can scan the repo. It's open source, free to use, and we'd appreciate the feedback and love. Thank you so much, and see you guys on the next one. is changing languages or whether the request should have been escalated or maybe a smaller model would have handled it if the runtime is checkpointing each of the state as it goes along, almost like an autosave, a command S, a control S in your agent. And once you have that, you can even close the loop. This conference is all about loops. So this is nothing different. You have a cohort of runs that you think maybe matter because maybe they're too expensive, they took too long. You replay a change, you diff it, you see what would have happened when you have the baseline, which you know what happened in the first place. And then you decide and you route and you ship back. That's closing the loop on your evals. It's basically evaluating using your production traces. So it's basically evaluating using your production checkpoints. So checkpoint, replay, diff, decide. And this is really the methodology that I've seen and I've seen others do, which has really scaled. For example, DoorDash. DoorDash has a blog post on the 1st of June where they talk about having a simulated environment where they replayed customer bots and they've done what-if scenarios and seen how they could have made it better. And where it used to take them hours and hours to do this, now they've reduced it to five minutes with hundreds of simulations, have 90% less hallucinations, and there's still two points within what they've seen in production. So the simulations are pretty good because they're grounded in what's already happened. So we're going to just walk through this in an example and we're going to see how this works. For this demo, we're going to be using a tool called Kitaru. Kitaru is a very new tool that is launched by the team at ZenML. ZenML has been around for many years and is a player in the orchestration space. I'm one of the co-founders. And Kitaru is something we've launched recently, which allows you to have a runtime layer below your harness layer and also connect to your traces and do all the checkpointing that we were just talking about and then running replay scenarios. So you can see here, I have all of my... I have a support agent which looks at my customer requests and escalates it when it needs to to humans. You can see various things as you might expect. Every tool call, I can see a timeline view if I wanted. And here, the difference is that if I click on a particular checkpoint like a tool call, I can see the configuration where it ran, the code it took, and the artifacts which came in and out of it. And this combination of code and the artifacts that it created and the environment in which it ran in, whether it was a Docker image or a sandbox, those are all snapshotted in state here between the checkpoints. And you can see that here in this particular example, there was a few tool calls every time it went to the LLM, and I can actually see how long it took. But imagine I wanted to do something different. Imagine I wanted to change the model. Maybe at this point, I wanted to use a cheaper model. Would it have done the same stuff afterwards? So to do this in Kitaru, it's quite easy. All you have to do... Now, to do this in Kitaru is very easy. All you have to do is you have to take your execution and replay it at a particular point. So I'm just going to copy this over, and I'm going to put it in my terminal and see what happens. So here what I'm doing is essentially I am saying, okay, after this particular tool call, change the model to GPT-5-nano, which is obviously a bit cheaper. And then what's going to happen here is there's going to be a new execution that's launched from 71, and you can see the first three checkpoints are skipped. So they're all skipped because Kitaru already has the state of all the checkpoints before that. All it needs to do is just change this particular checkpoint, and then start executing from here. And this is really cool because now I can see what would have happened in this scenario if there was a cheaper model. So this looks pretty similar to me, but what if I wanted to do an even different change? What if I wanted to mock a tool or if I wanted to change a tool call? Well, in order to do that, you can mock up, for example, the lookup policy tool. And here it's also very easy. I can just go back here, replay. And this time, rather than changing the model, I am changing the lookup policy. And I'm mocking it with another function in my code base, which returns a different lookup policy. And I'm just trying to understand if the policy had changed, what would have happened? This time I'm holding the model constant. Now, the interesting thing here is that because I have the code, it's very easy for me to do tool calls and to change these particular things and do more experiments than I would have had if I was completely disconnected from the code base. And here you can see that the code base is a little bit different. So I see my logs here. They look a bit different. And here you can see that I got a slightly different artifact, maybe from the tool call and it published the thing. So now again, I have three runs. I have the original run and I have the two replays. Now, what if I wanted to see them side by side, right? So kitaru actually has this very handy div command that lets me give an original ID of an execution and then allows me on the other side to actually see them side by side so I can see what happened. So there's a few warnings here about some artifacts, but in a second it will go and give me a URL. I can copy this URL and I can actually put it directly here. And now I have a very nice comparison of the original and the two forks. I can see a bunch of things here, but I think what's really interesting is this view. So in this view, you can see that the first part of the baseline is the same, right? So these things were skipped. They were exactly as it is. The state is exactly where it was. But now right after that in the third replay, it's a little bit different, right? The tool call happened a bit differently because we used a different policy and then something changed after here. Here it took a little bit longer and you can really start looking at the final result. And if I click on the final result, I can actually see the artifact side by side. I can see what decision it actually ended up making. So you can see that here it restricted the account charge. Here it also restricted it. Here it also restricted it. The first two, it actually needs review. And the third, because we changed the policy, is safe to answer. So this might or might not be good in your scenario, but is this what you expected to see? Well, maybe, because these two were cheaper at the end of the day. For the number of tokens consumed, they were cheaper, input and output. And you can see a bunch of detail here in the UI, which might be useful for these analyses. Okay, but this is just one point, right? What if I wanted to do this across a cohort? What if I wanted to have a bunch of runs, right, which actually may be sorted by cost. So maybe I took all of my extensive ones. And I want to do one change across the entire cohort. So change all the tool calls that happened for this particular set of configurations across the cohort. Or change the model itself and use a cheaper model across the cohort. So this gives me a bigger distribution of information and replays that I can use. And now I can just replay this. And the way you would replay it is you can use the replay many command, where you actually start from a particular point and you do the same as we did before, just across the cohort. This is going to take a while, so I'm not going to do it now. But once you execute it, in this particular case, I just emitted it to JSON. So I have a very nice JSON here. And this JSON gives me a bunch of things. So this is a bit hard to do with the UI, right? So there's a lot of things going on. You can't just do many, many, many comparisons here. But what you could do, and what I love to do personally, is use the Kitaru MCP server. And here what you do is, you can just say, hey, read this JSON report and do an analysis on what you think I should be doing across the cohort. And I think this is what is really important, is to be using agents and LLMs to analyze these cohorts across a plethora of data. Because at some point, I mean, 10 is probably easy to do, but what if you have thousands? And doing thousands and thousands is hard. And this is where skills and MCP servers get really relevant. And having the runtime be queryable and go into your execution and fetch the artifacts is very important. So it's going to be doing a lot of things. So it's trying to run an analysis around these decisions. And it's going to flag any red flags that happen. So while this is going on, maybe we can go back and continue our presentation. So here you can see, yeah, yes, if you change the model, it can get very cheap. And of course, you don't want to do it across just one sample, but across many. Because then you can see a bigger variation. I think one thing which is important here is, and to be very honest, is that while we've been doing this with our users and customers, what I've personally seen, is that having a naive model swap usually or oftentimes doesn't work. So just changing, for example, to a cheaper model and just looking at the cost one single dimensionally, obviously, it could be that you're spending a lot less money. But what happens if your support bot is not resolving the requests, right? So this is a study from Braintrust, excellent study where they actually looked at that. And they saw that there could be a false economy if you do a naive model swap, because it might look on paper that you're faster and you're cheaper. But at the end of the day, you have to look at the value created, right? So it's a trade off between how much money you want to spend and the result. And also, if you look at the Tau bench, what we must understand is that a model that passes 60% of the time is only self-consistent about a quarter of the time. So, which basically means that one replay is just an anecdote and having a cohort analysis is way, way, way better. And because then you can really see across a population and across scale, what would have happened, not just looking at one estimate. This can get very expensive, of course, and this is where you have to be really smart about what you replay and have tooling that really helps you. And this you can really bake into your production process as well, right? So, again, if you take a step back and look at the playbook, you can start from real runs, not synthetic, but real runs, real production, checkpointed state, build cohorts that matter. Maybe take the expensive ones, maybe take the long ones, maybe take the risky ones. So, the first thing is never ship anything by just replaying one or two things and just do this at scale and ship route and hold and try to automate that loop as much as possible. Maybe there's an agent that's doing that for you. That's even better. But you just have a human in the loop at the end. Or you just have a human in the loop at the end. Now, let's see where our cohort analysis has gotten. Okay, so we're done. So, the verdict is don't ship. So, even though it looked like from a single replay that it was cheaper to do and we reached the same result, across a bunch of those support cases, you actually saw that our agent concludes that you shouldn't be using a cheaper model in this particular case for your data. But this might be different for your data. In conclusion, if you want to be replaying your agent executions and answering the questions, what if I had done something different while designing this agent? Or what if the agent could be driven to do something different? Well, you can do this if you model your agent with your harness in a runtime that can checkpoint state and is able to replay that state from code with different scenarios. If you want to use Kitaru, the tool that I showed that allows you to do it, you can scan the repo. It's open source, free to use, and we'd appreciate the feedback and love. Thank you so much and see you guys on the next one.