Your Agent Failed in Prod. Good Luck Reproducing It. - Tisha Chawla & Susheem Koul, Microsoft
Description
When an autonomous agent fails in production and corrupts an enterprise data record, it rarely repeats the exact same execution trajectory twice. Standard application logs reveal what broke but completely fail to explain why, leaving platform teams unable to reproduce non-deterministic failures on demand. While durable execution engines excel at keeping an agent loop alive through state recovery, durability is fundamentally distinct from debuggability. State recovery reconstructs the present; it does not allow an engineer to re-enter the precise historical run that caused an erratic state mutation. This session introduces the record and replay pattern for autonomous workflows, bringing the core engineering philosophy behind low level systems tools like Mozilla rr straight into the agent loop. By capturing every model invocation, tool execution payload, memory boundary read, and intermediate state transition into an append only event log, engineers can deterministically replay a failed execution trace for true postmortem root cause analysis. This architectural pattern moves entirely beyond basic API mocking or simple response caching. Attendees will leave this session knowing how to architect a framework agnostic recording layer, identify the exact state mutations required to guarantee replay determinism, understand where this approach complements durable execution architectures, and learn how to transform an unreproducible production anomaly into an execution path they can step through line by line. Speakers: - Tisha Chawla (Microsoft): Tisha Chawla is a Software Engineer at Microsoft working within the Commerce and Ecosystem Data Platform team, where she builds agentic systems designed to hold up against real production data. Her technical work spans core internal platform initiatives across Spec Driven Development, SRE Agent adoption, and enterprise SWE Agents, focusing on deterministic execution frameworks and agentic software development lifecycles. Alongside
Summary
Generated by claude-sonnet-4-5At-a-Glance
- Verdict: Watch fully
- Core thesis: Chasing bitwise determinism in production AI agents is futile; instead, instrument boundary-level recording (input/output pairs at each node) to capture execution traces, then replay and stub those traces offline for deterministic debugging and regression testing without model calls.
- Why it matters: Ken's systems will fail non-deterministically in production—this shows a concrete engineering pattern (Chronicle) to debug and test agent failures you can't reproduce via prompt re-runs.
- Best use: Study the Chronicle boundary annotation pattern, adopt trace-based regression tests for agent ops, and incorporate the 'replayability not determinism' mindset into your agent CI/CD workflows.
Executive Summary
Tisha Chawla and Susheem Koul from Microsoft present a hard-won lesson: when an agent misbehaves in production—selling 1,000 shares instead of $1,000 worth—engineers reflexively lower temperature to zero and try to reproduce the exact output. They show this fails because temperature=0 doesn't guarantee determinism at the system level due to GPU floating-point variance, batch invariance in inference servers, and mixture-of-experts routing dynamics. Running the same prompt 1,000 times can still yield dozens of different outputs even at temperature zero, making traditional 'pull logs, rerun locally' debugging impossible.
The speakers reframe the problem: the question isn't 'how do I make the model deterministic?' but 'how do I debug a run I can't reproduce?' They introduce Chronicle, a proof-of-concept framework that annotates code boundaries (LLM calls, tool calls, retrieval steps) to record input/output pairs, model versions, and metadata as a trace. When a production failure occurs, Chronicle freezes the entire agent run—including all intermediate node states—so teams can replay the exact execution offline, stub the non-deterministic LLM nodes with their recorded outputs, run the modified code (e.g., new guardrails) live against the frozen context, and assert correctness without calling the model again. This turns every production trace into a free, rerunnable regression test.
The talk distinguishes deterministic testing (guardrails, tool logic) from behavioral testing (tone, trajectory), arguing Chronicle solves the former by kicking probability out of the test loop. Key takeaways: stop chasing bitwise determinism through APIs; log session variables (model version, build ID, RAG chunks); capture the full envelope not just prompts; use replays to debug and verify fixes; and keep temperature variation alive in production because that's what enables agency. The speakers provide a QR code for Chronicle code and articles.
Key Takeaways
- Claim: Setting temperature to zero does not make agent outputs deterministic at the system level. | Evidence: Hard data from Reddit/Hacker News threads shows running the same prompt 1,000 times at temperature=0 still returns dozens of different responses due to GPU non-determinism, floating-point math non-associativity, batch invariance (requests grouped with random traffic alter matrix ops), and mixture-of-experts routing (tokens rerouted if batches overflow expert capacity). | Caveat: Temperature zero guarantees greedy argmax sampling, but underlying scores can shift run-to-run; the speakers do not quantify how many different outputs occur per 1,000 runs, only that variance exists. | Implication: Ken cannot rely on temperature=0 + prompt replay for debugging; he needs trace-based observability to capture what the agent actually did, not what it should deterministically output. | Timestamp: timestamp unavailable
- Claim: The wrong question is 'how do I make the model deterministic?'; the right question is 'how do I debug a run I can't reproduce?' | Evidence: Teams waste weeks chasing bitwise determinism and conclude 'the system is unknowable.' The speakers argue determinism was never the north star—debugging was—and replayability (rebuilding a run well enough to debug it) is achievable via recording, not via freezing the model. | Caveat: No specific team or timeline cited for the 'weeks burned' claim; this is anecdotal positioning. | Implication: Ken should prioritize instrumentation and trace capture over temperature tuning or prompt engineering for reliability; the engineering investment belongs in observability, not in futile determinism chasing. | Timestamp: timestamp unavailable
- Claim: Chronicle annotates code boundaries (methods) to record input/output pairs and metadata, freezing the entire agent run as a trace that can be replayed offline with zero model calls. | Evidence: Demo shows a stock-selling agent with three boundaries annotated: initial planning (LLM), place_order tool, and finalize_agent (LLM). When the agent mistakenly sold 1,000 shares (quantity) instead of $1,000 worth, Chronicle captured the bad tool call JSON and all node I/O. In the test replay, Chronicle stubbed the LLM's bad output but ran the updated guardrail code live, asserting the order was blocked this time. | Caveat: Chronicle is a proof-of-concept; no mention of production maturity, performance overhead, or integration with existing observability stacks (e.g., OpenTelemetry). The speakers do not discuss how Chronicle handles streaming outputs, async workflows, or distributed traces across services. | Implication: Ken can adopt the boundary annotation pattern in his agent code to auto-generate regression tests from production failures; every production trace becomes a free, deterministic test case once the model calls are stubbed. | Timestamp: timestamp unavailable
- Claim: Deterministic testing (guardrails, tool logic) should use Chronicle-style trace replay; behavioral testing (tone, trajectory) should use LLM-as-a-judge techniques. | Evidence: Speakers distinguish two test categories: deterministic nodes where Chronicle 'kicks probability out of the window' by stubbing LLM outputs, and subjective behavioral checks where replay is insufficient and evaluator models are better suited. | Caveat: No examples or tooling provided for the behavioral side; the talk focuses exclusively on deterministic replay and does not integrate the two testing modes into a unified workflow. | Implication: Ken should use Chronicle for regression/guardrail tests and separate LLM-as-a-judge pipelines for qualitative checks; the two are complementary, not substitutes. | Timestamp: timestamp unavailable
- Claim: Record at the boundary (what enters/leaves each node), not at the network layer, because half your agent never touches the network (local retrieval, in-process tools, memory). | Evidence: The speakers state that network-layer logging misses local steps and breaks under streaming/async; boundary recording captures the meaning of each step, not the packets. | Caveat: No quantification of 'half your agent' or discussion of how to handle async boundaries, partial failures, or non-serializable state (e.g., open file handles, DB cursors). | Implication: Ken's instrumentation strategy should annotate agent code boundaries—not just API calls—to capture the full execution context; relying on HTTP middleware or LLM provider logs will leave blind spots. | Timestamp: timestamp unavailable
Detailed Brief
Why temperature=0 fails to deliver determinism
- Claims: Temperature zero means always pick argmax, but underlying scores shift run-to-run.; Floating-point math is non-associative; timing shifts in matrix ops flip the winning token.; Batch invariance: requests grouped with random traffic alter computation paths.; Mixture-of-experts routing: tokens rerouted if batches overflow expert capacity.
- Evidence: Reddit/Hacker News threads show 1,000 runs at temperature=0 yield dozens of different outputs.; Running the same matrix multiply alone on GPU 1,000 times yields identical bits, proving it's not a concurrency bug.; The real culprit is batch invariance—your request's computation depends on what else hit the server that millisecond.
- Caveats: No hard numbers on variance magnitude (e.g., 5% different? 30%?).; No distinction between open-source local inference (more controllable) vs. hosted API (less controllable).; The speakers do not discuss whether batch size=1 or single-request inference eliminates variance.
- Implications: Ken cannot debug agent failures by re-running prompts locally; he must capture the original execution state.; Hosted APIs (OpenAI, Anthropic) will never offer bitwise determinism; self-hosted inference might reduce but not eliminate variance.; Teams should abandon temperature tuning as a debugging strategy and invest in trace-based observability instead.
Chronicle: boundary annotation and trace replay
- Claims: Chronicle annotates methods with @boundary to record input/output pairs and metadata (model version, build ID).; Each production run is frozen as a trace; nodes can be stubbed during replay to isolate code changes.; Replay mode enables deterministic CI: rerun the exact failure offline with zero model calls.
- Evidence: Demo: stock-selling agent with three boundaries (planning LLM, place_order tool, finalize LLM).; Production failure: agent sold 1,000 shares instead of $1,000 worth; Chronicle captured the bad tool call JSON.; Test replay: Chronicle stubbed the LLM's bad output, ran the updated guardrail live, and asserted the order was blocked.; Speakers show JSON hyper-detail for each node: input schema, output, metadata (model version, sampling params).
- Caveats: Chronicle is a proof-of-concept; no discussion of production readiness, performance overhead, or storage costs.; No mention of how Chronicle handles streaming responses, async/concurrent nodes, or distributed traces across microservices.; No integration examples with existing observability stacks (OpenTelemetry, Honeycomb, Datadog).; The speakers do not address how to version or manage large trace datasets or how to decide which traces to keep.
- Implications: Ken can auto-generate regression tests from production failures by annotating agent code boundaries.; Every production trace becomes a free, rerunnable test case once the model calls are stubbed.; Chronicle's boundary pattern is orthogonal to framework (LangGraph, Autogen, custom); it wraps any Python method.; Teams can iterate on guardrails/tools without burning API credits or waiting for model inference during test runs.
Deterministic vs. behavioral testing
- Claims: Deterministic testing applies to guardrails and tool logic; Chronicle freezes LLM outputs to kick probability out of the test loop.; Behavioral testing measures subjective qualities (tone, trajectory) and requires LLM-as-a-judge techniques.
- Evidence: Chronicle demo shows deterministic assertion: tool output == 'blocked'.; Speakers state 'tone of the agent or whether the trajectory it took was right' is more subjective and better suited to evaluator models.
- Caveats: No examples or tooling for behavioral tests; the talk does not show how to integrate LLM-as-a-judge into the Chronicle workflow.; No discussion of when to use which test type or how to balance coverage.
- Implications: Ken should use Chronicle for regression/guardrail tests and separate LLM-as-a-judge pipelines for qualitative checks.; Deterministic tests are fast, free (no model calls), and repeatable; behavioral tests are slow, expensive, and probabilistic.; A complete agent test suite requires both modes; Chronicle solves only the deterministic half.
Five key takeaways from the speakers
- Claims: Stop chasing bitwise determinism through the API—fundamental API design makes it impossible.; Log session variables: LLM version, build ID, RAG chunks.; Capture the full envelope, not just the prompt—there are many ingredients in the final response.; Use replays to debug, fix failures, then use the same trace as a test case.; Keep generation-time variation alive; don't pin temperature to zero—that's what brings agency to the agent.
- Evidence: Speakers provide a QR code for Chronicle code and articles.; Final slide: 'We hope the traces you put in today make your on-call cycle tomorrow much better.'
- Caveats: No discussion of how to balance temperature variation (for creativity) with reliability/safety requirements.; No guidance on which session variables matter most or how to surface them in dashboards.
- Implications: Ken should treat temperature/top_p as inference-time knobs, not debugging levers.; Session variable logging (model version, code version, RAG chunks) should be instrumented from day one; retrofitting is painful.; Trace-based debugging is the path to SLAs and on-call sanity for agent systems.
Notable Concepts & Terms
- Bitwise determinism vs. replayability: Bitwise determinism = same input, same output (controllability); replayability = rebuild a run that already happened well enough to debug it (observability). Speakers argue you don't need the former, you need the latter.
- Boundary annotation: Chronicle's core abstraction: annotate any method (LLM call, tool call, retrieval) to record input/output pairs and metadata; boundaries become the unit of replay and stubbing.
- Batch invariance: The real culprit for non-determinism: your request's computation depends on what else hit the inference server that millisecond, altering matrix operation order and final token scores.
- Mixture-of-experts (MoE) routing: Experts have strict capacity limits; if a batch overflows a subnetwork, tokens get rerouted. Whether a token makes the cut depends on the traffic you got batched with, introducing non-determinism.
- Deterministic CI: Chronicle's replay mode: rerun the exact production failure offline with zero model calls by stubbing LLM nodes; enables fast, free regression tests.
- Deterministic vs. behavioral testing: Deterministic tests (guardrails, tool logic) use Chronicle trace replay; behavioral tests (tone, trajectory) use LLM-as-a-judge. Both are necessary; Chronicle solves only the former.
Operator Notes / Why Ken Should Care
- Ken's agent systems will fail non-deterministically in production; Chronicle's boundary annotation pattern is the most concrete mitigation strategy in this talk—adopt it for trace-based debugging.
- The 'temperature=0 still varies' claim is critical for agent ops: if Ken's team is burning cycles tuning temperature for reproducibility, redirect that effort to instrumentation and observability.
- Chronicle is proof-of-concept, not production-ready; Ken should evaluate whether to fork/extend Chronicle or build similar boundary recording into his own agent framework (LangGraph, Autogen, custom).
- Trace-based regression tests (replaying production failures with stubbed LLM outputs) are a major unlock for agent CI/CD; every production failure becomes a free test case.
- The speakers do not address cost/storage of traces, which matters at scale; Ken should plan retention policies and sampling strategies (e.g., keep all failures, sample 1% of successes).
- Chronicle's pattern is orthogonal to LLM provider and agent framework; it wraps any Python method, so it can be retrofitted into existing codebases.
- The talk implies Chronicle is open-source (QR code for code access); Ken should check if it's on GitHub and whether Microsoft is actively maintaining it or if it's a research artifact.
- For Ken's investing/GTM use cases: if agents interact with financial APIs (brokerage, trading), the $1,000 vs. 1,000 shares bug is a real risk; Chronicle-style guardrails + replay tests are essential for compliance and customer trust.
Watch Map
- timestamp unavailable: Intro: production agent failure scenario (sell $1,000 becomes sell 1,000 shares at $190/share = $190k disaster)
- timestamp unavailable: Why temperature=0 fails: GPU non-determinism, batch invariance, MoE routing, floating-point math
- timestamp unavailable: Reframe: wrong question (determinism) vs. right question (debugging/replayability)
- timestamp unavailable: Chronicle overview: boundary annotation, trace recording, metadata capture
- timestamp unavailable: Live demo: stock agent failure trace, JSON hyper-detail, replay with stubbed LLM + live guardrail, assertion passed
- timestamp unavailable: Deterministic vs. behavioral testing distinction
- timestamp unavailable: Five key takeaways, QR code for Chronicle code/articles
Source/Metadata
- Title: Your Agent Failed in Prod. Good Luck Reproducing It. - Tisha Chawla & Susheem Koul, Microsoft
- Transcript words: 2747
- Duration seconds: 849
- Timestamp note: Timestamps not present in transcript; watch_map is conceptual
Transcript
Imagine something your agent didn't produce was wrong. Coil the wrong tool, it wrote the wrong thing, and now suddenly your team is on call rotation to figure out what actually went wrong. Pretty common, right? Now, as per the standard engineering response, your gut will tell you to pull the raw prompt from the telemetry logs, pass it to the same model using the same prompt and run it locally to isolate the bug. Which we'll all do. And surprisingly, it will work as well. Run it again, it will work again. You run it 10 more times, it will be just perfect every time. But now let's talk about that one run which costed you, and that will be gone. You can reproduce it. And if you can reproduce it, you can debug it. And if you can debug it, you can promise it won't happen to your next customer or user, right? Now, I am Tisha. I have Sushin with me as my co-presenter. We won't run agents against real production backends, the kind of place where a bad write isn't, oh well, run it again. It's you on a call with a customer explaining where the data actually went. This whole talk is going to be about that one thing. You lose the second an agent goes, hey, buy in production, which is being able to reproduce it. That will be a not start for the next 10 minutes to follow. Now let's look at how this actually blows up. You've got an agent hooked to a broker API, which is the scenario I'm taking. The user says, hey, sell $1,000 of stock. Now comes the interesting part. Instead of doing the math, the agent sells the raw number 1,000 and dumps it straight into the quantity field. Guess what? It sells 1,000 shares instead. Now at $190 a share, $1,000 in 10 will become how much? $190,000 disaster, right? And the very side part is that the API on my infrared returned a clean 200 OK in 30 milliseconds. We got zero exceptions, zero alerts. If you see the trade, it's completely wrong. But your dashboards are sitting there, perfectly green, perfectly flawless. When such a scenario as we last discussed comes up, what's the first thing which you will do to try and fix this? The reflex here is to just turn the model temperature down to absolute zero. Assuming greedy decoding will make everything deterministic, right? But that's a complete misconception. Setting the temperature to zero doesn't fix a broken reasoning path. It just means the model is going to make the exact same logical error, the exact same way, at the exact same time. And honestly, even worse than that. To back up the scenario we just discussed, look at the engineering threads on Reddit and Hacker News. The hard data shows that temperature zero isn't even truly deterministic on a hardware level. Running the same prompt a thousand times can still return dozens of completely different responses just due to the underlying GPU non-determinism and the MOE architectures which are there. So to understand why this actually happens, we'll have to look at it from first principles. It comes down to four simple things. One, sampling determinism isn't system determinism. Temperature zero just means always take the argmax, but it doesn't guarantee that the underlying scores stay identical run to run. Two, floating point math isn't associative. The order you add your decimal matters, right? But a timing shift in matrix operation alters the final logic and which in turn will flip the winning token. Three, it's not a concurrency issue. Run the same matrix multiplication alone on a GPU a thousand times and I'll guarantee you'll get the exact same bits. So the real culprit is batch invariance here because a request gets grouped with whatever else hits the server that millisecond. Four, mixture of experts. Routing has the exact same bottleneck. Experts have strict capacity limits. If a batch overflows a specific subnetwork, tokens get rerouted. Whether a token makes the cut depends entirely on the traffic you got batched with. So the ultimate takeaway here is that chasing text output is a losing battle. Completely. We don't need the model to return the exact same token back every time. We just need our system to execute the exact same state transition. Which means we've been asking the wrong question all along, right? The wrong question is how do I make the model deterministic? And I've seen teams burning weeks on that and walk away deciding the systems just unknowable. The right question is how do I debug and retest a run I can't reproduce? Because determinism was never the north star. Debugging was. Two words we keep mixing up, which I'll talk about now, is bitwise determinism and replayability. Bitwise determinism is same input, same output. That's controllability. You're not getting it from a hosted API and you don't actually want it. Because the randomness is what makes the model good. Once the model explores more, you'll get more creative answers. The other one is replayability, which is rebuild a run that already happened well enough to debug it. That's observability. You don't need the model deterministic. You need the run recorded. And you don't freeze the model. You capture what it did. Another question which we are all thinking about is where do you record? First, you're not at the network layer. Because half your agent will never touch the network. The local retrieval, the in-process tools, the memory and the parts that do not shred under streaming and async. Record at the boundary instead. Because you need to capture what enters each node and what leaves it. The meaning of each step and not the packets. What replay adds here is a deterministic CI where you stop the model, you'll rerun the exact failure offline with zero model calls. Let's talk about the loop end-to-end now. It starts with annotation, recording, visualization, understanding, fixing. Then the part we're working on, which is replaying. And finally, verify. All right. Let's now see how to bring the workflow we just discussed into action. [SPEAKER_00] So we've established that replayability is a core tenet of productionizing any AI agent. [SPEAKER_00] But how do we build this in code? The meaning of each step and not the packets. What replay adds here is a deterministic CI where you stop the model, you'll rerun the exact failure offline with zero model calls. Let's talk about the loop end-to-end now. It starts with annotation, recording, visualization, understanding, fixing. Then the part we're working on, which is replaying. And finally, verify. All right. Let's now see how to bring the workflow we just discussed into action. [SPEAKER_00] So we've established that replayability is a core tenet of productionizing any AI agent. [SPEAKER_00] But how do we build this in code? [SPEAKER_00] As a proof of concept for this, what we have done is we have built something called Chronicle. [SPEAKER_00] At the heart of Chronicle lies the concept of a boundary. [SPEAKER_00] Think of a boundary as a bounding box around any node in your agentic workflow. A node can be a tool call. It can be a call to an LLM or a retrieval from a RAG. It doesn't matter. As long as it's a method, it can be annotated with the boundary annotation. Now, what does this annotation do? It ensures that anything that goes into the method and comes out of the method gets recorded. So any input and output pair will get recorded. On top of that, you can define parameters like your model version or the version of the code that is running. So that the entire state during which the agent run happened gets frozen and saved as a trace. Now, let's see this in action. We've been talking about this stock selling agent, which went haywire in production. This is a representation of the same. You have your initial planning step, which takes into account the user input. It can use the place order tool to do the actual selling and buying of the stocks. And then finally, it delegates to the finalized agent, which generates a succinct response for the end user. We have annotated all three of these methods with the boundary annotation. The first one is a tool and the second and the third one is an LLM. Let's run this. So here's what happened. You gave a user request to sell about $1,000 worth of ACME stocks. You have the three nodes. You have the input and the output recorded for every one of them. You can see that the LLM mistook the thousand as the quantity and generated a tool call of place order with the symbol ACME and quantity thousand. And the place order tool obviously executed this input and sold thousand dollars or thousand units of ACME stock at $190 per piece. This is where the problem started. Now you have a trace for it. But is this all you can see? No. We record much more details than this. So you can go and see the hyper detail JSON for each of the nodes that the call went into or went out of. You can see the metadata like the model version, the sampling versions that was there in the LLM call. You can see the input and the output. For example, this is for the place order tool. You can see the input went in as a sell of thousand quantity on ACME. And the output was obviously that it sold that quantity. And if you go one step back, you can see the agent one node creating that tool call that caused this entire problem. It created a tool call to place order with the symbol ACME and quantity thousand. Now this is all good. You have a recording. You have a trace. You figure out what the problem is. Now what do you do with it? You figure out that the problem is the LLM created a wrong tool call. You cannot control the LLM. That's what this entire discussion is about. You cannot enforce bitwise determinism. What you can do is you can put guardrails on your tools to enforce some level of credibility on your production agent. But how do you test this? Once you have built the guardrails, how do you test it? That is another tenet that Boundary offers. Since the Boundary annotation is already providing a bounding box around your methods, it can be used to stub your methods during testing. So think of it like this. You have a run which is recorded. That run recorded every input and output for every node. Now you have fixed your code at let's say the tool level, but you want the rest of the nodes to be stubbed so that the entire exact stack trace remains the same. How do you do it? You run a test suite with the same trace that was recorded earlier. You stub every node other than the node that you changed. And you let Boundary handle the rest. So let's see this one in action. This is your test case. You have loaded the trace and you have enabled the replay mode on Boundary, which allows Boundary to stub every node that you want. So for example, in this case, you want to stub the first agent call which generated the tool output. But you want the tool to be run live so that you can test out your code changes. Once you do this, you run your agent. Now another good thing about Boundary is that since it's already capturing the output of your tool as well, you can use it to run assertions. So you can take the output and assert that this time around the tool call got blocked. So let's run this once. Perfect. So you can see that agent 1, which was the LLM, was stubbed. You can see the same input-output as the recorded trace, but the tool and the LLM ran live. And you can see that for the tool, this time the output is blocked. And your assertion on the tool has passed because the order has blocked. So this is the power of merging your replayability traces with auto-generated testing and stubbing and assertions. So we just saw how Chronicle not only records your agent executions, but it also uses those recordings as test cases. Now, when we talk about testing in AI agents, I want to draw a very clear distinction here. There are two ways of testing AI agents, and both of them are equally important. There's the deterministic testing and then the behavioral testing. The deterministic testing applies to, obviously, the deterministic nodes of your agent graph. This could be your guardrails or your tool calls. Now, this is exactly where Chronicle shines. Because, as we just saw, Chronicle freezes the entire agent run as a context. So you can use the LLM nodes context to stub the LLM outputs. This essentially kicks the probability out of the window. I want to draw a very clear distinction here. There are two ways of testing AI agents, and both of them are equally important. There's the deterministic testing and then the behavioral testing. The deterministic testing applies to obviously the deterministic nodes of your agent graph. This could be your guardrails or your tool calls. Now, this is exactly where Chronicle shines. Because, as we just saw, Chronicle freezes the entire agent run as a context. So you can use the LLM nodes context to stub the LLM outputs. This essentially kicks the probability out of the window, and your entire agent run can become a test case. This is rerunnable, and since it never calls the model, it is free. On the behavioral side of things, you measure things like the tone of the agent, or whether the trajectory it took was right. This is more subjective, and this is where techniques like LLM as a judge are better off. Now, at this point, I want to talk about the key takeaways. First, stop chasing bitwise determinism through the API. The fundamental principles on which the APIs are built today do not make this possible. Second, know what are variables for your session, for example, your LLM version, or your build ID, or your rack chunks, and make sure that you are logging these. Third, capture the full envelope. Don't focus on just the prompt. There are a lot more ingredients that go into that final response. Fourth, use the replays to debug. Make sure that you find the issues, fix the failures, and then finally, use the same trace as a test case. Fifth and final, keep the generation-time variation alive. Don't try to pin the temperature to zero. After all, that is what brings the agency into your agent. So, the QR on your screen, you can scan that to get access to the code for Chronicle and a bunch of nicely written articles. And finally, thanks again for your time, and we hope that the traces that you put today in your agent make sure that your on-call cycle tomorrow is much better. Thank you. Second, know what are variables for your session, for example, your LLM version, or your build ID, or your rack chunks, and make sure that you are logging these. Third, capture the full envelope. Don't focus on just the prompt. There are a lot more ingredients that go into that final response. Fourth, use the replays to debug. Make sure that you find the issues, fix the failures, and then finally, use the same trace as a test case. Fifth and final, keep the generation-time variation alive. Don't try to pin the temperature to zero. After all, that is what brings the agency into your agent. So, the QR on your screen, you can scan that to get access to the code for Chronicle and a bunch of nicely written articles. And finally, thanks again for your time, and we hope that the traces that you put today in your agent make sure that your on-call cycle tomorrow is much better. Thank you.