Open Reader

Your Agent Didn't Fail. Your Harness Did. — Vinoth Govindarajan, OpenAI

completed 18:25 Jul 29, 2026 Watch on YouTube

Current Status

completed

Video ID

BInpv7lGp1o

RAG / Chat

Enabled
Your Agent Didn't Fail. Your Harness Did. — Vinoth Govindarajan, OpenAI
Description

Two runs touch the same session, the second write silently erases the first, and the agent keeps answering with total confidence from stale state. Nothing crashed and the model did not hallucinate, so this is a harness failure, the kind that lives in the system around the model rather than in the weights. Using OpenClaw as a public case study, Vinoth Govindarajan walks the usual suspects: state that was never persisted, overlapping writers with no single writer lane, a tool call that never returns because nothing set a deadline, and an approval that outlived the action it was supposed to authorize. The through line is that a model only proposes; the harness has to commit, and a receipt has to prove it. A transcript shows what the agent said, but a receipt is the evidence that survives: it records the mutation, the authority used, and whether the message actually reached the user, since an internal success that never becomes visible proof is its own failure. You leave with a run receipt audit to run on your own agents, five questions per incident: what woke it up, what state did it inherit, what authority did it use, what executed, and what evidence survived. Speaker info: - https://x.com/iamvinoth - https://www.linkedin.com/in/vinothgovindarajan/ - https://theagentstack.substack.com/ Timestamps: 0:00 - Introduction: harness failures vs model failures 1:32 - Delivery can succeed while the truth fails 2:46 - A model proposes, the harness commits, the receipt proves 4:14 - How events enter and state is rehydrated 5:48 - Idempotency, locks, and ordering 7:22 - Ownership: who persists the turn 8:28 - Single writer lanes and overlapping writes 10:09 - Time, deadlines, and cancellation 11:23 - Approval drift and bounded authority 13:05 - Internal success vs user-visible proof 14:08 - The run receipt audit: five questions

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Most production agent failures are harness failures—not model failures—because the system around the model must own state, serialize mutations, constrain authority, bound work, and prove user-visible outcomes.
  • Why it matters: This supplies a practical control-plane and reliability framework for action-taking agents, especially where silent state loss, concurrent writes, stuck tools, stale approvals, or unverified delivery can create confident but incorrect behavior.
  • Best use: Use it as an architecture review checklist for any agent workflow that retains memory, invokes tools, performs external actions, or runs across asynchronous channels.

Executive Summary

Vinoth Govindarajan argues that evaluating an agent primarily through model quality misses the production boundary that actually determines reliability. The model proposes an action, but the harness must assemble context, commit state transitions, authorize tool use, manage execution lifecycle, and retain evidence. His governing contract is: “a model proposes, the harness commits, and the receipt proves it.”

The talk uses publicly visible OpenClaw issues to illustrate five failure shapes: a delivered reply that was never durably recorded; concurrent writers silently overwriting each other; a tool call waiting forever for a result; an expired approval surviving as if still valid; and an internally successful send that never renders for the user. In each case, the model can remain coherent and the UI can look superficially healthy, making harness defects both more dangerous and harder to diagnose than obvious crashes.

The operational prescription is to give every fact a named system of record and replay path, use one ordered commit path per mutable state boundary, give external work explicit terminal states, bind approvals tightly to a specific action and lifetime, and construct a durable run receipt. The receipt should trace trigger, inherited state, authority, execution attempt, and evidence at the user-relevant edge—not merely record the model transcript or a tool's self-reported success.

This is a high-signal systems-design talk rather than a product pitch. It is particularly useful for designing or reviewing agent orchestration layers because it reframes familiar distributed-systems controls—idempotency, locking, timeouts, transactions, auditability, and ownership—for probabilistic planners that act across many event and tool surfaces.

Key Takeaways

  • Claim: A successful user-facing response is not reliable if the corresponding state cannot be replayed on the next turn. | Evidence: In the cited OpenClaw state-hole issue, a Telegram reply could be delivered while the router turn was not written to active context or the transcript; the user saw success, but the next turn lacked a durable record of the exchange. | Implication: For every fact an agent may later rely on, Ken should require a named owner/system of record and a defined replay path; storage location alone is not ownership. | Caveat: The issue is not that all delivery must be blocked on every persistence operation; the requirement is that the system clearly establish and verify the durable state boundary needed for future behavior.
  • Claim: Concurrency is compatible with agent systems, but each mutable state boundary needs one ordered commit path. | Evidence: The overlapping-writers example is a load-modify-save race: two callers read the same old state, each makes a locally correct change, and the second save silently erases the first—causing effects such as a dismissed commitment returning or duplicate follow-up. | Implication: Serialize only commits to shared mutable state, using a queue, transaction, lock, or mutex appropriate to the boundary; treat ordering as a user-experience property, not merely backend hygiene. | Caveat: The speaker explicitly does not recommend serializing the entire system: sub-agents, parallel reads, independent retrieval, and separate sessions may run concurrently.
  • Claim: Every external or asynchronous operation needs an explicit terminal outcome because silence is not a valid terminal state. | Evidence: In the dangling-tool-call case, a session contains a tool call with no matching result after a process death, dropped connection, or unrecorded timeout; later messages queue behind work that will never complete and the agent appears stuck. | Implication: Implement deadlines, cancellation, watchdog visibility, error results, maximum attempts, and recovery commands that can bypass the lane blocked by the failed work.
  • Claim: Model capability and a user's prior click do not constitute authority; approval must be a scoped, expiring execution state. | Evidence: The approval-drift example describes expired approved callback state that remained retrievable across restart and blocked later channel work: the button click existed, but valid authority did not. | Implication: Bind approval to the approver, session, run, specific tool, exact arguments, lifetime, outcome, and receipt; pair this with least-privilege tool surfaces and scoped credentials. | Caveat: Approval data can become unsafe or unverifiable when retries, replays, or callbacks lose its binding fields.
  • Claim: Internal tool success is insufficient evidence for actions whose purpose is a user-visible external effect. | Evidence: In the missing-edge-proof example, a message tool reported success in web chat or TUI, but the message did not render even though normal assistant replies appeared; the internal path accepted the request without proving the user saw it. | Implication: Do not let agents claim “already sent” or “completed” based solely on tool acknowledgments; collect confirmation—or an explicit failure to confirm—at the edge users actually care about. | Caveat: The required proof boundary depends on the product outcome: a calendar event, updated ticket, changed file, or rendered message each requires evidence from its relevant system edge.
  • Claim: A production agent should emit a durable receipt that makes every run inspectable across trigger, context, authority, execution, and outcome. | Evidence: The speaker's five-question audit asks: what woke the run, what state it inherited, which authority it used, what executed, and what evidence survived. He distinguishes a transcript (what the model said) and a tool result (what one component claimed) from a receipt that captures the verified chain. | Implication: Instrument agent runs as causally traceable state transitions, including trigger identity, context/policy/tool versions, action arguments, attempt number, idempotency key, external results, and edge confirmation.

Detailed Brief

Reference harness blueprint

  • Claims: The speaker presents a common architecture across personal and coding agents: events enter from chat, webhooks, timers, heartbeats, or external systems; a control plane maps each event to a session key; a session lane governs mutable state; runtime invokes models and tools; policies and approvals mediate actions; and an audit rail captures the run receipt.; The model is stateless in the relevant operational sense: each turn receives a harness-assembled working set that can include transcript, session state, memory, policy, and tool definitions.; A coherent answer does not demonstrate that context was complete or fresh; it can be fluent over stale or missing state.
  • Evidence: The speaker summarizes the blueprint as event, session key, throttle, tools, and audit.; He analogizes the model to a car engine and the harness to steering, brakes, road rules, dashboard, and a black box: capability without control is a liability.
  • Caveats: The talk is a system-design framework illustrated through OpenClaw issues, not a claim that OpenClaw uniquely has these problems or that its model is at fault.; The transcript's latter half substantially repeats the prior material, so the video has limited incremental content after its first full pass through the five failure modes.
  • Implications: Review context construction as a production dependency: missing transcript entries, stale policy versions, or inconsistent tool definitions are reliability defects even when model outputs appear sensible.; Separate the planner's intelligence from the control plane's responsibility for state transitions and side effects.

Focused production-path audit

  • Claims: The recommended starting point is one real production path in one agent system, rather than an abstract audit of every agent.; Causality requires identifying the trigger and its identity before reasoning about deduplication, ordering, or authorization.; The execution record should capture actual side-effect details rather than a post hoc natural-language summary of what the agent intended.
  • Evidence: Suggested trigger classes are user message, webhook, timer, tool result, sub-agent, and replay.; Suggested execution fields include tool/API call, arguments, attempt number, idempotency key, and external result.
  • Caveats: A receipt is only as useful as the evidence it retains; ending it at an internal component acknowledgment leaves the user-facing outcome unproven.
  • Implications: Use the audit to identify which failure class is structurally possible in a workflow before attempting prompt or model changes.; Prioritize receipt completeness for the most consequential action boundaries first, especially financial, messaging, ticketing, file, and calendar operations.

Notable Concepts & Terms

  • Model proposes, harness commits, receipt proves: The talk's central responsibility split: models suggest plans, while deterministic system infrastructure controls state transitions and generates evidence.
  • State ownership: A fact is reliable only when a named system of record can replay it; this distinguishes ownership from merely knowing where bytes are stored.
  • Session key / session lane: The session key defines the state boundary, while the lane provides a mechanism for one active ordered writer over that mutable state.
  • One ordered commit path: The narrow concurrency invariant: permit parallelism broadly but serialize changes at each shared mutable-state boundary.
  • Approval drift: A failure in which a historical human approval survives or is reused after its valid scope, action binding, or expiry has disappeared.
  • Run receipt: Durable evidence linking trigger, inherited state, authority, attempted execution, and confirmed external outcome; stronger than a transcript or tool result.
  • Edge proof: Confirmation at the boundary the user cares about—such as a rendered message or existing calendar event—rather than an internal acknowledgment.
  • Bound the work before the work bounds you: A lifecycle principle requiring timeout, cancellation, maximum attempts, and explicit terminal states for external operations.

Operator Notes / Why Ken Should Care

  • Select one high-consequence agent workflow and produce a sample end-to-end receipt for an actual run; identify any missing trigger identity, state version, authority binding, idempotency key, or edge confirmation.
  • Define a system-of-record matrix for agent-accessed facts: conversation turns, preferences, tasks, tickets, calendars, files, and external messages should each have one owner and an explicit replay method.
  • Audit per-session mutation paths for read-modify-write races; add a narrow transaction, queue, or lock at the commit boundary rather than reducing system-wide parallelism.
  • Add operational states and recovery paths for incomplete tool calls, ensuring retries and recovery commands cannot be indefinitely blocked behind the original stuck session work.
  • Treat approval tokens/callbacks as action-specific, expiring capabilities; reject approvals that cannot prove their binding to the current session, run, arguments, and tool.
  • Prevent user-facing completion claims unless the agent receipt contains the appropriate external confirmation or explicitly reports that confirmation is unavailable.

Source/Metadata

  • Title: Your Agent Didn't Fail. Your Harness Did. — Vinoth Govindarajan, OpenAI
  • Transcript words: 4131
  • Duration seconds: 1105
  • Timestamp note: No timestamps or chapter markers were present in the supplied transcript. The latter portion substantially repeats the main talk content.

Transcript

2474 words en Processed in 120.4s

Thank you for choosing to spend this session with me. My goal today is simple. I want to convince you all that most of the production failures are not, most of the agent failures are not model failures. Those are harness failures. So let's start with one production incident. The user saw the reply, the system forgot it happened. This is a failure shape I want to start with. Not a hallucination, not a crash, not a bad answer. The user's visible edge looked healthy, while the durable record had a hole. In this example, the user asked the agent to remember a refund for a customer. The assistant said it recorded the fact for the next turn. The interface looked normal, no red screen, no obvious failures. But the next turn cannot reconsider the fact. The user experienced success, the system inherited incomplete reality. Why this matters? The crash is annoying, but at least it gives you a boundary. You know something stopped. You usually see an error. You can often start from the last known good point. Silent success gives you a lie. Delivery can succeed while the persistent fails. The user has no reason to doubt the reply. The operator has no obvious alarm. The next turn can still sound confident because the model is coherent. But it is coherent over a broken history. That is why agent reliability matters. And agent reliability cannot stop at model quality. Hi, I'm Vinod. I work on core data and AI infrastructure at OpenAI. Before that, I worked on distributed systems at Apple and Uber. Outside of work, I write the agent stack, where I try to explain how production AI and data systems work under the hood. I'm not a member of OpenClaw, and this is not an OpenClaw product pitch. This is a pure system design talk. I'm using OpenClaw as a public case study because its issues, code, and docs make the harness around the agent unusually visible. Here's the production contract for the talk. A model proposes, the harness commits, and the receipt proves it. The model may suggest a message, a tool, an edit, or a command. But the model is not the production boundary. The harness owns the state transition, the authority check, the audit commit, and the receipt is the evidence that survives the turn. OpenClaw is the case study, and the contract is the takeaway. If you remember only three things from this talk, make it these. Own the state, order the mutation, and prove the action. A fact needs only one owner and one replay path. Shared mutable state needs one ordered commit path. And transcript is not the proof. A transcript tells you what the agent said. A receipt tells you what the system allowed, attempted, executed, and what the user-visible edge confirmed. To create a simple mental model, I created this car analogy of the harness. The model is the engine. It matters a lot. But nobody buys a production car by just looking at the horsepower alone. You also care about steering, brakes, road rules, dashboard, and a black box. The model gives you capability, but the harness gives you control. A powerful engine with no brakes is no autonomy. It is a liability with good acceleration. Here is the harness blueprint I wanted to discuss today. Every agent we know of, personal agents such as OpenClaw or Hermes, coding agents such as Codex, Cursor, OpenCode, or Claude Code, uses the same underlying architecture. Events enter from many surfaces: a chat, webhook, timer, heartbeat, or another external system. The control plane maps the events to a session key, and the session key determines the state boundary. The session lane gives you one active writer for the mutable state. The runtime calls the models and tools. Tools act through approvals and policies, and the audit rail becomes the run receipt. This is the blueprint: event, session key, throttle, tools, audit. The blueprint is the talk, and the incident is the proof that each boundary matters. Context is assembled. In agent runtime, it does not usually remember in a human sense. It is stateless. The harness rebuilds the working state for each turn. The working set may include the transcript, session state, memory, policy, and tool definitions. The model only sees what the harness supplies. If one input is missing or stale, the answer may still sound coherent. Coherence does not prove the working set was complete. These failures are not new. We already know about timeouts, retries, idempotency, locks, ordering, and state ownership. What changes is the agent setting. Now these failures sit around a probabilistic planner with dynamic plans. It's rebuilding the context for every turn. There are more event sources, and it can act through more action surfaces. So these failures are familiar. Agents make them easier to trigger and harder to explain. That's why agent harness matters. Let's talk about the first failure mode. This is the same failure mode I started this talk with. The user sees a success. The source of truth cannot replay it. Delivered is not remembered. In this state hole OpenClaw issue, a Telegram reply could succeed while the router turn was not written to the active context or transcript. The user saw the response. The log looked healthy, but the next turn had no durable record of the text change. A successful send proves delivery. It does not prove the future context. That distinction matters because the model can answer fluently over an incomplete record. The missing boundary was not intelligence. It was state ownership. By owner, I do not mean a person. I mean the system of record whose persistent state becomes the truth. A calendar event belongs to the calendar system. A support status belongs to the ticketing system, while a code change belongs to a workspace or repository. A conversation turn belongs to a session transcript. And a user preference belongs to a memory store. Storage tells you where the bytes live. Ownership tells you who can reconsider the reality. A replay is not reliable memory until a named owner can replay it. A system has to persist the turn. It has to name the owner or system of record. And it has to make the replay possible. The key question is simple. For every fact the agent might use later, who owns it and how would you replay it? If no owner can replay the fact, the system did not reliably remember it. Once we know who owns the state, the next question is who's allowed to change it and in what order. Two correct writes can still produce one wrong outcome. And last writer wins is not a consistency model. In this overlapping writer OpenClaw issue, a load, modify, save race is described. Two callers load the same old state. Each changes a different record. The second save silently erases the first. The user may see a dismissed commitment return or a severe duplicate follow-up. Neither writer is malformed. Both operations are locally correct. The missing boundary is serialization around the commit. The invariant is not no concurrency. That would be too slow, and it would miss the point. You can fan out sub-agents. Parallel reads are fine. Independent retrieval is fine. Many sessions can also run at once. The rule is narrower and simple. One ordered commit path for one mutable state boundary. This mechanism may be a queue, a transaction, or a lock. You can use locks or mutexes across sessions, and queues or transactions within a session. Be conservative at commit time, not across the whole system. Users do not see queues or locks. They see behavior. A lock correction feels forgetful. A stuck lane feels dead. And completion before delivery feels confused. Ordering is a product feature because users experience ordering bugs as personalities. Next, let's talk about time. In production, silence cannot be neutral. Let's review the next failure mode. The run waits for an event that cannot arrive. Silence is not a terminal state. In this dangling tool call issue, the session contains a tool call but no matching tool result. A process may have died. The connection may have dropped. A timeout might have happened before the result was recorded. The exact cause matters for debugging. The production failure is much simpler. The run is waiting for an event that will never arrive. New messages queue behind that silence. To the user, the agent simply looks stuck. Runs need deadlines and cancellation. A deadline bounds the wait. A watchdog makes the stuck work visible. Tools need time modes and error results. Channels need recovery commands that do not wait behind the stuck work they are trying to fix. Every external boundary needs an ending: success, failure, timeout, cancel, or max attempts. Most importantly, the receipt records the terminal outcome so the next step does not have to guess. Bound the work before the work bounds you. Now let's move on from state to authority, because a chat becomes risky when it becomes an action. Capability is not execution. The model can request an action. Requestability is not authority. Approval needs a shape. In this approval drift issue, expired approved callback state was still retrievable. The state stayed durable, survived restart, and blocked later channel work. The button click existed. The valid authority did not. This is the mistake: treating approval as a vague memory that the human was near the system or clicked yes. Approval is a scoped execution state. It must stay bound to the action it authorized. An expression must terminate rather than loop. A useful approval object answers who approved, in what session and run, for which tool and for which arguments, for how long, and with what outcome. It also points to the receipt. If those fields fall off during a retry, replay, or a channel callback, the harness can no longer prove the action being executed is the action being approved. The general lesson is simple. Capability is not execution. Least privilege narrows the tool surface. Scoped credentials ensure the right identity is used for the action. Approval and audit decide what happens before and after execution. The model can reason about the boundary, but it should not be the boundary. The model can request, but the system decides. Finally, even if the tool says success, the user-visible edge may disagree. The internal component reports success. The user-visible surface shows nothing. This is the inverse of the opening incident we saw. In this missing edge proof issue, the message tool reported success for a web chat or TUI run, but the message did not render. Normal assistant replies still appeared. The tool proved that the internal path accepted the request. It does not prove the user saw the result. That difference changes the conversation. The agent may later say, I already sent it, and the user may truthfully say, I never saw it. Internal success is not external proof. Proof is a chain, not a claim. The model proposed something, policy allowed or denied it, execution attempted it, and the user-visible edge confirmed or failed to confirm the outcome. The receipt preserves the chain. A transcript records what the agent said. A tool result records what one component claimed. A receipt records what the agent can verify at the boundary that matters. Let me recap all the incidents. Here are the five failure shapes to look for: a state hole, overlapping writers, dangling tool call, approval drift, and missing edge proof. For each one, ask the same question. What did the user see? Which boundary did it break? And what would the receipt have shown? Here is the audit I want you to run when you get back to your team. Pick one agent system, not all of them, one. Trace one real production path and ask for the receipt. The audit has five questions. What woke it up? What state did it inherit? Which authority did it use? What executed? And what evidence survived? These questions expose causality. They turn a fluent conversation into an inspectable production run. First, what woke it up? A user message, webhook, timer, tool result, sub-agent, or a replay. Name the trigger and its identity. Without that, you cannot reason about deduplication, order, or authorization. Second, which state did it inherit? Transcript, session state, memory snapshot, policy version, and tool surface. The model only reasons over the working set the harness assembled. Third, which authority did it use? Record the actor, session, run, arguments, scope, and lifetime. A model request is not permission. Authority should bind to one pending action. Fourth, what executed? Record the tool or API call, arguments, attempt number, idempotency key, and external results. This is a side-effect boundary, not the post-summary of what the agent intended. Fifth, what evidence survived? Did the ticket get updated? Did the message get rendered? Did the file get changed? Did the calendar event exist? The receipt should end at the boundary the user usually cares about. Now let's apply the same audit to the opening incident. What woke it up? A user message. What state it owned? That was the broken boundary. What executed? The channel send. What evidence survived? Delivery. What did not survive? The durable turn. Delivery survived while the state did not. That gap is the harness failure. The agent did not need a better model. The model did not need a better prompt. The system needed a better harness with a complete receipt. Let me recap the same three things I asked you to remember from the start of my talk. Own the state, order the mutation, and prove the action. A better model helps inside the turn. Ownership, ordering, life cycle, authority, and proof keep the system sane across turns. A model proposes, the harness commits, and the receipt proves it. Once text can become an action, the useful question changes. Do not only ask whether the model can reason, but ask whether the system can own the state, order the mutation, bound the work, constrain authority, and preserve evidence. A loop can answer a turn, and a harness can serve production. If you want to go deeper, scan the QR codes. The first points to the OpenAI Agents SDK, where all these harnesses are already built in so that you can use it to build your own agents. The second points to the agent stack, where I write about production agent systems in more detail. I'll be at the OpenAI booth after this talk if you want to talk about your harness design. Thank you. So these failures are familiar. Agents makes them easier to trigger and harder to explain. That's why agent harness matters. Agents makes them easier to understand. Let's talk about the first failure mode. This is the same failure mode I started this talk with. The user sees a success. The source of truth cannot replay it. Delivered is not remembered. In this state whole open claw issue, a telegram reply could succeed while the router turn was not written to the active context or transcript. The user saw the response. The log looked healthy, but the next turn had no durable record of the text change. A successful send proves transcript. It does not prove the future context. That distinction matters because the model can answer fluently over an incomplete record. The missing boundary was not intelligence. It was the state ownership. By owner, I do not mean a person. I mean the system of record whose persistent state becomes the truth. A calendar event belongs to the calendar system. A support status belongs to the ticketing system, while a code change belongs to a workspace or repository. A conversation turn belongs to a session transcript. And a user preference belongs to a memory store. Storage tells you where the bytes live. Ownership tells you who can reconsider the reality. A replay is not a reliable memory until a named owner can replay it. A system has to persist the turn. It has to name the owner or system of record. And it has to make the replay possible. The weak question is simple. For every fact the agent might use later, who owns it and how would you replay it? If no owner can replay the fact, the system did not reliably remember it. Once we know who owns the state, the next question is who's allowed to change it and in what order. Two correct writes can still produce one wrong outcome. And last writer wins is not a consistency model. In this overlapping writer open claw issue describes a load, modify, save race. Two callers load the same old state. Each changes a different record. The second save silently erases the first. The user may see a dismissed commitment return or a severe duplicate follow-up. Neither writer is malformed. Both operations are locally correct. The missing boundary is serialization around the commit. The invariant is not no concurrency. That would be too slow and it would miss the point. You can fan out sub-agents. Parallel reads are fine. Independent retrieval is fine. Many sessions can also run at once. The rule is narrower and simple. One ordered commit path for one mutable state boundary. This mechanism may be a queue, a transaction or a lock. You can use locks or mutex across the sessions and queues or transactions within a session. Be conservative at the commit time and not across the whole system. Users do not see queues or locks. They see behavior. A lock correction feels forgetful. A stuck lane feels dead. And completion before delivery feels confused. Ordering is a product feature because users experience ordering bugs as personalities. Next, now let's talk about time. In production, silence cannot be neutral. Let's review the next failure mode. The run waits for an event that cannot arrive. Silence is not a terminal state. In this dangling tool call issue, the session contains a tool call but no matching tool result. A process may have died. The connection may have dropped. A timeout might have happened before the results but it's recorded. The exact cost matters for debugging. The production failure is much simpler. The run is waiting for an event that will never arrive. New messages queue behind that silence. To the user, the agent simply looks stuck. Runs needs deadlines and cancellation. A deadline bones the weight. Watchdog makes the stuck work visible. Tools needs time modes and error results. Channels needs recovery commands that do not wait behind the stuck work they are trying to fix. Every external boundary needs an ending. Success, failure, timeout, cancel, or max attempts. Most importantly, the receipt records the terminal outcome so the next step does not have to guess. Bound the work before the work bounds you. Bound the work. Now let's we can move on from state to authority because a chat becomes risky when it becomes an action. Capability is not execution. The model can request an action. Requestability is not authority. Approval needs a shape. In this approval drift issue, expired approved callback was stated as retrievable. The state called back, state durable, served restart and blocked later channel work. The button click existed. The valid authority did not. This is the mistake. Trading approval as a vague memory that the human was near the system or clicked yes. Approval is a scoped execution state. It must stay bound to the action it authorized. An expression must terminate rather than loop. A useful approval object answers who approved in what session and run for which tool and for which arguments and for how long and with what outcome. It also points to the receipt. If those fields fall off during a retry, replay, or a channel callback, the harness can no longer prove the action was being executed as the action being approved. The general lesson is simple. Capability is not execution. Least privileges narrows the tool surface. Scoped credentials ensures the right identity is used for the action. Approval and audit decides what happens before and after the execution. The model can reason about the boundary, but it should not be the boundary. The model can request, but still the system decides. Finally, even if the tool says success, the user visible burl may disagree. Internal component reports success. The user visible surface shows nothing. This is the inverse of the opening incident we saw. In this missing edge proof issue, the message tool reported success for a web chat or TUI run, but the message did not render. Normal assistant replies still appeared. The tool proved that the internal path accepted the request. It does not prove the user saw the result. That difference changed the conversation. The agent may later say, I already sent it, and the user may truthfully say, I never saw it. Internal success is not external proof. Proof is a chain, not a claim. Model proposed something, policy allowed or denied it, execution attempted it, user visible edge confirmed or failed to confirm the outcome. The receipt preserves the chain. A transcript records what the agent said. The tool results records what one component claimed. A receipt records what the agent can verify at the boundary that matters. Let me recap all the incidents. Here are the file failure shapes you to look for. A state hole, overlapping writers, dangling tool call, approval drift, and missing edge proof. For each one, let's ask the same question. What did the user see? Which boundary it broke? And what would the receipt have got? Here is the audit I want you to run when you get back to your team. Pick one agent system, not all of them, one. Trace one production, one real production path and ask for the receipt. The audit has five questions. What woke it up? What state did it inherit? Which authority did it use? What executed? And what evidence survived? These questions expose casualty. They turn a fluent conversation into an inspectable production run. First, what woke it up? A user message, webhook, timer, tool result, sub-agent, or a replay. Name the trigger and its identity. Without that, you cannot reason about deduplication, order, or authorization. Second, which state did it inherit? Transcript, session state, memory snapshot, policy version, and tool surface. The model only reasons over the working set the harness assembled. Third, which authority did it use? Record the actor, session, run, arguments, scope, and lifetime. A model request is not permission. Authority should bind to a one pending action. Fourth, what executed? Record the tool or API call, arguments, attempt number, item potency key, and external results. This is a side effect boundary, not the post summary of what the agent intended. Fifth, what evidence survived? Did the ticket get updated? Did the message got rendered? Did the file got changed? Did the calendar even exist? The receipt should end at the boundary the user usually cares about. Now let's apply the opening incident, the same audit to the opening incident. What woke it up a user message? What state it owned that was the broken boundary? What executed the channel sent? What evidence survived delivery? What did not survive the durable turn? Delivery survived while the state did not. That gap is the harness failure. The agent did not need a better model. The model did not need a better prompt. The system needed a better harness with complete receipt. Let me recap the same three things I asked you to remember from the start of my talk. Own the state, order the mutation, and prove the action. A better model heads inside the turn. Ownership, ordering, life cycle, authority, and proof keep the system sane across turns. A model proposes the harness commits and the receipt proves it. Once text can become an action, the useful question changes. Do not only ask whether the model can reason, but ask whether the system can own the state, order the mutation, bound the work, constraint authority, and preserve evidence. A loop can answer a turn, and harness can serve a production. If you want to go deeper, scan the QR codes. The first points to the OpenAI agents SDK where all these harness are already built in so that you can use to build your own agents. And second points to the agent stack where I write about the production agents systems in more detail. I'll be at the OpenAI booth after this talk if you want to talk about your harness design. Thank you.