Thank you for choosing to spend this session with me. My goal today is simple. I want to convince you all that most of the production failures are not, most of the agent failures are not model failures. Those are harness failures.
So let's start with one production incident. The user saw the reply, the system forgot it happened. This is a failure shape I want to start with. Not a hallucination, not a crash, not a bad answer. The user's visible edge looked healthy, while the durable record had a hole. In this example, the user asked the agent to remember a refund for a customer. The assistant said it recorded the fact for the next turn. The interface looked normal, no red screen, no obvious failures. But the next turn cannot reconsider the fact. The user experienced success, the system inherited incomplete reality.
Why this matters? The crash is annoying, but at least it gives you a boundary. You know something stopped. You usually see an error. You can often start from the last known good point. Silent success gives you a lie. Delivery can succeed while the persistent fails. The user has no reason to doubt the reply. The operator has no obvious alarm. The next turn can still sound confident because the model is coherent. But it is coherent over a broken history. That is why agent reliability matters. And agent reliability cannot stop at model quality.
Hi, I'm Vinod. I work on core data and AI infrastructure at OpenAI. Before that, I worked on distributed systems at Apple and Uber. Outside of work, I write the agent stack, where I try to explain how production AI and data systems work under the hood. I'm not a member of OpenClaw, and this is not an OpenClaw product pitch. This is a pure system design talk. I'm using OpenClaw as a public case study because its issues, code, and docs make the harness around the agent unusually visible.
Here's the production contract for the talk. A model proposes, the harness commits, and the receipt proves it. The model may suggest a message, a tool, an edit, or a command. But the model is not the production boundary. The harness owns the state transition, the authority check, the audit commit, and the receipt is the evidence that survives the turn.
OpenClaw is the case study, and the contract is the takeaway. If you remember only three things from this talk, make it these. Own the state, order the mutation, and prove the action. A fact needs only one owner and one replay path. Shared mutable state needs one ordered commit path. And transcript is not the proof. A transcript tells you what the agent said. A receipt tells you what the system allowed, attempted, executed, and what the user-visible edge confirmed.
To create a simple mental model, I created this car analogy of the harness. The model is the engine. It matters a lot. But nobody buys a production car by just looking at the horsepower alone. You also care about steering, brakes, road rules, dashboard, and a black box. The model gives you capability, but the harness gives you control. A powerful engine with no brakes is no autonomy. It is a liability with good acceleration.
Here is the harness blueprint I wanted to discuss today. Every agent we know of, personal agents such as OpenClaw or Hermes, coding agents such as Codex, Cursor, OpenCode, or Claude Code, uses the same underlying architecture. Events enter from many surfaces: a chat, webhook, timer, heartbeat, or another external system. The control plane maps the events to a session key, and the session key determines the state boundary. The session lane gives you one active writer for the mutable state. The runtime calls the models and tools. Tools act through approvals and policies, and the audit rail becomes the run receipt.
This is the blueprint: event, session key, throttle, tools, audit. The blueprint is the talk, and the incident is the proof that each boundary matters. Context is assembled. In agent runtime, it does not usually remember in a human sense. It is stateless. The harness rebuilds the working state for each turn. The working set may include the transcript, session state, memory, policy, and tool definitions. The model only sees what the harness supplies. If one input is missing or stale, the answer may still sound coherent. Coherence does not prove the working set was complete.
These failures are not new. We already know about timeouts, retries, idempotency, locks, ordering, and state ownership. What changes is the agent setting. Now these failures sit around a probabilistic planner with dynamic plans. It's rebuilding the context for every turn. There are more event sources, and it can act through more action surfaces. So these failures are familiar. Agents make them easier to trigger and harder to explain. That's why agent harness matters. Let's talk about the first failure mode. This is the same failure mode I started this talk with. The user sees a success. The source of truth cannot replay it. Delivered is not remembered.
In this state hole OpenClaw issue, a Telegram reply could succeed while the router turn was not written to the active context or transcript. The user saw the response. The log looked healthy, but the next turn had no durable record of the text change. A successful send proves delivery. It does not prove the future context. That distinction matters because the model can answer fluently over an incomplete record. The missing boundary was not intelligence. It was state ownership.
By owner, I do not mean a person. I mean the system of record whose persistent state becomes the truth. A calendar event belongs to the calendar system. A support status belongs to the ticketing system, while a code change belongs to a workspace or repository. A conversation turn belongs to a session transcript. And a user preference belongs to a memory store. Storage tells you where the bytes live. Ownership tells you who can reconsider the reality. A replay is not reliable memory until a named owner can replay it.
A system has to persist the turn. It has to name the owner or system of record. And it has to make the replay possible. The key question is simple. For every fact the agent might use later, who owns it and how would you replay it? If no owner can replay the fact, the system did not reliably remember it. Once we know who owns the state, the next question is who's allowed to change it and in what order. Two correct writes can still produce one wrong outcome. And last writer wins is not a consistency model.
In this overlapping writer OpenClaw issue, a load, modify, save race is described. Two callers load the same old state. Each changes a different record. The second save silently erases the first. The user may see a dismissed commitment return or a severe duplicate follow-up. Neither writer is malformed. Both operations are locally correct. The missing boundary is serialization around the commit.
The invariant is not no concurrency. That would be too slow, and it would miss the point. You can fan out sub-agents. Parallel reads are fine. Independent retrieval is fine. Many sessions can also run at once. The rule is narrower and simple. One ordered commit path for one mutable state boundary. This mechanism may be a queue, a transaction, or a lock. You can use locks or mutexes across sessions, and queues or transactions within a session. Be conservative at commit time, not across the whole system.
Users do not see queues or locks. They see behavior. A lock correction feels forgetful. A stuck lane feels dead. And completion before delivery feels confused. Ordering is a product feature because users experience ordering bugs as personalities. Next, let's talk about time. In production, silence cannot be neutral. Let's review the next failure mode. The run waits for an event that cannot arrive. Silence is not a terminal state.
In this dangling tool call issue, the session contains a tool call but no matching tool result. A process may have died. The connection may have dropped. A timeout might have happened before the result was recorded. The exact cause matters for debugging. The production failure is much simpler. The run is waiting for an event that will never arrive.
New messages queue behind that silence. To the user, the agent simply looks stuck. Runs need deadlines and cancellation. A deadline bounds the wait. A watchdog makes the stuck work visible. Tools need time modes and error results. Channels need recovery commands that do not wait behind the stuck work they are trying to fix. Every external boundary needs an ending: success, failure, timeout, cancel, or max attempts. Most importantly, the receipt records the terminal outcome so the next step does not have to guess. Bound the work before the work bounds you.
Now let's move on from state to authority, because a chat becomes risky when it becomes an action. Capability is not execution. The model can request an action. Requestability is not authority. Approval needs a shape. In this approval drift issue, expired approved callback state was still retrievable. The state stayed durable, survived restart, and blocked later channel work. The button click existed. The valid authority did not. This is the mistake: treating approval as a vague memory that the human was near the system or clicked yes.
Approval is a scoped execution state. It must stay bound to the action it authorized. An expression must terminate rather than loop. A useful approval object answers who approved, in what session and run, for which tool and for which arguments, for how long, and with what outcome. It also points to the receipt. If those fields fall off during a retry, replay, or a channel callback, the harness can no longer prove the action being executed is the action being approved.
The general lesson is simple. Capability is not execution. Least privilege narrows the tool surface. Scoped credentials ensure the right identity is used for the action. Approval and audit decide what happens before and after execution. The model can reason about the boundary, but it should not be the boundary. The model can request, but the system decides. Finally, even if the tool says success, the user-visible edge may disagree. The internal component reports success. The user-visible surface shows nothing. This is the inverse of the opening incident we saw.
In this missing edge proof issue, the message tool reported success for a web chat or TUI run, but the message did not render. Normal assistant replies still appeared. The tool proved that the internal path accepted the request. It does not prove the user saw the result. That difference changes the conversation. The agent may later say, I already sent it, and the user may truthfully say, I never saw it.
Internal success is not external proof. Proof is a chain, not a claim. The model proposed something, policy allowed or denied it, execution attempted it, and the user-visible edge confirmed or failed to confirm the outcome. The receipt preserves the chain. A transcript records what the agent said. A tool result records what one component claimed. A receipt records what the agent can verify at the boundary that matters.
Let me recap all the incidents. Here are the five failure shapes to look for: a state hole, overlapping writers, dangling tool call, approval drift, and missing edge proof. For each one, ask the same question. What did the user see? Which boundary did it break? And what would the receipt have shown? Here is the audit I want you to run when you get back to your team. Pick one agent system, not all of them, one. Trace one real production path and ask for the receipt.
The audit has five questions. What woke it up? What state did it inherit? Which authority did it use? What executed? And what evidence survived? These questions expose causality. They turn a fluent conversation into an inspectable production run. First, what woke it up? A user message, webhook, timer, tool result, sub-agent, or a replay. Name the trigger and its identity. Without that, you cannot reason about deduplication, order, or authorization. Second, which state did it inherit? Transcript, session state, memory snapshot, policy version, and tool surface. The model only reasons over the working set the harness assembled.
Third, which authority did it use? Record the actor, session, run, arguments, scope, and lifetime. A model request is not permission. Authority should bind to one pending action.
Fourth, what executed? Record the tool or API call, arguments, attempt number, idempotency key, and external results. This is a side-effect boundary, not the post-summary of what the agent intended. Fifth, what evidence survived? Did the ticket get updated? Did the message get rendered? Did the file get changed? Did the calendar event exist? The receipt should end at the boundary the user usually cares about.
Now let's apply the same audit to the opening incident. What woke it up? A user message. What state it owned? That was the broken boundary. What executed? The channel send. What evidence survived? Delivery. What did not survive? The durable turn. Delivery survived while the state did not. That gap is the harness failure. The agent did not need a better model. The model did not need a better prompt. The system needed a better harness with a complete receipt.
Let me recap the same three things I asked you to remember from the start of my talk. Own the state, order the mutation, and prove the action. A better model helps inside the turn. Ownership, ordering, life cycle, authority, and proof keep the system sane across turns.
A model proposes, the harness commits, and the receipt proves it. Once text can become an action, the useful question changes. Do not only ask whether the model can reason, but ask whether the system can own the state, order the mutation, bound the work, constrain authority, and preserve evidence. A loop can answer a turn, and a harness can serve production. If you want to go deeper, scan the QR codes. The first points to the OpenAI Agents SDK, where all these harnesses are already built in so that you can use it to build your own agents. The second points to the agent stack, where I write about production agent systems in more detail.
I'll be at the OpenAI booth after this talk if you want to talk about your harness design. Thank you. So these failures are familiar. Agents makes them easier to trigger and harder to explain. That's why agent harness matters.
Agents makes them easier to understand. Let's talk about the first failure mode.
This is the same failure mode I started this talk with. The user sees a success. The source of truth cannot replay it. Delivered is not remembered. In this state whole open claw issue, a telegram reply could succeed while the router turn was not written to the active context or transcript. The user saw the response. The log looked healthy, but the next turn had no durable record of the text change. A successful send proves transcript. It does not prove the future context. That distinction matters because the model can answer fluently over an incomplete record. The missing boundary was not intelligence. It was the state ownership.
By owner, I do not mean a person. I mean the system of record whose persistent state becomes the truth. A calendar event belongs to the calendar system. A support status belongs to the ticketing system, while a code change belongs to a workspace or repository. A conversation turn belongs to a session transcript. And a user preference belongs to a memory store. Storage tells you where the bytes live. Ownership tells you who can reconsider the reality.
A replay is not a reliable memory until a named owner can replay it. A system has to persist the turn. It has to name the owner or system of record. And it has to make the replay possible. The weak question is simple. For every fact the agent might use later, who owns it and how would you replay it? If no owner can replay the fact, the system did not reliably remember it.
Once we know who owns the state, the next question is who's allowed to change it and in what order. Two correct writes can still produce one wrong outcome. And last writer wins is not a consistency model.
In this overlapping writer open claw issue describes a load, modify, save race. Two callers load the same old state. Each changes a different record. The second save silently erases the first. The user may see a dismissed commitment return or a severe duplicate follow-up. Neither writer is malformed. Both operations are locally correct. The missing boundary is serialization around the commit.
The invariant is not no concurrency. That would be too slow and it would miss the point. You can fan out sub-agents. Parallel reads are fine. Independent retrieval is fine. Many sessions can also run at once. The rule is narrower and simple. One ordered commit path for one mutable state boundary. This mechanism may be a queue, a transaction or a lock. You can use locks or mutex across the sessions and queues or transactions within a session. Be conservative at the commit time and not across the whole system.
Users do not see queues or locks. They see behavior. A lock correction feels forgetful. A stuck lane feels dead. And completion before delivery feels confused. Ordering is a product feature because users experience ordering bugs as personalities.
Next, now let's talk about time. In production, silence cannot be neutral. Let's review the next failure mode. The run waits for an event that cannot arrive. Silence is not a terminal state. In this dangling tool call issue, the session contains a tool call but no matching tool result. A process may have died. The connection may have dropped. A timeout might have happened before the results but it's recorded. The exact cost matters for debugging. The production failure is much simpler. The run is waiting for an event that will never arrive. New messages queue behind that silence. To the user, the agent simply looks stuck.
Runs needs deadlines and cancellation. A deadline bones the weight. Watchdog makes the stuck work visible. Tools needs time modes and error results. Channels needs recovery commands that do not wait behind the stuck work they are trying to fix. Every external boundary needs an ending. Success, failure, timeout, cancel, or max attempts. Most importantly, the receipt records the terminal outcome so the next step does not have to guess. Bound the work before the work bounds you. Bound the work. Now let's we can move on from state to authority because a chat becomes risky when it becomes an action.
Capability is not execution. The model can request an action. Requestability is not authority. Approval needs a shape. In this approval drift issue, expired approved callback was stated as retrievable. The state called back, state durable, served restart and blocked later channel work. The button click existed. The valid authority did not. This is the mistake. Trading approval as a vague memory that the human was near the system or clicked yes. Approval is a scoped execution state. It must stay bound to the action it authorized. An expression must terminate rather than loop.
A useful approval object answers who approved in what session and run for which tool and for which arguments and for how long and with what outcome. It also points to the receipt. If those fields fall off during a retry, replay, or a channel callback, the harness can no longer prove the action was being executed as the action being approved. The general lesson is simple. Capability is not execution. Least privileges narrows the tool surface. Scoped credentials ensures the right identity is used for the action. Approval and audit decides what happens before and after the execution. The model can reason
about the boundary, but it should not be the boundary. The model can request, but still the system decides.
Finally, even if the tool says success, the user visible burl may disagree. Internal component reports success. The user visible surface shows nothing. This is the inverse of the opening incident we saw. In this missing edge proof issue, the message tool reported success for a web chat or TUI run, but the message did not render. Normal assistant replies still appeared. The tool proved that the internal path accepted the request. It does not prove the user saw the result. That difference changed the conversation. The agent may later say, I already sent it, and the user may truthfully say, I never saw it.
Internal success is not external proof. Proof is a chain, not a claim. Model proposed something, policy allowed or denied it, execution attempted it, user visible edge confirmed or failed to confirm the outcome. The receipt preserves the chain. A transcript records what the agent said. The tool results records what one component claimed. A receipt records what the agent can verify at the boundary that matters. Let me recap all the incidents. Here are the file failure shapes you to look for. A state hole, overlapping writers, dangling tool call, approval drift, and missing edge proof. For each one,
let's ask the same question. What did the user see? Which boundary it broke? And what would the receipt have got? Here is the audit I want you to run when you get back to your team. Pick one agent system, not all of them, one. Trace one production, one real production path and ask for the receipt. The audit has five questions. What woke it up? What state did it inherit? Which authority did it use? What executed? And what evidence survived? These questions expose casualty. They turn a fluent conversation into an inspectable production run. First, what woke it up? A user message,
webhook, timer, tool result, sub-agent, or a replay. Name the trigger and its identity. Without that, you cannot reason about deduplication, order, or authorization. Second, which state did it inherit? Transcript, session state, memory snapshot, policy version, and tool surface. The model only reasons over the working set the harness assembled. Third, which authority did it use? Record the actor, session, run, arguments, scope, and lifetime. A model request is not permission. Authority should bind to a one pending action. Fourth, what executed? Record the tool or API call, arguments, attempt number, item potency key, and external results.
This is a side effect boundary, not the post summary of what the agent intended. Fifth, what evidence survived? Did the ticket get updated? Did the message got rendered? Did the file got changed? Did the calendar even exist? The receipt should end at the boundary the user usually cares about. Now let's apply the opening incident, the same audit to the opening incident. What woke it up a user message? What state it owned that was the broken boundary? What executed the channel sent? What evidence survived delivery? What did not survive the durable turn? Delivery survived while the state did not.
That gap is the harness failure. The agent did not need a better model. The model did not need a better prompt. The system needed a better harness with complete receipt. Let me recap the same three things I asked you to remember from the start of my talk. Own the state, order the mutation, and prove the action. A better model heads inside the turn. Ownership, ordering, life cycle, authority, and proof keep the system sane across turns. A model proposes the harness commits and the receipt proves it. Once text can become an action, the useful question changes.
Do not only ask whether the model can reason, but ask whether the system can own the state, order the mutation, bound the work, constraint authority, and preserve evidence. A loop can answer a turn, and harness can serve a production. If you want to go deeper, scan the QR codes. The first points to the OpenAI agents SDK where all these harness are already built in so that you can use to build your own agents. And second points to the agent stack where I write about the production agents systems in more detail. I'll be at the OpenAI booth after this talk if you want to talk about your harness design. Thank you.