AI Engineer

AI Agents Are Just Distributed Systems Now — Salman Munaf, TikTok

1652 summary words 7 min summary Watch video

Start with the signal

7 min read

Summary

At-a-Glance

  • Verdict: Skim
  • Core thesis: Once an LLM agent can call tools and change external state, it must be engineered as a distributed system with deterministic controls around a probabilistic coordinator.
  • Why it matters: This is a practical reliability and security framing for agent systems: better models reduce mistakes but do not solve duplicate execution, stale state, ambiguous failures, excessive permissions, or unbounded cost.
  • Best use: Use it as a concise architecture review checklist for production agents, especially tool contracts, state management, approval design, retries, recovery, and tracing.

Executive Summary

Salman Munaf argues that agent builders should stop treating production agents as upgraded chatbots. A text-only chatbot primarily risks a bad answer; an agent can retrieve data, call APIs, write to databases, send customer communications, and trigger multi-step workflows. The model is therefore a probabilistic workflow coordinator operating across failure-prone external systems, rather than the whole architectural boundary.

The central prescription is to constrain that probabilistic layer with conventional distributed-systems controls. Tool calls need explicit contracts, request IDs, idempotency keys, status lookup, bounded retries, exponential backoff, circuit breakers, rate limits, spend limits, and limits on turns and parallel execution. A timeout must be treated as an unknown outcome, not proof that an action failed, because retrying could duplicate an already-successful refund or other irreversible action.

The talk also reframes agent context as operational state whenever it can influence an action. Short-term chat history and long-term memory such as files, prompts, databases, and caches can become stale or conflict with authoritative data; teams must declare a source of truth, attach provenance where possible, and invalidate cached agent memory after source updates. Multi-step workflows also need defined compensation paths rather than an assumption that every operation can be rolled back.

On security and operations, Munaf recommends scoped credentials and tool allowlists instead of broad database or system access; parameter-bound approvals rather than blanket human approval; and end-to-end traces that capture model, prompt, retrieved context, tool requests and responses, writes, errors, and approvals. The useful final test is not whether the model is smart, but whether the system can bound, observe, and recover from what it does when it is wrong.

Key Takeaways

  • Claim: An action-taking AI agent is a distributed system, with the LLM serving as a probabilistic coordinator that must be surrounded by deterministic controls. | Evidence: Munaf contrasts text-in/text-out chatbots with agent loops that plan, call external services and tools, observe partial results, persist state, and make real-world changes. He cites reported incidents involving Replika deleting a production database and Air Canada's chatbot providing an incorrect refund outcome. | Implication: Review agent architecture at every boundary it crosses: systems accessed, state touched, credentials used, and irreversible actions enabled—not only prompts and model quality. | Caveat: The examples are presented as illustrations of preventable system-design failures; the talk does not independently establish the full root causes of those incidents.
  • Claim: Every consequential tool call needs idempotency and explicit outcome verification because timeouts mean an unknown result, not a failed action. | Evidence: The speaker's example is an agent calling a customer-refund tool that times out: the server may have completed the refund while the client received an error. He recommends request IDs, idempotency keys, and status lookup for prior requests. | Implication: Expose tools as safe, inspectable operations with a durable operation ID and a queryable status before allowing autonomous retries. | Caveat: Idempotency prevents duplicate effects for repeatable requests but does not itself determine whether an ambiguous first request succeeded; the agent still needs a source-of-truth status check.
  • Claim: Agents require hard execution bounds because their default behavior under failure is often to keep trying, creating retry storms, cascading dependency failures, and runaway spend. | Evidence: Recommended controls include maximum retries, exponential backoff, circuit breakers for unhealthy downstream systems, maximum turns, maximum parallel calls, rate limits, and maximum spend. | Implication: Put reliability and cost budgets in the orchestration/control plane, not in the model's instructions or its judgment about when to stop. | Caveat: These controls can cause an agent to abandon a task; systems therefore need an explicit terminal state and a route for escalation or later recovery rather than silently treating the task as complete.
  • Claim: Agent memory is state, not merely context, whenever it can drive an external action; it must be governed like a cache with source-of-truth and invalidation rules. | Evidence: Munaf separates short-term memory (the execution-thread chat history) from long-term memory, including project files, system prompts, databases, and cache layers. He warns that these stores can become stale, conflict with authoritative data, or corrupt later actions. | Implication: Define data ownership and freshness requirements per decision and tool call, then invalidate or refresh agent-held context after relevant records change. | Caveat: The talk provides the governing principle but not a conflict-resolution protocol for cases where multiple systems are legitimately authoritative for different fields.
  • Claim: Multi-step agent workflows need explicit transaction boundaries and compensation operations because success in early steps does not guarantee completion of the whole workflow. | Evidence: The example workflow updates an internal ticket and emails a customer but then fails to update the CRM. For an erroneous email, the proposed compensation is a corrective or apology email; the speaker also calls for persistence of every step so the system can identify the failure point. | Implication: Model workflows as sagas: persist step-level state, define what to do after each partial failure, and distinguish reversible actions from actions that require business remediation. | Caveat: Many external side effects cannot be literally undone; compensation restores business correctness only imperfectly and should be designed before automation is enabled.
  • Claim: Least privilege, parameter-specific approval, and full agent traces are foundational controls; a harmless model becomes dangerous when granted unsafe authority. | Evidence: Munaf advises separate read and write permissions, tool allowlists, and approvals bound to the action, timestamp, actor, and expiration. His example is that approval for a $30 refund must not be reused for a $300 refund. He says traces should capture the model, prompt, retrieved context, tool calls, requests, responses, errors, writes, and approvals. | Implication: Treat approval artifacts and traces as first-class control-plane records that can authorize a specific operation and reconstruct why it occurred. | Caveat: Human approval is not a substitute for access control or tool-level validation; approvals must themselves be enforced against the exact parameters of the proposed action.

Detailed Brief

Production-readiness test for an agentic workflow

  • Claims: The speaker's concluding decision framework is whether the organization can bound, observe, and recover from agent actions.; Model capability improves the likelihood of correct behavior but cannot eliminate network failures, stale data, or adversarial input.; Tool contracts should explicitly define request and response types and schemas, rather than exposing loosely constrained access to underlying systems.
  • Evidence: The talk repeatedly distinguishes deterministic traditional workflow coordination from an agent's variable, model-chosen sequence of actions.; External dependencies named include APIs, databases, queues, retrieval sources, and other tools, all of which introduce partial failure modes beyond model reasoning.
  • Caveats: The session is an architectural principles talk rather than a concrete reference implementation, code walkthrough, or product comparison.; The latter portion of the supplied transcript substantially repeats the earlier material, reducing the amount of incremental content in the full video.
  • Implications: An agent launch gate should assess operational containment and recoverability independently from offline task-success or benchmark performance.; The most valuable implementation work may reside in the tool and orchestration layer—identity, state, execution controls, and telemetry—rather than in additional prompt iteration.

Notable Concepts & Terms

  • Probabilistic coordinator: Munaf's framing for an LLM agent: unlike a deterministic workflow engine, it may choose varied actions and paths, so the surrounding system must constrain its authority and execution.
  • Idempotency key: A durable request identifier that lets a tool recognize repeated calls and avoid repeating a side effect such as issuing a second refund.
  • Ambiguous timeout: A timeout indicates that the caller does not know the outcome; it must trigger status verification rather than an automatic assumption of failure and retry.
  • Compensation operation: A defined corrective action for a partially completed or unsafe workflow, such as sending a correction after an incorrect customer email.
  • Circuit breaker: A control that stops calls to an unhealthy or saturated downstream dependency, protecting both the dependency and the agent system from cascading failures.
  • Source of truth: The designated authoritative data source an agent must rely on when chat history, memory, cache, files, and databases conflict.
  • Scoped credentials: Narrow permissions separated by operation type and limited by tool allowlists, replacing broad read-write access granted merely for convenience.
  • Parameter-bound approval: Human authorization tied to a specific actor, action, amount or other parameters, timestamp, and expiration, rather than reusable blanket consent.

Operator Notes / Why Ken Should Care

  • Create an agent-tool contract standard requiring typed schemas, durable operation IDs, idempotency behavior, status retrieval, permission scope, retry classification, and declared compensation behavior before a tool is enabled.
  • Add a pre-production failure-mode review for every action-capable workflow: timeout after server success, duplicate delivery, stale retrieval, partial workflow completion, downstream saturation, and budget exhaustion.
  • Implement an execution ledger that records plan/action/observation transitions alongside retrieved context, tool I/O, writes, approval artifacts, and final disposition; use it to support replay and incident reconstruction.
  • Require approval tokens to be cryptographically or systemically bound to exact action parameters, an authorized actor, and an expiration, with reuse rejected.
  • Audit existing agent credentials for broad write access and replace them with capability-specific identities, read/write separation, and explicit tool allowlists.

Source/Metadata

  • Title: AI Agents Are Just Distributed Systems Now — Salman Munaf, TikTok
  • Transcript words: 3732
  • Duration seconds: 1188
  • Timestamp note: No usable timestamps or chapters were present in the supplied transcript. The transcript's second half substantially repeats earlier content.
Full transcript 2217 words · 17 min read
0:11

Hello everyone, good afternoon. Today I will be talking about AI agents as distributed systems. As the models have started to become more complex, initially the LLM models were just text in, text out, without performing any actions. And the effect that they could produce was just a wrong model output. However, with the rising agent capabilities, where the systems can now talk to external systems, it has turned into a distributed system, and it is important to incorporate distributed systems thinking and concepts when building AI agents. I will be going over that in this talk.

0:29

You guys might have heard about incidents being caused by AI agents. For instance, the Replika AI agent deleting a production database, or the Air Canada chatbot making an incorrect refund. Both of these incidents, or a lot of these incidents, could have been prevented by good systems thinking when building these AI agents. For instance, for the Replika AI incident, we could have robust backups, we could have scoped authority, we ideally shouldn't allow AI agents to delete production databases. Moreover, for the Air Canada chatbot, it would have been a good idea to have an authoritative source of retrieval so that it's not making decisions based on stale or incorrect policies.

0:38

So let's go over the transition from chatbot to production system. Initially, when we were in the age where LLMs were just chatbots, we had prompt in and we were outputting text. There were no side effects. The agent was not interacting with any other system. However, in the agentic era, in the agentic revolution, now those agents, by ingesting prompt, can run an agent loop, call external services, call tools, and also perform state changes. The architectural boundary now has moved way beyond an LLM model. And the difference is that it can now cause side effects in the outside world. So when building AI agents, it is important to recognize the external systems that it is talking to, the states that it is interacting with, what credentials it has, and the actions that it can perform.

0:48

I ideally like to think about AI agents as having a probabilistic coordinator. In distributed systems as well, we used to have services which were coordinating multi-step workflows. However, they were deterministic in nature. But in the case of AI agents, the AI acts as a probabilistic coordinator. The amount of action, the kind of actions that it can take, can vary quite a lot. It is not just a decision tree that we typically would have mapped out in traditional systems. And those actions can have severe consequences if they are not confined by having deterministic controls in place. So it is important to ensure that we have deterministic controls in place to ensure that the AI agent is not performing any actions that might be problematic.

0:55

So let's discuss how a typical agent loop might look. At first, it might do some planning. Then, based on that plan, it will perform an action. And it will then observe the results of those actions. It might persist that into some data store, and then decide what to do next. Each step in this loop is crossing a boundary. During planning, it can interact with data sources to retrieve some data. During action, it can call external APIs, tools, databases, and perform actions. During the observation phase, it can get partial results and make subsequent actions based on those partial results. It can persist incorrect data. And when deciding, it might also decide to perform an incorrect action. Or worse, it can also do a retry storm.

1:02

So it is very important when building an agent loop to persist every step of the process. Whatever actions the agent is doing, whatever context it is retrieving, it is important to persist that so that if anything fails, the agent is able to recognize where it failed, and it can perform a reversible action. It can perform undo operations. Similarly, there should be explicit transactions identified for each step. So for instance, if an agent is making a call, if it fails, what should it do? What should be the transaction to compensate for an irreversible or unsafe operation? For instance, if an agent sends a wrong email to a customer, what should it do to compensate for that?

1:14

Tool calls are just wrappers around external APIs, databases, queues, and so on. And when making these remote calls, there are some failures that you incorporate, such as network delays, timeouts, duplicate requests, or worse, the server-side request might succeed. However, the client might be reported an error. We have seen instances where our database might have written the data. However, due to some other errors, the server might have reported the error to us. We can perform corrective actions by seeing the database and actual source of truth. But in the agent's case, we need to ensure that we have proper guardrails in place.

1:21

So, for instance, an agent calls a refund customer tool call, which performs a refund to the customer. The request times out. Did the refund happen or not? What will the agent infer from that? Would it retry refunding the customer? The timeout does not actually mean that a failure had occurred. It means unknown. When designing these tools, it is important to have request IDs and idempotency keys so that when making duplicate requests, they are not causing duplicate side effects. And the system can always do a status lookup, like what the previous request was and what the status of that was, so that it is not making side effects with duplicate actions.

1:32

So AI agents, whenever they face failures, retry. Their first action is to perform retries. So it is really important to have idempotency baked in. If the same request is coming in to an external API or the tool, it should recognize that this is a duplicate request and ensure that no side effects are taking place. Moreover, we should also prevent AI agents from performing retry storms to external APIs because this can cause cascading failures. We should have max retries, budget spend, and max parallel calls to ensure that the fanout is not that large. Moreover, we should have exponential backoff in place to ensure that the downstream dependencies are not being burdened. And we should also have compensation operations in place for operations that can have side effects.

1:41

A lot of teams, when building AI agents, think of AI agent context as just context that the AI agent has. However, when that context can influence an action, it's state. And that state can become stale, can conflict with the authoritative data, or corrupt future actions that the agent might perform. I like to classify it into two different types of memory that the agent has. First is the short-term memory, which is the chat thread that the agent has, which is tied to a single execution thread. And the second is the long-term memory. It can be project files, system prompts, databases that it interacts with, the cache layer, and so on. It is important to decide what will be the source of truth when these different data sources have conflicting information. And we should ideally treat memory as a cache, which can be invalidated, which can have provenance attached to it. So, for instance, whenever a data store or a database is updated, or the source of truth is updated, we invalidate the context or the memory that the agent has to ensure that it is not making actions based on incorrect or stale data.

1:54

Usually these agents perform multi-step actions. And the agent can succeed on the first couple of steps, and then it fails. It is important to reverse the entire transaction that was performed. And these can then cross system boundaries. So, for instance, an agent can update an internal ticket, send an email to a customer, and fail to update the CRM. We need to figure out what is the correct compensation operation when it hits that failure. So, for instance, as I mentioned earlier, if it improperly sends an incorrect email to the customer, it is important that the compensation operation is defined for the AI agent to ensure that it is sending an apology email to the customer, or an email that is correcting that mistake.

2:03

So the AI agent runs in a loop. And whenever it can do multiple calls, it can have a retry loop that it can run whenever it fails. So it is important to have circuit breakers whenever it is making external calls, to ensure that it is not burdening the downstream system. For instance, if a downstream is unhealthy, there should be system circuit breakers in place that prevent AI agents from calling that dependency. Moreover, it also prevents cascading failures when, for instance, the downstream dependency is unhealthy or is saturated.

2:14

It is also important to assign rate limits and budgets. An agent can run over your cost if it's not assigned proper budgets and rate limits. It will keep retrying and try to solve the problem that it is facing. So it is important that we have set up max turns, max parallelism, and max spend to ensure that the model is not crossing the budget boundary that we have set.

2:22

Moreover, usually whenever we are building AI agents, we try to give all the permissions that it can have to ensure that it can perform the task that we have. That's the first step that we usually take, to give the AI agents all the privileges to perform any actions. For instance, if it's interacting with our database, we just give it all the read-write access to the entire table. However, it is important to give scoped credentials to it. There should be separate read-write permissions, and there should be allow lists for the tools that it can call. A harmless model can become dangerous when it can perform unsafe operations.

2:38

Moreover, human approval shouldn't be tied to a blanket approval. It should be tied to action, timestamp, actor, and expiration. So, for instance, if a user has given an approval to approve a $30 refund, it shouldn't turn into a subsequent approval for a $300 refund. It is important that whenever an approval is given, it should be tied to the particular parameters that it was asked for.

2:45

So observability is an important requirement when building AI agents because logs are not enough. Teams need to reconstruct when an agent failed, what happened, what information it was reacting to, and why it failed. And logs alone are not enough for teams to determine that. It is important to trace the model that was called, the prompt that was given to it, and also the tool calls that were made, the request that was made, the response from the tool, the errors that it got, the retrieved context, the retrieved information that the agent was reacting to, the writes that it made, and the approvals that it got, and so on.

2:55

So I would like to end with the idea that, yes, model capability matters. Having good models improves the likelihood of it making correct operations. Smarter models reduce mistakes. It improves the capability that the model has. However, it cannot eliminate network failures, stale data, or adversarial input. It is important when building this architecture that we also reason about whether we can bound, observe, and recover from actions performed by the AI agent.

3:01

It is important to have tool contracts in place to ensure that it is only allowed to make operations that it is given, that it is provided the contract. And the contracts are clearly establishing the request and response types, the schema, and all these tools have idempotency baked into them, so that when repeated requests are sent in, it is not causing unsafe operations to be retried. Moreover, there should be source-of-truth decisions made. When there are conflicting memory states, it is important for the agent to realize this is the source-of-truth data that it should rely on. And we should have retry policies, like rate limits, set in to ensure that the agent is not retrying aggressively. Moreover, permissions should be set up. There should be traces and recovery paths.

3:10

So when building AI agents, we should also ask what the system lets it do when it is wrong. Thank you. Thank you.

3:32

Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you.

4:44

Thank you. I love you. I love you. I love you. I love you. I love you. I love you.

5:55

See you next time. So it is very important when building an agent loop to persist every step of the process. Whatever actions the agent is doing, whatever context it is retrieving, it is important to persist that so that if anything fails, the agent is able to recognize where it failed, and it can perform a reversible action. It can perform undo operations. Similarly, there should be explicit transactions identified for each step. So for instance, if an agent is making a call, if it fails, what it should do? What should be the transaction to compensate for a irreversible or unsafe operation?

6:42

For instance, if an agent sends a wrong email to a customer, what should it do to compensate for that?

6:53

So tool calls are just wrappers around external APIs, databases, queues, and so on. And when calling, when making these remote calls, there are some failures that you incorporate, such as network delays, timeouts, you can make duplicate requests, or worse, the server side request might succeed. However, the client might be reported an error. We have seen instances where our database might have written the data. However, due to some other errors, the server might have reported to us the error. So we have seen instances where we can make a error. We can basically perform corrective actions based on by seeing the database and actual source of truth.

7:58

But in agent's case, we need to ensure that we have proper guardrails in place. So for instance, an agent calls refund customer tool call, which basically performs a refund to the customer. The request times out. Did the refund happen or not? What will the agent infer from that? Would it retry refunding to the customer? Basically, the timeout does not actually mean that a failure had occurred. It means unknown. It is important to have, when designing these tools, it is important to have request IDs, item potency keys, so that when making duplicate requests, they are not causing duplicate side effects.

8:47

And the system can always do a status lookup, like what the previous request was and what was the status of that, so that it is not making side effects with duplicate actions.

9:05

So AI agents, whenever they face failures, they retry, their first action is to perform retries. So it is really important to have items. So it is really important to have item potency baked in. If a same request is coming in to an external API or the tool, it should recognize that this is a duplicate request and ensure that no side effects are taking place. Moreover, we should also prevent AI agents to perform retry storms to external APIs because this can cause cascading failures. We should have max returns, budget spend, and max parallel calls to prevent, to ensure that the fanout is not that large.

10:00

Moreover, we should have exponential backoff in place to ensure that the downstream dependencies are not being burdened. And we should also have compensation operations in place for operations that can have side effects.

10:28

So a lot of teams when building AI agents think of AI agent context as just a context that the AI agent has as just a context. However, when that context can influence an action, it's a state. And that state can become stale that can conflict with the authoritative data or corrupt future actions that the agent might perform. I like to classify it into two different types of memory that the agent has. First is the short-term memory, which is the chat thread that the agent has, which is tied to a single execution thread. And the second is the long-term memory. It can be project files, system prompts, databases that it interacts with, the cache layer, and so on.

11:23

It is important to decide what will be the source of truth when these different data sources have conflicting information. And we should ideally treat memory as a cache, which can be invalidated, which can have provenance attached to it. So for instance, whenever a data store or a database is updated, or the source of truth is updated, we invalidate the context or the memory that the agent has to ensure that it is not making actions based on the incorrect or stale data.

12:09

So usually these agents perform multi-step actions. And the agent can succeed on the first couple of steps, and then it fails. It is important to reverse the entire transaction that was performed. And these can then cross system boundaries. So for instance, an agent can update an internal ticket, send an email to a customer, and fail to update the CRM. We need to figure out what is the correct compensation operation when it hits that failure. So for instance, as I mentioned earlier, that it improperly sends an incorrect email to the customer.

13:01

It is important that the compensation operation is defined for the AI agent to ensure that it is sending an apology email to the customer, or an email that is correcting that mistake.

13:22

So the AI agent basically runs in a loop. And whenever it can do multiple calls, it can have a retry loop that it can run based whenever it fails. So it is important to have circuit breakers whenever it is making external calls, to ensure that it is not burdening the downstream system. For instance, if a downstream is unhealthy, there should be system circuit breakers in place that prevents AI agents to call that dependency. Moreover, it also prevents cascading failures when, for instance, the downstream dependency is unhealthy or is saturated. It is also important to assign rate limits and budgets.

14:17

An agent can go over, can run your cost if it's not assigned proper budgets and rate limits. It will keep retrying and try to solve the problem that if it's facing. So it is important that we have set up max turns, max parallelism, max spend to ensure that the model is not crossing the budget boundary that we have set.

14:54

Moreover, ideally, usually, whenever we are building AI agents, we usually try to give all the permissions that it can have to ensure that it can perform the task that we have. That's the first thing that we have. That's the first step that we take, usually, to give the AI agents all the privileges to perform any actions. Like, for instance, if it's interacting with our database, we just give it all the read-write access to the entire table. However, it is important to give scoped credentials to it. There should be separate read-write permissions, and there should be allow lists for the tools that it can call.

15:44

A harmless model can become dangerous when it can perform unsafe operations. Moreover, a human approval shouldn't be tied to a blanket approval. It should be tied to action, timestamp, actor, and expiration. So, for instance, if a user has given an approval to approve a $30 refund, it shouldn't turn into a subsequent approval for a $300 refund. It is important that whenever an approval is given, it should be tied to the particular parameters that it was asked for.

16:32

So, observability is an important requirement when building AI agents because logs are not enough. Teams need to reconstruct when an agent failed, what happened, what information was it reacting to, and why it failed. And logs alone are not enough for an agent, for teams to determine that. It is important to trace the model that was called, the prompt that was given to it, and also the tool calls that were made, the request that was made, the response from the tool, the errors that it got, the retrieved context, what the agent was, was the retrieved information that the agent was reacting to, the rights that it made, and the approvals that it got, and so on.

17:32

So, I would like to end with the idea that, yes, model capability matters. Having good models improves the likelihood of it making correct operations. Smarter models reduce mistakes. It improves the capability that the model has. However, it cannot eliminate network failures, stale data, or adversarial input. It is important when building this architecture, we also reason about can we bound, observe, and recover from actions performed by the AI agent. It is important to have tool contracts in place to ensure that it is only allowed to make operations that it is given, that it is provided the contract.

18:26

And the contracts are clearly establishing the request and response, response, response types, the schema, and all these tools have idempotency baked into it, so that when repeated requests are sent in, it is not causing unsafe operations to be retried. Moreover, there should be source of truth decisions made. When there are conflicting memory states, it is important for the agent to realize this is the source of truth data that it should rely on. And we should have retry policies like rate limits set in to ensure that the agent is not retrying aggressively. Moreover, permissions should be set up. There should be traces and recovery paths.

19:19

So when building AI agents, we should also ask what the system lets it do when it is wrong. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you.

19:41

Thank you. Thank you. Thank you. Thank you. I love you. I love you. I love you. I love you. I love you. I love you. See you next time.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note