Hello everyone, good afternoon. Today I will be talking about AI agents as distributed systems. As the models have started to become more complex, initially the LLM models were just text in, text out, without performing any actions. And the effect that they could produce was just a wrong model output. However, with the rising agent capabilities, where the systems can now talk to external systems, it has turned into a distributed system, and it is important to incorporate distributed systems thinking and concepts when building AI agents. I will be going over that in this talk.
You guys might have heard about incidents being caused by AI agents. For instance, the Replika AI agent deleting a production database, or the Air Canada chatbot making an incorrect refund. Both of these incidents, or a lot of these incidents, could have been prevented by good systems thinking when building these AI agents. For instance, for the Replika AI incident, we could have robust backups, we could have scoped authority, we ideally shouldn't allow AI agents to delete production databases. Moreover, for the Air Canada chatbot, it would have been a good idea to have an authoritative source of retrieval so that it's not making decisions based on stale or incorrect policies.
So let's go over the transition from chatbot to production system. Initially, when we were in the age where LLMs were just chatbots, we had prompt in and we were outputting text. There were no side effects. The agent was not interacting with any other system. However, in the agentic era, in the agentic revolution, now those agents, by ingesting prompt, can run an agent loop, call external services, call tools, and also perform state changes. The architectural boundary now has moved way beyond an LLM model. And the difference is that it can now cause side effects in the outside world. So when building AI agents, it is important to recognize the external systems that it is talking to, the states that it is interacting with, what credentials it has, and the actions that it can perform.
I ideally like to think about AI agents as having a probabilistic coordinator. In distributed systems as well, we used to have services which were coordinating multi-step workflows. However, they were deterministic in nature. But in the case of AI agents, the AI acts as a probabilistic coordinator. The amount of action, the kind of actions that it can take, can vary quite a lot. It is not just a decision tree that we typically would have mapped out in traditional systems. And those actions can have severe consequences if they are not confined by having deterministic controls in place. So it is important to ensure that we have deterministic controls in place to ensure that the AI agent is not performing any actions that might be problematic.
So let's discuss how a typical agent loop might look. At first, it might do some planning. Then, based on that plan, it will perform an action. And it will then observe the results of those actions. It might persist that into some data store, and then decide what to do next. Each step in this loop is crossing a boundary. During planning, it can interact with data sources to retrieve some data. During action, it can call external APIs, tools, databases, and perform actions. During the observation phase, it can get partial results and make subsequent actions based on those partial results. It can persist incorrect data. And when deciding, it might also decide to perform an incorrect action. Or worse, it can also do a retry storm.
So it is very important when building an agent loop to persist every step of the process. Whatever actions the agent is doing, whatever context it is retrieving, it is important to persist that so that if anything fails, the agent is able to recognize where it failed, and it can perform a reversible action. It can perform undo operations. Similarly, there should be explicit transactions identified for each step. So for instance, if an agent is making a call, if it fails, what should it do? What should be the transaction to compensate for an irreversible or unsafe operation? For instance, if an agent sends a wrong email to a customer, what should it do to compensate for that?
Tool calls are just wrappers around external APIs, databases, queues, and so on. And when making these remote calls, there are some failures that you incorporate, such as network delays, timeouts, duplicate requests, or worse, the server-side request might succeed. However, the client might be reported an error. We have seen instances where our database might have written the data. However, due to some other errors, the server might have reported the error to us. We can perform corrective actions by seeing the database and actual source of truth. But in the agent's case, we need to ensure that we have proper guardrails in place.
So, for instance, an agent calls a refund customer tool call, which performs a refund to the customer. The request times out. Did the refund happen or not? What will the agent infer from that? Would it retry refunding the customer? The timeout does not actually mean that a failure had occurred. It means unknown. When designing these tools, it is important to have request IDs and idempotency keys so that when making duplicate requests, they are not causing duplicate side effects. And the system can always do a status lookup, like what the previous request was and what the status of that was, so that it is not making side effects with duplicate actions.
So AI agents, whenever they face failures, retry. Their first action is to perform retries. So it is really important to have idempotency baked in. If the same request is coming in to an external API or the tool, it should recognize that this is a duplicate request and ensure that no side effects are taking place. Moreover, we should also prevent AI agents from performing retry storms to external APIs because this can cause cascading failures. We should have max retries, budget spend, and max parallel calls to ensure that the fanout is not that large. Moreover, we should have exponential backoff in place to ensure that the downstream dependencies are not being burdened. And we should also have compensation operations in place for operations that can have side effects.
A lot of teams, when building AI agents, think of AI agent context as just context that the AI agent has. However, when that context can influence an action, it's state. And that state can become stale, can conflict with the authoritative data, or corrupt future actions that the agent might perform. I like to classify it into two different types of memory that the agent has. First is the short-term memory, which is the chat thread that the agent has, which is tied to a single execution thread. And the second is the long-term memory. It can be project files, system prompts, databases that it interacts with, the cache layer, and so on. It is important to decide what will be the source of truth when these different data sources have conflicting information. And we should ideally treat memory as a cache, which can be invalidated, which can have provenance attached to it. So, for instance, whenever a data store or a database is updated, or the source of truth is updated, we invalidate the context or the memory that the agent has to ensure that it is not making actions based on incorrect or stale data.
Usually these agents perform multi-step actions. And the agent can succeed on the first couple of steps, and then it fails. It is important to reverse the entire transaction that was performed. And these can then cross system boundaries. So, for instance, an agent can update an internal ticket, send an email to a customer, and fail to update the CRM. We need to figure out what is the correct compensation operation when it hits that failure. So, for instance, as I mentioned earlier, if it improperly sends an incorrect email to the customer, it is important that the compensation operation is defined for the AI agent to ensure that it is sending an apology email to the customer, or an email that is correcting that mistake.
So the AI agent runs in a loop. And whenever it can do multiple calls, it can have a retry loop that it can run whenever it fails. So it is important to have circuit breakers whenever it is making external calls, to ensure that it is not burdening the downstream system. For instance, if a downstream is unhealthy, there should be system circuit breakers in place that prevent AI agents from calling that dependency. Moreover, it also prevents cascading failures when, for instance, the downstream dependency is unhealthy or is saturated.
It is also important to assign rate limits and budgets. An agent can run over your cost if it's not assigned proper budgets and rate limits. It will keep retrying and try to solve the problem that it is facing. So it is important that we have set up max turns, max parallelism, and max spend to ensure that the model is not crossing the budget boundary that we have set.
Moreover, usually whenever we are building AI agents, we try to give all the permissions that it can have to ensure that it can perform the task that we have. That's the first step that we usually take, to give the AI agents all the privileges to perform any actions. For instance, if it's interacting with our database, we just give it all the read-write access to the entire table. However, it is important to give scoped credentials to it. There should be separate read-write permissions, and there should be allow lists for the tools that it can call. A harmless model can become dangerous when it can perform unsafe operations.
Moreover, human approval shouldn't be tied to a blanket approval. It should be tied to action, timestamp, actor, and expiration. So, for instance, if a user has given an approval to approve a $30 refund, it shouldn't turn into a subsequent approval for a $300 refund. It is important that whenever an approval is given, it should be tied to the particular parameters that it was asked for.
So observability is an important requirement when building AI agents because logs are not enough. Teams need to reconstruct when an agent failed, what happened, what information it was reacting to, and why it failed. And logs alone are not enough for teams to determine that. It is important to trace the model that was called, the prompt that was given to it, and also the tool calls that were made, the request that was made, the response from the tool, the errors that it got, the retrieved context, the retrieved information that the agent was reacting to, the writes that it made, and the approvals that it got, and so on.
So I would like to end with the idea that, yes, model capability matters. Having good models improves the likelihood of it making correct operations. Smarter models reduce mistakes. It improves the capability that the model has. However, it cannot eliminate network failures, stale data, or adversarial input. It is important when building this architecture that we also reason about whether we can bound, observe, and recover from actions performed by the AI agent.
It is important to have tool contracts in place to ensure that it is only allowed to make operations that it is given, that it is provided the contract. And the contracts are clearly establishing the request and response types, the schema, and all these tools have idempotency baked into them, so that when repeated requests are sent in, it is not causing unsafe operations to be retried. Moreover, there should be source-of-truth decisions made. When there are conflicting memory states, it is important for the agent to realize this is the source-of-truth data that it should rely on. And we should have retry policies, like rate limits, set in to ensure that the agent is not retrying aggressively. Moreover, permissions should be set up. There should be traces and recovery paths.
So when building AI agents, we should also ask what the system lets it do when it is wrong. Thank you. Thank you.
Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you.
Thank you. I love you. I love you. I love you. I love you. I love you. I love you.
See you next time. So it is very important when building an agent loop to persist every step of the process. Whatever actions the agent is doing, whatever context it is retrieving, it is important to persist that so that if anything fails, the agent is able to recognize where it failed, and it can perform a reversible action. It can perform undo operations. Similarly, there should be explicit transactions identified for each step. So for instance, if an agent is making a call, if it fails, what it should do? What should be the transaction to compensate for a irreversible or unsafe operation?
For instance, if an agent sends a wrong email to a customer, what should it do to compensate for that?
So tool calls are just wrappers around external APIs, databases, queues, and so on. And when calling, when making these remote calls, there are some failures that you incorporate, such as network delays, timeouts, you can make duplicate requests, or worse, the server side request might succeed. However, the client might be reported an error. We have seen instances where our database might have written the data. However, due to some other errors, the server might have reported to us the error. So we have seen instances where we can make a error. We can basically perform corrective actions based on by seeing the database and actual source of truth.
But in agent's case, we need to ensure that we have proper guardrails in place. So for instance, an agent calls refund customer tool call, which basically performs a refund to the customer. The request times out. Did the refund happen or not? What will the agent infer from that? Would it retry refunding to the customer? Basically, the timeout does not actually mean that a failure had occurred. It means unknown. It is important to have, when designing these tools, it is important to have request IDs, item potency keys, so that when making duplicate requests, they are not causing duplicate side effects.
And the system can always do a status lookup, like what the previous request was and what was the status of that, so that it is not making side effects with duplicate actions.
So AI agents, whenever they face failures, they retry, their first action is to perform retries. So it is really important to have items. So it is really important to have item potency baked in. If a same request is coming in to an external API or the tool, it should recognize that this is a duplicate request and ensure that no side effects are taking place. Moreover, we should also prevent AI agents to perform retry storms to external APIs because this can cause cascading failures. We should have max returns, budget spend, and max parallel calls to prevent, to ensure that the fanout is not that large.
Moreover, we should have exponential backoff in place to ensure that the downstream dependencies are not being burdened. And we should also have compensation operations in place for operations that can have side effects.
So a lot of teams when building AI agents think of AI agent context as just a context that the AI agent has as just a context. However, when that context can influence an action, it's a state. And that state can become stale that can conflict with the authoritative data or corrupt future actions that the agent might perform. I like to classify it into two different types of memory that the agent has. First is the short-term memory, which is the chat thread that the agent has, which is tied to a single execution thread. And the second is the long-term memory. It can be project files, system prompts, databases that it interacts with, the cache layer, and so on.
It is important to decide what will be the source of truth when these different data sources have conflicting information. And we should ideally treat memory as a cache, which can be invalidated, which can have provenance attached to it. So for instance, whenever a data store or a database is updated, or the source of truth is updated, we invalidate the context or the memory that the agent has to ensure that it is not making actions based on the incorrect or stale data.
So usually these agents perform multi-step actions. And the agent can succeed on the first couple of steps, and then it fails. It is important to reverse the entire transaction that was performed. And these can then cross system boundaries. So for instance, an agent can update an internal ticket, send an email to a customer, and fail to update the CRM. We need to figure out what is the correct compensation operation when it hits that failure. So for instance, as I mentioned earlier, that it improperly sends an incorrect email to the customer.
It is important that the compensation operation is defined for the AI agent to ensure that it is sending an apology email to the customer, or an email that is correcting that mistake.
So the AI agent basically runs in a loop. And whenever it can do multiple calls, it can have a retry loop that it can run based whenever it fails. So it is important to have circuit breakers whenever it is making external calls, to ensure that it is not burdening the downstream system. For instance, if a downstream is unhealthy, there should be system circuit breakers in place that prevents AI agents to call that dependency. Moreover, it also prevents cascading failures when, for instance, the downstream dependency is unhealthy or is saturated. It is also important to assign rate limits and budgets.
An agent can go over, can run your cost if it's not assigned proper budgets and rate limits. It will keep retrying and try to solve the problem that if it's facing. So it is important that we have set up max turns, max parallelism, max spend to ensure that the model is not crossing the budget boundary that we have set.
Moreover, ideally, usually, whenever we are building AI agents, we usually try to give all the permissions that it can have to ensure that it can perform the task that we have. That's the first thing that we have. That's the first step that we take, usually, to give the AI agents all the privileges to perform any actions. Like, for instance, if it's interacting with our database, we just give it all the read-write access to the entire table. However, it is important to give scoped credentials to it. There should be separate read-write permissions, and there should be allow lists for the tools that it can call.
A harmless model can become dangerous when it can perform unsafe operations. Moreover, a human approval shouldn't be tied to a blanket approval. It should be tied to action, timestamp, actor, and expiration. So, for instance, if a user has given an approval to approve a $30 refund, it shouldn't turn into a subsequent approval for a $300 refund. It is important that whenever an approval is given, it should be tied to the particular parameters that it was asked for.
So, observability is an important requirement when building AI agents because logs are not enough. Teams need to reconstruct when an agent failed, what happened, what information was it reacting to, and why it failed. And logs alone are not enough for an agent, for teams to determine that. It is important to trace the model that was called, the prompt that was given to it, and also the tool calls that were made, the request that was made, the response from the tool, the errors that it got, the retrieved context, what the agent was, was the retrieved information that the agent was reacting to, the rights that it made, and the approvals that it got, and so on.
So, I would like to end with the idea that, yes, model capability matters. Having good models improves the likelihood of it making correct operations. Smarter models reduce mistakes. It improves the capability that the model has. However, it cannot eliminate network failures, stale data, or adversarial input. It is important when building this architecture, we also reason about can we bound, observe, and recover from actions performed by the AI agent. It is important to have tool contracts in place to ensure that it is only allowed to make operations that it is given, that it is provided the contract.
And the contracts are clearly establishing the request and response, response, response types, the schema, and all these tools have idempotency baked into it, so that when repeated requests are sent in, it is not causing unsafe operations to be retried. Moreover, there should be source of truth decisions made. When there are conflicting memory states, it is important for the agent to realize this is the source of truth data that it should rely on. And we should have retry policies like rate limits set in to ensure that the agent is not retrying aggressively. Moreover, permissions should be set up. There should be traces and recovery paths.
So when building AI agents, we should also ask what the system lets it do when it is wrong. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you.
Thank you. Thank you. Thank you. Thank you. I love you. I love you. I love you. I love you. I love you. I love you. See you next time.