AI Engineer

Agents Are Where Microservices Were in 2015 — Roberto Milev & Uday Kanagala, Navan

1850 summary words 8 min summary Watch video

Start with the signal

8 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Production agent systems are at an early maturity stage analogous to microservices circa 2015, so teams should build a reliable single-agent foundation with modular skills, lifecycle-aware runtime, observability, evaluation, and authorization before pursuing multi-agent complexity.
  • Why it matters: Navan offers a practical enterprise reference stack and several reusable operating patterns for running stateful, tool-using agents at scale, while clearly identifying unresolved problems in cost control, replay/debugging, testing, and standards.
  • Best use: Use this as an architecture and operating-model benchmark for agent-control-plane decisions, especially skill packaging, trace design, policy enforcement, and the threshold for introducing multi-agent systems.

Executive Summary

Navan frames the current agent stack as a replay of the microservices transition: useful infrastructure and conventions are emerging, but the industry has not yet solved the hard operational problems. Its central warning is to resist multi-agent orchestration until a team can build, test, secure, and operate a single agentic loop reliably.

Their proposed production stack has five practical layers: a stateful agent runtime; memory that progresses from conversational state to longer-term and episodic knowledge; context management through dynamically loaded skills; operational instrumentation and evaluation; and authorization/guardrails around tool use. Navan runs on AWS AgentCore but has built additional session persistence and rehydration capabilities where the managed runtime was insufficient.

The most operationally useful point is that conventional logs are inadequate for long, non-deterministic agent trajectories. Navan instruments session and tool-call hooks, emits OpenTelemetry traces through Braintrust, and records decision signals including current goal, rationale, belief state, tool calls, confidence, and whether an answer was inferred. It evaluates multi-step work through trajectories rather than expecting deterministic step-by-step graphs.

The speakers view runtime, memory, MCP-based tool invocation, and basic scaling as comparatively mature. They identify cost predictability, model-routing/fallback policies, replay and debugging, agent-specific observability, and immature standards such as A2A as the areas that still need deliberate internal engineering rather than blind dependence on vendors.

Key Takeaways

  • Claim: Treat multi-agent orchestration as an advanced optimization, not a default architecture: first make one agentic loop reliable. | Evidence: The speakers explicitly adapt the microservices maxim: if a team cannot build a well-structured monolith, it should not adopt microservices; similarly, if it cannot build a single agent, it should not build a multi-agent system. Navan uses a single master agent that progressively loads sub-skills. | Implication: Ken should make a single orchestrator plus independently testable capabilities the default OpenClaw pattern, and require a concrete boundary, ownership model, or protocol need before adding autonomous peer agents. | Caveat: They recognize a legitimate multi-agent/A2A use case when separate organizational teams or system boundaries need formal contracts to communicate; this is distinct from splitting one workflow into agents prematurely.
  • Claim: Skills should be the atomic unit for context, execution, testing, and reuse. | Evidence: Navan defines a skill as both domain/task instructions and its associated tool-execution behavior. It dynamically composes an agent's context from pluggable domain skills and uses progressive disclosure: begin with limited context, then expand through skill metadata only when needed. | Implication: Design capabilities as versioned skill modules with scoped instructions, permitted tools, metadata, and isolated tests instead of putting all tool descriptions and domain context into a universal system prompt. | Caveat: The transcript does not provide a concrete skill schema, retrieval policy, or measurement showing that progressive disclosure outperforms alternative context-packaging approaches.
  • Claim: Agent runtimes must be designed around stateful sessions and lifecycle management, unlike stateless API services. | Evidence: The speakers say agents inherently require persistent sessions, isolation, and a lifecycle different from conventional APIs. Navan relies heavily on AWS AgentCore Runtime but built its own session-persistence and rehydration layer to fill a gap. | Implication: Separate agent session state, checkpointing/rehydration, and isolation from request-serving infrastructure; do not assume a stateless web-service deployment model is sufficient for durable agent work. | Caveat: Their confidence that runtime is 'pretty much solved' reflects an AWS-centric deployment and may not transfer cleanly to self-hosted, cross-cloud, or highly long-running agent environments.
  • Claim: For agent operations, traces at session and tool-call interception points are more useful than raw logs. | Evidence: Navan argues that agents generate too much reasoning output for log-centric debugging. It intercepts pre-session, post-session, pre-tool, and post-tool hooks; emits OpenTelemetry traces through Braintrust; and captures current goal, reasons, belief status, tool calls, confidence, and inferred-answer status. | Implication: Implement an event and trace contract around every tool invocation and state transition. Prioritize actionable decision metadata and policy outcomes over retaining unrestricted chain-of-thought-style logs. | Caveat: The speakers acknowledge that OpenTelemetry is not yet a fully settled fit for agentic calls; it can be made to work, but agent-native observability conventions remain immature.
  • Claim: Agent evaluation must assess the quality and completeness of a non-deterministic trajectory, not require an identical sequence of steps on every run. | Evidence: For workflows involving 20 to 30 decisions, Navan says an agent may choose different steps each run, making a deterministic graph unrealistic. It therefore relies heavily on trajectory evaluations that measure progress from source toward goal, and uses inferred-answer signals to identify regressions and trigger correction or human review. | Implication: Build evals around outcome quality, tool-use safety, constraint satisfaction, and path efficiency, with regression labels for uncertain/inferred outputs; do not make exact trace matching the primary release gate. | Caveat: The talk does not define a trajectory scoring formula, benchmark set, pass threshold, or process for adjudicating valid but unusual paths.
  • Claim: Authorization and guardrails must be enforced at tool boundaries because agents act under delegated authority rather than as simple users or service accounts. | Evidence: The speakers use the example, 'book me a flight whenever it's cheaper than $200': the agent executes a purchase based on a user's standing assertion. Navan applies checks before and after every tool call to block or make informed policy decisions. | Implication: Model agent identity and delegation explicitly, apply fine-grained policy checks to each consequential action, and make pre-tool authorization a non-optional control-plane function rather than a prompt instruction. | Caveat: The transcript does not specify how Navan represents delegated intent, approval limits, auditability, revocation, or liability for actions completed under an agent's authority.
  • Claim: Industry convergence is real around MCP and managed runtime/memory primitives, but cost, replay/debugging, and A2A standards remain material unsolved risks. | Evidence: The speakers call MCP a de facto protocol and note widespread tool-calling support, while also saying cost is difficult to predict or guardrail, cheaper-model routing and reliable fallbacks are unresolved, replay/debugging is hard, and A2A is young and vendor-driven. | Implication: Adopt interoperable interfaces such as MCP where useful, but preserve a provider-independent cost, routing, replay, and policy layer rather than assuming vendor platforms will solve the economics or operational controls. | Caveat: The statement that vendors are incentivized to make customers consume more tokens is the speakers' interpretation, not evidence of vendor intent.

Detailed Brief

Memory model: beyond basic RAG

  • Claims: RAG arose because an agent cannot carry unlimited context, but retrieval alone is not the full memory architecture.; Memory should accumulate across short-term conversational state, application-managed long-term semantic memory, and episodic memories of instances that succeeded or failed.
  • Evidence: The speakers describe a provider-style pipeline of ingestion, extraction, consolidation, and retrieval.; Navan uses AWS AgentCore Memory while tailoring its implementation to its own use case.
  • Caveats: No retention, privacy, provenance, conflict-resolution, or memory-quality evaluation method is described.; Better frontier models may reduce some memory pressure, but do not eliminate the need for curation and retrieval decisions.
  • Implications: Treat memory as a layered product subsystem with distinct data classes and write/retrieval policies, rather than a single vector store.; Episodic success/failure records can become a feedback asset for agent improvement if they are linked to evaluations and trace data.

Maturity assessment of the agent platform stack

  • Claims: The speakers view runtimes, scaling, memory primitives, MCP, and tool invocation as relatively mature components.; Testing patterns are becoming better defined even though agents remain unreliable and difficult to test.; Observability, orchestration practices, standards, economics, and replay/debugging have not reached equivalent maturity.
  • Evidence: They state that cloud providers AWS, GCP, and Azure all provide versions of an agent runtime.; They describe MCP as evolving toward stateless operation and A2A as a young protocol promoted by particular vendors.; They suggest agents themselves may eventually help reduce the cognitive burden of debugging agent behavior.
  • Caveats: This is a practitioner maturity assessment from a conference presentation, not a comparative benchmark or vendor-neutral market study.; Their conclusion that scaling is not a problem may apply mainly to brute-force LLM execution capacity, not to cost-efficient or latency-sensitive production workloads.
  • Implications: Allocate engineering effort disproportionately to control-plane capabilities—evaluation, trace/replay, policy, and spend controls—rather than differentiating mainly through another agent framework.; Avoid locking the architecture to immature multi-agent standards until contracts, governance, and adoption stabilize.

Notable Concepts & Terms

  • AgentCore Runtime: AWS's agent runtime, used heavily by Navan; Navan supplemented it with session persistence and rehydration.
  • AgentCore Memory: AWS memory capability that Navan uses as part of a broader, use-case-specific memory architecture.
  • Skills: Navan's reusable unit of work combining domain context/instructions with tool-execution behavior; used for dynamic context composition.
  • Progressive disclosure: Start an agent with minimal relevant skill context and reveal additional detail through metadata only as its task requires it.
  • Trajectory evals: Evaluation approach for non-deterministic multi-step agents that measures progress, efficiency, or completeness toward a goal rather than exact path replication.
  • OpenTelemetry (OTel): The proposed instrumentation substrate for tracing sessions and tool operations, though the speakers consider its agent-specific fit unfinished.
  • MCP: Model Context Protocol, described as the emerging de facto standard for invoking tools/services from agents.
  • A2A protocol: Agent-to-agent communication protocol positioned for contracts across organizational or team boundaries, but characterized as early-stage.

Operator Notes / Why Ken Should Care

  • Define a minimum production-agent control-plane checklist: durable session checkpointing, skill registry/versioning, pre- and post-tool policy hooks, trace events, trajectory evals, and spend limits.
  • Adopt a delegated-authority model before enabling transactional tools: capture user intent, constraints such as price/amount thresholds, approval requirements, expiry/revocation, and an auditable execution record.
  • Instrument tool boundaries with structured fields for goal, selected tool, inputs/outputs subject to redaction, policy decision, confidence/uncertainty, execution cost, latency, and failure category.
  • Create a regression suite of representative long-horizon tasks and score outcome constraints plus trajectory quality; do not gate releases on deterministic step sequences.
  • Require explicit model-routing and fallback policies for each skill, including maximum token/cost budgets and escalation conditions, because vendor runtimes do not solve predictable unit economics by default.
  • Use MCP for tool interoperability where it reduces integration friction, but defer broad A2A adoption until there is a concrete cross-boundary contract and governance need.

Source/Metadata

  • Title: Agents Are Where Microservices Were in 2015 — Roberto Milev & Uday Kanagala, Navan
  • Transcript words: 4469
  • Duration seconds: 1167
  • Timestamp note: No timestamps or chapters were present in the supplied transcript. The latter portion substantially repeats earlier material.
Full transcript 2566 words · 20 min read
0:13

Hello everybody. Welcome to our talk. My name is Roberto Milev. I am the chief architect at Navan, and I have Uday here, who is also part of the architecture team. Navan is a travel and expense management company, and we'll share with you some of our learnings around how we run AI and what we have discovered.

0:20

So if you've been long enough in this industry, you remember that over time there are a few paradigm shifts, and we all tend to jump on a bandwagon and try to do things right. Last time was when we all jumped on the microservices bandwagon, and out of that a lot of good things came out, like container orchestration, Kubernetes. Then we had service mesh, circuit breakers, all of those good things. But it didn't happen overnight. It took a long time. It took some time for us to learn how to do these things.

0:27

So one of the quotes from there is, if you can build a well-structured monolith, why even try to build microservices? It translates today because if you can't build a single agentic loop, why go in and try to build a multi-agent orchestrated system?

0:33

So over time, just like previously, a reference architecture is emerging. We have learned a few things by doing in production. We have a lot of agents, a lot of tokens per day being used. And as I said, there are a few layers that have standardized, that have crystallized around what we need to run agentic flows reliably in production: runtime, memory, context management, all around operational cross-cutting concerns, and around orchestration as well. So today we'll go over some of these layers, all of these layers actually, and we will show where the industry is, what we have done, what we have learned, and so on.

0:50

So starting at the runtime layer, we've talked a lot and we've built a lot of services in order to scale them statelessly before. And now we're in a new world where agents are stateful by nature. They need to have persistent sessions, they need to have isolation, their life cycle is different than the life cycle of a traditional API service, and so on. So the cloud providers have jumped in and tried to fill this gap. AWS, GCP, Azure, they all have some incarnation of an agentic runtime. If you scan the QR code for this slide and for the following slides, you will see a comparison of some of the features and how different cloud providers try to approach this.

0:57

So, at Navan, we run everything on AWS. AWS has an agent core runtime. We heavily use that, but we have filled some gaps around that, like the session persistence and rehydration is something that we have built. And we also run a bunch of other SDKs for writing agents. And part of these runtimes is typically that they are framework agnostic, although they all prefer their native framework in a way.

1:04

The next layer in the stack is around memory. We started with RAG. RAG was a big thing for a while. We were driven to that out of necessity because you cannot fit an unlimited amount of context into an agent. And over time, all of these cloud providers and the industry have implemented a pipeline where memory is automatically generated by following a workflow of ingestion, extraction, and then consolidation and retrieval. And there are parts of RAG that are built in, things like a long-term memory that inherently has some semantic characteristics. But memories build up over time from short-term conversational memory to long-term memory that you manage yourself, then episodic memories about instances that worked well and didn't work well, and so on. We at Navan, again being an AWS shop, utilize their agent core memory, but we are also doing it in a way that matches our use case.

1:12

And then the next thing is context management. It's a hot topic. It was a hot topic, and it's still a hot topic. Context windows are growing bigger, but there's never enough context. Or if there is too much context, again, agents struggle with that because you lose focus, and so on.

1:19

What we found working is focusing on skills as a unit of context. And I'll explain what I mean by that. We look at skills as both having context, meaning instructions and setup about a certain domain or a task. And there's also the second part of the skill, which is the tool execution and the agentic part. And we compose context dynamically out of skills that we use as units of work that are pluggable, that we can test independently, and that we can reuse.

1:25

So, for example, when we have an agent, we have skills that are specific to a domain. And based on that, we compose them. And we rely on the progressive disclosure, which is a feature of the skills itself, to start with the limited scope of context and then expand by included metadata further down the line. I'll hand it out to Ude now to walk us through the rest of this. Thanks, Ruta. All right. Can I have a quick show of hands here? Who had built an agent which failed halfway through a multi-20-step or 30-step process and been able to figure out quickly or reason about why the agent failed?

1:53

So, again, logs. Traditionally with microservices, we all are familiar with logs. There's logs out there and then we go check out the logs. But this changes everything the moment we switch to agents. Agents output a lot of thinking. There's too much to consume. So that's not the right way to do it. Traditionally, that was the way, but our thought has to be changed now.

1:59

In the way the cloud, as an example, when we take cloud as an example for an agent, there's hooks, and we can intercept everything that cloud as an agent does at that level. So what kind of tool it calls, what kind of decision it's making. So before pre-tool and post-tool call, or pre-session or post-session, all of that are points in time for us to intercept and make a decision and either do a blocking operation or to log a metric or emit a metric. So this is a critical place where we can emit OTL traces. At Navan, we use one of our provider Braintrusts to emit these OTL traces. And through these traces, we should be able to figure out the spans, the traces, and at what point in time the agent is stuck, which gives much more confidence into how we operate and build the agent.

2:09

This is a day-two operational challenge. Building an agent is this. There's so many frameworks, but how you navigate building and operating an agent later is a primary concern now. And moreover, the reasoning chain, the thought process, and critical signals that we emit here.

2:16

As part of the trace captures, we emit a few primary signals here. What is the current goal the agent is going through? The reasons behind its operations, and the belief status and the tool calls that it's making. So this gives us judgment pointers in the traces. And when the agent makes a decision, there is a confidence score, how confident it is when it makes this judgment. So whether there are multiple paths that lead to this choice, or whether this is an inferred answer. These are signals that give us confidence later to review. If this is an inferred answer, there could be a human in the loop to guide through and tweak the agent to perform a little better.

2:22

Again, can I have a raise of hands again to see how confident you are, 100% confident, in testing pipelines with your agents? Right? So this is one of the other critical aspects today. Because agents are non-deterministic. We've all been used to programming and writing much more deterministic flows. And we know how it works. When I ask an engineer, the engineer can come and tell me how the algorithm, the sequence of operations, everything is programmed in our mind, everything is expectations. But now the agents come in a non-deterministic way. And how do we test them? So that is very critical here.

2:41

And yeah, we are also struggling. We've started building agents. The day-two operations was challenging, and then we failed in a lot of steps. How do we course-correct? The moment we change something, something else breaks. So how do we do that?

2:48

One approach that we took, this is from research papers around the concept of trajectory. In a multi-step orchestration, when an agent makes 30 steps or decisions to reach a goal, if that is a program, that's a different story. But this is not a program. This is a non-deterministic way of doing it. It makes up its own steps every time differently. So how can we chart a deterministic graph here? Is it possible? No. Can we have a trajectory of it starting from an end to a goal and then see how much, how far it went in the trajectory and how far it went from the source to the destination is what we can compute to evaluate the efficiency or the completeness of the agent evaluation?

2:59

So we heavily rely on trajectory evals. And there are a few other signals, as I briefly spoke around in the previous slide, around the inferred signal. If the answer is from an inferred answer, how can we loop that in and make a signal on how we can classify that this is a regression and make fixes towards the agent? So next is the guardrails. Where? Is this the one? Yeah. So guardrails and authorization. This displays a critical role in enterprise AI. A lot of information is being piped to models. There could be sensitive information that goes into it without our knowledge. And we as leaders, how can we put in this governance layer to stop this is very critical here.

3:32

And the concept of authentication and authorization is taking up a different approach here. Traditionally, we've seen a user or a service account. But now, what is an agent? An agent can be acting on behalf of users. There are so many use cases there. Hey, book me a flight whenever it's cheaper than $200. So we just tell this assertion, and then the agent goes and figures out and does this action on behalf of me. So is it me making this purchase, or is it agent me making it on behalf of me? So the line is being blurred here, and we need to make fine-grained authorization decisions here, and the policy layer, that's where the guardrails and authentication authorization plays a critical role.

3:41

And in Navan, what we employ here is before every tool call, pre-tool and post-tool, we have these guardrails to check and block and make informed decisions. And this single-agent versus multi-agent, again, this is orchestration wars. You can think of whether to build a single agent or a multi-agent. Again, as Roboto briefly hinted, if you can't perfect and build a single agent, why go towards multi-agent? So learn from our failures, experiences, and build towards that.

3:57

At Navan, the approach that we have taken is a single master, and then we adopted sub-skills. There are sub-agents within it. So it's a single agent that can progressively load the skills and understand decisively what needs to be loaded into the context and then make this navigation through the use case.

4:07

But there are other patterns that are also emerging. There are different classes of use cases here. One is agent-to-agent communication. So if you take a large-scale organization and there are so many of these teams that are acting as boundaries and they don't talk to each other, how do we communicate? There are two agents on either side, right? How do we do it? So there is A2A protocol which can help us establish the contracts in terms of skills, and we can use A2A as a protocol there, which is a boundary between the teams. Yeah, over to your food.

4:23

All right. So as we went through the stack, it's obvious that some components of the stack are in a more mature state, and we already have good answers for them. As I said, the runtime, I think it's pretty much solved. We are so advanced in orchestration, and we are running LLMs in a very brute-force way, so scaling is not a problem. Also, memory, I think as the frontier LLMs get better and as our practices get better, we will find a way to cover the majority of the use cases, and there is good maturity around the cloud providers.

4:37

MCP has emerged as a de facto protocol, and tool calling is now a feature that everybody supports. So we are seeing some industry convergence around that as well, and MCP as a standard is also evolving. Now it's becoming stateless. We are reaching a point where we know how to invoke services and tools with agents. In some areas, things are happening, but there's still a lot of unknown around observability. There is a push towards Otel, but does Otel really work for agentic calls? Yeah, you can make it work, as Uday was saying.

4:51

Also, we are getting more comfortable around the testing patterns. It's very hard to test, but we have found a way to give customers quality experiences even with the unreliability of agentic systems, and I think that's getting in a state that is more or better defined. Orchestration is another one where we have patterns. We can build bigger agents, smaller agents. As we said previously, probably the right answer is to not over-engineer. So we are learning there, and a pattern of school thought is also emerging.

5:09

Where we're all struggling with, and the previous talk was about this from the developer AI-assisted development perspective, but also we are seeing these issues from our production agents, it's very hard to predict cost and it's very hard to manage cost and put guardrails and solve this in a way where there is a reliable fallback maybe, or have agents be using cheaper models for certain tasks. This is all driven by the big AI vendors who I think, their interest is for us all to spend more tokens.

5:18

Replay and debugging, Uday talked about that. That's also a big, big issue. It's very hard to understand, but I think this is also something that is going to be solved because we can now use agents to get over the cognitive overload of trying to debug what they do. And then standards. Standards are emerging by the community. Hotel, as I mentioned, agent to agent is young. It's pushed by certain vendors, but I think over time we will get there. With all of this said, we know what we need, and it's up to us to go ahead and build it. Thank you, everybody. Thank you. Thank you.

5:56

And we compose context dynamically out of skills that we use as units of works that are pluggable, that we can test independently, and that we can reuse. So, for example, when we are we have an agent, we have skills that are specific to a domain. And based on that, we compose them. And we rely on the, you know, the progressive disclosure, which is a feature of the skills itself, to start with the limited scope of context and then expand by included metadata further down the line. I'll hand it out to Ude now to kind of walk us through the rest of this. Thanks, Ruta. All right. Can I have a quick show of hands here?

6:50

Who had built an agent which failed halfway through a multi-20 step or 30 step process and be able to figure out quickly or reason about why the agent failed?

7:05

So, again, logs, we've generally been traditionally with microservices. We all are familiar with logs. There's logs out there and then we go check out the logs. But this changes everything the moment we switch to agents. Agents output a lot of thinking. There's too much to consume. So, that's not the right way to do it, right? So, traditionally, that was the way, but our thought has to be changed right now. In the way the cloud, as an example, when we take cloud as an example for an agent, there's hooks and we can intercept everything that cloud as an agent that does at that level. So, what kind of tool

7:41

it calls, right? What kind of decision it's making. So, before pre-tool and post-tool call or pre-session or post-session. So, all of that are a point in time for us to intercept and make a decision and either block do a blocking operation or to log a metric or emit a metric, right? So, this is a critical place where we can emit OTL traces. At now on, we use one of our provider brain trusts to emit these OTL traces. And through these traces, we should be able to figure out the spans, the traces, and at what point in time where the agent is stuck, which gives much more confidence into how we operate and build

8:25

the agent. This is a day two operational challenge. Building agent is this. There's so many frameworks, but how do you navigate building and operating an agent later is a primary concern now. And moreover, the reasoning chain, the thought process, and critical signals that we emit here. As part of the trace captures, we emit a few primary signals here. What is the current goal the agent is going through? The reasons behind its operations and the belief status and the tool calls that it's making. So, this kind of gives us a judgment pointers in the traces. And when we make, when the agent makes a decision,

9:06

there is a confidence score, how confident it is when it makes this judgment, right? So, whether there are multiple paths that it leads to this choice, or whether this is an inferred answer. So, basically, these are signals that gives us confidence later to review. If this is an inferred answer, there could be a human in the loop to guide through and tweak the agent to perform a little better.

9:32

Again, can I have a raise of hands again to see how confident are you, like, 100% confident in testing pipelines with your agents? Right? So, this is one of the other critical aspect today. Because agents are non-deterministic. We've all been used to program and write much more deterministic flows. And we know how it works. When I ask an engineer, the engineer can come and tell me how the algorithm, the sequence of operations, everything is programmed in our mind, everything is expectations. But now the agents come into a non-deterministic way. And how do we test them? Right? So, that is very critical here. And, yeah, we are also struggling. We've started

10:19

doing building agents. The day-to operations was challenging. And then we failed in a lot of steps. How do we course correct? The moment we change something, something else breaks, right? So, how do we do that? One approach that we took, this is from research papers around the concept of trajectory. Like, in a multi-step orchestration, when an agent makes 30 steps or decisions to make to reach to a goal, if that is a program or that's a different story. But this is not a program. This is non-deterministic way of… It makes up its own steps every time differently. So, how can we

11:00

chart a deterministic graph here? Is it possible? No. Can we have a trajectory of its starting from an end to a goal and then see how much, how far it went in the trajectory and how far it went from the source to the destination is what we can compute to evaluate the efficiency or the completeness of the agent evaluation? So, we heavily rely on trajectory evals. And there are a few other signals, as I briefly spoke around in the previous slide, around the inferred signal. If the answer is from an inferred answer, how can we loop that into and make a signal that on how can we classify that this is a regression and make fixes towards the agent?

12:02

So, next is the guardrails. Where?

12:11

Is this the one? Yeah. So, guardrails and authorization. This is critical… displays a critical role in enterprise AI. A lot of information is being piped to models. There could be sensitive information that goes into it without our knowledge. And we as leaders, how can we put in this governance layer to stop this is very critical here. And the concept of authentication and authorization is taking up a different approach here. Traditionally, we've seen a user or a service account. But now, what is an agent? Agent can be acting as on behalf of users. There

13:00

is so much of things, so many of use cases there. Hey, book me a flight whenever it's cheaper than $200, right? So, we just tell this assertion and then agent go figure out and does this action on behalf of me. So, is it me making this purchase or is it agent me making on behalf of me? So, there is… Agent acts as a… on behalf of user or agent user service account as well. So, the line is being blurred here and we need to make fine-grained authorization decisions here and the policy layer, that's where the guardrails and authentication authorization plays a critical role. And in Navan,

13:37

what we employ here is before every tool call, pre-tool and post-tool, we have these guardrails to check and block and make informed decisions. And this single agent versus multi-agent, again, this is kind of orchestration wars you can think of whether to build a single agent or a multi-agent. Again, as Roboto briefly hinted, if you can't perfect and build a single agent, why go towards multi-agent, right? So, learn from our failures, experiences and build towards that. At Navan, yeah, the approach that we have taken is a single master and then we adopted sub-skills. There are sub-agents within it. So,

14:30

it's a single agent that can progressively load the skills and understand decisively what needs to be loaded into the context and then make this navigation through the use case. But there are other patterns that are also emerging. There are different class of use cases here. One is agent-to-agent communication. So, there are, if you take a large-scale organization and there are so many of these teams that are acting as a boundaries and they don't talk to each other, let's say, how do we communicate? There are two agents on either of this side, right? How do we do it? So, there is A2A protocol which can help us

15:11

establish the contracts in terms of skills and we can use A2A as a protocol there, which kind of is a boundary between the teams. Yeah, over to your food.

15:28

All right. So, as we went through the stack, it's obvious that some components of the stack are in a more mature state and we already have good answers for them. As I said, the runtime, I think it's pretty much solved. We are so advanced in orchestration and we are running LLMs in kind of a very brute force way. So, scaling is not a problem. Also, memory, I think as the frontier LLMs get better and as our practices get better, we will find a way to cover the majority of the use cases and there is good maturity around the cloud providers. MCP has emerged as a de facto protocol and tool calling is now a

16:16

feature that everybody supports. So, we are seeing some industry convergence around that as well and MCP as a standard is also evolving. Now, it's becoming stateless. We are reaching a point where kind of we know how to invoke services and tools with agents. In some areas, things are happening, but you know, there's still a lot of unknown around observability. There is a push towards Otel, but does Otel really work for agentic calls? Yeah, you can make it work as Uday was saying. Also, we are getting more comfortable around the testing patterns. It's very hard to test, but we have found a way to give

17:02

customers quality experiences even with the unreliability of agentic system and I think that's kind of getting in a state that is more or better defined. Orchestration is another one one where, you know, we have patterns. We can build, you know, bigger agents, smaller agents. As we said previously, probably the right answer is to not over-engineer. So, we are learning there and a pattern of school thought is also emerging. Where we're all struggling with, and the previous talk was about this for the developer AI assisted development perspective, but also we are seeing these issues from our production agents. It's very hard to predict cost and it's very hard to manage cost

17:58

and put guardrails and solve this in a way where there is a reliable maybe fallback or have agents be using cheaper models for certain tasks. This is all driven by kind of the big AI vendors who I think their interest is for us all to spend more tokens. Replay and debugging, Uday talked about that. That's also a big, big issue. It's very hard to understand, but I think this is also something that is going to be be solved because we can now use agents to get over the cognitive overload of trying to debug what they do. And then standards. Standards are emerging by, you know, the community. Hotel, as I mentioned,

18:52

agent to agent is young. It's kind of pushed by certain vendors, but I think over time we will get there. With all of this said, you know, we know what we need and it's up to us to go ahead and build it. Thank you, everybody. Thank you. Thank you.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note