Open Reader

FinOps for AI Agents: Who Spent All the Tokens? — Tisha Chawla & Susheem Koul, Microsoft

completed 21:24 Aug 22, 2026 Watch on YouTube

Current Status

completed

Video ID

GJX19pNhmSw

RAG / Chat

Enabled
FinOps for AI Agents: Who Spent All the Tokens? — Tisha Chawla & Susheem Koul, Microsoft
Description

Turning the full policy suite on cut average agent spend by about 78% across benchmark runs on two open source repos, and lifted the share of runs that actually completed from 67% to roughly 96%. That second number is the point. Simple throttling holds the bill down by killing runs, so Tisha Chawla and Susheem Koul built their control plane to steer instead. Their framing is that every software era grew a control surface, usage caps in SaaS, autoscaling policies in cloud, and the agentic era still has none at the layer where code calls a model. Gateways can hard cap and downgrade, but nothing sits between your code and the spend it triggers. The design splits in two. In your code an annotation marks a boundary around methods you already have, floating attribution up without a rewrite, while a governor holds the list of actions you authorized so the control plane cannot do whatever it likes to your agent. The control plane groups runs into segments by any dimension you emit, sets budgets against a time window, and attaches policies. Actions come in two flavors. Halt is a circuit breaker. Steer is the interesting one: a cost guard watches both how much of the budget is gone and how fast it is going, and when it predicts an overrun it injects an instruction to keep outputs succinct rather than killing the run. They demo it in preview mode first, policies evaluating with enforcement off, which is how you would actually introduce this to a production agent. Speaker info: - https://www.linkedin.com/in/tisha-chawla - https://dev.to/tisha - https://www.linkedin.com/in/susheemkoul - https://susheemk.substack.com Timestamps: 0:00 - From token maxxing to value maxxing 1:58 - Every era got a control surface, this one has not 4:30 - First principles at the model call boundary 6:19 - Why gateways are not enough 8:57 - The SDK side: boundary, ledger, governor 13:13 - Segments, budgets, actions, policies 14:53 - Halt versus steer 16:36 - Demo: preview mode, then enforcement 18:1

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Agent cost control must move beyond per-request gateway caps to a run-level control plane that attributes cumulative token spend and steers agent behavior before resorting to termination.
  • Why it matters: Multi-step agents create costs through loops, sub-agent spawning, retrieval volume, and expanding context, none of which can be governed effectively by model routing or request-level limits alone.
  • Best use: Use this as a design reference for an agent FinOps/control-plane layer: instrumentation boundaries, attribution dimensions, budget policies, safe rollout, and steering interventions.

Executive Summary

Tisha Chawla and Susheem Koul frame AI FinOps as a shift from "token maxing" to "value maxing." Their argument is that token spend is created at the model-call boundary but must be governed at the agent-run boundary, where a system can see the cumulative cost and behavior of an entire workflow rather than isolated API requests. The motivating failure modes are runaway loops, uncontrolled sub-agent fan-out, excessive retrieval, and growing context windows that can exhaust budgets before conventional cloud-style controls notice.

Their proposed TokenOps architecture uses an out-of-band control plane connected to agent code through a bridge. Developers add attribution data and boundary annotations around methods or model objects; these send run-level telemetry and ledger entries upward while also providing a channel for approved control-plane actions to flow back into the running agent. A developer-configured governor limits which mutations are permitted, preventing the control plane from arbitrarily changing agent behavior.

The central distinction is between halting and steering. A hard budget cap acts as a circuit breaker, but the presenters prefer first attempting in-place interventions: compact context, reduce tool outputs, constrain retrieval chunks, inject instructions to be more concise, or otherwise mutate approved components. Their cost guard combines budget consumed with token-consumption velocity to predict a likely overrun and intervene before the run becomes unrecoverable.

The presentation is strongest as a control-plane pattern rather than as independently validated performance evidence. The team reports benchmarks on Browser Use and MetaGPT showing roughly 78% lower average spend and completion improving from 67% under simple throttling to about 96% with its full policy suite, but it does not provide task definitions, baseline configurations, quality measurements, statistical variance, or details sufficient to validate those results. The practical takeaway is to build run-level observability and safe steering first, then treat claimed savings as a hypothesis to test in Ken's own workloads.

Key Takeaways

  • Claim: Per-request model gateways are insufficient for controlling agent costs because the economically meaningful unit is the full agent run, including its iterative calls, tools, sub-agents, and context growth. | Evidence: The speakers contrast tools such as LiteLLM, Portkey, and Cloudflare, which they characterize as request-level systems offering routing or hard caps, with the missing ability to control the agent-tool-agent loop, sub-agent spawning, and expanding context across a run. | Implication: Instrument and budget workflows by run, task, user, tenant, cohort, or business use case—not merely by individual LLM request or model endpoint. | Caveat: Request-level gateways remain useful as a lower-layer guardrail for provider routing and absolute caps; the claim is that they do not replace run-level governance.
  • Claim: Cost governance requires attribution before enforcement: every model-related trace needs to be linked to the agent run and to relevant usage dimensions. | Evidence: The proposed bridge attaches agent-run IDs and attributes to boundary records, while the control plane can segment spend using dimensions such as a cohort tag (the example is "AI 2026") and apply budgets at run, agent, or cohort level. | Implication: A usable cost ledger should support both fine-grained debugging of an expensive run and roll-up accountability for users, teams, customers, experiments, or product cohorts.
  • Claim: Budget exhaustion should be the last resort; a control plane should first steer an agent back within budget through approved, in-place interventions. | Evidence: The policy catalog includes context compaction, tool-output reduction, loop and progress detection, and steering actions such as allow, mutate, and inject; halting is a separate kill action. Their retrieval example reduces an RAG tool from 20 chunks to five when later chunks are not being used. | Implication: Design degradation ladders for every agent: reduce retrieval breadth, summarize state, limit iterations, downgrade verbosity, or alter plans before killing a potentially valuable task. | Caveat: Steering can change output quality, completeness, or task behavior, so allowed interventions must be explicitly constrained and evaluated per workflow rather than universally enabled.
  • Claim: Predictive cost guards can intervene before a hard budget breach by considering both current spend and the rate at which tokens are being consumed. | Evidence: In the demo, the cost guard predicts that a run will exceed its allocation based on budget consumed plus consumption velocity, then injects a system instruction asking the model to produce more succinct, summarized outputs. | Implication: Monitor burn rate and projected end-of-run cost, not just cumulative spend; trigger graduated controls early enough for them to affect subsequent calls. | Caveat: The transcript does not explain prediction accuracy, false-positive handling, or whether a prompt-injection-style instruction consistently produces savings without reducing task success.
  • Claim: Governance should be introduced in preview mode before enforcement to tune policies safely against production-like agent behavior. | Evidence: Preview mode runs the policies and records their results in the dashboard but does not execute associated actions. The presenters position it as a way to test guardrails, observe policy behavior, and finalize thresholds before enabling enforcement. | Implication: Adopt a shadow-policy rollout: log would-have-fired actions, compare counterfactual risk and expected savings, then progressively enable low-risk steering before hard stops.
  • Claim: The proposed implementation minimizes framework coupling by treating instrumentation boundaries as a generic wrapper around methods or model-provider objects. | Evidence: A boundary annotation records method inputs and outputs as ledger entries and receives control-plane actions; a separate "wrap complete" helper covers object-based model APIs. The presenters say this works regardless of frameworks such as LangChain and that the control plane can live in the customer's own tenant. | Implication: Standardize an internal agent-runtime adapter that exposes trace context, approved action hooks, and policy receipts across frameworks rather than building FinOps logic into each agent. | Caveat: Despite being described as out-of-band, the design still requires adding annotations/wrappers and a governor configuration, so it is not zero-integration operationally.
  • Claim: The presenters report that steering policies can cut spend while preserving more completions than simple throttling. | Evidence: Across multiple iterations, stress tests, and simple and hard scenarios on Browser Use and MetaGPT, they report nearly 78% lower average spend with TokenOps' full policy suite and completion increasing from 67% under simple throttling to roughly 96%. | Implication: Use the figures to justify a controlled evaluation, with task-quality and latency metrics alongside token savings and completion rate, rather than using them as an ROI assumption. | Caveat: These are vendor-presented benchmark claims without disclosed workloads, task-quality scoring, cost model, policy settings, sample sizes, or variance; they should not be treated as generalizable production results.

Detailed Brief

Reference architecture and policy model

  • Claims: The control plane is organized around segments, a run ledger, budgets, actions, and policies.; Policies combine a budget threshold with an action and target either an attributed segment or an individual agent run.; The governor is the local safety boundary: it accepts only developer-declared action types and applies them non-destructively where possible.
  • Evidence: Segments can be created from arbitrary attribution dimensions and support coarse-grained cohort controls as well as run-level controls.; The ledger consolidates the traces associated with one agent run.; The architecture separates the agent runtime, bridge, and control plane; the bridge is bidirectional rather than telemetry-only.
  • Caveats: The presentation does not address authorization for who can create or modify policies, auditability of runtime mutations, tenant isolation mechanics, or failure behavior if the bridge/control plane is unavailable.
  • Implications: Treat the policy engine as production infrastructure with change control, policy versioning, audit logs, and fail-open/fail-closed decisions—not merely as an observability dashboard.; Separate a global safety ceiling from workload-specific policies; different customer tiers or task classes can legitimately have different economic envelopes.

Future direction: a self-learning policy layer

  • Claims: The intended end state is a control plane that learns from its continuously updated ledger.; The system would identify uncaught runaway-cost failure modes and either create new policies or refine existing policy parameters.
  • Evidence: The presenters explicitly propose using accumulated ledger data to ask which failure modes remain uncaught, then generate policies or tune parameters.; Existing coverage spans spend management, context management, loop detection, and progress detection.
  • Caveats: Automatically generated policies could create regressions, silently lower task quality, or optimize against easily measured token spend rather than business value; no validation or approval workflow is described.
  • Implications: If pursuing adaptive policies, require offline replay, shadow evaluation, quality gates, and human approval before a learned policy is allowed to mutate live agent behavior.

Notable Concepts & Terms

  • Token maxing to value maxing: The talk's framing shift: token consumption should be judged by attributable task value rather than celebrated as raw usage.
  • Agent-run control plane: A governance layer operating across the full lifecycle of one agent task, rather than independently governing each model API request.
  • Boundary annotation: A generic method/object wrapper that emits input-output telemetry with run attribution and receives control-plane actions.
  • Governor: The local runtime component that constrains which control-plane actions are permitted and how they are applied.
  • Ledger: The accumulated, attributed record of traces and spend for an agent run, used for debugging, budgeting, and future learning.
  • Steer-type actions: Interventions intended to preserve completion while lowering spend, including mutation, instruction injection, context compaction, and output reduction.
  • Preview mode: Shadow enforcement: policies execute and are observed, but their actions are not applied to the live run.
  • Cost guard: A predictive policy that uses consumed budget and token burn velocity to trigger earlier corrective action.

Operator Notes / Why Ken Should Care

  • Define a canonical cost-attribution schema for every agent run: run ID, workflow/agent version, tenant/customer, user, task type, environment, model/provider, and business outcome where available.
  • Build a policy ladder for each production workflow: observe only, warn/inject, reduce retrieval or context, limit loop/fan-out, and terminate only at a defined safety ceiling.
  • Run candidate policies in preview mode and review a sample of would-have-steered and would-have-killed runs for task-quality regressions before activating them.
  • Benchmark any steering implementation against a no-governance baseline and a hard-throttle baseline using completion, answer quality, latency, tool-call count, and cost—not completion rate alone.
  • Establish an explicit action allowlist in the runtime and an audit trail recording the policy version, trigger, mutation, and resulting run outcome.
  • Treat adaptive or self-generated policy proposals as offline recommendations until validated through replay or controlled experiments.

Source/Metadata

  • Title: FinOps for AI Agents: Who Spent All the Tokens? — Tisha Chawla & Susheem Koul, Microsoft
  • Transcript words: 5042
  • Duration seconds: 1284
  • Timestamp note: No usable timestamps or chapters were present in the supplied transcript; substantial material near the latter half is duplicated.

Transcript

3485 words en Processed in 137.5s

. Okay, good morning everyone. I am Tisha, and I have Sushim with me as my co-presenter. All right, so we will be talking about the most expensive question in AI today. I think a lot of you would have come across the scenario that when you opened an AI bill through your agentic workflows, you could not actually trace back where that bill was actually coming from. And I do not think that is a problem right now because right now the industry is valuing token maxing, that is spending the most amount of tokens for exploration, for all of those purposes. And people are proud to call themselves token billionaires, and I think that is all right. But this talk is the shift from token maxing to value maxing. How do we get there? And we will talk about it from this question: who spent all the tokens? And if anyone spent all the tokens, there has to be value associated with this, and that is the talk about. All right, now in order to minimize the gap from token maxing to value maxing, we will see the patterns which the existing past software evolution era has had. For instance, when we talk about the SAS era, the interface was UI, and the control was in the form of usage gaps, like the seat limits or tier-based policies. Now, when we moved on to the cloud era, the control surface again changed, the model became pay as you go, and the control moved in the form of auto provisioning and auto scaling policies. Now, we are in the agentic era. And now how the cost is calculated here is in the form of model calls, how the code calls your model. But what we have observed is that there is not a proper control plane in place for that. We do have a control plane in place as model gateways, where there are hard caps or there is model routing to downgrade the model. But the part where the code calls the model, that is what we will be talking about today. And we also see that in the last year, we have seen a lot of unbounded consumption happening. If you read the news, there was news about the AI budget for Uber getting exhausted within four months, and there were companies who ran into hundreds of millions of dollars within just months or days, and there were a lot of other news in place as well where these runaway loops led to a very massive increase in the cost. And there were not proper mechanisms to control it. So, when we see all of this, the first thing that comes to our mind is, is there a tool or is there a product to save us? But we will instead talk about the first principles of how we can design a system which is actually true enough to solve the problem from the very root. So, for that, let us dive on to the principles. First of all, let us talk about token being the unit of cost. We are charged in terms of token. So, now we have to see value also in terms of token. Next, we all know that cost is created at the LLM model call boundary. So, that is what we have to track. And if we do not have proper attribution, if we do not know what agent run made that particular call, we cannot control it. We just know the broad picture of what went wrong, but we cannot trace it back or narrow it down. So, that is why attribution is a very important element to have. And once you know which particular run or which particular agent is actually attributing to the cost, you should have proper policies in place to actually stop it. Let us take an example that if you have a loop which is running very excessively and which is not required, or if your context is growing very out of range, you should have policies in place which can solve that particular thing there and then instead of halting that. And as the last resort only, a halting should happen from a budget gap. So, these are the first principles. Now, let us see how we can define an ideal platform on top of these principles which we talked about. All right. So, one thing which is very important, which matters here, is that when we talk about the existing frameworks for token ops or for token management, most of them monitor the model request. They are model gateways which do model routing or hard budget capping. But what we need right now is something which monitors you at the run instead. If you see, we need something which can control the loop between the agent call, between the tool and the agent, something which can see or control the spawning of multiple sub agents happening from one main agent, or something which can control the growing of context. So, that is the need of the hour, and that is what we need. So, for all of this, we are proposing a platform which first of all has a cumulative budget across the attribution runs which happened. And then where enforcement actually happens in call path rather than a separate thing. For example, if something goes wrong, if your context is just growing heavily, then in-place compaction should happen, or in-place caching or something like that should happen. And after exhausting the list of all in-place policies, only the budget cap halt should happen at the very last. So, that is something which we are proposing. But if we look at the landscape today, if you see the tools like LiteLLM, Portkey, Cloudflare, all of those, they will happen at the request level. If you see, halting is there, routing is there for some of those, but all of this again is at a request, and you cannot control the cost at the request layer, at the model layer. So, this is the missing piece, which is basically navigating it at the level of the agent run layer. So, for that we have token ops, which is a runaway token governance for AI agents, and this is the architecture for that. So, first of all, one thing I would want to highlight is the intentional decision we took here was an out-of-bound plane. So, it does not interfere with your code at all. So, if you see here, that out-of-bound plane has three modules, which I will be talking about, the first one being instrumentation. It is a common observability layer where you will have the basic telemetry, the open telemetry, the cost in microns, the enrichment layer, and basically the attribution, what caused that particular run. So, if you see here, that is the accounting where you will basically accumulate it in a kind of a ledger, the total runs which are happening. And finally, we have this enforced layer, which has two main purposes. One is steering it through the policies which we have defined, which I think we will cover later, and then we have halt in place as the final thing if your budget is getting exhausted. So, yeah, that is there. Now, when we again look at the landscape, this kind of will solve a lot of problems which happened when we look at the previous tools or products there because it is happening at run and it is helping you solve the problem from the very root by steering it in place. All right. So, with this I would like to hand it over to Sashim for the demo. So, yeah, we have established the principles behind token ops till now. Now let's shift gears, talk about the design part of it, and maybe get into the code and eventual demo. So, what I have behind me on the screen is the bird's-eye view of what token ops looks like today. It's three layers. We will go left to right and top to bottom. So, on the leftmost, you have your own agent run time which you are trying to instrument and manage the cost for. The middle layer is what we are calling the bridge that basically shuffles data between your agent and the control plane, and the control plane is where the mind of the system lies. Right. So, let's talk about the bridge layer very briefly. If we go from top to bottom, you have the attribution on top. So, what we are trying to do here is every agent run that you do, it's attributed to some user dimensions. So, the idea is everything that you do, every run of the agent, is accounted to some usability or some usage. This comes in handy later, we will talk about it. The second part, which is the boundary annotation that you see, this is pretty much the heart and soul of this middle layer. So, the idea behind the boundary annotation is that you take any method, it doesn't matter what framework you are using, you might be using, let's say, Langchain, Langspot, whatever. If you have a method, you can annotate it with boundary. What this annotation is going to do is, it's going to do two things. First, it's going to track the input and the output, and it's going to flight that up If we go from top to bottom, you have the attribution on top. What we are trying to do here is, every agent run that you do is attributed to some user dimensions. The idea is everything that you do, every run of the agent, is accounted to some usability or some usage. This comes in handy later. We will talk about it. The second part, which is the boundary annotation that you see, is pretty much the heart and soul of this middle layer. The idea behind the boundary annotation is that you take any method. It doesn't matter what framework you are using. You might be using, let's say, Langchain, Langspot, whatever. If you have a method, you can annotate it with boundary. What this annotation is going to do is, it is going to do two things. First, it is going to track the input and the output, and it is going to flight that up to the control layer and record it there as a ledger entry. Now, this will be annotated with the further agent run ID and the other attributes, and so on. The second thing the boundary annotation does is, it acts as a channel through which the control plane can push actions down to the agent. This is where the entire intelligence lies. We do not have a single-directional highway. We want the control plane to be able to tweak the behavior of the agent on the fly to ensure that we are able to squeeze in more runs inside our budget gap. Right? Now, let's say the control plane pushes down an action. Let's take a small example. Let's say you have a RAG retrieval tool which is generating 20 chunks every retriever for every call, and that's eating up your budget. And let's say the LLM is not even using the chunks that are after five because they are not relevant. Right? They are sorted by relevance. So, let's say the control plane observes this and it wants to limit the output to just five chunks. So, it can push down an action, but that action has to be received by boundary and then has to be executed by something. That is where the third node, the governor node, comes in. The governor knows what actions are allowed on your agent by you as a developer, and it receives those actions from the control plane and knows how to apply it in a non-destruction. In an interactive way. So, that's the first three. The fourth one, the wrap complete, is essentially just a helper method. As we know, most of the agent providers or the model providers provide objects rather than methods for their LLMs. Right? So, wrap complete is just another way of applying boundary on objects rather than methods. Let's shift right to the control plane. On the control plane, the first layer is the segment. Now, this is where the attribution that we talked about earlier comes into picture. So, any dimensions that you float from the attribution layer. Let's say you have a preview agent that you share with everyone in this room, and your agent is floating a dimension saying that cohort is AI 2026. Right? So, you can create a segment which is a cohort of users, which is based on this tag-like dimension being AI 2026. Right? And you can apply your budgets at this cohort level. So, you don't necessarily have to restrict everything at an agent level or a run level. You can do rollups. You can do fine-grain or coarse-grain control. Right? So, that's the segmentation part of it. Ledger, as I mentioned, is just one agent run, all the traces in one place. Then you have budgets. Budgets are basically just the static thresholds that work across a time window against a particular segment or an agent run. And then you have actions. On the actions part, we have broadly two flavors. First is the halt-type actions, which basically just kill your agent if it exceeds a budget. The second part, where we are adding value, is the steer-type actions. So, here we do not kill the agent. Instead, we try to steer the behavior of the agent or the components of the agent to try and fit that particular run within the allotted budget. Right? And then the policies layer is where it all comes together. You basically group the budgets, the actions, and then set your policies against certain segments or agent runs, and that is where it executes. Right? So, moving on. What changes in your code? That is the boundary annotation that we just talked about. As Tisha mentioned earlier, this is all out of band. So, you do not have to change your code. You just have to apply the annotation on the methods that you have. This boundary annotation will take care of floating all the information up to the control plane. And the control plane lies in your own tenant. So, you do not need to worry about any data leaks or anything. Then, if I talk about the governor. For the governor, you just have to create an instance. You just have to pass it to your own configs. These configs will basically declare what sort of actions are allowed for those agents. Right? So, your control plane cannot just willingly do any random things on your agents. So, before we move on to the demo, I will just briefly touch upon the testments that we are going to use today. It is a simple two-agent workflow. We have a research agent which has access to a search tool. You give it a question. It is allowed to look up on the web as many times as it wants. And once it knows that it has all the data, it passes the findings on to the second agent, which is a summarizer, which creates a research report. Right? So, with that out of the way, let's just quickly walk over to the demo. For the demo, we have three different scenarios that we are going to talk about. For the first one, we are going to run the token ops in what we call preview mode. In preview mode, what happens is that all the policies run as is, but the enforcement doesn't happen. So, if you see, we ran a particular run over here which completed, but we did not see any sort of failures there. The policies executed, but the actions that were associated with those policies were not allowed to be executed. So, we are just going to load the dashboard screen here. Yeah. So, this is the governance output. Governance is off, the run completed, but in the dashboard you can see the policies have executed. So, you can see the cost budget, the cost guard, and so on. Right? So, this was the first scenario. For the second scenario, what we are going to do is, we are going to turn on the governance now. While that is happening, I just want to touch upon why this is important. So, if you want to include this product into your production agents, you want to have a safe environment or a safe way to firstly put it in your production environment, test the guardrails, tweak the guardrails, see what the policies are doing, and then finalize the thresholds. Right? So, this is the second one where we have now enforced the governance, and you can see in the dashboard that the pre-call cost cap has exceeded. So, you had a budget allotted for this run, but the agent exceeded the budget and it was killed immediately. So, that is the simple circuit-breaker sort of methodology. So, this is the halt behavior. And now, let us see the steer behavior, which is where we are trying to add value to this entire cost management scenario. So, this time we are going to run the second prompt. The budget allotted for this one is slightly higher, but it is still not high enough for the agent to complete in time. So, what instead happens is, there is something called cost guard which kicks in. This cost guard takes into account two things. First, how much of your allotted budget have you consumed? Second, what is the velocity at which you are consuming tokens? Now, based on these two things, if it predicts that you are going to run out of your tokens or your allotted budget by the end of the run, it is going to inject something into your system instructions. That something could be as simple as, hey, you are running out of budget, so make sure that the LLM outputs are more succinct or more summarized, right? So, that is the way we are doing this tier. Now, this was a very simple test bench to show you how this works on working code. We have also benchmarked it on a couple of open source reports. So, we have benchmarked it on browser use as well as meta GPT. We ran it across multiple iterations, across stress tests, across simple scenarios, hard scenarios, and everything. And the results we see are, the average spend goes down by almost 78% with token ops enabled, with the full policy suit that we have today. On the completion part, when we compare it with throttling, just simple throttling, your simple throttling is going to kill your agent runs no matter what, right? So, with the reduced average spend, what you get is, you get an uplift in that completion percentage from 67% to roughly 96%. So, that is the value add that token ops is doing here. Now, this is the policy catalog that we run this benchmark against. This is what we support today. We research what are the different failure modes that are there today out in the wild and tried to cover most of them here. So, we have benchmarked it on browser use as well as Meta GPT. We ran it across multiple iterations, across stress tests, across simple scenarios, hard scenarios, and everything. And the results we see are: the average spend goes down by almost 78% with token ops enabled with the full policy suite that we have today. On the completion part, when we compare it with throttling, just simple throttling, your simple throttling is going to kill your agent runs no matter what, right? So, with the reduced average spend, what you get is an uplift in that completion percentage from 67% to roughly 96%. So, that is the value add that token ops is doing here. Now, this is the policy catalog that we run this benchmark against. This is what we support today. We researched what are the different failure modes that are there today out in the wild and tried to cover most of them here. So, you have things across spend management, you have things across context management like context compaction, tool output reduction. You have things across loop detection and progress detection and stuff like that. So, this is the entire set of policies that we support. And at the bottom, you can see the actions. So, as I mentioned earlier, we have two flavors. You have the halt-type actions and then the steer-type actions. So, for the steer, we can do allow, mutate, inject, and so on. And for the halt, it can be a simple kill. But this is not the end state that we envision for this. The end state is, we have a lot of data, right? We have a ledger that is continuously being updated. So, what we want to try is a self-learning module within the token ops plane, within the control plane, which can look at this ledger and ask this question: hey, why or what is the failure mode that I am still not able to catch? And then based on that, it can do two things. One is it can enhance. It can generate new policies on the fly based on the missing or the still runaway costs. Or it can refine the existing parameters for the existing policies that are there, so that the runaway costs are managed more effectively in the future. So, with that, I think that is all we have for you guys today. Thank you so much for your time. And you can scan this QR code. That is the public wiki. We are updating it almost regularly. So, you can scan this and stay up to date. And Tisha and I are around. So, if you guys have any questions or if you want to discuss more about it, just let us know. That's it. Thank you. Instead, we try to steer the behavior of the agent or the components of the agent to try and fit that particular run within the allotted budget. Right? And then, the policies layer is where it all comes together. You basically group the budgets, the actions and then set your policies against certain segments or agent runs and that is where it executes. Right? So, moving on. What changes in your code? That is the boundary annotation that we just talked about. As Tisha mentioned earlier, this is all out of band. So, you do not have to change your code. You just have to apply the annotation on the methods that you have. This boundary annotation will take care of floating all the information up to the control plane. And the control plane lies in your own tenant. So, you do not need to worry about any data leaks or anything. Then, if I talk about the governor. So, for the governor, you just have to create an instance. You just have to pass it to your own configs. These configs will basically declare what sort of actions are allowed for those agents. Right? So, that your control plane cannot just willingly do any random things on your agents. So, before we move on to the demo, I will just briefly touch upon the testments that we are going to use today. So, it is a simple two agent workflow. We have a research agent which has access to a search tool. You give it a question, it is allowed to look up on the web as many times as it wants. And once it knows that it has all the data, it passes the findings on to the second agent which is a summarizer which creates a research report. Right? So, with that out of the way, let's just quickly walk over to the demo. So, for the demo, we have three different scenarios that we are going to talk about. For the first one, we are going to run the token ops in what we call preview mode. So, in preview mode, what happens is that all the policies run as is, but the enforcement doesn't happen. So, if you see, we ran a particular run over here which completed, but we did not see any sort of failures there. The policies executed, but the actions that were associated with those policies were not allowed to be executed. So, we are just going to load the dashboard screen here. Yeah. So, this is the governance output. Governance is off, the run completed, but in the dashboard you can see the policies have executed. So, you can see the cost budget, the cost guard, and so on. Right? So, this was the first scenario. For the second scenario, what we are going to do is, we are going to turn on the governance now. While that is happening, I just want to touch upon why this is important. So, if you want to like include this product into your production agents, you want to have a safe environment or a safe way to firstly put it in your production environment, test the guardrails, tweak the guardrails, see what the policies are doing, and then finalize the thresholds. Right? So, this is the second one where we have now enforced the governance and you can see in the dashboard that the pre-call cost cap has exceeded. So, you had a budget allotted for this run, but the agent exceeded the budget and it was killed immediately. So, that is the simple circuit breaker sort of a methodology. So, this is the halt behavior. And now, let us see the steer behavior, which is where we are trying to add value to this entire cost management scenario. So, this time we are going to run the second prompt. The budget allotted for this one is slightly higher, but it is still not high enough for the agent to complete in time. So, what instead happens is, there is something called cost guard which kicks in. This cost guard, it takes into account two things. First, how much of your allotted budget have you consumed? Second, what is the velocity at which you are consuming tokens? Now, based on these two things, if it predicts that you are going to run out of your tokens or your allotted budget by the end of the run, it is going to inject something into your system instructions. That something could be as simple as, hey, you are running out of budget, so make sure that the LLM outputs are more succinct or more summarized, right? So, that is the way we are doing this tier. Now, this was a very simple test bench to show you like how this works on a like working code. We have also benchmarked it on a couple of open source reports. So, we have benchmarked it on browser use as well as meta GPT. We ran it across multiple iterations, across stress tests, across simple scenarios, hard scenarios and everything. And the results we see are, the average spend goes down by almost 78% with token ops enabled with the full policy suit that we have today. On the completion part, when we compare it with throttling, just simple throttling, your simple throttling is going to kill your agent runs no matter what, right? So, with the reduced average spend, what you get is, you get an uplift in that completion percentage from 67% to roughly 96%. So, that is the value add that token ops is doing here. Now, this is the policy catalog that we run this benchmark against. This is what we support today. We kind of research what are the different failure modes that are there today out in the wild and tried to cover most of them here. So, you have things across spend management, you have things across context management like context compaction, tool output reduction. You have things across loop detection and progress detection and stuff like that. So, this is the entire set of policies that we support. And at the bottom, you can see the actions. So, as I mentioned earlier, we have two flavors. You have the halt type actions and then the steer type actions. So, for the steer, we can do allow, mutate, inject and so on. And for the halt, it can be a simple kill. But this is not the end state that we envision for this. The end state is, we have a lot of data, right? We have a ledger that is continuously being updated. So, what we want to try is, we want to try a self-learning module within the token ops plane, within the control plane, which can look at this ledger and ask this question, hey, why or what is the failure mode that I am still not able to catch? And then based on that, it can do two things. One is it can enhance. It can generate new policies on the fly based on the missing or the still runaway costs. Or it can refine the existing parameters for the existing policies that are there. So, that the runaway costs are managed more effectively in the future. So, with that, I think that is all we have for you guys today. Thank you so much for your time. And you can scan this QR code. That is the public wiki. We are updating it almost regularly. So, you can scan this and stay up to date. And Tisha and I are around. So, if you guys have any questions or if you want to discuss more about it, just let us know. That's it. Thank you.