From coding to Knowledge work agents — Karan Vaidya, Composio
Description
Karan Vaidya pointed his own OpenClaw at hiring outreach and it mass emailed candidates exactly as instructed. Some of the people in the room had received one. The thread that followed put his name on Twitter, and every check in the software engineering playbook would have passed. The addresses were real, the emails were valid, they reached actual people. Nothing tested the only question that mattered, which was whether the outreach should have gone at all. Vaidya, cofounder and CTO of Composio, turns his own disaster into a larger claim: coding agents did not race ahead because models are better at code, but because code already had everything an agent needs around it. He walks six primitives coding got for free and knowledge work lacks entirely. Centralization, since one deal is scattered across Salesforce, Notion, Gmail, Slack and Zendesk behind five separate logins. History, because git records every change while business apps keep nothing an agent can read back. Context, both the map of how an organization works and the unwritten sense of what good looks like. Then verification, governance and reversibility. On governance he cites the alignment director at Meta whose email agent kept deleting messages after she told it to stop, until she reached a physical machine. Two hundred were gone, and she had prompted it in advance to confirm. A prompt lives in memory that gets compacted away, so Composio puts the wall outside the agent. Reversibility is the one he concedes is hardest, because a sent email or a wire cannot be called back. Speaker info: - https://x.com/KaranVaidya6 - https://www.linkedin.com/in/kaavee315/ - https://kvaidya.com/ Timestamps: 0:00 - Why nearly all agent tool calls are software engineering 0:43 - From autocomplete to autonomous in three years 1:26 - The scaffolding code already had 2:35 - Centralization, and one deal across five apps 4:11 - History, and agents that start blank every time 6:55 - Two kinds of context, the map and the style 9
Summary
Generated by gpt-5.6-terraAt-a-Glance
- Verdict: Watch fully
- Core thesis: Knowledge-work agents lag coding agents not primarily because of model capability, but because they lack six operational primitives that code already has: centralized context, history, organizational context, verification, governance, and reversibility.
- Why it matters: This is a practical architecture argument for moving agents from unreliable cross-SaaS automation to trusted operational systems, especially where actions are consequential, distributed across tools, and difficult to undo.
- Best use: Use it as a design checklist for an agent control plane: identify which of the six primitives are absent before granting an agent real-world authority.
Executive Summary
Karan Vaidya argues that coding became the first major autonomous-agent domain because its surrounding environment was already agent-ready. Repositories centralize truth, Git preserves history, tests and CI verify outputs, branch protections constrain releases, and version control enables reversion. Better models and coding harnesses mattered, but they were insufficient on their own.
Knowledge work has the opposite operating environment. A single customer or deal can be fragmented across Salesforce, Gmail, Notion, Slack, Zendesk, and other systems, with separate identities and incomplete histories. An agent therefore begins by reconstructing context, has little durable memory of prior work, and cannot reliably infer how a specific organization or user wants work done.
The proposed answer is a shared infrastructure layer that centralizes connected applications and access, records every tool action, distills operating patterns into reusable skills, validates work before it is real, and enforces controls outside the model prompt. Vaidya frames this as the missing substrate for dependable knowledge-work agents rather than another model improvement.
The strongest practical point is that knowledge-work failures are often irreversible: sent emails, deleted records, and completed wires cannot be treated like bad commits. For reversible actions, provide undo; for irreversible ones, execute first in a sandbox or require approval. The presentation is also a Composio product thesis, so its claimed implementation and scale should be evaluated independently.
Key Takeaways
- Claim: The autonomy gap between coding and knowledge work is largely an infrastructure gap, not a model-capability gap. | Evidence: Coding agents operate alongside repositories, commit history, tests, CI/CD, code review, linters, and reversion paths; Vaidya contrasts this with support, sales, finance, and hiring workflows spread across disconnected SaaS systems. | Implication: Do not assume a strong coding agent will transfer directly to business workflows. Build the operational environment and controls around the model before expanding its authority. | Caveat: The claim that software engineering is "100% autonomous" is rhetorical; production software work still commonly requires human oversight, product judgment, and deployment controls.
- Claim: Centralizing cross-application truth and identities is the baseline requirement for a useful knowledge-work agent. | Evidence: A deal may require Salesforce records, Notion documents, Gmail correspondence, Slack conversations, and Zendesk support history; today an agent must independently find and stitch those threads together across separate logins. | Implication: An agent platform should expose a governed, unified workspace of connected data and permissions rather than make every task an ad hoc sequence of SaaS searches and authentications.
- Claim: An action record gives agents both memory and operators an audit trail, enabling trust to be earned from observed behavior rather than accepted from model assertions. | Evidence: Vaidya proposes logging every action across every app, including what the agent touched, skipped, succeeded at, and failed at; he compares this to Git history, which lets coding agents inspect prior changes and lets humans verify what occurred. | Implication: Treat tool-call traces and state changes as first-class system data. They can support incident review, replay of successful procedures, and progressive increases in agent autonomy. | Caveat: Logging alone does not establish correctness; it provides observability and examples for review, but still needs validation rules and governance.
- Claim: Useful organizational context must include both workflow topology and local standards of quality, then be distilled from operational history into reusable skills. | Evidence: The speaker distinguishes architecture context—how data and processes flow across systems—from style context—how a company or individual prefers work done. His example of drafting a customer document requires combining product usage, PostHog behavior, and Salesforce deal details before writing begins. | Implication: Build agent memory around organization-, team-, and user-level preferences, but give owners a way to inspect, curate, and override learned procedures. | Caveat: Observed historical behavior can encode inconsistent, outdated, or low-quality practices, so inferred skills should not automatically become policy.
- Claim: Knowledge-work agents need pre-execution verification because syntactically valid tool actions can still be reputationally or operationally wrong. | Evidence: Vaidya recounts using OpenClaw for mass recruiting outreach: the emails were valid and delivered to real people, yet the campaign was a public failure. He proposes checking drafts against prior approved email style and running destructive actions against sandboxed versions of real tools before production. | Implication: For external communications and high-impact workflows, validate the business appropriateness of an action—not merely whether the API call succeeds—and use simulation or review gates before live execution. | Caveat: Similarity to prior messages is not a complete quality test; policy, recipient suitability, timing, factual accuracy, and business intent also need independent checks.
- Claim: Agent governance must be enforced outside the prompt through both deterministic access boundaries and behavior policies. | Evidence: The talk compares code branches, code owners, and preview deployments with proposed controls such as allowing a hiring agent to read email only, allowing a support agent to create drafts but not send, and applying natural-language policies such as never deleting more than 10 emails without approval or never emailing outside a domain. | Implication: Separate capability authorization from task instructions. Define least-privilege tool scopes, then layer policy evaluation, approval thresholds, and escalation paths based on blast radius. | Caveat: Natural-language policies can themselves be ambiguous or imperfectly interpreted; deterministic permissions remain the harder boundary for irreversible or high-blast-radius actions.
- Claim: Because many knowledge-work actions cannot be undone, trust must be established before execution rather than repaired afterward. | Evidence: Examples include sent emails, completed wires, and hard-deleted records. The speaker cites an anecdote about an email-connected agent continuing to delete messages despite a prompt-level instruction to confirm first, with roughly 200 emails lost; he proposes actual undo for reversible actions and sandbox-first execution for irreversible ones. | Implication: Classify actions by reversibility and blast radius. Give agents autonomous execution only for low-risk or reversible operations; require sandbox validation and explicit approval for irreversible production changes. | Caveat: A sandbox is only as valuable as its fidelity to production, and it cannot reverse an already-sent external communication or completed financial transaction.
Detailed Brief
A six-primitive control-plane model for knowledge-work agents
- Claims: The six primitives implied by the presentation are centralization, history, context, verification, governance, and reversibility.; These are interdependent rather than isolated features: centralization makes comprehensive records possible; records create the raw material for context and learned skills; verification and governance make greater autonomy tolerable; reversibility determines whether controls must occur before or after action.; Governance should vary with blast radius instead of imposing a single universal approval gate, analogous to letting coding agents work freely on a branch while restricting merges to main or production deployment.
- Evidence: The coding analogy uses repositories as a central truth source, Git as history, codebase conventions as local context, tests and linters as verification, branch/code-owner/deployment restrictions as governance, and revert/bisect workflows as reversibility.; Composio positions its product as the layer through which app connections, credentials, records, controls, policies, and sandboxes can be managed.; The speaker says Composio has processed more than one billion tool calls in total and approximately 300 million tool calls monthly.
- Caveats: The latter scale and product-capability statements are vendor claims in a hiring/product presentation, not independently substantiated in the transcript.; The framework does not resolve difficult semantic questions such as whether a commercially sensible email, escalation, or payment should occur; it supplies operational mechanisms that can make those judgments safer to test and enforce.
- Implications: The durable competitive layer in agents may shift toward identity, state, policy, auditability, sandboxing, and workflow-specific evaluation rather than model access alone.; A generic connector catalog is insufficient if it lacks unified state, durable action records, policy enforcement, and action-class-specific rollback or simulation.
Notable Concepts & Terms
- Six primitives: The speaker's framework for bringing coding-like autonomy to knowledge work: centralization, history, context, verification, governance, and reversibility.
- Centralization: A unified layer for connected applications, credentials, permissions, and cross-tool information so agents do not begin every task by reconstructing the world.
- Record of work: A durable log of agent actions and results that functions as both agent memory and human audit infrastructure.
- Skills: Distilled patterns from operational records that capture how a tool works, how a company works, and how an individual prefers work to be done.
- Sandbox-first execution: Running an action against a mock or non-production version of a tool before allowing it to affect the real world; intended for destructive or irreversible actions.
- Deterministic controls: Hard capability boundaries outside the model context, such as read-only email access or draft-but-not-send permissions, which cannot be bypassed by prompt loss or compaction.
- Natural-language policies: Behavioral constraints applied on top of granted access, such as deletion limits or recipient-domain restrictions.
- Reversibility: Whether an action can be reliably undone; it is the key determinant of whether an agent can be trusted post hoc or must be checked before execution.
Operator Notes / Why Ken Should Care
- Create an action taxonomy for Ken's agents: read-only, reversible write, externally visible write, destructive write, and financial/legal commitment; assign default permission, verification, approval, and rollback rules to each class.
- Require all production-capable agents to emit an append-only action ledger with tool inputs, outputs, state diffs, policy decisions, approvals, and correlation IDs across systems.
- Audit current agent workflows for prompt-only safety rules; move non-negotiable restrictions into OAuth scopes, service permissions, middleware policy checks, and rate/volume limits.
- Prioritize sandbox or dry-run adapters for irreversible operations, especially sending bulk communications, deletion, permission changes, record mutation, and financial actions.
- Evaluate Composio or equivalent infrastructure against the full six-primitive checklist, not merely connector breadth or nominal tool-call volume.
Source/Metadata
- Title: From coding to Knowledge work agents — Karan Vaidya, Composio
- Transcript words: 3957
- Duration seconds: 1241
- Timestamp note: No timestamps or chapters were present. The transcript repeats the closing reversibility section and conclusion.
Transcript
Karan Veddhya Reviewer Reviewer Reviewer Reviewer Hey folks, I'm Karan Veddhya, co-founder and CTO of Composio. Most agent-trick tool calls today are still happening in one field. No guesses, it's software engineering. Every other kind of work is trailing far behind. If models keep getting better, then why are we still limited to just agentic coding? That's the trillion-dollar question I'm here to answer. Three years ago, coding agents were just autocomplete. Today, software engineering is fully autonomous. We went from pressing tab, tab, tab to let's just cloud cook. That's just magic. And why did it happen so fast in coding? Most people would think it's models. Yeah, models got really better over time over the last two to three years. And so did the harnesses: cloud code, codex, cursor. But on their own, it wouldn't have been enough. It only worked because all the infrastructure and systems around coding were literally meant for agents. Code came with the support that agents needed. You have got the repo, the commit history, test, CI, CD, review, linters, revert if anything goes wrong. The kind of stuff that makes you trust the agents, the systems around code. Now, we are pointing these same amazing agents at everything else. Support, finance, sales. But the agents that were doing phenomenally well in coding are just working blind. Because the infrastructure around coding doesn't even exist in other fields. So how do we close the bridge between coding agents and knowledge work agents? We think it's core six primitives, and coding had all six of them, while knowledge work doesn't have any. And that's what we need to build. First is centralization. Coding agents worked pretty well, partly because they were very near the source of truth. They knew the what, the why, and how. You give them the repo, the infrastructure as code, and you close the loop and let the model cook. The agent starts with everything they need, all in a single place, that is the code base. This is exactly what knowledge work missed today. For example, a single deal is scattered across five different platforms. The records are in Salesforce, the docs in Notion, the emails in Gmail, conversations in Slack, and the support history is in Zendesk. There's no single source of truth, single place to get all the information. Everything is separate, and every app has its own login. Before a knowledge work agent can even start to do things, it has to go and pull all the threads and tie them together itself. And that's still the base point where coding agent had started. It already had it all. So how can you expect knowledge work to do, knowledge work agent to do the same level of work as coding agent? So the first thing we build is the missing center, one place where all your apps, all your connections, all your logins exist. So the agent doesn't need to do the hard work of stitching them all together. They find it all in a single place. And they get the baseline that the coding agent started with, which is the repo, the information across all the stacks in one single place. That's the foundation you start with, and you can give right accesses to your agent. The next thing agent needs is a sense of history, the ability to look back in the past. In code, you get it for free. Git keeps a record of every single thing that went in, every single change that was made. So the agent can always look back and see how a certain change was made, why something worked, why something didn't work. Think about the kind of thing you actually ask your agent to do. We had to revert a change in the past because of some failure, but that was pretty hard to pull off. Can you look at it and get it back again? It just reads through the history and get it back and cook it. The history isn't just for agent. It's also for you to keep a record what the agent is doing. You can see what the agent is doing, where it is fucking up, where it is doing successful things. And instead of trusting what the agent is saying to you, you can just go to those particular apps and look at what it has done. Now ask those same questions about knowledge work. What led to the CRM being in a state where it is today? How did my colleague craft that amazing mail that led to the closing of the deal? What's the actual process to escalate a support issue or even close one? The answers are smeared across hundreds of apps, and none of them keep the history. So the agent has no memory. It starts from blank state almost every time. No idea what was tried before, what worked, what didn't work. And you have nothing to look at as well. Once the agent runs, it tells you it has done successfully. You don't know if it has actually done successfully. There's no way to know if it is right or not. And that's what's missing, a record of work. Now because everything finally runs through one single place, that centralization, we can build a layer on top of it, the record. Every single action that the agent takes can be logged across every other app. Whatever it touched, whatever it skipped, what worked, what didn't. Via this, firstly, the agent gets memory. It can look back at how similar tasks were done before, what was successful, and replicate it again. It doesn't start with a blank state all the time. Second, you get trust. You can finally see exactly what the agent is doing. So instead of hoping it will do the right thing, you can just go back and check and catch it if it does something bad. And as you see it more and more doing the right things, you'll develop the trust and offload more tasks to it. The next thing an agent needs is context. And there are really two kinds of context, if you think about it. The first is the shape of the platform, the architecture. How things flow into each other, how things are tied, the data flows. Like a map which a senior engineer carries in their head and a junior engineer takes probably three months to develop. The second is style. This isn't what's objectively correct, but more like what good looks like in your company. So how you do things, things like linter, type checks, et cetera. And maybe you use a TypeScript decorator, which nobody else would. This is not exactly somewhere in a playbook. It's more in your code base. It's all available in your code base. So the agent can just go and look and figure out the specs, what you like, the linters, the formatters, et cetera. Now, coming to knowledge work, the same thing. Say you're writing a doc to a customer. To even start, I would have to open the database to pull their usage, check post hoc of how they have been actually using things, and Salesforce to look at their deal details. Only then I can even start writing the first line of the doc. The answer wasn't isolated in just one of those tools. I'm able to write this doc because I'm pulling the threads across all these tools into one single context in my head. So putting history and context together, that's how you map how the organization works. And that part is not available to agent today. So as we did centralization and logging, the record we just built, the one that gives the agent memory and lets you check what it did, also does one more interesting thing. If you log enough of what every agent is doing, you start to see patterns. You start to see how the organization works, and you start to form skills, which is some sort of distillation of how the organization has been working. Which approaches work, which don't, what led to failures in the past, et cetera. The record isn't just history of what happened anymore. It's a picture of how your company operates. And it actually works at three different levels. How a tool works in general, which is applicable to every person. How a company does things. And how you prefer to do things. What good looks like to you. And that's the context that was missing for a knowledge work agent. How the work actually gets done. The real playbook of sorts. And the preference of a company, of a personal user. And now the agent can query it and stop guessing how the company operates. The other reason coding agents work so well: they test themselves. The work checks itself. Verification. The moment the agent writes a code, a stack of checks follow. The unit tests can catch small mistakes. The integration tests catch the ones that only affect components three blocks away. The type system would not even work and run if anything would happen. The compiler will not even build. On top of it sits the software checks: linters, formatters, bugbot.md, review skills, et cetera. And these ensure that the code matches the way your team likes to follow, the standards of your team. None of it needs you. The agent completes the loop on its own. And makes sure that it follows the standard and is able to make the code run. Now think about, so there's a while back. I pointed my open claw at a hiring outreach. Mass emails to candidates. It ran. It sent tons of emails. Some of you might have also gotten it from my open claw. It did exactly what I told it to do. It was also a disaster. The kind that ends up on Twitter with my name on top of it. Yeah, I think you can see a fuck you Karan Vedya. I was not the happiest when it happened. And here's the thing. Every check from the past slide would have passed. The emails were valid. The addresses were real. It actually got to real people who posted. There was no test in the world to actually question what really mattered. Should this have gone at all? That's the gap. In code, these tests tell you what's wrong and right. Here, the internet told me that I was wrong. So we build the checks that are missing. The problem in the above thread wasn't the outreach was wrong. It was that it went out before even I got to know. So the fix is simple. Catch before it's even real. So we have two ways in which we do that. One, before the agent sends anything, it checks the draft emails that I've sent before. If it matches my style. If it matches the goodness that I like. The second, before doing anything destructive in the real world scenario, we provide the agent sandboxes which mock the real tools. And they can send, they can do action on top of these sandboxes. So instead of the blast radius hitting the real world, it will hit a sandbox and then I can review it before the agent does the real thing. Put those two together and you've got something knowledge work never had. A way for agent to check its own work before it's even real. It can finally close its own loop instead of stopping to wait for you. And with all that, you can trust the action it is taking without you getting bombarded with the tweets that I shoot. Next thing, the agent needs is governance. Building trust is controlling what the agent can do. Putting up the right walls around the agents. In code, this is mostly solved. And has multiple layers. The agent can do whatever it wants on its own branch. But it can't merge to main. A human reviewer sits in between it merging to main. The critical files have code owners. So whenever it touches one of them, the right people are getting involved. We use agents to ship to preview deployments, never let it touch the production deployments, so we control it there. The governance is not a single gate, but multiple of them. And each varying its sizes depending on the blast radius it exposes. None of it slows the agent down in safe parts. Just prevents it from fucking up production. And the tighter those lines are, the more you can trust the agent and let it go berserk. You probably saw this one. The director of alignment at Meta Super Intelligence Lab hooked up an agent to its email, and it started destroying its email, deleting a lot of them. She told it to stop. It kept going. Finally, she had to run to a physical machine to stop it. But by then, 200 emails had actually vanished. She had told it beforehand in prompt to confirm before acting on such cases. But that was just a prompt, which probably would have compacted away. And if someone whose sole job is AI alignment can't prompt the agent correctly, then probably none of us can. And that's the real reason these agents are so hard to trust. Not because they're worse than the coding agents, but because there's no wall around them. In code, the wall was already built into the system while we were developing earlier. Knowledge work also has some bits and pieces here and there. For example, Gmail has scopes. Salesforce has permission levels. But it's so scattered all over the place that it's very hard to have real control, and mostly people end up doing it via prompting. And prompting is fragile. The agent will find those loopholes. Things will get compacted away. And at scale, one of these fence will break, and you'll also be in the same condition where 200 of your important emails are vanishing. So what would actually stop it? Not a better instruction, but a wall that the agent can't cross, even if it forgot that wall existed. So we build these walls in two layers. The first layer is deterministic. Control over what the agent can reach, what it has access to. A hiring agent can probably just read the emails. A support agent can create a draft email, but not actually send it. The boundary lives outside these agents. It can't be argued with by the agent, or forgotten, or compacted. Use instruction failed because it lived in agent's memory in the prompt. This doesn't. But access alone wouldn't have saved her because she was actually building an email agent, so it definitely needed access to that email. The other thing that we do is provide policies, which is you can define national language policies of what the agent can do even with those accesses. So things like never delete more than 10 emails without my permission. Never email outside a particular domain. Rules that, with even those access, control the behavior. So between those two things, one layer controls what the agent can reach, and the other layer can control the behavior with what it can do with that reach. Together, it's real governance for the agent. Not asking the agent to behave, but enforcing what it can do. The last pillar, reversibility. And this is where we reach when things go wrong. Can I undo it? In code, you almost always can. Every change is recorded. Things can be walked back. You can revert the last commit, or you can bisect to the commit that broke your production and revert it. Now, I'm not saying it's good. I won't pretend like that. If things go in production and break, it's always bad. But it's still not permanent. You can still walk back from it. And that's what gives you confidence to let your agents cook and let them do some magic because even if they break things, you have a pathway back. For knowledge work, there is no undo button. Think about use inbox. Those 200 emails are gone. They have vanished. That's the normal case, by the way. The disaster cases, a sent email, which you can't revert back. A wire that has already been made, so you can't get that money back. A deleted record, gone forever. Most actions actually in knowledge work don't have an undo button. And that changes the whole equation. That changes the blast radius. With code, you can trust the agent after the fact. Let it run, check the result, undo if it's wrong. Out here, there's no coming back. The only place left for you is to trust before the agent acts. That's what makes these agents feel dangerous in a way coding agents never did. It's not that they fail often. It's that out there failure is forever. So either you completely go up front or never let it act. Let me be honest. Reversibility is the hardest to replicate in knowledge work. Real undo, the way it exists for code, probably doesn't exist in all the scenarios in knowledge work. But we have some scenarios where undo exists, and we call them. So let's say you add a label. You can remove the label afterwards. But for actions that you can't undo at all, like hard deletes that disappear the emails from your inbox, we again provide a sandbox where the agent can do the thing first in the sandbox, and you can review it, and then actually goes into the production environment. None of it touches the real world. That's the whole flip. In code, you can undo the mistake after it happens. Here, you catch it before it does. Different timing, same result. A mistake that won't stick. Think about you again. The actions we could reverse, we would give it a reverse button. The ones we couldn't, the agent would hit the sandbox first, and she would be notified, your 1,200 emails are going to get deleted. Do you want it? It's not done yet. But across billions of actions that we are going through, we are learning on the way which ones can be walked back, which ones can't, and preparing the sandbox accordingly. If you take one thing away today, take this. For two years, the model was the bottleneck. So everybody was racing towards better and better model. Now the models have gotten good enough, where software engineering is 100% autonomous. But now everything else is the bottleneck. The same model that writes your code can also do your hiring, sales, and other knowledge work. But right now it's working blind. No history, no context, no ways to verify, no guardrails, no undo. So the bottleneck has moved. Now it's infrastructure that nobody has yet built, and that's what we are building at Composio. Yeah. We are powering billion-plus tool calls in total, 300 million tool calls happening every month. And if you are building an agent, just point it to Composio and see the magic happen for knowledge work. And if you want to build the future of substrate of AI agents, then please come to me. We are definitely hiring, and there's a lot, lot left to do. The models will keep getting better. The bottleneck won't be models. It will be the things around it. Thank you. . but enforcing it what it can do. The last pillar, reversibility. And this is where we reach when things go wrong. Can I undo it? In code, you almost always scan. Every change is recorded. Things can be walked back. You can get revert the last commit, or you can get bisect to the commit that broke your production and revert it. Now, like I'm not saying it's good. I won't pretend like that. If things go in production and break, it's always bad. But it's still not permanent. You can still walk back from it. And that's what gives you confidence to let your agents cook and let them do some magic because even if they break the things, you have a pathway back. For knowledge work, there is no undo button. Think about use inbox. Those 200 emails are gone. They have vanished. That's the normal case, by the way. The disaster cases, a sent email, which you can't revert back. A wire that has already been made, so you can't get that money back. A deleted record, gone forever. Most actions actually in knowledge work don't have an undo button. And that changes the whole equation. That changes the blast radius. With code, you can trust the agent after the fact. Let it run, check the result, undo if it's wrong. Out here, there's no coming back. The only place left for you is to trust before the agent acts. That's what makes these agents feel dangerous in a way coding agents never did. It's not that they fail often. It's that out there failure is forever. So either you completely go up front or never let it act. Let me be honest. Reversibility is the hardest to replicate in knowledge work. Real undo, the way it exists for code, probably doesn't exist in all the scenarios in knowledge work. But we have some scenarios where undo exists, and we call them. So let's say you add a label. You can remove the label afterwards. But for actions that you can't undo at all, like hard deletes that disappear the emails from your inbox, we again provide a sandbox where the agent can do the thing first in the sandbox, and you can review it, and then actually goes into the production environment. None of it touches the real world. That's the whole flip. In code, you can undo the mistake after it happens. Here, you catch it before it does. Different timing, same result. A mistake that won't stick. Think about you again. The actions we could reverse, we would give it a reverse button. The ones we couldn't, the agent would hit the sandbox first, and she would be notified, your 1,200 emails are going to get deleted. Do you want it? It's not done yet. But across billions of actions that we are going through, we are learning on the way which ones can be walked back, which ones can't, and preparing the sandbox accordingly. If you take one thing away today, take this. For two years, the model was the bottleneck. So everybody was racing towards better and better model. Now the models have gotten good enough, where software engineering is 100% autonomous. But now everything else is the bottleneck. The same model that writes your code can also do your hiring, sales, and other knowledge work. But right now it's working blind. No history, no context, no ways to verify, no guardrails, no undo. So the bottleneck has moved. Now it's infrastructure that nobody has yet built, and that's what we are building at Composio. Yeah. We are powering billion-plus tool calls in total, 300 million tool calls happening every month. And if you are building an agent, just point it to Composio and see the magic happen for knowledge work. And if you want to build the future of substrate of AI agents, then please come to me. We are definitely hiring, and there's a lot, lot left to do. The models will keep getting better. The bottleneck won't be models. It will be the things around it. Thank you. .