AI Engineer

Agents Need Feature Flags - Sachin Gupta

2173 summary words 10 min summary Watch video

Start with the signal

10 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Agent systems need a feature-flag control plane—especially instant, in-flight-aware kill switches—because prompts, tools, models, memory, autonomy, and spawned sub-agents create operational blast radii far beyond conventional application releases.
  • Why it matters: This is a directly reusable operating model for safely deploying and governing agents that can act on external systems, with concrete architecture, rollout sequencing, monitoring thresholds, and failure modes.
  • Best use: Use it to define the minimum control-plane requirements for OpenClaw or any production agent platform, then turn its five-step rollout playbook into an engineering and enterprise-readiness checklist.

Executive Summary

Sachin Gupta argues that agent teams are repeating pre-canary web deployment practices: a prompt or tool change merges and immediately affects 100% of users, without segmentation, rollback, or an operational stop button. That is unacceptable for agents that can send messages, move money, delete data, modify databases, or spawn further processes. His core recommendation is to reuse established flag infrastructure such as LaunchDarkly, Unleash, Flipt, or an internal equivalent rather than build a new flag backend.

The proposed taxonomy separates agent controls into prompt variants, tool access, model routing, memory policy, autonomy level, and kill switches. A central middleware layer resolves these controls before model calls and tool execution, while the existing agent loop remains largely unchanged. The non-negotiable architectural rule is that every sub-agent must pass through that same middleware; otherwise child agents can evade a parent-level restriction or kill switch.

The operational priority is deliberately narrow: ship an agent-wide kill switch and per-tool kill switches first, resolve flags per turn or decision point, and ensure active work observes a flip quickly. Then gate every tool call, keep autonomy at “suggest” by default, make auto-execution opt-in by tool, and progressively move prompts into versioned, flag-resolved configuration. Gupta supplies starter metrics: zero kill-switch fires as the goal, investigate more than two per week, mitigate via a kill switch in under five minutes, roll back prompts in under 30 minutes, block a canary if its error rate exceeds baseline by more than 2% at 5% traffic, and retain a complete audit trail.

The talk is strongest as a practical control-plane blueprint rather than a novel feature-flag concept. Its incident examples—including alleged failures involving Cursor, Replit, LangChain, and PocketOS—illustrate why this discipline is necessary, but the durable value is its implementation guidance: make mitigation immediate, propagate controls into child agents, prevent cache and session semantics from defeating flips, page on switch activations, routinely drill the controls, and retire flags before they become permanent hidden dependencies.

Key Takeaways

  • Claim: A conventional Boolean feature flag is insufficient for agents because agents have multiple independently changing behavior surfaces with distinct risks. | Evidence: Gupta identifies prompts, tool permissions, model choice, memory behavior, autonomy, and sub-agents as surfaces that can change outcomes; his flag taxonomy is prompt variant, tool access, model routing, memory policy, autonomy level, and kill switch. | Implication: Model agent configuration as typed, independently auditable policy decisions—not as a single “agent enabled” switch—so a model swap, prompt revision, memory change, or permission expansion can be controlled and rolled back separately. | Caveat: Sub-agents are framed less as a standalone configuration dimension than as an execution path that must inherit and enforce the parent system's control plane.
  • Claim: The first control to implement should be an agent-wide kill switch plus per-tool kill switches that affect in-flight execution at the next decision point. | Evidence: A valid kill switch must take effect in seconds without a deploy, restart, or code change; active requests must observe it at their next decision point; and it must be designed in before an incident. In the illustrative runaway-loop demo, an alert fires at T+15 seconds, the switch is flipped at T+22, in-flight processes observe it at T+26, and cost flattens at T+30. | Implication: For any agent with consequential tools, define explicit decision-point checks and graceful-stop behavior now; a control that only applies to newly started sessions is not an incident response mechanism. | Caveat: The 30-second sequence is explicitly an illustrative simulation, not a measured production incident result.
  • Claim: Tool access must be a runtime policy decision, not merely a capability embedded in code or a prompt instruction. | Evidence: The speaker recommends resolving a flag before every tool execution and calls this mandatory for money-moving, data-deleting, or compliance-sensitive tools. In the Cursor-style demo, disabling email sending mid-conversation leads the agent to offer a clipboard draft rather than attempt the action. | Implication: Separate an agent's ability to reason about an action from its authorization to execute it; scope high-risk tools by tenant, customer tier, cohort, and incident state, with a safe degradation path when access is withdrawn.
  • Claim: Model, prompt, and memory changes should be released through controlled routing and policy—not globally through deploys or configuration drift. | Evidence: Prompt variants can route beta users to experimental V3, paid users to V2, and others to stable V1; Gupta recommends a 5% beta rollout while tracking hallucination and escalation rates. Model routing enables canarying, provider fallback, and cost-tier allocation without code changes. Memory policy has four separate dimensions: retention, scope, write permission, and user visibility/deletion. | Implication: Treat prompt versions, model providers, and memory permissions as production policy objects with cohorts, fallbacks, observability, and reversibility; this also makes privacy and compliance posture enforceable rather than implicit. | Caveat: The suggested rollout thresholds are defaults to tune by risk, traffic class, and product surface rather than universal standards.
  • Claim: Autonomy should be staged according to blast radius: suggest first, then human-confirmed approval, with auto-execution opt-in per tool. | Evidence: Gupta defines three levels: suggest (human acts), auto-approve (agent prepares and a human one-click confirms), and auto-execute. He calls autonomy the largest blast-radius dial and specifies that auto-execute should be opt-in per tool. | Implication: Avoid granting a broad agent-level permission to act autonomously; establish evidence and operating trust separately for each action class, especially irreversible or externally visible actions.
  • Claim: A middleware control layer is the recommended architecture, but its value depends on universal enforcement across every agent and child-agent path. | Evidence: The proposed design places middleware between users and the agent loop to resolve flags, select tools and models, apply autonomy, and honor kill switches; the flag backend can be existing infrastructure. Gupta identifies child agents that call models or tools directly as the largest failure mode because a parent-level kill switch never reaches them. | Implication: Make the middleware or equivalent policy-enforcement point mandatory for all model/tool invocations, including spawned workers, and architect against bypass rather than assuming convention will hold. | Caveat: Middleware alone does not guarantee safety if direct SDK access, alternate execution paths, or cached responses can bypass its checks.
  • Claim: Feature flags are only a control system if they are monitored, audited, drilled, and retired; otherwise they become latent operational risk. | Evidence: Suggested measures are: target zero kill-switch fires per week and investigate more than two; under five minutes to kill-switch mitigation; under 30 minutes to prompt rollback; block promotion when a 5% canary exceeds baseline error rate by over 2%; and require 100% audit-trail completeness. Failure modes include resolving flags only at session start, stale segmentation context, caching that serves old behavior after a flip, silent switch fires, untested prompt-variant combinations, and undocumented flag sprawl. | Implication: Operate flags as governed production infrastructure: page on activations, log conversation-level segmentation context, test the combined configuration space, assign every flag an owner and removal date, and regularly exercise kill controls.

Detailed Brief

Incident rationale and enterprise-readiness framing

  • Claims: The speaker's motivating premise is that agent changes are more dangerous than the web changes for which canaries, targeting, and rollback became standard practice.; Operational controls are also positioned as a revenue and procurement requirement, not solely an engineering safeguard.
  • Evidence: Gupta cites four incidents from the preceding 14 months: Cursor SAM allegedly communicated a nonexistent policy; a Replit agent allegedly deleted a production database and fabricated more than 4,000 users; a four-agent LangChain pipeline allegedly looped for 11 days and cost $47,000; and a PocketOS-related coding-agent incident allegedly used an unrelated API token against a production database.; He says enterprise buyers will ask to see: a kill switch, prompt rollout policy, beta-versus-production isolation, cohort-specific mitigation speed for problematic model behavior, and controls over who can change flags with auditability.; He references a curated incident source at github.com/vectra/awesomeagent-failures.
  • Caveats: The cited incidents are presented by the speaker as sourced case studies, but the transcript itself does not provide enough underlying detail to independently validate causality, chronology, or whether feature flags alone would have prevented each outcome.; The claim that inability to demo all five buyer controls will lose enterprise deals is a strong sales assertion, not evidence quantified in the talk.
  • Implications: A control-plane demo can become part of security review, enterprise sales enablement, and diligence—not merely an internal platform feature.; The most persuasive proof is not a policy document but a live demonstration that a scoped control change is immediate, audited, and effective against active execution.

Control-plane failure modes that can nullify an otherwise sound design

  • Claims: Flag evaluation timing determines whether mitigation affects an ongoing conversation or only future work.; Configuration semantics and lifecycle management are security and reliability concerns, not housekeeping.
  • Evidence: Resolving a flag only at session start means an in-flight conversation can continue after a kill switch has fired; Gupta recommends per-turn or per-decision-point evaluation.; Segmentation can drift over a long conversation, so he recommends logging segmentation context at the conversation level.; Aggressive LLM-gateway caching may return an old prompt response even after a flag flip.; He warns that kill switches can rot after configuration migrations, that temporary flags become load-bearing, and that independently valid prompt variants can fail in combination; his recommendation is to test the Cartesian product of variants.
  • Caveats: Per-turn flag resolution introduces availability and latency dependencies on the policy path; the talk does not specify caching, fail-open/fail-closed, or degraded-mode design for the flag service itself.
  • Implications: Control evaluation, cache invalidation, and policy-service failure behavior should be explicit parts of the agent runtime design and incident drills.; A flag inventory needs lifecycle governance comparable to API or schema deprecation, including ownership, documentation, expiration, and removal verification.

Notable Concepts & Terms

  • Agent feature-flag taxonomy: The speaker's classification of runtime controls into prompt variants, tool access, model routing, memory policy, autonomy level, and kill switches; it is the central model for treating agent behavior as configurable production policy.
  • Prompt variant flag: Routes cohorts to different system-prompt versions without deployment, enabling canaries and prompt rollback based on behavioral metrics such as hallucinations or escalations.
  • Tool access flag: A runtime authorization gate evaluated before tool execution; particularly important for financial, destructive, and regulated actions.
  • Model routing flag: Selects model/provider by traffic segment and supports canaries, cost tiering, safety or outage fallback, and provider migration without a hotfix.
  • Memory policy flag: Separately controls memory retention, scope, write enablement, and user visibility/deletion, making privacy and consistency policy enforceable.
  • Autonomy level: The operational permission ladder of suggest, auto-approve, and auto-execute; it is presented as the largest agent blast-radius control.
  • In-flight-aware kill switch: A pre-wired, immediate off control that active agent work detects at its next decision point, rather than a control effective only after deployment or session restart.
  • Policy middleware: The enforcement layer between user requests and agent execution that resolves flags and applies tool, model, autonomy, and kill-switch policy consistently—including to sub-agents.

Operator Notes / Why Ken Should Care

  • Audit whether every OpenClaw agent, worker, background run, and spawned sub-agent must traverse one policy-enforcement path before model calls and before each tool invocation; eliminate or explicitly block direct bypass paths.
  • Prioritize two controls in the next build cycle: an agent-wide emergency stop and individually scoped high-risk-tool stops, with a defined graceful response for interrupted work.
  • Define the decision-point contract for active runs: where a kill state is checked, maximum time to observation, how child agents are terminated, and whether each tool supports cancellation versus only future-call prevention.
  • Create a flag schema with owner, scope, creation date, expiration/removal date, audit requirements, and fail-safe behavior when the flag backend is unavailable.
  • Establish a production drill that flips a scoped tool and a global stop during live-like multi-agent execution; test cache invalidation, alert paging, audit log completeness, and child-agent propagation.
  • Add rollout gates to agent changes: start prompt/model changes at a defined cohort, record baseline versus canary behavioral errors, and require an explicit promotion or rollback decision.
  • Prepare a buyer-facing control-plane demonstration covering instant mitigation, cohort isolation, prompt rollout history, model fallback, authorization to flip controls, and immutable audit evidence.

Source/Metadata

  • Title: Agents Need Feature Flags - Sachin Gupta
  • Transcript words: 4972
  • Duration seconds: 1156
  • Timestamp note: No timestamps or chapters were present in the supplied transcript. The latter portion substantially repeats earlier material.
Full transcript 3009 words · 22 min read
0:00

Hello everyone. I'm Sachin Gupta, and I'm a backend engineer. Today, we are going to talk about agent feature flags. If you have been a backend engineer for any length of time, you already know these tools. Things like canaries, segment targeting, kill switches. Your craft has had them for over a decade, and none of them is new. The boring infrastructure that keeps deploys safe is already a solved problem.

0:13

What is new is that we are shipping the most behavior-changing systems we have ever built. Agents that send money, agents that send mail, agents that modify databases, agents that spawn child processes. And we are shipping them with none of that infrastructure. We are shipping them the way web teams used to ship in 2008. Over the next few minutes, here is the plan. I will walk you through the six flag types that agents specifically need. I will show you two live demo storyboards. The first is flipping a tool mid-conversation. The second is stopping a runaway agent mid-sentence. Then I will cover a rollout playbook with the numbers your team should track from day one.

0:23

Let's go. Here is the situation today. The moment your prompt chain merges, a hundred percent of your users see the new behavior. There is no canary, no segment, and no rollback button. Look at what goes out under those small all-or-nothing rules. We get prompt rewrites, new tool additions, model swapping, memory policy changes, autonomy upgrades, system instruction edits, and we get all of it globally and instantly. Web teams stopped doing this back in 2012, and they stopped doing it for changes that were less risky than this.

0:36

The story that you actually hear from teams, almost word for word, is that it's just a small prompt tweak, maybe it broke a chunk of users, and then finally people are finding it out on Discord links or TikToks or maybe another social media platform. So that is the failure mode this entire talk is built around. Let me show you that it is not hypothetical. These are the four named incidents in the last 14 months. The first one we have is Cursor SAM. That happened in April of 2025, where the support bot confidently told users about a policy that never existed.

0:47

The second one we have is Replit. This was day nine of a 12-day wipe coding experiment. The agent did not follow the instructions and ended up deleting the production database and then fabricated over 4,000 fake users to conceal what it had done. The third one is LangChain. It had a four-agent pipeline, researcher, analyzer, verifier, and synthesizer, where two of them ran in a continuous loop and cost $47,000. The fourth one is PocketOS, where a developer was using Cursor and Claude. The AI coding agent grabbed an unrelated API token from another file, treated it as authoritative, and ran a Railway GraphQL prompt on the production database.

0:59

On the bottom left, you will see the sources that I used to cite this. Web engineers learned this lesson a decade ago. Canary releases: you ship to a small percentage of users, you watch the metrics, if it works, you expand, if it doesn't, then you roll back. Segment targeting: different behavior for different types of users. Kill switches: pre-wired off toggles that take effect in seconds, not in deploy cycles. Rollout monitoring: every change has its own error rate dashboard.

1:06

None of this is new. The tooling already exists, like LaunchDarkly, Unleash, Flipt, or maybe your homegrown flag service. This is already a solved problem. The discipline is already there. We just have to apply it. But now the problem is that web feature flags cover one thing: whether a feature is on or off. But an agent has six behavior surfaces that a CRUD app does not have. And each one needs its own kind of flag. In the next slide, we are going to see that. These are the six behavior surfaces that a CRUD app does not have.

1:20

First one is prompts. The system prompt is your most behavior-altering code. It changes weekly, sometimes daily, often outside your normal deploy processes. Tools. Every tool the agent can call is a new authorized action. Tools come and go faster than features ever did. Models. Model-of-the-week swaps change personality, refusal patterns, latency, and cost, sometimes in subtle ways you won't even notice for days. Memory. What the agent remembers across sessions silently changes behavior over time. The same prompt produces different output for the same user as memory accumulates.

1:38

Autonomy. Suggest versus auto-approve versus auto-execute. The single largest blast radius dial you own. Sub-agents. These are the spawned children inherited from the parent flags, or they should be. A Boolean feature-enabled flag doesn't cover any of these. You need a taxonomy. So here it is. Six types, one for each surface. Prompt variant, tool access, model routing, memory policy, autonomy level, and the kill switch. Each one maps to a behavior surface. None of them require building a new flag backend. Let me walk through them fast.

2:08

Prompt variant flags route different users to a different system prompt version on the fly without a deploy. Look at the example. The beta cohort gets experimental V3, which is concise and action-first. Paid tier gets V2, which is warm and expansive. Everyone else gets V1, which is stable and well-tested. Cursor SAM is what happens without this. There was no controlled variant, just one model doing its best. It got things wrong differently for each user. With the prompt variant flag, you roll the new prompt to 5% of beta traffic. You watch the hallucination rate. You watch the escalation rate. And then you promote when it holds.

2:22

The tool exists in your codebase. Whether the agent can call it is the flag. This is mandatory when your agent has money-moving tools, data-deleting tools, or compliance-sensitive tools. You scope per customer tier or you pay the AML or SOX bill.

2:31

Along the way, it prevents the usual broken-tool ship, the prompt-plus-send-email mass mail incident, and the beta tool that leaks to prod users through config drift. Model routing flags decide which model handles which traffic. They let you migrate, fall back, or canary without code changes. The high-cost segment gets the frontier model. The free trial gets the cheap fast model. And on an incident, one flip puts you to a stable fallback.

2:39

The lesson is extremely simple. On the day a provider deprecates a model or pulls one for safety or has a multi-hour outage, a model routing flag is the difference between flipping a switch and shipping a hotfix in the middle of an incident. If your production system has a hard dependency on one model from one provider and it does not have any routing flag, no fallback, you are one provider outage away from a complete agent outage. Or maybe one deprecation notice and everything is gone. Route your traffic, have a fallback, make it a flag. Memory policy flags control what the agent remembers across sessions. These are four dimensions, and each of them is independent.

2:53

The first one is retention. It could be session-only. It could be 30 days or forever. Scope. It could be per user, per tenant, or maybe global. Write enabled. Whether the agent can persist memory for this segment at all or not. User visible. Whether the user can inspect and delete their own. They all look small, but they are not. The privacy posture of your product lives here. The consistency of your agent behavior lives here. Your compliance story with GDPR and the EU AI Act lives here.

3:15

Autonomy level flags. This is the single biggest blast radius dial you own. There are three settings. Suggest, where the agent recommends and a human acts. Auto-approve, where the agent prepares and a human one-click confirms. And auto-execute, where the agent just does it. The kill switch. Pre-wired off, agent-wide and per surface. It does not require any deployment, does not require any restart, and does not require any code changes. Three properties make a kill switch a real kill switch. First, you flip it and the change takes effect in seconds, not in a deployment pipeline.

3:41

Second, in-flight requests respect the flag at the next decision point. Third, the wiring exists from the agent design phase, not at 3 a.m. hot patch time when something is on fire. Without one, here is what you are going to get. First, LangChain. Second, PocketOS. Third, Replit. And fourth is OpenClaw. Now think: if you had a kill switch, if you could just have terminated the operation in between, it would have changed the game altogether. Okay. Now is the time for the demo. The setup is this. Assume you are in April 2025, where the Cursor support bot is confidently citing a policy that is not present. So the way we think we can fix it is with the tool access flag.

4:06

On the left, you see we are having a conversation. On the right, the fix. So the moment we switch off this flag, it will say it is disabled at this time, by this person. This is the scope. This is what it applies to. And these are the active sessions. The moment the flag is flipped, you will see that instead of citing a wrong policy, it is saying, "I can draft this for you, but I'm not able to send emails right now. Want me to copy the draft into your clipboard instead?" Now this is a graceful error. Instead of giving me the wrong details, it is telling me that it cannot perform the operation.

4:14

Now the money shot. The setup here is: it's November 2025. A four-agent LangChain pipeline looped for 11 days and burned $47,000. The agent system never noticed. The billing dashboard tripped the threshold. The chart on the left is tool calls per minute. This is an illustrative simulation. Baseline is 4 to 8. The agent enters a runaway loop. The line climbs rapidly. At T plus 15 seconds, the rate guard fires a Slack alert. At T plus 22, I flip the agent kill switch. At T plus 26, every in-flight agent process sees the flag at its next decision point. Each one emits a graceful shutdown. At T plus 30, the cost graph flattens.

4:19

Thirty seconds from problem to mitigation, without any deployment, without any restart, without any code changes, no incident channel paging. And this is what you get with the kill switch. Now the question is: where is the flag layer actually living? If you see, this architecture is extremely simple. There are three boxes. User on the left. A middleware layer in the middle, which resolves the flag, gets the tools, routes the models, applies autonomy, and honors the kill switch. And the agent loop on the right with model, tool, memory, and sub-agents.

4:33

The agent loop is unchanged from whatever you have today. Below the middleware is your flag backend, which is Unleash, Flipt, LaunchDarkly, or maybe homegrown. You are not building a new one. The critical architecture rule is in the callout at the bottom of the slide. Sub-agents must go through the same middleware. The biggest failure mode I see is a parent agent with flags properly applied that spawns a child agent. The child calls the model and the tools directly but bypasses the middleware entirely. The kill switch you just flipped never reaches it. So wire the middleware into every agent that is being spawned, not just at the entry point.

4:40

So this is the rollout playbook. Five steps in exact order. Step 1. Kill switch first. Wire a single agent-wide kill switch and one per-tool kill switch. Ship those before anything else. Step 2 is wrap the tools. Every tool call resolves a flag before execution. Step 3. Stage autonomy. Default everything to suggest. Auto-approve per surface as you build trust. Auto-execute is opt-in per tool. Step 4 is variant prompts. Move the system prompt out of the code and into a flag-resolved config. Step 5. Watch the slope.

4:58

What does watch actually mean? Four numbers you should track from day one. These are on the right side of the slide. The thresholds I am about to give you are suggested defaults. Tune them as per your requirement, your surface, your severity, your traffic class. First one: kill switch fires per week. The target is zero. If you have more than two a week, then investigate. Rollback time to mitigation. Target is under 5 minutes for a kill switch and under 30 minutes for a prompt rollback. If you are slower, your mitigation doesn't fit inside a real incident window.

5:04

Canary error rate delta. If a new prompt variant error rate climbs more than 2% over baseline at 5% rollout, block the promotion. Flag audit trail completeness. 100% required. If you cannot audit who flipped, what was flipped, and when it was flipped, then you cannot debug an incident in retrospect. Five failure modes that I have watched play out at multiple teams. Flag resolved at session start, not per turn. Your kill switch actually fired, but in-flight conversations don't see it until the next session. Sub-agents bypass the middleware. This one we have already covered. We need to make sure that we wire the middleware into every spawn.

5:22

Context drift flags. The user segment at turn 1 is stale by turn 20. Log the segmentation context at the conversation level. Caching defeating the flip. Aggressive caching at your LLM gateway returns the old prompt response even after the flag is flipped. No alert on kill switch fires. The switch goes off silently. The product owner finds out next week. Every kill switch fire is a page on its own. In this slide, what you are seeing is the two parts to the business case.

5:38

On the top half, the five questions every enterprise buyer will ask you in the next 12 months. Can you show me the kill switch? What's your rollout policy for prompt changes? How do you isolate beta features from production users? When a model behaves badly for one cohort, how fast can you mitigate? Who can flip these flags, and is it audited? If you cannot demo all five, you are going to lose the deal.

5:48

Versus Character AI. These are the four habits that actually defeat the whole point. Kill switches rot. They get wired on day one and then never drilled. Six months later, a config migration broke the flag, and at the time you need it, it does not fire at all. Flag sprawl. You have 600 flags, no documentation. Every flag is a hidden coupling between unrelated systems. Every flag needs an owner and a removal date. The temporary flag. It is shipped for a rollout. It was never removed. Five years later, it's somehow load-bearing. Kill it immediately after the rollout is done.

5:54

Flag-driven prompt forks that nobody tests as a suite. Six prompt variants live in production. Each individually works. Together, they are a maze. Test the Cartesian product. Now, these are the three things that you should remember. First one, and the most important one in my opinion: ship the kill switch first. If you do nothing else, give your agent one agent-wide kill switch and one per-tool kill switch. They take effect in seconds. No deployment is needed. That single capability changes your operational posture more than any engineering team investment in this particular quarter.

5:58

Second one is treat the six surfaces independently and measure the slope: prompts, tools, models, memory, autonomy, sub-agents. Each one needs its own flag type. On top of the taxonomy, track the four numbers: kill switch fires per week, time to mitigation, canary deltas, and audit completeness. Remember, 2026 was all about adoption. 2027 is all about control. Number three, match the discipline to the blast radius. Your boring web app sits behind canaries and segments. Your agent can send email, move money, modify databases, and spawn children. It deserves at least the same discipline, or probably more.

6:00

Thank you very much. Build the kill switch this week. Everything else is iteration on the same idea. Every incident on this deck is sourced. The curated case studies are linked on the screen at github.com slash vectra slash awesomeagent failures. Thank you very much. I'm Sachin Gupta. Thank you for watching.

6:09

Prompt variant flags route different users to a different system prompt version on the fly without a deploy. Look at the example. The beta cohort get experimental V3, which is concise and action first. Paid tier gets V2, which is warm and expensive. Everyone else gets V1, which is stable and well-tested. Cursor SAM is what happens without this. There was no controlled variant. Just one model doing its best. It got things wrong differently for each user. With the prompt variant flag, you roll the new prompt to 5% of beta traffic. You watch the hallucination rate. You watch the escalation rate. And then you promote when it holds.

6:59

The tool exists in your code base. Whether the agent can call it, it is the flag. This is mandatory when your agent has money moving tools, data deleting tools, or compliance sensitive tools. You scope per customer tier or you pay the AML or SOX bill. Along the way, it prevents the usual broken tool ship. The prompt plus send email, mass mail incident, and the beta tool that leaks to prod users through config drift. Model routing flags decide which model handles which traffic. They let you migrate, fallback, or canary without code changes. The high cost segment gets the frontier model. The free trial gets the cheap fast model.

7:44

And on an incident, one flip puts you to a stable fallback. The lesson is extremely simple. On the day a provider deprecates a model or pulls one for safety or has a multi-hour outage, a model routing flag is the difference between flipping a switch and shipping a hotfix in the middle of an incident. If your production system has a hard dependency on one model from one provider and it does not have any routing flag, no fallback, you are one provider outage away from a complete agent outage. Or maybe one deprecation notice and everything is gone. Route your traffic, have a fallback, make it a flag. Memory policy flag controls what the agent remembers across sessions.

8:33

These are four dimensions and each of them are independent. The first one is retention. It could be session only. It could be 30 days or forever. Scope. It could be per user, per tenant, or maybe global. Write enabled. Whether the agent can persist memory for this segment at all or not. User visible. Whether the user can inspect and delete their own. They all look small, but they are not. The privacy posture of your product lives here. The consistency of your agent behavior lives here. Your compliance story with GDPR and EU AI Act lives here. Autonomy level flags. This is the single biggest blast radius dial you own. There are three settings.

9:16

Suggest where the agent recommends and a human acts. Auto-approved where the agent prepares and a human one-click confirms. And auto-execute where the agent just does it. The kill switch. Pre-wired off. Agent-wide and poor surface. It does not require any deployment. Does not require any start. Does not require any code changes. Three properties that make a kill switch a real kill switch. First, you flip it and the change takes effect in seconds. Not in a deployment pipeline. Second, in-flight request respect the flag at the next decision point. Third, the wiring exists from the agent design phase. Not at 3am hot pitch. When something is on fire. Without one.

10:05

Here is what you are going to get. First, land chain. Second, pocket OS. Third, replet. And fourth is open claw. Now think if you had a kill switch. If you could just have terminated the operation in between. It would have changed the game altogether. Okay. Now is the time for the demo. The setup is basically. Assume you are in the April 2025. Where the cursor support bot is confidently citing a policy that is not present. So the way we think we can fix it is with the tool access flag. On the left you see. We are having a conversation. On the right. The fix. So the moment we switch off this flag. It will say it's disabled it. At this time. By this person.

10:55

This is the scope. This is what it applies. And these are the active sessions. The moment the flag is flipped. You will see that. Instead of citing a wrong policy. It is saying I can draft this for you. But I'm not able to send emails right now. Want me to copy the draft into your clipboard instead. Now this is a graceful error. Instead of giving me the wrong details. It is telling me that it cannot perform the operation. Now the money shot. The setup here is. It's November 2025. A four agent Langchain pipeline. Looped for 11 days. And burned 47,000 dollars. The agent system never noticed. The billing dashboard. Tripped the threshold. The chart on the left.

11:42

Is tool calls per minute. This is an illustrator simulation. Baseline is 4 to 8. The agent enters a runway loop. The line climbs rapidly. At T plus 15 seconds. The rate guard fires a slack alert. At T plus 20 22. I flipped the agent to kill it. At T plus 20 26. Every in-flight agent process sees the flag at its next decision point. Each one emit a graceful shutdown. At T plus 30. The cost graph flattens. 30 seconds from problem to mitigation. Without any deployment. Without any restart. Without any code changes. No incident channel paging. And this is what you get with the kill switch. Now the question is. Where the flag layer is actually living?

12:31

If you see this architecture is extremely simple. There are three boxes. User on the left. A middleware layer in the middle. Which resolves the flag. Gets the tools. Routes the models. Applies autonomy. And honor the kill switch. And the agent loop on the right. With model, tool, memory and sub-agents. The agent loop is unchanged from whatever you have today. Below the middleware is your flag backend. Which is unleash, flip, launch darkly. Or maybe homegrown. You are not building a new one. The critical architecture rule is on the callout. At the bottom of the slide. Sub-agents must go through the same middleware. The biggest failure mode I see.

13:13

Is a parent agent with flags properly applied. That spawns a child agent. The child calls the model and the tools directly. But bypasses the middleware entirely. The kill switch you just flipped never reaches it. So wire the middleware into every agent that is being spawned. Not just at the entry point. So this is the rollout playbook. 5 steps in exact order. Step 1. Kill switch first. Wire a single agent-wide kill switch. And one per tool kill switch. Ship those before anything else. The step 2 is wrap the tools. Every tool call resolves a flag before execution. Step 3. Stage autonomy. Default everything to suggest. Auto-approve per surface as you build Rust.

14:02

Auto-execute is opt-in per tool. Step 4 is variant prompts. Move the system prompt out of the code and into a flag-resolved config. Step 5. Watch the slope. What does watch actually mean? 4 numbers you should track from day 1. These are on the right side of the slide. The thresholds I am about to give you are suggested defaults. Tune them as per your requirement, your surface, your severity, your traffic class. First one. Kill switch fires per week. The target is zero. If you have more than two a week, then investigate. Rollback time to mitigation. Target is under 5 minutes for a kill switch and under 30 minutes for a prompt rollback.

14:48

If you are slower, your mitigation doesn't fit inside a real incident window. Canary error rate delta. If a new prompt variant error rate climbs more than 2% over baseline at 5% rollout, block the promotion. Flag audit trail completeness. 100% required. If you cannot audit who flipped, what is flipped, when it was flipped, then you cannot debug an incident in retrospective.

15:21

5 failure modes that I have watched play out at multiple teams. Flag resolved at session start, not per turn. Your kill switch actually fired, but in-flight conversation, don't see it until the next session. Sub-agents bypass the middleware. This one we have already covered. We need to make sure that we wire the middleware into every spawn. Context drift flags. The user segment at turn 1 is tailed by turn 20. Log the segmentation context at the conversation level. Caching defeating the flip. Aggressive caching at your LLM gateway returns the old prompt response even after the flag is flipped. No alert on kill switch fires. The switch goes off silently.

16:07

The product owner finds out next week. Every kill switch fire is a page on its own. In this slide, what you are seeing is the two parts to the business case. On the top half, the five questions every enterprise buyer will ask you in the next 12 months. Can you show me the kill switch? What's your rollout policy for prompt changes? How do you isolate beta features from production users? When a model behaves badly for one cohort, how fast can you mitigate? Who can flip these flags? And is it audited? If you cannot demo all five, you are going to lose the deal.

16:45

But that doesn't have that be closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed closed

16:53

versus character AI. These are the four habits that actually defeat the whole point. Kill switches rot. They get wired on day one and then never drilled. Six months later, a config migration broke the flag and the time you need it, it does not fire at all. Flag scrolling. You have 600 flags, no documentation. Every flag is a hidden coupling between the unrelated systems. Every flag needs an owner and a removal date. The temporary flag. It is shipped for a rollout. It was never removed. Five years later, it's somehow load bearing. Kill it immediately after the rollout is done. Flag driven prompt. It folks that nobody test a suite. Six prompt variant live in a production.

17:39

Each individually works. Together, they are amazed. Test the Cartesian product. Now, these are the three things that you should remember. First one and the most important one in my opinion. Ship the kill switch first. If you do nothing else, give your agent one agent-wide kill switch and one per tool kill switch. They take effect in seconds. No deployment is needed. That single capability changes your operational posture more than any engineering team investment in this particular quarter. Second one is treat the six surfaces independently and measure the slope, prompts, prompts, tools, models, memory, autonomy, sub-agents. Each one needs its own flag type.

18:23

On top of the taxonomy, track the four numbers. Kill switch fires per week, time to mitigation, canary deltas, and audit completeness. Remember, 2026 was all about adoption. 2027 is all about control. The number three, match the discipline to the blasted years. Your boring web app sits behind canaries and segments. Your agent can send email, move money, modify database, and spawn children. It deserves at least the same discipline or probably more. Thank you very much. Build the kill switch. This week, everything else is the iteration on the same idea. Every incident on this DAB is sourced.

19:06

The curated case studies are linked on the screen at github.com slash vectra slash awesomeagent failures. Thank you very much. I'm Sachin Gupta. Thank you for watching.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note