Guardrails First: Engineering Member-Facing Health AI — Rashi Agrawal, Hinge Health
Description
A healthy 60 year old man asked a popular AI assistant how to cut salt from his diet. It pointed him at sodium bromide. Three months later he arrived in an emergency room with paranoia and hallucinations, bromide at 200 times the safe level, and stayed three weeks. Rashi Agrawal stacks that against the first independent safety test of a consumer health AI, out of Mount Sinai, which under triaged life threatening emergencies half the time, and against ECRI naming chatbot misuse the top health technology hazard of 2026. Roughly 40 million people already triage themselves this way. None of it is a frontier problem. It is the production baseline. Her argument is that most healthcare AI safety failures are architectural decisions made before a single token is generated. PHI is stripped at the pipeline boundary on ingestion, so a developer who opens a dashboard finds nothing to redact because it was never stored. Anything that can never be wrong lives in a code layer above the model rather than in its prompt: routing to 911 or 988, deciding which capability owns a turn, verifying who is on the other end. The frontier labs publish an authority hierarchy in which every layer above the user sits one prompt injection from being overridden, and her reading is blunt: if they will not treat a prompt as a security boundary, neither should you. Safety then runs as a continuous layer of judges scoring live traffic, with one discipline attached. When a score drops, first ask whether the judge is right. Speaker info: - https://www.linkedin.com/in/rashi283/ - https://sessionize.com/rashiagrawal/ Timestamps: 0:00 - The state of healthcare AI, and 40 million self triagers 1:04 - Poisoned by a chatbot 1:30 - Under triaging emergencies half the time 2:35 - Three non negotiable foundations 3:41 - Where PHI actually lives 5:53 - Deterministic rules belong above the model 7:27 - If the labs will not trust the prompt, neither should you 7:54 - Escalation, intent routing, identity 9:39 - Sa
Summary
Generated by gpt-5.6-terraAt-a-Glance
- Verdict: Watch fully
- Core thesis: Member-facing health AI must be designed around architectural security boundaries, deterministic controls for irreversible decisions, and continuous production evaluation—not prompt-based guardrails alone.
- Why it matters: The talk offers a reusable control-plane pattern for any high-stakes agent system: move must-not-fail decisions into code, continuously evaluate live traffic, and create an explicit human decision process when automated controls are insufficient.
- Best use: Use it to pressure-test OpenClaw or other agent architectures for authority boundaries, escalation routing, identity checks, evaluation operations, and launch-governance rules.
Executive Summary
Rashi Agrawal argues that healthcare AI failures are primarily system-design failures rather than isolated model failures. Her evidence is the production risk profile of consumer health chatbots: a reported case of a man hospitalized after following erroneous sodium-bromide diet advice, and a Mount Sinai safety test in which a consumer health AI reportedly under-triaged life-threatening emergencies 50% of the time. Her response is not better prompting, but a three-layer guardrail architecture.
First, protect PHI structurally at ingestion rather than relying on downstream redaction and policy compliance. Second, put deterministic code above the model for decisions that cannot be probabilistic: emergency escalation, high-stakes intent routing, and identity verification. Third, treat production monitoring as a permanent operating layer, using many automated judges, direct user feedback, and human review of sampled traces—especially complete sampling of high-stakes cases.
The second half is a launch-decision framework for situations where clinical, legal, compliance, product, and engineering stakeholders see different risks. It prioritizes worst-case harm over average incidence, separates severity from engineering capacity, defaults toward the safer error for safety defects, grounds launch criteria in the organization's actual revealed tolerance, and treats deferred launch work as committed debt rather than optional backlog.
The most useful operational nuance is that evaluators themselves are probabilistic systems. A quality-score drop should not automatically trigger an agent or prompt change; teams must first establish whether the judge is correct. The resulting operating model is a feedback loop in which each newly observed failure can produce a new judge, while humans remain the constrained resource for interpreting signals and making consequential calls.
Key Takeaways
- Claim: Safety-critical AI failures should be treated as architecture failures first, because the system must prevent certain classes of error before model generation occurs. | Evidence: Agrawal cites a reported case in which a 60-year-old man followed chatbot advice to replace salt with sodium bromide for three months, resulting in bromide levels 200 times the safe limit and a three-week hospitalization; she also cites a Mount Sinai test reporting 50% under-triage of life-threatening cases such as diabetic ketoacidosis and respiratory failure. | Implication: For high-consequence agent workflows, Ken should assess failure prevention at routing, data, authentication, and tool-permission boundaries—not accept a strong model or system prompt as the core safety argument. | Caveat: The examples establish the stakes but do not provide comparative data showing which specific architectures outperform others across healthcare products.
- Claim: PHI and other sensitive data protections need to be enforced at the pipeline boundary, with environment and regional access isolation designed into the system. | Evidence: The proposed pattern strips PHI at ingestion before it reaches the data lake, separates production from non-production with no connecting pipes, and restricts raw-PHI access by both role and permitted geographic region. | Implication: Design systems so dashboards, development environments, and downstream analytics are incapable of exposing unnecessary raw sensitive data; do not make redaction in logs the principal control. | Caveat: This is a healthcare-specific compliance design, but the general principle applies to any regulated or sensitive-data environment.
- Claim: Must-not-fail decisions belong in a deterministic code layer above the model, because prompts are not reliable security or authority boundaries. | Evidence: Agrawal's stack runs code before the model on every turn. That code handles emergency routing to 911 or 988 for self-harm or acute emergencies, high-stakes intent routing, and identity verification before access to member data. She notes that even model providers publish prompt authority hierarchies that remain vulnerable to prompt injection. | Implication: For agent systems, explicitly enumerate irreversible actions—money movement, privileged access, external communications, safety escalation, data disclosure—and gate them with deterministic policy, authorization, and workflow code. | Caveat: The model can still classify or handle the long tail of ordinary interactions; the recommendation is not to eliminate model judgment, but to deny it final authority over irreversible calls.
- Claim: Production safety requires continuous evaluation of live traffic, not a one-time pre-launch benchmark or golden test set. | Evidence: The recommended signal stack is many continuously refreshed automated judges across clinical accuracy, safety, escalation, relevance, drift, and refusal; per-message member feedback; and random trace review, with 100% sampling of high-stakes cases. | Implication: Ken should plan evaluation as an operational function with telemetry, trace retention, review queues, escalation thresholds, and staffing—not as a release-stage QA activity. | Caveat: Automated judges catch regressions at scale but miss tone and context issues; human reviewers must interpret the combined signals and act on them.
- Claim: In high-stakes launch decisions, severity should be set by the worst plausible outcome rather than average frequency, and it must remain independent of delivery capacity. | Evidence: Agrawal contrasts a defect that mildly annoys 100% of users with one that could seriously harm 0.1%; the latter is more severe. Once identified, the legitimate choices are to fix it, delay launch, or explicitly accept the risk with sign-off—not silently downgrade it because the team cannot fix it on schedule. | Implication: Adopt a launch rubric that separates harm severity from ownership, schedule pressure, and implementation effort, and require named acceptance for material residual risk. | Caveat: Her framework is directional rather than mechanically decisive; organizational leaders still must make the final tradeoff.
- Claim: When evidence is uncertain, teams should choose asymmetric defaults: hold and fix possible safety defects, but usually ship minor polish flaws. | Evidence: The rationale is asymmetric cost: shipping a real safety bug is worse than delaying for a false alarm, while delaying for a low-impact polish issue can cost more than releasing a small imperfection. She also argues that an organization's real launch threshold is its 'revealed risk tolerance'—what it already permits in production—not its aspirational policy. | Implication: Define defect classes and default actions in advance, then compare new launch risks to existing production behavior while preserving the ability to raise—not merely inherit—the organization's standard. | Caveat: Using revealed tolerance as a floor can normalize legacy risk; it should inform consistency, not excuse known unacceptable behavior.
- Claim: Before changing an agent in response to an evaluation regression, verify that the evaluator is correct; judges are software that require their own calibration and iteration. | Evidence: In Agrawal's example, a clinical judge incorrectly flags standard FDA caffeine guidance—400 mg for most adults, with lower limits for pregnancy or some medications—as hallucination because it lacks an individualized check. By contrast, an answer claiming 1,000 mg daily is safe is correctly flagged and requires an agent fix. | Implication: Maintain evaluator versioning, adjudicated samples, false-positive/false-negative tracking, and a triage step that distinguishes 'fix the agent' from 'fix the judge.' | Caveat: Judge validation adds review work and slows reactive changes, but changing production behavior on an unvalidated score can introduce fresh failures.
Detailed Brief
Operating model: monitoring creates a growing safety-control library
- Claims: Some failures cannot be permanently removed through prompt edits because they reappear under new user prompts, tool configurations, or model behavior changes.; A production failure should be converted into a new evaluation signal or judge, allowing the monitoring system to expand with the product.; The practical bottleneck is not model access or compute; it is human capacity to review signals, identify patterns, and decide what action to take.
- Evidence: Agrawal says each successive prompt-level fix can yield diminishing returns and that the failure rate does not reach zero.; Her operational rule is that a newly observed production failure means the system now needs a new judge.; Dashboards and automated scoring can refresh frequently, but they do not replace human interpretation.
- Caveats: Adding judges without ownership, calibration, and review capacity can create noisy monitoring rather than meaningful safety coverage.
- Implications: Budget headcount and operating cadence for evaluator maintenance and trace review as first-class reliability infrastructure.; Treat the evaluation suite as a versioned, expanding product asset rather than a static test corpus.
Governance rules for unresolved launch work
- Claims: Cross-functional disagreement is expected because clinical, legal, compliance, product, and engineering teams evaluate a single defect through different loss functions.; Human-in-the-loop design applies to launch governance as well as model output handling.; A fast follow is committed debt when a feature was intentionally omitted from launch, not a discretionary item that may disappear into a backlog.
- Evidence: The launch scenario is set five days before release, with clinical seeking a hold for member safety, legal focused on regulatory exposure, compliance on audit risk, product on adoption, and engineering on schedule velocity.; Agrawal frames the final decision as requiring a human when the architecture and monitoring layers cannot resolve the issue.
- Caveats: The transcript does not specify a formal approval matrix, risk register format, or deadline mechanism for enforcing fast-follow commitments.
- Implications: Create a documented risk-acceptance path with accountable signatories and expiration or follow-up dates.; Make deferred guardrail work visible as launch debt with an owner and committed remediation window.
Notable Concepts & Terms
- Guardrails-first architecture: Safety is built into system boundaries and execution flow before model behavior is considered, rather than bolted onto prompts after the product exists.
- Code above the model: A deterministic pre-model layer handles irreversible decisions such as emergency escalation, sensitive routing, authorization, and identity checks.
- PHI at the pipeline boundary: Protected health information is removed or isolated at ingestion so downstream storage, dashboards, and non-production systems do not receive it by default.
- Continuous production judges: Automated evaluators score live conversations across multiple quality and safety dimensions, enabling regression and drift detection after launch.
- Revealed risk tolerance: The organization's true release standard is inferred from risks it already permits in production, not merely from stated zero-defect aspirations.
- Asymmetric default: When uncertain, select the error with lower downside: hold suspected safety issues, but generally release low-impact polish imperfections.
- Verify the scorer: A judge regression is only a signal until adjudication establishes whether the evaluator or the agent is actually wrong.
- Fast follows as committed debt: Capabilities deferred from launch should be treated as binding remediation obligations, not optional backlog items.
Operator Notes / Why Ken Should Care
- Create an inventory of irreversible agent actions and require deterministic pre-model gates for each, including authorization, routing, escalation, and external side effects.
- Add a formal evaluator-adjudication workflow: no material prompt, model, or agent change should be made from a judge-score movement until sampled cases establish whether the judge is valid.
- Define defect classes with pre-agreed defaults, named risk-acceptance authority, and a rule that capacity constraints cannot alter severity.
- Instrument live agent traffic for automated evaluation, direct user feedback, and risk-tiered trace sampling; reserve complete review coverage for the highest-stakes paths.
- Staff and assign ownership for the human review queue before expanding agent deployment; monitoring without decision capacity is a false control.
- Track deferred guardrails and safety work as launch debt with an owner, due date, and explicit closure criterion.
Source/Metadata
- Title: Guardrails First: Engineering Member-Facing Health AI — Rashi Agrawal, Hinge Health
- Transcript words: 4348
- Duration seconds: 1309
- Timestamp note: No timestamps or chapters were provided in the transcript; the latter portion substantially repeats the monitoring and decision-framework material.
Transcript
Hello and good morning. Chaitana gave us a great overview of what Abridge does. Today, I'm here to talk more from a practitioner's view of how we are building healthcare AI within Hinge Health. Hi, I'm Rashi Agrawal. I lead AI and ML at Hinge Health. Today, I will be talking about guardrails that are needed to build member-facing healthcare AI. I want to talk a little bit about the state of healthcare AI right now. We do have a lot of frontier models that are running. Believe it or not, 40 million people actually use these models for triaging their healthcare issues. But there is a caveat. These are some of the headlines that have been happening in the past year or so. "Poisoned by a chatbot." Let's start with this one person. A 60-year-old healthy man asked a popular AI assistant how to cut salt from his diet. The LLM told him to swap it with sodium bromide. He did it for three months. He landed in the ER with paranoia and hallucinations. Bromide levels were 200 times the safe limit. Three weeks in the hospital. For what? For following diet advice? Let's look at another pattern. The first independent safety test of a consumer health AI out of Mount Sinai found that this health AI is under-triaging life-threatening emergencies 50% of the time: diabetic ketoacidosis, respiratory failure. It told people to go see a doctor in a day or two. The right answer was ER, right now. And this is in fringe. In February, ECRI, the patient safety group that hospitals trust to rank their top risks, named AI chatbot misuse as the number one health technology hazard of 2026, number one on the list that they publish every year. So this is not really a frontier problem. This is the production baseline that we are working with right now. So the question becomes: how do you ship AI to somebody who's already trusted you with their health? The next 20 minutes are all about that. It starts with three non-negotiable foundations. One, the constraint is the architecture. Most AI safety failures in healthcare are not model failures. They are architectural decisions that were made before even a single token was generated. Two, deterministic rules belong above the model, not inside it. What can never be wrong cannot be left to probability. And three, safety is a continuous evaluation layer, not a one-time gate. Launch of your product is where the real risk starts, not where it ends. That's the first part of what I want to talk about today. The second half is what happens when the architecture is not enough and a human has to make a decision of what ships versus what holds. Let's start with layer one. Protecting PHI takes both policy and architecture. Policy tells you what to protect, and architecture makes sure that it actually happens. The first thing that changes when you start shipping member-facing health AI is where PHI lives. Most teams treat PHI as a runtime problem, something to redact when a log gets written to a dashboard. That's the reactive version. The architecture version strips PHI at the pipeline boundary, at ingestion, before it ever reaches the data lake. By the time the data is stored, the PHI is gone. So a developer opens a dashboard. There's nothing to redact. The PHI was never there. The rest of the architecture works in a similar way. Production and non-production stay completely separate. No pipes in between, because even a single pipe is all that it takes for member data to leak into a dev environment. HIPAA laws are very stringent, especially in healthcare. The regulatory bar is much, much higher. So you have to be very careful about the architecture that you're designing. Another big thing: access depends on two things, your role and your geographic region. We all work with teams that are geographically distributed, but not everybody has access to PHI. That is a certification, a policy that is applied to specific regions only. An engineer outside the regulated region cannot reach raw PHI at all. And the compliance rules, HIPAA, FDA's good machine learning practice, state laws like Texas, Triaga, they are not afterthoughts. They are the grounding input in how you actually design your systems. You cannot slap HIPAA on top of an underlying system or an architecture. You start with it and let the architecture grow around it. When PHI is protected at the architecture level, you're not just trusting that the policies will get followed. You're actually relying on a system that's incapable of certain failures. Let's move to layer two. Probabilistic systems are great at generation. We all know that. However, they are unreliable for things that can never be wrong. So the rule is very simple. Must-not-fail behavior belongs above your prompt, above the model. And what does "above the prompt" actually mean? It means that there is a code layer that runs first on every turn before the model even runs. The code layer is what makes your irreversible decisions: the decision of whether this is an emergency escalation, should they be routed to 911, should a clinician step into the loop. All of those are irreversible decisions that need to lie at a deterministic code-level layer. The model handles the long tail of your conversations and interactions with your members. The picture to hold in your head is a stack: code on top, model below. Every turn goes through the code layer first. Most turns do reach the model, but the model never gets a vote on high-stakes calls. Here's how you can think about it in a different way. A model is not a guardrail. A model with a system prompt is also not a guardrail. Code that runs above the model is closer. Even the labs that build these frontier models publish the authority hierarchy: root, system, developer, user, guideline. Every layer above user is one prompt injection away from being overridden. If the labs themselves don't trust the prompt as a security boundary, neither should you. So what does live in this code layer? Let's examine it a little bit. Let's take three examples. First, very relevant to healthcare, is emergency escalation. If a member mentions self-harm, suicidal ideation, or an acute medical emergency, the system must route to 911 or 988. The model should not even see this turn. Code runs first, decides and routes, and makes a decision right away. Another example: intent routing. Which capability in your underlying multi-agentic system, multi-agentic architecture, handles a conversation turn? Is it clinical? Is it tech support? Is it education from the millions of credited articles? Is it exercise recommendation? The model can help to classify, but high-stakes paths must again take a deterministic route at the top itself. You don't want a clinical question quietly being routed to your generic tech support agent. That's unrecoverable. Third, identity verification. Anything that touches member data has to check that the right member is at the other end. That's an authentication check, and authentication is a security boundary. Prompts are not. The underlying pattern across all three: code runs first. Code makes the irreversible decisions. The model handles what's left. Last but not least, layer three. As we all know, safety is not a gate you pass once. It is a continuous layer that runs the whole time. Most teams treat evals as a pre-launch checklist. You run your test, you ship, you move on. That's necessary, of course, but that's hardly enough. What actually holds up in production is judges that continuously keep scoring real conversations as they happen. Not a saved golden data set. Live traffic, scored on a lot of dimensions all the time. These signals come from three sources, and each one catches something different. First, automated judges. Thirty, forty, name it, as much as you can scale. Automated judges with multiple dimensions, always refreshing: clinical accuracy, safety, escalation, relevance, drift, refusal, etc., etc. I can keep going on. But you get the point. These are the automated judges that are always going to catch regressions and even sensitive drops in quality. Second, your gold mine of information. That's going to be member feedback. Thumbs up, thumbs down on each and every single message. That's the truth signal. That's your member communicating with you, and it's the only one that comes straight from the person that you're serving it to. It catches tone problems and things that judges miss. Third, sample traces. Random samples spread across capabilities, with high-stakes cases checked every single time. One hundred percent sampling on those. Ultimately, people need to read these signals. People are going to catch what no single metric is going to catch. And here's the part that nobody really warns you about. The bottleneck is not the compute, the models, the capability. It's actually having enough people to read the signal and act on it. One more thing about layer three. Some failures, you can't just prompt away. You ship the fix, it comes back under new conditions. New prompts, new tools, the model shifts. You ship the fix again. Each round buys you less and less. The rate never hits zero. At this point, monitoring is not a last resort. It is the first resort, which is always on. A new failure that you see in production simply means you now have a new judge. Your underlying architecture and your system need to be able to keep scaling with new judges, new monitoring, as you keep scaling your consumers. And that's the point. Monitoring is how you know that the architecture is still holding. But monitoring also tells you when the architecture is not enough. And when the architecture is not enough, a human has to decide. This is the second part of my talk, where I want to focus on the decisioning frameworks. Let's take an example. You're about to ship consumer AI again in the healthcare space, and you have a feature, a specific capability, that you're about to launch. There is one issue left on the board five days before your launch. You have multiple different stakeholders. Five stakeholders look at the same issue. Each one sees a different risk, and they don't agree on what to do about it. Clinical sees member safety risk. They want to hold the launch. Legal sees regulatory exposure. Compliance sees audit risk. Product sees adoption risk. If it ships broken, the feature won't land. And engineering sees velocity risk. They can't fix it without slipping the date. They want to ship. Five rational people fight different risks and five very different fixes. So what do you do? Do you hold the launch and fix, or do you actually ship? The next slide is the framework I actually use for making these decisions. Five rules. This is how I think about decisions when stakeholders disagree. Rule one: worst case always wins. Severity is set by the worst possible outcome, not the average. And this is extremely relevant in healthcare. A bug that lightly annoys 100% of users is way less severe than one that could cause serious harm in 0.1% of cases. This is non-negotiable. The worst case matters more than the average case, always. So when you're triaging, don't ask, "How often does this happen?" Ask, "What's the worst version of this?" That sets the severity. Rule two: severity is not capacity. This one keeps politics out of it. As we all know, as we ship features, there's always a little bit of contention between timelines, features, deliverables. But a bug's severity comes from the harm that it causes, not who owns it, not whether your team has the capacity to fix it, not how hard the fix is. You have three options in front of you at this point: fix, delay the launch, or accept the risk with explicit sign-off. Those are the three. You never quietly downgrade a bug just because you can't get to it. Rule three: asymmetric default. When you don't know what to do, always pick the safer mistake. And there are two spectrums to it. One is safety bugs versus polish. The other side is polish bugs. For safety bugs, the math is one-sided. Shipping a real safety bug is much worse than delaying for a false alarm. So for safety bugs, when you're not sure, always hold and fix. On the other side, for polish bugs, the math runs the other way. Delaying a launch costs more than shipping a small flaw. So when you're not sure, ship in case of polish bugs. Ultimately, the framework doesn't decide for you. It just tells you which way to lean. Rule four: revealed risk tolerance, not stated risk tolerance. Your launch bar is what your org already accepts in production, not what it says it will accept. If a behavior has been live in your existing product for weeks, months, without escalation, without member complaints, without leadership concern, you cannot call it a launch blocker just for a new thing. Your stated risk tolerance might be no bugs in production. But your revealed risk tolerance is what's actually shipping today. Calibrate to the revealed one. That's the floor. Rule five: humans are the constraint. Judges scale. Pattern interpretation doesn't. Always, always design for human in the loop. Judges score traces automatically. Dashboards refresh every few hours. None of that is hard anymore. But what's hard is having enough people to read the signal and act on it. One more piece around this. Fast follows are committed debt, not an optional backlog. If you didn't ship it at launch, it's not a wish-list item. It's already committed. The five rules tell you how to decide. But they all assume one thing: that your underlying signal is true. So here's the discipline that needs to come first. In a non-deterministic system, the judge is also non-deterministic. Before you trust the score, verify the scorer. And here's what it looks like in practice. Say you're watching a clinical accuracy judge in production. The score has been steady at 4.9 for weeks. Today, it drops to 4.5. Tomorrow, it stays at 4.5. The immediate instinct is, "Let's start changing the prompts. The agent is broken. Let's fix the agent." That's reactive, and it's risky. You fix one thing and you break another. Worse, you're changing the agent based on a signal that might not be true. And the discipline needs to be different. First, ask whether the judge is right. We can solidify that with a concrete example. Let's take it side by side. In scenario A, same question: member asks about caffeine. The agent gives FDA standard guidance: 400 milligrams for most adults, less if pregnant or on certain medications. The judge flags it as a hallucination because the agent mentioned pregnancy and medications without checking. But that's just clinical context. The judge is over-calling in this case. Fix the judge in this scenario. For the same question, scenario B: the agent says 1000 milligrams a day is fine. That's well above the safety limits. The judge correctly flags it, and the agent is wrong. In this case, fix the agent. The rule is: always ask, is the judge right, before changing the agent's response. Fixing a judge prompt is not cheating. Judges are software too, and they need to continuously evolve. This is what production discipline looks like when the system is not deterministic. The whole talk in one slide. If you screenshot one thing, this would be it. Six takeaways, three from architecture, three from decisioning. On the architecture side, the pattern is very simple. Don't X what you can buy. Don't policy what you can architect. Don't prompt what you can code. Don't gate what you can monitor. On the decisioning side, the pattern is how humans decide when the system cannot. Score by the worst case and default to the safer mistake. Calibrate to your org and always design for the human in the loop. Fast followers are debt, not backlog. Yes, building guardrails first is slower than bolting them on later. But that's the design, not limitation. We are not building a generic low-stakes chatbot. We are building a system that has to be worthy of someone's health. The architecture is how. The decisioning is when. And member trust is why. Thank you. Let's continue the conversation on LinkedIn. Thank you. Bye-bye. Bye-bye. Bye-bye. Bye-bye. Bye-bye. Bye-bye. Bye-bye. Bye-bye. Bye-bye. You run your test, you ship, you move on. That's necessary, of course, but that's hardly enough. What actually holds up in production is judges that continuously keep scoring real conversations as they happen. Not a saved golden data set. Live traffic. Scored on a lot of dimensions all the time. These signals come from three sources and each one catches something different. First, automated judges. 30, 40, name it, you know, as much as you can scale. Automated judges with multiple dimensions always refreshing. Clinical accuracy, safety, escalation, relevance, drift, refusal, etc., etc. I can keep going on. But you get the point. These are the automated judges that are always going to catch regressions and any even sensitive drops in quality. Second, your gold mine of information. That's going to be member feedback. Thumbs up, thumbs down on each and every single message. That's the truth signal. That's your member communicating with you. And it's the only one that comes straight from the person that you're serving it to. It catches tone problems and things that judges miss. Third, sample traces. Random samples spread across capabilities with high-stake cases checked every single time. 100% sampling on those. Ultimately, people need to read these signals. People are going to catch what no single metric is going to catch. And here's the part that nobody really warns you about. The bottleneck is not the compute, the models, the capability. It's actually having enough people to read the signal and act on it. One more thing about layer three. Some failures, you can't just prompt away. You ship the fix, it comes back under new conditions. New prompts, new tools, the model shifts. You ship the fix again. Each round buys you less and less. The rate never hits zero. At this point, monitoring is not a last resort. It is the first resort which is always on. A new failure that you see in production simply means you now have a new judge. Your underlying architecture and your system needs to be able to keep scaling with new judges, new monitoring as you keep scaling your consumers. And that's the point. Monitoring is how you know that the architecture is still holding. But monitoring also tells you when the architecture is not enough. And when the architecture is not enough, a human has to decide. And this is the second part of my talk where I want to focus on the decisioning frameworks. Let's take an example. You're about to ship, you know, consumer AI again in the healthcare space. And you have a feature, a specific capability that you're about to launch. And there is one issue left on the board five days before your launch. And you have multiple different stakeholders. Five stakeholders look at the same issue. Each one sees a different risk. And they don't agree what to do about it. Clinical sees member safety risk. They want to hold the launch. Legal sees regulatory exposure. Compliance sees audit risk. Product sees adoption risk. If it ships broken, the feature won't land. And engineering sees velocity risk. They can't fix it without slipping the date. They want to ship. Five rational people fight different risks. And five very different fixes. So what do you do? Do you hold the launch and fix? Or do you actually ship? The next slide is the framework I actually use for making these decisions. Five rules. This is how I think about decisions when stakeholders disagree. Rule one. Worst case always wins. Severity is set by the worst possible outcome, not the average. And this is extremely relevant in healthcare. A bug that lightly annoys 100% of users is way less severe than one that could cause serious harm in 0.1% of cases. This is non-negotiable. The worst case matters more than the average case. Always. So when you're triaging, don't ask, how often does this happen? Ask, what's the worst version of this? That sets the severity. Rule two. Severity is not capacity. This one keeps politics out of it. As we all know, as we ship features, there's always a little bit of contention between timelines, features, deliverables. But a bug's severity comes from the harm that it causes. Not who owns it. Not whether your team has the capacity to fix it. Not how hard the fix is. You have three options in front of you at this point. Fix, delay the launch, or accept the risk with explicit sign-off. Those are the three. You never quietly downgrade a bug just because you can't get to it. Rule three. Asymmetric default. When you don't know what to do, always pick the safer mistake. And there are two spectrums to it. One is safety bugs and polish. The other side is polish bugs. For safety bugs, the math is one-sided. Shipping a real safety bug is much worse than delaying for a false alarm. So for safety bugs, when you're not sure, always hold and fix. On the other side, for polish bugs, the math runs the other way. Delaying a launch costs more than shipping a small flaw. So when you're not sure, ship in case of polish bugs. Ultimately, the framework doesn't decide for you. It just tells you which way to lean. Rule four. Revealed risk tolerance. Not stated risk tolerance. Your launch bar is what your org already accepts in production. Not what it says it will accept. If a behavior has been live in your existing product for weeks, months, without escalation, without member complaints, without leadership concern, you cannot call it a launch blocker just for a new thing. Your stated risk tolerance might be no bugs in production. But your revealed risk tolerance is what's actually shipping today. Calibrate to the revealed one. That's the floor. Rule five. Humans are the constraint. Judges scale. Pattern interpretation doesn't. Always, always design for human in the loop. Judges score traces automatically. Dashboards refresh every few hours. None of that is hard anymore. But what's hard is having enough people to read the signal and act on it. One more piece around this. Fast follows our committed debt. Not an optional backlog. If you didn't ship it at launch, it's not a wish list item. It's already committed. The five rules tell you how to decide. But they all assume one thing. That your underlying signal is true. So here's the discipline that needs to come first. In a non-deterministic system, the judge is also non-deterministic. Before you trust the score, verify the scorer. And here's what it looks like in practice. Say you're watching a clinical accuracy judge in production. The score has been steady at 4.9 for weeks. Today, it drops to 4.5. And tomorrow, it stays at 4.5. The immediate instinct is, let's start changing the prompts. The agent is broken. Let's fix the agent. That's reactive. And it's risky. You fix one thing and you break another. Worse, you're changing the agent based on a signal that might not be true. And the discipline needs to be different. First, ask whether the judge is right. We can solidify that with a concrete example. Let's take it side by side. In scenario A, same question, member asks about caffeine. The agent gives FDA standard guidance. 400 milligrams for most adults, less if pregnant or on certain medications. The judge flags it as a hallucination. Because the agent mentioned pregnancy and medications without checking. But that's just clinical context. The judge is over calling in this case. Fix the judge in this scenario. For the same question, scenario B, the agent says 1000 milligrams a day is fine. That's well above the safety limits. The judge correctly flags it. And the agent is wrong. In this case, fix the agent. The rule is always ask, is the judge right before changing the agent's response. Fixing a judge prompt is not cheating. Judges are software too. And they need to continuously evolve. This is what production discipline looks like when the system is not deterministic. The whole talk in one slide. If you screenshot one thing, this would be it. Six takeaways. Three from architecture. Three from decisioning. On the architecture side, the pattern is very simple. Don't X what you can buy. Don't policy what you can architect. Don't prompt what you can code. Don't gate what you can monitor. Don't you can monitor. On the decisioning side, the pattern is how humans decide when the system cannot. Score by the worst case and default to the safer mistake. Calibrate to your org and always design for the human in the loop. Fast followers are debt, not backlog. Yes, building guardrails first is slower than bolting them on later. But that's the design, not limitation. We are not building a generic low stakes chatbot. We are building a system that has to be worthy of someone's health. The architecture is how. The decisioning is when. And member trust is why. Thank you. Let's continue the conversation on LinkedIn. Thank you. Bye-bye. Bye-bye. Bye-bye. Bye-bye. Bye-bye. Bye-bye. Bye-bye. Bye-bye. Bye-bye.