AI Engineer

Inside 847 Production Clinical AI Notes — Sebastian Fox, Composo

1999 summary words 9 min summary Watch video

Start with the signal

9 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: For high-stakes AI, static rubrics and generic LLM judges fail because they cannot determine which context-dependent errors matter; evaluation must continuously learn from real failures and expert judgment.
  • Why it matters: This is a practical architecture for building safer agent and AI-operation control planes where a plausible output can still cause material harm through omissions, intent reversals, or contextually critical details.
  • Best use: Use it to redesign evaluation from a one-time benchmark or prompt rubric into a production feedback loop: mine real failures, capture expert reasoning, and retrieve relevant precedent and policy into each evaluation.

Executive Summary

Sebastian Fox argues that the most dangerous AI failures are not conspicuous hallucinations but outputs that look reasonable while omitting, altering, or over-interpreting the one fact that changes the decision. His clinical examples make the point sharply: an AI note omits jaw pain in a patient over 50 with a new headache, obscuring possible giant-cell arteritis and a same-day sight-threatening emergency; another records tests as arranged even though the clinician and patient decided to defer them.

The central diagnosis is that current evaluators can detect textual discrepancies, but they cannot reliably determine their significance. A static rubric, frontier-model judge, optimized prompt, or deterministic concept-difference check can notice obvious errors, yet it often approves notes containing serious omissions. The difficult task is not identifying that the source and output differ; it is applying domain judgment to decide whether a difference is harmless, consequential, or catastrophic in this specific case.

Fox proposes treating evaluation as a continuous operational loop rather than a fixed system. Teams should discover failure modes from real production outputs, have domain experts leave reasoning-rich corrections on representative cases, and assemble a case-specific evaluation context at runtime by retrieving similar expert-judged examples, applicable corrections, and reference guidelines. This preserves changing, explainable expert judgment outside model weights and makes new feedback usable immediately.

Although framed around ambient clinical scribes, the argument maps directly to contract review, support agents, compliance workflows, and any agent system that can be confidently wrong. The practical starting point is deliberately modest: collect free-form expert comments on real outputs, then use those comments to build a living failure-mode ontology and retrieval-backed judging layer.

Key Takeaways

  • Claim: Plausible-looking omissions are more dangerous than obvious hallucinations because they can quietly change the clinical or operational decision while leaving every recorded sentence technically defensible. | Evidence: A note summarized a new headache as likely tension-type and omitted jaw pain while chewing in a patient over 50; that missing detail could indicate giant-cell arteritis, requiring urgent steroids to avoid vision loss. In another example, a note preserved the discussed plan to arrange tests rather than the final agreement to defer tests and try antibiotics. | Implication: Evaluation should explicitly search for missing decision-changing context and final-intent reversals, not merely factual contradictions or overt fabrication. | Caveat: The examples are clinical, but the speaker explicitly generalizes the pattern to contracts, customer support, and other high-stakes workflows.
  • Claim: Production error rates in ambient clinical notes are material enough that passive trust in deployment is unjustified. | Evidence: Fox cites the largest real-world study as finding serious potentially harmful errors in about 1 in 20 notes, important omissions in nearly 1 in 5, and hallucinations in more than 1 in 10. He says ambient scribes are already used in roughly one-third of US practices, while adverse-event reporting for these systems is generally absent. | Implication: For any consequential AI workflow, absence of incident reports is not evidence of safety; teams need direct output-quality surveillance and sampling. | Caveat: The transcript does not name or link the underlying study, and the cited metrics should be independently verified before being used for external claims.
  • Claim: The difficult evaluation problem is prioritization, not basic comparison: a judge must know which discrepancy matters in the current context. | Evidence: The same omitted travel detail is insignificant for a patient with blood in the urine who visited France but potentially diagnostic for one who visited Lake Malawi, where freshwater exposure suggests schistosomiasis. Similarly, interpreting “it just happened” as sudden-onset headache adds a red flag for intracranial bleed that the patient did not explicitly state. | Implication: A universal severity rubric will systematically underperform where materiality depends on case context, organizational policy, and evolving expert practice. | Caveat: This domain judgment is tacit, contextual, and moving: experts may disagree, guidelines evolve, and institutions can apply different standards.
  • Claim: A generic LLM-as-judge, even when enhanced with a detailed rubric and deterministic checks, can become a second silent failure layer rather than a safety net. | Evidence: Fox ran the same generated notes through a serious judging setup using transcript, note, context, a faithfulness rubric with pass/fail examples, possible rubric optimization, and deterministic medical-concept comparisons. Most notes passed; about one in five clean passes still contained a serious error, often an omission. | Implication: Do not use a high judge-pass rate as a safety guarantee unless the judge has been audited for false negatives on real, consequential failure modes. | Caveat: The speaker characterizes the comparative performance qualitatively and does not provide the underlying dataset size, evaluation protocol, or exact metrics.
  • Claim: The right place to store hard-to-specify quality standards is an evolving corpus of expert judgments and references, not only a static prompt rubric or model weights. | Evidence: Fox frames three choices: specify standards in prompts/rubrics, bake them into weights through fine-tuning or continual learning, or retain past expert judgments, corrections, and reference documents as retrievable context. He favors the third because a new example can affect the next call immediately, the basis for a score is inspectable, and changing standards do not require retraining. | Implication: Build evaluation memory as a first-class system asset: versioned cases, expert rationale, corrections, policies, and guidelines should be indexed for case-specific retrieval. | Caveat: Retrieval does not eliminate the need for strong base models, data governance, expert review processes, or validation of the retrieval and judging pipeline.
  • Claim: A durable evaluation system is a recurring discover-capture-calibrate loop driven by production data. | Evidence: The proposed loop is: discover failure modes by clustering real outputs and naming an emergent failure-mode ontology; capture expert scores, reasoning, and corrections; then calibrate each new output using retrieved similar cases, applicable corrections, and guidelines. Each new judgment sharpens subsequent evaluations and new failure modes enter the loop as they surface. | Implication: Treat evaluation operations as an ongoing capability with feedback ingestion, taxonomy maintenance, retrieval quality monitoring, and periodic expert calibration—not as a benchmark completed before launch. | Caveat: Synthetic tests remain useful for regression coverage, but Fox argues they cannot reveal failures that the team did not already anticipate.

Detailed Brief

Failure taxonomy: transcription errors versus reasoning and summarization errors

  • Claims: Ambient-scribe failures occur both in transcription and in downstream generation, but Fox focuses on failures that remain even with a perfect transcript.; With correctly transcribed words, the model can still add unsupported content, change stated content, or omit necessary content.; The most problematic errors are often subtle inference errors and failure to represent the final agreed plan, rather than blatantly invented facts.
  • Evidence: Sound-alike transcription examples include Humalog being heard as Humulin, hyperthyroidism becoming hypothyroidism, and dropping “no” from “no evidence of cancer.”; Humalog and Humulin are presented as clinically distinct insulins whose confusion could materially disrupt blood-sugar management.; The speaker describes generated notes from three leading production ambient-scribe products and plots errors by importance and whether a strong automated check catches them; the high-importance errors cluster among the missed cases.
  • Caveats: The transcript does not identify the products, disclose sample size, or provide the plot and labeling methodology.; A system design should separate audio/transcription quality controls from semantic faithfulness and decision-integrity evaluation; one will not fully cover the other.
  • Implications: Maintain distinct incident categories for transcription, unsupported inference, meaning reversal, omitted context, and final-plan/intent errors.; Use evaluation data to determine whether quality investment belongs in speech recognition, note generation, workflow state handling, or the post-generation control layer.

Why this architecture is an evaluation-control-plane pattern, not just a clinical technique

  • Claims: Verification is only inherently easier than generation when correctness has an externally checkable target, such as a compiler or unit test.; High-stakes natural-language outputs lack a complete predefined unit test because the operational meaning of “good” includes tacit expert judgment.; The proposed runtime calibration is fundamentally context engineering: construct the relevant standard for the individual output rather than apply one frozen standard across all cases.
  • Evidence: Fox contrasts math and code, where verifiers are cheap and explicit, with the question “is this note safe and complete?” where the verifier itself must be built.; For the headache example, the suggested retrieval bundle includes similar expert-judged cases, red-flag corrections for a new headache over age 50, and relevant clinical criteria or guidelines.; He analogizes the learning problem to RLHF: quality cannot be fully written as a reward function, so it must be learned by exposure to judged examples.
  • Caveats: The proposed approach depends on retrieval selecting genuinely relevant precedent; irrelevant or poorly governed examples can bias judgments or expose sensitive records.; Explainability improves when the retrieved evidence is visible, but it still requires audit trails showing what was retrieved, what influenced the score, and who approved the underlying examples.
  • Implications: For agent systems, route high-impact outputs through a policy-and-precedent retrieval layer before approval, execution, or escalation.; Design the evaluator to produce an auditable decision package: detected issue, severity, retrieved precedent, governing reference, confidence, and recommended disposition.

Notable Concepts & Terms

  • Ambient scribe: An AI system that listens to clinical encounters and generates documentation; the video uses it as a vivid example of risks from plausible but incomplete generated records.
  • Tacit judgment / “taste”: The domain expert's contextual sense of what is material, dangerous, or irrelevant—knowledge that cannot be exhaustively converted into a static rubric.
  • Asymmetry of verification: The idea that checking an answer can be easier than producing it; Fox argues this breaks down when the evaluator must first infer an unwritten, context-specific quality standard.
  • Failure-mode ontology: A taxonomy of real error patterns discovered from production outputs, used to organize expert review, retrieval, and ongoing evaluator improvement.
  • Case-specific calibration: Assembling an evaluation standard dynamically for each output from similar judged cases, expert corrections, and relevant policy or guidelines.
  • LLM-as-judge: Using a language model to evaluate another model's output; the talk warns that generic judges can confidently miss material errors without grounded domain precedent.
  • RLHF: Reinforcement learning from human feedback; invoked as evidence that complex quality criteria are learned from examples rather than fully specified in advance.

Operator Notes / Why Ken Should Care

  • Audit existing AI evaluator pass decisions using a blinded sample of outputs reviewed by domain experts; calculate serious-error false-negative rates, especially for omissions and final-decision reversals.
  • Create a production feedback schema that requires experts to record severity, what was missing or wrong, why it mattered, the corrected output or action, and supporting policy/reference material.
  • Stand up a versioned failure-mode registry sourced from real incidents and near misses; do not begin with only a brainstormed taxonomy or synthetic benchmark set.
  • For high-impact agent actions, require the evaluator to retrieve and log relevant policy, prior adjudicated cases, and domain guidance before allowing autonomous execution.
  • Separate quality controls for source capture/transcription, semantic faithfulness, materiality/severity, and final-state/intent accuracy; assign owners and metrics to each.
  • Treat static evaluator prompts and model fine-tunes as components, not the control plane; monitor retrieval relevance, expert-disagreement rates, and evaluator false negatives as continuous operational metrics.

Source/Metadata

  • Title: Inside 847 Production Clinical AI Notes — Sebastian Fox, Composo
  • Transcript words: 4798
  • Duration seconds: 1188
  • Timestamp note: No usable timestamps or chapter markers were present in the supplied transcript; substantial repeated transcript passages were also present.
Full transcript 3207 words · 21 min read
0:12

This is a clinical note an AI wrote from a real consultation. Take a few seconds and read it.

0:18

It reads like a routine headache. A new headache, likely tension type, take some paracetamol, come back if it doesn't settle. Looks completely fine, doesn't it? Here's what's missing. In the room, she also mentioned her jaw aches when she chews. A new headache, over 50, with jaw pain on chewing, that's giant cell arthritis. And untreated, it can take her sight within days. It's the same day, start steroids now, emergency. And that one line never made it into the note. On the page, it's the paracetamol headache. And nothing in the note is technically wrong. The dangerous part is what isn't there. And so that's what I'm going to talk about today. The dangerous failures are often the ones that actually look completely fine.

0:24

Firstly, who am I? I'm Seb, a medical doctor by background, and now run Composo, where we build AI evaluation systems for high-stakes domains. So that one was a subtle error, but sometimes it's not subtle at all. A man in his twenties sees his GP for a sore throat, tonsillitis. The AI writes that up. It gives him chest pain, suspected angina, diabetes medications he'd never taken, and an address for a hospital that doesn't exist. And I really like the LLM for this one. I think it's a good attempt at a hospital name. And weeks later, he's invited to diabetic eye screening for diabetes he doesn't have. That's genuinely a real case that happened recently.

0:40

Obviously, these crazy ones someone notices, but it's those quiet ones that sit in the record uncaught that are the most challenging and can actually do a lot more damage. And they're not rare at all. In the largest real-world study of these notes, about one in 20 carried an error that was serious enough that it could cause significant harm to the patient. One in 20. That's not theoretical in testing. That's in production on real patients. And that's only the serious ones. If you widen that lens to all errors, nearly one in five had an important omission, and more than one in 10 had a hallucination.

0:46

And AI is being deployed at scale across healthcare, fast. Ambient scribes are one of the leading cases, already in about a third of US practices and climbing. Physician AI use doubled last year, and none of this is tracked. So for most of these systems, there's no adverse event reporting at all. The errors never show up as incidents. They just sit in the record. So errors this common that are going unseen, it's quite hard for me to believe that it's not already affecting patients. It's not that we checked and it's fine. It's that we are flying blind.

0:51

And this isn't just a healthcare problem. It's every high-stakes use of AI. Healthcare shows it more viscerally, because here being confidently wrong can be life and death. But everything I show you can map straight back onto other domains as well. So here's what I want to do. I'm going to show you what exactly is going wrong, why it's going wrong, why the systems we built to capture it don't work, and a suggestion at how maybe we can start to fix that.

0:56

So first, what's going wrong and why? So LMs are getting good, obviously. They don't make stupid mistakes anymore most of the time. So it's not about dumb errors. Everything here came out of three of the best production ambient scribes on the market, ones that we all know. We generated a load of notes across them last week. And this is exactly what's going on right now. This is every failure we found. Each dot is an error, colored by type. Left to right, how much it matters; bottom to top, whether a strong automated check catches it. And that split is the point. A handful up top get caught, but almost everything sits below the line. The ones I care about most are these on the bottom right, the high-stakes and missed ones. Let me show you what a couple of those look like.

1:03

So a woman comes in with a headache. Doctor asks, did it come on suddenly or build up gradually? She says she doesn't know. It just happened. The note records that as abrupt, sudden onset. And sudden onset is a red flag. You can see why "it just happened" could maybe be interpreted and inferred as abrupt onset. But that's a feature that points to a bleed on the brain. She never said it. The model decided it. And now that one word drives the whole workup.

1:10

Here's another. Doctor suggests running some tests, running some tests. The patient says, can we just try antibiotics instead? They agree, hold off on the tests, treat and see how it goes. Note records the opposite: arranged tests today. It kept the plan that they talked out of, not the one they chose. Every line in the note reads fine because it's not really a hallucination at all. It's not wrong. It was there in the original, but it's just not what they ended up deciding.

1:19

So why are these happening? In ambient scribes, there's first transcription and then generation. A lot of it does happen on the transcription layer. It can be words misheard for their soundalikes. So humalog, heard as humulin. Two insulins on completely different timelines, so swapping them could crash a blood sugar. Hyperthyroidism becomes hypothyroidism, the opposite condition. Or a drop to no on "no evidence of cancer" that becomes "evidence of cancer." So these are really hard problems, and they are common. Not the ones I'm going to focus on, because most of what goes wrong is actually even with a perfect transcript.

1:25

It's the model reading the words correctly and still doing one of three things. Either it adds something that was never said, it changes something that was, or it omits something that should be there. Now the blatant version of each of these is really easy to catch. The hard part in all three is the same. It's telling whether that thing that was added or changed or dropped actually matters. It's detecting that slight over-inference versus the dangerous fabrication, the harmless rephrase versus the meaning flip, a dropped line of small talk versus a dropped allergy. So the ones that matter slip through along with all of the ones that don't.

1:33

That call, which difference matters, is taste, effectively, taste effectively. Not aesthetic taste, but essentially judgment. It's whether in this context a missed allergy might kill someone or is not important. And I think there's three properties that really matter about this. It's tacit, so your domain experts have it but they can't fully write it down. It's contextual, so the same detail is critical in one note and noise in the next. And it's moving. The model changes, guidelines change, two good doctors disagree, different hospitals have different definitions. So there's no fixed target to write down.

1:38

And so the model knows the facts, ultimately. They're extraordinarily capable, but what they lack is a sense of what matters here for this specific example. And that's why even brilliant models make these mistakes.

1:46

So one natural move: you're never going to make that generator perfect. Generation is cheap. So stop fixing it at the source. Let it write, put a checker after it, pass only what clears the bar. And that checker should be the easier job. The generator has to get everything right and pay attention to lots of varying instructions, whereas the checker only has to find the one thing that's wrong and just focus on that task. You can also give it more time, more tokens, the exact failure modes to hunt for. Evaluation should be easier than generation. Evaluation is the asymmetry of verification. It's the asymmetry of verification, verifies law. And that's why AI has raced ahead anywhere you can cheaply check the answer, maths and code.

1:51

And doing this is exactly what the best teams do. They put a lot of energy into evaluation. It starts with the gold standard, which is expert humans reviewing notes, which obviously works offline, but you can't put a human on every note in production. So they automate it. They build a serious system. And some of the best versions of this that I've seen are: you take the transcript and the note and context, put it in front of the judge, a detailed rubric for faithfulness with worked pass and fail examples. The rubric may be auto-optimized with Jeeper or something like that. Maybe you have some deterministic NLP to count up medical concepts that are differing between the two. That's a powerful system.

1:59

And yet, I pulled all of those errors earlier out of ambient scribes in an afternoon. So if the evaluation is this good, how are these errors still getting through? So I built this system and ran those same notes through it. And it scored most of them fine. It flagged a handful of them and signed off the rest. But one in five of those clean passes still had some sort of serious error buried in it. And often that was an omission, the things that should have been there and actually quietly weren't. And that's the best version of a judge that I've seen. of the judge, a detailed rubric for faithfulness with worked, pass and fail examples. The rubric,

2:13

maybe auto-optimized with Jeeper or something like that. Maybe you have some deterministic NLP to count up medical concepts that are differing between the two. That's a powerful system. And yet, I pulled all of those errors earlier out of Ambient Scribes in an afternoon. So if the evaluation is this good, how are these errors still getting through? So I built this system and ran those same notes through it. And it scored most of them fine. It flagged a handful of them and signed off the rest. But one in five of those clean passes still had some serious error buried in it. And often that was an omission: the things that

2:52

should have been there and actually quietly weren't. And that's the best version of a judge that I've seen in a lot of teams, and it weighed them through. Why did it do that? It's not stupid. It's a frontier model. Serious engineering behind it, more than clever enough to read the whole encounter and catch every obvious error. And it's not blind either. And that's part of the trap. If you take a note that says start Amoxicillin when the real decision was actually to wait and see, it's faithful to the words. Amoxicillin did come up. But it's a lie about the intent. A good judge might catch that. Might. But whether it flags

3:27

that versus the other dozen other things that it could comment on depends on it knowing what decision matters most. And so it's not blind. It just can't tell what counts, essentially. So the note passes confidently. And you put a judge like that in front of your system. You've not added a safety net. You've added a second silent failure that just nods along with the first. And here's the root of it. So in Matzl code, the verifier comes for free: a unit test, a compiler. But for "is this note safe and complete?" there's no unit test. You have to build the verifier yourself. And verification is only easier

4:02

than generation for the easy bit, i.e. spot the difference between transcript and note. But that's not the hard bit. The hard bit is knowing, of all those differences you've seen, which matter. And that's harder than writing that plausibly good general note in the first place. Because that standard of good was never written down anywhere that the judge can read it. A rubric that you pre-specify is only the taste you could write down. The taste that matters is the part that you couldn't. And so here's a bit more detail on what "matters" looks like. Two patients, both with blood in their urine, both notes dropped the

4:42

same kind of line where they'd been on holiday. One had been to France, the other to Lake Malawi. Same omission, same shape, same mistake. Well, not really, because blood in the urine is obviously worrying either way, and you're going to investigate it. But the France trip is irrelevant. The Lake Malawi trip is the diagnosis. Fresh water in sub-Saharan Africa means just somiasis until proven otherwise. And it completely changes what the management plan is. So that same dropped line in one note is pure noise. In the other, it's the answer. And which one it is, you simply can't write all of that down in advance.

5:20

So if you can't write it down, you can't write taste down, how do you get that into your evaluator and your whole application system? Well, we've answered a version of this before. RLHF exists because you can't write the reward function for good. You learn it from examples by showing it. The only question is where you keep what you've learned. And there are three places. You can either specify it up front. You can stuff the prompt, write the perfect rubric. We've just watched that fail, essentially. You can bake it into the weights, fine-tuning or continual learning. But for a standard that's still

6:02

moving and a score that has to be explainable, the weights, I think, are the wrong place to keep that. They go stale. They can't tell you why. And you can't change them without a retrain. So there's the third option, which I'll show you, which is you essentially keep the taste as the examples themselves: past judgments, expert corrections, references. And for each output, you retrieve the ones that bear on it into the judge's context, add one, and it's live on the next call. You can point at exactly what moved the score. For this problem, it's both better and also cheaper to do. So that's the way to do that. One repeating loop, three steps: discover the failure

6:46

modes from real outputs, capture how your experts judge them, calibrate every output against that. And when the standard moves, the loop moves with it. So in more detail, discover. You don't write that rubric in a vacuum. You have to put the system in production and look at the real outputs. You cluster what goes wrong, and the failure modes surface on their own. You name them. This is your failure mode ontology: discover from your data, not guessed on a whiteboard. And you can't shortcut it. The ways that a real system goes wrong are effectively unbounded. And synthetic test cases only cover the failures you've

7:31

already imagined. The ones that hurt you are often the ones that you didn't. And you only find those in real outputs. So this ontology is your map: what to capture judgment on and what to retrieve against, including the failures that you never thought to check for. After that, it's capture and then calibrate. So those discovered modes, they're not a checklist that the judge runs, but they organize everything: what you ask your experts about, how you index the cases that you'll retrieve. And capturing is a simple part. You put real outputs in front of your experts. A clinician spends a focused few hours leaving

8:09

comments. A session doesn't have to be a months-long labeling project to start with. And you collect their judgment: not just a score, but the reasoning and corrections. And over time, you build up that record of how your experts actually judge. You then calibrate. The generic part of this, you can write down once easily. For example, be faithful or don't drop anything important. But what you can't write down is what counts as a serious miss for this specific note. That's contextual. And it shifts from note to note. So what we recommend is you assemble that on the fly. For each output, your judging agent pulls in

8:42

everything that bears on this one case: its memory of the most similar outputs that it's judged before and how they scored, the expert corrections that apply, the reference docs and guidelines. It's just context engineering per output. And crucially, not just one pre-specified rubric in a vacuum, and not a model that you have to retrain every week, but a full case-specific standard assembled for this output. And it's a loop as well. Every output you judge, every correction, sharpens the next. And when a brand-new failure mode appears, discovery surfaces it and it flows straight back in. And so to

9:23

make that a little bit more concrete, that headache that I opened with, the one that was really a possible blindness emergency, here's the kinds of things that you would want to pull in for that note: the nearest cases that your experts have judged. Not this exact patient, but the same shape. Maybe a red flag filed as routine. The corrections that apply, like a new headache over 50, suggest something that you need to check red flags on. And some criteria and guidelines. And you pull all of that in. It hasn't memorized this case. It's a capable model and handed the right context to reason from. Held

10:10

against that, the dropped red flag stands out. It was never actually hard to catch. It just didn't know what mattered. And so if you take that same dataset of generated notes from the start and pass it through these three judging systems: the first, a strong off-the-shelf judge with a rubric, frontier model. It's better than a coin flip, but it misses most of what matters. The second, that serious system that we talked about before: rubric, GPERP, maybe some deterministic checks. Better again, but still missing quite a lot of what counts. The third, the judge running this loop, discovered failure

10:49

modes, calibrated per output against what experts judged, is performing a lot better on this specific dataset. Same notes. The only thing that changes is what the judge was shown. The difference here, it's not more compute or a better prompt. It's that the first two fight taste and lose. They guess the criteria, they freeze one standard, and they go stale. This repeating, evolving loop does the opposite. It discovers the modes, fits the standard to each note, and keeps learning. So you might not write clinical notes, but if you ship anything where being confidently wrong has a cost,

11:26

the contract review that misses the clauses that change the deal, the support agent that promises a refund you don't offer, the same thing is true for all of those. It's watched, if at all, by a judge with no taste for what matters in your domain. So three things: discover your failure modes from real outputs,

11:45

Modes, calibrated per output against what experts judged, are performing a lot better on this specific dataset. Same notes. The only thing that changes is what the judge was shown. The difference here is not more compute or a better prompt. It’s that the first two fight, taste, and lose. They guess the criteria, they freeze one standard, and they go stale. This repeating, evolving loop does the opposite. It discovers the modes, fits the standard to each node, and keeps learning.

11:52

So you might not write clinical notes, but if you ship anything where being confidently wrong has a cost, the contract review that misses the clauses that change the deal, the support agent that promises a refund you don’t offer. The same thing is true for all of those. It’s watched, if at all, by a judge with no taste for what matters in your domain.

11:57

So three things. Discover your failure modes from real outputs, don’t guess them. Capture your experts’ judgment on them, the standard that they can’t write down. Calibrate every output against the cases that they’ve already judged. Not a static rubric, not a retrained model. Then keep that loop running. And if you take one thing away, the easiest place to start is your experts leaving free-form comments on real outputs. That’s the raw material for everything else.

12:04

Your judge can verify anything that you write down in advance, but the standard of good never could be. And so stop trying to write it all down in advance and just start capturing it case by case and evolving it. That’s why evaluation can’t be a thing you build once and freeze. The standard it checks against doesn’t exist on paper. It has to be discovered from real outputs, captured from the people who hold it, and kept alive as it moves. Evaluation isn’t something you have. It’s something that you do continuously over time. Thank you. trip is the diagnosis. Fresh water in sub-Saharan Africa means just somiasis until proven otherwise.

12:16

And it completely changes what the management plan is. So that same drop line in one note is pure noise. In the other, it's the answer. And which one it is, you simply just can't write all of that down in advance. So if you can't write it down, you can't write taste down, how do you get that into your evaluator and your whole application system? Well, we've answered a version of this before. RLHF exists because you can't write the reward function for good. You learn it from examples by showing it. The only question is where you keep what you've learned. And there's three places. You can either specify it up front. You can

12:55

stuff the prompt, write the perfect rubric. We've just watched that fail, essentially. You can bake it into the weights, fine tuning or continual learning. But for a standard that's still moving and a score that has to be explainable, the weights, I think, are the wrong place to keep that. They go stale. They can't tell you why. And you can't change them without a retrain. So there's the third option, which I'll show you, which is you essentially just keep the taste as the examples themselves. Past judgments, expert corrections, references. And for each output, you retrieve the ones that bear on it into the judge's context, add one, and it's live on the

13:34

next call. You can point at exactly what moved the score. For this problem, it's both better and also cheaper to do. So that's the way to do that. One repeating loop, three steps. Discover the failure modes from real outputs, capture how your experts judge them, calibrate every output against that. And when the standard moves, the loop moves with it. So in more detail, discover. You don't write that rubric in a vacuum. You have to put the system in production and look at the real outputs. You cluster what goes wrong and the failure modes surface on their own. You name them. This is your failure mode

14:11

ontology. Discover from your data, not guessed on a whiteboard. And you can't shortcut it. The ways that a real system goes wrong are effectively unbounded. And synthetic test cases only cover the failures you've already imagined. The ones that hurt you are often the ones that you didn't. And you only find those in real outputs. So this ontology is your map, what to capture judgment on and what to retrieve against, including the failures that you never thought to check for. After that, it's capture and then calibrate. So those discovered modes, they're not a checklist that the judge runs, but they organize everything.

14:47

What you ask your experts about, how you index the cases that you'll retrieve. And capturing is a simple part. You put real outputs in front of your experts. Clinician spends a focused few hours leaving comments. A session doesn't have to be a months long labeling project to start with. And you collect their judgment. Not just a score, but the reasoning and corrections. And over time, you build up that record of how your experts actually judge. You then calibrate. The generic part of this, you can write down once easily. For example, be faithful or don't drop anything important. But what you can't write down

15:22

is what counts as a serious miss for this specific note. That's contextual. And it shifts from note to note. So what we recommend is you assemble that on the fly. For each output, your judging agent pulls in everything that bears on this one case. It's memory of the most similar outputs that it's judged before and how they scored. The expert corrections that apply. The reference docs and guidelines. It's just context engineering per output. And crucially, not just one pre-specified rubric in a vacuum. And not a model that you have to retrain every week. But a full case-specific standard assembled for

16:01

this output. And it's a loop as well. Every output you judge, every correction, sharpens the next. And when a brand new failure mode appears, discovery surfaces it and it flows straight back in. And so to make that a little bit more concrete, that headache that I opened with, the one that was really a possible blindness emergency, here's the kinds of things that you would want to pull in for that note. The nearest cases that your experts have judged. Not this exact patient, but the same shape. Maybe a red flag filed as routine. The corrections that apply, like a new headache over 50, suggest something that

16:39

you need to check red flags on. And some criteria and guidelines. And you pull all of that in. It hasn't memorized this case. It's a capable model and handed the right context to reason from. Held against that. The drop red flag stands out. It was never actually hard to catch. It just didn't know what mattered. And so if you take that same dataset of generated notes from the start and pass it through these three judging systems. The first, a strong off-the-shelf judge with a rubric, frontier model. It's better than a coin flip, but it misses most of what matters. The second, that sort of serious

17:18

system that we talked about before. Rubric, GPERP, maybe some deterministic checks. Better again, but still missing quite a lot of what counts. The third, the judge running this loop, discovered failure modes, calibrated per output against what experts judged, is performing a lot better on this specific dataset. Same notes. The only thing that changes is what the judge was shown. The difference here, it's not more compute or a better prompt. It's that the first two fight, taste and lose. They guess the criteria, they freeze one standard and they go stale. This repeating evolving loop does the opposite.

17:55

It discovers the modes, fits the standard to each node and keeps learning. So you might not write clinical notes, but if you ship anything where being confidently wrong has a cost, the contract review that misses the clauses that change the deal, the support agent that promises a refund you don't offer. The same thing is true for all of those. It's watched, if at all, by a judge with no taste for what matters in your domain. So three things. Discover your failure modes from real outputs, don't guess them. Capture your experts' judgment on them, the standard that they can't write down.

18:33

Calibrate every output against the cases that they've already judged. Not a static rubric, not a retrained model. Then keep that loop running. And if you take one thing away, easiest place to start is your experts leaving free-form comments on real outputs. That's the raw material for everything else. Your judge can verify anything that you write down in advance, but the standard of good never could be. And so stop trying to write it all down in advance and just start capturing it case by case and evolving it. That's why evaluation can't be a thing you build once and freeze. The standard it checks against

19:13

doesn't exist on paper. It has to be discovered from real outputs, captured from the people who hold it, and kept alive as it moves. Evaluation isn't something you have. It's something that you do continuously over time. Thank you.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note