AI Engineer

Shipping AI to a Million Patients Without an A/B Test — Jared Joselowitz, Ufonia

2032 summary words 9 min summary Watch video

Start with the signal

9 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: For patient-facing clinical agents, safe deployment cannot rely on A/B tests, rollback, or vendor benchmarks; it requires a traceable evidence system that simulates defined hazards, evaluates behavior at scale, optimizes against clinical cost functions, and expands autonomy only as evidence accumulates.
  • Why it matters: This is a concrete blueprint for replacing reactive production experimentation with a pre-deployment evaluation and governance loop in domains where an erroneous agent response is irreversible and potentially harmful.
  • Best use: Use it as a reference architecture for high-stakes agent evaluation: hazard taxonomy, synthetic edge-case generation, LLM-as-judge validation, cost-sensitive optimization, version traceability, and staged deployment gates.

Executive Summary

Jared Joselowitz describes how Ufonia evaluates DORA, a regulated voice agent that performs clinical conversations such as post-operative follow-ups and pre-operative checks. Ufonia has completed roughly 200,000 UK clinical calls across 20 hospitals and plans to reach one million patients within two years. The central operating constraint is that conventional software-release safety nets do not apply: patients cannot ethically be randomized into an inferior variant, a spoken clinical call cannot be rolled back, and a model vendor's benchmark claims provide no defense after an incident.

The company therefore begins with harm rather than model capability. It identifies clinical hazards with clinicians—such as missing sudden vision loss or severe pain, hallucinating medical advice, or failing to acknowledge distress—and builds tests specifically intended to trigger those failure modes. Its Matrix framework uses a simulated patient agent, PatBot, to generate large volumes of clinically grounded conversations with DORA. Ufonia tested whether this simulator was plausible through a patient-and-public-involvement study, where participants frequently judged simulated-patient conversations as more realistic than real-patient examples.

At scale, the system relies on an LLM judge, Bev Judge, rather than engineers manually reviewing transcripts. The judge receives the dialogue, expected behavior, and clinician-defined hazards, then returns a pass/fail result and explanation. Against a 240-example corpus labeled by 10 clinicians from 10 specialties, the cited top model, Gemini 2.5 Pro at the time of the paper, achieved 0.96 F1 and near-perfect sensitivity—an intentionally prioritized metric because missing a true red flag is much worse than over-escalating one.

Simulation is not presented as proof of real-world benefit. It is the fast inner loop that earns the right to progress into supervised real-patient evaluation and monitored deployment. Ufonia combines real-call data with manufactured rare events and transcription failures, automatically optimizes prompts with a cost-sensitive metric, reruns the simulation gate, and deploys in stages. Its governing principle is that the regulated artifact is not merely the model or prompt: it is the evidence chain linking every call, data set, prompt version, judge verdict, and deployment decision back to a specific hazard.

Key Takeaways

  • Claim: Reactive release practices are unsuitable for patient-facing agents because their safety model assumes an organization can afford a limited number of failures. | Evidence: Joselowitz identifies three broken safety nets: A/B testing can be unethical or illegal when it exposes patients to an inferior change; a spoken call cannot be undone after delivery; and model-card benchmark scores are not adequate evidence in a post-incident review. Even a 5% rollout can expose hundreds or thousands of patients to unproven care. | Implication: Ken should treat high-consequence agent releases as an evidence-gated assurance problem, not as an observe-and-roll-back product experiment. | Caveat: The argument is specifically strongest where outputs can cause irreversible or material harm; lower-stakes products may still appropriately use controlled production experimentation.
  • Claim: The evaluation program should be organized around a clinician-defined hazard model, not generic model accuracy or broad benchmark performance. | Evidence: The talk names concrete hazards: failure to detect red-flag symptoms such as sudden vision loss or severe pain, fabricated medical answers, and failure to respond appropriately to patient distress. Ufonia reports documenting dozens of such hazards and grounding test scenarios in specific clinical workflows rather than abstract prompts. | Implication: Build a failure taxonomy before selecting models or metrics; each system capability, test case, and deployment control should map to a distinct operational harm.
  • Claim: Synthetic patient simulation enables ethically safe, scalable testing of rare and dangerous conversational cases before real patients are exposed. | Evidence: Matrix conditions its PatBot simulated patient on a target scenario—for example, asking whether the agent is human or AI—and runs conversations within a defined clinical context. Ufonia chose simulation over hired actors because actors would not scale with rapid system iteration. In a four-pair patient/public realism study, the majority judged the simulated patient more realistic in three of four comparisons. | Implication: Use synthetic personas and adversarial scenario generation as the pre-production inner loop, especially to manufacture infrequent harms that production data will not surface quickly enough. | Caveat: Realistic simulation is not equivalent to clinical validity: synthetic patients can miss real-world behaviors, so passing simulations only earns the right to conduct carefully supervised testing with real people.
  • Claim: LLM-as-judge evaluation can scale safety review only if it is validated against domain experts and tuned toward the relevant error asymmetry. | Evidence: Bev Judge evaluates simulated dialogues against expected behaviors and clinician-defined hazardous scenarios, producing pass/fail outcomes and failure reasons. Ufonia evaluated it on 240 examples with labels from 10 clinicians across 10 specialties; the cited Gemini 2.5 Pro judge achieved 0.96 F1 and almost perfect sensitivity. | Implication: Before relying on automated evaluation, create an expert-labeled adjudication set and measure the judge on the error type that matters most—often recall/sensitivity for severe hazards rather than average agreement. | Caveat: The reported score is tied to Ufonia's 240-example validation corpus and the then-current model, not a general proof that any LLM judge is clinically reliable.
  • Claim: Prompt optimization should be automated and driven by an explicit cost function, because manual prompt tuning is brittle, subjective, and difficult to audit. | Evidence: Joselowitz cites results where formatting changes shifted benchmarks by 76 percentage points and few-shot example ordering moved performance from near-random to near state-of-the-art in some cases. Ufonia uses JEPA (Genetic Pareto), associated with the DSPy creators: it identifies failed examples, has a strong model reflect on them and update prompts, and retains a Pareto frontier of candidate prompts. The reported iteration time falls from hours or days of manual work to roughly 30–60 minutes. | Implication: Make the objective function—not hand-authored prompt intuition—the core control surface; preserve reproducible prompt versions, optimizer inputs, and evaluation outcomes. | Caveat: Automation does not remove clinical judgment: clinicians must still define the data, hazards, and success metric that the optimizer is allowed to pursue.
  • Claim: Safety optimization must be cost-sensitive because different mistakes have radically different clinical consequences. | Evidence: For red flags, correctly detecting a true red flag is rewarded highly, while missing one is treated as potentially catastrophic. A false positive is framed as a relatively mild burden—such as asking a patient additional questions—so the system can deliberately optimize for sensitivity rather than a flat aggregate accuracy score. | Implication: Define asymmetric loss matrices for agent decisions; do not accept a single aggregate quality number where severe false negatives can be hidden by high performance on routine cases. | Caveat: Higher sensitivity can increase false positives and workflow burden, so the selected trade-off must be clinically and operationally acceptable for the specific use case.
  • Claim: The deployable product is an evidence-backed operating system, with autonomy expanded only through staged proof rather than a one-time model launch. | Evidence: Ufonia's loop combines real-call data, synthetic rare symptoms and mistranscriptions, prompt optimization, and Matrix simulation gating before deployment. It then moves through user testing, supervised clinical evaluation with clinicians in the loop, and monitored deployments. Every call, data set, pinned prompt, and judge verdict is intended to trace to the hazard it addresses. | Implication: Implement release gates that tie allowed agent autonomy to accumulated, auditable evidence, and treat post-deployment data as input to the next safety-evaluation cycle rather than as permission to rely on passive monitoring. | Caveat: The work is never complete: new modalities and languages introduce new hazards, requiring ongoing data collection and reassessment.

Detailed Brief

DORA's operational scope and regulatory framing

  • Claims: DORA performs clinical conversations that clinicians would otherwise conduct, rather than being positioned as a replacement for doctors.; Providing symptom guidance and answering medical questions makes the system a medical device and brings it within a regulated safety obligation.; The speaker reduces the regulatory problem to three questions: what the software does, what can go wrong, and how the organization ensures those failures do not occur.
  • Evidence: The demonstration concerns a cataract-surgery follow-up in which DORA asks clarifying questions about blurry near vision, explains that it should improve in the first days after surgery, and advises avoiding swimming for a month.; Ufonia reports deployment across 20 UK hospitals, approximately 200,000 completed clinical calls, an intended one-million-patient scale-up over two years, and US launch in two live clinics with six further clinics signed across four states.
  • Caveats: The transcript gives operational scale claims but does not provide comparative clinical outcomes, incident rates, or independent validation of DORA's patient benefit.
  • Implications: A conversational interface becomes a substantially different governance category when it can answer domain questions or advise users, even if its intended role is administrative or supportive.

Evidence lineage as the regulatory deliverable

  • Claims: The speaker's central governance standard is full traceability from observed or simulated behavior back to the hazard being controlled.; New deployment data is not just telemetry; it feeds a recurring improvement flywheel alongside deliberately constructed edge cases.
  • Evidence: The stated lineage includes every call, data set, pinned prompt, judge verdict, and the exact hazard addressed.; The testing corpus incorporates real-call data plus synthetic cases such as rare symptoms and mistranscriptions that may not naturally appear at sufficient frequency.
  • Caveats: The presentation does not detail the specific documentation format, approval authority, retention policy, or thresholds used to approve a deployment gate.
  • Implications: For auditability, agent teams need provenance across data, prompts, evaluators, model versions, and releases—not only a final quality score or a static model card.

Notable Concepts & Terms

  • DORA: Ufonia's voice-based clinical conversational agent for workflows such as post-operative follow-up and pre-operative checks.
  • Matrix: Ufonia's simulation framework for generating and evaluating clinically grounded agent-patient dialogues before real-patient exposure.
  • PatBot: The LLM-based simulated patient in Matrix, conditioned on specific test scenarios and personas.
  • Bev Judge: The LLM-based evaluator that assesses conversations against expected behaviors and clinician-defined hazards, returning structured failure judgments.
  • Sensitivity: The priority safety metric for detecting true hazards; the speaker favors overcalling potential red flags over missing dangerous ones.
  • JEPA (Genetic Pareto): A prompt-optimization approach that iteratively reflects on failures, generates revised prompts, and preserves a Pareto frontier of high-performing candidates.
  • Cost matrix: An explicit representation of asymmetric error costs used to optimize agent behavior for the consequences that matter clinically rather than for average accuracy.
  • Simulation inner loop / real-patient outer loop: A staged assurance model: simulation supports rapid pre-deployment iteration, while supervised real-world evaluation supplies the only meaningful evidence of actual patient benefit.

Operator Notes / Why Ken Should Care

  • Create a hazard register for each consequential agent workflow, with severity, expected response, test scenario, responsible domain reviewer, and release-blocking status.
  • Build an expert-labeled evaluation set before adopting an LLM judge; measure false-negative behavior separately from overall agreement or F1.
  • Require prompt, model, evaluator, test-data, and policy versions to be pinned and linked to each release decision.
  • Use an asymmetric loss function for safety-critical workflows, then explicitly approve the added false-positive and human-escalation load that higher recall will create.
  • Establish a staged autonomy ladder: synthetic evaluation, supervised user testing, monitored limited deployment, then broader operation only after predefined evidence thresholds are met.
  • Add manufactured rare cases, ambiguity, distress signals, and transcription faults to regression suites rather than waiting for production incidents to generate coverage.

Source/Metadata

  • Title: Shipping AI to a Million Patients Without an A/B Test — Jared Joselowitz, Ufonia
  • Transcript words: 3415
  • Duration seconds: 1154
  • Timestamp note: No timestamps or chapters were provided in the transcript.
Full transcript 3615 words · 15 min read
0:00

JARED BORISTA- Hello, everyone. It's really nice to see you all. My name is Jared, and I'm going to share some of the work that we do on shipping health care AI safely. So just a little bit about me. I come from South Africa, where I actually studied electrical engineering. And before making the very unique decision to transition to AI a few years ago. I now work as a research engineer for Euphonia, which is a health care company based in the UK. And the work we do, our work within the science team, is we build the safety and evaluation stack behind DORA, which is a clinical conversational agent. And my job and our job within the science team

0:51

is proving that the product is safe before a patient ever actually hears it. So shipping to patients takes away the normal safety nets you would normally ship with. Three of them could be that you can't actually A-B test on patients, of course. Randomizing patients into a worse variant is unethical and often illegal. You can't undo a call. Once DORA says it, it's been said, and there is no rollback. And very importantly, the model card won't save you. You can't claim some model vendors said that they have 92% on some benchmark. It's not a defense at a post-instant review. So everything that we do at Euphonia comes from living within these three constraints.

1:33

So what is DORA? A little bit about DORA. DORA is a voice AI agent that phones patients and has real clinical conversations, post-op follow-ups or pre-op checks. These are calls that a clinician would usually have to make themselves. But DORA doesn't... So DORA obviously doesn't replace doctors, but it just takes a hugely time-consuming task off their plate and gives them their time back. So far, we've done around 200,000 real clinical calls within the UK across 20 hospitals. And we are contracted to scale to a million patients in the next two years. We also released our product in the US this year, and we are live in two clinics.

2:09

And we've signed up for six more clinics across four different states. So I'm just going to... For a bit of better context, I just want to do a quick demo of what the experience with DORA is actually like. That's not uncommon after cataract surgery. Are you having trouble seeing clearly at a distance, up close, or both? At a distance is fine. It's just up close that it's a bit... I'm having a bit of trouble with that. So, when exactly did you first notice it was a bit blurry when trying to read or see things up close?

2:48

Well, really, just after the surgery. And I was wondering how long it would last. Thanks. So, you're asking about the duration? That's understandable. Blurry vision should improve rapidly in the first few days after surgery. Okay. Also, is it all right if I go swimming again? You should avoid swimming for a month after surgery. That's just to be on the safe side and help your eye heal properly. All right. Okay. Awesome.

3:22

So, as you just heard, sometimes DORA asks about symptoms. It gives advice. Sometimes it answers questions. And when it does this, this makes it what's called a medical device. And therefore, it has to, of course, be regulated. And we have to prove that it's safe. And regulation essentially bottles down to three main questions. What does your software do? What could go wrong? And how do you ensure that it doesn't? For normal software, it's quite tractable, the question. But for a model that talks to actual patients, what could go wrong is quite huge. So, where do we start? We start from what could go wrong. We start from the harm. What could actually harm a patient?

3:58

And let's look at some examples. DORA could miss a red flag symptom, such as sudden vision loss or severe pain. A patient could ask a medical question and DORA invents an answer, hallucinates something. The patient could be distressed and DORA just ignores it and carries on without actually acknowledging the distress. There's many, many, many documented hazards of these, 20, 30, 40. And we have to ensure that none of them actually happen in real life. So, how would we actually normally catch a problem like this before it actually spreads? We would lean usually on the playbook that most software ships on. You ship to a small percentage of people. You watch the dashboard.

4:37

You roll back if it breaks. And you iterate from there. This is a very good playbook. It's reactive. It's fast. It's very safe. And it's how the industry usually de-risks the launch. But there's a hidden assumption here that it only works because you can afford to be wrong. For an instance, a bad change hits a few users. You can quickly catch it. You can roll back. And no one's actually literally harmed. This is the one assumption is why that it breaks when the actual user is a patient. For 5%, that could be hundreds if not thousands of patients that have got unproven changes and undue care. Roll back. You can't really roll back. The call has already happened.

5:13

The person has already been harmed. By watching the dashboards, the dashboards going red means that a patient was actually hurt. So, the reactive loop is actually gone now. So, how do you iterate at all when you can't touch a patient until you're sure? Well, for this, we started looking at examples from other high reliable industries. The most obvious one is self-driving cars. Obviously, we're in SF now. There's a lot of Waymos driving around. They've only just come to London, unfortunately, very late to the party. But what did self-driving cars do? Well, they didn't just drive around, crashing into walls and say, we won't do that again, and then doing another RL loop.

5:51

They put millions of miles of simulations first before they actually got any passages into the car. For us, we believe in the same thing. Simulation is only the real ethical option we can go with. You can't run all the hazards I just mentioned on real people as a first grasp. So, for our clinical history taking, we built a simulation framework called Matrix. And I'm going to work through how it works and how we use it to prove that our product is safe. And the paper's on archive if you want to read it along with some of the other research that we do. At its core, Matrix recreates a real clinical conversation but with no real patient in it.

6:31

We use an LLM to play the patient. We call it PatBot. And what does it do? We use a simulated patient and not a hired actor because hired actors don't scale. If we want to iterate very fast and simulate different things at the same time while also updating our system, hiring actors would just be too slow of a process. So, as a first version, we just use a simulated patient. The simulated patient is conditioned on the actual scenario we want to test. The scenario defines exactly what the patient should try and do when talking to our agent. For example, asking whether the agent is a human or an AI.

7:09

PatBot then has a conversation with Dora, our target system, and then generates simulated dialogues. Very importantly, this all happens under a very specific clinical use case context. So, the scenarios are grounded in real clinical workflows and not abstract situations. So, how do we actually make sure that the patient is realistic? If PatBot is sounding robotic, the tests aren't really worth much. So, we had to, of course, validate it. The first thing we did was just a pure script adherence check. If we told PatBot to do something, does PatBot do it? Yes or no? This helped us filter out a lot of maybe weaker models that didn't listen to instructions properly.

7:51

But just purely following instructions does not make a realistic patient. We want a patient that flows more realistically like a real person. So, we set up what's called a PPI study, a patient and public involvement study. We took real patients and we showed them two sets of conversations. One conversation was between a real doctor and a real patient. And one conversation was between Dora and PatBot within our matrix framework. And we showed them these two examples side by side and said, looking at the patient, can you tell which one is the real person and which one is the simulated person?

8:24

So, I'm going to just wait for a few seconds here if you guys want to quickly read the two conversations. Maybe we can do a hands up. Who thinks conversation A is the real person? Who thinks conversation B is the real person? Okay. I think us as engineers sometimes are pretty good at finding these things. But it was actually much more difficult than we thought. And we did this with four conversation sets. In three out of the four, the majority of people actually thought that the simulated patient was more realistic. But the most important thing that we found was, of course, there is no single realistic patient. That doesn't really make sense.

9:02

Some people prefer to speak more verbosely, a lot of ums and ahs. Some people are more straight to the point, a lot of yeses and nos. But the point is that we actually want to simulate all these different scenarios. We want to simulate people with very diverse personas. So, I'm going to wait for a few seconds here if you guys want to quickly read the two conversations. Maybe we can do a hands up. Who thinks conversation A is the real person? Who thinks conversation B is the real person? Okay. I think we as engineers sometimes are pretty good at finding these things. But it was actually much more difficult than we thought. And we did this with four conversation sets.

9:38

In three out of the four, the majority of people actually thought that the simulated patient was more realistic. But the most important thing that we found was, of course, there is no single realistic patient. That doesn't really make sense. Some people prefer to speak more verbosely, with a lot of ums and ahs. Some people are more straight to the point, with a lot of yeses and nos. But the point is that we actually want to simulate all these different scenarios. We want to simulate people with very diverse personas. But what it did show us is that at least our PatBot was realistic enough for this simulation. Okay.

10:14

So now you've got thousands and thousands of simulated dialogues. Are we as the engineers going to read through them one by one and see if a hazard happened? Of course not. It doesn't scale at all, for one. And number two, we're not clinicians. So we don't actually know if an actual hazard has really occurred. So we use another LLM as a judge, of course, and we call it Bev Judge. It takes the simulated dialogue, a set of expected behaviors, and the hazardous scenarios that we talked through with clinicians, and it makes a judgment: pass or fail. If it fails, it gives us a reason why it gave that answer.

10:52

So we get a structured output of which hazards were triggered and what actually went wrong in that scenario. So how did we validate Bev Judge? We validated Bev Judge against expert clinicians. We created a corpus of 240 examples. We had a ground truth of whether a hazard existed in these conversations, yes or no. Then we got 10 clinicians from 10 clinical specialties to label them for whether they had a hazard or not, and we did the same thing with the judge. And the results showed that our judge is at least on par, if not slightly better, than the real expert clinicians.

11:26

The top model, which as of a year ago when we wrote the paper was Gemini 2.5 Pro, now we've maybe updated the models, achieved an F1 score of 0.96. And even more importantly, it achieved almost perfect sensitivity. Sensitivity being a very important metric to healthcare and to clinicians, of course, because you want to make 100% sure, almost, that no hazards appear in a conversation. You would rather over-call hazards that aren't there than under-call hazards that are there. So now we have an automated judge that performs at expert level, and this is actually what makes this whole process scalable. Okay.

11:56

So now Matrix can grade thousands of conversations, but grading isn't technically improving the product. A pile of pass-fails tells you where Dora breaks and where it's not safe, but doesn't actually make the product better. So how do you do this without experimenting on the patient? The answer is that maybe very long ago in our world, eight months ago, we would manually prompt engineer this. We would look at which agents are going wrong or which prompts are going wrong, and we would have to manually prompt engineer. But we know that prompt brittleness is real, and it's quite absurd.

12:25

Formatting changes alone have been seen to swing benchmarks by 76 percentage points, and reordering few-shot examples flips a model from near-random, so near 50%, to near state-of-the-art on some benchmarks. And hand-tuning can't survive that. It's very subjective. It's not reproducible, and very importantly, it's extremely time-consuming. So over the last year or so, these prompt optimizers have started to come out, and we've focused on those. The one that we use the most is JEPA, which stands for Genetic Pareto. It comes from the same people who made DSPY, if anyone knows about them. And how does JEPA work?

13:09

You essentially define a metric for what good is, which I'll get into a bit later. Then you pass your data through JEPA, and it tells you which examples failed. Then you get a very strong LLM to reflect on the failures and update the prompt automatically. You do this over and over and over again, and it keeps what they call a Pareto frontier of the best prompts until your budget has been exhausted, and you've now come up with what JEPA considers the best prompt. So we believe this is a much better process from both a time-consuming perspective. It takes manual prompt engineering from on the order of hours to days to an hour of minutes.

13:31

Normally, between 30 minutes and an hour, you get an optimized prompt. And very importantly, it's reproducible, and there's a very clear audit trail and clear feedback loop. And if anything goes wrong, it's purely now a data science problem. It's mainly focused on the data, how to make your data robust, the feature engineering, and you define the actual metric along with the clinicians. So how do you actually know what good is? It's not a flat accuracy score. You don't want just an average of how your whole data set did. You give it a cost matrix. So let's go back to our sensitivity metric.

14:13

Let's say it's very important for clinicians to understand when and where a red flag is present. If a red flag is present and you correctly catch it, that's good. If you miss it, it could be catastrophic. If there's no actual red flag and it overcalls that there's a red flag there, it's just mildly annoying to the patient. They may need to answer a couple extra questions, but it's not a catastrophic harm situation. So what can we do? We can optimize for sensitivity. We can work with the feedback metric and we can make it give a higher reward for finding the red flags and a lower reward for missing them. So you can optimize for certain metrics.

14:51

You can also optimize for something else that a clinician might want. They might want to optimize for accuracy or might want to optimize for some other metric. All you have to do is recompile the prompt and then you've got a new optimized prompt. So remember earlier when we said we feel like the reactive loop is gone? You ship, watch, and roll back.

15:19

This is what we believe replaces it. We take real calls, real data. We then use synthetic edge cases which may not come up in real calls, such as rare symptoms or mistranscriptions. We get an optimized prompt through JEPA or some other prompt optimizer. Then we pass it through something like Matrix as a simulation safety gate. If anything fails or anything doesn't look right, we can redo that whole process, redo the data or relabel or get more data. And then after that, when you're happy with that, you can do some gated deploy, which we'll get into a bit. The most important thing here is that it's a flywheel.

15:42

Every single deployment and every new call produces more call data, so your system is consistently improving. So as we said with our Matrix framework, we use simulated patients. But however realistic you think they are, of course they are not real patients. Passing every test in simulation doesn't prove that Dora actually helps someone in real life. Things might come up in real life that you can't get in simulation. It only earns the right to actually try carefully. Simulation is the inner loop. It's fast. It's free. You can do thousands of runs before anyone real is actually exposed. But real patients are the outer loop. And that's where the only real proof is.

16:28

So simulation is necessary, but it's not sufficient. Simulation earns the right to test on real people and then eventually real patients. But you don't just flip a switch. You cross it in stages, and each stage earns the right for the next. After you've done your simulations and you're happy with the results, you might do a round of user testing. Then you get supervised clinical evaluation based on those tests, and you base it on real patients. You can do some voice actors, but of course the most realistic is to get real patients. But in this step, it's very important that there are clinicians at every step in the loop.

16:59

Then you can do some deployments, but it's still monitored. And how much autonomy you allow the system depends on your evidence. As the system gets more evidence, you can give it more independence. And underneath all of this, every call, every data set, and every pinned prompt, every judge verdict, traces back to the exact hazard that it addresses. That's the real deliverable. The important thing is that you don't ship the model. You ship the evidence when trying to regulate. So what can you take back to your own stacks, both in healthcare and other areas? You first have to define exactly what harm is for your product.

17:37

You have to manufacture your rare but dangerous cases. Don't wait for them to happen naturally. Make your evaluation metric your real cost function to optimize. Pin your prompt versions and keep the traces. These are the important things. The important thing also is that the work is never done. As you move into new modalities or new languages, there will always be new hazards that start to arise. And how much autonomy you allow the system to do depends on your evidence. As the system gets more evidence, you can give it more independence.

18:30

And underneath all of this, every call, every data set, and every pinned prompt, every judge verdict, traces back to the exact hazard that it addresses. That's the real deliverable. The important thing is that you don't ship the model. You ship the evidence when trying to regulate. So what can you take back to your own stacks, both in healthcare and other areas? You first have to define exactly what harm is for your product. You have to manufacture your rare but dangerous cases. Don't wait for them to happen naturally.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note