Open Reader

Don’t be data poor — Anuj Iravane, Anterior

completed 16:45 Aug 19, 2026 Watch on YouTube

Current Status

completed

Video ID

XAsb7MIAzm8

RAG / Chat

Enabled
Don’t be data poor — Anuj Iravane, Anterior
Description

Roughly 70% of medical communication still moves by fax. What reaches Anterior is scanned fax bundles that can run past 300 pages, carrying handwriting, checkboxes, tables and images across one patient's entire clinical trajectory. Anuj Iravane calls it an observation through a fuzzy lens over a lifespan. It is exactly the data his evals need, and the data he is least allowed to keep: their contracts rule out retaining it, deriving from it, or holding redacted or anonymized copies. Nothing survives into a dataset. In a domain where 95% accuracy is not good enough, that is a real problem. So they generate it, by running the inference workflow backwards. The forward task takes unstructured data plus a policy, follows a reasoning trace and arrives at a label. Reversed, you sample a label, sample a reasoning trace, then build the record that would have produced it. That works because Anterior already models policies explicitly as decision trees, so traces come from a far more uniform distribution than a model asked to invent variety, which tends to collapse onto the same few cases. A coarse to fine pipeline layers patient invariants into a journey of provider encounters, then fans out into documents, with a consistency eval catching contradictions between documents written in parallel. Because generation starts from the label, labels are correct by construction and ground truthing disappears. Clinicians own the pipeline as skills rather than code. Roughly 90% of their datasets are now synthetic, and in a blind review clinicians separated synthetic from real only about 60% of the time. Speaker info: - https://x.com/anujiravane - https://www.linkedin.com/in/anujiravane/ - https://www.anterior.com/ Timestamps: 0:00 - Policy guided decisions over highly unstructured data 1:05 - Most medical communication still arrives by fax 2:11 - Why 95% is not good enough 2:37 - The data you need most is the data you cannot keep 3:05 - Betting on generating it instead 3:55 - Why one s

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: For high-stakes workflows where real data is protected, scarce, or expensive to label, generate evaluation data by reversing the production decision process: sample policy outcomes and reasoning paths first, then synthesize the documents that would support them.
  • Why it matters: This is a reusable architecture for building robust eval sets and pre-deployment edge-case coverage without retaining sensitive customer data or relying on naturally occurring production examples.
  • Best use: Use it as a design reference for policy-driven agent evals: formalize the decision logic, sample scenario coverage intentionally, generate artifacts hierarchically, and let domain experts own the scenario and skill layers.

Executive Summary

Anuj Iravane describes Anterior's response to a healthcare-specific but broadly applicable data problem: the most valuable source material—unstructured patient records—is protected health information that contracts often forbid the company from retaining, anonymizing, or deriving reusable data from. Because healthcare workflows demand much more than nominal 95% accuracy and include a long tail of rare cases, small samples of real customer data cannot establish reliable coverage.

Anterior's central move is to invert its normal inference workflow. Rather than asking an LLM to one-shot generate a long medical record, it starts from a sampled policy outcome and a deterministically sampled reasoning trace from an explicit symbolic decision-tree representation of the policy. That trace becomes a diverse conditioning input for synthetic record generation, giving the team more controlled scenario coverage than an LLM's native generation distribution or a limited production sample.

The record-generation pipeline is coarse-to-fine: define stable patient attributes, create a longitudinal patient journey, plan documents for each provider encounter, generate the documents in parallel, then run refinement and consistency checks. Since the pipeline starts with the intended labels and policy logic, it can perform a round-trip validation: run the target task against the generated record and verify that its evidence and result match the originally sampled case. This produces labels "by construction" and reduces dependence on manual ground truthing.

The operational lesson is as important as the synthetic-data design. Anterior exposes both human steering and the workflow's modular skills to clinicians, who can turn production failures or anticipated cases into new test scenarios and add document-type support without engineering changes. The company says roughly 90% of its datasets are synthetic today, though used for evaluation rather than model training, and reports clinicians could distinguish synthetic from real records only about 60% of the time in a blind review.

Key Takeaways

  • Claim: Synthetic evaluation data is most useful when it is generated from the underlying policy and decision path, not by directly prompting an LLM to invent complete examples. | Evidence: Anterior reverses the forward task of applying a policy to an unstructured record: it samples a label, samples a compatible reasoning trace from the policy, and generates the medical record backward from those inputs. | Implication: For agent systems with rules, approval criteria, routing logic, or compliance requirements, define the scenario space from the control logic first so eval coverage is intentional rather than an accidental byproduct of available data. | Caveat: This approach depends on having a sufficiently explicit, structured representation of the governing policy; it is less directly applicable to tasks whose correct outcome cannot be represented as inspectable decision logic.
  • Claim: Explicit symbolic policy models solve a key diversity and edge-case-coverage problem in LLM-generated data. | Evidence: For a CPAP medical-necessity workflow, Anterior models approval/rejection conditions as decision trees, then deterministically samples different reasoning paths for a given outcome. The speaker contrasts this with a customer sample of 200 cases and a 95% evaluation score, which says little about rare cases absent from the sample. | Implication: Separate coverage testing from production-distribution estimation: oversample rare but consequential branches for robustness, while retaining a separate distribution-aware evaluation when prevalence or calibration matters. | Caveat: Uniformly sampling paths may improve test coverage but does not itself prove that the synthetic scenario frequencies match real-world prevalence.
  • Claim: Long, realistic unstructured records should be synthesized hierarchically, following the real-world process that created them. | Evidence: Instead of generating a 300-plus-page medical record in one pass, the pipeline creates patient invariants, an ordered patient journey, encounter-level document plans, and then individual documents conditioned on prior history. This mirrors how documentation is produced during provider encounters. | Implication: For complex agent test fixtures—case files, ticket histories, audit trails, customer accounts, or multi-document workflows—use a stateful coarse-to-fine generator rather than a monolithic prompt. | Caveat: Parallel document generation creates a material consistency risk across records, which requires a reconciliation stage.
  • Claim: A synthetic-data pipeline can validate its own label-data alignment through round-trip evaluation, reducing manual labeling cost. | Evidence: Because the pipeline begins with the intended labels and reasoning trace, Anterior compares those inputs with the final generated record to check that the target task reaches the same outcome and that supporting data is concordant. It also runs LLM-based cross-document consistency checks for contradictions and inaccuracies. | Implication: Build eval generators with testable invariants and independent checks; do not treat internally consistent synthetic fixtures as a substitute for periodic validation against expert-reviewed real cases. | Caveat: Round-trip checks establish consistency with the encoded policy and generator, not independent proof that the policy is complete, clinically correct, or representative of production reality.
  • Claim: Keeping generation and evaluation in a normalized text representation can be more efficient than generating visual document formats. | Evidence: Anterior generates and evaluates records as plain text/Markdown rather than rendered PDFs, arguing that capable PDF parsers already convert complex documents into Markdown and that rendered PDFs add little value for its workflows. | Implication: Use text-domain synthetic data for semantic-policy evaluation, but add rendered-document and OCR tests separately when visual extraction is part of the production failure surface. | Caveat: This assumption only holds when the target system is not materially sensitive to layout, handwriting, scans, form geometry, or OCR failure modes.
  • Claim: Domain experts should own the scenario-generation and workflow-configuration interface, rather than merely reviewing outputs produced by AI engineers. | Evidence: Anterior lets clinicians intervene at each generation stage and models patient journey generation, document generation, enrichment, and evaluation as modular skills in an internal agent harness. A clinician can add a new intake-form document skill for a customer without engineering changes. | Implication: A skills layer is a practical control-plane boundary between engineering and domain operations: experts can encode new cases and failure modes directly, while engineering maintains the harness, safety constraints, and observability. | Caveat: Clinician autonomy requires versioning, approval controls, regression tests, and auditability; otherwise domain-authored changes can silently alter evaluation standards.
  • Claim: Just-in-time synthetic datasets can shorten deployment cycles by allowing teams to simulate edge cases before customer data arrives. | Evidence: Anterior says most of its datasets are created just in time for customer deployments, that about 90% of its datasets are synthetic, and that clinicians could identify synthetic versus real records in blind review only about 60% of the time. The company currently uses synthetic data for evaluation. | Implication: The immediate business value is faster pre-production testing and failure-mode rehearsal, not a blanket claim that synthetic data can replace real-world validation or training data. | Caveat: The reported fidelity metric is limited: 60% discrimination is only modestly above chance and does not demonstrate task-level equivalence or safety for training production models.

Detailed Brief

The healthcare data constraint that motivates the design

  • Claims: Healthcare administrative workflows are policy-guided decisions applied to highly unstructured evidence rather than clean tabular inputs.; Medical-record data has unusually high variability because it represents different patients' longitudinal clinical trajectories and appears as handwriting, tables, checkboxes, key-value fields, images, and scanned fax bundles.; Retaining, reusing, redacting, anonymizing, or creating derivative datasets from protected health information may be prohibited by customer contracts.
  • Evidence: The speaker estimates that around 70% of medical communication still occurs via fax.; Anterior applies agents to health-plan workflows including prioritization, payment integrity, HEDIS measures, and medical-necessity review.; The speaker characterizes healthcare accuracy requirements as high enough that 95% is not sufficient.
  • Caveats: The talk presents a healthcare deployment perspective and does not provide comparative benchmark results against alternative synthetic-data methods or a detailed error analysis.
  • Implications: Data governance constraints can be treated as an architectural input: build persistent evaluation assets from policy and expert knowledge rather than trying to preserve prohibited source records.; The same strategy is relevant where data is ephemeral, sensitive, proprietary, or costly to label, even outside healthcare.

Why one-shot LLM record generation is rejected

  • Claims: LLMs are useful synthetic-data generators but are poorly suited to one-shot generation of long, diverse records at scale.; The speaker attributes generation homogeneity partly to limited exposure to this niche data type during pretraining and training objectives optimized for helpfulness rather than creativity or diversity.
  • Evidence: Medical records in the target setting can exceed 300 pages.; The speaker compares asking for a full record in one pass to asking an LLM to write a novel in one shot.; He cites Cynthia as a related example of sampling scenarios from a symbolic causal-state representation.
  • Caveats: The stated explanation for mode collapse is a practitioner hypothesis in the talk, not a demonstrated causal result.
  • Implications: Generation quality should be designed through decomposition, constraints, and scenario priors rather than pursued solely through more elaborate prompts or a larger generator model.

Notable Concepts & Terms

  • Reverse inference workflow: The central generation strategy: sample the desired decision outcome and policy reasoning trace, then construct the evidence record backward from them.
  • Symbolic policy representation: A decision-tree-like formalization of policy conditions that supports deterministic reasoning-path sampling, consistency, and coverage control.
  • Reasoning trace: The sampled chain of policy conditions that explains how a chosen label should be reached; it conditions record generation.
  • Coarse-to-fine generation: A hierarchical pipeline from patient invariants to journey, encounter plans, and documents, designed for long-context scalability and realism.
  • Round-trip check: A validation that the generated record, when processed by the target task, supports the initially sampled label and reasoning inputs.
  • Skills-based workflow: Modular agent-harness components for generation, enrichment, and evaluation that domain experts can configure or extend.
  • PHI: Protected health information; the retention restrictions on this data are the primary reason Anterior needs a synthetic alternative.
  • HEDIS: A health-plan performance-measurement domain cited as one of Anterior's production administrative workflows.

Operator Notes / Why Ken Should Care

  • For each policy-driven agent, create a machine-readable decision graph with outcomes, required evidence, branch conditions, and known exception paths before expanding the eval corpus.
  • Add a scenario sampler that can deliberately oversample high-cost, low-frequency branches; maintain a distinct real-distribution suite for calibration and operational forecasting.
  • Implement fixture generation as modular skills with versioned inputs, generated artifacts, evaluator outputs, and regression baselines so domain-authored changes are reviewable and reversible.
  • Require two validation layers for generated cases: cross-artifact consistency checks and an independent execution of the target workflow against the expected outcome.
  • Do not eliminate real-data validation: maintain an expert-reviewed holdout process, especially where source-document layout, OCR, or unmodeled policy ambiguity affects production behavior.
  • Assess whether a text-only fixture strategy is sufficient for Ken's systems; if ingestion depends on screenshots, PDFs, forms, or browser state, add a representation-level test suite rather than assuming parsers remove that risk.

Source/Metadata

  • Title: Don’t be data poor — Anuj Iravane, Anterior
  • Transcript words: 4599
  • Duration seconds: 1005
  • Timestamp note: No timestamps or chapters were present in the supplied transcript. The latter portion substantially repeats earlier content.

Transcript

2689 words en Processed in 89.2s

Hello everyone, welcome to Don't Be Data Poor. My name is Anuj. I lead AI at Interior. Just a bit about Interior: we are a clinician-led AI company built for health plans, backed by Sequoia and NEA. What we do is run AI transformations for health plans, as part of which we build agents for several high-stakes healthcare administrative workflows in production. Things like prioritization, payment integrity, HEDIS measures, etc. It's okay if you're not familiar with any of these workflows, because a lot of the work that we do can actually be summarized in the same way. It's policy-guided decision-making over highly unstructured data. The unstructured data looks something like this, right? You have these scanned fax bundles containing medical records full of patient information. A not-so-fun fact is that I think around 70% of medical communication still happens via fax. Fortunately or unfortunately, this is the data that we end up working with the most. It is very rich and information-dense data that we see here. The data distribution here comes from a very long tail of rare cases with nuanced scenarios. It models an entire clinical trajectory for a patient, and every single person's journey is very different. It also presents itself in varied formats. You have things like bad handwriting, tables, checkboxes, key-value pairs, images, a lot of tough data to deal with. But I personally think it's a very fascinating source of data that we see here. It's like an observation through a very fuzzy lens over an entire person's lifespan. It's really unique. I'm sure you must have heard this enough times today already, but in healthcare, the baselines for accuracy are exceptionally high. Ninety-five percent is not good enough. At Interior, this is why we invest very deeply in data sets and evals. These unstructured medical records are a stable source of data for these evals, and we work with this kind of data in almost every workflow that we try to automate. But the problem is we can't really keep this data. It's PHI. It's highly protected. We can't retain it, we can't reuse it, and we can't even derive information from it. Most of our contracts prohibit us from doing anything like that. Even things like redacting it, anonymizing it, and keeping derivative copies are a strict no-no, completely off the table. So nothing really survives in any sort of data set that we want to persist over a period of time. So what this talk is about is: what do you do when the data set you most need is also the data you're least allowed to keep? The answer that we put our bets on is that we can synthetically generate this data ourselves. There's been a lot of focus on synthetic data recently. You have Frontier Labs striving to generate synthetic data for continued pre-training, for RL, for computer use, for agents. So it's a hot topic, and it's a hot topic on our minds as well. The moment you say generate, the first thing that comes to mind is, okay, can we try to use an LLM to generate synthetic data? I think you can. I personally believe LLMs are a fantastic tool to generate synthetic data, and several teams have already demonstrated this. There have been some papers in the healthcare space and outside the healthcare space. People have successfully used LLMs to generate synthetic data for different purposes. There are some known challenges in trying to use these LLMs to create data, especially if you're trying to one-shot the whole process. It's really hard to generate diverse, realistic-looking synthetic records, and this is even more of a problem when you're trying to do this at scale. Oftentimes, these medical records are over 300 pages long. It's like imagining that you wouldn't ask an LLM to write a novel for you in one shot, right? It's the same reason why you wouldn't use an LLM to just one-shot a synthetic record for you. LLMs seem to suffer from this very strange mode collapse problem when it comes to generating diverse data, creative data. I think there are two main reasons for it. The first one is, as Aish mentioned in stock earlier, there's very little exposure to this data source in the pre-training data corpus. Today's objectives for pre-training and post-training are largely not incentivized for creativity or diversity, really. They're incentivized to be helpful systems. So with these challenges in mind, I'll walk you through one of our approaches in how we managed to build a pipeline to generate synthetic data. Earlier, I mentioned our forward tasks look something like this, right? You have workflows and tasks that start with some unstructured data and a policy, and you execute your policy against that data. You follow this reasoning trace through it, and you arrive at some sort of an outcome, which is your label. So this is our forward task. The idea we had was to try and reverse this process. Can we actually start by sampling a random label, figuring out a reasoning trace for that label, and then trying to generate data backwards from that? The idea here is that if you can actually sample these two things with enough diversity, we will be able to generate data that's conditioned on a diverse set of inputs, allowing us to circumvent the diversity problem a little bit. So just a quick aside on policies. We've talked about policies a bit, but let me just clarify what these really mean, right? This is an example policy we have for a CPAP device for patients. This particular one is for a medical necessity review workflow, and it outlines all these diverse sets of conditions that a patient might have in which a CPAP device should be approved or rejected. In this policy, as well as many other policies, you can think of these as essentially decision trees that outline all these sorts of conditions that dictate how some outcomes are met. At NTIR, we actually spend a lot of time and energy trying to model these policies explicitly as decision trees. We work with symbolic representations similar to decision trees, and it helps us achieve a better accuracy and consistency score when executing them in an LLM-based workflow. The reason why I'm bringing this up is that by having this sort of symbolic representation of a policy, you actually have a way to deterministically sample different reasoning traces for a given outcome. So back to the idea of reversing the process, right? This sampling of reasoning traces from the policies is what helps us get that diverse conditioning input to then generate medical records from. The key idea here is that the distribution we sample from is a much more uniform and effective prior distribution than what you'd normally get from an LLM. One added benefit of sampling this way is that, in theory, you're able to test for far more scenarios than you would likely get from production data sources. So what I mean by that is, say you get a sample of 200 cases from your customer, and you try to have an eval that measures performance against that, and you get a 95% score. It doesn't really tell you what your performance would be in those rare edge cases that are not in that data set. There will always be rare edge cases that are outside that distribution just because of the fact that our data is so highly variant. So for those familiar with Cynthia, they follow a similar pattern of sampling scenarios from a symbolic causal state representation. There are a few folks in the space who are working with these symbolic representations to generate diversity in synthetic data generation. So let me walk you through the rest of the pipeline. Once we have this diverse set of samples as a conditioning input, what we did was build an LLM-based pipeline that follows a coarse-to-fine pattern to progressively build up a medical record layer by layer. Here we first start with creating some patient invariants like biological sex, birth date, blood group. We use that, along with the reasoning trace, with an LLM again to produce an ordered list of events and provider encounters that a patient might have had, and we call this the patient journey. This is a high-level overview of what a patient might have gone through in their lifespan, captured by a list of events in natural language. In the real world, it is actually only during these provider encounters that documentation is really generated, at least for the data that we get. Most of our source data is generated during these provider encounters. So we model exactly that in our pipeline. We first generate a document plan for each encounter, and then based on that and the preceding history of the patient, we fan out into generating the actual documents to hydrate them with actual synthetic information. This coarse-to-fine layering is what allows us to keep the different prompt payloads in the pipeline very token-efficient from both input and output perspectives. This also helps enable scaling across longer patient journeys. You can scale this pipeline, have a much longer patient journey, and just fan out and generate documents that way without overloading the context windows of your LLMs. Finally, we have this sort of refinement loop at the end that uses a set of evals to provide feedback to improve specific parts of the generated documents. For example, one of the evals we have is an LLM-based check for consistency between all documents. This makes sure that there are no contradictions, inaccuracies, or conflicting information between two generated documents. This is important because we have a parallel fan-out process that is used to generate these documents independently. Because we started with the labels for this particular pipeline run, what we also have is an ability to use those labels and compare those against the generated medical records to see if the tasks that we originally used actually match and the data is in concordance with the task inputs and outputs. So we can do this sort of round-trip check to ensure that our data is actually in sync and, by default, get correct labels by construction. In theory, this is a really nice property to have. You can basically skip the expensive ground-truthing process you need for your data. One thing to clarify here is that so far, all the generation has been happening just in plain text and markdown text. It is possible to go from that to a rendered PDF, but we don't really see much value in doing that because we have state-of-the-art PDF parsers today. They are available to everyone, and they allow you to convert any sort of complex PDF into a nice markdown representation. So all of this synthetic generation and evaluation happens in the text domain. This is just an example of a pipeline that we created from scratch, and it's very easy to build. It's largely fully LLM-based. But who came up with this, right? Who am I to know anything about what a good medical record looks like? So how do we know if this is any good? I think this has been mentioned a few times today already, but you really don't. No AI engineer ever would. You want your domain experts to be the ones telling you what's good and what's not good, which is why we believe that it is of great value to empower your domain experts to own your whole data pipeline. Specifically, we do this in two ways, right? We enable our clinicians to interject at each point in the generation process with a human-in-the-loop mechanism. At any point, a clinician can steer the generation process to make a medical record the way they want it. We often see our clinicians use this to first look at cases that happen in production, get some interesting ideas, and then use those ideas along with the steering in this pipeline to make cases that look similar to what we might see in production or what they have seen in production. This is what makes the data generated from this really useful, right? You can actually model your failure cases beforehand or even after you see them in production. Secondly, and I think most importantly, we let our clinicians also own the whole logic of the pipeline. We do this by modeling the whole pipeline as a skills-based workflow running on a genetic agent harness that we built internally. Every section you see here, all the way from the patient journey to the document generation to the document enrichment to the evals, all of these things are skills that run on our agent harness. As an example, if a clinician wanted to add support for a new document type, let's say for a new customer, and they wanted their intake forms to look a certain way, they could easily just make a new skill file for it, attach it to the pipeline, and voila, there wouldn't be any engineering changes required. So it's completely clinician-owned from that perspective. Just as an aside generally, I feel like skills are really an amazing interface between AI engineers and domain experts, especially in vertical AI. We see this being modeled in several of our other workflows, both for internal use cases and in production as well. So some results from this. Even though we only really use synthetic data for evaluation at the moment, there are already a lot of merits that we get from it. Roughly 90% of our data sets are already made of synthetic data. This helps us maintain a very high production accuracy score across many customer deployments. The pipelines that I just showed you are already able to achieve very high fidelity on this generated data. In a blind review, clinicians were only able to distinguish synthetic from real about 60% of the time. So there is room for improvement, but it's close, and it's quite a promising avenue for us to invest more here. The fact that is most interesting to me, and what I'm really excited about, is that most of our data sets today are created just in time for these customer deployments, right? When you have the ability to create data from scratch so quickly, you don't need to depend on waiting for data from your customer. You can just model all your edge cases, simulate them, and test your workflows before you go live in production. So some takeaways. If you're looking to build your own synthetic data pipeline in healthcare or even another domain, try reversing your inference workflow. Diversity should always be sampled from an appropriate distribution for your use case. Try to emulate the process in which the data was actually generated. Like I showed you, we were using LLMs while trying to emulate how our medical records might actually be generated during patient encounters. I would highly recommend you try doing that. The fourth, most important thing, I think, is when you're making a data pipeline like this, it's really important to give your domain experts the keys, because these are the people who know about your data, and they will help you drive toward recursive self-improvement, not the AI engineers. Cool. So you don't need a PHI problem for this. Anywhere the data you need is ephemeral, sensitive, or even expensive to label, you can think about generating data yourself, and hopefully you won't be data poor. Thank you, everyone. Thank you. Thank you. Thank you. Thank you. Thank you. You follow this reasoning trace through it. And you arrive at some sort of an outcome, which is your label. So this is our forward task. And the idea we had was to try and reverse this process. Can we actually start by sampling a random label, figuring out a reasoning trace for that label, and then trying to generate data backwards from that? The idea here being that if you can actually sample these two things with enough diversity, we will be able to generate data that's conditioned on diverse set of inputs, allowing us to kind of circumvent the diversity problem a little bit. So just a quick aside on policies. We've talked about policies a bit, but let me just clarify what these really mean, right? So this is an example policy we have for a CPAP device for patients. This particular one is for a medical necessity review workflow. And it sort of outlines all these diverse set of conditions that a patient might have in which a CPAP device should be approved or rejected. So in this policy, as well as many other policies, you can think of these as essentially decision trees that outline all these sorts of conditions that dictate how some outcomes are met. And at NTIR, actually, we spend a lot of time and energy in trying to model these policies explicitly as decision trees. We work with symbolic representation similar to decision trees, and it helps us achieve a better accuracy and consistency score when executing them in LLM-based workflow. And the reason why I'm bringing this up is that by having this sort of symbolic representation of a policy, you actually have a way to kind of deterministically sample different reasoning traces for a given outcome. So back to the idea of, like, reversing the process, right? This sampling of reasoning traces from the policies is what helps us get that diverse conditioning input to then generate medical records from. And the key idea here is that the distribution here that we sample from is a much more uniform and effective prior distribution than what you'd normally get from an LLM. One added benefit of sampling this way is that, in theory, you're able to test for far more scenarios than you would likely get from production data sources. So what I mean by that is, like, say you get a sample of 200 cases from your customer, and you try to, like, have an eval that measures performance against that, and you get a 95% score. It doesn't really tell you about what your performance would be in those rare edge cases that are not in that data set. There will always be rare edge cases that are outside that distribution just because of the fact that our data is so highly variant. So for those family with Cynthia, like, they follow a similar pattern of sampling scenarios from a symbolic causal state representation. There's a few of the folks in the space who are working with these symbolic representations to generate diversity in synthetic data generation. So let me walk you through the rest of the pipeline. All right? So once we have this diverse set of samples as a conditioning input, what we did was we built an LLM-based pipeline that follows a course-to-find pattern to progressively build up a medical record layer by layer. So here we first start with creating some patient invariants like the biological sex, the birth date, the blood group. We use that, along with the reasoning trace, with an LLM again, to produce an ordered list of events and provider encounters that a patient might have had, and we call this the patient journey. So this is a high-level, you can think of it as a high-level overview of what a patient might have gone through in their lifespan, captured by a list of events in natural language. And in the real world, it is actually only during these encounters, provider encounters, that documentation is really generated, at least for the data that we get. Most of our data source data is generated during these provider encounters. So we model exactly that in our pipeline. We first generated a document plan for each encounter, and then based on that and the preceding history of the patient, we fan out into generating the actual documents to hydrate them with actual synthetic information. And this course-to-find layering is actually what allows us to keep the different prompt payloads in the pipeline very token-efficient from both input and output perspective. This also helps us enable to scale across longer patient journeys. So you can scale this pipeline, you can have a much longer patient journey, and you can just fan out and generate documents that way without overloading the context windows of your LLMs. Finally, we have this sort of refinement loop in the end that uses a set of evals to provide feedback to improve specific parts of the generated documents. For example, one of the evals we have is an LLM-based check for consistency between all documents. So this makes sure that there's no contradictions or inaccuracies or conflicting information between two documents that are generated. And this is important because we have a parallel fan-out process that is used to generate these documents independently. And because we started with the labels for this particular pipeline run, what we actually also have is an ability to kind of use those labels and compare those against the generated medical records to see if the tasks that we originally used actually matches and the data is in concordance with the task inputs and outputs. So we can do this sort of round-trip check to ensure that our data is actually in sync and by default get correct labels by construction. So in theory, this is a really nice property to have. You can basically skip the expensive ground-truthing process you need for your data. One thing to clarify here is that so far, all the generation has been happening just in plain text and markdown text. It is possible to go from that to a rendered PDF, but we don't really see much value in doing that because we have state-of-the-art PDF parses today. They are available to everyone, and they just allow you to convert any sort of complex PDF into a nice smart-down representation. So all of this synthetic generation and evaluation happens in the text domain. So this is just an example of a pipeline that we created from scratch, and it's very easy to build. It's largely fully LLM-based. But who came up with this, right? Like, who am I to know anything about what a good medical record looks like? So how do we know if this is any good? And I think this has been mentioned a few times today already, but, like, you really don't. Like, no AI engineer would ever would. Like, you want your domain experts to be the ones telling you what's good, what's not good, which is why we believe that it is of great value to empower your domain experts to own your whole data pipeline. And specifically, we do this in two ways, right? We enable our clinicians to kind of interject at each point in the generation process with a human-in-the-loop mechanism. So at any point, a clinician can steer the generation process to make a medical record in the way they want it. We often see our clinicians use this to first look at cases that happen in production, get some interesting ideas, and then use those ideas along with the steering in this pipeline to make cases that look similar to what we might see in production or they have seen in production. And this is what makes the data generated from this really useful, right? Like, you can actually model your failure cases beforehand or even after you see them in production. And secondly, I think most importantly, we let our clinicians also own the whole logic of the pipeline. We do this by modeling the whole pipeline as a skills-based workflow running on a genetic agent harness that we build internally. So every kind of section here you see, all the way from the patient journey to the document generation to the document enrichment to the evals, all of these things are skills that run on our agent harness. As an example, if a clinician wanted to, say, maybe add support for a new document type, let's say for a new customer, they wanted their intake forms to look a certain way, they could easily just make a new skill file for it, attach it to the pipeline, and voila, there wouldn't be any engineering changes required. So it's completely clinician-owned from that perspective. And just an aside generally, I feel like skills are really an amazing interface between AI engineers and domain experts, especially in vertical AI. We see this being modeled in several of our other workflows, both for internal use cases and in production as well. So some results from this, right? So even though we only really use synthetic data for evaluation at the moment, there's already a lot of merits that we get from it. Roughly 90% of our data sets are already made of synthetic data. This helps us maintain a very high production accuracy score across many customer deployments. The pipelines that I just showed you, already we're able to achieve a very high fidelity on this generated data. In a blind review, clinicians were only able to distinguish synthetic from real about 60% of the time. So room for improvement, but it's close. And it's a quite promising avenue for us to invest more here. And the fact that is the most interesting to me and what I'm really excited about is that all of these data sets, well, most of our data sets today then are created just in time for these customer deployments, right? When you have the ability to create data from scratch so quickly, you don't need to depend on waiting for data from your customer. You can kind of just model all your edge cases, simulate them, and test your workflows before you go live in production. So some takeaways, if you're looking to build your own synthetic data pipeline in healthcare or even another domain, try reversing your inference workflow. Diversity should always be sampled from an appropriate distribution for your use case. Try to emulate the process in which the data was actually generated. So like I showed you, we were trying to sort of like, we were using LLMs where we were trying to emulate how our medical records might actually be generated during patient encounters. So I would highly recommend you try doing that. And the fourth most important thing, I think, is when you're making a data pipeline like this, it's really important to give your domain experts the keys because these are the people who know about your data and they will help you drive towards a recursive self-improvement and not the AI engineers. Cool. So you don't need a PHI problem for this. Anywhere the data you need is ephemeral, sensitive, or even expensive to label, you can think about generating data yourself and hopefully you won't be data poor. Thank you, everyone. Thank you. Thank you. Thank you. Thank you. Thank you.