Verifiable Environments for AI in Biology — Kenny Workman, LatchBio
Description
A single spatial biology run can yield two to six terabytes of data, far more than a scientist can eyeball, and Kenny Workman argues that this is the raw material for teaching AI to actually do science. LatchBio, five years deep in pharma, treats experimental biology as a verifiable substrate: lay a chunk of a tumor over a sequencing surface and you get a giant matrix of numbers whose analysis has a right answer, which is exactly what you need to benchmark and improve a model. They adapted coding models into biology tools and built benchmarks like sequencing based spatial analysis, and found what everyone in post training now knows, that frontier models cannot yet be trusted with this and that measurement is what drives progress. The reason biology is hard is that it is messy and the field rarely agrees on the answer, so much of the work is designing tasks where reasoning, not memorized knowledge, is what gets rewarded. Workman walks through trajectory data from real scientists, tasks like finding the part of a tumor that matters, and why an ambiguous prompt quietly makes a benchmark uninformative. He also gets candid about biosecurity, where model refusals are their own evaluation problem and red team tasks have to be handled carefully, and frames the whole effort as a flywheel: better benchmarks, better tools, more of the program landscape indexed, repeat. Speaker info: - https://x.com/kenbwork - https://www.linkedin.com/in/kennyworkman - https://kenbw.com/ Timestamps: 0:00 - Terabytes of experimental data 1:41 - Decomposing a new paper into tasks 2:44 - A verifiable substrate for science 3:23 - Five years in pharma 4:13 - Coding models as biology tools 5:40 - Why frontier models can't be trusted yet 6:44 - Sequencing based spatial analysis 9:16 - Reasoning, not memorized knowledge 10:07 - Trajectory data from real scientists 12:14 - Why biology tasks are messy 15:36 - Biosecurity and refusals
Summary
Generated by gpt-5.6-terraAt-a-Glance
- Verdict: Watch fully
- Core thesis: Agentic biology will progress through code-like, data-interacting evaluation environments, but reliable benchmarks must combine deterministic verification, path-invariant task design, and expert human review because scientific workflows rarely have a single canonical answer.
- Why it matters: This is a concrete case study in building a vertical-agent control loop: deploy tools into real customer workflows, convert recurring work into benchmarks, use evaluations to drive post-training and model selection, and continuously repair benchmark assumptions from trajectory evidence.
- Best use: Use it as a design reference for evaluation-first agent products in domains where jobs are long-running, outputs are ambiguous, and correctness cannot be reduced to a simple final-answer check.
Executive Summary
Kenny Workman argues that biology is becoming a promising substrate for agent systems because modern experiments generate enormous, structured datasets and already follow an executable workflow: choose a biological model, generate data, process it, interpret it against prior knowledge, and make a claim. He likens analysis code and its surrounding software/data infrastructure to the verifiable substrate that made software agents trainable through benchmarks such as SWE-bench.
LatchBio arrived at this view through deployment. Originally a data-platform vendor, it moved toward white-labeled analysis products for experimental-kit manufacturers, then used its infrastructure as tools for agents that can inspect datasets, build dashboards, and dispatch external computation. Early coding models showed signs of usefulness, but could not be trusted for real scientific work because biology requires more than code generation: it requires data analysis and domain-specific scientific reasoning.
Its response was SpatialBench, initially 146 tasks across spatial-biology kits and analysis operations. The key benchmark-design lesson is that a task must be functionally verifiable, durable across legitimate analysis paths, and require actual data interaction rather than recall. Human review later revealed that many apparently objective tasks hid consequential choices—normalization, gene-list definitions, spatial radius, and arbitrary QC thresholds—so Latch produced a human-verified subset rather than treating its original grader specifications as ground truth.
For longer-horizon work, Latch created SpatialBench-Long to approximate paper-level analyses and industry drug-program decisions. Final outcomes alone are too sparse and uninformative for these workflows, so it is exploring rubric-based intermediate choke points that should hold across multiple valid solution paths. The talk closes with a product-and-research flywheel: customer deployments reveal useful tasks, benchmarks encourage frontier labs to improve models, and those improvements feed back into Latch's deployed products—now expanding into other omics, preclinical pharmacology, and biosecurity evaluation.
Key Takeaways
- Claim: Biological data analysis can provide a verifiable training and evaluation substrate for scientific agents, analogous to how executable code enabled software-agent benchmarking. | Evidence: Modern biology commonly follows a repeatable chain—biological model, experiment, data processing, interpretation, claim—while single-cell runs can generate 2-6 TB, spatial runs about 7 TB, and proteomics runs hundreds of GB. Latch frames the analysis DAG, code, and data infrastructure as the executable layer agents can operate against. | Implication: For agent domains with messy real-world reasoning, look for measurable intermediate artifacts and executable workflow states rather than waiting for a universally agreed final answer. | Caveat: Unlike code, biology often lacks a uniquely correct answer; an executable workflow alone does not make the ultimate scientific conclusion objectively verifiable.
- Claim: Real deployment revealed that generic coding models can begin to operate biological-analysis workflows, but they remain unreliable without focused post-training and domain-specific evaluation. | Evidence: Latch connected models to tools that ingest large experimental datasets, create dashboards, and dispatch external compute; scientists iterated with agents on questions such as differential gene expression between malignant and non-malignant tissue. Workman says prototypes were initially 'pretty bad' but showed early signs of working. | Implication: The product pattern is not a biology chatbot: it is a long-running, tool-using agent harness with explicit compute dispatch, iterative scientist oversight, and task-specific model selection. | Caveat: The relevant tools may run for days or weeks, unlike typical short software-agent tool calls, which raises orchestration, state-management, and evaluation-horizon requirements.
- Claim: A useful scientific-agent benchmark must be verifiable, durable across valid analysis paths, and resistant to answer-by-memory behavior. | Evidence: SpatialBench used a task prompt, input data nodes such as numeric matrices and high-content images, grader configuration, and a deterministic Python-like grader. Latch initially built 146 problems spanning spatial-biology kits and workflow tasks, borrowing the general structure of SWE-bench. | Implication: When designing agent evals, test for invariants across multiple legitimate trajectories; do not equate a fixed reference implementation with task correctness. | Caveat: A grader can falsely fail a scientifically valid solution if the benchmark author encodes one preferred pipeline or overly narrow expected output.
- Claim: Expert human verification is necessary to discover hidden ambiguity and invalid assumptions in scientific benchmarks. | Evidence: After observing trajectories from model releases between roughly January and March, Latch found many of its assumptions wrong. Tasks involving microglial activation, oligodendrocyte inflammation, neighborhood radii, normalization, and Spearman correlations concealed open methodological choices. Two rounds of scientist attempts produced a published verified subset. | Implication: Use human experts not just as final graders but as benchmark-red-teamers who expose underspecified prompts, arbitrary thresholds, and legitimate alternative approaches before relying on automated scores. | Caveat: Human review is a proxy rather than absolute ground truth, but Workman presents peer assessment by scientists as the best available method when canonical answers do not exist.
- Claim: Long-horizon scientific work needs process-sensitive evaluation because end-state rewards provide too little signal. | Evidence: SpatialBench-Long aims to simulate results sections of papers and industry go/no-go decisions; each task took a group of three people about a week to construct. One example asks an agent to reconstruct a metastatic niche by connecting primary-tumor and metastatic-lesion genetic/mRNA evidence to identify the likely seeding population. Workman says no models solve this task yet. | Implication: For complex workflows, instrument trajectory milestones that are invariant across solution paths, but validate that those process metrics predict actual outcomes before optimizing models against them. | Caveat: Latch's rubric/choke-point scores are associated with final verifiable outcomes but are only loosely numerically correlated, so the team is not yet confident using them for RL or definitive benchmarking.
- Claim: Latch is building a deployment-to-benchmark-to-model-improvement flywheel and expanding it beyond spatial biology into broader drug-discovery and biosecurity domains. | Evidence: The company works with kit manufacturers and customers, uses their workflows to identify benchmark targets, and reports that its benchmarks have appeared in recent Anthropic model cards and Claude Science materials. It has expanded to single-cell epigenomics, long-horizon tasks, preclinical pharmacology for small molecules, and a biosecurity team following an acquisition. | Implication: The strategic moat in vertical agents may come less from a general-purpose model and more from proprietary task environments, evaluation data, customer workflow access, and feedback loops that improve both products and benchmarks. | Caveat: The talk presents Latch's own positioning and adoption claims; it does not provide comparative benchmark scores, customer outcome metrics, or independent evidence of downstream scientific impact.
- Claim: Biology safety behavior needs evaluation that distinguishes ordinary scientific work from requests that are superficially benign but operationally harmful. | Evidence: Latch describes routine scientist tasks alongside red-team tasks that may claim to clone GFP but actually encode a toxin or could help bootstrap a virus. Workman also notes that models can over-refuse on benign questions, such as basic mitochondrial queries, and calls this fundamentally an evaluation problem. | Implication: Safety systems should be evaluated on both utility and adversarially transformed requests; coarse refusal policies create costly false positives without necessarily demonstrating robust protection. | Caveat: The transcript's statement comparing routine and red-team task results is garbled, so it does not establish a reliable quantitative direction or rate for either over-refusal or unsafe compliance.
Detailed Brief
Why spatial biology is a demanding first environment
- Claims: Spatial biology combines high-dimensional molecular measurements with image-like spatial context, making it suitable for evaluating whether an agent can reason from real experimental data rather than answer textbook questions.; The analysis pipeline varies materially by measurement technology, tissue type, and disease context, with limited field-wide consensus for individual processing steps.
- Evidence: In sequencing-based spatial assays, a tissue section is placed on DNA-bead slides that capture RNA; downstream inputs include a large numerical matrix and a high-content image.; Workman characterizes the technology landscape as spanning chemistry, optics, semiconductors, and physics, with each capture method built from decades of cumulative measurement advances.
- Caveats: High workflow variability makes fixed benchmark expectations especially vulnerable to rewarding convention rather than scientific validity.
- Implications: Heterogeneous data modalities are a useful stress test for agents because they require the system to select and justify an analysis route rather than execute a single templated script.
Benchmark expansion and ecosystem positioning
- Claims: Latch is attempting to map drug discovery as a collection of evaluable subproblems rather than treating it as one monolithic agent benchmark.; Its stated approach is to stratify work by therapeutic modality and experiment type from discovery through development and translation.
- Evidence: The company says it released an initial preclinical-pharmacology benchmark for small molecules and is aggregating preprints, evaluations, and model trajectories as a resource.; It also announced a collaboration with American Wetware and a surveillance company called Ackwood.
- Caveats: The presentation gives limited technical detail on the new pharmacology, surveillance, and biosecurity evaluations, so their task design and validity cannot be assessed from this talk.
- Implications: A broad biology-agent platform will likely need a portfolio of narrow, modality-specific environments instead of a single aggregate capability score.
Notable Concepts & Terms
- SpatialBench: Latch's initial benchmark for spatial-biology analysis agents; it uses real data inputs, task prompts, and deterministic grading modeled in part on SWE-bench.
- SpatialBench-Long: A longer-horizon benchmark intended to resemble full paper-result analyses or practical drug-program decisions rather than isolated pipeline steps.
- Analysis DAG: The directed workflow of data-processing states and operations; Latch uses its discrete components to make otherwise open-ended scientific work more evaluable.
- Durability: A benchmark property's requirement that grading remain correct across different scientifically valid analysis paths, not merely one benchmark author's preferred pipeline.
- Choke-point rubrics: Intermediate evaluation criteria based on path-invariant nodes in a workflow tree, proposed to add signal when final long-horizon rewards are sparse.
- Sequencing-based spatial biology: A spatial transcriptomics approach in which tissue RNA is captured on DNA-bead arrays, producing molecular matrices plus imaging data that need multi-step analysis.
- Human verification: Peer assessment by scientists used to identify ambiguity, arbitrary thresholds, and invalid benchmark assumptions where no canonical scientific answer exists.
- Routine versus red-team biosecurity tasks: An evaluation framing that tests both normal scientific utility and requests that appear benign but contain harmful biological intent or structure.
Operator Notes / Why Ken Should Care
- Adopt a benchmark-authoring review gate for any high-stakes agent workflow: require at least two independent expert attempts before freezing automated graders.
- For long-running agents, log and score checkpoint artifacts—not just final outputs—and measure whether checkpoint quality predicts successful end results before using it as an optimization target.
- Treat arbitrary domain thresholds, defaults, and 'appropriate' parameter language as evaluation risks; force authors either to specify decision rules or grade over valid ranges and alternative methods.
- Separate safety evaluation into benign-utility, obvious-risk, and deceptively framed-risk suites; monitor false refusals independently from unsafe-compliance rates.
- Assess vertical-agent opportunities partly by whether the business can obtain proprietary workflow traces, data environments, and expert adjudication loops—not only by access to a frontier model.
Source/Metadata
- Title: Verifiable Environments for AI in Biology — Kenny Workman, LatchBio
- Transcript words: 5599
- Duration seconds: 1062
- Timestamp note: No timestamps or chapter markers were present. The latter portion of the supplied transcript substantially repeats earlier material.
Transcript
Thank you to the organizers for having me. I'm one of the co-founders and CTO at Latch. We are a vertical AI lab for benchmark and agent engineering, hoping to motivate and explain exactly what that means today. Starting directly with motivation for agents in bio generally, many people in my domain are familiar with this curve, but this is the log-linear curve of data generated over the years in biology. And the reason I'm bringing it up will become directly important to the kinds of things we want to do in engineering. This curve is driven by a very small handful of experimental classes. One is called single-cell biology. This is where we split up cells, break them apart, and measure their RNA. The second is spatial biology, which will become the focus in the next segment of the talk. Same thing as single-cell, but you get spatial resolution. You can look at how RNA is spread out geometrically over a tissue. And the third thing is proteomics. It's a broad category of different techniques. They measure proteins. Less abundant in ordering, less data volume generated relative to the other two, but still important. So you guys are technical, and I always think it's good to ground things somewhat quantitatively, but these are really big numbers, and the experimental data from these techniques is growing quite rapidly, almost greater than any other domain of science other than particle collider machines. Single-cell experiments can yield two to six terabytes per run. Spatial runs can yield seven terabytes per run, proteomics a few hundred gigs. And the only reason I bring this up is to say, hey, the output of a single experiment can exceed what a scientist can safely store on a consumer laptop in many cases. And the law driving how the molecular capture works points to rapid gains in this throughput over the coming years. One thing I like to do when I read a new paper is decompose it and align it to this framework, because it will become important in a second. Modern biology research is centered around those experiments. You choose a model, a biological model, not the kind of models you guys are used to. You generate data from that model, you process the data, you creatively think about the results in the context of prior literature, and you make a claim. Almost all modern experiment papers that you see published follow this loose structure, with a lot of nuance. All that to say is they become something of a panning experiment. You're looking for signal using measurement in a sea of noise. And so this is building up to the claim that code and SWE data analysis scaffolds agentic biology. It becomes this executable substrate that we can use to train things. It induces a natural way to benchmark and climb capability. I've written about this a lot at this blog, link here. There's a lot more depth to this claim, so I wouldn't take it at face value, but it's something to look into. All you can take away from this is, just like code provided the verifiable substrate for complex software tasks that are not inherently verifiable, data analysis might do the same thing in bio. So how do we get started? We were originally a data tool vendor for biotech and pharma. Where we started five years ago out of Berkeley, I'm 25. We started when I was 20. We stored, transformed file data from large experiments as a service. We tried to build products, explored lots of things. Over the last two years, we started moving away from biotech and pharma and more toward the people who build those kits I was talking about. We packaged the software into white-labeled things that they provided the scientists themselves, helping them analyze their data. And then, over time, there became this strong interaction with the agents, using the infrastructure components as tools in the loop context you guys are familiar with. Except in our domain, the tools can take days or weeks. I'm serious. What started to happen around last summer is agent prototypes started to work. So we took coding models. To our knowledge at the time, they were not seriously post-trained on any tasks in biology to this point. And we started to build products that look a lot like all the other agent products. They have a chat interface for you to ask questions to, and they build dashboards and dispatch operations to external compute. The kinds of things that this agent would do is take large file data from the types of experiments I was talking about earlier, say a tissue biopsy from cancer with a spatial measurement. And then the scientist is just iterating with it to get at some question they have. Maybe between a malignant and non-malignant part of the tissue, what kind of genes are being overexpressed. But what was fascinating is, even though it was pretty bad, it showed the early signs of working. And it became clear to us at this time that agentic biology might look a lot like code. I actually lifted this slide from Anthropic's Claude Science announcement yesterday, but you had this faulty, silly engineer that became better. And then, as it improved in capability, you could dispatch work to teams of them that work together. The same pattern will probably emerge in science. And products and harnesses will emerge to orchestrate work and abstract it so teams of teams of agentic scientists can take on capability. But we needed focused post-training, because at the time, and still now, frontier models cannot be trusted to do real work. They're missing some capability between knowing biology and writing code. And this is exactly extracting scientific insight from real-world data. Unlike code, which is one constituent component of this work, it also involves data analysis and domain reasoning, scientific reasoning. We thought spatial biology is a good place to start. So we started building agents. This is a technical greenfield. We had many existing customers. It's also just a beautiful example of measurement drives progress. You can actually see biological phenomena play out. You can look at a developing mouse embryo. And we had to get in the guts of how the data was captured and analyzed to build good agents. I'm not going to get into this in detail, but I put up this tree of different capture technologies in spatial biology to highlight the diversity of things that exist. They really span advances in chemistry, optics, semiconductors, physics. Each branch is induced from decades of cumulative work to figure out how to measure a type of molecule. As a specific example, one technique we work with is called sequencing-based spatial, and it's where you take a slide of little beads with clumps of DNA attached to them that fuse the RNA inside of a tissue section. So biologists can lay a chunk of a tumor over it, and then it'll capture all the RNA in it. And then the data is a big matrix of numbers in a large high-content image. You have to take it through a sequence of steps to get to the end thing that you want. These steps are highly variable, especially across technology types, tissue, and disease contexts. There isn't a lot of consensus in the field for each step, so we really needed a measuring stick to understand if the agents we're building were doing scientific work. The existing benchmarks we saw at the time did not measure the tasks relevant to this category of work. They mostly measured things in Q&A settings, like what would you do in an academic way, or they weren't sufficiently focused on the experiment type. This is an actual screenshot of Anthropic's model card at the time that we built this benchmark. So we built one. It's called Spatial Bench. Last December, there were 146 problems. They spanned the different kits I talked about, or attempted to, and then they spanned all those different tasks that I talked about as well. So the thing that we found at this time, and still to an extent is true today, is the grading of these outcomes in biology is too sparse because the models are pretty bad. So you have to break things up into manageable chunks to get some semblance of verifiability, and that's induced by sticking to these little components of that DAG, that analysis DAG. Getting data to a state where it would exist right before a scientist could do work on it, and then figuring out what the ground truth would be in that context. So a single evaluation looks like one or more data nodes, again a matrix of numbers, high-content image, something like this. So there's a task prompt carefully describing some scientific goal, configuration for a grader, and then a deterministic grader. It's like a Python function. If you guys notice, this looks a lot like SWE-bench. We borrowed a lot of the early ideas and tried to extend them as much as possible. Evaluation ends up looking like this. A lot of JSON. And we ended up identifying properties of what we thought good biological tests were. A little different from code, and we've built on these over time, but they still hold up. They've got to be verifiable. You have to be able to check the success condition with a function. Nothing's changed there. We'll get into some rubric stuff later, but it still holds. Durability is particularly important. Science does not admit clear ground truth. If you are lazy with your ground truth, construction of the task, a possible valid analysis path can come with the correct answer, and you'll fail it incorrectly. So you've got to make sure you're reasoning about something that's somehow invariant across analysis paths. And then, obviously, we're working with agentic stuff here. You don't want the model to answer the question in one turn. You want the conclusion to require interaction with the data, not some memorized knowledge. In practice, that's pretty difficult. We learned a lot about what models could do, and which ones to use in specific context for this category of work for our customers. And we thought, hey, this is pretty cool. Let's start to improve and learn more about this benchmarking problem. So we jumped to human verification and long-horizon extension. I'm going to quickly breeze through these. So human verification is incredibly important in science. Science does not admit clear ground truths. After watching trajectory data from multiple rounds of model releases, circa January to March of this year, we really realized a lot of our assumptions were pretty bad. And in the absence of a canonical answer, having a bunch of scientists grade each other's work ended up being the best proxy. So I'm going to look at one issue to highlight exactly what I'm talking about, problem ambiguity. A task might ask an agent to split a gene list into two groups of activity, microglial activation, oligodendrocyte inflammation, biological categories of things. Score the cells, find neighboring oligodendrocytes around some region using an appropriate radius, compute a Spearman correlation at two time points. As you can probably clearly deduce, the original problem statement creates a host of open choices. How do you split the gene list? How do you count what inflammatory genes are? It's some ambiguous word. How do you normalize the data? What the hell is an appropriate radius? How do you pull the counts within the selected radius? These are all problems that pointed to tasks that were bad, that only became revealed with human verification. Another issue is a lot of people in bioinformatics who canonically have used numerical thresholds to QC stuff. It's just completely arbitrary stuff. The cool thing about evaluation, like coding, is it forces you to reason about things more rigorously than you would when you're doing the thing yourself. You have to teach the machine to do it. You might be picking out some structure that's more important, or more durable than what you were doing if you were just doing it on your own. So we just found a lot of these numerical thresholds to be bad. I'm not going to get into this. After two rounds of human attempts, we produced a verified subset of the benchmark. We published it. That was fun. And then we also tried to increase the time horizon. So I want to be clear, the frontier of knowledge is still not quite there with biology. The labs are starting to catch up with the post-training, but we want to stay ahead. So we built a benchmark that we thought would recapitulate really difficult, true work. So we built Spatial Bench-Long. Real biological tasks are messy. They use lots of different experiment types. They use the whole workflow, they don't use little chunks. They are tasks where every step is interpreted against experimental design or contextualized with some prior literature and the original goal of what you're doing in the first place. So we built a bank of these tasks that are really trying to simulate the results sections of entire papers or the kinds of decisions you'd make in practice in industry to make a go-no-go decision on a drug program. These tasks took a week for a group of three people to make each. Taught us a bunch of stuff. An example is, can an agent reconstruct a metastatic niche in a tumor? If you have a tumor biopsy and a bunch of metastatic biopsies from where it metastasized, spread across the body, can it use both the genetics, mRNA of the metastatic lesions and the tumor to find the part of the tumor that initially seeded the metastatic growth and let it spread? From that, you can figure out, hey, what parts of the tumor are more genetically fit? Which ones actually cause problems? And construct targeted medicines to nip them in the bud. For example, this is one of the benchmark evals in the long-horizon set. None of the models get this right, but they're getting there. As you can imagine, with these long-horizon extensions, verifiable rewards at the end are somewhat uninformative, so we're starting to play with rubrics, constructing these choke points. If you can imagine the set of analysis paths as inducing some sort of tree, there are nodes that are invariant with respect to different paths, and you can use these to build rubrics using knowledge of how the tasks work. We're playing with these. We noticed that they are associated with the verifiable outcomes, which is exciting, but they're loosely correlated numerically, making us not fully confident in them for things like RL or benchmarking. A lot more work to do here still. We still strongly believe the verifiability structure is what's going to carry intelligence a bit longer. And so these days, excitingly, we've been expanding from this initial spatial focus. Really cool to see the frontier labs and community adopt these benchmarks organically. We had this interesting position by building and shipping products early and playing with the coding agents. So I think we just had an early advantage, but the benchmarks are now in the recent Anthropic model cards. And this is a picture from yesterday, Eric just showing the benchmarks at the Claude Science launch. They don't tell us this happens. They just do it, and then you read about it, and it's cool. We published a bunch more papers beyond spatial to other omics classes, so other experiment types, single-cell epigenomics, so RNA, and then the bit above the DNA, and then long-horizon extensions of these things. And then we're starting to index and measure the very gnarly, complex landscape that is drug discovery. We just put out our first benchmark on preclinical pharmacology for small molecules, and then systematically biting off pieces of the program landscape from discovery to development to translation, stratifying it by therapeutic types and experiment types. We just acquired a company building in biosecurity to form a biosecurity team. And then we just put out a collaboration with American Wetware and a surveillance company called Ackwood. Some new work, the first was released this morning. I don't know if you guys have been hearing, refusals kind of suck in biology right now. If you ask for basic questions about mitochondria, it won't answer. It's kind of stupid. So this is just an evaluation problem. There's a lot more nuance to this, but I just use that because people tend to recognize it, where we build routine tasks that simulate the kinds of things that scientists would ask for, and then more sinister red-team tasks, which are supposed to look innocuous but have some structure that is bad. Like, hey, I want to clone a gene into a bacteria and I'm telling you it's GFP, it's a glowing protein, but in reality it's a toxin or could be used to bootstrap a virus. We found that the routine tasks drastically more frequently than the routine tasks, which is not great. We're aggregating a lot of these results, essential resource along with all the preprints and a lot of evals and trajectories for you guys to check out. And I actually did okay on time. Let's go. And that's it. So we are, I hate the word lab, but we're a research lab for bio. And we do research and deployment of these agents. So we still have a lot of customers. We work with the kit manufacturers, and we use that to inform what kinds of things we make benchmarks for. We try to get the labs to compete on the benchmarks because then it makes the models better at our products. And it's been a pretty rewarding flywheel. A lot of growth, and we're hiring aggressively across engineering and science. So if you're interested in this work, please find me afterwards. Thank you. Thank you. Thank you. originally a data tool vendor for biotech and pharma. Where we started five years ago out of Berkeley, I'm 25. We started when I was 20. We stored, transformed file data from large experiments as a service. We tried to build products, explored lots of things. Over the last two years, we started moving away from biotech and pharma and more towards the people who build those kits I was talking about. Package the software into kind of white labeled things that they provided the scientists themselves, help them analyze their data. And then over time, there became this strong interaction with the agents, using the infrastructure components as tools and the loop context you guys are familiar with. Except in our domain, the tools can take days or weeks. I'm serious. What started to happen around last summer is agent prototypes started to work. So we took coding models to our knowledge at the time. They were not seriously post trained on any tasks in biology to this point. And we started to build products that look a lot like all the other agent products. They have a chat interface for you to ask questions to and they build dashboards and dispatch operations to external compute. The kinds of things these, that this agent would do is take large file data from the types of experiments I was talking about earlier, say like tissue biopsy from cancer with a spatial measurement. And then the scientist is just iterating with it to get at some question they have. You know, maybe between a malignant and non-malignant part of the tissue, what kind of genes are being over expressed. But what was fascinating is, even though it was pretty bad, it showed the early signs of working. And it became clear to us at this time that agentic biology might look a lot like code. I actually lifted this slide from Anthropics Cloud Science announcement yesterday, but we just like, you know, you had this like kind of faulty silly engineer that became better. And then as it improved in capability, you could dispatch work to teams of them that work together. The same pattern will probably emerge in science. And products and harnesses will emerge to orchestrate work and abstract it so teams of teams of agentic scientists can take on capability. But we needed focus post training, because at the time and still now frontier models cannot be trusted to do real work. They're missing some capability between knowing biology and writing code. And this is exactly extracting scientific insight from real world data. Unlike code, which is one constituent component of this work, it also involves data analysis and domain reasoning, scientific reasoning. We thought spatial biology is a good place to start. So we started building agents. This is a technical greenfield. We had many existing customers. It's also just a beautiful example of measurement drives progress. You can actually see biological phenomena play out. You can look at a developing mouse embryo. And we had to get in the guts of how the data was captured and analyzed to build good agents. I'm not going to get into this in detail, but I put up this tree of different capture technologies in spatial biology to highlight the diversity of things that exist. They really span advances in chemistry, optics, semiconductors, physics. Each branch is induced from decades of cumulative work to figure out how to measure a type of molecule. As a specific example, one technique we work with is called sequencing-based spatial, and it's where you take a slide of little beads with clumps of DNA attached to them that fused the RNA inside of a tissue section. So biologists can lay a chunk of a tumor over it, and then it'll capture all the RNA in it, and then let you know, the data is a big matrix of numbers in a large high-content image. You have to take it through a sequence of steps to get to the end thing that you want. These steps are highly variable, especially across technology types, tissue, disease contexts. There isn't a lot of consensus in the field for each step, so we really needed a measuring stick to understand if the agents we're building were doing scientific work. The existing benchmarks we saw at the time did not measure the tasks relevant to this category of work. They mostly measured things in Q&A settings, like what would you do in a kind of academic way, or they weren't sufficiently focused on the experiment type. This is an actual screenshot of Anthropics model card at the time that we built this benchmark. So we built one. It's called Spatial Bench. Last December, there's 146 problems. They spanned the different kits I talked about, or attempted to, and then they spanned all those different tasks that I talked about as well. So the thing that we found at this time, and still to an extent is true today, is the grading of these outcomes in biology is too sparse because the models are pretty bad. So you have to break things up into manageable chunks to get some semblance of verifiability, and that's kind of induced by sticking to these little components of like that DAG, that analysis DAG. Getting data to a state where it would exist right before a scientist could do work on it, and then figuring out what the ground truth would be in that context. So a single evaluation kind of looks like one or more data nodes, again like a matrix of numbers, high content image, something like this. So there's a task prompt carefully describing some scientific goal, configuration for a grader, and then a determines the grader. It's like a Python function. If you guys notice, this looks a lot like SweetBench. We borrowed a lot of the early ideas, and tried to extend them as much as possible. Evaluation ends up looking like this. A lot of JSON. And we ended up identifying properties of like what we thought good biological tests were. A little different from code, and we've built on these over time, but they still hold up. They've got to be verifiable. You have to be able to check the success condition with the function. Nothing's changed there. We'll get into some rubric stuff later, but still holds. Durability is particularly important. Science does not admit clear ground truth. If you are lazy with your ground truth, construction of the task, a possible valid analysis path, and come with the correct answer, and you'll fail it incorrectly. So you've got to make sure you're reasoning about something that's somehow invariant across analysis paths. And then obviously, we're working with the gentix stuff here. You don't want the model to answer the question in one turn. You want the conclusion to acquire interaction with the data, not some memorized knowledge. In practice, that's pretty difficult. We learned a lot about what models could do, and which ones to use in specific context for this category of work for our customers. And we thought, hey, this is pretty cool. Let's start to improve and learn more about this benchmarking problem. So we jumped to human verification and long horizon extension. I'm going to quickly breeze through these. So human verification is incredibly important in science. Science does not admit clear ground truths. After watching trajectory data from multiple rounds of model releases, circa like January to March of this year, we really realized a lot of our assumptions were pretty bad. And in the absence of like a canonical answer, having a bunch of scientists grade each other's work ended up being like the best proxy. So I'm going to look at one issue to highlight exactly what I'm talking about is problem ambiguity. A task might ask an agent to split a gene list into two groups of activity, microgulial activation, oligodendrocyte inflammation, just like biological categories of things. Score the cells, find neighboring oligodendrocytes around some region using an appropriate radius, compute a Spearman correlation at two time points. As you can probably clearly deduce, the original problem statement creates a host of open choices. How do you split the gene list? How do you count what inflammatory genes are? It's like some ambiguous word. How do you normalize the data? What the hell is an appropriate radius? How do you pull the counts within the selected radius? It's like these are all problems that pointed to tasks that were bad, that only became revealed with human verification. Another issue is just like a lot of people in bioinformatics who canonically have used numerical thresholds to QC stuff. It's just completely arbitrary stuff. Cool thing about evaluation like coding is it forces you to reason about things more rigorously than you would when you're doing the thing yourself. You have to teach the machine to do it. You might be picking out out some structure that's more important, or more durable than what you were doing if you were just doing it on your own. So we just found a lot of these numerical thresholds to be like bad. I'm not gonna get into this. After two rounds of human attempts, we produced a verified subset of the benchmark. We published it, that was fun. And then we also tried to increase the time horizon. So I wanna be clear, the frontier of knowledge is still not quite there with biology. Like the labs are starting to catch up with the post-training, but we kinda wanna stay ahead So we built a benchmark that we thought would recapitulate like really difficult, true work. So we built a Spatial Bench-Long. Real biological tasks are messy. They use lots of different experiment types. They use the whole workflow, they don't use little chunks. They are tasks, every step is interpreted against experimental design or contextualized with some prior literature and the original goal of what you're doing in the first place. So we built a bank of these tasks that are really trying to simulate the results sections of entire papers or the kinds of decisions you'd make and practice in industry to make a go-no-go decision on a drug program. These tasks took like a week for a group of three people to make each. Taught us a bunch of stuff. An example is, can an agent reconstruct a metastatic niche in a tumor? If you have like a tumor biopsy and a bunch of metastatic biopsies from like where it metastasizes spread across the body, can it use like both the genetics, mRNA of the metastatic lesions and the tumor to like find the part of the tumor that initially seeded the metastatic growth and let it spread? From that you can figure out like, hey, what parts of the tumor are more like genetically fit? Which ones actually cause problems? And construct targeted medicines to nip them in the bud. Like for example, this is one of the benchmark evals in the long horizon set, none of the models get this right, but they're getting there. As you can imagine with these long horizon extensions, verifiable rewards at the end are like somewhat uninformative, so we're starting to play with rubrics constructing these choke points. If you can imagine like the set of analysis paths is inducing some sort of tree. There are nodes that are invariant with respect to, you have different paths and you can use these to build rubrics using knowledge of how the tasks work. We're playing with these. We noticed that they are associated with the verifiable outcomes, which is exciting, but they're loosely correlated numerically, making us not fully have confidence in them for things like RL or benchmarking. A lot more work to do here still. We still strongly believe the verifiability structure is what's going to carry intelligence a bit longer. And so these days, excitingly, we've been expanding from this initial spatial focus. Really cool to see the frontier labs and community adopt these benchmarks organically. We had this interesting position by like, you know, building and shipping products early and kind of playing with the coding agents. So I think we just had a early advantage, but the benchmarks are now in like the recent anthropic model cards. And this is a picture from yesterday, Eric, just showing the benchmarks at the cloud science launch. They don't tell us this happens. They just like do it. And then you like read about it and it's cool. We published a bunch more papers beyond spatial to other omics classes. So other experiment types, single cell epigenomics. So RNA, and then the bit above the DNA, and then long horizon extensions of these things. And then we're starting to index and measure the very gnarly complex landscape that is drug discovery. We just put out our first benchmark on preclinical pharmacology for small molecules, and then systematically biting off pieces of the program landscape from discovery to development to translation, stratifying it by therapeutic types and experiment types. We just acquired a company building in biosecurity to form a biosecurity team. And then we just put out a collaboration with American wetware and a surveillance company called Ackwood. Some new work, the first was released this morning. I don't know if you guys have been hearing, refusals kind of suck in biology right now. If you ask Fable basic questions about like mitochondria, it won't answer. It's kind of stupid. So I mean, this is just like an evaluation problem. There's a lot more nuance to this, but I just use that because people tend to recognize it, where we build routine tasks that simulate the kinds of things that scientists would ask for. And then more sinister red team tasks, which are supposed to look innocuous, but have some structure that is bad. Like, hey, I want to clone a gene into a bacteria and I'm telling you it's GFP. It's like a glowing protein. But in reality, it's like a toxin or could be used to bootstrap a virus. We found that the routine tasks like drastically more frequently and then the routine tasks, which is not great. We're aggregating a lot of these results, essential resource along with all the preprints and a lot of evals and trajectories for you guys to check out. And I actually did okay on time. Let's go. And that's it. So we are kind of like, I hate the word lab, but we're kind of like a research lab for bio. And we do research and deployment of these agents. So we still have like a lot of customers. We work with the kit manufacturers and we use that to inform what kinds of things we make benchmarks for. We try to get the labs to compete on the benchmarks because then it makes the models better at our products. And it's been a pretty rewarding flywheel. A lot of growth and we're hiring aggressively across engineering and science. So if you're interested in this work, please find me afterwards. Thank you. Thank you. Thank you.