Open Reader

State of Data — Sean Cai, Independent / State of Data

completed 18:22 Jul 26, 2026 Watch on YouTube

Current Status

completed

Video ID

ZyIoTOAbRfs

RAG / Chat

Enabled
State of Data — Sean Cai, Independent / State of Data
Description

Data Quality Research at Prime Intellect and State of Data Author. Prior investor at Hummingbird and Costanoa. Speaker: Sean Cai — CEO, Independent / State of Data X: https://x.com/SeanZCai Website: https://www.seancai.com/

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: The durable value in AI data is shifting from scalable but contrived expert-generated tasks to live, process-level enterprise workflows and the infrastructure that continuously turns those workflows into post-training and RL environments.
  • Why it matters: This is directly relevant to agent systems: long-horizon agent performance depends less on generic benchmark gains and more on real task trajectories, reliable verification, environment design, model routing, and retraining across changing base models.
  • Best use: Use the talk as a strategic framework for evaluating AI-data vendors, designing agent evals and RL environments, and identifying control-plane opportunities around enterprise-owned intelligence.

Executive Summary

Sean Cai argues that conventional data labeling is no longer the important frontier. The scarce input is process-based data: the sequence of decisions, tool calls, state changes, corrections, and outcomes that show how professionals actually produce work. He distinguishes Type 1 data—captured from real workflows—from Type 2 data—expert-created examples in artificial settings—and argues that the industry often misrepresents the latter as the former.

His market framework is verification. A domain becomes trainable and commercially viable when correctness can be cheaply decomposed and checked, when people agree on what correct means, and when fresh verified examples are plentiful. Coding matured first because repositories, tests, and commits provide all three. Higher-value domains such as finance, law, healthcare, biology, security, and taste are harder because their work is long-horizon, less objectively verifiable, and locked within private enterprise systems.

Cai is especially skeptical of benchmark-driven data markets. He says vendors can manufacture difficult-looking tasks, select failures, package them as benchmarks, then sell data designed to improve those same benchmarks. A headline score under one agent scaffold is therefore an unreliable measure; harness and infrastructure choices can produce hidden false positives and false negatives. His finance examples suggest models with similar aggregate scores can have opposite failure profiles, such as arithmetic strength versus methodological strength.

The investment and product implication is that data businesses must become enterprise AI infrastructure businesses. Enterprises that want to own rather than rent intelligence will need a layer for routing models by cost/latency/performance, migrating RL datasets and post-training across base-model changes, and translating messy business context into executable training and evaluation environments. Cai calls these bespoke translation systems “Antikythera mechanisms.”

Key Takeaways

  • Claim: The key AI-data bottleneck is no longer generic annotation volume; it is high-fidelity process data that captures real professional work from initial context through decisions and final outcomes. | Evidence: Cai contrasts state-based data such as ERP rows or saved files with process-based data: trajectories, reasoning traces, and sequences of decisions. He characterizes the legacy labeling market as roughly $10–15 billion per lab annually but calls it the least interesting part of the market. | Implication: For agent products, prioritize instrumenting live workflows—tool use, intermediate states, recoveries, approvals, and independently observed outcomes—rather than treating static document corpora or isolated prompt-response pairs as the core proprietary asset. | Caveat: Type 2, contrived expert-created data remains useful as an initial training mechanism; Cai's claim is that it is insufficient for moving models from competent performance toward robust expert performance.
  • Claim: AI data supply chains are permanently fragmenting because quality does not scale linearly with quantity, allowing specialists to outperform vertically integrated providers at sourcing, environment design, rewards, and evaluation. | Evidence: He says labs increasingly mandate diversification across roughly 20–30 vendors because they distrust any single vendor's ability to scale quality. He contrasts this with the prior model in which firms such as Mercor, Handshake, Surge, and Scale AI had to operate as vertically integrated providers. | Implication: Avoid assuming a single data or RL-environment vendor can be a full-stack strategic dependency. Build a vendor portfolio and retain internal ownership of task definitions, quality standards, data schemas, and evaluation logic. | Caveat: The assertion is presented as Cai's market observation, not as independently substantiated vendor-revenue or procurement data in the transcript.
  • Claim: Verifier's Law predicts which agent markets mature first: tasks are easier to train when their outcomes are decomposable, their definitions of correctness have consensus, and fresh verified examples occur frequently. | Evidence: Cai attributes the principle to researcher Jason Wei and divides verifiability into asymmetry of verification (checkable steps), veracity of verification (consensus on correctness), and proliferation of verification (the supply of new verified examples). He explains coding's early maturity through unit tests, shared standards for working code, and abundant GitHub commit histories. | Implication: Score prospective agent verticals before committing: select initial wedges with clear step-level checks, observable outcomes, recurring work, and enough real examples, then use those workflows as a foundation for harder long-horizon tasks. | Caveat: High verifiability is not equivalent to high business value; it explains trainability and speed of product maturity, while many economically valuable workflows remain difficult to verify.
  • Claim: Most synthetic expert benchmarks are structurally vulnerable to Goodharting and do not establish long-horizon agent reliability. | Evidence: Cai describes a Type 2 loop in which vendors ask experts to generate plausible tasks via chat, solve them, cherry-pick cases where models fail, market them as difficult benchmarks, and sell training data intended to improve those exact tests. He argues these benchmarks usually test isolated in-distribution questions, not dependent episodes of work. | Implication: Separate task/eval ownership from data-vendor incentives. For agent systems, test multi-step episodes with constrained state transitions, non-interchangeable tool calls, forced recovery from failure, and outcome-based grading—not merely standalone task completion. | Caveat: The talk does not claim all benchmarks are invalid; its narrower argument is that a single benchmark, especially when controlled by the same party selling the associated data, is an inadequate evidence standard.
  • Claim: Aggregate model leaderboard scores conceal meaningful model-specific failure modes, so evaluation must be rubric-level and scaffold-aware. | Evidence: Cai says cross-harness and cross-infrastructure differences cause benchmark divergence, including hidden false-positive and false-negative rates in SWE-style benchmarks. On private finance tasks—ARR waterfall reconciliation, an LBO evaluation memo, and a long-short pair trade—he reports that Opus 4.8 could underperform 4.7 on some rubrics, while GPT 5.5 and Opus 4.8 scored within three points overall but failed in opposite ways: GPT handled arithmetic better and Opus methodology better while missing arithmetic. | Implication: Do not route models using a general leaderboard alone. Build a workload-specific scorecard that isolates relevant dimensions—calculation, planning, policy compliance, tool execution, recovery, and domain methodology—and validate each model in the production harness. | Caveat: These are Cai's internal/private evaluations, with no task specifications, sample sizes, scoring protocol, or independently reproducible results provided in the transcript.
  • Claim: Data-market purchasing can serve as an upstream signal for forthcoming frontier-model application priorities, but only where the underlying modality and verification setup are sufficiently settled. | Evidence: Cai says Anthropic spent heavily on cybersecurity data in January and on biology data in March–April, followed two to three months later by Cyber and Bio/Life Sciences product efforts. His screen for robust long-horizon work includes enforced step length, heterogeneous non-interchangeable tool calls, state transitions that constrain subsequent actions, mandatory recovery, sequential decisions on a single entity, expert action at each step, and independently recorded outcomes. | Implication: Monitor enterprise data procurement and vendor formation as leading indicators, but avoid investing heavily in a data modality until its training representation, verifier design, and operational use case are clear. | Caveat: He cites robotics as a counterexample: choices among egocentric data, teleoperation, and UMI remain unresolved research questions, and he considers many robotics-data vendors unsophisticated. Data demand alone is therefore not a sufficient market signal.
  • Claim: The enduring enterprise opportunity is a model-agnostic control plane that turns proprietary work into continuously maintained AI capability; pure data businesses will be pushed toward services and application-layer ownership. | Evidence: Cai says successful data companies, including Mercor and Handshake by his account, are already generating substantial enterprise revenue. He names needs to route smaller models by cost, latency, and performance; retain RL datasets through base-model migration by rerunning post-training instead of restarting; and create “Antikythera mechanisms” that translate business context into training/evaluation systems. He also argues that model alternatives such as GLM 5.2 demonstrate application-layer companies can decouple from a single foundation-model provider. | Implication: The defensible architecture is not a one-model application or a static dataset; it is a portable operating loop: capture work, evaluate it, train or adapt models, route among providers, and repeat as models and workflows change. | Caveat: Claims about particular vendors' enterprise revenue and GLM 5.2's relative performance are unsupported in the transcript and should be independently checked before informing a market or investment conclusion.

Detailed Brief

The long-horizon opportunity and why live enterprises matter

  • Claims: The largest remaining automation opportunity lies in deeply dependent, long-horizon white-collar work rather than the short-horizon tasks that are currently addressable.; A dataset has a limited useful life at the frontier because models eventually absorb it; a live business relationship creates a renewable stream of work trajectories and outcomes.; RL-environment companies should be viewed primarily as research accelerators, and become venture-scale only when their generalized infrastructure supports a real enterprise application.
  • Evidence: Cai depicts task horizon on the horizontal axis and share of white-collar work on the vertical axis, arguing that the addressable threshold shifts rightward when a real-world data pipeline is built.; He calls startup codebases a frequently purchased but ultimately finite source of data, contrasting them with ongoing enterprise workflows.; He warns against vendors offering large volumes of weakly relevant video—for example, '100,000 hours of iPhone video'—as a substitute for a coherent robotics-data strategy.
  • Caveats: The speaker does not provide empirical estimates of the size of the long-horizon market, the cost of acquiring workflow rights, or the economics of enterprise data partnerships.; Capturing actual workflows raises unaddressed issues around privacy, IP ownership, consent, auditability, and security controls.
  • Implications: A durable data-rights strategy should target recurring operational systems where outcome quality can be observed over time, rather than one-time corpus purchases.; The hard problem is not merely collecting traces; it is preserving the context and state needed for those traces to remain useful as executable agent environments.

Antikythera mechanisms as the missing enterprise translation layer

  • Claims: Generic foundation models cannot directly consume messy enterprise context as a reliable training signal; organizations need bespoke systems that translate business rules, tools, constraints, and outcomes into evaluable environments.; Cai positions his newly announced work as building these systems and delivering RL as a service with companies that monetize their data assets.
  • Evidence: He repeatedly labels the mechanism an “Antikythera mechanism,” borrowing the image of a specialized machine that converts complex inputs into a structured output.; The talk identifies enterprise-owned intelligence as the end state, replacing permanent dependence on rented frontier-model capability.
  • Caveats: The presentation does not define the full five-job architecture; it explicitly names model routing, RL-data management across base-model migration, and Antikythera mechanisms, while leaving the remainder unspecified.; This portion is also a product announcement, so its framing should be read as strategic positioning as well as analysis.
  • Implications: The control plane for enterprise agents should model workflow semantics explicitly: permissions, tools, intermediate artifacts, state transitions, acceptance criteria, exception paths, and feedback sources.; The enterprise vendor likely to retain value is the one that operationalizes these semantics across model changes, rather than the one that merely brokers annotators or sells a snapshot dataset.

Notable Concepts & Terms

  • Type 1 data: Real workflow capture with minimal artificial reward shaping, such as code commits or session replays; Cai considers it the route to realistic expert-level model behavior.
  • Type 2 data: Contrived examples manufactured by hired experts in an artificial setting; useful for bootstrapping but vulnerable to benchmark gaming and weak transfer to real work.
  • Process-based data: Decision trajectories, reasoning traces, tool usage, and workflow states rather than final records; it is the central data asset for long-horizon agents.
  • Verifier's Law: Jason Wei's framing that a task's trainability rises with ease of verification; Cai uses it to predict which AI application domains will mature first.
  • Asymmetry, veracity, and proliferation of verification: Cai's three-part diagnostic: whether work can be checked stepwise, whether correctness has consensus, and whether verified examples are continuously generated.
  • Cross-harness / cross-infrastructure differencing: Performance variation caused by the agent scaffold and execution setup rather than the base model alone, making single benchmark results unreliable.
  • Antikythera mechanisms: Cai's term for bespoke enterprise systems that convert messy operational context into usable evaluations, rewards, and RL environments.
  • Neolabs: Data companies evolving into enterprise AI infrastructure, service, and application companies because durable value lies beyond data collection.

Operator Notes / Why Ken Should Care

  • Create an internal agent-evaluation policy requiring workload-specific, multi-step tests in the exact production harness before changing models or vendors; prohibit model selection from a single public benchmark.
  • Build an eval suite around real operating episodes with explicit state, required tools, recovery paths, independent outcome signals, and separate rubrics for methodology versus arithmetic/execution.
  • Treat workflow telemetry as a product capability: determine which agent interactions, human corrections, artifacts, approvals, and outcome signals can be legally captured and reused for evaluation or post-training.
  • Design model abstraction now: keep prompts, tools, evals, training data, and routing policies portable across providers and open-weight bases, with a defined revalidation process after each base-model change.
  • Due-diligence data/RL vendors for provenance, whether tasks are live or synthetic, evaluator independence, harness sensitivity, long-horizon specifications, and rights to use operational data; do not accept benchmark difficulty as proof of transfer.
  • Monitor cybersecurity, biology, finance, healthcare, law, and other high-wage domains for evidence that verification and enterprise-workflow access have improved; treat raw data spending as a lead signal rather than a commitment trigger.

Source/Metadata

  • Title: State of Data — Sean Cai, Independent / State of Data
  • Transcript words: 5374
  • Duration seconds: 1102
  • Timestamp note: No timestamps or chapter markers were present. The transcript contains a substantial duplicated segment and repeated extraction noise near the end.

Transcript

3257 words en Processed in 140.3s

[SPEAKER_00] Good to give this talk, and I just want to say right off the bat, this is probably going to be a little different from what you've seen so far at AI conference. I'm here to not deliver an agenda on any company's behalf, but just to expose a lot of alpha and data markets, but also just tell you what's really going on behind the scenes in a very murky landscape where nobody seems to know how McCore, Handshake, and a lot of these folks actually produce data. So, quick reframe before we start: when people hear data markets, they picture Scale around 2019. They picture these rooms of annotators in Manila labeling images, and that's real, and it's maybe $10 to $15 billion a year per lab, but it's the least interesting part. The models work now. What's scarce and badly priced is the data that takes them from generalist competence into real expertise. We knew this since 2024 when Scale got acquired and we were spamming a lot of GPQA data sets. So, look, let me start with this framing. Data is to the white color revolution what coal and iron is to the Victorian age and the information age right now. I have this piece called the TAM isn't a vertical, it's all of labor. And so the supply chain is just doing what every industrializing supply chain does. It unbundles. Two years ago, one vertically integrated giant like a McCore or a Surge or a Scale AI, if you guys are unfamiliar, they're massive data companies. They had to do all of it, but that was because that's the only way that unit economics worked in an immature market. Today, that's increasingly not the case. Specialists out-compete the giants at a lot of steps: sourcing the people, building environments, designing rewards, running evals. The fragmentation, I would say, is pretty permanent. It's not transitional. And quality increasingly does not scale linearly with quantity, which leads to a cottage industry in data land right now where you have labs literally mandate vendor diversification on a scale of 20 to 30 different vendors because they inherently distrust their ability to scale quality with quantity. So to orient you, here's the argument as it developed this year. In January, we saw a lot of industrialization and unbundling. You'll hear me come back to this mechanism I call the Anticothera Mechanisms, which is these bespoke systems that translate messy business context into evals. Increasingly important in labs, hungry quest for real-world seed data, the seed envs. I also explored this concept called type 1 versus type 2 data, contrived versus non-contrived, for the more researcher types in the audience, and why the GPQA cell playbook that worked in 2024 for data acquisition falls apart on long-horizon, non-verifiable work. So, going through all these topics, all of which are online, I'll actually just skip to the more interesting part, but I'm bringing this up here in case any of the particular topics I talk about are particularly interesting and you want to dive in more. So model improvement is a function of three inputs. We all see this commonly expressed in DOJ: compute, data, and talent. I put together this very rudimentary chart online in one of my pieces, in a very rudimentary equation, just to express the fact that if there's any imbalance in this compute, data, and talent, then you start seeing an inefficiency in producing what I call generalized AI model performance, right? But also we see an inefficiency in capex spend that is a result of this equation right now. Exponentially increasing capex spend, but AI revenues are falling far behind. Data is the underfunded leg here. It's the one that turns a generalist model into a real expert. It's the one that's actually, I believe, quite lacking in this equation and thus, because of the imbalance penalty parameter concepts that I'm expressing, presents a whole opportunity. But, going back to what I said at the start, what is data actually? So most people picture state-based data, the rows in an ERP, which is the final output or saved file. That's the 2023 next-token prediction model, and it's mostly personal data wrapped in privacy law. What is actually available nowadays is process-based data, which is the trajectory, the reasoning trace, the sequence of decisions. So it's what gets a professional from a blank page to a finished work output and delineates how the work gets done. So on top of that, it's a quality axis. This is the vocabulary I'll use all talk. Type one data is a pure capture of real workflows like GitHub commits or session replays with minimal reward shaping by non-experts. And type two is contrived data where you hire experts, sit them in an arbitrary setting, and have them manufacture examples. Type two, the right place to start. Models, when they were reading at a first-grade level, I think anybody in the world could teach them as a third-grade teacher. But type one is what gets you from 20 to 80%, so to say, because the realism is inherited from the work itself. And the structural reason why it matters is that data is the most appreciable asset there is. A data set is only available insofar as frontier data markets, as the frontier moves. So the only durable supply of it, technically, is a live business you partner with, not a dead set. Startups' code bases, like so many data companies out there are buying today. The dirty secret of the industry, though, is that everybody sells type two and bills it as type one. So now let me talk about the central axis of how you can think about verification and deciding which application-layer domains are maturing first and why. There's a reason why Claude Science came out so far in the future after models got good, before a lot of other application-layer advances. So Jason Wei is a researcher online whose blog is great. You should read him. He has this law. It's called Verifier's Law. The ease of training a model to do a task is proportional to how verifiable the task is. So I break verifiability into three axes here. And once you have them, a huge amount of this market stops being mysterious. First, asymmetry of verification. How hard is it to decompose the task into checkable steps? Veracity of verification. How much consensus is there about what correct even means? And then thirdly, proliferation of verification. How often does the real world hand you fresh examples of verified work? If you think about why coding is the first mature AI app-layer market, that's really no accident because we were blessed to have something called GitHub from Web 2.0, which solved all three of these at once. Unit tests give you objective, decomposable correctness. The community agrees on what working code means generally. And there are effectively infinite public examples with these commit messages as basically free reasoning traces. So they score high, high, and high on these three axes. Now, look where the money is trying to go now. Biology, security, taste, finance, healthcare, law. These sit pretty low on veracity, pretty low on verification, and the verification examples are locked in these enterprise workflows no Web 2.0 system ever captured, really, or very sparingly few. That's why you've probably been approached by Mercore to buy your Slack logs or your logs as of late if you work at an AI company. And that's the whole game, right? It's predictive, it's not just descriptive. Classify any professions' tasks on these three axes, and you can tell which markets mature most. So it's no surprise that after code, we went to search, and after search, we went to finance, and after finance, we went to healthcare and law, and after healthcare and law, well, I'd say cyber, biological, and scientific discovery, and maybe even taste, which is probably the most unverifiable out of all these. So verification is a bottleneck. Let me talk about why the industry sells a lot of snake oil here, and why most benchmarks you see are quietly fake. The dominant Type 2 recipe is: let's hire domain experts, let's have them use chat to generate plausible tasks, let's have them solve those tasks, and let's cherry-pick the ones where the model diverged, let's package them as a hard North Star benchmark. And then this is the perverse part: let's sell the data that he'll climb that same benchmark. It's basically, and maybe some of you guys here in SF have heard this a little too much, Goodhart's law with a profit motive. The moment your measure becomes a target, and then the target is set by people who aren't true domain experts, it stops measuring anything real. And so the whole market, I like to call it, sits in a fog of war. Labs, vendors, and enterprises, they're all guessing which data actually improves the model. The dominant Type 2 recipe is: let's hire domain experts, let's have them use chat to generate plausible tasks, let's have them solve those tasks, and let's cherry-pick the ones where the model diverged, let's package them as a hard North Star benchmark. And then this is the perverse part: let's sell the data that will help it climb that same benchmark. It's Goodhart's law with a profit motive. The moment your measure becomes a target, and then the target is set by people who aren't true domain experts, it stops measuring anything real. And so the whole market, I like to call it, sits in a fog of war. Labs, vendors, and enterprises, they're all guessing which data actually improves the model. Nobody can see this clearly; there's a structural tell. Contrived benchmarks only ever test a single isolated in-distribution question. They can't test whether a model sustains correct reasoning across a long-dependent episode, because these tests were all meant to be solved in isolation. In May, I spent a lot of time on this. Increasingly, in a lot of Anthropic blog posts, but also a lot of other researchers have noted that cross-harness differencing and cross-infrastructure differencing is the primary cause for a lot of benchmark divergences and performance. A lot of you are probably very familiar with the frontier SWEs of the world and the deep SWEs of the world benchmarks. There are issues that all of these benchmarks have related to that specifically, in particular just high false positive and false negative rates that are not immediately obvious when you look at the overhead stats on the benchmark. So, a single benchmark number under a single scaffold is basically one sample from a distribution that nobody measured. And that's why there's so much benchmark psychosis today. It's a pretty noisy sample. Really briefly, on an example I took, I basically have an internal version of Val's AI and a lot of private benchmarks. So, I took three finance tasks from three real-world vendors, like an ARR waterfall reconciliation, an LBO evaluation memo, and a long-short pair trade for hedge fund trading. So, long-horizon, non-verifiable finance tasks, relatively robust, deterministic verifiers paired with LLM as a judge applied correctly. So, when you actually run these, I think you'll notice very clearly Opus 4.8 is worse than 4.7 on a lot of these rubrics. You'll start noticing things if you actually do rubric analysis at the 4.7 to 4.8, except over-engineered self-reflection. You'll notice that GPT 5.5 and Opus 4.8 score within three points on the same task. They fail in exactly opposite directions, whereas GPT nails the arithmetic, but Opus nails the methodology, but loses arithmetic, which is to say it's clear, if you actually do very good agnostic benchmarking, you can tell where the post-training directions for a lot of these teams went. And subsequently, it informs a lot of the data that comes from this. So, going back to the start, a single benchmark number is a sample from a distribution nobody measured, and it's basically like taking a Swiss Army knife and using the screwdriver bit to cut cheese and concluding the knife is broken. So, the receipts from the last slide, this goes into an RL environments report that is sent to labs, but no single leaderboard number ever shows you this part of the data, which some of the most sophisticated RL environment companies, some of which are actually here at this conference, would show you. So, how do we read which domain is next instead of chasing hype? You can actually use this as a proxy: data markets as an upstream indicator of what next application-layer products labs will come out with. In January, Anthropic was spending a lot on cybersecurity data from new vendors. And in March and April, they were spending a lot on biological data from associated vendors. And what do you know happened two to three months after? Well, Mythos and Cyber and Cloud Bio/Life Sciences today. So, if you want to use this checklist I use, classify the profession's tasks under three axes: apply to long-horizon bar, the ones real labs use, enforced step length, heterogeneous tool calls that aren't interchangeable, state transitions that genuinely constrain future actions, mandatory failure recovery. You'll notice a pretty mixed bag of how long-horizon is even defined from the vendor perspective in terms of specs. And then, three, you want to look at the raw data for five signals: sequential decisions against one entity and inferable action, expert action per step, outcomes recorded by independent parties. But moreover, just economically available high-wage work. So, subsequently, one counterexample to keep you honest: robotics. The modality is not settled: ego versus teleop versus UMI. But moreover, I find just generally a huge degree of unsophistication within robotics data vendors today. A huge degree of unsophistication. Some of you are probably robotics researchers in a crowd. I don't know how many times people have come up to you and they're like, oh, here's 100,000 hours of iPhone video from my friends in India. Do you want to buy this for ego data? So, in that domain, your vendor's choice is entangled with an unsolved research question. At the end of the day, you have to realize all environment companies, if they actually succeed, are more research accelerators. It's a boutique industry. If it's venture-scalable, it's because the infrastructure they're building agnostically in-house helps an enterprise application-layer use case rather than assuming that data markets today will stay as they are forever. So, don't die on a modality hill. So, one last exhibit from my work. This is the map under everything I just described. Look, the share of the white-collar work is on the vertical axis. Task horizon is on the horizontal. Short horizon is very addressable right now. But you think about this long tail on the right side, this deep dependent long-horizon work. That's where the real economic value and data buildout both live. And this threshold line moves rightward every time somebody builds a real-world data pipeline. Now I want to talk a bit more about the model layer. It changes who needs to build what. So, historical fact: no pioneer of an infrastructure technology has actually held more than 10% of the market in the long run. I'm not saying this applies to Anthropic, but those who forget history are condemned to repeat it. Railroads built company towns. They charged tyrannical rents. They got nationalized as soon as the automobile moved in. AWS and Google consolidated the infrastructure layer and still never captured the application layer. And right now, OpenAI and Anthropic are carving out these fiefdoms. But all these pressures from anti-distillation, export appeals, enterprise exclusivity, they're basically the equivalent of railroad rents, and they erode. Like, the automobile, in some cases, has already arrived. GLM 5.2, surpassing GPT on a lot of real-world rubrics, is pretty hard proof that a lot of app-layer companies can decouple themselves from the model layer. And because models differ on efficiency and modality, they're not fungible like electricity. So, it's not exactly leading to nationalization, but it's definitely not heading to durable lock-in either. And the whole question on what to build next hinges on whether a general enterprise can decouple models from foundation model labs. But luckily, for a lot of data companies, this is actually where they're headed. I came here to give a talk on data markets. I'm here to tell you that the successful data companies nowadays are all pivoting to enterprise. I maybe shouldn't say this in a public setting, but Mercore and Handshake, people don't notice an incredibly large amount of their revenues are enterprise now, enterprise in ways you wouldn't expect a data business to do. Once enterprises stop renting a lab's intelligence and start owning their own, you need an entire abstraction layer that doesn't exist yet. There's five jobs: serve and route small models targeted by cost, latency, and performance profiles; and two, manage your RL datasets across base model migration. So, when you swap to a new open-source base, you rerun post-training automatically instead of starting over. Three, the antikythera mechanisms that I talked about before. And in the interest of time, happy to talk more about emerging infrastructure needs afterwards if you want to. So, let me bring it all together. I think data companies all realize that they have to be neolabs. Data businesses do not stay data businesses because the durable value accrues to the services and app layer of actual work. Enterprise in ways you wouldn't expect a data business to do. Once enterprises stop renting a lab's intelligence and start owning their own, you need an entire abstraction layer that doesn't exist yet. There's five jobs: serve and route small models targeted by cost, latency, and performance profiles. And two, manage your RL datasets across base model migration. So, when you swap to a new open source base, you rerun post-training automatically instead of starting over. Three, the antikythera mechanisms that I talked about before. And in the interest of time, happy to talk more about emerging infrastructure needs afterwards if you want to. So, let me bring it all together. I think data companies all realize that they have to be neolabs. Data businesses do not stay data businesses because the durable value accrues to the services and app layer of actual work. So, two takeaways: if you're a researcher, stop outsourcing your definition of realism to the same vendors you buy your evals and tasks from. That's just letting the test writer grade the task. And if you're a builder, your moat's not the data. It's the pipeline into real-world work, plus the infra to keep retraining on it as the models improve underneath you. I'll close on this. I'm building, I'm working on something new. This is the first public announcement of it. I'm building antikythera mechanisms. I'm working with a lot of real-world companies that monetize their data assets, but moreover, help implement RL as a service with a lot of the enterprises in the world, mitigating a lot of the pitfalls of a lot of companies I mentioned. If you guys want to talk about it afterwards, I'm on Twitter. I always write a lot on Twitter and Substack. And I'm around afterwards, too. Thank you. and I'm going to buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and it stops sort of measuring anything real. And so the whole market, I like to call it, sits in a sort of fog of war. Labs, vendors, and enterprises, they're all kind of guessing which data actually improves the model. Nobody can see this clearly, there's a structural tell. Contrived benchmarks, they only ever test a single isolated in distribution question. They can't test whether a model sort of sustains, like, correct reasoning across a long-dependent episode, because these tests were all sort of meant to be solved in isolation. In May, I spent a lot of time in this. Increasingly, in a lot of anthropic blog posts, but also like a lot of other researchers have noted, that cross harness differencing and cross infrastructure differencing is the primary cause for a lot of benchmark divergences and performance. A lot of you are probably very familiar with, like, the frontier SWE's of the world and the deep SWE's of the world benchmarks. There are issues that all of these benchmarks have related to that specifically. In particular, just high false positive and false negative rates that are not immediately obvious when you look at the above head stats on the benchmark. So, a single benchmark number under a single scaffold is, like, basically one sample from a distribution who's basically with nobody measured. And that's why there's so much what I call benchmark psychosis today. It's like a pretty noisy sample. So, you know, really briefly on an example I took, I always have, like, I basically have an internal version of Val's AI and a lot of private benchmarks. So, I took three finance tasks from three real-world vendors, like an ARR waterfall reconciliation and LBO evaluation memo and a sort of like long short pair trade for hedge fund trading. So, long horizon, non-verifiable finance tasks, relatively robust, like deterministic verifiers paired with LLM as a judge applied correctly. So, when you actually run these, I think you'll notice very clearly Opus 4.8 is worse than 4.7 on a lot of these rubrics. You'll start noticing things if you actually do rubric analysis at the 4.7 to 4.8, except over-engineered self-reflection. You'll notice that GPT 5.5 and Opus 4.8 score within three points on the same task. They fell in like exactly opposite directions, whereas GPT nails the arithmetic, but Opus nails the methodology, but loses arithmetic, which is to say it's clear if you actually do very good agnostic benchmarking, you can tell where the post-training directions for a lot of these teams went. And subsequently, it informs, I think, like a lot of the data that comes for this. So, you know, going back to the start, a single benchmark number is a sample from a distribution nobody measured, and it's basically like taking a Swiss Army knife and using the screwdriver bit to cut cheese and like concluding the knife is broken. So, you know, the receipts from the last slide, this goes into an RL environments report that is sent to labs, but no single leader board number ever shows you this part of, you know, data, which some of the most sophisticated RL environment companies, some of which are actually here at this conference, would show you this. So, how do we read sort of which domain is next instead of chasing hype? And you can actually use this as a proxy data markets as an upstream indicator of what next application layer products that labs will come out with. In January, Anthropic was spending a lot on cybersecurity data from new vendors. And in March and April, they were spending a lot on biological data from associated vendors. And what do you know happened like two to three months after? Well, Mythos and Cyber and Cloud Bio slash Life Sciences today. So, if you want to use this checklist I use, classify the professions tasks under three axes, apply to Longhars and Barr, the ones real labs use, enforced step length heterogeneous tool calls that aren't interchangeable, state transitions that genuinely constrain future actions, mandatory failure recovery. You'll notice a pretty mixed bag of how Longhars and is even defined from the vendor perspective in terms of specs. And then three, you want to look at the raw data for five signals. So, sequential decisions against like one entity and inferable action, expert action per step, outcomes recorded by independent parties. But moreover, just economically available high wage work. So, subsequently, one counter example to keep you honest, like robotics, the modality is not settled. Ego versus teleop versus UMI. But moreover, I find just generally a huge degree of unsophistication within robotics data vendors today. A huge degree of unsophistication. I don't know. Some of you guys are probably robotics researchers in a crowd. I don't know how many times like people have come up to you and they're like, oh, here's like 100,000 hours of like iPhone video from my friends in India. Do you want to buy this for Ego data? So, in that sort of domain, your vendor's choice is entangled with an unsolved research question. At the end of the day, you have to realize like, are all environment companies, if they actually succeed, they're more so research accelerators. It's a boutique industry. If it's venture scalable, it's because the infrastructure they're building agnostically in-house helps an enterprise application layer use case rather than assuming that data markets today will stay as they will forever. So, don't die on a modality hill, basically. So, one last exhibit just from my work. This is the map under everything I just described. Look, the share of the white collar work is sort of on the vertical axis. Task horizon is on the horizontal. Short horizon is very addressable right now. But you think about this long tail on the right side, this deep dependent long horizon work. That's where the real economic value and data build out sort of both lives. And this threshold line that kind of moves rightward every time somebody builds a real-world data pipeline. Now I want to talk a bit more about the model layer. It changes who needs to build what. So, historical fact, no pioneer of an infrastructure technology has actually held more than 10% of the market in the long run. I'm not saying this applies to Anthropic, but, you know, just those who forget history are condemned to repeat it. Railroads built company towns. They charged tyrannical rents. They got nationalized as soon as the automobile moved in. AWS and Google, they consolidated the infrastructure layer and still never captured the application layer. And right now, OpenAI and Anthropic are carving out these fiefdoms. But, like, all these pressures from anti-distillation, export appeals, enterprise exclusivity, they're basically the equivalent of railroad rents and they erode. Right? Like, the automobile, in some cases, has already arrived. Like, GLM 5.2, surpassing GPT on a lot of real-world rubrics, is pretty hard proof that a lot of app layer companies can decouple themselves from the model layer. And because models differ on efficiency and modality, they're not fungible like electricity. So, it's not exactly leading to nationalization, but it's definitely not heading to durable lock-in either. And the whole question on what to build next hinges on whether a general enterprise can decouple models from foundation model labs. But luckily, you know, for a lot of data companies, this is actually where they're headed. I came here to give a talk on data markets. I'm here to tell you that the successful data companies nowadays are all pivoting to enterprise. I maybe shouldn't say this in a public setting, but Mercore and Handshake, people don't notice an incredibly large amount of their revenues are enterprise now. Enterprise in ways you wouldn't expect a data business to do. Once enterprises stop renting a lab's intelligence and starts owning their own, you need an entire abstraction layer that doesn't exist yet. There's like five jobs, serve and route, small models targeted by cost latency and performance profiles. And two, manage your RL datasets instead of like across base model migration. So, when you swap to a new open source base, you rerun post training automatically instead of starting over. Three, the antikythera mechanisms that I talked about before. And in the interest of time, happy to talk more about like emerging infrastructure needs afterwards if you want to. So, let me bring it all together. I think data companies all realize that they have to be neolabs. Data businesses do not stay data businesses because the durable value sort of accrues to the services and app layer of actual work. So, two takeaways, you know, if you're a researcher, stop outsourcing your definition of realism to the same vendors you buy your evals and tasks from. That's kind of just letting the test writer grade the task. And if you're a builder, your moat's not the data. It's the sort of pipeline into real world work. Plus the infra to keep retraining on it as the models improve underneath you. I'll close on this. I'm building, I'm working on something new. This is the first public announcement of it. I'm building antikythera mechanisms. I'm working with a lot of real world companies that monetize their data assets. But moreover, help implement RL as a service with a lot of the enterprises in the world. Mitigating a lot of the pitfalls of a lot of companies I mentioned. If you guys want to talk about it afterwards, I'm on Twitter. I always write a lot on Twitter and Substack. And I'm around afterwards, too. Thank you. and I'm going to buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and buy a box and