Mercor Head of Product on Revenue Concentration from Frontier Labs
Description
Osvald Nitski is the Head of Product at Mercor, the AI-training and expert-data marketplace powering frontier-model development. Mercor last raised a $350 million Series C at a $10 billion valuation, and is reportedly in discussions for a new round at a $20 billion valuation. Mercor crossed $2BN in ARR in June; doubling from $1 billion in only four months. ----------------------------------------------- Timestamps: 00:00 Intro 01:12 Does Open-Source Cannibalize Mercor's Core Business? 02:29 Why 90% of Enterprise Workflows Can't Be Done With Open Models 07:47 Do We Have an Enterprise AI ROI Problem? 09:05 Balancing Token Spend vs Performance 10:07 Salesforce Spends $300M on Anthropic 12:13 AI Makes the PM Role Harder, Not Easier 14:55 The Biggest Product Mistake 22:42 Why RL Environments Are the Fastest-Growing Data Type Right Now 28:28 How Mercor's Product Teams Are Structured 33:54 The Three Secrets to Scaling Supply on the Marketplace 38:04 Mercor Has High Revenue Concentration 43:47 Why Data Projects for Enterprises Are So Operationally Intense 51:02 Why Cybersecurity Data Will Never Hit the 90% Sufficiency Ceiling 52:15 Hiring in SF: Brutal Talent War & What Mercor Looks For 54:15 Quick-Fire Round ---------------------------------------------------------------------------------------------- Subscribe on Spotify: https://open.spotify.com/show/3j2KMcZTtgTNBKwtZBMHvl?si=85bc9196860e4466 Subscribe on Apple Podcasts: https://podcasts.apple.com/us/podcast/the-twenty-minute-vc-20vc-venture-capital-startup/id958230465 Follow Harry Stebbings on X: https://twitter.com/HarryStebbings Follow Osvald Nitski on X: https://twitter.com/OsvaldNitski Follow 20VC on Instagram: https://www.instagram.com/20vchq Follow 20VC on TikTok: https://www.tiktok.com/@20vc_tok Visit our Website: https://www.20vc.com Subscribe to our Newsletter: https://www.thetwentyminutevc.com/contact ----------------------------------------------- #20vc #harrystebbings #founder #entrepreneur #me
Summary
Generated by gpt-5.6-terraAt-a-Glance
- Verdict: Watch fully
- Core thesis: Mercor's product leader argues that frontier-model progress expands rather than commoditizes the market for high-quality human data, especially for long-horizon agent tasks, enterprise-specific evaluations, and RL environments that mirror deployment.
- Why it matters: The interview provides a useful operating view of the emerging agent-data stack: evals become the specification and optimization target, services bridge an enterprise AI-skills gap, and product teams must preserve judgment while AI removes execution bottlenecks.
- Best use: Use it to pressure-test an AI-data or agent-services investment thesis and extract operating patterns for building eval-driven, human-in-the-loop agent products.
Executive Summary
Oswald Nitsky frames Mercor's core market as the frontier of model capability rather than a fixed set of current enterprise workflows. His rebuttal to the idea that open models will handle 90% of work is that demand estimates omit latent, long-horizon automation: agents that run procurement, legal, or other processes for weeks or months with limited oversight. Open models may eliminate spending on already-solved tasks, but they raise the baseline and leave customers pursuing new capability gaps.
The crucial market shift is from generic annotation toward high-fidelity evaluation and reinforcement-learning environments. Mercor defines environments as simulated applications plus a rich starting "world" of files and state, in which agents perform realistic tasks. The claim is that agents must train and be evaluated in conditions close to deployment; this makes data production operationally hard, sensitive to edge cases, and difficult to commoditize through small founder-led annotation shops.
On enterprise deployment, Nitsky sees services as a temporary but necessary answer to an acute knowledge-distribution problem. Companies lack enough people who can deploy agents, build evals, and operate AI-first workflows. Over a roughly decade-long transition, he expects that knowledge and product maturity to reduce the need for bespoke forward-deployed teams, though deployment-oriented engineers remain valuable now because they combine communication, systems simplification, and technical skill.
Internally, Mercor is using AI to increase engineering velocity while actively resisting product sprawl. Nitsky says the bottleneck has shifted toward product judgment: deciding which user problems and experiments matter commercially. The company has learned that maximal flexibility in an annotation platform creates chaos; it is now imposing guardrails around supported workflows, favoring senior, high-agency hires, and treating the preservation of human decision-making as essential even when AI performs execution.
Key Takeaways
- Claim: Open-source model gains do not inherently cannibalize Mercor because human data is valuable at the moving frontier of desired capability, not for tasks models already perform adequately. | Evidence: Nitsky says each customer buys training and eval data to close a specific capability gap; Kimi K3 or another open model merely means customers should not pay for data to solve what that model already does. He disputes the premise that 90% of enterprise workflows are currently covered because it overlooks unattempted long-horizon use cases such as an agent autonomously operating procurement for months. | Implication: For agent businesses, model commoditization is not sufficient reason to dismiss proprietary eval/data infrastructure. The value shifts to defining and improving the remaining business-specific capability frontier. | Caveat: This depends on a continuing supply of economically valuable unsolved tasks; it is a strategic assertion from a data vendor rather than independently validated market sizing.
- Claim: Binary task-completion percentages are the wrong metric for many high-value workflows; some work has uncapped quality rewards rather than a clear sufficient threshold. | Evidence: Mercor's Apex benchmarks put top models at roughly 50% on long-horizon workflows. Nitsky contrasts sufficiency tasks such as updating a CRM with legal arguments and medical advice, where output can continually improve and there is no meaningful "done" threshold. | Implication: Ken should distinguish automation markets that can saturate quickly from domains where evaluation quality, reliability, and optimization remain durable spend categories. | Caveat: The Apex benchmark methodology and the definition of a long-horizon workflow are not described in the transcript.
- Claim: Enterprise-specific evals and training data are likely prerequisites for specialized models, because what counts as better performance differs by company and objective. | Evidence: Nitsky agrees with Fireworks founder Lin Qiao's prediction of specialized models per company, arguing that each such model needs data showing how it performs in that enterprise setting. He later characterizes eval sets as both the PRD for what a customer wants and the optimization objective for improving the model. | Implication: A defensible agent control plane should make enterprise evaluation assets first-class: task suites, rubrics, failure taxonomies, and outcome metrics are more strategically durable than a generic model wrapper. | Caveat: The extent of specialization will depend on whether the incremental ROI justifies the cost; Nitsky explicitly says only an increasing subset of use cases will clear that bar over time.
- Claim: RL environments are the fastest-growing data category because training agents requires realistic simulations of the tools and state they will encounter in production. | Evidence: Mercor describes an environment as a simulation of an application, a rich initial world that may contain hundreds or thousands of files, and tasks requiring the agent to use those tools. The Salesforce example is explicit: an agent learning Salesforce needs a high-fidelity mock that behaves like Salesforce during training and evaluation. | Implication: The implementation bottleneck for agents is likely to be environment construction, state management, task design, and reliable scoring—not only model selection or prompt engineering. | Caveat: Nitsky acknowledges that "RL environments" is a hype term with inconsistent definitions across companies, and he cannot disclose specific customer work.
- Claim: AI raises engineering output but makes product discipline more important: teams should reduce product surface area and move the bottleneck toward customer understanding and business judgment. | Evidence: Mercor saw people push multi-thousand-line PRs and rapidly expand the product because features became cheap to build. Its annotation product became overly flexible while supporting hundreds of heterogeneous projects; Nitsky says the company should have imposed guardrails earlier. He expects a higher PM-to-engineer ratio as coding agents reduce engineering bottlenecks. | Implication: For Ken's agent systems, enforce a narrow supported-workflow contract, explicit exception paths, and product-level constraints before adding configurable capabilities merely because agents make them inexpensive to ship. | Caveat: Higher PM leverage only helps if PMs can make sound prioritization decisions; more rapid execution can otherwise amplify product complexity and bad bets.
- Claim: AI services and forward-deployed engineering are a near-term deployment mechanism, not necessarily a permanent substitute for product. | Evidence: Nitsky says agent deployment skills are concentrated among a relatively small group, particularly in San Francisco, and most enterprises cannot yet hire the requisite internal talent. He expects AI deployment knowledge to diffuse and products to become more self-sufficient over a potentially decade-long transition. | Implication: A services layer can be strategically justified when it produces reusable deployment knowledge, eval templates, and product abstractions. Treat pure bespoke delivery without a productization path as a risk. | Caveat: This is a long-term forecast; services can also persist where workflows remain highly bespoke, regulated, or operationally complex.
- Claim: Mercor's revenue concentration in frontier labs is a known risk, and its proposed remedy is self-serve, lower-touch human-data projects for enterprises. | Evidence: Nitsky says the company wants every enterprise to efficiently run eval and training-data projects, supported by AI project managers. Today, lab projects are white-glove because customer intent, guidelines, edge cases, expert feedback, quality assurance, and changing data formats require fast coordination and "paranoia" from operations. | Implication: For an investment view, concentration should be evaluated against concrete evidence of repeatable enterprise onboarding, lower implementation labor per project, and whether AI project management genuinely resolves edge-case coordination. | Caveat: The same complexity that creates Mercor's current value may make self-serve democratization difficult; the interview does not provide adoption, retention, or unit-economics evidence for the enterprise expansion.
Detailed Brief
Mercor's operating model: supply quality, project economics, and competitive dynamics
- Claims: Mercor treats expert experience as a supply-side moat rather than simply paying the highest short-term rate.; The company believes data spending is relatively resilient when a dataset directly improves an evaluation target tied to customer revenue.; Small specialist data vendors may win narrowly scoped work but struggle to scale throughput, operational systems, and quality management beyond founder-led production.
- Evidence: Nitsky attributes marketplace supply to timely and transparent payment, dignified and interesting work, visibility into future projects, skill development, strong communication, referral incentives, and a sourcing team able to find rare skills globally.; He says project costs are increasingly split between payments to human experts and LLM spend for synthetic-data generation and automated quality control.; His characterization of the competitive field is a cottage industry of founders personally creating annotations, often subsidized by venture capital; labs benefit from the low effective price, but vendors encounter limits when customers want 10x more throughput or projects.; Mercor uses vendor bake-offs and customer-shared competitive intelligence as feedback on where its service quality must improve.
- Caveats: Claims about Mercor's cash generation, market leadership, and competitors copying its releases are management assertions with no audited financials or external customer evidence in the interview.; Data projects whose value does not map clearly to measurable model or revenue gains may be more price-sensitive than the speaker suggests.
- Implications: Assess data vendors as managed-marketplace and operations businesses as much as model-data businesses: expert retention, project management, QA, and rare-skill sourcing determine whether they can fulfill frontier demand.; A data provider's strongest pricing position is where it owns a credible link between an eval metric, an intervention dataset, and an economically important customer outcome.
Hiring, product management, and decision quality in an AI-first organization
- Claims: Mercor is biasing toward people who can connect work to business outcomes rather than candidates selected primarily for familiarity with many software tools.; It has moved away from extensive take-home assignments: it uses one artifact-building exercise with an agent to establish AI fluency, then emphasizes whiteboarding, experiment design, statistics, and systems design.; Nitsky's operating principle is to delegate execution to models but not judgment or decisions, because over-delegation degrades human reasoning and creates false confidence.; The company prioritizes agency and ownership over polished interpersonal traits, arguing that communication habits can be improved more readily than intrinsic motivation.
- Evidence: He says Mercor has more than 10x'd headcount since he joined and employs roughly 500 people while trying to retain startup behaviors: paranoia, office presence, and speed.; Product areas are organized as small pods: roughly two to three PMs per area, dedicated data scientists, and a small design function that flexes between teams. The marketplace and "studio" annotation/eval platform have separate product teams.; Each pod plans sprints independently but participates in a weekly product-area meeting to align roadmaps and maintain accountability; the scaling challenge is increasing cross-product communication as organizational nodes multiply.; For talent, Nitsky says the preferred seniority range is not former executives but hungry operators who have done the target job for several years and are entering their prime, which he places around ages 25 to 35.
- Caveats: The speaker does not provide measurable evidence that the proposed PM-to-engineer ratio or interview process produces better outcomes.; An aggressively performance-oriented culture can create retention and coordination costs if high agency is not balanced by clear ownership boundaries and psychological safety.
- Implications: Design agent-enabled teams around explicit human decision rights: humans select goals, success metrics, experiments, and acceptance criteria; agents execute, summarize, and generate alternatives.; As build capacity rises, measure product organizations on validated learning and business impact rather than feature throughput or volume of AI-generated code.
Future opportunities: cyber and robotics
- Claims: Cybersecurity is structurally attractive for agent/eval data because it is adversarial: offensive and defensive capabilities continually move the target, preventing simple task saturation.; Mercor expects real-world physical data and robotics to become material over the next three years, but Nitsky expects adoption to look more like the gradual city-by-city spread of robotaxis than the immediate software distribution of ChatGPT.
- Evidence: Nitsky describes security as an AlphaGo-like competitive setting with uncapped performance rewards and says demand for cyber-defensive capabilities through new data types is increasing rapidly.; He says the company sometimes retains exceptional talent ahead of anticipated demand and can create off-the-shelf datasets during low-demand periods for later resale.; In defending robotics progress, he compares current constrained demonstrations to the long period in which self-driving cars still carried human operators before Waymo became a commonly used service in San Francisco, Austin, and Phoenix.
- Caveats: He cannot disclose the specific cyber data types customers seek, limiting the actionability of the claim.; The robotics thesis is speculative and explicitly includes physical-world scaling, distribution, and regulatory constraints.
- Implications: Cyber merits attention as a domain where continuous evals, red teaming, simulation environments, and adaptive defenses may support recurring data demand.; For robotics, favor businesses with credible physical-data acquisition, environment fidelity, and deployment pathways rather than relying on demo-driven narratives.
Notable Concepts & Terms
- Apex benchmarks: Mercor's referenced measurement of long-horizon workflow performance; Nitsky says leading models are approaching about 50%, though the methodology is not explained.
- Sufficiency-based workflow: A task with a practical completion threshold, such as updating a CRM, where additional model quality has little incremental value.
- Continuous uncapped rewards: A framing for domains such as legal reasoning, medicine, and cyber where performance can continually improve instead of passing a binary completion bar.
- Evals as PRD: The idea that an enterprise evaluation set specifies what the business wants a model to do and simultaneously defines the target against which it should be optimized.
- RL environments: High-fidelity simulated applications, initial states, files, and tasks used to train and evaluate agents in conditions resembling real deployment.
- World: Mercor's name for the rich initial state within an RL environment, potentially including hundreds or thousands of files analogous to the data on a user's machine.
- Forward-deployed engineer (FDE): An engineer who works closely with customers to deploy AI systems; Nitsky views the role as important while enterprise AI expertise remains scarce.
- Talent-only versus managed service: Mercor's two delivery modes: supplying experts for customers to operate themselves, or staffing experts through Mercor's platform and delivering the completed dataset.
Operator Notes / Why Ken Should Care
- Build an explicit eval asset strategy for every agent workflow: define production-like tasks, initial state, edge cases, scoring criteria, and the business outcome each score is intended to predict.
- Audit current agent products for uncontrolled configurability; identify workflows that should be unsupported, templated, or routed into an exception process rather than generalized into the product.
- Separate AI spend accounting into R&D, internal productivity, and customer-facing unit economics. Require outcome metrics for the latter instead of applying a single organization-wide token-spend percentage.
- If using services or forward deployment, mandate a productization artifact after each engagement: reusable connector, environment template, rubric, runbook, or automated project-management step.
- Evaluate data-vendor concentration risk using leading indicators: percentage of revenue outside frontier labs, implementation labor per customer project, self-serve completion rates, and repeat enterprise project volume.
- For hiring, test candidates on experimental design, systems reasoning, and decision quality under AI assistance—not merely on whether they can produce polished AI-generated artifacts.
Source/Metadata
- Title: Mercor Head of Product on Revenue Concentration from Frontier Labs
- Transcript words: 11606
- Duration seconds: 3762
- Timestamp note: No usable timestamps or chapters were provided. The transcript also contains repeated passages and a long extraction artifact near the end.
Transcript
Even Figma, we're moving away from it in favor of cloud design more and more. We end every week with so much more money in the bank. The business is very healthy, and we can't spend money fast enough to surface all of the demand that we have. Today we have Oswald Nitsky, CPO at McCaw, in the hot seat. Today, it's a really open conversation in a way that I don't think has been had with someone from McCaw before about what happens if Frontier Models actually do what the data providers are going to do. How does synthetic data cannibalize their business? Does open source help or hurt data providers because their biggest customer? Oh yeah, it's the Close Frontier Models. This and so much more in our conversation with Oswald today. Get a real internship as soon as possible because whatever you learn in school is probably going to be updated quickly. I don't think there's an ROI problem right now. I think we're in a period of... Ready to go? Oswald, it is so good to have you on the show, dude. I've heard so many good things from Brandon. So thank you so much for making this happen, man. Thanks for having me. Super excited. Dude, I am seeing open, open, open. Everyone claiming that we will see the mass migration from Frontier closed to open. Kimi very recently came out with their new model. And I literally, I didn't really know. Does open cannibalize McCaw's core business? I wouldn't say that open source model improvements cannibalize our core business because data is most valuable on the frontier of model performance. And so each of our customers has their own unique goals and is purchasing eval and training data sets to fill gaps in current model capabilities. Open models just raise the floor of what people are interested in. As long as customers still have new capabilities that they want to get better at, our business still continues to grow. Open source models just mean that nobody is buying anything that Kimi K3 can already do. So if 90% of enterprise workflows can be done with open models, which more and more people say they can be, and that 10% is really where you serve your customers and provide data, I'm naive. Does that not make it harder and harder to make huge amounts of revenue if that 10% in Frontier moves further and further away? I'm not convinced that 90% of enterprise workflows can be handled by open models or Frontier models right now. We think that these calculations might be based off of existing demand or things that come top of mind when current model users are thinking of what models can do, but there's a whole category of latent demand that people aren't even... These are things that people aren't even trying to do with models yet. Most commonly, we think these are long-horizon tasks, like setting up a procurement agent to fully automate your procurement team for months on end. You only check on it maybe once a week. We think that's just not even captured in these calculations when someone says enterprise workflows are being handled because nobody's trying to do these things yet. The market for data to support those use cases is growing, and that's where we see a lot of the leaders moving to. Okay, so we see a lot of leaders moving there and seeing new capabilities that they never thought existed. But then we have Alex Karp, and I thought a rather sedate performance. Normally, he jumps up and down much more, but it was still rather energetic. He said about the incredible skepticism we see from large enterprises towards data and sharing data with the Frontier model providers. To what extent do you see skepticism and fear from large enterprises in working with Frontier model companies? We see it depend on the specific workflow and how core it is to the business. Things that are just general things that every company needs to do, like HR, procurement, can be less sensitive, and enterprises are more open to putting these workflows on proprietary models. It's the core work that the company is doing that's vital to its business, that differentiates it from competitors, where we see more sensitivity. So you can imagine this being the actual legal services that a law firm provides. Like what are the actual memos that it's writing? What is the advice that it's giving to its clients? Am I the only one who sees the irony in we put the sensitive data on open source, most likely Chinese models, and we put the HR and procurement data on the closed model? Am I a moron? Well, it depends on where you run the open models, right? Whether or not that's a bad idea. So the beauty about open weights models is that the inference can happen in multiple places. So you could make mistakes using them, but you have more control. When you look at that dispersion, what do you think is inaccurate? You said you don't really believe the 90-10. What do you believe a more accurate representation is? Well, in our Apex benchmarks, we're getting closer to around 50% of long-horizon workflows. Top models are scoring around that much. But I think that there's a class of workflows that are just sufficiency-based where you do it and it's done and you're good. This is something like updating a CRM. You can't really get much better at it. And then there's a class of workflows that we shouldn't even be thinking about in terms of binary, like can the models do it or not? And these can be things like legal arguments or, to an extent, medical advice where you could always get better. And in those cases, I think that the percentage framing is just totally off and we need to be thinking more about continuous uncapped rewards. When we think about it could be better, I had Lin Qiao, the founder of Fireworks, on the show the other day. And she was like, exactly that is why we'll have specialized models for every single company. Because it could be better depends entirely on the company. One company wants to focus on growth. One company wants to focus on margin. Another wants to focus on, I don't know, if we're in Europe, work-life balance. And so you need individual specialized models for every company. Do you buy that we will have specialized models for every company? Or is that a little bit self-serving towards Fireworks? I buy it. I think it's also self-serving towards Mercore. In that we think that every specialized model will need enterprise-specific eval and training data to show the model how it's performing in its setting. And I think it depends on... I think the diversity and the market for this depends on the value that customers can get from these specialized models. So there will be cases where the ROI is really justified. And I think that those cases will increase over time. But we certainly believe in this future. Switching back, Alex Karp's second point in that show was ROI questionability. You mentioned the word ROI. That's what made me think of it. It's very present. And enterprises maybe have questionability around the ROI that they're getting. Do you think we have an enterprise ROI problem with AI today? I don't think there's an ROI problem right now. I think we're in a period of exploration and experimentation where there's more tolerance, more patience to get that ROI calculation right now. There's a lot of different projections around where token prices will go, where performance will go. And right now, we're starting to see some amount of tightening of the screws on spend here and there. But I think the paradigm we're in is still let's see what happens because things are moving so quickly that the ROI calculation might shift too dramatically still. Two ways I want to go in this. I'll take the first way. We saw Aaron from ClickHouse say that he's 6x to spend. And that's what they need to do because we need to be at the frontier. And then you see Uber and Microsoft and some forms of, I think it was Grok or X or one of Elon's companies, put budgets on per-user head. What do you think is the right way to be navigating this cycle? If I'm a founder listening, what would your advice be on how I should think about optimizing the balance between performance and budget? It totally depends on the use case. So I've mostly worked at hyper-growth companies where growth matters at all costs, right? And there is your willingness to spend for growth as long as the unit economics are fine. When you're looking at coding agent spend for your software engineers, that's not always cogs for your work. If that's really high, that could still be giving you compounding gains. If you're looking at a customer service agent that has massive token spend and the revenue you're getting from the customers being served is way lower than the token spend, then you're definitely in a bad position. But my experience has been just these growth-stage companies. And I think for a lot of founders, considering their token spend, if it's for growth, if it's for improving the efficiency of your headcount, that's just what you need to do to service large amounts of demand when you're starting up. Mr. Benioff from Salesforce said that he spends 300 million a year on Anthropic, which works out to be about 3.8% of developer salaries if you average the salaries. Do you think that is the going rate moving forward? Do you think that will be 20%? Or do you think it'll be 100%? Or will it be way less? If you're looking at a customer service agent that has massive token spend and the revenue you're getting from the customers being served is way lower than the token spend, then you're definitely in a bad position. But my experience has been just these growth-stage companies. And I think for a lot of founders, considering their token spend, if it's for growth, if it's for improving the efficiency of your headcount, that's just what you need to do to service large amounts of demand when you're starting up. Mr. Benioff from Salesforce said that he spends 300 million a year on Anthropic, which works out to be about 3.8% of developer salaries if you average the salaries. Do you think that is the going rate moving forward? Do you think that will be 20%? Or do you think it'll be 100%? Or will it be way less? I hope that we can move towards a future of better accounting of the outcomes being driven by token spend. Because even here, I think in a company like Salesforce, we have so many, a company of that size, certainly you should have different spend profiles depending on what the team is doing. Again, here you have teams that might be more like solutions engineering or forward deployed, where you have to think in terms of unit economics. And teams doing R&D, where you can have more tolerance for spend. So I think at the large companies, you have to consider which parts of your organization are doing what and how much tolerance you should have in different areas. I think, macro, the percentage will increase over time to more than 3%. Huh. Do you want to hear something funny? Brandon said on the show that it would hit 100%. And he said that you already spend more today than you do on salaries. Yeah. Yeah, yeah. We do. And 100% sounds reasonable. So, as I said, I've only worked at hyper-growth companies, and that's what Mercore is and continues to be, more so every day as the growth just accelerates. And for us, it makes sense because the demand that we have is so high. The company is more than 10x in headcount since I joined. The revenue has also commensurately increased. We're just in a race nonstop to service our insatiable customer demand. So for us, it makes sense because we can't spend money fast enough to service all of the demand that we have. Dude, do we just build 10x more products quicker? Help me understand. Do we have smaller engine product teams? Do we just build much more than we ever used to? How do you think about that? I think this paradigm makes the job of product management a lot harder because we're trying not to build 10x more product surface area. It makes things incredibly chaotic. We have moments in time where product surface area rapidly expands because people think, oh, I can make all these features really quickly. This is like, let me just push these multi-thousand-line PRs. But we're constantly in this battle to try to simplify our product surface area and find the interactions and the workflows that are most scalable. So the trend that we see is we're, as a product team, constantly fighting to reduce surface area and simplify things. And we also see a higher ratio of PMs to ENG because engineering is less bottlenecked. So there's much, much more work to be done in understanding the workflows of users, the needs of users, and what products actually drive revenue the most becomes the bottleneck now to servicing more demand for us. If we think about the pre-AI era, how has what it takes to be a great PM changed for this new world? There's two major changes. One is that you don't really need to learn as many tools anymore. You just have to be able to use coding agents. A couple tools will do everything you need. Even Figma, we're moving away from it in favor of cloud design more and more. So less tool diversity for us. And then the other is everyone needs to up-level a lot and think about business impact much more. I think that all work is starting to look like higher level. So the minutiae and the details get sorted out way faster. And all the PMs at Mercor have to think way more about, is what I'm focusing my time on the right thing? I can do things very quickly now. Skill issues have almost gone away. So now it's all about judgment. And am I doing what is going to drive the most business value? Dude, I have to ask. You said there a core job is retaining simplicity and deciding what to do versus what not to do. What did you do in product that, with the benefit of hindsight, you wish you hadn't done? And what did you learn? So one interesting thing that happened this year was our annotation platform serves a lot of different workflows. And the demand for human data is so large and it's so heterogeneous that, and our delivery team is so good at delivering projects and selling projects, that we supported, I think, too many workflows for human data projects. And we built a tool that was extremely flexible in supporting all sorts of different research experiments that customers might want to do. So the shape of data has changed a lot since it started with InstructGPT for Gen AI, from supervised fine-tuning to preference ranking to all these environment-type projects. There's a lot of multimodal projects that have totally different formats. And your annotation tool needs to support these and different workflows. And customers will ask for all sorts of stuff. We tried to serve every ask. We made a tool that's maximally flexible, has all sorts of... We had hundreds of different projects running on it. That's just chaos to manage. And what we needed to do sooner was to put guardrails on the type of services that we support and work closer with our operations team to say, hey, here's the best practices. Customers are going to ask for everything. We can do it. But should we do it? If there's no enduring demand for certain workflows, maybe it's not worth the investment. So putting guardrails, narrowing down the services that we support, was something we should have done a lot sooner, though we did it recently. How do you determine enduring demand? This is what makes Mercore a hyper-growth company, is that we're incredibly tapped into the market and the ecosystem. It's really judgment from leadership, I think. It's very hard to say what will data look like in a year or two. And the best way to figure it out is to stay in constant touch with leaders from a diverse set of labs and constantly be validating hypotheses. I think Brandon does it very well. I think our operations team does it very well. But ultimately, it's a guess. Which lab has the most advanced and sophisticated data team? I can't speak too much to customer details, but they're all super good. Everyone is sophisticated. Everyone blows me away in different ways. That is such an unfair question. Okay, I totally agree. The other question to ask is which has the worst team? No, I'm joking. My question to you was, you mentioned another element, though, which actually didn't shock me, but I thought it was interesting, was the movement away from Figma. Can you talk to me about that? Because I hear more and more companies doing the same. As a product leader today, how do you think about that? And what was the thinking there? The team do whatever is best for them. And this is a trend I've just observed amongst almost everybody, is that cloud design has done a great job. People really like using it. It's easy to use. It's easy to use. And we've just had a natural movement towards it. It's also a bit easier to not have too many tools, not manage too many licenses. And because cloud is making all these other great features, people just gravitate towards it. And then it's a bit less friction to have the procurement team issue licenses for Figma for every single person. We were talking about ROI earlier for enterprises. And we're seeing Microsoft set up a services department. We're obviously seeing Palantir skyrocket, and services becoming an increasing part of everyone's business. Is that the future of AI enterprise deployment? And how do you think about the incredible rise of services in deployment? Yeah. So I have a bit of a hot take here. I think it's the future for the short term, as the knowledge of how to use AI gets disseminated throughout industry. We have basically a concentration of a bunch of people in San Francisco who really know how to deploy agents, eval agents, be AI-first in engineering and in other areas. And that knowledge just isn't out there yet. And eventually it will be. And maybe you won't need, at that point, teams to go and set things up, set up AI agents for every enterprise. And it'll become more of a job function similar to software engineering. And so in the short term, it enables deployment. In the long term, products become more and more sophisticated, and they're able to do it themselves. Because Matan from Factory said to me, you know what? Fuck this. Services, they're just an excuse for crap product. I think that it's a knowledge dissemination problem. So I think that that's one way to look at it. I think it's the future for the short term as the knowledge of how to use AI gets disseminated throughout industry. We have a concentration of a bunch of people in San Francisco who really know how to deploy agents, eval agents, be AI-first in engineering and in other areas. And that knowledge just isn't out there yet. And eventually it will be. And maybe you won't need, at that point, teams to go and set things up, set up AI agents for every enterprise. And it'll become more of a job function similar to software engineering. And so in the short term, it enables deployment. In the long term, products become more and more sophisticated, and they're able to do it themselves. Because Matan from Factory said to me, "What? Fuck this. Services, they're just an excuse for crap product." I think that it's a knowledge dissemination problem. So I think that that's one way to look at it. The other way is, why not hire someone to just do this agent deployment at your own company? And I just don't think the skill is out there yet. I don't think there's enough. I don't think the talent is available for every enterprise to have their own expertise in it at this point in time. But that'll change over the long run. This is, I think, maybe a decade-long change. Question. Do good engineers really want to be FDs, though? I think there are a lot of different types of good engineers. There's a lot of ways to be a good engineer. And one way to be a good engineer is being a great communicator and cutting through to the source of a problem and simplifying. And I think that those engineers are great fits for FDs. And I think that those engineers are also great fits to eventually become founders. And I think that that is a different profile of person who's incredibly valuable. And that's what a lot of people are looking for when they're looking for FDs. And it's also a profile that we look for generally, which is why we have so many alumni go off and start companies. Do you like that? I suppose, Brandon, about this, but is it a good thing to have the McCall Mafia? Because you also want to retain talent. I'm proud that, of the people I work closest with on my teams, I've only had attrition to founding. And we've had quite a bit of it. It's a lot better to lose someone to starting a company than to taking another job. It's interesting from a personal level because I like these people. I wish the best for them. I really enjoy seeing it. It is tough, though. It makes the job of management a lot harder because we just have so many high-agency people who are very ambitious. And it's difficult, but I like it. And I'd rather be in an environment like this than one where everyone's soft and, "I don't want to work." Oh no, I'm just kidding. No wonder you left Europe. How has hiring changed in a post-AI new world? When you look at the people that you add to your team today, especially in product, what do you ask today or look for today that you didn't before? I think, touching on the earlier point of everybody needing to up-level and think closer to business impact, we've biased towards more senior hires who are better at understanding what drives the business forward, really grokking how we operate, how we make more revenue, how we deliver better services to our customers, how we keep our customers happy. I found that more senior candidates just get that a lot faster. And like I said, all of these tool, "Can you use the tool? Can you do all these other more junior things?" are becoming less relevant. So the hiring for us is biased towards more senior candidates. Do you worry that you're just falling for the classic, I'm so sorry to be in fast-growth founder mode, which is your VCs come in and say, "Oh, you need to hire this person from Facebook," and you get the seasoned operator who fits exactly that rubric, and it never works? It never works. We're not quite doing that. Seasoned here is a spectrum, right? I'm not saying we're hiring people who are formerly in executive positions. We are treating everything as an executive search where we want to find someone who's at the sweet spot. They're still hungry. They've done the job that we want them to do for a few years. And they're right in really hitting their prime. So that's— When do you think people hit their prime? I think 25 to 35. Oh, I just turned 30. I'm bang in the middle. Perfect. Good timing for you. Perfect timing for me. OK. In terms of the questions, what we look for in the take-home assignments, has that changed? We've moved away from take-home assignments. We do one take-home assignment, which is, "Can you just use an agent to go—you’re on your own for a bit of time—go use an agent, produce this artifact for me, and we'll look at it?" Do that once. You know that the person's AI-fluent. And then we move towards a lot of whiteboarding because we want to avoid— Like, we will do one round where we know, where we find out if the person is just familiar with AI tools. Don't laugh. OK, so cool. We do that. I'm familiar with AI tools. And now you're like, "Come into my room. We've got a whiteboard." What do you want? What are we going to do? What do you want to see? What would impress you? We care a lot about being able to set up good experiments and understanding statistics, having good judgment, and systems design as well. The reason is these are just skills that are so easy to be bad at. I'm so sorry. I'm so sorry to interrupt you. Good experiments and systems design, it feels quite wordy. What does that actually mean? We ask people— Well, I don't want to give away too much about our interview process. But we need to run a lot of experiments as a product team. We need to make sure that our team knows how to run a good experiment that actually reveals information and isn't just totally fudged. And with AI tools, it's very easy to offload a lot of thinking, a lot of judgment. We want to make sure that people still have the ability to have good judgment and know what they're doing and not just regurgitate what comes out of cloud. That's so interesting. I completely agree with you. I have it with my team, which is we do scripts for content, for reels. I do all questions myself. I would never use AI, and I'm very concerned about it because you lose the muscle to me. Can I ask you, how do you retain thinking, thought, creativity when so many people are so freaking hooked already? I think it's kind of—I tell my team it's kind of like phones. They fry your brain and they turn it into goop. But I do a lot of stuff on my phone. I use my phone all the time. You just have to learn personally where that boundary is of when is a good time to scroll through reels and when's a bad time. In a meeting, try not to scroll through it. For work, that boundary, I think, is between the judgment and decision-making and the execution, right? So I want very carefully never to delegate judgment or decision-making to models because it makes you think that it's doing the right thing. But you have to be paranoid with them still, right? You still have to double-check everything. And that's what I tell my team, that don't delegate your decision-making, your actual job, to a model because you're going to lose that ability. And then you're going to get psychosis. Totally agree with that. When you look at the experiments that you've run, does the data correlate to the outcome? I often think in investing, sometimes I do no work and no diligence and I make loads of money. And sometimes I do lots and I make terrible investments that lose all the money. Do the inputs correlate to the outputs? It varies. It varies because we run a lot of experiments. But sometimes they do, sometimes they don't. We want to get more that actually show good results and move the business forward. And that's really the job of the team, is to find the right experiments to run and make the narrative around these changes to our product having impacted the business in a positive way. That's a lot of the core job right now. So it's week to week, month to month. We get different results, but we try to trend in the right direction over time. I often think in investing, sometimes I do no work and no diligence and I make loads of money. And sometimes I do lots and I make terrible investments that lose all the money. Do the inputs correlate to the outputs? It varies. It varies because we run a lot of experiments. But sometimes they do, sometimes they don't. We want to get more that actually show good results and move the business forward. And that's really the job of the team, is to find the right experiments to run and make the narrative around this. These changes to our product have impacted the business in a positive way. That's a lot of the core job right now. So it's week to week, month to month. We get different results, but we try to trend in the right direction over time. And people start to learn the dynamics of the product, learn the dynamics of the user better and better to improve over time. When you think about product and running experiments that you mentioned there, and running good experiments, how do you structure the teams today? And what does that meeting look like? We have a few different groups that do experimentation. So we have two major product areas where this is most relevant. Our marketplace, which matches experts to jobs. And our annotation and eval platform, which is where experts log in to do annotation for eval or training data sets. Where our operations team also logs in to run those projects. And our customers will log in to see their data and run evals. So annotation platform, we call it studio; marketplace, call the marketplace. These two groups, they're kind of self-contained in trying to make their individual product offering better. And we have two main modes of engagement within human data. Talent only, which is when we just send experts to our customers and they'll run the project. So this is a lab that needs a doctor or a lawyer or whatever. And they're like, we're just going to use them. Thanks for finding the best person for the job. You'll need to pay them, performance-manage them, but we'll run the project. And then a managed service project where we give our customers data. So for the talent-only model, we just use the marketplace. For the managed service, we use the marketplace to send people to our annotation platform. And then we'll run the project and give them the whole data set. Totally. Are they two separate product teams? There are two separate product teams. How big are the product teams? Around two to three per product area, with also data scientists and a design team, data scientists dedicated to each and a design team that flexes between them of just a few. So you have pods of four or five? That's fair. Yeah. Got you. Totally. Okay. That makes absolute sense. Will those ratios change over time? Do you think between PMs and design, or will that stay the same? I think that the ratio of PM to ENG will change over time to have fewer engineers per PM, as engineering velocity increases with better coding agents. And we will be bottlenecked by understanding business needs, user needs, and that's more of a PM job. We need to be very careful as a hyper-growth company to grow the teams in lockstep because, as the headcount has increased more than 10x in the last year, we just want to be careful not to grow one faster than the other. The trend will be to a higher PM to ENG ratio, though. So when we talk about the good experiments and making sure that we're running a really tight process, what does that look like in terms of the meetings? You have a weekly product team meeting. What is the right way to approach cadence of product team meetings and how to run them today? We break it down into, so within these product areas, we'll have a whole PA weekly in this product and ENG and a lot of other stakeholders as well. And this one is just run, kind of like, it's broken down into pods. So as I mentioned, within that product area, there might be, let's say, three product managers that all have a pod of these parts of the product that we can naturally segment work into. Our marketplace, for example, has an expert-facing side and a hiring-manager-facing side. These are naturally two distinct pods. There are some other pods within here as well, like managing the expert experience, making sure that everyone has great customer support. There's never any issues with working firmware core. Each of these pods will do their own sprint planning. They'll come together in the weekly product area meeting. And we try to keep it efficient, but maintain a lot of visibility between the pods because they all need to have their roadmaps well aligned. But we need to have weekly meetings to maintain accountability. So we do them on Friday, a bit later in the day, make sure no one's leaving early on the weekend. Love it. What do you not do in your product meetings that you should do to make them better? It varies by product area. So the challenges in the marketplace versus the annotation platform are a bit different. The main challenge as we grow quickly is having the right amount of communication and feedback from other teams. So our marketplace and our studio team need to get information from each other, right? There are cases where something's wrong in one and it's popping up in the other. Something's wrong with one product and it's affecting the expert experience when they're on the other one somehow. And that communication just, because the headcount and the team's grown so quickly, the communication channels just explode very quickly. So we need to do more cross-product-area collaboration. Keeping it efficient is just really hard as the team grows because the nodes just keep moving around and there's more of them. What has been the secret to scaling supply on the marketplace side so efficiently? That's fucking hard. How have you guys done that so well? I'd probably put it down to three things. The first one is a great expert experience. So experts get paid on time. They get paid well, transparently. Everybody involved in what the expert experiences cares deeply about whether or not they're having any challenges and whether or not the work is dignified and well paid and fairly paid. And that is a requirement for a great referrals program because nobody's going to refer their friends, their colleagues, to some kind of job that sucks, right? So everybody caring about expert experience drives a great referral program and, additionally, a great sourcing team that's able to find people in every corner of the world with very specific skills helps us fill the gaps when we have spiky demand for a specific skill set. But are people as short-sighted as just wanting to be paid the most? I've heard that McCaw pays the most. Is that just a secret? I wouldn't say it's short-sightedness because we want to retain the top experts as well. Right? So if you get paid a lot on one project and there's some kind of crazy bonus payouts and stuff for short-term sprints, that doesn't get you to come back as much as a great experience with a lot of work. Visibility into what future work is coming up. The feeling of, I'm growing my skill set. I have the ability to pick between a few different jobs. I'm doing interesting work. I have great communications from the people running the project. It's really hard to sign up for online work and then get hit with this 100-page instruction document. It's a very foreign kind of job. That's part of the experience as well. And knowing that you're going to get paid highly for a long time for something that you can do for a long time is what keeps people interested. Have you seen your margin improve over time, or is it one where actually margin is relatively fixed given the complexity? So, margins are an interesting thing in this business. We try to think about, as a product team, how do we deliver the best value for our customers? And that is independent of how do we price the project. So there are cases where you could have automatic quality control and synthetic data improvements to make the delivery better. There are situations where you could think about the staffing on the project to change the cost of the service. All of that, as a product team, we want to make sure that we can deliver the best value to our customers and we can have the best experience for our experts. Margins are decided after the fact based on consideration of costs. And now, for a lot of these projects, the costs are driven equally from paying experts and LLM spend on things like synthetic data and automatic quality control. It's not revenue, it's not revenue, shouting from the crowds, throwing peanuts. Does that annoy you? And is there anything there that hasn't been said that you think people are just not getting? We try to think about, as a product team, how do we deliver the best value for our customers? And that is independent of how do we price the project. So there are cases where you could have automatic quality control and synthetic data improvements to make the delivery better. There are situations where you could think about the staffing on the project to change the cost of the service. All of that, as a product team, we want to make sure that we could deliver the best value to our customers and we can have the best experience for our experts. Margins are decided after the fact based on consideration of costs. And now, for a lot of these projects, the costs are driven equally from paying experts and LLM spend on things like synthetic data and automatic quality control. Does the it's not revenue, it's not revenue shouting from the crowds throwing peanuts, does that annoy you? And is there anything there that hasn't been said that you think people are just not getting? It doesn't annoy me, no, because we end every week with millions more in the bank. It's funny how you can have... I've been at other companies where I've seen interesting financial engineering and accounting. People can have all these different metrics, but we end every week with so much more money in the bank. The business is very healthy, and we can't spend money fast enough. So what people want to call it is up to them. But the cash flow is insane. Does it matter that you have such high revenue concentration? The frontier model providers, your biggest customers by far, some would say, whew, that's a lot of concentration. How do you think about that? So I can answer this from a how it affects the product team. Yeah. We would love to move down market. Our biggest challenge is moving down market so that every single enterprise can efficiently run human data projects for eval and training. And that'll diversify our revenue for sure because there are many more enterprises than there are labs. And that's a harder product to build. And that's the direction that we are taking our products, taking the company, is to be able to self-serve, run these projects very efficiently, have AI project managers so that it's a lot easier to do this work for smaller customers. Because running a human data project for a lab is incredibly hard. It's a white glove service that requires a lot of people on the operations team. As we make that more efficient with better products, better processes, we can do smaller projects that are more heterogeneous for more customers. It's the direction we have been heading, which has reduced concentration. And it's the direction that we'll continue to head as every enterprise begins to have human data work for their proprietary use cases. What's so hard about it? Making it really simple? Explaining it? What is the challenge with not dumbing down, but democratizing? Running a human data project is just hard. There is so much information that needs to be transmitted from the customers, the end users of our customers, to experts. And all the edge cases matter, right? So people will try to write a guideline that says, here's how you make a data point. But the experts will have some edge case that gets bubbled up. And what you do on that edge case matters a lot. So the process of making a human data project is basically continually surfacing these edge cases, which requires insanely fast alignment between customers, maybe their customers, maybe other experts in the field, and the experts who are doing the annotation. And it also requires a huge amount of paranoia from the operations team to make sure that every data point is perfect. It fits whatever guidelines the customers have. And the project's running on time. All the bottlenecks are removed. It's just an operationally intense process because it necessarily deals with edge cases and things that haven't seen before and are outside of model capabilities. The data types also change very frequently. So we've moved from supervised fine tuning to preference ranking to rubric-based annotation to now RL environments across a whole bunch of different modalities. There's a lot of complexity within each project and then between projects. So I would boil it down to those two things, the need for paranoia and the need for very crisp communication, that make it challenging. What data type is not hugely in demand today that you think will be hugely in demand next year? The data type that's growing the fastest for us is environments. You might have seen a lot about these RL environments on Twitter. It's a hype term. Every company has a different definition for it. But we are certainly the leader in the category and view it as basically these simulations of apps that you might want your agent to use. And also a rich start state, which we call the world, that is basically representative of all the data you might have on your machine, like your laptop. And then we have tasks that train agents how to use those tools to accomplish something that's useful. It's a bit of a complicated annotation process because the agent has to interact with this simulated world. We have to make that start state, which can be hundreds of files, thousands of files. And the shift here is that the data that the models are now, the agents are being evaled and trained on, looks a lot closer to what they see in deployment. Right? So if you want to learn how to use something like Salesforce, you need a pretty high fidelity mock that acts exactly like Salesforce in your eval and training. And it's complicated to get this set up. Just like years ago, preference ranking was really hard to get set up. SFT was really hard to get set up when InstructGPT first came out. So this is the frontier right now. Labs are figuring out. Neo labs are figuring it out. Eventually, it'll get so smooth that enterprises can do it too. Are labs price sensitive on data acquisition? By data acquisition? When they go on a project with you, are they price sensitive? Are they haggling, going, oh, well, Edwin at Surge gave me a 10% discount. Can I have that? Or are they like, just give me the fucking data? Well, there's always the aspect of negotiation and the procurement team trying to get a better deal. But we've chosen a great business where our work directly affects the business outcomes of our customers. So we have a great setup where if you're making an eval set in your lab, you're evaluating something that your customers want to do. If you could just do it better, you would make more revenue. If they're buying a training set, they're now hill climbing that eval set that they've said represents what their customers want to do. So as long as the amount of money they're spending on data is less than the revenue that they're going to get, they're happy to crank the lever. People want to crank it harder and harder because spend on record directly translates to more revenue for our customers. Do you think we'll have an unbundled data provider world? I'm a venture investor, and I see so many people that are like, oh, we're like, we're like McCaw, but for domestic robotics. And you're like, okay, cool. Good. Okay, I get it. But do you think we will see this kind of specialized data provider world where niches have thousands of players? To an extent, we're already in this world. I wouldn't... it's not that successful, though, for the small players always. So how I would describe it is we're facing what looks like a cottage industry of founders doing annotation themselves. Right? So you have all of these small startups where, as the skill bar for annotation gets higher and higher as models get better, you have startups where the founders are actually just making the data. Right? And labs love this because it's just totally mispriced. They get someone raises a bunch of money, they have loads of cash to blow, and they go to these labs and they're like, I need to get your business. Please let me work for you. And then they're smart people. They're founders. They're formerly great technical employees. But they're running the projects themselves. They're doing the annotation themselves. And this is just VC subsidized work that labs love. The problem is scaling it beyond a few data points or what one founder or full-time employees can do. And this is the position that we're in, is we're having to compete against basically founder-led annotation, where some of them are even running it as cash flow businesses and they're just taking the profits home themselves. It doesn't scale, though. And vendors, our customers know this, that it won't scale when you want to 10x the throughput, 10x the amount of projects. But it is indicative of the direction the field's heading in, in that we need higher skilled experts. We need the best people in the world to be doing the sanitation. Don't laugh. I have a bit of an ego. And so I like to feel like a special snowflake. And what I mean by that is I would be like, oh, when Meta or OpenAI or you name your large company is buying data from multiple people, it feels like you're being promiscuous and cheating on me. Do you mind? And do you monitor budget and percent of budget that gets spent with you versus another provider? Of course, we do a lot of competitive intelligence. It doesn't scale, though. And vendors, our customers know this, that it won't scale when you want to 10x the throughput, 10x the amount of projects. But it is indicative of the direction the field's heading in, in that we need higher-skilled experts. We need the best people in the world to be doing the sanitation. Don't laugh. I have a bit of an ego. And so I like to feel like a special snowflake. And what I mean by that is I would be like, oh, when Meta or OpenAI or you name your large company is buying data from multiple people, it feels like you're being promiscuous and cheating on me. Do you mind? And do you monitor budget and percent of budget that gets spent with you versus another provider? Of course, we do a lot of competitive intelligence. And our customers like us, so they'll often share information with us. But everybody just wants models to get better, right? So we're happy to have this kind of competitive pressure that tells us where to go. If someone else is able to do something better than us, we'd love to hear about it and then do it better than them. Right. It's healthy to have vendor bake-offs. It pushes us to make our services better. We do stay on top of it because we want to deliver better services to our customers. We want to know who's doing better than us, and then we want to surpass them. So it's a totally healthy thing to happen as long as Mercor is winning. You said models getting better there. We said frontier earlier. Yeah. I'm an investor in Lagora, and everyone's like, no, your real competition is actually Anthropic. And I'm like, if Anthropic go after legal and winning Cooley and Goodwin, something's gone very wrong with the world because they should be solving cancer and climate change. To what extent am I right, and how do I balance between Anthropic coming for Lagora and Figma and Anthropic also working on the frontier problems that humanity faces today? I would look to precedents from other big tech companies who've had a lot of different efforts, like Google, Microsoft, who coincidentally also try to solve climate change and cancer. But it's not their main business. And they have their hands in a lot of different areas. And they have their own business, but competitors still emerge. So you've seen, you remember Google Plus, right? Yeah. That didn't go anywhere, right? It probably maybe freaked some people out when it happened. You probably remember Threads. I don't know the current state of Threads, but the— Apparently 400 million users, according to their marketing team. That's very interesting. I won't comment too much on that because I— How fascinating. I'd love to see the engagement. I genuinely don't know anything about this. But yeah, so I think if you look to precedents here, large companies often try to make new bets, diversify, but they lose to companies that have intense focus on their market. So we'll see how it plays out. But I would wonder if there's anything to learn from history with Google and Microsoft having many business units, many efforts, but a core business that has driven all of their revenue. I love Brendan. I remember texting him when there was the hack. It's tough when there's a hack because you're like, I don't know what to say, but I'm here for you. Thumbs up. And I felt like such a VC because you're like, I'm here for you. Good luck. Fuck all help that is. My question to you: how did that change your mindset and approach to product? It's a really hard thing to go through. I remember you were under intense pressure and stress. No, seriously, I am sorry for that because it's horrible to go through. How did it change your product mindset? I'm not an expert in security, but we hired a lot of experts in security, and I listened to them. And that's the main change, just larger investment and learning from the experts that we've brought in-house. Are we entering a golden age for cyber? And what I mean by that is we're seeing a huge amount of AI-generated code, which in a lot of cases has holes. We're seeing a Lovable and a Replit and a you-name-it produce a huge amount of output. Are the threats going to increase much more significantly than we're anticipating? Most likely, yes. Where we see it the most is it's an interesting data type because it's competitive, and you can have these AlphaGo-type situations for cyber offense and defense. Where you can have uncapped rewards and performance, and the field is constantly moving. So we love this kind of stuff because it's like a game from a data perspective. And we see very rapidly increasing demand for cyber defensive capabilities via data and very interesting data types. And this is an example of something where sufficient— Wait, can you help me understand what data types do people want around security that they maybe didn't want before there was this explosion in demand? I have to be careful not to reveal too much about customer work. The category is growing very quickly, and the nature of a lot of security work is that it's adversarial, right? So it's not this sufficiency-style work like update a CRM and then you're good. There's a constant cat-and-mouse game between the offensive capabilities and the defensive capabilities, which, to our point earlier about the 90% of enterprise workflows that can already be completed, there's never going to be that 90% for security because the goal posts are always going to move. So cyber as a category is growing, and the nature of the data types is it's much more uncapped, evolving, adversarial in terms of where the goal posts are. Can you help me out here? You're Estonian by heritage. I say to European founders, SF is the worst place to start a company. It is impossible to acquire talent. It is impossible to afford it. And then it's impossible to retain it. Is the talent war in SF as brutal as it seems? Yeah, it's pretty brutal. It is very difficult to hire. It is difficult to retain. It's difficult. I think it's harder than before. But it's easy when you're on a rocket ship, right? It's always easy. When you're on a rocket ship to get someone, it's hard to make the right decisions about who you want to hire. But yeah. When you made a bad hire, what did you not see that you wish you'd seen? It's really hard to assess agency and ownership in the interview process. I am super freaking talented. I'm super talented. I'm a bit of an asshole. I'm not a total asshole, but I'm a bit of a douche. Are you okay with that? If you're super talented, yeah. The company culture here is of high agency, high performance, high ownership. Personalities can change. You can learn how to work with people better. But we care about growth, and we care about, we want to hire people who give a shit. That's a lot harder to coach into someone than smoothing it out with your colleagues, making sure that we have happy hours, people get along. That's easy to work out. You can have a couple assholes, they get drinks together a few times, and then you smooth it out. It's really hard to make someone give a shit. Yeah. Also, if you hire multiple assholes, they can just hang out together. It's fine. That's a group of that. We don't hire a lot of assholes. No, I can be also happy hours. Really? We had a great offsite just recently, actually, with our annotation team. We went to Tofino in Canada. It's on the west coast of Canada. It's the only place you can surf. And everyone did surfing lessons. We went to a floating sauna. And it was a great time. I thought it was actually great for the team, and it was a great use of money. And everybody loved it. And I think that doing these outdoor activities where people are being active is good. Are you ready for a quickfire round, dude? Sure. Yeah. What have you changed your mind on most in the last 12 months? Honestly, I think it's probably the environment, RL environment market. Because when we were starting it off last year, it was so complicated to do these deliveries. And it was so hard to get it to work that I just thought it wasn't going to work out. I thought it wasn't going to scale, but then it did. So I was pretty surprised. What changed? The demand was very high, and we got it to work, right? We just had to try a lot of different things to get environments to actually improve model performance. So we just kept going at it, and it ended up working. I'm your little brother, and I'm studying computer science at university today. You sit me down and say, little brother, you should know this. What should I know? Get a real internship as soon as possible, because whatever you learn in school is probably going to be outdated quickly. Interesting. Where should I get a real internship? I know that sounds stupid, but should I start my own company? Should I join a fast-growing company? Because when we were starting it off last year, it was so complicated to do these deliveries. And it was so hard to get it to work that I just thought it wasn't going to work out. I thought it wasn't going to scale, but then it did. So I was pretty surprised. What changed? The demand was very high, and we got it to work, right? We just had to try a lot of different things to get environments to actually improve model performance. So we just kept going at it, and it ended up working. I'm your little brother, and I'm studying computer science at university today. You sit me down and say, little brother, you should know this. What should I know? Get a real internship as soon as possible, because whatever you learn in school is probably going to be outdated quickly. Interesting. Where should I get a real internship? I know that sounds stupid, but should I start my own company? Should I join a fast-growing company? Should I join a super established company where there's adults in the room, so to speak? Maybe I'm biased, but join a fast-growing company in San Francisco. It doesn't need to have adults in the room, but somewhere on the frontier that's indicative of where the field is going. A bit larger than 10 people, not super early, just to filter out the companies that might not go anywhere. Would you say that you're too late for me? No. No, we still act like a startup. How many people do you have? Maybe 500. What's your favorite, Andrew? Culturally, we're a startup. We're paranoid. We're in office all the time. We're fast-moving. We want to hold on to that as long as possible. I love it. That's amazing. Totally. Absolutely. Yes. Which competitor do you most respect, and why then? I don't think about competitors too much. They're all even in that they're all behind Mercore. It's a bit of a non-answer, but we really try not to think about them as much as we try to think about our customers. So I respect our customers a lot. I love the work that they're doing. We stay on top of what competitors are doing. But every time I look at one of their websites, they're just doing something we did a week or a month ago. We write a blog. Someone else writes a blog a week later that's the exact same thing. We make an update to our website. Someone else makes an update to their website that's the exact same thing. So we... I spend a lot more time... Would you say that about Surge? It's happened before. They're a bit out there. We honestly... I don't spend that much time thinking about them because I spend more time thinking about customers. We've seen it... They're a bit out there in that they don't copy us as much. And they do seem a bit different from others in the field. Hard to say why. They're very secretive. Yeah. Are you kidding me? Yes, absolutely. I totally get that. Can you please paint the bull case for how Mercore is a $200 billion company? It looks like we sell services. We have... We're a tech-enabled services company. Our services are incredibly valuable in driving revenue gains for our customers, primarily through better model capabilities. Evals and training data are the primary bottleneck to model performance right now. If every enterprise needs to have specialized proprietary models, even if the capabilities start to saturate, the evals serve as the PRD for exactly what you want, but also the optimization objective for better performance. As long as better models are valuable to the economy, there will be demand for eval sets and training sets. If we can make that process faster and faster, we can serve a growing demand for human data for eval and training. And then we also have a growing agent deployment enterprise arm as well. What line of revenue do you not have today that you think will be very significant in three years' time? I think that real-world, physical data is going to grow significantly over the next three years. Robotics is an interesting area for us. The data market for robotics is nascent relative to Gen.AI, relative to autonomous vehicles as well. And we think that's going to grow a lot. Do you scale supply ahead of demand? We, at times, retain exceptional talent to do work that might be valuable in the future. And we can do off-the-shelf data creation to make use of supply when demand is low and then resell that data later. In that case, we do. Otherwise, we don't. What's the best piece of advice you've ever been given? I got a lot of advice to join small companies, join startups, move to San Francisco. I grew up in Canada. I went to school in Toronto. I followed that advice. I think it was great. I've loved living out here, and I like small companies. I like fast-growing companies. It's been super fun and great for my career. Final one for you. What are you most excited about that you don't think enough people are talking about? Probably the same answer as before in that the three-years-out opportunity of robotics. I think there's a lot of discussion around robotics on Twitter and in certain- I'm sorry, dude. Can you just help me out here? And this is where I get in trouble. It's Friday afternoon. It's past six. Fuck it. I can say what I want. I don't get it. Okay. Whenever you watch a robotics demo, you see this terribly moving robot around a home. And then after watching it take one water out of a fridge in 15 minutes, it goes, and Brendan was in the other room all along. And you're like, are you fucking kidding me? I had this absolute spacko in my kitchen for 15 minutes getting a water, and Brendan was in my back room doing it. That's where we're at. What am I not seeing? Help me get excited. Yeah. I think if you go back a decade or so, self-driving cars had the people in them all the time. You would see crews driving around San Francisco, and there was a person in it for years, for years, right? But now I take Waymo more than I take Uber. I'm thrilled for you. Welcome to London. We still have these people in cars. I love it, but it's in one city. It can't deal with very ambiguous data. It's still pretty irrelevant. It's made leaps and bounds in the past, at least in San Francisco and Austin, Phoenix. It's tough because, yeah, I guess the distribution is unequal, but it's an incredible service here in San Francisco, and people here use it a lot. So technically it works, and there might be regulatory challenges or other challenges with scaling. But you think we'll hit a ChatGPT moment with robotics, which will cause an inflection in usage and adoption? I think so. Yeah, I think so. But I think it might play out similar to driverless cars, where it's really hard to scale physical things as opposed to software. So it might be more of a Waymo, robotaxi, Cruise-type moment than a ChatGPT moment. But I think the progress will be there. Yeah. Dude, you have been fantastic. Thank you so much for putting up with this incredibly wayward, poorly structured conversation, which was brilliant. And I so appreciate you putting up with it. Thanks for having me. Yeah, it was super fun. Thanks, and I'm going to find a wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful Thank you. I love it, but it's in one city. It can't deal with like very ambiguous data. It's still like pretty irrelevant. It's made leaps and bounds in the past, at least in San Francisco and Austin, Phoenix. It's tough because, yeah, I guess the distribution is unequal, but it's an incredible service here in San Francisco and people here use it a lot. So like technically it works and it might be, you know, there might be like regulatory challenges or other challenges with scaling. But you think we'll hit a chat GPT moment with robotics, which will cause an inflection in usage and adoption? I think so. Yeah, I think so. But I think it might play out similar to driverless cars where it's really hard to scale physical things as opposed to software. So it might be more of like a Waymo, robotaxi cruise type moment than a chat GPT moment. But I think the progress will be there. Yeah. Dude, you have been fantastic. Thank you so much for putting up with this like incredibly wayward, poorly structured conversation, which was brilliant. And I so appreciate you putting up with it. Thanks for having me. Yeah, it was super fun. Thanks, and I'm going to find a wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful wonderful Thank you.