Open Reader

Cursor Has Data Worth Billions (And They Don't Know It Yet)

completed 1:06 May 04, 2026 Watch on YouTube

Current Status

completed

Video ID

B6gsRkqcbMk

RAG / Chat

Enabled
Cursor Has Data Worth Billions (And They Don't Know It Yet)
Description

Every AI company is running out of training data — and the most valuable dataset left is sitting inside developer tools like Cursor. Sam Lessin and Dave Haber break down why traces (the actual prompts and responses between engineers and AI) are the new oil, and why selling a company like Cursor for its product is selling it for scrap. Topics: AI training data, Cursor, Anthropic, software AI, developer tools, LLM data We’re also on ↓ X: https://twitter.com/moreorlesspod Instagram: https://instagram.com/moreorless Spotify: https://podcasters.spotify.com/pod/show/moreorlesspod Connect with us here: 1) Sam Lessin: https://x.com/lessin 2) Dave Morin: https://x.com/davemorin 3) Jessica Lessin: https://x.com/Jessicalessin 4) Brit Morin: https://x.com/brit

Summary

Generated by claude-haiku-4-5-20251001

Summary: Cursor Has Data Worth Billions

Main Topics

  • Computer use model effectiveness and data scarcity
  • The value of training data for AI model companies
  • Cursor's untapped data advantage in software engineering
  • The economics of data as a compounding asset

Key Points

Current State of AI Models

  • Computer use models are currently less than 40% effective
  • All major model companies (including Anthropic) are experiencing severe data shortages
  • Human-generated data is critical for improving model performance

The High-Value Niche: Code Generation

  • The highest revenue potential comes from models that can:
  • Write code
  • Use tools (Unix, Bash commands)
  • Generate software using operating systems
  • Improving code-writing abilities requires specific training data that's difficult to obtain

Cursor's Data Goldmine

  • Cursor operates in the direct "workstream" of software engineers
  • The company has access to invaluable traces data, which includes:
  • Actual prompts used by software engineers
  • Model responses to those prompts
  • Real-world usage patterns
  • This represents one of the largest private datasets of this type

Notable Quotes

  • "Every single model company is out of data."
  • "The only place you can get that is from being in the workstream of software engineers"
  • "This has extreme value. It's a compounding data asset. It's always going to be refreshed."

Takeaways

  • Cursor possesses a highly valuable, underutilized asset in its traces data from software engineers
  • This data has exponential value compared to typical datasets—it compounds over time and continuously refreshes
  • Model companies desperately need this type of data, making it a potential significant revenue or valuation multiplier for Cursor
  • The organization may not fully realize the strategic and financial value of the data they're collecting

Transcript

187 words en Processed in 34.7s

Computer use models are less than 40% effective right now. So the only way we're going to get there is with some kind of traces data. Every single model company is out of data. They need more human generated data to improve their models. Anthropic, we all know their revenue numbers. They've gone through the roof. This is based on models that can write code and use tools, Unix, Bash, being able to generate software using the Unix operating system, the whole thing driving the most value. They still need more data to improve the model's ability to write code. The only place you can get that is from being in the workstream of software engineers, specifically the traces, the actual software engineer used as their prompt, what the response of the model was. There is a private set of data that is extremely valuable to the model companies right now, and Cursor has a huge one. But when you're selling companies for their data sets, you're selling them for scrap. Not in this category. This has extreme value. It's a compounding data asset. It's always going to be refreshed. So the only way we're going to get there is with some kind of traces data. Every single model company is out of data. They need more human generated data to improve their models. Anthropic, we all know their revenue numbers. They've gone through the roof. This is based on models that can write code and use tools, Unix, Bash, being able to generate software using the Unix operating system, the whole thing driving the most value. They still need more data to improve the model's ability to write code. The only place you can get that is from being in the workstream of software engineers, specifically the traces, the actual software engineer used as their prompt, what the response of the model was. There is a private set of data that is extremely valuable to the model companies right now, and Cursor has a huge one. But when you're selling companies for their data sets, you're selling them for scrap. Not in this category. This has extreme value. It's compounding data asset. It's always going to be refreshed.