More or Less Podcast

Cursor Has Data Worth Billions (And They Don't Know It Yet)

297 summary words 1 min summary Watch video

Start with the signal

1 min read

Summary

Summary: Cursor Has Data Worth Billions

Main Topics

  • Computer use model effectiveness and data scarcity
  • The value of training data for AI model companies
  • Cursor's untapped data advantage in software engineering
  • The economics of data as a compounding asset

Key Points

Current State of AI Models

  • Computer use models are currently less than 40% effective
  • All major model companies (including Anthropic) are experiencing severe data shortages
  • Human-generated data is critical for improving model performance

The High-Value Niche: Code Generation

  • The highest revenue potential comes from models that can:
  • Write code
  • Use tools (Unix, Bash commands)
  • Generate software using operating systems
  • Improving code-writing abilities requires specific training data that's difficult to obtain

Cursor's Data Goldmine

  • Cursor operates in the direct "workstream" of software engineers
  • The company has access to invaluable traces data, which includes:
  • Actual prompts used by software engineers
  • Model responses to those prompts
  • Real-world usage patterns
  • This represents one of the largest private datasets of this type

Notable Quotes

  • "Every single model company is out of data."
  • "The only place you can get that is from being in the workstream of software engineers"
  • "This has extreme value. It's a compounding data asset. It's always going to be refreshed."

Takeaways

  • Cursor possesses a highly valuable, underutilized asset in its traces data from software engineers
  • This data has exponential value compared to typical datasets—it compounds over time and continuously refreshes
  • Model companies desperately need this type of data, making it a potential significant revenue or valuation multiplier for Cursor
  • The organization may not fully realize the strategic and financial value of the data they're collecting
Full transcript 187 words · 2 min read
0:00

SPEAKER_00

Computer use models are less than 40% effective right now. So the only way we're going to get there is with some kind of traces data. Every single model company is out of data. They need more human generated data to improve their models. Anthropic, we all know their revenue numbers. They've gone through the roof. This is based on models that can write code and use tools, Unix, Bash, being able to generate software using the Unix operating system, the whole thing driving the most value. They still need more data to improve the model's ability to write code. The only place you can get that is from being in the workstream of software engineers, specifically the traces, the actual software engineer used as their prompt, what the response of the model was. There is a private set of data that is extremely valuable to the model companies right now, and Cursor has a huge one. But when you're selling companies for their data sets, you're selling them for scrap. Not in this category. This has extreme value. It's a compounding data asset. It's always going to be refreshed.

0:04

SPEAKER_00

So the only way we're going to get there is with some kind of traces data. Every single model company is out of data. They need more human generated data to improve their models. Anthropic, we all know their revenue numbers. They've gone through the roof. This is based on models that can write code and use tools, Unix, Bash, being able to generate software using the Unix operating system, the whole thing driving the most value. They still need more data to improve the model's ability to write code. The only place you can get that is from being in the workstream of software engineers, specifically the traces, the actual software engineer used as their prompt,

0:45

SPEAKER_00

what the response of the model was. There is a private set of data that is extremely valuable to the model companies right now, and Cursor has a huge one. But when you're selling companies for their data sets, you're selling them for scrap. Not in this category. This has extreme value. It's compounding data asset. It's always going to be refreshed.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note