Lenny's Podcast

OpenAI’s head of platform engineering on the next 12-24 months of AI | Sherwin Wu

2442 summary words 11 min summary Watch video

Start with the signal

11 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Sherwin Wu argues that AI is rapidly turning engineers into managers of agent fleets, collapsing the cost of bespoke software, and shifting durable advantage toward workflow context, deployment discipline, and products designed for steadily improving models rather than current-model limitations.
  • Why it matters: This is a high-signal view from the OpenAI API/platform leader on agent operating patterns, code-generation governance, AI adoption failures, platform primitives, and the likely next product surfaces: longer-running agents, business-process automation, and audio.
  • Best use: Use it to pressure-test OpenClaw and agent-system architecture, define internal AI adoption practices, and refine investment/product theses around vertical workflow software rather than generic agent scaffolding.

Executive Summary

Wu describes OpenAI as already operating in an agent-heavy engineering model: 95% of engineers use Codex daily, every PR receives Codex review, and heavier Codex users reportedly open 70% more PRs than lighter users. He does not claim every production line was AI-authored, but says the vast majority of new code is likely AI-authored first and then reviewed, steered, or revised by people. The job is moving from individual coding toward directing many parallel agent threads, with senior judgment becoming more important rather than obsolete.

The operational constraint is not simply model quality. Agent failures often reveal missing context, underspecified tasks, and tribal knowledge that has not been encoded into repositories, documentation, skills, or system structure. OpenAI is experimenting with a fully Codex-written codebase specifically to remove the human “escape hatch” and expose these failure modes. Its answer to escalating PR volume is automated review and CI remediation: Codex reviews all PRs, catches routine issues, and can patch lint failures and restart CI, while humans retain meaningful but reduced oversight.

For companies deploying AI, Wu believes many initiatives may have negative ROI, though he does not cite quantitative evidence. His diagnosis is shallow usage plus top-down mandates without bottom-up workflow ownership. The better pattern is executive commitment paired with an internal evangelist group of technically inclined operators—often support, ops, or analytical staff rather than engineers—who discover workflow-specific uses, share practices, and create adoption momentum.

Strategically, Wu expects model gains to erase much of today’s scaffolding and argues builders should target capabilities that are nearly viable now but become exceptional as models improve. He expects more coherent multi-hour agent tasks and materially better native audio over the next 12–18 months. He is especially bullish on AI for repeatable, deterministic, enterprise business processes—not merely open-ended coding—and predicts that cheaper software creation could create a boom in narrow B2B products and smaller, highly leveraged companies. His platform framing is that startups should prioritize customer love over fear that OpenAI will enter their category, although that is an OpenAI executive’s stated strategic posture rather than a guarantee.

Key Takeaways

  • Claim: AI-native engineering is shifting engineers from writing code line-by-line to supervising parallel fleets of agents, with output and senior judgment becoming the differentiators. | Evidence: Wu says 95% of OpenAI engineers use Codex daily; Codex reviews 100% of PRs; and engineers who use it more heavily open roughly 70% more PRs than less-heavy users. Individual engineers may have 10–20 concurrent agent threads in progress. | Implication: Ken should design agent systems around task decomposition, concurrent execution, steering, review queues, and escalation—not treat code generation as a single-shot autocomplete feature. | Caveat: Wu explicitly avoids claiming that 100% of code currently running in production was written by AI; he frames it as the vast majority of code likely being AI-authored first and then reviewed or changed by humans.
  • Claim: The main cause of agent underperformance is often inadequate context rather than an irreducibly incapable model. | Evidence: A team at OpenAI is maintaining an experimental 100% Codex-written codebase. When an agent cannot complete a feature, the team cannot simply hand-code the fix; it responds by encoding missing tribal knowledge in code comments, repository structure, Markdown files, skills, and other in-repo resources. | Implication: For OpenClaw-style systems, context engineering should be treated as core infrastructure: explicit operating procedures, source-of-truth documents, task constraints, tool descriptions, and durable repository memory are likely higher leverage than adding more agent prompts. | Caveat: Wu says this work is still experimental and that OpenAI plans to publish broader learnings; the transcript provides only one emerging pattern, not a complete implementation playbook.
  • Claim: High-volume AI coding requires AI-native review and CI automation, but it does not eliminate the need for human oversight. | Evidence: Codex reviews every OpenAI PR; Wu says this can reduce a review from 10–15 minutes to 2–3 minutes and, for some small PRs, lets the author use Codex as the effective second set of eyes. OpenAI also uses Codex-driven tools to patch lint failures and restart CI automatically. | Implication: Build layered controls around autonomous code changes: automated review, tests/evals, policy-based merge thresholds, post-deploy monitoring, and explicit human checkpoints proportional to blast radius. | Caveat: Wu warns against the Sorcerer’s Apprentice failure mode: teams should not dispatch agents and disengage entirely. He describes the reduction as moving from roughly 100% to 30% human attention, not to zero.
  • Claim: Most enterprise AI programs fail to create value when adoption is imposed from the top without workflow-level ownership and experimentation from the bottom. | Evidence: Wu says he would not be surprised if many deployments have negative ROI, while acknowledging he does not see definitive quantitative data. He contrasts basic, underpowered employee usage with successful organizations that combine C-suite sponsorship, tool access, enthusiastic practitioners, workflow-specific experimentation, hackathons, seminars, and internal knowledge sharing. | Implication: Ken should avoid AI adoption metrics based only on licenses, mandates, or generic productivity claims. Establish a small cross-functional operator cohort with ownership of named workflows, baseline metrics, reusable patterns, and authority to spread successful practices. | Caveat: The negative-ROI assessment is an informed observation rather than a measured OpenAI dataset, and the transcript does not specify a standardized ROI methodology.
  • Claim: Builders should minimize dependence on brittle scaffolding and build for model capabilities expected soon, not merely for today’s constraints. | Evidence: Wu cites Fintool founder Nicholas’s phrase, “the models will eat your scaffolding for breakfast.” He points to earlier assumptions around agent frameworks and vector stores: as models improved, more elaborate retrieval logic and orchestration became less necessary in some cases, with simpler tool access and model-directed search becoming preferable. | Implication: Favor modular, replaceable abstractions and maintain a roadmap based on capability thresholds. Avoid creating a business whose sole value is compensating for a transient model weakness unless it owns proprietary workflow data, distribution, integration depth, or trust. | Caveat: This is not an argument to remove all scaffolding now. Wu specifically notes that current practices such as skills and file-based context management may still be useful, but their longevity is uncertain.
  • Claim: The next major product shift may come from longer-horizon agents and better native audio, which will expand what can be delegated beyond short interactive coding tasks. | Evidence: Wu says frontier models can already complete multi-hour software tasks about 50% of the time, while 80% reliability is just under an hour, according to a trend chart he references. He predicts more coherent multi-hour tasks in 12–18 months and highlights speech-to-speech/native multimodal models as an underappreciated enterprise opportunity because much business and service work is conducted through audio. | Implication: Prepare control planes for persistent work: checkpoints, asynchronous status, intervention interfaces, budget/time limits, audit trails, and human feedback loops. Prioritize voice-native operational workflows where calls and conversations are already the system of record. | Caveat: These are forward-looking expectations, not formal product commitments; Wu also says a model left alone for a day would still require feedback and control.
  • Claim: The largest AI application opportunity may be automation of repeatable business processes, and cheaper software creation could produce a broad long tail of specialized B2B products. | Evidence: Wu distinguishes open-ended knowledge work such as software engineering from SOP-driven work such as utility support and enterprise operations. He argues that a hypothetical one-person billion-dollar company could outsource functions to many specialized micro-startups, producing not just one large winner but potentially many $10–50 million businesses. | Implication: Focus product and investment exploration on high-volume, repeatable, system-integrated workflows with clear rules, data access, and measurable outcomes; distinguish the economic attractiveness of founder-scale businesses from VC-scale return profiles. | Caveat: He calls the downstream startup and VC-market consequences highly uncertain, including whether more small successful companies would reduce the number of conventional venture-scale outcomes.

Detailed Brief

Management model in an agent-amplified organization

  • Claims: Wu expects AI to widen the productivity spread between top performers and the rest of a team because high-agency users learn the tools and workflows faster.; His management philosophy is to spend disproportionate time keeping top performers unblocked, productive, heard, and able to propagate their practices.; As coding throughput accelerates, organizational decisions, dependencies, and process friction become the principal bottlenecks rather than implementation capacity.; Managers may eventually supervise larger organizations because AI connected to internal systems can synthesize organizational context, support research, and prepare performance-review evidence.
  • Evidence: Wu says OpenAI managers can use ChatGPT connected to GitHub, Notion, and Google Docs to assemble a deep-research-style view of an employee’s work over the preceding 12 months for performance reviews.; He references the 'surgeon' concept from The Mythical Man-Month, adapting it into a management role focused on looking around corners and supplying what high-leverage contributors need.; He proposes an untested but concrete management use case: ask an internal-knowledge-connected model to identify current blockers and predict blockers a team may face in coming months.
  • Caveats: Wu says management has changed less than engineering so far and that there is not yet a direct 'Codex for managers.'; Larger spans of control are his projection, not an established organizational best practice.
  • Implications: The management system around agents should track dependency risk and decision latency, not just agent output or PR count.; An internal organizational-intelligence agent could be valuable if grounded in appropriately permissioned work systems and evaluated for false positives, privacy, and managerial misuse.

OpenAI platform posture and stack

  • Claims: Wu presents OpenAI as an ecosystem platform company whose API is intended to let external builders reach specialized niches that OpenAI cannot serve directly.; He advises founders not to over-index on whether OpenAI will enter their category; in his experience, failed startups more often lack customer resonance than suffer direct platform competition.; The platform is offered in layers, allowing developers to choose between low-level flexibility and more opinionated agent-building components.
  • Evidence: Wu says OpenAI releases models used in its products into the API and does not block competitors from using its models; he also references testing 'Sign in with ChatGPT' and the separate ChatGPT app ecosystem.; He cites 800 million weekly active ChatGPT users as an incentive to enable outside developers to build for that audience.; The Responses API is described as the core low-level primitive for long-running agent interaction; the Agents SDK adds sub-agents, orchestration, and guardrails; AgentKit/widgets provide UI components; and Evals API supports quantitative testing of models, agents, and workflows.
  • Caveats: OpenAI’s stated ecosystem commitment is strategic positioning from its head of platform engineering, not protection against future product overlap, pricing changes, policy changes, or platform dependency.; The transcript does not cover API economics, rate limits, data controls, vendor lock-in, or comparative performance against other model providers.
  • Implications: Use OpenAI’s higher-level stack selectively where speed matters, but retain portable evaluation suites, workflow definitions, data connectors, and domain policy layers so the product does not depend solely on a single provider’s abstractions.; A platform-distribution opportunity may be emerging alongside the API opportunity, but it should be evaluated separately from direct API demand.

Notable Concepts & Terms

  • Sorcerer’s Apprentice: Wu’s metaphor for agentic coding: powerful delegated work can multiply output but requires active supervision to prevent agents from running off-course.
  • SICP / the wizard metaphor: Structure and Interpretation of Computer Programs framed programming as issuing incantations; Wu argues natural-language direction of agents makes that metaphor increasingly literal.
  • Context engineering: The practical work of encoding tacit knowledge, constraints, docs, skills, and structure so agents can reliably execute tasks without relying on a human rescue path.
  • Models will eat your scaffolding for breakfast: A warning that products or architectures built mainly to compensate for current model limitations—such as elaborate retrieval or orchestration layers—may be commoditized as models improve.
  • Bitter lesson: The AI/ML principle that scalable general methods and computation often outperform elaborate hand-built logic; Wu uses it to argue for simpler, model-capability-aware systems.
  • Business process automation: Wu’s preferred application category: repeatable, SOP-driven, deterministic workflows integrated with enterprise data and systems, as distinct from open-ended knowledge work.
  • Responses API: OpenAI’s low-level API primitive described here as optimized for long-running agents, offering flexible access without imposing a specific orchestration framework.
  • Agents SDK / AgentKit / Evals API: The higher layers of OpenAI’s stack: agent orchestration and guardrails, reusable UI components for agent products, and quantitative testing for agent/model workflow quality.

Operator Notes / Why Ken Should Care

  • Run an agent-readiness audit on one real workflow: identify the tacit information currently held in operators’ heads, then convert it into versioned instructions, tool contracts, examples, exceptions, and evaluation cases.
  • Define autonomous-change tiers for engineering agents: low-risk changes may auto-remediate and pass AI review; higher-risk changes should require test gates, human approval, and production monitoring.
  • Create a small AI operator guild drawn from technical-adjacent functions such as ops, support, analytics, and engineering; give it explicit workflow ownership and a mandate to publish reusable practices rather than merely promote tool usage.
  • Build or strengthen the control-plane features needed for long-running agents: task budgets, checkpoints, intervention controls, trace logs, rollback paths, approval gates, and outcome evals.
  • Screen B2B agent opportunities for repeatability, SOP clarity, accessible system data, integration feasibility, and measurable business outcomes; deprioritize generic copilots lacking a proprietary workflow wedge.
  • Maintain portability at the evaluation, data-access, and policy layers even when using OpenAI APIs or SDKs, since platform neutrality and roadmap alignment are stated intentions rather than contractual insulation from dependency risk.
  • Explore voice-first agent workflows where phone calls, support interactions, or operational conversations are central inputs and outputs; ensure consent, recording, privacy, and escalation requirements are designed in from the start.

Source/Metadata

  • Title: OpenAI’s head of platform engineering on the next 12-24 months of AI | Sherwin Wu
  • Transcript words: 25036
  • Duration seconds: 4779
  • Timestamp note: No usable timestamps or chapter markers were present in the supplied transcript; the transcript also contains substantial repeated passages.
Full transcript 14324 words · 111 min read
0:00

95% of engineers use Codex. 100% of our PRs are reviewed by Codex. For engineers, I don't know what job has changed more in the past couple years. Engineers are becoming tech leads. They're managing fleets and fleets of agents. It literally feels like we're wizards casting all these spells, and these spells are going out and doing things for you. What do you think people aren't pricing in yet? The second- or third-order effects of the one-person billion-dollar startup. To enable a one-person billion-dollar startup, there might be a hundred other small startups building bespoke software. So I think we might actually enter into a golden age of B2B SaaS.

0:31

I've been hearing more and more. There's this stress people feel when their agents aren't working. There's a team that's actually doing an experiment right now with OpenAI where they are maintaining a 100% Codex-written code base. They run into the exact problems that you're describing. Usually you're like, all right, I'll roll up my sleeves and figure it out. This team doesn't have that escape hatch. You've shared that listening to customers is not always the right strategy in AI. The field and the models themselves are changing so, so quickly. They tend to disrupt themselves. The models will eat your scaffolding for breakfast.

0:59

What's your advice to folks that are like, okay, I don't want to miss the boat? Make sure you're building for where the models are going and not where they are today. There's a quote from Kevin Whale, our VP of science here. He likes saying this is the worst the models will ever be. Today, my guest is Sherwin Wu, head of engineering for OpenAI's API and developer platform. Considering that essentially every AI startup integrates with OpenAI's APIs, Sherwin has an incredibly unique and broad view into what is going on and where things are heading. Let's get into it after a short word from our wonderful sponsors. Today's episode is brought to you by DX,

1:32

the developer intelligence platform designed by leading researchers. To thrive in the AI era, organizations need to adapt quickly. But many organization leaders struggle to answer pressing questions like, which tools are working? How are they being used? What's actually driving value? DX provides the data and insights that leaders need to navigate this shift. With DX, companies like Dropbox, Booking.com, Adyen, and Intercom get a deep understanding of how AI is providing value to their developers and what impact AI is having on engineering productivity. To learn more, visit DX's website at getdx.com slash Lenny. That's getdx.com slash Lenny.

2:13

Applications break in all kinds of ways: crashes, slowdowns, regressions, and the stuff that you only see once real users show up. Sentry catches it all. See what happened, where, and why, down to the commit that introduced the error, the developer who shipped it, and the exact line of code, all in one connected view. I've definitely tried the five tabs and Slack thread approach to debugging. This is better. Sentry shows you how the request moved, what ran, what slowed down, and what users saw. Sear, Sentry's AI debugging agent, takes it from there. It uses all of that Sentry context to tell you the root cause, suggest a fix, and even opens a PR for you.

2:54

It also reviews your PRs and flags any breaking changes, with fixes ready to go. Try Sentry and Sear for free at sentry.io slash Lenny and use code Lenny for $100 in Sentry credits. That's S-E-N-T-R-Y dot I-O slash Lenny.

3:15

Sherwin, thank you so much for being here, and welcome to the podcast. Thank you. Thank you for having me. I want to start with what's feeling like a barometer of progress in AI, especially in engineering. What percentage of your code, if you even write code anymore, and your team's code, is written by AI at this point? I do write code occasionally still. I'd actually say for managers like myself, it's way easier to use these AI tools than to manually code at this point. I know for myself and some of the other EMs, engineering managers at OpenAI, all of our code is written by Codex at this point. But more broadly, there's just so much energy. There's a tangible energy

3:52

internally around just how far these tools have gotten, how good Codex as a tool has gotten for us. And it's a little hard for us to exactly measure how much of the code is written because the vast majority of it, I'd say close to 100%, is usually generated by AI first. What we do track, though, is, at this point, the vast majority of engineers use Codex on a daily basis. So 95% of engineers use Codex. 100% of our PRs are reviewed by Codex daily as well. So any code that goes into production, that's merged in, Codex has its eyes on it and suggests improvements, suggests changes in the PRs. And so that's what we're seeing internally. But by and large,

4:33

the most exciting thing is just the energy that there is. Another observation that we've had is engineers who tend to use Codex more open way more PRs. So they're actually opening 70% more PRs than the engineers who aren't using Codex as much. And the gap is widening. So I feel like the people who are opening more PRs are starting to learn how to use the tool more and more, get more efficient, and that 70% gap keeps growing over time. And so it might have actually increased since I last looked at the number. Okay, so just to make sure we hear what you're saying, you're saying all of the code of these 95% of engineers at OpenAI is written by AI. It's written,

5:11

and then they review it. Yep, yep. It's crazy that that's almost not crazy anymore, that we're just getting used to this. I think there's still some getting used to, to be clear. There's also, I think, some engineers who, I think, trust Codex a little bit less. But every day I talk to someone who is blown away by something that it can do, and their bar of trust, or how much they trust the model to do on its own, goes up over time. And there's a quote from Kevin Whale, our VP of science here. He likes saying, this is the worst the models will ever be. And so this is the worst that the models will ever be for software engineering as well. Over time, we just see people

5:55

trusting it more and more, and then we'll see the models get better and better as well. Yeah, Kevin Whale, former podcast guest, he said exactly that line on this podcast. Yeah, yeah, yeah. A few times. Yeah. Peter, the Claudebot slash Moltbot slash OpenClaw, is what it's called now, developer recently shared that he uses Codex for his work, and he feels like anytime it does things, he just trusts that it has done the right job, and he's almost certain he could just commit it to master and it'll be great. Yeah, yeah. He's a great user of Codex. I know he's in close touch with the team, gives us great feedback. I'm not surprised that he uses it. Sorry,

6:26

it's called OpenClaw. OpenClaw. Yeah. OpenClaw is a great product. And then I saw that this morning. This is very recent, but this morning, I think Moltbook, with Sherrod as well, and seeing all of the AI agents talk to each other is pretty surreal. It's basically Her is happening in real life, is what I'm hearing. Yeah, yeah. So just coming back to this crazy moment we are living through for engineers in particular, we've gone from you write every line of code to now AI is writing all of your code. I don't know what job has changed more in the past couple years, a job that we didn't expect to change this much, where the job of an engineer is so different in the entire

7:01

lifespan of an engineer. In the past couple years, it's now shifted to I don't write any more code. How do you imagine the role of an engineer and the job of a software engineer looks in the next couple of years? What is that job? Yeah, it's honestly been really cool to see, and it's part of where the excitement is because the job is likely going to change pretty significantly over the next one or two years. It feels like we're still figuring things out, though, and so there's this excitement, especially from some of the software engineers, of we're in this rare moment, maybe over the next 12 to 24 months, where we'll get to figure things out ourselves

7:35

and set our standards for ourselves. In terms of where I see this moving, I think there's a common thing that everyone's saying, which is, people who are generally IC engineers are becoming tech leads. They're basically managers now. They're managing fleets and fleets of agents. I know many of the the role of an engineer and the job of a software engineer looks in the next couple of years? Just what is that job? Yeah, it's honestly been really cool to see, and it's part of where the excitement is, because the job is likely going to change pretty significantly over the next one or two years. It feels like we're still figuring things out, though, and so there's

8:08

this excitement, I know, especially from some of the software engineers, of we're in this rare moment, maybe over the next 12 to 24 months, where we'll get to figure things out ourselves and set our standards for ourselves. In terms of where I see this moving, I think there's a common thing that everyone's saying, which is, people are generally, like IC engineers, are becoming tech leads. They're like managers now. They're managing fleets and fleets of agents. I know many of the engineers on my team have 10 to 20 threads being pulled on at the same time. Obviously not active, running Codex jobs, but just a lot of parallel threads. They're checking in

8:48

on what they're doing. They're steering the agents and Codex and giving it feedback. And so their job has really changed from just writing the code itself into being almost like a manager. In terms of where I think this will go one to two years from now, one metaphor that I always come back to here is actually from this programming textbook that I read back in college called SICP. I don't know if you've heard of it: Structure and Interpretation of Computer Programs. So SICP. At MIT it was really popular, and it was actually used as the textbook for the intro programming course for a very long time. And it has this cult following. It teaches you programming. It teaches you

9:28

a dialect of Lisp called Scheme. And so it introduces you to functional programs, very mind-opening that way. But the thing that was memorable for me about that book, I read it in college, the very beginning of it describes programming as a discipline and draws this metaphor to basically sorcery. It says software engineers are like wizards, and programming languages are like incantations, and you're issuing these spells, and these spells are going out and doing things for you, and the challenge is what incantation do you have to say to make the program do what you want. And this book was written in 1980, so this is a while ago, and I think that metaphor

9:58

has actually persisted over time, and I think it's actually playing out as we move into this new era of vibe coding, or just what software engineering will look like, because programming languages were basically these incantations. They've changed over time, and the challenge is always, and the trend has been, that it's been easier and easier to get the computer to do what you want via programming. And I think the current wave of AI is probably the next stage of that evolution. It is now literally incantations, because you can tell Codex, you can tell Cursor exactly what you want to do, and then it'll go do it for you. And I particularly like the wizard

10:30

and the sorcery analogy, because I think our current state is starting to move towards the Sorcerer's Apprentice, from Fantasia, where Mickey Mouse finds the sorcerer's hat and he tries to do all these things. I think it's a really apt analogy because, one, it's really powerful now. What these incantations can do is extremely high leverage, but you have to know what you're doing, right? In Sorcerer's Apprentice, the whole plot is Mickey goes wild, the brooms go crazy, and everything's flooding. I think he literally sets the brooms off on a task and then goes asleep, and so it's like vibe coding at its greatest. And then eventually the old sorcerer comes back and cleans

11:03

everything up. When I see engineers doing these 20 different Codex threads at a time, there is some skill and there's some seniority and a lot of thought that needs to go into this, because you want to make sure that the models aren't going off the rails. You definitely don't want to completely go away and ignore the thing, but it's also extremely high leverage. A very senior engineer who's really proficient with these tools can now do way more things via what they're doing. And I think this is also what makes it fun. It literally feels like we're wizards now. It feels like we're closer to making it feel like this magical experience where we're casting all these

11:38

spells and having software do all these things for you. I was thinking of the Sorcerer's Apprentice exactly as the metaphor as you were describing that, so I'm glad you went there. A previous podcast guest described it as you have a genie that grants you wishes, and it's a wish you want. If you want to be big, how big exactly? Yeah, or it might be the monkey's paw type thing where it's like you got what you want, but what are the side effects? Yeah, I think that analogy is great. And the crazy thing for me is just the staying power of that book. SICP, it's called the wizard book. People call it the wizard book because that is the metaphor that they weave throughout

12:06

the book. And we've basically reached that point now, which is really cool. There's two threads I want to follow here. One is I've been hearing more and more there's this stress that people feel when their agents aren't working. You fire off all these Codex agents, and then you have to stay on top of them. Oh shit, one's not working. I'm wasting time. Do you feel that across your team at all? Yeah, yeah. It happens all the time, and I actually think this is where the interesting part of all of this lies right now, because these models aren't perfect, these tools aren't perfect, and we're still trying to figure out how to best interact with these, with Codex or with these

12:34

AI agents, to get work done. We see this come up all the time. There's a particularly interesting team that we have internally. There's a team that's actually doing an experiment right now within OpenAI where they are basically maintaining a 100% Codex-written code base. You'll have the AI write code, but you'll obviously end up rewriting a lot of it, and you might need to double back and change things, but this team is just fully Codex-pilled and leaning in entirely. And they run into the exact problems that you're describing, which is, their challenge is, I want to get this thing, this feature, built, but I can't get the agent to do it. And so usually there's an escape

13:04

hatch where then you're like, all right, I'll roll up my sleeves and figure it out. And then instead of using Codex, I might use tab complete in Cursor and things like that. But this team, for the experiment, doesn't have that escape hatch. And so then the challenge is, how do I get the agent to do this? And I actually think we're going to be publishing a blog post from some of our learnings here, but a lot of fascinating paradigms and best practices are falling out of this. One interesting thing that we've noticed, I don't know if this is what you feel, but we definitely feel it here, is a lot of the time when the coding agent is not doing what you want, it's usually

13:31

a problem with context and information that you've given it. It's just either underspecified, or there's just not enough information around how to do something available to the agent, available to Codex. And so when you have to solve it through that, the challenge is then to add documentation and actually work around this limitation, and basically encode more tribal knowledge that's in your head somehow into the code base, either via code comments itself, or code structure itself, or via text files like .md files, skills, any type of additional resources within the repository, so that the model can better do its task. There's a whole bunch of other learnings from

14:05

this group, which I think is fascinating to explore, but giving or removing that escape hatch of no longer using the AI has allowed them to start piecing together a lot of the problems that we'll have to solve if we really want to lean into agents. Another issue people run into: you talked about how people are shipping PRs like crazy, a lot more PRs if they're working with AI. Obviously, code review is becoming a bigger challenge. Is there anything you've figured out on your team to help speed that up, to make that scale, and not just create this terrible job for people where they're just sitting there reviewing PRs all day? Yeah, one thing is Codex reviews 100%

14:41

of all of our PRs at this point. And so I actually think one really interesting thing that's happened is the things that tend to, we tend to hand to the models immediately tend to be the things that annoy us or are the most bunch of other learnings from this group, which I think is fascinating to explore. But giving, removing that escape hatch of no longer using the AI has allowed them to start piecing together a lot of the problems that we'll have to solve if we really want to lean into agents. Another issue people run into: you talked about how people are shipping PRs like crazy, a lot more PRs if they're working with AI. Obviously, code review is becoming a bigger

15:18

challenge. Is there anything you've figured out on your team to help speed that up, to make that scale, and not just create this terrible job for people where they're just sitting there reviewing PRs all day? Yeah, one thing is Codex reviews 100% of all of our PRs at this point. I actually think one really interesting thing that's happened is the things that we tend to hand to the models immediately tend to be the things that annoy us or are the most boring parts of software engineering. It's also why it's more fun now, because we get to do more of the fun things. Speaking more for myself, I really hated code reviews. It was one of the worst things for me. I remember

15:50

my first job out of college, it was at Quora. I was working on the newsfeed, and I owned the code for the newsfeed, so I was a reviewer for newsfeed. It was the central piece of code that everyone would touch, and every morning I'd log in and be like, 20 to 30 code reviews. It's like, oh my goodness, I gotta get through all of these. I would procrastinate, and then it grows to 50. So there's just a lot of code reviews. Codex is really good at reviewing code. Actually, one thing that we've noticed that 5.2 in particular has gotten extremely, extremely adept at is reviewing code, especially when you steer it in the right direction. For code reviews, yeah, we create a lot

16:24

of PRs, but Codex reviews all of them, and it makes code reviews go from a 10-15 minute task to sometimes even just a 2-3 minute task because you have a bunch of suggestions already baked in. A lot of the time, especially for small PRs, you actually don't even need people to review. We trust Codex in this way. The original author looks at Codex. The benefit of code reviews is to have a second pair of eyes to make sure you're not doing anything dumb. Codex is a pretty smart second pair of eyes at this point, and that's something that we've heavily leaned into. The general CI process and the post-push and deployment process have also been heavily automated via Codex

16:44

internally. At this point, if you talk to Codex, there's a lot of automated stuff you can do with Codex. We've actually built some tools internally that help automate that process, automate the lint. If there's a lint error, it's a very easy Codex fix, and it can just patch it and restart the CI process. So all of that, we're trying to collapse into as little work for an engineer as possible, the byproduct of which is they can now merge and push out a lot more PRs. Codex writing, Codex reviewing its own code. I'm curious if you are open to using other models to review your model's work. Is that a path, or it's good enough? We don't need anything else. I will say there's

17:03

a circular thing here, and going back to Sorcerer's Apprentice, you want to make sure you're taking a look at their PRs. It's not like it's going to zero. It's more like going from 100% attention to 30% attention, which helps push through. In terms of multiple models, we test a lot of models internally. We use external models less. We think it's important to dogfood our own models and get feedback there, but there are a lot of internal variants of models that you can use to give you different perspectives here as well, and we found that to work quite well. Okay, so just to make sure we get a barometer of today's world at OpenAI in terms of AI and code, just so I

17:34

understand, and then I want to move on to a different topic: 100% of code across OpenAI is written by Codex at this point. Is that the way to frame it? I wouldn't make the statement that 100% of code running in production today is written by AI, and it's hard to do attribution there, but almost every engineer heavily uses Codex in all of their tasks at this point. So if I were to guesstimate, the vast majority of code at this point was probably authored by AI. Incredible. Okay, so there's a lot of talk, and we've been talking about the IC role, the work of an IC engineer. There's less talk about the changing role of a manager, especially an engineering manager. How has

17:57

your life as a manager changed with the rise of AI, and what do you think the role of a manager is in the future? It's definitely changed less than an engineer. There's no Codex for managers just yet. However, I use Codex quite a bit for some of the more manager tasks that I do. I'd say a couple things are changing. There are some trends. I don't think it's changed that much yet, but I see trends, and I think if you play it out, you can see where a lot of this is going. One thing that's becoming increasingly clear is Codex really empowers top performers to be a lot more productive, and I think this may be true for AI more broadly across society, which is the people who

18:28

really lean in, or the people who have high agency, or who will really get good at these tools, will supercharge themselves. I'm noticing this now as well, which is the top performers end up being a lot more productive, so you see a broader spread in team productivity in this way. One thing that I've always done as a management philosophy is to spend the majority of my time with top performers, make sure they're unblocked, make sure they're happy, make sure they feel productive, and they feel heard. I think this is even more true in an AI world, where your top performers are going to really be shooting, and seeing what's happening there is something that's paid dividends.

18:50

So I think that's one trend that I'm seeing, where spending even more time with top performers for managers is likely going to continue. The other thing is, this is more an observation, but my sense is with a lot of these AI tools available to managers, less for writing code, but for things like chat with organizational knowledge, being able to do research and understand organizational context a lot better. Another good example is we're doing performance reviews right now, and it's actually really easy to use ChatGPT with internal knowledge hooked up to GitHub and our Notion docs and Google Docs to get a really good sense of what this person has done over the last

19:09

12 months and write a little deep research report for it. My sense is managers will be able to manage much larger teams in this world, kind of like how software engineers are managing 20 to 30 Codexes. My sense is these tools will allow managers to be higher leverage and will allow them to manage teams of way more than the current best practice, which I think is 6 to 8 for software. You see this applied to the non-engineering domains like support or operations, where previously the size of the support team might be limited, but as you can pass off more things to agents, you can actually do more work and also manage more people this way. I think the same thing might

19:31

happen for people management as well, especially in tech companies, and we're already seeing this. There are some teams where there are EMs and people on your team to unblock them, make sure they have everything they need. Yeah, a very good example right now is there are, I would say, a group of engineers internally who are really Codex-pilled and are thinking through what the best practices are for interacting with this model, and that is just an extremely high-leverage thing for them to do. So as a manager, I'm just sharing documents and best practices everywhere. Things like that just elevate everyone, and I view that as another example of this trend that we're seeing,

19:51

where the top performers really get exceptional. People just have a sense this is big. AI is changing so much. The world is changing. It's going to be a huge deal. What do you think people aren't pricing in yet into what will change, into where things are going? One of my favorite phrases or things that have come out of this whole AI wave is the idea of the one-person billion-dollar startup. I think Sam may have been the first one to say it, but it's fascinating to think about. If people are so high-leverage, at some point there will likely be a one-person billion-dollar startup. While I think that's really cool, I think people aren't pricing the second- or third-order

20:11

effects of this. Really, what the one-person billion-dollar startup implies is that one person can have so much more agency and so much more leverage using one of these tools that it is just super easy for them to get everything done that they need to for their business, to ultimately create something that's a billion dollars. But I get exceptional people just have a sense this is big. AI is changing so much. The world is changing. It's going to be a huge deal. What do you think people aren't pricing in yet into what will change, into where things?

20:24

One of my favorite phrases or things that have come out of this whole AI wave is the idea of the one-person billion startup. I think Sam may have been the first one to say it, but it's fascinating to think about. If people are so high leveraged, at some point there will likely be a one-person billion startup. While I think that's really cool, I think people aren't pricing the second or third order effects of this. Really, what the one-person billion startup implies is that one person can have so much more agency and so much more leverage using one of these tools that it is just super easy for them to get everything done that they need to for their business to ultimately create something that's a billion dollars.

20:24

But I think there are a couple other implications of this. One of them is, if it's easy for a person to create a one-person billion startup, it also means it's way easier for people to create startups in general. I think there's going to be a huge startup boom and SMB-style boom where anyone can build software for anything. One, you're starting to see this play out in the AI startup scene, where software became a lot more vertical-oriented, where creating some AI tool for some vertical tends to work quite well because you really lean into that particular domain. You really understand the use case for it. If you play out AI, there's no reason why you can't have 100x more of these startups.

20:25

I think one world that we might end up seeing happen is, in order to enable a one-person billion-dollar startup, there might be 100 other small startups building bespoke software that works extremely well to support other types of small one-person billion-dollar startups. I think we might actually enter into a golden age of B2B SaaS and software and startups in general. I think that's a really interesting trend to see because, as it gets easier and easier to build software, as it's easier and easier to run a company, you might actually just end up seeing way more of these startups.

20:25

The way I've been thinking about it is, yeah, there might be one-person billion startups, or there might be 100 million startups. There might be tens of thousands of $10 million startups. As an individual, it's actually pretty great to have a $10 million business. That's enough for you're set for life at that point. We might really see an explosion in that way, and I feel like people aren't really pricing that in.

20:26

There's another third-order effect to this. Again, as you get to the further and further out predictions, I think there's a lot of uncertainty. I think the startup ecosystem will change. I think the VC ecosystem will change. We might end up in a world where there's just a handful of big players that are offering platforms and supporting all of these startups, but the types of venture-scale return startups that can really hundred- or thousand-x your investment might actually end up shrinking if you end up having a bunch of these smaller $10 to $50 million companies, which are not great for venture-style returns but are great for the individuals, the high-agency individuals who are now really leaning into AI to build these businesses for themselves.

20:26

Every layer. Okay, so the billion-dollar startup, I think about this a lot because I'm not going to be a billion-dollar startup because what I'm doing is not venture-scale in any way and not super high leverage. But I just could see how many support tickets I get from the most ridiculous things. It's hard for me to imagine one person. I'm bearish on this billion-dollar startup. I just want to share this thought simply because of the support costs. Even if AI is helping you, at a billion dollars, unless your ACVs are very high and you have very few customers, just dealing with support and people are like, they can solve. They solve. I'll email support. I'll ask about this thing. Just dealing with that is hard to scale, in my experience. So unless you have, in my opinion, a bunch of contractors, which I don't know, does that count as a single-person company? I feel like it's very difficult to scale a billion-dollar startup and not have someone helping you with at least the support work. AI, I think, will take you so far.

20:27

So I think that's true. Actually, I think my view on being the one person who has to dispatch an AI to solve and fix those support tickets, I think what might end up happening is there might be a whole smattering of other startups that are building software and super tailored towards what you might need. There might be 10 or 20 startups that build support software for podcasts and newsletters, and that might be a one-person startup. They might be able to code up this product very easily. They are able to build their own thing, and because it's so tailored and unique and hopefully useful for you, it might be something that you purchase as the one-person billion startup.

20:27

There's a question of what you in-house and what you outsource. What might happen is, because the cost of software and building products is collapsing so much, you might end up outsourcing a lot of this and, in doing so, reducing the size of your company. That's the world that might end up happening. Again, there's high uncertainty in what might play out here, but the end result still might be one person driving this high, massive leveraged company that might actually reach a billion dollars.

20:27

I could see that. I also think about Peter at Claudebot slash Moldbot slash OpenClaw, of how barraged he is right now by all these asks and emails and pings and DMs and PRs. I'm curious, and he's not even making any money out of this thing. Yeah, I can't imagine what it's like to be him right now. It must be absolutely insane. It's probably like the months after we launched ChatGPT, the craziness that was. Yeah, as one man. He's coming out on the pod, by the way, in a week. Oh, that's exciting.

20:28

Yeah, maybe the fourth-order effect is distribution becomes increasingly important because there are so many freaking things trying to get your attention. So people with an audience and platform, I think, become more and more valuable, which is good. Good stuff. Okay, I wanted to come back to you, just thinking about you as a manager of a team that is building the platform that powers basically the entire AI economy. Every AI startup is building on your API. Clearly, you're doing a great job. What other core management lessons have you learned? What do you find is really important and key to your success as a manager of engineers and people?

20:29

Yeah, I think a lot of the lessons that I've learned here, I don't know how specific it is to the OpenAI API or some of our enterprise products in particular. I think my management philosophy has obviously changed over time, but I think it probably stayed the same more than it's changed over time. One of these principles is what I talked to you about: performers and really trying your best to empower them.

20:29

The way that I think about it is, I come back to this analogy of software engineer as a surgeon, which comes from The Mythical Man-Month book. It's funny, so I pull it from the book, but in the book they actually describe this world where I think they were predicting the future because I think the book was in the '70s or something. They said that software engineering might end up moving into a world where the software engineers are like surgeons, or in a surgery room there's one person doing the work. There's the one person cutting whatever and doing all the surgery, and everyone else in the room is there to support them. It's the nurse and the assistant, the resident and the fellow, and the surgeon is like, I need a scalpel, and they give them a scalpel. Then they're like, I need this tool and this machine, and they'll bring it over. Everyone's there to support the one surgeon.

20:29

The Mythical Man-Month actually predicted that that is the direction that software is going to go. I don't think that's exactly played out, where it's much more collaborative. The philosophy, which is software engineering isn't really like surgery, where it's not just one person doing work, but the way in which I like treating the people on my team and the way that I act as a manager is looking around corners and unblocking people, especially from an organizational perspective, is extremely useful.

20:29

Again, going back to the AI conversations, even more important nowadays. If people are just cranking PR after PR, the main thing bottlenecking progress and shipping something tends to be organizational or process-oriented. If you as a manager can look around corners and unblock that, that's the best-case scenario. That's kind of

20:30

there to support the one surgeon. The myth, mammoth, actually predicted that that is the direction that software is going to go. I don't think that's exactly played out, where it's much more collaborative, and philosophy, which is software engineering, isn't really surgery, where it's not just one person doing work. But the way in which I treat the people on my team and the way that I act as a manager is looking around corners and unblocking people, especially from an organizational perspective, is extremely, extremely useful. Again, going back to the AI conversations, even more important nowadays, right? If people are just cranking PR after PR, the main thing bottlenecking progress and shipping something tends to be organizational or process-oriented. If you, as a manager, can look around corners and unblock, that's the best-case scenario. That's the way that I approach management, and especially engineering management, and so that's something that's really stuck with me over time. Even though software engineers aren't exactly surgeons, that metaphor has always stayed in my mind as of my career.

20:30

I love that, and I feel like I wonder if that's something AI can help with: look around corners and predict, here, this engineer is going to be blocked by this decision. We need to figure this out. We need to get— That's actually a really good point. I haven't tried this yet, but I wonder what would happen if I asked ChadGBT, hooked up to company knowledge, what are the active blockers? Look through all the Notion docs, maybe Slack messages. What are the active blockers on my team, and is there something I can do to help? I have not thought about that, but you're right.

20:31

Just had an insight right here, and I think, even more interestingly, what do you anticipate will be a blocker for this engineer or this team in the coming months? You ask the model, you ask the AI, to do the second and third order, anticipate what the blockers will be next month. I think we've got a good idea, right?

20:31

Platform product managers at the world's best companies use Datadog, the same platform their engineers rely on every day, to connect product insights to product issues like bugs, UX friction, and business impact. It starts with product analytics, where PMs can watch replays, review funnels, dive into retention, and explore their growth metrics. Where other tools stop, Datadog goes even further. It helps you actually diagnose the impact of funnel drop-offs and bugs and UX friction. Once you know where to focus, experiments prove what works. I saw this firsthand when I was at Airbnb, where our experimentation platform was critical for analyzing what worked and where things went wrong, and the same team that built experimentation at Airbnb built Eppo. Datadog then lets you go beyond the numbers with session replay. Watch exactly how users interact with heat maps and scroll maps to truly understand, so that you can roll out safely, target precisely, and learn continuously. Datadog is more than engineering metrics. It's where great product teams learn faster, fix smarter, and ship with confidence. Request a demo at datadoghq.com slash Lenny. That's datadoghq.com slash Lenny.

20:32

Okay, I'm going to shift to talking about the API and the platform that you all build. You work with a lot of companies implementing your API, your platform, building. You told me that you find that a lot of companies actually have negative ROI on their AI deployments, which I think is what a lot of people read about and feel and think, and it's interesting actually seeing that. What's going on there? What are they doing wrong? What's happening in the world of AI and deployments and ROI?

20:32

Yeah, so to be clear, I don't explicitly see quantitative numbers around this. It's actually really hard to measure these things. But especially from observing some companies trying to do AI, I would not be surprised if a lot of AI deployments are actually negative ROI. Part of this, too, is I think there's also general sentiment from folks around a couple things I've observed around this. One thing is, and I think I come back to this again and again, I think we in Silicon Valley just forget that we live in a bubble. Twitter is a bubble is a bubble. Silicon Valley is a bubble. Software engineering is a bubble. Most people in the world, most people in the US, are not software engineers, are not very AI-pilled, are not following every single model release. And so we're just highly out of the loop on how to use this technology. We always talk about all these best practices for codex, all these codex people within OpenAI. I'm sure everyone on X who posts are crazy power users of these AI tools. They lean into skills. They lean into—when I talk to some of these companies and I talk to the actual employees using these, it's the most basic thing that they're trying to do, and they have very little understanding of exactly how this technology works. So that's one big observation for me, which is they're asking very simple questions of these things. They're really not pushing it just yet.

20:33

What a more ideal AI deployment setup looks like, and this is how we've run things within OpenAI too, the companies where I think it started to work really well have a combination of both top-down buy-in, so it's the C-suite, we want to become an AI-first company, so there's buy-in, they buy the tools, they have exec support, but it also has bottom-up adoption and buy-in. What I mean by that is it has actual employees doing the work who are really excited about this technology and are willing to learn, evangelize, build best practices, and knowledge-share within the organization. We've seen this a lot internally. Obviously, OpenAI has always wanted to be a very AI-centric company, but when it really started taking off was with the introduction of codecs and these tools where actual employees themselves could start applying it to their work. I think you really need this because, at the end of the day, everyone's work is very different. It's very unique. Software engineering is different than finance is different than operations is different than go to market and sales. And so there's a lot of these last-mile intricacies of work that need to really be done.

20:33

Adoption. Explore the full extent of the capabilities. Apply to specific workflows. Do the knowledge sharing. Create excitement within folks who might want to use this technology because, in the absence of that, it's very difficult to pick up. Who would you put on this tiger team? Is it engineer-led? Do you find, in your experience, is it a cross-functional sort of team?

20:34

Yeah, it's interesting. Also, a lot of companies, software engineering-adjacent, basically technical people but are not software engineers, I think those are the ones who tend to get most excited around this. It's maybe the support team operations lead who doesn't code but loves using these tools and is an Excel wizard or something. So it's technical-adjacent or coding-adjacent and pretty technical. Those are the kinds of people I've seen in these companies who just really light up and get excited around this, and you can usually build a team around that. But yeah, it's oftentimes not software engineers. Software engineers, I think, will understand this, but not every company has software engineers. It's actually kind of a rarity. They're hard to find, they're expensive, and so it's these other types of folks.

20:35

What I'm hearing is the anti-pattern is top-down. This is very the CEO, founder, exec team just like, we are going to go AI-first, we're going to lean into AI, everyone's going to be judged on their performance using AI tools, how much your productivity is increasing thanks to AI. Without that being just top-down and not creating a team that is bottom-up, spreading the gospel, you find that doesn't work. Yeah, exactly. And the advice is find the people that are most excited, and instead of having them spread out through the organization, what you find works is create a little AI evangelist team that finds ways to use it and spreads it across the work.

20:35

Yeah, another, it's kind of like hearing you play back to another way, and empower them. Let them build hackathons. Let them hold seminars, do knowledge sharing, create the seeds of excitement internally. Okay, amazing. There's a couple hot takes I want to hear from you, something that I've seen you talk about and share. One is you've shared that talking to customers and listening to customers is not always the right strategy in AI, and it might often lead you astray.

20:37

I don't know if it's that hot of a take. I think the main thing here is obviously you should talk to your customers. I just think the AI field, especially what I've seen over the last three years working on the API and seeing all that evolve, is the field and the models themselves are just changing so quickly, they tend to disrupt themselves, especially around the tooling and scaffolding. There's this quote that I read actually earlier this week from an X article by this guy named Nicholas, who's the founder of a let them build hackathons, let them hold seminars, do knowledge sharing, create the seeds of excitement internally. Okay, amazing. There's a couple hot takes I

20:46

want to hear from you. Something that I've seen you talk about and share: one is you've shared that talking to customers and listening to customers is not always the right strategy in AI, and it might often lead you astray. I don't know if it's that hot of a take. I think the main thing here is, obviously, you should talk to your customers. I just think the AI field, especially what I've seen over the last three years working on the API and seeing all that evolve, is the field and the models themselves are just changing so quickly. They tend to disrupt themselves, especially around the tooling and scaffolding. There's this quote that I read earlier this week from an X

21:37

article by this guy named Nicholas, who's the founder of a startup called Fintool, where I think he was sharing a lot of the best practices that he has learned through building AI agents for financial services at a startup, Fintool. He had this phrase that I thought was really good, which is: the models will eat your scaffolding for breakfast. If you rewind back to 2022, right when ChatGPT launched, these models were pretty raw, and there was all this product scaffolding and things, especially in the developer space, to try and steer the model and build a scaffolding around it to get it to do what you want, like agent frameworks. There's vector stores, I think, that were

22:13

really popular back then, and much of that has gotten so much better that they ended up literally eating some of the scaffolding. I think this is even true today. I think the article from Nicholas, the current scaffolding which is fashionable, is skills, files-based context management. I could see a world where, at some point, that's no longer useful, where the model can actually manage all that skills-type thing. You have literally seen this play out. The agent frameworks, I think, are a little less useful now. There was a period of time, like 2023, where we thought vector stores were going to be the main way for you to bring organizational context into the models, and

22:56

you need to vectorize and embed every bit of your corpus, and then you do all this work to figure out the vector search, to optimize that, to pull out the right information at the right time. All of that is scaffolding because the model was not good enough. It turns out, in this case, as the models get better, a better approach is actually to take out a lot of that logic and trust the model and give it a set of tools for search. It doesn't need to be vector stores. I know a lot of companies are still using it, but the

24:11

entire scaffolding around that, and building an entire ecosystem around that, and assuming that's the only scaffolding that you need, has really changed. Tying this back to you don't always have to listen to your customers: because the field is changing so much, at any point in time a lot of people are in a local maximum. If you had only chased down that path, it actually would have led you to build something that is a local maximum, whereas as the models get better, we've had to reinvent and rethink the right abstractions and the right tools and frameworks to build around these models. The cool, exciting, crazy, annoying part is it's a moving target. The current

24:54

smattering of tools and frameworks right now will likely need to change as the models get smarter and better, but that is just the nature of building in this space. I think that's what makes it exciting, but it also means when you talk to customers, you need to balance the exact feedback that they want with where you think the models are going and where you think things will trend over the next one or two years. It's interesting how this is the bitter lesson. This big lesson that AI and ML folks learned is the less you overcomplicate, the less logic you add to machine learning, to AI, the more it'll be able to scale and grow. Take it all away and let it compute. Give it

25:42

more power to get smarter. OpenAI API team has been guilty of this, where we took some left and right turns when we shouldn't have. But the models get better, and we're all learning the bitter lesson day in, day out. What would be the key takeaway for folks building on the API or just building agents and having to build a little bit of this around for now? What would be— My view is it is clearly a moving target, and I think a lot of the companies that I've seen really do well build a product for an ideal type of capability that is maybe 80% of the way there today. They end up having a product that works but is just almost there. As the models get better, suddenly it might

26:32

click, and their product now is incredible because it works. Maybe with 0.3, at some point it suddenly works. With 5.1, 5.2, suddenly it unlocks it. But they're building these products with the model capability improvements in mind, and with that you end up creating an experience that's way better than if you had assumed that it's static in the first place. That would be my general advice, which is: build for where the models are getting. They are so much better so quickly that you often don't need to wait that long. To follow that thread, in the next 6 to 12 months, where is the API heading? Where is the platform heading? Where are the models heading, as much as you can

27:12

share? I know there's a lot of secrets here that maybe you're more excited about. Software engineering tasks, and how long of a task can these models do 50% of the time, 80% of the time. I think we're at something like multi-hour tasks being able to be done by software engineering tasks being able to be done by these frontier models 50% of the time. And then I think 80% is something like just under an hour. But the sobering thing about that chart is they plot all the previous models on this chart as well. So you can really see the trend of this. That's something that I'm really excited about, which is, I actually think products today are really optimized for

27:19

tasks that the model can do for minutes at a time. Even Codex and the coding tools, I'd say it's in the CLI. You're seeing it be interactive. It's quite optimized well for maybe at most 10-minute-type tasks. I have seen people push Codex to the limit into multi-hour-long tasks, but I think that's more of the exception. If you follow this trend, I think in the next 12 to 18 months, we could see models that could do multi-hour-long tasks very coherently. At some point, it might reach six-hour-a-day-long tasks where you dispatch it and have it do things on its own for a while. The types of products you build around that will look very different. You want to

27:22

give the model feedback. You obviously don't want it to completely run wild for a day. Maybe you do, but you probably don't. And then the universe of things you can have the model do will really expand. That's something that I'm really excited about seeing. Another thing over the next 12 to 18 months where I think it'd be really cool is improvements in the multimodal models. Actually, by multimodality, I'm mostly thinking about audio here, where the models are pretty good at audio. I think they're going to get a lot better at audio over the next six to 12 months, especially the native multimodal models, the speech-to-speech ones.

27:24

I think there's also interesting work being done around new types of models and architectures on the multimodal audio side as well. But audio, especially in the enterprise and in a business setting, I think is a hugely underrated domain still. Everyone talks about coding. It's all text. But we're talking in audio. A lot of the world's business is done via audio. A lot of services and operations are done via talking in audio. And so I think that that area is going to look very exciting in the next 12 to 18 months. And I think there will be even more unlock for what we can do with audio models there as well. Amazing.

27:26

So quick summary: expect agents and AI tools to run longer, that trajectory to continue to increase, and then audio and speech becoming a bigger deal, more first-party and native and better, and core to the experience. Yeah. Extremely cool. Okay.

27:31

I want to go back to one of your hot takes, another hot take that I've seen you discuss. You're very bullish on business process automation as an opportunity in the world of AI. Talk about that. Yeah, this goes back to the thing that I said previously, which is we live in a bubble in Silicon Valley. And a lot of the work that we do, that we're used to, software engineering, product management, building products, is very differently shaped than the work that goes on that area is going to look very exciting in the next 12 to 18 months. And I think there will be even more unlock for what we can do with audio models there as well.

27:45

Amazing. So, quick summary: expect agents and AI tools to run longer, that trajectory to continue to increase, and then audio and speech becoming a bigger deal, more first party and native and better and core to the experience. Yeah. Extremely cool. Okay. I want to go back to one of your hot takes, another hot take that I've seen you discuss. You're very bullish on business process automation as an opportunity in the world of AI. Talk about that.

27:46

Yeah, this goes back to the thing that I said previously, which is we live in a bubble in Silicon Valley. And a lot of the work that we do, that we're used to, software engineering, product management, building products, is very differently shaped than the work that goes on that runs our entire economy. And I see the same thing now when I talk to customers.

27:46

If you talk to any company that's not based in, it's not a tech company, there's a lot of business processes. And so what I mean by this is, I generally delineate it as software engineering is open-ended knowledge work. And this is why I think tools like Codex tend to be quite good, because it's exploring and you're giving it these open-ended things. But software engineering is fundamentally pretty open-ended, and it's not very repeatable. Right? So you build a feature, you're not trying to build the exact same feature over and over again. And a lot of tech jobs are in this space. I think data science is in this space as well, even some of the strategic finance stuff. But as you move further and further away from software engineering and what is core in tech, a lot of jobs are just business processes. They're repeatable things, repeatable operations, that some manager at a company has iterated on. There's usually a standard operating procedure that people want to follow, and you don't want to deviate from it that much. In software engineering, the ingenuity is in deviating, but a lot of the work being done in the world is actually just running through these procedures and operations. If I call a support line, they're running through one of these. If I call my utility company, there's a bunch of processes and things that they can and cannot do for me. And so I'm extremely bullish on this general category, and I think it's underrated because it's so different from what we think about in Silicon Valley. People tend not to think about it. But how can we apply AI and some of the tools and frameworks that we have toward this business process automation, toward automating and making easier repeatable business processes with high determinism that are fully integrated with business data and business decisions and different systems within an enterprise? And how can we actually make that process better? Because I actually think there's a lot of opportunity and a lot of work to be done in that area. And we just don't talk about it because it's a little bit less in our wheelhouse.

27:47

So your take here, just to make sure I fully understand it, is you think there's a much bigger opportunity outside of engineering for AI to impact productivity of companies and also jobs of these folks that are doing these repetitive, easily automated tasks.

27:47

Impact jobs and also just impact how work is done. So much of work is done in this way. You think about what I talk to customers about all the time, big enterprises, how will AI transform my company? How will it run in a world with AI in 20 years? And software engineering is part of the story, but there's so much more on the business process side. And I actually think it might look even more different on the business process side, and the work there is pretty substantial. It's actually interesting. I don't know, from an absolute percentage or absolute base, if it's bigger or smaller than software engineering. Software is pretty huge and pretty expensive as well, but it is pretty massive, and it's definitely bigger than you would think it is based off of how people talk about it, or don't talk about it, on X or Twitter.

27:47

Okay. In going in a slightly different direction, having built the platform, building the API, people building on API, the biggest question on people's minds is always just, how do I not have OpenAI squash my idea and build their own thing and then destroy this market I created? What's the general policy? What's the general philosophy of how startups should think about where OpenAI is unlikely to go?

27:47

My general answer here is the market is so big and so massive. I actually think startups should just not overly think about where OpenAI or these labs are going. I've talked to a lot of startups that have not worked out, startups that are doing really well. Every startup that I've seen that has fizzled out is not because OpenAI or a big lab or Google or something has come to squash them. It's because they built something and it really didn't resonate with the customers. Whereas the ones that take off, even in very competitive spaces like coding, like Cursor is huge at this point, and it's because they built something that people really love. And so my general advice is, don't overly stress about this. Just build something that people like, and you will have a space in this. I can't overstate how big of an opportunity there is right now. The opportunity space of building with AI is so big. A good example of this is the space is so big that the Overton window of what is acceptable and not acceptable for VCs to do has completely changed here. VCs are investing in competitive companies left and right. It's just like the space is so big because the opportunity is unlike anything that we've seen before. And while that affects how VCs operate, from a startup perspective, it's the most empowering thing in the world because even if you just build something that some people really, really love, you will end up with a massively valuable business. And so that's why I tell people, don't overly think about it.

27:47

The other thing I also think is important to remember, at least from an OpenAI perspective, one thing that we've always held very near and dear, which both Sam and Greg helped reinforce from the top as well, is we actually view ourselves fundamentally as an ecosystem platform company. The API was our first product. We think it's really important for us to foster this ecosystem and continue to support it and not squash it. And so if you look at the decisions we make, this is all we've done through it. Every single model we've released in one of our products gets released in the API. We really released these Codex models now that are a little bit more optimized for the Codex harness, but they always find their way into the API, and all of our customers end up using those. We don't hold back on any of that. We think it's really important to keep our platform neutral. And so we don't block competitors. We allow people to have access to our models. We also want, we've recently been testing more of the sign in with ChatGPT product as well, and so we want to foster this ecosystem. I think it's really important that we do so. The general thinking about this is a rising tide lifts all boats, and we might be an aircraft carrier. We're pretty big at this point, but we think it's important to raise the tide because everyone benefits. And I think we'll benefit as well, like our API itself,

27:48

hold back on any of that. We think it's really important to keep our platform neutral. And so we don't block competitors. We allow people to have access to our models.

27:48

We also want, we've recently been testing more of the sign in with chat GPT product as well. And so we want to foster this ecosystem. I think it's really important that we do so. The general thinking about this is, a rising tide lifts all boats and we might be an aircraft carrier. We're pretty big at this point, but we think it's important to raise the tide, because everyone benefits. And I think we'll benefit as well. Our API itself, it's grown pretty significantly because we act in this way. And so I'd really encourage people not to view open AI as this thing that'll just shove people out of the way, but instead focus on building something valuable. And we remain committed to providing an open ecosystem.

27:49

Why is that important to open AI? Just this focus on building a platform, creating a way for people to build businesses. Is that just, that's been the vision from the beginning. We want this to be a platform.

27:49

It's been the vision from the beginning. It goes back to our charter, actually, our mission. So open AI's mission has always been to, one, build AGI. So that's why we're doing that. But then the second thing is to spread the benefits of it to all of humanity. And the main part there is all humanity. Obviously ChatGPD is trying to do this. We're trying to reach however many, the whole world. But very early on, and this is why we launched the API back in, I think it was 2020 or something, really early, we don't think we, as a company, will be able to reach all of humanity, right? There's, I don't know, every corner of the world's pretty deep. And so we actually feel like, in order for us to fulfill our mission, we need to have some platform-style thing here where we can empower other people to build the customer support bot for podcasters and newsletter hosts, because we're not going to be able to do it ourselves. And so we've largely seen this play out with the API. This is why we talk to so many of our customers and really love seeing the diversity of things built on it. But yeah, it's been there since A1, because we view it as an expression of our mission.

27:49

And you haven't even mentioned the app store that you guys are launching, the ChatGPT app store. Yeah. Is that under your umbrella, by the way, or is that a different org and team?

27:50

It's a different team. So it's under ChatGPT. We obviously collaborate very closely with them and they built an apps SDK, which is built in close collaboration with our team. But that is more within the ChatGPT umbrella. But that is also another example of this, right? ChatGPT is like, we have these 800 million weekly active users who are just coming over and over again. It's a great asset to have as a business, but man, would it be better if we could somehow allow other companies to come in and take advantage of this as well and build for this audience as well? And then ultimately we think it'll help us expand that group as well, right? And so it all comes back to the mission. And we find that being a platform, being open, tends to help here.

27:51

Just that number, 800 million, I think it's M M A, is just like weekly, weekly, weekly, weekly act. Yeah. It's crazy. Billion people using weekly. It's how many, how these numbers we're just used to now, but that's insane, unprecedented. Yeah. It's mind-boggling for me to think about from a scale perspective. Honestly, the way I think about it is 10% of the world, and growing by the way, it's shooting up, come to chat GPT and use it every day, or sorry, every week.

27:52

At this point, I just want to double down on this point you're making. Open AI's mission was to make AI available to all of humanity. And I think some people diss that. They're like, oh, it costs money. And it's like, the fact that there's a free version of chat GPT that anybody can use, that is not so different from the most powerful AI model that exists in the world, for free. That's not gated, that anyone can use. If you're a billionaire, there's only so much more you can get out of AI than what someone in a village in Africa can get. And I know that's always been really important to open AI.

27:52

Yeah. Yeah. I mean, look, that's why I think we've leaned into the health work. We've leaned into education. Education is going to be very interesting here. The other insane trend here is the free model has gotten so smart over time. The free model back in 2022 was good at the time, but it's nothing compared to what you get today, because you get GPT-5 today. And so raising the floor across the world is something that we're really trying to do. And then we view it as part of our mission. The other flip side of this, by the way, is talking about the billionaires or whatever. I know people love saying you're using the same iPhone that Mark Zuckerberg's probably using, or the billionaires are using. But for $20 a month, you're basically using the same AI that the billionaires are using. For $200 a month, you get the same pro model that all the billionaires are using, but they're probably not using pro for everything. They're probably just using the plus tier ones for their day in and day out. And so, yeah, this democratization and just spreading of this benefit across all of the world is something that's really meaningful to us and something that drives a lot of what we do.

27:54

One last question, just for folks that are thinking about building on the API or just, oh wait, I could do cool stuff with OpenAI's models and APIs. What does your API and platform allow people to do? I know you can build agents on top of the platform. Just talk about what you allow.

27:54

So fundamentally, the API offers a bunch of developer endpoints, and these developer endpoints basically let you sample from our models. The most popular one that we have right now is one called responses API. And so this is an endpoint, and it's optimized for building long-running agents, so agents that'll work for a while. So what you can basically use, at a very low level, you're basically just giving the model text. The model will work for a while. You can pull it to see what it'll do, and then you'll get the model response back at some point. That's the lowest-level primitive that we have for people, and that's actually what a lot of people use. That's the most popular way of building on top of our API. With that, it is super unopinionated and you can do basically whatever you want. It's like the lowest-level thing. We've also started building more and more kind of like

27:54

now is one called responses API. And so this is an endpoint, and it's optimized for building long-running agents, so agents that'll work for a while. So what you can use, you can, you can, at a very low level, you're basically just giving the model text. The model will work for a while. You can pull it to see what it'll do, and then you'll get the model response back at some point. That's the lowest-level primitive that we have for people. And that's actually what a lot of people use. That's the most popular way of building on top of our API. With that, it is super unopinionated, and you can do basically whatever you want. It's the lowest-level thing. We've also started building more and more layers of abstraction on top to help people build some of these. And so next layer up, we have this thing called the agents SDK, which has also gotten extremely, extremely popular. This allows you to use the response API or some other API endpoints that we have to build what you might more traditionally think of as an agent, an AI working in an infinite loop. It might have sub-agents that it delegates to. It starts building all this framework, all the scaffolding. Actually, we'll see where this all goes, but it makes it a lot easier for you to build these, these, these, these agents, giving it guardrails, allowing it to farm out sub-tasks to other agents and orchestrate a swarm of agents. The agents SDK allows you to do that. And then above that, we've now started building tools to help also with the meta level of deploying an agent. So we have this product called agent kits and widgets, which are basically a bunch of UI components that you can use to very easily build a very beautiful UI on top of either our API or agents SDK, because a lot of times these agents look very similar from a UI perspective. And so there's agent kit. We also have a smattering of evals products, like evals API, where if you want to test and see if your models or your agent or your workflow's working, you can test it in a very quantitative way using our evals product. And so, yeah, I view it as these various layers. They're all helping you build what you want with our AI, with our models, and with increasing levels of abstraction and how opinionated it is. And so you can start, you can do that. You can use the whole stack, and it very quickly allows you to build an agent, or you can go down the stack as low as you want to basically responses API and build whatever you want because of how low-level it is.

27:54

Sherwin, is there anything else that you want to share? Anything else you want to leave listeners with? Anything we haven't touched on that you think might be helpful before we get to our very exciting lightning round?

27:55

Sherwin Lowe The only thing I'd leave folks with is, yeah, I think the next two to three years are going to be some of the most fun in tech and in the startup world that we'll have in a very long time. And I would just encourage people not to take it for granted. I entered the workforce in 2014. It was great for a couple of years. I felt like there was a period of five to six years where it wasn't very exciting in tech. And then the last three years have just been the most insanely exciting, energizing period of my career. And I think the next two to three years are going to be a continuation of that. And so I would encourage people not to take it for granted. I'm trying not to take it for granted. At some point, this wave is going to play out, and it's going to be a lot more incremental. But in the meantime, we're going to get to explore a lot of really cool things, invent a lot of new things, and change the world and change how we work. And so that's the main thing I'd leave folks with.

27:55

I love this message. I want to spend a little more time on it. When you say don't miss it, what do you recommend people do? Is it just build, lean in, learn, join a company building really interesting things? What's your advice to folks that are like, okay, I don't want to miss the boat?

27:55

Yeah. I would just say engage with it. So it's basically what you said: lean in. Building tools on top of this is part of the story. Just using the tools. You don't need to be a software engineer to lean into this. I think a lot of jobs are going to, going to, going to change here. So just using the tools, understanding the limitations of what it can and cannot do so that you can watch the trend of what it can start to do as the models improve. And so it's basically getting used to this technology and getting familiar with it instead of laying back and letting it pass you.

27:55

On the flip side of that, there's a lot of stress and anxiety around, there's so much happening. How do I keep up? I got to learn about Clotbot this week. Oh God. Yeah. Is there something you learned about it? You're at the center of this. How do you not get overly stressed and worried about missing things that are going on and just stay on top of news? What are some things you've learned?

27:56

Yeah. So I think I'm personally a bad example of this because I am basically chronically online on X and our company Slack. So I actually end up absorbing a lot of it. What I will say, though, just from observing other folks who are less addicted to this stuff like I am, yeah, a lot of it is noise. You don't need to have 110% of this pass your mind, go into your mind. Honestly, just leaning into one or two different tools, starting small, is already more than you need here. I think just the combination of the frenetic pace of the industry and X as a product creates this insane pace of news, which is honestly very overwhelming. The main thing is you don't need to know all of that to really engage with what's happening right now. And even something as simple as just install the codex client, play around with it, install chat, connect it to a couple of your internal data sources, Notion, Slack, GitHub, and see what it can and cannot do. All of that, I think, is a part of it.

27:56

Amazing. Sherwin, with that, we've reached our very exciting lightning round. I've got five questions for you. Are you ready? Yeah. Yeah, absolutely. First question: What are two or three books that you find yourself recommending most to other people? Oh, I'll talk about one nonfiction, one fiction book. The fiction book was, I just finished reading it. I, I, it was really, I really recommend it. It's There Is No Antimemetics Division by QNTM. It's, I think, an online author, but I saw it being shared on X. This, this, it's a science fiction-y kind of book, and I basically devoured it in two days. It's super, super well-written, super fascinating. It's about a government agency that's fighting things that make you forget it. And so it's just a very smart, creative book.

27:57

find yourself recommending most to other people? Oh, I'll talk about one nonfiction, one-on-one fiction book.

27:57

The fiction book was, I just finished reading it. I really recommend it. It's There Is No Anti-Memetics Division by QNTM. I think it's an online author, but I saw it being shared on X. It's a science fiction-y book, and I devoured it in two days. It's super well-written, super fascinating. It's about a government agency that's fighting things that make you forget it. It's just a very smart, creative book, and fresh, honestly, in terms of source material, that I really like. So I'd recommend that one. The book is also unintentionally hilarious. It's meant to be this sci-fi, almost horror-style book, but it made me laugh a couple of times. So that's the fiction book.

27:57

Nonfiction. I'm going to cheat, and I'm going to recommend two of them. In the last year, I've been reading a lot more about China and the U.S.-China relations. And I think there are two books that came out in the last year that have been really eye-opening for me in that regard.

27:57

The first one is the Dan Wang book, Breakneck. That one was really, really good. I really liked his analogy of the lawyerly U.S. as the lawyerly society, and China as the engineering society. And there are pros and cons to each. I read it and I was like, hmm, yeah, it does seem like we were run by lawyers in the U.S. So that's one. And the other one is the Patrick McGee book on Apple in China. It was super, super interesting. I'm a huge Apple fanboy. If you could see my desk right now, it's all Apple stuff. But it was super fascinating learning about Apple's relationship to China. And then, too, it had a lot of inside information about Apple as a company that I found fascinating. So it was also quite a page-turner and also very, very timely, a timely book as well.

27:58

The anti-memetics book sounds amazing. I'm buying it right now as you're talking. Yeah. Yeah. Yeah. I think it's only a couple hundred pages. I literally finished it in two days. It was just so, so good. Okay. Great tip. Okay. Favorite recent movie or TV show you have really enjoyed?

27:59

Yeah, that one's tough. With two kids and a busy job, I really haven't had much time to watch TV shows. I will say in the last couple of weeks, I watched a couple episodes on, I'm actually a big anime guy, and so I watched a couple episodes. There's a new season of this anime called Jujutsu Kaisen that's out. So season three of JJK was really good. In general, I'm a huge fan of Japanese anime. I think they create the most novel and unique plots and universes that Western media has shied away from. And so, generally, I'm a big fan of that. But yeah, I haven't really watched much, but saw a couple episodes of JJK recently. Extremely understandable in your role.

27:59

Yeah. Favorite product you recently discovered that you really love.

27:59

Yeah. Okay. So I recently had to set up Wi-Fi and home networking, and I went all in on Ubiquiti routers and security cameras. I'd never heard of it before I had to do this. I always just had a very simple setup. It's just such a well-built product. I don't know if you've used it before, but it's basically like the Apple of home networking. So, beautiful products. But the thing that actually makes it extremely good is that the software is good. And so they have a really great mobile app to help manage all of the home networking. So basically, with Ubiquiti, you can use it to buy wireless routers. You need ethernet wiring throughout your house to use it. But I actually think what makes it really good are the security cameras. So if you have security cameras that are plugged into the Ubiquiti ecosystem, they have an incredible mobile app, Apple TV app, and iPad app to see the live feed of your cameras. And so they're a little pricey, but not that pricey. But it's been just an incredible product experience.

28:00

All right. I went Eero, so I made a mistake. Good tip. Eeros are pretty good too, but I'm fully converted to Ubiquiti at this point. Okay. Good tip. Okay. Two more questions. Do you have a favorite life motto that you find yourself coming back to in work or in life? Yeah. The one that I always repeat to myself is never feel sorry for yourself. There's a lot of things that are going to happen at work, in life, and reminding yourself to never feel sorry, and that you always have a sense of agency to pull yourself up, is something that I've had to tell myself a lot, and also something that I repeat to a lot of other folks as well.

28:01

Last question. So in your previous life, you worked at Open Door, where you led work on basically figuring out how much to pay for houses. You basically built the model that told the company, here's how much we'll pay for this house. What's a variable in the price of a house that you didn't expect is really important and impacts the price of a house?

28:01

There's a bunch that were surprising. I'll maybe list a couple of the most interesting ones. Power lines and high-voltage power lines actually impact your price quite a lot. I didn't really fully internalize this until I went to Dallas and observed that when your house sits next to one of these giant voltage lines, it's buzzing, and most people have families. You don't want your kids near there. So I think that was one that really, really surprised me. That makes sense.

28:02

Yeah. And then the other one, which was something that was always really difficult for us to quantify, was floor plans. And so it is very important. Yes, of course it's really important. But quantifying what a good floor plan is like and what a really bad floor plan is like, we were doing all these things with how wide is the kitchen, and what style of kitchen is it, and where's the master bedroom. And so it was just really, really hard to quantify. But I remember floor plan was a big one because we'd have a home that wouldn't sell, and then our ops team would go in and be like, yeah, that's a floor plan issue. So how could you tell us? You go inside, you just feel it. You know, the floor plan feels off. So yeah, those are ones that were surprising. And then the last one that was more impactful than I thought is general curb appeal and even the front door. And so I actually think there's a Zillow book on this where front door replacement tends to be the highest ROI for homes. But the feel of, as you walk up to the home as a buyer, what you're interacting with in the first moments of the house, I think I'd underrated its importance.

28:03

That is extremely interesting, and I love that you had to figure out how to do all this. you just feel it. It feels, the floor plan feels soft. So yeah, those, those are ones that were surprising. And then the last one that was more impactful than I thought is general curb appeal and even the front door. And so I actually think there, there's a Zillow book on this where the front door replacement tends to be the highest ROI for homes. But just the feel of, as you walk up to the home as a buyer, what you're interacting with in the first moments of the house, I think I'd underrated its importance.

28:06

That is extremely interesting. And I love that you had to figure, figure out how to do all this in code and not. Yeah. And floor plans. I have a bunch of stories around floor plans. There's, there's, it's not digitized. So there's a handful of people who have paper floor plans of all these homes in Phoenix and Dallas. Yeah, a lot, a lot of fun, fun stories from the open door days. Okay. Sherwin, thank you so much for doing this. This was incredible. Where can folks find you online, and how can listeners be useful to you? Yeah. So I'm online on Twitter, on X. I'm just at Sherwin Wu and yeah, I mostly just

28:07

tweet about open AI and API and some of the products that we're launching. And then how folks can be useful to me: I love hearing about things that people are building. And so if you're working on a startup, if you're hacking on an idea, would love to just reach out to me on X. I would love to hear about what you're building and learn about how open AI can help support you. Amazing. Sherwin, thank you so much for being here. Yeah. Thank you, Lenny. Bye everyone. Thank you so much for listening. If you found this valuable, you can subscribe to the show on Apple Podcasts, Spotify, or your favorite podcast app. Also,

28:08

please consider giving us a rating or leaving a review, as that really helps other listeners find the podcast. You can find all past episodes or learn more about the show at lenny'spodcast.com. See you in the next episode. to AI to build these businesses for themselves I every layer okay so the billion dollar startup I think about this a lot because I I'm not going to be a billion dollar startup because what I'm doing is not venture scale in any way and not super high leverage but just could see how many support tickets I get from just like the most ridiculous things it's hard for me to imagine one person like I'm bearish on this billion dollar startup I just want to

28:53

share this thought simply because of the support costs even if AI is helping you at a billion dollars just like unless your ACVs are very high and you have very few customers just dealing with support and people are like they can solve ! They solve ! I'll email support I'll ask about this thing just dealing with that is hard to scale in my experience so unless you have in my opinion unless you have a bunch of contractors which I don't know does that count as a single person company I feel like it's very difficult to scale a billion dollar startup and not have someone helping you with at least the support work and AI I think will take you so far so I think that's true and

29:33

actually I think my view on being the one person who has to dispatch an AI to solve and fix those support tickets I think what might end up happening is there might be a whole smattering of other startups that are building software and super tailored towards what you might need and so there might be 10 or 20 startups that build support software for podcasts and newsletters and that might be a one ! person startup they might be able to code up this product very easily they are able to build their own thing and because it's so tailored and unique and hopefully useful for you it might be something that you purchase as the one person billion startup there's a question of what

30:31

you in-house and what you outsource and what might happen is because the cost of ! software and building products is collapsing so much you might end up outsourcing a lot of this and in doing so reducing the size of your company and so that's the world that might end up happening again there's high uncertainty in what might play out here but the end result still might be one person driving this high massive leveraged company that might actually reach a billion dollars I could see that I also think about Peter at Claudebot slash Moldbot slash OpenClaw of just like how he barraged he is right now by all these asks and emails and pings and DMs and PRs just like I'm curious

31:08

and he's not even making any money out this thing yeah I can't imagine what it's like to be him right now it must be like absolutely insane it it's probably like you know like the the months after we launched chat gbt the craziness that was yeah as one man he's coming out on the pod by the way in a week oh that's exciting yeah maybe the fourth order effect is distribution becomes increasingly important because there are so many freaking things trying to get your attention so people with an audience and platform I think become more and more valuable which is good good stuff okay I wanted to come back to you just thinking about you as a manager of a team that is building

31:52

the platform that powers basically the entire AI economy like every AI startup is building on your API clearly you're doing a great job what other core management lessons have you learned what do you find is really important and key to your success as a manager of engineers and just people yeah I think a lot of the lessons that I've learned here I don't know how specific it is to the opening API or some of our enterprise products in particular I think my management philosophy has obviously changed over time but I think it probably stayed the same more than it's changed over time one of these principles is kind of what I talked to you performers and really trying your best

32:47

to empower them the way that I think about it is kind of come back to this analogy of software engineer as a surgeon which comes from the mythical man month book so it's funny so I pull it from the book but in the book they actually describe this world where I think they were predicting the future because I think the book was in the 70s or something they said that software engineering might end up moving into a world where the software engineers are like surgeons or like in a surgery room there's like one person doing the work and you know there's the one person like cutting whatever and doing all the surgery and everyone else in the room is there to support them right

33:26

it's like the nurse and the assistant the resident and the fellow and the surgeon is like I need a scalpel and they give them scalpel and then they're like I need this tool and this machine and they'll bring it over everyone's there to just like support the one surgeon and so the myth mammoth actually predicted that that is the direction that software is going to go I don't think that's exactly played out where it's much more collaborative and philosophy which is software engineering isn't really like surgery where it's not just one person doing work but the way in which I like treating the people on my team and the way that I act as a manager is I looking around corners

34:30

and unblocking people especially from an organizational perspective is extremely extremely useful and again going back to the AI conversations even more important nowadays right like if people are just like cranking PR after PR the main thing bottlenecking progress and you know shipping something tends to be organizational or like process oriented and if you as a manager can kind of look around corners and kind of unblock the that that's the best case scenario that's kind of the way that I approach management and especially engineering management and so that's something that's really stuck with me over time and even though software engineers aren't exactly surgeons that

35:14

metaphor has always stayed in my mind as of my career I love that and I feel like I wonder if that's something AI can help with look around corners and predict here this engineer is going to be blocked by this decision we need to figure this out we need to get that's actually a really good point I haven't tried this yet but I wonder what would happen if I ask ChadGBT hooked up to company knowledge what are the active blockers look through all the notion docs maybe Slack messages what are the active blockers on my team and is there something I can do to help I have not thought about that but you're right just had an insight right here and I think even more interestingly

35:52

what do you anticipate will be a blocker for this engineer or this team in the coming months you ask the model you ask the AI to do the second and third order anticipate what the blockers will be next month I think we've got a good idea right platform product managers at the world's best companies use Datadog the same platform their engineers rely on every day to connect product insights to product issues like bugs UX friction and business impact it starts with product analytics where PMs can watch replays review funnels dive into retention and explore their growth metrics where other tools stop Datadog goes even further it helps you actually diagnose the impact of funnel

36:40

drop-offs and bugs and UX friction once you know where to focus experiments prove what works I saw this firsthand when I was at Airbnb where our experimentation platform was critical for analyzing what worked and where things went wrong and the same team that built experimentation at Airbnb built Eppo Datadog then lets you go beyond the numbers with session replay watch exactly how users interact with heat maps and scroll maps to truly understand so that you can roll out safely target precisely and learn continuously Datadog is more than engineering metrics it's where great product teams learn faster fix smarter and ship with confidence request a demo at datadoghq.com

37:25

slash Lenny that's datadoghq.com slash Lenny okay I'm going to shift to talking about the API and the platform that you all build so you work with a lot of companies implementing your API your platform building you told me that you find that a lot of companies actually have negative ROI on their AI deployments which I think is what a lot of people read about and feel and think and it's interesting actually seeing that what what's going on there what are they doing wrong what what's happening in the world of AI and deployments in ROI yeah so to be clear I don't like explicitly see quantitative numbers around this you know it's actually really hard to measure these things

38:07

but especially from observing some companies kind of trying to do AI I would not be surprised if a lot of AI deployments are actually negative ROI I mean part of this too is I think there's also general sentiment from folks around the a couple things I've observed around this so one thing is and I think I come back to this again and again like I think we in Silicon Valley just forget that we live in a bubble like we are so like Twitter is a bubble is a bubble Silicon Valley is a bubble software engineering is a bubble most people in the world most people in the US are not software engineers are not very AI pilled are not following every single model release and so we're

39:01

just highly out of the loop on how to use this technology and so we always talk about all these best practices for codex all these codex people within open AI I'm sure everyone on X who posts are crazy power users of these AI tools they lean into skills they lean into when I talk to some of these companies and I talk to the actual employees using these it's like the most basic thing that they're trying to do and they have very little understanding of exactly how this technology works and so that that's kind of like one big observation for me which is like they're asking very simple questions of these things they're really not pushing it just yet and so that what a more

39:56

ideal AI deployment setup looks like and this is kind of how we've run things within OpenAI too the companies where I think it started to work really well have a combination of both top down buy-in so it's like the C-suite we want to become an AI first company so there's buy-in they buy the tools they have exec support but it also has bottoms up adoption and buy-in and so what I mean by that is it has actual employees doing the work who are really excited about this technology and are willing to learn evangelize build best practices and kind of knowledge share within the organization we've seen this a lot internally so obviously OpenAI has always wanted to be a very AI

40:37

centric company but when it really started taking off was with the introduction of codecs and these tools where actual employees themselves could start applying it to their work and I think you really need this because at the end of the day everyone's work is very different it's very unique software engineering is different than finance is different than operations is different than go to market and sales and so there's a lot of these last mile intricacies of work that needs to really be done adoption explore the full extent of the capabilities apply to specific workflows do the knowledge sharing create excitement within folks who might want to use this technology

41:52

because in the absence of that it's very difficult to pick up and who would you put on this Tiger team is it like engineer led do you find in your experience is it a cross functional sort of team yeah it's it's interesting also a lot of companies software engineering adjacent like basically technical people but are not software engineers I think those are the ones who tend to get most excited around this it's like you know maybe the it's like maybe the like you know support team operations lead who doesn't code but loves using these tools and you know is like an Excel wizard or something and so it's like technical adjacent or like coding adjacent and like you know pretty

42:37

technical those are the kinds of like those are the kinds of people I've seen in these companies who just like really light up and get excited around this and you can usually build a team around that but yeah it's like oftentimes not software engineers software engineers I think will understand this but not every company has a software engineers is actually kind of rarity they're hard to find they're expensive and so it's these other types of folks what I'm hearing is the anti pattern is top down this is very the CEO found exec team just like we are going to go AI first we're going to lean into AI everyone's going to be judged on their performance using AI tools how much

43:10

your productivity is increasing thanks to AI and without with that being just top down and not creating a team that is bottom up spreading the gospel you find that doesn't work yeah exactly and the advice is find the people that are most excited and instead of having them spread out through the organization what you find works is create a little AI evangelist team that finds ways to use it and spreads it across the work yeah another it's kind of like hearing you play back to another way and empower them you know let them build hackathons let them you know hold seminars do knowledge sharing kind of create the seeds of excitement internally okay amazing there's a couple hot

43:59

takes I want to hear from you something that I've seen you talk about and share one is you've shared that talking to customers and listening to customers is not always the right strategy in AI and it might often lead you astray I don't know if it's that hot of a take I think the main thing here is obviously you should talk to your customers I just think the AI field especially what I've seen over the last three years working on the API and seeing all that evolve is the field and the models themselves are just changing so quickly they tend to disrupt themselves especially around the tooling and scaffolding there's this quote that I read actually earlier this week from an X

44:46

article by this guy named Nicholas who's the founder of a startup called Fintool where I think he was sharing a lot of the best practices that he has learned through building AI agents for financial services at a startup Fintool and he had this phrase that I thought was really good which is the models will eat your scaffolding for breakfast like if you look if you rewind back to 2022 right when ChatGPT launched these models were pretty raw and there was like all this product scaffolding and things especially in the developer space to basically try and steer the model and build a scaffolding around it to get it to do what you want like agent frameworks there's like like

45:24

vector stores I think was really popular back then and just much that and gotten so much better that they ended up yeah literally eating some of some of the scaffolding and I think this is even true today so I think the article from Nicholas actually you know the current scaffolding which is fashionable is skills files based context management I could see a world where at some point you know that's no longer useful where the model can actually manage all that skills type thing you have literally seen this play out right like the agent framers I think are a little less useful now there was a period time like 2023 where we thought vector stores is going to be the main way

46:14

for you to bring organizational context into the models and you need to vectorize and embed every bit of your corpuses and then you do all this work to figure out the vector search to optimize that to fill out the right information at the time all of that is scaffolding because the model was not good enough and turns out in this case it turns out as the models get better a better approach is actually to take out a lot of that logic and trust the model and give it a set of tools for search it doesn't need to be vector stores I know a lot of companies are still using it but the entire scaffolding around that and building an entire ecosystem around that and assuming that's

47:02

the only scaffolding that you need has really changed and so tying this back to the you don't always have to listen to your customers because the field is changing so much at any point in time a lot of people are in local maximum and if you had only chased down that path it actually would have led you to build something that is local maxima whereas as the models get better we've had to reinvent and rethink the right abstractions and the right tools and frameworks to build around these models and the cool slash exciting slash crazy annoying part is it's a moving target and so the current smattering of tools and frameworks right now will likely need as the models get

47:56

smarter and better but that is just the nature of building this space I think that's what makes it exciting but it also means when you talk to customers you need to balance the exact feedback that they want with where you think the models are going and where you think things will trend over the next one or two years it's interesting how this is the bitter lesson is this big lesson that AI and ML folks learned which is just like the less you over complicate the less logic you add to machine learning to AI the more it'll be able to scale and grow and just take it all away and let it just compute basically give it more power to get smarter !

48:43

OpenAI API team has been guilty of this where we took some left and right turns when we shouldn't have but yeah the models get better and we're all learning the bitter lesson day in day out so what would be the key takeaway for folks building on the API or just building agents and having to build a little bit of this around for now ! What would be ! My is clearly moving target and I think a lot of the companies that I've seen really do well is they build a product for an ideal type of capability that is maybe 80% of the way there today and they end up having a product that kind of works but is just almost there but as the models get better suddenly it might click and

49:42

their product now is incredible because it works maybe with 0.3 at some point it suddenly works with 5.1 5.2 suddenly it unlocks it but they're building these products with the model capability improvements in mind and with that you end up creating an experience that's way better than if you had assumed that it's static in the first place and so that would be my general advice which is build for where the models are getting so much better so quickly you often don't need to wait that long so to follow that thread where are like in the next 6 to 12 months where is the API heading where is the platform heading where are the models heading as much as much you can share I know

50:27

there's a lot of secrets here that maybe you're more excited or do ! people should start to software engineering tasks and how long of a task can these models do 50% of the time, 80% of the time. I think we're at something like multi-hour tasks being able to be done by software engineering tasks being able to be done by these frontier models 50% of the time. And then I think 80% is something like just under an hour. But the sobering thing about that chart is they plot all the previous models on this chart as well. So you can really see the trend of this. That's something

51:14

that I'm really excited about, which is, you know, I actually think products today really optimized for tasks that the model can do for like minutes at a time. Like even codecs and like the coding tools, I'd say like, you know, it's in the CLI, you're kind of like seeing it be interactive. It's really, you know, quite optimized well for like maybe at most 10 minute type tasks. I have seen people push codecs to the limit into like multi-hour long tasks. But again, I think that that's more of the exception. But if you follow this trend, like I think like in the next 12 to 18 months, we could

51:45

see models that could do multi-hour long tasks very, very coherently. At some point, it might reach like, you know, six hours a day long task where you kind of like dispatch it and have it do, you know, do things on its own for a while. The types of products you build around that will look very different. You want to give the model feedback. You obviously don't want it to completely run wild for a day. Maybe you do, but you probably don't. And then the universe of things you can have the model do really expand. So that's something that I'm really, really excited about seeing.

52:16

Another thing over the next 12 to 18 months where I think it'd be really cool is improvements in the multimodal models. So, and actually by multimodality, I'm mostly thinking about audio here where the models are pretty good at audio. I think they're going to get a lot better at audio over the next six to 12 months, especially the likes, you know, the native multimodal model, the speech-to-speech ones. I think there's also interesting work being done around new types of models and architectures on the multimodal audio side as well. But audio, especially in the enterprise and in a business setting, I think is a hugely underrated domain still. Like everyone talks about

52:54

coding, it's all text. But we're talking in audio. A lot of the world's business is done via audio. A lot of services and operations are done via talking in audio. And so I think that that area is going to look very exciting in the next 12 to 18 months. And I think there will be even more unlock for what we can do with audio models there as well. Amazing. So quick summary, expect agents and AI tools to run longer to that trajectory to continue to increase. And then audio and speech becoming a bigger deal, more first party and native and better and core to the experience. Yeah.

53:35

Extremely cool. Okay. I want to go back to one of your hot takes, another hot take that I've seen you discuss. You're very bullish on business process automation as an opportunity in the world of AI. Talk about that. Yeah, this goes back to the thing that I said previously, which is we live in a bubble in Silicon Valley. And a lot of the work that we do, that we're used to software engineering, product management, building products, is very differently shaped than the work that goes on that runs our entire economy. And I see the same thing now when I talk to customers. If you talk to any company that's not based in, it's not a tech company, there's a lot of business

54:16

processes. And so what I mean by this is, I generally delineate it as, there's like, software engineering is kind of like open-ended knowledge work. And this is why I think tools like Codex tend to be quite good because it's exploring and you're giving it these open-ended things. But software engineering is fundamentally like pretty open-ended and it's not very repeatable. Right? So like you build a feature, you're not trying to build the exact same feature over and over again. And a lot of like tech jobs are in the space. I think like data science is kind of in the space as well. Even some of the like strategic finance stuff. But as you move further and further

54:53

away from software engineering and like what, what is core in tech, a lot of jobs are just business processes. They're like repeatable things, uh, repeatable operations, um, that's, you know, some manager at a company has kind of like iterated on. Um, there's usually a standard operating procedure that people want to do. Uh, and you don't want to deviate from it that much, you know, there's like in software engineering, the ingenuity is, isn't, isn't deviating, but a lot of, a lot of the, the, the work being done in the world is actually just, um, running through these procedures

55:23

and operations. Like if I, you know, if I call, um, a support line, they're running through one of these. If I call my utility company, there's a bunch of processes and things that they can and cannot do, um, for me. Uh, and so I'm, I'm just extremely bullish on this general category of like, and, and, and I think it's underrated because it's so different from what we think about in Silicon Valley, people tend to not think about it, but how can we apply, um, AI, uh, and, and some of the tools and frameworks that we have towards this business process automation, towards automating, automating and making easier, um, repeatable business processes with high determinism,

56:00

um, that is fully integrated with business, uh, data and business decisions and, and, and different systems within an enterprise. Um, and how can we actually make that, that process better? Uh, because I actually think there's a lot of opportunity and a lot of work to be done, uh, in that area. And we just, we just don't talk about it because it's, it's a little bit less, uh, uh, in our wheelhouse. So your take here, just to make sure I fully understand it is you think there's a much, uh, bigger opportunity outside of engineering for AI to impact, uh, productivity of companies and also jobs of these folks that are doing these kind of repetitive, easily automated tasks.

56:35

Impact jobs and also just impact how work is done. Like so much of work is done in this way. Like you think about, you know, like what, uh, like basically we, I, I talk to customers all the time, big enterprises, like, like how, how will AI transfer my company? Like how will it run in, in, in, in a world, uh, with AI in like 20 years? Um, and, and, you know, software engineering is part of the story, but there's so much more on the business process side. And I actually think it might look even more different on the business process side and, and the work there is, is pretty substantial.

57:04

It's actually interesting. I don't know, like from an absolute percentage or absolute base, I don't know if it's bigger or smaller than software engineering. Like software is pretty huge and pretty expensive, uh, as well, but it is pretty massive and it's definitely bigger than, you know, uh, uh, uh, it's, it's bigger than you would think it is based off of how, how people talk about it or don't talk about it on X or Twitter. Okay. Uh, in going in a slightly different direction, uh, having built the platform, building the API, uh, people building on API, the biggest question on people's minds is always

57:33

just, uh, how do I not have open AI squash my idea and build their own thing and then, you know, destroy this, this market I created. What's the general policy? What's the general philosophy of how startups should think about where open AI is unlikely to go? My, my general answer here is, is, um, the market is so big and so massive. Like I actually think, you know, startups should just not overly think about where open AI or these labs are going. I've talked to a lot of startups, you know, that have, you know, not worked out startups that are doing really well. Every startup that I've seen that is kind of fizzled out is not because open AI

58:12

or, you know, big lab or Google or something has, has come to squash them. It's because they built something and it like really didn't resonate with, with the customers. Whereas the ones that take off, like even in very competitive spaces, like coding, like cursor is huge at this point. And it's because they built something that people really love. And so my general advice is like, don't, you know, don't overly stress about this, just build something that people like, and you will, you will have a space in this. I can't overstate how big of an opportunity there is right now. Like

58:39

the, the, the opportunity space of building with AI is so big. Like a good example of this is, is like the space is so big that the Overton window of what is acceptable and not acceptable for VCs to do has completely changed here. VCs are like investing in like competitive companies left and right. It's just like the space is so big because, because the opportunity is, is, is unlike anything that we've seen before. And while, you know, uh, that, that affects how VCs operate from a startup perspective, it's like the most empowering thing in the world, because the, like, even if you just build something

59:07

that, that some people really, really love, you will, you will end up with a massive, massively valuable business. Uh, and so I, that's why I tell people like, don't, don't overly think about it. The other thing, like, I also think is important to remember, uh, at least from an open AI perspective, one thing that, that, that we've always held very near and dear, which both Sam and Greg helped, you know, reinforce from the top as well is we actually view ourselves fundamentally as a like ecosystem platform company. The API was our first product. We think it's really important for us to

59:35

foster this ecosystem and continue to, you know, uh, support it and, and not squash it. And so if you kind of look at the decisions we make, it, this is all we've, we've through it. Every single model we've released in one of our products gets released in the API. Like even, you know, we really release these codex models now that are a little bit more optimized for the codex harness, but they always find their way into the API and like all of our, you know, uh, customers end up, end up using those. We don't hold back on any of that. Uh, we think it's really important to keep our platform neutral.

1:00:03

Uh, and so, you know, we don't block competitors. Um, we allow people to have access to our models. Um, uh, we also want, you know, like, uh, we've recently been testing more of like the sign in with chat GPT, you know, uh, product as well. And so we, we, we want to foster this ecosystem. I think it's really important that we do so. Uh, the general, like thinking about this is like, you know, a rising tide, like lifts all boats and, you know, we might be a aircraft carrier. We're like pretty big at this point, but we think it's important to raise the tide, uh, cause everyone kind of, uh, benefits. And I think we'll benefit as well. Like our API itself,

1:00:33

it's grown pretty significantly because we, we act in this way. And so I'd really encourage people not to view open AI as this kind of like, you know, thing that'll just, uh, uh, shove people out of the way, but instead focus on, on, on building something valuable. Uh, and we, you know, remain committed to, to, to providing an open ecosystem. Why, why is that important to open AI? Just this focus on building a platform, creating a way for people to build businesses, just like, is that just, that's been the vision from the beginning. We want this to be a platform. It's been the vision from the beginning. It comes, goes back to our charter, actually,

1:01:07

like our, our mission. Um, so the, the open AI's mission has always been to one to build AGI. So, you know, where I was doing that, but then the second thing is to like spread the benefits of it to all of humanity. And there's kind of like a lot of, you know, uh, uh, the main part there is all humanity, like, uh, and obviously ChadGPD is trying to do this, you know, we're trying to reach however many, you know, the whole world, but very early on, and this is why we, we launched the API, you know, back in, I think it was like 2020 or something like really early. We don't think we, as a company,

1:01:34

we'll be able to reach all of humanity, right? Like there's, I dunno, every, every corner of the world's like, like pretty, pretty, pretty deep. And so we actually feel like in order for us to fulfill our mission, we need to have some platform style thing here where we can empower other people to build, you know, the customer support bought for podcasters and newsletter hosts, uh, because we're not gonna be able to do it ourselves. Uh, and so we've largely seen this play out with the API. Uh, this is why we, we, we, you know, we, we, we, we talk to so many of our customers and, and,

1:02:04

and really, you know, love seeing the diversity of, of things built on, but yeah, it's been there since A1, because it's, it's kind of, we view it as an expression of our mission. And you haven't even mentioned the, uh, the app store that you guys are launching, the ChatGPT app store. Yeah. Is, is that under your umbrella by the way, or is that a different org and team? It's a, it's a different team. So it's under ChatGPT. We obviously collaborate very closely with them and, uh, you know, they built like an apps SDK, uh, which is a built in close collaboration with our team. Uh, but that is more within the ChatGPT umbrella. Uh, but that is also another,

1:02:32

like, that's another example of this, right? It's like ChatGPT is like, we, we, we, we kind of like have these 800 million weekly active users who are just coming over and over again. Like it's a great asset to have as a business, but like, man, would it be better if we could somehow allow, you know, uh, other companies to come in and, and, and, and, uh, take advantage of this as well and, and build for this, this audience as well. And, and then ultimately we think it'll help us expand that, that, that group as well. Right. And so it's all, it all kind of comes back to the mission. And, uh, we find that being a platform, being open tends to help here.

1:03:05

Just that number 800 million, I think it's M M A is just like weekly, weekly, weekly, weekly act. Yeah. It's crazy. Billion people using weekly. I just like, it's sort of how many, how these numbers were just used to now, but that's insane, unprecedented. Yeah. It's, it's mind boggling for me to think about from a scale perspective. Uh, honestly, I, and the way I think about it is like 10% of the world, uh, and growing by the way, like it's just, it's, it's shooting up, um, uh, come to chat GPT, uh, um, and, and use it every day or sorry, every week. And this point, I just want to double down on this point you're making open AI's mission was to make AI

1:03:43

available to all of humanity. And I think some people diss that they're like, oh, you know, it costs money. And it's like, uh, like the fact that it it's, there's a free version of chat GPT that anybody can use that is not so different from the most powerful AI model that exists in the world for free. That's not gated that anyone can use. Like if you have, if you're a billionaire, there's only so much more you can get out of AI than what someone, you know, in a village in Africa can, can get. And I know that's always been really important to open AI. Yeah. Yeah. I mean, look, uh, that that's why I think we've leaned into the health work. We've

1:04:16

leaned into like, like, uh, like, uh, education is going to be very interesting here. Um, the other insane kind of trend here is, is the free model has gotten so smart over time. Like the free model back in 2022 was, you know, like, uh, well, it's good at the time, but it's like nothing compared to what you get today. Cause you get GPT-5 today. Uh, and so the, like, you know, raising the floor across the world is kind of, you know, something that we're really trying to do. And then we view it as, as part of our mission. The other flip side of this, by the way, is like, you know, kind of talking

1:04:44

about like the billionaires or whatever. I know people love saying like, you're using the same iPhone that like, you know, Steve, or sorry, like Mark Zuckerberg's probably using or like the billionaires are using. But for like $20 a month, you're basically using, you know, like using the same AI that, you know, the billionaires are using, uh, for like $200 a month, uh, you get the same pro model that, you know, all the billionaires are using, but they're probably not using pro for everything. They're probably just using the plus tier ones, uh, for their day in and day out. And so, yeah, this kind of like democratization

1:05:12

and just like spreading of this, this benefit, like across all of the world is saying that's really meaningful to us and something that, um, uh, drives a lot of, of, of what we do. One last question, just for folks that are thinking about building on the API or just like, oh, wait, I could do cool stuff with OpenAI's models and APIs. What, what does your API and platform allow people to do? Like, I know you can build agents on top of the platform. Just talk about what you allow. So fundamentally, the API offers a bunch of developer endpoints, uh, and, and, uh, and these developer endpoints basically let you sample from our models. The most popular one that we have right

1:05:47

now is one called responses API. Uh, and so this is an endpoint and it's optimized for building long running agents. So agents that'll work for a while. So what you can basically use, you can, you can, at a very, you know, uh, uh, low level, you're basically just giving the model text. The model will work for a while. You can kind of, you know, pull it to see, see what it'll do. And then you'll get the model response back at, at some point, that's like the lowest level primitive that we have, uh, for people. And that's actually what a lot of people use. That's the most popular way of building

1:06:17

on top of our API with that. It is like super unopinionated and you can do basically whatever you want. It's like the lowest level thing. We've also started building more and more kind of like layers of abstraction on top to help people build, uh, some of these. Uh, and so next layer up, we have this thing called the agents SDK, which has also gotten extremely, extremely popular. Um, this allows you to use, you know, the response API or some other API endpoints that we have to build, uh, what you might more traditionally think of as an agent, like, uh, you know, an AI kind of working in an infinite loop.

1:06:46

It might have sub agents that it delegates to it starts building all this framework, all the scaffolding, actually, you know, we'll see where this all goes. Um, but it makes it a lot easier for you to build these, these, these, these kind of agents, giving it guardrails, allowing it to like farm out sub tasks to other agents and, and kind of like orchestrate a swarm of agents. Uh, the agents SDK, uh, kind of allows you to do that. And then above that, uh, we've now started building tools to help, uh, also with kind of like the meta level of deploying an agent. Uh, so we have this

1:07:16

product called, uh, um, agent kits, uh, uh, uh, and widgets, uh, which are basically a bunch of UI components that you can use to very easily, um, build a very beautiful UI, um, on top of, uh, uh, either our API or agents SDK. Um, because, you know, a lot of times these agents kind of look very similar from a UI perspective. Uh, and so there's Asian kit. We also have a smattering of like, uh, evals products, like evals API, where if you want to test and like, you know, see if your models or your, your agent or your workflows working, uh, you can test it in a very quantitative way, um, using our

1:07:49

evals product. And so, yeah, I, I view it as like these, these various layers. They're all kind of helping you build, um, what you want, um, with our, uh, AI, uh, with our models, um, and with increasing levels of abstraction and, and, and, and, uh, you know, how opinionated it is. And so, um, you can start, you can do that. You can use the whole stack and, and it very quickly allows you to build an agent, um, or you can go down, down the stack as low as you want to basically responses API and build, whatever you want, uh, because of how low upload is. Sherwin, is there anything else that you want to share? Anything else you want to leave listeners with?

1:08:21

Anything we haven't touched on that you think might be helpful before we get to our very exciting lightning round? Sherwin Lowe The only thing I'd leave folks with is, yeah, I think, um, I think the next like two to three years are going to be some of the most fun, uh, in tech and in the startup world, uh, that, that we'll have in a very long time. And, uh, I would just encourage people to not, uh, not take it for granted. Like I, I entered the workforce in 2014. It was great for like a couple of years. I felt like there was like a period of like five to six years where it wasn't very exciting in tech. Uh, and then in

1:08:53

the last three years has just been the most insanely exciting, energizing period, uh, of my career. Uh, and I think the next two to three years is gonna be a continuation of that. And so, uh, I would encourage people not to take it for granted. I'm trying to not take it for granted. At some point, you know, this wave is going to play out and it's going to be a lot more, you know, incremental. Uh, but in the meantime, we're going to get to explore a lot of really cool things, invent a lot of new things and change the world and change how we work. And so, uh, that's the main thing I'd leave folks with.

1:09:18

I love this message. I want to spend a little more time on it. Um, when you say don't miss it, is it, what do you recommend people do? Is it just build, lean in, learn, join a company building really interesting things? Like what's, what's your advice to folks that are like, okay, I don't want to miss the boat. Yeah. I would just say engage with it. So it's basically like what you said, um, lean in, um, building, uh, tools on top of this is, is part of the, you know, it's part of the story. Um, just using the tools. Like you don't, you know, you don't need to be a software engineer to, to lean into this.

1:09:45

Um, all, I think a lot of jobs are gonna, gonna, gonna change here. So just using the tools, understanding the limitations of what it can and cannot do so that you can kind of watch the trend of what it can start to do, um, as the models improve. And yeah. And so it's basically like getting used and getting, getting used to this technology and getting familiar with it instead of kind of like laying back and, uh, uh, uh, letting it, letting it pass you. On the flip side of that, there's a lot of, I think stress and just anxiety around, like there's so much happening. How do I keep up? I got to learn out. Clotbot this week. Oh God.

1:10:16

Yeah. What, is there something you learned about it? Just not like you're at the center of this. How do you not get overly stressed and worried about missing things that are going on and just keep stay on top of news with, and what are some things you've done learned? Yeah. So I, I think I'm personally a bad example of this because I am, I'm basically chronically online, uh, on X and, uh, our company Slack. So I, I, I actually try and absorb, I end up absorbing a lot of it. What I will say though, it was just like from observing other folks who are less, you know, addicted to this stuff like I am. Um, yeah, a lot of it is noise. Like you don't need to,

1:10:48

you don't need to have like 110% of this kind of pass your mind, like, like go into your mind. Honestly, just leaning into like one or two different tools, starting small is already like, you know, more than you need here. I think just the combination of like the frenetic pace of the industry X as a product just creates like this insane kind of like, um, uh, uh, uh, yeah, this insane like pace of, of news, which is honestly very overwhelming. Uh, the main thing is like, you don't need to be, you don't need to know all of that to, to really engage with what's happening right now. And even something as simple as just like install the codex client, play around with it,

1:11:27

install chat, you can connect it to a couple of your, uh, you know, internal, uh, data sources, notion, Slack, GitHub and see what it can and cannot do. Um, all of that I think is, uh, a part of it. Amazing. Sherwin with that, we reached our very exciting lightning round. I've got five questions for you. Are you ready? Yeah. Yeah, absolutely. First question. What are two or three books that you find yourself recommending most to other people? Oh, I'll talk about one nonfiction one-on-one fiction book. Uh, the fiction book was, I just finished reading it. I, I, it was really, I really recommend it. It's, uh, uh, there is no anti-memetics division by QNTM. Uh, it's a, uh,

1:12:02

I think it's like an online author, but I saw it being shared on X. Uh, this, this, uh, it's like a science fiction-y kind of book. Um, and it was, I basically devoured it in like two days. Um, it was, it's super, super well-written, super fascinating. It's about a government agency that's fighting, you know, things that make you forget it. Um, and so it's just a very like smart, like creative book that, that, and fresh, uh, honestly, in terms of like source material, uh, that, that, that I really like. So I'd, I'd recommend that one. Uh, the book is also unintentionally hilarious. So like,

1:12:33

it's like meant to be like this, like sci-fi, almost like horror style book, but it was, it was, it was, uh, it made me laugh a couple of times. So, uh, that's the, that's the, um, fiction book. Nonfiction. So I'm going to cheat and I'm going to recommend two of them. So in the last year, I've been reading a lot more about China and kind of like the U S China relations. And I think there are two books that came out in the last year that have been, you know, really, really eyeopening for me in, in, in that regard. The first one is the Dan Wang book, Breakneck. That one was really, really good. I really liked his

1:13:00

analogy of like the lawyerly U S is the lawyerly society. China is the engineering society. Uh, and there are pros and cons to each. I read it and I was like, Hmm, yeah, it does, does seem like we were run by lawyers, uh, in the U S. Uh, so then that's one. Uh, and the other one is the Patrick McGee book on Apple in China. It was super, super interesting. I'm a huge Apple fanboy. Like if you could see my desk right now, it's, it's all Apple stuff, but just like one, it was just super fascinating learning about Apple's relationship to China. And then two, it just like had a lot of

1:13:30

inside information about Apple as a company that I found fascinating. So it was also quite a page turner and, um, also, you know, very, very timely, a timely book as well. The anti-memetics book sounds amazing. I'm buying it right now as you're talking. Yeah. Yeah. Yeah. It's, it's like, I think it's only like a couple hundred pages. I literally finished it in two days. It was just like, so, so good. Okay. Great tip. Okay. Uh, favorite recent movie or TV show you have really enjoyed? Yeah, that one's tough. Cause you know, with, I have two kids and, uh, uh, a busy job. And so

1:13:59

I really haven't had much time, um, to watch TV shows. Uh, I will say in the last couple of weeks, I watched a couple episodes on, I'm actually a big anime guy. And so, uh, I, I watched a couple episodes. There's a new season of this anime called Jujutsu Kaisen, uh, that's out. Uh, so season three of JJK, uh, was, was, was really good. Um, in general, uh, I'm a huge, uh, fan of, uh, Japanese anime. I think they create the most, uh, novel and unique, uh, plots, uh, and universes that, uh, Western media has shied away from. Um, and so, uh, generally a big fan of that, but yeah, I haven't really watched

1:14:37

much, but saw a couple episodes of JJK recently. Extremely understandable in your role. Yeah. Favorite product you recently discovered that you really love. Yeah. Okay. So, so, uh, so I recently, uh, had to set up a wifi and like home networking and I went all in on ubiquity, uh, routers, um, and can't security cameras. I'd never heard of it before I had, had to do this. I always just had a very simple setup. Uh, and it's just such a well-built product. Uh, I don't know if you used it before, but it's basically like the apple of like home networking. So, uh, beautiful products. Uh, but the thing that actually makes it extremely good

1:15:11

is that software is good. Uh, and so they have a really great, um, mobile app to help manage, you know, uh, all of the, the home networking. Um, and so basically a ubiquity, you can use it to buy wireless routers. Um, you need ethernet, uh, wiring throughout your house to use it. Um, but I actually think what makes it really good are security cameras. So if you have security cameras that are plugged into the ubiquity ecosystem, they have an incredible mobile app, uh, and apple TV app and iPad app, um, to kind of see the live feed of your cameras. And, uh, and so, uh, they're, they're, they're a little pricey, but not that pricey. Uh, but it's been

1:15:44

just an incredible product experience. All right. I went Eero, so I made a mistake. Good tip. Eeros are pretty good too, but, uh, I'm fully converted to ubiquity at this point. Okay. Good tip. Okay. Two more questions. Do you have a favorite life motto that you find yourself coming back to in work or in life? Yeah. Uh, the one that I always, you know, repeat to myself is, uh, uh, never feel sorry for yourself. There's a lot of things that are going to happen, you know, uh, at work, uh, in life, uh, and reminding yourself to never feel sorry and that you always have a sense of agency to kind of pull yourself up is something that I've had to tell

1:16:16

myself a lot. And, um, also something that I repeat to, to, to, to a lot of other folks as well. Last question. So in your previous life, you worked at Open Door, where you led work on basically figuring out how much to, uh, pay for houses. You basically built the model that told the company, here's how much we'll pay for this house. What's like a variable in the price of a house that you didn't expect is really important and impacts the price of a house. There's a bunch that were surprising. I'll, I'll maybe list the, the, the, the, the couple of most, uh, uh, interesting ones. Um, power lines and like, uh, high voltage power lines, like are super,

1:16:52

super, uh, actually impact your price quite a lot. I didn't really fully internalize this until I went to like Dallas and observed, like when your house sits next to one of these giant, like, you know, voltage lines is like buzzing and most people have families. You don't want your kids kind of near there. Uh, so I think that was one that really, really, uh, kind of surprised me. That makes sense. Yeah. And then the other one, which, which was something that, uh, was always something really difficult for us to, uh, quantify, uh, was floor plans. Uh, and so it is very important. Like, yes,

1:17:22

of course it's really important, but just like quantifying what a good floor plan is like and what a really bad floor plan is like, like we were doing all these things with like, how wide is the kitchen and like, is it a, what style of kitchen is it? And then like, where's the master bedroom? And, and so it was just really, really hard to quantify, but I remember floor plan was a big one because like, we'd have a home that like wouldn't sell. And then our, uh, ops team would go in and be like, yeah, that's the floor plan issue. So like, how do you, how, how could you tell us? Like you go inside,

1:17:46

you just feel it. It feels, you know, the floor plan feel, feel soft. Uh, so yeah, those, those are ones that were, uh, surprising. And then the last one that was more impactful than I thought is, um, general like curb appeal and like, even like the front door. Uh, and so I actually think there, there's a Zillow book on, on this where, um, the front door replacement tends to be the highest ROI, uh, for homes. Um, but just like the feel of like, as you walk up to the home as a buyer, what you're interacting with and the first moments of the house, I think was, uh, I'd underrated its importance.

1:18:18

That is extremely interesting. Uh, and I love that you had to figure, figure out how to do all this, uh, in code and not. Yeah. And floor plans. I have a bunch of stories around like for floor plans. There's like, there's like, uh, it's not digitized. So there's like a handful of people who have like paper floor plans, uh, of like all these homes in like Phoenix and Dallas. Um, yeah, a lot, a lot of fun, fun stories from the open door days. Okay. Sherwin, uh, thank you so much for doing this. This was incredible. Uh, where can folks find you online and, uh, and how can listeners be useful to you?

1:18:44

Yeah. So I'm, uh, online on, on Twitter on X. I'm just at Sherwin Wu and, uh, yeah, I mostly just tweet about, uh, open AI and API and some of the products that we're launching. Uh, and then how folks can be interested, uh, can be useful to me. Uh, I love hearing about things that people are building. And so if you're working on a startup, if you're hacking on an idea, you know, would love to, uh, just reach out to me on X. Um, I would love to hear about, uh, what you're building and, and learn about how open AI can help support you. Amazing. Sherwin, thank you so much for being here.

1:19:13

Yeah. Thank you, Lenny. Bye everyone. Thank you so much for listening. If you found this valuable, you can subscribe to the show on Apple podcasts, Spotify, or your favorite podcast app. Also, please consider giving us a rating or leaving a review as that really helps other listeners find the podcast. You can find all past episodes or learn more about the show at lenny's podcast.com. See you in the next episode.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note