Open Reader

AI predictions: Job markets, Codex beats Claude, and the death of org charts | Dan Shipper

completed 1:34:06 May 24, 2026 Watch on YouTube

Current Status

completed

Video ID

4D3hDmGhFhA

RAG / Chat

Enabled
AI predictions: Job markets, Codex beats Claude, and the death of org charts | Dan Shipper
Description

Dan Shipper is the co-founder and CEO of Every, a media and software company that’s become a living laboratory for the future of work. Everyone at his company of about 30 people is an AI early adopter; from editors to ops people, they use AI to do much of their work, giving Every a unique lens into where the world is heading. A year ago on this show, Dan predicted that people were sleeping on Claude Code for nontechnical work, which proved to be remarkably prescient. Today he’s back with another set of calls: the SaaS apocalypse is dumb, CLIs are over, the forward deployed engineer is the most valuable new hire, and the only thing you need to do to stay employed is ride the models. *Dan’s predictions:* 1. The future of work will happen inside Codex or Claude Code. 2. Every company will have one “super-agent” inside their Slack that every employee talks to regularly. 3. SaaS is not dead—in fact, Dan is bullish on SaaS stocks. His contrarian take: “I would buy SaaS stocks right now.” 4. SaaS economics will shift: users will bring their own AI tokens into apps, which actually improves SaaS margins. 5. PMs will thrive in the AI era. 6. Full-stack designers will become superheroes. 7. The AI job apocalypse is not happening. 8. Forward deployed engineer is the new most essential role. 9. CLIs are over. 10. Automation is a lie. 11. We will read way more AI-generated writing and we will like it. 12. We’ll be building software for humans and agents to use together. *Brought to you by:* WorkOS—Make your app enterprise-ready, with SSO, SCIM, RBAC, and more: https://workos.com/lenny Vanta—Automate compliance, manage risk, and accelerate trust with AI: https://vanta.com/lenny *Episode transcript:* https://www.lennysnewsletter.com/p/the-ai-paradox-dan-shipper *Archive of all Lenny's Podcast transcripts:* https://www.dropbox.com/scl/fo/yxi4s2w998p1gvtpu4193/AMdNPR8AOw0lMklwtnC0TrQ?rlkey=j06x0nipoti519e0xgm23zsn9&st=ahz0fj11&dl=0 *Where to find Dan Shipper:* • X: https://x.com/

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Dan Shipper predicts that AI will reorganize work around company-level agents and agent-native desktop workspaces, expanding human leverage and changing roles more than it eliminates jobs.
  • Why it matters: The discussion offers directly reusable operating patterns for agent systems: centralized agent ownership, human-agent shared work surfaces, agent-ready SaaS design, approval and rollback controls, and a forward-deployed engineering function.
  • Best use: Use this as a strategic operating-model briefing for designing AI-native workflows, product interfaces, and team roles; treat the employment and market calls as informed but unvalidated predictions.

Executive Summary

Shipper’s central prediction is that knowledge work will settle into two complementary agent modes. First, companies will deploy a shared, general-purpose “super agent,” likely accessible through Slack, which can answer common questions and execute delegated asynchronous work. Second, individuals will conduct much of their active work inside agent-native computer environments such as Codex, Claude Cowork, or comparable successors, where the agent can observe the user’s work, use the browser and local tools, research, and take action alongside them.

His most operationally useful argument is that agents are not autonomous organizational replacements yet: they need dedicated humans who own their context, reliability, permissions, behavior, and ongoing improvement. Rather than giving every employee a fragile personal agent, he expects companies to begin with a centrally maintained agent, then add specialized team agents as the underlying systems mature. This creates a durable “forward-deployed engineer” role: people who build the systems that let less technical colleagues safely use powerful agents.

For software companies, Shipper rejects the SaaS-apocalypse view. He expects agents to increase demand and usage volume for SaaS rather than replace it, while shifting the product requirement from “embed an AI assistant” to “make the core application usable by a human and an external agent at the same time.” That requires a new control-plane UX: visible agent activity, approval queues, logs, rapid rollback, synchronized state across CLI and GUI, and infrastructure designed for agent-scale request volumes.

On labor, Shipper argues that models commoditize yesterday’s competence but do not eliminate the human work of identifying the right problem, reframing a bad plan, judging quality, integrating output into a coherent whole, and creating differentiated work. He is especially bullish on product managers and full-stack designers who can pair product or aesthetic judgment with AI-enabled implementation. His practical advice is to “ride the models”: repeatedly test newly capable models against real work, including problems they could not solve in prior releases.

Key Takeaways

  • Claim: The near-term enterprise agent architecture will start with one shared company-level agent, not a personal agent for every employee. | Evidence: Shipper says Every initially expected a personal-agent “parallel org chart” after OpenClaw emerged, but reversed his view because individual agents are too fiddly to operate. He cites Shopify and Ramp as companies using a shared agent model and says agent deployment should begin broadly, then specialize into team agents. | Implication: Build a centrally owned agent service first, with an accountable operator or small team, before distributing specialized agents across functions. | Caveat: He expects personal agents to become more viable later as models and agent harnesses require less maintenance; this is a near-term operating recommendation rather than a permanent architecture.
  • Claim: Every useful production agent currently needs an invested human owner; automation shifts work into maintenance, steering, and system design rather than removing it. | Evidence: Shipper repeatedly argues that once users stop maintaining an OpenClaw-style agent, its utility deteriorates. At Every, an internal consulting agent called Claudie is actively managed by an AI engineer who spends substantial time diagnosing behavior and improving it through Slack and coding tools. | Implication: Treat agent ownership as a real operating function with staffing, evaluation, reliability targets, escalation paths, and continuous context management—not as a one-time software installation. | Caveat: The speaker believes this human-maintenance burden will decline as agents improve, but offers no fixed timeline for when independent agents become trustworthy.
  • Claim: The dominant individual work surface may invert from AI-inside-SaaS to SaaS-inside-an-agent workspace. | Evidence: Shipper uses Codex as his daily work environment: he maintains a thread per project, opens an in-app browser to work in Proof, and has the agent observe his document activity, research, and act through his computer. He reports staying at inbox zero for 10 days by having Codex and Every’s Quora email agent gather emails, generate a page, and execute requests based on his spoken instructions. | Implication: Design workstations and internal tools around persistent project context plus agent access to the same applications and artifacts the human sees, rather than isolated chat prompts. | Caveat: This workflow depends on broad computer and browser access, which materially raises permission, data-handling, and action-approval requirements that the conversation acknowledges but does not resolve in detail.
  • Claim: SaaS is more likely to be amplified than displaced by agents, but products must become agent-native collaborative surfaces. | Evidence: Shipper says Every’s SaaS spend rose year over year despite pervasive AI use and predicts agents will create more SaaS users and much higher usage volume. In his Proof example, users bring their own agent/model tokens, reducing the vendor’s direct inference cost while leaving Proof as the collaborative application layer. | Implication: Prioritize agent-readable and agent-operable products, but avoid assuming that an expensive vendor-hosted chatbot must become the primary interface or cost center. | Caveat: Higher agent activity creates significant infrastructure and pricing pressure; he points to GitHub’s scaling challenges as agents make large numbers of requests quickly.
  • Claim: Agent-native product UX needs governance primitives that traditional single-user SaaS often lacks. | Evidence: Shipper argues that agents can make many changes in seconds, so products need visibility into what the agent is doing, approval workflows or an inbox of proposed/completed actions, activity logs, and fast rollback. He distinguishes the emerging model from a standalone CLI: the human uses the web UI while the agent may use CLI or APIs, and both must remain synchronized. | Implication: For any AI-operable application, make change review, provenance, reversible execution, state synchronization, and rate-limit resilience first-class product requirements.
  • Claim: AI increases the throughput of creation so sharply that the scarce work becomes judgment, integration, deletion, and quality control. | Evidence: Shipper notes that nontechnical employees at Every now submit pull requests, while OpenClaw maintainer Peter Steinberger reportedly receives thousands of pull requests daily, uses 50,000 Codex instances to sort them, and merges about 1,000. Shipper’s framing is that the key question is no longer merely whether something can be built, but whether it coheres with the overall system and what should be removed. | Implication: Shift team design and metrics from output volume toward intake triage, architecture consistency, quality gates, deletion discipline, and mechanisms that prevent low-quality work from reaching core systems. | Caveat: High-volume AI output can become “slop” unless teams create filtering systems; simply adding generation capacity can make specialist roles worse if review infrastructure is absent.
  • Claim: PMs and full-stack designers are unusually well positioned because AI lowers implementation barriers while increasing the value of problem selection and differentiated taste. | Evidence: Shipper describes Every’s Spiral lead Marcus, a lightly technical former Axios PM, using Cursor and then Claude Code to ship rapidly by combining technical fluency with product and user judgment. He similarly sees designers directly shipping pull requests and creating interactions that generic vibe-coded output cannot distinguish from commodity “slop.” | Implication: Develop product judgment, user research synthesis, technical literacy, and visual/interaction taste alongside hands-on agent usage; these are likely to compound more than narrow execution-only specialization. | Caveat: The speaker’s job-market conclusion is anecdotal and based heavily on Every’s unusually AI-forward environment; he notes that designer hiring has not yet broadly increased.

Detailed Brief

Why benchmarks overstate replacement risk

  • Claims: Benchmark gains measure performance on tasks that humans have already framed, articulated, and scored; they do not capture the upstream human work of noticing that the stated task is wrong.; The relevant comparison is not unaided human versus AI, but one human using AI effectively versus another human using AI less effectively.; A competent senior human often reframes the assignment rather than complying with it literally, which remains a meaningful distinction even when models perform long autonomous tasks.
  • Evidence: Shipper built a private “senior engineer benchmark” from his production problems after Proof, an app he vibe-coded, suffered recurring outages after launch.; He says models before GPT-5.5 scored roughly 30/100 on rewriting the product from first principles; GPT-5.5 scored 62/100 when using an Opus 4.7-generated plan, while two senior engineers using AI scored in the high 80s to low 90s.; His example of the gap: when instructed to fix a list of defects, models tend to patch the listed issues; a senior engineer may inspect the system, conclude the codebase needs a risky rewrite, and force that higher-level decision.
  • Caveats: This is an internal, nonstandard benchmark with a custom prompt and scoring framework, not independently validated evidence of general engineering capability.; Shipper nevertheless expects model progress to make the current benchmark obsolete, because he can redefine the test once models solve its present tasks.
  • Implications: Do not use autonomy-duration or benchmark headlines alone for workforce planning.; Evaluate models in real production workflows where task formulation, diagnosis, architectural judgment, and recovery from failure matter.

Writing, planning, and multi-agent collaboration

  • Claims: Teams will increasingly consume AI-generated internal documents, plans, and emails—and should judge them by whether the accountable human understands and endorses them, not by whether AI drafted them.; Two agents can outperform one because the user’s primary agent can carry rich personal and project context into a conversation with an application or company agent.; Long-form guidance can be made more operational by writing it for both humans and agents: humans receive the core narrative while agents ingest the detailed reference material for future application.
  • Evidence: Every used Notion agents for quarterly planning: a company strategy informed agent-led interviews with teams about prior results, goals, metrics, and alignment, producing planning reports that Shipper then reviewed for quality and cross-team dependencies.; Shipper says most of his email is drafted by GPT-5.5/Codex and recounts an investor email that Codex sent without review but that he judged equivalent to what he would have written.; Every labels AI-assisted published writing; Shipper says its guides are intended for agents as well as human readers because an agent can ingest extensive material and surface it later during a relevant decision.
  • Caveats: Shipper draws a hard line against sending AI-generated material that the sender cannot explain or defend line by line.; He maintains that human writing remains important, particularly for public-facing work, even though he rejects a blanket aversion to AI-assisted drafting.
  • Implications: Establish authorship and review norms: AI drafting is acceptable when a named human owns the reasoning and final message.; Create durable, agent-readable institutional knowledge so agents can retrieve and apply detailed guidance at the point of work.

Adoption method: ride the models through repeated real-world experimentation

  • Claims: The defensible individual strategy is not passive awareness of AI but continual experimentation with each newly released model against relevant work.; The frontier is not confined to model labs or San Francisco; it exists wherever a practitioner discovers how a new capability changes a concrete domain workflow.; Play and personal usefulness are better adoption drivers than fear of displacement.
  • Evidence: Shipper’s practice is to “turn the rock over again” whenever a model improves—retrying tasks that previously failed, such as his senior-engineer benchmark.; He recommends trying active workflows in Codex or Cowork and experimenting with agent products including OpenClaw, Hermes, Victor, and Every’s 1Plus1.; His suggested test is to choose an irritating problem in work or life and attempt to solve or build it with AI; the target is a personal “moment of joy” where the capability becomes self-evident.
  • Caveats: Enterprise employees may be blocked from using frontier models by security or company policy, which can slow individual learning and produce organizational capability gaps.
  • Implications: Maintain a recurring model-evaluation habit tied to your actual workflows rather than generic prompt experimentation.; Leadership should provide governed access to current tools or risk employees learning unevenly and outside controlled company systems.

Notable Concepts & Terms

  • Reach test: Every’s adoption heuristic: a tool matters when people organically reach for it when they begin work, rather than using it only because it is mandated or novel.
  • Super agent: A centrally maintained, company-wide agent that handles broadly useful delegated work before an organization proliferates personal or departmental agents.
  • Forward-deployed engineer: The emerging role responsible for making agents useful and safe in a real operating environment—maintaining behavior, context, tools, permissions, and workflows for other employees.
  • Computer errands: Every COO Brandon Gell’s term for personal-agent tasks such as handling groceries or other routine computer-mediated chores, distinct from workplace agents.
  • Human-agent shared work surface: A product paradigm in which a person uses a GUI while an agent interacts through browser, CLI, or APIs on the same underlying work, requiring synchronized state and shared visibility.
  • Allocation economy: Shipper’s earlier concept that humans increasingly work like managers: allocating work to AI, checking performance, improving systems, and making higher-level choices.
  • Bring-your-own-agent / bring-your-own-tokens: A SaaS model where the customer’s own agent/model performs AI work within a product, potentially preserving vendor margins versus the vendor funding inference itself.
  • Ride the models: Shipper’s recommended skill-building loop: test each new model release against real tasks, especially tasks that failed on earlier models, to discover newly viable workflows.

Operator Notes / Why Ken Should Care

  • Assign a named owner and service-level expectations to any company-wide agent; do not launch self-service personal agents without a maintenance and escalation model.
  • Prototype a shared-work-surface pattern in one core workflow: human UI, agent execution path, change preview, explicit approvals, immutable logs, and one-click rollback.
  • Audit product/API/HTML/CLI surfaces for agent usability and synchronization; prioritize machine-readable state, clear action semantics, rate limits, and immediate reflection of agent actions in the human UI.
  • Create a high-throughput AI-output intake process for code, analysis, content, and operational requests, with automated triage plus accountable human architectural review.
  • Establish an AI-authorship policy requiring a human owner to verify and defend AI-assisted internal plans, customer communications, and public materials.
  • Run a recurring frontier-model evaluation against unresolved internal tasks rather than relying on published benchmarks; document what newly works, what fails, and what governance it requires.
  • Develop or hire forward-deployed engineering capability before scaling agents across teams, particularly for data access, permissions, observability, and workflow reliability.

Source/Metadata

  • Title: AI predictions: Job markets, Codex beats Claude, and the death of org charts | Dan Shipper
  • Transcript words: 17434
  • Duration seconds: 5646
  • Timestamp note: No usable timestamps or chapters were present in the supplied transcript.

Transcript

17700 words en Processed in 1687.9s

The last time you were on this podcast, you had this hot take that people were sleeping on Claude Code. You were so unbelievably right. The premise of this episode is we're going to go through what else you predict will happen. The AI jobpocalypse is not really a thing. I am super, super bullish on PMs and full-stack designers. You guys have doubled in people in the past year, which is not what people would have expected from a company that is so AI-forward. I'm simultaneously extremely AI-pilled and very bullish on humans. Automation is a lie. Every agent needs a human. We have so much automation, so much AI, and I also work way more. Creativity. It just feels like it's going to be more and more valuable to stand out from all the slop that people are shipping and launching constantly. What models do in general is they make yesterday's human competence cheap. And so it becomes commoditized. It's not valuable anymore. What humans do is we go in there and we're like, yeah, we have all this frozen human competence from yesterday. How do I use this to make something new and interesting? What are some predictions for how the way we work is going to change? It's going to bifurcate in two main ways. One is everyone's going to have at least one agent that they talk to that they can offload work to. Second is that most of the work that you do is actually going to happen on your computer in an environment like Codex or CloudCowork. What you're predicting here is the SaaS tools will run within Codex or CloudCode. I think the SaaS apocalypse is dumb. I would buy SaaS stocks right now. What agents do is increase the number of users of SaaS, not get rid of it. A lot of people are moving to a CLI and trying to work from the terminal. We speed-ran the CLI era. It was nice while it lasted, but I think CLIs are over. Today, my guest is Dan Shipper, CEO and founder of Every. Dan and his team are building maybe the most AI-forward startup out there. And as a result, are very much living in the future of how work is going to look as AI becomes a bigger and bigger part of our day-to-day. Everybody at their company, including every non-technical person, uses Codex and Cowork and CloudCode to get much of their work done. And this is why, way before anybody else, Dan saw the rise of CloudCode and what is now Cowork, which he predicted almost a year ago when he was on the podcast last time. So I asked Dan to come back on the podcast to share his current biggest predictions for how work is going to change over the coming year for most people. We chat about what work will look like at most companies at the end of this year, how the shape of the work we do will change, and who will do best in this coming future, slash what you need to be working on right now. Hint, hint: product managers and designers are going to do very well. Dan makes a lot of bold predictions and many quite contrarian takes that I was not expecting him to say. And we are going to revisit this conversation exactly a year from today to see how much he got right. Before we get into it, do not forget to check out Lenny's productpass.com for a free year of the hottest and most well-crafted AI products in the world, available exclusively to Lenny's newsletter subscribers. With that, I bring you Dan Shipper. Dan, thank you so much for being here. Welcome back to the podcast. Thanks for having me. Always a pleasure to be with you. The last time you were on this podcast, you had this hot take that people were sleeping on Cloud Code, and in particular, Cloud Code for non-engineering work, for fixing files, sorting your hard drive, all these things that people hadn't thought about. Nobody was talking about this. This was a year ago. You were so unbelievably right about this. It's unreal what has happened since then. They built Cowork, which was this whole build on this very specific idea using Cloud Code for non-technical work. Codex is getting into this now. I imagine you've been seeing this. They're leaning into this non-technical use of basically coding agents. I feel like this has also been a big part of Anthropic's success over the past year, just how non-technical people use this stuff. So you were just so ahead on this stuff. I even wrote a newsletter post building on this idea. I'm like, hey, this is interesting. I should dig into this. I asked people how to use Cloud Code for non-engineering work, and I just had so many examples. And it's my second most popular posting. So clearly you have a unique glimpse into where things are heading. So the premise of this episode is we're going to go through what else you predict will happen in the future, how things will change for people building products. And I think it would be helpful to start with giving people a brief glimpse into just how you operate and how your team operates. That gives you this unique lens into where things are going. So just give us a sense of how you work. Thank you. I really appreciate the introduction. And yeah, I think one of the things about predicting the future, or the way that we think about predicting the future at Every, is that what you don't want to do is prognosticate. What you want to do instead is just live in it together. So everybody at Every is an AI early adopter. We're almost 30 people now. I think when we did our interview, we were 15. So we've doubled in size in the last year. We're all early adopters. And we have engineers. We have designers. We have writers. We have editors. We have salespeople. We have customer service people. And everybody has a little bit of that. Whatever that thing is where you're just like, I like to explore. I like to experiment. I'm very curious. And I'm super all-in on AI. And what that does, I think, is it creates this little pocket of the future where we're all living in it together. And we get to be a little bit further ahead because at any other company, there's a mix of people. There's early adopters. There's the middle-of-the-pack people. And there are people who are very anti. And another thing that happens, which is really cool, is we get to, because of our role, reviewing models and being a little bit of a tastemaker in AI, get access to stuff before it comes out. So we get to beta-test and alpha-test and help steer the direction of where things are going a little bit, which is very, very cool. And so when I think about predicting the future, when you create an environment like that, it's actually just about noticing what's going on. And I think a core part of it, too, is writing about it. I think articulating what you're noticing, articulating the future, kind of brings it about in this way that makes it real for you and your team and then anybody else who's on the internet who's reading it. And so the Reviewing models and and being a little bit of a taste maker in AI, we get access to stuff before it comes out. So we get to beta test and alpha test and help help steer the direction of where things are going a little bit, which is very, very cool. And so when when I think about predicting the future, it's actually, when you create an environment like that, it's actually just about noticing what's going on. And I think what a core part of it too is writing about it, I think. Articulating what you're noticing, articulating the future, kind of brings it about in this way that makes it real for you and your team and then anybody else who's on the internet who's reading it. And so the cloud code thing, it was this, it's this very organic thing where, for us, we tried cloud code when it came out. That's our job. We tried, we tried all the new stuff from all the new, all the new model, all, we tried all the new stuff from the model companies. And at the time, it was a little bit early. But right around, I think, Sonnet 3.5 or Sonnet 3.7, we were testing that to do our vibe check on it. And we were like, holy shit, this is crazy. This is like really, you can, they got rid of the code editor. And so from that point on, we just, we run, at this point now, we run six products, software products internally. At that time, we ran maybe two or three. And from that point on, we just started shifting to a world where everybody was, no one was looking at the code. Everybody was talking to their computer in English using cloud code in the terminal. And so I was able to see, ooh, this is starting to happen. And then because my job is a little bit just push and play with stuff, I was like, I wonder if I could use this for my writing, how could I do that? And then it just starts to unfold. And you're like, okay, this is not ready yet. But it's obviously useful for me. My, one of the things that we talk about internally is what I call the reach test, which is, do you just, when you wake up in the morning, do you reach for it organically? I love this combination of you are using the latest stuff. And I think this is, as you said, maybe an underrated skill, you're, you're good at being self-aware of here's what's weird and new and different and interesting. So that's a really cool combination, partly because you have to write about it and you write about it. So I think that's the perfect recipe for someone having a sense of where things are going. This episode is brought to you by our season's presenting sponsor, WorkOS. What do OpenAI, Anthropic, Cursor, Vercel, Replit, Sierra, Clay, and hundreds of other winning companies all have in common? They are all powered by WorkOS. If you're building a product for the enterprise, you've felt the pain of integrating single sign-on, skim, RBAC, audit logs, and other features required by large companies. WorkOS turns those deal blockers into drop-in APIs with a modern developer platform built specifically for B2B SaaS. Literally every startup that I'm an investor in that starts to expand upmarket ends up working with WorkOS. And that's because they are the best. Whether you are a seed stage startup trying to land your first enterprise customer or unicorn expanding globally, WorkOS is the fastest path to becoming enterprise-ready and unblocking growth. It's essentially Stripe for enterprise features. [SPEAKER_01] Visit WorkOS.com to get started, or just hit up their Slack, where they have actual engineers waiting to [SPEAKER_01] answer your questions. WorkOS allows you to build faster with delightful APIs, comprehensive docs, [SPEAKER_01] and a smooth developer experience. Go to WorkOS.com to make your app enterprise-ready today. [SPEAKER_01] Okay. So the way that I'm going to structure this conversation, there's going to be three [SPEAKER_01] buckets of predictions. One is how the way we work is going to change in the coming years. Two is how, [SPEAKER_01] what the shape of the work we're going to be doing is going to look like and change. And then three is [SPEAKER_01] who is going to be most successful in this future slash what should you be doing and working on now to be successful in this future? Lenny, my only ask is we come on a year from now and then you score it. I want to score. Okay. So this is a year from now. Okay. Okay. So is this, that's actually, is this your predictions for in a year, this is what it's going to look like, or this is the emerging future? I think, I don't, I will probably say I don't have an exact timeline. I think most of the stuff that I'm going to talk about will be pretty apparent within a year, but it may, it may take longer than that. Okay. But I think it will, it should within at least a year be not obviously wrong. It should seem like it's moving in that direction to count. Okay. May of 2027, we will review your predictions. Amazing. Okay. I love this. Okay. So let's dive in. What are some predictions for how the way we work is going to change in the coming year? One of my favorite questions, because I think if you look at benchmarks, you're just looking at, okay, AI is going to just take all of our jobs, basically. I meter has this really cool benchmark where it measures how long it can, the newest models can do tasks autonomously. And it's like, oh, it's like, it can, what's it called? Oh, like mythos preview, the big anthropic model that everyone's so worried about. It can do tasks of 17 hours at 50% accuracy. It's like, holy shit, that's crazy. And I think it is real. It's true. And the progress, model progress, is going up exponentially. And my experience and my feeling is that we will look back in a year and say, we actually have a lot more work to do. Humans have a lot more work to do even as models get better at doing work. And there's a really interesting paradox there. And my prediction for how work will, my big prediction of how work will change, or how you will be doing work in a year, is it's going to bifurcate in two main ways, how you use agents. One is you're going to be doing, I think, what we figured you would be doing five years ago when we thought about how work with AI works, which is everyone's going to have, at least in their company, at least one agent that they talk to that can do work that they can offload work to. And we'll talk about what that looks like, but it's essentially like open claw. Second is that most of the work that you do is actually going to happen on your computer What will change, or how you will be doing work in a year, is it's going to bifurcate in two main ways: how you use agents. One is you're going to be doing, I think, what we figured you would be doing five years ago when we thought about how work with AI works, which is everyone's going to have, at least in their company, at least one agent that they talk to that can do work, that they can offload work to. And we'll talk about what that looks like, but it's essentially like OpenClaw. Second is that most of the work that you do is actually going to happen on your computer in an environment like Codex or Cloud Cowork that becomes the operating system for how you do all of your work, whether that's your email, the documents you create, all that stuff. It's going to be on that kind of a surface. That's becoming the clear competitive landscape. So I want to go in order of those two. So the first one is you're going to have agents you delegate to, probably in Slack, but anywhere. The first thing that's interesting about that one is it's not clear what the architecture is going to be for that. Is everyone going to have an agent? Is every team going to have an agent? Is it going to be just one agent? Do agents specialize? There's this parallel shadow org chart. And when OpenClaw first came out, everyone internally adopted it. I was very convinced that it would be everyone has their own agent. And there's some really interesting things about that world of a parallel org chart. Agents in that world become little reflections of you, which is really cool and really interesting. Did you ever read The Golden Compass? It's like having a little daemon on your shoulder, that's a little part of your soul. Yeah. I really think that's what it looked like was happening. So I was very into personal agents, and I have completely flipped. I really think that the model for now is going to be a super agent, one agent for the entire company. You're starting to see this in some companies. So, Shopify very famously has one, Ramp has one now, and I think there's some really interesting reasons for that. I actually still think that the personal agent thing is coming, but what we found is there's all this hype with OpenClaw. Everyone's like, I'm going to set it up. It's so cool, or whatever. Then everyone realizes it's way too much work. This thing breaks all the time. I have to fumble around with it. I have to be able to SSH into my server and blah, blah, blah. Most people, to do work at least, just don't want to spend that time or can't. And the fundamental underlying thing that drives that is, whether it's OpenClaw or any other harness, in order for an AI agent to be useful right now, it really needs a human who cares about it. It really needs a human personal connection with someone who's watching what it does and making sure that it's doing the right thing and that it's useful for people. The minute you sever that connection, so the minute someone's like, I don't want to maintain this dumb OpenClaw, is the minute the agent is not really that useful anymore. That's why I think it has started to shift to a more one-agent-per-company model, because for now, the ideal is [SPEAKER_00] you basically set up a forward-deployed engineer [SPEAKER_00] or someone with that sort of profile who's responsible for making sure that that agent [SPEAKER_00] is working for the whole company. Then maybe you have some little team agents. [SPEAKER_00] I think as the models get better at being more independent, that will shift down, and [SPEAKER_00] it'll be more likely that we'll have more personal agents because we don't have to [SPEAKER_00] fuck around with all the internals. But the model that I see working for us and for a lot of other [SPEAKER_00] companies, including the model companies themselves, is [SPEAKER_00] when it comes to the sort of async agents, it's really a [SPEAKER_00] you have one agent at the [SPEAKER_00] top that's doing, sometimes, everything. A lot of times it's a particular kind of job [SPEAKER_00] that you've decided that everyone in the company needs an agent for, like data requests. [SPEAKER_00] Then I think it starts at the top and then starts to trickle down, [SPEAKER_00] where you make it more specialized agents and teams and all that kind of stuff. The mechanism is: [SPEAKER_00] agents need people who care about them. That is so interesting. That point about [SPEAKER_00] you need to garden your agent because there's context you have to keep adding to it. It breaks, as you said, and once it's just too much work, you're like, okay, forget this thing. I'm going to go back to Codex or Cloud or something like that. Exactly. Okay, cool. So this is a cool opportunity. So the idea, what you're predicting here is companies will have this super agent that everyone can talk to. You said Shopify has got River, I think it's called. What's the Ramp one called? I can't remember. Okay. It's probably got a funny name. Okay, so that's the prediction. Okay. That's the first prediction. We will start with agents at the top that are more general and are used by more people in the company, and then it will start to grow down as people get more used to these [SPEAKER_01] use cases, they get more specialized, and agents become less fiddly. [SPEAKER_01] They just work better. Is this mostly going to be in Slack? Do you predict? For work, yeah. It seems to make sense. I think people love having the green bubbles on OpenClaw. Sorry, the blue bubbles on OpenClaw. If you can use it with your iPhone, I think there's this little thing in people's heads where they really like to keep their personal and work agents separate. I think there's a whole territory. Our COO, Brandon Gell, calls this computer errands. There's this whole territory of using personal agents for your computer errands. It's like, order my groceries, or whatever. There's so much of that that I think this is going to be huge for. But we focus mostly on [SPEAKER_01] the work stuff, and I think that's going to happen mostly in Slack. Sweet. Go Slack. [SPEAKER_01] I think there's this little thing in people's heads where they really like to keep their personal and [SPEAKER_01] work agents separate. And I think there's a whole, there's a whole territory. Our, our COO, Brandon Gell calls this computer errands. There's this whole territory of using personal agents for your computer errands. It's order my groceries or whatever. And there's so much of that that I think this is going to, it's going to be huge for, but I focus, we focus mostly on the work stuff. And I think that's going to happen mostly in Slack. Sweet. Go Slack. Should we, do you want to talk about the other work surface? Absolutely. Codex co-work. Okay. This is the, let's do it. I'm so excited about this. I think it's the coolest thing. So what happened was Anthropic realized at some point that with cloud code, if you put an agent on your [SPEAKER_01] computer and it runs on your computer, it has everything and has access to everything that [SPEAKER_01] you have access to. It uses the terminal, so it has basically superpowered access to it. And [SPEAKER_01] not only that, these agents really understand how to use the terminal because there's so much content online about that. And it created this super powerful coding paradigm, which is, Anthropic was really doing it first. OpenAI for a while was, I, I, I, in my [SPEAKER_01] opinion, very, very behind on this, and then in my opinion has surpassed them recently. It's really [SPEAKER_01] interesting. But they were very early on this when people were still thinking about coding agents [SPEAKER_01] or coding models as being pair programmers. They were among the first to be like, no, and do it successfully. There were people before them, like Devin, who I think had had a big, had the big cloud environment, and OpenAI tried this too, but the real adoption seems [SPEAKER_01] to have happened when you put it on your computer. So they figured that out. And then I think they figured out, along with their community, that once you have a coding agent on your computer [SPEAKER_00] that can build anything, it's actually really good for any kind of work you want to do. And people started [SPEAKER_00] just hacking cloud code, essentially, to do all of their work. So Anthropic then built Cowork, [SPEAKER_00] which is a little bit of a nicer wrapping around cloud code, but it's fundamentally the same thing. And then I think OpenAI made a couple of different bets, but their main bet on a programming agent was the, the, the, the earlier versions of Codex were very technical, and they [SPEAKER_01] were super smart, but they were a little bit autistic. It was a little hard to, [SPEAKER_01] they didn't quite get what you meant. They got exactly what you said. [SPEAKER_01] And I think maybe three or four months ago, around the time that they launched 5.3, [SPEAKER_01] they started to move in this direction of, oh no, we get it. This model is fast. It's [SPEAKER_01] really good for general-purpose knowledge work type tasks. And then they launched the Codex desktop app. And I think the Codex desktop app takes, if you look at all the lessons that Anthropic learned, they went from cloud code to cowork, and you can kind of see that in the tabs on the Anthropic desktop app UI. I think OpenAI was just like, we, we see where this is going. Let's just skip to that. And so I think Codex right now, this is a horse race. They're going to have different positions. But I think Codex right now is my daily driver. I spend all, all my time in it. I flip to Claude every once in a while, but I think they're getting the paradigm right. And it's clear to me that whoever is in the lead, because I, again, I think it'll change, whoever's in the lead, it feels very obvious to me that all of the work that you do is [SPEAKER_01] going to be in one of those surfaces where, for example, when I'm writing a document, Codex has a [SPEAKER_01] browser in the app. It has an in-app browser. And when I'm writing a document, I just go into one of my Codex threads, which I have one thread for every project. And I just open the in-app [SPEAKER_01] browser. I go to the document. I usually do it in Proof, which is this online markdown editor [SPEAKER_01] I built. And then I just have Codex running and watching me in Proof, and Codex can see what I'm [SPEAKER_01] doing. I can see what Codex is doing. It's all kind of in one place, which is an extension of [SPEAKER_01] the same thing that made cloud code work really well originally. And I basically feel like I have this [SPEAKER_01] parallel work buddy that not only can it respond and write in the document, but then it can go [SPEAKER_01] do research. It can go, it can use my computer to basically do anything that I can do on my computer. And that's incredibly powerful. And I do this with everything. I've been in, I've been at inbox zero for like 10 days straight now, which if you know me is crazy. I never like this. And that's because I literally just have Codex gather all my emails with Quora, which is our email agent. And then it renders a little page. And I think I showed you this at the In Thrive, the Anthropic event. It renders a little page, and I just monologue into it and just talk at each email. I'm like, okay, go, go research this. Oh, here's a question from our lawyers. Can you go collect all of the documents from the last four years and then put them into a report and send them? And it just does it. And so all the stuff that I would procrastinate on, I don't really procrastinate on anymore. And so I feel like there's this, for a long time, we thought, I thought too, that the optimal experience of AI was going to be take AI and put it in a browser. And I think the reverse is actually starting to happen and be really, really valuable in a way that I did not expect, which is take the AI agent that you use all the time on your computer and put a browser in it so it can see everything you're doing. And that is just a magical combination that I think is very uncommon now. You can't even do this in cloud code because they don't let you browse external websites inside of cloud code. So it's very uncommon now, but I think it will be super common in a year. This is more profound than it may even sound. What I'm hearing is instead of AI being baked into SaaS tools, what you're predicting here is you will, the SaaS tools will run within Codex or Cloud Code. So that is one really important second-order effect of this. Okay, so yeah, I'm using Proof or really any website, maybe PostHog or whatever. a magical combination that I think will be, is very uncommon now. You can't even do this in cloud code because they don't let you browse external websites inside of cloud code. So it's very uncommon now, but I think it will be super common in a year. This is more profound than it may even sound. What I'm hearing is, instead of AI being baked into SaaS tools, what you're predicting here is, the SaaS tools will run within Codex or Cloud Code. So that is one really important second-order effect of this. Okay, so yeah, I'm using proof or really any website, maybe post hog or whatever. And I'm doing it inside of my agent, and the agent has access to the website. So it has access to everything that I have access to. And it has access to my whole computer. When I run the agent on that website, I'm using my tokens. I'm not using the vendor's tokens. I'm not using the app's tokens. And so it puts SaaS back in this place where, yeah, you want to make it friendly for an agent, and everyone's got a CLI. Now you want to make the HTML really, really usable. You want to make sure that anything that happens in the CLI shows up for the user immediately, all that stuff. There are a lot of issues to deal with. But once you do that, you actually don't really need to think about having an AI surface that's primarily going to be the thing that users use, in the sense that you don't need to build an agent necessarily into your product. I think you can. And there's another really interesting bifurcation of this that we should talk about, which is that having two agents is better than one. But I think for now, there's this really cool thing where with proof, for example, anyone who uses it, I don't pay for tokens because they bring their AI to proof. And so it changes what you build as a SaaS company. And you build it now for both humans and agents to use at the same time. And it changes your margins back to, well, I don't really have to pay for tokens anymore, because the user is going to bring the AI. So I think this is a huge deal. So what you're describing is more and more work that we do, more and more professional work. Is it just going to happen within Codex or Cloud Code? Where does Cursor fit into this? Is that one of the, is there potential there? That's a good question. I think that Cursor sees a lot of the same stuff. And there, in some ways, they have some of the same stuff, but it's better. I think that Cursor's cloud implementation is better than either Open AI or Anthropix and is more advanced. And I think that Cursor has, at least so far, more distinctly chosen a lane. They're more distinctly choosing to be for programmers. And that may limit how far they get in here. I think the definition of programmer is expanding enough that they'll have a big market, but I don't know that they're going to jump into, okay, use this to make a slide deck or whatever. But it is really clear that every model company is starting to realize how important it is to have a harness to get the most out of the model. And so where all the platforms are moving is to a world where you're not just doing prompt and response. When you call the model on the Open AI platform, the Anthropic platform, they're literally running the model on a computer that is in the cloud that they run and then giving you the result out of it. And they know that, in order to get the best results of the model, they need to offer that. And so you see, Anthropix got cloud-managed agents. Open AI does not have a response yet, but I assume that that's going to happen. And now Cursor was just essentially acquired by SpaceX. It's not a full acquisition, but it's close. So I think people are starting to realize, I can't just do the model part of it, I have to have this harness above it. And I think the ultimate form of that harness is, I can do any kind of knowledge work. Cursor itself feels like one of the things that it's going to be a hard decision for them, whether to stay just for coders or not. So people building products that aren't Open AI or Anthropic, if this proves to be true, the prediction here is they're going to be using your product over time inside of one of these agents. Is there something you would do if you're one of those companies to prepare for that future? I would just prepare for that. So, for example, every more classic piece of productivity software, whether it's Slack or Word docs or PowerPoints or whatever, it's really mostly meant for a human to use. And now people are doing CLI. So it's meant for an agent to use independently of a human. And we're moving into this new paradigm, I think, where the human and the agent are on the same piece of work together. And they're both doing things, and you need to have, I need to have visibility into what the agent is doing. The agent has to have visibility into what I'm doing. We have to go back and forth in this seamless way. And the kind of software that you make for that is going to be very different. So, for example, there's a lot of stuff that proof doesn't have. I don't have to have a lot of the Word document formatting, or page breaks, or making tables, or whatever, because the agent just does it. I don't need to worry about that. It can do all the formatting for me. So you can make the products a lot simpler and faster to start than the legacy products are. And then there's all these other affordances that you need to start to have because the way agents interact with software is very different. So, for example, agents can do a lot at once. They can just do a billion different things to your document or your slide deck or your code base or whatever. And how you display that to the user is going to be very different than the way you might display a human being concurrent in your document and doing stuff. You need approval. You need an inbox that summarizes, here's all the stuff that's going to happen or has happened. You need logs and the ability to roll it back real quick. So there's all those kinds of considerations that change the actual product. And then the underlying UX of it or the underlying infrastructure you need is different too, because agents can make a billion requests in three seconds. So how are you going to deal with that, right? This is exactly why GitHub is having problems right now, because the number of people using GitHub is skyrocketing exponentially. And it's really just people's agencies and GitHub. [SPEAKER_00] So I think it's this whole new world that is just starting. You're just starting to see a little [SPEAKER_00] peek of it. But there's so many cool things about it. So, for example, in proof, and some of our other [SPEAKER_00] products too, when someone has a problem, they don't email support, their agent sends a bug report. [SPEAKER_00] underlying UX of it or the underlying infrastructure you need is different too, because agents [SPEAKER_00] can make a billion requests in three seconds. So how are you going to deal with that, right? This is [SPEAKER_00] exactly why GitHub is having problems right now, because the number of people [SPEAKER_00] using GitHub has skyrocketed exponentially. And it's really just people's agents and GitHub. So I think it's this whole new world that is just starting. You're just starting to see a little [SPEAKER_00] peek of it. But there's so many cool things about it. So for example, in Proof, and some of our other [SPEAKER_00] products too, when someone has a problem, they don't email support. Their agent sends a bug report. [SPEAKER_00] And an agent bug report is way better than a human bug report. It has, here's exactly what I did. [SPEAKER_00] Here's the exact repro steps. Here's, Proof is open source, so here's what I think is going on in [SPEAKER_00] the code base. And then we just get that. It becomes a GitHub issue. And then we can just send off an [SPEAKER_00] agent to fix it. And you can't do that with everything, but it's so much better. And you can [SPEAKER_00] see the glimmers of this, this very fast, closed loop between I ran into something, [SPEAKER_00] a paper cut, a little feature, I want a little bug, and my agent just goes off and talks to the company [SPEAKER_00] agent. And then the company agent just goes and fixes it. That, I think, is incredibly cool. [SPEAKER_00] So as a part of this, a lot of people are moving to CLI and trying to work from the terminal. [SPEAKER_00] Is the part of this prediction that people shift away from that and back to actual UX with agents [SPEAKER_00] running alongside them? CLIs are over. We speed-ran the CLI era. It was nice while it lasted. But I think it's pretty clear. It's not that CLIs are going to completely go away. Obviously, they've been around for the last 30 years or 40 years or 50 years or whatever. They will continue to be around. And I think there is this moment when cloud code was so popular, or when cloud code is really starting to gain in popularity, that people were like, the thing that's working is the fact that it's the CLI. And I don't think that's what it is. And when you move into an actual UI for this, you start to realize we made GUIs for a reason. And it's just nicer to be in a GUI. And you can get all the same benefits inside of a GUI, especially for [SPEAKER_00] non-programmer work. But I would estimate that definitely the majority of the technical people inside of every are not using CLIs anymore as their main work surface. I think a lot of programmers are still flipping into it every once in a while, but it's more or less they're using codecs, cloud code, cursor, that kind of thing. Awesome. Okay, I definitely wanted to make that part clear. So coming back to the big picture of the prediction here, there's these two modes of work that you're anticipating. One is this super agent within a company that you chat with through Slack, most likely, that can go off and do work and answer questions. And then there's on your computer running codecs or cloud code. And within that, all the work that you normally do on your computer is now going to be living within codecs or cloud code, or maybe some third party [SPEAKER_01] that emerges that we're not even aware of yet. Yes. And you're going to use apps inside of the [SPEAKER_01] internal browser of those tools. Wow. Okay. Listening to you talk about it, [SPEAKER_01] it may not feel as profound as it is, because this is a big change to how we work. We don't [SPEAKER_01] currently have an AI that we talk to regularly in Slack. And we also don't work currently mostly in [SPEAKER_01] codecs or cloud code. So this is actually a pretty massive shift. I think so. [SPEAKER_01] Is there anything else along these lines before we get into our next prediction? Well, a few things. [SPEAKER_01] I am definitely not an agent maximalist. I really think we're going to have a lot of different agents that we use. It seems pretty clear to me. And I really do think that two agents are better than one. [SPEAKER_01] So that's a good example. When I have codecs interact with another agent, [SPEAKER_01] it can give so much more context about me and what I want than I would be able to type. [SPEAKER_01] And it can go back and forth talking about things that would take a long time for me to express [SPEAKER_01] directly to an agent. You get this speed-up effect when you assume that your users are using [SPEAKER_01] codecs or cloud code or co-work as their basic way they access your app. And a really simple example: [SPEAKER_01] we have this hosted OpenClaw product, which we had on waitlist. We actually had to pause it because [SPEAKER_01] we started taking all the waitlist. And OpenClaw is just a very hard agent harness to make work. [SPEAKER_01] It's moving so incredibly fast. And if you're a platform for it, it's like, [SPEAKER_01] when things break, you can't fix it. It's very hard. But one of the things that we learned in that [SPEAKER_01] process is if you're, let's say, building an agent product or any new software experience, [SPEAKER_01] what you would assume, let's say, to set up an agent is you need to build a little web [SPEAKER_01] interface or a little Slack workflow that asks people about, okay, who are you? And what are you going to use this for? And what's your ideal dream outcome, or whatever the things are that you would put on an onboarding checklist? If instead, you just make a hard line of we are only going to service users who use Codex or co-work, what happens is you just paste a prompt into Codex or co-work. It goes and talks to the app, and the app can be either just a regular server or it can be its own agent. And Codex has so much information about you that it can just give it, like, here's all the stuff I've been working on with Dan. Here's all the ways that he might [SPEAKER_01] want to use this app, and then bring it back to me. And it's this very custom experience. And also, for a technical product like an agent, when something goes wrong, I can just tell Codex, go fix it. And Codex will go talk to the app and figure out what's going on for me. And so I think the whole paradigm starts to change when you assume that everyone's got an agent and those agents are talking to other agents in this really magical and important way. There's a couple more things I want to touch on before we get started. There's so much to talk about. One is you made this point about SaaS tools not using, you can use tokens from the model companies, basically, when using a SaaS tool. Talk a bit more about that because that may change the business model for SaaS companies in the future. That feels like a big deal. Well, I think it actually may save their margins. Because right now all these companies are rushing to add an agent to their offering. And thinking, the whole paradigm starts to change when you assume that everyone's got an agent and those agents are talking to other agents in this really magical and important way. There's a couple more things I want to touch on before we get started. There's so much to talk about. One is, you made this point about SaaS tools not using, you can use tokens from the model companies when using a SaaS tool. Talk a bit more about that because that may change the business model for SaaS companies in the future. That feels like a big deal. Well, I think it actually may save their margins because right now all these companies are rushing to add an agent to their offering and thinking, oh, the agent is going to be the main way that people interact with me. And I think that costs tokens, obviously. And I actually think once I have Codex or co-work as my main work surface, I still want to use SaaS. So this is another good prediction. I would buy SaaS stocks right now. I think the SaaS apocalypse is dumb, and SaaS stocks will be up majorly in the next couple years. Not investment advice. But I would buy SaaS stocks. So I think it saves your margin because now the way that you're thinking then is not, I have to build AI into this. It's more like, I have to make a piece of software that humans and AI want to collaborate on together. And that's hard, but once you build it, it is a lot cheaper than assuming everyone's spending tokens. And I think it's a good business. And part of the reason I'm so bullish on SaaS is everybody internally here is, like I said, we have all got agents and we're all using Codex and whatever, and we still pay for a ton of SaaS, and our SaaS spend is up year over year. And we're not vibe coding every single little thing. And I think that what agents do is increase the number of users of SaaS, not get rid of it. And so I think SaaS companies are going to see an insane spike in the amount of demand that they have because there's going to be tons of agents using these products at a very high volume. And like I said, that's a huge infrastructure challenge. There's a lot of interesting pricing challenges, but it makes me very bullish on SaaS. I love that. If anything else comes out of this conversation, Dan Shipper, SaaS is the future of AI. B2B SaaS. Hashtag send tweet. I love just, yeah, this is quite contrarian. And the other interesting piece is that the fact that you guys are hiring, that you doubled in people in the past year, which is not what people would have expected from a company that is so AI forward. Talk about what your experience there of just, okay, we still actually need humans. Automation is a lie in the sense that every time you automate something, in order to make sure the automation is working well, you need a human on top of it, making sure that it's working well. And so I wrote this piece a couple years ago called The Allocation about the allocation economy, the idea that the way that humans are going to work with AI is going to be like being a manager. And the thing that you have to remember about managers is managers actually spend a lot of time working. Most managers are not on the beach. They're checking in with their employees all the time and trying to figure out, okay, how do we make this work good? How do we make it better? How's it doing? How's this person doing? All that kind of stuff. And I think there are some differences between being a human manager and being a model manager, but fundamentally, it still requires a lot of time and attention. And I think that we kind of missed that in the model discourse. And one of the reasons is benchmarks make it look like AI is more autonomous than it is. And by autonomy, I mean something specific. I'm going to try to express it. It's a little hard to express. But I learned this for myself because I've been feeling this paradox a little bit. I've been feeling, we have so much automation, so much AI, and I also work way more. And I think part of the paradox started to resolve for me a little bit when I made my own benchmark. So I made this senior, it's called the senior engineer benchmark. And it's like, how good is AI versus a human engineer? And the way that I built it is, again, I have this app, proof. I just vibe coded it on the side while running the rest of every. And when we launched it, because it was completely vibe coded, it just started going down, and I couldn't fix it. And it was very embarrassing. I had a lot of egg on my face. And the product worked. We tested internally, we had a lot of beta testers. But the day after launch, it was just every 10 minutes, the servers would go down and people were looking at me, and I'd be like, I don't know what's going on. Codex, fix it. And Codex is like, I don't know what's going on. Or really, Codex is like, I do know what's going on, I fixed it. And then it would cause four other errors. And then you're just going around in a circle, and I wasn't sleeping. And I vibe coded so hard, I got bursitis on my elbow. So there's a life lesson in there. Vibe code or elbow? Yeah. So anyway, I got actually two different senior engineers to fix it independently. So I have two different rewrites of the code base that tells me how they did it. Right. And so what I get to do is, when we get new models, I just give the new model a prompt. I say, this is vibe coded slap. If you wanted to rewrite it from first principles, how would you write it? Go do it. And all the models until GPT 5.5 got a 30 out of 100. And a human senior engineer gets high 80s, low 90s out of 100. So there's a lot to go. And then I tried GPT 5.5, and it got a 62. And mind you, the 60 score was GPT 5.5 using an Opus 4.7 plan. Opus 4.7 plans are very good. GPT 5.5 is the only model, though, that has the sense of agency and confidence to just rip out old code and actually rewrite from first principles. Other coding models, they try, they end up papering around the edges. And they're like, oh, this is a big job. I'll just do a little patch. And you're like, no, I specifically told you not to. So GPT 5.5, there's a 30-point bump in the score, 60 out of 100. It's Plans are very good. GPT 5.5 is the only model, though, that has the sense of agency and confidence to just rip out old code and actually rewrite from first principles. Other coding models, they try, they end up papering over the edges or around the edges. And they're like, "Oh, this is a big job. I'll just do a little patch." And you're like, "No, I specifically told you not to." So GPT 5.5, there's a 30-point bump in the score, 60 out of 100. It's very clear that in a year or less, it's going to be senior engineer level, right? And that gives you a certain picture in your mind, especially based on how I named the benchmark, which I think a lot of benchmarks do. And I can tell you that when we get to that point, it will be very easy for me to change the benchmark to zero out the current model. So that gets zero out of 100. And so, for example, it seems like there's no skill or no thought into the prompt, which is this is vibe-coded slap, like, "Fix it from first principles." But actually, it took me a while to get to a prompt that didn't give away the answer, but got the model to reveal what it's capable of. And the original prompt I gave it was the original prompt that I gave it when I was trying to fix the issue and production was going down, which is, I'd woken up in the morning, and I was like, "Okay, we had four or five reported issues yesterday. I want you to go through all the issues and then make a plan for how to resolve all of them. And go do it," right? And every coding model on the market, and I'm pretty sure this, here's a prediction, I'm pretty sure every coding model on the market will still do this in a year. Every coding model on the market will take that instruction seriously. And if I tell it, "Here's a bunch of issues, go fix it," they will just go try to fix the issues. What an actual human senior engineer does is they go look at the code base, and they're like, "This is a piece of shit. This guy doesn't know what he's doing." And then they say, "We're going to have to actually rewrite a lot of this. And it's going to be hard and risky. I know you don't want to hear that, but we're going to have to do that." And if you ask the model, "Hey, should we do that?" it'll probably get there. But it's not going to do it on its own. And there's a [SPEAKER_00] lot of incentives pushing against it doing that. And even if it does that, there's always a [SPEAKER_00] higher frame for us to go. And so I think it's really important, when we think about benchmark [SPEAKER_00] progress, to think about it from that perspective, which is benchmarks rise on problems that we've framed, that we can articulate, that we can score. And there's a lot of work that's human work that can't be scored until you write it down. But the act of thinking to prompt it or write it down is something that you can't measure, but means that even if the benchmarks get saturated, it doesn't mean the same thing as you totally replace all senior engineers. And I think it's why, even though the models are getting better at automation, I still hire engineers. I am so excited to tell you about this season's supporting sponsor, Vanta. Vanta helps over 15,000 companies like Cursor, Ramp, Duolingo, Snowflake, and Atlassian earn and prove trust with their customers. Teams are building and shipping products faster than ever thanks to AI. But as a result, the amount of risk being introduced into your product and your business is higher than it's ever been. Every security leader that I talk to is feeling the increasing weight of protecting their organization, their business, and not to mention their customer data. Because things are moving so fast, they are constantly reacting, having to guess at priorities, and having to make do with outdated solutions. Vanta automates compliance and risk management with over 35 security and privacy frameworks, including SOC 2, ISO 27001, and HIPAA. This helps companies get compliant fast and stay compliant. More than ever before, trust has the power to make or break your business. Learn more at Vanta.com slash Lenny. And as a listener of this podcast, you get $1,000 off Vanta. That's Vanta.com slash Lenny. One thing I mentioned recently on the podcast, I heard that, speaking of the code that you have of humans writing code, data labeling companies are buying code that was written before 2021, 2022, before AI became a thing, is very valuable data. Our Kisnel human code. Yeah, exactly. That's exactly right. And it's so interesting that that's exactly the kind of code used to build this model. Well, what's interesting, so I want to clarify there. So I did not have a human that can write the code all by hand, because I actually think that that's sort of, it feels silly to me. I don't really care because I know if an engineer is not using AI, I'm not going to work with them. I don't really care. It's sort of like, am I going to race a human against the car? I probably wouldn't do that. But I would race a human in a car versus another human in a car and say which one's better. And in this case, the way the benchmark is structured is, yeah, these human engineers used AI, but they used it in a way that I could not because I didn't understand it. And I didn't have time, and I didn't really want to go in and try to understand the code base, to be honest. And I think that's a really important thing when we think about benchmarks, is AI is a broadly distributed technology that any human can use. And when we are benchmarking AI against humans, we're actually really always talking about one human using AI versus another human using AI because AI doesn't use itself. It may be able to in this slightly somewhat [SPEAKER_00] recursive way, but in any real use case, there's always a human pretty close to it, [SPEAKER_00] making sure that it's working. Okay, I want to try to wrap up our first bucket. There's so much to talk [SPEAKER_00] about. I've made a little list of things that I think people should do based on your predictions to be [SPEAKER_00] successful. We'll talk about this at the end too, but just a few things. One is start using codex or [SPEAKER_00] clock code more and more for the work you're doing, and especially the browser use tools inside of it. [SPEAKER_00] Two is allow agents to use your products. If you're building a SaaS tool, [SPEAKER_00] make it easy for agents to be a user, essentially. Three is start thinking about some Slack bot that you can [SPEAKER_00] work with, try out tools. I know Slack has their own Slack bot that I think is really good About. I've made a little list of things that I think people should do based on your predictions to be successful. We'll talk about this at the end too, but just a few things. One is start using codex or clock code more and more for the work you're doing. And especially the browser use tools inside of it. Two is allow your agents to use your products. If you're building a SaaS tool, make it easy for agents to be a user, essentially. Three is start thinking about some Slack bot that you can work with, try out tools. I know Slack has their own Slack bot that I think is really good too. And I haven't played with it, but people really like it. So look for, I guess, a tool that could become the AI agent within your company. Buy SaaS stock ASAP, not investment advice. I think that's totally right. My slight tweak is when you're thinking about building your software for agents, the current model is I'm building a CLI that an agent uses, but they're using it in a delegated task to the agent. The agent's using the CLI. [SPEAKER_01] And where I think it's going is you and the agent are using the app together. The agent's probably using [SPEAKER_01] CLI, but you're using the web interface, and they both need to be in sync. And that is, [SPEAKER_01] I think, a new challenge. That's really interesting. Awesome. Anything else before we get to our next category? Buy SaaS. That's the title. Oh man. Okay. So the second category of predictions is around just the shape of the work that we're going to be doing is going to change. What do you predict? There's all this interesting stuff in terms of the shape of work. Once you're in this land where you've got these, you've got async agents off that you delegate work to, then you've got your codex, cloud code work surface that starts to happen. So one thing that we see a lot internally, and you also see this in the big model companies, is the number of pull requests that you get skyrockets. We have people in consulting or in ops roles or whatever, or editors, just making pull requests. And that's really cool. And that's a very different shape of work where you can expect that a higher percentage of your company or your users are going to be doing things that previously only technical users could do. And what that does is it creates all this pressure on the other end for the people who have to deal with all of the new code for how to deal with that. And so I think there's a lot of interesting things that happen with that. So, for example, open claw, I mentioned that earlier. Pete gets thousands of pull requests a day on open claw. And then he just spins up 50,000 codex instances and then sorts through them and then merges a thousand of them. Pete gets, it's really crazy. I actually think that that's going to be more and more common. There's, it brings up a lot of really interesting questions around which pull requests should you merge. Pete gets, it's really hard to build things. And when you add capacity in one part of your process, it breaks things. It used to be really hard to build things, and now it's very easy. So the point is not, can we build it? It's, would it make sense with the rest of what we've built, and how do we keep a sense of a coherent whole? And also what do we delete? I think Anthropic does this really well. They delete a lot of stuff from cloud code to make sure that it's not bloated. So I think there's a lot of that going to happen on one side. There's a lot of non-technical people can do technical work, and then technical people are in charge of making sure that that work gets into a product or into a process in a cohesive, coherent way. And also their product people are going to be doing that too. And I think that's quite cool. Something I'm hearing from people is that now that everyone can do everything, engineers can design, PMs can code, marketing people can ship stuff, [SPEAKER_00] there's just this confusion about what the hell is my job anymore? [SPEAKER_00] Yeah. What am I responsible for exactly? Am I supposed to be shipping stuff? Am I still a marketing person? [SPEAKER_00] And it's just creating a lot of confusion and uncertainty in the world. [SPEAKER_00] I think that's for real. And one of the things that I think is special about every is everyone is a generalist and really loves having their fingers in a lot of different pots, or whatever the metaphor is. I think that'll probably settle down at some point and it'll feel more normal. Marketing people are still going to do marketing, even if they're touching the website. That's just part of marketing now. But I also think that you can get a lot further being generalists now. And that's really cool, especially for smaller companies. The other thing that I think is interesting is that there are definitely some new job roles that are a thing. And the thing that is becoming really clear is the whole forward deployed engineer concept, I think, is for real. And it comes out of every agent needs a human. You go to the big model companies, they have these agents that run internally. They have teams of people that run these agents, and I don't think those teams are going away. The models are going to get more powerful, the agents are going to get more powerful, and the number of agents is going to grow, but people are still going to manage them. And so that looks like a very specific kind of person. And we have a couple of those people internally here. And it's the people who are in charge of making sure your agents are working and doing the right thing. [SPEAKER_01] We also do consulting. So we lend that out to people. [SPEAKER_01] And I think that's a big thing that people want. [SPEAKER_01] And it's another one of those places where you're like, hmm, automation was supposed to take away jobs, but it looks like it just created one, or many. [SPEAKER_01] And there's a specific type of engineer that really loves, Nitesh, who's one of our, who fits this. He's an AI engineer, and he fits the forward deployed category, and he's on our team. [SPEAKER_01] He spends most of his time actually talking to one of our agents in Slack. [SPEAKER_01] We have an agent internally called Claudie, which runs our whole consulting practice. [SPEAKER_01] And he spends a lot of time in Slack. There is code. [SPEAKER_01] And he is using cloud code and other things like that. But a lot of it is just talking to it and being like, why did you do this dumb thing? Let's fix that. And so there are certain kinds of engineers that I think love that and love having their hands on the latest thing. And also love making this being that's in a workspace. And it looks a bit different than building more traditional software. [SPEAKER_01] We have an agent internally called Claudie, which runs our whole consulting practice. [SPEAKER_01] And he spends a lot of time in Slack; there is code. [SPEAKER_01] And he is using cloud code and other things like that. But a lot of it is just talking to it and being, why did you do this dumb thing? Let's fix that. And so there are certain kinds of engineers that I think love that and love having their hands on the latest thing. And also love making this being that's in the works, in a workspace. And it looks a bit different than building more traditional software. And your sense there is, we're not near a place where these agents don't need a human. You've said that so many times now, that agents need a human, and there's the setup part. And then there's that maintaining-it-forever part. And it feels like both are important. Is what I'm hearing. This is gonna be a job for a long time. AI is not gonna get smart enough to just automate it. It's fully automated for awhile. Yes. I'm simultaneously extremely AI-pilled, extremely, and very bullish on humans and the role of humans in making sure that AI is working well. Interesting. Okay. So the two buckets here that you're talking about, one is the way I think I hear what you described earlier is this: the pace of shipping software and everything is just increasing, which also means there's so much more work reviewing all this sloppy output. I was just talking to a data science friend, and he was saying how his team, a data science team, their job used to be to do analysis, answer questions, see if this experiment was positive. Now everyone's doing that, and they're sharing their results, and they're like, no, this is not correct. And most of their job is now reviewing bad data science work. Which is a problem. And the same thing is happening with engineers. And it means that you need more; you actually need engineers for this, and you need data scientists. And it means that you haven't set up the appropriate systems or agents to help you with this. So the way that it works inside of the big model companies, for example, at least one of them has literally a data science bot that every single person in the org can query, that is hooked up to their data warehouse, that knows who's who. So that it knows at the warehouse level who has permission to access what. And so all of the basic questions, because there's a team that sets up this bot, all of the basic questions that people might want to ask, that it sometimes gets, that might get wrong, they're constantly making sure it's getting it right. And so the data science team doesn't have to answer all the bullshit questions because there's another team building an agent that is set up to do that really well. But if the team didn't exist, the data scientists would hate their lives. Yeah. It does, though, make the job maybe less fun because you're just sitting there gardening people's sloppy work. Well, that's what I think is, it can actually make the job better because, for the data scientists, you are now not dealing with all the silly requests. [SPEAKER_01] You're dealing with the deeper questions that are harder for the team who's dealing with all the basic requests and building an agent to do that. [SPEAKER_01] It's filtering all that stuff out so you can focus. Here's a question I've been thinking about. I was not planning to talk about this, but it's something that I've been thinking about. [SPEAKER_00] The question is, which product tech role is the least changed now? So engineers, 100% of code AI now. It's a completely different job. Product management, a lot of the PRDs are, you don't have to write as much. You can ship code. You don't have to wait for people. Design, the whole design process dead, according to a recent guest. There's no time to do the whole design process. Very different role. Data science, very different work now. There's marketing, there's sales. So here's the question. What do you think is the least fundamentally changed role so far? Well, one interesting thing is, I don't know if this counts, but CEOs and investors, it seems still very, very optional whether or not they use this stuff. It seems that way. I think the opposite is actually true. My experience, and we do a lot of this with senior executives and senior leadership teams, my experience is that your company is only going to go as far as your CEO goes in AI, and it's not something you can delegate. You have to have your hands in it because otherwise you don't have an intuition for it. But for a long time, it has seemed like, yeah, that's something that the people who are doing the work have to do. But I don't have to do that. I'll just tell them what to do. And so I think if you're a CEO, you can get away with your day looking very similar. [SPEAKER_00] I think that will change rapidly at some point, where it'll be like, oh no, I'm way behind. [SPEAKER_00] But for now, or maybe even middle managers, those kinds of people, I think it is fairly similar. [SPEAKER_00] I think maybe sales, because I think so exactly, so in person. [SPEAKER_00] Yeah, that's my vote. [SPEAKER_00] It's creeping up in the kind of BDR. [SPEAKER_00] We can deal with a lot of BDR-type queries. [SPEAKER_00] You're only talking to people who actually want it, and you can do it for sales. [SPEAKER_00] It's so useful to do research; one of my favorite Codex experiences [SPEAKER_00] is we're hiring a head of L&D, and we always put out a job post, whatever. [SPEAKER_00] But I was like, I feel like there's this company called General Assembly in New York, and they've done really good technology education for a long time. [SPEAKER_00] And so I was like, I feel like someone who worked at General Assembly and is now into AI would be really good. And I just literally typed it into Codex and then went off and was doing something else. And I came back, and it found the perfect guy. It was like, worked at General Assembly, was an instructor, is super AI-pilled, and follows me on Twitter. So I just DM him, and then I had dinner with him. [SPEAKER_00] And it's like, that's crazy. [SPEAKER_00] That would have taken so long before, and super valuable for sales, for recruiting, all that kind of stuff. [SPEAKER_00] Yeah. [SPEAKER_00] Sales is where my mind went. [SPEAKER_00] And so I was, I feel someone who worked at General Assembly and is now into AI would be really good. [SPEAKER_00] And I just literally typed it into Codex and then went off and was doing something else. [SPEAKER_00] And I came back and it found the perfect guy. [SPEAKER_00] It was worked at General Assembly, was an instructor, is super AI-pilled, and follows me on Twitter. [SPEAKER_00] So I just DMed him, and then I had dinner with him. [SPEAKER_00] And it's, that's crazy. That would have taken so long before and is super valuable for sales, for recruiting, all that stuff. Yeah. Sales is where my mind went. The top of funnel, AI is helping a lot with sourcing and qualifying, things like that. It feels like the work of a salesperson is not fundamentally different. Yeah. And customer support has fundamentally changed. So that's interesting. Sales so far. Very good for those folks. Yeah. Okay. So maybe just summarizing some of the predictions in this bucket of the shape of the work, how it's going to change. What I'm hearing so far is there should be a lot more reviewing of other people's output as a part of the work. And then two, there's going to be a lot more almost babysitting of AI agents to make them do the thing you want them to do for deploying, and then just gardening them along the way. Make sure they continue to do their work. Anything else before we get into our third bucket? I would split it into less babysitting agents and more your forward-deployed team is trying to build a whole system that makes it so that people who have less knowledge can use that system without doing something dumb. And that's a really interesting engineering challenge. I think babysitting makes it feel like, yeah, you're just waiting for it to fuck up and then fixing it or whatever. And that can be the case. But I think a lot of it is this extremely interesting engineering challenge of building a system to enable everybody else in your organization to do what used to be a technical job. And then if you're not one of those people, you're the data scientist or whatever, you can go a lot deeper with AI into really important questions that eventually probably filter into the work that the forward-deployed engineering team is doing, but is more generative and more new, and you're dealing with harder questions. [SPEAKER_01] One last thing that I think is really interesting is I think that we will be reading way more AI-generated writing in documents and emails, and we will like it. [SPEAKER_01] And I think we are already doing this in coding, where we read plan documents. [SPEAKER_01] I don't want an engineer to handwrite a plan document. That would be very silly. It would be obviously silly. And I think the same is true when we did our quarterly planning at the end of 2025. We did it all with Notion agents. And we had one Notion agent. And then we had a top-level company strategy. And then we had everybody in the company just talk to an agent, and it asked them about what happened last year. How did it go? What were your goals? What do you want to do this year? What are your metrics? It pushed back, and then it was like, how does this relate to the overall company idea? All that stuff. And then I got these incredibly good AI-generated strategy reports or quarterly plans for each team. And then I could go in and be like, okay, who needs to talk to each other? Which teams need to talk to each other that don't know they need to talk to each other? And which one of these is actually low quality, or which one of these is high quality? All that stuff makes it a lot easier to process. And I see that all the time now. I consistently get AI-generated stuff, and there is a difference between an AI-generated document that's slop and one that's not. And the slop one is they took less time to make it than it takes me to read it, and they don't stand behind every line. So my expectation is, if you send me an AI-generated document, I think that's great. And if we talk about it and it's clear you have no idea what's in it, big no-no, not allowed to do that. And I think we have this aversion to AI-generated stuff that will go away because the kind of strategy document that GPT 5.5 can write when it's directed well by someone on my team is way better than them just dinking and dunking their fingers on the keyboard. Right. Most people are really bad at writing strategy documents. The bar is low. Yeah. Yeah. And same thing with email. Most of my email is written by GPT 5.5 and Codex right now. And I would honestly prefer it to say that it's coming from GPT 5.5, and I may change it to do that. But I had this experience the other day where I had to send an email to one of our investors, and I asked Codex, go do it. And Codex knows to ask me, and it usually does. But this time it didn't. And it just sent the email, and I didn't look at it at all. [SPEAKER_01] And I was like, fuck. [SPEAKER_01] And so I went to my sent folder and looked at it, and I was like, oh, this is exactly what I would have sent. [SPEAKER_01] And so it's pretty close to that. A lot of the time it can be a little over-formal. And there's a couple of things that, when you really think about it, most of your email is kind of rote. It's kind of prosaic. [SPEAKER_01] I definitely want to be the one to think about what it should say. [SPEAKER_01] But the actual sentences don't matter that much to me. [SPEAKER_01] Usually. Sometimes they do, a lot. [SPEAKER_01] And this is coming from a writer. I care a ton about writing. [SPEAKER_01] I think that human writing is incredibly important. [SPEAKER_01] And I expect we only publish human writing. [SPEAKER_01] Well, actually, we publish a mix of human and AI writing, but we always label it. [SPEAKER_01] Sometimes it's nice to have an AI co-author on certain things. [SPEAKER_01] I absolutely think that human writing is important. [SPEAKER_01] And I think that the reaction or aversion to AI writing is silly. It's such an interesting lens on that, because when people think about AI writing, they think about social media and videos. And your point is internally, if you're just working on planning and documents and email and things like that, that is much less scary that it's AI-written. To your point, people are already doing this. [SPEAKER_01] And I expect we only publish human writing. [SPEAKER_01] Well, actually, we publish a mix of human and AI writing, but we always label it. [SPEAKER_01] Sometimes it's nice to have an AI co-author on certain things. [SPEAKER_01] I absolutely think that human writing is important. [SPEAKER_01] And I think that the the the reaction or the aversion to AI writing is silly. It's such an interesting lens on that, because when people think about AI writing, I think about social media and videos. And your point is internally, if you're just working on planning and documents and email and things like that, that is much less scary that it's AI-written. And to your point, people are already doing this. You almost prefer it a lot of times because people are really bad. Totally. Anyway, we have this too for external stuff. We publish all these guides, and the guides are often agent-assisted. They're agent-assisted, and the agent is a co-author, and they're intended to be read both by humans and by agents. And that's because if you're writing a huge informational thing, you do this all the time. In order to really apply it, the best way to do that is just have your agent ingest it. And remember the next time I'm doing pricing to remind me of this guide, and we'll go through it together or whatever. It allows you to operationalize the ideas much better, and it allows you to go much deeper because agents can read 10,000 pages in a second. And so you talk to the human about the story and the stuff that matters and the core ideas, and the agent has all the details that it can then apply for you when you need it. Awesome. Anything else in this category before we get into our final category? No. [SPEAKER_00] Okay, let's do it. [SPEAKER_00] So the final bucket is just who will be successful in this AI future that we are approaching, slash what should people be working on to be successful in this next year or two? [SPEAKER_00] I am super, super bullish on PMs. [SPEAKER_00] And I know that your audience will probably love that. [SPEAKER_00] But my anecdotal case that has convinced me of this is we have this guy internally, his name is Marcus, and he runs Spiral, which is our writing app. [SPEAKER_00] Marcus is a PM by training. [SPEAKER_00] He previously ran Axios's writing product and was a PM and had a big team, and it got to tens of millions in revenue in ARR. [SPEAKER_00] And he took a year off that job and just got super AI-pilled and just learned how to use Cursor really well. [SPEAKER_00] Now, I think he uses cloud code, but he was extremely Cursor-pilled for a long time. [SPEAKER_00] And he's, I would call him, lightly technical. Knows what a database migration is. If he has to look at the code, I think he can understand it. But we never could have hired him to do this job even a year ago. But the coding models have gotten good enough that he can pair the technical knowledge that he does have with his really spiky product sense and sense for writing and sense for users. And it's so dangerous. He ships faster than almost anyone on the team. And he has such an eye for every single user, every single conversation. What does it mean, and how do we collect it into a story about where we want to go next, and what are the issues we need to fix, and all that kind of stuff? And I think that he feels liberated because he doesn't have to organize a whole team of people to do that. He can just do it. And it's super impressive. And it makes me very, very bullish on any PM who gets really a adult music to my ears, Dan. You're making a lot of very happy listeners here. I've been saying this for a long time, too. It's just the skills you need to build are the things. The building now is done for you. What do you need to be good at? Figuring out what to build, figuring out if it's great, figuring out problems to solve. So I love that you're actually seeing this come to fruition. I really believe it. This could be the highest-rated podcast episode of my whole podcast. I love it. Hell yeah. It's going to be okay. Stas is back. PMs are back. [SPEAKER_01] This is the most contrarian episode I've ever done. [SPEAKER_01] Oh my God. [SPEAKER_01] So, okay. [SPEAKER_01] So the other, the other people that I think are going to be super, super power people. [SPEAKER_01] And I, again, this is because we see this internally, is full-stack designers. [SPEAKER_01] If you're a designer and you're in these tools all the time, you're so used to, okay, make this beautiful interaction. [SPEAKER_01] And then the engineer just doesn't want to do it, or it doesn't happen the way I think it should happen. Or there's all this stuff. And I see so many designers for us internally or externally where they now feel so empowered to go build stuff. Because they're like, I have all these ideas to make things look amazing and these interesting interactions. And that's the exact thing that it's really hard to do with vibe coding because it just all looks the same. So it all looks like slop, and they can make stuff that looks so different. And now they can actually build it. And what you see when we work with them internally is now they're just making pull requests. They don't need to hand it off as much. Sometimes they do, but a lot of times they just make pull requests, and it's like the thing is built, and that's it. And I think it's incredible for the way that companies work, but it's also, there's a huge opportunity for those people to become entrepreneurs and start their own thing. Because they can make stuff now. And I think designers are such creative people. And I think AI is a super tool for anyone like that. I so agree. Even though there's cloud design, there's all these AI designing tools. Once you see it, you're like, that's definitely cloud design. They're like, the creativity, to your point, is skin. It just feels like it's going to be more and more valuable to stand out from all the slop that people are shipping and launching constantly. So I completely agree. It's interesting that designer roles, I do research on the job market, and interestingly, designer roles have not grown in a while. So I'm waiting to see if that becomes a big trend. Just like, we need more designers. Hmm. That is really interesting. We'll see. We'll see. We'll see. That might be a way to predict this, is are people hiring more designers? I don't know. [SPEAKER_00] That is interesting. [SPEAKER_00] Yeah. All right. So that's so PM designer thriving, killing designer thriving. So I completely agree. It's interesting that designer roles, I do research on the job market, and interestingly, designer roles have not grown in a while. So I'm waiting to see if that becomes a big trend. Just like we need more designers. Hmm. That is really interesting. We'll see. We'll see. We'll see. That might be a way to predict this: are people hiring more designers? I don't know. That is interesting. Yeah. All right. So that's so PM designer thriving, killing designer thriving. I also just think generally the AI job apocalypse is not really a thing. Absolutely. We see companies starting to reorganize. And I think that makes a lot of sense. I think, to be honest, a lot of the reorganization, you can say it's AI, but it's like we overhired and the company's not doing as well. And all that was coming, and this is a good excuse. But the mass unemployment thing I think some AI CEOs are talking about, I think that's not going to happen. The pattern that I see so far, and again, I don't have a total crystal ball, but I do feel like we've seen enough of the new model drops to have some sense of how this is going, is that what a new model drop does, or what models do in general, is they make yesterday's human competence cheap. So what I mean by that is they ingest all this data of what has happened already, and they make it really cheap to deploy that in whatever situation you want as your own. Right. And what happens then is this is a new power that everyone has. [SPEAKER_01] So it gets adopted super rapidly, and suddenly that stuff is everywhere. [SPEAKER_01] It's like suddenly anyone can make a landing page. [SPEAKER_01] There's new landing pages everywhere. [SPEAKER_01] Suddenly everyone can write. There's slop tweets everywhere. [SPEAKER_01] But what's interesting is because it's all coming from these models, and everyone's using basically the same models, [SPEAKER_01] it all looks the same if you use it in the most default basic way. [SPEAKER_01] And so it becomes commoditized. [SPEAKER_01] It's not valuable anymore. [SPEAKER_01] And what humans do is we go in there and we're like, yeah, we have all this frozen human competence from yesterday. [SPEAKER_01] How do I use this? Like make something new and interesting. And I really think that structurally, because of the way the models work, because of the financial incentives of model companies to make them compliant and aligned structurally, they're always going to be trailing behind those people who are taking the models and using them to make new expertise or make new things that haven't been done that way before for their very, very particular situation. And that stuff is going to get incorporated into the models. But again, it will create room for people to push further ahead. And I think that you see this in a small way in pretty much all the jobs. Engineers, suddenly everyone's an engineer. That doesn't mean we fire the engineers. There's way more demand for engineers because you need the engineers to figure out, okay, this is all slop. How should this actually go in our code base? And I think that's something that the benchmarks rising doesn't really capture. And it feels like a thing that will take a long time to change. People may be hearing in this prediction here of just, okay, the job apocalypse, people are not going to be all fired. There's going to be human jobs remaining for quite a while. It may be almost too comforting because you probably have to change the way you operate to still have a job in the future. Do you have any sense of just, here's what you need to do to not be one of these layoffs? Yes. And I think that is actually super important. The only thing you need to do is ride the models. And that means use them for whatever it is that you do. We've talked about how codecs and co-work are becoming the standard operating system for work. If you're just doing that. And when new models come out, you're trying them and figuring out, okay, how can I, now there are new powers. How can I use them instead of just being like, I'm going to try to ignore it because it makes me afraid, which I think is honestly rational. It's a reasonable response. And also, if you ride on top of them, they extend your powers in a way that doesn't leave you behind. Like you're part of the future and part of the way work happens. And I think that we're going to need people doing that for a very, very long time. I like this term, ride the model. So what's it like, say a new model comes out. What do you think someone, say, working at Salesforce, say a PM at Salesforce, what should they do to ride the model? Well, one of the things that's really interesting is a lot of companies handicap their employees from even doing this because I don't know what model, I don't know if you can use the latest models in Salesforce. A lot of times you have to wait or it's whatever. So maybe you have to do it in your off time. But the thing that I really like to do with new models is play. And there are certain things where I know it can't quite do it yet. But when a new model comes out, I always turn the rock over again to be like, can I do it now? You know? So I could not do the senior engineer benchmark last time, and I turned the rock over again. And now it's at a 60 out of a hundred, which is really good. So the way to ride the models is not one specific thing because they're always changing, but it is to be curious and playful, to apply the new model to whatever it is that you care about, whether that's your job or something outside of your job, and to keep turning over rocks because it may not work now, but it may work eventually. It probably will work eventually. And the way that you use it matters. So what's really cool is that I think people think of the edge of AI as being in San Francisco. And I actually don't think that that's where it is. I think the edge of AI is wherever AI meets a real human doing something because the people in San Francisco, they're making it, but they don't actually know a lot about how to use it. They don't know, or at least they don't know everything about how to use it. They need to see how other people use it. [SPEAKER_00] It probably will work eventually. [SPEAKER_00] And the way that you use it matters. [SPEAKER_00] So what's really cool is that I think people think of the edge of AI as being in San Francisco. [SPEAKER_00] And I actually don't think that that's where it is. [SPEAKER_00] I think the edge of AI is wherever AI meets a real human doing something, because the people in San Francisco, they're making it, but they don't actually know a lot about how to use it. [SPEAKER_00] They don't know, or at least they don't know everything about how to use it. [SPEAKER_00] They need to see how other people use it. [SPEAKER_00] And so, whenever a new model comes out, you get to be one of the first people in the world to discover what it might be useful for. [SPEAKER_00] And that's, it's like a new discovery. [SPEAKER_00] And I think that's why, for example, we're in Brooklyn, but I really think of us. [SPEAKER_00] And I think we are quite far ahead of people in San Francisco because we just use them for everything. [SPEAKER_00] And if people do that consistently, I think it's going to be very hard to lose. That is one of the most amazing things about AI right now, is no matter how much money you have, or how little money you have, you have access to the most advanced AI model. It's not free, so you need some money, but you can get it immediately when it comes out. [SPEAKER_00] Maybe the only people that have an advantage are the people working at OpenAI or Anthropic. [SPEAKER_00] But otherwise, it's just available. [SPEAKER_00] I know I was at their event with you, their Code with Claude event with you last week, or a couple weeks ago. [SPEAKER_00] And they're all using mythos, and I'm like, God damn it. It's so annoying. [SPEAKER_00] But I think that's totally true. [SPEAKER_00] That is, if IBM had invented AI, you can bet it would not be like this. [SPEAKER_00] And it would be a bajillion dollars, and only the top companies could use it, and they would be using it in the weirdest, most uninteresting ways. And I think it's really important that AI was built in America and in the Silicon Valley culture. [SPEAKER_00] That's like, we want to make intelligence too cheap to meter. [SPEAKER_00] That's not the default stance. [SPEAKER_00] And it means that everyone has this broadly accessible tool that they can use. [SPEAKER_00] And I think that's amazing. It's such a good point. And interestingly, it's also created the fastest-growing companies in history, the biggest companies in history. That's true. Not a way to win. Those Silicon Valley guys, they're smart. If I zoom out on the conversation, it's really interesting. There's these two sides to the coin. One is, not a lot is actually, so much was not changing. SaaS continues. Jobs are not disappearing. We're still emailing each other. We're still working in Slack. A lot of the work has not changed. On the other hand, every role transformed. Engineers don't write code. PMs don't write PRDs. Design and design. It's so interesting how much has changed, how much has not changed. I don't know. It's interesting that people think it's going to be this whole new world, but in many ways it's okay. It'll continue the way it is, with a lot of stuff around the edges. That's how I feel. I'm simultaneously so excited, and it feels like everything has changed, and I'm so bullish on it. And the progress that we're making, all that kind of stuff. And yeah, I just feel like there are these things where they're going to be pretty similar to how they are. And that's probably good. [SPEAKER_00] And I think generally our intuitions about the future, the model that I have of what our intuitions are about the future, is the intuitions that people had in the Middle Ages about what happened at the end of the horizon. [SPEAKER_00] It's like, are there dragons? [SPEAKER_00] Does it drop off into nothingness or whatever? [SPEAKER_00] A lot of people have a lot of deep intuition that there's something terrible going to happen over the horizon. [SPEAKER_00] And also that some people are like, there's something incredible. [SPEAKER_00] It's going to change everything. [SPEAKER_00] We're all going to be happy as a utopia. [SPEAKER_00] And what happens is you get there and you're like, there's some really cool things. [SPEAKER_00] There's some not cool things. [SPEAKER_00] And it's just another horizon. [SPEAKER_00] And I think that's the way to think about the future. [SPEAKER_00] And until you get to that place where you're starting to see it, and I think we get to see it because we get to see it internally all the time. [SPEAKER_00] It's important not to let your mind get away from you and be like, this is going to happen and this is going to happen and whatever, because you're going to tell a story that sounds so real in the moment. [SPEAKER_00] But later on, you're like, actually, it's much more complex than that and somewhere it's a both. [SPEAKER_00] Everything's changed and nothing has. [SPEAKER_00] And once you get there, I think you're starting to see, like, oh yeah, this is a real thing. [SPEAKER_00] Part of it is the AI companies are very good at scaring us about what might happen in the future. [SPEAKER_00] And I think that's actually shifting. [SPEAKER_00] I think they've realized maybe we should not freak everybody out about that. PR strategy just does not make any sense to me. [SPEAKER_00] I do think that it's genuine, but it's so ineffective. And I think it's also wrong. Hmm. How about we end with maybe just a few things listeners should do to be successful over the next year with the way the world is moving. Ride the models. I would try all of your workflows in codex or cowork and see how that works. And if your company doesn't let you do it on your own time, I would try out some of these agent products like open claw or Hermes or, for less technical people. There's Victor. We have one plus ones. I would get comfortable with both of those ways of working and try to have fun. I think there's too much of I'm doing this because I have FOMO. [SPEAKER_00] It might, I might lose my job, or I might miss out on this big thing or whatever. [SPEAKER_00] And the best way to actually figure out interesting, useful things to do with AI is to do something enjoyable. [SPEAKER_00] We had Nikhil Singal on the podcast, and the way he described it is you got to find your moment of joy with AI. [SPEAKER_00] Once you find, wow, I can't believe AI did this for me. There's Victor. [SPEAKER_00] We have one plus ones. [SPEAKER_00] I would get comfortable with both of those ways of working and try to try to have fun. [SPEAKER_00] I think there's too much of, I'm doing this because I have FOMO. [SPEAKER_00] It might, I might lose my job, or I might miss out on this big thing or whatever. [SPEAKER_00] And the best way to actually figure out interesting, useful things to do with AI is to do something enjoyable. We had Nikhil Singal on the podcast, and the way he described it is you got to find your moment of joy with AI. Once you find, wow, I can't believe AI did this for me. This is awesome. We're going to keep building stuff. Yeah. I agree. If you haven't seen that yet, then it's just try find, try solving it. The thing I hear a lot is just find a problem in your life or work and see if AI can do it. Go to lovable, go to plug code, go to replicate, just try to build the thing. Yeah. And often it's, holy shit, this is so cool. [SPEAKER_01] Dan, is there anything else that we haven't covered? [SPEAKER_01] We've gone deep on so much. [SPEAKER_01] Is there anything else you wanted to share? [SPEAKER_01] Anything else you want to predict or just say before we get to a very exciting lightning round? [SPEAKER_01] I think we covered it. [SPEAKER_01] We did a lot. [SPEAKER_01] This is awesome. [SPEAKER_01] And I'm very excited to see how well or poorly I do in a year. [SPEAKER_01] And I hope that you hold me to it. [SPEAKER_01] We're going to have an AI score us. [SPEAKER_01] How about that? [SPEAKER_01] Well, great. [SPEAKER_01] Look at the world. [SPEAKER_01] Like a dance prediction. [SPEAKER_01] See where he goes. [SPEAKER_01] Well, with that, Dan Schipper, we've reached a very exciting lightning round. [SPEAKER_01] I've got five questions for you. Are you ready? [SPEAKER_00] I'm ready. What are two or three books that you find yourself recommending most to other people? Obviously Annie Dillard. Everyone at every has to read The Writing Life. [SPEAKER_00] When you join, you get a copy and you have to read it. You only have to read the last chapter, though. [SPEAKER_01] I think the last chapter is incredible. [SPEAKER_01] And it is at the intersection of writing and technology and the future. [SPEAKER_01] And it's relationship to the future and to time. And I think that's everything about every wrapped up into a very tight chapter. It's so good. And I think Annie Dillard just generally is fantastic. What else do I recommend? I'll just tell you a couple of things that I've read that I really liked recently. And whenever I like something, I always just tell everyone about it. So I have recommended these a lot. I've been reading one of the things I learned, which I didn't know, is Churchill is a really good writer. And he has a whole history of World War II that he wrote. And it's a combination history and memoir. And I think that's so cool because he was there, he did it.