Open Reader

WATCH ME WRITE LIVE WITH CHATGPT VOICE

completed 1:57:48 Jul 30, 2026 Watch on YouTube

Current Status

completed

Video ID

DXknnapb_GU

RAG / Chat

Enabled
WATCH ME WRITE LIVE WITH CHATGPT VOICE
Description

I'm writing the definitive history of Codex and live streaming my process.

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Voice-mode agents can become a practical writing and research control plane when the human retains editorial judgment, uses the agent for retrieval and iteration, and actively corrects its failure modes.
  • Why it matters: The stream provides a live, unscripted view of an agent-assisted knowledge-work loop: spoken ideation, document operations, sub-agent delegation, source retrieval, drafting, critique, and recovery from tool and context failures.
  • Best use: Watch as an operating-pattern study rather than a source on Codex history: extract the interaction design, delegation patterns, and guardrails for using voice agents in high-context work.

Executive Summary

This is a two-hour live writing session in which Dan Shipper drafts part of a reported history of OpenAI Codex using ChatGPT voice mode and a document product called Proof. The immediate writing assignment is to explain test-time compute, chain of thought, and why reasoning models enabled a resurgence of Codex. But the more valuable content is the demonstration of voice as a work interface: he speaks rough thoughts, asks the agent to preserve them as raw material, requests research and drafting alternatives, and uses the agent to read, navigate, annotate, and modify his draft.

The workflow is explicitly human-led. Shipper uses the system to reduce friction around recall, transcription, document navigation, terminology, source lookup, and producing alternative prose. He does not accept generated prose blindly; much of the session consists of rejecting weak metaphors, identifying when the model has turned notes into unauthorized prose, correcting attribution, and forcing it back to the exact document text. His core working rule is to verbalize intent and constraints, then iterate until the model’s output becomes useful.

The transcript also contains a clear technical argument about reasoning models. Earlier language models had to turn their computation immediately into user-visible output, while reasoning models can spend additional inference-time compute generating intermediate reasoning, trying tools or paths, incorporating results, and only then answering. Shipper frames this as the shift from asking a model to reason step by step publicly to systems that can use those steps deliberately before responding.

Operationally, the session exposes gaps that matter for any agent-control-plane design: permissions block edits, document-tab state is unreliable, the agent summarizes when asked to read verbatim, it invents or overextends prose, it struggles to retrieve the right context, and it occasionally misattributes material. The productive pattern is therefore not autonomous writing; it is a tightly supervised loop with explicit artifact boundaries, read-backs, verification, and human ownership of the final narrative.

Key Takeaways

  • Claim: Voice interaction is most useful for writing when it captures unformed thinking without forcing the author to stop and translate every idea into polished keyboard prose. | Evidence: Shipper repeatedly dictates fragments, asks the agent to place them in a separate "free write" section, and then uses the stored raw material to shape a test-time-compute explanation. He describes voice mode as having "truly changed how I work" after completing roughly 1,000 words during the session. | Implication: For high-context work, build voice workflows around capture-first artifacts and explicit promotion into the canonical draft, rather than letting a model silently convert spoken exploration into final copy. | Caveat: The value depends on preserving a distinction between rough notes and approved manuscript text; the agent initially violated this by turning dictated rough material into additional prose.
  • Claim: The best division of labor is human editorial direction plus agentic retrieval, transformation, and option generation—not delegated authorship. | Evidence: Shipper asks for alternative metaphors, a sub-agent’s independent technical pass, Bill Bryson-style structural inspiration, research sources, interview quotes, word-count checks, proof-document edits, and verbatim read-backs. He rejects many results as "no good" or "so bad" and continually narrows the request. | Implication: Treat an agent as a responsive editorial room: ask it to surface options and evidence, but keep taste, framing, acceptance criteria, and final synthesis with the operator. | Caveat: Style imitation requests require care; the agent itself declines to write in Annie Dillard’s exact style and instead offers an original passage using a similar high-level movement.
  • Claim: Test-time compute is framed as the crucial technical shift from immediate answer generation to allocating additional compute after a specific difficult request arrives. | Evidence: The working explanation contrasts early models, which had to work out an answer "one word at a time, immediately and in public," with reasoning models that can generate additional hidden intermediate tokens, explore a longer chain, call tools, incorporate results, and delay the visible response. The model summarizes this as "inference-time scaling." | Implication: For agent systems, model selection and orchestration should distinguish quick-response tasks from tasks that warrant long-horizon rollouts, tool execution, verification, and a larger inference budget. | Caveat: The transcript’s technical explanation is a writer’s simplified framing, not a rigorous model-architecture tutorial; it also acknowledges that code-model RL and test-based evaluation existed before reasoning models.
  • Claim: Chain of thought became materially more useful when it shifted from a user-requested prompting trick to a trained, deliberate capability supported by inference-time compute. | Evidence: Shipper’s draft states that early users found models often performed better when prompted to "think step by step out loud." He labels this chain of thought—showing intermediate working before the final answer—then contrasts it with O-series reasoning models that could take intermediate steps before responding. | Implication: When designing workflows, do not rely only on prompting models to explain themselves. Use reasoning-capable models and verification environments where their extra inference can be rewarded by task outcomes. | Caveat: The transcript recognizes that the model’s intermediate paths are not necessarily human demonstrations; training can combine demonstrations with rewards from verifiable outcomes, with the exact mix varying by system.
  • Claim: Coding is especially compatible with long-horizon reasoning because outcomes can be checked, allowing models to learn and operate through iterative attempts rather than a single plausible response. | Evidence: The conversation distinguishes code tasks where tests and other outcomes are verifiable, notes that models can run commands and read results back into their context, and cites the observed shift toward models successfully handling "longer and longer-horizon tasks." | Implication: Prioritize agent deployments where the environment supplies cheap feedback loops—tests, linters, previews, simulations, acceptance checks, or measurable business outcomes—because these enable safe productive iteration. | Caveat: The transcript does not provide a detailed benchmark or causal proof that test-time compute alone caused coding-product success.
  • Claim: A voice agent needs reliable artifact access and explicit verification because seemingly small control-plane failures quickly undermine high-context work. | Evidence: The agent cannot add a comment due to permissions, cannot see a new Proof document tab, incorrectly condenses text when asked to read it verbatim, duplicates content, needs repeated share-token guidance, and later requires correction for a quote attribution. | Implication: Require document identity, scope confirmation, exact-read modes, change previews, permission recovery, provenance tracking, and attribution checks before allowing an agent to edit core work products. | Caveat: Some failures may reflect the specific live configuration rather than general limitations of all voice-agent systems.
  • Claim: The strategic product thesis is that reasoning-model gains pointed teams toward asynchronous, long-running agent products rather than merely better chat interfaces. | Evidence: A reported interview quote attributed to Alex says, "If we could increase test-time compute, we got better performance," followed by the product conclusion: "design around asynchronous models, long-running models." Another reported observation says users were showing models that had been working for 48 hours. | Implication: For Ken’s agent systems, plan for job-based execution, resumability, observability, user interruption, and completion handoff—not only synchronous conversational UX. | Caveat: These are reported statements being assembled into an unfinished article, and the transcript contains no independently verifiable interview materials or timestamps.

Detailed Brief

The live writing loop: useful interaction patterns

  • Claims: The author repeatedly uses read-back as a cognitive reset: he asks the system to reread the current section before dictating the next thought.; He separates two task modes: preserve spoken material exactly in a free-write log, versus make approved edits to the proof document.; He requests independent alternatives from both the main assistant and a named sub-agent, then compares them rather than treating one output as authoritative.; He uses the model as a conversational research aide to identify terminology such as "test-time compute," "inference-time scaling," and "chain of thought."
  • Evidence: The agent is instructed to create a dated free-write section for the "thinking" passage rather than placing exploratory language into the PCF or final draft.; Shipper requests a GPT-5.6 sub-agent pass on the physical and architectural description of reasoning, alongside the main agent’s version.; He asks for the most useful reasoning-model explanation for a technically sophisticated reader and is pointed toward OpenAI’s "Learning to Reason with LLMs" and the DeepSeek R1 paper.; His questions move from sentence-level refinement to conceptual tests of whether the explanation is actually understandable to a general reader.
  • Caveats: The agent repeatedly claims successful edits or searches before the author can verify them, so verbal status messages alone are insufficient evidence of completed work.; Live-stream distractions and audience comments interrupt the writing loop; a production workflow should separate operating focus from community engagement.
  • Implications: A robust voice-work product should expose explicit modes such as capture, brainstorm, research, draft, revise, and apply edit, with durable links between each mode’s artifacts.; Parallel sub-agents are most useful when assigned genuinely independent approaches and when a human has a clear evaluation criterion.

Reported Codex-market narrative being drafted

  • Claims: The unfinished essay frames Codex as an incumbent comeback after Anthropic’s Claude Code appeared to establish a lead in the coding-agent category.; Its proposed explanation is organizational and technical: a small internal team could build a disruptive product fast enough to become undeniable, aided by AI leverage and reasoning-model capabilities.; The author treats the merger of a disruptive coding product back into ChatGPT as a rare example of a large technology incumbent disrupting itself.
  • Evidence: Drafted, unverified figures include Claude Code reaching more than $500 million annualized revenue within three months and more than $2.5 billion annualized by summer 2026; Codex is described as passing 5 million weekly active users before a merge and the combined app adding more than 1 million users per day.; Shipper reports that most of the Every team switched from Claude Code back to Codex and ChatGPT over the prior three months.; The session mentions interviews with people identified only as Alex/Alexander, Greg, Thibault, and Mia.
  • Caveats: These claims are draft material from a forthcoming reported essay, include future-dated references, and are not independently substantiated inside this transcript.; The livestream transcript is duplicated in its latter portion and contains speech-recognition errors, including product-name distortions such as "Broadcode," "flawed code," and "Chat GPP."
  • Implications: Use the market claims as leads for source checking, not as investment-grade evidence.; The more durable lesson is that a technical capability shift can require a product-surface and operating-model change, not merely a model upgrade.

Notable Concepts & Terms

  • Test-time compute: Compute allocated after a user submits a task, allowing the model to spend more time generating intermediate work before it commits to a visible answer.
  • Inference-time scaling: Another name used here for scaling performance by increasing compute during use, rather than only increasing training compute.
  • Chain of thought: The intermediate step-by-step working that users once explicitly prompted models to produce; the video presents reasoning models as making such intermediate work more deliberate and often private.
  • External processors: Shipper’s writing metaphor for early language models: they effectively had to perform their reasoning through the same stream of text delivered to the user.
  • Long-horizon tasks: Tasks requiring many sequential actions and feedback cycles; the transcript treats growing success on these tasks as a key reason to build long-running coding agents.
  • Proof: The document environment used in the demonstration for a live draft, comments, free-writing logs, and agent-directed edits.
  • Asynchronous models: A product orientation implied by long reasoning and tool-use runs: systems should work over time and return results later rather than requiring every task to fit a synchronous chat turn.
  • Codex merge: The transcript’s reference to combining Codex with ChatGPT; it is positioned as evidence that the coding-agent product became central enough to reshape the larger product.

Operator Notes / Why Ken Should Care

  • Design voice-agent workflows with a hard separation between raw dictation, proposed copy, and approved document edits; require explicit promotion between stages.
  • Add an "exact read" command that disables summarization, plus document/tab identifiers and edit diffs so the operator can verify the agent is acting on the correct artifact.
  • Route long-horizon coding and research jobs through asynchronous execution with checkpoints, tool-result ingestion, and a clear completion/approval handoff.
  • Instrument permission and context failures as first-class workflow events; do not leave agents stuck in vague retry loops when access tokens, tabs, or comments fail.
  • Require provenance for research and quotations: source location, speaker attribution, and confidence before a system inserts factual claims into a canonical draft.
  • For content production, use parallel agents for independent framing or research passes, but make the human explicitly select, reject, or synthesize outputs rather than merging automatically.

Source/Metadata

  • Title: WATCH ME WRITE LIVE WITH CHATGPT VOICE
  • Transcript words: 14918
  • Duration seconds: 7068
  • Timestamp note: No usable timestamps or chapters were present in the supplied transcript. The transcript contains a substantial duplicated closing segment and transcription errors.

Transcript

12576 words en Processed in 1155.7s

All right. I'll just put the link in. Stream link. She should be live. All right. Here we go. Ready when you are. So, just FYI, I'm live streaming right now, the process of doing this writing. Can you just take a look? The live stream should be up on Dia. Just take a look and see if there's anyone on the stream. Okay, I'll check that for you. Checking now. All right. While you're doing that, I'm just going to tell people on the stream what's even happening here. And you can just interrupt me whenever you figure it out. Okay. So I have been taking a week off to write a history, a definitive history of Codex. And during this week, I got accidentally extremely voice pilled. I feel like voice is so good. And I figured out a bunch of ways of writing and just using the computer with voice mode that I think are really, really cool. So I figured I'd spend a little bit of time today, since I'm writing anyway. I'm just going to live stream this process. And one thing I really like is watching people work. So I'm pretty much going to try to forget that this is even being live streamed. And I'm just going to work the way that I would normally work on this piece. Hopefully it'll be an interesting thing for you guys to see. And also, hopefully I'll get my words, and I've done about 500 words so far today, and I want to get another 500 in. So I'm also going to just play. Let's play some music. I don't even know if this will work, if you'll be able to hear it. Codex, or Chetjit, did you figure out if anyone's watching this? Almost there. Hold on. Yeah, someone's there. The user at Glisalia has commented twice. LFG and what's our favorite voice, Dan? Also, it looks like 87 people are in the audience right now. Okay, great. Awesome. Well, if anyone says any comments that you think I should know about, just let me know. Okay, so let's start with, I want a refresher on where I am, and I want to do a 45-minute writing sprint. So we should finish by about one. And let's start with a refresher on where I am. So we just did a section on the gap and why the gap is really hard to close. And now we're going into test and compute. So just to get me started, can you just reread that section, not the opening scene, the one right after that, for me? Sure, I'll pull that up and read it. At the time, I thought the gap was probably fatal. The danger in a new product category isn't simply falling behind, it's running the wrong race. If you misunderstand the category from the start, each step locks you deeper into that mistake until changing course feels impossible. And for a while, OpenAI hit real turbulence for the first time since I'd started covering them in the GPT-3 days. Claude Code surged with startling speed. Anthropic said it hit over $500 million in annualized revenue within three months of launch. By summer 2026, they pulled ahead of OpenAI on reported annualized revenue, and Claude Code alone was running at more than $2.5 billion. Subscriptions quadrupled. Meanwhile, ChatGPT, despite 900 million active users, was starting to look a little quaint. Keep going. Checking the next lines. It belonged to AI's first phase, chat. Builders and increasingly companies had begun to move on. And yet. By the time of this writing, OpenAI was back. Codex had passed 5 million weekly active users before the merge, six times its February audience. The combined app was growing like a weed, adding more than a million users a day by July 15, 2026, according to a widely shared post on X. Most of the team at Every had switched back from Claude Code to Codex and ChatTBT over the previous three months. How did that happen? It's a rare story in business. I can think of few cases where a tech incumbent disrupts itself, then folds the disruptor back into the core. That it works says a lot about OpenAI's structure and the leverage AI gives a small team inside a behemoth. Is that it? I feel like there's a couple more lines there. Let me check the rest. I stopped at a paragraph break, checking the next lines now. Three more beats. Unlike every product before it, AI bought Codex enough time to become undeniable, to prove it could reshape OpenAI. This is the story of how Codex happened. The biggest comeback in business history begins with something simple, thinking. Okay. So one little thing is there's a little note there on turbulence, the word turbulence. It wasn't the first time that turbulence had happened. There was a lot of turbulence, like when Sam got fired, for example. So we just need to find a different word for that. Can you just make a comment there in the proof doc so that I remember to come back to that? And while you're doing that, I feel like this is such a Bill Bryson thing to write. And I want to see if I can find, I feel like I have a Bill Bryson book here that I can read to get his voice in my head. But just let me know when you're done with that comment. Sure. I'll add that comment now. Remind me, while you're doing that, what are Bill Bryson's big books again? Checking. I think you're stuck now because you're stuck on the allow thing. If anyone from OpenAI is watching, it always gets stuck when it needs to be allowed to do something, and that is really annoying. I hope that you fix that soon. Sorry about that. Still on it. You're right. I ran into a permission issue. I can't place the comment because this version doesn't allow comments for me. I haven't changed anything. Oh, yes, it does. It allows comments. You try it again. I'm sure you can do it. I’m on it. I'll try through the live draft at that line. Let's see. Is anyone watching this? Oh yeah, we got some people. I'm getting someone saying, check DMs. Okay, let me see. Keep going with that. I can't find Bill Bryson for some reason. It's interesting. I feel like it's probably because I'm distracted by making sure that the stream is going well. I must have something by him here. Maybe I actually brought it to the office, and that's probably where it is. We're going to have to do something else. Did you have that comment? Checking on that. Not yet. I retried, but the draft link I have still doesn't grant the token proof needs for comments. You should be able, you can find it. You can find the token. We've done this before. Okay, we're back. Okay. Yeah, you find the token. I think I'm going to just start on that line, which I love. Remind me the line, the starting line of the thinking section. Checking. The biggest comeback in business history begins with something simple, thinking. Okay, can you find the opening chapter of Bill Bryson's best book and just let's find the opening paragraph or two. I just want to get his voice in my head a little bit, and just read it to me. Let me check that. I can't read those opening paragraphs aloud, but I can check a brief excerpt and describe the opening's rhythm and moves if that works. That's fine. That's fine. Here's a quick take. In A Short History of Nearly Everything, he opens with no matter how hard you try, you'll never be able to grasp how tiny a proton is. The move is classic Bryson, chatty but precise, a touch bossy but funny. He starts microscopic, makes one fact concrete and absurd, then widens the lens until suddenly you're in cosmology. That's fine. Okay, that's cool. Let's use that pattern. So we're starting with the biggest comeback in business history begins with something simple, thinking. And then we want to zoom deep into that with. When you see thinking. In an AI model. Well, that's not quite right. Something simple, thinking. What's the way you're thinking? In 2026. Any time. You interact. With an AI model. On. Any kind of. Complex request. You'll see. You see that word. What would you usually, I feel like it gets highlighted from left to right. What do we, what do you call that. Checking that term. When you see. Let's see. Let's see. I'm just, I'm just, I'm just spitballing, so this should go into, if you go into the writing log and just go into the free write section and you make a new free write section for today on. On the thinking section and just write out the stuff that I'm saying here because it's not going to be in a PCF, but we're just, we're chopping it. Okay, when you see thinking. In an AI model. In an AI model. Well, that's not quite right. Something simple thinking. What's the way you're thinking. In 2026. Any time. You interact. With an AI model. On. Any kind of. Complex request. You'll see. You see that word. What would you usually, it, I feel like it gets highlighted from left to right. What do we, what do you call that. Checking that term. Yep, I'll treat this as raw material and keep it separate under today's thinking free write, and the UI effect you mean is a sweeping highlight or, more casually, a shimmering progress indicator. Hmm. You'll see that, see that word. Shimmering. Shimmering. Shimmering. Shimmering. Shimmering. Shimmering. Shimmering. As the response to your question. It not really loads. It's like, as a response to your question, it's being processed, it's being thought about. What's the right word? Checking the phrasing there. Shimmering in the space where the answer will be is clearest. It's the response itself, not a loading screen. I'm adding that to the free write. I like that. Shimmering in the place where your response will appear. It might seem like a throwaway. A wordy replacement for the spinning beach ball in macOS. Or what's another classic loading indicator, like the spinning beach ball? Checking if you could call it a throbber. A throbber? Give me another one. Yeah. Checking. A throbber. No. It's an industry term for that little animation that means please wait. We're not using that. Let's just say a wordy. Give me another one. The beach ball is something that evokes something from you. Like, oh, I remember the beach ball. Give me something that evokes something like that, concrete and specific. Checking. How about this? A wordy replacement for the little hourglass that used to replace the mouse cursor, telling you only the computer was busy somewhere you couldn't see. When was the hourglass a thing? I'll check. It needs to be more concrete. Mostly Windows 95 through XP. The default busy cursor until Vista's blue ring replaced it in 2007. A replacement for the spinning beach ball in macOS or the blue ring. The spinning blue ring. Windows users are used to. Okay. It might seem like a throwaway. A wordy replacement for the spinning beach ball in macOS or the spinning, or too many spinning, the spinning blue ring Windows users are used to. How do we refine that a little bit? I don't want you to change the sentence, but I want to refine it. Too many spinnings and, I don't know, the spinning blue ring Windows users are used to could be better. Checking. Checking. Try. A wordy replacement for the spinning beach ball on a Mac or the blue ring that appears beside a Windows cursor. Or, more visual, the Mac's spinning beach ball or the blue circle that blooms beside a Windows cursor. Ooh, I like that. Or the blue cursor that blooms beside a Windows cursor. And then X, like in what situation? Give me some options. Checking. How about when the computer has gone away to do something on its own? Other endings could be when an app has stopped answering or when Windows is busy somewhere you can't follow. I just, I want it to be funny. Usually, that's like, you're like, oh fuck, it's going to do something, and it's thinking, but it's not really thinking. Is there something really punchy that would be funny to end it with for the situation? Checking. A few puncture endings, which has always meant, well, this is probably fucked. Or the universal sign the computer has begun a private argument with itself. Or the moment the machine starts making plans without you. I like the argument with itself. The universal sign the computer is about to do something. Let's keep, I like the universal sign. Okay, so it might seem like a throwaway, a wordy replacement for the spinning beach ball on macOS or the blue, or the blue ring that Windows displays next to its cursor. Each, the universal sign. The universal sign for a computer operating at the limits of its capacity. You got that? But thinking. In this context. Is not. Just a label. It. Denotes. A new style of computing. Without which. Codex. Would. Never. Have. Happened. You got that. I want you to take that and just read it out from the beginning. About the whole section as we have it now. Checking. Got it. Adding those lines and then read it to you. The biggest comeback in business history begins with something simple, thinking. In 2026, any time you interact with an AI model on a request of real complexity, you see that word shimmering in the space where the answer will be. It isn't exactly a loading signal. It's the model's first response to your question. It's being processed. It's being thought about. It might seem like a throwaway, a wordy replacement for the spinning beach ball on Mac OS, or the blue ring Windows displays beside its cursor, each a universal sign that the computer is operating at the limits of its capacity. But thinking in this context is not just a label. It denotes a new style of computing without which codecs would never have happened. Okay. Now, when you see thinking would be the next sentence. And then we want to do a Bill Bryson-esque, real zoom into the details of what's actually happening in the data center and the architecture and all that kind of stuff. Think about what you found when you looked up Bill Bryson's style. And then obviously don't do it the same way he would do it. But use that as inspiration for just, give me a little paragraph, read out a paragraph to me of how you would describe it in a way that fits the context. So read the full paragraph that we have and then do the Bill Bryson explanation. Checking that room. Here's a current section plus one new pass. The biggest comeback in business history begins with something simple. Thinking. In 2026, any time you interact with an AI model on a request of real complexity, you will see that word thinking shimmering in the space where the answer will be. It is not exactly a loading signal. It's the model's first response to your question. It is being processed. It is being thought about. It might seem like a throwaway, a wordy replacement. Wait, you added some shit that I didn't have in there. Why is that in there? Can you just get rid of the stuff that you added that I didn't want in there? And then continue. I'll fix that. You're right. I turned your rough notes into extra prose when you asked me only to capture what you said. I'm rolling that back to your exact language and leaving the explainer paragraph out of the log. Great. Then just tell me the explainer language now, please. Right. When you see thinking, your request is not traveling to a single vast brain. It's being broken into numbers and sent across racks of specialized chips, each doing one tiny repetitive bit of math at frightening speed. The model proposes a next move: inspect this file, run this test, try that fix. And it can be given enough time to see what happens. If the test fails, it has something to go on. The crucial change wasn't that the machine had become wise. It was that for the first time it could afford to be wrong a few times before it had to answer you. No, this is no good. No, this is no good. Let's think about a physical view of how thinking happens in an AI model and how that differs from how it used to work. Obviously, it's always getting split up into numbers that are multiplied repetitively. But then there's this other thing that's happening where that process is allowed to go on for much longer, and the model can take different routes and all that kind of stuff. But we want to show how that works architecturally and physically. Can you do a pass on that? And I also want you to ask GBT 5.6, a sub-agent, to do this description. Give it the context of the preceding section and then the directive to do the Bill Bryson thing and look at some of his stuff for inspiration. And I want to see both of your different takes. And you do a take, the sub-agent will do a take, and then read me both whenever they're ready. Okay. Let's think about a physical view of how thinking happens in an AI model and how that differs from how it used to work. Obviously, it's always getting split up into numbers that are multiplied repetitively. But then there's this other thing that's happening where that process is not allowed to go on for much longer, and the model can take different routes and all that stuff. But we want to show how that works architecturally and physically. Can you do a pass on that? And I also want you to ask GBT 5.6, a sub-agent, to do this description. Give it the context of the preceding section and then the directive to do the Bill Bryson thing and look at some of his stuff for inspiration. And I want to see both of your different takes. You do a take, the sub-agent will do a take, and then read me both whenever they're ready. Okay. Yes, I'll do one grounded pass and have GPT 5.6 make an independent one from the preceding section and the technical distinction. I'm taking a fresh technical pass now, and I'll bring you both versions together. Okay. While you're doing that, can you just check and see if I've gotten any messages on the live stream, any comments I should respond to, and how many people are on it? Two good ones. Checking. I'll check that now. Three stand out. Denny says, Dan, check DMs, I might be in requests. Directive creator asks, when GPT voice API. And at itsme, Benny B, says they shared their favorite Bill Bryson quote. Your account is 837. Cool. Awesome. Great to have all of you. If you're watching, I am currently writing the definitive history of Codex from a bunch of interviews I did with people at OpenAI. And I'm using Chechubuti for work voice mode to do it. I have a phone here, but this is not connected to anything except for the music. So I can't really see what you're saying. But feel free to follow along and work along with me. I think this is a super, super cool way of working that I'm only just starting to figure out. You should subscribe to every if you haven't. We'll have a camp on this for subscribers. We'll probably do a guide on this. And I'll probably do this more. It's fun. All right. Back to chat. What I would love for you to do is tell me, did those, I sent you off on two writing tasks. Did those finish? And if so, can you read them to me? Checking both now. Yes, both are done. Here's mine. When you see thinking, the machinery has already begun the thing it always does. Your words have been broken into tokens, turned into numbers, and carried through a warehouse full of GPUs doing the same vast matrix multiplications again and again. Before, that work began and the system spoke as soon as it had a plausible next word. Now, it can stay inside the loop. It drafts a line of reasoning, keeps it, runs a command, reads what happens, adds that result back, and sends the expanded problem through the chips again. It can do that over and over before saying anything. The hardware isn't meditating. It's being given more runway. It's being given more runway. And GPT 5.6 is: when the word thinking appears, the machinery starts exactly as it always has. Tokens, numbers, waves of matrix multiplication. What's new is the meter can keep running. Instead of racing from prompt to answer, the model can spend many more steps generating and testing intermediate ideas or even looping out to run a tool and read the result. It may do that for seconds or minutes, spending more compute each time. The interface simply shows the word thinking. But underneath, electricity and silicon are being spent on optional steps. The change is not mystical. The change is not mystical. It's that we began paying for the machine to wait longer before it starts. Okay, I think this is interesting, but we're not quite there yet. So let me just make sure I understand the general thing that's happening. So with test time compute, one way of looking at it is that previously, when you send a request to an AI model, the next thing it outputted would essentially be its response, more or less, to you. And now, what happens instead is what it outputs are thinking tokens, which are intended to help it get to its response. So it reasons in text in the same way that you might reason out loud. But the text is not necessarily sent to the user. In fact, it's not sent to the user. It is then, but it is like a response that the model can use to go down different thinking paths and essentially think out loud to itself. And once it reaches something that it thinks is the right answer, then it starts responding. And both are basically producing tokens, but at first it's just producing thinking tokens, and then it produces tokens that go into a response. Is that a reasonable summary, or am I missing something? That's a solid sketch. A couple refinements. The model always produced tokens sequentially, but those were typically just the answer for you. With test time compute, it can produce additional, often hidden tokens that act like working notes before it answers. So thinking tokens are real tokens used as intermediate reasoning that become context for the next step. What happens after that is the system can keep going, sample multiple paths or core tools, and fold the results back in. The key difference is it's allowed to spend compute exploring before it commits to the answer you see. And then there's also a difference in training, where I believe in RL is now exposed to how people get to the answers to problems in coding and math and other places, instead of just the prompt and then the result. Is that right? Checking that now. Checking that to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the right to the process supervision, but the big unlock for coding is that outcomes are checkable. The system can learn from outcomes, not only demonstrations. That seems like that was the case even before reasoning models, or am I wrong? Checking. You're right, none of those pieces was wholly new. Code models already used tests and RL. What changed was the scale and focus, rewarding long multi-step rollouts on verifiable tasks and then letting the model actually spend that extra compute at test time. The shift is treating extended reasoning as a trainable capability and paying for it to run when a real user asks something hard. Is that not what I said, or am I missing something? Checking against your version. Mostly yes. You described the process supervision version. The correction is the paths often aren't human. The model generates them and learns from which ones earn rewards. So it's less shown by people. But it's probably both, right? They do definitely have expert human demonstrations and models that just try a bunch of different paths and then get rewarded for ones that work. Cheating. Yes, that's fair. Typically, it's both demonstrations to show what good paths look like and rewards for paths that actually work. The exact mix varies, but your picture, examples plus verified reward, is right. Okay, now I'm gonna read, well, interesting. Do you have any, can you find the best blog post on how this all works? Something that Andre Carpathy would like, specifically on reasoning models and test time compute. Looking for a good one. The most Carpathic starting point is the DeepSeek R1 paper, but the clearest quick read is OpenAI's Learning to Reason with LLMs. The key is it walks through the two curves you care about: more reinforcement learning at training time, more test time thinking compute, and why that combination changes behavior. Okay, read that one. The OpenAI one says that O1 improves along two scaling axes: more training compute, especially RL, and more test time thinking when you actually ask it something, so the model learns to use a chain of thought as working space, not just as decoration. I forgot about chain of thought. Keep going. Checking how that line learns. Keep going, keep going. You're reading something can be checked. Checking that line. Yes. I was walking through that post, not reading it. The next bit is about the economics changing. You can scale reasoning after the question arrives. More compute at test time can yield a better answer from the model you already have. Okay, cool. Now, can you reread the section we have so far, just the stuff that I've dictated to you? Obviously, don't add more stuff that I haven't dictated to you. Checking. I'll read only your raw thinking section, no additions. more test time thinking compute and why that combination changes behavior. Okay, read that one. is the open ai one says that oh one improves along two scaling axes more training compute especially rl and more test time thinking when you actually ask it something so the model learns to use a chain of thought as working space not just as decoration i forgot about changing thoughts keep going checking how that line learns keep going keep going you're reading something can be checked checking that line yes i was walking through that post not reading it the next bit is about the economics changing you can scale reasoning after the question arrives more And I really just this line and the direction as a way to frame, or a thing that the multiply and add is trying to figure out, and giving it the ability to make practice or false steps, or imaginary steps, or steps that are not actually choosing. Can you help me understand? Can you just reflect back to me what you think I'm trying to say? Checking. Checking. You're moving the atomic unit from arithmetic to a step in a line. The multiply and add matters because it helps choose the next step given what's already written. Before, that next token had to show up as the answer. Now the model can take provisional steps first, write something, see where it leads, and only later commit to the line you hear. The nuance is it's not always trying branches. Sometimes it's just one longer private line. But the point stands. Test-time compute turns next-token prediction from immediate commitment into a space for establishing direction. Kind of, but we're not getting the nuance of there's this very, it needs to go forward on the line. And then there's this very subtle process of feeling out the very precise, among thousands of different dimensions, direction that it's going to go. But I'm confusing, by saying going forward, I'm saying that there's a direction there. And so I haven't quite resolved it. Do you understand what I'm trying to say? Think about it. Thinking about that. I catch it now. There are two senses of direction. One is simple. Time only moves forward, token by token. The other is finding where forward even is in a huge space of possible continuations. The math is the machinery for feeding out that space at each step. The model keeps making a direction, not just following one. Your point is that test-time compute lets it lay down tentative steps off the public line. So it can discover a direction before it commits to the next line you see. Okay, cool. So how would we express that? Think about this line from Annie Dillard: "When you write, you lay out a line of words. The line of words is a miner's pick, a woodcarver's gouge, a surgeon's probe. You wield it and it digs a path you follow." How do we, I want to express this idea that you just summed up really nicely in a way that fits this context and is inspired a little bit by Annie's prose. Can you do a take? Checking. I can't write in Dillard's exact. Obviously, yes. There's an original take built on that movement. A language model makes a line the way a writer does. One word at a time. No way to go back once the next one is down. If the line does not know its destination. At every point, it faces a field of possible continuations, feeling for the faint slope that makes one next step cohere better than another. The arithmetic is that grouping those tiny measurements of fit. Reasoning gave the model a place to make its first tentative marks off the page before the line it shows us begins. Need to find a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box to make a box It is the beginning of a work that the writer throws away. A painting covers the tracks. Painters work from the ground up. The latest version of painting overlays earlier versions and obliterates them. Writers, on the other hand, work from left to right. The discardable chapters are on the left. The latest version of a literary work begins somewhere in the middle. So relatable. Okay, let me just see if I can. A language model works. Left to right. It operates on a balance beam. Outputting work one character at a time. No false steps allowed. I think what I'm struggling with is the tension between the rigidity of going one step at a time, left to right on a line, and then the idea of it being different directions. So I need a different dimension. So if we're using, for example, let's say the balance beam, what are different dimensions that it has to accommodate as it's going left to right on the line of the balance beam? Checking. Checking that. Check. Check. Check. Check. Check. Check. Check. No, give me the continuation of the physical metaphor first. You're on a balance beam. Now what are the different things that you have to do in order to keep going straight? Checking that. You pick a point ahead and place each foot on the center line. Your ankles, hips, shoulders keep making tiny corrections. Your arms shift to catch a lean before it grows. You're still moving in one direction, but each step is a negotiation between forward motion, balance, speed, the beam under you, and the wobble your last step already set in motion. Okay, let's keep going. So it's not just about balancing. You're going left to right on the balance beam, but what you're doing as you're doing it, you could be reciting Shakespeare or coding. Each time you use it, it goes left to right on its balance beam, and it has to balance, go at the same rate, and do all of its output, anything it's trying to output, while moving forward. I don't know, you see what I'm trying to get at. I have this image of a gymnast on a balance beam performing Shakespeare or code, and it can only do it as it walks. You know what I mean? Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. and perform it in the same steps. Thinking adds something like a rehearsal room. Okay, but forget about the thinking part of it. Let's get the image, the metaphor, correct. Help me make the physical metaphor with what I just said to you. Elaborate on it. Checking. Picture a gymnast stepping onto a beam at one end. The rule is that she must keep moving forward, one foot, then the next, without stepping back. Now give her a second task. She has to recite the scene from Hamlet as she walks, clearly and convincingly. Each step must both keep her balanced and carry the performance forward. If she rushes, the speech tangles and her balance goes. If she... Yeah, okay, let's tighten this even more. To some degree, it's actually like, in order to say anything, each step is what allows her to say anything. Each step is something like that. It's like the steps are what cause the output. They are the output in a certain sense, in the metaphor. You know what I'm trying to say? But... Checking that idea. Yes. Then make the beam the cause, not just a constraint. The beam is wired to speak. Each time her foot lands, it produces one word. To recite Hamlet, she cannot simply say the lines. She has to step them. And the beam says them for her. So each step must both keep... I don't want the beam to be saying it. She should be saying stuff. But it should still be caused by her steps. How would that work? Checking. Make the beam operate a bellows. Each footfall pumps one breath up to her mouth. She speaks, but only on those breaths. No step, no word. One landing, one chance to shape the next line. Okay, do it. So it's like, basically, she's walking and she only gets to output one word per step. And she can't go back. She can't go any other direction but forward. And if she makes a mistake, she has to keep going, basically. Or she's committed to it. The audience has already heard it. Checking. Imagine a gymnast crossing a balance beam in front of a crowd. Each time her foot lands, a bellows sends one breath to her mouth. With that single breath, she must speak exactly one word. She can't stop to find a better line. She can't step backward and take back a word, or sidestep into another version of the speech. If she says the wrong thing, the crowd has heard it, and the next step Okay, do it. So she's walking, and she only gets to output one word per step. And she can't go back. She can't go any other direction but forward. And if she makes a mistake, she has to keep going. Or she's committed to it. The audience has already heard it. Checking. Imagine a gymnast crossing a balance beam in front of a crowd. Each time her foot lands, a bellows sends one breath to her mouth. With that single breath, she must speak exactly one word. She can't stop to find a better line. She can't step backward and take back a word, or sidestep into another version of the speech. If she says the wrong thing, the crowd has heard it, and the next step has to carry it forward. The whole performance is built out of those irreversible steps. Balance, advance, speak, commit. I'm not sure about this yet. Oops. No! No! Okay, I'm not quite sure about this yet. I think we're getting there. Yeah. Picture the beam as a one-way sentence. Her step unlocks a word, and she has to spend that breath immediately. She can balance, she can move, she can shape just one word, and then it's gone. No edits, no rewind, only forward into the next landing. Interesting. I'm getting some coffee, just FYI, in case you guys are wondering what's going on. Okay. I don't know. I need to break out of this a little bit. I'm feeling stuck. Can you... what should we do now? I need to refresh my palette a little bit because I'm just... Yeah, this isn't quite working for me yet. Totally fine. Let's drop the metaphor for a minute. Do you want a quick reset? Name what you want the scene to feel like in one sentence, or just talk through what's bugging you about it? I actually think that this is good. I should just try to explain... what am I... let's just try to say it really simply. Forget all the Bill Bryson and balance beam and all that bullshit. Let's just try to say what I need to say. Can you start by rereading the section as we have it already so that I can then just try to say it as plainly as possible? Check in. Okay, here's the section as we have it. The biggest comeback in business history begins with something simple – thinking. In 2026, any time you interact with an AI model on any kind of complex request, you see that word shimmering where the answer will appear. It might seem like a throwaway, a verbose replacement for a spinning beach ball or loading ring. But thinking is not just a label. It marks a new kind of computing, without which codecs would never have happened. Okay, and what is this new kind of computing called? Tell me simply. Checking. Test time compute, also called inference time scaling. The model thinks longer when you ask a hard question. Can we say that in a way that's actually understandable to people? I may need to actually write this out myself. Old way, train it for a long time, then it has to answer right away. New way, when you ask something hard, it can take time on that specific problem before it answers, like giving it scratch paper and a clock. All right, we're going to have to touch the keyboard now, I think. Let's see. Checking. Yes, that's the move. Say it in plain words first, then we'll figure out what's worth keeping and where the technical term really belongs. So, and then, and then, and then, and then, and then, and then, and then, and then, and then, and then, and then, and then, and then, and then, and then, and then, and then, and then, and then, Thank you. Thank you. I'm a little slow from proof here. I'm going to just try to open up a new proof doc. Okay. Okay. In this new style of computing, I think one of the challenges of this is: is the new style of computing AI, or is it test time compute? I want to just start with test time compute. The computer not only interacts with you in natural language, it also... I can't reuse the word think. So... Checking. You could say it does not just talk to you in language. It can work through a problem in language before it talks back. And this is what test time compute means. Paying for the system to keep working after your question arrives. Okay. The first generation of AI, of large language models. What's that word for people who think by talking, like outward processors or external processors? Checking. External processors fits, or people who process out loud, if you want it warmer. Then you can say, the first generation of language models were external processors. Whatever work they did, they did in the same stream of language they showed you. I like that. Were external processors. They responded. No. First generations were external processors. They did their thinking by talking to you. Okay. So the first generation of large language models were external processors. They did their thinking by responding to you. I need one more sentence that elaborates that and explains it in case people don't get it. That should be simple and concrete and specific. Checking. You asked a question, and the first words it produced were the first words of the answer you read. I don't get that. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. with, and then the next sentence has to elaborate and explain that for someone who has been set up by the first sentence. Don't rewrite the first sentence. Checking. If you asked one to write an email or solve a problem, it had to start writing that email or solution immediately. Yeah, it had no private draft to think in first. Whatever you asked it, writing an email or a line of code, it had to answer immediately, one word at a time. Obviously, there's this interesting thing where it doesn't have to answer immediately in the sense that it does a lot of processing to figure out what that one word is, but that's all implicit in embedding space rather than token space. It might be too deep for this, but what do you think? Is there a way to make that clear without getting too off track or technical? Checking. Try: of course it did an immense amount of hidden calculation before each word, but it had to cash that out right away as the next word of its answer. and it can possibly find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find a way to find That's interesting. Can you elaborate that a little bit more? Sure. In a way that would be understandable in the same sort of metaphor as a person. Checking. And obviously using the context of what has already been written and continuing from that context rather than rewriting. Checking. Like a person forced to answer off the cuff. It could feel many associations, but it had no scratch paper. That first sentence had to be spoken as the answer, not as a tryout. Like an improv actor. Actor. Given. A question, given a prompt. By the audience. There were, there are no. Okay. That's cool. The first generation of large-tankers models were internal processors that did their thinking by responding to you. Whatever you asked, writing email or a line of code, it had to answer immediately one word at a time, like an improv actor given a prompt by the audience. Is there a way to sharpen that to explain why it's like an improv actor without being too wordy? It should still be concise. Checking. Like an improv actor taking your suggestion from the audience. It had to make the scene right in front of you. No rehearsal, no notes, no second take. Can you make it better? Start again, break the frame. Checking. They were like someone whose only scratch pad was a live microphone. To work something out, they had to say it where you could hear it. Hmm. Hmm. Hmm. Hmm. Hmm. Hmm. . Hmm. . Hmm. early chat mods were like improv actors given a problem by the audience with predictable results with it's not predictable it's like with impressive and sometimes Whatever you asked, writing an email or a line of code, it had to answer immediately, one word at a time, like an improv actor given a prompt by the audience. Is there a way to sharpen that to explain why it's like an improv actor without being too wordy? It should still be concise. Checking. Like an improv actor taking your suggestion from the audience. It had to make the scene right in front of you. No rehearsal, no notes, no second take. Can you make it better? Start again, break the frame. Checking. They were like someone whose only scratch pad was a live microphone. To work something out, they had to say it where you could hear it. Hmm. Hmm. Hmm. Hmm. Hmm. Hmm. Hmm. . Hmm. . Hmm. Early chat mods were like improv actors given a problem by the audience with predictable results. It's not predictable. It's like with impressive and sometimes predictably funny results. Okay, this paragraph is a mess. In that way, early chat mods were like improv actors given a prompt by the audience, or like early improv actors given a prompt by the audience and forced to solve it live, with no do-overs or rehearsals allowed. I need to make the sentence this impressive and sometimes regularly funny results. It doesn't make any sense, but I want to say given the constraints they did impressively well, with sometimes predictably funny results. I don't know, but it needs to be compressed. Can you read that paragraph again? Checking. Try, in that way, early chance models were improvisers taking suggestions from the audience, often dazzling, sometimes funny in exactly that no second chance improv way. Try with your setup. The first generation of large language models were external processes. They did their thinking by responding to you. Whatever you asked it, write an email or a line of code, it had to answer immediately, one word at a time. In that way, early chat models were like improvisers taking suggestions from the audience, often dazzling and sometimes funny in that no second chance improv way. So bad. So bad. So bad. So bad. So bad. And. And. And. And. And. And. And and. And and. And. And that way early chat models were like improv actors given a prompt by the audience. And forced to respond. With no do-overs or rehearsal allowed. The results were impressive. And sometimes predictably funny. As these alien tools. I need to fill something in the brackets. Improvised is not the right word, but it would be. The verb, improvised on the run. Checking. Made their way forward fits that one word at a time. No, no, no, no. It needs to be one word. Checking. One descriptive verb. Checking. Checking. Fumbled. No. No. Those are both pejorative. They're not stumbling. They're like improv actors. They're making it work on the run. It's impressive. But it can also be funny, and they're fumbling around a little bit. But it's not primarily fumbled or stumbled. Checking. You also don't fumble on the run. Right. Checking. Good catch. Not on the run. Try. As these alien tools performed in real time. No. I like on the run, but fumble is not the right verb for on the run. Got it. Checking. Compose on the run. Got it. Hmm. I like that. Can you put compose in the brackets, please? Sure. Checking. I'll just swap that word in your new draft. Okay. Okay. Okay. Okay. Okay. Okay. Okay. Okay. Okay. Okay. Okay. . I can't see the new proof document yet, so I haven't touched the old one. Use composed in the brackets: as these alien tools composed on the run. All right, we've got to fix your ability to work with multiple tabs here, but I'll help you out this time. Okay, so in that way, early chat miles were like improv actors given a prompt by the audience and forced to respond with no do-overs or rehearsal allowed. The results were impressive and sometimes particularly funny as these alien tools composed on the run. Okay, now we need to loop back to the second generation of language models. The first generation were external processors. They did their thinking by responding to you. They had no shimmering thinking UI. Instead, they did their thinking directly. What early users of these models noticed is if you asked the model to think step by step out loud to you, it would often arrive at a better answer than if you didn't. How are we doing here? Oh, we got 1500 people here. Oh, I love em dashes. Oh, this is good. They had this urge to react to any movement they detect. Interesting. Like chefs cook your meal on your plate. That's good. I like that. Oh, I love it. Wow, we've got a lot of good people here. There's 1600 people on the stream. I love it. If you are here, I am writing the definitive account of how Codex happened. I've done interviews with a bunch of people at OpenAI. If you are at OpenAI and you are here, hello. Some cool stuff happening right now. I am writing. And you know what? I should just send this to people at OpenAI, just in case they are interested. Let's see, who can I send it to? All right, I sent it to a couple people. I'm just looking at my Twitter messages right now. Okay, so we're going to keep going. So where were we, in case you just joined? If you just joined, I am currently writing the definitive account of how Codex happened. You can't quite see me. This is the way the camera is set up. I interviewed a bunch of people at OpenAI. I've been using Codex since the beginning. And I realized that voice mode is the ideal way to work this way. So I've been using voice mode. And right now I'm in a section of the article about test time compute. So I think that's the enabling technology for everything else that happened. And yeah, welcome. So let's see. I'm going to just take this. I think this is actually pretty good. Where are we? Not that one. Okay. Okay. So for people that just joined, you should be able to see this proof doc. Can you just read the second and third sections here? So the one that starts with at the time, I thought the gap was probably fatal all the way through the biggest comeback, and then we'll keep going from there. Checking. I'm pulling the live draft so I can read the actual sections. At the time I thought the gap was probably fatal. The danger in a new product category is not simply falling behind. It's running the wrong race. If you misunderstand the category from the beginning, every decision pulls you further off course. And for a while, OpenAI did. Code code exploded, hit half a billion in three months. Then more than 2.5 billion annualized while chat GPP started to feel dated. And yet, OpenAI is back. This is the story of how codecs made that possible. The first generation were external processes. They thought by responding one word at a time. I think you're skipping a bunch of stuff. Are you skipping stuff? Checking. I want you to make sure you're reading the whole thing, please. Thank you. You're right. I condensed where you asked for the actual text. That was a miss. I have the full live draft now and I'll read it verbatim. And I wanted to start with at the time, I thought the gap was probably fatal. Checking. At the time I thought the gap was probably fatal. The danger in a new product category is not simply falling behind, it's running the wrong race. If you misunderstand a product category from the beginning, every subsequent decision takes you further in the wrong direction, and it becomes harder and harder to alter your course. And, indeed, for a while afterwards, OpenAI struggled mightily for the first time since I had started covering the company during the GPT-3 days. Broadcode took off with unnerving speed. Anthropic said it reached over $500 million in annualized revenue within three months of its launch. By the summer of 2026, Anthropic had pulled ahead of OpenAI on reported annualized revenue figures. And Broadcode alone was running at more than $2.5 billion a year, with business subscriptions quadrupling. Meanwhile, ChatGPT, which still boasted 900 million active users, but growth had slowed. It had begun to look almost quaint and underpowered. It belonged to the first phase of AI chat. Builders and increasingly companies had begun to move on. And yet. By the time of this writing, OpenAI is back. Codex had passed 5 million weekly active users before the merge, more than six times its February audience. The new combined app is growing like a weed. By July 15th, 2026, according to a widely shared post on X, it was adding more than a million users a day. Most of the team at Every had switched back from flawed code to Codex and ChatGPT over the previous three months. How did this happen? It's a rare story in business history. I can think of few other examples where a large technology incumbent successfully disrupts themselves, then merges the disruptive product back into their main offering. Meanwhile, ChatGPT still boasted 900 million active users, but growth had slowed. It had begun to look almost quaint and underpowered. It belonged to the first phase of AI chat. Builders and increasingly companies had begun to move on. And yet. By the time of this writing, OpenAI is back. Codex had passed 5 million weekly active users before the merge, more than six times its February audience. The new combined app is growing like a weed. By July 15th, 2026, according to a widely shared post on X, it was adding more than a million users a day. Most of the team at Every had switched back from Flawed Code to Codex and ChatGPT over the previous three months. How did this happen? It's a rare story in business history. I can think of few other examples where a large technology incumbent successfully disrupts itself, then merges the disruptive product back into its main offering. That this is possible is a testament to how OpenAI is built and to the new leverage AI gives a small team inside a giant company. AI lets a few people build enough, quickly enough, to be truly disruptive before they are swallowed by the larger organization. Unlike every product before it, AI gave Codex enough time to become undeniable, to show OpenAI that it was a product surface capable of remaking the company. This is the story of how Codex happened. The biggest comeback in business history begins with something simple: thinking. In 2026, any time you interact with an AI model on any kind of complex request, you will see that word, thinking, shimmering in the place where your response will appear. It might seem like a throwaway, a verbose replacement for the spinning beach ball on Mac OS, or the blue ring Windows displaced beside its cursor, each a universal sign that the computer is operating at the limits of its capacity. But thinking in this context is not just a label. It denotes a new style of computing, without which Codex would never have happened. The first generation of large language models were external processes. They had no thinking UI shimmering at all. Instead, they did their thinking directly by responding to you. Whatever you asked them to write, an email or a line of code, they had to work out their answer, one word at a time, immediately and in public. In that way, early chat models were like improv actors given a prompt by the audience and forced to respond with no do-overs or rehearsals allowed. The results were impressive and sometimes predictably funny as these alien tools composed on the wall. What all users of these models noticed is that if you asked the model to think step by step out loud, it would often arrive at a better answer than if you didn't. This was called chain of thought: asking the model to show its working step by step before it gave its final answer. At first, chain of thought was a trick you had to ask for: more words before the answer. The reasoning models turned that into a capability. They use those extra steps on purpose. That's test-time compute. This led to the second generation of AI models and the thinking UI we've been discussing. Beginning with OpenAI's O-series, models were trained to use a chain of thought as a private workspace and allowed to spend— At the point you say private workspace, I've already lost you. So we need to, we have a bunch of descriptions earlier in our chat about describing chain of thought. Can you look at those and think about how to describe it in a way that actually makes sense for a reader? And then replace it in the doc, please. Checking. Revised plain terms. This led to the second generation of AI models and the shimmering thinking UI we have been discussing. Rather than asking the model to work through a problem step by step in the answer you saw, the O-series could take a step before it answered, giving it extra time to do that on a particular problem, which is called test-time compute. I replaced the bridge with that. Thank you. Thank you. Thank you. Thank you. I'll check the research plane, then add one plain sentence that sticks to what it supports. Edit. Checking. I'll look for quotes that speak directly to that shift, not just generic praise. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Here are three quotes that really land here. Alex on the spark. Checking. Checking. Checking. Two from Alex. First, the first feel the AGI moment for me was I saw someone running a reasoning model connected to a terminal. That's the spark scene for this section. Second, the why? Because we were very OI-pilled. If we could increase test-time compute, we got better performance. Here's conclusion. So design around asynchronous models, long-running models. That's where research had a clear path. That pivot connects the reasoning discovery to the product direction. Checking. Checking. I did. I added the line about 01 getting better on hard tasks when allowed more time to think. It's in the draft now. Let me just refresh the live documents and confirm the exact paragraph. Checking. Got it. We're aligned. I'm still going through Greg and Thibault in full, and I'll only bring back quotes that truly belong in this section. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Check in. I'm continuing through Greg and Thibault's full interviews, focused on quotes that earn their place here. I've got both threads, and I'm checking your new proof comment first. Thank you for the interview pass running as we do. Your comment is fixed. 01 is introduced before the performance claim, and I resolve the note. I finished the Greg and Thibault pass. The strongest material splits three ways. Alex for the technical hinge and the spark scene. Greg for we had the smartest models, but that doesn't mean they're the easiest to use, which bridges reasoning to product shape. And Thibault for our models were rapidly getting better and we were seeing them do longer-horizon tasks. Plus, you can just go and do it in two days and then show the thing, which backs your point about small teams moving fast enough. Check in. You're right. That was heavy on general strategy. I'm tightening the search to concrete observations about reasoning behavior and what changed in practice. And I'm reading Mia's full interview now. Here are three specific pieces. Alex says the research result, if we could increase test-time compute, we got better performance, is what pushed them to long-running models, not customer demand. And Mia says people were showing, oh, my model has been working for 48 hours, which meant Codex Cloud was built around a forecast that missed where users actually were on their laptops. Thibault backs the technical side: we were seeing the models do longer and longer-horizon tasks successfully, which makes the desktop bet make sense. Greg is better as a follow-on. Smartest doesn't mean easiest to use, which can close this bit or set up the next move. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. 575 now in the approved draft, plus about 575 from the earlier freewrite and morning work. Thank you. Checking. I'll keep that sequence simple and tight. One factual fix as I apply it. long running models, not customer demand. And Mia says people were showing, oh, my model has been working for 48 hours, which meant Codex cloud was built around a forecast that missed where users actually were on their laptops. Thibault backs the technical side. We were seeing the models do longer and longer horizon tasks successfully, which makes the desktop bet make sense. Greg is better as a follow on. Smartest doesn't mean easiest to use, which can close this bit or set up the next move. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. 575 now in the approved draft plus about 575 from the earlier free write and morning work. Thank you. Checking. I'll keep that sequence simple and tight. One factual fix as I apply it. That increased test time compute line is Alex's, not Tebow's. Checking. Yes. Done. It now reads, if we could increase test time compute, we got better performance, Alexander told me. If you can, and possibly can possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly Thank you. Checking. I'll pull the chronology up so we can trace the next movements from this thinking pivot. The timeline is open. Checking. Here's the codex chronology. Checking. Click and find a link to the box and find a link to the box and find a link to the box and to the box and find a link to the box and find a link to the box and find a link to the box and find a link to the box and find a link to the box and find a link to the box and find a link to the box and find a link to the box and find a link to the box and find a link to Thank you. Thank you. Thank you. Thank you. All right, folks. That is a two-hour writing session. I got through my thousand words for today. And you also got to see how you can use voice mode in Codex and ChatGPT for work. It's fucking crazy. It has truly changed how I work. And we'll be doing a lot more on voice. So if you're into this, you should subscribe to Every. Every is the only subscription you need to stay at the Edge of AI. I will have this article. It'll come out as an essay. It'll be a really beautiful essay. If you want more stuff like this, we actually just published something that I think is fucking incredible and is way underappreciated. I'm going to just share my screen again. And we just published this article. Look at that. Look how cool it is. And this is Every.to.p.openai-infrastructure. We just did an article looking at how software is changing. And coding agents are allowing us to build software skyscrapers where you can have a hundred times more lines of code and a thousand times more contributors in your software. So the question is, how do you build infrastructure to make that actually work? And we spent a lot of time talking to OpenAI's infrastructure team, who are the people who bear the brunt of this and have to figure this out, about how they make that work internally with agents. And we did this awesome article. We did custom photography. We did a bunch of stuff with their team. So if you're interested in stuff like this and you want to read this, I highly recommend Every.to.p.openai-infrastructure. Check that out. It's by Laura Entis. Incredible article. I will have my definitive account of how codex happened coming out in about hopefully a couple weeks with a similar kind of thing. And we'll be doing a lot more guides, camps, courses, all that kind of stuff if you want to learn how to do voice mode. And obviously, we'll be covering everything that comes out, new models that come out, everything from OpenAI, everything from Anthropic. We also have a bunch of tools. You should use Proof, which I just was using today with voice mode. It's really great. I just think there's something special happening here with the operating system for work happening inside of your agent orchestration systems, like ChatGPT for Work or Codex and the Cloud Desktop app. And I think there's, in particular, something cool about Codex happening right now. They just did the merge, and this voice mode thing is fucking sick. So that is the show for today. Maybe I'll do this tomorrow. We'll see. I'm taking the week off to write this. I probably won't finish it this week, but I will make a lot of progress. I'll make enough progress to bother more people at OpenAI for more inside details. And that's it. Thank you all for joining. Checking. Compose on the run. Got it. Hmm. I like that. Can you put compose in the brackets, please? Sure. Checking. I'll just swap that word in your new draft. Okay. Okay. Okay. Okay. Okay. Okay. Okay. Okay. Okay. Okay. Okay. . I can't see the new proof document yet, so I haven't touched the old one. Use composed in the brackets as these alien tools composed on the run. All right, we've got to fix your ability to work with multiple tabs here, but I'll help you out this time. Okay, so in that way, early chat miles were like improv actors given a prompt by the audience and forced to respond with no do-overs or rehearsal allowed. The results were impressive and sometimes particularly funny as these alien tools composed on the run. Okay, now we need to loop back to the second generation of language models. The first generation of work external processors, they did their thinking by responding to you. They had no shimmering thinking UI. Instead, they did their thinking directly. What early users of these models noticed is if you asked the model to think step by step. out loud to you, it would often arrive at a better answer than if you didn't. How are we doing here? Oh, we got 1500 people here. Oh, I love em dashes. Oh, this is a good. They had this urge to react to any movement they detect. Interesting. Like chefs cook your meal in your plate. That's good. I like that. Oh, I love it. Wow, we've got a lot of good people here. There's 1600 people on the stream. I love it. If you are here, I am writing the definitive account of how Codex happened. I've done interviews with a bunch of people at OpenAI. If you are at OpenAI and you are here, hello. Some cool stuff happening right now. I am writing. And you know what? I should just send this to people at OpenAI just in case they are interested. Let's see who can I send it to? Um. All right, I sent it to a couple people. I'm just looking at my Twitter messages right now. Okay, so we're going to keep going. So where were we in case you just joined? If you just joined, I am currently writing the definitive account of how Codex happened. You can't quite see me. This is the way the camera is set up. I interviewed a bunch of people at OpenAI. I've been using Codex since the beginning. And I realized that voice mode is like the ideal way to work this way. So I've been using voice mode. And right now I'm in a section of the article about test time compute. So I think that's like obviously the enabling technology for everything else that happened. And yeah, welcome. So let's see. I'm going to just take this. I think this is actually a pretty good. Where are we? Not that one. Okay. Okay. So for people that just joined, can you, you should be able to see this proof doc. Can you just read the second and third sections here? So the one that starts with at the time, I thought the gap was probably fatal all the way through the biggest comeback and then we'll keep going from there. Checking. I'm pulling the live draft so I can read the actual sections. At the time I thought the gap was probably fatal. The danger in a new product category is not simply falling behind. It's running the wrong race. If you misunderstand the category from the beginning, every decision pulls you further off course. And for a while, OpenAI did. Code code exploded, hit half a billion in three months. Then more than 2.5 billion annualized while chat GPP started to feel dated. And yet, OpenAI is back. This is the story of how codecs made that possible. The first generation were external processes. They thought by responding to one word at a time. I think you're skipping a bunch of stuff. Are you skipping stuff? Checking. I want you to make sure you're reading the whole thing, please. Thank you. You're right. I condensed where you asked for the actual text. That was a miss. I have the full live draft now and I'll read it verbatim. And I wanted to start with at the time, I thought the gap was probably fatal. Checking. At the time I thought the gap was probably fatal. The danger in a new product category is not simply falling behind, it's running the wrong race. If you misunderstand a product category from the beginning, every subsequent decision takes you further in the wrong direction, and it becomes harder and harder to alter your course. And, indeed, for a while afterwards, OpenAI struggled mightily for the first time since I had started covering the company during the GPT-3 days. Broadcode took off with unnerving speed. Anthropic said it reached over $500 million in annualized revenue within three months of its launch. By the summer of 2026, Anthropic had pulled ahead of OpenAI on reported annualized revenue figures. And Broadcode alone was running at more than $2.5 billion a year, with business subscriptions quadrupling. Meanwhile, ChatGPT, which still boasted 900 million active users, but growth had slowed. It had begun to look almost quaint and underpowered. It belonged to the first phase of AI chat. Builders and increasingly companies had begun to move on. And yet. By the time of this writing, OpenAI is back. Codex had passed 5 million weekly active users before the merge, more than six times its February audience. The new combined app is growing like a weed. By July 15th, 2026, according to a widely shared post on X, it was adding more than a million users a day. Most of the team at Every had switched back from flawed code to Codex and ChatGPT over the previous three months. How did this happen? It's a rare story in business history. I can think of few other examples where a large technology incumbent successfully disrupts themselves, then merges the disruptive product back into their main offering. That this is possible is a testament to how OpenAI is built and to the new leverage AI gives a small team inside a giant company. AI lets a few people build enough, quickly enough, to be truly disruptive before they are swallowed by the larger organization. Unlike every product before it, AI gave Codex enough time to become undeniable, to show OpenAI that it was a product surface capable of remaking the company. This is the story of how Codex happened. The biggest comeback in business history begins with something simple, thinking. In 2026, any time you interact with an AI model on any kind of complex request, you will see that word thinking shimmering in the place where your response will appear. It might seem like a throwaway, a verbose replacement for the spinning beach ball on Mac OS, or the blue ring windows displaced beside its cursor, each a universal sign that the computer is operating at the limits of its capacity. But thinking in this context is not just a label. It denotes a new style of computing, without which Codex would never have happened. The first generation of large language models were external processes. They had no thinking UI shimmering at all. Instead, they did their thinking directly by responding to you. Whatever you asked them to write an email or a line of code, they had to work out their answer, one word at a time, immediately and in public. In that way, early chat models were like improv actors given a prompt by the audience and forced to respond with no do-overs or rehearsals allowed. The results were impressive and sometimes predictably funny as these alien tools compose on the wall. What all users of these models noticed is if you ask the model to think step-right as an analogy, it would be open right to the other person. Okay, this is the technique called... Technique called chain of thought. Okay, we need a description of chain of thought. Can you insert it like a really simple one that works in that context? Checking. I'll add a simple line right after the step-by-step bit, explaining that chain of thought just means showing the intermediate reasoning instead of only the answer. I tried that live draft. Yeah, I think you may not have the share token, but you should be able to get the share token. It should be in our chat history. Oops. Checking. I'll look through the local thread. Here, here, I got it for you. Checking. I checked the thread artifacts and the proof tabs I can see. I sent it to you. I sent it to you. Checking. Okay, so... Okay, do it. Okay, so I'm just... Okay, do it. Okay, do it. I'm just going to think out loud. Okay, what early years of these models notice is if you ask the model to think step by step out loud, it would often arrive at a better answer than if you didn't. This is the prompting to do an equal chain of thought. After you insert that, then we need to bridge into thinking, and test time compute. Done. It now says, this was called chain of thought. Asking the model to show its working step by step before it gave its final answer. And the bridge is, at first, chain of thought was a trick you had to ask for. More words before the answer. The reasoning models turn that into a capability. They use those extra steps on purpose. That's test time compute. It needs to parallel like the second generation, like the first generation needs to parallel, you know? This led to the second of AI models. And the thinking UI we've been discussing. Okay, so see what I said? This led to the second generation of AI models and the thinking UI we've been discussing. Now, put in a couple sentences there that you think would make sense and explain the O-series models, basically. Basically. Checking. I'll add a short bridge into the second generation right after the chain of thought paragraph, leaving your wording untouched. Added. This led to the second generation of AR models and the shimmering thinking UI we have been discussing. Beginning with OpenAI's O-series, models were trained to use a chain of thought as a private workspace and allowed to spend- Okay, at the point you say private workspace, I've already lost you. So we need to, we have a bunch of descriptions earlier in our chat about like describing chain of thought. Can you look at those and think about how to, how to describe it in a way that actually makes sense for a reader? And then replace it in the doc, please. Checking. Revised plain terms. This led to the second generation of AI models and the shimmering thinking UI we have been discussing. Rather than asking the model to work through a problem step by step in the answer you saw, the O-series could take a step to before it answered, giving it extra time to do that on a particular problem which is called step-time computer. I replace the bridge with that. and possibly can possibly possibly possibly Thank you. Thank you. Thank you. Thank you. I'll check the research plane, then add one plane sentence that sticks to what it supports. Edit. Checking. I'll look for quotes that speak directly to that shift, not just generic praise. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Here are three quotes that really land here. Alex on the spark. Checking. Checking. Checking. Two from Alex. First, the first feel the AGI moment for me was I saw someone running a reasoning model connected to a terminal. That's the spark scene for this section. Second, the why? Because we were very oil pilled. If we could increase test time compute, we got better performance. Here's conclusion. So design around asynchronous models, long running models. That's where research had a clear path. That pivot connects the reasoning discovery to the product direction. Checking. Checking. I did. I added the line about 01 getting better on hard tasks when allowed more time to think. It's in the draft now. Let me just refresh the live documents and confirm the exact paragraph. Checking. Got it. We're aligned. I'm still going through Greg and Thibault in full and I'll only bring back quotes that truly belong in this section. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Check in. I'm continuing through Greg and Thibault's full interviews focused on quotes that earn their place here. I've got both threads and I'm checking your new proof comment first. Thank you for the interview past running as we do. Your comment is fixed. 01 is introduced before the performance claim and I resolve the note. I finished the Greg and Thibault pass. The strongest material splits three ways. Alex for the technical hinge and the spark scene. Greg for we had kind of the smartest middles, but that doesn't mean they're the easiest to use, which bridges reasoning to product shape. And Thibault for our models were rapidly getting better and we were seeing them do longer horizon tasks. Plus, you can just go and do it in two days and then show the thing which backs your point about small teams moving fast enough. Check in. You're right. That was heavy on general strategy. I'm tightening the search to concrete observations about reasoning behavior and what changed in practice. And I'm reading me as full interview now. And I'm reading me as full interview now. I'm reading you as well. Here are three specific pieces. Alex says the research result. If we could increase test time compute, we got better performance is what pushed them to long running models, not customer demand. And Mia says people were showing, oh, my model has been working for 48 hours, which meant Codex cloud was built around a forecast that missed where users actually were on their laptops. Thibault backs the technical side. We were seeing the models do longer and longer horizon tasks successfully, which makes the desktop bet makes sense. Greg is better as a follow on. Smartest doesn't mean easiest to use, which can close this bit or set up the next move. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. Checking. 575 now in the approved draft plus about 575 from the earlier free write and morning work. Thank you. Checking. I'll keep that sequence simple and tight. One factual fix as I apply it. That increased test time compute line is Alex's not Tebow's. Checking. Yes. Done. It now reads, if we could increase test time compute, we got better performance, Alexander told me. If you can, and possibly can possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly possibly Thank you. Checking. I'll pull the chronology up so we can trace the next movements from this thinking pivot. The timeline is open. Checking. Here's the codex chronology. Checking. Click and find a link to the box and find a link to the box and find a link to the box and to the box and find a link to the box and find a link to the box and find a link to the box and find a link to the box and find a link to the box and find a link to the box and find a link to the box and find a link to the box and find a link to the box and find a link to Thank you. Thank you. Thank you. Thank you. All right, folks. That is a two-hour writing session. I got through my thousand words for today. And you also got to see how you can use voice mode in Codex and ChatGPT for work. It's fucking crazy. It has truly changed how I work. And we'll be doing a lot more on voice. So if you're into this, you should subscribe to Every. Every is the only subscription you need to say at the Edge of AI. I will have this article. It'll come out as an essay. It'll be a really beautiful essay. If you want more stuff like this, we actually just published something that I think is fucking incredible and is way underappreciated. I'm going to just share my screen again. And we just published this article. Look at that. Look how cool it is. And this is Every.to.p.openai-infrastructure. We just did an article looking at how the basically software is changing. And coding agents are allowing us to basically build software skyscrapers where you can have a hundred times more lines of code and a thousand times more contributors in your software. So the question is, like, how do you build infrastructure to, like, make that actually work? And we spent a lot of time talking to OpenAI's infrastructure team, who are the people who kind of, like, bear the brunt of this and have to figure this out about how they make that work internally with agents. And we did this, like, awesome article. We did custom. We did, like, photography. We did a bunch of stuff with their team. So if you're interested in stuff like this and you want to read this, I highly recommend Every.to.p.openai-infrastructure. Check that out. It's by Laura Entis. Incredible article. I will have my definitive account of how codex happened coming out in about hopefully a couple weeks with a similar kind of thing. And we'll be doing a lot more guides, camps, courses, all that kind of stuff if you want to learn how to do voice mode. And obviously, we'll be covering everything that comes out, new models that come out, everything from OpenAI, everything for Anthropic. We also have a bunch of tools. You should use Proof, which I just was using today with voice mode. It's really great. I just think there's something special happening here with the operating system for work happening inside of your agent orchestration systems, like Chachypt for Work or Codex and the Cloud Desktop app. And I think there's in particular something cool about Codex happening right now. They just did the merge, and this voice mode thing is fucking sick. So that is the show for today. Maybe I'll do this tomorrow. We'll see. I'm taking the week off to write this. I probably won't finish it this week, but I will make a lot of progress. I'll make enough progress to bother more people at OpenAI for more inside details. And that's it. Thank you all for joining.