AI Engineer

The Golden Age of AI Engineering — Alexander Embiricos & Romain Huet & Peter Steinberger, OpenAI

2105 summary words 9 min summary Watch video

Start with the signal

9 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: AI engineering is shifting from manually supervising individual coding agents to directing persistent, event-driven teams of agents through a chat-first control plane with optional hands-on inspection.
  • Why it matters: The presentation directly outlines OpenAI's emerging agent stack and Peter Steinberger's practical manager-worker orchestration pattern, including persistent context, delegation, triggers, isolated execution, review gates, and OpenClaw integration.
  • Best use: Use it as an architecture and product-design reference for OpenClaw control planes, long-running agent workflows, human approval boundaries, and the transition from interactive coding assistants to autonomous execution systems.

Executive Summary

OpenAI argues that stronger coding agents will expand rather than eliminate engineering because engineering is fundamentally problem selection, judgment, design, and delivery—not code entry. The capability progression has moved from completion and inline prediction to agents that edit, test, and pursue long-running goals, making the engineer's role increasingly about setting direction and evaluating completed work.

The proposed product model has two complementary modes: a universal conversational interface for delegating work and a collaborative surface for inspecting, steering, or directly editing when necessary. OpenAI believes this should be chat-first rather than code-first: users normally communicate intent and let agents work, then descend into implementation details only for exceptional or high-value decisions.

OpenAI presents Codex as an extensible stack rather than a closed application. Its own products reportedly use the same Responses API, open-source Codex harness, Agents.md instructions, App Server, and plugin interfaces offered to developers. The stack supports alternative models, third-party applications, existing Codex subscriptions, and integrations including OpenCode, Pi, Droid, OpenClaw, Xcode, and JetBrains.

Peter Steinberger supplies the strongest operating lesson: manually juggling many terminals is not orchestration but human polling. His newer model uses one persistent manager that delegates to workers, reacts to triggers, and returns decision-ready artifacts such as a PR, diff, video, or running build. Agents own the inner execution loop; the human owns the outer loop of direction, exceptions, and approval. As tokens and compute become easier to supply, scarce human attention becomes the primary constraint.

Key Takeaways

  • Claim: The relevant transition is from AI-assisted code generation to agents that can execute complete computer-based goals, including the work before coding and after implementation. | Evidence: The speakers trace the progression from completion and inline prediction to command-based edits, self-testing, and agents pursuing long, difficult goals until completion. They contrast a 2024 drone demo, where the model could not run or verify its code, with a 2025 demonstration in which a model controlled and tested a live camera and lighting system. | Implication: Ken should evaluate agent systems by end-to-end work landed—not code generated—and include intake, environment selection, testing, review, deployment, and evidence collection in the orchestration design. | Caveat: The claim that agents can perform any task available on a computer is aspirational and broad; reliability, permissions, environmental access, and task-specific evaluation are not quantified.
  • Claim: The durable agent interface will combine chat-first delegation with an optional, deeply inspectable workspace rather than forcing users to supervise every execution step. | Evidence: OpenAI compares the experience to managing a team: most work begins through conversation and proceeds independently, while users occasionally enter the details to point at a specific element, request a change, or edit it themselves. The Codex app is positioned as this combination of simple chat and progressively deeper inspection. | Implication: OpenClaw and related control planes should optimize the default experience around intent, status, exceptions, and outputs while retaining traceability and direct intervention for high-risk work. | Caveat: The speakers do not address which security-sensitive or irreversible tasks should require hands-on supervision rather than default delegation.
  • Claim: OpenAI is treating Codex as a layered developer platform whose public primitives are also used by its first-party products. | Evidence: The described layers include models through the Responses API, API-level context compaction, the open-source Codex harness, Agents.md, an open-source App Server, and application plugins. The harness does not hard-code OpenAI models, and the same harness is reportedly included in model post-training so models learn tool use in an inspectable environment. | Implication: Ken can use OpenAI's stack as a reference implementation while preserving model and runtime portability at the harness boundary rather than coupling OpenClaw to one front end or model family. | Caveat: Open sourcing interfaces and harness code does not remove dependency on OpenAI subscriptions, hosted models, undocumented service behavior, or future platform changes.
  • Claim: Agent economics should be optimized for completed value rather than token consumption, with model cost and speed changing which orchestration patterns are feasible. | Evidence: OpenAI calls the objective "value maxing." It claims GPT 5.6 Terra delivers GPT 5.5-level intelligence at half the cost, gives Luna pricing of $1 per million input tokens and $6 per million output tokens, and previews GPT 5.6 Sol on Cerebras at 750 tokens per second. The speakers suggest using this speed to run five or six approaches in parallel and select the best. | Implication: Model routing should account for expected task value, success probability, elapsed time, and verification cost—not merely benchmark rank or token price—and speculative parallelism should be reserved for decisions where diversity has measurable value. | Caveat: These are vendor-reported preview claims without workload-specific quality, latency-distribution, or total-agent-cost data; parallel search can erase per-token savings.
  • Claim: The distinction between local and cloud agents should disappear behind a runtime scheduler that selects the appropriate environment for each job. | Evidence: The speakers note that users currently leave laptops open so agents can continue working and argue instead for many parallel tasks running in isolated boxes. Codex Cloud is due for upgrades, while the stated target is an agent that decides which available local or remote environment fits the work without making the user choose manually. | Implication: A serious agent control plane needs environment-aware routing based on credentials, data sensitivity, required tools, compute, test isolation, cost, and task duration, with runtime placement hidden from ordinary users but auditable by operators. | Caveat: No scheduling policy, data-residency model, credential strategy, or isolation specification is provided.
  • Claim: Persistent context, delegation, and event triggers are the three primitives that turn multiple agents into an autonomous operating loop. | Evidence: Steinberger says server-side compaction made long-running work reliable enough to stop optimizing around sessions; coordination allowed one thread to create and steer projects; and automation allowed the same manager to wake when events occurred. His example starts with a newly filed issue, which a manager compares with project goals before assigning implementation and review workers. | Implication: Ken should treat memory continuity, hierarchical delegation, and triggers as core orchestration services, then add explicit budgets, stop conditions, idempotency, escalation rules, and audit trails before increasing autonomy. | Caveat: The transcript does not explain conflict resolution, loop termination, budget enforcement, or protection against a manager repeatedly spawning low-value work.
  • Claim: As autonomous execution scales, human attention—not tokens or compute—becomes the limiting resource, so agents should return compact decision packages rather than streams of intermediate reasoning. | Evidence: Steinberger describes moving from token constraints to compute constraints, then using separate test boxes to shift the bottleneck to attention. Instead of watching code generation, he wants the manager to return a PR, original issue, proposed diff, video, or VNC-accessible build; the human reviews once and approves or comments while agents continue the loop. | Implication: The outer-loop interface should prioritize provenance, test evidence, risk flags, material diffs, and approval decisions while sampling enough underlying execution data to prevent polished summaries from masking failures. | Caveat: High model competence does not guarantee that compressed summaries expose hidden security, architectural, or product errors.

Detailed Brief

Platform extensibility and ecosystem evidence

  • Claims: OpenAI says new capabilities required by Codex are first converted into reusable platform primitives, with long-context compaction offered as the representative example.; The App Server was created to provide one control interface for the Codex harness across first-party surfaces such as the VS Code extension and Codex app.; Application-level innovation is intended to happen through shared extension points rather than only through OpenAI's roadmap.
  • Evidence: Toma, known as Dimalian on X, used App Server to build the native Codex Monitor application before OpenAI launched the Codex app; he later joined OpenAI and built Codex for iOS.; The OpenCode team reportedly inspected the reference implementation, including ChatGPT signing behavior, and selectively reused or changed its components.; Browser use and computer use were implemented as plugins using the same extension points exposed to developers; role-specific data-science and design plugins are also open source.
  • Caveats: The talk offers no compatibility guarantees, governance process, versioning policy, or long-term commitment for these extension points.; Subscription reuse across third-party applications raises authentication and delegated-access questions that the presentation does not explore.
  • Implications: The strategic value of the stack may come less from the Codex user interface than from becoming a common harness, authentication path, and extension substrate across competing agent products.; Third-party builders can differentiate at the workflow and experience layers without reimplementing the entire execution loop, but platform-risk controls remain necessary.

Product development changes inside AI-native engineering teams

  • Claims: The speakers argue that agent leverage improves product judgment as well as engineering throughput because teams can test more ideas before committing.; OpenAI says it uses Codex to build Codex, creating a feedback loop between model capability, developer tooling, and first-party product requirements.
  • Evidence: The product team reports prototyping more ideas and spending more time with users rather than treating generated code volume as the main benefit.; The speakers cite community experiments on X as a recurring source of previously unanticipated Codex use cases and future product direction.
  • Caveats: Faster prototyping can increase roadmap noise unless teams maintain strong criteria for evidence, user value, and abandonment.; The talk does not quantify changes in cycle time, defect rates, user contact, or shipped-product outcomes.
  • Implications: The organizational advantage may accrue to teams that convert cheaper experimentation into better selection mechanisms, not simply to teams that generate the most prototypes.; Community-discovered workflows can function as external R&D and should be monitored systematically rather than through ad hoc social-media browsing.

Notable Concepts & Terms

  • Value maxing: Optimizing agents for useful completed outcomes rather than maximizing token usage, raw intelligence, or generation volume.
  • Inner execution loop: The agent-owned cycle of investigating, implementing, testing, reviewing, and continuing work without continuous human supervision.
  • Outer loop: The human-owned layer for setting direction, allocating attention, resolving exceptions, and approving consequential outputs.
  • Server-side compaction: A mechanism for compressing long context so a manager agent can persist across extended tasks without being organized around short-lived sessions.
  • Agents.md: A shared instruction-file convention intended to work across agent implementations rather than locking project guidance to Codex.
  • Codex harness: The open-source execution loop connecting models, tools, and environments; it can use non-OpenAI models and is also used during OpenAI model post-training.
  • App Server: OpenAI's open-source control layer for embedding and operating the Codex harness across different applications.
  • Attention bottleneck: The constraint that emerges after token availability and execution compute scale: humans cannot inspect every agent stream and must review decision-ready outputs selectively.

Operator Notes / Why Ken Should Care

  • Prototype a manager-worker workflow around one bounded repository event, such as issue intake, with separate implementation and review agents plus a mandatory human merge gate.
  • Define a standard decision-package schema containing the originating request, goal alignment, material diff, tests, risk flags, screenshots or video, deployment impact, and requested approval.
  • Add explicit spawn budgets, time limits, retry ceilings, cancellation behavior, and duplicate-event protection before enabling trigger-driven agent loops.
  • Separate execution onto disposable test boxes or sandboxes and prohibit workers from inheriting production credentials by default.
  • Test the Codex App Server and harness boundaries for OpenClaw interoperability, but maintain an adapter layer so authentication, model choice, and runtime placement remain replaceable.
  • Benchmark model tiers using cost per verified completion and human-review minutes, including whether parallel candidate generation actually improves acceptance rates.
  • Create an operator dashboard that surfaces stalled jobs, repeated retries, unusual tool use, budget overruns, and approval queues rather than exposing every intermediate agent message.
  • Review third-party Codex subscription sign-in flows for token scope, revocation, auditability, and tenant isolation before treating subscription portability as an enterprise integration path.

Source/Metadata

  • Title: The Golden Age of AI Engineering — Alexander Embiricos & Romain Huet & Peter Steinberger, OpenAI
  • Transcript words: 5414
  • Duration seconds: 1513
  • Timestamp note: No timestamps or chapters were present. The supplied transcript contains duplicated passages and appears to end during Peter Steinberger's segment.
Full transcript 4441 words · 24 min read
0:00

[SPEAKER_02] Good morning, everyone.

0:12

SPEAKER_02

I'm Roman. Hey, everyone. I'm Alexander. Wow, this room is incredible. There's over 7,000 AI engineers here today with us. And it's not just about who's talking about this technology. It's also about who's actually using it and pushing the frontier every day. So we couldn't be more proud to be here with all of you today. And when we were thinking about this event with Alex, we kept coming back to the World's Fair. And the World's Fair may actually be the future visible to everyone by building it in public. Ideas that previously sounded impossible were actually suddenly there. People could see them, they could walk into them, and they could even start to believe in them.

0:45

SPEAKER_02

And honestly, this event has the exact same energy. The future of engineering is not arriving from somewhere else. It's really being built here by the people in this room and much faster than most expected. And that's why it's surprising that people keep saying that engineers are going away. The argument is that coding is abstracted away, and therefore, eventually, we won't need engineers. Well, in fact, we think it's quite the opposite. Software ate the world. And then AI ate software. But now what we're here to say is that AI engineers are aiding the world. AI engineers are the people here pushing the frontier. Yes.

1:13

SPEAKER_02

And you all are figuring out how this new capability can reach everyone. And there has never been a better time to be an engineer, in fact, because engineering was never about writing code. Engineering has always been about solving problems for yourself and for other people as well. It's about taking the latest science and combining it with design, with taste, with judgment, and most of all, imagination to make something that people can actually use. And in that sense, it's not the end of engineering. We think it's a return to the roots of engineering. And the technology we're building on is accelerating, getting faster and faster.

1:34

SPEAKER_02

[SPEAKER_00] For example, we used to ship a new model every 15 months or so, and now it's about roughly every six weeks. [SPEAKER_00] And in case you missed it, last week we launched a preview of the 5.6 series, and we're super excited to get it into all of your hands. [SPEAKER_00] Now, building on top of all these models, the rate of product progress is relentless. [SPEAKER_00] And as a result, I don't have to tell you, the engineering feels completely different.

1:47

SPEAKER_02

[SPEAKER_00] So just to go over a couple years of what for me were successive mind-blowing experiences, obviously for a long time we've had completion, and then we went to inline prediction, and then finally we had command K where you could ask a model to make a change, but they wouldn't test the work. [SPEAKER_00] Then models started testing the work, and now we have models taking on long, hard goals until they're done. [SPEAKER_00] And for me, each of these phases, I remember the first time, was just mind-blowing, and then obviously afterwards, you just get used to it and you're trying to get your work done.

1:56

SPEAKER_02

Yeah. In fact, I can't believe that build and test loop was not even part of the models just two years ago. This was a picture of me at Dev Day 2024, and I used 01 in preview at the time to build a mini-drone interface from scratch. And the slightly insane part is the model could not actually run the code or verify its own work, and I knew the demo would work most of the time, but surely not all of the time. So I had to cross my fingers. You can see here that I was pretty nervous, but that's me.

2:13

SPEAKER_00

[SPEAKER_02] I only do live demos, so I never know what's actually going to happen each time. [SPEAKER_02] Luckily, it did work, and by Dev Day of last year in 2025, I was confident enough now that the model could test their own work to control an entire camera system and lighting system live. [SPEAKER_02] But we've come a long way. [SPEAKER_02] Yeah, so we refer to Roman as the demo god, and before the demo, I'll ask him, so hey, how often does demo work? And he'll be like, three times out of four, and we're like, all right, good luck. Obviously, we've come miles since then. And this year alone has been crazy.

2:34

SPEAKER_00

So what we're putting up here are all the things that we've shipped so far, and actually not even all, a selection of the things that we've shipped so far this year. My favorite things that we've shipped are Codex app, goal mode, remote. These are things that really change how it feels to do work. Obviously we couldn't do these things if we didn't use Codex to build Codex. But I think to me what is most interesting is that now Codex can do, and agents can do, any task that you can do on your own computer.

2:48

SPEAKER_00

And so that means they're not just helping you with the coding, but they're helping you with what happens before the coding, and they're helping you with what happens after the coding. And this is really key, right? I think there's been a lot of talk, there will be a lot of talk today about loops. And if you can connect the agent to not only the work that you have to do, but why it has to be done, that's how you can get the agent to start to begin much more work. And then if you can connect it to what you do afterwards, review and deploy, that's how you help it land much more work. So with all of this, of course we can move much faster.

3:00

SPEAKER_00

But to me as a product person, the most exciting thing is actually that we make better decisions around what to build. For instance, we prototype many more ideas, and we spend much more time with users. So yes, that's all of you. So I wanted to pause and just give you all a big thank you, both for the love and the constructive feedback.

3:11

SPEAKER_02

[SPEAKER_00] I would say it's safe to say that we, Codex, and actually the entire industry wouldn't be here without you. [SPEAKER_00] Yeah, thank you so much for all the feedback. [SPEAKER_00] We're constantly listening to all of you. Thank you.

3:22

SPEAKER_02

[SPEAKER_00] So with all of this, of course we can move much faster. [SPEAKER_00] But to me as a product person, the most exciting thing is actually that we make better decisions around what to build. [SPEAKER_00] For instance, we prototype many more ideas, and we spend much more time with users. [SPEAKER_00] So yes, that's all of you. [SPEAKER_00] So I wanted to pause and just give you all a big thank you, both for the love and the constructive feedback. [SPEAKER_00] I would say it's safe to say that we, Codex, and actually the entire industry wouldn't be here without you. [SPEAKER_00] Thank you so much for all the feedback.

3:44

SPEAKER_02

[SPEAKER_00] We're constantly listening to all of you. Thank you. Okay. So the models are getting really good. I would say if you pick a medium length computer task, and you give me and the model the same amount of time to get that task done, probably, at least in my case, the model will do a better job than me for the average task. And so, okay, we're getting these models. In some ways they're smarter than us.

3:58

SPEAKER_00

[SPEAKER_02] They can do almost anything. [SPEAKER_02] How should we shape that? [SPEAKER_02] What should the products that we use feel like? [SPEAKER_02] And so to answer that, we look to our mission. [SPEAKER_02] The part of it here that I've got up is, AGI that benefits all of humanity. [SPEAKER_02] And I think in order to do this, there are two main questions that I think about now. [SPEAKER_02] One is, how do we set up the agents to actually do things in the world? [SPEAKER_02] So what can they do? [SPEAKER_02] Gradually agents are getting connected to more and more things. [SPEAKER_02] And where do they run? [SPEAKER_02] More on that later.

4:34

SPEAKER_00

[SPEAKER_02] And then the other question is, how do we use these agents? [SPEAKER_02] For us, what should the product feel like around them? [SPEAKER_02] And for us, the goal is squarely not to automate engineers. [SPEAKER_02] Instead, the product shape that we want is one that maximally empowers engineers. [SPEAKER_02] So, if we think about what that product shape is, we actually think it's pretty simple. [SPEAKER_02] I read a lot of sci-fi and watching superhero movies, and I actually think that the simple ideas in there are approximately right. [SPEAKER_02] So there are two modalities roughly. [SPEAKER_02] Chat.

4:55

SPEAKER_00

[SPEAKER_02] I actually think I know some people think chat is dead. [SPEAKER_02] I think chat is underrated. [SPEAKER_02] And some kind of hands-on experience. [SPEAKER_02] So what you want is a single entity that you can ask for help with anything, anywhere. [SPEAKER_02] And then you want a powerful collaborative UI that you can use when you want to inspect, steer, or shape things yourself. [SPEAKER_02] And so I had Codex image-gen me an illustration of this to help understand when you might want to use these things. [SPEAKER_02] And so my analogy for you, yes, I hope you like the image-gen. [SPEAKER_02] My analogy for you would be it's just like working with a team.

5:28

Most of the time, you're just talking about stuff, and your team is just doing stuff. You don't actually want to watch over the shoulder or have to walk over to the workbench of your teammate for every single unit of work. Mostly you just want to talk and let them cook. [SPEAKER_02] And then every now and then, you want to dig in and really dig in all the way to the weeds of things and dig into that problem together. [SPEAKER_02] And for us, as we build product, we have this idea that we want to make it so that you can retain this feeling of mastery of the work that we're doing. [SPEAKER_02] Because that's really powerful.

5:51

SPEAKER_00

[SPEAKER_02] We don't want to make it feel like actually it's really hard to get to the details and disassemble the hardware in this case. [SPEAKER_02] So the way that we're bringing this to life is just the beginning, but this is why we built the Codex app. [SPEAKER_02] You get a very simple chat interface that you can use for coding and for anything else. [SPEAKER_02] And you can have a conversation and then go as deep as you want. [SPEAKER_02] So in the case here, we have Romain's predicted score of this upcoming World Cup match. [SPEAKER_02] I hope I'm right. [SPEAKER_02] We'll see. [SPEAKER_02] Okay.

6:14

SPEAKER_00

[SPEAKER_02] And what you can do here is you can go in and you can point at a very specific thing and say, hey, I want you to make this change. [SPEAKER_02] Or you can make this change yourself. [SPEAKER_02] And a fun story here is that actually I remember pitching some of you who I know are in the audience this idea before we started. [SPEAKER_02] And I was told squarely, I will never use such a tool. [SPEAKER_02] I will never leave my terminal or Vim or Emacs. [SPEAKER_02] But actually those people are now using it. [SPEAKER_02] And even internally, within our team, there were a lot of questions. [SPEAKER_02] Why should we build this? [SPEAKER_02] People love the CLI.

6:44

SPEAKER_00

[SPEAKER_02] They love the IDE. [SPEAKER_02] And it's a little subtle, but our take is that you can't really build that collaborative interface for any kind of work in a CLI. [SPEAKER_02] It's mostly chat. [SPEAKER_02] And then in the IDE, the order is wrong. [SPEAKER_02] So you're starting with the code. [SPEAKER_02] But now it's time to transition to working with teammates where you chat first and you dig in when you need it. Totally. And we're moving really fast on this, on the product surface and the model layer, of course. But we're also trying to keep pace with all of you, right?

7:19

SPEAKER_00

Honestly, half the time I open X, I see someone in this room doing something that I had not realized Codex could actually do. [SPEAKER_02] And honestly, this is what pioneers do. [SPEAKER_02] You guys experiment. [SPEAKER_02] You set up tools for yourself, for your team. [SPEAKER_02] And in turn, we get inspired. [SPEAKER_02] We learn from you. [SPEAKER_02] And eventually, everyone benefits. [SPEAKER_02] And so we are helping, you are helping us figure out what to build next and also what the future of engineering should look like. [SPEAKER_02] But for that to work, one thing that we really care about is that Codex cannot be a closed product that only OpenAI can improve.

8:03

SPEAKER_00

[SPEAKER_02] [SPEAKER_00] Honestly, half the time I open X, I see someone in this room doing something that I had not realized Codex could actually do. [SPEAKER_02] And honestly, this is what pioneers do. [SPEAKER_02] You guys experiment. [SPEAKER_02] You set up tools for yourself, for your team. [SPEAKER_02] And in turn, we get inspired. [SPEAKER_02] We learn from you. [SPEAKER_02] And eventually, everyone benefits. [SPEAKER_02] And so we are helping, you are helping us figure out what to build next and also what the future of engineering should look like.

8:25

SPEAKER_00

[SPEAKER_02] But for that to work, one thing that we really care about is that Codex cannot be a closed product that only OpenAI can improve. [SPEAKER_02] So we've intentionally designed Codex as a set of layers that anyone can build on. [SPEAKER_02] And we want to show you a little bit of that stack today and how it manifests. [SPEAKER_02] First, it starts with the model. [SPEAKER_02] And Alexander showed how quickly we're progressing on models. [SPEAKER_02] And you guys use these models through the responses API. [SPEAKER_02] And guess what? [SPEAKER_02] This is how we build the Codex app, right? [SPEAKER_02] We use the same models through the same API.

8:51

SPEAKER_00

[SPEAKER_02] And we actually are building on the same thing that we give to developers. [SPEAKER_02] And when Codex needs something new, we always try to bake it into the API first so you can benefit as well. [SPEAKER_02] One example recently was compaction. [SPEAKER_02] Codex needed a way to compact long contacts for long-running tasks. [SPEAKER_02] And so we'll build that into the API. [SPEAKER_02] So that means your agents can use the same primitives that we build for ourselves. [SPEAKER_02] Moving on to the next layer, the Codex harness is also open source. [SPEAKER_02] So you can inspect it, you can fork it, you can adapt it.

9:17

SPEAKER_00

[SPEAKER_02] And we also took the same approach with agents MD.

9:24

SPEAKER_02

Instead of reinventing a new file format for Codex to follow instructions, we thought let's pick a name that other agents can actually use as well. The models are the default in the harness, the models from OpenAI, but they are not hard-coded in there. So if you want to use an open model and keep the same agent loop, you can. And we also bring this Codex harness into the post-training process of our models. So that means the models can learn to call tools and navigate an environment that's actually something that's open source. Now take the open code team, for instance.

9:33

SPEAKER_02

They actually were able to inspect how we have this reference implementation, and they could reuse the parts that make sense to them or change entirely all the rest and make different choices. I know, for instance, they were trying to see how we did signing with ChatGPT, and so they could look at the code and learn from it. And we think it's better than having developers reverse engineering how it builds and how we launch. But now let's say speaking of subscriptions that we want to go a level higher, and how you bring this harness into an app? And how do you let people sign in with their existing Codex subscription, for instance?

10:00

SPEAKER_02

Well, it turns out we had the same problem ourselves, because we wanted to build a VS Code extension and the Codex app. And we wanted to have a unified way to actually control this harness. So we built App Server, and we also made that open source. And the App Server is not a community adapter. It's really the path that we use for our own products, and you can use it too. Toma, for instance, here, aka Dimalian on X, he built his own native app for Codex called Codex Monitor before we even launched the Codex app, because he could build that using the App Server. And now he works on our team, and he actually built Codex for iOS.

10:22

SPEAKER_02

And moving up the stack at the app layer, we also want to make sure that innovation is not blocked on our own ideas. And so we build extensible primitives here, the in-app browser that we showed on the screen, and plugins. So if you take, for instance, browser use and computer use, these were built as plugins using the same extension points that we have available for all of you. And lastly, we also recently built role-specific plugins for Codex, to make it easier to customize for people who work in data science or design, for instance. And these plugins are also open source. You can see under the hood and get inspired from them if that's useful.

10:42

SPEAKER_02

So our goal is really to keep making this as open and flexible as we can. And the best part is, people can use their existing subscription in more and more places, from OpenCode, Pi, Droid, OpenClaw, to even Xcode and JetBrains as IDEs. And you can see how they're becoming quite a meaningful part of how people use these tools. And that's really why we want to care about building this open ecosystem with all of you. So really, if there's one thing I wanted to take away from this section and this stack, it's this. [SPEAKER_00] [SPEAKER_02] We're not building one system for OpenAI and a second system that's simplified for developers.

11:08

SPEAKER_02

[SPEAKER_00] [SPEAKER_02] At every layer, we actually use the thing that we give to you. [SPEAKER_00] [SPEAKER_02] And we want to thank all of you, because every time you fork the harness, every time you find the edge of capabilities of the models, it means we get to learn and improve. [SPEAKER_00] [SPEAKER_02] And honestly, with 7,000 of the finest AI engineers in this room today, I'm confident that all of you will define a lot of how we will experience AI and how the world will experience AI in the future. [SPEAKER_00] [SPEAKER_02] So thank you. [SPEAKER_00] [SPEAKER_02] I want to give a shout out to whoever over there is injecting energy.

11:32

SPEAKER_02

[SPEAKER_00] [SPEAKER_02] That's you. [SPEAKER_00] [SPEAKER_02] Okay, thank you so much. [SPEAKER_00] So with all of your help, we are making agents explosively useful. [SPEAKER_00] And so now the question is, how do we get value out of them? [SPEAKER_00] And that's not token maxing. [SPEAKER_00] We have a term for this that you use as well. [SPEAKER_00] I don't know, is it on screen? [SPEAKER_00] Value maxing. [SPEAKER_00] So when we talk to engineering leaders, most of the conversation is about themes relating to the idea of value maxing.

12:06

SPEAKER_02

[SPEAKER_00] So we're going to walk you through a few common topics that come up, some things where we've already made a lot of progress, and some things where actually there's a lot more progress to still be made. [SPEAKER_00] So the first one of these is cost efficiency. [SPEAKER_00] Everyone wants frontier intelligence. [SPEAKER_00] And so now the question is, how do we get value out of them? [SPEAKER_00] And that's not token maxing. [SPEAKER_00] We have a term for this that you use as well. [SPEAKER_00] I don't know, is it on screen? [SPEAKER_00] Value maxing.

12:43

SPEAKER_02

[SPEAKER_00] So when we talk to engineering leaders, most of the conversation is about themes relating to the idea of value maxing. [SPEAKER_00] So we're going to walk you through a few common topics that come up, some things where we've already made a lot of progress, and some things where actually there's a lot more progress to still be made. [SPEAKER_00] So the first one of these is cost efficiency. [SPEAKER_00] Everyone wants frontier intelligence. [SPEAKER_00] Pick your favorite eval. [SPEAKER_00] You want the best model. [SPEAKER_00] So with terminal bench here, for instance, that's GPT 5.6 Sol. [SPEAKER_00] And we can't wait for you to have it.

13:32

SPEAKER_02

[SPEAKER_00] But okay, you also want as much intelligence as you can get. [SPEAKER_00] And that's where efficiency comes in. [SPEAKER_00] Cost efficiency has been a focus for us for quite some time. [SPEAKER_00] And the results are continuing to pay off. [SPEAKER_00] So for example, GPT 5.6 Terra, I think it's in dark blue in there, brings GPT 5.5 level intelligence, but at half the cost. [SPEAKER_00] And Luna there, beats some pretty notable models in this eval.

14:13

SPEAKER_02

[SPEAKER_00] But at only $1 per million input tokens and $6 per million output tokens. [SPEAKER_00] I'll leave it up to you to compare those costs, but that is insane value. [SPEAKER_00] Yeah, we really can't wait to see all of you build with GPT 5.6 and these new family of models.

14:19

SPEAKER_00

[SPEAKER_02] [SPEAKER_00] Now, the next thing I want to touch on is speed, right? [SPEAKER_02] [SPEAKER_00] GPT 5.3 Codex Spark showed you what speed can unlock. [SPEAKER_02] [SPEAKER_00] But we also know that you all want frontier intelligence. [SPEAKER_02] [SPEAKER_00] You don't want to have a model that's not as great as what you can operate at the very best. [SPEAKER_02] [SPEAKER_00] Well, this is GPT 5.6 Sol running on Cerebris. [SPEAKER_02] [SPEAKER_00] The frontier intelligence at 750 tokens a second. [SPEAKER_02] We can't wait to see what you can build with this next month.

14:43

SPEAKER_00

[SPEAKER_02] And honestly, with that perspective, this is having a pretty substantial PR written in 10 seconds. [SPEAKER_02] And it's not just about getting one answer faster, right? [SPEAKER_02] It's about what can you do with that speed? [SPEAKER_02] You can think about an agent taking different approaches, maybe five or six in parallel. [SPEAKER_02] Maybe coming back and picking the best one in the time you would have taken to not even generate just one. [SPEAKER_02] So we really can't wait to see what that can unlock when you have frontier intelligence, the very best at that speed.

14:59

SPEAKER_00

[SPEAKER_02] It really starts to feel less like waiting for an AI to respond and much more like a co-worker that's already showing you the results as it goes. [SPEAKER_02] Speaking of working with co-workers, can I get a show of hands? [SPEAKER_02] Who is familiar with this kind of site in offices? [SPEAKER_02] Okay. Okay. [SPEAKER_02] [SPEAKER_00] Well, a lot of you are very well behaved. [SPEAKER_02] [SPEAKER_00] At least you see some people up front. [SPEAKER_02] [SPEAKER_00] So yeah, a lot of people are keeping their laptops open so that agents can keep working.

15:31

SPEAKER_00

[SPEAKER_02] [SPEAKER_00] And this is funny, but you know, what we really want is to be able to shut our computers. [SPEAKER_02] [SPEAKER_00] And we want to be able to run many tasks in parallel, isolated on their own box. [SPEAKER_02] [SPEAKER_00] Now, we've been actually aiming at this from the start. [SPEAKER_02] [SPEAKER_00] Our first major launch was Codex Cloud. [SPEAKER_02] [SPEAKER_00] And it is due for some major upgrades coming soon.

15:54

SPEAKER_02

[SPEAKER_00] But better yet, as we think about this, the future shouldn't have this awkward distinction between a local task and a cloud task and you have to decide where to run everything. [SPEAKER_00] Really what you should have is going back to what I was saying earlier. [SPEAKER_00] You should just have an agent, you talk to it wherever, whenever, about anything. [SPEAKER_00] And it should figure out, okay, what do I need to do, which environment is right for my work, and use whatever is available. [SPEAKER_00] In fact, Theo made this prediction over the weekend on this very topic. [SPEAKER_00] And it's a pretty acute tweet.

16:19

SPEAKER_02

[SPEAKER_01] [SPEAKER_00] Alex, what do you think? [SPEAKER_01] [SPEAKER_02] Sooner or later than six months? [SPEAKER_01] [SPEAKER_02] I think maybe not the exact details of the vibe of this tweet, much sooner than six months. [SPEAKER_01] [SPEAKER_02] I mean, at the pace at which everything is going, I would not be surprised if it's sooner indeed. [SPEAKER_01] [SPEAKER_02] Well, so now, you might be wondering, where's the live demo today? [SPEAKER_01] [SPEAKER_02] Well, for this AI engineer, we wanted to do something a little different this time around.

16:50

SPEAKER_02

[SPEAKER_01] [SPEAKER_02] And we think it's a very unique moment for all of us to reimagine how we work and how we build. [SPEAKER_01] [SPEAKER_02] And so we wanted to bring a special guest who has been at what's possible with agents and really has pushed us to be more AGI-pilled at OpenAI. [SPEAKER_01] [SPEAKER_02] So with that, please welcome to the stage the claw father, Peter Steinberger.

16:59

SPEAKER_00

[SPEAKER_01] [SPEAKER_02] Peter, take it away. [SPEAKER_01] [SPEAKER_02] Okay, thank you all. [SPEAKER_01] [SPEAKER_02] Thanks, Al. [SPEAKER_01] [SPEAKER_02] Good morning, everyone. [SPEAKER_01] [SPEAKER_02] And we think it's a very unique moment for all of us to reimagine how we work and how we build. [SPEAKER_01] [SPEAKER_02] And so we wanted to bring a special guest who has been exploring what's possible with agents and really has pushed us to be more AGI-pilled at OpenAI. [SPEAKER_01] [SPEAKER_02] So with that, please welcome to the stage the claw father, Peter Steinberger. [SPEAKER_01] [SPEAKER_02] Peter, take it away. [SPEAKER_01] [SPEAKER_02] Okay, thank you all.

17:29

SPEAKER_00

[SPEAKER_01] [SPEAKER_02] Thanks, Al. [SPEAKER_01] [SPEAKER_02] Good morning, everyone. [SPEAKER_01] [SPEAKER_02] I love this picture because it reminds me how much has changed in a few months. [SPEAKER_01] I was juggling 10 or more terminal windows, always waiting for one of them to finish so I could steal the agent and queue new work. [SPEAKER_01] In January, that felt like peak productivity. [SPEAKER_01] Today, it feels a little bit silly. [SPEAKER_01] I thought I was orchestrating.

18:00

SPEAKER_02

[SPEAKER_01] Really, I was polling. [SPEAKER_01] I was the scheduler, the router, and the memory. [SPEAKER_01] At first, I paired with one agent. [SPEAKER_01] With 10 terminals, I was no longer pairing. [SPEAKER_01] I was managing 10 direct reports. [SPEAKER_01] Now, I mostly talk to a long-running manager, which delegates work to a team. [SPEAKER_01] For tricky work, I can still drop down and pair directly with the worker. [SPEAKER_01] But my default changed.

18:50

SPEAKER_02

[SPEAKER_01] I managed the manager of a small company of agents. [SPEAKER_01] Three changes made that possible. [SPEAKER_01] Number one, server-side compaction made long-running tasks reliable enough that I stopped optimizing around sessions.

18:56

SPEAKER_02

[SPEAKER_01] Coordination lets one thread create and steer the right projects. [SPEAKER_01] And third, automation can make the same manager when something happens. [SPEAKER_01] So we have persistent context, delegation, and triggers.

19:16

SPEAKER_01

There's your loop. And once the loop starts working, you discover the next problem. The bottleneck keeps moving. Last year, I was primarily constrained by tokens. Now, I fixed that by joining OpenAI. I know, the strategy does not scale. Then, my constraint shifted to token compute. All these threads run at the same time, and my MacBook starts sounding like a jet engine. That's mostly fixed by using test boxes, so agents can run tests on a separate machine. Now, I'm primarily constrained by attention. And unlike tokens or compute, I can simply add more of it. So the most important skill today is deciding where to spend it.

20:02

SPEAKER_01

Are you still staring at the agent while the code flies by? I know, it feels cool, but with the earlier models, this was necessary. You see the agent going in a direction you don't like. You hit escape. You steer it. You steer it back. But the latest generation of models is so good at understanding intent that it's a waste of time to watch the agent generate code. Imagine someone files an issue on one of my open source projects. The manager wakes up, reads it against the project's goals, notes, and vision, and decides whether it might be a fit. If it does, it creates a worker.

20:53

SPEAKER_01

That worker investigates, implements the change, runs the tests, and another agent can review the result. I don't need to watch those agents work or consume every intermediary message. When the manager needs me, it returns a PR, the original issue, the proposed diff, maybe a video or even a running build I can VNC into. I review once, I leave a note, I maybe approve, the loop continues, and it can land after the checks pass.

21:19

SPEAKER_01

The agent runs the inner execution loop. I set the direction, and I make decisions in the outer loop. Paul Salt is already running a version of this. He pinned his chief of staff. That worker investigates, implements the change, runs the tests, and another agent can review the result. I don't need to watch those agents work or consume every intermediary message. When the manager needs me, it returns a PR, the original issue, the proposed diff, maybe a video or even a running build I can VNC into. I review once, I leave a note, I maybe approve, the loop continues, and it can land after the checks pass. The agent runs the inner execution loop.

21:57

SPEAKER_01

You know, last year, I was primarily constrained by tokens.

22:09

SPEAKER_01

Now, I fixed that by joining OpenAI. I know, I know, the strategy does not scale. Then, my constraint shifted to token compute. All these threads run at the same time, and my MacBook starts sounding like a jet engine. That's mostly fixed by using test boxes, so agents can run tests on a separate machine. Now, I'm primarily constrained by attention. And unlike tokens or compute, I can simply add more of it. So the most important skill today is deciding where to spend it. Are you still staring at the agent while the code flies by? I know, I know, it feels cool, but with the earlier models, this was necessary. You know, you see the agent going in a direction you don't like.

23:20

SPEAKER_01

You hit escape. You steer it. You steer it back. But the latest generation of models is so good at understanding intent that it's a waste of time to watch the agent generate code. Imagine someone files an issue on one of my open source projects. The manager wakes up, reads it against the project's goals, notes, and vision, and decides whether it might be a fit. If it does, it creates a worker. That worker investigates, implements the change, runs the tests, and another agent can review the result. I don't need to watch those agents work or consume every intermediary message.

24:12

SPEAKER_01

When the manager needs me, it returns a PR, the original issue, the proposed diff, maybe a video or even a running build I can VNC into. I review once, I leave a note, I maybe approve, the loop continues, and it can land after the checks pass. The agent runs the inner execution loop. I set the direction, and I make decisions in the outer loop. You know, Paul Salt is already running a version of this. He pinned his chief of staff. That worker investigates, implements the change, runs the tests, and another agent can review the result. I don't need to watch those agents work or consume every intermediary message.

24:52

SPEAKER_01

When the manager needs me, it returns a PR, the original issue, the proposed diff, maybe a video or even a running build I can VNC into. I review once, I leave a note, I maybe approve, the loop continues, and it can land after the checks pass. The agent runs the inner execution loop.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note