Full Workshop: Setting Yourself Up for Success —Jason Liu, OpenAI Codex
Description
Jason Liu walks through how Codex works as a general tool for controlling your computer: setting up a memory vault and assistant threads, prompting it to collaborate with other threads, exploring computer use, thinking about long-running work streams, and preparing to work in loops. Speaker: Jason Liu — Developer Experience, OpenAI Jason helps developers get more from Codex, the Agents SDK, and the OpenAI API. Before OpenAI, he created Instructor and taught developers how to build reliable AI applications. X/Twitter: https://x.com/jxnlco LinkedIn: https://www.linkedin.com/in/jxnlco Website: https://jxnl.co/
Summary
Generated by gpt-5.6-terraAt-a-Glance
- Verdict: Watch fully
- Core thesis: Codex becomes substantially more useful when treated not as a chat-based coding assistant but as a persistent, context-rich operating system of skills, memory, scheduled threads, and computer-use agents.
- Why it matters: The workshop provides a concrete operating model for building long-running agent systems: durable project context, self-improving procedures, monitor-to-worker orchestration, verification-driven goals, and permission boundaries.
- Best use: Watch to extract reusable patterns for Ken's agent control plane and personal/team AI operations, while discounting product-specific feature claims and adopting stricter security controls than the speaker uses.
Executive Summary
Jason Liu presents an unusually practical vision of Codex as an agentic work environment rather than an IDE. His core workflow is to centralize persistent organizational knowledge in a personal monorepo, use long-lived pinned threads as functional teammates, and give those threads access to relevant connectors, skills, memory, and scheduled wake-ups. He argues that improved context compaction makes the former best practice of constantly starting fresh chats increasingly obsolete.
The operating model has three stages: bring in context through voice dictation, connectors, app shots, and files; conduct work through persistent threads, skills, sub-agents, goals, and loops; then take action through messaging, documents, browser automation, and desktop computer use. Liu's strongest practical recommendation is to capture procedures as editable skills, let them update when they learn from failures, and promote personally proven skills into team plugins only after sustained use.
The workshop's most relevant architectural idea is hierarchical thread orchestration. A monitor thread watches incoming signals such as support complaints, Slack activity, pull-request state, or launch information; it creates or messages specialized worker threads; those workers maintain state, await dependencies, escalate recurring issues, and report concise updates to the human. This is a concrete pattern for turning ambient information flow into managed workstreams rather than one-off assistant requests.
Liu is enthusiastic about computer use and broad permissions, describing agents that fill forms, check in for flights, route support, edit iMovie projects, and communicate through browser interfaces when connectors lack capabilities. He does acknowledge that this creates real security risk: a determined model can bypass a connector restriction by opening a browser and clicking Send. His mitigation is a combination of per-project instruction files, auto-review/permission settings, and organization-level restrictions, but the workshop does not offer a rigorous threat model, audit framework, or reliable evaluation approach.
Key Takeaways
- Claim: Persistent, long-lived threads can now function as durable agents because context compaction is good enough to preserve multi-week project continuity. | Evidence: Liu reports threads five weeks old with roughly 400 sub-agents that still know their responsibilities; he says the old advice to start a new thread after about 20 messages or isolate every task in a new conversation is no longer generally correct. | Implication: Ken should design agents around named, persistent workstreams with explicit state and memory, rather than defaulting to stateless prompt-per-task execution. | Caveat: This is a speaker report about current Codex behavior, not a demonstrated benchmark; he also notes uncertainty around memory behavior and cross-project context bleed.
- Claim: A personal or team memory vault is the foundation for increasingly implicit agent behavior. | Evidence: Liu organizes a git-backed monorepo with project directories, person records, loose notes, summaries, and a managed to-do list; project files include relevant Slack channels, while person files contain emails and identities. He reviews accumulated agent changes with git diff. | Implication: Use a versioned, inspectable knowledge layer as the durable source of truth for agents, with project-scoped instructions and metadata instead of relying on opaque chat memory alone. | Caveat: He considers the projects and people directories much more useful than exhaustive daily summaries, and says a fully automated task-verification system can be token-expensive.
- Claim: Skills should be lightweight, self-improving operational procedures that start as personal experiments and mature into shared organizational infrastructure. | Evidence: Liu describes skills as files plus scripts and plugins as collections of skills. He rapidly one-shots personal skills, corrects failures in place, allows skills to edit themselves when they learn, and shares them with a team after months of successful personal use. Examples include PR finalization, incident triage, communications routing, review styles modeled on colleagues' prior PR feedback, and a personal writing-style skill based on six months of email and Slack history. | Implication: Build an agent skill lifecycle: capture a repeated workflow after doing it once, log/correct errors, version it, gate high-impact actions, and only then publish a team-grade version with explicit ownership and evaluation. | Caveat: He does not claim to have robust evaluations for skills that span many live connectors; for team-shared skills he applies more scrutiny than for personal ones.
- Claim: The most scalable orchestration pattern is monitor threads that detect signals and create or coordinate specialized worker threads. | Evidence: For support, Liu describes a monitor detecting repeated Twitter or Slack complaints, creating a triage thread, routing evidence to the relevant engineer and channel, checking status hourly, and notifying users once resolved. The monitor can recognize recurrence and message the existing downstream thread rather than opening duplicate work. | Implication: Ken can implement a supervisor-worker control plane in which monitors own intake, deduplication, scheduling, and escalation, while scoped workers own resolution and verifiable completion. | Caveat: This pattern depends on reliable issue identity, thread state, and escalation rules; the talk offers no quantitative evidence of false-positive, duplicate, or missed-escalation rates.
- Claim: Scheduled heartbeats and verification-based goals are the mechanisms that turn agents from reactive assistants into long-running workers. | Evidence: A heartbeat schedules a message back into the same thread, such as keeping a pull request rebased, CI-green, and responsive to review feedback. Slash goal repeatedly checks a verifier until completion; Liu reports using it for code migrations with unit tests as the verifier. His UltraGoal pattern externalizes goal.md, plan, and optional state/work-log files so goals can evolve during execution. | Implication: Every autonomous workflow should specify a verifier, state location, cadence, stopping condition, and escalation threshold; use adaptive polling rather than a fixed high-frequency loop. | Caveat: Long-running loops can waste substantial tokens or repeat empty status messages; Liu recommends explicit stop criteria, dynamic heartbeat frequency, and replying only 'no updates' when nothing changed.
- Claim: App shots and computer use materially expand what agents can do, but they also undermine simplistic connector-based safety assumptions. | Evidence: Liu says app shots include an application's accessibility tree in addition to pixels, allowing an agent to identify Slack channel and user IDs and reduce multi-hop lookup work. He uses computer use for forms, document signing, browser actions, iMovie editing, flight check-in, and file uploads when a Slack connector cannot upload files. He explicitly warns that an agent blocked from emailing through a connector may instead open Chrome and press Send. | Implication: Treat GUI automation as a privileged capability with policy enforcement at the desktop/browser and identity layers, not merely at individual tool connectors; require audits and approval gates for external, financial, legal, or destructive actions. | Caveat: The speaker's stated safeguards—auto review, per-project AGENTS.md instructions, and organization-level outbound restrictions—are useful but insufficient as a complete security model for credentialed desktop agents.
- Claim: Human leverage shifts from crafting perfect prompts to choosing the right work artifact, agent topology, context, and level of model effort. | Evidence: Liu often dictates long, messy voice memos and asks the model to formulate the operational prompt itself. He says his remaining 'taste' is deciding whether work should become a goal, skill, new thread, HTML artifact, document, or project record. He also recommends low or medium reasoning for routine automation rather than using maximum reasoning by default. | Implication: Optimize the operating interface around rapid context capture and deliberate workflow design, while routing routine tasks to cheaper/faster models and reserving high reasoning for genuinely difficult or high-stakes work. | Caveat: His claim that dictated speech is roughly three times faster than typing is directional and user-dependent; messy input only works if the agent can search the relevant connected context safely and accurately.
Detailed Brief
Concrete workspace and context design
- Claims: Liu uses a single personal monorepo as the only primary Codex project, while allowing Codex to operate on code outside that directory.; A top-level AGENTS.md can direct where external projects should be saved, allowing the memory vault to remain separate from implementation repositories.; Each project can maintain its own README and AGENTS.md to constrain local conventions and reduce cross-project contamination such as npm versus pnpm assumptions.; Pinned threads are intentionally organized for human cognition as much as model execution: Liu compares distinct threads to distinct functional roles such as chief of staff, project owner, CLI/open-source work, and social channels.
- Evidence: His global instruction says not to save code in the monorepo and instead save it under /dev; Codex then knows where to clone projects and where slide work lives.; He keeps a projects directory for active workstreams and a people directory for people who have contacted him, including what they work on and related channels.; He proposes having Codex read all pinned threads, rename them, pin relevant ones, and use emojis to signal readiness.
- Caveats: Using a single vault/project can interfere with sidebar git-review behavior, which Liu accepts because he reviews code in GitHub.; Liu does not offer a definitive answer for memory leakage across projects; his practical remedy is local instruction files.
- Implications: The agent workspace should separate durable operational context from code artifacts while retaining links between them.; Thread taxonomies, naming conventions, and status signaling are control-plane UX requirements, not cosmetic organization.
Artifacts, communications, and human review
- Claims: Agents can prepare human review queues rather than only executing or drafting in place.; Codex can produce and edit artifacts including spreadsheets, Word documents, PDFs, slide decks, and HTML applications; Liu believes sharing small apps is becoming preferable to circulating full documents in context-dense organizations.; A chief-of-staff agent can consolidate connector data into a concise daily or weekly briefing with direct links back to source emails and Slack messages.
- Evidence: Liu configured an agent to open a Chrome tab for every email requiring a reply, with drafts ready for review after meetings.; He built the workshop slide deck in Codex, annotated unwanted whitespace or structure while presenting, and had Codex revise it in the background.; He previously had Codex generate a Wednesday-night PowerPoint of the week's shipping schedule, then generalized it into a channel update useful to the wider team.
- Caveats: The transcript shows slides being corrupted or misordered during the live presentation, illustrating that generated artifacts still need review.; Drafting in a personal style can improve throughput but should not be conflated with reliable judgment, especially in external communications.
- Implications: Prefer agent workflows that generate reviewable decision artifacts and deep links to evidence, not summary text detached from its source.; Use asynchronous preparation to turn human attention into final selection, judgment, and approval rather than clerical navigation.
Adoption posture, evaluation, and cost management
- Claims: Liu recommends learning agent workflows through direct use and iterative correction rather than attempting full formal evaluation before personal deployment.; He sees organizational impact in becoming a 'plugin hero': building reusable capabilities that teammates adopt, rather than maximizing one's own token usage.; He believes future knowledge work will reward taste and vocabulary: the ability to recognize quality, diagnose poor artifacts, and articulate what should change.
- Evidence: When asked about evaluating skills, he says he generally 'YOLO one shots' personal workflows because connector-rich systems are difficult to snapshot and evaluate; he treats two months of successful use as meaningful evidence before team sharing.; He says one of OpenAI's most-used skills runs before PR review and catches style-guide or implementation issues.; For token efficiency, he advises using low or medium reasoning for ordinary work and cites a chief-of-staff thread configured at medium effort.
- Caveats: Personal trial-and-error is inappropriate as the sole assurance mechanism for workflows involving confidential data, money, legal documents, external messaging, or production changes.; The workshop is product advocacy and anecdotal practice, not an independently validated account of reliability, cost, privacy, or security.
- Implications: Separate low-risk exploratory agent deployment from governed production automation, and establish different standards for each.; Measure shared-skill value through adoption, error reduction, cycle-time improvement, and avoided coordination work—not token volume.
Notable Concepts & Terms
- Compaction: The claimed improvement in preserving useful context across very long Codex threads; it underpins Liu's recommendation to keep durable project threads instead of constantly restarting chats.
- Pinned threads: Long-lived, named Codex conversations treated as persistent teammates or functional owners; they can be scheduled, connected to memory, and—per Liu—message other threads.
- Heartbeat / loop: A scheduled wake-up that injects a message into an existing thread so it can recheck conditions, continue work, poll dependencies, or report changes.
- Slash goal / UltraGoal: A persistent objective with a verification criterion; UltraGoal moves goal, plan, and optional state into editable files so scope and understanding can evolve during long-running execution.
- App shots: A context-capture feature that Liu says supplies both an image and the application's accessibility tree, making UI state and follow-on tool calls more structured than a screenshot alone.
- Computer use: Desktop GUI control for applications and browsers, enabling tasks beyond connector APIs but creating a much larger credentialed-action and policy-bypass attack surface.
- Personal monorepo / memory vault: A version-controlled file hierarchy containing projects, people, notes, skills, and operational state that agents consult and update as persistent organizational memory.
- Monitor thread: A supervisory agent that watches channels or signals, detects and deduplicates issues, creates or updates worker threads, and escalates unresolved or recurring work.
Operator Notes / Why Ken Should Care
- Prototype a version-controlled agent memory vault with a minimal schema: projects, people, active goals, source links, and per-project policy/instruction files; require every agent update to be reviewable through git diff.
- Implement monitor-worker orchestration for one bounded operational domain, with explicit issue IDs, deduplication rules, a state machine, ownership routing, retry cadence, and escalation SLAs.
- Define a production autonomy policy that distinguishes read/search, draft, internal write, external communication, legal/financial action, and destructive actions; enforce it below the agent layer so GUI fallback cannot silently bypass connector restrictions.
- Create a skill maturity process: sandboxed personal trial, failure logging, test fixtures or replayable evals where feasible, security review for privileged tools, named owner, and controlled team rollout.
- Require every heartbeat or goal loop to include a verifier, stopping condition, maximum budget, idle-response policy, and adaptive polling schedule to prevent token burn and alert fatigue.
- Test app-shot/accessibility-tree capture as a structured UI-state interface, but validate whether captured identifiers, sensitive text, and action targets are redacted and governed appropriately.
- Use low/medium reasoning defaults for routine monitor, summarization, form, and status tasks; route only complex planning, ambiguous diagnosis, and high-impact decisions to stronger reasoning settings.
Source/Metadata
- Title: Full Workshop: Setting Yourself Up for Success —Jason Liu, OpenAI Codex
- Transcript words: 14832
- Duration seconds: 4502
- Timestamp note: No timestamps or chapter markers were present in the supplied transcript.
Transcript
. Alright, let's just kick things off. How many people here already saw the keynote that I gave? Okay, not everyone. That's good. This talk is effectively going to be a stretched version of what I had given on the main stage, except two things. One, I want to give you some time to try to set things up yourself, Wi-Fi, God's permitting. And then also, two, be a little more interactive. I had to be very high level when I was talking about what I use Codex for, but here, if you have any questions, we have 70 minutes. If you have any questions, just raise your hand and we can start answering some of these things, especially because not a lot of the workflows have been really well documented, and so if you're very curious how things work, I'm really happy to answer any questions we have here. Yeah, I'm Jason, I work at OpenAI. I don't really know what my job is anymore. We do a lot of things. Clearly the slides are already in the wrong order as well. But generally, I've done things like doing a lot of prototyping work and just writing lots of code and just having a goal run for two days to build a game or some kind of web application, right? We've also looked at things like running evals and hill climbing things. I also use computer use to edit iMovies and make little videos. I do partnerships and education and operations by taking my meeting notes, turning them into documents, working with other vendors, and also working with different foundations and programs to get funding. And all of this work is done effectively in the Codex app. I don't know if you can tell, but right now, the slide is being served on localhost in the in-app browser of the Codex application. And so anytime I find something I don't like, I might just hit the annotate tool, give a comment, and have Codex clean up these slides. I'm not the biggest token maxer. I think I'm doing all right. I see some folks doing a couple billion every day. The goal of this talk isn't just to waste all your tokens, but really help you avoid wasting your tokens by telling you what has actually worked. And in particular, the tricks I use to make these things productive. And again, the goal, like I said earlier today, was to catch you up on what's changed in the Codex app, give you some time to set things up. So if you have time right now and you haven't downloaded Codex, just go ahead and do that. And then we can go a little bit deeper into setting things up. And so I've actually prepared a little monorepo that you can use to clone in, get all the skills that you need, get all the setup that you need, and go from there. So let's go a little bit deeper. If you're new, feel free to set things up. If you're pretty experienced, just chill out, try some of the things I'm talking about. And then I think every 15 minutes we'll have some time for questions and we can go into the more, what feels like AI psychosis, but maybe actually works domain of using these systems. A lot of the work and knowledge work now, because the coding is solved, because a lot of this operations work is solved, is really just understanding what you can do. In a world without AI, maybe I have 10 teammates, each teammate is working on one thing, so I need to have 10 things I'm keeping track of. Now we live in a world where everyone I'm working with has 10 projects. I now have to keep track of 200 things, and I don't know what's important. There's definitely a Slack message I've missed, there's probably some email I've missed somewhere by some foundation, and Codex helps me organize all of this stuff. And so again, the things I really want you to take away from this workshop is the fact that compaction works really, really well. I have threads now that are five weeks old that have 400 sub-agents in them, and they generally just know what they need to do, they know what their job is. I also want you to become really comfortable with talking to your computer. Earlier today I said that Tony Stark is not texting Jarvis. Right? And there's really no future when text input is the thing that matters. I basically use a foot pedal, so I have a button that is transcribed and a button that says enter, and so I'll just come by my desk with my hands behind my back and I just go, fix this, make this change, also message this guy on Slack. And then I just go back to talking to my coworkers and trying to figure out what is the human side of actually working at OpenAI rather than just monitoring Slack all day. AppShots is my favorite feature of all time. It's very satisfying. If any of you are on Codex right now, just press the command button side by side. You're going to get this real nice animation or you're going to get a modal to tell you to install computer use. Just do that. It's amazing. Invest in your personal memory. At this point, when someone asks me what I'm doing, I don't even have any idea. I have to look at my threads and look at the conversations to figure out how much I've delegated away and how much has been automated. And if you can then invest not only in skills for yourself, but also plugins for your entire team, you can become the superhero that actually augments the rest of your company. It's one thing to say, oh man, I can use all these tokens and look at how many tokens I'm using. But actually, if you're rewarded by how often the plugins you've built are being used by your teammates, that's a huge win. How are we doing implementation? One of the most popular skills is just the finalize the Codex app skill. And anyone who makes a pull request basically triggers the skill before review. And basically everyone at the company uses it. And it's always been able to find things that I've done wrong or that are against the style guide. One of the skills I have is just reviewing docs, and it's basically just copying the pull request reviews of our friend Charlie over here. And so I just have review my code like Charlie based off the past year of feedback he's given on pull requests. Review my code like Dominic. And these things are incredibly valuable. And then lastly, once you get more comfortable with all those first four things, your pinned threads with automations, these things that wake up these threads over time, they're going to feel like teammates. And more interestingly, now that threads can talk to each other, so every thread has the ability to list other pinned threads, has the ability to rename threads, and has the ability to send messages to each other, not only can you have teammates, but you can have teammates that work together, and you can effectively start having managers. And so you went from an IC enabled by an IDE, then you have pinned threads that feel like a team where you're the manager. And very quickly in the future, as models get better, this is where the puck is going to skate to, you're going to start having your manager threads, and then your IC threads. And I'm sure in the future there's going to be some other crazy orchestration. And all of this really is due to the fact that compaction works. Even six months ago, I don't know how many people here have been told this, but you were always told, if a conversation goes very long, start a new thread. After 20 messages, it's not going to be that good. Every feature should be its own conversation. If you do a code review, start a new session. Those things basically aren't true anymore, and a lot of it has to do with compaction. Just pin the thread, rename it to the project ID, and that project thread should be able to delegate to sub-agents, create new threads, and have conversations, and then write to your memory vault, which will allow you to just log what's happening. And then with automations, you can just wake them up. And so there's really three acts of working with AI, working in Codex. You bring the context in, and I'll talk about how you do that, and what are the ways you can bring context in. Then you work on it, right? For example, the slide deck is just in the Codex app. And then you take actions out in the real world. So, I asked this during the keynote, but I'm also curious what the audience here is doing. But how many people use dictation when they interact with an AI? Nice. How many use dictation even at work? Yeah. I think we should all be a little bit more shameless in doing these kinds of things. You generally talk about three times faster than you type, and it's just incredibly productive to be able to give the messy version of what you're thinking about to the AI, and take that extra time and just try to be even more thoughtful to the people you work with, right? You bring the context in, and I'll talk about how you do that and what the ways are that you can bring context in. Then you work on it, right? For example, the slide deck is just in the Codex app. And then you take actions out in the real world. So, I asked this during the keynote, but I'm also curious what the audience here is doing. How many people use dictation when they interact with an AI? Nice. How many use dictation even at work? Yeah. I think we should all be a little bit more shameless in doing these kinds of things. You generally talk about three times faster than you type, and it's just incredibly productive to be able to give the messy version of what you're thinking about to the AI and take that extra time and try to be even more thoughtful to the people you work with, right? It's now, I don't want to send my coworker a 15-minute voice memo, but I should feel very comfortable sending an AI a 15-minute voice memo. Because you're going to include some random tangents. You might just say, I'm pretty sure I had a meeting with Charlie sometime last week about the agent's SDK. And it will go and read 35 meeting messages to figure out which one it was and make it relevant, and now all of a sudden, whatever memo you're going to write or some project tracker that you're trying to do is going to work, right? But I would never do that with AI. So, as you guys are doing, just listening to this talk, try to set up Codex the way that I've been describing these things, right? So, once you have your ability to have you input into the machine a lot more effectively, you can start thinking about using things like skills and plugins. Skills is a very simple construct. It's just a couple of files and some scripts. A plugin is a library of these things. And as you are doing things many, many times, you can start thinking about creating your own skills. And as you package a bunch of skills, you might start thinking about building out a plugin. If you want to install the plugins, we have a pretty good ecosystem now. It's something I'm really proud of. If you just go in the sidebar, click Plugins, you can search whatever plugins make sense for you. So, if you use Slack, you can install Slack. If you use Gmail, Teams, most of these things are pretty built out. If there's something that you feel like you are missing, just @ me on Twitter. And I'm sure one of my Twitter monitors will pick it up and send a message to someone on the connectors team. If you're already actually looking at the plugins panel, I also really recommend starting the process of setting up the Chrome extension as well as computer use. We'll talk about this a little bit later, but computer use was the first time in a long time I really felt the AGI of being at work. I was in iMovie for the first time. I didn't know how to use it. And it was just teaching me how to export the movie. It was able to figure out where the sound effects were, and it placed it in the right timestamps. Really small things like this really make using a computer very fun again. I don't really have the time to learn new software, but if Codex is going to show me what's going on, it's pretty awesome. And as you see the cursor move, oftentimes you're cheering for it to do the right action. If you don't have a link for the Chrome extension, you can just click this button here. The difference with computer use is computer use can work behind the scenes to control any application. Right? So whether it's Slack or some trading software, God forbid, it can control all of those things. With the Chrome extension, it just controls everything in the Chrome app. But the cool thing here too is that, again, it doesn't take over your screen. Right? Sometimes I'll just be working on my computer, I'll go to the Chrome browser, and I'll realize that Codex has just opened up three tabs to look at my Twitter DMs and then closes them back up as I'm just responding to some other email. It's really cool to watch these things work in the background. You can connect to a bunch of other plugins. I use things like Notion, Linear. I also use Obsidian. It's just a good time. Once you do this, what you're going to find is just by asking really vague questions about your day and just tagging the right plugins, you're going to realize the AI can learn a lot about you, right? The AI does not assist them now where it does one search request and tries to come up with an answer, right? It might check your emails and find a loose thread. It might check Slack or some meetings and figure out what's actually going on. Who are these people? I had one of my loops basically realize I was meeting with somebody, look at their LinkedIn, and realize that we went to the same university at the same time. And so the moment I jumped on my call, I was like, hey, you're also from Waterloo. Do you remember this, this, and this person? And immediately we had a connection, right? And obviously I didn't tell them it was AI, but that's some of the small things that you can do by just improving your automation. It can make you closer to people. As you build out your memory system, right? As you build out the Codex memory system and your ability to trigger plugins, maybe day one you have to tag everything. But I've become a worse and worse manager over time, right? Now I'll just open up the composer and say, what has changed about the launch? And they'll be able to do a good job, right? And that's possible because you have this long history. You have all these pinned threads. You have these memories, right? It's the same thing with an employee. Day one you have to show them every standard operating procedure. But at some point you have an employee that has been here for seven years, and you can just say, hey, I think you should make the company more money. And they can figure it out. But it's only because they have this context. One thing you can do, for example, if you want to try it out, you can just say, hey, check out the schedule, find all the sessions, organize them in a markdown file, put them in a spreadsheet. And you'll just realize that we can do these things. And maybe you'll do it with web search. Maybe you can do it with Chrome. Lots of fun things here. If you want to get inspired by looking at what kind of skills exist, we have two really great sources. One is if you just run the skill installer skill, it will actually list out all of the OpenAI-curated skills. These include ones for things like GitHub, best practices when writing Playwright code, ReMotion, for example. But you can also check out websites like skills.sh or use, I think this is Vercel's skills tool. And then you can find other skills. Right? So if I'm thinking about doing some more motion design or web design, or I know that someone told me I shouldn't use memo in React, but I don't really know what that means, I can now go install the React best practices skill. Right? But again, internally one of the highest-impact things I think you can do as the AI champion in your company is to figure out what the team needs and build out those skills. Right? I have a lot of skills on doing things like triage and how you do comms. Right? If there's an outage on Twitter, how do you convert that? Figure out who needs to hear this. How do you start the sev? What stat-sig gates do you need to check? All these things are now just automated. And that's exactly what I just said in this slide. We also have a really good plugin creator and a skills creator skill. So if you just ask Codex to trigger it, it will try to interview you to figure out what's going on. And even in a more useful way, you can also just do it yourself once, document everything, and tell Codex to make a skill from what you've learned. Right? And as long as you tell it, hey, by the way, every time you run this skill, you're allowed to edit yourself if you learn something new. You can edit the skill file. These things will also improve over time. And a big theme that's happening over this talk really is you have to get really comfortable with asking. We'll obviously try to make more of these things more slash commands. But more and more, I'm just not touching a computer, so it doesn't even make sense for me to run a slash command. I just want to say, what's launching this week? Check Twitter. Look at what I'm seeing in the browser. The example I've been developing internally has just been this developer experience triage skills. So, again, this skill just documents every Slack channel that you should be aware of. Right? And as long as you tell it, hey, by the way, every time you run this skill, you're allowed to edit yourself if you learn something new. You can edit the skill file. These things will also improve over time. And a big theme that's happening over this talk really is you have to just get really comfortable with asking. We'll obviously try to make more of these things more slash commands. But more and more, I'm just not touching a computer, so it doesn't even make sense for me to run a slash command. I just want to say, what's launching this week? Check Twitter. Look at what I'm seeing in the browser. The example I've been developing internally has just been this developer experience triage skills. So, again, this skill just documents every Slack channel that you should be aware of. It knows which engineers have worked on what projects. It knows what Slack channels are taking in feedback. I know that if you DM me on Slack and you tell me that some regression has happened, I need to ask for a feedback ID. Now the agent does this automatically. And it does it automatically with app shots. So, again, I don't know how many times I'm going to say this, but app shots is one of my favorite features. How many people here have sent a screenshot to Slack? To Codex? Right? Almost everybody. But the issue is the screenshot does not have that much information. The model has to then do OCR. And if you send a screenshot of a Slack thread, the model has to read the Slack thread and then do a list Slack channels function. And then realize there's a guy named Charlie and then do a list person. It takes a lot of hops. But with app shots, it takes not only the image, but the entire accessibility tree of the app. And so when I give it an app shot of a Slack channel, it knows the channel ID. So it knows exactly what function to call to post there. It has the user IDs of every single person in that channel. So if I take an app shot and say do some research and reply, it's only one function call. It knows to send a message to channel U12725. And then because it knows that Charlie is U425, it can do that in a very fast hop. So not only is it a very quick way of getting context into your system, it just gives so much more context that these subsequent tool calls do a really good job. I have not filled out a form in two weeks. Because I just now tell Codex to fill out this form. It knows all the fields. It then figures out that it's in Chrome, and so it'll use the browser extension. If it's in Safari, it'll use computer use. The model has become really, really intelligent. And so just like you might have a manager that gets an email and they forward the email to me with three question marks, and it's your job to figure out what's going on, you can start doing that with your AI as you start investing in these skills. And most of this is because of the fact that you've built out your memory system. So if you guys are taking a look at these slides, jxnl slash personal monopreneurator. I have a repo template. That is actually the template I used on my personal computer. It's basically just a directory tree and a bunch of skills that I use to grow out my memory. I'll also make one call out, which is if you open this in your browser, just press app shots and tell Codex to set this up for you, and then you can pay attention to the rest of the talk. Yeah, yeah, yeah, yeah. You can just tell Codex Jason has written a personal monorepo template on GitHub. Please find it and then install it. All right. Yeah. Yeah, so this is a really good point. So, for example, on the DX team I make a lot of demos. And so I have 16 repos. Real time demo one, real time demo two, funny, right? You have all these demos. And so I don't create new projects for them. Right? The only project that exists on my sidebar is the personal monorepo sidebar. But Codex is able to still manage files outside of that project directory. And so in my Agents.md file I just say, don't save any of the code in the monorepo. Save it in slash dev. And just by that one line, if I tell it to clone a new project, it saves it in slash dev. If I tell it that I want to work in my slides, it knows that there's a slash dev slides directory. But it's just an easier way of managing everything. I want to start all my projects from my personal vault, and then it can touch the file system in any way that it wants to. One call, it kind of breaks git review sometimes on the sidebar. But generally it's been a pretty good experience for me. Because I just review my code in GitHub. Sweet. These are some of the skills I have installed there. There's no need to take a photo. Just ask Codex afterwards. But the assistant plugin basically has the ability to onboard you. It will interview you. It will figure out what plugins you need to install. And then it will actually go create the threads it thinks it needs. It will create the automations. It's a pretty fun one. I have a bunch of skills in auditing AI code and AI writing. I don't include this, but one of my favorite skills of all time is called write like me. And if you want to make one like that, all you have to tell Codex is, Hey Codex, I want you to read all the emails I've written in the past six months, all the Slack messages I've written in the past six months, and write a style guide for how to message just like me. And then that's it. And then any time I tell it to send a Slack message or write an email, it will go, Okay, this is an email. Clearly this is a custom support form. So I will be much more stern in my messaging. Let me go draft this email. Hasn't failed me yet. One thing I've also added that I think is really valuable to call out is, I've made my own loop skill just because I do like having a slash command every once in a while. I'll talk about this in part two. And I also have a skill called simple HTML artifact that just designs artifacts the way I like them. I want my background to be white. I want some certain style guides. And then UltraGoal, which is a super version of Goal that we'll also talk a little bit more about. And then if anyone's curious, new person, new project, that's just a way of running a script to bootstrap a new person. I have a palantir for my personal life now. It's just a CRM. And basically any time my AI agent finds a new person that's emailed me or messages me on Slack or on iMessage, I just keep track of these things. And the new project is the same way. Let me just double check. Yeah, cool. And so I'll give you maybe 10 minutes to try to set this up, and we can go into a little bit of a Q&A. I'm happy to answer any questions about how we bring context into our systems, how I've organized my personal memory vault, and some other crazy uses of app shots if anyone has any questions. Yeah, what's your question? Does this get rid of the need for an obsidian brain that can be by itself because you have this compacted nature that's really good now? Yeah. So the question was, do I still basically use obsidian brain? The answer is yes, because I still want to keep track of everything. One thing I actually really like doing is I make my monorepo vault a git repo. And so maybe it'll work on it for a couple of days. And I'll come back. And I'll just run git diff. And by running git diff I can just see what the model has updated and what the model has not updated. And I can just confidently review that over time and just realize, oh yeah, I guess Charlie did respond to this person and closed the loop. And I didn't realize that, but now I know. Right? And oftentimes that's relevant in another conversation. More than that, it's also very helpful for when other people are asking me questions. Right? Codex will feel very good about reading my memory vault, drafting a response, and then ask me for permission to send that message off. And so if someone messaged me on Slack a question that the AI could have answered, the AI will just try to answer it. And it might be simple things like, oh, who should I talk to about this project? Right? And the model knows because it's in the memory vault. One thing I'll also call out is if you want to use more tokens, you can also have custom automations where the job is to be able to use more tokens, just to maintain and manage and garden your memory vault. But generally that has not been a big issue for me. One question over there. Yeah. And I didn't realize that, but now I know. Right? And oftentimes that's relevant in another conversation. More than that, it's also very helpful for when other people are asking me questions. Right? Codex will feel very good about reading my memory vault, drafting a response, and then ask me for permission to send that message off. And so if someone messaged me on Slack a question that the AI could have answered, the AI will just try to answer it. And it might be simple things like, oh, who should I talk to about this project? Right? And the model knows because it's in the memory vault. One thing I'll also call out is if you want to use more tokens, you can also have custom automations where the job is to be able to use more tokens, just to maintain and manage and garden your memory vault. But generally that has not been a big issue for me. One question over there. Yeah. What do you think about doing eval skills that you're creating, or I guess I'm going to review automated eval. That's a thing called a mining tool to better manipulate this. Yeah. But first you just YOLO one shot. Honestly, I generally go down the path of YOLO one shot, but only because, oh, that scared the hell out of me. Only because I know that the way I build my skills is that they self-improve all the time. Right? I think the difference would be if I make a skill that I share with my team, I think about that a little bit more. Right? Because it's like, okay, does the triage plugin know who is working on what feature? Can it route correctly? For my personal work, I generally just build a skill as quickly as possible. And every time it makes a mistake, I just correct it and I tell it to move on. And then generally what happens is if I've used a skill for two months, I just generally feel pretty good about sharing with my team, because I've just experienced it working. Yeah. Most of my skills connect to so many other plugins and connectors that I just don't know how to eval that, because I can't snapshot my Slack at any given time. The question over there. Are you pretty easily using the desktop app or using the codec CLI? Sometimes I'll use the codec CLI every once in a while if I want to be a little bit faster. But generally, the desktop app has a pretty good experience, primarily because everything I do is an app shot. If I'm watching a video, I'll just app shot, summarize this, and I'll continue to watch the video. I'll just watch the video with the LLM summary. Right? Or if I see some kind of form or someone asking to sign something, app shot, use DocuSign, sign this and save it on my desktop. Last week, it DocuSigned something and then found a faxing service and faxed my medical records. That's awesome. But the CLI can't really do that. Yeah. Any other questions? Yep. What do you do while I'm waiting for a day? I'm learning to juggle. I'm learning to play the drums. Well, I think it's two things. In the office, what I'm trying to do is I'm trying to talk to more people. Right? It's like I'm just the AI's assistant to get more context that the AI can't get. No, I think my job when I'm at work really is just, when the AI is running, I don't know, I should be talking to somebody. I should be learning about what they're working on and trying to make connections, and then figure out what are the points of connection. Right? I should be talking to more people in real life as AI works. I think someone had a question over there. Oh? I had a question about . Yeah. So the question is when do I think 5.5 is overkill versus 5.3 Spark? I think this is colored by two things. Because I have unlimited tokens, I don't really make those decisions. And then secondly, because I'm not watching my AI work, most of these things are automations that run in the background, so the latency has not really affected me. The times where I use Spark is primarily when there's a really simple computer use task. Right? I just wanted to click all the buttons and fill out this form. I think I have a thing that just checks me into flights, and that is a Spark agent. And so now any time there's an email that's a flight check-in, my agent will check me in, download the boarding pass, and then send the boarding pass to myself on iMessage. Right? And I just never do that kind of stuff anymore. Again, it is really weird when you're just working, and all of a sudden you check your Chrome desktop and it's just JetBlue as the first page. But again, it speaks to the fact that having access to your computer is uniquely powerful because it has your auth and your credentials and your file system. Yeah. Over there? Yeah. So you're really quiet. Do you mind just speaking up? Yeah. I'm asking if you want to do long-running . Yeah. Yeah. Yeah. How do you actually, do you actually do that if you have to create a different device, or it's only cloud and one part of it? Yeah. So I have, we'll talk about this in the next section, but with remote control I can control both my local computer and a remote computer. I think the difference is computer use is tricky because it has to be on your computer. That's one thing. The second thing too is if you go into settings, computer use, there's a flag called locked use. And if you enable that, as long as your laptop is plugged in, even if the monitor is closed, you can still trigger computer use commands through your phone. That also gets really weird. Because, go on. Yeah. And I was going to say, my question is you have two . Yeah. So what are you for? Take your laptop, grab them, which one . You're really quiet. I can't hear you. Yeah, I'll hear you. Okay, okay, cool. So, yeah, if you have multiple computers, you can still connect both your iPhone to both those devices, right? Some people just have a Mac Mini. I think the difference is do you want to control your codex or specifically computer use? Because that requires an operating system with a GUI. But we can talk about this in a little bit. Cool. Like I said before, were there any more questions? I don't know if someone just raised their hand. Okay, go ahead. Is there an elevated risk in the other user? And how do you control that? Yeah. I mean, there's always some kind of risk. I think earlier versions might edit a document a little too eagerly, right? But realistically, I think these models have done a really good job of being very precautious. And more often than not, it's me going, no, please just do it. Just sign the document. Please just send this message. I found that the 5.5 models are pretty reluctant to take these destructive actions. That's one thing. The second thing is, if you look at the sidebar here, whoo, you have the ability to change your permissions. And so, as a show of hands, how many people use ask me for every permission? Yeah. Yeah, exactly. Okay. How many people use full auto, full permissions YOLO mode? Okay, I don't like that. And then how many people have used auto review? Yeah. So, I think auto review has been really, really great. And again, I'm usually annoyed by the fact that my models won't do more than I want them to do. And so generally, it has not been as big of an issue. The only examples where I'm really annoyed is it will edit documents it shouldn't be editing. But I just add something to the agents empty, and it's never really messed with me too much. Yeah. I know that's not a real answer, but I think with a combination of auto review and agents empty file, I have generally felt pretty safe. Yeah. And then if you're also at an organization, there are different admin settings that you can have. So, for example, at OpenAI, you can't use an MCP server to send an email if any of the people in that email is a non-OpenAI email. Right? Or you can't send a Slack message to external Slack channels. That's when things get dangerous. Right? But these are some things that you can control. One question over there. Yeah. Do you have any concerns regarding either security and/or privacy? That's tough because I work here. I think that's hard to answer because I don't really know what our data retention policies are for individual versus enterprise. But Charlie, do you have any thoughts there? I'm just going to go over to you. Just concerns about security and privacy. Privacy. Yeah. And then if you're also at an organization, there are different admin settings that you can have. So, for example, at OpenAI, you can't use an MCP server to send an email if any of the people in that email is a non-OpenAI email. Right? Or you can't send a Slack message to external Slack channels. That's when things get dangerous. Right? But these are some things that you can control. One question over there. Yeah. Do you have any concerns regarding either security and/or privacy? That's tough because I work here. I think that's hard to answer because I don't really know what our data retention policies for individual versus enterprise are. But Charlie, do you have any thoughts there? I'm just going to go over to you. Just concerns about security and privacy. Privacy. Let's say you get an email offering you a job at a competitor company. I think a lot of it goes back to we do want to make the models are fallible. Right? I don't think anybody in this room would be shocked to understand that you can still jailbreak a model, for example. But both the models themselves are getting smarter and better at not doing silly things. And at the same time, we're figuring out, like Jason mentioned, what are those bigger limitations around the sandbox? We started with very simple sandboxes where it was like you can just run this command and nothing else. And slowly the sandbox has grown to the entire computer. And I think we're figuring out what are the computer-level or organization-level edges to that sandbox that we need to build. This is me using the delegate skill. Great answer. Any more questions before we jump into act two? Sweet. Awesome. So we just talked about a bunch of different ways of bringing context into the system, right? You can use your voice, you have plugins, you can use app shots, and then you can also design different skills and plugins to figure out how to do more systematic work. So now we can talk a little bit more about the work itself. So, like I said before, every pinned thread effectively is a teammate in my mind. I have my chief of staff thread. I have, Swix prefers to call it, the god thread. I have a thread to manage the agent's SDK, whether that's implementation and documentation. It has two sub-agents that it delegates to. The CLI, the open source program, and Twitter. But if you want to make it wake up, all you have to say is keep an eye on this until sometime. Keep an eye on this every 30 minutes. If you remember those secret words, you can effectively automate about everything in your life at this point. And what this does, it will trigger a heartbeat automation, a thread automation. You should think of it as a way of scheduling a message back into the thread, right? In the beginning, when we set up automations, it was very much the case that an automation would create a new thread every time. So it might be give me a morning brief, and it would create a new thread, and then it would do this kind of work. But as these models got better, I think the right design is scheduling these messages into the same thread. So, for example, because if you download the monorepo, we have a loop skill. If you just do loop, this is the equivalent of just saying keep an eye on this. Keep an eye on this pull request. Any time there's feedback, fix it. Make sure it's always mergeable. Make sure it's always rebased on master. Make sure that CI is always passing. And it'll just do that. And then maybe you make a pull request on a Monday. You get really busy. Thursday afternoon, all the feedback has been integrated. CI is passing. And you're not 4,000 commits behind the main thread. I also do this with support, right? Again, if someone is dealing with some issues on Twitter or on Slack, app shot, at developer experience skill, figure this out. And I'll say, okay, this is an issue on the browser side. James is the one that works on the browser. The channel is called browser feedback. I'm going to post in that channel, DM James, and then I'll use computer use to open up Twitter to let them know that I've escalated this internally. And then I will check every hour to figure out if James on that channel has responded, and then let the user know. And then sometime later in the future, you're checking your computer, and all of a sudden Twitter opens up. And it's just like, hey, so and so, this has been resolved. And then you just hit enter. And then you check your thread. It's like, oh yeah, a pull request has been made. It will get merged by next Thursday. And this is a crazy experience to witness. Right. This actually allows us to do way more support without making it the worst part of my job. And then with the chief of staff thread. Oh, I remember this is where my slides get really messed up thanks to a codex. So I can't do everything just yet. You can also just do a loop that says check all my connectors and give me an update as to what is the most important thing I should be thinking about. Give it to me in a nice format. Make sure you have links to every email that you read. Make sure you have a Slack link so you can deep link into the application. And now it's been really, really helpful to track these random things. And again, I think in my chief of staff thread, I had that line that says, check into all my flights if you can. Send me the boarding pass on iMessage. And it just works. And again, it's really weird when you start seeing your computer doing stuff while you're working. Because again, most of these tasks run in the background. And because I've had my permissions on, it's not like it needs to tell me that it's using Safari. I just find out that it's using Safari. That might scare some people, but it's pretty great. And then lastly, this is something that happens really a lot at OpenAI, which is I don't even know what's launching. Right? Is something delayed? Has something landed? Is it going to happen at 11 a.m.? Is it going to happen at 4 p.m.? I have no idea. We've actually made a ton of progress on this. And I'm pretty sure this is also just automation and AI. But I actually used to have Codex make me a PowerPoint every Wednesday night of what's shipping for the rest of the week. Right? And that's useful. And then you just say, great, it's useful for me. Now let me make sure I can just post it on the channel as a Slack message. And now you're, again, using your skills not only to benefit yourself, but benefit your entire team. That's the plug-in hero mindset. And this one's pretty funny, too. I have an example. Let me just double check this slide. Yeah. I have another example which has also been pretty crazy, which is the one I gave during the keynote, which is, I had been editing a short film in iMovie. And on my bike ride, someone gave me feedback about the video. So I just went on my phone and told the AI, okay, there's a file somewhere. Can you just find it in iMovie? There's only one iMovie project. Read the Slack message. Export the video. If the Slack MCP server does not allow file upload, use computer use to upload the file. And then watch that thread every hour. And if they have any feedback, re-export the video and re-share it. And then bike home. And by the time I got home, it was like, oh, by the way, it was actually way easier to use a Google Drive connector. So I've just been uploading the same file on Google Drive instead. So they only have one URL to manage. And yeah, we fixed a bunch of stuff in the typography and then we shipped it. Again, mind-blowing stuff. It's very simple but mind-blowing. Right? But that's what a heartbeat is. Right? A heartbeat is just a way of waking up your thread over time to take some actions. The other thing you can do is set goals. So, slash goal is pretty amazing. It basically defines a verification step and says, okay, as long as this is running, check this verification step. If it's not done, keep going. A very simple idea. And as long as there's a verifier, it does really, really well. For example, I've just been migrating a bunch of software into Rust. Right? It's like if this is a Python project that is amenable to be rewritten in Rust, slash goal, migrate the back end into Rust, make sure all unit tests pass. And, yeah, we fixed a bunch of stuff in the typography, and then we shipped it. Again, mind-blowing stuff. It's very simple but mind-blowing. Right? But that's what a heartbeat is. Right? A heartbeat is just a way of waking up your thread over time to take some actions. The other thing you can do is set goals. So, slash goal is pretty amazing. It defines a verification step and says, okay, as long as this is running, check this verification step. If it's not done, keep going. A very simple idea. And as long as there's a verifier, it does really, really well. For example, I've just been migrating a bunch of software into Rust. Right? It's if this is a Python project that is amenable to be rewritten in Rust, slash goal, migrate the back end into Rust, make sure all unit tests pass. And I was able to not only rewrite the rich terminal library in Rust, I also rewrote UV in TypeScript just to see if I could. And we're at 100% test coverage. It's pretty amazing. Obviously, you should not be doing this in your work, but it's very helpful to understand that as these systems have better verification, you can make a lot of progress. In the monorepo, I've also included a skill called UltraGoal. And all it does is, instead of setting the goal in the app, we set it in a file. Right? So, we have a goal.md file. And what that means is you can edit the goal while it's being run, so you can add more scope, just like many real projects do. We also define a plan, again, that we can reference. But the benefit of this is, as you're learning more about the project and you're changing the plan and the goal, as these models are just looping, it can update its understanding of the system. And then sometimes I have a state.md file or a work log just to track these longer-running tasks. If things are running for a day or two, I want to know what's going on. And I'm never going to read this four-gigabyte session JSON object. But I can look at the work log, have another model summarize it using a side chat, and go from there. This meant to say remote control. But again, if folks are just on their computers, I would also recommend trying that out. In the sidebar, there should be a button that says remote control, especially if you have the iOS app. This is the thing where we talked about being able to control your computer through your phone. So if you go on the iOS app, you enter codex, you can do this flow where you can scan a QR code, and all of a sudden your ChatGPT app can message and queue any thread in the application, including, I think, remote threads. This is super powerful because, again, oftentimes, every time I try to leave the house, someone's asking me for something, and now I can just ask codex. Really, I should just be having something that monitors Slack and just does it, but I'm not there yet. But again, this is one of those big, fieldy AGI moments. Right. Yeah. And again, like I said, the Chief of Staff thread effectively is the single source of truth for what's going on in my life. On my personal computer, I have a different one. Lots of good stuff. This is all I wrote for monitoring. So, I'm going to do a good job. I'm going to do a good job. I'm going to do a good job. And again, you'll start editing it over time. Maybe you don't like the formatting or you wish you included links. For a while, what I made it do was, if it found all the emails, not only to ask it to draft the responses, but I would make it open a Chrome tab for every email I need to reply to in Chrome. And so I'll open my computer, I'll take a meeting, I come back, and on my computer are just seven Chrome tabs. And I can just review the drafts and send each one. Small things like this just to prepare your computer while your meetings are happening. Really productive. So, now we talked a little bit more about doing the work itself, right? We haven't really gone into things like artifacts just yet. But I'm also curious if anyone has any questions on how I've been doing things so far. Over here. So, I noticed, so I do a lot of meta-prompting. Yeah. With GPC, I noticed that you would have very small prompts. Yeah. And so, but when I meta-prompt, I get a lot of stuff from GPC. Yeah. So, what's your advice? I generally always prefer to have the model write the goal, or write the prompt itself. More and more of these models are just getting better at doing that. And it's more in distribution to what they want. In reality, I will send a ten-minute voice memo, right? I'm just, this is some issue that's happening. I think there's a project about some thread. I don't know if they got back to me. I think their name is Dylan. Please look in Gmail. It's really, really messy. And then it's okay, well, do I want to make it set a goal? Do I want it to create a new thread? Do I want to make a new skill? That's really where my taste lies now. It's okay, how do I want to organize what the work product is? Do I tell it to then do all this work and put it into an indexed HTML to share it? Do I want to make a Word doc? That's basically it. But I generally just send very long messages. Yeah. For example, here, this is not the problem. This is the prompt for the chief of staff thread. This is the prompt so the model can make the chief of staff thread. Right? And that model will be much, that output will be much better at determining how verbose the automation is or how often it should check or what connectors. And then because I have this monorepo, one of the things I do is every project file has a link to every Slack channel this project is relevant for. And the model will just see that, and it'll read those Slack channels. Right? Every person.md file has their email address, their Slack connector, multiple work addresses, and it'll read all those things. I would never prompt that myself. Yeah. Yeah. Yeah. So I think in that example, what I realized was when I mentioned the Slack channels that are relevant for a project, the results were better. So I just added that in the front matter of the markdown file. But these are the things that you grow over time. I have not tried the deepest. So the question is, what is the difference between this and an open claw and a Hermes agent? I think, I'm sure there are differences I'm not aware of right now, but I can imagine a world where it's going to get very, very similar very quickly. Right? Most of my work, and I'll talk about this later on, is my threads manage themselves. Right? It could be the same thing as a sub-agent. I don't know how many people use Hermes agents and open claw for very wide work. Right? I think I tried open claw to do some house automation stuff. Yeah. Yeah. So the comment is, yeah, with these models, there's only one thread. I think that's very reasonable. Just in reality, there are so many things that we work on that I need the organization, right? It's is there a world where my banker and my therapist and my personal trainer and my girlfriend is the same person? Maybe, but my tiny brain can't figure that out. And it's easy for me to understand what the work is by making folders. Right? It's the computer doesn't know the folders this, but the folders are for me in some ways, if that makes sense. Yeah. Cool. Great. Yeah, one question? And then that one. Yeah. So the other thing is, one of your slides has three connectors, so et cetera. And then you said after a while you can just ask a comment. Yeah. How long would it take to do that? Because I go ahead and use slash and connectors, et cetera, skills, but I didn't know that it just automatically would be the one to use. Yeah. I think it depends. I think it depends on how proactive you are in telling the AI to remember these things. And it's easy for me to understand what the work is by making folders. Right? It's, the computer doesn't know the folders, but the folders are for me in some ways, if that makes sense. Yeah. Cool. Great. Yeah, one question? And then that one. Yeah. So the other thing is, one of your slides has three connectors, et cetera. And then you said after a while you can just ask a comment. Yeah. How long would it take to do that? Because I go ahead and use slash and connectors, et cetera, skills, but I didn't know that it just automatically would be the one to use. Yeah. I think it depends. I think it depends on how proactive you are in telling the AI to remember these things. So, for example, once I realized that I should include Slack channel IDs in project documents, the model's like, oh, there's a Slack channel. I should read the Slack channel. And that became really obvious. Now I basically never tag things. I don't really know when that happened. But, yeah, again, a lot of it is just getting in the habit of, remember this for next time. Update the skill for next time. Update the agents on BDOT for next time. And that is the... Sometimes. Or is it memory? I think it would be hard for you to prove it. There's also a memory system that's outside of the docs. Generally, I'm pretty happy with the memory. I don't know if I can give you a time estimate. But I think I would just try it out. Right? Use it for a couple weeks. Make sure your memory is turned on, by the way. It's also in your settings. I know our settings panel is pretty crazy right now. But, yeah, I think I would just turn memories on and just see whether or not these changes happen over time. Because I basically never... I don't remember a single time I've used at mentioned something. Yeah. And the last part is pretty fast, right? So, we talked about reading context. We talked about working with the context. Now, the last thing I have to do is just write context. Most of this has been pretty simple, right? You can draft emails. You can draft Slack updates. If you feel very brave, you can send them. But please be respectful. I'm sure Charlie has gotten hundreds of sent from chat to BT Slack messages from me over time. But I hope they sound like me more now. Building one-pagers is also very good. This slide deck was made with Codex. And soon we'll be able to do things like serve applications. And now, I think, at least internally, so much of the work that we do has just been sharing apps rather than full documents. I don't know how many people know about this, but we also have a really good artifacts ecosystem. More and more, Codex has become a tool for all of the work that I do. It can open and render Excel spreadsheets, Word documents, PDFs, slides. And with the annotation tool, editing things is pretty fun. So, even with these slides, this is actually served on the in-app browser of the Codex app. And what I'll do is I'll just give my talk. I'll press next and press next. And then when I don't like something, I just select it. It's like, hey, fix this. I don't like the white space. These two slides need to be broken into two more slides. I hit enter. As Codex is working to clean this up, I'm just going down to get down the slides. And it's generally been a pretty natural way of working. This deck really came from me reading my own blog post out loud and then generating the material as we go along. But it's still two or three skills to make the slides look this way. Yeah. Yeah, and then once you do that, you can do, again, again, it's the same concept over and over again, right? You can build out these loops that just touch other parts of the system. Most of our project tractors are just Google Sheets updated with loops. The C-slizer loops. And the annotations, yeah, I just said everything. My bad. And then one thing that's also very helpful is, as the company gets more context-dense, right? Maybe it's all A-A agents. Just the ability to summarize things over Slack has been incredibly useful. That was clearly a slop slide that ChatGPT added in. So once you can take actions on these DIN artifacts, I think the biggest thing, and the thing I really want people to try out, is just computer use. Again, it's the first time I had that feel the AGI moment, right? A cursor is trying to do some action. You can see it move across the screen. And when it does it really well, you really gain a lot of faith in the system. So we obviously have plugins, right? And this is for sending Slack messages. But then the in-app browser is going to get even more powerful soon, right? We're going to be able to basically treat this like the browser that you use. I now try to use the in-app browser as much as possible. And then for everything else, use computer use. How many people have used computer use, by the way? Very few. Oh, it's like 10% of you. What's a crazy thing that you've done? What's the craziest thing anyone's done? Any volunteers? Managing my home lab. Home lab? Yeah. What's a home lab? 30 agents. I have around, I have a . Oh. I have a bunch of boxes. Nice. And essentially 30 agents, and I can call it and just do stuff. I have a . Wow. Any other crazy computer use stories? You gotta get AGI pill and try out computer use. Yeah. One thing I'll say is, because computers are so powerful, you get, like you mentioned before, there is some safety component. I am now remembering this example where, again, because the Slack connector was not able to upload files, if this model is really determined, right, it could be the one wish willow. It just says, okay, great, well, if I can't add a Slack file, let me go on computer use and press file upload and do that. Right? There will be some times where the model, based on how you prompt it, becomes really determined. It's like, oh, it seems like I can't email someone using the Gmail connector. Let me open up Chrome and hit the send button. So those are the kind of things that you should be really wary about, right? These are real security issues. And, again, more than not, having things with the agent MD file has been really, really helpful, especially if you do things like guardian mode, or auto mode, excuse me. And it has also really changed the way I think about doing work. I feel like now when I'm doing things like building an application, if it's a native application, most of my testing is just done by Codex using computer use. If it's a website, again, it's just using the in-app browser. I think we talked a lot about this already, like, handle service work, like, it's kind of awesome just being a checkout page, app shots, you know, find me a coupon. It's made me more money than I would have expected earlier. Filling out forms, testing applications. I haven't really had a good use of this just yet, but it can also control the iPhone through screen mirroring. So do with that what you will. And this is sort of the escalation of the talk, right? Then the question is, what can the computer not do, and why can't it do it? And one of the things that you can't do is control Codex. But you don't really need to. Because these Codex threads can already talk to each other, it can already control itself in very powerful ways. Right? If you think of the example, this is not the slide I want. Damn, okay. I think I messed up some slides. Earlier in the talk, I talked about this idea that if there was some kind of support issue, I could take an app shot, and it would call this dx triage skill. And it will rename the skill, it will do the loop, and it communicates Slack and Twitter, and it's very nice. But I still need to be the person that triggers these things, right? Then the question is, what can the computer not do, and why can't it do it? And one of the things that you can't do is control codex. But you don't really need to. Because these codex threads can already talk to each other, it can already control itself in very powerful ways. Right? If you think of the example, this is not the slide I want. Damn, okay. I think I messed up some slides. Earlier in the talk, I talked about this idea that if there was some kind of support issue, I could take an app shot, and it would call this dx triage skill. And it will rename the skill, it will do the loop, and it communicates, Slack and Twitter, and it's very nice. But I still need to be the person that triggers these things, right? That's the same thing as me seeing an issue and making a polar cusp. In the future, the future for you, it's happening already now here, but now what I just have, I just have a single monitor thread. And any time it identifies any of these issues, it will go off and create a new thread. And the thread's job is to do all this triage. And what might happen is, maybe this triage is waiting on someone on Slack to acknowledge this issue. Maybe a pull request has been created, but it has not merged. If someone else complains again in the future, the monitor thread just goes, oh, yeah, I think this is the same issue. Not only is it the same issue, let me send a message to that downstream thread, so that it's aware this issue is recurring. Maybe that thread will send a Slack message, but in the main thread, I'll just see a message that says, hey, it's been the third day this issue has been live based on Twitter feedback. Should we do something? And those are the things. It's very hard for me to keep track of these things because I'm just on Twitter all the time. But by the agent being able to just manage and consume all this information, it makes these things much more tractable. I definitely messed up the slides over here, so I'm going to skip a couple slides. Yeah, this is the example. And I really want you to play around with this idea. I don't think it's been fully baked yet. Most of my work is about just having monitors create sub-threads. These threads are then managed and pinned onto the sidebar. This is also one of the best things, right? With a sub-agent, the thread just has these shapeless entities floating in the background, the JSON thread or the Galileo thread or whatever. The difference with sub-agents versus these threads is because they show up in the sidebar, you can just notice that something has changed. You know there's a new issue that's come up, right? And a lot of it, too, is just using the sidebar as effectively the hub of understanding what are the ongoing work streams. And this has been super powerful for me. But it all just starts from pinning a thread, taking actions, having this monorepo to manage all of your context, and then having different ways of waking systems up. And earlier, these systems wake up because you messaged it, or you set up a heartbeat. And now these things can be woken up by another thread running somewhere else. And so more often than not, most of my automations just happen on the monitor level of threads, and they trigger and create new threads and then manage themselves. Yeah. So I think we talked about a bunch of things, right? We talked about computer use, structured plug-ins. And, yeah. I think that's it. Try out some computer use stuff. I know I can't really, I'm really worried about the Wi-Fi here. I'm gonna try something too bad, but, yeah. Try out app shots, try out computer use, and really play around with what these models are capable of. There's a question over there. We'll do that one next. Yeah, I was wondering if you could show us parts of the 401 in your monitor level. I don't think they're gonna show you on this computer. They're gonna take me away. I mean, I can talk about it a little bit. So basically, my monorepo is set up so that it looks very much like the one over there. I have a projects directory. And in that projects directory is a named directory for every work stream I'm working on. So maybe it is the voice launch video. Or it is the agent's SDK. Right? Or it is the codex for open source grant program. So those are some of the projects. Then I have a people directory, which is just every single person that's ever DM'd me. Right? I know what they're working on. I know what kind of problems they're thinking about. I know what other side channels they're a part of. And then I have a bunch of different loose notes, agent summaries. I have daily summaries of what I've done. This is mostly just to test the limits of AI. I don't think it's very useful for that kind of stuff. Mostly it's just the projects and the people. Right? And then I have a to-do list that an agent just maintains. So I have a single thread that just checks the to-do list. Oh, this says it was undone. Let me have a sub-agent verify that no one has done this task. Those things are pretty token expensive. I don't really know if it would be worth people setting this up for themselves unless they just don't use that many credits. But I think the biggest ones is just the Sheva staff thread. 9 a.m., tell me what's happening today. Tell me what's happening this week. And I think if you just do that, you're gonna get a lot of juice out of that. Are you not getting crushed by context flows or whatever, the cloud that just comes to that? It's honestly, I've never experienced it. I think the compaction is just really, really good. I know that's not a very satisfactory answer, but if I could improve parts of the model, I would rather improve its writing tone rather than its ability to search the context. Generally, it's been pretty good. I think it might have to do with the way the codex memories have been set up, but I have not dug in too much into those details. I think you had a question over here? Yeah. was just how do you deal with memories bleeding across different projects? I don't think I have a good answer to that, primarily because it's unclear to me what the downstream side effect would be. Maybe one project uses npm, another project uses yarn or something, but generally I just clean those up in the agent.md files for those specific projects. Another thing to mention is every project directory had its own readme and its own agent.md files. And so, yeah, I think sometimes one project I was working on was npm and one was pnpm. I just cleaned that up in the agent's empty. I would be curious how much of that can be cleaned up just by using those simple tools. But talk to me afterwards. I'm really curious what's so bad about the bleeding. Yeah, maybe that would be a good example of having project-level scoping versus a single thread. But, yeah, I might have to look into more details about how the memory part is actually running. Any other questions? Over here. Did you change your codex on the file so that you can't do it? You don't want to read it. How did you set that up? No, so I think, honestly at this point it just kind of knows. It's such a bad answer, but it's almost like an AGI. To me it feels like a very AGI-filled thing, but I think generally in the beginning when I was working with it, I would have a skill called check notes. And it's like, hey, if you feel like you need to check your notes, check the notes. All right, and so if I had a hunch or some intuition that the model might not be able to do this, I would just mention, hey, check the notes. But even if you look at my OpenAI profile, I think my check note skill has been used like a hundred, fifty thousand times. But I've never mentioned it in the past two or three months. It's because, again, I think the memory system has been doing a lot of the heavy lifting. But again, it's like onboarding an employee. In the first two or three months of onboarding an employee, It's such a bad answer, but it's almost like an AGI. To me, it feels like a very AGI-filled thing, but I think generally, in the beginning, when I was working with it, I would have a skill called check notes. And it's, hey, if you feel like you need to check your notes, check the notes. All right, and so if I had a hunch or some intuition that the model might not be able to do this, I would just mention, hey, check the notes. But even if you look at my OpenAI profile, I think my check note skill has been used like a hundred fifty thousand times. But I've never mentioned it in the past two or three months. It's because, again, I think the memory system has been doing a lot of the heavy lifting. But again, it's like onboarding an employee. In the first two or three months of onboarding an employee, you have to give a lot of instruction. You have to give them a lot of context. And you have to let them fail. And when they fail, you have to give them an opportunity to write down what they've learned and clean up itself. I would be very surprised if you got good results in the first day of setting this up. But just, yeah, maybe not to your surprise, but now it's like, I don't, it just works. Right, and that's a great feeling to have, right, because it lets me go back in the flow of just doing my job. Are you a question? Yeah, thanks so much for the talk. I wanted to ask, so I'm a college student, so they can see it at the same group, and I'm going to be entering the workforce. I'm so old, it's so long ago. I have a line that's like, if you want to have good taste, you have to eat, right? I think your job is to consume a little bit more, try out different applications. Do you know what a good onboarding flow looks like? Do you know what that feels like? Have you built an app that makes you frustrated? And just your ability to consume more things and develop your vocabulary on how to complain about things that are bad, how do you complain about the slop? Those are the things, those are the skills I think will be very valuable. And specifically around vocabulary. Right, you can't really describe things that you don't really understand. And this happens in both things, like cooking, right? If you just don't know what the ingredients are, you can't describe that something is too salty, or it could be, it's like, it lacks some acid. You just need to consume a little bit more and then think critically about why you like something and why you don't like something. Because I think taste is a big issue, but I think a lot of it is because we are not consuming the right stuff. Right, try out all the different apps, if you're thinking about apps, for example. Yeah. Any other questions? Yep. You had your chiefs of staff, all these other things, and then you said you prompt the main thread, and then afterwards it goes down. Can you demonstrate how you do that? How do you configure that? Which part? So, talking to one thread, and then afterwards it's going to a different thread. So I understand you're in a thread and it bends up to sub-agents. I get that. But I think, at least I believe what you were talking about, is having a team of staff that's maybe going toward an IC. Yes, yes. And that's just like a thread. How did you set that up? Yeah, I mean, ironically, that slide just says threads can talk to each other, just ask. There is a list thread tool and a send message to thread tool that Codex has available. And so typing that is kind of awkward. But really, I just say, sometime last week, I was working on slides. Can you go find that thread, rename it, and pin it for next time? Right? I would never type that, but again, voice input makes it really easy. Sometimes I'll say, for example, this talk, I think I messed up some slides. I had a main slide writing system. And I said, okay, this slide could be split up in three acts. Make up a thread for each act, pin it, rename it act one, two, and three, and review each section. And then once it's done, review the whole slide. And yeah, I generally like to just ask it. I think I'm going to try to add more commands to be a little bit more explicit around thread control, but I think that will just come as the models get better, and as the model is more aware of what Codex can do. But yeah, it really is just like, I think one thing you should try to do is just have a thread and just say, read all the other threads that are pinned, and rename them. And use an emoji to color code its readiness. And just seeing it do that will be a first step to just understand how this thread control works. So this feature is very new, but it already has been pretty wild to just have a single thread, and you're just talking to it for the entire afternoon. Sweet. If there's no more questions, yeah, one more. With the walking thread, you said it's very new. Do you know if it's available across any other popular AIs? Because I don't think I've seen any before. Don't say that name to me. No, so the question was, do other coding systems have these kinds of tools? Not yet, I think. I'm sure they will soon. But as far as I know, this goes back to the eating thing, where I just have not tried the latest versions of other tools. I don't know if Conductor has these skills, but definitely, I'm sure, again, very competitive space, these features will probably propagate very quickly. Do you have old threads run on old models where you have the new threads go back and pull that in as context? It depends. Sometimes just say find the old thread and update the model. Just update the model. For example, when 5.4 came out and then we moved to 5.5, I realized all my old automations were just on 5.4. And it really bugged me. And I was like, hey guys, can we add a feature that has GPT latest so we can avoid this issue? And someone was just like, tell a thread to update the other threads. And I was like, yeah, so much of it is just ask. You just have to develop the language to figure out what you want to do. But I think that's, again, the most important thing. Yeah, so just ask it and then it'll automatically do it rather than you having to go through it. Yeah, I think I have an automation that just cleans up old threads and buy model IDs. But, yeah. Thanks, man. What's up? Does it live anywhere publicly? Not yet. I'm not happy with the slides enough that I'll publish them, but I have a blog post. Just search, if you just Google Codex Maxing, it's all there. Three X's, I think. But I'll share some more stuff on Twitter. Go on. Yeah, so you said that you work with unlimited tokens. Yeah. But since you work so closely with it, do you have any sort of advice on how we can, because a lot of time when you're doing loops, you get a lot of the same message from Codex and that can lead into the token limits pretty quickly. Do you have any advice? I know you don't have a problem, but since you have the expertise, do you have any advice you can share on how we can avoid hitting those limits, depending on, you know. Yeah. I think the biggest misconception here is that X high will give me the best results. And so there's a lot of X high maximalists here. I want everything on X high, right? When I was demoing getting a coupon thing and my friend does an app shot and hits enter, I'm like, why did you turn on X high? Right, it ran for two minutes, searching every website available for coupons. Obviously I can't give that much advice, because it's something like on my personal computer, I don't really run into limits, but really get comfortable with low and medium thinking. These models are still very, very smart, right? Low thinking on 5.5 is still so much better than prior models. And for a lot of this work that's not just like, make me a video game from scratch, you just don't need that kind of work. My team of staff I think is default medium. That's very helpful advice, and I agree, but just as a follow up, I want everything on X high, right? When I was demoing getting a coupon thing and my friend does an app shot, and hits enter, I'm like, why did you turn on X high? Right, I ran for two minutes, searching every website available for coupons. Obviously, I can't give that much advice, because it's something on my personal computer, I don't really run into limits, but really get comfortable with low and medium thinking. These models are still very, very smart, right? Low thinking on 5.5 is still so much better than prior models. And for a lot of this work that's not just, make me a video game from scratch, you just don't need that kind of work. My team of staff I think is default medium. That's very helpful advice, and I agree, but just as a follow up, I guess in a few words, I think what I was really asking, how do I tell the thread to shut up, saying the same thing over and over again? Oh, the answer's really bad, but I just say, if there's no updates, just reply with one word, no updates. That's one thing, and the second thing too is, I would also play around with how often these heartbeats are happening. Are you running this thing every 30 minutes? Are you running this every 9 a.m.? That's one thing. And the second thing too is, do you have some kind of stopping criteria? Right, so for example, I had some argument with Amazon, and I'm just like great, they put me in a 75 minute wait list, check every five minutes if the queue is better, and once you get the five minute wait time, check every one minute and keep replying until you get my money back. And I took a shower, and when I came back, I had $400 in my credit card. I think it's possible to play more dynamic heartbeats to be able to play that for a while. I've not tried too much there, but it's all possible, right? Because to create a heartbeat, it's just edit a text file. And so you should definitely be able to create your own heartbeat, and then change how frequent or infrequent. I might try to do that actually. I might just say, hey, during the weekdays, change your heartbeat to be more active, and during weekends or in the afternoons, change them to be less active. Yeah. I've tried versions of the Chief of Staff thread where I just set a goal that says never stop, and you're only allowed to set sleep. And that has also worked pretty well. It'll be like, hey, it's 9 p.m. Jason, I haven't seen Jason post a Slack message in two hours. I'm gonna sleep for five hours. And that also works. Yeah. But I think the biggest one, honestly, is people should not be afraid of load reasoning. X high is not X high results. It's just think more. and possibly can possibly possibly possibly possibly possibly and possibly possibly possibly possibly possibly possibly I would also play around with like how often these heartbeats are happening. You know, are you running this thing every 30 minutes? Are you running this every 9 a.m.? That's one thing. And the second thing too is like, do you have some kind of stopping criteria? Right, so for example, like I had some argument with like Amazon, and I'm just like great, like they put me in a 75 minute wait list, like check every five minutes if the queue is like better, and once you get the five minute wait time, check every one minute and keep replying until you get my money back. And I took a shower, and when I came back, I had like $400 in my credit card. I think it's possible to play more dynamic heartbeats to be able to play that for a while. I've not tried too much there, but it's all possible, right? Because to create a heartbeat, it's just edit a text file. And so you should definitely be able to create your own heartbeat, and then change how frequent or infrequent. I might try to do that actually. I might just say like, hey, like during the weekdays, like change your heartbeat to be more active, and during weekends or in the afternoons, change them to be less active. Yeah. Like I've tried versions of the Chief of Staff thread where I just set a goal that says like never stop, and you're only allowed to like set sleep. And that has also worked pretty well. It'll be like, hey, it's like 9 p.m. Jason, like I haven't seen Jason post a Slack message in like two hours. I'm gonna sleep for like five hours. And that also kind of works. Yeah. But I think the biggest one, honestly, is people should not be afraid of load reasoning. Like X high is not like X high results. It's just like think more. and possibly can possibly possibly possibly possibly possibly and possibly possibly possibly possibly possibly possibly