Every

OpenAI’s Codex Workflows for Knowledge Work

2182 summary words 10 min summary Watch video

Start with the signal

10 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: OpenAI positions the combined ChatGPT/Codex environment as an operating system for knowledge work: agents can progressively absorb a worker's recurring processes, operate across tools and files, and convert manual workflows into governed, reusable skills and internal applications.
  • Why it matters: The most useful material is a concrete operating pattern for agentic knowledge work: start with a real recurring workflow, capture context and judgment while completing it normally, codify discrete skills, and retain human approval until reliability compounds.
  • Best use: Watch for the implementation and governance lessons from the strategic-finance case study, rather than for product-launch claims; translate its project-scoped instructions, skill library, audit loop, and approval gates into Ken's agent/control-plane design.

Executive Summary

The video argues that increasingly autonomous models shift knowledge work from asking a chatbot for isolated outputs to supervising several delegated workstreams. OpenAI speakers describe Codex and ChatGPT as complementary surfaces in one workspace: users can research or brainstorm in ChatGPT, hand the resulting thread into Codex to build or execute, and let the agent use connected tools, browser access, local files, and organizational systems.

Its strongest example comes from OpenAI strategic finance. A non-engineer describes building, with Codex over roughly a month, a custom compute-cost and monthly-close application spanning a data lake, spreadsheets, Google Sheets, slide outputs, usage data, and qualitative commentary. The system reportedly compressed a process that formerly took nearly five days into about five hours, while leaving people responsible for final commentary, auditing, and release.

The practical message is explicitly not to one-shot an entire business process. Run the process normally while having the agent collect context from files, Slack, and source systems; identify recurring chunks; encode them as skills; and maintain project-specific instruction files. Over recurring cycles, errors and preferences are converted into better instructions, plugins, and workflow behavior.

The speakers do not present autonomy as inherently trustworthy. They compare onboarding an agent to onboarding an analyst: initial work may be only 70% correct, then improve through review and correction toward 90-95%. Their governance pattern is to permit first drafts and background preparation, use cross-document auditing skills to find inconsistencies, and impose explicit policies such as never sending to Slack or an executive channel without human review.

Key Takeaways

  • Claim: The intended workflow is a unified, thread-connected workspace in which research, planning, execution, and coding can pass between ChatGPT and Codex without manual copy-paste. | Evidence: Roman describes starting feature ideation or deep research in ChatGPT, then @-mentioning that conversation when ready to build in Codex; Codex is also presented as a dedicated space inside the ChatGPT app. | Implication: For agent systems, preserve task lineage and shared context across planning and execution surfaces rather than treating research chats, coding agents, and operational automations as disconnected tools. | Caveat: The product claims are made by OpenAI employees in a launch-oriented discussion and are not independently validated in the transcript.
  • Claim: The most valuable autonomous-work pattern is delegated execution with verification, not merely faster drafting. | Evidence: Dom says the model can use computer access and a Chrome extension to complete tasks end to end and cross-check its own work; he reports personally running five or six tasks in parallel. In one example, the model inspected a video frame by frame, inferred interactions, and rebuilt interactive visualizations from it. | Implication: Ken should distinguish agents that can plan, act, and test against an artifact or source system from assistants that only generate a proposed answer; the former require orchestration, observability, and approval design. | Caveat: Self-verification reduces supervision burden but does not eliminate the need for external validation on consequential workflows.
  • Claim: Recurring, high-friction business operations are better targets than generic one-off tasks because the upfront training effort compounds over each cycle. | Evidence: Kyle's strategic-finance team used Codex to automate a complex monthly compute allocation and close process across disparate systems. He says the workflow evolved from roughly two days of accounting allocation plus one to two days of analysis and artifact production into a custom application that compressed nearly five days of work into about five hours. | Implication: Prioritize agentization candidates by recurrence, process pain, availability of source artifacts, and the value of retaining institutional judgment—not simply by whether the task appears easy to automate. | Caveat: The reported time compression followed a month of build work and continued iteration; it was not achieved through a single prompt.
  • Claim: The recommended build method is progressive process capture: execute the workflow conventionally once, let the agent gather context, then codify repeatable substeps as skills. | Evidence: Kyle advises against one-shotting a giant system. Instead, he recommends feeding the agent the files, process, and Slack context while doing the work “old-school,” then turning recurring chunks into skills, larger skills, plugins, and improved agent instruction files. | Implication: Ken should architect agent rollout as a learning loop with versioned procedures and post-run refinements, not as a one-time prompt-engineering exercise. | Caveat: This requires disciplined documentation and repeated use; it is initially additional work rather than immediate labor elimination.
  • Claim: Trust should be earned through staged delegation, explicit review boundaries, and reusable audit procedures rather than assumed from stronger models. | Evidence: Kyle acknowledges both hallucinations and missed source material. His response is to correct errors so they become future instructions, initially use the agent for first passes, and employ a Google Drive audit skill that traces a metric across a memo, slides, Excel, and Google Sheets to identify inconsistencies. | Implication: For sensitive workflows, build source-tracing and reconciliation into the agent's workflow, measure failures by type, and gate external actions behind human approval until proven reliability is task-specific. | Caveat: The speaker frames reliability as improving from roughly 70% initially to 85%, then 90-95% over one or two months; even at higher reliability, final review remains part of the described finance workflow.
  • Claim: Project-level memory and policy are the practical control plane for multi-workstream agents. | Evidence: Kyle describes a general agent instruction file plus separate project-folder instruction files and relevant skills. An inbound Slack item can be routed to the appropriate project thread, where the agent uses that project's context; he sets an explicit policy that the agent must never send a Slack response before he reviews it. | Implication: Ken should model each agentic project as a bounded workspace with scoped context, role-specific skills, routing rules, and action permissions—not as one undifferentiated long-running chat. | Caveat: The transcript does not specify enforcement mechanisms, access controls, audit logs, or how conflicting instructions between global and project scopes are resolved.
  • Claim: Agent-built internal applications can replace static reporting artifacts when the workflow requires exploration, explanation, and team collaboration. | Evidence: The finance example moved from data-lake/Excel/Google Sheet/slide handoffs to a custom hosted dashboard with hierarchy views, P&L and GPU-cost detail, generated narrative slides, collaborative commentary, and a conversational mascot that guides users through the dashboard. Google Slides remained an export format for board sharing. | Implication: Where reports are repeatedly rebuilt and interrogated, the target product may be an agent-maintained operational application with export capability, rather than a better spreadsheet or slide-generation workflow. | Caveat: Static artifacts were not eliminated entirely because external sharing constraints still required export to Google Slides.

Detailed Brief

Knowledge-work capabilities and interface claims

  • Claims: OpenAI says every internal team—from finance and recruiting to sales—has been using Codex daily, reflecting a shift beyond conventional software development.; The speakers argue that nontechnical users should not need to interact with visible code even when code is used behind the scenes.; They present computer use as capable of handling vaguely specified intent, navigating apps or Chrome tabs, and allowing the user to continue other work while the task runs.
  • Evidence: A documentation-refresh example used a screenshot as input; the model reportedly cross-referenced the codebase and rebuilt the component as actual HTML rather than retaining a screenshot.; A Chrome “site chat” feature is described as connecting browser context with the ChatGPT application and local filesystem access, enabling work across a webpage or Google Doc and local files.; For developers, the launch mentions inline code editing, PR review in Codex, and a higher reasoning-budget mode called “Sol Ultra.”
  • Caveats: These are vendor statements about newly launched features; the transcript offers demos and anecdotes but no benchmark methodology, reliability statistics, pricing, security model, or enterprise deployment details.; “Computer use is basically solved” is a speaker assertion and should not be treated as evidence that browser-based agents are safe for unrestricted production actions.
  • Implications: The meaningful product transition is not simply a better coding model; it is agent access to work context and execution surfaces that traditionally sit outside the IDE.; A usable agent system must make its work legible—through explanations, visualizations, and inspectable artifacts—because user understanding becomes the bottleneck as autonomy rises.

Finance workflow design details

  • Claims: The finance application was designed around a hierarchy of compute usage and costs that existing generic visualization tools reportedly did not represent adequately.; The team uses generated reporting artifacts as checkpoints or extracts rather than as the primary place where analysis occurs.; Skills can be shared across projects while remaining specialized at the project layer, enabling a portfolio of reusable operational capabilities.
  • Evidence: The cited compute workflow incorporates external consumer and enterprise usage, internal research usage, product-level behavior, P&L effects, GPU drivers, and qualitative close commentary.; The dashboard's conversational mascot is described as a skill-powered Q&A layer that can answer questions and direct less familiar teammates through the application.; The workflow reportedly starts with incoming Slack context, routes the work into the relevant project, opens a project thread, and begins a governed first pass.
  • Caveats: The finance data model, allocation rules, and underlying controls are not shown, so the case study demonstrates a workflow pattern rather than a reproducible accounting-control design.; The system depends on access to multiple sensitive sources; the transcript does not address least-privilege design, data retention, auditability, or segregation of duties.
  • Implications: The highest-value automation may be a living decision interface that combines data transformation, narrative generation, review workflows, and user guidance.; For internal agent products, adoption can improve when the interface helps users navigate and understand an unfamiliar application rather than merely exposing an agent chat box.

Notable Concepts & Terms

  • Codex as an operating system for knowledge work: The video's framing for a persistent agent environment that coordinates research, files, tools, browser activity, project context, and execution—not just code generation.
  • Skills: Reusable encoded procedures for recurring chunks of work; the speakers treat skills as the mechanism by which a user's process and corrections compound into future automation.
  • Agent MD file: A described instruction/configuration file for an agent, used both globally and per project to supply persistent context, behavioral constraints, and orchestration guidance.
  • Project folder: A scoped workspace containing thematic work, project-specific instructions, relevant skills, and threads; it is presented as the unit through which the agent routes and executes work safely.
  • Progressive process capture: The recommended implementation method: complete a difficult recurring workflow normally while capturing inputs and decisions, then codify pieces over time instead of attempting a one-shot autonomous system.
  • Google Drive audit skill: A cross-artifact verification procedure that traces a metric through linked memos, slides, spreadsheets, and Sheets to flag inconsistencies before distribution.
  • Watch-and-replay / replay-and-record: A described training feature through which the agent observes how a user performs a task or corrects an error, then updates the applicable skill with that context.
  • Staged delegation: A trust model in which agents begin with research and first drafts, humans review outputs and approve consequential actions, and autonomy expands only after repeated demonstrated performance.

Operator Notes / Why Ken Should Care

  • Select one recurring, cross-system workflow with measurable cycle time and an existing human owner; run it in parallel for at least two cycles before replacing the incumbent process.
  • Create a project-scoped agent contract containing allowed sources, authoritative source hierarchy, prohibited actions, required validation checks, approval gates, and escalation conditions.
  • Build a reconciliation skill for any financial, operational, or investment artifact that traces key metrics back to authoritative systems and flags conflicting values across documents.
  • Keep outbound communications and high-impact system changes in draft-only mode by default; promote permissions only after task-level reliability evidence is collected.
  • Instrument the learning loop: log misses, hallucinations, source omissions, reviewer corrections, and time saved, then convert recurring corrections into versioned skills or project instructions.
  • Evaluate whether recurring reports should become an internal application with queryable evidence and exports, rather than continuing to automate spreadsheet-to-slide production alone.
  • Validate product/security claims independently before granting browser, local-file, Slack, email, or finance-system access; the discussion provides no access-control or compliance specifics.

Source/Metadata

  • Title: OpenAI’s Codex Workflows for Knowledge Work
  • Transcript words: 8388
  • Duration seconds: 1769
  • Timestamp note: No usable timestamps or chapter markers were present in the supplied transcript. The transcript also contains substantial duplicated passages near the latter half.
Full transcript 5849 words · 37 min read
0:00

I think this combination, ChatGPT, Codex, and GPT 5.6 is the gold standard to me for knowledge work. It is the place where I spend all my time, except for the times when I use Fable, which we'll get to. I use it for coding, I use it for writing, I use it for a lot of different things. And I think it has this combination of power, performance, usability, and speed that just makes it great for collaborative work. And we've got two very special people joining us from OpenAI, Dom and Roman. Welcome. Please introduce yourself and tell us what you do at OpenAI. Yeah. Hi, I'm Dom. I work on developer experience for Codex, and now, I guess, ChatGPT. Sweet.

0:21

Hey, everyone. I'm Roman. I lead developer experience working with Dom. But yeah, exciting day. Tell us how things are going. So this has been a bit of a different model launch because it's the first model that I remember where we knew it was coming before it came out.

0:38

Yeah, no, it's a super exciting day. So now GPT 5.6 Sol, which is our new frontier model, is available to everyone. We worked hard to make it available to everyone. So it's an exciting day for people to now use it. And yes, we also noticed with the rise of Codex over the past few weeks how many people were using Codex as an app, not just for software development, but also for anything around the code, right? And that has been true at OpenAI for many weeks now. We saw every single team at the company actually leveraging Codex every single day, from finance to recruiting to sales. Everyone was actually trying to use this product to make amazing things with it. And so that's why, on one hand, we have this model. It's really frontier intelligence for coding. We can talk more about this, but also for any kind of knowledge work. And at the same time, we wanted to make sure the surface that people use to access these capabilities was simpler, right? If you're working in sales or in finance, you don't really want to see code, even if there is code behind the scenes. And so that's why we now have ChatGPT work with this new work agent that gives so much more power. You can connect to all of your tools, the services that you use every day, and now you have, as you saw in the livestream, the ability to do any kind of work on your computer from ChatGP.

0:43

Amazing. And if I'm a Codex user, and I'm trying to think about, okay, how do I now use this one app? How do I think about it, and what are the differences? When do I get into the work tab, and when do I go into the Codex tab? Yeah.

0:52

Yeah, it's a great question. Honestly, for Codex users and developers, it's the same Codex that you know and love, right? We're just making it better. It has its own dedicated space within the ChatGPT app now. But if Codex works for you, you can just stay there for all of your work. And in fact, we as engineers have many tasks to do every single day that are not coding tasks, right? Sometimes we write documents, we write documentation, or sometimes we have to explain our work, we have to keep up with projects from Slack and different sources. So all of that already works well in Codex. I think what's quite powerful now too, with ChatGPT and Codex coming together, is that we had feature requests from builders telling us, for instance, hey, I'm ideating with ChatGPT. I'm running a deep research, for instance, about some ideas for features I should build. But then I have to copy and paste that over to Codex, which is strange, right? And now what's nice is with this one app, one surface, I can start brainstorming with ChatGPT, even with a deep research. And when I'm ready to build, I can @ mention that thread and get that conversation going right away. So this interplay is very interesting because now it's one surface, and all of these tasks and threads connect to one another.

0:57

I think also don't worry too much about that quote-unquote toggle in the top corner. You can literally just stay in Codex. Or if you're curious, while a task is running, just quickly switch between the two. And you can see all the tiny changes that are happening in the UI if you're switching between the two. But you're not losing out if you're doing knowledge work in Codex, for example. Quite the opposite. I've been doing plenty of knowledge work this week on Codex.

1:03

And by the way, we've been shipping every week now on Thursday for the past 10 weeks or so. So today is no exception, even for developers using Codex. We have a ton of new features today, of course, with Sol Ultra to get max reasoning budget for your most ambitious tasks. But we also have new things like inline code editing, reviewing PRs directly within the Codex space in the ChatGPT app, and many more things. So in case you were wondering, we're just getting started for developers as well.

1:09

I love that. One of the things that we've been talking about is there's a very hot thing right now about loops, where if you're a developer, instead of doing the work yourself, you're building the system that does the work. And that's the loop the agent is running in, more or less. And one of the things I feel in using this model over the last month or so is that it's the first time where a loop workflow is available for non-developers, where you can actually go and delegate tasks and actually have it do a lot of your work for you while you're tending above it. It sounds like there's no work to do. Actually, there's a lot of work to do. Making the loop work is a lot of work, but you have a lot more leverage. So for example, I have not, and I think Ramon, as we've been emailing, you may have seen some of this, I have not done any of my email in the last two months because it's just basically 5.6 in Codex that knows all my preferences and knows how I would respond to stuff. And I'm giving it little tweaks here and there, but mostly it's doing my email for me.

1:15

But I'm curious, is that something that you guys have been seeing internally? Do you have examples of how it has changed how you've worked? Because I really do think that this is one of the things we've been saying. And I'm sorry to say this, but I think Fable is a better overall programmer, but it's less usable. It's a less usable model. And I think 5.6 is still very powerful, but it's way more usable so that someone who's non-technical can use it for delegated work in a way that is very new. So I'm curious what your experience is.

1:22

Yeah, I think the interesting thing with 5.6 is it naturally is more, I would almost call, independent, especially if you have computer use and the Chrome extension. It just gets the job done, and it will do what has to be done, right? So if you're using computer use, Chrome extension, all of those aspects, even for knowledge work, there are a lot of things where it can just get it done end to end and verify its own work. And you might feel like it might take a bit longer on a task, but that's because by the time that you come back, it's actually going to be done. I've noticed this myself where yesterday I was running five or six tasks in parallel because it was just getting the job done at a level where I'm like, I don't have to tend this at all. And you don't even necessarily have to think about it, especially as someone who hasn't been AI-pilled like we all have. You don't have to think about it as a loop, right? It's just that the model naturally cross-checks its work and makes sure that things are done. And it does it in incredible ways. I'm going to show you a demo.

1:28

Please. So as part of this whole merge situation, we also revamped the docs. We moved the Codex docs out of developers.openai.com into its new learn.chatgpc.com. And we wanted to have a lot of delightful moments. And so one of the things that we ended up doing here is, for example, a lot of these components are no longer screenshots. This is actually, we gave Codex a screenshot, and 5.6 Sol actually just went ahead, someone who hasn't been AI pill, we all have, you don't have to think about it as a loop, right? It's just the model naturally cross-checks its work and makes sure that things are done. And it does it in incredible ways. I'm going to show you

1:55

a demo. Please. So as part of this whole merge situation, we also revamped the docs. We moved the Codex docs out of developers.openai.com into its new learn.chatgpc.com. And we wanted to have a lot of delightful moments. And so one of the things that we ended up doing here is, for example, a lot of these components are no longer screenshots. This is actually, we gave Codex a screenshot, and it, like five, six, Sol actually just went ahead, cross-referenced it with the code base, and rebuilt the whole thing in actual HTML. Wow. And where this became absolutely wild for me was last night. I added, we added a new feature

2:27

called visualizations in the app, which works regardless of whether in chatgpc work or Codex, or even on the web. And you can ask it to visualize concepts for you. And we had, let me show it in the blog post, actually. We have in the blog post here a couple of examples of what this looks like. It's really cool. It gives you these interactive demos, where you can play around with things to visualize ideas. I gave Sol this blog post draft. And I'll show you the task in a second. But it actually was able to build me the full interactive things from the video. So it actually inspected the video using the Chrome extension. So cool. All of the necessary parts took

3:09

screen caps of different frames and then figured out how do the interactions actually work? What are the things we're trying to show here? And so I was able to completely rebuild this by inspecting the frames of the video. It absolutely blew my mind. And obviously, I had not even seen this. This is awesome. But we would not take the time to build this if we did not have a model like Sol at our disposal, right? Because this would be too much effort or too many turns. But the model being so autonomous is so cool. If you're looking at this, you're like, I don't know when I would use a visualization. I think one

3:49

thing to be really clear on is, as the models get more powerful, and they're able to do more autonomously, the bottleneck becomes, can I understand what it just did? Do I know, really, do I have a real mental model of what's going on? And so its ability to tell stories like that in a visual way that clicks, I think is actually, if you're thinking about what is the next frontier, I think that kind of thing is a huge, huge unlock. Totally. And that's why I think ultimately, it comes back to the DNA of OpenAI, right? We have, we are a research and deployment company. And the two sides are very critical. You want to make the very

4:23

best frontier intelligence with research, but you also want to make it usable for people. That's why we spend so much time on the harness, which is open source, the Codex harness, leveraging all of this, and able to work for a long time, very reliably check its own work, but also really be delightful in the way it explains what it did. So I think that's all important. And by the way, one thing that's magical in GPT 5.6, if you've not tried it yet, is really computer use. Computer use is basically solved at this point. It's so much faster. And the fact that you can spin up agents to delegate tasks that can be quite nebulous, not even quite

4:58

precise, but the model understands the intent, starts navigating your own apps or your Chrome tabs with its own cursors. You can continue doing the work and check back when it's done. It's really delightful to see. I love it. I use it all the time. I am often in a meeting and my computer does something on its own, and I'm like, oh shoot, Codex is doing work. So guys, I know that we're out of time and that you have a very busy schedule. Thank you so much for coming on. Any final things you want to leave us with before you head off? I think there's one small feature that I would highly recommend checking out, which is if you have

5:38

the Chrome extension installed in the latest version, you actually have a site chat now. So you can actually open the chat inside Chrome directly. And this is incredible because it actually connects to your ChatGPT app, which means that you have local file system access, and it allows you to do things like looking at a website or Google Doc, and it can directly reference files that are on your file system and interact with it too. And that's just incredible for things like knowledge work. Totally. And my parting word would be, these models have now reached such a level of capability that no idea is impossible. Just give it your most ambitious projects, all of the

6:14

ideas you've been putting off, you've not tried before, the hardest bugs. And we can't wait to see what happens because it's a really great model. I love it. Thanks for joining, guys. Thanks for the work you're doing. We actually have a special guest, Kyle Kober from OpenAI. Kyle, welcome. How you doing? Good. Nice to meet you guys. Tell us about what you do on OpenAI and your thoughts on the model. Yeah. So I would say very well aware of your thesis where Codex is becoming an operating system for knowledge work. And that's 100% true. I am not a software engineer, yet now, since I'd say February, I've been doing software engineer-type work, just

6:50

integrating this into my natural workflows. And it's a very exciting day with Sol and the merge of Codex and Chaxx. I feel like now everyone is going to get a taste of that Codex-style workflow, where if you've just been using chat and not Codex as the knowledge worker, and I jokingly will say don't box the knowledge worker in, it only takes a few months of doing it to where you're able to do almost anything. But it's really exciting that everyone's going to get a taste of this. And it's something that we've been doing internally now for months. Codex is 100% my operating system, runs basically everything for me today. Can you tell us more about some of those

7:29

things that it's doing for you that people may not think immediately, oh my god, yeah, you can use it for that. Yeah, it's just way more proactive, where it's integrated into all your systems, where if it's connected to Slack, Gmail, Outlook, if you're on a Windows computer, or basically any of the Microsoft suites for Excel, PowerPoint, Google Slides, it's just unbelievable where it can gather context. So every morning for me, my chief of staff is basically checking Slack and core channels that I'm very involved in, seeing where I'm potentially behind, giving me context links from Codex directly to where I need to go, pre-drafting

8:19

stuff, which is extremely powerful. It just allows you to be always connected to what's going on. And for me too, I would say there's so much stuff going on in forecast files, closed files, and these G Suite artifacts that we use, and being able to have it run through comments, draft stuff, get my take, and be truly the operating system where you start with Codex and it enables you to get to that 70, 80%. And then after a few weeks or months of using it, it gets to 90% because it learns your style. It just is a huge time saver and productivity boost. Take me through that. Give me a specific, you're talking about closed files and forecasting,

9:16

I want the crazy finance stuff that I might not know about that you use it for. Can't share anything live. So this is a Google slide that has a GIF, which was neat where Codex made this slide for me from looking at my hosted app or site that I created. So basically everything Codex is taking a first pass at, even creating this live demo GIF thing, right? But the real world messiness, I would say, is each month compute is extremely complex. We use it externally for folks like yourself where people are paying users, both for consumer enterprise. We also use it internally, obviously, for research.

9:53

So there's a very hierarchy view to it, which didn't really exist in any of the existing software applications like Tableau or other visual things. So we basically went ahead and built our own. I want the crazy finance stuff that I might not know about that you use it for. Can't share anything live, so this is a Google Slide that has a GIF, which was neat, where Codex made this slide for me from looking at my hosted app or site that I created. So everything Codex is taking a first pass at, even creating this live demo GIF thing, right? But the real-world messiness, I would say, is each month

10:30

compute is extremely complex. We use it externally for folks like yourself, where people are paying users, both for consumer and enterprise. We also use it internally, obviously, for research. So there's a hierarchy view to it, which didn't really exist in any of the existing software applications, like Tableau or other visual things. So we went ahead and built our own. And what's really neat is this is something now that is a part of our workflow every month, where we close the books with our computer accounting team. And also what's really neat is this actually does a whole allocation on the backend where Codex can meet your team where you're at. So

11:02

let's say you have a very strong data science or data engineering team, where everything is in the data lake. Amazing, because Codex can interact with that fast. If not, everything's in there and some of it's in a G sheet or an Excel. Great. Doesn't matter. Codex can do the work of pairing data across these different systems for you. So in this case, this months ago basically started out where it was automating an allocation that maybe took two days of accounting time. And then my team, we spend a few days auditing all the detail, going way deeper into the product usage, since we care about what a pro user is doing with Codex or image gen for

11:40

the month to understand how the business is behaving from the strategic finance lens. Right. So we would spend a day or two there and then do the data lake, our Excel to a G sheet, outputs in the G sheet, link those to slides. And that was the old way of automating it. And we can still use G sheet and slides to gut check or check artifacts. They almost become extracts now for us, but right now we go from data lake straight to a custom hosted app, which is built from the Sites product that we have. So it's a way that we compress almost five days of stuff into five hours, and it's not perfect. There's still a lot of work that went into

12:20

building this. You can't just one-shot doing an entire workflow for you, right? But you can teach it over the course of a week or two weeks. And then as it runs this monthly process for you, it just gets better and better. You codify the learnings with it. It learns the skills, you improve that, and then now where we're at is we're taking what we did for compute, doing it for other areas of the business where people are now able to, in strat fin, use all the skills that went into building this. We're creating our own internal plugins. And that's something that, too, will make its way to the enterprise. So it's

13:03

exciting as a finance person, getting to do software engineering work and getting to actually improve the product and improve the experience, not only for our strat fin team, but eventually, too, for the people like yourself using it. So yeah, don't box the knowledge worker in, is my joke, but yeah, here's a perfect example of what you're talking about. That's really interesting. So help me understand who built this. Did you build it? Did someone in your team build it? Are you pairing it? Yeah. I built this with Codex, basically just working outside of my normal work hours, probably into the evening, all of March, for a month.

13:39

And so the first opportunity we had for this to dual-track this with our normal work was in March. And now we've had a few months of it, and it now is humming, where out of the box, we're 95 to 98% of the way there. And it's really coming in and just finalizing some of the qualitative side of your monthly close. And I would say also what's neat is whatever you can imagine is what you can build. So in this demo, you can see it going through, you have an overview tab. Again, we have the hierarchy view of how we think of compute. You can then ask it or see, hey, how does this hit the P and L?

14:20

And that goes to more of the compute GPU specifics, like what were the GPUs that drove this cost? Or what does it take for image gen to run on a plus account, for example? And then because it has all this information and built the backend of this data set, it can actually take a great first pass at the qualitative side of your close slides or whatever slides are doing. These are virtual slides built into this dashboard. We're not using the traditional slides anymore, right? And then what's really neat, you'll see this little mascot, bottom right. This is an idea I had, inspired by Codex pets. If you guys like your Codex pets,

14:54

for people that are maybe on your team that are less familiar with this or how to navigate a new piece of software you built, there's a mascot you can just talk to, and it will take you throughout the dashboard and answer questions. And this is all based on a skill where it runs a Q&A factory and gives this mascot a way to answer and direct you across the dashboard, which is really neat, which is something that, I mean, you couldn't even think of building any of this. And now you're just having fun building software that helps your team understand what's going on with your business. I love it. I'm going to give people on the team a chance to ask questions. But before I

15:38

flip it over, one final question I have for you is, this is a very complicated piece of software. Speaker 2 And it sounds like you built this over the course of a month or two. What did you learn about how to make something like this that, if you were going to start over, and you could tell yourself a couple things from a couple months ago, would have helped you avoid some of the pitfalls? Because I know you can make this without coding, but there's a lot to do to actually make it good. Yeah, that is a very good question. And I would say it's perpetual for us, right? This is not done. Every month, we're trying to make this

16:12

better. I would say the big learning is, and what some people will try to do when they use Codex for the first time, or chat for work, they'll try to one-shot stuff, where it's like, I would say, take this as one of the hardest processes we have each month, which is why we thought to even try this. I would say take it to your hardest process, do it the old-school way, but have it gather all the context on all the files, the process, the Slack channels. And then as you go along the way, you say, okay, this is a recurring way that we're doing this. Let's codify it with a skill. So you're almost reverse-engineering the process over time. And then as it learns

16:56

that, you can either build a larger skill or create a plugin for your team or yourself, or enhance your agent MD files along the way. I think one learning, too, is for the knowledge worker to don't be scared to try to figure out how all this stuff works, where people will say, oh, my agents are doing this. The way we think about agency, or some people think about agency, or myself, is your Codex has an agent MD file. And then with each of your projects, which this compute metrics is one of my projects, it has its own agent MD file, so that my Codex is able to orchestrate across all these different projects. So stay organized on the

17:32

backend, try to recursively improve your skills. You can create skills that improve your skills specific to certain projects, or just overall reflect and refine, which is a way that we think about a skill most people have or build themselves here, which actually integrates really nicely with Chronicle, if you've tried or use that on the backend of Codex. So it becomes super powerful. But I would say, don't

17:51

Oh, my agents are doing this, the way we think about age, or some people think about agency, or myself. Your Codex has an agent MD file. And then with each of your projects, which this compute metrics is one of my projects, has its own agent MD file. So my Codex is able to orchestrate across all these different projects. So stay organized on the back end, try to recursively improve your skills. You can create skills that improve your skills specific to certain projects, or just overall reflect and refine, which is a way that we think about is a skill most people have or build themselves here, which actually integrates really nicely with Chronicle, if you've tried or used that on the back end of Codex. So it becomes super powerful.

17:58

But I would say, don't try to take the time. Treat it as almost an extension of yourself. You're training a version of yourself to take this away. And yeah, the investment is maybe it won't feel like a one-shot, and maybe it takes a week, or it could take days. It just depends on how complex the thing is. This is extremely complex data from disparate resources, and different teams are working on this. So if this takes you a month, the nice thing is I never have to make these slides again. Each month, all the data is magically put together. Every slide just updates, the drop-down for the new month comes online, and then we're basically at that last mile, just trying to go through and edit commentary.

18:02

And now it is collaborative. So once you build this and push it, our whole Strapin team is coming in and editing commentary. So it actually becomes a control point where we're releasing the final stuff before we then have Codex actually export all this to Google Slides, since there's some shareability, we can't have our board accessing internal sites. But anyway, that's some of the advice and learnings. A lot of people are intimidated, and I would say just ask it and work with it and treat it as an extension of you, but also something that can uplevel you. And don't be scared just because it has Codex in the name or chat for work. It is limitless in terms of what you can do. I love it.

18:07

So I would say maybe a summary is, don't just try to one-shot this gigantic system all in one. Do it just in time. So as you have something come up in your workday that is part of a big project, see if you can make a skill that helps you do a chunk of it and build that up over time. And it sounds like you parallel track things. So as you're going through the process this month, or a couple months ago, you were building a system that would help you for the next time, but not necessarily this time. So all those things are awesome. I know we only have you for one minute left. Does anyone else on the team have a question?

18:18

Yeah, Kyle, I wanted to ask a question. I work a lot with finance teams in hedge funds and other financial firms. And the big friction point they always have is, how do I trust it? Because I have all this data, and they notice two issues. And this is with all AI models, but either one, it hallucinates sometimes, and that problem has gone way down, but it still occurs sometimes. So it makes up facts that weren't in the actual data. And then the other is just, this is way more common, but it misses things. So when there's a lot of data, it might not have looked at that file, or it didn't catch that email. So can you walk me through how you build trust in this piece of software as you're building it? What are some steps you take to evaluate?

18:23

Yeah, that's a great question. And it's actually probably a recurring theme we see when I get to join some deployment calls with some of our enterprise customers. But I would say it's a process. It's the same way if you hire an analyst, and they're coming up to speed. They're going to get stuff wrong. It won't be perfect.

18:28

I will say five, six soul with ultra turned on is a big step up. And the sub-agent use is built in, which is just incredible, where it can multitask and do different pieces of the process at the same time. But it all comes down to using it periodically. Let's say, Mike, it did get something wrong for you. You have to teach it not to do that ever again. And that's how that process goes, where it's like, hey, I didn't like the way you output this or designed this, maybe if you're doing more of building a build, or if you're auditing materials.

18:34

For example, I have a Google Drive audit skill, which is pretty neat, where before we ship anything, we're able to audit a document and look for any inconsistencies between a metric that's maybe mentioned on a memo that's linked to a Google Drive, to a slide that's linked to an Excel or linked to a G sheet. You basically trace the metric, and then it tells you what is wrong. The only way it's able to do that, though, is because the first few times, I was like, hey, this is how I think about tracing things. This is how I think about going through these files. Watch me.

18:39

And now there's a watch-and-replay or replay-and-record feature where you use that. It's watching how you think about doing stuff if it got something wrong, so that it gains the context, updates your skill, and then that's how you gain the confidence that that won't happen again. So again, it's this one-shot mindset where you shouldn't expect it necessarily to be perfect the first time, but you should know as you're bringing it along, it's a compounding curve where that first time maybe it's 70% there, where you're catching multiple mistakes. The next time it's 85, and then basically a month or two from now, it's at 90, then at 95. The payoff is eventually there. It's not less work initially because you need the time to build it and teach it, but the payoff does come, and that's how you get more comfortable.

18:44

Now, for me, I allow it to take first stabs at work where it is basically trying to incubate work from reading Slack. And I'm not having it just fire off the end results, but it's getting there. I get to review, figure it out, and it's already working on stuff before I even get to read it, which is kind of neat. But to get that comfortable, you have to be using it for months. So I'd say in the Codex journey, it's something that's fun and exciting to get along, but that's how I would answer that.

18:47

And the surface area there, that's great, by the way, but the surface area there, is it one long-running thread for that type of task that you keep putting more information in, or is it more that you're training the skills?

18:52

Yeah. So create a project folder for thematic stuff. That's where it's happening. That project folder is getting better and better skills that it can use across all projects, but are more specific to what is being used. You can constantly update the agent's MD file for that project folder. So let's say an inbound comes from Slack. Your automation of a skill for your chief of staff folder sees it. It knows, oh, this is for a new piece of data and analysis for this project that Mike's team is working on. It basically starts a thread in that project. And then because there's an MD file and skills for that project, it starts a pass, and you can govern. Like, I never want you to send something back on Slack until I get to read it. It will listen. It's only as good as the skills you make it. So however you invest the time to give it the instructions, it'll behave.

18:58

So that's how I get very confident, where it can just be running, and I know it's not going to just ship something to our CFO channel. Then Sarah Fryer is going to ping me like, hey Kyle, what is going on? Sweet. Thank you, Kyle. I know you're out of time. Thank you so much for joining and for

19:08

This is for a new piece of data and analysis for this project that Mike's team is working on. It starts a thread in that project. And then because there's an MD file and skills for that project, it starts a pass and you can govern. I never want you to send something back on Slack until I get to read it. It will listen, and it's only as good as the skills you make it. So the way you invest the time to give it the instructions, it'll behave. That's how I get very confident that it can just be running. And I know it's not going to just ship something to our CFO channel. Then Sarah Fryer is going to say, "Hey Kyle, what is going on?" So. Sweet. Thank Kyle. I know you're out of time. Thank you so much for joining and for sharing. We'd love to have you on again. Yes. Thanks. Love to see you guys. Thanks Kyle.

19:13

the qualitative side of your closed slides or whatever, you know, slides are doing like these are virtual slides built into this dashboard. We're not doing we're not using like a good like the traditional slides anymore, right. And then what's really neat, you'll see this little mascot bottom, right. This is an idea I had kind of inspired by Codex pets. If you guys like your Codex pets, but for people that are maybe on your team that are less familiar with this, or like how to navigate a new, you know, piece of software you built, basically, there's a mascot you can just talk to,

19:39

and it basically like will take you throughout the the dashboard and answer questions. And this is all based on like a skill where it basically like runs a q&a factory and basically gives this like mascot, like basically like a way to answer and direct you across the dashboard, which is really neat, which is something that I mean, you couldn't even think of building any of this. And now you're just having fun building software that helps your team, like understand what's going on with your business. I love it. I'm going to give people on the team a chance to ask questions. But before I

20:09

before I flip it over, one final question I have for you is, this is a very complicated piece of software. Speaker 2 And it sounds like you built this over the course of a month or two. Um, what did you learn about how to make something like this, that if you were going to start over, and you could tell yourself a couple things from a couple months ago would have helped you avoid some of the pitfalls? Because I know it's like, I know you can make this without coding, but there's a there's a lot to do to like actually make it good. Yeah, that is a very good question. And I would say

20:45

it's it's a perpetual for us, right? Like this is not done every month, we're trying to make this better. I would say the big learning is and what some people will try to do is when they use codecs for the first time, or, you know, chat for work, they'll basically try to one shot stuff, where it's like, I would say take this is like one of the hardest processes we have each month, which is why we thought to even try this, I would say take it to your hardest like process, have it like, like you do it the old school way, but have it gather all the context on all the files, the process, the like slack channels. And then as you go along the way, you basically say

21:19

like, Okay, this is like a recurring way that we're doing this, like, let's codify it with a skill. So you're almost like reverse engineering the process over time. And then as it kind of learns that you can either build like a larger skill or kind of create like a plugin for your team or yourself, or basically enhance your agent MD files along the way. I think one learning to is for the knowledge worker is to don't be scared to try to figure out how all this stuff works, where people will say like, Oh, my agents are doing this, like the way we think about age, or some people think about agency

21:48

or myself, it's like, basically, your codecs has an agent MD file. And then with each year of your projects, which this compute metrics is kind of one of my projects has its own agent MD file, so that kind of my like, codecs is able to orchestrate across all these different projects. So stay organized on the back end, try to recursively improve your skills, you can create skills that improve your skills specific to certain projects, or just overall, like reflect and refine, which is kind of a way that we think about is a skill most people have or build themselves here, which actually integrates really nice with Chronicle,

22:22

if you've tried or use that on the back end of codecs. So become super powerful. But I would say, like, don't, don't try to like, take the time treated as almost like an extension of yourself, you're like training a version of yourself to take this away. And yeah, the investment is maybe it won't feel like a one shot, and maybe it takes like a week, or it could take days, it just depends on how complex the thing is, this is extremely complex data from disparate resources, and different teams working on this. So if this takes you a month, but the nice thing is, I never have to make these slides again, like each month, all the

22:54

data is magically as it's kind of put together, every slide just updates the drop down for the new month comes online. And then we're basically at that last, last mile, just trying to go through and edit commentary. And now it is collaborative. So it's like, once you kind of build this and push it, my our whole strap in team is coming in and editing commentary. So it actually becomes a control and a point where we're have releasing the final stuff before we then kind of have codecs actually export all this to Google slides, since there's some like, you know, shareability, like we can't have our board

23:24

right accessing like internal sites. But anyway, that's kind of I would say some of the advice and learnings, but a lot of people are just intimidated. And I would say just ask it and work with it and treat it as kind of this, like an extension of you, but also something that can like uplevel yourself and don't be scared just because it's has codecs in the name or chat for work. It's like, it is limitless in terms of like what you can do. I love it. So I would say like maybe a summary is like, don't just try to like one shot this like gigantic system all in one do it just in time. So as you have something come

23:59

up in your workday, that is like part of the part of a big project. See if you can make it make a skill that helps you do a chunk of it and build that up over time. And and sort of it sounds like you parallel track things. So you were kind of like, as you're going through the process this month, or a couple months ago, you were like building a system that would help you for the next time, but not for necessarily this time. So all those things are awesome. I know we only got you for one minute left. So anyone else on the team have a question? Yeah, Kyle, I wanted to ask a question. I work a

24:29

lot with finance teams in hedge funds and other financial financial firms. And the big friction point they always have is, how do I trust it? Right? Because I have all this data, and they noticed two issues. And this is with all AI models, but either one, it hallucinates sometimes, and that problem has gone way down, but it still occurs sometimes. So it makes up facts that like weren't in the actual data. And then the other is just like, this is way more common, but like it misses things. So, you know, when there's like a lot of data, it might just like not have looked at that file, or it didn't catch

25:06

that email. So can you walk me through like, how, how do you build trust in this piece of software, like as you're building it? Like what are some steps you take to evaluate? Yeah, that's a great question. And it's actually probably a recurring theme we see when I get to join some deployment calls with some of our enterprise customers. But I would say it's a process, right? It's the same way if you hire an analyst, and they're kind of coming up to speed, they're going to get stuff wrong, it won't be perfect. I will say five, six soul with ultra turned on is extremely it's a big step up. And the sub agent use is kind of built in, which is just incredible to

25:42

where it can kind of multitask and do different pieces of the process all at the same time. But it all comes down to kind of using it periodically where let's say, Mike, it did get something wrong for you, you got to teach it, like basically not to do that ever again. And kind of it, that's how that's how that process goes, where it's like, hey, I didn't like the way you output this or design this, maybe if you're doing more of like a building a build, or if you're kind of auditing materials, for example, I have a kind of, we call it kind of like a Google Drive audit skill, which is pretty

26:11

neat, where before we ship anything, we're able to basically audit a document and look for any inconsistencies between a metric that's maybe mentioned on a memo that's linked to a Google Drive, like to do a slide that's linked to an Excel or link to a G sheet, you basically trace the metric, and then basically tell you like what is wrong. The only way it's able to do that, though, is because I like the first few times like, hey, this is how I think about tracing things. This is how I think about going through these files, like watch me. And now there's a kind of a watch and replay or

26:42

replay and record feature now where you kind of use that where it's like watching how you think about doing stuff if it got something wrong, so that it kind of gains the context, updates your skill, and then that's how you gain the confidence that, hey, that won't happen again. So again, it's kind of like this like one shot mindset where you shouldn't expect it necessarily to be like perfect the first time, but you should know as you're bringing it along, I would say it's like a compounding curve where that first time maybe it's like 70% there, where you're like catching multiple mistakes. The next time

27:10

it's 85, and then basically like a month or two from now, it's at like 90, like that's at 95. Like the payoff is eventually there. It's not less work initially because you need the time to build it and teach it, but the payoff does come and that's how you get more comfortable. Where like now for me, I allow it to take first stabs at work where it is basically trying to incubate work from reading Slack. And I'm not having it just fire off the end results, but it's kind of getting there and it's like, oh, I get to review, figure it out and kind of it's already working on stuff before I even get to read it,

27:40

which is kind of neat. So, but to get that comfortable, you have to be using it for months. So I'd say in the codex journey, it's something that's like fun and exciting to get along, but that's, that's how I would answer that. And the surface area there, that's great by the way, but the surface area there is, is it like one long running thread for that type of task that you keep putting more information in, or is it more that you're like training the skills? Yeah. So create a project folder for like thematic stuff that are, that's where it's happening. Where like that project folder is getting better and better skills that can use, can use across all

28:14

projects, but are like more specific to that are being used. You can constantly update the like MD, like the agent's MD file for that project folder. So that let's say an inbound comes from Slack, your kind of automation of a skill for your chief of staff folder sees it, it knows like, oh, this is for, you know, a new piece of data and like analysis for this project that Mike's team's working on basically like starts a thread in that project. And then because there's a MD file and like skills for that project, it kind of starts a pass and you basically can govern, like, I never

28:49

want you to send something back on Slack until I get to read it. Like it will listen or like, it's only as good as the skills you make it. So like how you invest the time to give it the instructions, it'll behave. So that's how I get like very confident where it's like, it can just be running. And I know it's not going to just ship something to our CFO channel. Then Sarah Fryer is going to pay me like, Hey Kyle, what is going on? So. Sweet. Thank Kyle. I know, I know you're, you're out of time. Thank you so much for joining and for sharing and we'd love to have you on again. Yes. Thanks. Love to see you guys. Thanks Kyle.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note