Open Reader

Compound Engineering Now Works Better in a Multi-Model World

completed 55:39 Jul 23, 2026 Watch on YouTube

Current Status

completed

Video ID

9exWJmbKeMo

RAG / Chat

Enabled
Compound Engineering Now Works Better in a Multi-Model World
Description

Kieran Klaassen and Trevin Chow break down Compound Engineering v3.20. The latest release has a monster 120 merged PRs and adds four new skills to make it easier to get the better quality results, including ce-pov, ce-explain, and ce-commit-push-pr. Every is the most AI-native startup on the internet. Through ideas, software and education, subscribers get the tools to work at the frontier of AI. Start your free trial today: https://every.to/subscribe/plans?utm_source=youtube Follow Every: https://x.com/every

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Compound Engineering should operate as an outcome-oriented, repository-based learning system that uses human judgment at the start and end of work while routing planning, implementation, critique, and review across whichever models and harnesses perform best.
  • Why it matters: It offers a practical control-plane pattern for multi-agent engineering: preserve learned context in repo artifacts, use independent model perspectives to reduce single-model bias, and automate the operational tail of delivery such as PR review, CI repair, and rebasing.
  • Best use: Watch to extract reusable design patterns for Ken's agent systems—especially multi-model orchestration, durable organizational memory, human-in-the-loop requirements work, and autonomous PR operations—rather than as a generic coding-plugin tutorial.

Executive Summary

The speakers frame Compound Engineering as a philosophy rather than a fixed toolset: agents should generate durable artifacts in the repository that encode decisions, research, implementation knowledge, and lessons from failures. Those artifacts allow later agents to avoid rediscovering the same information, while humans reserve attention for defining the problem and judging the final result. They call this the "AI sandwich": concentrated human thinking up front, substantial agent execution in the middle, and deliberate human review and feedback at the end.

The principal release is a move from a single-model workflow to explicit multi-model and multi-harness orchestration. Rather than asserting that one model or IDE harness is universally best, the system can invoke another model for specific reasoning-heavy stages, delegate implementation into another harness, and use configurable defaults or explicit prompt directives. The stated goal is not minimizing token use in isolation, but minimizing cost per successful shipped outcome.

A particularly useful pattern is the CE POV "Oracle" workflow. The originating agent constructs an unbiased decision prompt, independently polls a panel of models, then evaluates consensus or disagreement through as many as three deliberation rounds. The speakers argue this avoids anchoring a second model on the first model's answer and frequently surfaces considerations the initiating model omitted.

The video also emphasizes operational scale. CE Explain converts repository activity into grounded explanations, learning exercises, and periodic reports of material changes; CE Babysit PR automates review-response, CI repair, safe rebasing, and PR monitoring for up to eight hours. For teams, the adoption argument is to demonstrate better repository artifacts and delivery outcomes rather than mandate a particular agent tool, model, or plugin.

Key Takeaways

  • Claim: Compound Engineering's core mechanism is compounding knowledge in repository artifacts so subsequent agents and humans can reuse prior reasoning rather than repeatedly rediscover it. | Evidence: The speakers describe generating files that capture learnings, research, decisions, and how a problem was approached; agents can later discover those files during the next workflow. They recommend automatically offering or running a CE compound step when an agent solves something non-trivial. | Implication: Ken should treat agent memory as versioned, inspectable operational knowledge in the repo or system of record—not as transient chat history or an opaque vendor memory layer. | Caveat: Stored knowledge needs maintenance: the speakers provide a Compound Refresh process to delete invalid documents, update links, and merge overlapping documents.
  • Claim: The recommended operating model is an "AI sandwich": human reasoning defines and refines the problem, agents consume significant compute for execution, and humans apply final judgment and convert failures into future learning. | Evidence: They advocate asking only a few adaptive questions at the outset rather than exhausting users with 30 questions, then letting the AI perform substantial work, followed by feedback or a compounding pass if the output falls short. | Implication: For Ken's systems, optimize the human interface around high-leverage problem framing, exception handling, and review—not continuous low-value supervision of agent execution. | Caveat: This is explicitly not hands-off automation: one speaker says increased auto-mode work had reduced his learning of new technologies and concepts.
  • Claim: Multi-model orchestration is valuable because models contribute different perspectives and capabilities, not because every task requires multiple models or multiple IDE harnesses. | Evidence: The release supports instructions such as using Fable for high-reasoning planning while an Opus session orchestrates the workflow, or using Codex for implementation from within Claude Code. It can also set per-harness defaults in config.yaml. | Implication: Ken should decouple the control plane from the execution environment: retain a preferred interface while routing individual stages to the model, subscription, or harness with the best marginal performance and availability. | Caveat: The speakers characterize the capability as experimental and say they will remove it if it is not used or does not prove valuable; harnesses remain materially different even as their feature sets converge.
  • Claim: Independent multi-model deliberation can improve decisions by avoiding the anchoring bias caused by showing one model another model's initial answer. | Evidence: CE POV Oracle first creates an unbiased problem statement, queries models such as Grok and Codex independently, and has the initiating agent synthesize results. When responses conflict, it can conduct up to three adjudication rounds and end in agreement, stalemate, or rejection. The speaker reports that more than half the time the originating model acknowledges it missed something after seeing independent feedback. | Implication: Use independent panel review for consequential architecture, product, and investment-relevant decisions; do not simply ask a second model to critique a prompt already framed by the first. | Caveat: The "more than half" observation is anecdotal rather than a reported benchmark, and an Oracle response can fail to complete external research, as shown in the Grok example.
  • Claim: Token efficiency should be measured as cost per successful shipped solution, not as raw token minimization. | Evidence: The speakers state they are "not token shy" and will spend more tokens if it improves the result, using the example that a $7 feature can be preferable to a $5 feature if the cheaper result creates additional manual work. They also say accumulated repo knowledge should reduce future web search and grounding overhead. | Implication: Ken should instrument cost against task completion quality, cycle time, and human rework; inexpensive inference that produces weak or unverifiable work is a false economy. | Caveat: They acknowledge the system should not spend extreme amounts—such as thousands of dollars—on simple work, and say more rigorous evaluations are still needed.
  • Claim: Brainstorming and planning serve different control functions: brainstorming is adaptive human-in-the-loop product discovery, whereas planning translates clarified work into an implementation path. | Evidence: The brainstorm skill asks right-sized, often prose-first questions rather than fixed multiple-choice prompts and can create rough web UI sketches to clarify directional choices. The plan and PRD previously existed as two documents but were merged into a unified plan. The speakers generally start with brainstorming unless the request is genuinely unambiguous or already comes from a well-developed product ticket. | Implication: Separate idea validation and requirements pressure-testing from technical decomposition in Ken's workflows, while allowing a fast path for clear, low-ambiguity tasks. | Caveat: The speakers acknowledge individual workflows vary; some users go directly to planning, particularly when requirements are already mature.
  • Claim: As agent throughput rises, delivery operations become a bottleneck; PR babysitting is positioned as an autonomous operational loop rather than an interactive coding aid. | Evidence: The Babysit PR skill responds to code-review feedback, resolves CI failures, safely rebases against a moving main branch, and can monitor a PR for eight hours. It is automatically invoked through the commit-push-PR flow and also runs as part of LFG, their end-to-end workflow. | Implication: Ken should prioritize closed-loop agent operations after code generation—review ingestion, CI remediation, merge conflict handling, and clear human escalation thresholds—because these determine whether agent output actually reaches production. | Caveat: The video asserts that this is better than native harness alternatives but provides no comparative evaluation, reliability metrics, or discussion of escalation controls for risky changes.

Detailed Brief

Repository comprehension and human learning layer

  • Claims: CE Explain is intended to preserve operator understanding when agents increasingly perform work autonomously.; The system can explain a concept in the context of the current repository, explain a PR or recent diff, or teach interactively by withholding the answer and asking the user to reason through examples.; Repository explanations are meant to make code review and management visibility more effective, not merely to audit whether an agent wrote correct code.
  • Evidence: A demonstrated command, "CE explain since last Thursday," produced a weekly report with merged PR count, the most important changes, and a timeline of when those changes landed.; The report highlighted an Oracle-panel change, a PR-babysitting improvement, and cross-harness model elevation.; For a question about Server-Sent Events versus WebSockets in a Ruby on Rails context, CE Explain reportedly found prior agent research in the repository explaining why SSE had not been selected for that server setup.; Commit-push-PR can add a concise explainer to the PR description when a pull request introduces an important new concept, data structure, or term.
  • Caveats: The explanation quality depends on repository artifacts and accessible source context being accurate and sufficiently complete.; The video does not establish whether automatically generated weekly reports accurately prioritize changes across large, high-churn repositories.
  • Implications: A scheduled change-intelligence report can become an executive interface for a fleet of agents and a lightweight artifact for Slack updates, leadership reporting, and onboarding.; Interactive learning modes are a potential guardrail against operator deskilling in highly autonomous development environments.

Adoption strategy in larger organizations

  • Claims: Teams should standardize on quality of committed artifacts and outcomes rather than force every engineer to use a single model, harness, plugin, or prompting style.; Compound Engineering can be adopted incrementally, including by individuals whose teams do not formally use its workflow.; The most persuasive internal adoption mechanism is showing useful results rather than campaigning for the tool.
  • Evidence: The speakers say some organizations git-ignore Compound Engineering folders if team norms prevent committing them, though they view resistance to useful repository documentation as a broader organizational problem.; They state that usage exists among individuals or teams at Google, Amazon, Meta, governments, and among solopreneurs, without quantifying adoption.; They describe a common pattern in which colleagues find helpful repository documents through their own agents and then ask how those documents were created.
  • Caveats: Claims of use at named companies and governments are anecdotal and do not imply official organizational endorsement.; Git-ignoring knowledge artifacts undermines the main compounding benefit because future collaborators and agents cannot reliably discover them.
  • Implications: For an enterprise rollout, define accepted artifact formats, review standards, retention rules, and security boundaries while allowing heterogeneous agent clients underneath.; Track adoption through measurable effects—less repeated investigation, faster review, fewer CI/merge interventions, and reusable design context—rather than tool-install counts.

Notable Concepts & Terms

  • Compound Engineering: An opinionated workflow philosophy in which agents create durable repository artifacts so problem-solving knowledge compounds across later tasks, agents, and people.
  • AI sandwich: The proposed division of labor: human problem framing first, high-token agent execution in the middle, and human judgment plus learning capture at the end.
  • Harness: The agent interface or execution environment, such as Claude Code, Codex, Cursor, or Grok CLI; it is treated as a user preference and capability surface distinct from the underlying model.
  • CE POV Oracle: A multi-model decision workflow that independently polls a model panel with an unbiased prompt and has an orchestrating agent reconcile agreement or disagreement.
  • CE Explain: A repository-grounded teaching and change-intelligence skill for explaining concepts, PRs, diffs, or activity since a specified date, including interactive learning mode.
  • LFG: An end-to-end Compound Engineering workflow that can move from brainstorming or planning through implementation, testing, PR opening, and PR babysitting, including cross-model routing.
  • Babysit PR: An automated post-PR loop that handles review feedback, CI failures, rebasing, and ongoing monitoring, addressing the operational work after code generation.
  • Compound Refresh: A maintenance routine for the accumulated knowledge base that removes stale documents, fixes links, and consolidates related material.

Operator Notes / Why Ken Should Care

  • Prototype a model-panel pattern for high-consequence decisions: independently solicit at least two model views from a neutral problem statement, then require a structured synthesis that records consensus, disagreement, and unresolved assumptions.
  • Implement lifecycle management for agent-created memory before scaling it: ownership, freshness checks, deduplication, deletion, and rules for what may enter version control or shared knowledge stores.
  • Measure agent economics by fully loaded cost per accepted outcome—model cost, human clarification, review cycles, CI failures, merge overhead, and post-ship correction—not token count alone.
  • Add a scheduled, repository-grounded change briefing for Ken's multi-agent projects, with explicit linkage to PRs, key decisions, and material system concepts.
  • Design autonomous PR remediation with bounded runtime, branch protections, allowed fix classes, audit logs, and mandatory escalation triggers; the video's operational vision is useful but does not supply these governance controls.
  • Avoid prematurely standardizing all users on one IDE or model. Standardize interface contracts, artifacts, evaluation criteria, and security policy while preserving routing flexibility.

Source/Metadata

  • Title: Compound Engineering Now Works Better in a Multi-Model World
  • Transcript words: 9477
  • Duration seconds: 3339
  • Timestamp note: No timestamps or chapter markers were present in the supplied transcript.

Transcript

9192 words en Processed in 391.6s

Okay, we are live. Let's see. Yes, we're live. Nice. We'll wait for just a couple of moments for people to arrive. We're also on YouTube, I think. Welcome, everyone. I am here with Trevin, and we're going to jam on some stuff. Or not jam. We're going to explain what we've been working on in the last weeks with Compound Engineering. We're here to ask you to ask us questions so we can answer, and you can listen to our answers. And we can do a fun thing together. So come up with your hardest questions. For example, to annoy Trevin, why do you need more harnesses than one harness? Can't we keep it simple? Great question. I saw it on Twitter. Someone asked, yes, Trevin answered… It activated me. There is something there. I think it's a great question. Because why overcomplicate things? Why do we add stuff? I think the theme of this is, hey, where are we with Compound Engineering? Also, maybe a short intro: what is new? What is the philosophy? Because it all comes back to the philosophy. And why do we add so many skills? Is that even good? Isn't my context rotting? Aren't we doing way too many things? Do we really need this? Do I really need to spend all my tokens on all of this? Yeah. Thank you. Why do you need more than one harness? That person is trolling now. Come on. Okay. I can put it on screen now. That person's got to be trolling. Why don't we start with… Why don't you kick off with the Compounding Philosophy and the AI Sandwich and then… Yes. Let's do it. Okay. So, I'll just start. Welcome, everyone. We're here talking about Compound Engineering, which was born out of a philosophy of… I was just building Quora, and I was like, well, AI is pretty cool. But it makes mistakes. And I just started creating some skills and things. And that turned into this entire ecosystem. And at some point, Treven was thinking the same from a product perspective. I was thinking, hey, actually, this is cool, this engineering stuff. But we need to brainstorm. Really, the problem now is not even the engineering because AI is pretty good at it for most work. So, how do we really understand the problem we're solving? And how do we communicate and capture that in the files? And really, the philosophy is that by generating these solutions or files within a system like your file system, you can compound knowledge. So, you can extract learnings and make sure that the next time you run through this flow, it will actually discover these. And, yeah, that's really the philosophy. Make sure that you learn from what you do and extract how you think about things inside a system so you can move on to the next layer where you can think more deeply. And where we think is at the start and the end. And Treven called this the AI sandwich, which is funny. But, basically, we think at the beginning really hard about what is the problem we're solving. And we activate our brain with questions from the AI. Not too many. Just the right amount. We don't want to exhaust ourselves with 30 questions. We want three questions that will actually get everything out that we need to know. And we hand that off to AI. The AI will do a bunch of tokens. Our philosophy is more tokens is better. We're not token shy. But we want to be efficient. So, the cost per task is important. It shouldn't cost us $2,000 to do a simple thing. But if we can spend a little bit more money to get a better result, we're pushing for that. That's what we want. And we're also optimizing Compound Engineering for the best models while also making sure it works with the less good models. But really, the focus is on just making the best, really pushing the edge. When the model comes back, it should be excellent. Ideally, it's great. Everything works. Everything looks good. If for any reason it doesn't, make sure then to comment on it and run Compound on it so it will learn for the next time. But the point at that stage is to push your own judgment even harder and make it even better. And that's the loop. We invented the loop before loops were loops or graphs. It's now graphs. I think we have a graph too. We have so many skills. You can make a pretty nice graph out of it. But that's where it comes from, what we're trying to do. And my goal always is to delete as much as possible. So it's funny that we have all these new releases. But I think the goal is to delete things we don't need anymore because the providers will do this for us. So that's a big thing. But at the same time, we want to increase the scope of the things we do and the problems we can solve. So, we really started as a software engineering thing. Now it's used for knowledge work. It can be used for almost anything from ideation to product marketing, really. You can even do AI research or really specific hill climbing things. You can learn, you can make educational videos with Compound Engineering if you want to do that kind of stuff. So there's a lot of stuff it covers now. And, yeah, long intro, but that's where we stand. I'm curious, Trevon, what you think, if I missed anything here. No, I think it's great. I think maybe a few things I'd add are, I think, really double down on the deletion part. We try and delete things. There's even skills we've had over a one-month period. I introduced a skill, then removed it a month later because I was like, oh, this actually isn't even necessary. There was a cloud optimization, a permission optimizer. Then cloud released it. And I was like, cool. We just delete the skill. Delete it. Yeah. And so I think this is the advantage coming from the open source side and from the non-loyalty side. Kira and I are indiscriminate about what's valuable. And obviously it's very opinionated with the way we work, the way we like to work. We see eye to eye on almost everything that you see in the plugin. But we don't have a loyalty to one model, one provider, which is actually part of the reason we did this release, where it's actually embracing kind of a new way to work, which is, I think, necessary right now to get the best results. So we can cover that more when we dive into a bit more about the skills. Great. One question before we dive in. I think this might be quite advanced. Any chance you could explain it to me like I'm an eight-year-old? We won't get into that. Sorry, this will be advanced. This will be for the people that actually know already what it is. We'll push this question, but there are a lot of things out that will explain it more for a beginner. We will also work on more documentation. There will actually be a video out soon for beginners as well that we'll do. But for now, let's focus on the people that know it. But check out the repository. There's lots of good information there as well. So sorry, but we're going to focus on the release. One tip you can do, Hugo, you asked, is literally point your agent at the repo. There's lots of documentation we've done for each skill. Ask it your question. Just say, hey, I'm new to this repo. Explain this to me like I'm eight. And look at this folder of documentation. That is probably the best way to start. And you can even get it to create a video for you probably and do something like that. Yeah. And actually there is one skill that's new that you can use for this. So first of all, you need to know what a plugin is. If you don't know that, just read this. You can install it anywhere you want. But there is a learn or explain skill now. Yeah. Yeah. See, explain. And you can see here in docs skills, you can see what it does. So for all the skills, you can say see, explain, and then point to this repository, for example. I've got an example actually, Kieran, of one use of it that I use quite a bit since you created that skill that I could show. Cool. Google Compound Engineering GitHub, you'll come here, or anything. Yeah. Okay. Let's start with one skill. So maybe, Trevyn, you can do, we celebrate, we release the new version, and then you show that skill. And then we go from there. Sorry. Did you have a skill in mind, or do you want me to just pick one? No. You said you had an example for the explanation. If you don't know that, just read this. You can install it anywhere you want. But there is a learn or explain, explain skill now. Yeah. Yeah. See, explain. And you can see here in docs skills, you can see what it does. So for all the skills, you can say see, explain and then point to this repository, for example. I've got an example actually, Kieran, of one use of it that I use quite a bit since you created that skill that I could show. Cool. Google, Compound Engineering, GitHub, you'll come here or anything. Yeah. Okay. Let's start with one skill. So maybe, Trevyn, you can do, we celebrate, we release the new version, and then you show that skill. And then we go from there. Sorry. Did you have a skill in mind, or do you want me to just pick one? No. You said you had an example for the explanation. Oh, yeah. I've got an example. I've got two examples actually. So let me start with the problem. Okay. So I didn't expect this release to actually have this skill when I had started thinking about this problem. But then it quickly became my favorite skill, even though I think a lot of the other work is going to overshadow it. But I think this will quickly become the most commonly used skill for a large percentage of people. So here's the problem you've got. You're building. Your agent then says, hey, we could do this or this. Which one do you want to do? You often sit there, and maybe half the time you have an answer. The other half of the time, you have no idea. You're like, I don't have enough information. I'm not sure which one's better. So you might say to the agent, I'm not sure. Help me decide. It then has its own biases and whatever. So what I wanted to do was I got tired of going back and forth. Actually, bring up the POV skill. We'll start with that one, and then I'll bounce to this one after. I got tired of copying and pasting around and trying to adjudicate the different models. And what I was doing was saying to the agent, hey, here's the point of view from Claude. You're Grock. What do you think of this? And the problem with that is that I've already biased Grock to know what Claude thinks. And so, depending, the way the models work is they get biased, and they start to evaluate that at a starting point. So then I started constructing better prompts, which was more unbiased. Hey, Claude, describe to me in an unbiased way what we're trying to decide so I can give it to another agent. I'd copy that. And then I would go back and forth, and I would copy Grock's response, give it to Claude, and say, this is what Grock said. So what we did was the first version of the skill looked at your repo to answer a question. Sometimes we might say, hey, should we adopt better off for the repo? So your agent can actually look at your repo, do outside research, and actually give you a point of view. But what I realized after I released this first version of the skill was to, and which was inspired by a friend of mine, Pishman, actually, who started talking to me about this and aligned with some work that I was doing. But I quickly then expanded this to say, well, why can't this do a better job actually giving me a multi-model point of view? So I then created this idea of the Oracle. So in a very quick way, you can actually get a grounded opinion, unbiased, from multiple models, and then have the model you're starting in actually orchestrate the replies. So what I'll do is let me share my screen here. Okay, so here is an example I started playing. So this is, I use or or orca IDE for my idea of choice here. So in this repo, this is literally the Compound Engineering repo. And I got this idea in my head about giving the ideation ideas labels, real-world labels, like when it came up with a crazy idea, to give it an analogy that you could easily think about. And so I started talking to Claude. Claude early in the conversation is like, yes, this is a great idea, or this is bad. So what I just did, I said slash CPOV Oracle this. And Claude then uses the skill. It constructs an unbiased prompt. It went to Grok. It went to Codex, and it basically came back, and it basically says this is the answer. And you can see here it says, hey, the basis grounding: no. And it explains why. It explains what's going on here, what it thinks. And you go down here, it explains its verified facts. It cites them. Grok said this, or actually Grok was unable to actually do the web portion for some reason. But then here Codex cited Apple naming guidance and actually went through it. And this is all stuff that Claude didn't do originally. So it's not that they are better. The whole point is that they're different. And what's amazing is I can't tell you, more than half the time probably, Claude will say I completely missed something and fill in the blank. Another model pointed this out. And now I've changed my mind. So what the skill does is it starts with an unbiased polling of the Oracle panel. It gets it back, it looks at it, and it sees, well, one, if we all agree, we're good. And it'll actually say this strengthens my position. Here's the recommendation. If they disagree, the orchestrating agent will actually adjudicate it and go back to the panel and say, hey, here's what I think. What do you think about that? So now the model, the agent's already been grounded in a point of view, and now you're challenging it. So you actually get up to three rounds of this. And in my evals and things, three rounds is actually a great amount of time going back and forth with it. And you end up with agreement, a stalemate, or a rejection where they'll say, hey, don't do this, or you should agree and do this. And so this is a skill that I use a lot. And so right now, a great way to use a skill is next time your agent asks you a question, you're not sure, invoke CEPOV and say Oracle this. Or you can even do things like ask Grok, get Grok's opinion, get Grok's opinion in cursor. And so there's lots of ways to invoke this. And also the other way around here in Codex, you can ask get Fable's answer from Claude. So that's a clever way to do this. So that's one thing that I'm really excited about. The other skill that I would say is, and we can come back to the things in Q&A, is Kieran's slick skill, which is the C.E. explains. That's where we started. So this skill is interesting. Kieran's got a really great teaching mindset where he will lean in on the I want to understand what's going on, even though I'm getting agents to do a bunch of work. And it's pushed me to really use the tools that he's given me through the plug-in. And I use it a bunch to explain what's going on. And the nice thing is it's used to explain PRs. Agents do a lot of work. It's really great to catch up on things. One thing I do a lot is I'll catch up on repos, some repos that I work in. One slick feature is C.E. explain since last Thursday. It will actually go in and explain to you what happened. So actually share this right here. Let me share and let me show you. See here. I should be able to share here. I ran since last Thursday on Compound Engineering. So what it does is it does a full explanation of your week taught back to you. It tells you how many PRs are merged. It identifies which ones are the most important. So it tells me about the Oracle panel that I added to common engineering point of view. It talks about the babysitting skill and how I account for time and wall clock time. And the last one was cross-harness model elevation, like how cloud code works and how it elevates. And what it does, interesting for those that work in teams, and especially those that are leaders, managers of teams, or whatnot, or managers of fleets of agents, it gives a timeline of what's happened. And it highlights, of the three things that are important, here's when they landed in order in the timeline. So I found this super impressive as a way to learn what's happening and follow up on things. And I think there's a lot of people probably out there that will empathize with this because with more agents, there's more work going on, more threads. A lot of people get overwhelmed with how do I follow what's going on without reading the chat logs of everything and manually going in. So a way to use this skill is, for those working teams, if you have lots of agents going on, run it on a loop, run on a schedule every week. And what it does, interestingly, for those that work in teams, and especially those that are leaders, managers of teams, or managers of fleets of agents, is it gives a timeline of what's happened. It highlights the three things that are important. Here's when they landed, in order, in the timeline. So I found this super impressive as a way to learn what's happening and follow up on things. And I think there's a lot of people probably out there that will empathize with this because, with more agents, there's more work going on, more threads. A lot of people get overwhelmed with how do I follow what's going on without reading the chat logs of everything and manually going in. So a way to use this skill is for those working in teams. If you have lots of agents going on, run it on a loop, run it on a schedule every week. Run it since last Monday, or whatever the date is, and get the report and just read the report. And you could archive those reports as well and check those in, share them in Slack to your team. Repurpose them for presentations to your broader management team is a great way as well for those that work in big companies. So those are the two that I would, I'll pause there, those are two that I would probably start with. Yeah, these are great. And where this came from was the idea of compounding knowledge is great. But if you keep compounding your database or in your projects, there's another point where you need to learn and compound, which is in your head. And the more I'm doing things on auto mode, I just realize I haven't learned as much. Yes, I ship a full product and I do a lot of work and I learn how to do agentic engineering and all of that. But actually new technologies or new concepts, I've learned less. And I hear this around me from other people that leaned into the agentic engineering more. And explaining, explain is kind of a way to compound knowledge, but then in our own brains, how to energize what we think ourselves and a way to teach about changes. So different things you can do here, it can be a concept. So you can say, hey, for example, I use lots of WebSockets in Quora, but obviously SSE is mostly used for streaming stuff in most places. So I'm like, oh, but why is SSE less used in Ruby on Rails, for example? And explain that concept to me. And obviously you can do this in a normal setting, but the cool part is it will ground it in the repository. And in this case, it actually found research that was done by an agent where it did research to use SSE for streaming for the agent. And it decided not to do it specifically because of the server setup. And I had no idea, I did not read that, but it brought it out of the context and it grounded it in external things, which is really cool. For the diff stuff, this is mostly what you said, hey, what was shipped? What was changed? Do I need to learn anything? I think that's interesting. What it will do also, if you run PR or commit push PR, which is another command, it will do a mini version of explain. So in your pull request description, it will, if there is a large new concept introduced in that pull request, explain what it is. And we'll say, hey, you can learn more by running this command on it. And it will explain these things. So in larger features where you haven't looked at all the details of the plan, this is great to make sure that you catch larger moving pieces and things, and you can integrate. It also helps your team members or people that work in the same repo. Yeah. This explainer that happens in the PR description is really clutch because oftentimes you read a PR, you don't really realize there's a new concept in this now being introduced, either a term, a general idea, a data structure, or something important to the system. So it's a really great way to pay it forward to people that are reviewing your PR, but also, again, agents read PRs too. So it helps plant little breadcrumbs to compound knowledge across the workflow. Yes. So it's part of a workflow. There is one other part. There's also a learning mode. So it can give examples and then not give the answer, but ask you to walk through it. So it's really story-driven. It's not about giving the answers, like, hey, this is how the thinking works, what do you think? So you want to learn stuff or prepare for a meeting. For example, if you want to present about something and they ask you details, or if you're a true Vyde coder and you have no idea what you've done, but they're going to ask you questions and probably you should know exactly why you did what you did. This is also the one to run where you will know all the answers to all the questions, hopefully. And I think this is super interesting. I'm actually already starting to write an article about this. So I will share more about the thinking behind it, but I think this is very important, especially if you're in the agentic engineering for a while, you'll realize, oh shit, I need to keep learning. And some people say read the code, and that's a version of this. I think CE explain is better. But if you hear people say read the code, this is kind of my variation on that thing. And the reason is not read the code because you want to check whether the AI did anything. You want to read the code because you want to learn. And I think CE explain is hitting that same solution. So there's a question someone asked, you want to highlight that question about token efficiency. We can talk about that. Oh yeah. Okay. Sorry. This is very quickly. I just said, hey, how does compounding work? CE explain uploads, and this is a document that comes out of it as well. So it does cool explainers as well. Cool. Let's do some questions, maybe. Yeah. That one there. This one, we talked about earlier, maybe before preamp, before you arrived. Our goal is to be token efficient, but not token deficient. So we are not trying to save tokens. We're trying to use tokens in the best way possible. And we will let it rain tokens if we feel it's better. And so we pay attention to how much tokens are being used. But we're not trying to save tokens for token's sake. That's the best way that I can say this. I think there are some people who are, there are others. I think there are other plugins and other skills that are on a different spectrum that care about different things. And rightly so, they have a different focus. We're not as focused on that. We're focused on outcome. And almost the way you think about it is token per unit of outcome or token to value. We're trying to maximize value, and we're willing to spend as many tokens as possible to reach that value. Yeah. Yeah, if we can get more out of it, but we do understand no one wants to pay hundreds of thousands of dollars per month doing simple things. Because in the end, the best thing is not looking at tokens, but it's how much does it cost to ship a solution. And I think that's what we're really trying to optimize for. Does it make sense to pay $5 for this feature? And is it better if we spend $7? Then I think it's worth spending $7 instead of $5 and then still having to do a lot of manual work. The goal is to just get the most absolute best thing out of the models. And at some point it just keeps running in loops and it doesn't improve anymore. And that's our goal. Are we always right? No, but that's what we're trying to do. And every model is different as well. So that's also, maybe that's a way to go into the multi-harness update because that's a big one. Yeah. So multi-harness. So this relates a little bit to the question that Seth asked about how about the right model for the job. I think it's not, there's a trend, I don't know, in my article I published this morning, eight weeks ago, I was primarily in Opus and Cloud Code. That was my workhorse, 90% of my time was in it. I was slowly using codex more. I think the OpenAI team is just killing it, crushing it with adding more features and behaviors in that. But I think the key is not we're trying to use a harness for a harness's sake. Harnesses are preferences. People like them better. But to be honest, I think a lot of the harnesses are used because they give you access to a model. So there's obviously exceptions. There are some harnesses like cursor that are multi-model. I know Microsoft has CoPilot multi-model. I think it's not, there's a trend. I don't know, in my article I published this morning, eight weeks ago, I was primarily in Opus and Cloud Code. Like, that was my workhorse, 90% of my time was in it. I was slowly using codex more. I think the OpenAI team is just killing it, crushing it with adding more features and behaviors in that. But I think the key is not we're trying to use a harness for a harness sake. Harnesses are preferences. People like them better. But to be honest, I think a lot of the harnesses are used because they give you access to a model. So there's obviously exceptions. There are some harnesses like cursor that are multi-model. I know Microsoft has CoPilot multi-model. But there's other ones that are not. Cloud Code is not. Codex is not. I know there are hacks and people will come out of the woodwork and say, you can do this and that. But by and large, those harnesses are not. And so the question is, the two most popular harnesses, Codex and Cloud Code, a lot of people use that. So the idea was born from, what if you want to reach across the aisle and you want to do work? And the most common initial one was, hey, Trevon, how can we get Fable to do planning on this? And the idea was, well, flip the model to Fable, run planning and brainstorming. Then remember, you better remember, flip it back to a different model before you start the work. Otherwise, you're going to run it fully in Fable or something like that and kill your credits or, when it was API billing, pay through the nose on money. So it was born on that idea. And then I started to think about, OK, wait a minute. What about beyond Fable? What else can we do? And so it really got interesting of creating a bit of what started as a bit then became a lot, which was a deterministic way for the skill to be able to do. And so it came up with a pretty good system, pretty happy with. I think there were going to be bugs and things, but here's what you can do. And it's hard to explain succinctly because it's so expansive. But I'll give some examples. Many people run CE plan and they give an idea and they want to plan with it. They don't brainstorm. They just plan with it. So what you can do is put use Fable at the end of the day. So what you can do is have your session in Opus in whatever reason level you want. You can run CE plan. It will be orchestrated by Opus, but it will use Fable for specific high reasoning steps to try to get the value of Fable while not needing to use Fable for the whole thing. So it's a way for us to manage tokens and things like that in a way. The other way you can do it. A lot of people will say is, hey, I have unused subscriptions that I want to use for implementation. And I really like those models too. But I like to be in cloud code. Awesome. CE work use codecs. So if you're in cloud code, it will then do the implementation in codecs and it will delegate the work there and it will use the plan that's created. It will paralyze as much as it can and it will go out and go figure that out. You can even do things like use Grok in the Grok CLI, use Grok in cursor, that will work too. There's lots of permutations here and the Kiran's LFG skill. I've bolted more into that. You can even do some things like LFG, the feature name, plan with Fable, implement with codecs. So you could be in cloud code and you could go do that. So this is the LFG implement with Fable. So that example he's highlighting here is a really common use that I do. I will brainstorm a feature, go through the interactive Q&A that the brainstorm does, look through the visual examples, really sweat the requirements and the scenarios. Then I'll say, great, I'm good with that. LFG implement with codecs. It will go run the plan. It will go do everything in cloud code in this case. And then when it comes to the implementation, it will implement with codecs and will paralyze units of work that it can. And the LFG is great because it will continue on and test it and open the PR and babysit it and get it ready to merge. So that's like, there's probably a lot more that I'm missing, but if you pop up a stack, you think the point of view skill gave you a way to leverage the knowledge and opinion of multi-model to adjudicate an opinion and point of view. The other skills now have been given powers to optionally reach across the aisle to use planning in different models and use different harnesses if you need to. The other skills now have been given to you in planning and implementation. And again, if you don't want to do it, good to go. It doesn't do anything. But there's also, for people that want to always do it, there's config.yaml file. We support that in the documentation. We actually have config changes that you can set your defaults and the skills will read the defaults. And there's a file in the repo that describes the config options that you can define even per harness. So you can say, when I'm in Claude, use codecs for the implementation and whatnot. But generally speaking, a lot of people end up using it with trigger words in the prompt because they want it to be more in control. So that's all. Pause there. I think that's where I would say there's a lot of power in this that I'm super interested to see how people end up using this. But I think that probably the most popular ways are going to be the point of view skill to get multimodal and then just planning with Fable and then implementing with codecs. It's probably what I suspect is going to be really common. And then Grok probably gets sprinkled in there, probably by a percentage of people. Yeah, great. And also, this kind of worked already in Cursor since Cursor is multi-agent. And this is great that it works now everywhere because some people just love one harness over the other. And that's like, here's the question, do we need more than one harness? We need more than one harness because people have preferences. It's kind of like saying, why do you need more than one browser or why do you need more than one? Yeah, exactly. Yeah. Anyone, the point is that we don't want to say cold code is the best so you can only use cold code because it is not the best. It is very good like any other. It's the one I would use, and I use Cursor a lot as well and I use Codex too. But it is whatever you want to use for what step as well. And the really cool thing is, you can now hand off things way easier. And you don't need to copy paste. It's just more flexibility. And I think enabling people to experiment with all these is actually, we can learn what works well and what doesn't work well. And if this is something no one uses and we don't use in two months, we'll just remove it. But I think this is the only way because we're not linked to anything. You can say, oh, why don't you just do open code or something like that because then you can do everything. I think the harmless is important and they do really work differently. Yes, they come closer and closer, but they are very different. And especially for people that are more engineering, they maybe prefer Codex or 5.6 or some other model. Maybe people want to plan or brainstorm. Maybe Cloth is better, like an Opus model there. And it's very nice to choose. And I think probably the biggest thing will be you just use a Fable and you're like, oh shit, I'm through. Let's use my OpenAI subscription or use my cursor subscription. And you can just use all the subscriptions. And the cool part is, even though one is better than the other, if we don't try out the different models, we will never know if it's better or not. Yeah, the way you think about it, everyone, it's a little bit selfish in a way. If you look at this, Kira and I are building combat engineering. We use it every day. So in many ways, we're building it for ourselves. And it just happens that there's a lot of people who also want to use it. So we also bounce between harnesses. So there's that too. And so I think we've obviously taken into account some broader needs and expanded in ways that probably we wouldn't use it. Let's use my OpenAI subscription or use my Cursor subscription. And you can just use all the subscriptions. And the cool part is, even though one is better than the other, if we don't try out the different models, we will never know if it's better or not. Yeah, the way you think about it, everyone, it's a little bit selfish in a way. If you look at this, Kira and I are building combat engineering. We use it every day. So in many ways, we're building it for ourselves. And it just happens that there's a lot of people who also want to use it. So we also bounce between harnesses. So there's that, too. And so I think we've obviously taken into account some broader needs and expanded in ways that probably we wouldn't use it. But I think it follows a through line of how we would tend to use things and what we recommend, which is why you feel like compound engineering is opinionated. It's because it is. It's our opinion about things. So, yeah. Yeah. This is a good question, Dave. So, I can take this one if you want. Brainstorm and plan. OK, so planning pre-existed. And I added brainstorm, as this is actually the first skill I think I added when I joined, and it's evolved. Yeah, here's the way that the evolution was. Plan started. There was an implementation plan. And what I realized was, hey, we need more requirements. We need more brainstorming, ideation work as a normal type of product work. So I had a brainstorm, and initially they wrote two different documents. One was a traditional PRD and one was the implementation plan. About a month ago or three weeks ago, I merged the documents. They're now a unified plan. Logistically, it's simpler to manage. It's easier for agents to keep updated and follow. But the way I do it is, one, there's lots of ways to do this. But here's what I would generally say: if there's no ambiguity at all and you're just creating, create a technical information plan. You just go straight to CE plan. Totally fine. I think that's rarer than most people actually think. Lots of your ideas are bad ideas, and lots of ideas could be better if slightly adjusted. Brainstorm is really the smartest product person in the world working with you to ask you some hard questions. Why are you doing this? Are you sure? But it doesn't just ask you cookie-cutter questions like why are you doing this and ask you every time. It's right-sized to the problem you have, and you can come with the craziest idea or the most basic idea. And some of the magic in there is it's adaptively asking you questions to help refine. It's trying to ask you in a way that's really effective. So we don't just bias you to pick from multiple choices. There's a very deliberate decision that I made originally. The first question is typically a question you have to answer in prose. Because rather than saying which one do you want, one, two, or three, that's too limiting at the beginning, especially the more ambitious the idea is. And then what I added recently was many times the brainstorm skill will need you to make a choice between something like, hey, would you rather do it this way or this way? The brainstorm skill will now offer you an ability to sketch it, spin up a web server, and actually show you web UI, not for final visuals, but more of a direction. Imagine you're designing a chat app and you're trying to show a notification somewhere, something like that. Sure, someone could explain it to you in words, but many times the easiest way is to actually see it. So the brainstorm is an incredible way to explore scenarios, requirements, and pressure-test you before you get into the okay, agent, let's figure out the right way to implement it. Do some work to figure out how would we break up the work, how would we approach this, and tackle this before you go into actual implementation. So that's where, I think, there's lots of ways to think about this, PRDs traditionally and whatnot. But the way you think about it is just zooming back, taking language and dogma out of what process is. I know the Rails community has a lot of shaping stuff that I know that Dave's talking about here. Get clarity on what you're trying to build and who you're trying to build it for. What is, are you being ambitious enough, or are you being too ambitious, and then refine from there? That's really what these two skills are for. So generally speaking, what I say is start with CE brainstorm and then move on to CE plan afterwards. I know a lot of people, my friend Matt, use only CE plan, and people do that for sure. But I think that there are many ways that can actually be limiting because you're not actually going through an exercise of maybe proving to yourself that this is actually a really great idea. Yeah, and how I look at it also, the brainstorm skill is human in the loop. It is supposed to be activating your brain and making you think, and CE plan is just you give it one thing and it builds it, basically. It's part of the first step in LFG. So I don't run plan ever. I run /LFG. So it's either that I brainstorm or I run /LFG, which includes the planning. And if you work in a team, you might get a ticket and a product person already thought about something. You just need to implement it. Maybe you can do the brainstorm again. I probably didn't miss something. But let's say you have a very smart product person and you're going to create something and they hand it off to you to implement. I think then it's really the plan. So it's also product versus engineering in the old sense, even though we do that now. Maybe this one is also, this question is, hey, I love CE. Use it on my personal things. Any tips for using it in larger projects' repositories where it isn't using CE as the workflow yet? Yeah, my thought, and this gets asked a lot, actually, a lot of people use it in large companies very successfully. And I think there's lots of ways to do it. There's not a one-size-fits-all. But generally what I would give advice, Kyle, is if your team is not open to docs and compound learning being committed to the repo, I think there's probably a larger issue at play. But what you could do, some people do, is they git-ignore the folders, and that's okay, too, because they don't know. But I think I would say aside from the artifacts that are being checked in, no one knows that you're using whatever skills you are. So I would go as far as to say, listen, you should be checking in documents and explaining to people, and they shouldn't care how you're doing the work. If they don't think that the documents you're checking in are useful or the code is good, that's something different. That is saying the skill or your method, your tool, isn't producing the result that your team expects. So I think when you frame it that way, most people don't need to then worry about it. But I do know in large companies, large teams, there are sometimes tricky dynamics around, oh, we don't use that plugin, we use this plugin. And I think a lot of teams are coming around to the fact that plugins, harnesses, and models are all a lot of personal preference. And there's a lot of mixing and matching going on. And so I think the best teams out there right now that are embracing agenda coding are open to that fact and treat that as a ground truth and figure out what are ways the repo should be set up and what are the team processes and culture that need to exist to embrace that. So everyone, so Kyle, you can use compound engineering. Someone else can use superpowers. Someone else can use no plugin and skills at all and just bare prompt, like trad prompting, I call it, on their own. And so all those things should be possible. And so I think that's probably the best way to do it. If the question is about how do you evangelize it, I would just evangelize it by doing. Show by doing. Just continue to do it and model it and share your success stories. And people tend to come around to tooling that gives outsized results. Yeah. And there's someone saying for so it's great, but corporate environment is hard to embrace and push it horizontally and vertically for context. And you've said here. Yeah, again, if you just start using it, that's really the best way. And just share how you use it in your workflows and how it helps you do better work. It could be a CE explain, or it could be you creating a document that extracts everything that has been shipped in a week. I call it on their own. And so all those things should be possible. And so I think that's probably the best way to do it. If the question is about how do you evangelize it, I would just evangelize it by doing, just continue to do it and model it and share your success stories. And people tend to come around to tooling that gives outsized results. Yeah. And there's someone saying, so it's great, but corporate environment is hard to embrace and push it horizontally and vertically for context. And you've said, here. Yeah, again, if you just start using it, that's really the best way. And just share how you use it in your workflows and how it helps you do better work. It could be a C explain, or it could be you creating a document that extracts everything that has been shipped in a week. It could be simple things like that as well, where you can start. One really good one is Trevon's CEDoc review, which is great. It's just grilling a document and asking hard questions about the document. And that one is a good starter as well. But I know it's used at Google, Amazon, and Meta. Not everyone, but there are teams or individuals that are pushing this within gigantic companies. I know it's used by governments. I know it's used by lots of solopreneurs like you. It's very useful in different settings. I think it's not as, you don't need to embrace it fully. You can just individually embrace it and inspire. I think that's the best way to do this, because if you say, okay, now you need to use this plugin, it's a hard thing to say to engineers. But it's easy to show them how cool it is or how things work or how it is different than superpowers or getting shit done or any others. One of the best, most common stories that I hear for people that use common engineering is they never mention it. The compounded learning ends up in the repo, and other people's agents find the documents, and then someone asks them, how did you create this document? Where did it come from? This is really helpful. How do I contribute documents like that? So that's another way to show by doing, so everything Karen said, everything I said earlier, I think really applies here. And so I think thinking about it from that lens of don't evangelize your tools, evangelize the results, and then the tools can be a means to an end. One other thing you can say as well is it's more token efficient. We need to get some hard, we need to run more evals, but we ran some evals, and I know some people that ran evals that showed if you have these learnings tokens, it will just be more efficient because it's not going out of its way to search the web and ground itself and everything. So once you have these documents in your repository, it will make decisions more quickly with more confidence. So it is actually more token efficient at a higher quality in the long term. So you could even sell it as it might be saving us tokens as well. Yeah, I think the way that you could explain to someone is saying, listen, your agents, all of your agents, are learning things as they go. They stumbled upon things, they got code review that they did something, they learned something that was non-trivial. Why are you forcing another agent to relearn that again when they work on the same area of code? So the compound skill creates documents to actually compound those learnings. In my agents MD file, I have essentially an instruction that says, and it's actually in the document, actually in our skill docs, where it basically says when you encounter something, when you solve something non-trivial, offer to run CE compound skill. And great. So often sessions go, it'll just run, and I actually have it auto-running it when it encounters it, and it'll become part of the commit. And so Dave, you had mentioned something about pruning docs. We do have a compound refresh skill, which essentially does this for the compound engineering documents. So I actually have that in some repos scheduled to run every week as well, where it'll run and do it and then commit essentially a cleanup so that it will delete docs that are no longer valid, update links, and also combine documents when they should be combined. Yeah, this one is great. Yeah, there are many, many skills. Even I sometimes am like, I don't even know everything we have. Because Trevon is adding so many cool things. Other people are adding things as well. I would say everyone should actually check out Combat Engineering Docs Skills and just browse through here and click on things where you're like, oh, this is interesting. So, for example, if you don't know what this is, you say, oh, what is this? And it's a nice explanation with the TLDR on the top for what it can do, what it does, and why. And yeah, go check all these things out because they're pretty cool. The configuration is one I need to use more. I know I can configure way more. Here you can set default models. Now you can also say, hey, use these skills or these reviews. So this is to customize it for your team if you want to standardize things inside a team. Oh, sorry. I'm not screen sharing here. There we go. Boom. Boom. So here you can see it. So configuration. This is the one I was telling about. And you can see here all the skills here have pretty good documentation. So you can just read everything. This is the Compound Refresh, for example. And it says, hey, what to do with it, when to use it, what it does. And sometimes there are different modes like headless mode or interactive mode. This is for the pipeline that we have. So yeah, read these through or have your agent go through these because you can learn a lot from these and discover other things you haven't even thought were in here. And if it's not in here, please contribute and edit. There's one skill. I know we're probably going to cut soon. There's one skill I will say that got an immediate upgrade that is so crazy to me. Why we haven't done this before is the babysitting PR skill. Oh, yeah. I spent a crazy amount of time on this skill. I thought this was going to be super easy. No problem. The amount of time I spent to get this working well is pretty crazy. So what it does, everyone does this. You've seen Claude, all the harnesses have it. This is better. I promise you, better than all of them. The best. What it will do is it will babysit a PR after you open it, respond and figure out all the code review feedback. It will resolve CI errors. It will auto-rebase safely and figure out if main is shifting under you, how to get it current. All the annoying things that you hate when your agent opens up a PR and you realize, oh, there's four pieces of feedback and main now moved. And how do I rebase this and all the stuff. So when you use the commit push PR skill, you will open a PR with an awesome PR description that is focused on value, the right thing, avoiding all the traps, all the bullshit agents do by default based on their training data of bad PRs. It will include an explanation from C explain, and it will then auto-babysit the PR, and so it will babysit the PR for eight hours if you want. And so you often don't have to do anything to enable babysitting. It will just automatically launch if you use a commit push PR skill. But if you don't want to use it, you can use babysit PR individually and just pass it a PR link or whatever. Some of the ways that people are using it is in their agents MD, they will have a direction and say when opening a PR, always use the C commit push PR skill, and with babysitting, and then that will always make sure your agents are following the skill, even if they try to do something different. So this is, I think, a little dark horse one that is like, I know I said POV skill is my most used. I think babysitting probably is the most used one, but not interactively. It just works. It just runs. Is it part of LFG. It will run as part of LG because it will get invoked out of the commit push PR. Yes. Okay, my favorite skill is this one, LFG. If you don't know it, check it out. Basically, you just say LFG brainstorm, LFG brainstorm, or LFG with a plan. You can now say on codex with this model or whatever you want. But basically it will do all of this. So if you're new, try this out and it will use everything under the hood. Also, if you don't call any of these, Cole says it kind of works. And it will just find things from compound engineering, and I know there are other unnamed libraries that will inject hooks to hijack everything. We don't do that. We just have descriptions. In my Hermes box, I only have competent engineering.