Every

How to Build a Multi-Agent Review Swarm

1904 summary words 8 min summary Watch video

Start with the signal

8 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: A coding agent becomes materially more useful when it operates from the company’s canonical context, executes remotely, and is governed by a multi-model, domain-specific review loop that it must resolve before opening a PR.
  • Why it matters: The video provides a concrete control-plane pattern for turning planning artifacts, codebase access, CI feedback, and code review into one asynchronous agent workflow rather than a sequence of manual handoffs.
  • Best use: Use it as a reference architecture for an agentic software-delivery loop: spec-to-task generation, remote execution, structured work journaling, CI remediation, and swarm-based quality review.

Executive Summary

A Notion AI engineer describes an internal workflow in which Notion is not merely a documentation layer but the primary interface and context store for coding agents. He starts tasks through unstructured voice input, asks the agent to explore the codebase through a remote development VM, and has it turn the resulting understanding into a detailed task with code pointers, requirements, and self-verification criteria.

The implementation agent then receives the linked task, builds the change in a cloud-hosted coding environment, maintains notes as it works, opens a PR, and monitors CI. In the demonstrated case, it started at 12:53 and completed at 2:27 while the engineer attended meetings and initiated other work; the stated goal is not one-shot perfection but removing execution busywork after careful human planning and constraint-setting.

The most reusable technical idea is the “review swarm” skill. It partitions a change set into domains such as front end and back end, assigns separate correctness and maintainability reviewers for each domain, runs both GPT and Opus over those reviews, aggregates findings through a top-level agent, and loops the implementation agent until the resulting issues are addressed.

The broader operating lesson is that consolidation matters: connecting the agent to the canonical task system, company knowledge, codebase, remote compute, and CI reduces context drift and app-switching. The speaker also argues that teams should build bespoke agent skills around their own standards rather than wait for an off-the-shelf workflow to fit.

Key Takeaways

  • Claim: Treating the work-management system as the agent’s native operating surface can turn organizational context into executable engineering tasks. | Evidence: The agent is asked to investigate a missing model-picker migration using access to the codebase, then creates a task in Notion’s M2 milestone project containing code pointers, requirements, and self-verification steps. | Implication: Ken should design agents to write back structured plans, provenance, and acceptance checks into the same system where work is assigned, rather than leaving analysis in ephemeral chat threads. | Caveat: This depends on the task system having reliable permissions, accurate source material, and tool access to both the relevant workspace context and engineering environment.
  • Claim: Voice-first, free-form task intake can yield richer agent instructions than tightly edited prompts because it captures reversals, nuance, and contextual detail. | Evidence: The speaker begins tasks by “yapping” through the situation and says models tolerate, and may benefit from, corrections such as “actually, never mind, do that instead.” | Implication: Use conversational intake for discovery and context capture, but have a planning agent normalize it into an executable task contract before an implementation agent acts. | Caveat: Free-form intake still needs a later conversion into explicit requirements and verification criteria before autonomous implementation.
  • Claim: Remote, preconfigured compute is a key enabler of asynchronous coding agents because it eliminates local-environment and repository-state friction. | Evidence: Notion AI can invoke a computer; instead of a bare Vercel sandbox, the workflow uses internal “Boxy” coding VMs with remote compute on demand and codebase access. The speaker contrasts this with managing whether a local repo, Codex installation, and environment are current and correctly connected. | Implication: For Ken’s agent systems, standardized ephemeral or remote workspaces should be treated as infrastructure, with deliberate environment bootstrap, scoped credentials, and auditability rather than ad hoc local-machine dependence. | Caveat: The transcript does not address isolation boundaries, credential handling, or how remote VMs are prevented from accessing unintended production resources.
  • Claim: The human’s highest-value role shifts toward planning, constraints, and tool provisioning, while the agent absorbs implementation and remediation busywork. | Evidence: After the engineer supplies the task and a review-swarm instruction, the agent builds the feature, opens a PR, and fixes CI issues autonomously. The demo reports a 12:53 start and 2:27 completion while the engineer was in meetings; the speaker explicitly says the value was not simply asking it to make no mistakes. | Implication: Measure agent leverage by reduction in engineer coordination and repair time after planning—not by one-shot completion claims—and retain human responsibility for scope, acceptance criteria, and release decisions. | Caveat: The example is a single successful internal change and does not establish a general reliability rate, review quality benchmark, or safe autonomy threshold.
  • Claim: A review swarm should separate correctness review from maintainability review, because bug detection alone does not protect the long-term architecture of large changes. | Evidence: The speaker’s skill slices a change set into domains such as front end and back end, then delegates to a correctness reviewer for bugs and a maintainability reviewer for reuse, existing patterns, and designs that could fail at scale. | Implication: Ken should encode architectural standards as a distinct review lane, not bury them in generic code review; domain decomposition also makes reviewer prompts and ownership more precise. | Caveat: The transcript does not specify reviewer prompts, severity thresholds, false-positive handling, or who adjudicates conflicts among reviewers.
  • Claim: Using multiple model families and a top-level synthesis agent creates a practical ensemble review loop rather than relying on a single model’s judgment. | Evidence: For every domain slice, the workflow runs both GPT and Opus; a top-level agent collects their findings into an action report, and the implementation agent iterates until the findings are resolved. | Implication: Build review orchestration around structured findings, deduplication, severity ranking, and termination criteria, rather than simply increasing the number of agents or model calls. | Caveat: Multiple-model review may increase cost, latency, and correlated false positives; “until things are good” requires an explicit stopping rule in production use.
  • Claim: Agents should maintain a living work journal while executing, making the task system both the specification source and the operational record. | Evidence: A review-swarm skill instructs the agent to take notes, and the speaker says Notion is updated with what the agent found, issues it encountered, and progress during feature construction. | Implication: Require agents to emit durable intermediate artifacts—decisions, discovered dependencies, failed attempts, and validation results—so humans and subsequent agents can audit or resume work without reconstructing context.

Detailed Brief

Workflow design: from exploratory request to PR

  • Claims: The demonstrated task began as an investigation of how a newly built model picker worked and how to migrate it into an alpha version of the product.; The agent can use shell and computer tools for work that exceeds its native tool set, including exploring code and generating files.; The workflow is intentionally asynchronous: the engineer can close the tab after dispatching a task and return to completed artifacts later.; The speaker’s personal task-management practice is minimal: a canonical company task board exists, but he tracks active work largely through pinned chats and focuses on one task at a time.
  • Evidence: The prior workflow used Codex connected to a repository, producing a Markdown file or using a Notion MCP integration to create a task; the native Notion workflow removes those app transitions.; The speaker describes Vercel sandboxes as useful but bare-bones compared with the internal prebuilt coding VM environment.; The agent’s output is positioned as a prebuilt task ready for implementation, rather than a vague summary of its codebase exploration.
  • Caveats: The presentation is a product-adjacent internal demo and does not show failed runs, task rejection, PR review by humans, deployment, rollback, or security controls.; The transcript repeats its final review-swarm segment, so it contains less distinct material than its stated word count suggests.
  • Implications: A robust agent workflow needs explicit state transitions across planning, execution, review, CI repair, and handoff; a chat interface alone is insufficient.; The canonical workspace can function as an operational memory layer when agent outputs are captured as tasks and journals rather than isolated conversations.

Custom skills as organizational policy

  • Claims: The speaker frames agent skills as highly malleable organizational tools rather than fixed vendor features.; His review swarm was created by asking Codex to analyze existing skills, preserve what he liked, reject what he did not, and build a new workflow around his own priorities.
  • Evidence: The stated priorities include maintainable feature and infrastructure changes, reuse of existing code patterns, and prevention of designs that may break under scale.; The speaker calls the resulting workflow the best review-loop flow he has personally experienced.
  • Caveats: A generated skill should be tested against representative historical changes before it becomes a required quality gate; the transcript offers no evaluation methodology.
  • Implications: Ken can treat skills/prompts as versioned operating policy: derive them from existing standards, test them against known edge cases, and evolve them where they produce poor findings or miss important risks.

Notable Concepts & Terms

  • Review swarm skill: A custom orchestration workflow that decomposes a code change, assigns specialized reviewers, aggregates their findings, and drives implementation iterations.
  • Correctness reviewer: A reviewer lane focused on bugs and functional defects in a defined slice of the change set.
  • Maintainability reviewer: A reviewer lane focused on reuse, adherence to established patterns, infrastructure quality, and potential scalability failures.
  • Boxy: Notion’s internal nickname for remote, on-demand coding VMs that provide a prepared development environment to the agent.
  • Canonical source of truth: Notion is positioned as the unified location for specifications, tasks, company knowledge, execution notes, and resulting work records.
  • Living work journal: Agent-maintained notes recording discoveries, problems, and progress during implementation, intended to preserve operational context.
  • Multi-model review: Running GPT and Opus independently over each domain slice before a higher-level agent synthesizes actionable findings.
  • Notion MCP: A prior integration path the speaker used to have Codex create Notion tasks, contrasted with a more native in-product workflow.

Operator Notes / Why Ken Should Care

  • Prototype a code-change control loop with explicit artifacts: intake transcript, normalized task contract, codebase findings, implementation journal, reviewer findings, CI status, and PR handoff.
  • Define a review matrix for each change domain: correctness, maintainability, security/auth where relevant, test adequacy, and operational impact; do not use a single generic “reviewer” prompt.
  • Add a finding schema and stop policy before deploying a swarm: severity, evidence location, recommended remediation, deduplication across models, owner, resolution evidence, and escalation conditions.
  • Require remote agent workspaces to use scoped credentials, isolated environments, reproducible setup, and logs of tool actions before granting repository-write or PR-opening privileges.
  • Evaluate any custom review skill on a benchmark of prior incidents and accepted PRs to measure missed defects, false-positive burden, model disagreement, latency, and cost.
  • Keep a human approval gate for scope expansion, sensitive files, production-impacting changes, and final merge until empirical reliability and governance are established.

Source/Metadata

  • Title: How to Build a Multi-Agent Review Swarm
  • Transcript words: 2407
  • Duration seconds: 721
  • Timestamp note: No timestamps or chapter markers were provided. The transcript repeats the final discussion of context consolidation and the review-swarm design.
Full transcript 1974 words · 11 min read
0:06

All right. Hi, and welcome. Thanks, Dan. Excited to be here. So you are a software engineer at Notion, working on Notion AI. Notion is one of our favorite companies, and we've built our entire company on Notion. And I think you're one of the most slept-on companies in the AI age, because you guys get it and get how valuable having a company brain that is connected to all your agencies. And just excited to have you on to show us a little bit about how you do your work, especially your coding work in Notion. Yeah. We're not a frontier lab or anything, but we're extremely AI-built, nonstop. AI for everything. I can tell. I can feel it from the product decisions.

0:51

So maybe jump into your demo and show us a little bit about how you work. Yeah. Yeah. So right now we're looking at a bunch of me yapping. Actually, I'll even open up a new thread here so you can get an idea of how this works. I start a new thread. For me, I start all of my tasks by just yapping. I'm a big believer in if I'm speaking, there's all these little bits of nuance and detail that I'm going to provide versus me just typing on the keyboard. If I'm just talking, it's completely free-form. And I'm including way more context and details.

1:16

And the models don't really care about how many times I say, or if I say, do this—oh, wait, no, actually, never mind, do that instead. The models don't care, and in fact, I think it might even help. But this demo that we're looking at here, this is actually something real I was working on today. We're working on the redesign of our model picker, and this version, this alpha version of our new AI, doesn't have the new model picker. And so I was like, cool, the guys have built that. They did a really great job. I want to bring that over, do some code sharing and whatever. And so I'm just like, blah, blah, blah, blah, blah, blah, blah.

1:48

And we shipped recently in Notion AI the ability to use Notion AI and add a computer. Right now we're using Vercel sandboxes. And so anytime you're asking Notion AI to do something that's too complicated with the tools it's already got, it would be like, cool, I'll just open up a computer. I'll start using bash or whatever and make PowerPoints, HTML files, do whatever. And that's great, but you're starting with a really bare-bones sandbox anytime you're doing that. And with what we've been building, we have this other internal project. We nicknamed it Boxy, which is our coding VMs, or remote compute on demand.

2:09

And so instead of running everything locally, I've got these other machines out in the cloud that I can connect to and write code on. We were like, wait a minute. What if we gave Notion AI, instead of a simple sandbox, our prebuilt Boxy VMs? And so now we have, in this little picker, Boxy as our computer. And so when I ask it to go and figure out how this new model picker that we're building works, it can actually go out and run all these computer commands and go and explore our code base. And at the end of it, in this instance, I'm actually asking it to create a task.

2:34

I'm like, cool, add it to our M2 milestone project, make the task, and include all the context and requirements for doing this migration, and how do you verify it yourself? And so it ends up producing this pretty detailed task with all the code pointers, requirements, verification steps. And this is prebuilt, ready to run. What I'd been doing before we built this was I'd actually use Codex. I still love using Codex. I would connect it to my repo, either with Boxy or locally, and then have it create a Markdown file, or I'd use the Notion MCP to create a task like this. But I'm hopping between apps and stuff.

2:57

And all of a sudden, I'm just talking into Notion, and with all these native tools to create pages and stuff, it just writes the page, no app hopping or anything. Plus, it runs in the cloud. So I send it a message, and I just close the tab, and I go to a meeting or something. I come back, and I'm like, cool, there's my task. And are you, is this all in a board somewhere? Yeah, we have our company canonical task board up here. I'm pretty averse to project management. I can't look at Kanban boards or anything. I'm just one task at a time. I've gotten into the habit of pinning my tabs or my chats as I go.

3:44

And so that's almost my de facto project management, like what tasks do I still have pinned? That's what I gotta be thinking about. That's really interesting. And does this also kick off reviewers or anything like that? Or how are you tracking it through the full life cycle? Well, yeah. So this is just the generation of the work that needs to be done. What I end up doing next is I spin up a new chat thread where I'm like, work on this. And actually, you can see up here, this is my actual workflow. So I will link the task. I'm like, meet all the requirements.

4:43

I have this review swarm skill because before I even look at the code, I want this thing to be hammering all of the generated code it did itself with a little bit in-depth review skill that I've built. And then once it's done iterating on all that, I want it to put up a PR. Then I want it to monitor all of our continuous integration. If there's a type error, test fails, that's fine, but just go fix it. I don't want to look at that. And this thing, I had it start at 12:53, and this is all from within Notion. Yeah. At the end of it, done at 2:27, PR is up, things pass. Earlier I was just clicking around, and it's perfect. Basically one shot at it.

5:12

And I don't want to glorify the one shot as in I just said, build a thing, make no mistakes. The part that's awesome for me is I'm spending all the time planning, and I'm thinking about what I want to achieve. What's the environment? What are the changes? What are the constraints? And then I'm giving it all the instructions and the tools it needs to be able to build this feature. And it's completely eliminated the busywork. That was literally an hour and a half, two hours of work that this did. I was in meetings. I kicked off other tasks. I think, to me, one of the killer reasons I love doing this within Notion now is that it will find an example.

5:49

We built this edit references feature a while ago, and here in part of my skill, I tell it, take notes. And so it's able to use Notion as the source of truth, not only for the original spec and the tasks, the requirements, but then it also turns into a living work journal. As it's building the feature, it's going and updating all the things it finds, problems it runs into, whatever. That's really interesting. I'm curious because we use Notion for a lot of planning and company information stuff, so it's a new thing for me to process. Oh, maybe I could do coding in here.

6:13

What are the non-obvious second-order effects of having your coding agent live inside of Notion that I might not realize? One of the things I'm thinking about is I have Codex search our meeting notes all the time, which are taken in Notion. Right. So I assume there's stuff like that that happens. What's your experience so far? Yeah, for me, I didn't think it'd be as big of a deal personally as it's been to just have less context for me to manage. Is my local repo up to date? Is Codex updated? Do I have the latest? Is it connected to the right environment or whatever? I think simply being able to eliminate more of those steps is really exciting.

7:03

But I think what's been really great for me is just the consolidation of essentially treating Notion as the source of truth for all context. But eliminating the hopping around has been pretty magical for me. That's awesome. Super cool. Any other last piece of the process you want to show us? I found we're making gigantic changes, especially with this feature that we're building, and I needed something that's not only going to look for bugs, which obviously I need caught, but I need something that's looking out for, am I writing maintainable features and changes and infrastructure? So I built this review swarm skill to essentially take a change set, slice it into domains.

7:32

So maybe you have a front-end change and a back-end change. And then within those changes, delegate out to a correctness reviewer and then a maintainability reviewer. And then each of those is going to look for bugs, and the other one's going to look for patterns. Is there code that we should be reusing? Are there existing patterns that we should be doing? Or did I build something in a way that's going to just completely implode once it goes to scale? And then for each of those slices, I also run all of this for both GPT and Opus.

7:53

And so they're each doing a pass, and then the top-level agent collects all of the findings and turns that into a report of things that need to be actioned. And then essentially I just have the agent loop on that until things are good. And it's, for me personally, been the best review loop flow I've ever experienced. And I think that's another really cool thing about the age that we're in, is how malleable this stuff is. I was like, cool, all of these other skills are really great, but I have my own things that I care about.

8:16

And instead of just being like, well, nobody's built it, I don't have the time, whatever, I just opened Codex and I was like, analyze these other skills, but here's the things I don't like about them. Here's the things I do like. Build me a new skill. And it just did it. And that's turned into my own thing. That's awesome. Well, I hope you open-source that skill so that we can all get the benefit of it. Ryan, this is amazing. Thank you so much for joining. Like one of the things I'm thinking about is I have my, I have codecs search our meeting notes all the time, which are taken in notion. Right. Um, so I assume there's like stuff like that, that happens.

8:56

What's your experience so far? Yeah, for me, I didn't think it'd be as big of a deal personally as it's been to, just have less. Like context for me to manage. Um, is my like local repo up to date is like codecs updated. Um, do I have like the latest, is it connected to the right environment or whatever? Um, I think just simply being able to eliminate more of those steps is like really exciting. But I think what's been really great for me is just the consolidation of. You know, essentially treating notion as like the source of truth for like all context. Um, but like eliminating the like hopping around has been pretty, pretty magical for me. That's awesome. Super cool.

9:47

Uh, any, any other last piece of the process you want to show us? Um, I found we're making like gigantic changes, uh, especially with this feature that we're building and I needed something that's like not only going to like look for bugs, which like obviously I need caught, but like, I need something that's looking out for like, am I writing maintainable features and changes and infrastructure? So I built this review swarm skill to essentially take a change set, slice it into domains. So maybe you have like a front end change and a back end change. And then within those changes delegate out to a correctness reviewer and then a maintainability reviewer.

10:31

And then each of those is going to like look for bugs and the other one's going to look for like, what are like patterns that we, uh, is there code that we should be reusing? Are there existing patterns that we should be doing or like, did I build something in a way that's going to just completely implode once it goes to scale? Um, and then for each of those slices, I also run all of this for both, um, GPT and Opus. And so they're each doing a pass and they, then the top level agent collects all of the findings and turns that into like a report of like things that need to be actioned. Um, and then essentially I just have the agent loop on that until things are, are good.

11:07

Um, and it's for me personally, uh, it's been like the best review loop flow I've ever experienced. Um, and I think that's another really cool thing. And just about the age that we're in is like how malleable this stuff is. I was like, cool. All of these other skills are like really great, but like I have my own things that I care about. And instead of just being like, well, nobody's built it. I don't have the time, whatever. I just like opened codex and I was like analyze these other skills, but here's the things I don't like about them. Here's the things I do like build me a new skill and it just did it. Uh, and that's turned into my own thing. That's awesome.

11:48

Uh, well, I hope you open source that skill so that we can all get a benefit of it. Um, Ryan, this is amazing. Thank you so much for joining.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note