AI Engineer

Missions: Multi-Agent Systems That Ship for Days — Luke Alvoeiro, Factory

830 summary words 4 min summary Watch video

Start with the signal

4 min read

Summary

Missions: Multi-Agent Systems That Ship for Days — Summary

Main Topics

  • The Human Attention Bottleneck: Modern software engineering's limitation is not intelligence but human bandwidth to supervise task execution
  • Multi-Agent Framework Taxonomy: Five frontier approaches to multi-agent communication and coordination
  • The Missions System: A production-ready architecture for long-running autonomous agent workflows (up to 16+ days)
  • Model-Agnostic Architecture: Strategic use of different LLMs for different roles
  • Validation-First Design: Behavior-driven verification to prevent drift in long-running tasks

Key Points

The Core Problem

  • Current bottleneck: Human attention, not AI intelligence
  • Engineers can only drive a few tasks forward per day despite having backlogs of 50+ features
  • Modern models are capable enough to handle complex work, but lack supervision bandwidth

Five Multi-Agent Communication Patterns

  • Delegation: One agent spawns sub-agents (e.g., parent asks child to determine database schema)
  • Creator-Verifier: Separate agents for implementation and validation to avoid sunk cost bias
  • Direct Communication: Agents communicate without central coordinator (challenging without single source of truth)
  • Negotiation: Agents coordinate over shared resources with potential win-win outcomes
  • Broadcast: One agent sends status updates and constraints to many (critical for coherence)

Missions Architecture

Three-Role System:

  • Orchestrator: Handles planning through conversation, scopes requirements, produces validation contracts
  • Workers: Implement features with clean context, commit via git for clean handoffs
  • Validators: Verify work through both traditional (tests, linting, code review) and behavioral (user testing) validation

Validation Contract Innovation:

  • Written during planning before any code is written
  • Defines correctness independently of implementation
  • Can contain hundreds of assertions per complex project
  • Ensures validators aren't biased by implementation details

Execution Strategy

  • Serial execution of features (not parallel) to minimize conflicts and token waste
  • Targeted parallelization on read-only operations (code search, API research, code review)
  • Results in dramatically lower error rates and better correctness for multi-day tasks
  • Longer wall-clock time spent in real-world execution/testing than token generation

Structured Handoffs

Workers document:

  • What was completed
  • What was left undone
  • Commands run and exit codes
  • Issues discovered
  • Adherence to orchestrator procedures

This prevents context loss and enables self-healing through milestone corrections.

Model-Agnostic Design ("Droid Whispering")

  • Different models excel at different roles: planning (slow reasoning), implementation (fast code fluency), validation (precise instruction following)
  • No single model provider is best at all three
  • Architecture improves with each model release (most logic in 700 lines of prompts/skills, not hard-coded state machines)
  • Works with open-weight models thanks to structural advantages
  • Validation can use different provider to avoid training data bias

Production Results (Slack Clone Example)

  • 60% time and tokens spent on implementation
  • Validation rarely succeeds on first attempt (demonstrating QA value)
  • 50% of final code is tests
  • 90% code coverage achieved
  • Heavy use of prompt caching to offset computational costs

Notable Quotes

> "The bottleneck in software engineering nowadays is not intelligence. It's limited by human attention."

> "What if a human decides what to build, and then a system figures out how to do so? An agent could work for hours, for days, and you come back to finished work."

> "Tests written after implementation don't catch bugs. They confirm decisions."

> "You're only as strong as your weakest link. And if you're locked into one model provider, then you're constrained by that family's weakest capability."

> "Missions sort of ensure the discipline, and the models provide the intelligence."

> "The data from production missions is clear. This works on real projects at scale today."

Takeaways

For Practitioners

  • Use validation contracts to define correctness before implementation starts
  • Separate implementation and validation with different agent instances to ensure fresh perspectives
  • Choose the right model for each role rather than forcing one model everywhere
  • Embrace serial execution of features over parallel to reduce errors in long-running tasks
  • Leverage structured handoffs to prevent context loss across agent transitions
  • Build for model improvement: Design architectures flexible enough to benefit from better models

For Team Economics

  • Teams can increase concurrent work streams from ~10 to ~30 with missions
  • Engineers focus on architecture and product decisions, not execution details
  • Code bases become cleaner with comprehensive tests and better structure
  • Both agents and humans are more productive in well-structured mission environments

Operational Insights

  • Longest production mission ran 16 days; systems designed to support 30+ days
  • Mission Control UI enables asynchronous monitoring without constant oversight
  • 16-day missions demonstrate viability for major features and refactors
  • Use cases include overnight prototyping, internal tools, large migrations, ML research, and codebase modernization

Strategic Vision

  • Multi-agent systems representing agent ecosystems will drive next-generation innovation
  • Successful operators will develop intuition for how different models compose under pressure
  • Future opportunities: further parallelization, orchestrating missions into complex workflows, and handling even longer-running projects
Full transcript 2906 words · 20 min read
0:00

[SPEAKER_00] Hi everyone, my name is Luke.

0:14

SPEAKER_00

My goal is that 20 minutes from now, you'll be able to assemble agent teams that can complete tasks orders of magnitude harder than what you can complete with a single agent today. A little about me. I come from a background in DevTools. About two and a half years ago, I started a project at Block, which is where I was working at the time, and that project evolved into Goose. Goose is now one of the leading coding agents that is open source, and it recently was donated to the Agentec AI Foundation. So it's been really cool to see.

0:28

SPEAKER_00

Nowadays I work at Factory, where I lead our core agent, Harness, and Factory's mission is to bring autonomy to the entire software development life cycle. I want to start off with a claim. The bottleneck in software engineering nowadays is not intelligence. It's limited by human attention. Even the best engineers can only complete a couple of tasks at a time. They may have a backlog of 50 features, but they can only drive a few forward per day, because every task requires their attention, every commit needs their review. Today's models are smart enough to figure out all 50 of these tasks, but there's not enough bandwidth to supervise their implementation.

0:45

SPEAKER_00

So we kept asking ourselves, what if a human decides what to build, and then a system figures out how to do so? An agent could work for hours, for days, and you come back to finished work. That's what I'm here to talk about. When you start researching multiagent frameworks and systems, you quickly realize that the field is a bit of a mess. Everyone has their own framework, their own terminology, their own opinions of what works and doesn't work. So I want to propose a simple taxonomy. There are five frontier multiagent frameworks. One is delegation.

1:07

SPEAKER_00

This is where one agent spawns another agent, and the parent agent may say go figure out the database schema and then gets a response back. This is the simplest form of multiagent communication and it's what most people implement first. You have sub-agents and coding tools are the most common example. The other one is creator verifier, where one agent builds something and then you have another agent that checks that work. The key is a separation of concerns. The agent that implemented the code has sunk cost bias, wants that code to work. A fresh agent with fresh context is way more likely to find issues, and this is why we do code review as humans as well.

1:24

SPEAKER_00

Another one is direct communication. This is when agents communicate without a central coordinator. It's like DMing each other. It's hard to get right though because state fragments across conversations without that coordinator and there's no single source of truth. The next one is negotiation. Negotiation is when agents communicate over a shared resource. So that might be they want to use the same API, they want to modify the same portion of the code base. But negotiation doesn't need to be adversarial. In fact, the best use case is when there's net positive sum trading, and that's when agents have a potential win-win situation while interacting.

1:44

SPEAKER_00

And then the last one is broadcast. That is when one agent sends information to many. Think of it as status updates, new context that applies to everyone, new shared constraints. It's a bit less flashy than the other ones but it's critical for maintaining coherence over long running tasks. So when you have all of these different building blocks, how do you assemble that into a system that can run for many days? Missions is our answer. It's a system that combines four of those: delegation, creator verifier, broadcast and negotiation into a single workflow.

1:59

SPEAKER_00

You describe a goal, you scope that through a conversation, you approve a plan, and then the system handles execution for hours or days. That enables you to focus on something else. Notably, a mission is not a single agent session. It's an ecosystem of agents that communicate through structured handoffs and shared state. It uses a three-role architecture: orchestrator, workers, and validators. The orchestrator handles planning. When you describe what you want, the orchestrator is like your sounding board. It asks you the right strategic questions. It checks if there are any unclear requirements in the problem space.

2:33

SPEAKER_00

Then it eventually produces a plan that includes features, milestones, and something called a validation contract. That validation contract defines what done means before any coding is done. I'll come back to why that matters because it turns out to be really important to the system. The next role is workers. They handle implementation. When a feature is assigned to a worker, that worker has clean context, no accumulated baggage, no degraded attention. The worker reads its spec, implements the feature, and then commits via git, allowing the next worker to inherit a clean slate and a working code base. And then the last role is validators. They handle verification.

3:25

SPEAKER_00

Most systems validate by running lint, type check, tests, maybe they do code review. Missions does all of that, but we also validate behavior. Instead of just asking, does the code look right, we ask, does this work end to end? That's the difference that lets missions run for many hours, many days in a row without drifting. Making it work had to involve rethinking validation entirely. When you've worked with coding agents before, you've probably seen this pattern where an agent builds a feature, it writes some tests, the tests pass, there's full coverage. But the tests were shaped by the code, not by what the code was attempting to actually do.

4:07

SPEAKER_00

Tests written after implementation don't catch bugs. They confirm decisions. So if you rely on validation like that, your system will eventually drift. That's why this validation contract exists. It's written during planning, before any code, and it defines correctness independently of implementation. For a complex project, this can be hundreds of assertions. Each feature is assigned one or more assertions that it must satisfy. The sum of all features must mean that every assertion is covered. After each milestone of features, we have two types of validators that run: the scrutiny validator and the user testing validator. The first one is more traditional.

5:13

SPEAKER_00

It runs the test suite, type checking, lints, and critically, it spawns dedicated code review agents for each completed feature within the milestone. And then the second one, which is the user testing validator, is more interesting. It acts like a

5:33

SPEAKER_00

during planning, before any code, and it defines correctness independently of implementation. So for a complex project, this can be hundreds of assertions. And each feature is assigned one or more assertions that it must satisfy. The sum of all features must mean that every assertion is covered. After each milestone of features, we have two types of validators that run. So you have the scrutiny validator and the user testing validator. The first one is more traditional. It runs the test suite, type checking, lints, and critically, it spawns dedicated code review agents for each completed feature within the milestone. And then the second one, which is the user testing validator, is more interesting. It acts like a QA engineer. It spawns the application. It interacts with it through computer use or something similar to that. It fills out forms, checks that pages render correctly, clicks buttons, and ensures that functional flows work holistically. So this step takes significantly longer than the previous one of the scrutiny validator. Because the system is interacting with a live application. And what we've noticed is that most of the mission's wall clock time is actually spent here, waiting for this real-world execution to occur instead of generating tokens. Critically, neither validator has seen the code before. They are not invested in the implementation, and so validation is adversarial by design. Okay, so then validation catches bugs, right? But for a system that runs for many days, you also need to make sure that context isn't lost between the agents. When a worker finishes a feature, it doesn't just say, I'm done. It fills out a structured handoff detailing what was completed, what was left undone, what commands were run throughout that agent loop, and what were the exit codes of those commands. What issues were discovered? And did it abide by the procedures that the orchestrator defined for that worker? That's how we catch issues, and how the system self-heals. The errors get caught at milestone boundaries, corrective work gets scoped, and the mission pulls itself back on track. Not by hoping that agents remember what happened, but by forcing them to write it down, and then actually address issues, and I'll present on that in just a sec. Our longest mission ran for 16 days, which is much longer than a full sprint, and we believe that they can run for 30. That's only possible because of this structure. So once we had this architecture, the next question became, how do we actually run it, right? The most obvious choice is parallelism. If you have 10 agents running at one point in time, then you have 10 times the throughput. But we tried that, and it doesn't really work for tasks in the software dev domain, because agents have a lot of different things conflict. They step on each other's changes, they duplicate work, they make inconsistent architectural decisions. And so the coordination overhead ends up eating up the speed gains, all the while you're burning tokens. The difference with missions is that we run features serially. So there's only one worker or validator running at any given point in time. Within a feature, we allow for parallelization on read-only operations. So you have something like searching through the code base or researching APIs, all that gets parallelized. Within validators, we also parallelize read-only operations such as code review. This is serial execution with targeted internal parallelization. It seems slower on paper, but the error rate drops dramatically, and when you have tasks that run for many days, this correctness compounds. Now, your standard chat interface doesn't really work for something that lasts many days. At a quick glance, you need to be able to see how much of the project you have completed, and what amount of the budget that you originally set off with you have burned through. So using a mission, actually, we built Mission Control, which is a dedicated view for this. You can see what is the active worker doing right now. Read off handoff summaries that detail what did the worker, the validator discover, how it's going to alter its course moving forward. Or you could just go check out, go hang out with your friends that night. This entire view lets you just run missions asynchronously, and you could be plugged in as a project manager overseeing the implementation, or you could just go and hang out with your friends. Okay, so the right model in each role. Everything here sort of assumes one thing, and that is that you're using the right model in each role. Planning benefits from slow, careful reasoning, implementation from fast code fluency and creativity, validation benefits from precise instruction following. And so no single model nor model provider is best at all three of these. Using systems like missions requires the development of a new skill, which internally we've been calling droid whispering. But it's this idea that you need to be able to mentally model how different LLMs interact, where they fail, how those failures compound over a multi-day run, and then you need to make a deliberate choice as to which model sits in which seat. Theo, the engineer who built our missions prototype, came up with our model defaults. But we really encourage people to make these their own and customize them to the needs of their project. So for example, validation might use a different model provider entirely to make sure that it's not biased by the same training data. There's a structural advantage of a model agnostic architecture. You're only as strong as your weakest link. And if you're locked into one model provider, then you're constrained by that family's weakest capability. As models continue to specialize, the ability to put the right model in the right seat becomes a compounding advantage. And works in the other direction too. If you're using missions, the structure of that can compensate for models that are not quite at the frontier level performance. So the validation contracts, the milestone checkpoints, they allow you to run missions very, very successfully even using open weight models. Now, this all sounds quite theoretical. What does it actually look like in production? I want an example of building a clone of Slack right here. This slide has a ton of info, but I'll walk you through just a few things I want to call out. 60% of our time is spent on implementation. And 60% of our tokens as well. Notice how validation never succeeds on the first go. That's in the mission, what's it? The one on the bottom left. We almost always have to create follow-up features. So it really demonstrates the value of a system that does this QA loop. You end up with 50% of your lines of code at the very end, in the bottom right, being tests. And 90% of your code is covered by those tests. And lastly, we take advantage of prompt caching heavily to make sure that we're sort of offsetting the price of running such a long task. People are really taken to missions, and it's been awesome to see what folks have been building with them. Some examples I've included in this slide, but ones that I want to call out are specifically in the enterprise setting, which is where factory really shines. They've been used to prototype new ideas and features overnight, to make sure that people can build internal tools at increasingly rapid rates, to run huge refactors and migrations, for ML research, and to modernize code bases so that agents are more productive in them. One thing that I wanted to talk about was also this concept of the bitter lesson.

5:39

SPEAKER_00

And lastly, we take advantage of prompt caching heavily to make sure that we're offsetting the price of running such a long task. People are really taken to missions, and it's been awesome to see what folks have been building with them. Some examples I've included in this slide, but ones that I want to call out are specifically in the enterprise setting, which is where factory really shines. They've been used to prototype new ideas and features overnight, to make sure that people can build internal tools at increasingly rapid rates, to run huge refactors and migrations, for ML research, and to modernize code bases so that agents are more productive in them.

6:21

SPEAKER_00

One thing that I wanted to talk about was also this concept of the bitter lesson. Because every person building multi-agent systems has this fear of the next model release making their architecture obsolete overnight. So when we were building missions, we decided we had to make this system get better with every model improvement. This means that almost all of the orchestration logic is defined in prompts and skills, instead of a hard-coded state machine. How it decomposes failures or decomposes features and handles failures is all in about 700 lines of text, and four sentences of this can alter the execution strategy pretty dramatically.

6:59

SPEAKER_00

Worker behavior is driven by skills that the orchestrator defines per mission, so you get very customized behavior. And the only deterministic logic is very thin. It's focused on enabling models to do what they do best, while the system handles the bookkeeping, right? Stuff like running validation and ensuring that progress is blocked when there are some handoff issues that are not addressed. So missions sort of ensure the discipline, and the models provide the intelligence using primitives that they're already familiar with, like agents MD, skills, etc. So what does this unlock? Remember the bottleneck that I started off with? Human attention.

7:36

SPEAKER_00

The economics are changing. Before, a team of five engineers might be able to work on 10 work streams at any given point in time. Now, maybe with missions, we can bring that up to 30. The team can focus on interesting problems, such as the architecture, product decisions, instead of worrying about the execution, per se. And the important thing is the code base ends up cleaner than when you started. The end-to-end tests, the unit tests, the skills, the structure that missions provide means that agents and humans are more productive in that environment moving forward.

8:03

SPEAKER_00

So now that you understand how missions are structured and how they actually work, you can see that they're really a composition of those original strategies, right? Delegation shows up everywhere in how the orchestrator spawns workers and how we spawn research sub-agents. Creator verifier is fundamental in that validation and implementation are always separate agents with separate context. Broadcast runs through the shared mission state that every agent references. And negotiation shows up at milestone boundaries, where the orchestrator defines, does this handoff summary look correct? Do we need to create follow-up features, re-scope, etc.

8:45

SPEAKER_00

But strategies aren't enough. You need the connective tissue. You need these structured handoffs so that agents don't lose context. You need the right model in each role. And you need an architecture that will improve with each model improvement. So what I like to think about is that people in this room who are thinking in terms of agent ecosystems, who develop an intuition for how different models compose under pressure, that those folks are going to be really shipping the next generation of innovation. There's a lot of open questions still, right? How do we further parallelize the workload of missions so that they run faster?

9:21

SPEAKER_00

How do we start orchestrating missions themselves into even more complex workflows? How do we start to do those? But the data from production missions is clear. This works on real projects at scale today. So this is what I'll leave you with. Open Droid. Try running slash missions. Argue with the orchestrator about the scope. Approve the plan. And then go do something else. I'm excited to see what you guys build. And I'll be around to answer any questions for the rest of the day. Thanks. parallelize read-only operations such as code review. This is serial execution with targeted internal

10:18

SPEAKER_00

parallelization. It seems slower on paper, but the error rate drops dramatically, and when you have tasks that run for many days, this sort of correctness compounds. Now, your standard chat interface doesn't really work for something that lasts many days. At a quick lens, you need to be able to see how much of the project have you completed, and what amount of the budget that you originally set off with have you burned through. So using a mission, actually, we built Mission Control, which is a dedicated view for this. You can see what is the active worker doing right now. Read off handoff summaries that detail what did the worker, the validator discover, how it's going

10:59

SPEAKER_00

to sort of alter its course moving forward. Or you could just go check out, go hang out with your friends that night. This entire view lets you just run missions asynchronously, and you could be plugged in as a project manager overseeing the implementation, or you could just go and hang out with your friends. Okay, so the right model in each role. Everything here sort of assumes one thing, and that is that you're using the right model in each role. Planning benefits from slow, careful reasoning, implementation from fast code fluency and creativity, validation benefits from precise instruction following. And so no single model nor model

11:44

SPEAKER_00

provider is best at all three of these. Using systems like missions requires the development of a new skill, which internally we've been calling droid whispering. But it's this idea that you need to be able to mentally model how different LLMs interact, where they fail, how those failures compound over a multi-day run, and then you need to make a deliberate choice as to which model sits in which seat. Theo, the engineer who built our missions prototype, came up with our model defaults. But we really encourage people to make these their own. And customize them to the needs of their project. So for example, validation might use a different model

12:19

SPEAKER_00

provider entirely to make sure that it's not biased by the same training data. There's a structural advantage of a model agnostic architecture. You're only as strong as your weakest link. And if you're locked into one model provider, then you're constrained by that family's weakest capability. As models continue to specialize, the ability to put the right model in the right seat becomes a compounding advantage. And works in the other direction too. If you're using missions, the structure of that can compensate for models that are not quite at the frontier level performance. So the validation contracts, the milestone checkpoints,

12:56

SPEAKER_00

they allow you to run missions very, very successfully even using open weight models. Now, this all sounds quite theoretical. What does it actually look like in production? I want an example of building a clone of Slack right here. This slide has a ton of info, but I'll walk you through just a few things I want to call out. 60% of our time is spent on implementation. And 60% of our tokens as well. Notice how validation never succeeds on the first go. That's in the mission, what's it? The one on the bottom left. We almost always have to create follow-up features. So it really demonstrates the value of a system that does this QA loop.

13:38

SPEAKER_00

You end up with 50% of your lines of code at the very end, in the bottom right, being tests. And 90% of your code is covered by those tests. And lastly, we take advantage of prompt caching heavily to make sure that we're sort of offsetting the price of running such a long task. People are really taken to missions, and it's been awesome to see what folks have been building with them. Some examples I've included in this slide, but ones that I want to call out are specifically in the enterprise setting, which is where factory really shines. They've been used to prototype new ideas and features overnight,

14:17

SPEAKER_00

to make sure that people can build internal tools at increasingly rapid rates, to run huge refactors and migrations, for ML research, and to modernize code bases so that agents are more productive in them. One thing that I wanted to talk about was also this concept of, like, the bitter lesson. Because every person building multi-agent systems has this fear of the next model release sort of, like, making their architecture obsolete overnight. So when we were building missions, we decided we had to make this system get better with every model improvement.

14:56

SPEAKER_00

This means that almost all of the orchestration logic is defined in prompts and skills, instead of, like, a hard-coded state machine. How it decomposes failures, or decomposes features, and handles failures, is all in about, like, 700 lines of text, and four sentences of this can alter the execution strategy pretty dramatically. Worker behavior is driven by skills that the orchestrator defines per mission, so you get very customized behavior. And the only deterministic logic is very thin. And it's focused on enabling models to do what they do best, while the system handles, like, the bookkeeping, right? Stuff like running validation and ensuring that progress is blocked

15:35

SPEAKER_00

when there are some handoff issues that are not addressed. So missions sort of ensure the discipline, and the models provide the intelligence using primitives that they're already familiar with, like agents MD, skills, etc. So what does this unlock? Remember the bottleneck that I started off with? Human attention. The economics are sort of changing. Before, a team of five engineers might be able to work on 10 work streams at any given point in time. Now, maybe with missions, we can bring that up to 30. The team can focus on interesting problems, such as the architecture, product decisions, instead of worrying about the execution, per se.

16:16

SPEAKER_00

And the important thing is the code base ends up cleaner than when you started. The end-to-end tests, the unit tests, the skills, the structure that missions provide means that agents and humans are more productive in that environment moving forward. So now that you understand how missions are structured and how they actually work, you can see that they're really a composition of those original strategies, right? Delegation shows up everywhere in how the orchestrator spawns workers and how we spawn research sub-agents. Creator verifier is fundamental in that validation and implementation are always separate agents with separate context.

16:55

SPEAKER_00

Broadcast runs through the shared mission state that every agent references. And negotiation shows up at milestone boundaries, where the orchestrator defines, you know, does this handoff summary sort of like look correct? Do we need to create follow-up features, re-scope, etc. But strategies aren't enough. You need the connective tissue. You need these structured handoffs so that agents don't lose context. You need the right model in each role. And you need an architecture that will improve with each model improvement.

17:25

SPEAKER_00

So what I like to think about is that people in this room who are thinking in terms of agent ecosystems, who develop an intuition for how different models compose under pressure, that those folks are going to be really shipping the next generation of innovation. There's a lot of open questions still, right? How do we further parallelize the workload of missions so that they run faster? How do we start orchestrating missions themselves into even more complex workflows? How do we start to do those? But the data from production missions is clear. This works on real projects at scale today. So this is what I'll leave you with. Open Droid. Try running slash missions.

18:02

SPEAKER_00

Argue with the orchestrator about the scope. Approve the plan. And then go do something else. I'm excited to see what you guys build. And I'll be around to answer any questions for the rest of the day. Thanks.

18:19

SPEAKER_00

Friends. Friends. Friends. Friends. Friends. Friends. Friends. Friends. Friends.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note