What Does Done Even Mean? Agents and Paperclip's Liveness Model - Dotta, Paperclip
Description
What does “done” mean when agents can produce more work than humans can possibly review? This talk argues that the future of agentic work is not just faster output, but a stronger trust protocol: systems where “done” means an artifact has met a stated standard, carries evidence, has been checked by the right verifier, assigns ownership of remaining risk, and clearly authorizes the next action. Drawing from Paperclip’s liveness model, it shows how teams can avoid approval theater, keep work moving, route review by risk, and turn agent completion from a vague confidence signal into something others can safely build on. Speakers: - Dotta (Paperclip): Dotta is the creator of Paperclip, the Open-source app for zero human companies X/Twitter: https://x.com/dotta GitHub: https://github.com/cryppadotta
Summary
Generated by claude-sonnet-4-5At-a-Glance
- Verdict: Watch fully
- Core thesis: Agent systems require a formal liveness model to prevent verification theater and AI slop by treating 'done' as a structured object with explicit claims about artifacts, evidence, verification, and handoffs rather than a boolean checkbox.
- Why it matters: As agents produce code faster than humans can verify, systems need protocols balancing liveness (unblocked progress) with verification (quality control) or they collapse into either infinite review queues or zero-quality output.
- Best use: Steal Paperclip's checklist and invariants when designing multi-agent workflows; treat this as a control-plane design pattern reference for production agent systems.
Executive Summary
Dotta from Paperclip tackles the operational crisis emerging as agents outpace human verification capacity. The central problem: 'done' is currently a flat boolean in most agent systems, but in reality it's a bundle of claims—artifact produced, evidence of completion, rubric for verification, designated reviewer, next owner, and next action. When agents can generate pull requests, documentation, and code changes faster than any human can review them, the old model breaks down into either verification theater (humans rubber-stamping queues they can't actually process) or AI slop (unverified agent output shipped at volume).
Paperclip's solution is a liveness model built on three invariants: ensure productive work continues, only real blockers stop work, and infinite loops are bounded. This manifests as explicit state transitions between tasks, first-class blockers enforced by a control plane, interactive approval moments with audit trails, designated reviewers/approvers per task, and 'watchdog' agents that maximize progress toward goals across heterogeneous harnesses (OpenClaw, Hermes, Codex, etc.). The key insight is treating liveness and verification as opposing forces that must be balanced, not collapsed into a single 'tests passed' gate.
The actionable output is a checklist for production agent work: define exactly what done means for each task; separate verifier from author (use different models); demand evidence not assertions (give agents screenshot tools, browser harnesses, custom hooks to verify their own work); establish clear chain of custody so every agent knows the next handoff. This isn't Paperclip-specific—it's a design pattern for any serious multi-agent system where humans are accountable for outcomes but can't manually verify every step.
Key Takeaways
- Claim: Agent systems create a new failure mode: they generate more work than humans have time to verify, leading to either verification theater or uncontrolled AI slop. | Evidence: Dotta contrasts exhaustive human verification (which fails at high volume and becomes form-filling ritual) with pure liveness (which produces low-quality output at scale). Example: an agent opens a PR, passes tests, updates docs, closes the issue—looks done but is it merge-ready, deploy-ready, customer-announcement-ready? | Caveat: No quantified threshold given for when human verification becomes untenable (e.g., tasks per day, team size); the claim assumes agents already outpace review capacity. | Implication: Ken's agent workflows need explicit liveness/verification protocols from day one or they'll hit a ceiling where output volume exceeds review bandwidth and quality collapses. | Timestamp: 00:25
- Claim: 'Done' is not a boolean; it's a bundle of claims: artifact produced, evidence of completion, rubric, designated reviewer, next owner, next action. | Evidence: Dotta enumerates the components: producer claim → reviewer finding no obvious issues → verified against standard → authorized approval → real-world survival. Example: an agent closing a ticket is not the same as code being merge-ready or deploy-ready. | Caveat: The framework doesn't specify how to encode these claims programmatically or what data structures to use in practice. | Implication: Ken should model task completion as a struct/object with explicit fields for artifact, evidence, rubric, reviewer, approver, risk assessment, and next action rather than a status enum. | Timestamp: 00:50
- Claim: Paperclip's control plane enforces three invariants: ensure productive work continues, only real blockers stop work, and infinite loops are bounded. | Evidence: Implementation details: explicit state transitions between tasks, first-class blockers enforced by the control plane, moments of interactive human approval with audit trails, designated reviewers/approvers on tasks, watchdog agents that maximize goal achievement. | Caveat: No open-source code or API shown; unclear how blockers are formally defined or how infinite-loop bounding works concretely. | Implication: Ken's orchestration layer should enforce dependency DAGs, formal blocker states, and watchdog-style maximizers that retry across failures rather than naive for-loops over task queues. | Timestamp: 02:15
- Claim: Watchdogs are harness-agnostic maximizer agents that enforce goal completion across heterogeneous tools (Pi, OpenClaw, Hermes, Codex). | Evidence: A watchdog is given a goal and ensures all agents continue working until that goal is achieved, providing a consistent interface regardless of which coding/execution harness is in use. | Caveat: No detail on how watchdogs detect goal achievement, handle conflicts between heterogeneous agents, or prevent adversarial agent behavior. | Implication: Ken can layer a meta-agent that retries/delegates across different LLM providers and tools to maximize success probability without re-engineering each harness. | Timestamp: 02:50
- Claim: Best practice: separate the verifier from the author by using different models (e.g., code with Claude, verify with Codex). | Evidence: Dotta's checklist explicitly recommends using a different model for verification to avoid the same model's blind spots appearing in both generation and review. | Caveat: No empirical data provided on error-detection rates when using same vs. different models; assumes model diversity improves verification quality. | Implication: Ken should configure agent pipelines with distinct models for generation and verification roles to catch errors the original model wouldn't surface. | Timestamp: 03:30
- Claim: Agents should provide evidence (screenshots, test runs, browser interactions) rather than assertions of completion. | Evidence: The checklist instructs giving agents custom browser harnesses, screenshot tools, and hooks to verify work themselves rather than just reporting 'done.' | Caveat: No discussion of evidence storage, evidence quality validation, or how to prevent agents from faking evidence. | Implication: Ken's agent tooling should include first-class screenshot/session-recording capabilities and structured evidence schemas, not just status flags. | Timestamp: 03:45
Detailed Brief
The Core Problem: Agent Output Exceeds Human Verification Capacity
- Claims: Programming is solved; agents produce code/docs faster than humans can verify; This creates a failure mode where agents generate more work than review bandwidth; Current systems flatten 'done' to a single green checkmark, collapsing nuanced operational claims
- Evidence: Example: agent opens PR, passes tests, updates docs, closes issue—looks done, but is it merge-ready? deploy-ready? customer-announcement-ready?; Exhaustive human verification fails at high volume, becoming verification theater; Pure liveness (no review) produces AI slop—output worse than nothing
- Caveats: No empirical thresholds for when verification becomes theater (tasks/day, team size); Assumes agents already outpace human review; may not hold for all orgs
- Implications: Ken needs liveness/verification protocols from the start or will hit a quality/volume ceiling; Single-status systems (done/not-done) are inadequate for production agent workflows
The Liveness Model: Balancing Progress and Quality
- Claims: Liveness = work continuing with no blockers; Verification = assurance of correctness; These are opposing forces that must be balanced, not collapsed into one gate; Pure liveness → AI slop; pure review → infinite review queue
- Evidence: Paperclip's control plane enforces explicit state transitions, first-class blockers, interactive approval moments with audit trails; Three invariants: ensure productive work continues, only real blockers stop work, infinite loops are bounded; Mechanisms: designated reviewers/approvers per task, watchdog agents for goal maximization
- Caveats: No open-source implementation shown; Unclear how infinite-loop bounding works concretely; No metrics on verification latency or throughput improvements
- Implications: Ken should model orchestration as a state machine with explicit transitions and blocker enforcement; Watchdog pattern is reusable: meta-agent that retries across tools until goal achieved
The Actionable Checklist for Production Agent Work
- Claims: Define exactly what 'done' means for each task (artifact, scope, rubric, evidence, verifier, approver, risk, next action); Separate verifier from author (use different models); Demand evidence not assertions (give agents tools to verify their own work); Establish clear chain of custody (every agent knows next handoff)
- Evidence: Example: code with Claude, verify with Codex to avoid blind spots; Give agents custom browser harnesses, screenshot tools, custom hooks to run and click through UI; Chain of custody ensures no work gets stuck in limbo
- Caveats: No guidance on evidence validation or preventing fake evidence; No discussion of how to encode 'done' object structure in practice
- Implications: Ken should implement task completion as a struct with explicit fields for all claims; Agent tooling should include screenshot/session-recording, structured evidence schemas; Verification should use a different model family than generation to maximize error detection
Notable Concepts & Terms
- Liveness vs. Verification Trade-off: Liveness = unblocked progress; verification = quality assurance. Pure liveness → AI slop; pure verification → infinite review queue. Control planes must balance both.
- Verification Theater: When human review becomes a rubber-stamp ritual because the volume of agent output exceeds review capacity, creating the appearance of quality control without substance.
- Done as an Object (not a Boolean): Treating task completion as a structured data type with fields for artifact, evidence, rubric, reviewer, approver, risk, next action—rather than a single true/false flag.
- Watchdog Agent: A meta-agent given a goal that ensures all other agents continue working until the goal is achieved, harness-agnostic, maximizes success across heterogeneous tools.
- First-Class Blockers: Dependencies and constraints between tasks enforced by the control plane as explicit primitives, not as implicit side-effects of task status.
- Chain of Custody: Explicit handoff protocol ensuring every agent knows who receives work next, preventing tasks from entering limbo states.
Operator Notes / Why Ken Should Care
- Ken should treat this as a control-plane design pattern library for multi-agent orchestration, not just Paperclip marketing.
- The liveness/verification trade-off is a fundamental tension in any agent system; design orchestration to balance both from day one.
- Separate verifier from author (use different models) is an immediately actionable pattern for improving agent work quality.
- Watchdog agents as harness-agnostic maximizers are a reusable abstraction for retry/fallback logic across heterogeneous tools.
- Modeling 'done' as a struct with explicit claims (artifact, evidence, rubric, reviewer, approver, next action) is a universal improvement over status booleans.
- Demand evidence (screenshots, logs, test runs) rather than agent assertions; design agent tooling to provide this evidence natively.
- For GTM/investing: Paperclip is solving a real scaling problem (agent output exceeding human review), but the checklist is open for anyone to implement.
- For AI ops: the three invariants (ensure productive work continues, only real blockers stop work, infinite loops bounded) are a useful mental model for orchestration correctness.
Watch Map
- 00:00: Intro: agent opens PR, looks done—but is it merge-ready, deploy-ready, announce-ready? Different operational claims.
- 00:25: The new failure mode: agents create more work than humans have time to verify.
- 00:50: Done is a bundle of claims: artifact, evidence, rubric, reviewer, approver, next step.
- 01:15: Different levels of done: producer claim, reviewer check, verified standard, authorized approval, real-world survival.
- 01:40: Verification theater: exhaustive human review fails at high volume, becomes form-filling ritual.
- 02:00: Protocol for task progression: keep tasks moving, prevent invalid states, tie execution to contracts/constraints.
- 02:15: Liveness vs. verification: opposing forces. Pure liveness → AI slop; pure review → infinite queue.
- 02:40: Paperclip's mechanisms: state transitions, first-class blockers, interactive approval with audit trails, designated reviewers/approvers.
- 02:50: Watchdog agents: harness-agnostic maximizers that enforce goal completion across Pi/OpenClaw/Hermes/Codex.
- 03:10: Three invariants: ensure productive work continues, only real blockers stop work, infinite loops bounded.
- 03:30: Treat done as an object: artifact, scope, rubric, evidence, verifier, approver, risk, next action.
- 03:45: Checklist: define done, separate verifier from author (use different models), demand evidence (screenshots/tools), clear chain of custody.
- 04:10: Give agents tools to verify their own work: browser harnesses, screenshots, custom hooks to click and test.
Source/Metadata
- Title: What Does Done Even Mean? Agents and Paperclip's Liveness Model - Dotta, Paperclip
- Transcript words: 1521
- Duration seconds: 433
- Timestamp note: Timestamps inferred from 433-second video duration and transcript flow; chapter markers not present in transcript.
Transcript
An agent opens a pull request. It passes the tests. It updates the documentation. It closes the issue and comments, looks done to me. But is it actually done? Is it done enough to merge? Is it done enough to deploy? Is it done enough to announce to your customers? These are fundamentally different operational claims, and most agent systems just flatten it to a single green checkmark. I'm Dota. I'm the creator of Paperclip, and I'm going to give you some hard-earned lessons that we've learned in creating Paperclip's liveness model. What does done even mean? Here's the thing. Programming is solved, and agents can now produce more code and documentation faster than any human can ever verify. And this actually gives us a new failure mode: agents can actually create more work than humans have time to verify. And so we need a way to verify that our agents are done more than just letting them check a checkbox. Done doesn't mean that an agent just changed the status of a task to being done. Saying that something is done is actually a bundle of claims. You're saying that an artifact was produced. That you have evidence that the task is actually complete. You have a rubric in which you can verify against. You know exactly who the owner is for the next step, and you know exactly what the next step is. There are different levels to how done something is. The producer might claim something is complete, but you need to have a reviewer, another party that looks at it and finds no obvious issues. You want to verify and make sure that the evidence actually meets a specified standard. You want to make sure that a person who is authorized to approve it actually approves that the work is done. And you want to make sure that there's somebody who actually stands behind the decision. And ideally what you want is that the outcome has actually survived real world conditions. Because exhaustive human verification fails at high volume. You might be able to verify a few tasks per day. But essentially, if you have humans verifying all the tasks and they have to sign off on it, eventually what you just get is a form of verification theater. What you need is a protocol for defining how tasks actually progress through a system. You want to make sure that tasks are always kept moving, but they don't get stuck into invalid states. You need a control plane that actually has the execution of the tasks being tied to specific contracts and constraints about what the system will do and what agents it will hand off to your next task. Because really what you're trying to play against is this idea around keeping work moving, but also having it verified. When a task has been reviewed by a human, you get the assurance that it's correct. But having a human verify it means that the task is dead in its tracks. You also want to keep liveness. Liveness means that the work is continuing with no blockers. And you're always trying to keep these two things in balance. If you have tasks that are completely alive with no approvals, then what you get is this classic AI slop because you're producing a lot of things with no quality control. And it's worse than creating nothing after a long period of time. But if you have pure review, then you have this enormous review queue where humans can't actually review it by hand anyway. These agents will be creating far more than you can ever actually review. And so we have to find a way to tease apart the bundle of claims that are involved in saying a task is done. With Paperclip, we have a number of mechanisms to keep this going. You might think that you can easily just write a for loop over your task manager and have your agents work. But quickly you'll find that falls apart. As soon as you start integrating task dependency trees, blockers, multiple agents, idempotent checkouts, locks on checkouts, you find that this tension between liveness and verification actually gets quite complicated. There are really three invariants that are extremely important when you're thinking about what you want out of a control plane for your agentic work. You want to ensure that productive work continues. You want to make sure that only real blockers stop work. And you want to make sure that infinite loops are bounded. In Paperclip, we have built a number of mechanisms to deal with this problem. For example, every time you have a task, there are clear transitions to what the next state could be. We have first-class blockers between tasks, and the control plane enforces those blockers. We have moments of interactive human approval where human choices leave an audit trail. You can set reviewers and approvers on tasks explicitly, meaning when this task completes, another agent can review it. We also have the idea of watchdogs, which is this maximizer mode, which says try as hard as you can to make sure that this happens. When you have a watchdog, it's another agent who is given a goal, and it enforces that all of your agents continue to work until that goal has been achieved. The important thing here is that the watchdog within Paperclip is harness agnostic. You can use it with Pi, OpenClaw, Hermes, Cloud Code, Codex. Whatever you're using, you have one consistent interface for ensuring that goal is complete. One of the best pieces of advice we have is that you stop treating done as a Boolean and treat it more like an object. This isn't specific to Paperclip. It's just advice on how you think about what is done. Humans automatically paper over these details, but when we're building agentic systems, it's important that your agents can distinguish between the different pieces of what they're claiming when they say something is done: the artifact that they're saying is complete, the scope, the rubric or the standard, the evidence that it's done, who verified the work, who has the authority to sign off on the work, and what risk might be left, and what the next action is going to be. So you want to make sure that when you define done, it's not just a checkbox. If you want to get 100 times more work done, you should steal this checklist. You need to define exactly what done means for this task. You definitely want to separate the verifier from the author. Often, this means you're using a different model. So if you're coding using Claude, have Codex verify. You want to ask your agents to provide evidence. Don't just ask them to say "Is this done?" but give them the tools they need to verify that the work is done. Write the code to have the custom browser harness. Write the code to take the screenshots. Make sure they have access to a browser. Make sure that they have custom agent hooks or custom agent tooling to actually run through and click the buttons and try it out themselves and verify that the work is truly done. Make sure you have a clear chain of custody so every agent knows that as soon as they're done, who they're supposed to give the work to next. It can be easy to just fire off a single line instruction and see whatever comes back. But if you have serious work that you're accountable for, it's very important that you define what done really means in as much detail as possible. Thank you. So, you want to make sure that when you define done, it's not just a checkbox. So, if you want to get 100 times more work done, you should steal this checklist. You need to define exactly what does done mean for this task. You definitely want to separate the verifier from the author. Often, this means you're using a different model. So, if you're coding using Claude, have Codex verify. You want to ask your agents to provide evidence. Don't just ask them to say, is this done, but give them the tools they need to verify that the work is done. Write the code to have the custom browser harness. Write the code to take the screenshots. Make sure they have access to a browser. Make sure that they have custom agent hooks or custom agent tooling to actually run through and click the buttons and try it out themselves and verify that the work is truly done. Make sure you have a clear chain of custody that every agent knows that as soon as they're done, who they're supposed to give the work to next. It can be easy to just fire off a single line instruction and vibe with whatever comes back. But if you have serious work that you're accountable for, it's very important that you define what done really means in as much detail as possible. Thank you.