Open Reader

What the Best Agents Share — Mardu Swanepoel, Flinn AI

completed 10:21 May 26, 2026 Watch on YouTube

Current Status

completed

Video ID

7CrPrHgoEYk

RAG / Chat

Enabled
What the Best Agents Share — Mardu Swanepoel, Flinn AI
Description

Harvey, Cursor, Manus, and Claude operate in completely different domains but share four patterns: focus modes that constrain the action space to improve output quality, transparent execution that surfaces tool calls and reasoning to build user trust, personalization that optimizes for speed to understanding rather than just speed to output, and reversibility that bounds the downside of mistakes so users take on higher value tasks. Mardu Swanepoel from Flinn AI breaks down how each company puts these into practice. Cursor lets you roll back changes at the line, file, or conversation level and run multiple model outputs in parallel from the same input. Harvey builds playbooks from a firm's legal methods so the agent works the way the firm would. Claude surfaces a live task list alongside every tool call's inputs and outputs so users can intervene before the agent goes further in the wrong direction. Speaker info: - https://www.linkedin.com/in/mardu-swanepoel-000/

Summary

Generated by claude-sonnet-4-5

30-second take

Mardu Swanepoel from Flinn AI argues that the best AI agents (Cursor, Claude Cowork, Manus, Harvey) share four design patterns that prioritize user control and clarity over raw speed. He explicitly rejects "do anything" agents, claiming they fail because they optimize for speed-to-outcome rather than speed-to-understanding. The four patterns—focus modes, transparent execution, personalization, and reversibility—are positioned as engineering solutions to make agents trustworthy, contextually aligned, and low-risk. The talk is a product design framework disguised as agent analysis, directly applicable for building production agent UIs.

Key takes

  • Focus modes beat general-purpose agents because they constrain action space, improve eval quality on smaller surfaces, and align user expectations (Cursor's planning/debug modes as examples). Implication: "Do anything" UIs confuse users and dilute performance.
  • Speed-to-understanding > speed-to-outcome—agents that produce fast but misaligned outputs waste time. Personalization (Harvey's playbooks, Claude's skills/memory) encodes user principles so agents "do the right thing, not just something." Implication: Generic outputs have no value without user context.
  • Transparent execution builds trust and reduces waste—showing reasoning, tool calls, inputs/outputs (Claude Cowork, Manus examples) shifts users from delegation to collaboration, enabling early intervention when agents go off-track.
  • Reversibility bounds mistake cost, making users willing to tackle higher-value tasks. Cursor's multi-granularity undo (line/file/conversation state) and parallel model outputs explicitly reduce downside risk. Implication: Users won't trust agents with important work unless they can safely revert.

Useful details

  • Cursor's focus modes: Planning mode doesn't write code, just plans and asks questions. Debug mode spins up a dedicated debug server and uses hypothesis-driven troubleshooting.
  • Claude Cowork: Shows a to-do list of completed/pending tasks, displays all tool calls with inputs/outputs, and lists context sources and skills used.
  • Harvey (legal agent): Uses "playbooks" (firm-specific contract review methods/principles) and memory across interactions. Word add-in integrates with native Word track-changes API for reversibility.
  • Cursor reversibility levels: Line-level accept/reject, file-level accept, conversation state rollback (undo last N messages), and parallel outputs across different models for comparison.
  • Picasso quote framing: "Good artists copy, great artists steal"—interpreted as studying deeply to make concepts your own, not literal plagiarism.

Caveats / counterpoints

  • No evidence presented: All claims about "best agents" are based on speaker opinion and product observation, not benchmarks, user studies, or comparative data.
  • Sample bias: Only looks at coding/legal/productivity agents (Cursor, Claude, Harvey, Manus)—doesn't address whether these patterns generalize to other domains (customer service, research, creative work).
  • Reversibility trade-offs not discussed: No mention of implementation complexity, performance costs of tracking state for undo, or how reversibility scales with agent autonomy.
  • Personalization risks ignored: Playbooks/memory could encode user biases or outdated practices; no discussion of when to override user preferences vs. defer to them.
  • Focus modes might fragment UX: Switching modes adds friction; speaker doesn't address when unified interfaces might be better or how to prevent mode proliferation.

Ken relevance

High relevance for agent product design and GTM strategy. If Ken is building agent systems (especially for content, business ops, or dev workflows), these four patterns are immediately actionable UI/UX principles. Transparent execution and reversibility directly address the trust gap in autonomous agents—critical for enterprise adoption. Personalization via playbooks/memory could differentiate Ken's agents in vertical markets (e.g., GTM playbooks for sales agents, investment frameworks for research agents). Focus modes suggest opportunity to specialize agents per task type rather than building one general agent. For investing/sourcing, this framework helps evaluate which agent startups have product depth vs. just LLM wrappers. If Ken is creating content, this talk provides a strong contrarian angle: most agent builders optimize the wrong metric (speed-to-outcome vs. speed-to-understanding).

Watch verdict

Watch fully. This is a dense, opinionated framework backed by concrete product examples that directly applies to building or evaluating agent products. The speaker's argument that most agents optimize the wrong thing (outcome speed vs. understanding speed) is a useful mental model Ken can use across multiple domains (product, content, investing). The 10-minute investment pays off if Ken touches agent systems in any capacity.

Transcript

1777 words en Processed in 73.8s

All right, so before I jump into sharing with you what I believe the best agents share, I actually want to share a quote with you. And this is a quote that I actually keep quite close to me when I personally develop agents. It's a quote that maybe a lot of you might be familiar with, although I do want to dive into a little bit of what Pablo Picasso meant when he said steal in this quote. He didn't necessarily mean stealing in the sense of taking something physically that is not your own and presenting it as your own. But instead, he referred to going and looking at something, studying it deeply, really understanding it and making it your own, and then using that to come up with something better and something unique that you wouldn't have been able to come up with having not done this process. And that is really what I want to do today in this talk. I want to have a look at four of what I believe potentially to be some of the best agents that we have access to at the moment, go and study them deeply, understand what they do, and see what we can learn from them in order to ourselves actually build agents in a much better way. I'm going to have a look at four specific patterns that these agents use. For each of these, I'm briefly going to touch on what exactly this pattern entails, importantly, what is the value that it adds to you using them, and then thirdly, show you quickly how that actually looks in real life in these agents. The first one is what I call focus modes. And focus modes is really where we put the agent in a specific mode where we constrain the action and the input space. So we go into a planning mode or a research mode. What do we get from this? Well, first of all, the biggest benefit is for us as engineers. We get the ability to improve the agent's output quality on this smaller constrained action space. We really can potentially go and say, let's drop a bunch of tools, let's really refine our system prompt, let's optimize our evals to do really well on this small space first before we just do anything. Secondly, what is also really valuable actually is from a user perspective. One thing in these kind of do anything, ask me anything agent UIs is the fact that the user doesn't necessarily know what to do to get the best result out of the agent, and they also have very big expectations. So by going into a specific mode, we actually say, let's align a little bit the user's expectations and also tailor their inputs and behavior specific to this mode. Cursor does this really well. So on the right-hand side, you can see the Cursor chat interface, and you can very easily switch between different modes by simply selecting a drop-down. And each of these modes then has specific behaviors and expectations that it sets for the user. It then does very specific things. So in the middle, we see planning mode. It actually doesn't write any code. It just comes up with a plan, and it asks you questions, and you should be fine with it because that's what you signed up for. In debug mode, it has a very specific hypothesis-driven approach towards, okay, what are the potential issues with your code? Let's spin up a dedicated debug server and push logs there and actually figure it out. So in my opinion, a really powerful way in which Cursor is using modes to actually do certain things really well. The second pattern is transparent execution. And what we're trying to do in this instance is really trying to make what the agent is doing and using and thinking extremely clear to the users. And the crux of what we're trying to achieve here is to shift from delegation to collaboration, to really making the user part of the process and not just letting the agent come up with an end result. The benefits we're getting here is, first of all, trust in the output. If I give you a task and you come back with just simply the results, I will have less trust in the results than if you were to actually share with me your process, share the thoughts you had, what did you read, what did you assume, what were the things you actually are uncertain about. So we really use this process of transparency to build trust in the eventual outcome that the agent comes up with. Additionally, it also enables the user to intervene at an earlier point in time if it sees the agent is really doing the wrong thing and thereby reducing waste. If at step two of the agent, we saw the agent has just read from notion, docs A and B, and I wouldn't have done that, then we can very easily say, hey, I think let's stop and take a different approach. This is something that Claude Cowork for me does quite well. Top right, it has a progress list or a to-do list of things that it has done and will be doing, so it makes it clear what's the step that it's about to take. It gives you a good idea of the context that it's using, the skills that it's drawing from. In terms of tool calls, it's actually showing you all of the tool calls that it's making and also the inputs and the outputs of those tool calls. And this really makes it quite clear to the user what is actually going on from an execution perspective of the agent. Manus does something very similar. You also have your task progress where you can see the tasks completed and to be done, and it also gives you a very good idea of what it actually looked at and what it made of those things. The third pattern is personalization, and this is really where we try and give the agent the thoughts and systems and knowledge and principles and patterns that we would have used if we were to do the task ourselves. And fundamentally, what we're trying to get to here is to optimize or rather increase the speed of understanding of the agent. And this is a point that I think quite a few agents doesn't really get right in the sense that they optimize for speed to outcome, but not speed to understanding, in the sense that it's very easy to just generate an output for a user, but if it's not really in line with what the user wants in terms of how they wanted it, it's going to be useless. So optimizing for speed to understanding in the sense of really understanding all of the nuances and implicit things from the user, how it would have approached it, is really critical for an agent to do the right thing and not just something. Personalization is for us a way of enabling a quicker speed to understanding for the agent and doing the right thing and not just something. This is something we get in various flavors and different agents. For me, two, which is quite nice, is the one is Harvey. Harvey has this idea of a playbook, and a playbook is for legal firms, typically the methods and principles that they use to, for example, review a certain contract. And you can create these playbooks in Harvey, and the agent would then do it in the same way as what your legal firm would have done it. Harvey also uses a fairly common concept, which is memory, so it actually creates memories as we go along and as we instruct the agent, and it can then draw from that in subsequent interactions. Claude, like many others, also has the idea of skills and connectors and systems that you can connect to in order to increase this knowledge base and improve the personalization of your agent. And the last one is then reversibility, and reversibility is really the ability for the user to be able to reverse or undo the actions that the agent has done. And the big thing that we are achieving from this is we're bounding the cost of our mistakes. So if we know what the worst-case outcome is, or at least what the downside cost could be for me, it makes the ROI calculation much easier for me to actually say, happy if you go and do that, versus there could be fairly big consequences. This then results in users being bolder and much more prone to actually taking risks and tackling higher-value tasks and use cases for the agent to actually do. This is done really well for me by Cursor as well. Cursor actually enables this reversibility on different levels of granularity. So top left, you can actually roll back or choose on a line level what you want to accept or reject based on what the agent did. Bottom left, you can accept on a file level. Bottom right, you can actually go back into certain points of your conversation state. So you can say, we've now had this conversation, but actually the last three messages, all of the changes you've done, undo those and jump back. And then also it actually gives the ability to really do multiple outputs with the same input in parallel using different models. And thereby the user basically is knowingly saying, we will undo all but ideally one of our outputs in order to actually reach something that is valuable. So Cursor makes it really easy for you to not have much downside and experiment with things and try things out, knowing that you can, worst case, just undo and carry on. Harvey also does this quite well. And they actually use, so in this product, it's a Microsoft Word add-in that runs in Microsoft Word. And they actually integrate with the native Word API in order to have this changing of your changes and viewing of your changes in Microsoft Word as a reviewer or editor would natively using Word. All right. Thanks a lot. That was, I think, quite a lot for a short amount of time. I hope it was useful. Please reach out if there's more questions. Thank you. And that is really what I want to do today in this talk. I want to have a look at four of what I believe potentially to be some of the best agents that we have access to at the moment, go and study them deeply, understand what they do, and see what we can learn from them in order to ourselves actually build agents in a much better way. I'm going to have a look at four specific patterns that these agents use. For each of these, I'm briefly going to touch on what exactly this pattern entails, importantly, what is the value that it adds to you using them, and then thirdly, show you quickly how does that actually look in real life in these agents. The first one is what I call focus modes. And focus modes is really where we put the agent in a specific mode where we constrain the action and the input space. So we go into a planning mode or a research mode. What do we get from this? Well, first of all, the biggest benefit is for us as engineers, we get the ability to improve the agent's output quality on this smaller constrained action space. We really can potentially go and say, let's drop a bunch of tools, let's really refine our system prompt, let's optimize our evals to do really well on this small space first before we just do anything. Secondly, what is also really valuable actually is from a user perspective. One thing in these kind of do anything, ask me anything agent UIs is the fact that the user doesn't necessarily know what to do to get the best result out of the agent, and they also have very big expectations. So by going into a specific mode, we actually say, let's align a little bit the user's expectations and also tailor their inputs and behavior specific to this mode. Cursor does this really well. So on the right-hand side, you can see the cursor chat interface, and you can very easily switch between different modes by simply selecting a drop-down. And each of these modes then has specific behaviors and expectations that it sets for the user. It then does very specific things. So in the middle, we see planning mode. It actually doesn't write any code. It just comes up with a plan, and it asks you questions, and you should be fine with it because that's what you signed up for. In debug mode, it has a very specific hypothesis-driven approach towards, okay, what are the potential issues with your code? Let's spin up a dedicated debug server and push logs there and actually figure it out. So in my opinion, a really, really powerful way in which Cursor is using modes to actually do certain things really, really well. The second pattern is transparent execution. And what we're trying to do in this instance is really trying to make what the agent is doing and using and thinking extremely clear to the users. And the crux of what we're trying to achieve here is to shift from delegation to collaboration, to really making the user part of the process and not just letting the agent come up with an end result. The benefits we're getting here is, first of all, trust in the output. If I give you a task and you come back with just simply the results, I will have less of trust in the results than if you were to actually share with me your process, share the thoughts you had, what did you read, what did you assume, what were the things you actually are uncertain about. So we really use this process of transparency to build trust in the eventual outcome that the agent comes up with. Additionally, it also enables the user to intervene at an earlier point in time if it sees the agent is really doing the wrong thing and thereby reducing waste. If at step two of the agent, we saw the agent has just read from, I don't know, so notion, docs A and B, and I wouldn't have done that, then we can very easily say, hey, I think let's stop and take a different approach. This is something that Claude Cowork for me does quite well. Top right, it has like a progress list or a to-do list of things that it has done and will be doing, so it makes it clear what's the step that it's about to take. It gives you a good idea of the context that it's using, the skills that it's drawing from. In terms of tool calls, it's actually showing you all of the tool calls that it's making and also the inputs and the outputs of those tool calls. And this really makes it quite clear to the user what is actually going on from an execution perspective of the agent. Manus does something very similar. You also have your task progress where you can see the tasks completed and to be done, and it also gives you a very good idea of what it actually looked at and what it made of those things. The third pattern is personalization, and this is really where we try and give the agent the thoughts and systems and knowledge and principles and patterns that we would have used if we were to do the task ourselves. And fundamentally, what we're trying to get to here is to optimize or rather increase the speed of understanding of the agent. And this is a point that I think quite a few agents doesn't really get right in the sense that they optimize for speed to outcome, but not speed to understanding, in the sense that it's very easy to just generate an output for a user, but if it's not really in line with what the user wants in terms of how they wanted it, it's going to be useless. So optimizing for speed to understanding in the sense of really understanding all of the nuances and implicit things from the user, how it would have approached it, is really critical for an agent to do the right thing and not just something. Personalization is for us a way of enabling a quicker speed to understanding for the agent and doing the right thing and not just something. This is something we get in various flavors and different agents. For me, two, which is quite nice, is the one is Harvey. Harvey has this idea of a playbook, and a playbook is for legal firms, typically kind of the legal expert, but as I understand the methods and principles that they use to, for example, review a certain contract. And you can create these playbooks in Harvey, and the agent would then do it in the same way as what your legal firm would have done it. Harvey also uses a fairly common concept, which is memory, so it actually creates memories as we go along and as we instruct the agent, and it can then draw from that in subsequent interactions. Claude, like many others, also has the idea of skills and connectors and systems that you can connect to in order to increase this knowledge base and improve the personalization of your agent. And the last one is then reversibility, and reversibility is really the ability for the user to be able to reverse or undo the actions that the agent has done. And basically, the big thing that we are achieving from this is we're binding the cost of our mistakes. So if we know what the worst-case outcome is, or at least what the downside cost could be for me, it makes the ROI calculation much easier for me to actually say, happy if you go and do that, versus there could be fairly big consequences. This number one then results in users being bolder and much more prone to actually taking risks and tackling higher-value tasks and use cases for the agent to actually do. This is done really, really well for me by Cursor as well. Cursor actually enables this reversibility on different levels of granularity. So top left, you can actually roll back or choose on a line level what you want to accept or reject based on what the agent did. Bottom left, you can accept on a file level. Bottom right, you can actually go back into certain points of your conversation state. So you can say, we've now had this conversation, but actually the last three messages, all of the changes you've done, undo those and jump back. And then also it actually gives the ability to really do multiple outputs with the same input in parallel using different models. And thereby the user basically is knowingly saying, we will undo all but ideally one of our outputs in order to actually reach something that is valuable. So Cursor makes it really, I would say, easy for you to not have much downside and experiment with things and try things out, knowing that you can, worst case, just undo and carry on. Harvey also does this quite well. And they actually use, so in this product, it's a Microsoft Word add-in that runs in Microsoft Word. And they actually integrate with the native Word API in order to have this change, you could say, doing of your changes and viewing of your changes in Microsoft Word as a reviewer or editor would natively using Word. All right. Thanks a lot. That was, I think, quite a lot for a short amount of time. I hope it was useful. Please reach out if there's more questions. Thank you. Thank you. Thank you.