Open Reader

Your agent architecture has a half-life of 6 months — Dan Farrelly, CTO, Inngest

completed 19:20 Jul 21, 2026 Watch on YouTube

Current Status

completed

Video ID

X1kp-ABIIxQ

RAG / Chat

Enabled
Your agent architecture has a half-life of 6 months — Dan Farrelly, CTO, Inngest
Description

A short history of the right way to build an agent: RAG, ReAct, prompt chaining, orchestrator-workers, MCP, CLI, MCP again... CLI again?? Every time you adopt a trend you rebuild your architecture. In this talk, Dan Farrelly, Inngest cofounder and CTO, is not going to tell you what comes next. He's going to show you how to build so it doesn't matter. He'll cover the core primitives that show up in every production agent, how bringing decisions closer to code provides more stack flexibility, and why the right execution layer unlocks faster iteration. ### Dan Farrelly CTO and Co-founder · Inngest [LinkedIn](https://www.linkedin.com/in/djfarrelly) Dan Farrelly is CTO and co-founder of Inngest, a platform for durable serverless functions, workflows and agent orchestration. He was previously CTO at Buffer and created developer tools including Timezone.io and MailDev.

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Because models, prompts, tools, and sandbox patterns churn rapidly, agent systems should isolate a durable execution layer—responsible for state, retries, orchestration, and traces—from the fast-changing context and compute layers.
  • Why it matters: This is a directly applicable architecture pattern for building long-running, multi-agent systems without repeatedly rewriting reliability and orchestration logic whenever the AI stack changes.
  • Best use: Use it as a design review framework for OpenClaw or other agent systems: identify where execution, context, and compute are coupled, then define a durable control plane with resumability, flexible invocation, and end-to-end observability.

Executive Summary

Dan Farrelly argues that agent architectures have a practical six-month half-life because the context layer—models, prompts, tool standards, memory approaches, and frameworks—changes continuously. Teams create avoidable rewrite cycles when they bind those volatile choices to execution concerns such as workflow state, retries, scheduling, coordination, and failure handling. His proposed mental model separates an agent system into execution (the brain), context (knowledge), and compute (hands).

The stable investment should be the execution layer: an infrastructure-independent system that manages the lifecycle of work across planning, model calls, code execution, sub-agent invocation, loops, retries, and coordination. It must preserve durable state outside transient work environments, support multiple triggering and invocation patterns, and provide a trace across the whole session rather than only LLM calls. This lets teams replace a model, prompt stack, toolset, or sandbox without redesigning core operations.

The talk is especially relevant for emerging long-running patterns: background agents, scheduled or continuous loops, delegated workflows, and agent factories. Farrelly illustrates a health-monitoring loop that runs every 30 minutes, triggers a deeper triage agent only when needed, and uses a weekly reviewer to inspect execution history and improve the system. His broader point is that evaluation should use execution and business outcome signals—such as whether an engineering team opened a PR after triage—rather than relying only on thumbs-up/down feedback.

Although the talk is also an Inngest product pitch, its architecture guidance is concrete and portable. The key decision is not necessarily to adopt Inngest; it is to ensure that whichever orchestration/control-plane choice Ken makes delivers durable resumability, composable triggers, externalized state, and full-session traces.

Key Takeaways

  • Claim: Treat agent architecture as three distinct layers: execution, context, and compute, with execution intentionally decoupled from the other two. | Evidence: Farrelly defines execution as flow, state, durability, and retries; context as models, prompts, tools, and memory; and compute as sandboxes, runtimes, and automated browsers. | Implication: Agent-system interfaces should make models, prompts, tools, and compute backends replaceable dependencies rather than embedding them in workflow control logic.
  • Claim: The context layer has the shortest useful life, while a properly designed execution layer can remain valuable for years. | Evidence: He estimates prompts may last only weeks, models months, and execution years; he attributes rewrite cycles to one layer's short half-life dragging coupled layers down with it. | Implication: Put the most design effort and long-lived abstractions into orchestration, state, reliability, and auditability—not into framework-specific agent chains. | Caveat: These are directional lifecycle estimates, not measured benchmarks; the architectural principle matters more than the specific durations.
  • Claim: Resumability requires durable state external to the agent's current process, memory, disk, or sandbox. | Evidence: For a failure at step 38 of a long run, the system should retry or continue from that point rather than restart and lose prior token spend, elapsed time, and completed work. Farrelly says a three-hour run cannot safely retain its state solely in memory or on local disk. | Implication: Design every meaningful agent step to be checkpointable and replay-safe; do not make a sandbox or a single request process the source of truth for progress. | Caveat: External state introduces its own requirements around idempotency, state-schema evolution, and data/security governance, which the talk does not address.
  • Claim: A real execution layer must support composable invocation patterns rather than forcing orchestration into agent prompts or framework chains. | Evidence: The required primitives cited are cron scheduling, event triggers, APIs, human-in-the-loop actions, sub-agents, dynamic workflows, synchronous and asynchronous calls, and delayed invocation. | Implication: Choose an orchestration substrate that can express scheduled loops, event-driven work, delegation, waits, and human escalation without separately rebuilding queues, polling, workers, backoff, and scheduling in application code.
  • Claim: End-to-end execution observability is necessary to operate and improve agents, especially asynchronous and background systems. | Evidence: Farrelly says traces must cover the trigger through the entire stack, including model/tool activity, database errors, permission problems, triggers, and performance. He notes that a background agent may run for hours, make hundreds of calls or roughly 200 tool calls, and is likely to experience at least one failure. | Implication: Instrument a session-level trace and correlation model spanning user input, orchestration decisions, sub-agent fan-out, infrastructure actions, and outcomes; LLM traces alone are insufficient.
  • Claim: Sandboxes should be treated as ephemeral compute, not as the durability mechanism for an agent. | Evidence: Farrelly calls snapshots or sandbox-held state an anti-pattern because sandboxes are stateless and ephemeral by design; in his framing, compute is the agent's hands while execution supplies its context, sequence, and durability. | Implication: Use sandboxes for code execution, browsing, file manipulation, and isolated runtime work, while persisting agent progress and coordination state independently. | Caveat: Some sandbox products may offer persistence features, but the architectural point remains that durable workflow state should not depend on a particular compute instance surviving.
  • Claim: Agent evaluation should connect traces to observable downstream outcomes, not depend solely on explicit user feedback. | Evidence: In the triage example, useful success evidence could be whether the engineering team opened a PR; for research, whether the report was saved and used. He proposes attaching later events to the originating execution session and scoring with complete trace context. | Implication: Define outcome events for each high-value workflow, link them to execution IDs, and use them to drive post-run scoring, prompt/tool changes, and system-level evaluation. | Caveat: Outcome proxies can be delayed or confounded—for example, a PR may be opened for reasons unrelated to the agent—so they need explicit attribution rules and complementary quality checks.

Detailed Brief

Illustrative self-improving operational loop

  • Claims: A production agent loop can remain conceptually simple even when its individual tasks are complex: detect, investigate, then review and improve.; The reviewer should be orchestration-aware, not merely output-aware, because it needs to inspect what the system actually executed and delegated.
  • Evidence: Farrelly's example runs a health check every 30 minutes, collects high-level metrics, asks an LLM whether conditions are healthy, and invokes a triage agent if they are not.; The triage agent can retrieve detailed metrics, loop through tools, spin up a sandbox, clone code, and analyze commits to identify a root cause.; A weekly reviewer examines the history of triage behavior: whether prompts need adjustment, metrics are adequate, the agent has sufficient data, and the system is overreacting or failing to act.
  • Caveats: The example abstracts away production design issues such as authority boundaries, alert fatigue, safe remediation, and escalation policies.; A reviewer that can change its own agent configuration requires versioning, approval controls, and rollback paths.
  • Implications: Separate detection, investigation, and review into independently observable workflows with clear contracts rather than one monolithic autonomous agent.; Make improvement agents read structured execution history, including calls, retries, delegation, and outcomes, instead of judging only final text responses.

Notable Concepts & Terms

  • Architecture half-life: The idea that parts of an agent stack decay in usefulness at different rates; volatile components should not dictate the lifespan of core operational infrastructure.
  • Execution layer: The durable control plane for an agent's lifecycle: state, retries, scheduling, coordination, workflow progress, and end-to-end traces.
  • Context layer: Models, prompts, tools, and memory—the knowledge and decision inputs expected to change most frequently.
  • Compute layer: Ephemeral operational environments such as sandboxes, runtimes, and browsers that perform work but should not own durable workflow state.
  • Resumability: The capability to continue a multi-step execution after failure from a durable checkpoint rather than restarting the entire run.
  • Outcome-based scoring: Evaluating an agent from downstream events tied to its execution, such as a triage run leading to a PR, rather than only user ratings or model-output evaluation.
  • Execution-aware reviewer: A scheduled meta-agent or workflow that inspects traces and delegation history to assess whether an agent system is performing appropriately and how it should change.

Operator Notes / Why Ken Should Care

  • Run an architecture audit of each long-running agent: map which components currently own state, retries, scheduling, model routing, tool calls, sandbox lifecycle, and trace correlation.
  • Set a non-negotiable rule that sandbox-local files, process memory, and request state are not the authoritative record of workflow progress.
  • Define a portable execution contract for agents: durable step IDs, idempotency expectations, event/cron/API triggers, wait states, delegated-run lineage, and session-level trace IDs.
  • For every material agent workflow, select one or more downstream outcome events and establish attribution windows before using those outcomes as quality scores.
  • Require human approval and rollback/version controls before allowing reviewer or evaluator agents to modify prompts, tool policies, workflow logic, or production thresholds.

Source/Metadata

  • Title: Your agent architecture has a half-life of 6 months — Dan Farrelly, CTO, Inngest
  • Transcript words: 2853
  • Duration seconds: 1160
  • Timestamp note: No timestamps or chapter markers were present in the supplied transcript.

Transcript

2825 words en Processed in 121.2s

Going good? All right. There we go. All right. Hello, everyone. All right, just checking voice. All right, good. All right, well, my name's Dan, and I'm here to talk to you about how your agent architecture has a half-life of six months. But first, who am I? I'm Dan. I'm the CTO and co-founder at Ingest. And why should you listen to me? Well, first, I lead an amazing team over at Ingest. We build a system that reliably executes anything from agents to workflows to context pipelines, whatever you want. I'm deep building AI infra every day. And on top of that, I'm also building agents, not just pontificating up here in theory about how you should build. So speaking of building agents, if you've been building agents for more than six months, you've likely rewritten something, maybe more than once. A new model, a new framework or framework version, a new tool calling standard, a new pattern. Suddenly, your architecture doesn't fit. It's not a complaint. It's just the reality. Right? Things move faster than ever. You can see if you've been to any of the sessions or keynotes. So I want you to think about what you shipped six months ago and how much of it still runs the way that you originally wrote it. Some parts of your code likely survived, but did they survive by accident, or did you design it that way? Did they survive by design? How did you actually architect your agent? Did you architect it at all? It's okay. Did you use a framework? Did you custom roll something? Was the architecture an intentional design that you really thought through, or did it just evolve into what you have now? These are all various ways of where people are these days. So a lot of folks have talked about harness architecture, building harnesses. But mostly, I think a lot of people draw diagrams and they talk about specific components. Draw nice lines. It looks very simple, simple. But I want to talk about the conceptual layers when you're building a system like this. This is maybe the mental model, not specific components. So in my opinion, there are three discrete layers. First, the execution layer. I think of this as the brain. It's where flow, state, durability, retries happen. Then there's the context layer. This is the knowledge. This is models, prompts, tools, memory. This is the layer that changes the most. And then there's compute. This is the hands, right? This is sandboxes, runtimes, browsers that you're automating. I think these layers are important to consider for the following reason. It's your half-life. And what is half-life? It's a scientific term. And it's a scientific term for the time that it takes for something to decay by half. And I think that your architecture has a half-life also. So prompts last weeks if you're lucky. Maybe a single week. The models that you use, months, again, if you're lucky. But I think that execution can last years, if you do it right. So the problem is that I think that most teams couple everything together. And what happens then is that one layer's half-life leaks and drags the other components down. It's technical debt by another name. So my thesis is: think in layers. Decouple them. So let's talk about what I mean. Teams run into a lot of issues with the approaches that they choose. You might choose a framework. It might feel super magical to get going. You might grab a pre-built harness from one of the frontier labs or wherever. Or you might custom roll the entire system yourself. And it might take months or weeks or whatever. But I think in a lot of these situations, the abstractions either are not there at all, or they're too high level, or the layers merge. Orchestration is buried deep inside the chain or inside the framework. And you don't know what the heck's happening. State might be kept in a sandbox. Retries get tangled with the prompt logic. And what's hard is you can't swap any of these things out without rewriting almost everything. So I think you need to embrace the change. Know that things are going to change. Know that things are moving fast. But I think also with that, I want to focus on a layer that I think is the stable layer here, where you can invest and think about, get your abstractions right, so you don't need to rewrite a large component every six months. So I'm talking about the execution layer. I don't think enough people talk about this. And I define the execution layer as being the system responsible for running your code reliably, managing how, when, or whether each piece of work completes. And that's independent of the infrastructure that it's on. So what does this look like in an agent architecture? So the execution layer manages the full life cycle. Plan, call a model, run code, invoke a sub-agent, loop, retry, coordinate. You can swap the model, swap the context, swap the sandbox. The execution layer should be able to remain the same. It's just the fundamentals. And that's how you, I think, decouple and build the system so that it can evolve for change. So let's talk about what the execution layer has to do. First, resumability. Your execution layer must enable your system to pick up after failures without restarting from the beginning. Agents are handling longer-running tasks than ever before. So they must be able to resume when an LLM call fails, a tool call fails, everything flakes out. We all know. So if you have a failed request at step 38, you should be able to retry, wait, continue onwards, instead of having to go from the beginning, and you're going to lose tokens, costs, time, work that your agent might have completed. So for this to work, a three-hour run cannot hold state in memory or on disk. The state must live outside of the work. So this means that the state must be durable and external. So without it, you might cobble together some manual checkpointing. You might come up with a system that has a log-based approach where you're going to be hydrating the state back if you have to pick up where you left off. But I think a lot of those abstractions start leaking into the other layers of your harness. Right? It gets harder to change. So next, a key aspect of the execution layer is that it must enable you to combine a lot of invocation patterns. You're going to need crons. You're going to need to trigger things with events. You're going to need APIs. Human loop. Sub-agents or dynamic workflows also must be possible. And you're going to need to be able to invoke these things synchronously, asynchronously, delaying invocation. And I think the key here is that flexible execution and orchestration primitives will enable you to build the pattern or the system, the architecture, that you actually need. So if not, I think, again, your harness logic starts absorbing other concepts like queues, workers, polling, backoff, scheduling. And now you end up a mess, bad abstractions. So lastly, execution needs to provide observability across your entire session, not just the LLM calls and the tool calls, but database errors, permissions issues, triggers, performance. So the full session trace across your entire run is essential. So if you can't see the entirety of a trace from the trigger through the whole stack, it's really hard to debug it, let alone improve your agent and keep evolving it. So what about sandboxes? Sandboxes are so hot. So agents need sandboxes. We've settled upon that at this point in time. They need to execute code. They need to browse. They need to manipulate files. They might work on disk. But a sandbox is ephemeral and stateless by design. So using it for durability, with snapshots or something in state, I think is an anti-pattern. I think it's a difficult thing where the state gets lost and you're trying to pull the pieces back together. So I think when you have the execution layer separate, the execution layer is what gives the sandbox its context, its sequence, its durability. So I think of the sandbox as the hands and execution as the brain. So what's next? What are the next six months of architectures that you're going to need to handle? So if you're paying attention, you're joining a lot of these sessions, you'll see what architectures are emerging and what each approach brings: new engineering requirements. You have background agents, dynamic workflows, autonomous loops, agent factories, whatever you want to call it, all these emerging trends that we're seeing the last couple days. So the point of view is they're all long-running, they're asynchronous, they're delegated. That means that you need to be able to observe them all down to the core, right? They need to be inspectable by human and by an agent. So the patterns, as these things emerge, must be mixed and combined together. So what does this all have in common? I think that execution is fundamental to building any of these systems. So let's just take one example here. Background agents, right? They're not request-response. There's no person just waiting there for the chatbot to respond. So it might run for minutes or hours. It's going to have maybe hundreds of calls, maybe 200 tool calls. You're probably guaranteed to have at least one failure in that. So you can't even debug your background agent that's running asynchronously without the right infrastructure, without the observability, without everything that's going on in that process or multiple processes. And another example that we have is loop architectures, right? And I think we're talking about slash loop commands and whatnot and coding agents. But I think where we're taking this is: how are you building actual systems in your products that employ these approaches? So what is a loop? Right? A loop is a system that basically runs continuously or on a schedule. And it's assessing the state of the system against the goals that you set or the criteria that you set and determines what to do next. So you're going to need crons, sub-agent delegation, history needs to be inspectable. And of course, it needs to be reliable because you don't know when things are running or how the system is continuing to evolve. So the frameworks of three months ago were not designed to handle this, right? You're going to need to design these systems yourself. So I think you're going to need a proper execution layer. So let's look at some examples of code. It's small on here, but in general, let's just look at how you might put it together. In your loop, you're going to need some sort of cron. Maybe that cron here is some sort of health check system. It runs every 30 minutes, pulls down some key high-level system metrics, and it looks, you pass to an LM, you say, is this healthy? Is this looking okay? Should we do more? And is it healthy or not? What should we do? And if it isn't healthy, maybe you just invoke a triage agent that goes and looks at that individual service and digs more. It's pretty simple-looking, right? Now you have this triage agent itself. You've just triggered it. This should maybe pull some more context, pull some detailed metrics. It's just context. And then you're passing to the LM to start that investigation. Putting it in a loop, calling some tools, trying to get to the root cause. This may run for a minute. This may run for multiple minutes. It needs to spin up a sandbox, maybe clone code, maybe analyze commits. It's going to be doing a lot of work here. And you don't know when this is going to run. It might run, it might not. How do you know? How do you observe that system later? And then to complete the loop, you may have a reviewer function, right? Is the system actually performing as expected? It's going to run every week and look at the history of what just happened and evaluate how the triage system is working. Do we need to adjust the prompts? Do we need to do better metrics? Does this have the data that it needs? Is it doing anything at all? Is it overreacting? So it needs to be execution-aware. And I think it needs to be execution-aware or orchestration-aware, because it needs to be able to pull the logs of the system and understand, all right, what was executed? What did this agent choose to execute? What sub-agents did it fan out? What workflows did it call upon? And then it needs to be able to analyze that and understand what happened, to come back to you, either make the change itself or tell you, this is where we performed or underperformed. This is what we can do, how we can improve the system overall. So this completes this improving loop. It's three functions. It's pretty simple by design. It's flexible to iterate on. I know things always get a lot more complex in production, but I think when you think about the layers this way, it makes things a little bit easier to think about, building such a complex system that feels like it's something that you might not be able to reach yet. So in the spirit of the reviewing functions, what do you think about measuring your agent, how your agent is performing? I think what's interesting is this execution layer sits between what your users' input is and what's happening throughout the system. So user feedback, actions, and the results of the sessions all flow through this execution layer. And I think this makes it an ideal place because it becomes the hub for observability, and it allows you to be able to score and understand what your agent is doing. So to iterate on your application, you need all this data connected, and it needs to be aware of how your code is actually executed. Right? So I think the linking of this is extremely important. So this is what we built Ingest for. We're durable execution for AI agents. We're the execution layer. You can plug in any context layer, bring a model, framework, tool, any compute layer, bring whatever sandbox you want, bring whatever browser you want, get durable steps, small primitives, right? Event triggers, scheduling, agent-to-agent coordination, full session traces, no infra to manage. And I think, like I just mentioned with scores, we also think that the middle of the execution layer, when your agent is running, is a key place to instrument and score your agents. When you have access to the data that you need that's all flowing through there or you've just executed something, it's a perfect place to run something after and defer some scoring with the whole trace information, maybe the inputs and the outputs. So again, that's just execution orchestration, delayed task defers. And since there's also data flowing through, you can also wait for additional events and attach them to the sessions that you're building and understand, did this actually work? Right? If you're running this triage, was this triage successful? Did it result in an action by the engineering team? Then that means it probably was a positive result. Instead of a thumbs up, thumbs down, it's like, did we open the PR? Right? If it's a research agent, was this research saved? Was it a good report? These are things that are events that you should be able to attach. And when you have a system that can connect all these pieces, I think it really makes doing a lot of those things on creating outcome-based scores a lot easier. So to wrap up, build your harness. Understand the layers of the architecture. We have to embrace the fast pace of change. And I think if you can get your execution layer right and think about the right primitives, everything else can quickly evolve around it. And the next six months, the next three months will all be a lot easier, especially as everything continues to change. You all are here, so thanks for listening. I'm Dan. Come and find me at the Ingest booth. We're right on the other side of this wall. We're very bright and orange. So thanks, everyone. Appreciate it. to find out. to find out. to find out. to find out. to find out. to find out. to find out. bright and orange. So thanks everyone. Appreciate it. to find out. to find out. to find out. to find out. to find out. to find out. to find out.