AI Engineer

Agent Frameworks Considered Harmful — Rémi Louf, .txt

1854 summary words 8 min summary Watch video

Start with the signal

8 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Agent frameworks are often the wrong primary abstraction: reliable background agents need a runtime built around event-driven execution, durable logs, reproducible prompt state, and typed boundaries rather than code-defined graphs and chat-session UX.
  • Why it matters: This is a concrete control-plane architecture for turning agents from supervised terminal sessions into dependable, auditable operational systems—directly relevant to agent orchestration, OpenClaw-style automation, model routing, and AI ops.
  • Best use: Watch for the architecture and failure-driven design rationale, then use it as a checklist for evaluating or building an internal agent runtime rather than adopting an agent framework wholesale.

Executive Summary

Rémi Louf, CEO of .txt, argues from a two-week internal build rather than theory. He wanted an unattended morning-briefing system that combined scheduled market monitoring with event-triggered processing of walking voice notes, CRM updates, emails, and code events. His central objection is not that frameworks are universally bad, but that code-centric, graph-oriented frameworks tend to make agents live inside framework abstractions while failing to provide the runtime guarantees needed for autonomous background operation.

His alternative is an event-driven runtime. Agents are simple declarative definitions—initially Markdown/YAML-like files—that declare accepted event types and emitted event types. Schedules publish events, agents subscribe to them, and delivery processes such as Slack posting also subscribe to events. This removes manually maintained graph edges and enables fan-in and fan-out through event contracts, while making workflows editable by non-engineers who understand the system's available events.

The practical insight is that autonomous agents quickly reproduce ordinary distributed-systems problems: duplicate execution, dropped work, retries, queues, state, observability, and version control. Louf's system evolved from real failures: duplicate daily briefs led to proper queue/retry handling, a vanished voice note led to an append-only causal event log, and an untraceable degraded prompt led to content-addressed prompt construction. The runtime, not the agent prompt, becomes the durable product.

The most technically useful contribution is prompt provenance. Instead of treating a visible chat transcript as ground truth, the system stores each prompt component—system prompt, skills, tools, user input—and outputs as hash-addressed objects. A run can therefore be reconstructed exactly, diffed against another run, replayed with a different model, and audited despite compaction or unavailable chain-of-thought. Louf also makes typed tool calls and typed inter-agent events non-negotiable, arguing the runtime should make invalid actions impossible rather than merely discourage them.

Key Takeaways

  • Claim: The valuable endpoint for agents is unattended background work, not increasingly convenient supervised chat or terminal control. | Evidence: Louf contrasts a terminal UI with a tractor mower that still requires the user to ride it, and mobile agent apps with a mower controlled by a remote. His working system automatically produces a Slack morning brief from market monitoring and processed walking voice notes. | Implication: Prioritize workflows that can run asynchronously and deliver completed artifacts or exception alerts; do not mistake a capable interactive coding agent for an autonomous operations system. | Caveat: The use case is primarily internal knowledge-work automation, not a demonstration that all high-stakes workflows should run without human approval.
  • Claim: Event subscriptions are a simpler and more extensible coordination primitive than explicitly maintained agent graphs for many workflows. | Evidence: A voice-note processor accepts a voice-note event, emits a processed-note event, and the daily-brief agent combines that output with schedule-generated market-watch events; Slack delivery subscribes to a message-post event. Louf says fan-in and fan-out become free because agents subscribe to event types rather than requiring graph edges. | Implication: Model agent interfaces as explicit event contracts and let topology emerge from publishers and subscribers; reserve explicit workflow graphs for cases where their control semantics are genuinely necessary. | Caveat: This is an architectural preference demonstrated on a roughly 20-agent internal deployment, not a benchmark proving event-driven designs dominate graphs in every long-running or highly stateful workflow.
  • Claim: Reliable agent orchestration is predominantly a distributed-systems/runtime problem, not a novel prompting problem. | Evidence: The initial implementation took about one day with Codex but immediately produced duplicate Slack briefs, lost a voice note, and unreproducible prompt regressions. Fixes required a durable log, proper queues and attempt tracking, and versioned prompt state. | Implication: Assess an agent platform by its guarantees around delivery, retries, idempotency, causal tracing, persistence, and debugging—not only model quality, tool calling, or orchestration syntax.
  • Claim: An append-only, causally linked event log is the system's essential memory and debugging surface. | Evidence: Louf describes a single append path to an events table where nothing is lost, every event is queryable, and events are causally linked so an operator can identify which event triggered a downstream event. He says debugging becomes difficult even with only three or four agents. | Implication: Require durable event history and parent/causation IDs from the first production agent workflow; without them, failures will be difficult to reproduce and operational trust will erode quickly.
  • Claim: Prompt and model-input provenance should be content-addressed so every result can be exactly reconstructed, compared, and replayed. | Evidence: The runtime stores the system prompt, skill definitions, tool descriptions, user messages, and model answers as hash-addressed components; rendered prompts are represented as lists of hashes. This enables run diffs, exact request reconstruction, compaction-friendly manipulation of structured prompt graphs, and replay against another model. | Implication: Treat prompts as versioned build artifacts rather than mutable strings in application code. This supports regression diagnosis, evaluation, model migration, auditability, and cache-aware optimization. | Caveat: Exact reconstruction of visible inputs does not reveal proprietary hidden reasoning traces, which Louf notes providers do not expose.
  • Claim: Typed tool calls and typed events are safety and reliability boundaries that should reject invalid actions before they propagate. | Evidence: Louf reports that Anthropic's structured-output behavior initially caused roughly 20% of events to be malformed and rejected. He frames the runtime's job as making bad actions impossible, using typed tool calls at the external boundary and typed events between agents. | Implication: Enforce schemas at every agent-to-tool and agent-to-agent boundary, while pairing them with authorization, policy checks, idempotency controls, and approval gates for consequential actions. | Caveat: Schema validation prevents malformed interfaces but does not establish that a well-formed action is semantically correct, authorized, or desirable.
  • Claim: The agent-infrastructure market remains unsettled; small technical teams should build a thin internal system before committing to a vendor or framework. | Evidence: After trying existing options, Louf built his own runtime through failure-driven iteration, deployed it internally for a month, and reports 20 agents including contributions from nontechnical staff. He later replaced third-party APIs with open-source models, including local laptop inference, for his use case. | Implication: Run a narrowly scoped internal build to discover required primitives and operational constraints, then decide which components to buy, build, or replace with local/open-source models. | Caveat: The open-source-model claim is explicitly scoped to his workloads; it is not evidence that local models meet frontier-quality, latency, security, or scale requirements broadly.

Detailed Brief

Declarative agent definitions as userland

  • Claims: Louf separates the runtime from agent definitions: the runtime schedules, isolates, journals, and executes agent processes, while Markdown-based definitions are merely one user-facing interface.; Keeping agent definitions declarative lets teams version, diff, and review operational behavior through normal repository workflows without embedding every prompt change in application code.; The system can support other frontends; Louf notes that one frontend does not use the Markdown format at all.
  • Evidence: He initially resisted YAML but found a file-based interface easier because a user can write a file, drop it into a folder, and have the runtime load it.; He compares the runtime/userland separation to an operating system kernel and processes, while acknowledging the operating-system analogy is overused.
  • Caveats: Declarative files improve accessibility and reviewability, but they do not eliminate the need for strong runtime semantics, governance, testing, and deployment control.
  • Implications: Design the agent authoring layer as replaceable configuration over stable runtime primitives.; Enable nontechnical contributors only where event schemas, permissions, validation, and review paths are understandable and controlled.

Model portability and cost as a replay problem

  • Claims: Observability exposes the speed at which agent costs can rise, particularly when automated workflows create repeated model calls.; Content-addressed prompt state turns model substitution into a controlled replay/evaluation operation rather than a rewrite of the workflow.
  • Evidence: Louf rebuilt prior requests from their component graphs, sent them to alternative models, and evaluated whether outputs remained satisfactory.; He reports ultimately removing third-party APIs for this system and using open-source models, including a local model on his laptop.
  • Caveats: The talk provides no comparative quality, latency, throughput, or total-cost measurements for the model migration.
  • Implications: Store production request fixtures so model-routing changes can be tested against real historical workloads before rollout.; Use cost visibility to identify agent tasks suitable for cheaper or local models instead of applying a single frontier model everywhere.

Notable Concepts & Terms

  • Event-driven agent runtime: The proposed control plane: schedules and system changes emit events, and agents react by consuming and producing typed events.
  • Append-only causal event log: Durable system memory in which every event is retained, queryable, and linked to its triggering event for traceability.
  • Content-addressed prompts: Prompt components are stored by hash, like Git or Nix artifacts, enabling exact provenance, diffs, and deterministic reconstruction.
  • Prompt replay / replace: Rebuilding a prior model request from stored components and resending it to another model or altered request configuration for evaluation.
  • Typed events: Schema-constrained messages between agents; a primary reliability boundary that prevents malformed inter-agent handoffs.
  • Typed tool calls: Schema-constrained interactions between an agent and external tools, intended to prevent calls to invalid or nonexistent interfaces.
  • Userland agent definitions: Declarative Markdown/YAML-like agent specifications kept separate from the runtime, allowing versioning and non-code authoring.
  • Build before buy: Louf's recommendation to construct a small internal system first because agent infrastructure products have not yet stabilized and requirements are workload-specific.

Operator Notes / Why Ken Should Care

  • Define a minimum viable agent-runtime contract: durable append-only events, causation IDs, retry/attempt records, idempotency handling, typed event schemas, and typed tool schemas.
  • For each existing or planned agent, replace direct agent-to-agent calls with declared input and output event contracts; document publishers, subscribers, ownership, and permissions.
  • Create a prompt-artifact store that versions system instructions, tool definitions, skills, retrieved context, user inputs, and outputs independently; make run-level diffs and replay mandatory for regressions.
  • Build a historical replay suite from real production tasks before changing models, routing policies, prompt components, or tool schemas.
  • Separate unattended, low-consequence background workflows from actions requiring approval; add policy and authorization gates because valid structured output is not equivalent to safe intent.
  • Run a focused internal prototype before adopting an orchestration framework, and evaluate vendors against operational primitives rather than graph-builder ergonomics.

Source/Metadata

  • Title: Agent Frameworks Considered Harmful — Rémi Louf, .txt
  • Transcript words: 7503
  • Duration seconds: 1228
  • Timestamp note: No timestamps or chapter markers were present in the supplied transcript. The transcript also contains a substantial repeated passage and trailing extraction noise.
Full transcript 3649 words · 33 min read
0:12

Hi, everyone. Originally, I thought I was going to give a very technical talk, but I saw I was in the leadership track, which I'm not sure what it means. But I was, okay, I'm going to do half high level and half technical. So it's more a story about what I did in January, because around December, agents became really good. There was a step function, something happened, I think it was Opus 4.6. And that's when I realized, and I work in AI, okay, this thing is really happening. And so I took two weeks out, I took two weeks away. So I'm the CEO of dot text, which is a 15-people company. I just told my CTO, okay, I'm just going to go away for two weeks. And I'm just going to dive into this thing and try to understand what we can get out of it, and how good it is. And so the story is the story of me scratching my own edge for two weeks and trying to figure out how we can actually use agents and what are good primitives to build agents and whether it already exists. This was a really clickbait title, but actually, it turns out to be a good title, even for this talk.

0:17

So what you can see here on the left of the castle is my office. That's true. I do rent an office in that castle. And the small thing with an arrow that you can see is this robot mower, which works unattended all day, every day. It just does its stuff in the background without anyone having to use a remote control or think about it or anything. And I wanted the same thing for my morning because my mornings are always the same thing. The first couple hours are browsing market news, review of linear, could be Jira, my CRM. And also, I like to walk for about an hour in the morning. And then the next hour is spent trying to process the really long voice note that was recorded while walking. And all I wanted was my morning briefing with my coffee. And that's what we've been told for a couple of years, what the future would be. But then, when you really start working with it, even if you're not coding, all you get is a TUI today. So it's amazing, you can code, you can actually, I started doing things that were not coding in it, they're great for this, agents are great for this. But it's the equivalent of having a robot, like a tractor mower, that you still have to stay on, even if it's driving by itself, right? It's very frustrating because you have to, it can do many things, but you still have to be on.

0:21

And so, of course, the labs didn't stop there. And they came up with apps, which I call SSH with vibes. That's great. But in this situation, when that came up, I was, well, that's awesome. I don't have to use a terminal like SSH on my phone anymore. Codex is great. However, I noticed I just started, and I was on my walk, and I was just instructing the agent to do things while I was walking. And so I wasn't thinking very clearly anymore. I just started running agents on my phone during my morning walk. And this is not great, because this is the equivalent of, you're midway. It's not the tractor that you have to stay on, it can actually do something without you being right next to it. But you still have this remote control that you have to change your trajectory with every now and then. That's useful. It's absurd when you think about it. And actually, when you look at people on their phone all the time, just playing, this is absurd, and it's clearly transitional. Surely it's not going to stop there.

0:30

And so I did a very dumb thing as a CEO, which is I started coding, don't tell my board. And I started to build the dumbest thing that could possibly work. And of course, it became a crazy rabbit hole. The repo is there if you want to take a look at it, the code is not amazing. But it works. So the first thing is that I started using frameworks. They're great frameworks, and I'm not going to name any frameworks because they're all good in their own way. And they will have flaws in their own way, which is fine.

0:37

But I spent all my time actually editing the prompt within the code. And I was, this is actually not very useful. So I'm like everyone here, I hate YAML like the next guy. But I still found that this was actually a lot easier to start implementing agents without code. You can version it, you can diff it, you can review it in the PR. And it's just so easy, you can just write your file, you drop it in a folder, and then it just magically appears once you have the runtime, and it just magically works. And then I needed my market watch to run every morning while I'm walking in the fields. And for that, we have things that have been around for a while, which is cron jobs. And schedules specify when the agents need to be run. And also, we'll see it's very important later, they publish events. And markdown and cron, obviously, it's much more complicated than that under the hood. But the interface is this, you don't write code. And that's the whole product so far. And honestly, it mostly worked at this point. I'll come back on mostly later.

0:44

And so this is actually a real picture of one of my morning walks. And so what I do is I record voice notes while I'm walking. But cron jobs, people would use cron jobs for this, because that's what's available in codex today. But they're not ideal, because they cover when, but this is just one point in time, it doesn't cover because this happened. And things that happened in our system, automatically, when you drop the voice note now in the system, it will emit an event, and an agent will react to that event. And it's the same thing when you have a new email, a new entry in the CRM, anything, a new PR that's open, a new PR that's merged, etc. It just reacts to events. It's not just a cron job.

0:48

And that means that the voice note processor agent is not a cron job. But here you have accepts and returns. So it just declares what it accepts, and what it returns as an event. And here it accepts a voice note, transcribes it, turns it into durable notes on the right, and it emits a new event. And for that, it uses structured outputs, we'll come back to this. And now we finally have the future we were promised, because that voice note agent emits voice note process. And then I have my daily brief agent that actually will take the output of the cron job, will take the output of the voice note agents, and will create my daily brief, which is posted as a Slack message. So the slack message.post event is actually a process that subscribes to this and sends a Slack message to me. This is a real thing. It's working, I can show you after on my phone.

0:54

And there are frameworks that are going to sell you the fact that you need graphs for this and code. You do not need graph. In this case, all you need is events. You have no edges to maintain, agents simply subscribe to events. Anyone can come in and edit this. You don't need to know how to code. You just need to know what events exist in the system. Finding and fan-out are free, no code, and it's just drop a file and the topology emerges, whatever the log says happened.

0:58

And then, of course, I tried to run it. So the first version took about, I cheated. I used codex. And it took about a day to write the first thing out of my week. But of course, I tried it and it broke. So these are real examples, actually, the dates are real examples. It's like the first day, daily brief was posted to Slack twice. On Wednesday, one of my voice notes completely vanished. And then towards the end of the week, I played with the prompts all week and the market brief was garbage. But I didn't version my changes. And I couldn't remember actually what I changed in the

1:06

are free, no code, and it's just drop a file and the topology emerges, whatever the log says happened.

1:11

And then, of course, I tried to run it. So the first version took about, I cheated. I cheated. I used, I used Codex. And it took about a day to write the first thing out of my week. But of course, I tried it and it broke. So these are real examples, actually, the dates know about real examples. It's the first day, daily brief was posted to Slack twice. On Wednesday, one of my voice notes completely vanished. And then, towards the end of the week, I played with the prompts all week and the market brief was garbage. But I didn't version, I didn't version my changes. And I couldn't remember actually what I changed in the prompt that made the thing completely useless. Now, if there are distributed or ex-distributed engineers in the room, you probably know this shopping list already. There is nothing new under the sun. And each failure, so each of these failure modes that you found actually led to building one piece of what turned out to be a runtime. So the lost note actually turned into a log. I just wanted everything to be saved forever so that I could go back to it and look into what happened.

1:17

The duplicates, it was because I was not following, it did several attempts and I was not following them. I didn't have a proper queue. I wasn't counting the attempts, etc., etc. And then, probably the most interesting part is the last prompt, I got into a really deep rabbit hole in there. And I just ended up building a content address system for this. A content address system, you can think Git, you can think of Nix and any other build system. And that was, and I didn't do this because I wanted to design a runtime. I mean, by that point, I still just wanted my agents to work. And I also liked the distraction. And I just paid off debt as it appeared, like errors as they appeared. I hope my board won't see this talk. So the log. The log is the system's memory. Nothing is lost and everything is observed. You only have one append on the events table. On the left, it's a real common line, like command, zeta events, and you get all the events. They are causally linked as well. You know which event triggered which event, which happens to be super useful when you're debugging. And even with three, four agents, you start having major debugging headaches. So that was super, super helpful.

1:25

And everything is queryable, which again, for debugging. The second thing is, okay, we have a log, so we can trace back things, etc. But it's still really hard to know what went into the, like what went to the model, what prompt was sent to the model again. Because what you see when you're using Codex, it's a lie. You have a live chat session with the model. And so you tend to think that, oh, that's what the model saw. And that's exactly so I can understand what happened. The truth is, that's not exactly what the model saw. There are many reasons for that. One is, compaction obviously is a big thing. But also, there are just quirks. Also, OpenAI doesn't share, or Anthropic for that matter, don't share the thinking with you, the thinking traces. So you have no idea. I mean, you kind of have an idea of what went in, but not completely either. And so you need something different. You need something different. And that was the big rabbit hole, which is trying to find a way to build a system where you can trace back to what the model saw internally. And so what I did was built, nothing new. This is basically how build system works. So you have different parts for a prompt. You have your system prompt. You have a description of your first skill, of a second skill. Then you have the description of your tools. You have your user message, which is the question of the model. Each one of those is stored and addressed and stored somewhere as an identifier, which is a hash. And so when we build a prompt, instead of building a piece of, before rendering the text, we actually represent the prompt as a list of these hashes. And so what that means is that down the line, when I have a model answer, which by the way is also stored in the same way, we can trace back to the prompt very easily. And then from that prompt, we can know exactly what went into the model's context, which actually matters a lot. I mean, it matters a lot for debugging, but it also matters, it makes compaction a lot easier. You're just manipulating a graph, right? You're not manipulating strings. It's just a lot easier. And it makes KV cache management a lot easier as well, indirectly. But I think that when, I guess probably the main advantage when you use that skill is really auditability. It's like, you can know exactly what happened with that agent and why it returned what it returned. And so I'm just going to go pretty quickly over this. What you get once you have this graph is you get diffs. You can say, okay, what changed between these two runs? Which components changed? Was it just my message? Did I give the model a different skill? Did I give it a different tool? So you can just, yeah, you can just run this function and it will show you the difference between the runs. So here you have three components that were identical. There's one which is the user message changed. And then you had all these other messages that were actually, that were continuing, a continuation of a single session. Then you have another thing for free, which is replace. Replace turned out to be really useful for me because after a while, when I saw the cost ramp up, the thing when you have observability is that you do realize that costs increase very quickly. I wanted to try with open source models. And so I wanted to rebuild all requests to eval and see if I got the same thing out, if I got something satisfactory, if I need to change anything. And turns out that once you have this content addressing system, you can rebuild the request from the graph and you can just replay it exactly the same. And you can resend, you can use a different model. You can use a different request if you want, you can, you can change it. And so, yeah, you get actually a lot of things for free, you need to implement the thing.

1:32

And so this is kind of different from what you find, what I found when I started doing this, it might be different today because it was a couple of months ago, is that out there, you had a lot of libraries. So it's just frameworks and frameworks just call code. Your agents live inside their abstractions. And I don't like analogies with operating systems. Okay, everyone has used that analogy. But okay, let's say a kernel runs processes, and your agent kind of is a process. It doesn't matter what it does, actually. But the system can schedule it, it's built to isolate it, it can isolate it, and journals it with the log. And the agent definition, so the markdown, is userland, you don't need to use it with that system if you don't want to. Actually, you have a frontend that doesn't use this markdown format at all. And okay, here's a very important point. And that's kind of a takeaway. And it's also what justifies me working on this, because disclaimer: structured outputs is our specialty, and we've been working on this for three years. And it just ended up being a big dogfooding project. And the reason why I did this at the beginning was not because I absolutely wanted to use our software, I necessarily wanted to fork llama.cpp to other software, etc. It's just because Anthropic was terrible at structured outputs, and so 20% of my events were wrong and were rejected by the system. So that's why I ended up doing this. And the goal, the job of the kernel, is actually to make bad actions impossible, not just unlikely. And so you have these two boundaries between agents and the external world. The first one is typed tool calls, the two tool calls, you don't want to call tools that don't exist, etc., etc. And also the boundary of other agents, which is typed events. And this is non-negotiable, I found. You can get a lot of errors just from this. I wrote a really long blog post about this. If you follow the QR code, you'll find it.

1:39

fork Lama CPP to other software, etc. It's just because Enfropic was terrible at structured outputs, and so 20% of my events were wrong and were rejected by the system. So that's why I ended up doing this. And the goal, the job of the kernel is actually to make bad actions impossible, not just unlikely. And so you have these two boundaries between agents and the external world. The first one is typed tool calls, the tool calls. You don't want to call tools that don't exist, etc., etc. And also the boundary of other agents, which is typed events. And this is non-negotiable. I found you can get a lot of errors just from this. I wrote a really long blog post about this. If you follow the QR code, you'll find it. And the result of that is I actually deployed it within the company after I built this. And now, today, after a month of deploying it, we have 20 agents on the left that are not just contributed by technical people, by the way, which is what Markdown gives you. And then on the right is, we deploy, it's called the intranet, there's the briefs, there's a bunch of things, as you can see.

1:46

As a conclusion, a few lessons. The first one is that well-executed background agents are really magical. They feel like this robot more than I had at the beginning. I really just sit down when I come back and have this morning brief that is probably even better than what I would have had just doing it manually. And it just appears in my inbox every day and processes my random thoughts. The difficulties that you meet doing this kind of thing, it's just good old engineering problems. There's really nothing new under the sun when it comes to orchestrating these things. It's just good old software orchestration. Open source models are there. They're good enough. I replaced so I don't have any third-party APIs anymore. Now I just use open source models. And even on my laptop, I use a local model. So it's good enough for what I do with it. I'm quoting, I don't know, but for what I do with this, it's good enough. The infra category is definitely unsettled. I tried a few things before I started building myself. And I would advise that today, definitely start building before you buy. So if you're a small company, if you're a tech CEO, it's an advantage because you can just do this with tasking engineers to do this and get them off track. But I would definitely try to build before I buy, just to know exactly what I need and the limitations of what exists. Also, I will say that to people building frameworks for this, please eat your own dog food. Sometimes it's pretty clear that people are building agent orchestration frameworks, etc., but not eating their own dog food. So please do. And the other thing is, I'm really glad I took these two weeks off to play with the field because that completely changed it. That changed the trajectory of the company. I know we're an AI company, we should be in it, etc. But business is such that you're always thinking about the next thing, the next thing, the next thing. And it's the same everywhere. But what I'm urging you to do is to stop and actually immerse yourself in this and try to see how useful it can be for your company. So you can steal the code. It's not a product that we sell, and we don't intend to sell this. You can read our blog as well. So I haven't explained this yet, but I will publish something about it. And thank you for your attention.

1:53

That yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay

1:58

market news, review of like linear, could be Jira, my CRM. And also, I like to walk for about an hour in the morning. And then the next hour is spent trying to process the really long voice note that, you know, was recorded while walking. And you know, all I wanted was my morning briefing with my coffee. And that's kind of what we've been told for a couple of years, like what the future would be. But then when you really start working with it, even if you're not coding, all you get is a TUI today. So it's amazing, you can code, you can actually, you know, I started doing things that were not coding in it, they're great for this, agents are

2:43

great for this. But it's kind of the equivalent of having a robot like a tractor mower that you still have to stay on even if it's driving by itself, right? It's kind of very frustrating because you have to, it can do many things, but you still have to be on. And so of course, the labs didn't stop there. And they came up with apps, which I call basically SSH with vibes. That's great. But in this situation, when that came up, I was like, well, that's awesome. I don't have to use like a term like SSH on my phone anymore. Codex is great. However, I noticed I just started, you know, and I was on my walk, and I was just instructing the agent to do

3:23

things while I was walking. And so I wasn't thinking, you know, very clearly anymore. I just started running agents on my phone during my morning walk. And this is not great, because this is the equivalent of this is you're kind of midway, you know, it's not the tractor that you have to stay on, it can actually do something without you being right next to it. But you still have this remote control that you know, you kind of have to change your trajectory every now and then. That's useful. It's kind of absurd when you think about it. And actually, when you look at people, like on their phone

3:52

all the time, just playing this is kind of absurd, and it's clearly transitional, like surely we're not, it's not it's not going to stop there. And so I did a very dumb thing as a CEO, which is I started coding, don't tell my board. And I started to build the dumbest thing that could possibly work. And of course, it became a really a crazy rabbit hole. The repo is there if you want to take a look at it, the code is not amazing. But it works. So the first thing is that, you know, I started using frameworks. I mean, they're great frameworks, and I'm going to name any frameworks because they're

4:32

all good in their own way. And they will have flows in their own way, which is fine. But I spent all my time actually editing the prompt within the code. And I was like, this is actually not very useful. So I'm like everyone here, I hate YAML, like the next guy. But I still found that this was actually a lot easier to start implementing agents without code. You can version it, you can diff it, you can review in the PR. But it's just, and it's just so easy, you can just, you know, write your file, you drop it in a folder, and then it just magically appears once you have the runtime,

5:07

and it just magically works. And, you know, then I needed like my market watch to run every morning while I'm, you know, while I'm walking in the fields. And for that, we have things that, you know, have been around for a while, which is cron jobs. And schedules specify, you know, when the agents need to be run. And also, we'll see it's very important later, they publish, they publish events. And, you know, markdown and cron, obviously, you know, it's much more complicated than that under the hood. But the interface is this, you don't write code. And that's the whole product so far. And honestly,

5:49

just mostly worked at this point. I'll come back on mostly later. And so this is actually a real picture of my one of my morning walks. And so what I do is I record voice notes while I'm walking. But cron, you know, cron jobs, I mean, people would use cron jobs for this, because that's what's available in codex today. But they're not ideal, because they cover when, but this is just one point in time, it doesn't cover because this happened. And, you know, things that happened in our system, like, automatically, when you drop the voice note now in the system, it will emit an event, and an agent will react to that event. And it's the same thing when you have a new email,

6:33

a new entry in the CRM, I mean, anything, a new PR that's open, a new PR that's merged, etc, just reacts to events. It's not just a cron job. And that means that, you know, agents of the voice note processor, it's just, you know, not a cron job. But here you have accepts and returns. So it just declare what it accepts, and what it returns as an event. And here it accepts a voice note, transcribes it, turn it into durable notes on the right, and it emits a new event. And for that, it uses structured outputs, we'll come back to this. And, you know, now we finally have the future we're promised, because

7:11

that voice note agent emits voice note process. And then I have my daily brief agent that actually will take the output of the cron job, will take the output of the voice note agents, and we create my daily brief, which is posted as a Slack message. So the slack message.post event is actually, is actually like a process actually subscribes to this and emits and sends a Slack message to me. It's actually this is a real, this is a real thing. It's working, I can show you after on my phone. And, you know, there are frameworks that are going to sell you the fact that you need graphs for this

7:52

and code. You do not need graph. In this case, all you need is events. You have no edges to maintain, agents simply subscribe to events. Anyone can come in and edit this. You don't need to, yeah, you don't need to know how to code. You just need to know what events exist in the system. Finding and found out are free, no code, and it's just drop a file and the topology emerges, whatever the log says happened. And, you know, then of course, I tried to run it. So the first version took about, I mean, you know, I cheated. I cheated. I used, I used codex. And it took about like a day to write like the first thing

8:34

out of my week. But of course, I tried it and it broke. So these are real examples, actually, the dates know about real examples. It's like the first day, daily brief was posted to Slack twice. On Wednesday, one of my voice notes completely vanished. And then, you know, towards the end of the week, I kind of like played with the prompts all week and the market brief was garbage. But I didn't version, I didn't version my changes. And I couldn't remember actually what I changed in the prompt that made the thing completely useless. Now, if there are distributed or ex-distributed engineers

9:13

in the room, you probably know this shopping list already. There is nothing new under the sun. And, you know, each failure, so each of these failure modes that you found actually led to building one piece of what turned out to be a runtime. So the last note actually turned into a log. I just wanted everything to be saved forever so that I could go back to it and look into what happened. The duplicates, it was because I was not following, you know, it did several attempts and I was not following them. I didn't have a proper queue. I wasn't counting the attempts, etc., etc. And then,

9:55

probably the most interesting part is the last prompt, I got into a really deep rabbit hole in there. And I just ended up building a content, like a content address system for this. A content address system you can think of Git, you can think of NICS and any other build system. And, you know, that was and I didn't do this because I wanted to design a runtime. I mean, by that point, I still just wanted my agents to work. And I also like the distraction. And I just paid off debt as it appeared, like errors as they appeared. I hope my board won't see this talk. So the log. The log is the system's memory.

10:38

Nothing is lost and everything is observed. You can, you know, you only have one append on the events table. On the left, it's a real common line, like command, zeta events, and you get all the events. They are codally linked as well. Like, you know which event triggered which event, which happens to be super useful when you're debugging. And, you know, even with three, four agents, you start having like major debugging headaches. So that was super, super helpful.

11:12

And everything is queryable, which again, for debugging. The second thing is, you know, okay, we have a log, so we can trace back things, etc. But still really hard to know what went into the, like what went to the model, what prompt was sent to the model again. Because what you see when you're using codecs, it's kind of a lie. Like you kind of have like a live chat session with the model. And so you tend to think that, oh, that's what the model saw. And, you know, that's exactly so I can understand what happened. The truth is, that's not exactly what the model saw. There are many reasons for that.

11:50

One is, I mean, compaction obviously is a big thing. But also, you know, there are just quirks. Also, you know, OpenAI doesn't share, or Anthropik for that matter, don't share the thinking with you, the thinking traces. So you have no idea. I mean, kind of have an idea of what went in, but not completely either. And so you need something different. You need something different. And that was the big rabbit hole, which is trying to find a way to build a system where you can trace back to what the model saw internally. And so what I did was basically built, I mean, nothing new. This is basically how build system works. So you have different parts for a prompt. You have your

12:29

system prompt. You have a description of your first skill, of a second skill. Then you have the description of your tools. You have your user message, which is the question of the model. Each one of those is stored and addressed and, you know, stored somewhere as a identifier, which is a hash. And so when we build a prompt, instead of building a piece of, I mean, before rendering the text, we actually represent the prompt as a list of these hashes. And so what that means is that down the line, when I have a model answer, which by the way is also stored in the same way, we can trace back to

13:07

the prompt very easily. And then from that prompt, we can know exactly what went into the model's context, which actually matters a lot. I mean, it matters a lot for debugging, but it also matters, I mean, it makes compaction a lot easier. You're just manipulating a graph, right? You're not manipulating strings. It's just a lot easier. And it makes KV cache management a lot easier as well, indirectly. But I think that when, you know, I guess probably the main advantage that's when you use that skill is really auditability. It's like, you can know exactly what happened with that agent and why it

13:45

returned what it returned. And so, you know, I'm just going to go pretty quickly over this. What you get once you have this graph is you get diffs. Like you can say, okay, what changed between these two runs? Like which components changed? Was it just my message? Did I like give the model a different skill? Did I give it a different tool? So you can just, yeah, you can just run this function and it will show you, you know, the difference between the runs. So here you have, you know, three components that were identical. There's one which is, you know, the user message changed. And then you had all these other

14:23

messages that were actually, you know, that were continuing, it's continuation of a single session. Then you have another thing for free, which is replace. Replace turned out to be really useful for me because after a while, I mean, when I saw the cost ramp up, like the thing when you have observability is that you do realize that costs increase very quickly. I wanted to try with open source models. And so I wanted to rebuild all requests for to eval and see if I got the same thing out, if I got something satisfactory, if I need to change anything. And turns out that once you have, you know, this

15:01

content addressing system, you can rebuild the request from the graph and you can just replay it exactly the same. And you can, you know, resend, you can use a different model. You can use a different request if you want, you can, you can change it. And so, yeah, you get actually a lot of things, I mean, for free, you need to implement the thing. And so this is kind of different from what you find, I mean, what I found when I started doing this, it might be different today because it was a couple of months ago, is that out there, you had a lot of libraries. So it's just frameworks and frameworks just call code. Your agents leave inside their

15:38

abstractions. And I don't like analogies with, you know, operating system. Okay, everyone has used that analogy. But okay, let's say a kernel like runs processes, and your agent kind of is a process, it doesn't matter what it does, actually. But the system can schedule it, it's built to isolate it, it can isolate it, and journals it with the log. And the agent definition, so the markdown is userland, like you don't need to use it with that system if you don't want to. Actually, you have a frontend that doesn't use this markdown, this markdown format at all. And okay, here's a very important point. And,

16:13

you know, that's kind of a takeaway. And it's also what justifies me working on this, because disclaimer structured outputs is our specialty, and we've been working on this for three years. And it just ended up being a big dog fooding project. And the reason why I did this at the beginning was not because I absolutely wanted to use our software, I necessarily want to, you know, fork Lama CPP to other software, etc. It's just because Enfropic was terrible at structured outputs, and so like 20% of my events were wrong and were rejected by the system. So that's why I ended up doing

16:45

this. And the goal, you know, the job of the kernel is actually to make bad actions impossible, not just unlikely. And so you have these two boundaries with the between agents and the external world. The first one is type tool calls, the two tool calls, you don't want to, you know, you don't want to call tools that don't exist, etc, etc. And also the boundary of other agents, which is type events. And this is non negotiable, I found like you can get a lot of errors just from this. I wrote a really long blog post about this. It's if you follow the QR code, you'll find it.

17:19

And yeah, and the result of that is actually deployed it within the company after I built this. And now today, after a month of deploying it, we have 20 agents on the left, that are not just contributed by technical people, by the way, which is kind of what Markdown gives you. And then on the right is, you know, we deploy, it's called the intranet, there's the briefs, there's a bunch of, I mean, there's a bunch of things, as you can, as you can see. Kind of like a few, you know, as a conclusion, a few lessons. The first one is that well executed background agents are really magical. They feel like this, you know, robot more than I had at the

17:58

beginning is I really just sit down when I come back, and have this morning brief, that is probably even better than what I would have had just doing it manually. And it just appears in my inbox every day, and processes my, you know, random thoughts. The difficulties that you meet doing this kind of thing, it's just good old engineering problems. I mean, there's really nothing new under the sun when it comes to orchestrating these things. It's just good old software orchestration. Open source models are there. They're good enough. I replaced so I don't have any third party APIs anymore. Now I just use open

18:33

source models. And even on my laptop, I use a local model. So it's good enough for what I do with it. I'm quoting, I don't know, but for what I do with this, it's good enough. The infra category is definitely unsettled. I tried a few things before I started building myself. And I would advise that today, like definitely start building before you buy. So if you're a small company, if you're a tech CEO, it's kind of an advantage because you can just do this with, you know, tasking engineers to do this and get them off track. But I would definitely try to build before I buy just to know exactly what I need.

19:10

And you know, the limitations of what exists. Also, I will say that to people building frameworks for this is please eat your own dog food. Sometimes it's pretty clear that people are building, you know, agent orchestration frameworks, etc, but not eating in their own dog food. So please do. And the other thing is, I'm really glad I took this two weeks off to play with the field because that completely changed it. I mean, that changed the trajectory of the company. I know we're an AI company, we should be in it, etc. But you know, business is such that you're always thinking about

19:43

the next thing, the next thing, the next thing. And it's the same everywhere. But what I'm urging you to do is to stop and actually immerse yourself in this and try to see how useful it can be for your company. So you can steal the code. It's not a product that we sell, and we don't intend to sell this. You can read our blog as well. So I haven't explained this yet, but I will publish something about it. And thank you for your attention.

20:20

That yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note