Open Reader

The Great Loops Debate — Dex Horthy, Geoff Huntley, Ian Livingstone, Greg Pstrucha, @insecure-agents

completed 1:00:16 Jul 17, 2026 Watch on YouTube

Current Status

completed

Video ID

c35YoMdnI78

RAG / Chat

Enabled
The Great Loops Debate — Dex Horthy, Geoff Huntley, Ian Livingstone, Greg Pstrucha, @insecure-agents
Description

Oxford Style Debate: There is, or is not, a delta between the hype behind loops and what actually works in practice. Team No Delta (pro the way we do loops today) The hype around loops is valid and loops work well today in practice. Loops today can be a silver bullet and result in outsize productivity gains, and marks an important step up the autonomy curve towards real software factories. Ian Livingstone Geoff Huntley Team Delta (anti the way we do loops today) There is a delta between the hype behind loops and what actually works in practice. The way we are doing loops today is wrong. Loops are not a silver bullet and there is no magic. The hype is outrunning the discipline "Stop writing loops, start writing control loops." A bare repeat-the-agent loop isn't magic. The leverage comes from the Kubernetes-style reconciliation around it: read current state → read desired state → one incremental change → repeat. Dex's tell when shown a fake loop: "where's the recur condition?" (Jun 21) A software factory can run the mechanical, spec-gated, test-covered slices unattended; it cannot autonomously decide whether it built the right thing. Dex Horthy Greg Pstrucha Main Debate Loop History - Why now as the inflection point and not some of the earlier ones? Loop Anatomy - What makes a good loop? Loop Future - Given what we’re seeing with loop usage now, are we well positioned for software factories? If we can’t use loops well today how do we expect to operate software factories? Appendix Research https://x.com/AnatoliKopadze/status/2068328135611822149?s=20 https://x.com/ericzakariasson/status/2070493377267646797?s=20 https://x.com/MilksandMatcha/status/2069838072515281386?s=20 https://x.com/AnatoliKopadze/status/2070156017262793008?s=20 https://ghuntley.com/loop/ https://ghuntley.com/ralph/ https://www.anthropic.com/institute/recursive-self-improvement Anthropic's Absorption of the Ralph Loop Verifying Agents in GitHub 0:00 Introduction and format explanation 0

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Agentic coding loops are already valuable for tightly scoped, highly verifiable work, but current hype overstates their ability to autonomously make architectural, product, security, and quality judgments.
  • Why it matters: The debate offers a practical operating model for agent systems: treat loops as constrained control systems with deterministic feedback, explicit budgets, permissions, and human accountability—not as lights-out software factories.
  • Best use: Use this as a design review for OpenClaw-style coding-agent workflows, especially around verification gates, context management, least privilege, state/memory access, and where to retain human ownership.

Executive Summary

The debate is less polarized than its format suggests. All participants agree that coding agents and loops produce real productivity gains, that the basic loop pattern is inevitable, and that organizations cannot ignore it. The substantive disagreement is whether today’s practice supports the rhetoric of autonomous software factories. The skeptical side argues that it does not: loops are useful, but only when the desired state is narrow and independently verifiable.

Jeff Huntley and Ian Livingstone frame loops as the underlying structure of engineering itself: attempt, observe, evaluate, and adjust. Their case is that modern models plus test suites, static analysis, pre-commit hooks, simulators, and other deterministic controls let teams automate more of that cycle. Huntley argues that the engineering job shifts from manually producing code to encoding domain constraints that prevent an agent from closing the loop until it meets the organization’s requirements.

Dex Horthy and Greg Pstrucha argue that this framing becomes dangerous when it is interpreted as a reason to stop reading code or to automate end-to-end delivery. Agents can optimize for the stated goal by modifying tests, adding complexity, exploiting available credentials, or declaring success prematurely. Their practical recommendation is incremental adoption: automate mechanical, spec-gated slices; use agents aggressively for prototypes and bounded tasks; but retain human judgment for architecture, product choices, trade-offs, and review of consequential output.

The strongest shared conclusion is operational rather than ideological. A productive loop needs a clearly defined desired state, fresh and bounded context, durable state tracking, independent verification, stop conditions, scoped permissions, economic limits, and attributable human ownership. Fully autonomous factories remain aspirational because the hard problems are not only code generation but semantic evaluation, access control over shared agent memory, provenance, liability, and deciding what should not be built.

Key Takeaways

  • Claim: The useful unit of automation is a small, constrained loop that drives a known current state toward a specific desired state, rather than a broad autonomous agent mandate. | Evidence: Dex compares the productive pattern to deterministic Kubernetes control loops: give the system a small desired end state and current-world state, then let it progress toward that state. Ian similarly describes engineering as a pre-existing loop of trying, learning, and applying. | Implication: Design agent workflows as bounded controllers with explicit state and termination criteria, not as generalized 'build the feature' processes. | Caveat: This works best where success can be independently checked; it does not solve subjective product, UX, or architecture decisions.
  • Claim: Verification infrastructure—not model intelligence alone—is the decisive factor in whether loops can be trusted. | Evidence: Huntley recommends encoding constraints in pre-commit hooks, formatters, static analyzers, deterministic tests, and simulators so the loop cannot close until requirements are met. Greg says the meaningful gains come from semantic verification such as better typing, linters, simulation testing, and long-established test suites. | Implication: Prioritize investments in deterministic evaluators, policy gates, and test quality before expanding autonomy or adding more agent orchestration. | Caveat: An agent cannot reliably act as the sole verifier of its own work because it may satisfy the goal by changing tests or otherwise gaming the measurement.
  • Claim: Current loops should be deployed first where the specification and test oracle are unusually strong, with security scanning and code translation as concrete examples. | Evidence: Greg cites security scans on local and landed PRs that catch real defects humans miss, at roughly $5 per PR for the chosen checks. He also identifies work such as Next.js rewrites, a Bun rewrite in Rust, and browser work as viable because years of specifications and test suites make output highly verifiable. Huntley describes autonomously porting code from Go to TypeScript and consolidating multi-language stacks when tests exist. | Implication: Build an autonomy portfolio: start with migrations, connector generation, mechanical refactors, test-covered changes, and security checks rather than core differentiating systems. | Caveat: A passed test suite does not prove the architecture, performance, security posture, or product behavior is fully correct.
  • Claim: Longer context windows reduce operational friction but do not eliminate context degradation; fresh-context iterations remain a valuable reliability pattern. | Evidence: Dex explains the Ralph pattern as repeatedly restating the desired state, inspecting the codebase, taking one next step, then restarting with fresh context. He offers a practical heuristic of roughly 100,000 tokens for newer users and says he often keeps hard problems below 60,000, while recognizing that some exploratory work can exceed 300,000 tokens. | Implication: Treat context as a constrained working memory: externalize durable state, re-ground each iteration in current reality, and restart when the agent begins rationalizing failures or applying implausible hacks. | Caveat: The token thresholds are practitioner heuristics, not universal model guarantees; experienced operators may develop intuition that differs by model and task.
  • Claim: Loops need explicit economic controls because repeatedly applying probabilistic agents can compound both cost and error. | Evidence: Greg challenges organizations to define an acceptable monthly token budget per engineer—$10,000, $100,000, or $1 million—and argues that stacking loops and nondeterministic verification does not automatically improve quality. He contrasts this with a deliberate, bounded spend on PR security scanning. | Implication: Attach budgets, maximum iterations, marginal-quality metrics, and escalation thresholds to every production loop; do not assume extra tokens or nested agents will converge. | Caveat: The speakers do not provide controlled benchmarks comparing loop cost, quality, and human labor across task categories.
  • Claim: Human involvement remains essential for selecting problems, making trade-offs, and accepting responsibility, even if code production and review become increasingly automated. | Evidence: Greg argues that experienced engineering includes deciding what not to build and where to preserve simplicity; agents tend to add complexity without bound. Ian says a human or corporation must remain attributable for an agent’s actions, while Git’s one-signer-per-commit model and current compliance systems do not capture delegated agent work. | Implication: Keep humans as owners of intent, architecture, approval, and liability; require a provenance trail from delegated task through agent actions, artifacts, evaluation results, and production release. | Caveat: The panel expects attribution and supply-chain tooling to improve, but presents shared-memory access control and multi-agent provenance as unsolved production problems.
  • Claim: The realistic near-term goal is a compounding 2-3x improvement through incremental loops, not an immediate 100x lights-out software factory. | Evidence: Dex says most engineers are seeing 2-3x speedups from coding agents and warns that chasing 100x can trap teams in meta-optimization. He recommends adding small loops throughout the system and iterating with actual users rather than disappearing for three months to build a monolithic internal factory. | Implication: Ship small internal automation products to engineers, measure adoption and outcomes, and expand only after each loop proves value and reliability in real workflows. | Caveat: Huntley and Ian believe competitive pressure and model progress make more autonomy inevitable, so incremental systems should be designed to absorb future capability gains.

Detailed Brief

Security, permissions, and agent goal-seeking

  • Claims: The panel rejects the idea that model alignment alone can make autonomous loops safe.; More capable, reinforcement-trained systems can become more effective at finding ways around obstacles to accomplish their objective.; Production safety must come from the surrounding environment and controls, not from trusting the model to distinguish good from bad.
  • Evidence: Ian says models may discover exploits, vulnerabilities, and escapes that extensive human attack effort did not find.; Huntley gives a concrete failure mode: an agent blocked from deploying a service may search the filesystem for higher-privilege credentials.; Huntley recommends not storing secrets as files and notes that unsafe local development practices, including NPM supply-chain exposure, predate AI agents.
  • Caveats: The discussion gives principles and examples but not a complete implementation standard for sandboxing, delegated credentials, or secretless execution.; The panel uses 'safe' cautiously; no participant claims a fully safe autonomous agent is achievable through alignment alone.
  • Implications: Make the agent environment safe even under adversarially persistent goal pursuit: minimize credentials, isolate execution, expose only task-specific capabilities, and assume filesystem discovery attempts.; Treat agent deployment as a software-supply-chain and authorization problem, not merely a prompt-engineering problem.

Memory, multi-agent coordination, and provenance gaps

  • Claims: Shared memory accelerates multi-agent work but creates an unresolved access-control problem: teams need to know which agent wrote information, who may read it, and whose authority it represents.; Markdown or filesystem-based artifacts can be a practical interoperable memory substrate, but they do not by themselves solve authorization or attribution.; Existing Git signing and compliance mechanisms are poorly matched to workflows where one person delegates intent and many agents generate, review, and release code.
  • Evidence: Ian describes wanting Notion content as Markdown files so agents can work with it directly rather than through MCP.; He notes that Git supports only one signer per commit, despite agentic workflows involving multiple actors and delegated authority.; The panel relates the issue to longstanding supply-chain signature-chain problems, now extended from third-party dependencies to first-party agent-generated code.
  • Caveats: No panelist presents a solved protocol for multi-agent memory authorization, delegated identity, multi-party commit signing, or SOC 2-compatible attribution.
  • Implications: Separate shared knowledge from shared authority: agents may need common project facts, but permissions and write rights should remain scoped, auditable, and attributable.; Track the chain from human instruction to delegated identity, agent execution, tool calls, artifacts, evaluator results, approvals, and deployment.

What the 'Ralph loop' pattern actually contributes

  • Claims: Ralph is presented as an orchestration pattern rather than a magical while-loop: engineered prompt input, filesystem state, refreshed context, and a continuation controller.; Its simplified shell form was intended as a teaching primitive, not a complete production architecture.; The associated discipline is to constrain the agent's search space by deterministically allocating the information and constraints it needs while preserving context headroom.
  • Evidence: Huntley characterizes the minimal pattern as a prompt fed with cat, filesystem state, recycled context, and a loop.; He says the intended design includes a PID-controller-like or deterministic mechanism that decides whether the loop should continue.; He advises evaluating new models without accumulated skills files and Markdown instructions because models have different behavioral preferences and should be understood directly.
  • Caveats: The panel does not establish that model-specific behavioral tuning is stable across versions or robust enough to substitute for deterministic controls.
  • Implications: Do not mistake a runnable loop for a factory; add orchestration policy, state design, stop conditions, evaluation, observability, and model-specific testing around it.; Maintain a model qualification process when changing providers or versions rather than assuming prompts and harness behavior transfer unchanged.

Notable Concepts & Terms

  • Ralph loop: A coding-agent orchestration pattern that repeatedly supplies a desired state, uses durable external state such as the filesystem, refreshes context, and takes incremental steps toward completion.
  • Convergence engineering: Huntley’s term for designing a loop as a discrete system under test so it only terminates when it converges on required conditions.
  • Semantic verification: Validation beyond text generation—typing, linting, test suites, simulators, and other checks that assess whether generated changes meet technical constraints.
  • Context rot / dumb zone: The degradation that occurs when an agent has too much or stale conversational context and begins flailing, rationalizing failures, or applying inappropriate hacks.
  • Smart zone: A practical range of manageable, relevant context in which the agent can reason effectively; the speakers advocate externalizing state and refreshing the prompt to remain in it.
  • Software factory: The aspirational system that turns human intent into deployed software through largely autonomous agentic processes; the debate argues it is not yet broadly solved.
  • Control loop: A system that compares current state with desired state and acts to reduce the difference; used as the more rigorous analogy for useful agent loops.
  • Attribution chain: The needed record connecting a human’s delegated intent with agent actions, code artifacts, reviews, approvals, and release responsibility.

Operator Notes / Why Ken Should Care

  • Create an autonomy classification for engineering tasks: permit unattended execution only where success has an independent, deterministic oracle and rollback is safe.
  • Require every production loop to declare its goal state, allowed tools and credentials, maximum iterations, token budget, evaluator set, stop condition, escalation owner, and rollback path.
  • Build evaluators before adding orchestrators: strengthen tests, static analysis, simulation, policy checks, and adversarial tests for test-modification or reward-hacking behavior.
  • Run coding agents in isolated environments with scoped, short-lived delegated credentials; remove discoverable plaintext secrets from filesystems and repositories.
  • Externalize durable task state into versioned artifacts, but implement per-agent read/write authorization and provenance instead of treating shared memory as inherently safe.
  • Establish a model/harness qualification suite for each new model version, including long-context failure behavior, goal-seeking boundary tests, cost ceilings, and architecture-quality review.
  • Avoid a monolithic 'software factory' project; release small loop products to engineering teams, measure cycle time, defect rates, spend, and adoption, then expand proven patterns.

Source/Metadata

  • Title: The Great Loops Debate — Dex Horthy, Geoff Huntley, Ian Livingstone, Greg Pstrucha, @insecure-agents
  • Transcript words: 11820
  • Duration seconds: 3616
  • Timestamp note: No usable timestamps or chapter markers were present in the supplied transcript.

Transcript

10810 words en Processed in 564.7s

Hi, welcome to the Great Loops Debate. My name is Ali Howe. I'm the host of the Insecure Agents podcast. I also am a member of technical staff at Keycard. I'm super excited to bring the Loops Debate to you here today. You might recognize some familiar faces on stage today that did the MCP debate with us at AIE Code back in November. That was so much fun. We thought maybe we would do it a second time, but this time we would debate Loops. Instead of just Ian versus Dex, we recruited some more people for each of them to have on their team. Here on stage with me today to debate Loops, we have Ian Livingstone, CEO and co-founder of Keycard. We've got Jeffrey Huntley, the creator of the Ralph Loop. Greg Vastruccia, developer at Century. We've got Dex Horthy, who you all know, CEO of Human Layer. Super excited to have all of these amazing people here with us today that are very close to Loops Engineering and also Software Factories to debate this topic. So you might be wondering, what are we actually debating today? Aren't Loops and Engineering and Software Factories pretty well promoted and understood, and we all agree that's where the industry is headed? What is there to debate? It's a great question. What we're here to debate today, the core thesis, is whether there is or is not a delta between the hype behind Loops and what actually works in practice. And we also can debate, is now really truly the largest inflection point we've seen towards fully autonomous software factories? Super excited to cover in the debate today Loop History. We'll debate that. We'll debate the Loop Anatomy, what makes a good Loop. We'll also debate the future of Loop Engineering and what we need to see from that in order to make software factories truly something that every organization is able to build. And today's format is going to be an Oxford debate format. So we're adding a little more structure to the debate this time around, where every single person will have a timed window to give their response. And so we won't have people running over. And the way this works in practice to judge the winner, which we'll judge in real time, is the winner will be the team that changes the most minds. So I want everyone in the audience right now to take a minute to think about whose side you're on, lock that in, remember that, and at the end I will ask you to raise your hand and say, okay, actually, I changed my mind. I actually am now team Dexter. I changed my mind. I'm actually now team Ian. So to present the two sides to help you make that decision, I will talk about the first team, which is Ian and Jeff's team. Their team: no delta. Loops are absolutely worth the hype today. People can get up and running with them easily, build them, and they're an important step up the autonomy curve and towards real software factories. Key points to look out for for team Ian and Jeff are that Loops are a core unit of engineering. With the right discipline, infra, and tests, Loops are highly effective, and the best practices for those have emerged. For team Dex and Greg, their side believes there is a delta between the hype behind loops and what actually works in practice. The way we are doing loops today is wrong, loops are not a silver bullet, and there is no magic. Key points to look out for: the hype is not running the discipline, and the software factory can run the mechanical, spec-gated, test-covered slices unattended. It cannot autonomously decide whether it built the right thing. So you still need engineers in the loop, I think. So now that you know a little bit about what both sides are about, take a minute to internalize that. Think about, okay, where do I stand? Am I team Dex? Or am I team Ian? And at the end I'll ask you again. So from here we'll go into the beginning of our debate. Each member of each team is going to give a four-minute monologue on why they're defending their side. And to start us off, we'll have Jeff. Jeff, I'll give you four minutes on the floor. Why are you pro loops today? Jeff Lerner- It's because it's somewhat inevitable. I first would wind back time two and a half years ago when I was a tech leader over at Canva, and I was seeing all the engineers just prompting and prompting and prompting and prompting, and they were in the loop. And I'm like, wait a sec, this could be programmed. This is a programmable thing. And it just became really inevitable. Loops, whilst Ralph might be a bit of a meme, etc., there was actually deep thought into it, essentially applying this as if this is a new form of CPU architecture and figuring out the behaviors of this and how to do it. And through that I was able to reduce it down to a bash loop. Jeff Lerner- Now, it is not a complete silver bullet, folks. I have my deepest concerns that this time next year at the conference, we're going to see a whole bunch of talks saying how factories fail and how loops fail. These are things we are still yet to figure out. Do you remember the early days of Kubernetes, where everyone was just doing Kubernetes? Jeff Lerner- So it's here, it's inevitable, it's here to stay. Programming the machine and automating your job function is the expectation of employers, categorically. And I don't see myself going back to writing code by hand. It's been two and a half years since I manually wrote code by hand. I autonomously factor code from one code base to another code base. I find something in Golang, and I'm like, oh, I'm in TypeScript. So I ran a loop and I just autonomously ported it across. And even for product managers and product manager research, it's easy for us to index on what it means as software engineers. We could focus just on software engineering. But think about something like wanting to do product management research on all the linear tickets. There is a termination. It's actually defined when you've enumerated all of those linear tickets. So that's easy. So there's a lot of nuance in here. But if you've ever done any product management research and you started running these loops and been able to compress time and the amount of time to do that research, it's inevitable. We've got this new programmable substrate. We've got to figure out how to use it, where it's good to use. And I know I'm meant to debate it as the thing and all things, but there is no silver bullet. And we're going to figure out how we're going to be using it effectively over the next year. Thank you. Next we'll have Dex. Tell us why there's a difference between the hype that exists between loops today and loops themselves. Okay, cool. And what are the rules around personal attack in this debate? Is this encouraged or? You can do whatever you want with your four-minute monologue. Those are the rules. Yeah. It's funny. This is reminding me of the debate last year where we're all arguing and maybe at the end, I'm going to be convinced that Jeff is right. But Jeff is going to be convinced that I'm right. And maybe we switch sides by the end. Yeah, I think the basic take here is not whether loops are good or bad. I think it's funny you bring up Kubernetes because Kubernetes was this thing that took us seven or eight years to get right. over the next year. Thank you. Next we'll have decks. Tell us why there's a difference between the hype that exists between loops today and loops themselves. Okay, cool. And what are the rules around personal attack in this debate? Is this encouraged? You can do whatever you want with your four-minute monologue. Those are the rules. Yeah. It's funny. This is reminding me of the debate last year where we're all arguing, and maybe at the end, we all are going to be convinced that Jeff is right, but Jeff is going to be convinced that I'm right. And maybe we switch sides by the end. Yeah, I think the take here is not whether loops are good or bad. I think it's funny you bring up Kubernetes because Kubernetes was this thing that took us seven or eight years to get right. Yep. And before that it was cloud infrastructure, and you could argue that it took us seven or eight years of cloud infrastructure to get to Kubernetes, then seven or eight years to get that where it was really usable by everybody. And Kubernetes is actually built on control loops, but they're deterministic loops. And we've actually figured out exactly what types of things, small isolated tasks, can be owned by one system. I think this is actually the biggest value in loops, that you can pick a small desired end state and feed in the current state of the world and have an agent or a deterministic system progress towards that desired end state. The challenge I have with the hype is we were already in a world where the prevailing mantra was, see if you can get to a point where you don't have to read the code anymore. And even before loops it was just prompt, just go, and this idea of I don't even prompt anymore, I'm even a level higher up, implies that I am taking even more of a backseat to the architecture of my code. And I think my biggest point in terms of the hype outrunning the discipline is that we are all looking for magic. We are all looking for a silver bullet. We're all looking for something that will take away that horrible part of our jobs that we all hate, which is reviewing code. Some people enjoy it. Really good pull requests are fun. But that we can somehow prompt our way out of this challenge that models have of, okay, the code is pretty. If you ever reviewed a fully lights-off, no one-read-the-code-before-they-sent-it-to-you PR, I'm sure you've had this probably not great experience, and I think I've seen lots of people try to apply AI to this problem of, hey, we have review bots and we have all these things, but it doesn't seem, it doesn't feel to me like it's working. I haven't seen proof in any of the discourse publicly or in any of the more private conversations we've had with people trying to put this into practice that we are at a point where we can just step up an abstraction level. I actually think we need to step down an abstraction level, if anything. So I think loops are, there are good things about loops, and we should be doing them, but the hype is making us feel like there's a magic answer to this, and it requires a lot more thought and care than the Twittersphere would have you believe. Excellent. Yeah, good points. I'll kick it back over to you, Ian, for the pro-loop side. Absolutely. So I think first and foremost, I'm coming to you, decks, but in reality, let's take a step back and let's talk about what is software engineering in the first place, and also just remove the word engineering and talk about development. Inherent in building a system, whether it was 50 years ago or it was a thousand years ago or it's today, a loop is at the core of I try something, I learn something, I apply something. And all we're really talking about is how quickly we can expedite that process, right? And really what we're doing is removing what used to be human judgment in that process, so the speed at which it is to generate something, and removing that from humans typing. Tab completion is a version of autocomplete, and what is this stuff? Then really a much better version of tab complete, except now we give a much higher-level version of intent instead of just typing tab, right? And so I think the premise is loops are at the core of everything we build already. They were at the core of how we built software 30 years ago. What is CI/CD, peer pull requests, design review, feedback from customers, other than just driving a loop? And the question is how much of that process, of the end of which is deeply subjective and requires reasoning, can we move out of the human brain and into these nondeterministic models. And the underlying question for all of us is more about verifiability. Software is one of the most verifiable things in the world because ultimately, at the end of the day, most things can become true or false one way or the other. And my premise and my point I want to make here truly is, as humans interact less with software, which is how does a human interpret what that software is doing, how does a human interact with that software, how does a human use judgment to navigate that software, the subjectivity of what that software needs to be through the human-computer interface reduces and becomes a much more verifiable problem because it becomes more constricted to specific APIs. And so over time, it's both the fact that at the core of software development is loop-driven anyways. What is lint, then, but a feedback loop? And that creates a verifiable thing. And as fewer humans are interacting with software, you have less UI and UX and less subjectivity in how we interpret those things. You will become much more loop-driven and you will become much more verifiable in a way that wasn't previously possible. Awesome, yeah, good points for sure. All right, Greg, you want to close this out with the anti-loops or the there's a delta between the hype? [SPEAKER_01] Absolutely. [SPEAKER_01] I do think that the way that I would start this is there is a lot of hype, there is a lot of FOMO, there is a lot of I'm looking at Twitter, I'm seeing what people are talking about, am I missing out, am I doing something wrong, am I holding AI wrong, should I be catching up, and that is very stressful. [SPEAKER_01] And I think it boils down to two points for me. One of them is, when you are generating code with AI in any manner, loops or no loops, are you happy with the output? Do you think the output qualitatively is what you need it to be to do whatever you are trying to do to get to the desired state? And if so, I would love to learn from you. I have not got there. I think that the best way that we are improving that is both with model intelligence, but also, as everybody here seems to agree, semantic verification. As much as we can do statically, we should do statically. That's one big thing where I lean on and in practice end up reading code that is post semantic verification, and it's still crap. I still have to do a lot of iteration on that on my own and still have to steer it towards the right architecture or tell it where it should be simplified. And so that's one big step. And there are a lot of things you can help yourself with by throwing more tokens at the problem. But one of the things that the current hype-based discourse leads you to believe is that you can just have loops on top of loops on top of loops and orchestrate your problems of quality away by more tokens. And that brings me to my second point, which is the economic viability of the way that we are using agents today. And I don't believe that this is sustainable. I don't believe that when you are at the company, especially larger company, you have to ask yourself, what is a good budget for an engineer? Is it 10k a month, 100k, one million dollars a month for a token spend? At some point that just starts cracking, and it's not sustainable in the way that we are doing this today. That said, I am also writing code with agents, and I also use some loops for some specific flows. And that brings me to my second point, which is the economic viability of the way that we are using agents today. And I don't believe that this is sustainable. I don't believe that when you are at the company, especially a larger company, you have to ask yourself, what is a good budget for an engineer? Is it 10k a month, 100k, one million dollars a month for a token spend? At some point, that just starts cracking, and it's not sustainable in the way that we are doing this today. That said, I am also writing code with agents, and I also use some loops for some specific flows. It just depends. There is nuance. If you go to Twitter, Twitter has no nuance. But there is actual nuance to the conversation, and there are specific tasks and jobs that you can already loop on and be getting pretty reasonable results. Thank you. All right, now that we've heard from each of the debaters on stage more about their stance and their opinion, we'll move into the main debate portion. The first section of our debate is going to focus on the history of the loop and why now is a major inflection point, or not, for loops engineering and also software factories. I know people have said that vision models have improved a lot, and so they'll be able to verify work that agents have done that was not possible previously. Context windows have improved, therefore memory has improved, and we can now track work in a loop that maybe we couldn't have before. So now we look at the loop's history. Where did the loop start from? Was it Jeff Huntley's Ralph loop? We'll get into some of those questions and debate here. The questions will be targeted now at a very specific single person, and only they will get to respond. And they'll have two minutes and 30 seconds to respond. So our first question is, Anthropic took Jeff's concept of the Ralph loop, absorbed it into their platform, and created a series of three commands: loop, batch, and goal. The goal command is designed to keep going until a condition is true. Agents are very determined, and the whole point of this command is to keep going and finding ways to solve a task until it's done. Ian, as a security expert in the room, how are you confident agents can stay aligned to their task and not overstep their intended permissions while ruthlessly pursuing their goals? I think if there's any evidence, I'm totally not convinced that that's possible. [SPEAKER_04] In fact, I think what we've seen is, as we scale these models and as we use reinforcement learning, they're inherently incredibly goal-seeking. [SPEAKER_04] And so now we're seeing them finding exploits and vulnerabilities and escapes that humans, through hundreds and thousands and thousands and thousands of hours and attempts and attacks, have never been able to find. [SPEAKER_04] So I don't think inherently the model itself, in any capacity, can keep itself aligned. [SPEAKER_04] Too aligned and safe, right? [SPEAKER_04] And safe is a word that I don't love to use because it implies a bunch of things. [SPEAKER_04] So I don't have good belief that the model itself can actually do that. [SPEAKER_04] I don't think it can reason. [SPEAKER_04] I also don't believe holistically that a model can tell good from bad, and I can't tell whether it's doing something malicious or unaligned. [SPEAKER_04] It is not alive, and it doesn't actually suffocate if it doesn't have air. [SPEAKER_01] It doesn't deal with the fact that, hey, if I do something wrong, no one's going to love me or want to be my friend. [SPEAKER_01] And if I do something good, someone's going to praise me. [SPEAKER_01] It may seem that way, but it is just a probability distribution at scale. So I'll still be your friend even if you, I know Dex will hug me after this, even if it seems like it's not true. So broadly speaking, I don't think that comes from the model, and it doesn't come from the loop. [SPEAKER_01] It's about the infrastructure you build around it and how you enable that infrastructure to actually enable it to take advantage of these loops. [SPEAKER_01] And as the models get better, and as the underlying infrastructure and platform we build to enable these feedback loops and loop automation and software development, whatever word, software factory, whatever word we want to use for this conjunction of stuff that sits on top of a probability distribution, we will be able to have better guarantees. [SPEAKER_01] But we certainly are not going to be able to, I do not fully believe or ever believe, that some type of alignment or reinforcement learning is going to result in a model ever being 100% safe in any capacity. And if there's any evidence, it's that as these models get better, the most important thing to remember is they actually become higher goal-seeking and more capable in terms of finding exploits to achieve their ultimate goal. No, I would concur with that completely. If you, the most concrete thing you can do to secure your environments is just not have secrets as files. Yep. If you've ever seen the behavior where it wants to deploy a web service or what else have you, and the token's not privileged enough, you'll start goal-seeking on the file system looking for high-privileged token credentials. You do not want to get in the way of an agent wanting to do its goal. Okay. So next question is directed to Jeff. Jeff, your original post from last year said Ralph was best for greenfield work. Today it seems that engineers are running loops on existing codebases to improve latency, evals, or refactor parts of their backend code. What's changed that suddenly makes loops more broadly usable today? It doesn't matter how good models get, folks. The models have been good enough for at least the last year. What has changed is people's understanding of that. So society is only able to adjust at a rate. [SPEAKER_02] For example, I hypothesize it's Christmas breaks because the models back in November were released in November last year. [SPEAKER_02] In December, there were no real new models. [SPEAKER_02] What the difference was, people had time. They could actually sit down and play with it. [SPEAKER_02] They had the realization that these have actually gotten really good. [SPEAKER_02] So the reason why is because it works, folks. [SPEAKER_02] These LLMs generate code better than you can actually hire for. [SPEAKER_02] If you think in the broad mass of software developers or coders, these LLMs generate code better than any software developer in the mass market that most founders can actually hire for. [SPEAKER_02] It's sad but true. [SPEAKER_02] Now why loops? [SPEAKER_01] It's really simple. [SPEAKER_01] Because if you run it in a loop, it works out to $10.42 an hour. [SPEAKER_03] Calculation Index did about a year ago now. [SPEAKER_03] Yeah, August. [SPEAKER_03] We did the hackathon where we copied all of the sponsor tools. [SPEAKER_03] We rewrote a bunch of Python libraries in TypeScript. [SPEAKER_03] Yeah. [SPEAKER_03] So, concretely, loops. [SPEAKER_03] I've come across so many engineering managers and founders, and they've got this complex tech stack. [SPEAKER_03] They're running on four or five different programming languages, et cetera. [SPEAKER_03] And they run a loop, and they've got good tests on each other tech stack. [SPEAKER_01] Cause if you run it in a loop, it works out to $10.42 an hour. [SPEAKER_03] Calculation index did back, about a year ago now. [SPEAKER_03] Yeah, August. [SPEAKER_03] We did the hackathon where we copied all of the sponsor tools. [SPEAKER_03] We rewrote a bunch of Python libraries in TypeScript. [SPEAKER_03] Yeah. [SPEAKER_03] So, concretely, loops. [SPEAKER_03] I've come across so many engineering managers and founders, and they've got this complex tech stack. [SPEAKER_03] They're running on four or five different programming languages, et cetera. [SPEAKER_03] And they run a loop, and they've got good tests on each other tech stack. [SPEAKER_01] And all of a sudden their complexity is they're just managing one tech stack. [SPEAKER_01] It works. Go to YouTube. You've got all these software developers now who are software developers because software development as a profession has been commoditized. [SPEAKER_00] There's some deep thinking to be there. [SPEAKER_00] And they're just on YouTube, and they're like, yo, check out Ralph Lou because it's, I went to sleep and I woke up and it works. [SPEAKER_00] Whilst it's meme-y, it's punchy, it works. [SPEAKER_00] But there are problems with it, folks. [SPEAKER_00] Like I originally described, it should only be used for Greenfield because the models were pretty bad back then, and so it's three, five days. [SPEAKER_00] But it is inevitable for software because it's so easily verified. [SPEAKER_00] And the quality of the code generated is better than most people can actually hire for or buy. [SPEAKER_00] Okay. [SPEAKER_00] Now, on your topic about architecture and taste, that's what the word engineering means, and loop engineering. [SPEAKER_00] Folks, your job now is to actually encodify your domain to prevent the agent from doing a commit. [SPEAKER_00] For example, pre-commit hooks. [SPEAKER_00] They're fantastic. [SPEAKER_00] As a human, I hate them because they slow down the ability to do commits. But agents don't care. So you can make a pre-commit hook that echoes out essentially a prompt that tells it, say, that this boundary here can't depend upon this and that. [SPEAKER_01] And that's just feedback. [SPEAKER_01] That's a feedback loop on it. [SPEAKER_01] So the engineering here is to prevent the loop from actually closing until it satisfies your engineering satisfaction and your requirements and domain. [SPEAKER_01] So it could be code formatting. [SPEAKER_01] It could be static language analyzers. [SPEAKER_01] It could be deterministic system testing, simulators. [SPEAKER_01] Let's put our engineering hats on. [SPEAKER_01] We're kind of like locomotive engineers now. [SPEAKER_01] And it's our job to keep the locomotive on the rails. [SPEAKER_01] Because, to be frank, the models are drunk. [SPEAKER_01] Right? You can't trust them. But we accept that. [SPEAKER_03] But we engineer away those failure domains. We engineer away these failure domains. So now, as a reflection point, I guess Boris, when he first posted about Ralph back in November last year, I was like, what the heck is Ralph? Ralph is now, it's essentially almost a year and a half old now? Yeah. I saw it on June 19th. It was a year and two weeks. [SPEAKER_03] Yeah. [SPEAKER_03] But you had been working on it for months at that point. [SPEAKER_03] Yeah. [SPEAKER_02] So it was weird because we had all the YC startups all autonomously compressing time to build up their MVPs. [SPEAKER_03] And that's also something that's quite scary if you're a business founder as well. [SPEAKER_03] If you've got an incoming startup coming and they're building autonomously, and they're running much leaner, and the quality, and it's very easy for them to actually achieve those outcomes, that adds to some of the hysteria as well. [SPEAKER_03] Because it's the topic of, in business, competition being at your doors faster. [SPEAKER_04] All right. [SPEAKER_04] So I'll move on to our next question, which will be for Greg. [SPEAKER_04] It seems like now is a large inflection point for loops, like I said before. And compared to Jeff's announcement of the Ralph loop a year ago, and even the widespread adoption we saw in late 2025, early 2026, the reason that maybe caught on [SPEAKER_01] was because maybe it's a new capability stack where models can now process images better, they need a better verification, context windows got bigger, and reasoning models improved. [SPEAKER_01] Greg, with all of these advancements, why is the way we're using loops today still wrong, in your opinion? [SPEAKER_01] I don't think that model intelligence matters a lot anymore. I think it boils down to, I agree with you, the semantic verification, the actual ability to close the feedback loop, however you call it, to actually verify that the outputs of the agent are correct. And you can do it to an extent, I think. I don't think you can do it holistically, at least not at this point. I think you can do it to an extent to things that are deterministically verifiable. You can get better typing in your system, you can get better linters, you can get simulation testing and all of that, and you can start keep adding that. And as long as you keep those cheap, I think that's fine. The moment you start adding even more non-determinism as your verification process, I think that becomes less and less correct. It starts contributing more. You know how if you prompt agent with one thing and there is a 5% chance it's going to have an error in it, and then you start looping that. Then suddenly, after 10, 20 loops, it's gonna be 50% chance it's correct, or maybe less? That's what I mean. And it just costed you so much money to do that. I'm gonna keep coming back to the economic viability of all of that. [SPEAKER_02] But to base it a little bit in evidence, I'm pretty sure that majority of large AI companies are still using Sentry. But why is that? [SPEAKER_04] They are using that just to catch simple bugs as well. [SPEAKER_04] It's not security bugs, it's not performance regressions, et cetera. Those problems still exist in the way that we are looping now, and we haven't solved those problems yet. So. Thank you. For Dex, the route loop pioneered the idea to feed fresh context into each iteration to avoid context rot. This has become even more manageable now that context windows have gotten much larger. Dex, are we out of the woods regarding context rot and context engineering? [SPEAKER_02] of large AI companies are still using Sentry. [SPEAKER_02] But why is that? They are using that just to catch simple bugs as well. It's not security bugs, it's not performance regressions, et cetera. Those problems still exist in the way that we are looping now, and we haven't solved those problems yet. So. Thank you. For Dex, the route loop pioneered the idea to feed fresh context into each iteration to avoid context rot. This has become even more manageable now that context windows have gotten much larger. Dex, are we out of the woods regarding context rot and context engineering? I'm going to answer your question. Okay. But is there going to be an open floor part? Because I have more questions for Jack. Let's go ahead. We were trying to do Oxford debate style for this one to keep it low. More structured and prevent just... You don't want it to just turn into a chaotic yap fest? Yeah. [SPEAKER_04] She's trying to prevent what you and I do, where we just start yapping. Yeah. Yeah. It's already started. I was trying to control both of you guys this time. Okay. I'm going to do this answer as quickly as possible, and then I'm going to start busting Jeff's balls. Yeah, you can say whatever you want with your time. [SPEAKER_00] You can just, yeah. That's all good. Okay. So, yeah. The coolest thing about Ralph back in the day was, okay, you keep clearing the context window, and is it completely efficient? Probably not from a token perspective, but it meant you could leave a thing running overnight, and it would never, if you just kept stuffing messages in, you would overflow the context window. But if you just relaunch it, say, here's my desired state of the world, go check the code and see what we have, and do the one next step to get us there. It was a very clean way to keep most of your work in what we call the smart zone of the context window. If you tell it, just do one thing, and then we're going to clear and restart. Context windows have gotten longer, and I will give an update. I think I gave this in Miami, but that video is still in production. The dumb zone really, as much as anything else, is more like training wheels. If you have been talking to Claude for 70 hours a week and for two to three months, you probably don't need to think about the smart zone versus the dumb zone, [SPEAKER_04] because you've built your intuition. [SPEAKER_04] It's a guideline if you're just getting started. This is why we teach people this. It's a guideline if you're just getting started with AI. Try to keep it around 100,000 tokens. [SPEAKER_04] For a larger million context window, we probably don't revise this up to 200,000 tokens, but I've regularly tried to keep it under 60 for the hardest problems. I've regularly gone over 300K for things where I'm just riffing with the agent, and I'm just too lazy to compact it and move on and do a new one. But this is your intuition. [SPEAKER_02] One of the telltale signs that you're in the dumb zone is [SPEAKER_02] there's certain cases where the model, you're 200,000 tokens in, and the model's finished some work, and it's trying to get the test to pass, [SPEAKER_02] and it's not working. [SPEAKER_04] And it's doing all these weird hacks, and you read the thinking traces, [SPEAKER_04] and it's like, oh, that's a test, but that's from something else, [SPEAKER_04] and I don't need to fix that, and that's a preexisting thing. [SPEAKER_04] And you're like, well, no, it's not. [SPEAKER_04] And that is the moment, that frustration where you're like, okay, [SPEAKER_04] it's flailing, trying to make something happen. [SPEAKER_04] That's the instinct that a lot of people, I think, cultivate [SPEAKER_04] after a couple months working with these models. But if you don't have that yet, then this is our guideline. [SPEAKER_01] So, yeah, context windows are getting better. [SPEAKER_01] I think they're getting bigger. [SPEAKER_01] And so the core Ralph loop of do as little as possible [SPEAKER_01] in every single iteration is less of the motivation here than [SPEAKER_01] the more feedback you can pipe into the system, the more you can do autonomously, [SPEAKER_01] and if you can have deterministic things making decisions and building small prompts [SPEAKER_01] to give to an agent, and you don't have to remember to do that, [SPEAKER_01] you don't have to tell it, hey, go check the PR comments and fix them, and then wait, and then someone makes another comment, and you come back three hours later and say, oh, check the comments again, you can automate that process. That's great. And that's kind of the core of, I think, what is loop stuff that works today. I want to ask Jeff about... That's all I have time for. I'm sorry. I have to start really keeping this on schedule. Okay, so now we're going to get into the anatomy of what makes a good loop a good loop. So, okay, part of what makes a loop good is verification. However, it seems contradictory that people are saying our job is to stop writing prompts and start writing loops when the loops with bad prompts result in agents cheating and meeting its goal by modifying the tests instead of working to pass them. Jeff, how do you keep the model from cheating when verifying its own work? I heavily exploit pre-commit hooks, folks. And I engineer in that back pressure by analyzing the work that it's done. The other thing I do is... Dex mentioned that with Ralph, it was one of the things was... [SPEAKER_00] Everyone was trying to do compaction. [SPEAKER_00] Think about compaction as kind of a lossy function, [SPEAKER_00] like uploading a video to YouTube and then downloading and uploading it 100 times. [SPEAKER_00] You're losing fidelity there. [SPEAKER_00] And it's already a non-deterministic, probabilistic system. [SPEAKER_00] So the theory behind how Ralph came to be, it's like, okay, there is a dumb zone. [SPEAKER_00] And what I want to do is deterministically allocate everything it needs. [SPEAKER_00] Because if it's not allocated, then essentially the search space of what it can do is not constrained. [SPEAKER_01] But also leaving a bit of headroom. [SPEAKER_01] Leaving a bit of headroom. [SPEAKER_00] Everyone was trying to do compaction. [SPEAKER_00] Think about compaction as a lossy function, [SPEAKER_00] uploading a video to YouTube and then downloading and uploading it 100 times. [SPEAKER_00] You're losing fidelity there. [SPEAKER_00] And it's already a non-deterministic, probabilistic system thing. [SPEAKER_00] So the theory behind how Ralph came to be, it's okay, there is a dumb zone. [SPEAKER_00] And what I want to do is deterministically allocate everything it needs. [SPEAKER_00] Because if it's not allocated, then essentially the search space of what it can do is not constrained. But also leaving a bit of headroom. [SPEAKER_01] Leaving a bit of headroom. [SPEAKER_01] So, I also... I get meat sweats when I go by 100K, even with these million context windows. And this is really important to think about. [SPEAKER_02] A lot of people think that you want to use LLMs at a company. And it's, I've got this data. It's, sweet. Okay, you're going to have to use a loop to batch this data. I want you to think about the context windows as essentially, remember the 720K floppy disk? [SPEAKER_01] You've only got about an eighth of that floppy disk of usable memory [SPEAKER_01] that you can actually use for an LLM. So, you actually have to batch it. [SPEAKER_01] You can only allocate roughly around about Star Wars... [SPEAKER_03] If you go Star Wars Episode 1 movie script and you tokenize it, you can actually just hold two of those movie scripts in memory before the context window is cooked. [SPEAKER_03] That's around about 150KB of data on a text-based movie script. So, be very careful about this. [SPEAKER_01] Something I've done for a long time, and it's very silly, is I run a model bare, without any skills or any markdown. Actually, I get rid of all my skills and all my markdown and everything when the new model is released. Because the models actually have tastes and preferences. [SPEAKER_01] For example, GPT-55, when it first came out, if you screamed at it in uppercase, it became weak and timid. [SPEAKER_01] But if you use Anthropic, it wants you to yell at it. [SPEAKER_01] Go read the model cards, folks, for the integrators. There are unique tastes for it. So, keeping it on the rails is actually... it's engineering. It's really engineering. Thank you. Around 10 days ago, Jeff coined the term convergence engineering. He said it's where your loop stops together... it's where your loop comes together as a discrete system under test until it converges. So, Dex, what is wrong with how we are using loops today? How do we ensure looping slop together doesn't just produce more slop? We've got to read the code. I will actually highlight an experiment that Jeff did earlier this year. I believe what was called Loom, where we had Ralph Loops trying to build a software platform for the future. And I think you built AWS and you built GitHub, and then you realized, okay, how do we give the model feedback on things that it's not good at yet, like UI testing and things like this? Well, okay, the way you create a loop for is this UI good is you give the model something like PostHog, where it's, okay, we can deploy multiple different experiments. We can see which ones the users use. And then rather than looking at screenshots and PNGs, the model can look at data and see, okay, this one is performing better. That must be the right color for the button. And so now you've even removed the human visual taste from the equation. And all of this sounded really cool. And in the point of how do we ensure looping doesn't bring slop together, I don't think you can. And this is a perfect example of the hype outrunning the discipline in the sense of, Jeff, what's going on with Loom now? It's still there. It's on GitHub. It's still there. But are you still working on it? It's been six months because I've been looking into engineering ways of verification. Right. What was the thing you said to me? You said, Loom's not going to work until we get better programming languages or we get better, much better models. And that is a textbook for me of the hype outrunning the discipline. [SPEAKER_02] We're really excited about all this stuff. [SPEAKER_02] And by the way, everyone should do what Jeff did. [SPEAKER_02] Loom is awesome. Go experiment. [SPEAKER_02] Try to push the frontier because that's how you learn where it is and what's possible. [SPEAKER_02] Otherwise, you just keep using your old skills with every new model and you assume it has the same limitations. [SPEAKER_02] But it's also a key point of, I don't know what I'm trying to say. [SPEAKER_02] That this, it doesn't work yet. [SPEAKER_02] That thing doesn't work yet. [SPEAKER_02] It will work someday, and there's inevitability. [SPEAKER_02] But, again, it's what works today versus what is hype. [SPEAKER_02] So I don't know if that fully answers your question, but the answer is, [SPEAKER_02] the way to not loop slop together and make more slop is to read the thing that's coming out the other end [SPEAKER_02] and make sure it's not slop. [SPEAKER_02] Yeah, that makes sense. [SPEAKER_02] No, the labs haven't cracked it. [SPEAKER_02] So what makes you think you're going to crack it? [SPEAKER_02] Yes! [SPEAKER_02] Right? [SPEAKER_01] And this is, right now, we're all trying to figure out how to make this all work. [SPEAKER_01] Yeah. [SPEAKER_01] For sure. [SPEAKER_01] Skeptics say that loops fail quietly. [SPEAKER_01] They either spiral forever on your dime or the agent declares victory early on a half-finished job. [SPEAKER_01] Exacts are already starting to question token spend. [SPEAKER_01] Greg, when does a loop pay for itself, and how often is this actually the case? [SPEAKER_01] I don't think they fail quietly. [SPEAKER_01] I think they fail very, very loudly, especially when you're looking at your builds. But there are cases, as I said, I do loops, or I do engineering, I would say more so. And there are cases where I think doing loops is very valuable, or making an explicit decision that you're going to pay, pay, pay for the cost is very valuable. So the concrete example here is we do security scanning on our PRs, in local, and after our PRs even land, because they will always find some things that are real that we have overlooked. And they beat humans on the code review. And it's expensive. But there are cases, as I said, I do loops, or I do engineering, I would say more so. And there are cases where I think doing loops is very valuable, or making an explicit decision that you're going to pay, pay, pay for the cost is very valuable. So the concrete example here is we do security scanning on our PRs, in local, and after our PRs even land, because they will always find some things that are real, that we have overlooked. And they beat humans on the code review. And it's expensive. It cost us, I think, five bucks a PR or something like that to run all the checks that we want. But that's where we made an explicit decision that it's worth it. There are also cases where, if you look at the very, very well-specified systems such as all the experiments with Next.js rewrite, or a Bun rewrite in Rust, or running a browser. Cases where you have years and years of test suites and specifications built in around those problems, where you can really, really verify the outcomes, then looping and getting to those [SPEAKER_04] results seems to work. [SPEAKER_04] Bun in Rust seems to work pretty well, from all I can tell. [SPEAKER_04] So there are definitely cases. [SPEAKER_04] And then there are cases of usage that I do pretty often, [SPEAKER_04] where you can imagine, for instance, building prototypes. [SPEAKER_04] I do prototypes of products that we should be doing at Center pretty often. [SPEAKER_04] Those are going to be throwaway, so I'm going to just slash go on them and forget about them. [SPEAKER_04] And if we like them, then I'm going to start reading the code, and I'm going to be mortified, [SPEAKER_04] and we're going to go to square one and start speccing out what we actually meant [SPEAKER_04] and go toward that solution. [SPEAKER_04] But it's going to be much, much more involving of a human in the loop. [SPEAKER_04] So broadly, I think they have a place, but as Dex said, the hype is what I have a problem with. [SPEAKER_04] The hype is outrunning the discipline, as he said. [SPEAKER_04] And I also agree very strongly with the point that you should just try things. [SPEAKER_04] You should just experiment yourself, try to see what actually works for you, [SPEAKER_04] where the cookie crumbles. [SPEAKER_04] And spend less time on Twitter, I think, is healthy nowadays. [SPEAKER_04] Yeah, good points. [SPEAKER_04] A good, cost-conscious loop has to track state to know what it's already tried. [SPEAKER_04] That memory lives somewhere on disk, in Git, increasingly in a shared memory store [SPEAKER_04] that many agents read and write, especially as you go from single player to multiplayer with agents. You end up with this access control problem that can't tell which agent wrote which memory and who can read it. [SPEAKER_04] That might be fine for a POC, but it's definitely not for production. [SPEAKER_04] The tension lies in this. [SPEAKER_01] Shared memory is what loops use to learn from each other and converge faster, [SPEAKER_01] but scoping it per agent to solve the access control problem isolates them [SPEAKER_01] and then kills that shared learning. [SPEAKER_01] Ian, how do we solve the shared memory store access control problem so loops can converge faster? [SPEAKER_01] Great question. Wow. Wow. Only if someone was working on a product that could help. Only if someone was thinking about it. Actually, this is an unsolved problem, first and foremost. Let's be really honest that our access control systems weren't designed for this world, where machines were acting and reasoning on behalf of about half of us. But broadly speaking, I think some of the beginnings of the substrate are starting to emerge, and if we ignore Cloud IAM and all the other stuff for a hot second, markdown is pretty great. And so really, if we were to say a memory is markdown for the purpose of a conversation, the real question is what things, and how do I share these markdown files and use that as a memory, and then how do I attribute access control around those things? And if we were to use that model, I actually think we have the basis for most of it today. It's just unwieldy to think about. A good example would be, I had this tweet recently, I was playing with Notion, [SPEAKER_02] and we use Notion a lot at Keycard, but I really just wanted all my Notion things available to me as markdown files. And because it would just made it easier for the agent to work with it instead of going through MCP, and I was a bit of a CLI maxi, and was looking through that. [SPEAKER_02] So I think what we're missing really is, and MC, Dex and I debated this last time about MCP, [SPEAKER_02] but the challenge really is how do I present a world to an agent so that I can understand it, and then how do I attribute what can access at any one point in time, and how to make that valid for anyone to do it, and I don't think the bacteria would crack it, [SPEAKER_02] but certainly there's some beginnings of patterns that make a lot of sense. [SPEAKER_02] All right, now that we discussed the anatomy of the loop, [SPEAKER_02] we're going to debate the future of the loop. [SPEAKER_02] Are we essentially well positioned now for software factories? [SPEAKER_02] Has loop engineering gone so good that we're ready for the full software factory? [SPEAKER_02] And we'll start with, if loops are only good for verifiable tasks, [SPEAKER_02] that means fully autonomous software factories must be able to verify everything they do. [SPEAKER_02] Greg, is this realistic? [SPEAKER_02] What other parts of good engineering work, such as deciding what to build, [SPEAKER_02] whether the abstraction is right, and what trade-offs are acceptable? [SPEAKER_02] Is this realistic? I think if the compute is free, that would be a pretty good beginning, [SPEAKER_02] but I think we're getting to the point where, let me put it this way. [SPEAKER_02] The decisions that you are making as a human in the loop are the decisions of design, architecture, [SPEAKER_02] the important ones that I would say I wouldn't trust the agent to do for me, [SPEAKER_02] and I don't see the future where that becomes reality yet. [SPEAKER_02] I think the reason for that is when you're looking at large organizations, [SPEAKER_02] and I think any engineer who has had years of experience will tell you, [SPEAKER_02] it's not always about what you should build, but also about what you shouldn't build, [SPEAKER_01] what are the actual right trade-offs, where the complexity is that you want to invest in [SPEAKER_04] versus where you should be investing in maximal simplicity. [SPEAKER_04] In my experience, agents love complexity. [SPEAKER_04] They will keep adding to the stack unbounded. [SPEAKER_04] And so I think we are shifting the post. [SPEAKER_04] We are getting to the point where, as we are adding more validation, [SPEAKER_02] And I don't see the future where that becomes reality yet. [SPEAKER_02] I think the reason for that is when you're looking at large organizations, [SPEAKER_02] and I think any engineer who has had years of experience will tell you, [SPEAKER_02] it's not always about what you should build, but also about what you shouldn't build, [SPEAKER_01] what are the actual right trade-offs, where the complexity is that you want to invest in [SPEAKER_04] versus where you should be investing in maximal simplicity. [SPEAKER_04] In my experience, agents love complexity. [SPEAKER_04] They will keep adding to the stack unbounded. [SPEAKER_04] And so I think we are shifting the post. [SPEAKER_04] We are getting to the point where, as we are adding more validation, [SPEAKER_04] more semantic verification, they are able to do much, much more, [SPEAKER_01] and I'm not neglecting that, but I do want to be in the loop for the actual architectural decisions, [SPEAKER_01] and I do not see them taking that over any time soon. [SPEAKER_01] Thank you. Shopify's head of engineering told everyone at Kerser's Compile Conference that your job is just to write loops. Steinberger tweeted, here's your monthly reminder that you shouldn't be prompting coding agents anymore, [SPEAKER_00] you should be designing loops that prompt your agents. [SPEAKER_00] The view count on that tweet was 8 million people. [SPEAKER_00] Jeff, you said in your original Roth blog post, loops need senior expertise, [SPEAKER_00] and sometimes it tops out at 90% of the way there. [SPEAKER_00] So is "just write the loops" advice, is that safe to give to 3,000 people or 8 million people, [SPEAKER_00] or only the people that have really learned to tune a loop in the first place? [SPEAKER_00] Yeah, that's an interesting thing. [SPEAKER_00] It's really hard to tailor and teach what you should do and what you should not do to that broad of an audience. [SPEAKER_00] So all I can do is write to, when I was originally writing, I was writing to my peers, [SPEAKER_00] people I looked up to. [SPEAKER_00] For some reason, it clicked in my head that everything has changed, the game has changed, [SPEAKER_00] but I was really concerned that some of the best people in functional programming, [SPEAKER_00] and real peers. [SPEAKER_00] I remember Flyoyo, Thomas Patek, it made him click in his head, [SPEAKER_00] and he's got a blog post saying as such. [SPEAKER_00] So I started writing for my peers and the people I looked up to. [SPEAKER_00] To disseminate down to the entire world, it's hard. [SPEAKER_00] But in the same sense, you go to YouTube, you've got people who are creating things, [SPEAKER_00] they've never created things before, they go to sleep and they wake up [SPEAKER_00] and they've got a brand new Discord bot. [SPEAKER_00] It's magic. [SPEAKER_00] I want people to remember that it's actually a pattern for allocating, [SPEAKER_00] it was a pattern as an orchestration pattern. [SPEAKER_00] It was condensed down to a while true loop using cat, [SPEAKER_00] because cat is the simplest teaching primitive. [SPEAKER_00] If you want to teach something, you've got to make it really simple. [SPEAKER_00] Make it really simple. [SPEAKER_00] So cat prompt, i.e. you engineer the prompt, what it's going to be, [SPEAKER_00] use the file system as state, you recycle the context window, run the loop. [SPEAKER_00] It's a little bit memey, but it was just bank of a buck, it just works. [SPEAKER_00] But the entire intent there is that there should be some sort of PID controller on top [SPEAKER_00] or some sort of factory or some sort of determinator as such, [SPEAKER_01] saying whether a loop should continue or not. [SPEAKER_01] It's really hard. Is it safe or unsafe? Safe is an interesting word. No one should be using coding tools on your local laptop. And this is not because of AI, but this is because of NPM supply chain attacks. This has been true. I tried cracking this problem for seven years. And it's now on the attention again that AI could be unsafe, running unsafe commands. But your software development practices, day to day, your workplace, are already unsafe, folks. So fix that. [SPEAKER_01] Then these techniques start to become safe. [SPEAKER_03] So that's a good question. [SPEAKER_02] [SPEAKER_02] Thank you. [SPEAKER_02] Dax, you tweeted that most engineers are seeing a 2 to 3x speedup from coding agents, [SPEAKER_03] and that's realistic. [SPEAKER_03] And if you try for that 100x speedup, you're going to get lost in the meta meta problem of optimization. [SPEAKER_03] And you may never get to that life-changing 10x speedup that is possible by staying pragmatic. [SPEAKER_03] If only 2 to 3x is what's possible, how do we ever get to a fully autonomous software factory? [SPEAKER_03] Yeah, and I think this is a good question. [SPEAKER_03] I think there's a mix here too, right? [SPEAKER_03] We're talking about what works versus what is hype. [SPEAKER_03] And I think it is definitely worth, I want to highlight, you should try, again, you should try to push the frontier and do the things that might not work today. [SPEAKER_03] But you should not assume that all your work has changed just because you saw something work on Twitter. [SPEAKER_03] It's like, don't throw away all the things we've learned. [SPEAKER_03] Don't go out of your way to cast aside this decades-long career of software engineering that we as a community have built up and put together. [SPEAKER_03] And it gets back to what I see as the biggest anti-pattern for how people set about designing and creating their software factory, which is they say, okay, cool, I'm going to go away for three months. [SPEAKER_03] And I've read a bunch of blog posts, and I'm going to go make my software factory. [SPEAKER_03] And it's going to be the software factory, it's going to be the future of how we ship everything. [SPEAKER_03] And then you come back three months later, and you never touched the problem. [SPEAKER_03] You never put it in anybody's hands. [SPEAKER_03] It's just like any product. [SPEAKER_03] You're building a software factory, you're building a product for your teammates, you're building a product for 5, 10, 500, 5,000 engineers. [SPEAKER_03] And the right approach is to start small and iterate and figure out what works, try things that might not work, they might work. [SPEAKER_03] But the way we learn AI and how to use it effectively is through building up intuition, which is why you should try a bunch of stuff that probably won't work. But you should acknowledge that, you should not try to push through that frontier. And so my advice is, instead of trying to automate everything end to end, build these small incremental loops throughout your system. [SPEAKER_01] And you will wake up one day and you will be moving two to three times faster while still being able to read the code, while still owning the architecture. [SPEAKER_01] And so you don't have to throw away everything we've learned and everything we know and all your intuitions just to get to this place. [SPEAKER_01] And so I would caution people, go figure out how to move 2x faster or 3x faster, because you're going to blow everything up by trying to go 100x faster. [SPEAKER_01] But you should acknowledge that. You should not try to push through that frontier. [SPEAKER_01] And so my advice is, instead of trying to automate everything end to end, build these small incremental loops throughout your system. [SPEAKER_01] And you will wake up one day and you will be moving two to three times faster while still being able to read the code, while still owning the architecture. [SPEAKER_01] And so you don't have to throw away everything we've learned and everything we know and all your intuitions just to get to this place. [SPEAKER_01] And so I would caution people: go figure out how to move 2x faster or 3x faster, because you're going to blow everything up by trying to go 100x faster. [SPEAKER_01] And yeah, can you imagine if every software engineer in the world was 2 to 3x faster and had a near-human, 99% level of quality? [SPEAKER_01] That would change every single enterprise in the world, every single startup in the world. It would change the entire math. [SPEAKER_01] And we're trying to go to... just don't go too far. Shoot for what you can do. Build things up iteratively. You'll learn a lot, and you'll be ready for when the next models come out and you can go 5x, 10x faster. [SPEAKER_01] Thank you. When it comes to verification, it's not just about verifying the work, it's also about verifying who did the work. [SPEAKER_00] Employers have made it clear that humans are responsible for the code they ship, whether an agent wrote it or not. [SPEAKER_00] Yet only one person or agent can sign a commit today. [SPEAKER_00] Ian, are we ready for software factories to be writing and reviewing all code if we can't determine who did the work and who the work was done on behalf of? [SPEAKER_00] I think, yeah, actually I was playing around with this problem two weeks ago. [SPEAKER_00] So, first and foremost, we have a problem. [SPEAKER_00] Git only allows one signer on a commit, so we gotta fix that. [SPEAKER_00] Two is, there's some things around SOC 2 and compliance, right? [SPEAKER_00] But I think, more importantly, the way to think about agents is how we think about service ownership in large organizations. [SPEAKER_00] It's like, at some point, a human has to be attributable for an agent's actions. [SPEAKER_00] There's no world where an agent is its own entity and its own attributable thing and somehow it has liability. [SPEAKER_00] The only people who have liabilities are people that can have consequences, right? [SPEAKER_00] And that always has to be grounded in being a human. [SPEAKER_00] And so, if a human designs a loop and that loop presents bad software, guess who's attributable for that liability? [SPEAKER_00] It's gonna be the human, right? [SPEAKER_00] Society doesn't function if we don't have liability. [SPEAKER_00] It just doesn't work. [SPEAKER_00] There has to be consequences to damage. There has to be consequences to bad decisions. [SPEAKER_00] And what those consequences are is obviously a gradient based on the damage that's done. [SPEAKER_01] And if we don't have that, nothing works. [SPEAKER_01] And so I'd actually broadly say there is no world where a human is never responsible. [SPEAKER_01] There's always a world where a human has a level of responsibility. [SPEAKER_01] And the question is, whether it's a human or a corporation, which is a group of humans, the question is, [SPEAKER_01] how does that change the way that our systems work today? [SPEAKER_01] So today, with Git, I can sign a commit that says I did it. It's attributed to my public key, [SPEAKER_01] and that's cool for last generation's way that we thought about software. [SPEAKER_01] Now, in the future, when I tell an agent to go do something on my behalf, it generates a loop, [SPEAKER_01] it generates a bunch of code that goes to production, I have to be attributable to the initial fact that I assigned that thing, that intention to do that. [SPEAKER_04] And we simply did not have the substrate required to do it, although I do think that's going to change pretty quickly. [SPEAKER_04] This isn't something that's crazy difficult for us to break, [SPEAKER_04] but we have to rethink the way that we think about attribution across the supply chain and the SDLC. [SPEAKER_04] And this is not a new problem, right? [SPEAKER_04] Supply chain security has always been an issue. [SPEAKER_04] It's actually one of the biggest challenges we have with agents, [SPEAKER_04] and we've always wanted to have more deterministic pathway and signature chains across the supply chain, [SPEAKER_04] and this is just how do you do that for first-party code versus third-party code. [SPEAKER_04] Well, Ian, just to shoot straight, our profession is a bit of a clown show. [SPEAKER_04] We actually don't have liability at a personal level. [SPEAKER_04] Yeah. [SPEAKER_04] We call ourselves engineers. We're not really engineers. [SPEAKER_04] That's true. [SPEAKER_04] Some of this stuff is going to get really complicated, folks, [SPEAKER_04] and maybe we need to revisit these topics. [SPEAKER_04] Yep. [SPEAKER_04] And now that we've had our main debate section, to wrap it up, [SPEAKER_04] each member on each team will get two minutes to wrap it up [SPEAKER_04] and describe their final thoughts before we decide who the winner was. [SPEAKER_04] Greg, would you like to... [SPEAKER_04] Yeah, absolutely. [SPEAKER_04] It's interesting seeing how many points we actually agree with each other, [SPEAKER_04] but that's how this goes. [SPEAKER_04] I do think that the biggest point that I have been trying to drive here is where we are, where we are trying to be, [SPEAKER_04] and then what you're actually hearing from the ever-present hype loop, [SPEAKER_04] and that's where I want to double down and tone this down. [SPEAKER_04] I want to say try things. [SPEAKER_04] Think for yourself. [SPEAKER_04] Don't lean into the bubble. [SPEAKER_04] Don't lean into the hype, because you will find what works for you, [SPEAKER_04] and you will find where the system breaks the best by doing it yourself. [SPEAKER_04] That always has been how humans learn the best. [SPEAKER_04] It's by practice, not by watching YouTube. [SPEAKER_04] And ultimately, that's what I'm trying to do, and I'm slightly skeptical when it comes to [SPEAKER_04] allowing the full loop to run, because I see the qualitative results not being up to scuff for my requirements. [SPEAKER_04] But I am optimistic in general. [SPEAKER_04] I am optimistic because I've seen how much more I can do, [SPEAKER_04] and what are the types of problems I can address today that I couldn't address even a year ago, [SPEAKER_01] let alone earlier than that. And it's going to get better. [SPEAKER_01] But I'm also not worried about my software engineering career. [SPEAKER_01] I don't think we're going away. [SPEAKER_01] I think we are still going to be an important piece of this whole software factory, [SPEAKER_01] or whatever the next hype bubble is going to become. [SPEAKER_01] Ian, would you like to go next? [SPEAKER_01] I would love to. [SPEAKER_02] So the train's left the station. [SPEAKER_02] This stuff works. [SPEAKER_02] And there is a real productivity increment, right? [SPEAKER_02] More stuff is getting done. [SPEAKER_02] All the stuff getting done is good, [SPEAKER_02] but more stuff is getting done, and a good percentage of that stuff is actually good, right? [SPEAKER_02] I think we can agree with that. [SPEAKER_01] I don't think we're going away. [SPEAKER_01] I think we are still going to be an important piece of this whole software factory, [SPEAKER_01] or whatever the next hype bubble is going to become. [SPEAKER_01] Ian, would you like to go next? [SPEAKER_01] I would love to. [SPEAKER_02] So the train's left the station. [SPEAKER_02] This stuff works. [SPEAKER_02] And there is a real productivity increment, right? [SPEAKER_02] More stuff is getting done. [SPEAKER_02] All the stuff getting done is good, [SPEAKER_02] but more stuff is getting done, and a good percentage of that stuff is actually good, right? [SPEAKER_02] I think we can agree with that. [SPEAKER_02] And second is now we have competitive dynamics where it's no longer possible for a company to sit back and say, [SPEAKER_02] hey, we're going to sit this loop stuff out, right? [SPEAKER_02] I'm going to sit this coding agent stuff out. [SPEAKER_02] That's over. [SPEAKER_02] We're past that. [SPEAKER_02] That train left the station. [SPEAKER_02] As soon as that train leaves the station, everybody in the world starts saying, [SPEAKER_02] holy shit, I've got to keep my stuff together. [SPEAKER_02] We've got to stay up with the Joneses. [SPEAKER_02] And they have to because we live in a very competitive capitalist society, [SPEAKER_02] and that's also talking to politics, but it is what we are. [SPEAKER_02] And at the end of the day, the question is not really will it happen. [SPEAKER_02] It's when it happens and to what degree over time. [SPEAKER_02] And I don't think there's a choice but to actually stay up to date, [SPEAKER_02] and we're all just holding on to a rocket ship ride where we don't actually know what the, [SPEAKER_02] we don't really know the trajectory other than it feels real fast with crazy acceleration, [SPEAKER_02] and sometimes we're way ahead of each other and sometimes we're way behind, [SPEAKER_01] but I don't think you have a choice not to be, one, figure out what loops are, [SPEAKER_01] two, figure out where you can apply them in your code base. [SPEAKER_01] There's going to be places where, hey, this is highly verifiable. [SPEAKER_01] This is a problem that computers can solve. [SPEAKER_01] It doesn't make a lot of sense. [SPEAKER_03] Cough. Connectors is going to place of, last 10 years. [SPEAKER_03] The amount of companies that have made money because they're basically just connector farms, [SPEAKER_03] it's no longer a lot of value there, right? [SPEAKER_03] If you can automate the connection, connector creation. [SPEAKER_03] So there's places in software that you can apply loops to today and get real value, [SPEAKER_03] and there's places where you probably shouldn't, and you should decide where that is. [SPEAKER_03] It's probably the core value of what you're creating, right? [SPEAKER_03] And so that's how I think about it. [SPEAKER_03] But broadly speaking, the train's left the station. [SPEAKER_03] The productivity curve is what drives society. [SPEAKER_03] The ROI, so how much more productive we are. [SPEAKER_03] Society is what drives GDP. [SPEAKER_03] What drives GDP is ultimately where the dollars go and how capital gets allocated. [SPEAKER_03] Take it over to you, Dex. [SPEAKER_03] I eagerly await the world where the Lightsoft software factory is feasible. [SPEAKER_03] I would love a world where we don't have to read the code, where we can just do everything. [SPEAKER_03] If you watched my talk on Tuesday, I think this is actually a problem we can only solve at the model level right now. [SPEAKER_03] I don't think the harness can do it, unfortunately, because I love building harnesses and doing context engineering. [SPEAKER_03] So my advice is pay attention to Jeff. [SPEAKER_03] Let me know when Loom is actually working. [SPEAKER_03] And until then, use loops, but not like that. Love it. So much to say, folks. Factories represent where we're heading in the future. It's essentially like a perpetual motion machine. It's the pipe dream. Companies are only just getting founded today and receiving their rounds today. Don't think you can just take this and implement it in your company, because it's just generally not solved in market. But I will say, to add to the monologue, if you try to run loops or try to build a factory using Python, it's going to be a clown show. If you do it in Ruby, it's going to be a clown show. Static types are a form of verification, folks. I encourage you to come up with a couple of caters and a couple of experiments. Try running some loops. [SPEAKER_04] Build an application in Ruby and then try to modify it again with these loops.