AI Engineer

The Agent Behind the Curtain: Building the Oz Cloud Agent Platform — Safia Abdalla, Warp

1691 summary words 8 min summary Watch video

Start with the signal

8 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: A cloud-agent platform should hide infrastructure complexity while exposing consistent, composable primitives for environments, harnesses, orchestration, artifacts, and human-in-the-loop quality gates.
  • Why it matters: It provides a credible control-plane pattern for moving coding agents from local single-prompt tools into durable, multi-agent systems that can operate across team infrastructure and software-development workflows.
  • Best use: Use it as an architectural and operating-model reference for agent platforms, especially for multi-harness execution, API-first orchestration, and agent-assisted issue/PR intake.

Executive Summary

Safia Abdalla explains Warp's evolution from a developer terminal into a cloud-agent platform through one design principle: good developer tools meet people in their existing workflows, grow with the work's complexity, and absorb platform complexity before it reaches users. The transition to cloud agents creates real infrastructure demands—isolated execution, long-running work, security, deployment constraints, and state—but Warp's objective is to present these as simple, stable primitives rather than operational burden.

The proposed platform stack consists of agent environments, multiple interchangeable coding harnesses, multi-agent orchestration, shared artifact and state management, and APIs for every major primitive. The important design constraint is not merely allowing choice between Claude Code, Codex, Warp's harness, or custom tools; it is normalizing their interaction with conversation state, generated files, issues, PRs, and environments so flexibility does not fracture the experience.

Warp's open-source workflow is the strongest operational example. Following its open-source launch, Warp says GitHub stars rose from roughly 20,000 to more than 60,000, alongside thousands of PRs and hundreds of contributors. Agents automatically triage incoming issues, research repository context, request missing information, help draft specifications and implementation work, and iteratively review PRs before human reviewers are notified. The intended effect is to reserve scarce human review capacity for higher-signal changes while improving agents from the accumulated repository feedback.

Abdalla rejects the impersonal framing of a 'software factory' in favor of a workshop model: a repeatable but adaptable system where humans remain closely connected to the work and outputs. For such a system to let more technical and non-technical people turn intent into software, it needs event-driven automation, observability, iterative improvement loops, verification points, and token-aware economics—not just code-generation agents.

Key Takeaways

  • Claim: The core platform responsibility is to absorb cloud and infrastructure complexity rather than expose it as a user problem. | Evidence: Abdalla contrasts local coding agents with cloud workloads that must handle isolated sandboxes, longer-running tasks, infrastructure management, security concerns, and team deployment practices; she states that platforms should take on complexity before it reaches the user. | Implication: A viable agent control plane should offer managed execution for fast adoption plus bring-your-own/self-hosted execution for enterprise fit, while presenting both through a unified behavioral model. | Caveat: Abstracting complexity does not mean imposing one hosting model: serious teams often need agents to run on infrastructure they bring and manage themselves.
  • Claim: Multi-harness support only becomes useful when the platform standardizes state, artifacts, and guardrails across harnesses. | Evidence: Warp intends to support its own harness alongside Claude Code, Codex, and custom harnesses, but requires each to use platform-native capabilities such as stored and rehydrated conversation state and structured outputs including PRs, issues, and generated files. | Implication: Treat the harness/model layer as replaceable execution infrastructure; preserve a canonical artifact, state, permissions, and lifecycle layer above it. | Caveat: Simply adding many harnesses risks a fragmented product where each agent has materially different workflows and capabilities.
  • Claim: Real software tasks require explicit multi-agent decomposition rather than a single all-purpose coding agent. | Evidence: The talk describes a typical sequence of one agent researching and planning, another implementing, and a third validating; different agents may use different harnesses and models to create a more adversarial, robust workflow. | Implication: Design agent systems around typed roles and checkpoints—research, implementation, validation/review—with an orchestrator mediating messages and work status, rather than relying solely on monolithic prompts.
  • Claim: API-first exposure of agent primitives is what turns an agent product into a platform. | Evidence: Warp exposes APIs for creating agents and sub-agents, attaching a sub-agent to a parent through configuration, managing environments and compute, and working with produced artifacts. Internally, non-engineering employees have used the SDK to build custom Slack bots. | Implication: A control plane should not force adoption through a single UI: stable APIs and SDKs allow teams to embed agent capabilities in Slack, operations tools, research workflows, and their own internal interfaces.
  • Claim: Agents can materially scale open-source intake when they provide structure and context before humans enter the loop. | Evidence: After Warp went open source, it reports growth from about 20,000 to over 60,000 GitHub stars, thousands of PRs, and hundreds of contributors. Its issue agent researches the codebase and repository context, asks submitters clarifying questions, and helps turn vague reports into actionable work. | Implication: High-volume contribution channels are a strong initial deployment target for agents: intake clarification and context assembly can remove recurring human coordination work even before autonomous implementation is trusted. | Caveat: The talk reports Warp's own results and process but provides no quantitative measurement of triage accuracy, review quality, merge rate, latency, or false-positive/false-negative rates.
  • Claim: Agent-managed PR review can serve as a quality filter rather than a replacement for human engineering judgment. | Evidence: Warp says every contributed PR passes through iterative agent-managed review, and human reviewers are not pinged until the agent has approved it; the team uses incoming PR and code examples to improve the reviewing agent over time. | Implication: Place agents upstream of human review as an iterative evidence and quality gate, and instrument the feedback loop from human overrides, defects, and accepted/rejected changes to improve the evaluator. | Caveat: Agent approval is a routing threshold, not evidence that a change is safe or correct; the speaker explicitly retains humans in the process.
  • Claim: The better metaphor for scalable agentic development is a software workshop, not a software factory. | Evidence: Using a potter's workshop as the analogy, Abdalla highlights specialized workstations, prepared inputs, verification and restart points, close observation of work, dozens of apprentices, and hundreds of handcrafted mugs per day. | Implication: Agent systems should be designed as adaptive production systems with human craft, observable execution, event-triggered automation, verification, and continuous refinement—not as a one-shot code-output machine.

Detailed Brief

Design requirements for an adaptive agent workshop

  • Claims: Automation should respond to real-world events rather than only user-initiated prompts.; Observability is a product requirement because teams must inspect how work proceeds and refine processes from actual behavior.; The workflow itself must evolve with users' goals and the outputs being produced; it cannot be treated as static infrastructure.; Cost effectiveness is part of reliability: reducing defective outputs without uncontrolled token consumption is a core system objective.
  • Evidence: The potter analogy includes reacting to failures such as a defective handle or a kiln problem, with explicit checks determining where work should restart.; Abdalla cites repetitive engineering toil such as reproducing bugs and monitoring production systems as opportunities for structured, repeatable agent-enabled workflows.; Warp's internal SDK use extends beyond engineering to social-mention analysis and response drafting, product-query assistance, and competitive research.
  • Caveats: The presentation offers design philosophy and examples but not a concrete reference implementation for permissions, credential isolation, audit logs, model routing, recovery semantics, or cost controls.; Non-developer participation is framed as enabled by structure and guardrails, but the talk does not specify the governance model needed to safely authorize production changes.
  • Implications: Evaluate agent platforms as workflow systems: event triggers, inspection, evaluation, rollback/retry points, and economics should be first-class requirements.; The same platform primitives that support software delivery can create a broader internal automation layer, but only if access and output-review policies are defined separately.

Notable Concepts & Terms

  • Cloud agent platform: The cloud-based execution and coordination layer needed when coding-agent work exceeds what can reliably run on an individual developer's laptop.
  • Sandbox / bring-your-own infrastructure: Warp's execution abstraction: offer managed compute for ease of use while allowing teams to run workloads on their own infrastructure for security and deployment compatibility.
  • Harness: The coding-agent interaction/runtime choice, such as Claude Code, Codex, Warp's own harness, or a custom alternative; it should be interchangeable without changing platform semantics.
  • Platform-native experiences: Normalized services shared by all harnesses, including persisted conversation state and structured artifacts such as PRs, issues, and generated files.
  • Prompt-based orchestration: A user asks an orchestrator to delegate a task; it handles sub-agent creation, communication, and status tracking behind a simple prompt interface.
  • Composable primitives: APIs and SDKs for agents, sub-agents, compute, environments, and artifacts that permit customers and internal teams to build interfaces and workflows beyond Warp's UI.
  • Agent-managed review gate: A PR workflow in which an agent iteratively reviews changes before escalating approved work to human reviewers.
  • Software workshop: Abdalla's human-centered alternative to 'software factory': a scalable, repeatable, observable, and adaptable production environment for translating intent into software.

Operator Notes / Why Ken Should Care

  • Define a canonical agent-run record that is independent of model or harness: parent/child relationships, environment, credentials boundary, conversation state, artifacts, evaluation results, cost, and escalation status.
  • Prioritize an upstream issue/PR intake pilot where agents must gather missing reproduction details, retrieve repository context, draft a structured specification, and escalate only when defined evidence thresholds are met.
  • Require each agent workflow to have explicit role separation and verification gates; do not let the implementing agent be the only evaluator for its output.
  • Instrument the human-versus-agent review loop from day one: measure escalation volume, human overrides, defects after approval, cycle time, and token cost per accepted change.
  • Before extending the platform to non-engineering operators, establish authorization boundaries for data access, external communications, code changes, and production actions.

Source/Metadata

  • Title: The Agent Behind the Curtain: Building the Oz Cloud Agent Platform — Safia Abdalla, Warp
  • Transcript words: 6635
  • Duration seconds: 1251
  • Timestamp note: No timestamps or chapters were present in the supplied transcript. The latter portion is substantially duplicated, so the unique substantive content is shorter than the stated word count.
Full transcript 3520 words · 29 min read
0:12

Hi, welcome to this session. It is mysteriously titled The Agent Behind the Curtain. It's about how the team at Warp build our cloud agent platform. Before I talk about the how, though, I want to talk about the why and share a little bit about my own background. I've spent the past eight years building developer tooling. I started off in open source in the Python and data science space on the Jupyter Notebook core team. I was a maintainer on the Interact project. That work continued in my time at Microsoft, helping build APIs and SDKs for web developers. And now I'm working on bringing AI agents to the cloud at Warp.

0:27

One of the lessons that I've learned in my time building developer tooling is that really good dev tools meet devs where they are and grow with them. A tool that you're going to be using every day should accommodate your workflows, but also be really adaptable as the complexity and the nature of the work changes. And it's important for us to build really great dev tools because dev tools have a compounding effect on the world. If you make a piece of software that helps a developer or a builder do great work, your ability to magnify how much great software exists in the world increases.

0:33

So I think it is super important that these dev tools accommodate people's workflows. And people's workflows are important because they love their preferences. You might have a preference for a shell, a language, a harness, a review process, and something that adapts to your workflows and preferences is going to be more enjoyable to use and more enjoyable to build software with. And it also becomes a part of how you think and build. So it's super important to be attuned to that.

0:39

This notion of tools meeting developers where they are and growing with them comes to light really clearly in the progression of AI tooling in this space. This is the story for Warp specifically, but it's also the story for a lot of developer tools. In the pre-AI era, we had the Warp terminal, which met people in their command-line workflows that they were used to. Then this fantastic thing happened where AI got introduced to developers, and now you had a whole new set of tools that was available to you locally on your machine. A lot of people started to interact with agentic coding patterns in their terminal, in their IDE, in their editor. Then eventually we realized that we had reached the limits of what we could do on our laptops. We wanted agents to do work that was more long-running, that was adaptive to different constraints. And that work needed to happen in the cloud.

0:45

Whenever you send anything to the cloud, you adopt a lot of complexity because running things in the cloud requires us to navigate a much messier stack of infrastructure concerns. And when I say we here, I mean the people building developer tools, because that is the person that I am. And it gets at one of the things that are a core principle in how we think about unlocking capabilities here and building good developer tools: platforms should take on complexity before it reaches the user. A really good experience should not expose any of the leaky complexity that it handles to you.

0:52

When we think about building our cloud agent platform, we try and structure it so that every primitive models this philosophy of hiding complexity from the user so they can focus on the work that matters to them. The first place that this shows up when you're building a cloud agent platform is: where does the agent run if it's not running on a developer's machine? It needs a place to do its work like any developer would. And that place is typically a sandbox. It's an isolated environment in the cloud where agents do this task.

0:59

When we started building these out for our cloud agent platform, our first intuition was to provide self-hosted sandboxes so that developers had a really easy on-ramp for getting into our cloud agent platform. You didn't have to think about where your compute lived. It was just there for you. But the reality is that for teams doing serious work, they're probably managing their own infrastructure. They probably have dev boxes that they need to interact with. And so something that is hosted or managed is usually not sufficient. You really need to be able to run agent workloads on infrastructure that people bring so it adapts to their security concerns, their deployment practices, their workflows, and preferences on their team. And so you add support for not only managed hosting but also self-hosting to the platform. And that is complexity that you abstract away from the user in how the behavior is modeled.

1:06

The next component is a little bit more personal to people, and it's the harness that they want to use. As we talked about earlier, people are really passionate about the tools that shape their workflows. One of those tools is the harness. Who here has a preference for cloud code as a harness locally? Codex. Something else entirely. Right? So much diversity in the room, and we want to meet people where they work. So you want to integrate multi-harness support that not only accommodates preferences but also gives people the ability to use the right tool for the job.

1:12

Flexibility isn't something that you just cram into a platform, because a real risk you run if you cram it in is that it becomes fragmented. Your experience with working with Claude is different from working with Codex versus a custom harness that you might have. And so one of the key properties is making sure that the platform provides structure and guardrails around the harness so that the experience is consistent. For us, this means that harnesses can interact with all the platform-native experiences. So just being able to store conversation state and rehydrate it, being able to interact with the artifacts and outputs that are produced by agents, whether they're PRs, issues, new files that are generated, all of that should be structured the same way.

1:17

Okay, cool. So we gave you a place for your agent to run, and we gave you a choice for what harness you use, including Warp's own harness and any other harnesses that you want to bring. What if one agent isn't enough to do work? That's the reality of most software engineering. I wish that I could just send off one prompt and solve all of the problems that exist in my software, but the reality is that real engineering work rarely fits inside one prompt.

1:23

In a typical workflow, you might need one agent to go research a problem and plan a solution. You might need another agent to implement it, and you might need to bring in a third to validate it. You might want each of these agents to use different harnesses and different models in order to have a real adversarial and robust approach. So we have built-in support for that. You can orchestrate agents across the stack.

1:29

With a lot of agent-based experiences, this orchestration happens via a prompt. So I say slash orchestrate, or I cue the agent via prompting that I wanted to delegate work across multiple sub-agents for a task that I have here. And this orchestrator agent will do all of the messy complexity of interacting with sub-agents, mediating messages between them, and tracking the work that's happening for me behind the scenes with a single prompt. We abstract complexity away from the user by giving them this experience.

1:34

This prompt-based model for interacting with agents and sub-agents is really powerful. An even more interesting one is the notion of interacting with agents and sub-agents via the API. So everything in our surface area is exposed via an API, and I can fire off a request to say that I want to run a sub-agent that is attached to a parent agent via configuration that I provide. And this API is super magical because this is the key component of a platform. It's exposing the primitives in a way that users can build on top of.

1:40

The thing about great APIs and SDKs is people can build on top of them, which means that they're not restricted to your UI or your opinion of how a particular experience should look. This is where composability becomes really powerful. And so we're trying to be intentional about exposing an API for every key component of the stack. So this is APIs for spinning up agents and sub-agents, for managing the environments and compute that these agents are running in, for working with the artifacts that they produce. All of that is exposed in an API that you can build on top of.

1:45

And this ability to build on top of these primitives that are exposed via an API becomes really useful because anyone can build tools that overlap on top of these agentic experiences. An interesting phenomenon that's happened for us internally is we have a bunch of non-engineering teammates at Warp who have been able to use our SDK and API to build custom Slack bots to do a bunch of things. So we have folks in our developer relations team who have actually built out tooling to help us manage all of our social mentions. So as tweets and Reddit posts and things are coming in, we have agents that will pick them up, do some sentiment analysis on them, try and understand what the user wants, and then propose a response that folks on our social media team should use in response to the original tweet or Reddit post or what have you.

1:52

And all of this is enabled by our SDK. And you see a plethora of these types of experiences internally at Warp. We have people who have used them to help answer queries about how our product is working, do competitive research, all sorts of interesting things. And these primitives became a really big deal for us specifically when we decided to go open source. As I mentioned earlier, Warp started off as a terminal, but it grew into an agentic development

2:05

So we have folks in our developer relations team who have actually built out tooling to help us manage all of our social mentions. As tweets and Reddit posts and things are coming in, we have agents that will pick them up, do some sentiment analysis on them, try to understand what the user wants, and then propose a response that folks on our social media team should use in response to the original tweet or Reddit post or what have you.

2:12

And all of this is enabled by our SDK. And you see a plethora of these types of experiences internally at Warp. We have people who have used them to help answer queries about how our product is working, do competitive research, all sorts of interesting things.

2:19

And these primitives became a really big deal for us specifically when we decided to go open source. As I mentioned earlier, Warp started off as a terminal, but it grew into an agentic development environment. And about three months ago, we decided to go open source. This was much anticipated, long awaited, long awaited. It was a huge success for us. The number of GitHub stars that we had, I think, catapulted from around 20,000 to over 60,000. We had thousands of PRs. I'll talk a little bit more about how we've been managing that, and hundreds of contributors who had been longtime users of the platform and were finally getting a chance to build on top of it.

2:25

And when we went open source, we wanted to be really thoughtful about how we could use agents to help us manage the repository. We didn't want this to be the kind of thing where agents are just writing code and firing off PRs. We want them to participate meaningfully in the structure that we use to triage issues that came into the repo, provide context around them, do implementation, do reviews, but still have the space for humans to participate in this loop.

2:31

And we did that. So if you go to the Warp open source repo right now, you'll notice that if you file a new issue with a bug report or a feature request, an agent will kick in and start to triage the issue automatically. It'll do research across the codebase and context in the repo to understand what you're trying to propose. It might ask you questions if it feels like your original query was a little abstract to get more information. And it will do the work that's historically been very hard for open source, which is somebody has a problem or a bug that they want fixed. They don't give you enough details, and it's hard to get to the clarity that you need to get to to drive the work forward. So we can use agents to help us meaningfully in that way.

2:36

They can also help draft initial specifications and work for tasks, do implementation, and provide a review gate. So all PRs that get contributed to Warp go through an agent-managed review process. And it goes through multiple iterations, and we don't actually ping any of the human reviewers on our team until an agent has approved our PRs, which helps manage the workload a lot for the team. So all of those thousands of PRs, the things that humans actually have to manage are only the high-signal, high-quality ones.

2:42

And one of the key principles is that we improve the agent as we get more PRs in the repo and we see more examples of code. One of the things that we believe is that self-improvement loops are a really important way for you to enhance the overall SDLC lifecycle that you're seeing. So we did this, and we had a light bulb moment because it unlocked something huge. We had this structured process that could accommodate a big influx of issues and PRs on the repo. And the agents were there to support anyone in bringing their idea or their bug request, bug feature request, or bug report to Warp and then getting it through to the actual product.

2:48

That key insight of agents providing structure and context was a really big thing for us because it meant that anyone could participate in translating their intent into implementation. And oftentimes, the people who have really interesting intents and goals are the ones who are using software in interesting ways. And it's not always the person that's building it. It's the person who's got domain knowledge in the space.

2:54

And we're lucky because we're developers building a developer tool, and that's a really unique niche to fill in. But most software is developers building tools for non-developers in situations where they don't have domain expertise. If we provide these structures and guardrails, though, people who are non-developers can have the necessary tools to ship serious software because the infrastructure to support them exists.

3:01

This is where things get buzzy. You might have heard this term, the software factory. People talk about it a lot as far as automating how software gets built, providing these systems for doing work, all of these fun things. I want to push back on this term a little bit. I actually hate it because I don't think it gets the point across, and it feels a little, where's the people in this? So I want to tell a story before I share what I think is actually the better word.

3:07

So this is a mug that I have. I bought this mug about two summers ago from a farmer's market, and I stopped by this booth at the farmer's market. And you could just tell the person who had crafted this, the potter, was just someone who's really passionate about their work and what they do. And so he was telling me about all of these interesting details in the mug, the specific curve of the handle and the way he had structured it to accommodate different people's hands.

3:14

He had this specific dimple at the top of the handle where you could rest your thumb because he felt like that was a key ergonomic detail of this mug. He had this glazing at the top so if your cup overflowed, it wouldn't dribble down the sides. The glazing would catch it. So he just spent so much time thinking about the details of this mug and crafting it.

3:19

And then I was talking to him about his workshop, how many potters do you have, how many of these mugs are you making, yada, yada. And he got more animated, even more animated, and he started talking about his workshop setup and how he had set up different stations for different components of the mug. He had talked about how he had a specific process for sourcing clay and preparing ahead of time. He talked about how he actually incorporated verification for different components of the mug. If the dimple wasn't the right size, what would you do? What part of the process would you restart?

3:25

All of this thought that he had put not into the mug itself, but into how the workshop existed to support the creation of the mug, and how he was able to scale this to dozens of apprentices in his shop and hundreds of mugs handcrafted per day, which is pretty impressive.

3:31

And this got me thinking, I love what he did with his workshop. He had this really great idea, and he developed a serious and repeatable system that allowed anyone to take the idea of a perfect mug and turn it into the actual existence of a perfect mug. Some people might think workshops are this quaint thing where it's a workspace for a single individual, but I think the story really shows that there are actually heavy-duty systems for doing work, and that they're malleable, and that they react to signals and how people are interacting with the thing they're building in the space they're building it.

3:39

And it also underscores the really close interaction loop that humans have with the spaces they work in and the outputs that are produced. And I'm synthesizing all of these ideas, and I think this is what really we're driving at when we talk about building software factories. We want to give more builders, and the definition of who a builder is is expanding to non-developers, serious systems for turning their ideas into code. And we as individuals who've been building dev tooling or have been in software engineering for a long time have an understanding of what that serious system looks like and what kind of support it needs to give to individuals.

3:43

We break this down into the same techniques that my potter friend had and the same methodologies that I talked about earlier about exposing primitives. We expose things like the ability for these agents to implement automations that react to events in the real world the same way that a human in a workspace might need to react to a real event of a kiln being astray or a handle being broken.

3:49

We need to make these systems observable. My potter friend talked about how he actually watched the way people worked in his space and refined the process over time. That doesn't come for free. Your system has to actually be something that you can inspect and look into. And it has to improve over time. The workspace is not the static component that doesn't change ever. It needs to react to what's going on and modify itself to amend to the goals of the people that are working in it and the product that it's producing.

3:55

And it needs to be cost-effective. You want to reduce the number of broken mugs that come out the other end. You want to reduce the amount of buggy software that comes out of the other end of your bug, of your factory. And you want to do this without compromising on cost, without spinning too many tokens. All of these principles work to achieve a shared goal. And that shared goal is building systems that remove toil and drudgery from our software process.

4:06

talked about how he actually watched the way people worked in his space and refined the process over time. That doesn't come for free. Your system has to actually be something that you can inspect and look into. And it has to improve over time. The workspace is not the static component that doesn't change ever. It needs to react to what's going on and modify itself to amend to the goals of the people that are working in it and the product that it's producing. And it needs to be cost effective. You want to reduce the number of broken months that come out the other end. You want to reduce the amount of buggy software that comes out of the other end of your bug, of your factory. And you want to do this without compromising on cost, without spinning too many tokens. All of these principles work to achieve a shared goal. And that shared goal is building systems that remove toil and drudgery from our software process so that more people have the ability to build. We've seen the way toil and drudgery have manifested. It could be all of the difficulty you might have reproducing a bug, the challenges of monitoring a production system. Those are things that are really hard to do. And we can finally start to think about the structure of how we do them and building systems that allow us to do them repeatably and transferrably to people who are in non-technical roles. If these ideas excite you about how we can build these robust and reliable systems for anyone to ship software, you could stop by the work booth to come talk to me and the crew. We're at UG20. You can also mention me on Twitter. I'm CaptainSafia on all social media, GitHub, Twitter, all of that fun stuff. Or just drop me an email. You can find my email on my personal site. Thanks for coming to this presentation. I hope you learned something interesting about some of the engineering philosophies that are driving the next set of work we do as far as agents, developers, and AI.

4:12

yay yay yay yay ! ! You be able to run agent workloads on infrastructure that people bring so it adapts to their security concerns, their deployment practices, their workflows and preferences on their team. And so you add support for not only managed hosting but also self-hosting to the platform. And that is complexity that you abstract away from the user and how the behavior is modeled. The next kind of component is a little bit more personal to people and it's the harness that they want to use. As we talked about earlier, people are really passionate about the

5:13

tools that shape their workflows. One of those tools is the harness. Who here has a preference for cloud code as a harness locally? Codex. Something else entirely. Right? So much diversity in the room and we want to meet people where they work. So you want to integrate multi-harness support that not only accommodates preferences but also gives people the ability to use the right tool for the job. And flexibility isn't something that you just be like crammed into a platform because a real risk you run if you cram it in is that it becomes fragmented. Your experience with working with Claude is different from working with Codex versus a custom harness that you might have.

5:55

And so one of the key properties is making sure that the platform provides structure and guardrails around the harness so that the experience is consistent. For us this means that harnesses can interact with all the platform native experiences. So just being able to store conversation state and rehydrate it, being able to interact with the artifacts and outputs that are produced by agents whether they're PRs, issues, new files that are generated. All of that should kind of be structured the same way. Okay cool. So we gave you a place for your agent to run and we gave you a choice for what harness you use including warpstone harness and any other

6:35

harnesses that you want to bring. What if one agent isn't enough to do work? That's the reality of most software engineering. I wish that I could just send off one prompt and solve all of the problems that exist in my software but the reality is that real engineering work rarely fits inside one prompt. In a typical workflow you might need one agent to go research a problem and plan a solution. You might need another agent to implement it and you might need to bring in a third to validate it and you might want each of these agents to use different harnesses and different models in order to have a real adversarial and robust approach.

7:14

So we have built-in support for that. You can orchestrate agents across the stack. With a lot of agent-based experiences this orchestration happens via a prompt. So I say slash orchestrate or I cue the agent via prompting that I wanted to delegate work across multiple sub-agents for a task that I have here. And this orchestrator agent will do all of the messy complexity of interacting with sub-agents, mediating messages between them, and tracking the work that's happening for me behind the scenes with a single prompt. We abstract complexity away from the user by giving them this experience. This sort of like

7:53

prompt-based model for interacting with agents and sub-agents is really powerful. An even more interesting one is the notion of interacting with agents and sub-agents via the API. So everything in our surface area is exposed via an API and I can fire off a request to say that I want to run a sub-agent that is attached to a parent agent via configuration that I provide. And this API is super magical because this is the key component of a platform. It's exposing the primitives in a way that users can build on top of. The thing about great APIs and SDKs is people can build on top of them, which means that they're not

8:40

restricted to your UI or your opinion of how a particular experience should look. This is where like composability becomes really powerful. And so we're trying to be intentional about exposing an API for every key component of the stack. So this is APIs for spinning up agents and sub-agents for managing the environments and compute that these agents are running in for working with the artifacts that they produce. All of that is exposed in an API that you can build on top of. And this ability to build on top of these primitives that are exposed via an API becomes really useful because anyone can build tools that overlap on top of these agentic experiences.

9:25

Like interesting phenomena that's happened for us internally is we have a bunch of non-engineering team mates at Warp who have been able to use our SDK and API to build custom Slack bots to do a bunch of things. So we have folks in our developer relations team who have actually built out tooling to help us manage all of our social mentions. So as tweets and Reddit posts and things are coming in, we have agents that will pick them up, do some sentiment analysis on them, try and understand what the user wants, and then propose a response that folks on our social media team should use in response to the original tweet or Reddit post or what have you.

10:05

And all of this is enabled by our SDK. And you see like a plethora of these types of experiences internally at Warp. We have people who have used them to help answer queries about how our product is working, do competitive research, all sorts of interesting things. And these primitives became a really big deal for us specifically when we decided to go open source. As I mentioned earlier, Warp started off as a terminal, but it grew into an agentic development environment. And about three months ago, we decided to go open source. This was like much anticipated, long awaited, long awaited. It was a huge success for us. The number of like GitHub stars that we had,

10:53

I think catapulted from around 20,000 to over 60,000. We had thousands of PRs. I'll talk a little bit more about how we've been managing that and hundreds of contributors who had been long time users of the platform and were finally getting a chance to build on top of it. And when we went open source, we wanted to be really thoughtful about how we could use agents to help us manage the repository. We didn't want this to be the kind of thing where agents are just writing code and firing off PRs. We want them to participate kind of meaningfully in the structure that we use to triage issues that came into the repo,

11:32

provide context around them, do implementation, do reviews, but still have the space for humans to participate in this loop. And we did that. So if you go to the warp open source repo right now, you'll notice that if you file a new issue with a bug report or a feature request, an agent will kick in and start to triage the issue automatically. It'll do research across the code base and context in the repo to understand what you're trying to propose. It might ask you questions if it feels like your original query was a little abstract to get more information. And it will kind of do the work that's historically been very hard for

12:09

open source, which is somebody has a problem or a bug that they want fixed. They don't give you enough details and it's hard to get to the clarity that you need to get to to like drive the work forward. So we can use agents to help us meaningfully in that way. They can also help draft initial specifications and work for tasks, do implementation and provide a review gate. So all PRs that get contributed to warp go through an agent managed review process. And it goes through multiple iterations and we don't actually ping any of the human reviewers on our team until an agent has approved our PRs, which helps manage the

12:45

workload a lot for the team. So all of those thousands of PRs, the things that humans actually have to manage are only the high signal, high quality ones. And one of the key principles is that we improve the agent as we get more PRs in the repo and we see more examples of code. One of the things that we believe is that self-improvement loops are a really important way for you to enhance the overall SDLC life cycle that you're seeing. So we did this and we had a light bulb moment because it unlocked something huge. We had this like structured process that could accommodate a big influx of issues and

13:22

PRs on the repo. And the agents were there to support anyone in bringing their idea or their bug request, bug feature request or bug report to AWARP and then getting it through to the actual product. That key insight of agents providing like structure and context was a really big thing for us because it meant that it could anyone could kind of participate in translating their intent into implementation. And oftentimes the people who have really interesting intents and goals are the ones who are using software in interesting ways. And it's not always the person that's building it. It's the person who's kind of got domain knowledge in the space. And we're lucky because

14:06

we're developers building a developer tool and that's a really unique niche to fill in. But most software is developers building tools for non-developers in situations where they don't have domain expertise. If we provide these structures and guardrails though, people who are non-developers can have the necessary tools to like ship serious software because the infrastructure to support them exists. This is where like things get buzzy. You might have heard this term of the software factory. People talk about it a lot as far as like automating how software gets built, providing these systems

14:44

for doing work, all of these fun things. I kind of want to push back on this term a little bit. I kind of actually hate it because I don't think it gets the point across and it feels a little, where's the people in this? So I want to tell a story before I share what I think is actually the better word. So this is a mug that I have. I bought this mug about two summers ago from a farmer's market and I stopped by this booth at the farmer's market and you could just tell the person who had crafted this, the potter, was just someone who's like really passionate about their work and what they do.

15:22

And so he was telling me about all of these interesting details in the mug, the specific like curve of the handle and the way he had structured it to accommodate different people's hands. He had this like specific dimple at the top of the handle where you could rest your thumb because he felt like that was like a key ergonomic detail of the smug. He had this glazing at the top so if like your cup overflowed it wouldn't dribble down the sides, the glazing would kind of catch it. So he just spent so much time thinking about the details of this mug and crafting it. And then I was

15:52

kind of talking to him about his workshop, like how many potters do you have, how many of these mugs are you making, yada, yada. And he got more, even more animated and he started talking about his workshop setup and how he had set up different stations for different components of the mug. He had talked about how he had a specific process for sourcing clay and preparing ahead of time. He talked about how he actually incorporated verification for different components of the mug. If the dimple wasn't the right size, what would you do? What part of the process would you restart?

16:22

All of this thought that he had put not into the mug itself, but how the workshop existed to support the creation of the mug and how he was able to scale this to dozens of apprentices in his shop and like hundreds of mugs handcrafted per day, which is pretty impressive. And this got me thinking, I love what he did with his workshop. He had this really great idea and he developed a serious and repeatable system that allowed anyone to take the idea of a perfect mug and turn it into the like actual existence of a perfect mug. You know, some people might think like workshops are this quaint thing where it's like a

17:06

workspace for a single individual, but I think the story really shows that there are actually heavy duty systems for doing work and that they're malleable and that they react to signals and how people are interacting with the thing they're building in the space they're building it. And it also underscores the like really close interaction loop that humans have with the spaces they work in and the outputs that are produced. And I'm synthesizing all of these ideas and I think this is what really we're driving at when we talk about building software factories. We want to give

17:37

more builders and the definition of who a builder is is expanding to non-developers serious systems for turning their ideas into code. And we as individuals who've been building dev tooling or have been in software engineering for a long time have an understanding of what that serious system looks like and what kind of support it needs to give into individuals. We break this down into the same techniques that my Potter friend had and the same methodologies that I talked about earlier about exposing primitives. We expose things like the ability for these agents to implement automations that react to events in the

18:17

real world the same way that a human in a workspace might need to react to a real event of you know a kiln being a stray or a handle being broken. We need to make these systems observable. My Potter friend talked about how he actually watched the way people worked in his space and refined the process over time. That doesn't come for free your system has to actually be something that you can inspect and look into. And it has to improve over time the workspace is not the static component that doesn't change ever. It needs to react to what's going on and modify itself to amend to like the goals of the people that are

18:54

working in it and the product that it's producing. And it needs to be cost effective. You want to reduce the number of broken months that come out the other end. You want to reduce the amount of buggy software that comes out of the other end of your bug of your factory. And you want to do this without compromising on cost without spinning too many tokens. All of these principles work to achieve a shared goal goal. And that shared goal is building systems that remove toil and drudgery from our software process. So that more people have the ability to build. We've seen the way like toil and drudgery have manifested.

19:33

It could be all of the difficulty you might have reproducing a bug. The challenges of monitoring a production system. Those are things that are really hard to do. And we can finally start to think about the structure of how we do them and building systems that allow us to do them repeatably and transferrably to people who are in non-technical roles. If these ideas excite you about how we can build these robust and reliable systems for anyone to ship software, you could stop by the work booth to come talk to me and the crew. We're at UG20. You can also mention me on Twitter. I'm CaptainSafia on all social media,

20:12

GitHub, Twitter, all of that fun stuff. Or just drop me an email. You can find my email on my personal site. Thanks for coming to this presentation. I hope you learned something interesting about some of the engineering philosophies that are driving the next set of work we do as far as agents, developers, and AI. yay yay yay yay ! ! You

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note