Open Reader

Emulated: The Data for Fully Autonomous Software Engineers and Companies — Joseph Wang

completed 16:32 Jul 31, 2026 Watch on YouTube

Current Status

completed

Video ID

zkX03APVj0M

RAG / Chat

Enabled
Emulated: The Data for Fully Autonomous Software Engineers and Companies — Joseph Wang
Description

To train an agent that can run production software, you need training data that looks like production, and that is what Joseph Wang's team at Emulated builds. Coming from network infrastructure backgrounds, they know what happens when something like a database goes down at scale, and they argue that current post training environments do not capture it. A real task is not a tidy code diff; it is fifty to a hundred turns of solving live traffic while distributed nodes fail, configs conflict, and unforeseen problems appear mid incident. So Emulated simulates whole companies. Imagine acting as an engineer inside a cloud provider or an infrastructure service, provisioning resources across VPCs, subnets, and security groups, meeting real bars around cost and deployment, and keeping a service alive as it grows, all inside a high fidelity environment rather than a stub. Wang's bet is that domain expertise plus faithful simulation is what lets agents learn the messy, end to end reality of infrastructure work, and he closes looking for people who have trained models or run real infrastructure to help push that fidelity further across more domains. Speaker info: - https://emulated.so/ Timestamps: 0:00 - Useful work over longer horizons 1:20 - Backgrounds in network infrastructure 2:26 - How environments shape capability 3:16 - Fifty to a hundred turn tasks 4:59 - Why real incidents are messy 7:11 - Real infrastructure isn't a code diff 7:40 - Acting as an engineer inside the cloud 9:37 - Deployment, cost, and scaling bars 13:29 - Why it's called Emulated 15:01 - Simulating full companies

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Skim
  • Core thesis: Fully autonomous software engineers will require post-training data generated in high-fidelity, multi-node company environments—not conventional single-container coding benchmarks.
  • Why it matters: The talk identifies a central limitation for AI operations agents: models can write code in repositories but are not trained to manage the organizational, deployment, reliability, security, and live-traffic consequences of operating real systems.
  • Best use: Use it as a concise strategic framing for why agent evaluation and training environments must evolve from code-task sandboxes into realistic operational control-plane simulations.

Executive Summary

Joseph Wang and Sid present Emulated's thesis that the remaining capability gap for autonomous engineering agents is principally a data-environment problem. Current coding benchmarks can ask an agent to produce a large PR over 50–100 turns, but they largely constrain work to a codebase. They omit the surrounding work that makes engineering operationally consequential: interpreting stale tickets and postmortems, responding to customers, deploying gradually, monitoring production, and managing failures without taking down a live service.

Emulated proposes placing the context of an entire software company into a sandbox. Their example is an etcd-based production service where an agent must reason across source code, incidents, customer history, rolling deployments, failing or deprecated nodes, live traffic, network faults, data corruption, clock skew, and blast radius. The intended training target is not merely successful code generation, but end-to-end operation of a service over long horizons.

The speakers argue that even sophisticated deterministic simulation inside a single container is insufficient for cloud-scale infrastructure work. A realistic cloud-service environment must include actual provisioning and lifecycle concerns—hosts, VPCs, subnets, security groups, APIs, throttling, auth, deployment and rollback, health monitoring under partitions, configuration changes, DNS/certificates, telemetry, billing, and administration. Their proposed direction is a multi-node “cloud in a box” sandbox that provisions real infrastructure resources.

The presentation is strongest as a problem statement and architecture direction, rather than an implementation guide. It flags unresolved constraints: spinning up stacks such as AWS Lambda may take hours, infrastructure environments are costly, and real-resource environments still do not naturally reproduce the customer traffic and scale-dependent failures that determine production fidelity. Emulated begins with infrastructure because the domain has clearer objectives and matches the founders' expertise, while expecting lessons to transfer to other company workflows.

Key Takeaways

  • Claim: Autonomous-agent reliability is constrained less by isolated coding ability than by the quality and scope of the environments used to generate training data. | Evidence: The speakers contrast agents that can complete large benchmark tasks and produce multi-thousand-line PRs over 50–100 turns with their inability to reason through infrastructure concerns such as database MVCC, corruption risk, and multiyear architecture consequences. | Implication: For agent systems intended to own production work, evaluate and train against operational outcomes and decision context—not repository-level code correctness alone. | Caveat: This is presented as Emulated's operating thesis rather than comparative experimental evidence or benchmark results.
  • Claim: Existing software-agent benchmarks are structurally too narrow because their tasks predominantly occur inside a codebase. | Evidence: The talk names SWE-bench Pro, TerminalBench, Frontier Code, and DeepSWE as examples, arguing that they exclude PM-style customer discovery, experimentation and performance testing, and ownership of infrastructure over months or years. | Implication: Ken should separate code-agent benchmarks from company-operation benchmarks when judging claims of “autonomous engineer” capability. | Caveat: The transcript does not assess the individual benchmark designs or claim that they are unsuitable for code-level evaluation; the criticism is specifically about their limits as proxies for full engineering autonomy.
  • Claim: A credible end-to-end infrastructure task must force an agent to manage changing operational context and protect live traffic while making changes. | Evidence: In Emulated's etcd example, the agent must use tickets, projects, postmortems, and customer reactions that may be outdated; execute rolling deployments; handle deployment conflicts, failing nodes, and deprecated nodes; and monitor a service whose blast radius makes downtime unacceptable. | Implication: Agent environments should include stateful operational artifacts, conflicting evidence, irreversible or costly actions, and service-level constraints rather than only a hidden test suite.
  • Claim: Single-node container sandboxes can simulate some distributed behavior but hit a ceiling for infrastructure-agent training. | Evidence: The speakers describe simulating a distributed cluster inside one sandbox with flapping nodes and lagging learners, but argue this cannot adequately represent resource provisioning or AWS-scale behavior; they call standard post-training pipelines homogeneous because everything runs in a single container. | Implication: A control plane for autonomous infra agents may need isolated multi-node environments and resource APIs, not just a stronger terminal sandbox. | Caveat: Multi-node infrastructure increases setup time, execution cost, and operational complexity, so greater realism is not free or automatically sufficient.
  • Claim: Training environments for cloud-service engineering need to represent the whole service lifecycle, including governance and commercial operations. | Evidence: Their cloud-service decomposition includes host provisioning, VPCs, subnets, security groups, customer-facing APIs, throttling, authentication, authorization, CloudTrail-like auditability, version management, gradual rollout and rollback, partition-aware health monitoring, DNS/cert management, inventory, fraud, admin consoles, telemetry, and billing. | Implication: The practical target is broader than an autonomous SRE or coding agent: it is an agent able to operate a service through the interfaces and constraints of a real company.
  • Claim: “Real infrastructure” remains an incomplete proxy for production unless the environment can reproduce scale-dependent conditions and customer behavior. | Evidence: The speakers explicitly ask how to manage the cost of real resources and note that a simulated-real gap remains even with real infrastructure because live customer traffic and problems visible only at scale are still absent. | Implication: Do not equate cloud-backed sandbox execution with production readiness; an agent-training stack still needs realism criteria, fault/traffic generation, and staged safety gates. | Caveat: The talk provides no proposed measurement framework for fidelity, nor a solution for synthesizing traffic, organizational behavior, or rare production incidents.
  • Claim: Infrastructure is Emulated's initial vertical because it combines strong domain expertise with comparatively explicit success criteria, but the company expects vertical lessons to generalize. | Evidence: They cite their backgrounds in network infrastructure, distributed databases, and sandbox infrastructure; compare clear DevTools objectives such as low-latency, low-cost GPU sandboxes and uninterrupted training runs with early-stage startups still seeking product-market fit; and state they are exploring going deep in one domain before scaling outward. | Implication: For a general agent-company thesis, start with domains that have observable operational objectives, known failure modes, and strong evaluators before tackling ambiguous product decisions. | Caveat: Transfer from infrastructure to less specified, product-oriented company work is asserted but not demonstrated.

Detailed Brief

Post-training infrastructure and unresolved execution constraints

  • Claims: Changing the sandbox from a single container to real multi-node infrastructure changes the design of the post-training pipeline itself.; The speakers suggest that a post-training pipeline can be placed inside the sandbox, opening further model-training and RSI-related possibilities, but do not elaborate on a concrete system design.; Environment reset, provisioning latency, and per-run cost become core bottlenecks when training on complete cloud-service stacks.
  • Evidence: They state that spinning up an entire AWS Lambda-like stack can take hours and ask how that could fit within a post-training rollout.; They characterize the intended environment as a multi-node sandbox with access to real cloud resources, or a “cloud in a box.”
  • Caveats: No throughput numbers, unit economics, reset mechanism, orchestration architecture, evaluation protocol, or empirical lift from these environments is provided.; The transcript repeats the audience-Q&A response near its end, so the effective amount of unique content is somewhat lower than the stated transcript length suggests.
  • Implications: Any effort to operationalize this thesis should treat environment lifecycle management—warm pools, snapshots, reset integrity, resource quotas, and cost attribution—as first-class training infrastructure.; The investment or partnership question is not whether the vision is plausible, but whether a provider can make high-fidelity environments reproducible and economical at post-training scale.

Notable Concepts & Terms

  • Full-company sandbox: An environment containing not just source code but organizational context, incidents, customer interactions, deployment systems, and production-like operational constraints.
  • Single-node sandbox: A conventional containerized agent environment that can imitate some distributed conditions but cannot naturally provision or operate real multi-resource cloud systems.
  • Multi-node sandbox / “cloud in a box”: Emulated's proposed environment type: isolated infrastructure with multiple nodes and access to real cloud resources for more realistic agent training and evaluation.
  • Operational blast radius: The scope of customer or service damage caused by an agent's change; the talk uses it to distinguish production operations from code tasks with disposable test environments.
  • MVCC: Multi-version concurrency control; cited as an example of infrastructure-level reasoning where mistakes can cause database corruption.
  • Simulated-real gap: The difference between sandbox behavior and production behavior that persists even when real resources are used, particularly around live traffic and scale-emergent failures.
  • RSI: Mentioned in connection with placing post-training pipelines inside sandboxes; the speakers do not define or explain its intended mechanism in the transcript.

Operator Notes / Why Ken Should Care

  • Add an evaluation tier for operational agents that requires reading imperfect organizational artifacts, executing a staged deployment, observing service health, and recovering from injected failures while satisfying an explicit blast-radius budget.
  • Before using cloud-backed sandboxes for post-training, define reset-time, provisioning-time, per-episode cost, resource-isolation, and fidelity metrics; otherwise realism will likely destroy iteration throughput.
  • Prioritize initial agent domains with crisp customer outcomes, reliable telemetry, and known failure modes—such as infrastructure operations—rather than ambiguous product-management workflows.
  • Treat claims that an agent can “own a company” as unproven until they include evidence of performance under live-traffic analogues, authorization boundaries, rollback requirements, and prolonged stateful operation.

Source/Metadata

  • Title: Emulated: The Data for Fully Autonomous Software Engineers and Companies — Joseph Wang
  • Transcript words: 2712
  • Duration seconds: 992
  • Timestamp note: No timestamps or chapter markers were present in the supplied transcript; the final audience-Q&A segment appears duplicated.

Transcript

2317 words en Processed in 90.2s

So I appreciate the intro. My name is Joseph, and this is my co-founder Sid. Emulated is a data lab focused on increasing the reliability and autonomy of AI agents. And if you've been at AI Engineer and you've watched the talks, seen the tracks, then there's probably one takeaway that all the talks have in common. And it's that we're headed towards a future where agents are able to perform useful work over longer and longer horizons with little to no supervision. So today we're going to answer some of the questions of what this means for the data and model layers. We're going to touch on some pretty cool things, so look out for them. How to simulate a company within a sandbox, or sandboxes for multi-node systems and distributed clusters. And if we have a little bit of time, we'll also go into some of the work that we're doing with post-training pipelines and how these new types of sandboxes are affecting post-training infra as well. So, where Sid and I come from, our backgrounds are in network infra, distributed databases, and sandbox infra. And these are all areas where the workloads are mission critical. We all saw a couple months ago that when something like DynamoDB goes down, so does us-east-one and half the internet. And working on these systems, we saw a model capability gap when it came to operating and building these systems at scale and thinking about the consequences of architecture and systems design over the course of years. Yeah, so it led to a pretty natural question, right? For such mission-critical services, why is it that my model or my agent is so proficient at handling the application layer, but struggles when it comes to reasoning through infrastructure complexities, for example, things like MVCC on a database engine, which can lead to corruption issues, which was one of the roots of the DynamoDB failure a few months ago? So, with everything in ML, the gap in models is usually a gap in data. Models typically are only as good as data is. And to really highlight this point, model capability has never regressed whenever you introduce more high-quality data. So with that being said, what is the data gap then? What does data look like right now? And how is this influencing the model capability gap here? So if you look at any of the frontier or recent benchmarks, like Sweebench Pro, TerminalBench, or something like Frontier Code and DeepSwee, the tasks only operate within the code base. The agent is given a pretty large task and, over the course of 50 to 100 turns, produces a couple-thousand-line PR. But it doesn't do all of the work that a human does. It doesn't do what a PM does with talking to customers, understanding their problems, what an engineer does with trying out different approaches, performance testing them, and owning the underlying infra for the code base over the course of not just months but years. And this is really the gap that we're closing. We've taken software engineering companies and we've put them into containerized environments. So this includes organizational context like projects, incidents, customer conversations. The agent also has to deal with issues that only appear at scale, like network failures between distributed nodes, data corruption, and clock skew. And through all this, we also want the agents to reason about orchestrating through distributed clusters and also thinking about things like operational blast radius while solving live traffic. And the result is that the tasks that these agents have to complete, or we want the agents to learn, are that environments are far more complex and long horizon than a simple code task. So let's just bring a picture into the mix because it tends to make things more interesting. Here's an example we've built of an etcd consensus cluster that a typical production service might rely on. So an early environment might tend to operate and work primarily on that little blue square entitled etcd source code in the bottom right there. But a lot of the fun and the model capability gap that results from it is really in everything that surrounds it. So you start with the tickets, projects, postmortems. What are the train wrecks? Why did they happen? How did customers feel about them? And oftentimes those aren't necessarily up to date. The agent has to incorporate all that when it's reasoning through the actual change that current environments have it make. After it makes that change, you need to kick off rolling deployments. Those deployment systems can oftentimes be complicated, have conflicts, may not work. And all through that, when you're finally migrating from old hardware onto new hardware, you run into unforeseen problems. So the agent has to reason through them in real time, just like a human would. You have failing nodes. You have stale, deprecated nodes. And while all of this is happening, the service can't go down because there is a blast radius to solving live traffic. You have to observe and monitor your service. All of these components in the system are really what exemplifies a full end-to-end infrastructure task. So what Sid is describing here is an environment in a single-node sandbox where we're simulating a distributed cluster with multiple nodes, flapping nodes, lagging learners in a single sandbox. And you can get pretty far with this, right? You can see that there's live traffic. There's a lot of operational issues that a real engineer would have to deal with. And you can make this pretty long horizon by just, say, doing multiple deployments instead of just one. But really what we're seeing is that this is not enough. This fits into standard post-training pipelines in the sense that a standard post-training pipeline is boring. It's homogeneous. Everything just runs harbor. Everything is a single sandbox containerized. But real infrastructure doesn't work like this. This isn't how real companies run. And even though you can use something like deterministic simulation to simulate network failures, it doesn't represent what you might run into if you're building an AWS-scale service. So, I did see, I think, a couple people at AWS. Somebody had Viceroy open on their laptop. Fun times. But let's imagine here that we are all AWS engineers or GCP engineers. Azure too. No shade, right? And we are building a cloud service. It can also be some infrastructure service like Datadog, Vercel, Supabase. All of these services run into the same problems. You start off with a shiny piece of software. And this piece of software can serve a single customer pretty well. Maybe it's running on your machine. If you're working for NLB, this would be a load balancer, right? If you're working for AWS Lambda, it would be some sort of serverless runtime. But it needs to actually run somewhere. So, if you're an infrastructure engineer, the next step is you get into resource provisioning. And this is already where the single-node sandbox starts breaking down. How do you provision resources within a single sandbox? You can't exactly simulate something like EC2 or Cloud Run, right? So, you get into this host provisioning. It also includes provisioning of other resources like VPCs, subnets, security groups. And you need to expose this through some sort of API because your customers are going to want to do things like, give me this shiny piece of software. Or, I don't want it anymore. It costs too much. I'm going bankrupt. Delete it, please. And so, you're going to need some sort of front-end API. And if you have enterprise-grade customers who really care about quality, then you're going to have to meet certain bars like throttling, authentication, authorization. You can't really go without these things, right? If you're AWS, then that's CloudTrail too. And then beyond this, software is living. People forget this all the time, especially investors, right? They'll be like, oh, you wrote it. You're done. But software is living, and you probably need some sort of software deployment component as well. Something to, whenever you have an update to roll out, roll it out. And God forbid something goes wrong, roll it back. You need to manage all the different versions and make sure your deployments are gradual to limit your blast radius. And we're just kind of getting started with this. There's all sorts of things that you need to think about, like health monitoring with awareness for network partitions. And then how do you communicate with your host so you can change configs on the fly? Maybe your customer actually wants to call your endpoints, so you need DNS and cert management. And then your service grows a bunch. You need to keep track of all your resources, what's going on, fraud and stuff. Then you need admin consoles, telemetry, billing if you're making money, all sorts of things. And with all of this, I think there's one more slide for scheduling. Yeah. I think the point is fairly clear at this point. Beyond a certain threshold, there is a critical mass at which sandboxing on a single node can only get you so far. And that's why we envision the future going towards a world where environments do provision real infrastructure. Yeah. So what this is is a multi-node sandbox with access to real infra, real cloud resources. We kind of put a cloud in a box, so cloud box could be another name for this. And as you can imagine, changing the sandbox type so drastically here affects post-training pipelines as well, which I think we might be running a little bit low on time, so we won't get too much into it. But yeah, one really cool thing too is you can put a post-training pipeline in the sandbox. And there's some cool stuff with model training and RSI that you can get into there. So then this begs the question, this is all cool stuff, Joseph. Thank you, Sid, for speaking. Why are you leaking all of this alpha, right? Why are you telling all your organizational secrets and telling everybody, oh, okay, how do you build a system like this? It's because we're really interested in these challenges here. We think they're very fun. We think they're really cool. We think that you guys are cool people. Or maybe I'm just lying. Who knows? And we want to share these challenges with you in case you're interested in working on them as well. As you can imagine, there's a lot of different problems that we haven't touched on here. For example, spinning up the entire stack for something like AWS Lambda takes hours. How do you fit that into a post-training rollout? And then there's cost as well. How do you efficiently manage this? How do you make sure the simulated-real gap, even with real resources, still exists, right? You still have to have live customer traffic. You still have to have problems that only appear at a certain scale. So, if you're a distributed systems engineer, if you've trained models before, if you think that this stuff is cool, then we'd love to talk. We'd love to talk, kind of see where your opinions are, hear what you've worked on. Maybe that's Kubernetes and you have opinions. Well, everybody has opinions on auto scaling and rolling deployments and whatever, but really niche opinions, right? Like at CDU. Yeah, we'd love to talk to you and hear what you have. Thank you. Yep. What's your problem with emulating? Yeah. That touches into why it's called Emulated in the first place, right? The real world is very, very complex. And how we as an industry emulate the real world is incredibly contrived and low fidelity. So Emulated's goal is really, how do you make these agents own systems like this, maybe beyond systems, entire companies, by emulating the real world with full fidelity? Yeah, go ahead. Next question is, are you predominantly focused on infra and painters, actually hardware, or are you also looking at what, yeah, of course. The question is infra is really cool. In 2026, all of a sudden, infra has come back and it's very sexy. Everybody wants to work on it, right? But there's other types of RL environments as well. There are other workflows that you really want to capture that aren't necessarily infra related. So the reason why we're starting with infra is a couple. Well, the first most important one is it speaks to our background the most. We think that domain expertise is something that informs how high quality your data can be. Especially with the boutique nature of data nowadays. And the second is that when we're simulating full companies, infra is the easiest to approach. If you think about any infra company out there, whether it be Supabase or Modal, or any DevTools company, the problem statement is pretty clear. Engineers know what they want. If you are working for Modal, you know that your users want GPU sandbox, very low latency, very low cost. You don't want to fail halfway through your training run. That's what you care about. So the problem statement becomes much easier. Whereas if you are a company in the YC summer 2026 batch, you're probably still trying to find product-market fit, right? Yeah. At the same time, there are also lessons learned that, going really vertical on a single domain like infrastructure, do translate into other horizontal domains. So we're also exploring going deep into one and scaling out that way. Yeah. All right. Really appreciate it. Appreciate the questions. We'll probably step out and we can take a couple more outside just to make sure the next speaker has room. Yeah. Thank you guys for listening. Appreciate it. Thank you. maybe beyond systems, entire companies, by emulating the real world with full fidelity. Yeah, go ahead. Next question is, um, are you predominantly focused on like infra and, um, , painters, like, actually hardware, kind of, or are you also looking at, like, cool, or, like, what, what, um, yeah, of course. Like the question is like, you know, infra is really cool. Um, in 2026, all of a sudden, infra has come back and it's very sexy. Everybody wants to work on it, right? But there's other types of RL environments as well. Um, there is, uh, other workflows that you really want to capture that aren't necessarily infra related. Uh, so the reason why we're starting with infra is, um, there's a couple. Well, the first most important one is it speaks of our background the most. Um, we think that domain expertise is something that informs how high quality your data can be. Uh, especially with the boutique nature of data nowadays. Um, and the second is that when we're simulating full companies, infra is the easiest to approach. Uh, if you think about, like, any infra company out there, whether it be Superbase or Modal, uh, or any DevTools company, the problem statement is pretty clear. Uh, engineers kind of know what they want. If you are working for Modal, you know that your users want GPU sandbox, very low latency, very low cost. You don't want to fail halfway through your training run. Um, that's what you care about. So, the problem statement becomes much easier. Whereas if you are, you know, a company in the YC summer 2026 batch, you're probably still trying to find product market fit. Right. Yeah. At the same time, there's also lessons learned that going really vertical on a single domain like infrastructure do translate into other horizontal domains. So, we're also, um, exploring going deep into one and scaling out that way. Yeah. All right. Uh, really appreciate it. Uh, appreciate the questions. Um, we'll probably step out and we can, like, take a couple more, uh, outside just to make sure the next speaker has room. Um, yeah. Uh, thank you guys for listening. Appreciate it. Thank you.