Open Reader

From fork() to Fleet: Designing an Agent Sandbox Cloud — Abhishek Bhardwaj, OpenAI

completed 44:33 Jul 13, 2026 Watch on YouTube

Current Status

completed

Video ID

OqM67QG_Ikk

RAG / Chat

Enabled
From fork() to Fleet: Designing an Agent Sandbox Cloud — Abhishek Bhardwaj, OpenAI
Description

Sandboxes unleash agents by giving them secure, fully functional computers where they can tackle diverse tasks with minimal setup. This talk explores the architectural challenges of building an agent sandbox cloud. We compare runtime isolation technologies and their trade-offs, examine persistence and storage as the next major unlock for agent capabilities, and discuss the key decisions involved in orchestrating and scaling sandboxes. Abhishek Bhardwaj works on Agent and Reinforcement Learning Infrastructure at OpenAI. He builds systems that enable large-scale model training in RL environments, as well as secure and scalable cloud sandboxes for OpenAI’s agents. Before joining OpenAI, he created Arrakis, an open-source sandbox for AI agents. Previously, he worked at Google on ChromeOS and foundational microVM technologies, and at Replit on core infrastructure and early versions of Replit Agent.

Summary

Generated by claude-sonnet-4-5

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Building production agent sandbox clouds requires micro-VMs for security, disk persistence for long-horizon tasks, and orchestration for scale—security must trump performance trade-offs at ChatGPT/Codex scale
  • Why it matters: OpenAI's Codex/ChatGPT team explains first-principles design for running untrusted AI-generated code safely at scale, revealing infrastructure patterns that will define agent deployment
  • Best use: Study as reference architecture for agent infrastructure; contrast security vs performance trade-offs; understand why OpenAI chose micro-VMs over containers

Executive Summary

Abhishek Bhardwaj from OpenAI's Applied Reinforcement Learning and ChatGPT infrastructure team presents the design evolution of OpenAI's agent sandbox cloud—the system that securely executes code generated by models like o1 and powers products like Codex and ChatGPT code execution. The talk addresses why sandboxes are necessary (models need verifiable rewards via code execution to excel at math/code), then systematically builds from fork/exec through containers, gVisor, and micro-VMs, arguing that hardware virtualization is the only acceptable security boundary despite performance costs.

The core security argument is that containers, seccomp filters, and user-space kernels (GVisor) all still expose the host kernel to attack—models can chain exploits or discover kernel vulnerabilities. Micro-VMs (via KVM/Cloud Hypervisor/Firecracker) run guest kernels in VMX non-root mode (ring 0 in isolated CPU context), making host compromise require multi-step hardware attacks. Bhardwaj advocates skipping the 'seven stages of grief' and starting with micro-VMs for any agent product, accepting latency overhead as preferable to security breaches that destroy trust.

The second pillar is disk persistence, framed as 'the next unlock' after compute. Without durable storage, agents lose state on node failures or restarts—unacceptable for multi-day goal-mode tasks in Codex or multi-turn ChatGPT sessions. The talk details two persistence models: always-on (write-through to distributed storage like GCS via custom block filesystems) and explicit snapshotting (copy-on-write with incremental block-level diffs using FIEMAP). Persistence enables reliability (restore on failure), long-horizon tasks (multi-day goal mode), and exploration (Monte Carlo tree search with checkpoint/backtrack).

Orchestration at scale involves cluster/node schedulers that route sandboxes near inference clusters for low latency, use snapshot lineage to place restores on nodes with cached layers, and employ warm pools or memory snapshots for sub-second sandbox starts. The design prioritizes security over performance, reliability over raw throughput, and treats persistence as critical infrastructure—not an afterthought. Bhardwaj's experience (Google ChromeOS CrosVM, now OpenAI) colors every choice: 'system tricks can cover performance issues, but they cannot hide security breaches.'

Key Takeaways

  • Claim: Models must execute code for verifiable-reward tasks (math, code) to 'hill climb' and outperform pure next-token prediction | Evidence: ChatGPT answers '3+3' correctly (trained on internet data) but fails 'how many Rs in strawberry' until given code execution; training loop passes model responses to harness, executes code, grades correctness, backpropagates weights | Caveat: This applies to training (RL loops) and product (ChatGPT/Codex); security is critical in both—rogue code can exfiltrate model weights in training or user data in production | Implication: Any agent product or training infra must safely run untrusted code at scale; sandboxes are not optional for AI systems with tool use | Timestamp: 00:00-05:00
  • Claim: Fork/exec and containers expose the host kernel to direct syscall attacks; even GVisor (user-space kernel in Go) can be chained to exploit the host | Evidence: Fork/exec: process in ring 3 directly syscalls ring 0 kernel, can exploit kernel vulnerabilities or get root. Containers: still native processes with same host kernel, seccomp reduces attack surface but blocks legitimate requests (bad feedback loop). GVisor: sentry intercepts syscalls in user-space, but sentry/gofer sit atop host kernel—two-step chain (exploit sentry → exploit host kernel) is feasible, especially with models like o1 analyzing bug reports | Caveat: Containers with seccomp and AppArmor provide 'decent' isolation; GVisor is better (attack surface in user-space Go code vs kernel C code). But neither is hardware-isolated; advanced models may discover/chain exploits | Implication: For high-stakes agent clouds (multi-tenant, sensitive data, model training infra), containers and GVisor are insufficient; hardware virtualization is the security floor | Timestamp: 08:00-18:00
  • Claim: Micro-VMs (KVM + Rust VMMs like Cloud Hypervisor, Firecracker, CrosVM) isolate guest code in VMX non-root mode; host runs in VMX root—hardware guarantee that guest ring 0 cannot touch host | Evidence: Guest kernel runs in ring 0 but in separate CPU context (VMX non-root); host kernel and hypervisor run ring 0 in VMX root. Even if guest gets root or kernel exploit, it's confined to guest context. Paravirtualized devices (virtio) exit to host VMM process on device access; VMM is Rust (memory-safe), devices are jailed (block device can't access network, vice versa). CrosVM was first Rust VMM (Google, for ChromeOS), later forked by AWS Firecracker (Lambda/serverless); Cloud Hypervisor is community-driven, general-purpose | Caveat: Performance penalty: CPU context switch on every device access (I/O, block, network). Memory sharing is harder (balloon driver for reclaim, reactive not proactive). GPU passthrough via VFIO is single-tenant, no multi-tenant GPU sharing yet. Micro-VM term is about VMM memory footprint and boot speed, not guest size | Implication: Start with micro-VMs for agent sandboxes despite latency hit. 'Security breaches lose trust once, hard to regain; performance can be optimized with system tricks.' Bhardwaj's advice: skip the 'seven stages of sandbox grief' (trying containers, GVisor, etc.) and go straight to micro-VMs | Timestamp: 18:00-30:00
  • Claim: Disk persistence is 'the next unlock' after compute—without it, agents are 'computers without disks' that lose all work on restart, unacceptable for multi-day tasks or production reliability | Evidence: Users run goal mode in Codex for 3+ days; ChatGPT sessions create presentations, GitHub repos. If node dies or model flakes, work is lost, wasting GPU tokens and user trust. Persistence enables: (1) Reliability—checkpoint periodically, restore on another node if failure; (2) Long-horizon tasks—goal mode spans days without state loss; (3) Exploration—harness can checkpoint, try solution path A, backtrack, checkpoint, try B (Monte Carlo tree search over days) | Caveat: Snapshotting must be incremental (save diffs, not full GB+ disk each time) and fast (sub-second API response) to avoid bankrupting cost or UX. Restore must also be fast for product latency. Two paradigms: always-on (write-through to cloud storage) vs explicit (harness calls save API) | Implication: Agent infrastructure teams should prioritize disk persistence as core feature, not afterthought. Incremental block-level snapshots + restore lineage routing (place sandbox on node with cached layers) unlock new agent capabilities (multi-day reasoning, failure recovery, exploration) | Timestamp: 30:00-42:00
  • Claim: Explicit snapshotting via copy-on-write (XFS/Btrfs) + FIEMAP for block-level diff detection enables zero-copy base images, incremental uploads, and async snapshot returns | Evidence: Base image (Codex/ChatGPT rootfs) is copied with zero latency (copy-on-write, no block changes). Writable layer on top captures changed blocks. On snapshot, FIEMAP syscall returns which logical blocks changed and their ranges; zip and upload to GCS/S3. Harness can get snapshot ID immediately (lie while uploading in background). On restore, pull lineage of snapshot layers, apply diffs on base, start micro-VM with exact block-level state | Caveat: Block-level snapshotting is more efficient than file-level (less write amplification), but requires filesystem support (XFS, Btrfs, ext4 with extents). Always-on persistence via custom NBD filesystem on GCS/S3 is POSIX-compliant but requires tiered cache (in-cluster + in-VM) for performance; NFS is not POSIX-compliant and less performant | Implication: Implement block-level snapshot storage with lineage tracking; route restores to nodes with existing layers to minimize download time. This architectural choice separates commodity sandboxes from production-grade agent clouds | Timestamp: 42:00-48:00
  • Claim: Orchestration at ChatGPT/Codex scale requires cluster/node schedulers that prioritize latency (route near inference clusters), reliability (snapshot-based restore on failure), and smart placement (node with most snapshot layers cached) | Evidence: Top-level control plane picks cluster by region, load, proximity to ChatGPT/model cluster. Within cluster, scheduler scores nodes by load, failure state, and snapshot layer availability. For low latency, systems can: (1) pre-warm sandbox pools (idle CPU/memory cost), (2) use memory snapshots to boot micro-VMs in milliseconds, or (3) hybrid (warm pool grown from memory snapshots). Snapshot lineage routing: if snapshot has 4 layers, scheduler routes to node already caching all 4 → faster restore | Caveat: Warm pools waste idle resources; memory snapshots reduce waste but add complexity. Kubernetes and standard orchestrators exist, but Bhardwaj presents first-principles challenges without committing to Kubernetes (implied OpenAI uses custom orchestration). GPU access via VFIO is single-tenant, blocking multi-tenant GPU sandboxes | Implication: Agent sandbox orchestration is not generic container orchestration—requires snapshot-aware scheduling, latency-first routing, and persistent storage integration. If building from scratch, design scheduler around snapshot lineage and proximity to inference clusters, not just CPU/memory bin-packing | Timestamp: 48:00-53:00
  • Claim: Throughput matters in research (many parallel rollouts), latency in product (user-facing ChatGPT/Codex); both require reliability (GPU tokens are expensive) and security (protect model weights + user data) | Evidence: Research training loops run many rollouts (one task, multiple attempts) in parallel to maximize GPU utilization. Product must start sandboxes and execute code in sub-second to avoid churn (users expect fast, magical experiences). If sandbox fails, wasted GPU tokens in training; in product, users churn. Security protects OpenAI infra (models, weights) in training; user data in production | Caveat: Research and product have different primary metrics (throughput vs latency) but share reliability/security needs. Design must not sacrifice one for the other—micro-VMs + persistence serve both | Implication: When evaluating sandbox solutions, assess against all four pillars (latency, throughput, reliability, security) for your use case. Research-only sandboxes might accept higher latency for security; product cannot. Persistence is the bridge between research (long training runs) and product (long goal-mode tasks) | Timestamp: 05:00-08:00

Detailed Brief

Why Sandboxes: Models Need Code Execution for Verifiable Rewards

  • Claims: Pre-trained LLMs answer '3+3=6' correctly (trained on internet) but fail 'how many Rs in strawberry' (not enough training data); Verifiable reward tasks (code, math—testable true/false outcomes) need code execution to hill climb to correct answers; Training loop: harness passes model response → model emits code → harness executes code → grader scores correctness → backprop changes weights; Product side (ChatGPT, Codex): same flow but harness is user-facing, no training loop—must be fast, reliable, secure
  • Evidence: ChatGPT pre-trained on internet data where '3+3=6' appears frequently; rare to see 'strawberry R count' explicitly; Tool calling unlock: models that can execute code excel at math/code by testing and iterating, not pure next-token prediction; Example: training loop gives task, model writes code, harness executes, grader judges, loop backpropagates
  • Caveats: Models can generate non-malicious code most of the time, but 'good security hygiene' requires assuming code could be malicious (intentional or overzealous help attempts); Both training and product need security: training protects OpenAI infra/model weights; product protects user data and multi-tenant cloud nodes
  • Implications: Sandboxes are mandatory for any AI system with tool use or code execution; First principles: if you want models to solve verifiable tasks reliably, you must safely execute their code; The sandbox is not a product feature—it's infrastructure that enables training and product viability

Security Spectrum: Fork → Containers → GVisor → Micro-VMs

  • Claims: Fork/exec: simplest, most performant (native speed), zero isolation—process directly syscalls kernel, can exploit kernel or get root; Containers: namespaces isolate resources (PID, mount, net), cgroups limit CPU/memory to prevent noisy neighbor. Still native processes on host kernel—can exploit kernel boundary; Seccomp filters reduce attack surface by limiting syscalls and arguments, but blocks legitimate requests unpredictably (bad feedback loop for agents doing 'magical things'); GVisor: user-space kernel (sentry in Go) intercepts syscalls, filesystem via gofer daemon. Exploit is in user-space ring 3, not kernel ring 0. But sentry/gofer still run atop host kernel—two-step chain (exploit sentry → exploit host) is possible; Micro-VMs (KVM + Cloud Hypervisor/Firecracker/CrosVM): guest kernel in VMX non-root (ring 0 isolated CPU context), host in VMX root. Hardware guarantees guest cannot touch host even with root/kernel exploit
  • Evidence: Linux execution model: thread is smallest unit, kernel provides privileged access via syscalls that switch CPU to ring 0. Attack vectors: get root (ring 3, highest privilege user) or kernel exploit (ring 0, can dump process memory, total compromise); Fork/exec: while-loop forking brings down node (noisy neighbor). Model can try kernel exploits directly; Containers: PID namespace example—inside sees PIDs 1,2,3; outside sees native PIDs. Still same host kernel underneath; GVisor: sentry implements Linux API in Go (process, memory mgmt). File access via gofer. Reduces kernel attack surface but doesn't eliminate host kernel boundary; Micro-VMs: VMM (QEMU or Rust VMMs) forks, exposes API, calls dev-kvm (hypervisor API). Guest runs as separate Linux system. Paravirtualized devices (virtio block, net) exit to host VMM on access. CrosVM first Rust VMM (Google, ChromeOS), Firecracker forked for AWS Lambda, Cloud Hypervisor is community general-purpose
  • Caveats: Containers with seccomp provide 'decent protection' but feedback loop is poor—can't predict what syscalls agent needs, blocks users, requires filter updates; GVisor chain is harder (two steps) but o1/5.6 models may analyze bug reports and chain exploits; Micro-VMs have performance overhead: CPU context switch on every device access (block, net I/O). Memory sharing is reactive (balloon driver, not proactive reclaim). GPU via VFIO is single-tenant, no multi-tenant sharing. Term 'micro-VM' refers to VMM memory footprint and boot speed, not guest size
  • Implications: For agent sandboxes at scale (multi-tenant, sensitive data), hardware isolation via micro-VMs is the security floor; Bhardwaj's rule: 'Security breaches lose trust once, hard to regain. System tricks can cover performance issues.' Accept latency hit for security; 'Seven stages of sandbox grief': teams try fork, containers, GVisor, V8, realize they need full Linux, then want security → micro-VM. Bhardwaj's advice: skip grief, start with micro-VMs; Rust VMMs (CrosVM, Firecracker, Cloud Hypervisor) are memory-safe, jail devices (block can't access net, vice versa), providing second security layer beyond hardware isolation

Disk Persistence: The Next Unlock After Compute

  • Claims: Giving models a computer with no disk is like closing laptop = losing all work. Not acceptable for production agents; Product tasks are longer-horizon: users make presentations, create GitHub repos in Codex/ChatGPT sandboxes. Node failure or model flake loses work → wasted GPU tokens, user churn; Persistence unlocks three capabilities: (1) Reliability—checkpoint periodically, restore on another node if failure, supports cluster upgrades/A-B testing; (2) Long-horizon tasks—goal mode in Codex runs 3+ days, needs state to survive restarts; (3) Exploration—harness can checkpoint, explore solution path A, backtrack, checkpoint, explore B (Monte Carlo tree search over days/weeks); Snapshotting requirements: incremental (save diffs, not full GB+ each time), fast API (cheap/fast so harness can snapshot frequently), fast restore (product latency), choice of always-on (auto write-through) vs explicit (harness calls save); Design choices: incremental vs full snapshots; snapshot entire rootfs vs specific folders (workspace, home); file-level (entire files) vs block-level (changed blocks, less write amplification)
  • Evidence: Goal mode in Codex: Bhardwaj's record is 3 days; trend is longer and longer tasks; Example: node dies, user loses presentation → bad UX, wasted GPU tokens. With checkpointing, restore on another node, continue work; Monte Carlo exploration: harness checkpoints, tries solution, backtracks, checkpoints, tries alternative—enables multi-day reasoning loops; Incremental snapshotting at ChatGPT scale: full snapshots bankrupt cost, slow UX. Must save diffs only; Linux storage basics: disks are block devices (logical blocks 0…N), filesystem (XFS, ext4) maps file inodes to logical blocks via extents. Disk firmware maps logical blocks to physical sectors/pages
  • Caveats: Always-on persistence (write-through to GCS/S3) requires custom filesystem (NBD-based) with tiered cache (in-cluster + in-VM) for performance. NFS is not POSIX-compliant and less performant—models are 'very good at POSIX-compliant standard' Linux; Block-level snapshotting is efficient but requires filesystem support (XFS, Btrfs copy-on-write, ext4 extents). File-level has higher write amplification; Restore speed depends on snapshot lineage (chain of diffs); routing restore to node with cached layers minimizes download time
  • Implications: Disk persistence is not optional for agent products or research—it's the infrastructure unlock for reliability, long-horizon tasks, and exploration; Design snapshot storage with incremental block-level diffs, lineage tracking, and async upload (return snapshot ID immediately, upload in background); Integrate persistence with orchestration: scheduler routes restores to nodes with most cached layers, minimizing latency; Bhardwaj's hope: 'If we make this infrastructure correct, we can help solve diseases, find new drugs because the model can keep going for longer and longer'—long-horizon reasoning at scale requires rock-solid persistence primitives

Explicit Snapshotting: Copy-on-Write + FIEMAP for Incremental Block Diffs

  • Claims: Copy-on-write (XFS, Btrfs): zero-latency copy of base image (no blocks change until write). Writable layer on top captures changed blocks; FIEMAP syscall returns which logical blocks changed and byte ranges. Zip changed blocks, upload to GCS/S3. Harness gets snapshot ID immediately (lie while uploading in background); Restore: pull snapshot lineage (chain of diffs), download artifacts, apply diffs on base image using extent info, start micro-VM with exact block-level state; Async upload: snapshot API returns fast, uploads happen in background → low latency for harness, frequent checkpointing without blocking
  • Evidence: Base image: Codex/ChatGPT rootfs. Create zero-copy writable layer. On write, blocks change in layer, not base; FIEMAP (file extent map) tells which blocks changed, enabling incremental block-level snapshot (not entire file, not entire disk); Diagram: base.img → writable layer → FIEMAP → zip diff → upload to cloud. Restore: download lineage → apply diffs → start VM
  • Caveats: Requires copy-on-write filesystem (XFS, Btrfs, ext4 with extent support). Older filesystems lack this; Lineage chain grows over time; may need periodic full snapshots or garbage collection; Async upload risks: if upload fails before completion, snapshot ID is invalid. Need retry/validation logic
  • Implications: Implement block-level snapshot storage with FIEMAP-based diff detection for efficient, incremental persistence; Design harness API to call snapshot frequently (cheap, fast) without blocking on upload; Integrate lineage tracking into orchestration: scheduler knows which nodes have which layers, routes restores optimally

Always-On Persistence: Custom Filesystem on GCS/S3 with Tiered Cache

  • Claims: Write-through to cloud storage (GCS, S3) via custom NBD-based filesystem. Inside sandbox, sees block device (not shared folder like NFS); Tiered cache: in-VM cache (hot blocks) → in-cluster cache (warm blocks) → object storage (cold blocks). Write-back from in-cluster cache to object storage; Block device abstraction is more efficient than file-level sharing (9p, NFS) because guest can cache blocks, only exits to host when block not in cache. File-level (9p) exits on every file op, very slow; POSIX-compliant filesystem on top of object storage. Models are pre-trained on POSIX-compliant Linux, expect standard semantics
  • Evidence: Diagram: two micro-VMs, each sees block device (.img). As they write, writes go to in-cluster cache, then cloud storage; Comparison: left side (9p file sharing) = exit on every file op, very slow. Right side (block device via virtio-blk) = cache in guest, exit only on block miss, much faster; NBD (network block device) protocol allows block device to be backed by network storage (GCS/S3)
  • Caveats: NFS is not POSIX-compliant and less performant. Custom filesystem adds complexity but matches model expectations; Tiered cache requires cache eviction policy, write-back timing, consistency guarantees (what happens if in-cluster cache fails before write-back?); Block device abstraction is performant but less flexible than file-level sharing for selective folder persistence
  • Implications: For always-on persistence, build custom block-level filesystem on object storage with tiered cache, not NFS; Trade complexity for performance and POSIX compliance—models expect standard Linux filesystem semantics; Combine always-on (automatic write-through) with explicit snapshotting (harness-controlled checkpoints) for best reliability and exploration support

Orchestration: Cluster/Node Scheduling with Snapshot-Aware Placement

  • Claims: Multi-region clusters grouped by region, load. Top-level control plane picks cluster near ChatGPT/inference cluster for low latency; Within cluster, scheduler scores nodes by load, failure state, snapshot layer availability. Routes sandbox to node with most cached snapshot layers → faster restore; Low-latency creation strategies: (1) pre-warm sandbox pools (idle CPU/memory cost), (2) memory snapshots (save guest RAM, boot in milliseconds just-in-time), (3) hybrid (warm pool grown from memory snapshots); Snapshot lineage routing: if restore needs 4 layers, scheduler routes to node with all 4 cached → minimal download time. Example: node A has layers 1,2; node C has 2,3; node B has 1,2,3,4 → route to B
  • Evidence: Diagram: top-level control plane → cluster scheduler → node. Control plane uses region, load, proximity to inference cluster. Scheduler uses load, failure, snapshot layers; Warm pool: pre-start sandboxes, pick one on request. Fast but wastes idle resources. Memory snapshot: save guest RAM once, boot copies in milliseconds. Hybrid: warm pool + memory snapshot growth; Snapshot lineage diagram: node scores by layer overlap, highest score wins. Node B has all layers → fastest restore
  • Caveats: Warm pools waste idle CPU/memory; memory snapshots add complexity (save/load RAM state, potential for stale snapshots); Kubernetes exists, but Bhardwaj presents first-principles challenges without committing to it (OpenAI likely uses custom orchestration or heavily modified K8s); GPU access via VFIO is single-tenant, blocking multi-tenant GPU sandboxes—no solution presented for shared GPU in micro-VMs
  • Implications: Agent sandbox orchestration differs from container orchestration: must be snapshot-aware, latency-first, proximity-aware (to inference clusters); If building from scratch, design scheduler around snapshot lineage and inference cluster proximity, not just CPU/memory bin-packing; Use memory snapshots or warm pools for low-latency product experience; balance cost (idle resources) vs latency (on-demand boot); Snapshot-based restore on failure and cluster upgrades is key to reliability at scale

Notable Concepts & Terms

  • Verifiable reward tasks: Problems (math, code) with testable true/false outcomes; models can execute code, test correctness, and hill climb to solutions via RL training loops
  • Micro-VM: Not about guest size—refers to VMM (virtual machine monitor) with small memory footprint and fast boot time, typically Rust-based (CrosVM, Firecracker, Cloud Hypervisor) vs legacy QEMU
  • VMX non-root / VMX root mode: CPU contexts for hardware virtualization (Intel VT-x). Guest kernel runs ring 0 in VMX non-root (isolated), host kernel/hypervisor in VMX root. Hardware guarantees guest cannot touch host even with kernel exploit
  • Paravirtualization (virtio): Guest drivers aware they're in VM, use efficient virtio protocol (virtio-blk, virtio-net) to exit to host for device access. More performant than full device emulation
  • Copy-on-write (COW) snapshots: Zero-latency file/disk copy (no blocks change) until write happens. XFS/Btrfs create writable layer on top of base image; on write, only changed blocks are copied. Enables incremental snapshotting
  • FIEMAP: Linux syscall (file extent map) that returns which logical blocks of a file changed and byte ranges. Used to detect diffs for incremental block-level snapshots
  • NBD (network block device): Protocol to expose block device over network. Allows custom filesystem on GCS/S3 to appear as local block device inside micro-VM
  • Goal mode (Codex): OpenAI Codex feature where agent runs long-horizon task (Bhardwaj's record: 3 days). Requires persistence to survive node failures and restarts
  • Seven stages of sandbox grief: Bhardwaj's pattern: teams try fork, containers, GVisor, V8, realize they need full Linux + security → micro-VM. Advice: skip grief, start with micro-VMs
  • Snapshot lineage: Chain of incremental snapshot diffs that reconstruct full disk state. Scheduler routes restores to nodes with most lineage layers cached for fast restore
  • Seccomp (secure computing mode): Linux kernel feature to filter syscalls and arguments a process can call. Reduces kernel attack surface but blocks unpredictable requests, bad feedback loop for agents
  • GVisor: Google's user-space kernel (sentry in Go) that intercepts syscalls, running in ring 3 instead of ring 0. Reduces kernel attack surface but sentry/gofer still atop host kernel—two-step exploit chain possible
  • CrosVM / Firecracker / Cloud Hypervisor: Rust-based VMMs (virtual machine monitors). CrosVM: Google, ChromeOS. Firecracker: AWS (forked CrosVM), Lambda/serverless. Cloud Hypervisor: community, general-purpose. All memory-safe, jail devices, boot fast

Operator Notes / Why Ken Should Care

  • OpenAI's Codex/ChatGPT sandbox architecture is a reference design for any agent product: if you're building agents that execute code, start with micro-VMs and disk persistence from day one
  • Security trade-off is non-negotiable: Bhardwaj argues breaches lose trust permanently, performance can be optimized. If evaluating agent infra vendors, ask: 'Do you use micro-VMs or containers?' If containers, probe on kernel isolation story
  • Persistence is the unlock for long-horizon reasoning and reliability. If building agent infrastructure, implement incremental block-level snapshots (FIEMAP, copy-on-write) with async upload and lineage-aware scheduling. This is the gap between demo and production
  • Orchestration for agents differs from container orchestration: must be snapshot-aware (route restores to nodes with cached layers), latency-first (use memory snapshots or warm pools), and inference-cluster-proximate (low latency to model API)
  • If investing in agent infrastructure startups, ask: 'How do you handle security (containers vs micro-VMs)?', 'How do you persist state (always-on vs explicit snapshots)?', 'How do you route restores (snapshot lineage caching)?'. These are the hard problems at scale
  • GPU in micro-VMs is unsolved for multi-tenancy: VFIO is single-tenant. If building GPU-accelerated agent sandboxes, you'll need custom solution or accept single-tenant GPU (expensive). This is an open research/product problem
  • The 'seven stages of grief' pattern is real: teams underestimate security risk, try cheap solutions, get breached or realize kernel exposure, then adopt micro-VMs. Save time and reputational risk by starting with micro-VMs
  • For GTM: agent products compete on 'magical experiences' (Bhardwaj's phrase). Magical = fast (low latency sandbox start), reliable (state persists through failures), safe (no data leaks). Micro-VMs + persistence + smart orchestration are table stakes for competitive agent UX
  • OpenAI's design philosophy: security > performance, POSIX compliance for models (pre-trained on standard Linux), incremental snapshots at block level (not file level), explicit + always-on persistence paradigms (cover different use cases)
  • If building agent harnesses/orchestrators, expose snapshot APIs to harness (let harness checkpoint frequently for exploration), use lineage-aware routing (minimize restore latency), and assume long-running tasks (days to weeks, not minutes)

Watch Map

  • 00:00-05:00: Why sandboxes: models need code execution for verifiable rewards (math/code), training vs product needs
  • 05:00-08:00: Security spectrum: fork/exec (bad), containers (namespaces/cgroups, still host kernel exposure)
  • 08:00-13:00: Containers deep dive: PID/mount namespaces, cgroups for resource control, seccomp for syscall filtering (limits attack surface but bad feedback loop)
  • 13:00-18:00: GVisor: user-space kernel (sentry in Go), file access via gofer. Better than containers but still two-step exploit chain to host kernel
  • 18:00-23:00: Micro-VMs: hardware virtualization (VMX non-root/root), guest ring 0 isolated from host ring 0. Linux execution model: ring 0 (kernel) vs ring 3 (user), syscalls cross boundary
  • 23:00-28:00: Paravirtualization (virtio): guest drivers aware of VM, efficient device access. VMM (QEMU, CrosVM, Firecracker, Cloud Hypervisor) calls dev-kvm hypervisor API. CrosVM history: Google ChromeOS, first Rust VMM; Firecracker forked for AWS Lambda; Cloud Hypervisor is community general-purpose
  • 28:00-30:00: Micro-VM trade-offs: hardware isolation (best security), jailed devices (Rust, memory-safe), but performance overhead (CPU context switch on device access), reactive memory sharing (balloon driver), GPU via VFIO is single-tenant. Bhardwaj's advice: 'seven stages of grief'—skip grief, start with micro-VMs
  • 30:00-35:00: Disk persistence: 'the next unlock after compute.' Without disk, agents are 'computers without disks'—lose work on restart. Enables reliability (restore on failure), long-horizon tasks (goal mode 3+ days), exploration (checkpoint/backtrack Monte Carlo search)
  • 35:00-40:00: Persistence requirements: incremental snapshots (save diffs, not full GB+ each time), fast snapshot API (cheap so harness can checkpoint frequently), fast restore (product latency), always-on (auto write-through) vs explicit (harness calls save). Design choices: incremental vs full, entire rootfs vs specific folders, file-level vs block-level (block = less write amplification)
  • 40:00-45:00: Linux storage basics: disks are block devices (logical blocks 0…N), filesystem maps file inodes to blocks via extents. Two access modes: file-level sharing (9p, NFS—exit on every op, slow) vs block device (virtio-blk—cache in guest, exit on block miss, fast)
  • 45:00-48:00: Always-on persistence: write-through to GCS/S3 via custom NBD filesystem with tiered cache (in-VM → in-cluster → object storage). POSIX-compliant because models are pre-trained on standard Linux. NFS is not POSIX-compliant and less performant
  • 48:00-50:00: Explicit snapshotting: copy-on-write (XFS/Btrfs) zero-latency copy of base image, writable layer on top. FIEMAP syscall returns changed blocks/ranges, zip and upload to cloud. Async upload: return snapshot ID immediately, upload in background (lie while uploading). Restore: pull lineage, download artifacts, apply diffs on base, start micro-VM
  • 50:00-53:00: Orchestration: multi-region clusters, control plane picks cluster near inference cluster (low latency). Scheduler scores nodes by load, failure, snapshot layer availability. Low-latency strategies: pre-warm pools (idle cost), memory snapshots (boot in milliseconds), hybrid. Snapshot lineage routing: route restore to node with most cached layers → minimal download. Example: node B has all 4 layers, highest score, fastest restore
  • 53:00-55:00: Conclusion: sandboxes are critical for agent training and product, persistence is next unlock, use micro-VMs from start (skip grief). Bhardwaj's hope: correct infrastructure enables long-horizon reasoning to 'solve diseases, find new drugs'

Source/Metadata

  • Title: From fork() to Fleet: Designing an Agent Sandbox Cloud — Abhishek Bhardwaj, OpenAI
  • Transcript words: 13731
  • Duration seconds: 2673
  • Timestamp note: Timestamps estimated from transcript flow; video duration 2673 seconds (~45 minutes). Chapters derived from content transitions, not explicit markers in transcript

Transcript

7254 words en Processed in 1789.8s

name name [SPEAKER_00] name [SPEAKER_00] name [SPEAKER_00] name name [SPEAKER_00] name [SPEAKER_00] enforcement learning specifically. [SPEAKER_00] And on the product side, we also develop infra that helps run untrusted code as part of ChatGPT, CodexWeb, securely and reliably at scale. This talk is called From Fork to Fleet, Designing an Agent Sandbox Cloud. [SPEAKER_00] I'll be very clear that there are a lot of words [SPEAKER_00] in the title that have OS and infra concepts, [SPEAKER_00] but this is a first principles talk. [SPEAKER_00] So we'll cover what sandboxes are [SPEAKER_00] and why they are needed from first principles. [SPEAKER_00] We will also try to cover design intuitions [SPEAKER_00] around designing an Agent Sandbox Cloud [SPEAKER_00] to run sandboxes securely and reliably at scale. [SPEAKER_00] So if some of these words don't mean anything, [SPEAKER_00] don't be worried. [SPEAKER_00] We'll explain from first principles and go from there. [SPEAKER_00] Last year I gave a talk called [SPEAKER_00] How to Build an AI Sandbox from Scratch. [SPEAKER_00] If you're interested, you can look at that talk as well. [SPEAKER_00] Think of this as a spiritual sequel to that talk. [SPEAKER_00] So now let's forget about sandboxes or clouds for a second. [SPEAKER_00] Let's just go back in time. [SPEAKER_00] ChatGPT came out. [SPEAKER_00] It's a very large pre-trained model. [SPEAKER_00] People ask all sorts of questions [SPEAKER_00] and it responds really, really well, [SPEAKER_00] really human-like answers. [SPEAKER_00] But when people ask questions like, [SPEAKER_00] what is 3 plus 3? [SPEAKER_00] Or how many Rs in strawberry? [SPEAKER_00] Sometimes it works and sometimes it doesn't. [SPEAKER_00] But it answers questions like, what is 3 plus 3? [SPEAKER_00] Quite well, because it's trained on the entire internet. [SPEAKER_00] And apparently people have written 3 plus 3 equal to 6 many, [SPEAKER_00] many times on the internet. [SPEAKER_00] So it gets it right. [SPEAKER_00] But people haven't asked, how many Rs in strawberry? [SPEAKER_00] Enough on the internet. [SPEAKER_00] And so it gets it wrong. [SPEAKER_00] And so it's obvious that for anything code or math related, [SPEAKER_00] or any problem which has a verifiable reward, [SPEAKER_00] which means that it can be tested, [SPEAKER_00] that it's true or false, [SPEAKER_00] the model needs something more. [SPEAKER_00] And the key unlock was that, [SPEAKER_00] given the model tool calling capability, [SPEAKER_00] or a way to execute code, [SPEAKER_00] the model gets these verifiable reward questions [SPEAKER_00] around code and math correctly. [SPEAKER_00] And thus from first principles came that, [SPEAKER_00] if we give the models the ability to execute code, [SPEAKER_00] it can hill climb and be very, very good at math, [SPEAKER_00] code and other domains that have verifiable rewards. [SPEAKER_00] Cut to 2026. [SPEAKER_00] And we are seeing the consequences of doing that at scale. [SPEAKER_00] So now how can it answer what is 3 plus 3? [SPEAKER_00] And how can it answer how many Rs in strawberry? [SPEAKER_00] Well, it can write code to do this. [SPEAKER_00] So if you see the diagram, [SPEAKER_00] we have a training loop, [SPEAKER_00] and the training loop gives it tasks or questions, [SPEAKER_00] and then the training and the harness [SPEAKER_00] then passes the response of the model, [SPEAKER_00] and the response might say, [SPEAKER_00] hey, execute code on my behalf. [SPEAKER_00] The harness is responsible for executing the code, and then a grader judges whether the answer is correct or not, and then the training loop back propagates and changes the weights. And that's how we train it to do two things. We train it to call code execution or tools on certain classes of problems. And secondly, we ensure that the code it executes actually solves the problem. So this is why tool calling is important on the training side. Now let's talk about the product side. In the previous slide, we showed how the models are trained to emit code in order to solve certain tasks. The agent executes the code, and we verify the reward. Well, all of this is useless if we don't support this model on the product side. So it's the same slide as before, but we don't have a training loop, and now the harness is passing the response and calling the tools and executing the code somewhere, right? So where is this code or tool being executed, right? This could be your laptop, with codex or any agent you want, or it could be a node on the cloud with codex web or chat GPT. While we can expect the model to generate non-malicious code, good security practices and hygiene mean that we want to protect our environment, whether it's a laptop or the cloud node. Attacks can be intentional or unintentional, and we want to protect against those. And it could be trying to get root on your system or trying to exploit a kernel vulnerability. The models are getting really, really big, and they might try to help you in an overzealous fashion and try to get root to do so. So we want to try and avoid and make sure it doesn't attack the node or where it's running. Thus, this is where the sandbox comes in. We need the sandbox to run this untrusted code and ensure that it can do its work, but it shouldn't be able to exploit any vulnerabilities and get root on your system. In the cloud, other sandboxes with other users' data might be running as well, and we don't want it attacking and getting data of other users out. So a sandbox is an environment in which you can run these tool calls and execute code on behalf of the model securely, and it could be on your laptop or it could be on the cloud. I think everyone's probably used OpenClaw or any agent like that. I think it's a very, very big peek into what's to come. A lot of these agents are running locally on your laptops, and I think it's a slap on the face for 20 years of cloud computing that everyone's running this locally on their laptops. and getting data of other users out. So a sandbox is an environment in which you can run these tool calls and execute code on behalf of the model securely, and it could be on your laptop or it could be on the cloud. I think everyone's probably used OpenClaw or any agent like that. I think it's a very, very big peek into what's to come. A lot of these agents are running locally on your laptops, and I think it's a slap on the face for 20 years of cloud computing that everyone's running this locally on their laptops. And if you see the image, I don't know if this is an actual product, but it's very funny because the lid of the laptop is open. It's because you don't want your agents to sleep. Well, we have a whole cloud. There are industries built on this thing, so I think the future is us running your agents in the cloud. They're persistent, long-running, and I really hope that moving forward, this is not something you see anywhere. But people rented VPSs. They ran OpenClaws on Hexner or Mac minis in the cloud. So I think OpenClaw was a very good peek into what might be for sandbox clouds in the future. We've discussed why sandboxes are important in both research and product, but they have slightly different needs. In research, we want to optimize for throughput. We want to run many training loops at scale and have many rollouts. A rollout is one version of a task. So what is three plus three? And you might have five answers to it, and one answer is a rollout. We want to take many shots on goal in parallel. And so throughput is very important in research. In product, latency is very important. I think any successful product in the last 20 years has been super fast. So if you don't start a sandbox in time and you don't execute code fast enough, people will churn from your product. Reliability is important on both sides. If you fail constantly, you've wasted GPU tokens on both sides, and GPU is gold right now. So you want to make sure the tokens you're getting are useful. And similarly, if your agents aren't reliable on the product side, it's game over. People will churn from your product. Security is also important on the research side. We are training on OpenAI infrastructure. So if a model gets root and it's not aligned, it can try to attack OpenAI infrastructure. It can exfiltrate and release our model rates or whatever. And then on the product side, it can exfiltrate other users' data and attack the infrastructure as well. So security is important for both research and product. So today we'll focus on these three pillars. There are many parts of a Sandbox Cloud, but we'll specifically focus on runtime. So how can we run Sandbox on one node securely? Secondly, we'll focus on persistence. I think compute was the first unlock. People realized you give Sandboxes a Linux computer, and they do crazy things because they're pre-trained on so much Linux data. But I think now if you help them, if you give them a computer with an actual disk that can be saved, then they become a true knowledge worker. And the last part is orchestration and how to run these at scale for many users at ChatGPT and Codex scale. So before we start, this is a first principles talk. So let's discuss how Linux executes code on your machine. On Linux, a thread is the smallest unit of execution. The kernel is the thing that provides privileged access to a thread via something called system calls or iOctils. So whenever the user space program wants to access some hardware or some privileged resource, it needs to talk to the kernel and call this operation that switches the hardware context to a more privileged context. And so let's look at what that looks like. So if you see here, your processor has different rings of execution. Based on which ring you're executing in, you get different privileges. So the kernel mode is executing in ring zero. It has the highest privilege. And anything in user mode is running in ring three. Whenever we want to access privileged resources, we call system call instructions that change the context of the CPU. And so there are two attack vectors in a Linux system. First is getting root. In getting root, you're still in ring three, but you're the highest privileged user on the system. So you can actually pretty much do anything on the system. You can read your SSH keys, encryption data, et cetera. And then the second version is actually running a kernel mode exploit, so running code in ring zero. So this is terrible. You can actually dump processes, memories, and I don't even want to say what all can happen. So you can get root, and you can have kernel exploits, and if you get kernel exploit, it's a New York Times article waiting to happen. So these are the two attack vectors on a Linux system. So with that background, let's design the simplest way to execute tools on a Linux system. For example, we can literally have an API server that your harness is calling, and for every tool call, it can fork a process, exec the tool that the model needs, and just have a fork and exec model. So now there are a couple of problems with this model. A, as we discussed the execution model, and if you get kernel exploit, it's a New York Times article waiting to happen. So these are the two attack vectors on a Linux system. So with that background, let's design the simplest way to execute tools on a Linux system. For example, we can literally have an API server that your harness is calling, and for every tool call, it can fork a process, exec the tool that the model needs, and just have a fork and exec model, right? So now there are a couple of problems with this model, right? A, as we discussed the execution model, the fork process can now directly talk to the kernel, and so the model can try to attack the kernel, get root, or try to get a kernel exploit, right? The second thing is, imagine you have a while loop and just forking processes in one tool call, right? So now you've become a bad neighbor or a noisy neighbor, and you've brought down the node, and no other tool calls can run right. So fork exec is the simplest thing you can do. It has one thing going for it. It's the most performant solution because it's as fast as just forking something. It's native performance, but everything else is very bad about this solution. So we saw that fork and execing was bad. There was a noisy neighbor issue, and there's a security issue. We now turn to something called containers on Linux. You must have heard of Docker and other things associated with the containers, but remember, this is the first principles talk, so let's go deep into what containers are in their very raw form. Containers rely on two concepts on Linux, namely namespaces and C groups. Namespaces are for resource isolation, and C groups are for controlling the amount of resources a set of processes can consume, right? So let's see in depth how they can help our problems. So in this diagram, you can see that there are different types of namespaces a set of processes can have. So on the left, you can see this container has a PID namespace. So inside that process namespace, it looks like we have a process hierarchy of PIDs 1, 2, and 3, but if you look at the PIDs from outside the container, they're just regular processes with other PIDs, right? So you've abstracted one resource, which is the process inside this container. Another example is a mount namespace. So if you see the bottom half of the diagram, you can mount any file system on top of any mount inside the container, but from outside, the original mount is still visible, right? So this is a nice view of isolating different resources in a container, and there are many namespaces like PID, network, mount, et cetera. So moving on to the second principle of containers, right, C groups. So now we have a way to isolate resources within containers via namespaces, but remember, I gave the example of someone doing a while with fork, right? How can we control that? Well, we can have C groups and control how much CPU and memory a container can consume, and that way we can make sure that one container doesn't bring down the entire node. So namespaces and C groups provide some amount of isolation and control on the system resources to not bring it down in case it's malicious or bad. If you've seen the previous diagrams, the fundamental problem with containers is that there are still native processes running on the host. So even if we abstract resources and two containers can't attack each other because their resources are isolated, a process in a container can still exploit the kernel boundary and try to get root or a kernel exploit, right? And once they can get root, they can attack other people's data from other sandboxes, exfiltrate data, and try to get control of the node like we discussed before. So security is a spectrum. You can still get a decent amount of protection from containers. One example to do this is something called seccomp. So you can actually have a filter on the amount of system calls your container can call, and you can also control the arguments each system call can take. So you can say that you're reducing the attack surface of the kernel for this container and have some amount of sanity on how it can attack the kernel. The problem with this is that many times you don't know beforehand what system calls some container might call, right? So now you're blocking requests for users, and later on you have to change the seccomp filter to allow something or not allow something. So the feedback loop is pretty bad for a product like OpenClaw or any agent that wants to do crazy things, right? You want to make magical experiences happen in the sandbox, and you don't want to restrict them, right? So, yeah, to sum it up, containers interact with the same host kernel, so they do have some protections, but at the end, it's the same host kernel they're trying to attack, right? And they can get root and kernel exploits, et cetera. So we went from forking to containers. Raw fork was obviously very bad for all the reasons. Containers are better than fork because they still have some protection, and you can still have seccomp for reducing the attack surface. Can we do better? Can we reduce the host kernel exposure a little bit more? Yeah, so the second thing we can discuss is to solve the attack surface problem is something called GVisor. The crux of the security boundary is the kernel API from before. So GVisor uses this point Raw fork was obviously very bad for all the reasons. Containers are better than fork because they still have some protection, and you can still have seccomp for reducing the attack surface. Can we do better? Can we reduce the host kernel exposure a little bit more? Yeah, so the second thing we can discuss to solve the attack surface problem is something called GVisor. The crux of the security boundary is the kernel API from before. So GVisor uses this point as its key security story. So GVisor implements a lot of syscalls in user space. You can think of it as an application kernel. Its sentry is a user space kernel written in Go, and the file system is accessed by another daemon called the Gopher. So the sentry implements the Linux API itself, including process management, workload management, and whenever you call a system call, it's intercepted and it's serviced by this user space program. So then you can argue that you're not implementing an attacking code that's running in ring zero or kernel mode. You're actually running in user space and ring three, and any exploit you have is still in user space. So it's better than you exploiting the kernel directly. But the fundamental problem is that the sentry and the Gopher, just going back, the sentry and the Gopher are still on top of the host kernel. So if there is an exploit on them, you can still have a chained two-step exploit. So you first exploit a problem in the sentry or the Gopher, and then you exploit from the Gopher to the kernel, right? You can still get to the host kernel eventually. And with these models of 5.6 and other types, you can have them figuring out bug reports and these things and trying to chain exploits, right? So you can still get to the host kernel, right? So, yeah, the chain is harder here because it's a two-step chain compared to the others, but it's still reachable. Can we do better than this? Yeah, that's the same slide. It's saying that in all of these solutions, you can eventually get to the host kernel, whether it's fork, Gvisor, or containers. So the key question is, can we have a way in which untrusted or malicious code can have exploits, but they can never or find it very hard to exploit the host, right? So even if I can get root, even if I can exploit the kernel, I still want my host to be protected, right? So that's the final goal we have. So do we have something that we can use for this? So if you've used, if you've known some of my background, you'll know where I'm going with this. So turns out Linux provides a very nice thing called virtualization, which is hardware-powered at the CPU level. And it provides an abstraction at the hardware level, which means that even if you can get root or execute in ring zero in the guest, your host is still protected. And this happens because the guest kernel runs in ring zero, but in a separate processor context called VMX non-root, while the host kernel and the hypervisor run in ring zero in something called VMX root mode. So ring zero gives the guest kernel full control inside the guest, but no control on the host. So you can exploit the guest all you want, but the host is still protected. The processor has to switch between the guest and the host whenever the guest wants to access some privileged resource. We'll see how that works. But that is a key trade-off here. There's a performance penalty you pay every time the CPU is switching back and forth between these two modes. Let's see with the diagram what I mean. So you can see in the old diagram, we showed ring zero and ring three for executing user space and kernel code. But here, the guest kernel and the guest user space are running in separate CPU context. So this is guaranteed at a hardware level. And so even if you can get to ring zero in the guest, your host is protected from it. Yeah. Now let's talk about what does para-virtualize and hardware-based virtualization mean. This is the key of all the VM sandboxes you might have seen on Hacker News or Reddit. So we go from step by step, right? So on the right-hand side, you see something called the VMM. Can anyone see that on the right-hand side? The VMM stands for Virtual Machine Monitor. This software is called the Virtual Machine Monitor, and you might have heard of QEMU when you search for Linux virtualization. of all the VM sandboxes you might have seen on Hacker News or Reddit. So we go from step by step, right? So on the right-hand side, you see something called the VMM. Can anyone see that on the right-hand side? The VMM stands for Virtual Machine Monitor. This software is called the Virtual Machine Monitor, and you might have heard of QEMU when you search for Linux virtualization. QEMU and other VMMs, their sole job is to talk to dev-kvm. So you see that arrow going down. The dev-kvm is the hypervisor API of the Linux kernel. All QEMU and other VMMs do is set up the kernel root fs, set up some, allocate some memory for this guest virtual machine and call into the dev-kvm API. As discussed inside, when the guest runs on the left-hand side, it's running as a completely different Linux system. It doesn't know what's happening on the host side, but it has these devices, like a block device or a net device, coming up as regular Linux devices, but when it actually tries to access the devices, they exit out to the host context, and if you see the block backend and the network backend, there's actually code running, emulating the devices in the host. So whenever the guest exits out, it's being serviced by the VMM process and these device processes. So what do we mean by para-virtualization? So para-virtualization means that we want good performance while we want to access hardware in the guest. So the guest drivers, when the user space talks to them, the guest drivers are aware that they're running in a virtual machine, and so they talk via something called virtio, which is a more efficient way of the guest and the host talking to each other. And so inside, there are just PCI devices like any other hardware, and when you talk to the PCI devices, you magically exit out onto the host. So 15, 20 years ago, a bunch of Linux wizards made this thing happen and made this performant, and it truly is magical how it works reliably and performantly across different hardware processors. And so from the host's point of view, when this is running, you just see a block thread. If you do a PS, you'll see a block thread, but that's running on the guest context. And when the thread exits out, the thread finally wakes up. It wakes up in this VMM process, and the process runs, and it's running inside. Yeah, so we had this seismic shift in 2023 where a bunch of new VMMs came. For a long time, it was just QEMU that was used for running virtual machines. But QEMU has a lot of craft. It supports many, many architectures. It has many devices, and it's written in C. And historically, many, many escape attacks were attacking the devices written in C. So the first Rust-based VMM was something called CrossVM, which our team at Google wrote when I was there to support Linux virtual machines on top of Chromebooks. And the key point here was that we don't need all the craft of QEMU, and we can use Rust to be a memory-safe implementation of this tricky system software. And secondly, we have these emulated devices that we can jail. So if you attack the block device, we've only given it permission to access block resources. So you can't access network resources. Similarly, if you attack the net device, we've jailed it, and we don't give it access to block resources. So then you can still have a second gate of security. So both Rust-based safety, like being memory-safe, and just jailing at a more granular level at the device, provides better safety than just QEMU. So you must have heard this word micro-VMs everywhere, and no one really answers why the word micro comes. So turns out it has nothing to do with what's running inside the guest. It's everything to do with the VMM itself. So all these new-age Rust-based VMMs, they have a much smaller memory footprint because they don't support as many devices, provides better safety than just QEMU. So you must have heard this word micro-VMs everywhere, and no one really answers why the word micro comes. So it turns out it has nothing to do with what's running inside the guest. It's everything to do with the VMM itself. So all these new-age Rust-based VMMs have a much smaller memory footprint because they don't support as many devices, and they also boot much faster because they don't have as much cruft. So a mixture of just less bloat and just booting up faster is why the industry has called them micro-VMs. One second. Yeah, sorry. And then you must have seen Firecracker and Cloud Hypervisor. There's a lot of confusion around who came before. CrossVM was the first Rust-based VMM that came before, and then Firecracker forked CrossVM, and it's used on Amazon for their Lambda and serverless load. And Cloud Hypervisor is a more general VMM that many, many companies contribute to, and when you historically see micro-VMs on the internet, it's powered by one of these VMMs. So it may seem complicated, but at the end, everything is APIs. Micro-VMs are no different. So here's a small view of how you can actually start a micro-VM. So your harness actually forks a Cloud Hypervisor binary process. When the process starts, it exposes an API over a Unix domain socket, and if you can see, in step two, we call the create API and give it the rootFS, kernel, CPU, and memory, and then we finally call start. And when we call start, you can see from previous diagrams that the VMM literally calls into dev-KVM, just like before, and it starts these guest micro-VMs. And then once the guest is running, in this case, the agent sandbox is running, it might be that I want to talk to something inside the sandbox to say that, hey, save your state, or I'm attaching some devices, or I'm doing some XYZ operation. So generally, in agent sandboxes, you have a PID one that exposes an API server, and your harness or something on the outside is talking to this API server. And then this is the way how you can use micro-VMs to be agent sandboxes that can run more securely than the other primitives that we showed. In this case, we are using vsoc, which is a socket that can help you communicate between the guest and the host, or you can use the IP stack on the node itself. And so there are no free lunches in systems, and so there are some trade-offs. So obviously, you get really good isolation using micro-VMs at a hardware level. You can still attack the host, but it's much harder. You have to attack the KVM stack, and then you have to attack the device. The chain is much harder to do, but it has been seen, and I'm sure it will be seen more with these models. You can jail the devices generally, so you can use seccomp and other security-hardening things for the devices. So even if you get compromised on one device, your whole system cannot be brought down. Like I said, there is a performance overhead. When you're exiting and entering the host and the guest context, it's a very heavy operation, and you pay in performance. Memory sharing is not as easy. There's something called a balloon driver, and so you have to actually ask the guest to give back memory and reclaim it. So it's always a reactive thing. You can't immediately reclaim and claim memory. And then a lot of sandboxes these days have GPU access, presumably for auto research or some sort of ML research type of agents. is not as easy. There's something called a balloon driver, and so you have to actually ask the guest to give back memory and reclaim it. So it's always a reactive thing. You can't immediately reclaim and claim memory. And then a lot of sandboxes these days have GPU access, presumably for auto research or some sort of ML research type of agents. This is not as easy with micro VMs. There's something called VertIO GPU, but that provides high-level graphic library type access. For direct metal access, there's something called VFIO, but it can only be shared by one sandbox at a time. It cannot have multi-tenant things. But given all of this, my view is security, like system tricks can cover performance issues, but they cannot hide security breaches. And as a company, you can lose trust once, and it's very hard to regain. So I always prefer the more secure solution and try to make up with system tricks for performance issues. And in my history of working on sandboxes, I've seen there are seven stages of grief, the seven stages of sandboxing. Like in the end, everyone always wants a VM because they tried everything. They tried containers, GVisor, V8, and then they realized, oh, I want a whole Linux box because I don't want to end up without XYZ functionality. And if I want a whole Linux box, I want to obviously be secure. So if you're a startup or a founder in this space, let me save you the story and two years of grief. Just please use micro VMs from the start. And then if it doesn't work, tell me and then we can talk about other things. So now we've discussed how to run untrusted code on one node. And I think now people have woken up to the fact that these models are very, very good drivers of Linux boxes. Like so if you give them a computer, they can just pretty much do magical things as we've seen with OpenClaw. They just pre-train on a lot of Linux, right? However, imagine if I gave you a computer without a disk. Every time you close the laptop, your data and your work goes away, right? That's not a fun world to live in. And somehow the agents, some in the cloud at least, are in this sort of world right now, right? So we want to give them durable storage. And so this part of the presentation is specifically working on disk storage, not memory persistence, but disk persistence. So let's see why it's important and how we can do it. So as shown in previous diagrams, your micro VMs have disks attached to them. There might be just files or other disk devices on the node that you pass through to the VM. And so from a product perspective, the tasks that the users are doing in these sandboxes are becoming much more complicated and much more longer horizon. So people are making presentations and entire GitHub repos are being created inside the sandbox. Now imagine if the node dies on the cloud or the model has a flake and you created this presentation and you just lost it, right? It's bad for us because we wasted a bunch of GPU tokens. It's obviously bad for the user because you did a lot of this work and you lost it. So not just from a product perspective, but from a good experience and utilization perspective, we need to have some way to save the disk state of the sandbox. And let's go into three big use cases on what persistence can unlock, right? So counterintuitively, persistence actually helps reliability and scale. They might seem like orthogonal concepts, but they're very much related. So for instance, you have a long running task in a sandbox that has many packages installed and you've created GitHub repos and presentations and things, right? If you keep checkpointing it periodically and if the node fails or the cluster fails, you can now restore the sandbox in the exact checkpoint state on another node. You can also do it intentionally if you want to upgrade a cluster or do some A-B testing. You have a long running task in a sandbox that has many packages installed and you've created GitHub repos and presentations and things, right? If you keep checkpointing it periodically and if the node fails or the cluster fails, you can now restore the sandbox in the exact checkpoint state on another node. You can also do it intentionally if you want to upgrade a cluster or do some A-B testing of nodes. The persistence has now let you reliably run sandboxes across your fleet, right? So it's a very good way to scale and be reliable. As I said, how many of you have used goal mode in Codex or know what it is? Yeah, amazing, right? So if you run goal mode, now it's my, I think three days is my record for running something, but now people are doing longer and longer tasks and I think this trend will continue in the cloud as well. So to support this, we obviously need to have checkpointing so that the model can save state, restore it on another node and you can keep going forward and forward, right? And so you're resilient to any failures across the infrastructure. This is the most interesting part actually, right? So if your harness wants to explore multiple solutions or sample spaces, it can actually checkpoint the sandbox state and it can do a Monte Carlo search and go ahead and backtrack, checkpoint again. So this way it can actually do rollouts over many, many days and come back with the actual solution, right? And I think my sincere hope is that if we make this infrastructure correct, we can help solve diseases, find new drugs because the model can just keep going on for longer and longer, right? But it doesn't happen until we have really, really rock solid primitives to do this. And so given the needs and the pillars we want to support, here are some of the things that this snapshotting solution should support, right? First, at ChatGPT or Codex scale, we want to do incremental snapshotting. And what that means is if you call snapshot twice, I'm just snapshotting the diff between the two snapshots. Otherwise, if I have to save gigabytes of data at every turn, I'm going to bankrupt the company and it's just a slow experience regardless, right? The snapshotting API itself should be very, very cheap and fast so the model and the harness can keep snapshotting and exploring very fast. And similarly, just like creation should be fast for products, restoring is just nothing but just recreation from a snapshot and so restoring should also be very, very fast for a good product experience. And then we have two paradigms we'll discuss. One is always saving in which the harness doesn't have to explicitly call a save API versus explicit saving where the harness is calling save, save, save. We'll see how we can implement both so I'll give a reference solution. And then there are some more design choices with disk snapshotting, right? So first, as I said, on the left-hand side, do you want incremental or full snapshots? I argue, I think at our scale we want incremental snapshotting. Secondly, do you want to snapshot the entire rootFS or do you want to have certain folders like workspace or mount something that you want to keep that configurable? And lastly, we can see how you can snapshot at a file system level or at a block level. So in Linux, disks are nothing but block devices and each file maps to different blocks. Or do you want to have certain folders like workspace or mount something, something that you want to keep that configurable? And lastly, we can see how you can snapshot at a file system level or at a block. So in Linux, disks are nothing but block devices and each file maps to different blocks. We'll see after this. And so you can do very, very efficient snapshotting by just zipping up the blocks that have changed for a file or you can do entire files that have changed for more write amplification. And the high-level flow of snapshotting is you figure out what's changed, you zip it up, put it to the cloud, and then when you restore, you pull it down and you restore the micro VM, right? And so this is a first principles talk and so I just wanted to go over how Linux storage works from first principles. So Linux represents disks as block devices. So you see on the right, it thinks of a block device as having logical blocks from zero to N. A file system maps directories and files into something called an inode data structure and the inode says, okay, offset zero in file F maps to this block, logical block in disk D and then on the disk itself, there's firmware running which says, oh, logical block one is sector 100 or sector 500 or page XYZ. So the hierarchy is on these logical blocks up to the file system and we leverage that for block based snapshotting. And then within a micro VM there are two ways of accessing storage. On the left hand side, you can think of this as sharing a folder like Google Drive but it's very, very inefficient because the VM is doing file system operations, exiting at every file system operation which is very inefficient. On the right hand side, you actually give a disk-like abstraction at a block. So you give a block device to the micro VM and it is way more efficient because you can use the caches inside the guest and you don't have to exit out as much. You only exit out when you truly need to access the block device and the host has to service you. This is the high-level diagram of how you will have always-on persistence in disk snapshotting. So if you have two micro VMs, they have a block device which they will see this dot image as a block device inside and as they are writing to the block device, we are writing through to the cloud. So there's some distributed file system that we mount inside that's giving this always-on persistence. And like I mentioned, for explicit persistence, we have an actual API called the save API. Calling the save API from the harness figures out what's changed between the last snapshot, it bundles up this diff into an artifact and returns a snapshot ID. Later on, you can give us this snapshot ID. We figure out the lineage of snapshots that make this snapshot ID and then we download them one by one and apply it on the node and you get a micro VM restored with this thing. So let's see how we can implement this. Give me one second. I want to see how we're doing on time. We have five minutes left, so we'll go fast. Yeah, so we can do explicit persistence using something called copy-on-write. So Linux has these XFS file systems, XFS-like file systems, so you How we can implement this. Give me one second. I want to see how we're doing on time. We have five minutes left, so we'll go fast. Yeah, so we can do explicit persistence using something called copy-on-write. So Linux has these XFS file systems, XFS-like file systems, so you can do pretty much have zero latency copies because you don't change any blocks when you copy, but when you actually change the blocks on a file, that's when you pay the penalty. And so in this design, we have a base dot image, which might be the codex base image or whatever, the chat GPT base image. We create a zero copy on top of it, which is a writable layer. And then when you write to it, now you're changing the blocks in this layer. And then when you want to snapshot, I use something called FIE map, which tells me what blocks have changed and what ranges. I zip that up and store it to the cloud. And I can do something very nifty here. I can actually lie to you while I'm uploading to the cloud. So the snapshot can happen, return very fast as I'm uploading in the background. I don't have to wait till I'm uploading. And then on the other side, I have this diff, I download this artifact, I figure out what extents have changed, and then I apply it back on top of the base image and I start the micro VM again. So now we have a restored sandbox with the exact same state at a block level. Now how can we do always-on persistence? So this is one way of doing this. I know there's NFS also and other distributed file systems, but NFS, for instance, isn't as performant and it's not POSIX compliant. And I think our models are just very good at anything POSIX compliant and standard. So you can write a file system actually on top of a GCS or S3 or durable block storage, and you can use something called NBD. So within the sandbox, you'll actually see a block device, but inside it will have literally a tiered cache of blocks that are persisted first to an in-cluster cache, and then the in-cluster cache is writing back to the block storage, the object storage. So you have this nice global tiered architecture where you're actually caching things at the block level and finally inside the micro VM and get a very performant file system inside. Yeah, so that's the persistence part of the presentation. And so the one takeaway I want you guys to think about is I think storage is the next unlock here. As you're working on sandboxes, think of what all you can snapshot and restore fast to give this new paradigm to harnesses so they can recover from failures and explore, do Monte Carlo-like searches. So we've discussed running things on one node, and of course, if one node dies, we are done. So we want to be able to run across many, many nodes across the world. So I'm not going to mention about Kubernetes or other acquisition things. This is the first principles talk, so we'll lightly hint about the challenges here. So ideally, we want multiple machines to support the runtime we discussed. So we can group nodes into clusters and spread the clusters across the regions. A top-level control plane chooses a cluster using region, load, and other factors, and it's not very different from the orchestrators you are familiar with. These, ideally it uses a cluster About the challenges here. So ideally, we want multiple machines to support the runtime we discussed. So we can group nodes into clusters and spread the clusters across the regions. A top-level control plane chooses a cluster using region, load, and other factors, and it's not very different from the orchestrators you are familiar with. Ideally it uses a cluster close to your ChatGPT cluster, so you can have fast access to the harness. Inside the cluster there's a scheduler also, and the scheduler tells you which node to pick based on the load and other factors, right? So if nodes are dying or failing, it won't choose that or will intelligently route the sandbox to the thing. And again, low latency and reliability remain very, very key north stars for this architecture. And so here we can use some micro-VM features to support low latency creation ideas. So a lot of systems cheat for low latency. They pre-warm sandboxes and they pick one, which is great. Another way to do this is you can actually take a memory snapshot of a micro-VM and just in time start it in milliseconds as the request comes. And so you can leverage this nice micro-VM property that you can save the guest memory and start from that. And the third one is a hybrid solution. So you can have a warm pool, but as it's growing, you can grow it from the memory snapshot so you can get best of both worlds. The trade-off for a warm pool is that you're consuming CPU and memory in idle state. Ideally, you don't want to do that, right? So there's a trade-off between one and three and two. And so here is a way where we can use snapshot to restore for better orchestration. So remember, we discussed that a snapshot can have a lineage of many, many layers. So once you want to restore from a snapshot and you find out you have, let's say four layers that you want to pull down and that makes the lineage, you can actually smartly route you to a node which has to download the least amount of stuff. So in this diagram, you can see node A has some layers, node C has some layers, but node B has all the layers that you need. So the scheduler then routes you and it gives it the highest score and routes it because it has all the snapshot layers. So you can use snapshot with orchestration to have faster creates and even more reliable orchestration. Yeah. This is the talk and hopefully it gives you some design intuition around sandboxes, why they're important, what's the next unlock, and I want to see more of you guys using it in a secure way. Thank you. exec the tool that the model needs, and just have a fork and exec model, right? So now there are a couple of problems with this model, right? A, as we discussed the execution model, the fork process can now directly talk to the kernel, and so the model can try to attack the kernel, get root, or try to get a kernel exploit, right? The second thing is, imagine you have a while loop and just forking processes in one tool call, right? So now you've kind of become a bad neighbor or a noisy neighbor, and you've kind of brought down the node, and no other tool calls can run right. So fork exec is the simplest thing you can do. It has one thing going for it. It's the most performant solution because it's as fast as just forking something. It's native performance, but everything else is very bad about this solution. So, like, we saw that fork and execing was bad. There was a noisy neighbor issue, and there's a security issue. We now turn to something called containers on Linux. You must have heard of Docker and other things associated with the containers, but remember, this is the first principles talk, so let's go deep into what containers are in their very raw form. Containers rely on two concepts on Linux, namely namespaces and C groups. Namespaces are for resource isolation, and C groups are for controlling the amount of resources, a set of processes can consume, right? So let's see in depth how they can help our problems. So in this diagram, you can see that there are different type of namespaces a set of processes can have. So on the left, you can see this container has a PID namespace. So inside that process namespace, it looks like we have a process hierarchy of PIDs 1, 2, and 3, but if you look at the PIDs from outside the container, they're just regular processes with other PIDs, right? So you've abstracted one resource, which is the process inside this container. Another example is a mount namespace. So if you see the bottom half of the diagram, you can mount any file system on top of any mount inside the container, but from outside, the original mount is still visible, right? So this is a nice view of, like, isolating different resources in a container, and there are many namespaces like PID, network, mount, et cetera. So moving on to the second principle of containers, right, like C groups. So now we have a way to isolate resources within containers via namespaces, but remember, I gave the example of someone doing a while with fork, right? Like, how can we control that? Well, we can have C groups and control how much CPU and memory a container can consume, and that way, like, we can, like, make sure that one container doesn't bring down the entire node. So namespaces and C groups provide some amount of isolation and control on the system resources to not bring it down in case it's malicious or bad. If you've seen the previous diagrams, the fundamental problem with containers is that there are still native processes running on the host. So even if we abstract resources and two containers can't attack each other because their resources are isolated, but a process in a container can still exploit the kernel boundary and try to get root or a kernel exploit, right? And once they can get root, they can attack other people's data from other sandboxes, exfiltrate data, and try to get control of the node like we discussed before. So security is a spectrum. Like, you can still get a decent amount of protection from containers. One example to do this is something called seccomp. So you can actually have a filter on the amount of system calls your container can call, and you can also control the arguments each system call can take. So you can say that you're reducing the attack surface of the kernel for this, like, container and, like, kind of have some amount of sanity on how it can attack the kernel. The problem with this is that many times you don't know beforehand what system calls some container might call, right? So now you're blocking requests for users, and later on, like, you have to, like, change the seccomp filter to allow something or not allow something. So the feedback loop is, like, pretty bad for a product, like OpenClaw or any agent that wants to do, like, crazy things, right? Like, you want to make magical experiences happen in the sandbox, and you don't want to restrict them, right? So, yeah, this is the fact, like, to sum it up, like, containers interact with the same host kernel, so they do have some protections, but at the end, it's the same host kernel they're trying to attack, right? And they can get root and kernel exploits, et cetera. So we went from, like, forking to containers. Like, raw fork was obviously very, very bad for all the reasons. Containers are better than fork because they still have some protection, and you can still have seccomp for reducing the attack surface. Can we do better? Can we reduce the host kernel exposure a little bit more? Yeah, so the second thing we can discuss is to solve the attack surface problem is something called GVisor. The crux of the security boundary is the kernel API from before. So GVisor uses this point as its key security story. So GVisor implements a lot of syscalls and user space. You can think of it as an application kernel. Its sentry is a user space kernel written in Go, and the file system is accessed by another daemon called the Gopher. So the sentry implements the Linux API itself, including process management, workload management, and whenever you call a system call, it's intercepted and it's, like, serviced by this user space program. So then you can argue that you're not implementing an attacking code that's running in ring zero or kernel mode. You're actually running in user space and ring three, and any exploit you have is still in user space. So it's better than you exploiting the kernel directly. But the fundamental problem is that the sentry and the Gopher, just going back, the sentry and the Gopher are still on top of the host kernel. So if there is an exploit on them, you can still have a chained two-step exploit. So you first exploit a problem in the sentry or the Gopher, and then you exploit from the Gopher to the kernel, right? You can still get to the host kernel eventually. And with these models of, like, 5.6 and other types, like, you can have, like, them, like, figuring out bug reports and these things and, like, trying to chain exploits, right? So you can still get to the host kernel, right? So, yeah, the chain is harder here because it's a two-step chain compared to the others, but it's still reachable. Like, can we do better than this? Yeah, that's the same slide. It's saying that in all of these solutions, you can eventually get to the host kernel, whether it's fork, Gvisor, or containers. So the key question is, can we have a way in which untrusted or malicious code can have exploits, but they can never or find it very, very hard to exploit the host, right? So even if I can get root, even if I can exploit the kernel, I still want my host to be protected, right? Like, that's the final goal we have. So do we have something that we can use for this? So if you've used, if you've, like, known some of my background, you'll know where I'm going with this. So turns out Linux provides a very, very nice thing called virtualization, which is hardware-powered at the CPU level. And it provides an abstraction at the hardware level, which means that even if you can get root or execute in, like, ring zero in the guest, your host is still protected. And this happens because the guest kernel runs in ring zero, but in a separate processor context called VMX non-root, while the host kernel and the hypervisor run in ring zero in something called VMX root mode. So ring zero gives the guest kernel full control inside the guest, but no control on the host. So you can exploit the guest all you want, but the host is still protected. The processor has to switch between the guest and the host whenever the guest wants to access some privileged resource. We'll see how that works. But that is a key trade-off here. There's a performance penalty you pay every time the CPU is switching back and forth between these two modes. Let's see with the diagram what I mean. So you can see in the old diagram, we showed ring zero and ring three for executing user space and kernel code. But here, the guest kernel and the guest user space are running in separate, like, CPU context. So this is guaranteed at a hardware level. And so if you can, even if you can get to ring zero in the guest, like, your host is, like, protected from it. Yeah. Now let's talk about what does para-virtualize and hardware-based virtualization mean. This is the key of all the VM sandboxes you might have seen on Hacker News or Reddit. So we go from step by step, right? So on the right-hand side, you see something called the VMM. Can anyone see that on the right-hand side? The VMM stands for Virtual Machine Monitor. This software is called the Virtual Machine Monitor, and you might have heard of QEMU when you search for Linux virtualization. QEMU and other VMMs, their sole job is to talk to dev-kvm. So you see that arrow going down. Like, the dev-kvm is the hypervisor API of the Linux kernel. All QEMU and other VMMs do is set up the kernel root fs, set up some, allocate some memory for this guest virtual machine and call, like, into the dev-kvm API. As discussed inside, when the guest runs on the left-hand side, it's running as a completely different Linux system. It doesn't know what's happening on the host side, but it has these devices, like a block device or a net device, coming up as regular Linux devices, but when it actually tries to access the devices, they exit out to the host context, and if you see the block backend and the network backend, there's actually code running, emulating the devices in the host. So whenever the guest exits out, it's being serviced by the VMM process and these device processes. So what do we mean by para-virtualization? So para-virtualization means that we want good performance while we want to access hardware in the guest. So the guest drivers, when the user space talks to them, the guest drivers are aware that they're running in a virtual machine, and so they talk via something called virtio, which is a more efficient way of the guest and the host talking to each other. And so inside, there are just PCI devices like any other hardware, and when you talk to the PCI devices, you magically exit out onto the host. So 15, 20 years ago, a bunch of Linux wizards made this thing happen and made this performant, and it truly is magical like how it works reliably and performantly across different hardware processors. And so from the host's point of view, when this is running, you just see a block thread. Like if you do a PS, you'll see a block thread, but that's running on the guest context. And when the thread exits out, the thread finally wakes up. It wakes up in this VMM process, and the process runs, and it's running inside. Yeah, so we had this seismic shift in 2023 where a bunch of new VMMs came. For a long time, it was just QEMU that was used for running virtual machines. But QEMU has a lot of craft. It supports many, many architectures. It has many devices, and it's written in C. And historically, many, many escape attacks were attacking the devices written in C. So the first Rust-based VMM was something called CrossVM, which our team at Google wrote when I was there to support Linux virtual machines on top of Chromebooks. And the key point here was that we don't need all the craft of QEMU, and we can use Rust to be a memory-safe implementation of this tricky system software. And secondly, we have these emulated devices that we can jail. So if you attack the block device, we've only given it permission to access block resources. So you can't access network resources. Similarly, if you attack the net device, we've jailed it, and we don't give it access to block resources. So then you can still have a second gate of security. So both Rust-based safety, like being memory-safe, and just jailing at a more granular level at the device, like provides better safety than just QEMU. So you must have heard this word micro-VMs like everywhere, and no one really answers like why the word micro comes. So turns out it has nothing to do with what's running inside the guest. It's everything to do with the VMM itself. So all these new-age Rust-based VMMs, they have a much smaller memory footprint because they don't support as many devices, and they also boot much faster because they don't have as much cruft. So a mixture of like just less bloat and just booting up faster is why the industry has called them micro-VMs. One second. Yeah, sorry. And then you must have seen Firecracker and Cloud Hypervisor. There's a lot of confusion around who came before, like CrossVM was the first like Rust-based VMM that came before, and then Firecracker forked CrossVM, and it's used on Amazon for their Lambda and serverless load. And Cloud Hypervisor is a more general like VMM that many, many companies like contribute to, and when you historically see micro-VMs on the internet, like it's powered by one of these VMMs, basically. So it may seem complicated, but at the end, everything is APIs. Micro-VMs are no different. So here's a small view of how you can actually start a micro-VM. So your harness actually forks a Cloud Hypervisor binary process. When the process starts, it exposes an API over a Unix domain socket, and if you can see, in step two, we call the create API and give it like the rootFS, kernel, CPU, and memory, and then we finally call start. And when we call start, you can see from previous diagrams that the VMM literally calls into dev-KVM, just like before, and it starts these like guest micro-VM. And then once the guest is running, in this case, the agent sandbox is running, it might be that I want to talk to something inside the sandbox to say that, hey, save your state, or I'm attaching some devices, or I'm doing some XYZ operation. So generally, in agent sandboxes, you have a PID one that exposes an API server, and your harness or something on outside is talking to this API server. And then this is the way how you can use like micro-VMs to be agent sandboxes that can run like more securely than the other primitives that we showed. in this case, we are using vsoc, which is a socket that can help you communicate between the guest and the host, or you can use the IP stack on the node itself. And so there are no free lunches in systems, and so there are some trade-offs. So obviously, you get really, really good isolation using micro-VMs at a hardware level. You can still attack the host, but it's much harder. You have to attack the KVM stack, and then you have to attack the device. The chain is much, much harder to do, but it has been seen, and I'm sure it will be seen more with these models. You can jail the devices generally, so you can use seccomp and other security-hardening things for the devices. So even if you get compromised on one device, your whole system cannot be brought down. Like I said, there is a performance overhead. Like when you're exiting and entering the host and the guest context, it's a very, very heavy operation, and you pay in performance. Memory sharing is not as easy. There's something called a balloon driver, and so you have to actually ask the guest to give back memory and reclaim it. So it's always a reactive thing. You can't immediately reclaim and claim memory. And then a lot of sandboxes these days have GPU access, presumably for auto research or some sort of like ML research type of agents. This is not as easy with micro VMs. There's something called VertIO GPU, but that provides high-level graphic library type access. For direct metal access, there's something called VFIO, but it can only be shared by one sandbox at a time. It cannot have multi-tenant things. But given all of this, my view is security, like system tricks can cover performance issues, but they cannot hide security breaches. And as a company, you can lose trust once, and it's very hard to regain. So I always prefer the more secure solution and try to make up with system tricks for performance issues. And in my history of working on sandboxes, I've seen there are like, I would call it the seven stages of grief, the seven stages of sandboxing. Like in the end, everyone always wants a VM because they tried everything. They tried containers, GVisor, V8, and then they realized, oh, I want a whole Linux box because I don't want to like end up without XYZ functionality. And if I want a whole Linux box, I want to obviously be secure. So if you're a startup or a founder, like in this space, like let me save you the story and two years of grief. Just please use micro VMs from the start. And then if it doesn't work, tell me and then we can talk about other things. So now we've discussed how to run untrusted code on one node. And I think now like people have woken up to the fact that these models are very, very good drivers of Linux boxes. Like so if you give them a computer, they can just pretty much do magical things as we've seen with OpenClaw. They just pre-train on a lot of Linux, right? However, like imagine if I gave you a computer without a disk. Every time you close the laptop, like your data and your work goes away, right? Like that's not a fun world to live in. And somehow the agents are some in the cloud at least are in this sort of world right now, right? So we want to give them durable storage. And so this part of the presentation is specifically working on disk storage, not memory persistence, but disk persistence. So let's see why it's important and how we can do it. So as shown in previous diagrams, your micro VMs have disks attached to them. there might be just files or other disk devices on the node that you pass through to the VM. And so from a product perspective, the tasks that the users are doing in these sandboxes are becoming much more complicated and much more longer horizon. So people are making like presentations and entire GitHub repos are being created inside the sandbox. Now imagine if like the node dies on the cloud or the model has a flake and you created this like presentation and like you just lost it, right? It's A, bad for us because we wasted a bunch of GPU tokens. It's obviously bad for the user because you did a lot of this work and you lost it. So not just from a product perspective, but just from like a good experience and like utilization perspective, we need to have some way to save the disk state of the sandbox. And let's go into like three big use cases on what persistence can unlock, right? So counterintuitively, persistence actually helps reliability and scale. They might seem like orthogonal concepts, but they're very much related. So for instance, you have a long running task in a sandbox that has many packages installed and you've created like GitHub repos and presentations and things, right? If you keep checkpointing it periodically and if the node fails or the cluster fails, you can now restore the sandbox in the exact checkpoint state on another node. You can also do it intentionally if you want to upgrade a cluster or do some A-B testing or nodes. The persistence has now let you reliably run sandboxes across your fleet, right? So it's a very good way to scale and be reliable. Like I said, like how many of you have used goal mode in Codex or know what it is? Yeah, amazing, right? So if you run goal mode, now it's like my, I think three days is my record for running something, but like now people are doing longer and longer tasks and I think this trend will continue in the cloud as well. So to support this, we obviously like need to have checkpointing so that the model can save state, restore it on another node and like you can keep going forward and forward, right? Like, and so you're resilient to any failures across in the infrastructure. This is the most interesting part actually, right? Like, so if your harness wants to explore multiple like solutions or sample spaces, it can actually checkpoint the sandbox state and it can like do a Monte Carlo like research and like go ahead and like backtrack, checkpoint again. So this way it can actually do rollouts over many, many days and come back with the actual like solution, right? And I think my sincere hope is that if we make this infrastructure correct, we can help like solve diseases like find new drugs because the model can just keep going on for longer and longer, right? But it doesn't happen until we have really, really rock solid primitives to do this. And so given the needs and the pillars we want to support, here are some of the things that this snapshotting solution should support, right? First, at chat, GPT, or codex scale, we want to do incremental snapshotting. And what that means is like if you call snapshot twice, I'm just snapshotting the diff between the two snapshots. Otherwise, if I have to save gigabytes of data at every turn, like I'm going to bankrupt the company and like it's just a slow experience regardless, right? The snapshotting API itself should be very, very cheap and fast so the model and the harness can keep snapshotting and exploring like very fast. And similarly, just like creation should be fast for products, like restoring is just nothing but just recreation from a snapshot and so restoring should also be very, very fast for a good product experience. And then we have two paradigms we'll discuss. One is always saving in which the harness doesn't have to explicitly call a save API versus explicit saving where the harness is calling save, save, save. We'll see how we can implement both so I'll give a reference solution. And then there are some more design choices with disk snapshotting, right? So first, like I said, on the left-hand side, do you want incremental or full snapshots? I argue, I think at our scale we want incremental snapshotting. Secondly, do you want to snapshot the entire rootFS or do you want to have certain folders like workspace or mount something, something that you want to keep that configurable? And lastly, like, we can see how you can snapshot at a file system level or at a block. So in Linux, disks are nothing but block devices and each file maps to different blocks. We'll see after this. And so you can do very, very efficient snapshotting by just zipping up the blocks that have changed for a file or you can do entire files that have changed for more write amplification. And the high-level flow of snapshotting is basically you figure out what's changed, you zip it up, put it to the cloud, and then when you restore, you pull it down and you restore the micro VM, right? And so this is a first principles talk and so I just wanted to go over how Linux storage works from first principles. So Linux represents disks as block devices. So you see on the right, like, it thinks of a block device as having logical blocks from zero to N. a file system maps directories and files into something called an inode data structure and the inode says, okay, offset zero in file F maps to this block, logical block in disk D and then on the disk itself, there's firmware running which says, oh, logical block one is like sector 100 or sector 500 or page XYZ. So the hierarchy is on these logical blocks up to the file system and we leverage that for block based snapshotting. And then within a micro VM there are two ways of accessing storage. On the left hand side, you can think of this as sharing a folder like Google Drive but it's very, very inefficient because the VM is doing file system operations, exiting at every file system operation which is very inefficient. On the right hand side, you actually give a disk-like abstraction at a block. So you give a block device to the micro VM and it is way more efficient because you can use the caches inside the guest and you don't have to exit out as much. You only exit out when you truly need to access the block device and the host has to service you. This is the high-level diagram of how you will have always-on persistence in disk snapshotting. So if you have two micro VMs, they have a block device which they will see this dot image as a block device inside and as they are writing to the block device, we are writing through to the cloud. So there's some sort of distributed file system that we mount inside that's giving this always-on persistence. And like I mentioned, for explicit persistence, we have an actual API called the save API. Calling the save API from the harness figures out what's changed between the last snapshot. it bundles up this div into an artifact and returns a snapshot ID. Later on, you can give us this snapshot ID. We figure out the lineage of snapshots that make this snapshot ID and then we download them one by one and apply it on the node and you get a micro VM restored with this thing. So let's see how we can implement this. Give me one second. I want to see how we're doing on time. We have five minutes left, so we'll go fast. Yeah, so we can do explicit persistence using something called copy-on-write. So Linux has these XFS file systems, XFS-like file systems, so you can do pretty much have zero latency copies because you don't change any blocks when you copy, but when you actually change the blocks on a file, that's when you pay the penalty. And so in this design, we have a base dot image, which might be the codex base image or whatever, like the chat GPT base image. We create a zero copy on top of it, which is a writable layer. And then when you write to it, now you're changing the blocks in this layer. And then when you want to snapshot, I use something called FIE map, which tells me what blocks have changed and what ranges. I zip that up and store it to the cloud. And I can do something very nifty here. I can actually lie to you while I'm uploading to the cloud. So the snapshot can happen, return very fast as I'm uploading in the background. I don't have to wait till I'm uploading. And then on the other side, I have this diff, I download this artifact, I figure out what extents have changed, and then I apply it back on top of the base image and I start the micro VM again. So now we have a restored sandbox with the exact same state at a block level. Now how can we do always-on persistence? So this is one way of doing this. I know there's NFS also and other distributed file systems, but NFS, for instance, isn't as performant and it's not POSIX compliant. And I think our models are just very good at anything POSIX compliant and standard. So you can write a file system actually on top of a GCS or S3 or durable block storage, and you can use something called NBD. So within the sandbox, you'll actually see a block device, but inside it will have literally a tiered cache of blocks that are persisted first to an in-cluster cache, and then the in-cluster cache is like writing back to the block storage, the object storage. So you have this nice global tiered architecture where you're actually caching things at the block level and finally inside the micro VM and get a very performant file system inside. Yeah, so that's the persistence part of the presentation. And so the one takeaway I want you guys to think about is I think storage is the next unlock here. As you're working on sandboxes, think of what all you can snapshot and restore fast to give like this new paradigm to harnesses so they can recover from failures and explore, do Monte Carlo-like searches. So we've discussed running things on one node, and of course, if one node dies, we are done. So we want to be able to run across many, many nodes across the world. So I'm not going to mention about Kubernetes or other like acquisition things. This is the first principles talk, so we'll lightly hint about the challenges here. So ideally, we want multiple machines to support the runtime we discussed. So we can group nodes into clusters and spread the clusters across the regions. A top-level control plane chooses a cluster using region, load, and other factors, and it's not very different from the orchestrators you are familiar with. These, like, ideally it uses a cluster close to your chat GPT cluster, so you can, like, have fast, like, access to the harness. Inside the cluster there's a scheduler also, and the scheduler tells you which node to pick based on the load and other factors, right? So if nodes are dying or failing, it won't choose that or will intelligently, like, route the sandbox to the thing. And again, low latency and reliability remain very, very key north stars for this architecture. And so here, like, we can use some micro-VM features to support low latency creation ideas. So a lot of systems cheat for low latency. They pre-womb, like, sandboxes and they pick one, which is great. Another way to do this is you can actually take a memory snapshot of a micro-VM and just in time start it in milliseconds as the request comes. And so you can leverage this, like, nice micro-VM property that you can save the guest memory and start from that. And the third one is a hybrid solution. So you can have a warm pool, but as it's growing, you can, like, grow it from the memory snapshot so you can get best of both worlds. Like, the trade-off for a warm pool is that you're consuming CPU and memory in idle state. Ideally, you don't want to do that, right? So there's a trade-off between one and three and two, basically. And so here is a way where we can use snapshot to restore for better orchestration. So remember, we discussed that a snapshot can have a lineage of many, many layers. So once you want to restore from a snapshot and you find out you have, like, let's say four layers that you want to pull down and that makes the lineage, you can actually, like, smartly route you to a node which has to download the least amount of stuff. So in this diagram, you can see node A has some layers, node C has some layers, but node B has all the layers that you need. So the scheduler then routes you and it gives it the highest score and routes it because it has all the snapshot layers. So you can use, like, snapshot with orchestration to just have faster, like, creates and even just more reliable, like, orchestration. Yeah. This is the talk and hopefully it gives you some design intuition around, like, sandboxes, why they're important, what's the next unlock, and I want to see more of you guys using it in a secure way. Thank you.