One Operator, Many Drones: Inside Skydio's Autonomy Stack — Suchet Bargoti, Skydio
Description
Suchet Bargoti opens with a live demo instead of slides. From a laptop on conference Wi-Fi, he launches a docked drone over San Mateo, starts a second one on a power line in Colorado, has one autonomously track a car, and then sends the whole fleet back to dock. Skydio, the largest US drone manufacturer, now has thousands of these docks deployed with utilities, police and construction companies, and about 16 million people live within two miles of one. Bargoti, Skydio's Director of Inspection and Mapping, explains the shift to "drones as infrastructure." He shows a drone spotting a utility pole burning from the inside, and SFPD following a stolen car without a high-speed chase. He then breaks down the autonomy stack. It splits intelligence between the edge and the cloud, uses maps as world models that the fleet keeps up to date, tracks objects through occlusion, and lets a VLM agent find and follow a "white Jeep" using tool calls instead of hand-coded rules. He also covers where fully end-to-end learning still falls short of the reliability physical systems need. Speaker info: LinkedIn: https://www.linkedin.com/in/sbargoti Related links: Skydio: https://www.skydio.com Timestamps: 0:00 Live demo: flying drones from a laptop 1:02 Drones as infrastructure 2:12 Launching a second drone in Colorado 3:07 Autonomous car tracking 3:42 From hobby toy to tool to infrastructure 5:01 Sending the fleet back to dock 5:26 A utility pole burning from the inside 6:01 Following a stolen car with SFPD 7:06 Reliability in Alaska cold and Texas heat 7:51 Why one pilot per drone doesn't scale 9:45 Full-stack autonomy 10:40 The data flywheel 11:30 Autonomy on the edge and in the cloud 12:05 High-quality video on low bandwidth 12:54 Maps as world models 13:49 Keeping maps up to date with the fleet 14:49 Tracking through occlusion 15:49 Heavier VLM tracking in the cloud 16:19 Semantic reasoning for infrastructure 16:58 Agentic "find and follow" with a VLM 17:48 End-t
Summary
Generated by gpt-5.6-terraAt-a-Glance
- Verdict: Watch fully
- Core thesis: Skydio’s route to one operator commanding many drones is a layered autonomy stack: safety-critical control on the drone, fleet coordination and heavier reasoning in the cloud, and high-level agent interfaces that replace pilot micromanagement with objective-based commands.
- Why it matters: It is a concrete production example of how agentic orchestration must be adapted for embodied, safety-critical systems—using constrained tools, stateful world models, observability, and explicit edge/cloud task partitioning rather than relying on an unconstrained end-to-end agent.
- Best use: Use it as a reference architecture and design review for multi-agent control planes, especially the division between local execution, cloud planning, fleet memory, tool APIs, reliability engineering, and human escalation.
Executive Summary
Bargoti frames drones as infrastructure rather than equipment: permanently docked, remotely available embodied agents that can launch, execute missions, and recover without a dedicated pilot at each site. The live demonstration shows one operator launching drones in San Mateo, Colorado, and Skydio headquarters concurrently, assigning a vehicle-tracking task, then returning all drones to dock. The operating premise is that a laptop—or the network connection from it—should not be a single point of failure once a mission is handed off.
The scaling constraint is cognitive as much as operational. A pilot-per-drone model fails as fleets and incident volume grow, so Skydio wants operators to state goals such as searching an area, finding a person, or following a specific vehicle while the autonomy stack handles flight, safety checks, route planning, tracking, and recovery. Bargoti expects the eventual interface to be even higher-level, potentially an incident-triggered Slack-style workflow where a drone dispatch is implicit.
The technical architecture is deliberately hybrid. Immediate safety-critical behavior and minimal tracking run on the edge device, while cloud infrastructure receives video and telemetry for fleet-wide coordination, heavier inference, semantic reasoning, long-horizon planning, and slower VLM-based decisions. Skydio’s data flywheel logs real-world deviations between intended and actual behavior, sanitizes customer data, retrains and evaluates models, then redistributes improvements across the fleet.
Crucially, Skydio is not presenting full end-to-end reinforcement learning as production-ready for the whole stack. In physical systems requiring “many nines” of reliability, opaque end-to-end policies make failures difficult to observe, diagnose, and guarantee against. Its nearer-term agentic design gives models access to controlled primitives and drone APIs—such as search, detect, track, follow, inspect, and thermal sensing—so high-level agents can reduce brittle hand-authored branching while the underlying safety envelope remains explicit.
Key Takeaways
- Claim: A scalable drone operation requires transitioning from direct piloting to objective-based fleet orchestration. | Evidence: The live demo launched and tasked drones in San Mateo, Colorado, and Skydio HQ at the same time; Bargoti argues that assigning one skilled pilot per drone breaks down as 9-1-1 calls and infrastructure alerts increase. | Implication: For any multi-agent system, the control plane should let an operator manage goals, priorities, exceptions, and fleet state—not manually sequence each agent action. | Caveat: The speaker describes the target state as potentially unsupervised operation, but does not establish that every mission type can already operate without human oversight.
- Claim: Robust embodied-agent systems need an explicit edge/cloud split rather than putting all intelligence in one place. | Evidence: Skydio keeps immediate autonomous actions and a minimal tracking capability on the drone, while cloud GPUs and inference engines process telemetry/video for heavier models, longer-term planning, map-level reasoning, and VLM decisions that can tolerate roughly one to two seconds of latency. | Implication: Ken should treat local execution as the fault-tolerant action layer and reserve central orchestration for work that benefits from broader context, cross-agent visibility, or more expensive inference. | Caveat: Cloud-side reasoning is constrained by bandwidth and latency; it cannot replace local safety-critical behavior.
- Claim: A shared, continuously updated world model is the coordination substrate for a fleet. | Evidence: Skydio combines prior building data with vector data such as roads and power lines to plan globally; when drones observe an unmodeled construction site, the observation feeds a map-sync process so later fleet iterations receive an updated representation. | Implication: Multi-agent orchestration benefits from a versioned shared state layer where agents both consume operational context and contribute validated updates that improve subsequent decisions. | Caveat: Maps become stale, so prior maps alone are insufficient; the system needs a governed update loop from live observations.
- Claim: Agentic behavior should be built from constrained, inspectable tools rather than from a free-form model directly controlling a physical system. | Evidence: In Skydio’s example, a user asks to find and follow a white Jeep; a VLM identifies candidate objects while accessing drone APIs and tools for trajectory, tracking, and follow behavior. For utility inspection, the system is building semantic primitives that agents can invoke to understand a scene and act within it. | Implication: Design agent systems around typed capabilities, state inspection, and bounded actions; make natural language the task specification layer, not the safety or actuation layer. | Caveat: The talk does not detail tool permissions, policy enforcement, or human approval gates, all of which are essential when tools can cause physical action.
- Claim: Real-world deployment data creates a defensible learning flywheel, but it must be paired with privacy controls and evaluation. | Evidence: Skydio logs flights to capture cases where the drone behaved differently from expectations, sanitizes data to remove private information, obtains clarity on what customers share, and feeds data into retraining and evaluation for reinforcement-learning ecosystems. | Implication: Operational telemetry should be treated as a first-class product asset, with explicit consent, data governance, failure labeling, offline evaluation, and controlled rollout—not merely as logs. | Caveat: The value of the flywheel depends on data-sharing permissions, effective sanitization, and rigorous regression testing before updated behavior reaches deployed systems.
- Claim: End-to-end learned control remains strategically attractive but is currently hard to certify and debug at the reliability level required for physical autonomy. | Evidence: Bargoti says Skydio is testing reinforcement learning from raw sensor data to outcomes, but identifies observability, failure diagnosis, guarantees, and “many nines” reliability as major challenges for a completely end-to-end approach. | Implication: For high-consequence agent workflows, optimize first for explainability, recoverability, and testable guarantees; use learning to replace complex decision branches incrementally rather than eliminating the control structure wholesale. | Caveat: This is Skydio’s engineering judgment rather than a universal proof that end-to-end control cannot work; the proposed answer is selective use of learned/world-model components where hand-coded branching becomes brittle.
- Claim: The infrastructure model is already producing operational value in public safety and utility inspection. | Evidence: A utility patrol identified a pole burning internally before it could become a fire risk; with SFPD, drones tracked a stolen car while suspects changed license plates and tinted windows, enabling safer police positioning instead of a high-speed chase. Skydio says about 16 million people live within two miles of deployed infrastructure and targets 99.99% reliability across environments including Alaska cold and Texas heat. | Implication: The strongest initial markets for autonomous fleets are recurring, time-sensitive workflows where travel time, staff exposure, or scarce expert operators are the principal bottlenecks. | Caveat: These are company-presented examples and deployment claims; the talk provides no independent outcome data, cost comparison, or incident-rate analysis.
Detailed Brief
Architecture constraints that determine whether cloud orchestration works
- Claims: Cloud-managed fleets are only viable if video and telemetry remain useful under poor network conditions.; The system must be resilient to operator-device failure after mission handoff rather than requiring the user’s laptop to remain connected.; Vertical integration across hardware, autonomy software, cloud services, user interface, networking, and fleet operations is presented as Skydio’s advantage in managing cross-layer trade-offs.
- Evidence: The live conference demonstration was conducted over conference Wi-Fi; Bargoti states that closing the laptop should not prevent safe mission completion.; Skydio invests in compressing video for transmission and decoding it into clearer output under the same low-bandwidth network conditions.; The company describes itself as controlling the drone hardware, software, cloud, UI, and autonomy stack.
- Caveats: The presentation does not quantify bandwidth requirements, fallback communications, cloud availability targets, or the exact autonomous behavior when connectivity is lost.; Full-stack ownership can improve integration but can also increase organizational complexity and reduce component substitutability.
- Implications: A credible orchestration platform needs degraded-mode design: local mission continuity, safe abort/recovery behavior, and a communications layer designed around information quality rather than raw data volume.; When evaluating agent-control platforms, inspect cross-layer ownership and interfaces because the critical failure modes often occur between perception, networking, planning, and actuation.
Model roles in Skydio’s autonomy stack
- Claims: Bargoti uses “model” to cover several distinct functions: map-like world representations, conventional perception/tracking models, VLM-driven semantic reasoning, and exploratory end-to-end reinforcement learning.; Object tracking requires a persistent representation that reasons through occlusion rather than simply reacting to visible pixels.; The desired future replaces large decision trees in workflows such as search and rescue with high-level agent access to sensing and action tools.
- Evidence: A target that disappears behind a building can be pursued by reasoning about where it may reappear and repositioning the drone accordingly.; The search-and-rescue example contrasts hand-authored branches—search trees, then another location, then thermal sensing—with agent-directed use of higher-level tools.; Skydio is extending the infrastructure approach beyond conventional quadcopters to smaller quadcopters and exploring fixed-wing systems.
- Caveats: The talk does not disclose benchmark results, model accuracy, false-positive rates, reinforcement-learning safety methods, or how models are validated across weather and mission classes.
- Implications: “Agentic orchestration” is not one model: it is a role-based assembly of memory, perception, planners, and action tools operating at different latencies and reliability requirements.; A reusable control plane should abstract heterogeneous agent types behind common capability APIs while preserving vehicle-specific constraints.
Notable Concepts & Terms
- Drones as infrastructure: A shift from manually deployed tools to permanently docked, always-available autonomous assets serving recurring public-safety, utility, and construction workflows.
- Embodied agents: The speaker’s framing of drones as agents whose decisions have direct physical consequences, raising the reliability and safety bar beyond software-only agents.
- Edge/cloud autonomy split: Immediate flight safety and fast local responses remain on the drone, while the cloud handles fleet context, expensive inference, semantic interpretation, and longer-horizon planning.
- World model: A map-like representation that fuses prior geographic/vector data with fleet observations so drones can plan globally and improve shared environmental knowledge.
- Learning flywheel: Operational flights generate data about mismatches between expected and actual behavior, which is sanitized, evaluated, used for retraining, and redeployed.
- Visual language model (VLM): A model used for semantic scene understanding, such as locating a white Jeep, while tool APIs translate that understanding into bounded drone actions.
- Many nines reliability: The reliability standard that drives Skydio away from opaque fully end-to-end control and toward systems with observable, testable, constrained components.
- Primitives/tools: Encoded capabilities that an agent can invoke to inspect, track, navigate, or sense, replacing brittle bespoke if/then mission logic with composable actions.
Operator Notes / Why Ken Should Care
- Use Skydio’s edge/cloud allocation as a design template: define which agent actions must execute locally under disconnection, versus which can wait for central planning or heavy inference.
- Require every high-impact agent capability to expose a bounded API, explicit preconditions, telemetry, abort behavior, and an auditable decision trail before permitting autonomous execution.
- Add a shared-state lifecycle to any fleet/control-plane design: source provenance, confidence, versioning, stale-data detection, validation, and propagation of updates to other agents.
- Evaluate agent projects against the physical-autonomy standard of observability and recovery, not just task success: ask how failures are detected, explained, contained, and regression-tested.
- Monitor network quality as a first-class control-plane variable; invest in selective compression and degraded-mode behavior rather than assuming continuous high-bandwidth connectivity.
Source/Metadata
- Title: One Operator, Many Drones: Inside Skydio's Autonomy Stack — Suchet Bargoti, Skydio
- Transcript words: 7759
- Duration seconds: 1247
- Timestamp note: No usable timestamps or chapters were provided; the transcript contains a substantial repeated segment.
Transcript
Thanks everyone for coming. So we'll be talking about how to use agentic orchestration to command many different drones to achieve the different tasks that we have ahead of us and we're gonna rethink this presentation a bit and not just jump straight into slides but instead we're gonna fly. So what we have here on the left-hand side is a real-time view of what we're gonna be doing right now. Am I showing anything? I'm flying. There you go. Sweet. Okay, we are somewhere here in the city down south, which is where our headquarters is. We have a few drones up and I'm gonna hit launch and these are drones that are in docked stations. So we're calling this drones as infrastructure where we actually have thousands of these drones now planted around the country with power utilities, with public safety, with construction companies. What I'm doing right now is using my keyboard to simply fly here. For those of you that know the area, this is San Mateo and perhaps if I zoom in here, we may see a very faint render of the SF skyline, although the fog is always there, so we might just say hello to the SFO airport here. Some planes are launching. If anyone has a plane tracker app, they can see this is real-time stuff there. So what we have been building is the full autonomy stack behind how a vehicle operates autonomously, how the cloud system operates here, how the cloud servers work, and where the different levels of intelligence and automation need to happen for us to make this robust such that when I am here and say hey, there's an incident happening here, I need to go respond, I could at the same time go back here and say hey, what about that other thing that's happening on the other side of the country? What if we launch that instead? So while that's happening, I'm now going to launch something in Colorado. So this is if some sort of fire incident has happened on the power lines and we wanted to go out and inspect those things, a dock is now opening up in Colorado while the first drone is still flying safely and I'm just gonna let it do all the safety checks that it needs to do before the launch. This is all happening over conference Wi-Fi traffic, so you can imagine I can shut down my laptop right now and everything needs to safely happen behind the scenes. So I'm gonna do this while the first part is still happening. Let me go back to the first drone and give it another set of instructions. Let's have a look around here on the first drone. The second drone has kicked off and maybe there's some cars going around there that we can perhaps track. So let's track that. What's this car doing here? All right, so we're now tracking this car. Maybe this is a runaway car that we needed to follow and my hands are off right now. This is the autonomous system taking over through the various interactions that I've given for it to do. At the same time, while this is happening, we can do the final thing, which is yet another drone in the system back at our HQ and we can say hey, why don't we run a third drone? So what we're building towards is how do we enable autonomy at scale? Where traditionally, how it started off historically 15 years ago, you might have a drone at home, it's a hobbyist drone and you might play around with it, tinker with it, work with the controlling software. Then about 10 years ago, drones started to become a lot more available and they started to become a tool. A lot of industries out there started to use them. They'll carry them with them in the truck, go out there, deploy it. The next year that we're working towards is drones infrastructure. Imagine these systems, and for the sake of this conference, I'm going to call them agents. These are physical embodied agents that are available at all times, anywhere, for the different use cases that we're interested in, that can automatically launch, execute, and do their tasks. What is the minimum amount of autonomy that needs to be baked in and what does the future interface look like? Today, everyone needs to be a dedicated pilot. I went through a certification exercise. I need to think about the safety standards here, but you can imagine in a few years time, the safety is going to be determined by the autonomous system and the interface becomes really high level. It could be a little Slack bot that says hey, something's happened here, why don't we send a drone and you might not even know that the drone launched. So while this is happening, I'm going to tell all the drones to pause and return to dock. So we'll write this. They're all going to start returning to dock. While that's happening, I'm going to start on the presentation. So what I've shown you here is not just concept. These are literally systems that are being used in production. As I've mentioned, we are the largest manufacturer of drones in the US and we want to give people superpowers through this technology. Let's take a quick look at what some of the people are doing here. This is in the Northeast coast of the US where our client has set up the system next to a power station and they use this to do normal patrols. As they flew around, they found that some of the poles were burning from the inside and it was really starting to show up here. This could have fallen at any time and create a fire risk that they would have otherwise not caught without having to send people there, which itself is quite expensive. Moving on to the next use case is on the public safety side. This is San Francisco. We work very closely with SFPD. Normally, when a car gets stolen, you'll see a high-speed chase happening in the city, quite dangerous. But what if you could deploy a drone instead? The way the people in the car don't even know that there's a drone following them. Here's a person who has stolen the car on the right-hand side and they're about to change their license plates. They go here to get out their tools, come back. Luckily, they point the license plate up so the drone can see it and we know exactly what they're doing. But they're now replacing the plate in the car with the new one and all this time they don't know that they're being chased. How they behave in public, how they behave out there is very different and it's a lot safer in how the operations are done. They're now going to go ahead and tint the windows, but this allows the police to strategically position themselves in the safest way to intervene at the right time rather than doing a high-speed chase outside. So these are being used across many different industries and we are really starting to treat this as infrastructure that can operate day in, day out, nighttime, rain, sunshine. We've got a few of these docks deployed in Alaska, so very cold weather. Few of these docks deployed in Texas, so extreme heat. These need to be reliable down to 99.99%, where we do our simulation testing to be able to prove that. Today, we have about 16 million people living within two miles radius of this infrastructure that the public safety, the power companies can use this technology to be able to respond to incidents without having to travel there. So what really happens when this happens at scale? Can I get a hands up of people that have flown a drone before? A couple of hands up. When I started to fly it, it took me a few hours to really figure it out. Then I put on the FPV and that was even more tricky, but it felt good to get that expertise up. But it's a skill that you develop over time and you say okay, for each skilled person, we're gonna put them next to a drone and they're gonna start working. But now you want more of these, so police companies, infrastructure companies need to start hiring these people more, and at some stage, this really starts to break down. The more 9-1-1 calls that come in, the more alerts that can come in—it doesn't really scale. So we're rethinking what this means in terms of this initial first-person viewing engagement with these systems to how do we convert this to a more strategic multi-agent view that you can command the entire fleet with an objective in mind without having to worry about the flight. Let me just go back and see if that was all working well. Great, they all landed. I am so pleased when that happens successfully, although it's meant to happen all the time. All right, let's go back to this. All right, we're just gonna carry on. So what does that mean when we start to launch different things? You saw me launch them. I'm still thinking about it. Okay, I need to launch this one, that one, that one, and even for myself in this, who's used to this, I'm gonna have some cognitive challenges. Where our vision is to be able to launch many of them in a potentially unsupervised way. So what commands these things? How do we get it to just hey, get out there, launch, search in this area, find a missing person or look out for this type of car and hold your position there? In order to do that, we really need to think about how do we get the many nines of reliability that we need in autonomous flight. That's where we get an edge in the industry because at Skydio, we're controlling the hardware, the software, the cloud, the user interface to be able to manage all that, and specifically the autonomy, allowing us to think about how our underlying vision So what does that mean when we start to launch different things? You saw me launch them. I'm still thinking about it. Okay, I need to launch this one, that one, that one, and even myself in this. I'm used to this. I'm going to have some cognitive challenges. Our vision is to be able to launch many, many of them in a potentially unsupervised way. So what sort of commands these things? How do we get it to, like, just, "hey, get out there, launch, search in this area, find a missing person or look out for this type of car and hold your position there"? In order to do that we really need to think about how do we get the many nines of reliability that we need in autonomous flight, and that's where we get an edge in the industry because at Skydio we're controlling the hardware, the software, the cloud, the user interface to be able to manage all that, and specifically the autonomy, allowing us to think about how our underlying vision system should work to see the environment, to behave in the environment correctly, whether it's at high altitudes, whether it's cloudy, whether it's at high speeds. How do we deal in the rain? On the bottom left, we're showing how do we navigate in cities, how do we plan large-scale and be able to do that. And for anyone that has worked with any sort of GPS device in the city, even our phones, they kind of suck. So how do we robustly do that? And how do we also do tracking when there's a lot of occlusion? These are all the different places where we're thinking about how to train AI systems, quote-unquote, models. The word model itself has different meanings in different places, and I'll discuss a little bit on what that means for us. But in order for us to really harness this, we are learning on the go. This is a learning flywheel that we're getting out there. We're collecting data, we're operating, and we're coming back and doing that so that each flight we can log the data, similar to Google Street View, where we need to think about sanitizing that data, make sure there's no private information left there, and make sure customers know exactly what they're sharing with us. But if we can do that, we have access to huge amounts of data that we can learn from. Every single time we instructed this but the drone did this. Every single time we thought this was going to happen but this happened. That can come back to our learning agents, our reinforcement learning ecosystems, to be able to retrain, evaluate, and send it back out there and continue that flywheel that allows us to have that robust framework. And the other added advantage that we have is it's not just about having autonomy on the drone. As I mentioned earlier, we have the luxury now to have autonomy on the edge device but also have autonomy on the cloud. What I was showing you earlier, all that video feed, all the telemetry, that's going through a cloud server. We could set up GPUs and we could set up inference engines there to be able to have that heavier lifting, maybe that longer-term planning there, whereas the immediate autonomous actions happen on the drone and we're constantly thinking about the trade-off that we need to make to make that successful. One thing is true: however, once you do start thinking about having your agents in the cloud, the amount of data coming to the cloud really matters. Here's a sort of investing in how we think about enabling the best video quality coming up through to the servers in low bandwidth areas, being able to optimize that, being able to encode that information into smaller sizes and be able to decode into something clear that allows us to have much higher quality in the exact same sort of network conditions. Here are some of the areas of investments that we make so that we can have this more cloud-based infrastructure to be able to manage this at scale. So once we do that, I want to explore at a high level some of the models that we have in our system that allows us to orchestrate all of this. Firstly, a model—this is a section about world models. I want to talk about world models from a context of maps, not dissimilar to how Waymo works. They have a map of the world and they navigate in that world. They do local perception but also think about global planning. If I want to go from part A of the city to part B, I can't just keep hitting every building and navigate around them. I need to think about what's the optimal path along the way. So we start off with a lot of prior information and we merge that with not just building data but maybe there's vector data such as where the power lines are, where the roads are, if you want different behavior in these areas, and we think about how to combine these resources to ultimately build a map that we can plan and navigate around. And the drone has knowledge of this map at all times to be able to go around that. But like any map, maps can go out of date. Luckily, we have so many eyes in the sky to think about how to maintain and update these maps along the way. On the left-hand side, we are rendering our knowledge of the world in points onto our video feed. However, that doesn't line up perfectly everywhere. There are some sections here that, if I sort of zoom out here, there are some sections here that a new construction site had set up that we did not know about. However, as a drone now starts to fly—and this is not just one drone but a fleet of drones—they're now observing these things that we can feed back into our map syncing process. That can come back, land, give the data, and now once in the next iteration, all the drones in the fleet have this most updated map of the world that they can do all the planning in. A second type of model is perhaps today a more conventional sort of machine learning inference model, which is the ability to track and understand objects in the scene, be able to track them, be able to track them behind occlusions. So the implicit representation behind the scene is some sort of world representation of the object that, "hey, it's gone behind this building and it might come out on the other side," so I should navigate myself so I can follow it there or I should move myself in a different direction to be able to do that. Maybe five years ago this would be more done in a more conventional way. You have to manually think about how to move, but now you can think about more reinforcement learning style techniques or more learned end-to-end approaches that can really help out here without you having to engineer all the edge cases that can go in. And obviously, how do we do this robustly in rain, snow, day, night, using vision only? One other thing that we're doing here on the tracking side is we have some minimal set of tracking that's happened on device, on the edge, but we can do some high-level tracking that happens in the cloud. So perhaps it can reason more about your entire map, perhaps it can reason more, use heavier models, use VLMs with lower latency which don't respond as quickly at a rate of, let's say, seven to ten Hertz but can give you feedback at a one to two second latency. But that's good enough for us to make broad decisions about where to move. For a lot of our infrastructure customers, we're doing a lot of semantic reasoning. So what is there in the scene? Here is an illustration of us thinking about utility poles, and we want to—when instruction comes in, "hey, go look at this line, there's something gone wrong"—the drone needs to go there, it needs to understand the scene, then it needs to take actions within that scene, and we're constantly looking at how to build these primitives, these tools, ultimately that we today code, but at any time an agent can access these tools to better understand what the drone is seeing and what it could do about it. So an example of the agentic sort of system in action is a visual language model. On the top left side, the user here is typing in, "look for a white jeep and find and follow it." It accesses a drone API to command a certain sort of trajectory that it should take. While that is happening, the detection head, the VLM is running here to say, "hey, what's in the scene? Am I looking for it?" It finds something, then it has access to the tools that allow the drone to track and follow, and that's without any specific coding of that law or rules, but instead having a more agentic approach, giving the agents all the tools that it needs to be able to understand the drone state and make decisions given the information in the context that's available there. And then there's always the sort of long-term vision that's often there in the self-driving community right now or any sort of robotic system: what if we could just give it raw sensor data and outcomes and get perfect results? Perhaps that's the actuation that happens. Perhaps it's where the drone is pointing. Perhaps it's where the drone goes. And we're definitely doing a lot of testing and trialing with reinforcement learning on what that looks like if we have multiple instantiations of this behavior. Does it get to the right end spectrum? The main consideration whenever we work with a physical system is that you're often looking at really high volumes of reliability. As I was saying, many nines of reliability. And doing a completely end-to-end system does have its challenges in the sense that the observability of what's going wrong and the guarantees and reliability is very difficult today. So while it's a direction that we're continuously taking and exploring, it's about figuring out which segments of your end-to-end chunk need to move to a more world model representation of it. community here right now or any sort of robotic system what if we could just give it raw sensor data and outcomes the perfect results perhaps that's the actuation that happens perhaps it's where the drone is pointing perhaps it's where the drone goes and we're definitely doing a lot of testing and trialing with reinforcement learning on what that looks like if we have multiple instantiations of this behavior does it get to the right end spectrum the main consideration whenever we work with a physical system is that you're often looking at really high volumes of reliability as I was saying many nines of reliability and doing a completely end-to-end system does have its challenges in the sense that the observability of what's going wrong and the guarantees and reliability is very difficult today so while it's a direction that we're continuously taking and exploring it's figuring out which segments of your end-to-end chunk need to move to a more world model representation of it and it's specifically the things where we always find ourselves okay I need to hand engineer this I need to code this in specifically I need to look at these rules like search and rescue is one of those things like oh go look for these trees but if you don't find the trees look under here or then turn on thermal or but then if you didn't find it here look there we're trying to get away from having to have all these if statements in the code and these branching strategies and let the agent have some high level tools to be able to instruct these very high level commands for the drone so I talked about a quadcopter today but we're doing this similar with a much smaller form factor quadcopter and we're now starting to look at this what this looks like from a fixed wing as well all in the world of infrastructure so they can be launched from anywhere recovered from anywhere and ultimately the sweet spot where we can start to very quickly build on the cloud to orchestrate these things and the way we're thinking about it is allowing these systems to have basic APIs of interactivity that the cloud agents can come in and tap into and make decisions on and ultimately allow for very high level thinking when we work with these drones so like many talks here we are hiring as I said we have full stack we do the end to end thing hardware software autonomy full stack front and back end wireless networking everything all the technologies on our mobile phones we're now making them fly so come say hi or look at our website would love to talk more thank you Sweet. Okay we are somewhere here in the city down south which where our headquarters is. We have a few drones up and I'm gonna hit launch and these are drones that are in docked stations so we're calling this drones as infrastructure where we actually have thousands of these drones now planted around the country with power utilities, with public safety, with construction companies and what I'm doing right now is using my keyboard to simply fly here. For those of you that know the area this is San Mateo and perhaps if I kind of zoom in here we may see a very faint render of the SF skyline although the the fog city is always there so we might just say hello to the SFO airport here. Some planes are launching if anyone has a plane tracker app they can sort of see this is real-time stuff there. So what we have been building is the full autonomy stack behind like how does a vehicle operate autonomously, how does the cloud system operate here, how do the cloud servers work and where does the different levels of intelligence and automation needs to happen for us to make this happen robustly such that when I am here and saying hey oh there's an incident happening here I need to go respond I could at the same time go back here and say hey what about that other thing that's happening on the other side of the country what if we launch that instead. So while that's happening I'm now going to launch something in Colorado. So this is a imagine some sort of fire incident has happened on the power lines and we wanted to go out and inspect those things so a dock is now opening up in Colorado while the first drone is still flying safely and I'm just gonna let it do all the safety checks that it needs to do before the launch and this is all happening in sort of conference Wi-Fi traffic so you can imagine I can shut down my laptop right now and everything needs to safely happen behind the scenes. So I'm gonna do this while the first part is still happening let me let's go back to the first drone might give it another set of instruction let's have a look around here on the first drone the second drone has sort of kick-started off and maybe there's some cars sort of going around there that we can perhaps track so let's track and what's this car doing here all right so we're now tracking this car maybe this is a runaway car that we needed to follow and my hands-off right now this is the autonomous system kind of taking over through the various interactions that I've given for it to do and at the same time while this is happening we can do the final thing which is yet another sort of drone in the system back at our HQ and we can say hey let's why don't we run a third drone. So what we're building towards is how do we enable autonomy at scale where traditionally how it started off historically 15 years ago that you might have a drone at home it's a hobbyist drone and you might play around with it you will tinker with it you will work with the controlling software and then about 10 years ago drones started to become a lot more available and they started to become a tool a lot of industries out there started to use them they'll carry them with them in the truck go out there deploy it and the next year that we're working towards is drones infrastructure imagine these systems and for the sake of this conference I'm going to call them agents these are physical embodied agents that are kind of available at all times at anywhere for the different sort of use cases that we're interested in that can automatically launch execute and do their tasks and what is the minimum amount of autonomy that needs to be baked in and what is the future interface look like. Today everyone needs to be a dedicated pilot I go I went through a certification exercise I need to think about the safety standards here but you can imagine in a few years time the safety is going to be determined by the autonomous system and the interface becomes really high level it could be a little slack bot that says hey something's happened here why don't we go send a drone and you might not even know that the drone launched so while this is happening I'm going to tell all the drones to pause and return to dock so we'll write this they're all going to start returning to dock and while that's happening I'm going to start on the presentation so what I've shown you here is not just concept these are literally systems that are being used in production as I've mentioned we are the largest manufacturer of drones in the US and we want to give people superpowers through this technology. Let's take a quick look at what some of the people are doing here this is in the Northeast coast of the US where our client has set up the system next to a power station and they kind of use this to do normal patrols and as they flew around they found that some of those the pole was burning from the inside and it was really starting to show up here and this could have fallen at any time and create a fire risk that they would have otherwise not caught without having to send people there which itself is quite expensive. Moving on to the sort of next use case is on the public safety side this is San Francisco we work very closely with SFPD. Normally when a car gets stolen you'll see a high-speed chase happening in the city quite dangerous but what if you could deploy a drone instead the way the people in the car don't even know that there's a drone following them. Here's a person who has stolen the car on the right-hand side and they're about to change their license plates on that so they go here to get out their tools they come back luckily they point the license plate up so the drone can see it and we know exactly what they're doing but they're now replacing the plate in the car with the new one and all this time they don't know that there's they're being chased. How they behave in the public how they behave out there is very different and it's a lot safer in how the operations are done. They're now going to go ahead and tint the windows but this allows the police to like strategically position themselves in the most safest way form to intervene at the right time rather than doing a high-speed chase outside. So these are being used across many different industries and we are really starting to treat this as infrastructure that can operate day in day out nighttime rain sunshine we've got a few of these docks deployed in Alaska so very cold weathers few of these docks deployed in Texas so extreme heat and these need to be reliable down to 99.99% where we do our sort of simulation testing to be able to prove that and today we have about 16 million people living within two miles radius of this infrastructure that the public safety the power companies can kind of use this technology to be able to respond to such as incidents without having to travel there. So what really happens when this happens at scale can I get a hands up of people that have flown a drone before? A couple of hands up when I started to fly it it was like it took me a few hours to like really figure it out then I put on the FPV that was even more tricky but it felt good to get that expertise up but it's kind of like a skill that you develop and you develop that skill over time and you say okay for each skilled person we're gonna put them next to a drone and they're gonna start working but now you want more of these so police companies infrastructure companies need to start hiring these people more and at some stage this really starts to break down the more 9-1-1 calls come that come in the more alerts that can come in it doesn't really scale so we're kind of rethinking as to what this means in terms of this like initial first person viewing engagement with these systems to how do we convert this to a more strategic multi-agent view that you can kind of command the entire fleet with an objective in mind without having to worry about the flight. Let me just go back and see if that was all working well. Great they all landed I am so pleased when that happens successfully although it's meant to happen all the time. All right let's go back to this. All right we're just gonna carry on. So what does that mean when we start to launch different things? You saw me launch them I still kind of thinking about it okay I need to launch this one that one that one and even myself in this like who's used to this I'm gonna have some cognitive challenges where our vision is to be able to launch many many of them in a potentially unsupervised way so what sort of commands these things how do we get it to like just hey get out there launch search in this area find a missing person or look out for this type of car and hold your position there and in order to do that we really need to think about how do we get like really like the many nines of reliability that we need in autonomous flight and that's where we get an edge in the industry because at Skydio we're sort of controlling the hardware the software the cloud the user interface to be able to manage all that and specifically the autonomy allowing us to think about how our underlying vision system should work to see the environment to behave in the environment correctly whether it's at high altitudes whether it's cloudy whether it's at high speeds how do we deal in the rain on the bottom left we're showing how do we navigate in cities how do we plan large-scale and be able to do that and for anyone that has worked with any sort of GPS device in the city even our phones they kind of suck so how do we robustly do that and how do we also like do tracking when there's a lot of occlusion these are the all the different places where we're thinking about how to train AI systems quote-unquote models the word model itself is has different meanings in different places and I'll discuss a little bit on what that means for us but in order for us to you like really harness this is we are learning on the go this is a learning flywheel that we're getting out there we're collecting data we're operating and we're coming back and doing that so that each flight we can log the data kind of like Google Street View where we need to think about sanitizing that data make sure there's no private information left there and make sure we don't sort of like the customers know exactly what they're sharing with us but if we can once we do that we have access to a huge amounts of data that we can kind of learn from every single time we instructed this but the drone did this every single time we thought this was going to happen but this happened that can come back to our learning agents our reinforcement learnings ecosystems to be able to retrain evaluate and send it back out there and kind of continue that flywheel that allows us to have that robust framework and the other added advantage that we have is it's not just about having autonomy on the drone as I mentioned earlier we have the luxury now to have autonomy on the edge device but also have autonomy on the cloud what I was showing you earlier all that video feed all the telemetry that's going through a cloud server we could set up GPUs and we could set up inference engines there to be able to have that heavier lifting maybe that longer term planning there whereas the immediate autonomous actions happen on the drone and kind of always thinking about the trade-off that we need to make to make that successful one thing is true however that once you do start thinking about having your agents in the cloud this the amount of data coming to the cloud really matters here's a sort of investing in how we think about enabling the best video quality coming up through to the servers in low sort of bandwidth areas being able to sort of optimize that being able to encode that information into smaller sizes and be able to decode into something clear that allows us to have much higher quality exact same sort of network conditions here and it's the air some of the areas of investments that we make so that we can have this more cloud-based infrastructure to be able to manage this at scale so once we do that I want to sort of explore at a high level some of the models that we have in our system that allows us to orchestrate all of this things firstly a model this is a section about world models I want to sort of talk about world models from a context of maps not too dissimilar to how Waymo works they have a map of the world and they kind of navigate in that world they do local perception but also think about global planning if I want to go from part A of the city to part B I can't just like keep hitting every building and kind of navigating around them I need to think about what's the optimal path along the way so we start off with a lot of prior information and we merge that with not just sort of building data but maybe there's a vector data such as where the power lines are where the roads are if you want different behavior in these areas and we think about how to combine these resources to ultimately build a map that we can plan and navigate around and the drone has knowledge of this map at all times to be able to go around that but like any map maps can go out of date luckily we have so many eyes in the sky to think about how to maintain and update these maps along the way on the left hand side we are rendering our knowledge of the world in the points onto our video feed however that doesn't line up perfectly everywhere there's some sections here that out if I sort of zoom out here there's some sections here that a new construction site had set up that we did not know about however as a drone now starts to fly and not this is not just one drone but your fleet of drone they're now observing these things that we can feed back into our map syncing process that can come back land give the data and now once in the next iteration all the drones in the fleet have this most updated map of the world that they can do all the planning in a second type of model is perhaps today a more conventional sort of machine learning inference model which is the ability to be able to track understand objects in the scene but be able to track them be able to track them behind occlusions so the this implicit representation behind the scene is some sort of world representation of the object that hey it's gone behind this building and it might come out on the other side so I should navigate myself so I can kind of follow it there or I should move myself in a different direction to be able to do that maybe five years ago this would be more done in a more conventional way you have to like manually think about how to move but now you can think about more reinforcement learning style techniques or more learned end-to-end approaches that can really help out here without you having to engineer all the edge cases that can go in and then obviously like how do we do this robustly rain snow day night using vision only one other thing that we sort of doing here on the tracking side is we have some minimal set of tracking that's happened on device on the edge but we can do some high-level tracking that happens in the cloud so perhaps it can reason more about your entire map perhaps it can reason more about use heavier models use vlms with lower which have which don't respond as quickly at a rate of like let's say seven to ten Hertz but can give you feedback at a one to two second latency but that's good enough for us to make broad decisions about where to move for a lot of our infrastructure customers we're doing a lot of semantic reasoning so what is there in the scene here is an illustration of us thinking about utility poles and we want to when instruction comes in hey go look at this line there's something gone wrong the drone needs to go there it needs to understand the scene then it needs to take actions within that scene and we're constantly looking at how to build these primitives these tools ultimately that we today code but at any time an agent can access these tools to better understand what the drone is seeing and what it could do about it so an example of the agentic sort of system in action is a visual language model on the top left side the user here is typing in look for a white jeep and it's doing a sort of find and follow it accesses a drone API to command a certain sort of trajectory that should take while that is happening the detection head is kind of the vl the vlm is kind of running here to say hey what's what in the scene am i looking for it finds something then it has access to the tools that allow the drone to track and follow and that's without any specific coding of that law rules but instead having a more sort of agentic giving the agents the all the tools that it needs to be able to understand the drone state and make decisions given the information in the context that's available there and then there's always the sort of long-term vision that's often there in the self-driving community here right now or any sort of robotic system what if we could just give it raw sensor data and outcomes the perfect results perhaps that's the actuation that happens perhaps it's where the drone is pointing perhaps it's where the drone goes and we're definitely sort of doing a lot of testing and trialing with reinforcement learning on what that looks like if we have multiple instantiations of this behavior does it get to the right end spectrum the main consideration whenever we work with a physical system is that you're often looking at really high volumes of reliability as I was saying many nines of reliability and the doing a completely end-to-end system does have its challenges in the sense that the observability of what's going wrong and the guarantees and reliability is very difficult today so while it's a direction that we're continuously taking and exploring it's kind of figuring out which segments of your end-to-end chunk need to move to a more like sort of world model representation of it and it's specifically the things where we always find ourselves okay I need to hand engineer this I need to code this in specifically I need to look at these rules like search and rescue is one of those things like oh go look for these trees but if you don't find the trees look under here or then turn on thermal or but then if you didn't find it here look there we're trying to get away from having to have all these if statements in the code and these branching strategies and kind of let the agent have some high level tools to be able to instruct these very high level commands for the drone so I talked about a quadcopter today but we're kind of doing this similar with a much smaller from form factor quadcopter and we're now starting to look at this what this looks like from a fixed wing as well all in the world of infrastructure so they can be launched from anywhere recovered from anywhere and ultimately the sweet spot where we can kind of start to very quickly build on the cloud to orchestrate these things and the way we're thinking about it is allowing these systems to have basic APIs of interactivity that the cloud agents can come in and tap into and make decisions on and ultimately allow for very high level thinking when we work with these drones so like many talks here we are hiring as I said we have full stack sort of we do the end to end thing hardware software autonomy full stack front and back end wireless networking everything all the technologies on our mobile phones we're now making them fly so come say hi or look at our website would love to talk more thank you and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes ! Thank you.