Open Reader

Tell the Robot What You Want — Sandhya Subramani, AWS

completed 17:22 Aug 29, 2026 Watch on YouTube

Current Status

completed

Video ID

S6aSoQ6_u5A

RAG / Chat

Enabled
Tell the Robot What You Want — Sandhya Subramani, AWS
Description

Sandhya Subramani asks her rover how many people it can see, a question nobody ever trained it to answer. It thinks for a second, checks its front camera, and reports two, one near the speaker monitor and one further right. Scout is a small four legged robot running a Raspberry Pi, reaching the internet over a SIM card and a 4G connection, and its whole personality comes from an agent layer sitting above the movement policies it already had. The demo goes exactly as live robot demos go. It falls over, gets coaxed upright, announces a signature performance and then just turns its headlights off. The idea underneath is neat and portable. If you give an agent software tools, you can also hand it a hardware tool and let it choose which preset policy to fire, which turns a robot trained for a fixed task list into something you can address in plain language. She wires it up in about five lines using an open source AWS framework that already covers 40 or so robots. Scout actually runs three agents at once, one thinking about the environment, one talking to her over Telegram, and a voice agent she disabled so it would stop interrupting her on stage. Her framing is that the agent decides what to do while the trained policy decides how, and the robot doubles as a rig for collecting the next round of training data. Speaker info: - https://www.linkedin.com/in/sandhyasubramani/ Timestamps: 0:00 - Meet Scout, running on a Raspberry Pi over 4G 1:27 - Answering a question nobody trained it for 3:46 - Giving an agent a hardware tool 5:02 - Five lines to hand a robot to an agent 6:09 - Three agents running at once 7:17 - The four layers, and how observations travel up 9:35 - Splitting work between cloud and edge 10:45 - Where policies go if they scale like language models 11:53 - Asking it to do something complex 14:12 - What is actually under the hood 15:21 - Mirror mirror, over Telegram

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Skim
  • Core thesis: An LLM-based agent can make a conventionally programmed robot more flexible by translating natural-language requests into calls to existing robot policies and hardware functions.
  • Why it matters: The talk offers a concrete embodied-agent control-plane pattern: separate high-level task selection by an agent from low-level execution by trained robot policies, with cloud and edge roles split for training and latency.
  • Best use: Use it as a short architecture reference for tool-using agents controlling physical systems, not as a production-ready robotics implementation guide.

Executive Summary

Sandhya Subramani demonstrates Scout, a Raspberry Pi-equipped rover connected over 4G, that can receive natural-language instructions, inspect its camera feed, speak responses, and invoke preset physical actions such as turning on headlights or spinning. The core proposition is not that the LLM itself controls motors directly: the agent selects among available robot tools and trained policies, while those policies determine the mechanical execution.

The proposed stack has four layers: an agentic layer that receives instructions and observations; a policy-provider layer containing traditionally trained policies or VLA models; a backend such as a simulator or hardware environment where those policies run; and the physical robot. The speaker compresses the division of labor into a useful rule: the agent decides what to do, while the policy decides how to do it.

Scout runs multiple concurrent Strands agents: a continuously assessing thinker, a communications agent connected to web and Telegram interfaces, and a voice agent. The system is described as hybrid cloud-edge: cloud resources train VLA models and policies using AgentCore, while edge execution enables faster runtime actions. The demo also positions deployed robots as data-collection systems, where manually guided episodes and agent reasoning can become future training data.

The presentation is high-level and demo-oriented rather than an engineering deep dive. Its own failures—ambiguous commands leading to minimal actions, imperfect visual analysis, and the rover falling during the demo—show why real deployments need hard action constraints, safety controls, evaluation, and reliable fallback behavior beyond an LLM prompt and tool list.

Key Takeaways

  • Claim: The key architecture pattern is to expose a robot's existing capabilities as tools and let an agent map natural-language goals to the appropriate tools or policies. | Evidence: Scout is instructed to turn on headlights, greet the audience, spin, and inspect the room; the speaker says an agent can be given a hardware tool—the robot—with access to preset functions or programmable policies. | Implication: For agent systems, physical automation should be modeled as a governed tool layer with explicit capabilities, rather than treating the foundation model as the source of low-level control. | Caveat: The agent can only invoke capabilities that the robot already exposes; the speaker's own example says 'pick up the red cube' works only if the robot has that capability.
  • Claim: High-level task reasoning and low-level robot execution should be separated: the agent chooses what to do, and a trained policy or VLA model determines how to carry it out. | Evidence: The speaker explicitly states, 'the agent decides what to do and the policy decides how it should be done,' and describes policy training through collected and simulated data before it becomes a VLA model. | Implication: This separation is reusable beyond robotics: retain deterministic, validated execution modules while using agents for intent interpretation, policy selection, and orchestration. | Caveat: This does not eliminate the need for robot-specific training today; the speaker frames agent orchestration as an interim step until more general VLA models become capable enough.
  • Claim: A hybrid cloud-edge control plane is presented as necessary to combine large-scale learning with low-latency physical execution. | Evidence: The architecture places Strands agents on both edge and cloud, uses cloud resources and AgentCore to train VLAs and policies, and calls policies directly at the edge for faster robot execution. | Implication: Any physical-agent design should explicitly classify which functions can tolerate cloud latency and which require local, bounded execution. | Caveat: The talk does not define routing criteria, latency targets, offline behavior, or failure handling when cloud connectivity is unavailable.
  • Claim: Scout uses multiple specialized agents concurrently rather than one monolithic conversational agent. | Evidence: The speaker identifies three Strands agents: a thinker that continually assesses the environment, a communications agent connected to a web app and Telegram, and a voice agent that was disabled during the stage talk. | Implication: Specialized agent roles can improve modality coverage, but a production system needs an explicit supervisor, shared-state design, and arbitration rules. | Caveat: The transcript does not explain how these agents coordinate state, resolve conflicting commands, or prevent one agent's observations from triggering unsafe actions.
  • Claim: An embodied agent can serve as a data-generation loop, not merely an endpoint application. | Evidence: The speaker says she can manually move Scout, create navigation training episodes, and capture how it responds and reasons about requests; this information can then be used to improve future performance. | Implication: Instrumenting agent actions, observations, policy choices, and outcomes is strategically valuable, but collected traces should be treated as evaluation and training assets requiring governance. | Caveat: The presentation does not specify labeling standards, data quality controls, consent/privacy handling for camera footage, or how reasoning traces are validated before training use.
  • Claim: Prompting and tool descriptions materially shape which robot action the agent selects. | Evidence: Scout's configuration reportedly uses Strands as the harness SDK, Anthropic Claude Opus 4.8 as the 'brain,' a system prompt defining rules and the purpose of each rule, OpenAI Realtime for voice, and additional safety/guardrail instructions. | Implication: Tool schemas, policy descriptions, action preconditions, and guardrails should be tested as product-critical control interfaces; vague natural-language autonomy is not a sufficient operating specification. | Caveat: The demo shows non-deterministic or weak behavior under vague requests: 'do something complex' led primarily to speech, headlight changes, attempted movement, and a fall.

Detailed Brief

Demo behavior and practical reliability signal

  • Claims: Natural-language interaction extends beyond direct commands into camera-grounded conversational responses.; Remote messaging can make a physical agent accessible when its operator is not nearby.
  • Evidence: Scout responds to questions about the stage and audience through its front camera, although it qualifies that its low ground-level viewpoint makes faces difficult to assess.; The communications agent is connected to Telegram and a web app; the speaker sends a Telegram instruction to spin, inspect the room, and describe people.; The rover is described as running on a Raspberry Pi located in New York City and communicating via a SIM-card-backed 4G connection.; During the live demo, Scout falls multiple times, does not always follow follow-up spin instructions, and gives a subjective, style-based answer rather than reliable person identification.
  • Caveats: The observed behavior is a stage demonstration, not evidence of robust perception, navigation, or autonomous recovery.; The speaker says Scout should be able to resume after a fall, but the demo does not establish dependable recovery.
  • Implications: Remote robot access expands the attack and failure surface: communications channels, identity, command authorization, observation privacy, and emergency stop controls need first-class treatment.; Visual-language output should not be confused with validated situational awareness; perception claims need task-specific accuracy and safety evaluation.

Framework scope and long-term framing

  • Claims: The speaker presents the framework as a thin integration layer over a broad set of robots and policy providers.; Her long-term vision is increasingly general VLA models that reduce the amount of robot-specific fine-tuning and programming required.
  • Evidence: The transcribed framework name is described as an AWS-built open-source framework and as supporting more than 40 robots across eight categories through simple robot tool calls.; The speaker characterizes the current approach as a stepping stone toward models that could have sufficiently broad physical-world knowledge to operate from far less bespoke training.
  • Caveats: The transcript does not provide a repository, compatibility matrix, benchmark, security model, policy-provider list, or evidence for the support and generalization claims.; The framework name is inconsistently transcribed as 'TransAgence,' 'Trans,' and 'trans package,' while the agent harness is clearly referred to as Strands.
  • Implications: Treat interoperability claims as a due-diligence item rather than an adoption conclusion; verify supported hardware, abstraction quality, and operational boundaries before committing to the stack.

Notable Concepts & Terms

  • Strands agents: The AWS-associated agent harness SDK used to connect an LLM to Scout's robot tools, communications interfaces, and specialized agent roles.
  • VLA model: A vision-language-action model that represents the learned policy layer translating a selected task into robot behavior.
  • AgentCore: Named as the cloud-side component used to train VLA models and policies in the proposed hybrid architecture.
  • Policy provider: The layer from which the agent selects a trained robot policy; it bridges agent intent with traditional robotics training and execution.
  • Hybrid cloud-edge execution: The design in which training and larger-scale processing occur in the cloud while latency-sensitive policy calls can run on the robot edge.
  • Thinker, communications, and voice agents: Scout's three functional agent roles, separating ongoing environment assessment, text-based remote interaction, and speech interaction.
  • OpenAI Realtime: The voice component named in Scout's configuration; it was disabled during the presentation because ambient speech would trigger interactions.
  • Claude Opus 4.8: The model named by the speaker as the high-level reasoning 'brain' behind the Strands configuration.

Operator Notes / Why Ken Should Care

  • Adopt the agent-versus-policy boundary for any embodied or high-consequence automation: agent selects approved operations; a constrained, validated executor performs them.
  • Require a command-arbitration design before using multiple concurrent agents, including priority rules for human, remote, autonomous, and safety commands.
  • Design an edge-safe mode with local execution, hard motion bounds, emergency stop, and deterministic fallback rather than relying on cloud connectivity or prompt compliance.
  • Log observation, intent, selected tool/policy, action parameters, outcome, and recovery state as a structured dataset; separate usable training data from unvalidated model-generated reasoning.
  • Before evaluating the referenced AWS framework, resolve the exact package/repository name and verify its claimed robot support, policy integrations, authentication model, and safety controls.

Source/Metadata

  • Title: Tell the Robot What You Want — Sandhya Subramani, AWS
  • Transcript words: 4786
  • Duration seconds: 1042
  • Timestamp note: No timestamps or chapters were present. The latter portion of the supplied transcript substantially repeats earlier material.

Transcript

2664 words en Processed in 89.8s

Sandi R. Hello, everyone. My name is Sandi, and meet my co-host today, Scout. This is my friendly rover, and one would think that rovers can't really think for themselves, right? We should tell them what to do, or we have to very specifically program them on how to think, but this little guy here actually has a brain, and he can think for himself. Let me show you my screen. Oh, no, it's going to the wrong screen. I'm going to see how I can stop this, and I'm going to see how I can move to my screen. Just a second. I'm going to see what Scout here is looking at, and Scout here has a small little brain, and Scout can understand what I'm saying in natural language. For example, if you can see my screen here, if I say, hey, Scout, I'm going to type to him. Say, hey, Scout, turn on your headlights, and say hi to everyone. He is actually going to be able to understand and do those actions, but how is he able to do these things, right? It's going to take him a couple of seconds to think about it, because how this has been set up is he's working with a Raspberry Pi installed, and that Raspberry Pi is actually in a small little location in New York City, and he's got a SIM card here, which is then connecting to that Raspberry Pi, and then he's getting signals. So this is via 4G connection, and so he's like, oh, hey, on it, and if you can see, I'm going to tell him, spin around and, oh, no, he's actually speaking to me. He's like, I see the stage, and it's looking great, big presentation screen to the left, bright stage lights, and rows of seats for the audience. How many people do you see? I've never trained him to answer this question, but let's see if he's getting it. He's thinking for a second or two. Let me take a good look. From the front camera, I see two people, one near the speaker monitor set up at the back of the stage, and another person further right. Nah, spin 360. And he's also saying, is this session about to start, right? So he's going to spin 360 any second now, and then he's probably going to be like, oh, wow, I did. So this little robot here is a next generation of robot where, ah, there we go. He is spinning 360 now, and he's probably going to tell me what he's seeing. And he's saying, let's spin, right? So this new generation of robots is different from my traditional robot training because I have given this guy a little brain. And what do I mean by I've given him a brain? I've given this robot an agent, and I've given it, it's called TransAgence, which is an open source framework which was built by AWS. And I'm going to say, what happens is we have these existing tools that the robot can do. He can take certain actions by himself, but only those actions by himself. So what we can do is we can add a layer of LLM, or even better, add a layer of agent to it so that the agent orchestrates which tool to call and how to really get the robot to start doing the things we want. So in traditional software with traditional AI, like AI engineering, we can give agents software tools. Similarly, we can give the same AI agent a hardware tool called a robot which has access to preset functions or programmable policies, and then the agent can decide which policy to implement when. So all it takes is one robot agent for us to be able to do new and numerous tasks and have it understand what we're teaching it in natural language. So how do we get started with it? All it takes is five lines of code. This is through the agent harness called strands, and all we have to do is import the strands of the agent and call the robot tool. And we say tools equals the robot, and then we say pick up the red cube and should be able to pick up the red cube, assuming that the robot has that capability. Yeah, now he's seen someone and he's like, oh, let me go towards that person. So he gets pretty excited. This guy is pretty special because he doesn't have just one agent. He's got three different agents. All three of them are strands, and all three of them are working simultaneously. One of them is the thinker agent, and that's the part of him that's constantly thinking and assessing the environment and, what do I do next? And that part of his brain is constantly thinking. Then there's the other communication part of it, and I'm going to show you that in a bit, right? And I've connected him to my Telegram app as well as to my web app, and so he is able to have a conversation with me in natural language and then take actions based on what I'm telling him to do, apart from him just perceiving and thinking and figuring out what he wants to do. And the third agent, the third type of agent that he's got access to, is the voice agent. I did have to disable it because every time I speak, he's going to think I'm speaking to him, and so he's going to keep chatting away with me, and it's just not going to be fun because we're going to have our co-host interrupting me all the time. So I disabled that feature for the time being, but essentially, all three of these agents work in tandem with this one robot, and thereby this gives him the ability to do way more than just what he's been trained to do, more than just the policies that he's learned. Now, what is a quick overview on this trans package itself? This trans package supports more than 40 different robots under eight categories, and all of these are just simple robot tool calls. And how is this all set up? Four different layers. The first one is the agent layer, the topmost one. And there are two parts to this. One is how the actions go in, and the second is how it observes and the observations go up. So if you notice, it's very bi-directional. So first, when we give it an instruction, we would be talking to this trans agent, which is the agentic layer. That would then decide which policy to call. And the policy provider, again, a trans agent supports a bunch of different policy providers, and we can then train our policy based on our traditional robot training. So in our policies, we would collect data, and then we would train on it, and we would create more simulation data, and that policy then becomes a VLA model, which then the robot would have access to, trans agents would have access to, and then it would invoke that specific policy based on the question that we're asking it or the command that we're giving it. And that policy needs to sit somewhere, right? So that sits in the back end, which could be your simulation environment, or it could be a real hardware chip, your hardware environment. That is the back end on which, that is the interface on which the policy is running. And finally, the output actually takes place in the physical hardware, which is the robot. And so the robot, ah, see? So now it's responding. This, even if he falls down, he's supposed to be fine. He technically shouldn't get hurt. He should be able to pick back up from where he stops. Ah, okay. So I'm telling him to go back a bit. Back off. Let's see if he actually backs off. So that is the four layers of how to get started with building this, right? And what's happening under the hood? A more picturistic view of what's the architecture of what's going on under the hood. We want everything is strands agents on the edge as well as on the cloud. We want to be able to train the VLA and the policies using agent core. And we want that to happen on the cloud. But we also want to be able to call it directly on edge so that our robot can execute functions and policies faster. So this is a hybrid model where a part of it happens on the cloud and another part of it happens on the edge. And strands can decide when to call which part of it. And so this helps with massive amounts of training as well when it's constantly collecting information and it's able to train on that information and learn from itself, but also just execute at runtime really, really quickly. Now, like I said, the agent decides what to do and the policy decides how it should be done. But he's pretty smart. He should be able to pick himself back up if he's not fully fallen down. And he should be able to continue moving along. So I think he's okay. Now, where does this leave us and why is this so special? We started off with very traditional robots. Robots have existed since forever, right? And they've always just been pre-programmed to be automated and do a certain set of tasks autonomously. But there is a future in this world where these robot policies, these VLA models, could be so advanced that we wouldn't even need to do this. They could be as large as our large language models so that, ah, wait. Hang on. He's falling back again. I'm going to see if I can get him to move back up. Good boy. Stop. And then he's fallen off again. We get to a point where these VLA models could be as large and as amazing as our larger language models, and they have all the information in the world. And we wouldn't even have to do this. We might just have to feed in one simple model, and then we could give it to him, and then he would know exactly what to do. But until that point where we don't have to fine-tune on top of existing VLAs and existing policies, we can do this. And this is a stepping stone towards a future where we don't need to train robots anymore. So now if we wanted to do more things than just the tasks it's trained on, give it an agent and see what it can do. And so let me quickly go back to my demo, and I'm going to show you how it's actually working. Okay. So this is my, so this is Trans here. This is Scout here. And I've been telling him to do a bunch of things. So I can say, hey, do something complex. That's not complex. He's going to be thinking now. He's going to fall off. So he's saying let's spin. Full 360 done. Still safely on the stage. I can see the bright stage lights and the audience seating area. All good. What's there? Ooh, a challenge. So he's speaking. I call this my signature performance. But he's not doing anything. What are you doing? He clearly seems to be speaking. But what are you doing? Please do something. He just turned off his headlights. Cool. Okay. Now he's calling. So do you see it saying calling over speak, which was the function that it called because I said do something complex? So now it spoke. But now I think it should have been attempting to do something. And it fell off because it tried doing something. I've actually seen it do a funky dance, this funky dance move. But he's got a mind of his own, right? Now what's going on under the hood here? A couple of things. The first thing is here I can use this. What is the point of creating him? I can use him to create my data sets. Because I'm able to also manually move him, I will get him to navigate in the direction that I want him to. And then I can create training episodes and I can get information on how he's responding and how he's reasoning based on the questions that I ask. And this is super good information for me to then be able to make him do a better job of it. So that's one part of this whole process and this experiment of giving him his own autonomy and getting him to do things so that I can create more data. But also, apart from that, this is my configuration. So over here, under the hood, strands agents, which is your harness SDK, is using currently Anthropic Cloud Opus 4.8 under the hood. So that is the brain. And then this is my simple system prompt where I'm telling it what it's supposed to be doing. And I'm telling it all of the rules, and I'm also giving it access to all of the rules that it's already got. So I'm telling it what each of these rules are meant for. And so that's how strands decides which tool to invoke based on what I'm asking it to do. And the voice that it's using is the one of OpenAI real time. And I've also given it more information for it to be able to, just safety and guardrails to ensure that it's doing really well. Now, these are two of the agents. The other thing that it can do is also chat with me on Telegram. This is amazing because when I'm not at home and I still want to get it to speak to me, I can say, hey, Scout. Turn around, spin around, analyze, don't fall off. Analyze the room and tell me who is the best looking person. This is always a hit or miss. Don't blame me if Scout decides to be biased. I tried playing with this because it makes me feel like Snow White and the Queen saying mirror, mirror on the wall, who is the fairest of them all? And so let's see if strands can give us that. So it's saying, ha, ha, on and on there. Spinning around to scope the room. Is he spinning? He is. Looks like he's making, and let's look at what he's looking at. Now, he's still looking at the stage. This is the other mirror. He's looking at me. Hi. Am I the prettiest? Oh, you flatter me. But no, turn again. Spin some more, please. He's decided not to spin more. Let's see what he says. Spin complete. Here is my analysis of the room. You can see my screen, right? I have no way of zooming this in. Spotted six to seven people total. A couple seated front center near the bright stage right. A few more in the audience. One person in a dark outfit on the stage. The rocking cocky jeans and sneakers. Thank you very much. Oh, no, no, no. The person in front of me. Cocky jeans and sneakers. Clearly walking with purpose on stage. Best looking. Honestly, for my low-angle ground view, it's tough to tell faces. But the person on the stage right now gets style points for the confidence stride and the dramatic blue curtain backdrop. Thank you very much. Every time I need a confidence boost, I can be the only person in the room and ask this guy who's the prettiest of them all, and he will always choose me as the answer. Thank you very much. Thank you very much. robot to start doing the things we want. So in traditional software with traditional AI, like AI engineering, we can give agents software tools. Similarly, we can give the same AI agent a hardware tool called a robot which has access to preset functions or programmable policies, and then the agent can decide which policy to implement when. So all it takes is one robot agent for us to be able to do new and numerous tasks and have it understand what we're teaching it in natural language. So how do we get started with it? All it takes is five lines of code. This is through the agent harness called strands, and all we have to do is import the strands of the agent and call the robot tool. And we say tools equals the robot, and then we say pick up the red cube and should be able to pick up the red cube, assuming that the robot has that capability. Yeah, now he's seen someone and he's like, oh, let me go towards that person. So he gets pretty excited. This guy is pretty special because he doesn't have just one agent. He's got three different agents. All three of them are strands, and all three of them are working simultaneously. One of them is the thinker agent, and that's the part of him that's constantly thinking and assessing the environment and like, what do I do next? And that guy is, that part of his brain is constantly thinking. Then there's the other communication part of it, and I'm going to show you that in a bit, right? And I've connected him to my telegram app as well as to my web app, and so he is able to have a conversation with me in natural language and then take actions based on what I'm telling him to do. Apart from him just perceiving and thinking and figuring out what he wants to do. And the third agent, the third type of agent that he's got access to is the voice agent. I did have to disable it because every time I speak, he's going to think I'm speaking to him, and so he's going to keep chatting away with me, and it's just not going to be fun because we're going to have our co-host interrupting me all the time. So I disabled that feature for the time being, but essentially, all three of these agents work in tandem with this one robot, and thereby this gives him the ability to do way more than just what he's been trained to do, more than just the policies that he's learned. Now, what is a quick overview on this trans package itself? This trans package has more than supports more than 40 different robots under eight categories, and all of these are just simple robot tool calls. And how is this all set up? Four different layers. The first one is the agent layer, the topmost one. And there are two parts to this. One is how the actions go in, and the second is how it observes and the observations go up. So if you notice, it's very bi-directional. So first, when we give it an instruction, we would be talking to this trans agent, which is the agentic layer. That would then decide which policy to call. And the policy provider, again, a trans agent supports a bunch of different policy providers, and we can then train our policy based on our traditional robot training. So in our policies, we would collect data, and then we would train on it, and we would create more simulation data, and that policy then becomes a VLA model, which then the robot would have access to, trans agents would have access to, and then it would invoke that specific policy based on the question that we're asking it or the command that we're giving it. And that policy needs to sit somewhere, right? So that sits in the back end, which could be your simulation environment, or it could be a real hardware chip, your hardware environment. That is the back end on which, that is the interface on which the policy is running. And finally, the output actually takes place in the physical hardware, which is the robot. And so the robot, ah, see? So now it's responding. This, even if he falls down, he's supposed to be fine. He technically shouldn't, he technically shouldn't get hurt. He should be able to pick back up from where he stops. Ah, okay. So I'm telling him to go back a bit. Back off. Let's see if he actually backs off. So that is the four layers of how to get started with building this, right? And what's happening under the hood? Like a more picturistic view of what's the architecture of what's going on under the hood. We want everything is basically strands agents on the edge as well as on the cloud. We want to be able to train the VLA and the policies on with using agent core. And we want that to happen on the cloud. But we also want to be able to call it directly on edge so that our robot can execute functions and policies faster. So this is sort of like a hybrid model where a part of it happens on the cloud and another part of it happens on the edge. And strands can decide when to call which part of it. And so this helps with massive amounts of training as well when it's constantly collecting information and it's able to train on that information and learn from itself but also just execute at runtime really, really quickly. Now, like I said, the agent decides what to do and the policy decides how it should be done. But he's pretty smart. He should be able to pick himself back up if he's not fully fallen down. And he should be able to continue moving along. So I think he's okay. Now, where does this leave us and why is this so special? We started off with very traditional robots. Robots have existed since forever. Right? And they've always just been programmed, pre-programmed to be automated and do a certain set of tasks autonomously. But there is a future in this world, where these robot policies, these VLA models could be so advanced that we wouldn't even need to do this. They could be as large as our large language models so that... Ah, wait. Hang on. He's falling back again. I'm going to see if I can get him to move back up. Good boy. Stop. And then he's fallen off again. We get to a point where these large language... The VLA models could be as large and as amazing as our larger language models. And they have all the information in the world. And we wouldn't even have to do this. We might just have to feed in one simple model. And then we could give it to him. And then he would know exactly what to do. But until that point where we don't have to fine tune on top of existing VLA's and existing policies, we can do this. And this is a stepping stone towards a future where we don't need to train robots anymore. So now if we wanted to do more things than just the tasks it's trained on, give it an agent and see what it can do. And so let me quickly go back to my demo. And I'm going to show you how it's actually working. Okay. So this is my... So this is Trans here. This is Scout here. And I've been telling him to do a bunch of things. So I can say, hey, do something complex. That's not complex. He's going to be thinking now. He's going to fall off. So he's saying let's spin. Full 360 done. Still safely on the stage. I can see the bright stage lights and the audience seating area. All good. What's there? Ooh, a challenge. So he's speaking. I call this my signature performance. But he's not doing anything. What are you doing? He clearly seems to be speaking. But what are you doing? Please do something. He just turned off his headlines. Cool. Okay. Now he's calling. So do you see it saying calling over speak, which was the function that it called because I said do something complex? So now it spoke. But now I think it should have been attempting to do something. And it fell off because it tried doing something. I've actually seen it do like a funky dance. Like this funky dance move. But he's got a mind of his own. Right? Now what's going on under the hood here? A couple of things. The first thing is here I can use this. What is the point of creating him? I can use him to create my data sets. Because I'm able to also manually move him, I will get him to navigate in the direction that I want him to. And then I can create training episodes and I can get information on how he's responding and how he's reasoning based on the questions that I ask. And this is super good information for me to then be able to make him do a better job of it. So that's one part of this whole process and this experiment of giving him his own autonomy and getting him to do things so that I can create more data. But also, apart from that, this is my configuration. So over here, under the hood, strands agents, which is your harness SDK, is using currently Anthropic Cloud Opus 4.8 under the hood. So that is the brain. And then this is my simple system prompt where I'm telling it what it's supposed to be doing. And I'm telling it all of the rules and I'm also giving it access to all of the rules that it's already got. So I'm telling it what each of these rules are meant for. And so that's how strands decides which tool to invoke based on what I'm asking it to do. And the voice that it's using is the one of OpenAI real time. And I've also given it more information for it to be able to, like, just safety and guardrails to ensure that it's doing really well. Now, these are two of the agents. The other thing that it can do is also chat with me on Telegram. This is amazing because when I'm not at home and I still want to get it to speak to me, I can say, hey, Scout. Who is turn around, spin around, analyze, don't fall off. Analyze the room and tell me who is the best looking person. This is always a hit or miss. Don't blame me if Scout decides to be biased. I tried playing with this because it makes me feel like Snow White and the Queen saying mirror, mirror on the wall. Who is the fairest of them all? And so let's see if strands can give us that. So it's saying, ha, ha, on and on there. Spinning around to scope the room. Is he spinning? He is. Looks like he's making, and let's look at what he's looking at. Now, he's still looking at the stage. This is the other mirror. He's looking at me. Hi. Am I the prettiest? Oh, you flatter me. But no, turn again. Spin some more, please. He's decided not to spin more. Let's see what he says. Spin complete. Here is my analysis of the room. You can see my screen, right? I have no way of zooming this in. Spotted six to seven people total. A couple seated front center near the bright stage right. A few more in the audience. One person in a dark outfit on the stage. The rocking cocky jeans and sneakers. Thank you very much. Oh, no, no, no. The person in front of me. Cocky jeans and sneakers. Clearly walking with purpose on stage. Best looking. Honestly, for my low angle ground view, it's tough to tell faces. But the person on the stage right now gets style points for the confidence stride and the dramatic blue curtain backdrop. Thank you very much. Every time I need a confidence boost, I can be the only person in the room and ask this guy who's the prettiest of them all, and he will always choose me as the answer. Thank you very much. Thank you very much.