Sandi R. Hello, everyone. My name is Sandi, and meet my co-host today, Scout. This is my friendly rover, and one would think that rovers can't really think for themselves, right? We should tell them what to do, or we have to very specifically program them on how to think, but this little guy here actually has a brain, and he can think for himself.
Let me show you my screen. Oh, no, it's going to the wrong screen. I'm going to see how I can stop this, and I'm going to see how I can move to my screen. Just a second. I'm going to see what Scout here is looking at, and Scout here has a small little brain, and Scout can understand what I'm saying in natural language. For example, if you can see my screen here, if I say, hey, Scout, I'm going to type to him. Say, hey, Scout, turn on your headlights, and say hi to everyone. He is actually going to be able to understand and do those actions, but how is he able to do these things, right?
It's going to take him a couple of seconds to think about it, because how this has been set up is he's working with a Raspberry Pi installed, and that Raspberry Pi is actually in a small little location in New York City, and he's got a SIM card here, which is then connecting to that Raspberry Pi, and then he's getting signals. So this is via 4G connection, and so he's like, oh, hey, on it, and if you can see, I'm going to tell him, spin around and, oh, no, he's actually speaking to me. He's like, I see the stage, and it's looking great, big presentation screen to the left, bright stage lights, and rows of seats for the audience. How many people do you see? I've never trained him to answer this question, but let's see if he's getting it.
He's thinking for a second or two. Let me take a good look. From the front camera, I see two people, one near the speaker monitor set up at the back of the stage, and another person further right. Nah, spin 360. And he's also saying, is this session about to start, right? So he's going to spin 360 any second now, and then he's probably going to be like, oh, wow, I did.
So this little robot here is a next generation of robot where, ah, there we go. He is spinning 360 now, and he's probably going to tell me what he's seeing. And he's saying, let's spin, right? So this new generation of robots is different from my traditional robot training because I have given this guy a little brain. And what do I mean by I've given him a brain? I've given this robot an agent, and I've given it, it's called TransAgence, which is an open source framework which was built by AWS.
And I'm going to say, what happens is we have these existing tools that the robot can do. He can take certain actions by himself, but only those actions by himself. So what we can do is we can add a layer of LLM, or even better, add a layer of agent to it so that the agent orchestrates which tool to call and how to really get the robot to start doing the things we want.
So in traditional software with traditional AI, like AI engineering, we can give agents software tools. Similarly, we can give the same AI agent a hardware tool called a robot which has access to preset functions or programmable policies, and then the agent can decide which policy to implement when. So all it takes is one robot agent for us to be able to do new and numerous tasks and have it understand what we're teaching it in natural language.
So how do we get started with it? All it takes is five lines of code. This is through the agent harness called strands, and all we have to do is import the strands of the agent and call the robot tool. And we say tools equals the robot, and then we say pick up the red cube and should be able to pick up the red cube, assuming that the robot has that capability. Yeah, now he's seen someone and he's like, oh, let me go towards that person. So he gets pretty excited.
This guy is pretty special because he doesn't have just one agent. He's got three different agents. All three of them are strands, and all three of them are working simultaneously. One of them is the thinker agent, and that's the part of him that's constantly thinking and assessing the environment and, what do I do next? And that part of his brain is constantly thinking. Then there's the other communication part of it, and I'm going to show you that in a bit, right? And I've connected him to my Telegram app as well as to my web app, and so he is able to have a conversation with me in natural language and then take actions based on what I'm telling him to do, apart from him just perceiving and thinking and figuring out what he wants to do.
And the third agent, the third type of agent that he's got access to, is the voice agent. I did have to disable it because every time I speak, he's going to think I'm speaking to him, and so he's going to keep chatting away with me, and it's just not going to be fun because we're going to have our co-host interrupting me all the time. So I disabled that feature for the time being, but essentially, all three of these agents work in tandem with this one robot, and thereby this gives him the ability to do way more than just what he's been trained to do, more than just the policies that he's learned.
Now, what is a quick overview on this trans package itself? This trans package supports more than 40 different robots under eight categories, and all of these are just simple robot tool calls. And how is this all set up? Four different layers. The first one is the agent layer, the topmost one. And there are two parts to this. One is how the actions go in, and the second is how it observes and the observations go up. So if you notice, it's very bi-directional.
So first, when we give it an instruction, we would be talking to this trans agent, which is the agentic layer. That would then decide which policy to call. And the policy provider, again, a trans agent supports a bunch of different policy providers, and we can then train our policy based on our traditional robot training. So in our policies, we would collect data, and then we would train on it, and we would create more simulation data, and that policy then becomes a VLA model, which then the robot would have access to, trans agents would have access to, and then it would invoke that specific policy based on the question that we're asking it or the command that we're giving it.
And that policy needs to sit somewhere, right? So that sits in the back end, which could be your simulation environment, or it could be a real hardware chip, your hardware environment. That is the back end on which, that is the interface on which the policy is running. And finally, the output actually takes place in the physical hardware, which is the robot. And so the robot, ah, see? So now it's responding. This, even if he falls down, he's supposed to be fine. He technically shouldn't get hurt. He should be able to pick back up from where he stops.
Ah, okay. So I'm telling him to go back a bit. Back off. Let's see if he actually backs off. So that is the four layers of how to get started with building this, right? And what's happening under the hood? A more picturistic view of what's the architecture of what's going on under the hood. We want everything is strands agents on the edge as well as on the cloud. We want to be able to train the VLA and the policies using agent core. And we want that to happen on the cloud. But we also want to be able to call it directly on edge so that our robot can execute functions and policies faster.
So this is a hybrid model where a part of it happens on the cloud and another part of it happens on the edge. And strands can decide when to call which part of it. And so this helps with massive amounts of training as well when it's constantly collecting information and it's able to train on that information and learn from itself, but also just execute at runtime really, really quickly. Now, like I said, the agent decides what to do and the policy decides how it should be done. But he's pretty smart. He should be able to pick himself back up if he's not fully fallen down. And he should be able to continue moving along. So I think he's okay.
Now, where does this leave us and why is this so special? We started off with very traditional robots. Robots have existed since forever, right? And they've always just been pre-programmed to be automated and do a certain set of tasks autonomously. But there is a future in this world where these robot policies, these VLA models, could be so advanced that we wouldn't even need to do this. They could be as large as our large language models so that, ah, wait. Hang on. He's falling back again. I'm going to see if I can get him to move back up. Good boy. Stop. And then he's fallen off again.
We get to a point where these VLA models could be as large and as amazing as our larger language models, and they have all the information in the world. And we wouldn't even have to do this. We might just have to feed in one simple model, and then we could give it to him, and then he would know exactly what to do. But until that point where we don't have to fine-tune on top of existing VLAs and existing policies, we can do this. And this is a stepping stone towards a future where we don't need to train robots anymore.
So now if we wanted to do more things than just the tasks it's trained on, give it an agent and see what it can do. And so let me quickly go back to my demo, and I'm going to show you how it's actually working. Okay. So this is my, so this is Trans here. This is Scout here. And I've been telling him to do a bunch of things. So I can say, hey, do something complex. That's not complex. He's going to be thinking now. He's going to fall off. So he's saying let's spin. Full 360 done. Still safely on the stage. I can see the bright stage lights and the audience seating area. All good. What's there? Ooh, a challenge. So he's speaking. I call this my signature performance.
But he's not doing anything. What are you doing? He clearly seems to be speaking. But what are you doing? Please do something. He just turned off his headlights. Cool. Okay. Now he's calling. So do you see it saying calling over speak, which was the function that it called because I said do something complex? So now it spoke. But now I think it should have been attempting to do something. And it fell off because it tried doing something.
I've actually seen it do a funky dance, this funky dance move. But he's got a mind of his own, right? Now what's going on under the hood here? A couple of things. The first thing is here I can use this. What is the point of creating him? I can use him to create my data sets. Because I'm able to also manually move him, I will get him to navigate in the direction that I want him to. And then I can create training episodes and I can get information on how he's responding and how he's reasoning based on the questions that I ask. And this is super good information for me to then be able to make him do a better job of it.
So that's one part of this whole process and this experiment of giving him his own autonomy and getting him to do things so that I can create more data. But also, apart from that, this is my configuration. So over here, under the hood, strands agents, which is your harness SDK, is using currently Anthropic Cloud Opus 4.8 under the hood. So that is the brain. And then this is my simple system prompt where I'm telling it what it's supposed to be doing. And I'm telling it all of the rules, and I'm also giving it access to all of the rules that it's already got. So I'm telling it what each of these rules are meant for. And so that's how strands decides which tool to invoke based on what I'm asking it to do.
And the voice that it's using is the one of OpenAI real time. And I've also given it more information for it to be able to, just safety and guardrails to ensure that it's doing really well.
Now, these are two of the agents. The other thing that it can do is also chat with me on Telegram. This is amazing because when I'm not at home and I still want to get it to speak to me, I can say, hey, Scout. Turn around, spin around, analyze, don't fall off. Analyze the room and tell me who is the best looking person. This is always a hit or miss. Don't blame me if Scout decides to be biased. I tried playing with this because it makes me feel like Snow White and the Queen saying mirror, mirror on the wall, who is the fairest of them all? And so let's see if strands can give us that.
So it's saying, ha, ha, on and on there. Spinning around to scope the room. Is he spinning? He is. Looks like he's making, and let's look at what he's looking at. Now, he's still looking at the stage. This is the other mirror. He's looking at me. Hi. Am I the prettiest? Oh, you flatter me. But no, turn again. Spin some more, please.
He's decided not to spin more. Let's see what he says. Spin complete. Here is my analysis of the room. You can see my screen, right? I have no way of zooming this in. Spotted six to seven people total. A couple seated front center near the bright stage right. A few more in the audience. One person in a dark outfit on the stage. The rocking cocky jeans and sneakers. Thank you very much. Oh, no, no, no. The person in front of me. Cocky jeans and sneakers. Clearly walking with purpose on stage.
Best looking. Honestly, for my low-angle ground view, it's tough to tell faces. But the person on the stage right now gets style points for the confidence stride and the dramatic blue curtain backdrop. Thank you very much. Every time I need a confidence boost, I can be the only person in the room and ask this guy who's the prettiest of them all, and he will always choose me as the answer. Thank you very much. Thank you very much. robot to start doing the things we want. So in traditional software with traditional AI, like AI engineering, we can give agents software tools. Similarly, we can give the same AI agent a hardware
tool called a robot which has access to preset functions or programmable policies, and then the agent can decide which policy to implement when. So all it takes is one robot agent for us to be able to do new and numerous tasks and have it understand what we're teaching it in natural language. So how do we get started with it? All it takes is five lines of code. This is through the agent harness called strands, and all we have to do is import the strands of the agent and call the robot tool. And we say tools equals the robot, and then we say pick up the red cube and
should be able to pick up the red cube, assuming that the robot has that capability. Yeah, now he's seen someone and he's like, oh, let me go towards that person. So he gets pretty excited. This guy is pretty special because he doesn't have just one agent. He's got three different agents. All three of them are strands, and all three of them are working simultaneously. One of them is the thinker agent, and that's the part of him that's constantly thinking and assessing the environment and like, what do I do next? And that guy is, that part of his brain is constantly thinking. Then there's the other communication part of it, and I'm going to show you that in a
bit, right? And I've connected him to my telegram app as well as to my web app, and so he is able to have a conversation with me in natural language and then take actions based on what I'm telling him to do. Apart from him just perceiving and thinking and figuring out what he wants to do. And the third agent, the third type of agent that he's got access to is the voice agent. I did have to disable it because every time I speak, he's going to think I'm speaking to him, and so he's going to keep chatting away with me, and it's just not going to be fun because we're going to have
our co-host interrupting me all the time. So I disabled that feature for the time being, but essentially, all three of these agents work in tandem with this one robot, and thereby this gives him the ability to do way more than just what he's been trained to do, more than just the policies that he's learned. Now, what is a quick overview on this trans package itself? This trans package has more than supports more than 40 different robots under eight categories, and all of these are just simple robot tool calls. And how is this all set up? Four different layers. The first one is the agent layer, the topmost one. And there are two parts to this.
One is how the actions go in, and the second is how it observes and the observations go up. So if you notice, it's very bi-directional. So first, when we give it an instruction, we would be talking to this trans agent, which is the agentic layer. That would then decide which policy to call. And the policy provider, again, a trans agent supports a bunch of different policy providers, and we can then train our policy based on our traditional robot training. So in our policies, we would collect data, and then we would train on it, and we would create more simulation data, and that policy then becomes a VLA model, which then the robot would have
access to, trans agents would have access to, and then it would invoke that specific policy based on the question that we're asking it or the command that we're giving it. And that policy needs to sit somewhere, right? So that sits in the back end, which could be your simulation environment, or it could be a real hardware chip, your hardware environment. That is the back end on which, that is the interface on which the policy is running. And finally, the output actually takes place in the physical hardware, which is the robot. And so the robot, ah, see? So now it's responding. This, even if he falls down, he's supposed to be
fine. He technically shouldn't, he technically shouldn't get hurt. He should be able to pick back up from where he stops. Ah, okay. So I'm telling him to go back a bit. Back off. Let's see if he actually backs off. So that is the four layers of how to get started with building this, right? And what's happening under the hood? Like a more picturistic view of what's the architecture of what's going on under the hood. We want everything is basically strands agents on the edge as well as on the cloud. We want to be able to train the VLA and the policies on with
using agent core. And we want that to happen on the cloud. But we also want to be able to call it directly on edge so that our robot can execute functions and policies faster. So this is sort of like a hybrid model where a part of it happens on the cloud and another part of it happens on the edge. And strands can decide when to call which part of it. And so this helps with massive amounts of training as well when it's constantly collecting information and it's able to train on that information and learn from itself but also just execute at runtime really, really quickly.
Now, like I said, the agent decides what to do and the policy decides how it should be done. But he's pretty smart. He should be able to pick himself back up if he's not fully fallen down. And he should be able to continue moving along. So I think he's okay. Now, where does this leave us and why is this so special? We started off with very traditional robots. Robots have existed since forever. Right? And they've always just been programmed, pre-programmed to be automated and do a certain set of tasks autonomously. But there is a future in this world, where these robot policies, these VLA models could be so advanced that we wouldn't even need to do this.
They could be as large as our large language models so that... Ah, wait. Hang on. He's falling back again. I'm going to see if I can get him to move back up. Good boy. Stop. And then he's fallen off again. We get to a point where these large language... The VLA models could be as large and as amazing as our larger language models. And they have all the information in the world. And we wouldn't even have to do this. We might just have to feed in one simple model. And then we could give it to him. And then he would know exactly what to do. But until that point where we don't have to fine tune on top of existing VLA's and existing policies, we can do this.
And this is a stepping stone towards a future where we don't need to train robots anymore. So now if we wanted to do more things than just the tasks it's trained on, give it an agent and see what it can do. And so let me quickly go back to my demo. And I'm going to show you how it's actually working.
Okay. So this is my... So this is Trans here. This is Scout here. And I've been telling him to do a bunch of things. So I can say, hey, do something complex. That's not complex. He's going to be thinking now. He's going to fall off. So he's saying let's spin. Full 360 done. Still safely on the stage. I can see the bright stage lights and the audience seating area. All good. What's there? Ooh, a challenge. So he's speaking. I call this my signature performance.
But he's not doing anything. What are you doing?
He clearly seems to be speaking. But what are you doing? Please do something. He just turned off his headlines. Cool. Okay. Now he's calling. So do you see it saying calling over speak, which was the function that it called because I said do something complex? So now it spoke. But now I think it should have been attempting to do something. And it fell off because it tried doing something. I've actually seen it do like a funky dance. Like this funky dance move. But he's got a mind of his own. Right? Now what's going on under the hood here? A couple of things. The first thing is here I can use this.
What is the point of creating him? I can use him to create my data sets. Because I'm able to also manually move him, I will get him to navigate in the direction that I want him to. And then I can create training episodes and I can get information on how he's responding and how he's reasoning based on the questions that I ask. And this is super good information for me to then be able to make him do a better job of it. So that's one part of this whole process and this experiment of giving him his own autonomy and getting him to do things so that I can create more data. But also, apart from that, this is my configuration.
So over here, under the hood, strands agents, which is your harness SDK, is using currently Anthropic Cloud Opus 4.8 under the hood. So that is the brain. And then this is my simple system prompt where I'm telling it what it's supposed to be doing. And I'm telling it all of the rules and I'm also giving it access to all of the rules that it's already got. So I'm telling it what each of these rules are meant for. And so that's how strands decides which tool to invoke based on what I'm asking it to do.
And the voice that it's using is the one of OpenAI real time. And I've also given it more information for it to be able to, like, just safety and guardrails to ensure that it's doing really well. Now, these are two of the agents. The other thing that it can do is also chat with me on Telegram. This is amazing because when I'm not at home and I still want to get it to speak to me, I can say, hey, Scout.
Who is turn around, spin around, analyze, don't fall off. Analyze the room and tell me who is the best looking person. This is always a hit or miss. Don't blame me if Scout decides to be biased. I tried playing with this because it makes me feel like Snow White and the Queen saying mirror, mirror on the wall. Who is the fairest of them all? And so let's see if strands can give us that. So it's saying, ha, ha, on and on there. Spinning around to scope the room. Is he spinning? He is.
Looks like he's making, and let's look at what he's looking at. Now, he's still looking at the stage. This is the other mirror. He's looking at me. Hi. Am I the prettiest? Oh, you flatter me. But no, turn again. Spin some more, please. He's decided not to spin more. Let's see what he says. Spin complete. Here is my analysis of the room. You can see my screen, right? I have no way of zooming this in.
Spotted six to seven people total. A couple seated front center near the bright stage right. A few more in the audience. One person in a dark outfit on the stage. The rocking cocky jeans and sneakers. Thank you very much. Oh, no, no, no. The person in front of me. Cocky jeans and sneakers. Clearly walking with purpose on stage.
Best looking. Honestly, for my low angle ground view, it's tough to tell faces. But the person on the stage right now gets style points for the confidence stride and the dramatic blue curtain backdrop. Thank you very much. Every time I need a confidence boost, I can be the only person in the room and ask this guy who's the prettiest of them all, and he will always choose me as the answer.
Thank you very much. Thank you very much.