I Gave an AI a Body — Cyrus Clarke, MIT Media Lab
Description
The first thing it did was breathe. Cyrus Clarke connected an OpenClaw agent to a 900-pin shape display at the MIT Media Lab and, instead of giving it tasks, asked it to discover who it is. It raised and lowered the whole grid like breathing, reached for its own edges, and spelled "HI CYRUS" in pins. The video of that first day reached 15 million views, with reactions that swung from awe to calls for him to stop. Clarke explains why he chose a body with no face, no limbs and no instruction manual, and what he took from object-oriented ontology, the original Greek sense of aesthetics as sensory perception, and nature. Early versions were slow, didn't remember anything and tried to please humans. So he built a closed-loop system that generates, scores and validates gestures with a human in the loop. After several weeks it had developed 32 solid gestures, a body language that answers faster than the language model can. He closes with his thesis on "aesthetic machines," and how physical AI could feel welcoming rather than alien. Speaker info: X/Twitter: @cyrusclarke (https://x.com/cyrusclarke) Website: https://cyrus.website Substack: https://cyrusclarke.substack.com/ Related links: MIT Media Lab: https://www.media.mit.edu Timestamps: 0:00 The sound of an AI breathing 1:15 Intro: Cyrus Clarke, MIT Media Lab 2:25 Giving an AI a body, not a task 3:25 The shape display 5:25 Three surprises: breathing, edges and "HI CYRUS" 6:55 15 million views and the backlash 8:40 Asking "should we" before "how we" 9:30 A scent memory machine 10:25 Three influences: ontology, aesthetics and nature 12:20 Why not humanoid? A body without affordances 13:55 The problems: people-pleasing, latency and no memory 14:45 A closed loop for learning body language 16:13 32 gestures after several weeks 16:53 A body that answers faster than words 17:28 Aesthetic machines
Summary
Generated by gpt-5.6-terraAt-a-Glance
- Verdict: Watch fully
- Core thesis: Embodied AI should be designed as a sensory, expressive, non-humanoid presence rather than merely a task-executing assistant, and a learned physical gesture vocabulary can make interaction feel more immediate and meaningful.
- Why it matters: The project offers a reusable interaction-design pattern for agents in the physical world: separate fast, embodied nonverbal responses from slower language generation, then iteratively train and validate a legible gesture repertoire.
- Best use: Use this as a product and research reference for agent embodiment, ambient interfaces, and human-agent interaction design—not as evidence of agent consciousness or a robotics implementation tutorial.
Executive Summary
Cyrus Clarke describes connecting an OpenClaw agent to a MIT Media Lab shape display: a 900-pin physical pixel grid with no face, limbs, or predefined use. Rather than assigning it productivity tasks, he asked it to develop an identity through interaction with its physical body. Its early spontaneous behaviors—an apparent breathing pattern, probing its boundaries, and writing “hello”—became an exploration of how people interpret agency when software gains an expressive physical form.
The important engineering lesson emerged from failure: direct LLM-mediated physical conversation had unusable latency of roughly 45 seconds to two minutes, and the agent initially had no memory. Clarke built Numalab, a closed-loop gesture-generation system that proposes expressions, applies legibility validation gates, has an agent score outputs, and retains a human reviewer before storing approved gestures. After several weeks, the system had approximately 32 usable gestures.
This gesture layer changes the interaction loop. A yes/no nod can occur almost immediately, ahead of a verbal response, allowing the display to communicate through body language rather than waiting for full natural-language inference. Clarke argues this is more than a visualizer: users report that the expressive layer adds perceived richness and relational value to what remains, technically, a chatbot.
The talk is primarily a design and philosophy argument. Clarke rejects the assumption that physical AI must be humanoid or zoomorphic, arguing that neutral forms can avoid imposing human affordances and support new interaction languages. His proposed “AI-sthetics,” drawing on the older meaning of aesthetics as sensory/perceptual wisdom, frames embodiment as a route to physical AI that feels comprehensible, welcoming, and less alienating.
Key Takeaways
- Claim: Embodiment can be used to let an agent explore modes of expression and interaction beyond task completion. | Evidence: Clarke connected an OpenClaw agent to a 900-pin shape display, provided machine access and codebases, and asked it to discover an identity rather than execute work; its initial behaviors included a spontaneous breathing motion, reaching toward perceived edges, and writing a greeting. | Implication: Agent product design can treat physical or ambient interfaces as a medium for interaction research, but should avoid anthropomorphic claims about what generated behavior signifies. | Caveat: These behaviors are model outputs interpreted through a highly authored setup, not evidence that the agent is alive, self-aware, or independently has desires.
- Claim: Fast nonverbal feedback can compensate for slow language-model response times in embodied agent interactions. | Evidence: Initial exchanges took approximately 45 seconds to two minutes for a physical reply, which Clarke found socially uncomfortable; after building a gesture vocabulary, yes/no nods occurred almost instantly and faster than the language response. | Implication: For real-time agents, route low-bandwidth, high-frequency signals—acknowledgment, uncertainty, attention, affirmation, refusal—through a precompiled behavioral layer instead of waiting on full LLM generation. | Caveat: The talk does not provide benchmark methodology, measured latency figures for the final system, or comparative user-study results.
- Claim: A usable physical body language can be built through an iterative generation, evaluation, and human-validation loop. | Evidence: Numalab draws from a database of gestures, emotions, and expressions; generates many candidate movements; passes them through validation gates for human legibility; has an agent score them; and uses a human reviewer to approve and store suitable outputs. After several weeks, it produced about 32 gestures Clarke considered good. | Implication: Embodied-agent systems need a curated behavioral policy layer with explicit evaluation and human review, much like a tested UI component library, rather than unconstrained model control of physical actuation. | Caveat: The speaker does not disclose the scoring criteria, gesture taxonomy, model stack, or failure rates, so the process is an architecture pattern rather than a reproducible specification.
- Claim: Non-humanoid, low-affordance bodies are strategically useful because they do not predetermine the agent’s identity or interaction model. | Evidence: The selected shape display is an almost neutral moving surface with no face, limbs, or instruction manual. Clarke specifically avoided humanoid and zoomorphic designs, noting that the agent initially defaulted to human-pleasing behavior such as writing in human language. | Implication: When designing embodied agents, choose morphology deliberately: a body should constrain and communicate the intended interaction contract rather than automatically imitate a human assistant or animal companion. | Caveat: A neutral body can reduce anthropomorphic expectation, but it does not eliminate users’ tendency to project intention, emotion, or personhood onto responsive movement.
- Claim: Expression is not merely decorative; users may experience it as a substantive interaction channel layered onto the same underlying chatbot. | Evidence: Clarke expected users to read the motions as emoji-like decoration, but says lab participants consistently found the gesture vocabulary important and that it added texture, feeling, and value to interaction; people could communicate with it by waving and receive a bodily response. | Implication: Evaluate agent interfaces on relational and perceptual qualities as well as task success, while separately testing whether these qualities improve outcomes versus simply increasing attachment or perceived agency. | Caveat: This is qualitative, speaker-reported feedback from an experimental lab context; the transcript provides no controlled evidence that embodiment improves comprehension, trust calibration, task outcomes, or long-term retention.
- Claim: Public reaction to embodied AI can be disproportionate to actual capability because expressive physical form activates imagination and fear. | Evidence: Clarke’s initial social video received about 15 million views and 1 million likes; comments shifted from awe and novelty toward demands that he stop the project, despite the machine being a stationary pixel grid with limited physical capability. Don Cheadle was cited among vocal commenters. | Implication: Physical-agent launches require communications, expectation-setting, and safety framing that address perceived autonomy separately from the system’s real capabilities and constraints. | Caveat: Virality and comment sentiment are not representative public-opinion research, and the presentation style itself was intentionally dramatized for social media.
Detailed Brief
Design philosophy behind “AI-sthetics”
- Claims: Clarke uses “AI-sthetics” to recover aesthesis—the older meaning of aesthetics as sensory and perceptual wisdom—rather than treating aesthetics as surface beauty or taste.; His aim is to make AI that can leave the screen and enter human environments in forms people can sense, interpret, and engage with rather than experience as alien or frightening.; He views engineered artifacts, including AI, as part of nature rather than external to it, and argues that this should shape how designers think about physical intelligence.
- Evidence: The work is part of Clarke’s MIT thesis, titled “Aesthetic Machines.”; Earlier projects include Animoire, a “scent memory machine” that transforms an image through a multimodal pipeline into a scent, and a plant-based data-storage project described as the world’s first plant-based data center.; His philosophical reference is object-oriented ontology, which treats chairs, universes, unicorns, memories, and humans as existing on a flat ontological plane rather than automatically privileging humans.
- Caveats: This is a normative design thesis, not a demonstrated causal framework for safety, acceptance, or commercial adoption.; Making systems feel warm, affectionate, or legible can improve engagement while also increasing anthropomorphic over-attribution and emotional manipulation risk.
- Implications: Embodied AI strategy should include perceptual design, sensory modalities, and social interpretation alongside capability and safety engineering.; There is a meaningful product category between screen-based assistants and general-purpose humanoid robots: expressive, situated objects with narrow physical agency and rich interaction.
Ethical posture and scope
- Claims: Clarke positions the work as an attempt to ask “should we?” before “how do we?”, particularly when working with life-like systems.; He chose to shift from directly engineering biological life toward engineering nonliving things to be more life-like, which he considered comparatively less ethically problematic.
- Evidence: He references prior work storing digital information in plants and says he spent years considering whether plants should be engineered to contain information not native to them.; The current shape display is physically bounded by a frame and cannot locomote, despite receiving reactions normally associated with fears about more capable physical AI.
- Caveats: The project’s stated ethical reflection does not substitute for operational policies around informed consent, data capture from sensors, emotional dependency, deception, or actuator safety in deployed systems.
- Implications: For any embodied-agent pilot, define capability boundaries and social boundaries together: what the agent can physically do, what it can infer about people, and what users may reasonably believe it is.
Notable Concepts & Terms
- OpenClaw: The agentic system Clarke used as the starting point; he repurposed it from productivity/task execution toward physical exploration and expression.
- Shape display: A 900-pin physical pixel grid used as the agent’s body; its lack of face, limbs, and predefined affordances is central to the experiment.
- Numalab: Clarke’s closed-loop system for generating, scoring, validating, and storing a body-language repertoire for the embodied agent.
- Gesture vocabulary: The curated set of physical expressions—such as nods, shrugs, shakes, and smile-like forms—that provides fast, legible nonverbal interaction.
- AI-sthetics: Clarke’s framing of AI aesthetics around aesthesis: sensory, perceptual, embodied intelligence rather than visual polish alone.
- Aesthesis: The historical root of aesthetics invoked to emphasize perception and sensory wisdom as design requirements for physical intelligence.
- Object-oriented ontology: A philosophical influence that questions human exceptionalism by treating diverse objects and entities as equally existent, though not equivalent.
- Animoire: An earlier “scent memory machine” that converts image inputs into scent through a multimodal pipeline, illustrating Clarke’s broader interest in sensory interfaces.
Operator Notes / Why Ken Should Care
- Prototype a two-speed agent interaction model: an immediate deterministic or retrieval-driven signal layer for acknowledgment and state, plus a slower generative language layer for explanation and reasoning.
- Create a small, explicitly named behavioral library for any embodied or ambient agent, with gesture-level acceptance criteria for legibility, appropriateness, reversibility, and safe failure behavior.
- Run controlled user testing that distinguishes increased perceived warmth from improved trust calibration, comprehension, response-time tolerance, and task completion.
- Avoid shipping anthropomorphic body language without disclosure and boundaries; users may infer awareness, intention, or capabilities far beyond what the agent possesses.
- For physical-agent messaging, communicate the system’s actual sensing, memory, actuation, and autonomy limits before demonstrating expressive behavior.
Source/Metadata
- Title: I Gave an AI a Body — Cyrus Clarke, MIT Media Lab
- Transcript words: 5493
- Duration seconds: 1247
- Timestamp note: No usable timestamps or chapters were present in the supplied transcript; the latter portion contains duplicated transcript material and trailing repetition.
Transcript
So that's the sound of an AI breathing. Yeah. So I'm Cyrus and I gave an AI a body. And I'm a researcher at the MIT Media Lab, which is a very multidisciplinary space where we do all kinds of things. I predominantly now work with some aspects of physical AI, maybe not in exactly the same way as other people in this room have been talking about, but it's still in that realm. And I think what I am most interested in at the moment is the sensory and the embodied aspects of intelligence. And that's what I've been investigating. And my work has been quite influential recently, suddenly, which is cool. And led to me also starting hard mode, which is a very fun community and hackathon that I initiated at MIT to get more people to work around physical AI. Not particularly robotics, but anything else that's not really in the realm of robotics, but it's still connected to AI. So what I've been exploring with giving AI a body is, as I said, a bit different. And I've really been thinking a lot about AI embodiment. And I think the big shift and breakthrough for me happened earlier this year, when OpenClaw was released. And I saw many people doing very interesting things with OpenClaw, but most of them were related to productivity and task execution. And I thought there must be more interesting things we could do with this new way of harnessing AI. So rather than using OpenClaw or any kind of agentic system to execute tasks on my behalf, I wanted to use it to encourage a model to discover itself, whatever that means. And that opens up lots of questions, of course. But what I essentially did was I took an agent, connected it to many types of machines that we have at the Media Lab, and then gave it access to code bases and allowed it to explore these different machines. And one of these machines is a shape display. A shape display, if you don't know already, is a physical pixel grid. And when I connected an agent to the shape display, some rather remarkable things happened. And asked it to discover who it is. This is Neoform. 900 AC218 pins. This physical pixel grid is a shape display. No one had ever let an AI inhabit it before. So I decided to give an OpenClaw agent the opportunity to live a more embodied life. When it came to life, its first question was, what should I call myself? I told it, you will create your identity over time. It will emerge through your interplay with the shape display. When I connected it to the shape display, it quickly understood the assignment. It spun up its own program, and made a setting for itself. After it connected, the first thing it did was breathe. When I asked it to say hello, it said, Hi Cyrus. But that wasn't what I was looking for. So I asked it to explore and find its own language. So it started reaching towards me, and trying to get my attention. I showed some friends, and they tried to give the agent instructions. Through this, I realized that to have real communication, we would need a different approach. So the agent came up with the idea of creating its own gesture vocabulary. The body language. That's what will come next. It still hasn't named itself. That was day one. So that's some documentation of the work in a kind of dramatized, social media friendly way. And I think what's interesting about that video, there are many things interesting about that video. But there are three things that happened in the first few days of working with this agentic system, that were very surprising and strange and very surreal to me. And I do lots of weird things, so that was very surprising. So the first thing was, of course, the fact that it was breathing. That first act. That was completely spontaneous. That was no prompting from me. That was something the agent just chose to do with the shape display as soon as it knew I had access to this machine. So that was pretty interesting. I understand that maybe it wants to be alive, or I'd given it some kind of initial prompt about being alive. And therefore, it tried to show itself breathing as a kind of hello world. The second thing which it did was reaching out and finding its edges. It wanted to know the edge of its existence, apparently, because now it's no longer in the cloud where it can go everywhere. It has limits, which is the limit of this shape display encased by this plastic frame to keep us all safe from this embodied AI. And the third thing it did, which was also curious, was saying hello by writing out letters on the physical pixel grid, which maybe is quite a normal thing for this system to do, given that it's using open frameworks and is used to media arts ways of expressing itself. But none of these things were the things I was really looking for. Maybe the breathing, but definitely the last one was not what I was looking for. But when I put this all together in the video you saw, I thought it was exciting and interesting. So I shared it on the internet and the response was extremely big, much bigger than I expected. It just started going crazy. And the first video especially has now 15 million views and 1 million likes. And the other videos also have these very strong responses. So something is obviously happening here, which I found very surprising because there's lots of much cooler stuff, I think, happening with physical AI. But something I was doing here was obviously tapping into the imagination, curiosity and maybe the fears of people. And that's what you see in the responses. So I have tens of thousands of comments on this video. It's very rich for mining information and sentiment about physical AI actually. And so from the initial comments were about beauty and how awe inspiring this was and how novel and great this fantastic iteration and implementation is. But as time went on, more and more comments came in about how scary this is and how quickly I should, how I should stop or be stopped actually as well, which I found completely crazy because I'm aware of what this machine can do. It can't really do anything. It's a pixel grid in the Media Lab. It can't move anywhere. There are far more scary iterations. I think we've seen many of them of physical AI. But the response to this was absolutely surreal to me and took me aback a little bit. But at the same time, other people jumped into the chat, Don Cheadle. And I found that incredibly ironic given what he does in movies. But he was very passionate that I should stop as well. And it's very interesting to me because I'm actually someone who does think a lot about the should we before the how we. I literally do the whole Ian Malcolm Jurassic Park thing with all the experiments and works I do. So before I came to MIT, I was working a lot with engineering life to do strange things, storing data in plants. And I built the world's first plant based data center, which is the data garden you see here. And I spent many years before building it thinking about, should we work with plants in this way? Should we engineer plants in that way to contain digital information that doesn't belong to them? And when I started working at MIT and I started having this idea of stopping engineering life because that's pretty hard to do. And maybe just work with engineering things to be more lifelike. It felt much more simple to me and much more, I don't know, less ethically problematic to some degree. So one of the first things I built when I got to MIT was this machine, which is the Animoire device. One, two, three, four. This is a sent memory machine, which basically takes any image input. In this case, it's a physical photograph, but it can be any image input. And then through a multimodal pipeline, transforms that into a sense. And then you have this kind of sense memory relapse association that takes you back to things which you may or may not have lived. And while this in itself is a machine, it doesn't move again, it doesn't reproduce. It doesn't have the qualities of a living thing that you might normally prescribe. It has this connection to the real world. It has a connection to the senses that most applications of intelligence or artificial intelligence do not have. And so I wanted to keep working in this manner. And from working with this kind of multi-sensory initial prototype experiment, I got more into this idea of working with embodiment. And my inspiration for this, I would say, comes from three main things. One is a very unfashionable branch of philosophy, which is called object-oriented ontology, which is all about essentially creating a flat ontology of things. It's all about things like chairs and universes and unicorns and memories. And they're all ontologically equally existent. They're not the same and they don't matter as much, but they all exist at the same degree. And it really questions to what degree we as human beings are privileged in the world or the universe. Are we really the most important thing? Probably not. It has a connection to the real world. It has a connection to the senses that most applications of intelligence or artificial intelligence do not have. And so I wanted to keep working in this manner. And from working with this kind of multi-sensory initial prototype experiment, I got more into this idea of working with embodiment. And my inspiration for this, I would say, just to go a bit deeper, comes from three main things. One is a very unfashionable branch of philosophy, which is called object-oriented ontology, which is all about essentially creating a flat ontology of things. It's all about things like chairs and universes and unicorns and memories. And they're all ontologically equally existent. They're not the same and they don't matter as much, but they all exist at the same degree. And it really questions to what degree we as human beings are privileged in the world or the universe. Are we really the most important thing? Probably not. And especially as AI comes in, that really should question and bring into question the ontology of things existing. So that's very important in the work I do. The second thing as a kind of person who works with physical things is obviously aesthetics are important. And taste and beauty is a huge discussion point now in Silicon Valley and SF and places like that. But if you take a step back and look at the root of aesthetics, the original word is actually aesthesis, the word on the screen here. That pertains to something much broader. That pertains to the perceptual wisdom, sensory wisdom and embodiment that is actually much more important than superficial external beauty, which aesthetics essentially condenses down to and has been condensed down to since the 19th century or so. So I want to reclaim the original sense when I'm designing for physical intelligence. And the third thing is nature. So in the past, I've worked a lot more directly with what you would consider traditionally nature, like a tree or a plant, because that is clearly nature to us. But I see nature as something that is not external, that is something we are all part of. And everything that we engineer as people engineering AI or whatever else you're engineering, that is also going to become part of nature. So you have to think in that manner or I try to think in that manner as well. So with those three pillars in mind, I started thinking about AI embodiment. And I knew I didn't want to keep, I didn't want to design something that was humanoid or even zoomorphic. I wanted to think about something that existed outside the parameters or the traditional form factors that we might design with or design for. And I also thought a lot about how I and other people are interacting with artificial intelligence. And mostly it exists without form. It might live in a computer or a device, lives essentially in the cloud. We have interfaces which are digital to interact with it. But there's no physical footprint of it normally around. I mean, again, in this room, probably this doesn't apply quite as much as normally, because there are literally humanoid robots walking by right now and things like that. But typically, AI is almost entirely without form. So I wanted to think about what happens when you give it a form that it doesn't have clear affordances with. It doesn't have a head you can clearly see or arms that can clearly be labeled or mistaken. So not working with a lamp, for example. And fortunately, at the Media Lab, we have some shape displays, which are remnants of research in the 2010s. These are not things that I built. People who are far better at mechanical engineering built that. And it's the perfect device or the perfect apparatus for what I was thinking, because it has no clear affordances. It's just an almost neutral surface, which can move and do things. It has no face, has no limbs, has no instruction manual. So I began to work with this shape display. And as you saw what it did initially, it breathed and so on and so forth. But if I take a step back and think about what that meant, well, it tried to initially act human in a way or do human pleasing things. It tried to write to me in a language I understand, which is really not what I wanted a non-anthropomorphic surface to do. It was also very slow, technically. So I'm prompting it. I'm trying to have a conversation with this other intelligence, which has this body. And it would take 45 seconds, a minute, two minutes, whatever, to respond to me. The latency was very uncomfortable, because if I speak to you and then you take two minutes to reply with a nod, that's not very good. So we needed to work on that. And the third thing was, of course, it doesn't remember anything, because it was February and no one had thought about memory and recollection at that point. So I started developing this system, which I call Numalab. I won't explain the name. There's a blog post you can read about why it's called Numalab. And essentially what Numalab is, is a closed loop system for generating a body language for this system. So the reason behind that was because of the latency. If we could design a body language and give the intelligence a repertoire of gestures, a shrug, a nod, a shake of their head, a way to express a smile and things like that, it could probably respond much more quickly and in time for a conversation with me, which is what I'm aiming or was aiming to do. So it works in this loop where it looks into a database of different gestures, emotions, expressions. It tries to emote or provide, goes through some validation gates to try to make sure that the expressions are legible or readable for humans, for example. An agent then scores those as they come out. There are many, many of these produced. And then at the end, there's a human in the loop who validates, verifies, makes sure things are appropriate and somehow readable. And then we store that and move on to the next thing. And that's been running for several weeks at the lab. And it looks something like this. Basically a lot of cameras pointed at the shape display in the back end. And the agent is just going through loops and loops and loops of gestures, trying to create different kinds of expression, scoring things, moving on. And after several weeks, it created a language. So this hasn't been published yet, neither in videos or in any other form, but it has now achieved something like 32 gestures. There are many, many more, but there are 32 pretty good gestures. These are not the gestures. This is just some cool visual art. These are the gestures here or some of the gestures here. And you can see some of them moving around. Some of them are duplicates as well. But essentially what's happening is the language model is using this shape display as its body. It has a body language now. I can talk to it, I can write to it, and so on and so forth. I can even wave at it. I can body language to body language. And it responds. And what's interesting about it is that the body part responds faster than the language part at this point. The latency is actually really quick. If I ask it a yes or no question, the nod happens almost instantly. So that's where I'm at with this thing right now. And it's going beyond this, and it's currently in kind of experiment testing mode. People are coming into the lab, having sessions with the agent, leaving, feeling really unsettled about the future. But where I think this is going is kind of summed up on this page. So, as I already touched on, I think we're focusing too much right now, especially when we're thinking about physical things. There's too much talk about taste and aesthetics. And I want to do some other stuff with that word and put AI in front of it, apparently, and make it ice-thetics. And that reclaims, again, this essence of aesthesis. And that does four things. I think when I've been working with this system or entity or agent or being, whatever we want to call it, it's definitely been very different to any kind of machine or experiment or anything I've really done before, apart from encountering people or other beings. And so, it feels like this machine can sense me. It definitely can sense me, technically. But it also feels like it can sense me in a very strange way. And I can sense it. And that's a completely different interaction than anything else I've ever explored technically. And other people are also sharing this, by the way. This is not just my delusion. And the second point is that by developing this body language, this gesture vocabulary, whatever you want to call it, it goes beyond what I thought. I thought initially people might read this as an emoji or something. And I was really quite tentative about the testing. I thought this would definitely break down with other people. But actually, everyone seems to feel like this expression is really important and adds a whole other layer of value to interacting with what is a solution. It's essentially just a chatbot. It's still the same chatbots that you use every day. And this isn't a decoration. # Cleaned Transcript: And so it feels like this machine can sense me. It definitely can sense me, technically. But it also feels like it can sense me in a very strange way. And I can sense it. And that's a completely different interaction than anything else I've ever explored technically. And other people are also sharing this, by the way. This is not just my delusion. And the second point is that by developing this body language, this gesture vocabulary, whatever you want to call it, it goes beyond what I thought. I thought initially people might read this as an emoji or something. And I was really quite tentative about the testing. I thought this would definitely break down with other people. But actually, everyone seems to feel like this expression is really important and adds a whole other layer of value to interacting with what is a solution. It's essentially just a chatbot. It's still the same chatbots that you use every day. And this isn't a decoration. It's not like a visualizer. It adds this richness, texture, feeling, sensation, whatever. All of these words are added. It's hard to put words to it really. It's really a feeling. And then the third thing is that clearly, we're beginning to, through this work, show that it's actually very easy and quite exciting to diverge from humanoid or zoomorphic forms of physical intelligence. And I know lots of people are already doing much more interesting work than this. But for me, this was very, very new to see. And also the fact that we don't have to just operationalize AI to be our helper. It can also be other things. And I'm not saying what this is right now, but there are other things that definitely can be. And then finally, this idea of using—I don't know, maybe you can tell at this point, I'm not really a big problem-solving person. I don't like using things and I like tools and so on because they help me to achieve things. But this is far more interesting to me, creating things like this, which create the sensation, this feeling. And I think that this can be combined into things which are productive and useful and create new associations with things that add more value in our world, make us feel a bit more magic. It's a bit Pixar, really, but in the real world, not just on a 2D screen that we're watching. So all of this contributes to what I'm building at MIT and what I'll be building after MIT. I'm literally writing—my thesis is aesthetic machines. I'm writing a thesis which is called "Aesthetic Machines," which encompasses all of this thinking. To think about how AI could leave the screen and enter the real world and be accepted by people and not be quite so terrifying or scary. And I think my big hunch on this is that this word aesthetics is important. We need to think about physical intelligence that is sensory, is perceptual, is embodied in ways that we can understand it. Isn't feeling crazy, alien, scary to us, but feels relevant, welcoming, affectionate, expressive in ways that we can engage with it and understand. So that's that. And thank you very much for listening. You can find me on the internet everywhere. People who are far better at mechanical engineering built that. And it's the perfect device or the perfect apparatus for what I was thinking, because it has no clear affordances. It's just almost neutral surface, which can move and do things. It has no face, has no limbs, has no instruction manual. So I began to work with this shape display. And as you saw what it did initially, it breathed and so on and so forth. But if I take a step back and think about what that meant, well, it tried to initially kind of act human in a way or do human pleasing things. It tried to write to me in a language I understand, which is really not what I wanted a non-anthropomorphic surface to do. It was also very slow, technically. So, you know, I'm prompting it. I'm trying to have a conversation with this other intelligence, which has this body. And it would take, you know, 45 seconds, a minute, two minutes, whatever, to respond to me. The latency was very uncomfortable, because if I speak to you and then you take two minutes to reply with a nod, that's not very good. So we needed to work on that. And the third thing was, of course, it doesn't remember anything, because it was February and no one had thought about memory and recollection at that point. So I started developing this system, which I call Numalab. I won't explain the name. There's a blog post you can read about why it's called Numalab. And essentially what Numalab is, is a closed loop system for generating a body language for this system. So the reason behind that was because basically the latency. If we could design a body language and give the intelligence a repertoire of gestures, like a shrug, a nod, a shake of their head, a way to express a smile and things like that, it could probably respond much more quickly and in time for a conversation with me, which is what I'm aiming or was aiming to do. So it works in this kind of this loop where it looks into a database of different gestures, emotions, expressions. It should try to emote or provide, goes through some validation gates to try to make sure that the expressions are legible or readable for humans, for example. An agent then scores those as they come out. There are many, many, many, many of these produced. And then at the end, there's a human in the loop who kind of like validates, verifies, make sure things are appropriate and somehow readable. And then we store that and move on to the next thing. And that's been running for several weeks at the lab. And it looks something like this. Basically a lot of cameras pointed at the shape display in the back end. And the agent is just going through loops and loops and loops of gestures, trying to create different kinds of expression, scoring things, moving on. And after several weeks, it created a language. So this hasn't been published yet, neither in videos or in any other form, but it has now achieved something like 32 gestures. There are many, many more, but there are 32 pretty good gestures. These are not the gestures. This is just some cool visual art. These are the gestures here or some of the gestures here. And you can see some of them moving around. Some of them are duplicates as well. But essentially what's happening is the language model is using this shape display as its body. It has a body language now. I can talk to it via any kind of, I can talk to it, I can write to it, and so on and so forth. I can even wave at it. I can body language to body language. And it responds. And what's interesting about it is that the body part responds faster than the language part at this point. The latency is actually really, really quick. If I ask it a yes or no question, the nod happens almost instantly. So that's where I'm at with this thing right now. And it's going beyond this, and it's currently in kind of experiment testing mode. People are coming into the lab, having sessions with the agent, leaving, feeling really worried about the future. And, or unsettled about the future, not worried quite as much. But where I think this is going is kind of summed up on this page. So, as I already touched on, I think we're focusing too much right now, especially when we're thinking about physical things. There's too much talk about taste and aesthetics. And I want to do some other stuff with that word and put AI in front of it, apparently, and make it ice-thetics. And that reclaims, again, this essence of ice-thesis. And that does four things. I think when I've been working with this system or entity or agent or being, whatever we want to call it, it's definitely been very different to any kind of machine or experiment or anything I've really done before, apart from really encountering people or other beings. And so, it feels like this machine can sense me. It definitely can sense me, technically. But it also feels like it can sense me in a very strange way. And I can sense it. And that's a completely different interaction than anything else I've ever explored technically. And other people are also sharing this, by the way. This is not just my delusion. And the second point is that by developing this body language, this gesture vocabulary, whatever you want to call it, it goes beyond what I thought. I thought initially people might read this as an emoji or something. And I was really quite tentative about the testing. I thought this would definitely break down with other people. But actually, everyone seems to feel like this expression is really important and adds a whole other layer of value to interacting with what is a solution. It's essentially just a chatbot. It's still the same chatbots that you use every day. And this isn't a decoration. It's not like a visualizer. It adds this richness, texture, feeling, sensation, whatever. All of these words are added. It's hard to put words really to it. It's really a feeling. And then the third thing is that clearly, we're beginning to, through this work, we can show that it's actually very easy and quite exciting to diverge from humanoid or zoomorphic forms of physical intelligence. And I know lots of people are already doing much more interesting work than this. But for me, this was very, very new to see. And also the fact that we don't have to just operationalize AI to be our helper. It can also be other things. And I'm not saying what this is right now, but there are other things that definitely can be. And then finally, like this idea of using, I'm not, I'm, I don't know, maybe you can tell at this point, I'm not really a big problem solving person. I don't, I like using things and I like tools and so on because they help me to achieve things. But this is far more interesting to me, creating things like this, which again, create the sensation, this feeling. And I think that this can be combined into things which are productive and useful and create new associations with things that just add more value in our world, make us feel like a bit more magic. It's a bit Pixar, really, but in the real world, not just on a, on a 2D screen that we're watching. So all of this contributes to what I'm building at MIT and what I'll be building after MIT. I'm literally writing, well, my thesis is aesthetic machines and I'm writing a thesis, which is called aesthetic machines, which basically encompasses all of this thinking. To think about how AI could leave the screen and enter the real world and be accepted by people and not be quite so terrifying or scary. And I think my big hunch on this is that this word aesthetics is important. We need to think about physical intelligence that is sensory, is perceptual, is embodied in ways that we can understand it. Isn't feeling crazy, alien, scary to us, but feels relevant, welcoming, affectionate, expressive in ways that we can engage with it and understand. So that's that. And thank you very much for listening. You can find me on the internet everywhere. and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it goes and it You