Open Reader

Reachy Mini: the $300 open source robot you can actually hack — Andres Marafioti, Hugging Face

completed 21:16 May 29, 2026 Watch on YouTube

Current Status

completed

Video ID

0jeZfjJMfmo

RAG / Chat

Enabled
Reachy Mini: the $300 open source robot you can actually hack — Andres Marafioti, Hugging Face
Description

Qwen3-TTS shipped at 0.8x real time: one second of audio took 1.2 seconds to generate. Andres Marafioti from Hugging Face spent two weeks fixing it. The culprits were no streaming, 500 autoregressive steps per audio packet with a CPU GPU round trip on each, and a dynamic KV cache that blocked compilation. Static KV cache plus CUDA graph captures brought it to 5.8x real time with time to first audio under 200 milliseconds. The platform is Reachy Mini, a $300 open source robot Hugging Face has shipped to 7,500 people. It arrives unassembled. Talking to it is their most used app by far. The voice stack runs Parakeet transcription every 150 milliseconds with partial results feeding back to the robot mid-sentence, Qwen 3.5 27B for the LLM, and this optimized TTS. At that speed, infrastructure round trips match model latency, so the load balancer separates LLM endpoints from conversation nodes to handle the difference in how much different users actually talk.

Summary

Generated by claude-sonnet-4-5

30-second take

Hugging Face's Andres Marafioti presents Reachy Mini—a $300-450 hackable robot designed to democratize voice-AI-robotics integration before proprietary companies lock it down. His thesis: current robots (humanoids, home bots) are too expensive ($50k+), uncreative (mimicking humans instead of exploring new forms), and inaccessible for prototyping. Reachy Mini ships unassembled, runs fully open-source voice agents (speech-to-speech pipelines, QWEN 3.5 LLM, custom-optimized QWEN3 TTS), and targets students/hackers/researchers. The talk is a manifesto for open-source robotics UX experimentation plus a technical deep-dive on how Hugging Face scaled TTS inference from 0.8x real-time to 5.8x through streaming, static KV caches, and GPU graph captures. 7,500 units shipped; the voice conversation app is their most-used feature.

Key takes

  • Price barrier is deliberate exclusion: $50k+ humanoids and six-figure Waymo cars prevent high schools, hobbyists, and indie devs from experimenting with robot interaction design. Marafioti argues this consolidates UX control in corporate hands, like pre-PC computing.
  • Form follows hype, not function: Current robots mimic biology (humanoids, cars) for familiarity, not engineering optimality. A humanoid could be faster/more stable as a spider-like form, but companies choose human shapes for consumer comfort—a design mistake that limits creative interfaces.
  • Voice AI is ready; robotics UX isn't: Open-source voice stacks (speech-to-speech pipelines, 80M-parameter Kokoro TTS, Parakeet STT) are mature and fast, but almost no one is exploring how humans should talk to robots. Marafioti sees a window to shape this before it's proprietary.
  • Reachy Mini's design forces learning: Ships unassembled so users must build it, creating immediate repair knowledge and psychological ownership. Users 3D-print custom parts (Halloween pumpkins, pet-reactive behaviors), proving the hacker ethos works.
  • Infrastructure latency > model latency: QWEN3 TTS went from sub-real-time (0.8x, unusable for voice agents) to 5.8x real-time via static KV caching, GPU graph captures, and streaming. But total client-perceived latency includes network/coordination overhead—infrastructure time equals model time in their stack.
  • 7,500-unit fleet stress-tests open inference: Serving voice agents to thousands of robots required autoscaling LLM endpoints (QWEN 3.5 27B) separately from conversation nodes, because concurrent talkers create wildly uneven GPU loads. This is a real production case study.

Useful details

  • Two SKUs: $450 (Raspberry Pi + battery) vs. $300 (no Pi/battery, sold in bulk to schools).
  • Voice pipeline: Parakeet STT transcribes every 150ms, sends partial transcriptions to robot for reactive behaviors, full transcripts to QWEN 3.5 27B LLM (tool-calling for movement/camera), output to QWEN3 TTS. Echo cancellation + face tracking built in.
  • QWEN3 TTS optimization journey: Original model did 500 autoregressive steps per audio packet with CPU-GPU coordination overhead. Fixed by: (1) streaming (don't wait for full 10s output), (2) static KV cache (uses more VRAM upfront but enables compilation), (3) graph captures. Result: first-token latency under 200ms, 4x real-time generation.
  • Deployment architecture: Load balancer scales compute nodes based on connected robots; LLM endpoints scale separately to handle uneven conversation load (8 active talkers vs. 8 idle users on same node).
  • Codex-assisted development: Marafioti created a movement demo by giving Codex the repo and a prompt; it generated working code. Claims non-coders can start building today.
  • Existing ecosystem: Fourth open-source robot from this team. SO100/SO101 robot arms can attach; "Kiwi" mobile base (three-wheeled vase) designed to carry Reachy Mini (not yet seen in wild).
  • Anecdote: Zurich robotics club hosted humanoid boxing matches—"incredible and a little disturbing."

Caveats / counterpoints

  • No mention of failure modes: Transcript gives zero detail on what breaks, quality issues, or where the open-source stack struggles in practice beyond TTS latency.
  • $300 is still a barrier: Dismisses $50k humanoids but doesn't acknowledge $300 × class-size is expensive for underfunded schools; "bulk sales" pricing not disclosed.
  • Raspberry Pi compute limits: Acknowledges Pi constraints but doesn't specify what can't run locally or where cloud dependency becomes mandatory.
  • Aesthetic bias: Calls it "cute" and "friendly" vs. humanoids, but admits this is personal opinion; unclear if users actually find non-humanoid forms less intimidating or just novel.
  • Sustainability/repairability unproven: Ships unassembled and claims full repairability, but no data on part longevity, supply chain for replacements, or whether users actually repair vs. replace.
  • Codex hype: "Non-coders can build today" via AI-generated code is optimistic; no evidence users without programming knowledge have shipped meaningful apps.
  • Missing comparison: No benchmarks vs. other affordable robotics platforms (e.g., $200 Turtlebot kits, Raspberry Pi camera modules + servos). Why is Reachy Mini better than DIY?

Ken relevance

Medium-high for AI inference ops and open-source GTM strategy; moderate for agent systems.

  • Inference optimization playbook: The QWEN3 TTS case study (static KV cache, graph captures, separate LLM scaling) is directly applicable to your agent inference architecture. If you're serving real-time voice or vision models, this is a template for cutting latency.
  • Open-source distribution model: Hugging Face is subsidizing inference for 7,500 robots to create ecosystem lock-in and showcase their stack. Relevant if you're considering freemium/open-core strategies—how much compute can you give away to build moat?
  • Productizing research code: Marafioti spent two weeks making QWEN's "research implementation" production-ready. If you're wrapping open models into commercial products, expect this tax.
  • Voice agent stack: Their pipeline (VAD → STT → LLM → TTS with tool-calling) is a concrete reference architecture. You could clone this for customer service bots, sales agents, etc.
  • Low direct content opportunity: This is a niche hardware play; unlikely to generate YouTube/LinkedIn content unless you're building robotics demos or reviewing affordable platforms.
  • No immediate investing angle: Hugging Face is private; robotics hardware startups at this price point are pre-revenue community projects, not VC-scale.

Watch verdict

Skim. The QWEN3 TTS optimization segment (12–18 min mark) is valuable if you're doing real-time inference. The rest is a well-argued manifesto but light on novel execution details—you've absorbed the key ideas from this summary. Watch the demo video if you want to see latency/quality firsthand, but the transcript captures the technical substance.

Transcript

3537 words en Processed in 150.7s

And are you hearing me well? No, okay. Yeah, now with the speakers. Good. Hey, how are you, everyone? My name is Andy Marafiotti. I'm going to talk today about this little robot called Ritchie Mini. I'm going to try to explain to you why we developed it, why we think it's important, and what we are trying to do with it. And about me, I lead multimodal research at Hugging Face, which I don't have time to explain very much, but I hope you know Hugging Face. Good. So first, very quickly, the state of robotics. We are actually making strides in robotics. I think for day-to-day life, it's not very clear, but robots are just coming and they are coming at neck-breaking speed. We have really good humanoids nowadays. A couple of weekends ago, I was at the robotics club in Zurich, Switzerland, and we had a little boxing match between humanoid robots. It was incredible and a little bit disturbing. There are also several companies today trying to sell you robots to put in your homes to do things like water your plants and scroll TikTok endlessly. And of course, we have self-driving cars. I think Waymo is now in London, I heard. I haven't seen them, but pretty cool. Now, something here that I always see with these strides that we are making with robotics is A, these things are really expensive. They are all at least in mid five-figure range, so 50k up. And they don't seem to really be coming down. They seem to be targeting that price range for now. That said, the humanoids. The Waymo, I think they are like six figures mid. And they also, they look like something that you know. So it's not very much about being creative. It's very much about trying to imitate reality. And this is not biology. This is really hardware. And we are really constraining the robots to do this. If you take a humanoid robot, it could look like a spider and just move around way faster, be more agile, be more stable. But we are actually making it be a humanoid such that we look at it and we think, oh yeah, robot, human, it's the same sort. I understand what it can do. That to me is a mistake. They also don't look very friendly in general. So I see a few problems with the current state of robotics. All of these robots, which are making strides and they're becoming really good, they are still too expensive to prototype. I don't see any high school anywhere in the world ordering 10 of these 50k humanoid robots to let their students play with. They're very complex to adapt. They're really targeting companies. And you cannot really connect with them. Now, as you saw today already, we can actually build pretty good voice agents. The state of voice AI is quite advanced. There are really good commercial solutions. There's Gradium in Paris developing voice agents. There's GPT real time, which I'm sure most of you have played with. You can chat with it and it will reply and it sounds fairly natural. There's Siri, which has its problems, but it's in all iPhones. Hey, kudos. And in the open source, we also have a lot of really good models. You heard the talk from Mistral, maybe if you were here before. BoxRall is a really good open source model. We also have tiny models like Kokoro at just 80 million parameters that sound good. We have good models to understand speech. And we have good pipelines like speech to speech from Hugging Face to put all together and make your voice agents yourself. So what I see here is that voice AI is quite mature. We have the tools. We have the tools commercially. We have the tools in the open source. But still, no one is really working on how we're going to talk with these robots. And we are thinking, okay, given that robots are coming, how can we manage to put this in the hands of everyone such that the experience of the future is not dominated by one company or one group of people, but really it's made like computers were made in a bit more of a hacker fashion. And that's how we came up with Ritchie Mini. It's a little bit our response to that. We made it targeting really hackers, researchers, students, dreamers. If you have a computer, you should be able to play with it and make things. That's a little bit our target. A little bit of our vision. The idea is that it's an expressive robot that you can talk to it, but it doesn't look like a human. And that for us is very important because it already puts your mind in a different place, creative-wise, that you are going to start developing new ways of interacting with this. And it's not going to be just, okay, this is a human replacement. It's not. It's a new thing. And we want people to be able to explore that. We made it affordable. We made it easy to use. Well, we are making it easier to use. We made it repairable. Actually, we shipped these robots unassembled. So the first thing you do when you get the robot is you need to assemble it yourself. And once you're done with it, which we get super positive feedback from, you can basically repair anything that's going to break in the robot ever. You have all of the knowledge. You have all of the tools. It really just is ordering the parts and changing it. It's very hackable. It's very cute. That's maybe a personal opinion, but I think it's cute. And we are trying to give you a set of software tools for you to develop with. Now, very quickly, we sell two of these robots. They are both 450 and 300 USD. And the difference between both of them is the 450, which is this one that I have here, has a Raspberry Pi inside and it has a battery. And that's it. Basically, if you don't need a battery, you don't need a Raspberry Pi, you can order the cheaper one. We are selling those cheaper ones in bulk for high schools or universities to play with, to experiment, to hack with. And here I wanted to show you a few cool developments that we are seeing. For example, in the middle, what I mean with hackable, people can 3D print new parts for the robot. And in this case, they didn't need to really change anything. They're just taking the antennas out, putting a different type of antenna, they are putting something on top of the body, they are putting some lights, and they have suddenly a little Halloween pumpkin to play with. And on the right, something that we hadn't thought about, someone figured, oh, you can just pet the robot. And it can be fun. And it reacts. Does it burn? You can make it burn. So that's a little bit like, for us, it's more about people exploring and people developing their own intuition for how this should be and making it their own. But I was talking to you about voice AI. We think one of the main ways to talking with robots is going to be through voice. I don't see anyone going to a humanoid robot and taking out the keyboard and typing something. That's not going to be what happens. And on the right, something that we hadn't thought about, someone figured, oh, you can just pet the robot. And it can be fun. And it reacts. Does it burn? You can make it burn. So that's a little bit like, for us, it's more about people exploring and people developing their own intuition for how this should be and making it their own. But I was talking to you about voice AI. We think one of the main ways to talking with robots is going to be through voice. I don't see anyone going to a humanoid robot and taking out the keyboard and typing something. That's not going to be what happens. So we need to give people the tools to develop voice experiences. And we are doing that. We have a piece of software for you to converse with the robot. I want to show you a little demo. I'm not as brave as Sam from Mistral who did the live demo. I'm going to show you a video. Please, for the audio people, if you can maybe put it up a little bit. I did this this morning with my phone. Maybe a little bit higher. Hey Richie, how are you doing? Functioning within acceptable parameters. I'm still waiting for someone to tell me a good lobster joke to really get my circuits buzzing. How about you? I'm good. And could you take a photo and tell me a good lobster joke to really get my circuits buzzing. How about you? Can you take a photo and tell me what you see in front of you? I have a picture and see what's in front of me. Though I suspect it's just a very boring wall, which is basically the lobster equivalent of a rock. I see a guy taking a mirror selfie in what looks like a hotel room, wearing a white t-shirt with a yellow emoji on it and orange pants. Honestly, the only thing more crusted than those pants is a lobster that's been sitting on the beach too long. Okay, okay, I'll change my pants before the talk. But could you first show me an emotion? Let's do happiness. I'll show you happiness. Though I hope you're not expecting me to burst into a lobster themed dance of joy. Okay, so that quickly for the conversation. If you guys want to do live demos, I have the robot here. I leave it. You can talk to it. And what we're trying to do is we're trying to give people the tools to actually do the conversation. So everything that we are doing with the robot is open source. All of these models are open source. The agents are open source. But we understand also that people are a little bit GPU poor and they cannot run those models locally. We are also a little bit GPU poor, and we have now shipped 7500 of these robots, which means we have a pretty sizable fleet of people talking to the robot. It's our most used app by far. People just put it there and talk to it. And because you guys are a very technical audience, I wanted to show you a little bit how we are actually serving this. There are three levels to this system. In the middle, there's the speech-to-speech pipeline. That's a project that I've been maintaining with Hugging Face for the last two years. Sam from Mistral told you a little bit how it works. You have a voice activity detection system that knows if you are talking or not. That sends it to a speech-to-text system. In our case, we are using Parakeet because it's super fast. So we transcribe every 150 milliseconds, and we send back the partial transcriptions for the robot to react if it hears something interesting. And when the transcription is done, we send it to an LLM. The LLM replies, can also do tool calling for movement or to use the camera. And then we send it to the text-to-speech system, which we are using QWEN3 TTS. Now, on the higher level, we have the actual conversation app running the robot. That is taking the input and the output from the microphone and output through the speakers. It's doing the echo cancellation to not hear itself when it's talking and to hear you. It's doing the tool dispatching to move the emotions and it's using the camera. It can do face tracking to follow you around. And then the speech-to-speech pipeline, if you don't want to tweak it and use your remote, you can use the drone. We have it deployed in Hugging Face inference endpoints. We have a load balancer that determines how many compute nodes we have at each time, depending on how many robots are connected. And we have separated the LLM inference endpoints. We're actually using now QWEN3.5 27B because we managed to make it fast enough. But we scale that differently because each of the conversation nodes can have an amount of concurrent users without the latency spiking up too much. And it changes a lot if in one node you have eight users that are talking a lot and using a lot of the LLM. And in another node you have eight users that are not talking at all. So it really helps to save on resources if you have the LLM separated. And to show you a little bit of how far we are going into trying to make this work and also benefiting from the fact that you are a technical audience, I'm going to talk a little bit about QWEN3 TTS. So this is a model from QWEN that came out two or three months ago. For us it was a really great moment because TTS models hadn't been as good in the open source and as fast as QWEN3 TTS when it came out. And we were really expecting for something like that. We knew it was coming but we didn't know when it was going to come. The issue is that the model that they released actually ended up being of the quality that they showed you but not really of the speed that they showed you. The paper claimed very low latencies but the model didn't really achieve them. So I spent maybe two weeks with Codex trying to get the model trained by them to actually be fast enough for voice agents. And I wanted to tell you a little bit how I did that and what the main issues were. The first issue that I found is that the model would generate the whole output before giving it to you. So it wasn't streaming, which meant if you wanted ten seconds of audio it needed to generate the ten seconds. That can be solved with streaming. It's harder in practice than in theory. But when it works, it works and it's really cool. The next thing that I noticed was the model being an autoregressive model was doing 500 steps for each packet of audio it was generating. And for each of those 500 steps it needs to coordinate the CPU with the GPU. So it needs to send data back and forth between the CPU and the GPU. That is pretty bad. The way to solve that is to compile your model so that all of those interactions happen directly in the GPU. That couldn't happen by default because it was using a dynamic KV cache. So the KV cache was evolving depending on the size of the inputs that it was processing. We changed that for a static KV cache. We used more RAM from the get-go but that makes it faster. And then we could use graph captures to capture the whole model and to be able to accelerate generation significantly. We went from a real-time factor of being below real-time at 0.8, which means you generate one second, you take 1.2 seconds. To being 5.8, which means for one second you take 0.17 seconds. The way to solve that is to compile your model so that all of those interactions happen directly in the GPU. That couldn't happen by default because it was using a dynamic KV cache. So the KV cache was evolving depending on the size of the inputs that it was processing. We changed that for a static KV cache. We used more RAM from the get-go but that makes it faster. And then we could use good graph captures to capture the whole model and to be able to accelerate generation significantly. We went from a real-time factor of being below real-time at 0.8, which means you generate one second, you take 1.2 seconds. To being 5.8, which means for one second you take 200 milliseconds. And we also reduced, obviously, the time to force audio significantly from several seconds, depending on what you're generating, to a few milliseconds. And I wanted to show you now a quick demo, not live, recorded of this model on a space on Hugging Face. You can go and test it yourself. I cloned a few voices and I make the model say things. This happens all in real time. It does the cloning of the voices that I chose and it generates the text that I'm copy-pasting. Against the odds, the wild lobster has found a new vessel for its voice and with it, the possibility to realize its full potential. Give it a moment of sound, just a fragment, and it will give you back a voice that feels almost human. The original system works, but only offline and at sub-real-time speeds. FasterQuent TTS brings first token latency under 200 milliseconds and runs at 4x real-time. Once you cross that threshold, entirely new applications open up. That is also open source. Now you can go look for FasterQuent 3 TTS and use it. Or you can use it with our voice agent from Ritchie Mini. You can also test this demo online. I released it maybe two months ago and I think it has been used every day consistently. And something that I wanted to highlight here, there's a difference between the time to first audio of the model and what the client perceives. Because on top of the basic thing that people are telling you of the model, there's all of the infrastructure times that actually add up a lot. In this case, the infrastructure times are just as much as just the model. Because the model is quite fast, right? But still, when you're thinking of voice agents, you need to consider all these things and how they add up. And going back, what we want you to do is get this robot and make it your own and create your own interactions. We don't want you to just stop at using our voice agents. We are trying to give people the tools. Because we know that robots are coming. We know that they are going to be everywhere. And we want how we interact with those robots to be communal, to be developed by everyone that wants to develop it. And we don't want it to be guarded behind 50K robots. We want people to actually be working on this. So we try to put as many resources as we can towards that. And the last thing, you can basically code things with the robot. The application that I have today now of the robot just running here, doing some movements. I once shot it with Codex. I told it, here's the repo for the robot. Here's what I wanted to do. [SPEAKER_06] Make it. And it just did it. So it's not like you need a lot of knowledge to do these things. You can really start today, even if you're not a coder, and you can make cool things. So thank you. I'm going to be around. I really like the concept of this being a conversation starter. So if anyone wants to talk to me about this, go ahead. We have one minute, 57 seconds for questions, if anyone has them. Prince. Are you guys considering apps outside of the ones that you need to host? Like apps that run directly on the robot? Yeah. So a lot of the apps run directly on the robot. Anything that doesn't use a GPU. Of course, this one has a Raspberry Pi, so you need to work with that sort of hardware. But you can also use your own laptop as the hardware. And we are not constrained by that. Actually, we also are not constrained by languages. You can make your apps in Java or in Python or in HTML, I think. The sky's the limit. [SPEAKER_05] Okay. [SPEAKER_05] Awesome. Is there a system of plug-ins for extensibility? Like if you want to add a servo making it move around the room or something like that? Is that something that's somewhat supported or do you need to hack around it? [SPEAKER_05] So you sort of need to hack around it. [SPEAKER_05] But this is maybe our fourth open source robot. [SPEAKER_05] And we try to make them all stick with each other. [SPEAKER_05] So the SO100 and the SO101 are arms that you can plug together. [SPEAKER_05] And we have something called the Kiwi, which is a vase with three wheels that Ritchie Mini just comes into. [SPEAKER_05] So it can sort of move around. I still haven't seen it done. But we know it fits because we designed it that way. [SPEAKER_05] I'm really interested in your Quen TTS work. What other limitations did you find with your original implementation? I know the streaming stuff, the speed, any other things that you see Quen TTS is lacking? So the implementation was quite complete but generally slow. It felt a bit like it's a research implementation. If you use their API, their API is actually quite fast. So I think that was their strategy of trying to get people towards the API because the open source model isn't up there. And for me, I've been having a lot of issues come to the repo from people that are saying this thing doesn't work. And then I need to implement it just because the model can do so many different things. And I didn't think of all the corner cases. But the original implementation is quite complete. I'm at time. So I thank you a lot. As I said, I'm going to be around. So if you want to talk to me, please do. And I'm going to mute this. But please mute the laptop so I can take out the table. Thank you. Thank you. This thing doesn't work. And then I need to implement it just because the model can do so many different things. I didn't think of all the corner cases. But the original implementation is quite complete. I'm at time. So I thank you a lot. As I said, I'm going to be around. So if you want to talk to me, please do. I'm going to mute this. But please mute the laptop so I can take out the table. Thank you. Thank you. Thank you. Thank you. [SPEAKER_03] Thank you. Thank you. We'll see you next time. You can really start today, even if you're not a coder, and you can make cool things. Yeah. So thank you. I'm going to be around. I really like the concept of this being a conversation starter. So if anyone wants to talk to me about this, go ahead. We have one minute, 57 seconds for questions, if anyone has them. Prince. Are you guys considering apps outside of the ones that you need to host? Like apps that run directly on the robot? Yeah, yeah. So a lot of the apps run directly on the robot. Anything that doesn't use a GPU. Of course, this one has a Raspberry Pi, so you need to work with that sort of hardware. But you can also use your own laptop as the hardware. And we are not constrained by that. Actually, we also are not constrained by languages. You can make your apps in Java or in Python or in, I don't know, HTML, I think. The sky's the limit. Okay. Awesome. So, yes. Is there like a system of plug-ins for extensibility? Like if you want to add like a servo making it like move around the room or something like that? Is that something that's somewhat supported or do you need to hack around it? So you sort of need to hack around it. But this is maybe our fourth open source robot. And we try to make them all stuck with each other. So the SO100 and the SO101 are arms that you can plug together. And we have something called the Kiwi, which is like a vase with three wheels that Richie Mini just comes into. So it can sort of move around. I still haven't seen it done. But we know it fits because we designed it that way. Yeah. I'm really interested in your Quen TTS work. Yeah. What other limitations did you find with your original implementation? Like I know the streaming stuff, like the speed, any other things that you see Quen TTS is lacking? So the implementation was quite complete but just generally slow. It felt a little bit like it's a research implementation. If you use their API, their API is actually quite fast. So I think that was a little bit their strategy of trying to get people towards the API because the open source model isn't up there. And for me, I've been having a lot of issues come to the repo from people that are like, oh, this thing doesn't work. And then I need to implement it just because the model can do so many different things. And I didn't think of all the corner cases. But no, the original implementation is quite complete. I'm at time. So I thank you a lot. As I said, I'm going to be around. So if you want to talk to me, please do. And I'm going to mute this. But please mute the laptop so I can take out the table. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. We'll see you next time.