Armin Reviewer Yeah, let's get started. I'm Armin. I'm the co-founder and CEO of Perceptron. We'll get a little bit into what we do, but primarily what I want to talk about today is our research stance that we want to move away from distinctions between VLMs, VLAs, world models, whatever you want to call it, to something that we call embodied foundation models.
So specifically what we do at Perceptron, kind of our North Star really, is we want to be able to build physical AI foundations that give us the ability to perceive, understand, and interact with the physical world in real time. And so the North Star mission is really to bridge the physical and digital worlds, fundamentally meaning that our goal is wherever there is a device, an instrument, a robot, a camera, a sensor, we're essentially there providing intelligence to it. And specifically when we talk about this paradigm of being able to perceive, being able to reason, being able to act, we view this as a unification of traditional multimodal modeling.
So I came from Fair. I was there for six years. And my target there was really to try to figure out how to scale up recipes for multimodal models. And so one of the early things that we started working on, and we've published a lot in this domain, was around early fusion. So this concept that you want to bring in all these modalities as early as you can. And the real complexity there is trying to figure out what is the correct way to actually properly represent all the different modalities, both on the input and on the output that you want to be able to represent holistically. And so VLMs have kind of become the standard when I talk about multimodal models,
you probably think of VLMs. So this is the ability to take in some image, video, and some text and be able to output some text essentially. And there's variations of this. There's models like the ER models, the embodied reasoning models, that are able to output maybe some grounding points, that are able to do spatial understanding or reasoning a little bit better. Then we have things like VLAs that extend the output domain from just text to also actions. And these are traditionally, of course, continue to be built on by standard VLM backbones, although there's been some efforts to try to migrate away from VLMs to things like role models
or role action models, although nothing that's been super fruitful yet. And then we get into more interesting and complex variants of multimodal models like role models, where you essentially are outputting video from some inputs, and the inputs can be either image or video, or image, video, and actions. And the last point that I'll talk about is something recent, which we call semantic role models, which is you don't really output anything, but you learn some type of representation that you think is useful in the future. What we view as an embodied foundation model is actually a framing that allows you to both do the standard perception, the embodied reasoning,
and then the North Star target of control all within one model. So being able to reason across different sense of modalities on the input and being able to do most of what I mentioned on the output, all within one single unified model. And so very quickly, I'm going to talk about two challenges, and these are fundamental research challenges that we face, and I'll talk about how our company has approached this and what other folks are doing in this area as well. So the very first thing is think about purely if you're going to try to model something like video, like one hour of video. Depending on the representation, you might have something like one million visual tokens
that are coming in. The truth is that there's not actually any ground truth that you can use effectively, right? So you can do things like, and people have done this, of course, like pull out the transcripts, predict the transcripts from the video, or maybe synthetically label some frames, ask some questions. And it turns out that this is a humongous waste, right? So if you think about what's going into your model, you have one million tokens going in, and you're essentially calculating the loss on something like 0.2 percent of all the tokens that are going in. And so this is very fundamentally problematic. And the truth is, any way you try to figure out how to fix this,
you're essentially injecting a wrong training signal. Either the signal is too sparse, it's too synthetic, or it's too indiscriminate. And so approaches beyond just synthetic enrichment have been, well, let's predict every single pixel. I mean, true, this is a very dense signal, but it actually very poorly allocates attention. You're comparing a background pixel with the same degree of importance as you're treating a grip or tip, or the contact points, or the specific failures, or the physics. And so what we do at Perceptron is, and this is some of our core IP, is we think about what does a natural perceptive objective look like?
So specifically, how can I predict the percepts that we think will matter in the future in a very, very automatic way? So as an example, you might hard code something like, well, if you have a robotic arm, well, the tip of the grippers turns out to be a very useful percept that you can predict into the future. And folks have started doing this. Like the MOMO Act folks from AI2 have done this. There are other VLAs that have done this. But this is still a hard coded percept. So the question is, can you figure out an automatic way that the model semantically is able to learn this very unique objective? And the truth is we have figured out a way.
We're not going to share how we do it here, but this is just hinting at how we approach the problem of sparsity. The second core problem that we've spent a lot of time focusing on is context bloat. So if you have always-on cameras, if you have robots that don't necessarily wait, they don't stop, you're essentially having to reason over a very long, very long amount of tokens. And so text is relatively dense and video is relatively sparse. And so the question is, are there architectural breakthroughs that allow you to be able to deal with this problem natively rather than just trying to figure out how to fix this imbalance?
And so there's a couple of things that you can do. One core first principle is you need to start training different modalities as completely different. So you can't treat text tokens as the same as image tokens, as the same as audio or video tokens. So one thing that you can start thinking about doing is focusing on spatial compression or token compression. And people do relatively simplistic things, and we started off doing the simplistic things, and it does work. You can start thinking about averaging patch-wise representations across a video or an image. And you can start getting some interesting compression rates,
up to 10x. But still, this is relatively a hack, and there aren't well-used architectural methods to actually solve this. So what we've done in the last couple of months, we've released what we think is our approach to dealing with varying degrees of sparsity, which is let the model figure out what tokens it should look at and what tokens it shouldn't. And so we released our data sparse mixture of experts paper, which essentially allows you to do this. There's a router in the model, it allows it to actually predict what token I should input, what token I should skip, and it allows us to do it for free through all the different layers.
And it turns out that if you just let the model learn, if you're not actually hard-coding any significant architectural priors, the model actually does learn. So if you end up visualizing the data sparse compute that our models use, you actually see that the models innately learn to start focusing on very high-density information or task-relevant information. So in upper right, you can see that the model decides to focus in on the graph, which is likely what's interesting within the figure. And it turns out that even task-dependent allocation ends up happening as well. So if you look
at the bottom left, if you just ask a very general question, you're going to see an attention graph that is throughout the whole image. So the model just doesn't know what the proper way to allocate compute is. At the same time, if you ask it to do something like segment out all the fruit, you can see that it's going to allocate more tokens to what it thinks are fruit tokens. And so this is a very nice and clever trick that we use and we've published and other folks are starting to use around embedding priors into the architecture that are useful to deal with the sparsity imbalances of your modalities, but not
too harsh to the point that the models aren't actually learning natively. And so we put all this together, and you guys might have seen the release, but we essentially
task-dependent allocation ends up happening as well. So if you look at the bottom left, if you just ask a very general question, you're going to see an attention graph that is throughout the whole image. So the model just doesn't know what the proper way to allocate compute is. At the same time, if you ask it to do something like segment out all the fruit, you can see that it's going to allocate more tokens to what it thinks are fruit tokens. And so this is a very nice and clever trick that we use and we've published and other folks are starting to use around embedding priors into the architecture that are useful to deal with the sparsity imbalances of your modalities, but not too harsh to the point that the models aren't actually learning natively.
And so we put all this together, and you guys might have seen the release, but we essentially released our model which was what we considered to be the first embodied foundation model a couple weeks ago. And this model is essentially frontier with respect to Gemini 3.1 Pro. It's actually better than Gemini embodied reasoning, and it's something like 15 times cheaper. And it's essentially trained on this one petabyte dataset that we've collected across literally everything. It's from internet crawls to our own custom mid-training recipes or synthetic data pipelines. We have this one petabyte of data across text, images, videos, trajectories, and these trajectories can be very general. It could be desktop use trajectories. It could be playing a video game trajectory. And it turns out that once you start doing these things, very interesting properties end up emerging. And so the biggest property that we saw, which is obvious in retrospect, is that you can actually start thinking of doing classical CV tasks as being an agentic task. And so in this case, we essentially reframed detection as an agentic task. So our model can write code. It can ask to zoom in into specific portions. It can change the contrast. And you can actually see here, this is a very hard problem. I think there's a whole Reddit subreddit of these problems of trying to find very hard objects in images. And our models essentially do this very well. But they do this in an agentic sense. So this isn't a classical detect this one box. This is the model actually deciding that it needs to tile things up. It needs to change the contrast. It proposes a box here. I think here, it increases contrast and it can find the bird. So again, very hard to do, even for a human, and humans are very good at perceptive tasks, this is a relatively tough thing to do. And this all comes from just having natively embodied models that actually understand how to look at different modalities. And so following up, how does this relate to the general physical AI stance around robotics? So I stole this slide from GDM folks. And so one thing we're starting to see from robotics agentic systems is this separation between what we call embodied reasoning models or orchestrators and tactile policy models. So you can think of problems as like, if I'm making coffee and that takes me three minutes to do that, one thing I can do is try to force my whole VLA to try to figure out how to do this individual task. Or what I can use, I can have an orchestrator model that breaks up these tasks into sub-tasks, and there's a tactile control policy that's running on top. And there's a full spectrum between full VLA only all the way to this agentic system. But the main thing I'm trying to highlight is that embodied reasoning is actually a very interesting and complex problem that is yet to be solved. That being said, our models continue to be frontier on embodied reasoning. And because they're frontier, we start seeing really cool things that we haven't seen before.
Here's a concrete example of doing very complex robotic data annotation. And you can see here the model is jumping around, looking at different portions of the video, clipping it, figuring out whether or not the captions are correct, self-verifying. And we can all do this because AR models are fast. They're significantly cheaper than anything else that's out there. So if you try to do this with Gemini, this video would probably cost you a couple of dollars, or for us it's probably in the cents. And these all emerged from being able to have these frontier embodied reasoning capabilities that we just previously have not seen from other models. Okay, probably going to share the biggest research breakthrough that we've had, and I think we'll share more of this in the upcoming weeks probably on Twitter. But one thing that we found is we've discovered new scaling laws for embodied foundation models. So these are models that you can jointly do control-based training, you can do trajectory training, you can do perceptive training, you can do embodied reasoning training. If you just figure out what the right way to mix this all together is and the right objectives to use, you actually start seeing very interesting levers that you maybe previously haven't been able to see before.
So the concrete lever that I'll talk about is this ability to trade general video pre-training data for tele-op data. So as we know, tele-op data is very expensive. It's on the orders of $100 per hour of data. For $100, I can collect significantly more video pre-training data. And so what this graph is showing is that if you're just training pure VLAs, pure policies, there's this band that you have. So you do actually still have scaling loss. So you do get benefits for more and more tele-op data. That being said, the benefits are not as substantial as if you are really training these unified embodied foundation models. And so this is what the bottom half of the graph is. And the really cool lever that we get is you can essentially trade 10x less tele-op data if you have 10x more video pre-training data. And so far this has held for the amount of compute that our company has. And we will be continuing to push the fold on how far you can push these embodied foundation models.
So here's a couple of videos of a policy that hopefully we will open source one of the smaller models in a couple of weeks. But this is all running natively within a single model that is capable of doing the embodied reasoning in order to figure out the task. It's actually outputting control tokens. You can see it's a little bit jittery, but that's okay. Hopefully it will be figured out at scale. And you can actually see very complex tasks that previously I think would be really tough for pure VLAs to do. So if you look at the right-hand video, this is requiring the model to actually read the title of the book, have the knowledge about what type of book this is, and then properly allocate it within one of the bins. So this is actually a multi-step task between perception and control that is really tough to do if you have a pure end-to-end control model that is not aware of the different perceptive tasks that it needs to accomplish in order to do this task.
Yeah, it's pretty cool. It also works zero-shot relatively well out of the box. So we're excited to get this in the hands of folks in a couple of weeks, sometime in July. Yeah, going to leave a couple of minutes for general questions, but if you guys are interested, let's connect. So one cool thing that we do with our company is we actually, for a limited set of partners, give access to our Mark 1 weights. We give access to our larger embodied foundation models weights. So email me, DM me on Twitter, whatever is easier. And then yeah, I'll open up. There's a couple minutes left for questions.
Really tough to do if you have a pure end-to-end control model that is not aware of the different perceptive tasks that it needs to accomplish in order to do this task. Yeah. It's pretty cool. It also works zero-shot relatively well out of the box. So we're excited to get this in the hands of folks in a couple of weeks, sometime in July.
Yeah. Going to leave a couple of minutes for general questions, but if you guys are interested, let's connect. So one cool thing that we do with our company is we actually, for a limited set of partners, give access to our Mark 1 weights. We give access to our larger embodied foundation models weights. So email me, DM me on Twitter, whatever is easier. And then, yeah, I'll open up. There's a couple minutes left for questions.
[SPEAKER_00]: Like the temporal and spatial aspects of these models, because that would have been a challenge for a very long time. And when I saw on your example from right, it's reading the catalog and now into what you're looking at, it has a lot of content. And then spatial already is obviously clear, and also I see some temporal. So can you talk a little bit about that?
Yeah, I think the question is how are we able to nail temporal understanding to this degree? It's a good question. I mean, to be honest, it's not nailed. So there's still a lot of work to actually get it to a place where you can reliably deploy. The core thing is how do you think about context management? So you have a relatively limited context. I think the models here have one million contexts, but that's relatively easy to fit in with high FPS video. And so you have to start thinking about, are there interesting things that you can do? I'll throw something out there. We used to do this, but we got past this. How do I think about key frames versus delta frames? How can I manage my context by training these two off? And then you start thinking about during your pre-training objective, how can I start natively ingesting things that I think will be useful for the robotics tasks?
So, for example, a very basic thing that even the Gemini models used to struggle at—I think the new ones are pretty good—but being able to tell cardinalities. So left and right is very hard to tell if you do internet scale crawls, because no one on the internet is necessarily labeling things as, you know, this object is to the left of this object, it's below this object. So really thinking about data distributions early on gives you this ability relatively quickly. And it's also embodied reasoning. This type of embodied reasoning was a very concrete focus with us, which is why I think we were able to surpass Gemini with relatively less compute. So with VLA and VLA, this robustness on products?
I was wondering whether you looked at, are your models that are potentially more robust than the kind of...
Yeah, so probably the coolest robustness that we've seen is that for VLA models specifically, if you go and you take one of the Chinese ones, and you try to fine tune it for a specific policy, if you just change the background of the table, if you change the background of the table, the policy will actually fail. What's really interesting, if you do this type of joint perceptive and control modeling, you're much more robust to these types of errors or if the light is hitting it a slightly different way. And I think we primarily view this as robustness to background in a way that I think traditional models don't necessarily have. That being said, I'm not going to over claim. I mean, it's still relatively hard. I think if I was going to go and shine a flashlight into one of the arms, it's probably not going to work. But being able to jointly model these things helps a significant amount. We also do a lot of online augmentation, so we do actually fake one of the arm cameras being off, right? We fake sunlight coming in from a certain direction, right? So we do these things during training to improve robustness. But the big gains come from taking this early fusion paradigm and then moving into the robotics domain.
Cool. Is it more useful to be a knowledge base using the K1 model? Oh, knowledge bases? It's useful for, I mean, if you want to caption images, videos, it's relatively well. We work with robotics partners for very complex egocentric annotation. So it's not necessarily building on an ontology, but being able to do very deep structured extraction, I think our models are very good at. By the way, everything that I showed here is public APIs, so you can go play around with it. The benchmarks are public. I think I'm out of time. They're cutting me off, so I can talk with folks outside. But thank you, guys.
it increases contrast and it can find the bird. So again, very hard to do if, like, even for a human, and humans are very good at perceptive tasks, this is a relatively tough thing to do. And this all comes from just having natively embodied models that actually understand how to look at different modalities. And so following up, kind of how does this relate to the general physical AI stance around robotics? So I kind of stole this slide from GDM folks. And so one thing we're starting to see from kind of robotics agentic systems is this separation between what we call kind of embodied reasoning
models or orchestrators and tactile policy models. So you can think of problems as like, you know, if I have a ‑‑ if I'm making coffee and that takes me three minutes to do that, I mean, one thing I can do is try to force my whole, you know, VLA to try to figure out how to do this individual task. Or what I can use, I can have an orchestrator model that breaks up these tasks into sub-tasks, and there's a tactile control policy that's running on top. And there's kind of a full spectrum between, you know, full VLA only all the way to this kind of agentic system. But the main thing I'm trying to highlight is that
embodied reasoning is actually a very interesting and complex problem that is yet to be solved. That being said, our models continue to be frontier on embodied reasoning. And because they're frontier, we start seeing really cool things that we haven't seen before. Here's a concrete example of doing very complex egocentric ‑‑ or not egocentric, but this is robotic data annotation. And you can kind of see here the model is jumping around, looking at different portions of the video, you know, clipping it, figuring out whether or not the captions are correct, self‑verifying. And we can all do this because AR models are fast. They're significantly cheaper than
anything else that's out there. So if you try to do this with Gemini, this video would probably cost you a couple of dollars, or for us it's probably in the sense. And these all kind of emerged from being able to have these frontier embodied reasoning capabilities that we just previously have not seen from other models. Okay, probably going to share the biggest research breakthrough that we've had, and I think we'll share more of this in the upcoming weeks probably on Twitter. But one thing that we found is we've discovered new scaling laws for embodied foundation models. So these are models, again,
that you can jointly do control‑based training, you can do trajectory training, you can do perceptive training, you can do embodied reasoning training. If you just figure out what the right way to mix this all together is and the right objectives to use, you actually start seeing very interesting levers that you maybe previously haven't been able to see before. So the concrete lever that I'll talk about is this ability to trade general video pre‑training data for tele‑op data. So kind of as we know, tele‑op data is very expensive. It's on the orders of, you know, $100 per hour of data. For $100, I can collect significantly more video pre‑training data.
And so what this graph is showing is that if you're just training pure VLA's, pure policies, there's this band that you have. So you do actually still have scaling loss. So you do get benefits for more and more tele‑op data. That being said, the benefits are not as substantial as if you are really training these unified embodied foundation models. And so this is what kind of the bottom half of the graph is. And the really cool kind of lever that we get is you can essentially trade 10x less tele‑op data if you have 10x more video pre‑training data. And so far this is kind of held for the amount of compute
that our company has. And it will be continuing to kind of push the fold on how far you can push these embodied foundation models.
So here's a couple of videos of a policy that hopefully will open source one of the smaller models in a couple of weeks. But this is all running natively within a single model that is capable of doing the embodied reasoning in order to figure out the task. Actually is outputting control tokens. You can see it's a little bit jittery, but that's okay. Hopefully it will be figured out at scale. And you can actually see very complex tasks that previously I think would be really tough for pure VLA's to do. So I think if you look at the right‑hand video, this is requiring the model to actually read the title
of the book, have the knowledge about what type of book this is, and then properly allocate it within one of the bins. So this is actually a multi‑step task between perception and control that is really, really tough to do if you have a kind of a pure end‑to‑end control model that is not aware of the different perceptive tasks that it needs to accomplish in order to do this task.
Yeah. It's pretty cool. It also works zero‑shot relatively well out of the box. So we're excited to get this in the hands of folks in a couple of weeks, sometime in July.
Yeah. Going to leave a couple of minutes for general questions, but if you guys are interested, let's connect. So one cool thing that we do with our company is we actually, for a limited set of partners, give access to our Mark 1 weights. We give access to our larger embodied foundation models weights. So email me, DM me on Twitter, whatever is easier. And then, yeah, I'll open up. There's a couple minutes left for questions.
JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF
like the temporal and spatial aspects of these models, because that would have been a challenge for a very long time. And like when I saw on your example from right, it's just like, it's reading the catalog and now into what you're looking at, like it has a lot of content. And then spatial already is obviously clear, and also I see some temporal . So can you talk a little bit about that? Yeah, I think it's, so yeah, so the question is how are we able to nail temporal and understanding to this degree? It's a good question. I mean, to be honest, it's not nailed. So there's still a lot of work to actually get it to a place
where you can reliably deploy. The core thing is how do you think about context management? So you have a relatively limited context. So I think the models here have one million contexts, but that's relatively easy to fit in with the high FPS video. And so you have to start thinking about, are there interesting things that you can do? I'll throw something out there. We used to do this, but we got past this. But like, how do I think about like key frames versus delta frames? How can I manage my context by training these two off? And then you start thinking about during your pre-training objective,
how can I start kind of natively ingesting things that I think will be useful for the robotics tasks? So, for example, a very basic thing that even kind of the Gemini models used to struggle at, I think the new ones are pretty good, but being able to tell cardinalities. So like left and right is very hard to tell if you do internet scale crawls, because no one on the internet is necessarily labeling things as, you know, this object is to the left of this object, it's below this object. So really thinking about data distributions early on gives you this ability relatively quickly. And it's also like ER, this type of embodied reasoning was a very concrete focus with us,
which is why I think we were able to kind of surpass a Gemini ER with relatively less compute.
So what issue we, with VLA and VLA, this robustness on products? I was wondering whether, you know, this is really, really cool, I was wondering whether you looked at, are your models that are potentially more robust than, you know, the kind of . Yeah, so probably the coolest robustness that we've seen is that for VLA models specifically, like if you go and you take one of the Chinese ones, and you try to fine tune it for a specific policy, if you just change the background of the, I don't know, even like in the table, if you change the background of the table, the policy will actually fail.
What's really interesting, if you do this type of joint perceptive and control modeling, you're much more robust to these types of errors or if the light is hitting it a slightly different way. And I think we primarily view this as robustness to background in a way that I think traditional models don't necessarily have. That being said, I'm not going to over claim, like, I mean, like, it's still relatively hard. I think if I was going to go and shine a flashlight into one of the arms, it's probably not going to work. But being able to jointly model these things helps a significant amount.
We also do a lot of online augmentation, so we do actually, you know, fake, I don't know, one of the arm cameras being off, right? We fake sunlight coming in from a certain direction, right? So we do these things during training to improve robustness. But the big gains come from taking this early fusion paradigm and then moving into the robotics domain.
Cool. Is it more useful to be a knowledge base using the K1 model? Oh, knowledge bases? It's useful for, like, I mean, if you want to caption images, videos, it's relatively well. We work with robotics partners for, like, very complex egocentric annotation like this. So it's not necessarily building on an ontology, but being able to do kind of very deep structured extraction, I think our models are very good at. By the way, everything that I kind of showed here is kind of public APIs, so you can go play around with it. The benchmarks are public. I think I'm out of time. They're cutting me off, so I can talk with folks outside. But thank you, guys.
I think I need to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to