From VLM/VLA's to Embodied Agents — Armen Aghajanyan, Perceptron AI
Description
Feed a model one hour of video and roughly a million visual tokens go in, yet the loss lands on about two percent of them, since the only ground truth is a transcript or a few labeled frames. Armen Aghajanyan calls that a humongous waste, and predicting every pixel treats a background pixel with the same weight as a gripper tip or a contact point. Perceptron's answer is a perceptive objective that learns which percepts will matter, rather than hardcoding the gripper. The second problem is context bloat from cameras that never switch off. Patch averaging buys ten times compression, but his fix is data sparse mixture of experts, a router that decides per layer which tokens to read and which to skip. Left alone, the model zooms into the graph in a figure and spends more tokens on fruit when asked to segment fruit. Put together, that produced the model his team released a few weeks earlier, trained on a petabyte spanning text, images, video, and trajectories from desktop use to video games, which he says beats a frontier lab's embodied reasoning model at a fraction of the cost. Detection turns into an agentic task, with the model tiling the image, raising the contrast, and proposing boxes until it finds the bird. The biggest result is a new scaling law: training jointly on perception, reasoning, and control lets ten times more video pretraining substitute for ten times less teleop data, which costs about a hundred dollars an hour. He closes with one model emitting control tokens to sort books by reading their titles, an open release promised for July, and questions on temporal context, background robustness, and structured extraction. Speaker info: - https://x.com/ArmenAgha - https://www.linkedin.com/in/armenag - https://perceptron.inc Timestamps: 0:00 - Perceptron's north star: one model that perceives, reasons, and acts 1:23 - Early fusion, and the VLM to VLA to world model ladder 3:40 - Challenge one: an hour of video has almost no ground truth 5:41 - Challenge tw
Summary
Generated by gpt-5.6-terraAt-a-Glance
- Verdict: Watch fully
- Core thesis: Perceptron argues that physical AI should converge on unified embodied foundation models that jointly perceive, reason over, and control the world, rather than treating VLMs, VLAs, world models, and policies as separate categories.
- Why it matters: The talk offers a concrete architecture-and-training thesis for long-context multimodal agents: learn task-relevant percepts and dynamically allocate compute, then use broad video pretraining to reduce reliance on costly action demonstrations.
- Best use: Use it to extract design patterns for multimodal agent control planes and to evaluate Perceptron as a potential model/data partner in robotics, video understanding, or structured visual-agent workflows.
Executive Summary
Armen Aghajanyan frames Perceptron's work as an attempt to build an "embodied foundation model"—one system that accepts heterogeneous sensory inputs and can produce language, grounded representations, trajectories, and control tokens. His argument is that the usual labels of VLM, VLA, and world model obscure the desired endpoint: a model that can perceive, reason, and act in real time across cameras, robots, sensors, devices, and digital interfaces.
The technical case rests on two bottlenecks. First, ordinary video-language supervision is extremely sparse: an hour of video can contain roughly one million visual tokens, while transcript or synthetic-label losses supervise only a tiny fraction. Pixel prediction is dense but indiscriminate, treating irrelevant background pixels as importantly as contacts, gripper tips, failures, and physical interactions. Perceptron says it instead learns automatically selected, future-useful perceptive objectives, though it does not disclose the method.
Second, long-lived camera and robot streams create context bloat. Perceptron's published Data Sparse Mixture of Experts approach uses routing to decide which tokens warrant computation at each layer, rather than uniformly processing text, image, audio, and video tokens. The claimed result is task-conditional allocation: broad attention for general questions, but concentrated allocation on relevant objects for a task such as fruit segmentation.
The strategic claim is that joint training across video, perception, trajectories, embodied reasoning, and control creates a more data-efficient and robust path to robotics than a pure end-to-end VLA policy. Perceptron reports that 10x more general video pretraining can substitute for 10x less teleoperation data, which it estimates costs about $100 per hour. It also reports better tolerance of visual background and lighting shifts, but acknowledges that reliable temporal understanding and deployment-grade robustness remain unsolved.
Key Takeaways
- Claim: The useful abstraction for physical AI is a unified embodied foundation model, not a collection of separately named VLM, VLA, and world-model components. | Evidence: Aghajanyan defines the target as one model that reasons across sensory modalities and can perform perception, embodied reasoning, and control; its inputs span text, images, video, and trajectories, while outputs can include language, grounding, and control tokens. | Implication: For agent-system design, separate planning, visual understanding, and action modules should be treated as an implementation choice rather than a fixed conceptual boundary; shared representations may improve cross-stage handoffs. | Caveat: This is a research framing and product direction, not evidence that a single model is universally preferable to a modular system for every safety-critical or latency-sensitive deployment.
- Claim: Sparse video-language supervision is a fundamental training mismatch for embodied intelligence, and reconstructing all pixels is not an adequate remedy. | Evidence: The speaker estimates that one hour of video may produce about one million visual tokens, while transcript prediction or synthetic frame-question labels compute loss on roughly 0.2% of those tokens. Pixel-level objectives are criticized for assigning equal importance to backgrounds and physically meaningful features such as gripper tips, contacts, failures, and physics. | Implication: Training data and objectives for visual agents should be evaluated for whether they reward causal/task-relevant state prediction, rather than merely adding captions, QA labels, or dense reconstruction losses. | Caveat: Perceptron withholds the mechanism behind its claimed automatic percept-selection objective, so its advantage cannot be independently assessed or reproduced from this talk.
- Claim: Context management for continuous multimodal streams should be learned through conditional compute allocation rather than handled solely by fixed compression heuristics. | Evidence: Perceptron's Data Sparse Mixture of Experts paper adds a router that chooses which tokens to process or skip through model layers. The speaker contrasts this with patch averaging, which can yield up to 10x compression but is characterized as a limited hack. Visualizations reportedly show the router selecting a graph for a figure question and concentrating on fruit regions when asked to segment fruit. | Implication: For OpenClaw-style multimodal workflows, route expensive visual processing based on task and salience instead of applying one uniform frame rate, token budget, or image-resolution policy to every step. | Caveat: The examples demonstrate qualitative behavior; the transcript provides no latency, throughput, routing-overhead, or reliability benchmarks for production workloads.
- Claim: Embodied reasoning can make perception an iterative agentic process rather than a one-shot computer-vision prediction. | Evidence: In a hard object-finding example, the model reportedly writes code, tiles or zooms image regions, changes contrast, proposes a bounding box, and finds a bird. In robotic video annotation, it jumps among clips, checks whether captions are correct, and self-verifies; the speaker claims this can cost cents versus a couple of dollars with Gemini. | Implication: High-value visual automation should expose tools such as crop, zoom, contrast adjustment, temporal clip selection, and verification, rather than demanding perfect perception from a single static model call. | Caveat: The cost comparison and capability claims are vendor-reported, with no workload specification, model configuration, or independent evaluation in the transcript.
- Claim: Joint embodied pretraining may sharply reduce dependence on expensive teleoperated robotics demonstrations. | Evidence: Perceptron claims a scaling relationship in which 10x more general video pretraining can trade for 10x less tele-op data; it places tele-op collection at roughly $100 per hour. The model is trained on a stated one-petabyte corpus spanning internet crawls, custom mid-training data, synthetic pipelines, text, images, videos, and trajectories including desktop use and videogame interaction. | Implication: For embodied-agent investment or build plans, broad passive visual data and interactive digital trajectories may be economically strategic complements to scarce real-world demonstrations, but the claimed substitution rate needs validation on target tasks. | Caveat: The reported scaling law has only held within Perceptron's available compute range, and the transcript does not provide task-level success rates, graph axes, or external replication.
- Claim: A practical robotics stack may separate long-horizon embodied reasoning/orchestration from low-level tactile control, even when both are trained within a unified foundation-model regime. | Evidence: Using a coffee-making example, Aghajanyan contrasts forcing a VLA to execute an entire three-minute task with an orchestrator that decomposes it into subtasks while a tactile control policy executes actions. He presents this as a spectrum from pure VLA to agentic orchestration. | Implication: Design long-horizon physical workflows with an explicit hierarchy: a stateful planner that manages goals, observations, and recovery, plus tight-loop control policies for manipulation and safety-critical execution. | Caveat: The speaker also shows a single model emitting control tokens, so the presentation supports a hybrid spectrum rather than a definitive architectural prescription.
- Claim: Joint perceptive-control modeling appears to improve robustness to ordinary visual distribution shifts, though not enough for unconstrained deployment. | Evidence: The speaker says a fine-tuned VLA policy can fail when only the tabletop background changes, whereas joint perception-and-control models are more tolerant of background variation and modest lighting changes. Perceptron also trains with online augmentation such as simulated camera failure and directional sunlight. | Implication: Treat multimodal pretraining and augmentation as robustness multipliers, not safety guarantees; validate under camera outages, lighting, backgrounds, occlusion, and temporal-state changes before operational use. | Caveat: Aghajanyan explicitly says the system would likely still fail under stronger perturbations, such as shining a flashlight into an arm camera; temporal understanding is likewise "not nailed" and not yet reliably deployable.
Detailed Brief
Spatial and temporal competence depends on the pretraining distribution, not architecture alone
- Claims: Long-context capacity alone does not solve high-frame-rate video understanding because even a one-million-token context can fill quickly.; The model must learn spatial relations and temporal state from data deliberately shaped around downstream embodied tasks.
- Evidence: The speaker cites key-frame versus delta-frame strategies as an earlier context-management direction that Perceptron moved beyond.; He highlights cardinality and spatial-relational concepts such as left/right and above/below as skills that internet-scale crawls may not label sufficiently, noting that earlier Gemini models struggled with such distinctions.
- Caveats: No concrete temporal benchmark, horizon length, failure-rate analysis, or comparison against other context architectures is supplied.
- Implications: A multimodal training program should explicitly audit whether source data supervises relations, object counts, temporal transitions, and action-relevant geometry rather than assuming web-scale video provides them.
Product and access signals
- Claims: Perceptron positions its model as both a robotics capability and a general visual-structure extraction system.; The company says its public APIs and public benchmarks can be used to test image/video captioning and deep structured extraction, while access to larger Mark 1 weights is limited to selected partners.
- Evidence: The speaker describes use with robotics partners for complex visual and egocentric-style annotation, rather than ontology or knowledge-base construction.; He says a smaller policy model may be open-sourced, while larger embodied-foundation weights are available through partner access.
- Caveats: The timing references are relative to the event and may no longer be current; no API capabilities, pricing, licensing, or benchmark methodology are specified.
- Implications: A low-commitment evaluation can begin with public APIs and benchmark inspection, while any partner conversation should focus on weight access, data rights, deployment mode, and evidence on the target visual-control distribution.
Notable Concepts & Terms
- Embodied foundation model: Perceptron's proposed unified model class for perception, reasoning, and control across physical-world sensory inputs and actions.
- VLM: Vision-language model: typically consumes images or video plus text and produces text; presented as insufficiently action-oriented on its own.
- VLA: Vision-language-action model: extends a VLM-like backbone to emit actions, but is presented as weaker when used as the entire long-horizon robotics stack.
- Embodied reasoning / ER model: A model capability for spatial understanding, grounding, task decomposition, and interactive perception before or alongside action.
- Data Sparse Mixture of Experts: Perceptron's routing-based architecture that selectively allocates computation to tokens deemed relevant instead of processing all multimodal tokens uniformly.
- Natural perceptive objective: Perceptron's undisclosed training objective intended to automatically identify and predict future-useful percepts rather than depend on sparse synthetic labels or full pixel reconstruction.
- Tele-op data: Human-operated robot demonstration data, described as expensive but potentially replaceable in part through large-scale general video pretraining.
- Orchestrator versus tactile policy: A hierarchical robotics pattern in which an embodied reasoning model plans and decomposes a long task while a lower-level policy handles fine motor control.
Operator Notes / Why Ken Should Care
- Run a small benchmark against Perceptron's public API using representative tasks that require iterative visual inspection: difficult object localization, temporal clip selection, image/video structured extraction, and self-verification.
- For any multimodal-agent pipeline, prototype a task-conditioned perception loop with explicit tools for crop, zoom, enhancement, clip navigation, and verification; compare its cost and accuracy against one-shot vision calls.
- When assessing embodied-model vendors, require task-level evidence for the asserted video-pretraining-to-tele-op-data substitution ratio, including target task, data volumes, compute, robustness conditions, and success-rate curves.
- Keep long-horizon planning separate from low-level actuation in the operating design until unified-control claims are validated under failures, latency limits, camera loss, lighting changes, and safety constraints.
- Monitor the availability and terms of Perceptron's smaller open model and partner-only Mark 1 weights; prioritize diligence on benchmark methodology and deployment readiness over headline comparisons to Gemini.
Source/Metadata
- Title: From VLM/VLA's to Embodied Agents — Armen Aghajanyan, Perceptron AI
- Transcript words: 6078
- Duration seconds: 1241
- Timestamp note: No timestamps or chapter markers were provided. The transcript contains substantial duplicated passages and extraction artifacts near the end.
Transcript
Armin Reviewer Yeah, let's get started. I'm Armin. I'm the co-founder and CEO of Perceptron. We'll get a little bit into what we do, but primarily what I want to talk about today is our research stance that we want to move away from distinctions between VLMs, VLAs, world models, whatever you want to call it, to something that we call embodied foundation models. So specifically what we do at Perceptron, kind of our North Star really, is we want to be able to build physical AI foundations that give us the ability to perceive, understand, and interact with the physical world in real time. And so the North Star mission is really to bridge the physical and digital worlds, fundamentally meaning that our goal is wherever there is a device, an instrument, a robot, a camera, a sensor, we're essentially there providing intelligence to it. And specifically when we talk about this paradigm of being able to perceive, being able to reason, being able to act, we view this as a unification of traditional multimodal modeling. So I came from Fair. I was there for six years. And my target there was really to try to figure out how to scale up recipes for multimodal models. And so one of the early things that we started working on, and we've published a lot in this domain, was around early fusion. So this concept that you want to bring in all these modalities as early as you can. And the real complexity there is trying to figure out what is the correct way to actually properly represent all the different modalities, both on the input and on the output that you want to be able to represent holistically. And so VLMs have kind of become the standard when I talk about multimodal models, you probably think of VLMs. So this is the ability to take in some image, video, and some text and be able to output some text essentially. And there's variations of this. There's models like the ER models, the embodied reasoning models, that are able to output maybe some grounding points, that are able to do spatial understanding or reasoning a little bit better. Then we have things like VLAs that extend the output domain from just text to also actions. And these are traditionally, of course, continue to be built on by standard VLM backbones, although there's been some efforts to try to migrate away from VLMs to things like role models or role action models, although nothing that's been super fruitful yet. And then we get into more interesting and complex variants of multimodal models like role models, where you essentially are outputting video from some inputs, and the inputs can be either image or video, or image, video, and actions. And the last point that I'll talk about is something recent, which we call semantic role models, which is you don't really output anything, but you learn some type of representation that you think is useful in the future. What we view as an embodied foundation model is actually a framing that allows you to both do the standard perception, the embodied reasoning, and then the North Star target of control all within one model. So being able to reason across different sense of modalities on the input and being able to do most of what I mentioned on the output, all within one single unified model. And so very quickly, I'm going to talk about two challenges, and these are fundamental research challenges that we face, and I'll talk about how our company has approached this and what other folks are doing in this area as well. So the very first thing is think about purely if you're going to try to model something like video, like one hour of video. Depending on the representation, you might have something like one million visual tokens that are coming in. The truth is that there's not actually any ground truth that you can use effectively, right? So you can do things like, and people have done this, of course, like pull out the transcripts, predict the transcripts from the video, or maybe synthetically label some frames, ask some questions. And it turns out that this is a humongous waste, right? So if you think about what's going into your model, you have one million tokens going in, and you're essentially calculating the loss on something like 0.2 percent of all the tokens that are going in. And so this is very fundamentally problematic. And the truth is, any way you try to figure out how to fix this, you're essentially injecting a wrong training signal. Either the signal is too sparse, it's too synthetic, or it's too indiscriminate. And so approaches beyond just synthetic enrichment have been, well, let's predict every single pixel. I mean, true, this is a very dense signal, but it actually very poorly allocates attention. You're comparing a background pixel with the same degree of importance as you're treating a grip or tip, or the contact points, or the specific failures, or the physics. And so what we do at Perceptron is, and this is some of our core IP, is we think about what does a natural perceptive objective look like? So specifically, how can I predict the percepts that we think will matter in the future in a very, very automatic way? So as an example, you might hard code something like, well, if you have a robotic arm, well, the tip of the grippers turns out to be a very useful percept that you can predict into the future. And folks have started doing this. Like the MOMO Act folks from AI2 have done this. There are other VLAs that have done this. But this is still a hard coded percept. So the question is, can you figure out an automatic way that the model semantically is able to learn this very unique objective? And the truth is we have figured out a way. We're not going to share how we do it here, but this is just hinting at how we approach the problem of sparsity. The second core problem that we've spent a lot of time focusing on is context bloat. So if you have always-on cameras, if you have robots that don't necessarily wait, they don't stop, you're essentially having to reason over a very long, very long amount of tokens. And so text is relatively dense and video is relatively sparse. And so the question is, are there architectural breakthroughs that allow you to be able to deal with this problem natively rather than just trying to figure out how to fix this imbalance? And so there's a couple of things that you can do. One core first principle is you need to start training different modalities as completely different. So you can't treat text tokens as the same as image tokens, as the same as audio or video tokens. So one thing that you can start thinking about doing is focusing on spatial compression or token compression. And people do relatively simplistic things, and we started off doing the simplistic things, and it does work. You can start thinking about averaging patch-wise representations across a video or an image. And you can start getting some interesting compression rates, up to 10x. But still, this is relatively a hack, and there aren't well-used architectural methods to actually solve this. So what we've done in the last couple of months, we've released what we think is our approach to dealing with varying degrees of sparsity, which is let the model figure out what tokens it should look at and what tokens it shouldn't. And so we released our data sparse mixture of experts paper, which essentially allows you to do this. There's a router in the model, it allows it to actually predict what token I should input, what token I should skip, and it allows us to do it for free through all the different layers. And it turns out that if you just let the model learn, if you're not actually hard-coding any significant architectural priors, the model actually does learn. So if you end up visualizing the data sparse compute that our models use, you actually see that the models innately learn to start focusing on very high-density information or task-relevant information. So in upper right, you can see that the model decides to focus in on the graph, which is likely what's interesting within the figure. And it turns out that even task-dependent allocation ends up happening as well. So if you look at the bottom left, if you just ask a very general question, you're going to see an attention graph that is throughout the whole image. So the model just doesn't know what the proper way to allocate compute is. At the same time, if you ask it to do something like segment out all the fruit, you can see that it's going to allocate more tokens to what it thinks are fruit tokens. And so this is a very nice and clever trick that we use and we've published and other folks are starting to use around embedding priors into the architecture that are useful to deal with the sparsity imbalances of your modalities, but not too harsh to the point that the models aren't actually learning natively. And so we put all this together, and you guys might have seen the release, but we essentially task-dependent allocation ends up happening as well. So if you look at the bottom left, if you just ask a very general question, you're going to see an attention graph that is throughout the whole image. So the model just doesn't know what the proper way to allocate compute is. At the same time, if you ask it to do something like segment out all the fruit, you can see that it's going to allocate more tokens to what it thinks are fruit tokens. And so this is a very nice and clever trick that we use and we've published and other folks are starting to use around embedding priors into the architecture that are useful to deal with the sparsity imbalances of your modalities, but not too harsh to the point that the models aren't actually learning natively. And so we put all this together, and you guys might have seen the release, but we essentially released our model which was what we considered to be the first embodied foundation model a couple weeks ago. And this model is essentially frontier with respect to Gemini 3.1 Pro. It's actually better than Gemini embodied reasoning, and it's something like 15 times cheaper. And it's essentially trained on this one petabyte dataset that we've collected across literally everything. It's from internet crawls to our own custom mid-training recipes or synthetic data pipelines. We have this one petabyte of data across text, images, videos, trajectories, and these trajectories can be very general. It could be desktop use trajectories. It could be playing a video game trajectory. And it turns out that once you start doing these things, very interesting properties end up emerging. And so the biggest property that we saw, which is obvious in retrospect, is that you can actually start thinking of doing classical CV tasks as being an agentic task. And so in this case, we essentially reframed detection as an agentic task. So our model can write code. It can ask to zoom in into specific portions. It can change the contrast. And you can actually see here, this is a very hard problem. I think there's a whole Reddit subreddit of these problems of trying to find very hard objects in images. And our models essentially do this very well. But they do this in an agentic sense. So this isn't a classical detect this one box. This is the model actually deciding that it needs to tile things up. It needs to change the contrast. It proposes a box here. I think here, it increases contrast and it can find the bird. So again, very hard to do, even for a human, and humans are very good at perceptive tasks, this is a relatively tough thing to do. And this all comes from just having natively embodied models that actually understand how to look at different modalities. And so following up, how does this relate to the general physical AI stance around robotics? So I stole this slide from GDM folks. And so one thing we're starting to see from robotics agentic systems is this separation between what we call embodied reasoning models or orchestrators and tactile policy models. So you can think of problems as like, if I'm making coffee and that takes me three minutes to do that, one thing I can do is try to force my whole VLA to try to figure out how to do this individual task. Or what I can use, I can have an orchestrator model that breaks up these tasks into sub-tasks, and there's a tactile control policy that's running on top. And there's a full spectrum between full VLA only all the way to this agentic system. But the main thing I'm trying to highlight is that embodied reasoning is actually a very interesting and complex problem that is yet to be solved. That being said, our models continue to be frontier on embodied reasoning. And because they're frontier, we start seeing really cool things that we haven't seen before. Here's a concrete example of doing very complex robotic data annotation. And you can see here the model is jumping around, looking at different portions of the video, clipping it, figuring out whether or not the captions are correct, self-verifying. And we can all do this because AR models are fast. They're significantly cheaper than anything else that's out there. So if you try to do this with Gemini, this video would probably cost you a couple of dollars, or for us it's probably in the cents. And these all emerged from being able to have these frontier embodied reasoning capabilities that we just previously have not seen from other models. Okay, probably going to share the biggest research breakthrough that we've had, and I think we'll share more of this in the upcoming weeks probably on Twitter. But one thing that we found is we've discovered new scaling laws for embodied foundation models. So these are models that you can jointly do control-based training, you can do trajectory training, you can do perceptive training, you can do embodied reasoning training. If you just figure out what the right way to mix this all together is and the right objectives to use, you actually start seeing very interesting levers that you maybe previously haven't been able to see before. So the concrete lever that I'll talk about is this ability to trade general video pre-training data for tele-op data. So as we know, tele-op data is very expensive. It's on the orders of $100 per hour of data. For $100, I can collect significantly more video pre-training data. And so what this graph is showing is that if you're just training pure VLAs, pure policies, there's this band that you have. So you do actually still have scaling loss. So you do get benefits for more and more tele-op data. That being said, the benefits are not as substantial as if you are really training these unified embodied foundation models. And so this is what the bottom half of the graph is. And the really cool lever that we get is you can essentially trade 10x less tele-op data if you have 10x more video pre-training data. And so far this has held for the amount of compute that our company has. And we will be continuing to push the fold on how far you can push these embodied foundation models. So here's a couple of videos of a policy that hopefully we will open source one of the smaller models in a couple of weeks. But this is all running natively within a single model that is capable of doing the embodied reasoning in order to figure out the task. It's actually outputting control tokens. You can see it's a little bit jittery, but that's okay. Hopefully it will be figured out at scale. And you can actually see very complex tasks that previously I think would be really tough for pure VLAs to do. So if you look at the right-hand video, this is requiring the model to actually read the title of the book, have the knowledge about what type of book this is, and then properly allocate it within one of the bins. So this is actually a multi-step task between perception and control that is really tough to do if you have a pure end-to-end control model that is not aware of the different perceptive tasks that it needs to accomplish in order to do this task. Yeah, it's pretty cool. It also works zero-shot relatively well out of the box. So we're excited to get this in the hands of folks in a couple of weeks, sometime in July. Yeah, going to leave a couple of minutes for general questions, but if you guys are interested, let's connect. So one cool thing that we do with our company is we actually, for a limited set of partners, give access to our Mark 1 weights. We give access to our larger embodied foundation models weights. So email me, DM me on Twitter, whatever is easier. And then yeah, I'll open up. There's a couple minutes left for questions. Really tough to do if you have a pure end-to-end control model that is not aware of the different perceptive tasks that it needs to accomplish in order to do this task. Yeah. It's pretty cool. It also works zero-shot relatively well out of the box. So we're excited to get this in the hands of folks in a couple of weeks, sometime in July. Yeah. Going to leave a couple of minutes for general questions, but if you guys are interested, let's connect. So one cool thing that we do with our company is we actually, for a limited set of partners, give access to our Mark 1 weights. We give access to our larger embodied foundation models weights. So email me, DM me on Twitter, whatever is easier. And then, yeah, I'll open up. There's a couple minutes left for questions. [SPEAKER_00]: Like the temporal and spatial aspects of these models, because that would have been a challenge for a very long time. And when I saw on your example from right, it's reading the catalog and now into what you're looking at, it has a lot of content. And then spatial already is obviously clear, and also I see some temporal. So can you talk a little bit about that? Yeah, I think the question is how are we able to nail temporal understanding to this degree? It's a good question. I mean, to be honest, it's not nailed. So there's still a lot of work to actually get it to a place where you can reliably deploy. The core thing is how do you think about context management? So you have a relatively limited context. I think the models here have one million contexts, but that's relatively easy to fit in with high FPS video. And so you have to start thinking about, are there interesting things that you can do? I'll throw something out there. We used to do this, but we got past this. How do I think about key frames versus delta frames? How can I manage my context by training these two off? And then you start thinking about during your pre-training objective, how can I start natively ingesting things that I think will be useful for the robotics tasks? So, for example, a very basic thing that even the Gemini models used to struggle at—I think the new ones are pretty good—but being able to tell cardinalities. So left and right is very hard to tell if you do internet scale crawls, because no one on the internet is necessarily labeling things as, you know, this object is to the left of this object, it's below this object. So really thinking about data distributions early on gives you this ability relatively quickly. And it's also embodied reasoning. This type of embodied reasoning was a very concrete focus with us, which is why I think we were able to surpass Gemini with relatively less compute. So with VLA and VLA, this robustness on products? I was wondering whether you looked at, are your models that are potentially more robust than the kind of... Yeah, so probably the coolest robustness that we've seen is that for VLA models specifically, if you go and you take one of the Chinese ones, and you try to fine tune it for a specific policy, if you just change the background of the table, if you change the background of the table, the policy will actually fail. What's really interesting, if you do this type of joint perceptive and control modeling, you're much more robust to these types of errors or if the light is hitting it a slightly different way. And I think we primarily view this as robustness to background in a way that I think traditional models don't necessarily have. That being said, I'm not going to over claim. I mean, it's still relatively hard. I think if I was going to go and shine a flashlight into one of the arms, it's probably not going to work. But being able to jointly model these things helps a significant amount. We also do a lot of online augmentation, so we do actually fake one of the arm cameras being off, right? We fake sunlight coming in from a certain direction, right? So we do these things during training to improve robustness. But the big gains come from taking this early fusion paradigm and then moving into the robotics domain. Cool. Is it more useful to be a knowledge base using the K1 model? Oh, knowledge bases? It's useful for, I mean, if you want to caption images, videos, it's relatively well. We work with robotics partners for very complex egocentric annotation. So it's not necessarily building on an ontology, but being able to do very deep structured extraction, I think our models are very good at. By the way, everything that I showed here is public APIs, so you can go play around with it. The benchmarks are public. I think I'm out of time. They're cutting me off, so I can talk with folks outside. But thank you, guys. it increases contrast and it can find the bird. So again, very hard to do if, like, even for a human, and humans are very good at perceptive tasks, this is a relatively tough thing to do. And this all comes from just having natively embodied models that actually understand how to look at different modalities. And so following up, kind of how does this relate to the general physical AI stance around robotics? So I kind of stole this slide from GDM folks. And so one thing we're starting to see from kind of robotics agentic systems is this separation between what we call kind of embodied reasoning models or orchestrators and tactile policy models. So you can think of problems as like, you know, if I have a ‑‑ if I'm making coffee and that takes me three minutes to do that, I mean, one thing I can do is try to force my whole, you know, VLA to try to figure out how to do this individual task. Or what I can use, I can have an orchestrator model that breaks up these tasks into sub-tasks, and there's a tactile control policy that's running on top. And there's kind of a full spectrum between, you know, full VLA only all the way to this kind of agentic system. But the main thing I'm trying to highlight is that embodied reasoning is actually a very interesting and complex problem that is yet to be solved. That being said, our models continue to be frontier on embodied reasoning. And because they're frontier, we start seeing really cool things that we haven't seen before. Here's a concrete example of doing very complex egocentric ‑‑ or not egocentric, but this is robotic data annotation. And you can kind of see here the model is jumping around, looking at different portions of the video, you know, clipping it, figuring out whether or not the captions are correct, self‑verifying. And we can all do this because AR models are fast. They're significantly cheaper than anything else that's out there. So if you try to do this with Gemini, this video would probably cost you a couple of dollars, or for us it's probably in the sense. And these all kind of emerged from being able to have these frontier embodied reasoning capabilities that we just previously have not seen from other models. Okay, probably going to share the biggest research breakthrough that we've had, and I think we'll share more of this in the upcoming weeks probably on Twitter. But one thing that we found is we've discovered new scaling laws for embodied foundation models. So these are models, again, that you can jointly do control‑based training, you can do trajectory training, you can do perceptive training, you can do embodied reasoning training. If you just figure out what the right way to mix this all together is and the right objectives to use, you actually start seeing very interesting levers that you maybe previously haven't been able to see before. So the concrete lever that I'll talk about is this ability to trade general video pre‑training data for tele‑op data. So kind of as we know, tele‑op data is very expensive. It's on the orders of, you know, $100 per hour of data. For $100, I can collect significantly more video pre‑training data. And so what this graph is showing is that if you're just training pure VLA's, pure policies, there's this band that you have. So you do actually still have scaling loss. So you do get benefits for more and more tele‑op data. That being said, the benefits are not as substantial as if you are really training these unified embodied foundation models. And so this is what kind of the bottom half of the graph is. And the really cool kind of lever that we get is you can essentially trade 10x less tele‑op data if you have 10x more video pre‑training data. And so far this is kind of held for the amount of compute that our company has. And it will be continuing to kind of push the fold on how far you can push these embodied foundation models. So here's a couple of videos of a policy that hopefully will open source one of the smaller models in a couple of weeks. But this is all running natively within a single model that is capable of doing the embodied reasoning in order to figure out the task. Actually is outputting control tokens. You can see it's a little bit jittery, but that's okay. Hopefully it will be figured out at scale. And you can actually see very complex tasks that previously I think would be really tough for pure VLA's to do. So I think if you look at the right‑hand video, this is requiring the model to actually read the title of the book, have the knowledge about what type of book this is, and then properly allocate it within one of the bins. So this is actually a multi‑step task between perception and control that is really, really tough to do if you have a kind of a pure end‑to‑end control model that is not aware of the different perceptive tasks that it needs to accomplish in order to do this task. Yeah. It's pretty cool. It also works zero‑shot relatively well out of the box. So we're excited to get this in the hands of folks in a couple of weeks, sometime in July. Yeah. Going to leave a couple of minutes for general questions, but if you guys are interested, let's connect. So one cool thing that we do with our company is we actually, for a limited set of partners, give access to our Mark 1 weights. We give access to our larger embodied foundation models weights. So email me, DM me on Twitter, whatever is easier. And then, yeah, I'll open up. There's a couple minutes left for questions. JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF JEFF like the temporal and spatial aspects of these models, because that would have been a challenge for a very long time. And like when I saw on your example from right, it's just like, it's reading the catalog and now into what you're looking at, like it has a lot of content. And then spatial already is obviously clear, and also I see some temporal . So can you talk a little bit about that? Yeah, I think it's, so yeah, so the question is how are we able to nail temporal and understanding to this degree? It's a good question. I mean, to be honest, it's not nailed. So there's still a lot of work to actually get it to a place where you can reliably deploy. The core thing is how do you think about context management? So you have a relatively limited context. So I think the models here have one million contexts, but that's relatively easy to fit in with the high FPS video. And so you have to start thinking about, are there interesting things that you can do? I'll throw something out there. We used to do this, but we got past this. But like, how do I think about like key frames versus delta frames? How can I manage my context by training these two off? And then you start thinking about during your pre-training objective, how can I start kind of natively ingesting things that I think will be useful for the robotics tasks? So, for example, a very basic thing that even kind of the Gemini models used to struggle at, I think the new ones are pretty good, but being able to tell cardinalities. So like left and right is very hard to tell if you do internet scale crawls, because no one on the internet is necessarily labeling things as, you know, this object is to the left of this object, it's below this object. So really thinking about data distributions early on gives you this ability relatively quickly. And it's also like ER, this type of embodied reasoning was a very concrete focus with us, which is why I think we were able to kind of surpass a Gemini ER with relatively less compute. So what issue we, with VLA and VLA, this robustness on products? I was wondering whether, you know, this is really, really cool, I was wondering whether you looked at, are your models that are potentially more robust than, you know, the kind of . Yeah, so probably the coolest robustness that we've seen is that for VLA models specifically, like if you go and you take one of the Chinese ones, and you try to fine tune it for a specific policy, if you just change the background of the, I don't know, even like in the table, if you change the background of the table, the policy will actually fail. What's really interesting, if you do this type of joint perceptive and control modeling, you're much more robust to these types of errors or if the light is hitting it a slightly different way. And I think we primarily view this as robustness to background in a way that I think traditional models don't necessarily have. That being said, I'm not going to over claim, like, I mean, like, it's still relatively hard. I think if I was going to go and shine a flashlight into one of the arms, it's probably not going to work. But being able to jointly model these things helps a significant amount. We also do a lot of online augmentation, so we do actually, you know, fake, I don't know, one of the arm cameras being off, right? We fake sunlight coming in from a certain direction, right? So we do these things during training to improve robustness. But the big gains come from taking this early fusion paradigm and then moving into the robotics domain. Cool. Is it more useful to be a knowledge base using the K1 model? Oh, knowledge bases? It's useful for, like, I mean, if you want to caption images, videos, it's relatively well. We work with robotics partners for, like, very complex egocentric annotation like this. So it's not necessarily building on an ontology, but being able to do kind of very deep structured extraction, I think our models are very good at. By the way, everything that I kind of showed here is kind of public APIs, so you can go play around with it. The benchmarks are public. I think I'm out of time. They're cutting me off, so I can talk with folks outside. But thank you, guys. I think I need to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to find the odds to