The Best Models Still Reason Like Toddlers — Andrew Dai, Elorian
Description
Show a frontier model part of a chessboard, ask how many white squares are visible, and it answers 32. Andrew Dai's diagnosis: it saw a chessboard, knew chessboards have 32 white squares, and hallucinated the rest. The pattern matching that makes these models superb at naming flowers hurts them once a question needs counting or spatial grounding: they miscount a Catan player's roads from the pieces left off the board, and they miss a robot arm lifting a lid because they cannot hold state across a long video. His test for understanding versus reasoning: if a person can answer in one second, so can the model; if it takes longer, the model falls apart. The benchmarks hide this: one popular reasoning suite uses 32 by 32 pixel images, and a multimodal science exam can mostly be answered without the image. What is missing, he says, is visual thinking. Video generators produce cartoonish explosions because their training data is Hollywood and game engines; detection models are robust but passive. Elorian's approach has four parts: visual reasoning data that does not exist online, a synthetic data flywheel of evals, agents, SFT, and RL, architectural changes on top of the transformer, and native visual chain of thought, where the model draws boxes around every hotel before narrowing to the red ones. Dai spent twelve years at Google Brain and DeepMind, first authored the paper that introduced pretraining and fine tuning, and co led GLaM, PaLM 2 pretraining, and Gemini data. He closes on robotics, construction sites where safety rules live in policy text, and mechanical design, where a testing platform takes thousands of engineering hours and frontier models fail on blueprints and CAD. Speaker info: - https://x.com/andrewdai - https://www.linkedin.com/in/andrewdai/ Timestamps: 0:00 - Frontier models versus human visual reasoning 0:56 - Chessboard hallucination: pattern matching that hurts 2:32 - Catan roads, robot arms, and context amnesia 4:09 - Understanding versus reaso
Summary
Generated by gpt-5.6-terraAt-a-Glance
- Verdict: Watch fully
- Core thesis: Frontier multimodal models remain strong at instant visual recognition but are unreliable at deliberate spatial, temporal, and causal reasoning, which blocks dependable deployment in robotics, construction, and engineering design.
- Why it matters: Any agent or control system that uses visual inputs for consequential decisions must treat current vision-language models as fallible recognizers rather than grounded world models, and must add task decomposition, verification, and domain-specific evaluation.
- Best use: Use the talk to calibrate product claims and system design for visual agents, especially where images, video, plans, CAD files, or physical actions are inputs to automated workflows.
Executive Summary
Andrew Dai argues that leading models such as Claude, ChatGPT, and Gemini can recognize familiar visual patterns but break on questions that require sustained inspection, counting, spatial grounding, object permanence, or temporal tracking. His core distinction is between visual understanding, which can often be answered at a glance, and visual reasoning, which requires System 2-style multi-step attention to an image or video.
The practical consequence is that seemingly simple visual prompts can produce confident errors: a partial chessboard is completed into a standard board in the model's imagination, a Catan board is miscounted by inference from nearby pieces, and a robot-video model loses track of consequential state changes such as lifting a lid and turning on a stove. Dai characterizes this as context amnesia across video and over-reliance on pattern matching.
He also challenges common multimodal benchmarks as poor proxies for real-world visual intelligence. ARC-AGI-style low-resolution tasks and MMMU science questions may reward text knowledge or coarse recognition without testing the geometric alignment, spatial relationships, object permanence, and policy-to-scene grounding required in physical environments.
The latter half is Elorium's product thesis: build visual-thinking models using proprietary multimodal reasoning data, synthetic-data flywheels, supervised fine-tuning and RL, architecture changes, and native visual chain-of-thought. The proposed commercial opportunity is an API layer that turns visual streams into action-relevant scene understanding for robot control, construction compliance, CAD/CAM quality assurance, and mechanical design, with simulator-based validation analogous to software unit tests.
Key Takeaways
- Claim: Frontier multimodal models should not be trusted for detailed visual reasoning merely because they identify objects well. | Evidence: Dai's chessboard example asks for the number of white squares in a partial board; models answer 32 by inferring a complete standard chessboard rather than counting what is visible. In a Catan example, models infer five blue roads from off-board pieces when the board contains seven. | Implication: For workflows involving counts, locations, dimensions, inventory, layouts, or compliance checks, require explicit visual verification rather than accepting a model's single-pass natural-language answer. | Caveat: These are illustrative adversarial examples from the speaker rather than a published comparative error-rate study, so they demonstrate a failure mode but do not quantify average performance by model or task.
- Claim: The useful boundary is not image versus video, but fast pattern recognition versus slow, multi-step visual reasoning. | Evidence: The speaker proposes a practical test: if a human can answer within roughly one second—such as identifying a game, flower, or three visible chess pieces—it is recognition; if the task requires inspection of several elements and relationships, it is reasoning. He links the latter to Kahneman's System 2 thinking. | Implication: Classify visual-agent tasks by required inspection and state tracking before selecting a model; reserve higher-risk use cases for pipelines that can inspect, decompose, and validate intermediate visual conclusions.
- Claim: Long-video understanding remains especially weak because models can lose physical state and causal context over time. | Evidence: In a robot manipulation example, state-of-the-art models reportedly miss that the arm lifted a lid and later miss that it is turning on the stove. Dai calls this 'context amnesia' caused by an inability to maintain consistency across long videos. | Implication: Do not let a general video model be the sole state estimator for robots, monitoring systems, or incident workflows; persist event logs and structured state outside the model, and reconcile each new frame against that state. | Caveat: The talk does not provide video lengths, models tested, or an evaluation protocol, so the claim should inform risk assessment rather than serve as a vendor-independent benchmark result.
- Claim: Widely cited multimodal benchmarks can overstate readiness for real-world visual deployment. | Evidence: Dai says ARC-AGI image tasks are commonly only 32×32 or 64×64 pixels, making high scores a weak proxy for complex real-world visual tasks. He similarly argues that many MMMU questions can be answered from text knowledge or coarse image recognition without materially using the image. | Implication: Evaluate visual systems on production-like data and failure modes—geometric alignment, spatial relations, object permanence, temporal continuity, and policy grounding—rather than using general multimodal leaderboard scores as deployment evidence. | Caveat: Benchmark criticism is directionally useful, but the presentation does not offer a replacement benchmark, baseline scores, or evidence that all ARC-AGI/MMMU items lack real-world relevance.
- Claim: Existing computer-vision components are robust for narrow perception but are fundamentally passive and insufficient for complex task reasoning. | Evidence: The talk cites Google Lens for classification, SAM3 for segmentation, YOLO for detection, and Mask R-CNN, describing them as widely deployed pixel-to-semantic-label tools that do not themselves answer higher-level relational or causal questions. | Implication: A dependable visual-agent stack should retain specialized detectors and segmenters as perception tools, then layer structured reasoning, policy interpretation, planning, and deterministic validators above them instead of expecting one model call to do everything.
- Claim: Visual chain-of-thought and executable validation are the proposed path from recognition to useful physical-world reasoning. | Evidence: Elorium's example asks how many red hotels are built in a scene: the model first identifies and boxes all hotels and other objects, then filters to red hotels. For mechanical design, Dai proposes validating generated designs in Siemens-style or other physical simulators, analogous to running unit tests on generated code. | Implication: The durable architecture pattern is inspectable intermediate representations plus external validators: visual grounding should create auditable objects, regions, counts, constraints, or state transitions that a rules engine or simulator can test. | Caveat: This is Elorium's stated approach and roadmap, not a demonstrated production result; the company says it plans an API release by year-end but provides no accuracy, latency, pricing, or integration evidence.
- Claim: The largest near-term opportunity for visual reasoning is in industries where errors have physical, regulatory, or engineering consequences and generic models currently fail. | Evidence: The speaker highlights robotics, construction safety, architecture, CAD/CAM quality control, and mechanical engineering. One mechanical-engineering prospect reportedly needs 100–200 human hours for a small robot-testing-platform component and 2,000–3,000 hours for the full platform; it had tried frontier models but found them error-prone on blueprints and 3D CAD/CAM context. | Implication: Prioritize verticals where visual mistakes are measurable and validation artifacts already exist—plans, regulations, bills of materials, simulation environments, inspection criteria—because those create clearer data and evaluation loops than open-ended image understanding. | Caveat: The labor estimates and customer experience are anecdotal and supplied by the company; they establish potential pain but not proven market-wide economics or technical feasibility.
Detailed Brief
Elorium's proposed model-development stack
- Claims: The company believes visual-reasoning-specific multimodal data cannot be sourced adequately from the open internet and must be collected or generated deliberately.; Its improvement loop combines evaluations, agents, supervised fine-tuning, reinforcement learning, synthetic data, and transformer-based architectural changes.; The intended differentiator is reasoning performed in the visual space rather than only textual chain-of-thought over an image embedding.
- Evidence: Dai presents a four-stage approach: proprietary multimodal data collection/generation, a synthetic-data flywheel using evals plus agents/SFT/RL, architectural improvements on transformer models, and visual chain-of-thought.; The visual chain-of-thought example uses boxes around candidate objects before applying an attribute filter, making the reasoning process spatially grounded and inspectable.
- Caveats: No technical architecture details, benchmark results, training-data scale, or independent comparisons are disclosed.; The assertion that no frontier model has visual chain-of-thought is a broad vendor claim and should not be taken as a settled industry classification.
- Implications: A visual-reasoning vendor should be assessed less on generic VQA demos and more on whether it exposes stable intermediate visual state that can be inspected, corrected, and audited.; Data flywheels should target the specific relations that downstream systems need, such as containment, distance, count, sequence, occlusion, and plan-versus-site conformance.
Deployment model and company positioning
- Claims: Elorium intends to integrate as a scene-understanding API rather than replace robot planning and control systems wholesale.; Construction is framed as a multimodal policy-grounding problem: interpret textual OSHA or site rules, observe a dynamic site, and assess whether the observed scene satisfies the rule.; Engineering design is framed as analogous to coding only when outputs can be checked programmatically or in simulation.
- Evidence: The company says its API is planned for release by the end of the year and would provide action-relevant visual understanding to existing robotics planning and control systems.; For construction, example questions include helmet compliance and whether work matches previously defined plans using current camera infrastructure.; Dai describes a team of roughly 20 and cites his prior work at Google Brain/DeepMind, co-founder Yim Fei's work leading Apple's MM1 research, and chief reasoning architect Dustin Tran's prior post-training role at xAI.
- Caveats: The API timing is prospective and not a shipping commitment substantiated by a product demonstration.; Robot and construction deployment additionally require safety certification, sensor reliability, latency guarantees, and robust handling of ambiguous or occluded scenes; these operational issues are not addressed in the talk.
- Implications: The strongest adoption wedge is likely decision support or exception detection with human review and simulator/rules-based gates, before autonomous physical execution.; For diligence, demand domain-specific evaluations using real camera conditions, explicit false-positive/false-negative costs, state-retention tests across long videos, and evidence that outputs integrate with existing control or compliance systems.
Notable Concepts & Terms
- Visual reasoning vs. visual understanding: The talk's central distinction: recognition is fast classification or coarse identification, while reasoning requires deliberate multi-step examination of objects, relationships, and changes.
- System 2 thinking: Kahneman's term is used to describe the slower, attention-demanding visual tasks on which current models allegedly fail.
- Context amnesia: The speaker's label for video models losing prior state and failing to preserve a consistent account of events over time.
- Object permanence: The ability to retain the identity and state of objects despite temporal progression, movement, or occlusion; presented as a missing requirement for physical-world AI.
- Visual chain-of-thought: A proposed reasoning process that operates through visual intermediate steps—such as bounding, grouping, and filtering objects—rather than only text-based explanations.
- Synthetic data flywheel: Elorium's proposed iterative training loop combining evaluations, agents, supervised fine-tuning, reinforcement learning, and generated visual-reasoning data.
- SAM3 / YOLO / Mask R-CNN: Examples of specialized segmentation and object-detection systems that are valuable for narrow perception but, in the speaker's view, do not independently perform causal or relational reasoning.
- Programmatic or simulation validation: The proposed analogue of software unit testing for physical design: test AI-generated mechanical outputs in a simulator or against formal constraints before relying on them.
Operator Notes / Why Ken Should Care
- Create a visual-agent risk tiering rubric: allow recognition-only tasks to proceed with standard confidence thresholds; require structured state, multi-step grounding, and verification for counts, layout, temporal events, and physical actions.
- For any video-based workflow, store external event and object state rather than relying on a model's implicit memory across frames or calls.
- Build evaluation sets from real operating images, videos, drawings, and policies; include adversarial partial views, occlusions, misleading familiar patterns, object counts, spatial relations, and sequence-of-events tests.
- Make intermediate visual outputs machine-checkable: bounding regions, object IDs, counts, coordinates, inferred constraints, and evidence frames should feed rules engines, human review, or simulators.
- If evaluating Elorium or comparable vendors, request independently reproducible results on CAD, construction, or robotics tasks, including temporal consistency, calibration, latency, failure recovery, and integration with downstream control systems before treating the year-end API roadmap as investable proof.
Source/Metadata
- Title: The Best Models Still Reason Like Toddlers — Andrew Dai, Elorian
- Transcript words: 4598
- Duration seconds: 1096
- Timestamp note: No timestamps or chapters were provided. The transcript includes a substantial repeated passage covering the latter portion of the presentation.
Transcript
Hi, thanks for coming everyone. I'm the co-founder and CEO of Elorium and I'm here to talk about some of the issues with current models, current frontier models, this includes Claude, ChatGPT, and Gemini, and how they handle visual problems. And this might be new to some of you who don't work in the visual space, but actually there's quite a big gap between how these models handle visual reasoning and how humans deal with it. And you will see that we're actually quite far away from any definition of AGI for visual reasoning. So here are some examples of how easy it is to find where models break down. And you can find these examples yourself, just takes a few minutes. In this first example, we have a chessboard hallucination. And we give the models this picture and ask how many white squares are in the image. And any ordinary person who doesn't hallucinate would probably not say it's 32. So 32, of course, the models say this because they see part of the chessboard and they hallucinate the complete board. And as a result, they give the wrong number. And you see this quite a lot that models rely heavily on pattern matching. That's what makes them so good at identifying plants and animals and flowers in the real world. But when it comes to complex questions, that part hurts them. So the pattern matching is actively hurting them in this case. So in there, what's going on is that the main reason is that, oh, this is a chessboard. Chessboards all have 32 squares. Therefore, this one must have 32 white squares too. On the right example, I'm a big board game player, have quite a collection. So you can see there's a board game theme going on here. And actually, you can reproduce this outside if you just go to outside the talks is a board game area. There, there are chessboards. I will bet if any of you place the pieces in some kind of random position, no frontier model will be able to tell you where those pieces are located, where all those pieces are located. And then another example here is Catan. Here, another very simple question. How many rows does the blue player have? These frontier models think extensively about this problem. One response I've seen is that, oh, the guy has 10 blue rows off to the side of the board. Therefore, there must be five blue rows on the Catan board. But obviously, that's not true. There's seven if you actually count. So these models, again, are great at guessing, great at pattern matching. But they are not really very spatially grounded. And they just can't handle any kind of detailed questions. And then finally, we have this example where it actually affects robots, where here you have a robot arm manipulating this cup and cooker, basically. And the state of the art models today, they miss the fact that the robot arm lifted the lid. And at the end, they also miss the fact that the robot is turning on the stove, right now. So there's essentially context amnesia happening. And this is because these models can't maintain consistency across long videos. And they very easily lose track of what's happening. And a very common question I get is, how do you define a visual reasoning problem versus a visual understanding problem? Or you could say visual thinking compared to visual understanding. I think a very simple way to do it is just ask yourself the same question. If you looked at an image or video, how long would it take you to answer the question? So, for example, in both of these cases, I doubt anyone in this room would be able to give an answer if they were only allowed one second to look at the image. So one second isn't enough to do these kind of complex questions, also called system two kind of thinking in Daniel Kahneman's book. But if I asked you, what game is this? Or similarly, what flower is this? Or if there are only three pieces on the chessboard, if I ask you how many pieces are there? Those questions, I'm sure all of you would be able to answer in less than a second. And similarly, all the frontier models would get that kind of question right. So that is the distinction that we make between what is understanding, what is pattern recognition versus what is reasoning, where you actually have to look in detail at the picture and look at various things. And this is exactly where frontier models fall apart today. So as you are designing your own systems, that's something to keep in mind. Keep these visual tasks very simple. Otherwise, you will have hallucinations and a lot of hallucinations. So this leads into evals, of course. Frontier models, there are already a bunch of multimodal reasoning evals or visual reasoning evals. Some that you might have heard of is Arc AGI. This is quite often brought up to people saying, oh, the frontier models are 85% or 90% on Arc AGI, therefore we are 90% of the way to AGI itself. But I think these people, they haven't really looked at any of the benchmark data, because if you actually look at the data, you will notice that the images are only 32 by 32 or 64 by 64 pixels. And I would challenge anyone to give me a real world complex task that can be reduced to a 32 by 32 pixel problem. I think you will very quickly realize almost no tasks, almost no interesting tasks can be reduced to that kind of resolution. Another eval that people commonly bring up is MMMU. This is the massive multi-discipline, multimodal understanding. This is a step up from MMVU because it has images rather than just pure text science questions. This is science questions based on images, but still images are a minor part of a lot of these questions. A lot of the questions you can just answer without looking at the image or just doing some pattern recognition, just knowing roughly what the image is about. So what we really need is new visual reasoning benchmarks in the industry that really target the things that people care about, like geometric alignment, spatial intelligence, object permanence. And these are really critical for AI to be deployed in these visual use cases. And you might have noticed that still in a lot of industries that are primarily visual, there isn't much uptake of AI, right? A lot of the AI uptake has been in the software engineering world and in the mathematician world, documents, document handling, et cetera. But there is actually a huge gap, huge opportunity that is just being looked over right now based on the interest in coding. And so the missing paradigm in visual AI is thinking. So we have generation models, very high quality generation models, made by dance, see dance model. So we have these very high fidelity models and they look great, but they lack actual physical grounding and causal logic. So you will probably notice that if you ask these models to produce a picture, a video of something blowing up, or a building falling down or these things, they look very cartoonish. They look Hollywood style kind of things. And that's because they are just outputting what was in the training data. And a lot of disaster videos, a lot of action kind of videos on the internet are just going to be from Hollywood or game engines. So they're working to reproduce that. And that's fundamentally a problem because that means they can be no better than those kind of videos. Understanding where we currently are is we have a lot of tools that can map pixels to semantic labels, Google Lens. It's obviously great to identify plants and flowers and I use that all the time. There's SAM3 for segmentation, YOLO for object recognition detection, Mask CNN. These are of course highly robust and they're used everywhere in the industry, but they're fundamentally passive. So there's no reasoning capability to them. So they can't answer more complex questions. And really where the frontier is, is with thinking, visual thinking models. These models will have active spatial and temporal intelligence. They can extract causal logic for planning, agentic workflows and physical execution. And so our approach is four stage. So we are collecting and generating our own multimodal data, visual reasoning specific data. This kind of data we found you just can't get online. We have a synthetic data flywheel using evals, agents, SFT and RL to improve the model. We're making some advances to the architecture in terms of various different improvements on top of the transformer based architecture. And we're also enabling visual chain of thought reasoning. And this is one of the key things that humans have that no frontier model has today. Since the frontier models are only textual chain of thought based. And this is one example of a visual chain of thought. So the question is how many red hotels are built in this photo? Then the model realizes, oh, first we need to identify all the hotels. So it draws boxes around hotels and other objects. And then it reduces that to the red hotels. So it's this multi-step process happening in the visual space natively. So about our company, I'm the co-founder and CEO. I spent the last 12 years at Google Brain and DeepMind. I developed a lot of the foundational techniques for the model, for modern LLMs. 11 years ago, I was the first author of the work that introduced pre-training and fine tuning. That's the work when combined with the transformer paper in 2017 led to the GPT series of models. So all the GPT papers cite our paper. I co-led the early MOE models. The first model that was state of the art called GLAM. And then more recently, I co-led the Palm II pre-training in architecture. And I was co-lead for the Gemini data area. And my co-founder, Yim Fei, he was at Apple and Google research. He led research for Apple's first public multi-modal model, MM1. And he has a lot of experience in visual reasoning across language as well. And this is our team. So we're roughly 20 people now. We also have a chief reasoning architect, Dustin Tran. Previously, he was lead of post-training at XAI. And we've hired a world-class team across many other companies like Apple, XAI, DeepMind, Amazon, and so on. And in terms of the use cases that I mentioned, robotics is one primary use case. So robots have really critical bottlenecks performing complex real-time physical actions in these kind of dynamic environments. But existing vision models, you probably realize, are trained from static images and very directed videos. They are not like, they don't have active physical interaction. So existing methods are over-engineered and brittle. And we are planning to release a model API available by the end of this year. And at that point, the API can be used to deliver action-relevant scene understanding into existing planning and control systems for these robots. Another important use case for visual reasoning is construction. So construction sites, they have these very complex zone-specific safety rules. And computer vision can't adapt fast enough to changing safety rules. They also can't interpret things like OSHA policy language and match that language to what's actually going on at the site. Or understand the spatial relationships that are important there. For example, how many of these workers are wearing helmets? Is the construction happening according to the plans that were defined earlier? And currently enforcing these rules require training separate models for different use cases. Because these models, as I said before, are very brittle. So you constantly have to do retraining. And our approach with the video understanding capabilities that we are building into the models allows you to ground these video streams in the safety regulations. And of course the safety regulations are in text. They're in language. So the model has to manipulate both language and vision very well. And this will allow these models to interpret site policies using the current camera infrastructure that they have. And then finally, architecture and design we think is also a very promising use case here. This is exactly the use case where you need to be very detail-oriented. So back to the board game example around counting spatial relationships. This shows up a lot in architecture and design. Like if you design a house with four bedrooms instead of three, that homeowner is going to be very angry. So obviously counting is actually important. And also just understanding these spatial constraints, real-world constraints is a very manual process these days. We spoke to a mechanical engineering company just a few weeks ago and they said to design one small part of a robot testing platform takes 100 to 200 hours of time. To design the entire testing platform, I believe it takes 2,000 to 3,000 hours of human time. And they've tried frontier models, but they just don't work for these use cases. They really struggle to understand visual context across these architecture blueprints, 3D CAD CAM files. And so there are lots of errors there. And similarly, we believe that this can be useful for other kinds of design as well. Not just architecture and engineering, but maybe designing for the web or fashion or other things. And our approach is we're using multimodal reasoning to extract this geometric logic that's important. We're allowing programmatic validation or simulation validation. Just like in code, you can run code against unit tests. You can also run mechanical devices through simulators that have been developed through Siemens and various other companies to see if something will work in the real world. So there's a lot of parallels actually between this kind of mechanical design and coding itself. But mechanical design is still relatively untouched by AI. And CAD CAM quality control is another potential use case. And ultimately, we believe that this is going to be a critical step to the future of mechanical design, where AI can make faster cars, more efficient rockets, better batteries. And all these things cannot be done just with code. People are not coding up the next iPhone or coding up the next SpaceX rocket. It's all fundamentally very visual. So you can find out more about us through our website, Lauren.ai, our Twitter page, x.com slash LaurenAI, or our LinkedIn page. And happy to take any questions. I'll be standing around here for a little bit. Thanks. very spatially grounded. And they just can't handle any kind of detailed questions. And then finally, we have this example where it actually affects robots, where here you have a robot arm manipulating this cup and cooker, basically. And the state of the art models today, they miss the fact that the robot arm lifted the lid. And at the end, they also miss the fact that the robot is turning on the stove, like right now. So there's essentially context amnesia happening. And this is because these models can't maintain consistency across long videos. And they very easily lose track of what's happening. And a very common question I get is, how do you define a visual reasoning problem versus a visual understanding problem? Or you could say, like, visual thinking compared to visual understanding. I think a very simple way to do it is just ask yourself the same question. If you looked at an image or video, how long would it take you to answer the question? So, for example, in both of these cases, I doubt anyone in this room would be able to give an answer if they were only allowed one second to look at the image. So one second isn't enough to do these kind of complex questions, also called system two kind of thinking in Daniel Kananen's book. But if I asked you, what game is this? Or similarly, what flower is this? Or if there are only three pieces on the chessboard, if I ask you how many pieces are there? Those questions, I'm sure all of you would be able to answer in less than a second. And similarly, all the frontier models would get that kind of question right. So that is the distinction that we make between what is understanding, what is like pattern recognition versus what is reasoning, where you actually have to look in detail at the picture and look at various things. And this is exactly where frontier models fall apart today. So as you are designing your own systems, that's something to keep in mind. Keep these visual tasks very simple. Otherwise, you will have hallucinations a lot and a lot of hallucinations. So this leads into evals, of course. Frontier models, there are already a bunch of multimodal reasoning evals or visual reasoning evals. Some that you might have heard of is Arc AGI. This is quite often brought up to people saying, oh, the frontier models are 85% or 90% on Arc AGI, therefore we are 90% of the way to AGI itself. But I think these people, they haven't really looked at any of the benchmark data, because if you actually look at the data, you will notice that the images are only 32 by 32 or 64 by 64 pixels. And I would challenge anyone to give me like a real world complex task that can be reduced to a 32 by 32 pixel problem. I think you will very quickly realize almost no tasks, almost no interesting tasks can be reduced to that kind of resolution. Another eval that people commonly bring up is MMMU. This is the massive multi-discipline, multimodal understanding. This is a step up from MMMU because it has images rather than just pure text science questions. This is science questions based on images, but still images are a minor part of a lot of these questions. A lot of the questions you can just answer without looking at the image or just doing some pattern recognition, just knowing roughly what the image is about. So what we really need is new visual reasoning benchmarks in the industry that really target the things that people care about, like geometric alignment, spatial intelligence, object permanence. And these are really critical for AI to be deployed in these visual use cases. And you might have noticed that still in a lot of industries that are primarily visual, which I will go into, there isn't much uptake of AI, right? A lot of the AI uptake has been in the software engineering world and in the mathematician world, like documents, document handling, et cetera. But there is actually a huge gap, huge opportunity that is just being looked over right now based on the interest in coding. And so the missing paradigm in visual AI is thinking. So we have generation models, very high quality generation models, like by dance, see dance model. So we have these very high fidelity models and they look great, but they lack actual physical grounding and causal logic. So you will, you probably notice that if you ask these models to produce a picture, a video of a, some, like something blowing up, like, or a building falling down or these things, they look very cartoonish. They look Hollywood style kind of things. And that's because they are just outputting what was in the training data. And a lot of disaster videos, a lot of like action kind of videos on the internet are just going to be from Hollywood or game engines. So they're working to reproduce that. And that's fundamentally a problem because that means they can be no better than those kind of videos. Understanding the, what we are, where we currently are is we have a lot of tools that can map pixels to semantic labels, like Google Lens. It's obviously great to identify plants and flowers and I use that all the time. There's SAM3 for segmentation, YOLO for object recognition detection, Mascar CNN. These are of course highly robust and they're used everywhere in the industry, but they're fundamentally passive. So there's no reasoning capability to them. So they can't answer more complex questions. And really where the frontier is, is with thinking, visual thinking models. These models will have active spatial and temporal intelligence. They can extract axonal logic for planning, agentic workflows and physical execution. And so our approach is a four stage. So we are collecting and generating our own multimodal data, visual reasoning specific data. This kind of data we found you just can't get online. We have a synthetic data flywheel using evals, agents, SFT and RL to improve the model. We're making some, we made some advances to the architecture in terms of various different type, various different improvements on top of the transformer based architecture. And we're also enabling visual chain of thought reasoning. And this is one of the key things that humans have that no frontier model has today. Since the frontier models are only textual chain of thought based. And this is one example of a visual chain of thought. So the question is like how many red hotels are built in this photo? Then the model realizes, oh, first we need to identify all the hotels. So it draws boxes around hotels and other objects. And then it reduces that to the red hotels. So it's this multi-step process happening in the visual space natively. So about our company, I'm the co-founder and CEO. I spent the last 12 years at Google Brain and DeepMind. I developed a lot of the foundational techniques for the model, for modern LLMs. 11 years ago, I was the first author of the work that introduced pre-training and fine tuning. That's the work when combined with the transformer paper in 2017 led to the GPT series of models. So all the GPT papers cite our paper. I co-led the early MOE models. The first model that was state of the art called GLAM. And then more recently, I co-led the Palm II pre-training in architecture. And I was co-lead for the Gemini data area. And my co-founder, Yim Fei, he was at Apple and Google research. He led research for Apple's first public multi-modal model, MM1. And he has a lot of experience in visual reasoning across language as well. And this is our team. So we're roughly 20 people now. We also have a chief reasoning architect, Dustin Tran. Previously, he was lead of post-training at XAI. And we've hired a world-class team across many other companies like Apple, XAI, DeepMind, Amazon, and so on. And in terms of the use cases that I mentioned, robotics is one primary use case. So robots have really critical bottlenecks performing complex real-time physical actions in these kind of like dynamic environments. But existing vision models, you probably realize, are trained from static images and very directed videos. They are not like, they don't have active physical interaction. So existing methods are over-engineered and brittle. And we are planning to release a model API available by the end of this year. And at that point, the API can be used to deliver action-relevant scene understanding into existing planning and control systems for these robots. Another important use case for visual reasoning is construction. So construction sites, they have these very complex zone-specific safety rules. And computer vision can't adapt fast enough to changing safety rules. They also can't interpret things like OSHA policy language and match that language to what's actually going on at the site. Or understand the spatial relationships that are important there. For example, like how many of these workers are wearing helmets? Like is the construction happening according to the plans that were defined earlier? And currently enforcing these rules require training separate models for different use cases. Because these models, as I said before, are very brittle. So you constantly have to do retraining. And our approach with the video understanding capabilities that we are building into the models allows you to ground these video streams in the safety regulations. And of course the safety regulations are in text. They're in language. So you have to be – the model has to manipulate both language and vision very well. And yeah, this will allow these models to interpret site policies using the current camera infrastructure that they have. And then finally, architecture and design we think is also a very promising use case here. This is exactly the use case where you need to be very detail-oriented. So back to the board game example around counting spatial relationships. This shows up a lot in architecture and design. Like if you design a – if you design a house with four bedrooms instead of three, that homeowner is going to be very angry, right? So obviously counting is actually important. And also just understanding these spatial constraints, real-world constraints is a very manual process these days. We spoke to a mechanical engineering company just a few weeks ago and they said to design one small part of a robot testing platform takes 100 to 200 hours of the time. To design the entire testing platform, I believe it takes 2,000 to 3,000 hours of human time. And they've – a lot of these places they've tried frontier models, but they just don't work for these use cases. They really struggle to understand visual context across these like architecture blueprints, 3D CAD CAM files. And so there are lots of errors there. And similarly, we believe that this can be useful for other kinds of design as well. Not just architecture and engineering, but maybe like designing for the web or fashion or other things. And our approach is we're using multimodal reasoning to extract this geometric logic that's important. We're allowing programmatic validation or simulation validation. Just like in code, you can run code against unit tests. You can also run mechanical devices through simulators that have been developed through Siemens and various other companies to see if something will work in the real world. So there's a lot of parallels actually between this kind of like mechanical design and coding itself. But mechanical design is still relatively untouched by AI. And yeah, CAD CAM quality control is another potential use case. And ultimately, we believe that this is going to be a critical step to the future of mechanical design, where AI can make faster cars, more efficient rockets, better batteries. And all these things cannot be done just with code. People are not coding up the next iPhone or coding up the next SpaceX rocket. It's all fundamentally very visual. So you can find out more about us through our website, Lauren.ai, our Twitter page, x.com slash LaurenAI, or our LinkedIn page. And yeah, happy to take any questions. I'll be standing around here for a little bit. Thanks.