Hi, thanks for coming everyone. I'm the co-founder and CEO of Elorium and I'm here to talk about some of the issues with current models, current frontier models, this includes Claude, ChatGPT, and Gemini, and how they handle visual problems. And this might be new to some of you who don't work in the visual space, but actually there's quite a big gap between how these models handle visual reasoning and how humans deal with it. And you will see that we're actually quite far away from any definition of AGI for visual reasoning.
So here are some examples of how easy it is to find where models break down. And you can find these examples yourself, just takes a few minutes. In this first example, we have a chessboard hallucination. And we give the models this picture and ask how many white squares are in the image. And any ordinary person who doesn't hallucinate would probably not say it's 32. So 32, of course, the models say this because they see part of the chessboard and they hallucinate the complete board. And as a result, they give the wrong number. And you see this quite a lot that models rely heavily on pattern matching. That's what makes them so good at identifying plants and animals and flowers in the real world. But when it comes to complex questions, that part hurts them. So the pattern matching is actively hurting them in this case. So in there, what's going on is that the main reason is that, oh, this is a chessboard. Chessboards all have 32 squares. Therefore, this one must have 32 white squares too.
On the right example, I'm a big board game player, have quite a collection. So you can see there's a board game theme going on here. And actually, you can reproduce this outside if you just go to outside the talks is a board game area. There, there are chessboards. I will bet if any of you place the pieces in some kind of random position, no frontier model will be able to tell you where those pieces are located, where all those pieces are located.
And then another example here is Catan. Here, another very simple question. How many rows does the blue player have? These frontier models think extensively about this problem. One response I've seen is that, oh, the guy has 10 blue rows off to the side of the board. Therefore, there must be five blue rows on the Catan board. But obviously, that's not true. There's seven if you actually count. So these models, again, are great at guessing, great at pattern matching. But they are not really very spatially grounded. And they just can't handle any kind of detailed questions.
And then finally, we have this example where it actually affects robots, where here you have a robot arm manipulating this cup and cooker, basically. And the state of the art models today, they miss the fact that the robot arm lifted the lid. And at the end, they also miss the fact that the robot is turning on the stove, right now. So there's essentially context amnesia happening. And this is because these models can't maintain consistency across long videos. And they very easily lose track of what's happening.
And a very common question I get is, how do you define a visual reasoning problem versus a visual understanding problem? Or you could say visual thinking compared to visual understanding. I think a very simple way to do it is just ask yourself the same question. If you looked at an image or video, how long would it take you to answer the question? So, for example, in both of these cases, I doubt anyone in this room would be able to give an answer if they were only allowed one second to look at the image. So one second isn't enough to do these kind of complex questions, also called system two kind of thinking in Daniel Kahneman's book. But if I asked you, what game is this? Or similarly, what flower is this? Or if there are only three pieces on the chessboard, if I ask you how many pieces are there? Those questions, I'm sure all of you would be able to answer in less than a second. And similarly, all the frontier models would get that kind of question right. So that is the distinction that we make between what is understanding, what is pattern recognition versus what is reasoning, where you actually have to look in detail at the picture and look at various things. And this is exactly where frontier models fall apart today. So as you are designing your own systems, that's something to keep in mind. Keep these visual tasks very simple. Otherwise, you will have hallucinations and a lot of hallucinations.
So this leads into evals, of course. Frontier models, there are already a bunch of multimodal reasoning evals or visual reasoning evals. Some that you might have heard of is Arc AGI. This is quite often brought up to people saying, oh, the frontier models are 85% or 90% on Arc AGI, therefore we are 90% of the way to AGI itself. But I think these people, they haven't really looked at any of the benchmark data, because if you actually look at the data, you will notice that the images are only 32 by 32 or 64 by 64 pixels. And I would challenge anyone to give me a real world complex task that can be reduced to a 32 by 32 pixel problem. I think you will very quickly realize almost no tasks, almost no interesting tasks can be reduced to that kind of resolution.
Another eval that people commonly bring up is MMMU. This is the massive multi-discipline, multimodal understanding. This is a step up from MMVU because it has images rather than just pure text science questions. This is science questions based on images, but still images are a minor part of a lot of these questions. A lot of the questions you can just answer without looking at the image or just doing some pattern recognition, just knowing roughly what the image is about.
So what we really need is new visual reasoning benchmarks in the industry that really target the things that people care about, like geometric alignment, spatial intelligence, object permanence. And these are really critical for AI to be deployed in these visual use cases. And you might have noticed that still in a lot of industries that are primarily visual, there isn't much uptake of AI, right? A lot of the AI uptake has been in the software engineering world and in the mathematician world, documents, document handling, et cetera. But there is actually a huge gap, huge opportunity that is just being looked over right now based on the interest in coding.
And so the missing paradigm in visual AI is thinking. So we have generation models, very high quality generation models, made by dance, see dance model. So we have these very high fidelity models and they look great, but they lack actual physical grounding and causal logic. So you will probably notice that if you ask these models to produce a picture, a video of something blowing up, or a building falling down or these things, they look very cartoonish. They look Hollywood style kind of things. And that's because they are just outputting what was in the training data. And a lot of disaster videos, a lot of action kind of videos on the internet are just going to be from Hollywood or game engines. So they're working to reproduce that. And that's fundamentally a problem because that means they can be no better than those kind of videos.
Understanding where we currently are is we have a lot of tools that can map pixels to semantic labels, Google Lens. It's obviously great to identify plants and flowers and I use that all the time. There's SAM3 for segmentation, YOLO for object recognition detection, Mask CNN. These are of course highly robust and they're used everywhere in the industry, but they're fundamentally passive. So there's no reasoning capability to them. So they can't answer more complex questions.
And really where the frontier is, is with thinking, visual thinking models. These models will have active spatial and temporal intelligence. They can extract causal logic for planning, agentic workflows and physical execution. And so our approach is four stage. So we are collecting and generating our own multimodal data, visual reasoning specific data. This kind of data we found you just can't get online. We have a synthetic data flywheel using evals, agents, SFT and RL to improve the model. We're making some advances to the architecture in terms of various different improvements on top of the transformer based architecture. And we're also enabling visual chain of thought reasoning. And this is one of the key things that humans have that no frontier model has today. Since the frontier models are only textual chain of thought based.
And this is one example of a visual chain of thought. So the question is how many red hotels are built in this photo? Then the model realizes, oh, first we need to identify all the hotels. So it draws boxes around hotels and other objects. And then it reduces that to the red hotels. So it's this multi-step process happening in the visual space natively.
So about our company, I'm the co-founder and CEO. I spent the last 12 years at Google Brain and DeepMind. I developed a lot of the foundational techniques for the model, for modern LLMs. 11 years ago, I was the first author of the work that introduced pre-training and fine tuning. That's the work when combined with the transformer paper in 2017 led to the GPT series of models. So all the GPT papers cite our paper. I co-led the early MOE models. The first model that was state of the art called GLAM. And then more recently, I co-led the Palm II pre-training in architecture. And I was co-lead for the Gemini data area.
And my co-founder, Yim Fei, he was at Apple and Google research. He led research for Apple's first public multi-modal model, MM1. And he has a lot of experience in visual reasoning across language as well. And this is our team. So we're roughly 20 people now. We also have a chief reasoning architect, Dustin Tran. Previously, he was lead of post-training at XAI. And we've hired a world-class team across many other companies like Apple, XAI, DeepMind, Amazon, and so on.
And in terms of the use cases that I mentioned, robotics is one primary use case. So robots have really critical bottlenecks performing complex real-time physical actions in these kind of dynamic environments. But existing vision models, you probably realize, are trained from static images and very directed videos. They are not like, they don't have active physical interaction. So existing methods are over-engineered and brittle. And we are planning to release a model API available by the end of this year. And at that point, the API can be used to deliver action-relevant scene understanding into existing planning and control systems for these robots.
Another important use case for visual reasoning is construction. So construction sites, they have these very complex zone-specific safety rules. And computer vision can't adapt fast enough to changing safety rules. They also can't interpret things like OSHA policy language and match that language to what's actually going on at the site. Or understand the spatial relationships that are important there. For example, how many of these workers are wearing helmets? Is the construction happening according to the plans that were defined earlier? And currently enforcing these rules require training separate models for different use cases. Because these models, as I said before, are very brittle. So you constantly have to do retraining.
And our approach with the video understanding capabilities that we are building into the models allows you to ground these video streams in the safety regulations. And of course the safety regulations are in text. They're in language. So the model has to manipulate both language and vision very well. And this will allow these models to interpret site policies using the current camera infrastructure that they have.
And then finally, architecture and design we think is also a very promising use case here. This is exactly the use case where you need to be very detail-oriented. So back to the board game example around counting spatial relationships. This shows up a lot in architecture and design. Like if you design a house with four bedrooms instead of three, that homeowner is going to be very angry. So obviously counting is actually important. And also just understanding these spatial constraints, real-world constraints is a very manual process these days.
We spoke to a mechanical engineering company just a few weeks ago and they said to design one small part of a robot testing platform takes 100 to 200 hours of time. To design the entire testing platform, I believe it takes 2,000 to 3,000 hours of human time. And they've tried frontier models, but they just don't work for these use cases. They really struggle to understand visual context across these architecture blueprints, 3D CAD CAM files. And so there are lots of errors there.
And similarly, we believe that this can be useful for other kinds of design as well. Not just architecture and engineering, but maybe designing for the web or fashion or other things. And our approach is we're using multimodal reasoning to extract this geometric logic that's important. We're allowing programmatic validation or simulation validation. Just like in code, you can run code against unit tests. You can also run mechanical devices through simulators that have been developed through Siemens and various other companies to see if something will work in the real world. So there's a lot of parallels actually between this kind of mechanical design and coding itself. But mechanical design is still relatively untouched by AI. And CAD CAM quality control is another potential use case.
And ultimately, we believe that this is going to be a critical step to the future of mechanical design, where AI can make faster cars, more efficient rockets, better batteries. And all these things cannot be done just with code. People are not coding up the next iPhone or coding up the next SpaceX rocket. It's all fundamentally very visual. So you can find out more about us through our website, Lauren.ai, our Twitter page, x.com slash LaurenAI, or our LinkedIn page. And happy to take any questions. I'll be standing around here for a little bit. Thanks. very spatially grounded. And they just can't handle any kind of detailed questions. And then finally,
we have this example where it actually affects robots, where here you have a robot arm manipulating this cup and cooker, basically. And the state of the art models today, they miss the fact that the robot arm lifted the lid. And at the end, they also miss the fact that the robot is turning on the stove, like right now. So there's essentially context amnesia happening. And this is because these models can't maintain consistency across long videos. And they very easily lose track of what's happening. And a very common question I get is, how do you define a visual reasoning problem versus a visual
understanding problem? Or you could say, like, visual thinking compared to visual understanding. I think a very simple way to do it is just ask yourself the same question. If you looked at an image or video, how long would it take you to answer the question? So, for example, in both of these cases, I doubt anyone in this room would be able to give an answer if they were only allowed one second to look at the image. So one second isn't enough to do these kind of complex questions, also called system two kind of thinking in Daniel Kananen's book. But if I asked you, what game is this? Or similarly, what flower is this? Or if there are only three pieces on the chessboard,
if I ask you how many pieces are there? Those questions, I'm sure all of you would be able to answer in less than a second. And similarly, all the frontier models would get that kind of question right. So that is the distinction that we make between what is understanding, what is like pattern recognition versus what is reasoning, where you actually have to look in detail at the picture and look at various things. And this is exactly where frontier models fall apart today. So as you are designing your own systems, that's something to keep in mind. Keep these visual tasks very simple. Otherwise, you will have hallucinations a lot and a lot of hallucinations.
So this leads into evals, of course. Frontier models, there are already a bunch of multimodal reasoning evals or visual reasoning evals. Some that you might have heard of is Arc AGI. This is quite often brought up to people saying, oh, the frontier models are 85% or 90% on Arc AGI, therefore we are 90% of the way to AGI itself. But I think these people, they haven't really looked at any of the benchmark data, because if you actually look at the data, you will notice that the images are only 32 by 32 or 64 by 64 pixels. And I would challenge anyone to give me like a real world complex task that can be reduced to a 32 by 32 pixel problem.
I think you will very quickly realize almost no tasks, almost no interesting tasks can be reduced to that kind of resolution. Another eval that people commonly bring up is MMMU. This is the massive multi-discipline, multimodal understanding. This is a step up from MMMU because it has images rather than just pure text science questions. This is science questions based on images, but still images are a minor part of a lot of these questions. A lot of the questions you can just answer without looking at the image or just doing some pattern recognition, just knowing roughly what the image is about.
So what we really need is new visual reasoning benchmarks in the industry that really target the things that people care about, like geometric alignment, spatial intelligence, object permanence. And these are really critical for AI to be deployed in these visual use cases. And you might have noticed that still in a lot of industries that are primarily visual, which I will go into, there isn't much uptake of AI, right? A lot of the AI uptake has been in the software engineering world and in the mathematician world, like documents, document handling, et cetera.
But there is actually a huge gap, huge opportunity that is just being looked over right now based on the interest in coding. And so the missing paradigm in visual AI is thinking. So we have generation models, very high quality generation models, like by dance, see dance model. So we have these very high fidelity models and they look great, but they lack actual physical grounding and causal logic. So you will, you probably notice that if you ask these models to produce a picture, a video of a, some, like something blowing up, like, or a building falling down or these things, they look very cartoonish. They look Hollywood style kind of things.
And that's because they are just outputting what was in the training data. And a lot of disaster videos, a lot of like action kind of videos on the internet are just going to be from Hollywood or game engines. So they're working to reproduce that. And that's fundamentally a problem because that means they can be no better than those kind of videos. Understanding the, what we are, where we currently are is we have a lot of tools that can map pixels to semantic labels, like Google Lens. It's obviously great to identify plants and flowers and I use that all the time. There's SAM3 for segmentation, YOLO for object recognition detection, Mascar CNN.
These are of course highly robust and they're used everywhere in the industry, but they're fundamentally passive. So there's no reasoning capability to them. So they can't answer more complex questions. And really where the frontier is, is with thinking, visual thinking models. These models will have active spatial and temporal intelligence. They can extract axonal logic for planning, agentic workflows and physical execution. And so our approach is a four stage. So we are collecting and generating our own multimodal data, visual reasoning specific data. This kind of data we found you just can't get online.
We have a synthetic data flywheel using evals, agents, SFT and RL to improve the model. We're making some, we made some advances to the architecture in terms of various different type, various different improvements on top of the transformer based architecture. And we're also enabling visual chain of thought reasoning. And this is one of the key things that humans have that no frontier model has today. Since the frontier models are only textual chain of thought based. And this is one example of a visual chain of thought. So the question is like how many red hotels are built in this photo? Then the model realizes, oh, first we need to identify all the hotels.
So it draws boxes around hotels and other objects. And then it reduces that to the red hotels. So it's this multi-step process happening in the visual space natively. So about our company, I'm the co-founder and CEO. I spent the last 12 years at Google Brain and DeepMind. I developed a lot of the foundational techniques for the model, for modern LLMs. 11 years ago, I was the first author of the work that introduced pre-training and fine tuning. That's the work when combined with the transformer paper in 2017 led to the GPT series of models. So all the GPT papers cite our paper. I co-led the early MOE models. The first model that was state of the art called GLAM.
And then more recently, I co-led the Palm II pre-training in architecture. And I was co-lead for the Gemini data area. And my co-founder, Yim Fei, he was at Apple and Google research. He led research for Apple's first public multi-modal model, MM1. And he has a lot of experience in visual reasoning across language as well. And this is our team. So we're roughly 20 people now. We also have a chief reasoning architect, Dustin Tran. Previously, he was lead of post-training at XAI. And we've hired a world-class team across many other companies like Apple, XAI, DeepMind, Amazon, and so on. And in terms of the use cases that I mentioned, robotics is one primary use case.
So robots have really critical bottlenecks performing complex real-time physical actions in these kind of like dynamic environments. But existing vision models, you probably realize, are trained from static images and very directed videos. They are not like, they don't have active physical interaction. So existing methods are over-engineered and brittle. And we are planning to release a model API available by the end of this year. And at that point, the API can be used to deliver action-relevant scene understanding into existing planning and control systems for these robots. Another important use case for visual reasoning is construction.
So construction sites, they have these very complex zone-specific safety rules. And computer vision can't adapt fast enough to changing safety rules. They also can't interpret things like OSHA policy language and match that language to what's actually going on at the site. Or understand the spatial relationships that are important there. For example, like how many of these workers are wearing helmets? Like is the construction happening according to the plans that were defined earlier? And currently enforcing these rules require training separate models for different use cases. Because these models, as I said before, are very brittle.
So you constantly have to do retraining. And our approach with the video understanding capabilities that we are building into the models allows you to ground these video streams in the safety regulations. And of course the safety regulations are in text. They're in language. So you have to be – the model has to manipulate both language and vision very well. And yeah, this will allow these models to interpret site policies using the current camera infrastructure that they have.
And then finally, architecture and design we think is also a very promising use case here. This is exactly the use case where you need to be very detail-oriented. So back to the board game example around counting spatial relationships. This shows up a lot in architecture and design. Like if you design a – if you design a house with four bedrooms instead of three, that homeowner is going to be very angry, right? So obviously counting is actually important. And also just understanding these spatial constraints, real-world constraints is a very manual process these days.
We spoke to a mechanical engineering company just a few weeks ago and they said to design one small part of a robot testing platform takes 100 to 200 hours of the time. To design the entire testing platform, I believe it takes 2,000 to 3,000 hours of human time. And they've – a lot of these places they've tried frontier models, but they just don't work for these use cases. They really struggle to understand visual context across these like architecture blueprints, 3D CAD CAM files. And so there are lots of errors there. And similarly, we believe that this can be useful for other kinds of design as well.
Not just architecture and engineering, but maybe like designing for the web or fashion or other things. And our approach is we're using multimodal reasoning to extract this geometric logic that's important. We're allowing programmatic validation or simulation validation. Just like in code, you can run code against unit tests. You can also run mechanical devices through simulators that have been developed through Siemens and various other companies to see if something will work in the real world. So there's a lot of parallels actually between this kind of like mechanical design and coding itself. But mechanical design is still relatively untouched by AI.
And yeah, CAD CAM quality control is another potential use case. And ultimately, we believe that this is going to be a critical step to the future of mechanical design, where AI can make faster cars, more efficient rockets, better batteries. And all these things cannot be done just with code. People are not coding up the next iPhone or coding up the next SpaceX rocket. It's all fundamentally very visual.
So you can find out more about us through our website, Lauren.ai, our Twitter page, x.com slash LaurenAI, or our LinkedIn page. And yeah, happy to take any questions. I'll be standing around here for a little bit. Thanks.