AI Engineer

FLUX, Open Research, and the Future of Visual AI — Stephen Batifol, Black Forest Labs

621 summary words 3 min summary Watch video

Start with the signal

3 min read

Summary

FLUX, Open Research, and the Future of Visual AI

Main Topics

  • Black Forest Labs (BFL) Overview: Company behind Stable Diffusion, Latent Diffusion, and Flux models
  • Flux Model Evolution: From Flux 1 through Flux 2 Klein, with progression from text-to-image to multi-modal capabilities
  • CellFlow Research: Novel approach to training multi-modal generative models without external encoders
  • Real-Time Generation: Klein model achieving near-instantaneous image generation and editing
  • Future Direction: Movement toward world models, robotics, and physical AI

Key Points

Product Releases & Timeline

  • Flux 1 (August 2024): Breakthrough text-to-image model, open-sourced, fastest at the time (~7-8 seconds), optimized for running on laptops
  • Flux Context: First open-source editing model combining text-to-image and image editing capabilities
  • Flux 2 (November): Multi-reference model with state-of-the-art image generation and editing; supports up to 10 simultaneous images
  • Flux 2 Klein (January): Real-time generation at 300ms and editing at 500ms

Technical Innovation: CellFlow

Problem with Traditional Approach:

  • Generative models don't inherently understand physical constraints (e.g., objects shouldn't pass through each other)
  • Current solution uses external encoders (DINO V2/V3) for "representation alignment"
  • Limitations: scaling ceiling, modality-specific encoders, misaligned objectives, unpredictable performance improvements

CellFlow Solution:

  • Self-supervised learning approach eliminating external encoders
  • Uses dual-noise strategy: high-noise student model and low-noise teacher model
  • Combines representation learning and generation in single flow
  • Enables seamless multi-modal training (images, video, audio, actions)

CellFlow Results:

  • Superior performance across all modalities compared to baseline flow matching
  • Faster convergence without plateauing
  • Improved text rendering accuracy in images
  • Better anatomy and physics in generated content
  • Eliminates visual artifacts like flickering in video

Practical Applications Demonstrated

  • Image Generation: Photo-realistic output with correct hand anatomy, veins, and accessories
  • Image Editing: Multi-image outfit composition, e-commerce visualization
  • Video Generation: Smooth motion with proper physics (e.g., correct pushup form)
  • Audio-Video Synthesis: Synchronized speech generation with video
  • Robotics/Actions: Robot arm manipulation with accurate object interaction

Klein Model Performance

  • Speed: 0.5 seconds for editing, 300ms for generation (vs. 15-20 seconds for competitors)
  • Quality: On-par or superior to larger models while being significantly faster
  • Real-time Applications: Enables interactive mock-up rendering at thinking speed

Notable Quotes

> "Our first operating principle is to release state-of-the-art models... we want to raise the bar on quality with every release we do."

> "We are a research company first, we publish things in the open, we publish papers, we really want to make sure that the field is moving forward with us."

> "Real-time generation... you render mock-ups as fast as you think. You don't have to wait 10 seconds, 20, a minute or two."

> "The reason is robots. That's why we care. That's why robotics and automation. This is where we're taking BFL."

Takeaways

Technical Advancements

  • CellFlow represents paradigm shift in multi-modal generative model training
  • Eliminates dependency on external specialized encoders
  • Enables unified training across diverse modalities simultaneously

Company Direction

  • Focus on state-of-the-art models with open research publication
  • Movement toward visual intelligence and real-time generation
  • Strategic pivot toward robotics and physical AI applications
  • Partnership model with major enterprises (Microsoft, Adobe, Canva, Mistral)

Future Vision

  • World Models: Training models to understand and simulate geometry, relationships, and physics
  • Interactive AI: Real-time visual engines for gaming and film
  • Automation: Application to safe autonomous driving and manufacturing
  • Scale: Leveraging generative worlds to train agents at scale

Key Implications

  • Real-time interactive design tools becoming feasible
  • Unified approach reducing complexity of multi-modal AI systems
  • Movement from pure generation toward embodied AI and robotics
  • Open research approach democratizing access to state-of-the-art models
Full transcript 4041 words · 32 min read
0:14

SPEAKER_01

Thank you for coming today. I really appreciate it. Thank you for coming to this talk which is Black Forest Labs, Flux, Open Research and the Future of Visual AI. I'm going to start quickly with a quick intro of myself. I'm Stefan Batifoul, sorry. I'm a Developer Relations Engineer at BFL and I want to start with two questions. First, who here knows about BFL? Raise your hand. Okay, who here knows Flux? About the same people actually but for the people that don't know BFL, I have a quick intro so you're not lost but BFL at a glance we are the team behind stable diffusion, latent diffusion and the Flux models as well. Our team has more than 200,000 academic citations and we don't only build models, we actually also work with enterprises and customers. So some of our customers are Microsoft, Adobe, Canva, Mistral and many more. And the way we started is we started in August 2024 with Flux 1. Flux 1 was the first breakthrough, that was the model that was the big competitor to stable diffusion back then and there was really the breakthrough where people were like, oh this is a really cool model. We released it in open source in the first place. So really that was the one that was text to image only and you could run it on your laptop. That was a game changer as well and the anatomy was really good in comparison to the other models and especially in comparison to other models that were way bigger. So this is where we really had a breakthrough and Clem from Hugging Face actually gave us a shout out back then. This is fairly old but Flux was actually the model that was the most liked on Hugging Face back then. This is not true anymore but back then that was really the thing and that was really good and really big actually for a company that was coming out of nowhere and just released this model. We then released Flux Context which was the first open source editing model in the world. That was the combination of text to image and image editing as well. This one now what I'm showing you here, it's obvious because now we have editing models everywhere but back then there was a big breakthrough where you could do both text to image and image generation at the same time. If I have an example here we have this input image and then you remove the snowflake from the face you know you can see you have the character consistency but you can also then move this person to be in Freiburg which is where our headquarters, she's chilling taking a selfie in the streets of Freiburg and then you can do some local editing where you can change the background to be snowy and then you also have snow on her face and everything. It was also the model that was really one of the fastest back then. If you remember this is the time where you had the first GPT image where it would take like 40-50 seconds to generate or edit images whereas Context if I remember correctly was like seven to eight seconds. It was also really useful to tell stories. I've seen a lot of use cases from our partners, from our customers where they would start with an image and then create a storyboard like we see here. We have the famous seagull here which has the VR headsets drinking a beer in a bar but then you can actually create other things. You can have a friend that is joining and that is them drinking with them. Now I guess they got a bit tipsy and they're wearing hats in the bar then they're going outside and then this is a story you could create and that was really useful actually for video model or for animation models. You would give those images as input frames or as end frames and then the video model could then create different content.

0:19

SPEAKER_01

In November we released Flux 2 which is our steps towards what we call visual intelligence. Flux 2 was still our best model. Those are samples which I don't know if you can see clearly but in my opinion they are really amazing samples and it's impossible to tell they are AI generated. If you look at the hands, if you look at the veins, if you look at the bracelets of the person on the left there's personally no way that it would tell it's AI generated. Same for the turtles you see on the right side, same for the dog or cat in the bath. That would be very hard to get the sample like this but those are AI generated.

0:25

SPEAKER_01

And then you have more as well so it's not only the people or animals you could do proper product photography. You can see it with a waffle here on the bottom right or you can make some very cute images like this person on the left that is driving the mopeds with some balloons and this is what we release in November but it's not only an image generation it's also an image editing model at the same time. You can see on the left we have six images that we give to the model and my prompt was literally create an outfit with those images and then the model is intelligent enough to actually make things that make sense like the jacket is worn properly same for the tie and on the right side it's a bit more of a simple use case where you have the sofa and then you have to imagine maybe you're an e-commerce website or you're a sofa maker and then you want people to imagine what it looks like in your flat or what it really would look like if you were to buy it. And those use cases are really important and those are the main use cases we have currently for Flux 2. But it also takes up to 10 images simultaneously so you can really edit a lot of images at the same time and then you can create magic things. It's very good at character product and style consistency.

0:30

SPEAKER_01

And what I want to make clear is that BFL as a company and as a research lab our first operating principle is to release state-of-the-art models. This is what we want to focus on, this is what we want to do as a company. We want to raise the bar on quality with every release we do. So we did it in the past with Flux 1 when it came out, we did it with context. Flux 2 was our best image model to date. It's the first one we released that was actually multi-reference as well. It was state-of-the-art in the open source world, it was state-of-the-art for text to image and image editing. In January we released Flux 2 Klein which is a step towards interactive editing and interactive generation. It generates and edits images in less than a second. I'll talk a bit more about it later on during the talk but I think the fastest it can do if I remember correctly is 500 milliseconds for editing and 300 milliseconds for generation. So basically real time. But this is not it, we also have more things that are coming and this is where I want to talk about today. So I mentioned it, we are a research company first, we publish things in the open, we publish paper, we really want to make sure that the field is moving forward with us. This is our big focus as well, so it's state-of-the-art model, publishing things in the open and that's what we want to do. But I want to first take a step back and tell you a

0:37

SPEAKER_01

talk but I think the fastest it can do if I remember correctly it's 500 milliseconds for editing and 300 milliseconds for generation. So real time. But this is not it, we also have more things that are coming and this is where I want to talk about today. So I mentioned it, we are a research company first, we publish things in the open, we publish papers, we really want to make sure that the field is moving forward with us. This is our big focus as well, so it's state-of-the-art model, publishing things in the open and that's what we want to do. But I want to first take a step back and tell you a bit about how do you train models and especially models that are generating content, generating images and everything. When they generate things, when you train them they actually don't understand what they're generating. They don't understand that my glass here should be actually on this table, I shouldn't go through it. Because you train them, you have images and then you add some random noise to those images and then you just try to denoise them. That's what you do, that's what those models are doing and when you denoise images you never learn that my glass shouldn't go through here, you never learn that you're sitting on the chair you shouldn't go through it. So what you do is that you use what is called representation alignment. So you use an external model that actually knows about this and that is an encoder that is an image encoder that is teaching our model, hey he is currently sitting on the chair he shouldn't go through it. And those models are external and they are trained to segment images whereas our models are trained to generate images or generate videos or generate audio. And you try to align them to be on the same objective so that our generative model actually learns okay you shouldn't go through the chair, my glass should stay on this table. And this is great because it really improves the way generative models are working. We can see here on the right it is 70 times faster to actually converge and to reduce the loss when you use this external alignment. So you're okay this is great but as usual if something is working well there are also counterparts to it. So the first one is that you have a scaling ceiling. You imagine you're working with a model that is external that has been trained it's a checkpoint you're not changing it anymore. What if you train a new model and you have a generative model that you want to scale up? You're still limited by this encoder that you have on the side. You're never actually scaling up fully with it. Also those are specialized in modalities. You have an encoder for example DINO V2 and the other one that I can't remember. It's specialized in images only. What if you want your model to generate images, audio, video and more? You would have to have encoders for all of those and you can imagine then you would have a very Frankenstein setup, nothing would really make sense. And the objectives also misaligned so I've shared it before. We want to generate content. We want to generate images or audio. The other one is here to segment things. So they have different objectives and you're trying to make them work together and it works great but it's also not perfect. Here you see on the right side we have DINO V2 and DINO V3. DINO V3 is a better model technically than DINO V2 but when you train your model you're actually getting worse performances. DINO V3 is here in red and green and so you're getting worse performances. So you're okay this is supposed to be a better model and yet when I do train a model to generate things then it gets worse. And there's also not really any rules as to why certain encoders should work or otherwise shouldn't. So how can we solve this? How can you teach a model representation directly without this external encoder? This is what we released about a month and a half ago now which is a research paper. You can read it. It's called Cellflow. It's in the open. We released it to make sure we're moving the field forward again and it's not only us benefiting from it. And it's basically a scalable approach to training multi-modal generative models. So they use self-supervised learning so you don't need any other models to train it. And I'm going to try to go into a tiny bit more details into it. But we combine representation learning and generation in the same flow. So you see here on the left you have videos, images, audio. You have different modalities. What do you do when you usually train a model? You add some noise, you add random noise, you try to denoise it and then you align it with the encoder. How do we do it then? We actually add two different kinds of noises that are both random and they're both different. The first one we're adding is we're adding a lot of noise to the asset. So this is the one you see at the top. And the other one we're adding is a low amount of noise. This is what you see at the bottom. And the idea is that then we have two models that are actually working together. We have the student one which is always getting the images with the most noise and is trying to denoise them. And then the teacher one which is basically a more stable version of the student is always getting the low noise images. And then the student one is actually trying to learn two things at the same time. It's trying to minimize the loss for the generation and the loss in representation. And this is how then you actually work across different modalities. This is then you only have one model. You don't have anything external. And if you actually scale up your model then you're scaling up your student. You're scaling up your teacher. And you don't have to worry about the encoder that you have on the side anymore. And this is where we're working on. This is something we are currently using for different models that we're training. And this is where we believe the future is going to be and to get rid of those encoders that we have. We actually train models. So those disclaimers these are research models. They're not meant to be released in production. But we released actually one model on all those modalities. On the left we're comparing flow matching which is the usual way of training models with ours. And you can see we are better in audio. So this is what you see in orange on the right side. And then we're also better in images. So the dashed lines is the baseline. And then we are the full line where we can see we are also better at images. And also better at video. So with this approach without having the encoder and the external model that you may struggle with. You actually get better at every modality that you're training your model on. It's also converging faster. You can see on the right. The baseline is converging. It's actually hitting a plateau. Whereas we are converging faster and we're still decreasing the loss. And I'm pretty sure that if we were to go towards two million steps.

0:46

SPEAKER_01

And then we're also better in images. So the dashed lines is the baseline. And then we are the full line where we can see we are also better at images. And also better at video. So with this approach without having the encoder and the external model that you may struggle with, you actually get better at every modality that you're training your model on. It's also converging faster. You can see on the right. The baseline is converging. It's actually hitting a plateau. Whereas we are converging faster and we're still decreasing the loss. And I'm pretty sure that if we were to go towards two million steps, the baseline would really plateau and then wouldn't really get any better. Maybe actually get worse. Whereas we would still go down in loss. And this is the difference between the two. If you use Flux in the past or if you use different models to generate, you may have noticed the text might not be perfect. Or things don't really make sense. This is what you see at the top. Where on the left it's the Flux. But you can see you have some letters that are missing. Or maybe you have two letters. Like on "worlds" for example you have two L's instead of one. Whereas with this approach now in the Cel-Flow approach you can see at the bottom everything makes sense. They learn representation. They learn that Flux then the letters should be like one next to the other. And the same on the mirror. Same on the tree. And this is where we believe this is the future. But on top of this we can see some comparisons here. On the left is a baseline. Where again the letters are wrong. And on the right you can see that the letters are correct. Here is the same for the anatomy. Where you see on the left you have a face. It's looking a bit odd. Let's put it that way. And on the right this is the one with Cel-Flow. And again this is not a production model. Where you expect the face to be perfect. But you can see that the anatomy is way better than what you have on the left. But I want to show you as well some different generation. If it loads. Yes thank you. So this is also possible. This is also possible for video generation. So this is the same model that has been trained on images. Now also can generate videos. On the left you see the baseline. It's a weird way to do a push up. Let's put it that way. Whereas on the right it's a perfect form. The arms are correct. The hair as well is correct. And nothing is wrong in it. And this is a way to actually fix all those artifacts that you may see usually in generations. It's the same here for the birds. My bird. Yes thank you. It's the same here for the birds where you see on the left side you have the baseline. There's a lot of flickering. There's a lot of weird things happening. Because the model was using this encoder was trying to align things. Whereas on the right side with Cel-Flow it just does it perfectly. And the bird is walking on the floor and there's no flickering or anything. But it's not only about images or videos or audio. You train those jointly. So you can also actually generate things jointly. We have here an example of a video and audio sample. Where the idea is that we have someone that is saying hello from the black forest. I will just play them and you will hear the difference. Again this is not a production ready model. So it's not perfect. But you can hear the difference hopefully. Hello from the black forest. Hello from the black forest. Hello from the black forest. So this one was the baseline. Where if you hear it correctly, if you try to pay attention to what he's saying, you hear "hello from the black forest." There's a bit of weird things at the end. From the black forest. Hello from the black forest. Whereas on the right side you can see the prompt is really just "say hello from the black forest." And then it ends here. And this is the same model that was trained on those images that we've seen before on video. And on video and audio. All from the black forest. Thank you. But this is cool and this is great. But what if you could also teach robots on how to use this. This is also the same model. This one is trained on actions. And not only on images, video or audio. So it can also predict actions. And what I'm going to show you now is a robot that is trying to pick up a can. And make it closer to us. On the left, this is a baseline. Again you see some flickering. You see the arm is doing weird things. Whereas on the right for the same amount of steps, you can see Cel-Flow. The robot is picking up the arm directly. And bringing it closer. And this is where we're going as well as the company. This is where we're really interested. It's not only image generation or video, but it's also doing actions. And doing more things towards physical AI. And there is more. It's also how do we make our models faster. Because this is really important for us. This is a demo of Klein, which is near real-time editing. You see it on the right side. This is generated with Klein on Korea. Where you see the edits. And this is not a video model. These are images that are always editing in real-time. Not only they are faster. They also actually are at least on par or better than other models. And I'm almost out of time. They added five minutes. So I don't know if I'm. Okay cool. So I'm not out of time. So I can relax. So yes here on the left, you can see we have Klein that is 4b and 9b. That is compared to the other open source models. So it's at least on par. While the latency is around 0.5 seconds, Klein is around 15 seconds. And if you are on par and you're way faster, then this is really good for us. Same for image to image. You can see the editing. For Klein 9b we are at a tiny bit more than 0.5 seconds. Whereas Klein is still at around 15 seconds. And then same for multi-ref. You're at it. But we're still at less than a second. Whereas Klein is more towards 20 seconds. And this is what is really important for us. Because you really want to actually generate things in real-time. This is where we believe there's also visual intelligence. This is where we're going as a company in the future. And why does it matter? I mentioned it. Real-time generation. So you can imagine you render mock-ups as fast as you think. You don't have to wait 10 seconds, 20, a minute or two. You do things in real-time. And you can guide them in real-time. This is also where we're going.

0:53

SPEAKER_01

And this is what is really important for us. Because you really want to actually generate things in real-time. This is where we believe there is visual intelligence. This is where we're going as a company in the future. And why does it matter? As I mentioned it. Real-time generation. So you can imagine you render mock-ups as fast as you think. You don't have to wait. You don't have to wait 10 seconds, 20, a minute or two. You do things in real-time. And you can guide them in real-time. This is also where we're going. You can think of interactive visual engines for gaming or films. Where you really render a movie as you prompt it. On top of this, there's also world models.

1:49

SPEAKER_01

The idea of world models and behind it. And why it matters for us. It's you train your models to understand and simulate geometry. The relationship and different interactions of the world. And you may be okay that's cool. From a research perspective. Why do we care? The reason is robots. That's why we care. That's why robotics and automation. This is where we're taking BFL. And that's why we also want to go towards world models. Is to train agents in those generative worlds. To scale safe driving. And automate every manufacturing. And I think that is it. Thank you very much.

3:01

SPEAKER_01

Do we take questions? Yeah. We can take questions. Can you share something where you say you're training on portable? You mean the data? Yes. Can't really. This is trade secret. Data is very sensitive as you can imagine. So I can't really share this. We're partnering with a lot of people though for it. How do you store the state of the world in those action prediction models? Well this is what the model is learning. It's basically those representations.

4:12

SPEAKER_01

The model is learning that in itself. As the state. And it has some kind of memory. And this is the way we do it. What is that memory? Is it in the context window? Or is it external in the galaxy? No, it's in the context window that you have. You train. And then you have the tokens. And then they'll be like, oh look. I've moved.

5:25

SPEAKER_01

Here is where I should be next. Can you then run it for long? Or does it do some convection? Define long. What do you call long? Indefinitely.

6:03

SPEAKER_01

Indefinitely. That I'm not sure. There's always going to be a limit. So you may have a sliding window. But this is the way we do it. Thank you. Thank you. It generates and edits images in less than a second. I'll talk a bit more about it later on during the talk but I think the fastest it can do if I remember correctly it's 500 milliseconds for editing and 300 milliseconds for generation. So basically real time. But this is not it, we also have more things that are coming and this is where I want to talk about today. So I mentioned it, we are a research company

7:13

SPEAKER_01

first, we publish things in the open, we publish paper, we really want to make sure that the field is moving forward with us. This is our big focus as well, so it's state-of-the-art model, publishing things in the open and that's what we want to do. But I want to first take a step back and tell you a bit about how do you train models and especially models that are generating content, generating images and everything. When they generate things, when you train them they actually don't understand what they're generating. You know they don't understand that my glass here should be actually on this table,

7:50

SPEAKER_01

I shouldn't go through it. Because you train them, you have images and then you know you add some random noise to those images and then you just try to denoise them. That's what you do, that's what those models are doing and when you denoise images you never learn you know that my glass shouldn't go through here, you never learn that you know you're sitting on the chair you shouldn't go through it. So what do you do is that you use, you do what is called like representation alignment. So you use an external model that actually knows about this and that is an encoder that is like an image encoder

8:23

SPEAKER_01

that is teaching our model, hey he is currently sitting on the chair he shouldn't go through it. And those models are external and they are really like trained to segment images whereas our models are trained to generate images or generate videos or generate audio. And you try to align them to be on the same objective so that our generative model actually learns okay you shouldn't go through the chair, my glass should stay on this table. And this is great because it really improves the way generative models are working. We can see here on the right you know it is 70 times faster to actually converge and to

9:00

SPEAKER_01

reduce the loss when you use this external alignment. So you're like okay this is great but as usual if something is working well they are also counterparts to it. So the first one is that you have a scaling ceiling. You imagine you're working with a model that is external that has been trained it's a checkpoint you're not changing it anymore. What if you train a new model and you have a generative model that you want to scale up? You're still like limited by this encoder that you have on the side you know. You're never actually scaling up fully with it. Also those are specialized in modalities. You have an

9:38

SPEAKER_01

encoder for example Dyno V2 and the other one that I can't remember. It's specialized in images only. What if you want your model to generate images, audio, video and more? You would have to have encoders for all of those and you can imagine then you would have like a very Frankenstein setup you know nothing would really make sense. And the objectives also misaligned so I've shared it before. We want to generate content. We want to generate images or audio. The other one is here to segment things. So like they have different objectives and you're trying to make them work together and it works great but it's also not perfect.

10:16

SPEAKER_01

Here you see on the right side we have Dyno V2 and Dyno V3. Dyno V3 is a better model technically per se than Dyno V2 but when you train your model you're actually getting worse performances. You know Dyno V3 is here in red and green and so you're getting worse performances. So you're like okay like this is supposed to be a better model and yet when I do train a model to generate things then it gets worse. And there's also like not really any rules as to why you know certain encoders should work or otherwise shouldn't. So how can we solve this? How can you teach you know a model representation directly without this external encoder?

10:58

SPEAKER_01

This is what we released about a month and a half ago now which is a research paper. You can read it. It's called Cellflow. It's in the open. We released it to really make sure you know we're moving the field forward again and it's not only us benefiting from it. And it's basically a scalable approach to training multi-model generative models. So they use self-supervised learning so you don't need any other models you know to train it. And I'm going to try to go into a tiny bit more details into it. But we combine representation learning and generation in the same flow. So you see here on the left you have videos, images, audio. You have different modalities.

11:39

SPEAKER_01

What do you do when you usually train a model? You add some noise, you add some random noise, you try to denoise it and then you align it with the encoder you know. How do we do it then? We actually add two different kind of noises that are both random and they're both different. The first one we're adding is actually we're adding a lot of noise to the asset. So this is the one you see at the top. And the other one we're adding like a low amount of noise. This is what you see at the bottom. And the idea is that then we have two models that are actually working together.

12:10

SPEAKER_01

We have the student one which is always getting the images with the most noises and is trying to denoise them. And then the teacher one which is basically a more stable version of the students is always getting the low noises images. And then the student one is actually trying to learn two things at the same time. It's trying to minimize the loss for the generation and the loss in representation. And this is how then you actually work across different modalities. You know this is then you only have one model. You don't have anything external. And if you actually scale up your model then you're scaling up your student. You're scaling up your teacher.

12:49

SPEAKER_01

And you don't have to worry about the encoder that you have on the side anymore. And this is where we're working on. This is something you know we are currently using for different models that we're training. And this is you know where we believe the future is going to be and to get rid of those encoders that we have. We actually train models. So those disclaimer those are research models. They're not meant to be released in production. But we released actually one model on all those modalities. On the left we're comparing flow matching which is the usual way of training models with ours.

13:23

SPEAKER_01

And you can see we are better in audio. So this is what you see in orange on the right side. And then we're also better in images. So the dashed lines is the baseline. And then we are like the full line where we can see we are also better at images. And also better at video. So with this approach without having the encoder and the external model that you may struggle with. You actually get better at every modality that you're training your model on. It's also converging faster. You can see on the right. You know the baseline is converging. It's actually hitting a plateau. Whereas we are converging faster and we're still you know decreasing the loss.

14:02

SPEAKER_01

And I'm pretty sure that if we were to go towards two million steps. You know the baseline would really plateau and then wouldn't really get any better. Maybe actually get worse. Whereas we would still go down in loss. And this is the difference between the two. If you use flux in the past or if you use different models you know to generate you may have noticed the text might not be perfect. Or you know things don't really make sense. This is what you see at the top. Where on the left it's like the futurist flux. But you can see you know you have like some letters that are missing. Or maybe you have two letters. Like on worlds for example you have two L's instead of one.

14:40

SPEAKER_01

Whereas with this approach now in the cell flow approach you can see at the bottom everything makes sense. There's like they learn representation. You know they learn that flux then for the letters should be like one next to the other. And the same on the mirror. Same on the tree. And this is where we believe this is the future. But on top of this we can see some comparisons here. On the left is a baseline. Where again the letters are wrong. And on the right you can see that the letters are correct. Here is the same for the anatomy. Where you see on the left you know you have like a face. It's looking a bit odd. Let's put it that way.

15:17

SPEAKER_01

And on the right this is the one with cell flow. And again this is not like a production model. Where you know you expect the face to be like perfect. But you can see that the anatomy is way better than what you have on the left. But I want to show you as well some different generation. If it loads. Yes thank you. So this is also possible. This is also possible for video generation. So this is the same model that has been trained on images. Now also can generate videos. On the left you see the baseline. It's a weird way to do a push up. Let's put it that way. Whereas on the right it's a perfect form. You know the arms are correct. The hair as well is correct.

15:57

SPEAKER_01

And nothing is wrong in it. And this is you know a way to actually fix all those artifacts that you may see usually in generations. It's the same here for the birds. Oops my bird. Yes thank you. It's the same here for the birds where Thank you. Where you see on the left side you have the baseline. There's a lot of flickering. There's a lot of like you know weird things happening. Because the model was using this encoder was trying to align things. Whereas on the right side with cell flow it just does it perfectly. And like the bird you know is walking on the floor and there's no flickering or anything.

16:35

SPEAKER_01

But it's not only about images or videos or audio. You train those jointly. So you can also actually generate things jointly. We have here an example of a video and audio sample. Where the idea is that we have someone that is saying hello from the black forest. I will just play them and you will hear the difference. Again this is not a production ready model. So it's not like perfect. But you can hear the difference hopefully. Hello from the black forest. Hello from the black forest. Hello from the black forest. So this one was the baseline. Where if you hear it correctly. If you try to pay attention to what he's saying. You hear like hello from the black forest.

17:16

SPEAKER_01

There's a bit of like weird things at the end. From the black forest. Hello from the black forest. Whereas on the right side you can see you know the prompt is really just say hello from the black forest. And then it ends here. And yeah this is the same model that was trained on those images that we've seen before on video. And on video and audio. All from the black forest. Thank you. But this is cool. And this is great. But what if you could also teach robots on how to use this. This is also the same model. This one is trained on actions. And not only on images, video or audio.

17:52

SPEAKER_01

So it can also predict actions. And what I'm going to show you now. It's a robot that is trying to pick up a can. And make it closer to us. On the left. This is a baseline. Again you see some like flickering. You see like the arm is doing weird things. Whereas on the right for the same amount of steps. You can see cell flow. The robot is picking up the arm directly. And like breaking it closer. And this is where we're going as well as the company. This is where we're really interested. It's like not only image generation or video. But it's also doing actions. And doing more things towards physical AI.

18:26

SPEAKER_01

And there is more. It's also how do we make our models faster. Because this is really important for us. This is a demo of Klein. Which is you know like near real-time editing. You see it on the right side. This is you know generated with Klein on Korea. Where you see the edits. And this is not a video model. This is those are images that are always editing in real-time. Not only they are faster. They also actually at least on par or better than other models. And I'm almost out of time. Oh they added five minutes. So I don't know if I'm. Okay cool. So I'm not out of time. So I can chill.

19:02

SPEAKER_01

So yes here on the left. You can see you know we have Klein. That is 4b and 9b. That is compared to the other open source models. So it's at least on par. While the latency. You know it's like 0.5 seconds. While Klein is like around like 15 seconds. You know. And if you are like on par. And you're like way faster. Then this is really really good for us. Same for image to image. You can see the editing. For Klein 9b we are at like a tiny bit more than 0.5 seconds. Whereas Klein is still at around 15 seconds. And then same for multi-ref. You're at it. But we're still at less than a second. Whereas Klein is more towards the 20 seconds.

19:40

SPEAKER_01

And this is what is really really important for us. Because you really want to actually generate things in real-time. This is where we believe there are also to visual intelligence. This is where we're going as a company in the future. And why does it matter? I was like I mentioned it. Real-time generation. So you can imagine. You render mock-ups as fast as you think. You know you don't have to wait. You don't have to wait like 10 seconds. 20. A minute or two. You do things in real-time. And you can guide them in real-time. This is also where we're going. You can think you know interactive visual engines for gaming or films.

20:17

SPEAKER_01

Where you really render a movie as you prompt it. On top of this. There's also world models. The idea of world models and behind it. And why it matters for us. It's you train your models to understand and simulate geometry. The relationship. And like different interaction of the world. And you may be like okay that's cool. From a research perspective. Why do we care? The reason is robots. That's why we care. That's why you know robotics and automation. This is where we're taking BFL. And that's why we also want to go towards world models. Is to train agents in those generative world. To scale safe driving. And automate every manufacturing. And I think that is it.

20:59

SPEAKER_01

Thank you very much.

21:05

SPEAKER_01

Do we take. Yeah. We can take questions. Can you share something where you say your training on portable? You mean the data? Yes. Can't really. This trade secret. I mean data is very sensitive as you can imagine. So I can't really share this. We're partnering with a lot of people though for it. How do you store the state of the world in those action prediction models? Well this is what the model is learning. It's basically like those representation. You know the model is learning that in itself. As like the state. And it has like some kind of memory. And this is the way we do it. What is that sometimes memory? Is it like in the context window?

21:44

SPEAKER_01

Or is it external in the galaxy? No it's. Yeah it's the context window that you have. You know you train. And then you have the tokens. And then they'll be like oh look. I've moved. Here is where I should be then next. Can you then run it for long? Or does it do some convection? Define long. What do you call with long? Indefinitely. Indefinitely. That I'm not sure. I mean there's always going to be a limit. So you may have you know like a sliding window. But this is the way we do it. Thank you. Thank you.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note