AI Engineer

How Transformers Finally Ate Vision – Isaac Robinson, Roboflow

728 summary words 3 min summary Watch video

Start with the signal

3 min read

Summary

How Transformers Finally Ate Vision

Main Topics

  • Evolution of Computer Vision Architectures: Comparison between CNNs and Vision Transformers (ViTs)
  • The Role of Pre-training: How massive, ViT-specific pre-training compensates for architectural limitations
  • Practical Deployment Challenges: Addressing the gap between powerful foundation models and real-world resource constraints
  • Future Directions: Unified multi-modal (video + image + text) pre-training approaches

Key Points

The Competition: CNNs vs. Transformers

  • Convolutional Neural Networks (CNNs): High inductive bias, N² computational complexity, excellent at capturing local spatial patterns
  • Vision Transformers (ViTs): No inductive bias, N⁴ computational complexity, but ultimately superior performance
  • The paradox: Why does the computationally expensive, less biased architecture win?

Architectural Evolution

  • ViT (Vision Transformer): Splits images into 16×16 patches and applies transformer attention
  • Swin Transformer: Adds window-based local attention to reduce complexity back to N², reintroducing locality bias
  • ConvNext: Applies transformer learnings to convolutional networks; initially outperforms ViT on ImageNet
  • Hara/Meta's Approach: Systematically strips architectural inductive biases and replaces them with learned ones through pre-training
  • Return to ViT: After applying Flash Attention optimizations from LLM research, ViT dominates again

The Pre-training Advantage

  • Masked Autoencoder (MAE): ViT-specific pre-training technique that cannot be applied to CNNs; model learns back lost inductive biases through self-supervised learning
  • DINOv2/DINOv3: Extreme ViT-specific pre-training produces rich, semantically meaningful feature representations
  • Key Insight: Pre-training compensates for architectural lack of bias—models learn spatial reasoning from data rather than architecture

Speed and Optimization

  • Flash Attention: LLM infrastructure breakthrough that makes ViT's N⁴ complexity practical
  • Speedup Impact: Hara showed advantages without Flash Attention, but gains disappeared once Flash Attention was applied
  • Result: Speed advantages of specialized architectures evaporate with proper optimization

Real-World Applications: SAM Series

| Model | Backbone | Approach |

|-------|----------|----------|

| SAM 1 | ViT + MAE | Original |

| Mobile SAM | Tiny ViT (CNN-Transformer Hybrid) | Specialized for deployment |

| SAM 2 | Hara + MAE | Faster alternative |

| SAM 3 | Massively pre-trained ViT | Gives up architecture ablation; scales up instead |

Notable Quotes

> "Massive VIT-specific pre-training, plus speedups from LLMs, plus pre-training compatible neural architecture search. And that's it."

> "We've kind of won."

> "You can't actually apply MAE to a convolutional network. How do you drop out a patch when you're doing this convolution that's invariant across patches? So it's a VIT-specific pre-training technique that adds inductive bias that would otherwise be missing from the structure."

> "No deployment flexibility means that we have these one size fits all models."

> "This combination of this huge advancement in foundation model pre-training and these strong deployment methods combined with an actual ability to deploy these in hardware-constrained environments ends up being the final nail in the coffin for these classical convolutional-based methods."

Takeaways

Primary Conclusions

  • ViTs + Massive Pre-training > Specialized Architecture: The combination of transformer simplicity and large-scale pre-training outperforms carefully designed inductive biases
  • Pre-training is the Real Innovation: Success isn't about architecture—it's about learning from massive unlabeled data using ViT-compatible techniques (MAE, DINO)
  • LLM Infrastructure Benefits Vision: Attention optimizations developed for language models (Flash Attention) made ViT's quadratic complexity practical
  • Deployment Flexibility is Critical: Foundation models like SAM 3 (800M parameters, 300ms latency) are powerful but impractical for edge deployment

Practical Solutions

  • RF100VL Dataset & RFDetter: Roboflow's approach uses neural architecture search with flexible knobs to generate efficient model families from foundation models
  • 40× speedup at same accuracy vs. fine-tuning SAM 3
  • 15× speedup with meaningful quality improvements

Key Challenge

  • Cost Problem: Heavy reliance on expensive pre-training means high computational cost for every deployment variant
  • Solution: Pre-training compatible neural architecture search bridges foundation models with resource-constrained environments

Future Direction

  • Unified video + image + text pre-training architectures
  • SAM 3 demonstrates this direction with multi-modal tracking and pre-training at scale
  • Video JEPPA and similar approaches still awaiting meaningful downstream adoption

Technical Insights

  • N² vs. N⁴ Complexity: ViT's quadratic attention scaling in patches (not pixels) is offset by superior learned representations
  • Inductive Bias Transfer: Pre-training allows models to learn CNN-like spatial biases without hard-coding them
  • Modular Deployment: Drop-in compatible architectural knobs enable hardware-aware model generation without retraining from scratch
Full transcript 2257 words · 19 min read
0:14

SPEAKER_00

Hi, I'm Isaac Robinson. I'm the research lead at Roboflow and I'm here to talk to you today about how Transformers finally ate vision. So I'm going to start off with a brief summary of the competition. Then we're going to go through an overview of the evolution of the transformer, why that ended up winning out, some consequences of that, and what's next. So where we started, convolutional neural networks. I'm sure everyone here is aware of how these work, but just to summarize, they have excellent inductive bias motivated by looking at how the eye works. So you convolve against your image and you have activations that light up the same way regardless of where in the image the thing is happening. Great inductive bias, a person in an image is a person regardless of whether they're in the upper left or the bottom right. We build these interesting hierarchical structures out of these res nets, et cetera. This is how we've done vision for a very long time. Then comes the transformer. Again, I'm sure everyone is aware of how a transformer works, but just to summarize, we have a set of tokens. We run a set-to-set operation. So there's no inductive bias. This is just an n-squared transformation. We inject the inductive biases into the transformer. So, for example, a classical autoregressive transformer, we add a causal mask to the attention matrix, and that gives us sequential modeling. For vision, we have a vision transformer. And this is very complicated. A lot of engineering went into this. We take our image. We split it into patches. 16 by 16 was the original. And we add a learned positional encoding. And then we throw that into a transformer. And that's it. So transformers, n-squared, set-to-set. We've got patches in our image. This is, we've got n over 16 patches for the side length n. And we end up actually with n to the fourth power with the resolution compute scaling. We have no inductive bias. The thing that is in the upper left can have a totally different activation pattern if it's in the bottom right. And so the question actually arises, which is better? The high inductive bias n-squared convolutional network or the no inductive bias and the fourth power of VIT? So as everyone would expect, it's the VIT. So how is this possible? And I'm going to make the argument that it is because of massive VIT-specific pre-training. And then we get to borrow a lot of speedups and infrastructure from the fact that LLMs are blowing up. So to trace this evolution, we're going to talk about, we just talked about the introduction of the VIT, then how people tried to say, OK, well, this cannot possibly be the best thing that we can do. How do we make this better? So we go to SWIN, then back to a convolutional-based network, ConvNext, then Hara, which I think has some really beautiful takeaways. And then, as always happens with machine learning, we come back, better lesson to the simple thing that scales well, the VIT. So first we're going to start with SWIN. So we have this patchify operation. And we take our patches and we split them. Instead of doing global attention across all the patches at once, we just say, OK, we're going to do attention in this window. If we just keep doing attention in this window, though, the tokens will not be able to interact with each other. So these two will never see each other. And so the next layer, in fact, we shift the window a little bit. So we've got these back and forth overlapping windows. And this looks very, very similar to what I described with the convolution. Yes, so we've got this similar looking operation that is happening on these sections of the image. And then we end up with like overlap between the filters, the locations that they get applied on. And this is how we proceed. And this actually gets us down to N squared if your window size is independent of your resolution. And it adds a locality inductive bias following the convolutional that. So that seems logical. That makes sense. Then we go to, someone said, OK, look, there's this transformer operation that has no inherent relationship with vision. Let's go back to the convolutional network. Let's take all the learnings that we've had from the vision transformers and just apply them into a convolution network and see what happens. So ConvNext says, OK, we're going to do a patchify operation. We're going to do a, I think it was a four by four patch instead of a 16 by 16 patch. And we're going to say our VIT was, as all transformers, a self-attention, feed forward, self-attention, feed forward, et cetera. And that self-attention is mixing your spatial information. OK. So for the convolutional network, what if we just say, OK, we're going to have the convolution mix the spatial information. We're going to do the same pattern: mixer, feed forward, mixer, feed forward, onwards. And we're going to borrow the same hierarchical structure that everyone has been using for these convolutional networks. And also throw in layer norm and a couple other innovations. And that's it. We're going to try that. Turns out that beats VIT and SWIN when you apply it on the standard ImageNet reference. That's great. Finally, we have something that makes a little bit of sense. Turns out that's not super fast. So someone, Meta, decided, OK, what are the actual important things here? The ConvNext has a bunch of these beautiful inductive biases. It's following this formula that we got from the transformer. Let's look at what those inductive biases are actually useful for. So we're going to take a really good inductively biased transformer model. We're going to strip out the biases one at a time. We're going to get a speed up because we don't have all the specialized equipment anymore for the inductive bias. And we're going to use pre-training to learn the bias instead. So this is a really great example of the balance between pre-training and inherent inductive bias, which ends up being how transformers ultimately went out. Here we're using MAE, Masked Autoencoder. For those of you who are not familiar, you take your image, you take your patches, you drop a bunch of the patches, and you ask the model to reconstruct what would have been in the patches just based on the context. Very similar to BERT, for those of you who come from the language space. You do this at scale, and it turns out the model actually learns back the inductive biases. But you can't actually apply MAE to a convolutional network. How do you drop out a patch when you're doing this convolution that's invariant across patches? So it's a VIT-specific pre-training technique that adds inductive bias that would otherwise be missing from the structure. So that's great. That's super interesting. That works nicely. It turns out it doesn't... You can take that to an extreme, and you throw in DINOv2, DINOv3, these really VIT-specific pre-training techniques. And you, at the end of it, don't

0:18

SPEAKER_00

language space. You do this at scale, and it turns out the model actually learns back the inductive biases. But you can't actually apply MAE to a convolutional network. How do you drop out a patch when you're doing this convolution that's invariant across patches? So it's a VIT-specific pre-training technique that adds inductive bias that would otherwise be missing from the structure. So that's great. That's super interesting. That works nicely.

0:28

SPEAKER_00

It turns out you can take that to an extreme, and you throw in Dynav2, Dynav3, these really, again, VIT-specific pre-training techniques. And at the end of it, you don't just have these inductive biases towards how to process an image. You actually have really, really, really rich feature maps out of the box. So for example, this is a PCA decomposition of the feature maps produced by a Dynav3 pre-trained VIT. And we see that the paws of the cat have different colors, and that it's tracing the paws correctly for each of the different cats. The satellite imagery is decomposed in a way that is semantically meaningful. And in fact, the self-supervised learning objective is catching up with the best that we have from supervised learning. And this is via linear probe. So you have your frozen features. You're just probing into them. You're not training anything specifically on your data, except for that linear projection at the end. And you're getting very, very, very close to the best that we know how to do with fully supervised learning.

0:34

SPEAKER_00

Okay, but what about the speed issue? This is still n to the fourth power. It turns out that people care a lot about attention. So in the LLM world, we start introducing these tools, Flash Attention, the biggest one. And Hera explicitly showed a speedup for the same accuracy versus VIT. But they have a note in their paper where they say, okay, we see the speedup, and we're not going to measure with Flash Attention.

0:46

SPEAKER_00

So then you add back in Flash Attention, and suddenly it doesn't really matter. You're back to this very silly n to the fourth power thing that benefits from this VIT-specific pre-training method. And that's it. We've kind of won. So this is an evolution of the backbones. How it is that the n to the fourth thing ended up beating out everyone that tried to beat it. What does this mean in practice and application? So SAM is a very famous series of models. Again, I don't know where people come from, but SAM was, from my perspective, one of the most important foundation model series in vision period. And we actually see the same pattern. So SAM to Mobile SAM to SAM 2 to SAM 3, if you look at the backbones that are underlying it, it's a VIT trained with MAE. Then someone says, okay, well, surely this cannot be the best that we can do. So Mobile SAM actually uses a specialized convolutional transformer hybrid called Tiny VIT and replaces the VIT backbone. And then SAM 2 actually uses Hera with this MAE pre-training. And then SAM 3 just gives up on the architecture ablation and just says, okay, well, we've got this massively pre-trained backbone. Let's just stick it in. That's the best that we can do.

0:56

SPEAKER_00

So that's all good and fine. But where does that actually leave us?

1:03

SPEAKER_00

This is really expensive if we're relying on these huge pre-training strategies in order to recover the performance that is lost due to the fact that our architecture is not biased towards the subject at all. That means we have to spend a huge amount of money every time we want to do a deployment. So no deployment flexibility means that we have these one size fits all models. So SAM 3 is this very, very powerful thing, but it's also 800 million parameters. It takes 300 milliseconds to run on a T4 GPU. So it's not actually usable in a lot of cases, especially because vision historically has been focused on these very low power edge devices, these resource-constrained deployment scenarios.

1:08

SPEAKER_00

So at RoboFlow, what we have done is we've attempted to say, how do we actually take these fixed foundation models and transform them into something that has flexibility? So we introduced a dataset, RF100VL, that measures how well the foundation models transfer to downstream diverse tasks with respect to object detection, which is one of the canonical vision-centric tasks. And we see about a 40x speedup for the same accuracy versus fine-tuning SAM 3. And for merely a 15x speedup, we get a meaningful improvement.

1:13

SPEAKER_00

So this combination of this huge advancement in foundation model pre-training and these strong deployment methods combined with an actual ability to deploy these in hardware-constrained environments ends up being the final nail in the coffin for these classical convolutional-based methods. These, at the time of our publication of RFDetter, these were the best real-time instance segmentation models, and we outperformed them in a meaningful way. It might be.

1:25

SPEAKER_00

And so this is all those models of that line use actually the same foundation model. We just modify the foundation model using neural architecture search, such that we generate an entire family of high-performance models in one go. To do this, we actually introduce a bunch of flexible knobs that are drop-in compatible with existing foundation model infrastructure. And by mixing and matching them in a way that is dependent on target data and target hardware, we can resolve the issue that these foundation models do not have deployment flexibility. Quality.

1:41

SPEAKER_00

So massive VIT-specific pre-training, plus speedups from LLMs, plus pre-training compatible neural architecture search. And that's it. Do we have architectures that support unified video plus image plus text pre-training already? Is someone working on that?

2:04

SPEAKER_00

So there are a lot of people working on a huge amount of different combinations of things. I actually think SAM 3 is a good example of that. Well, in terms of vision-specific video processing, it does video processing from the perspective of tracking objects through video. So they do massive scale pre-training. They do the perception encoder pre-training for their backbone, and then they do a huge amount of downstream pre-training. Or I guess you would just call it training at that point. What about JEPPA?

2:21

SPEAKER_00

Yeah, the video JEPPA and the VJEPPA, those are another variety of foundation model. I think in terms of single image-centric pre-training, JEPPA doesn't seem to outperform a lot of the other ones. Video JEPPA, I haven't seen anyone use it meaningfully in a video context for downstream transfer yet, but we'll see. Any other questions? Cool. I think in terms of single image-centric pre-training, JEPPA doesn't seem to outperform a lot of the other ones. Video JEPPA, I haven't seen anyone use it meaningfully in a video context for downstream transfer yet, but we'll see. Any other questions?

3:09

SPEAKER_00

So how is this possible? And I'm going to make the argument that it is because of massive VIT-specific pre-training. And then we get to borrow a lot of speedups and infrastructure from the fact that LLMs are blowing up. So to trace this evolution, we're going to talk about, we just talked about the introduction of the VIT, then how people tried to say, OK, well, this cannot possibly be the best thing that we can do. How do we make this better? So we go to SWIN, then back to a convolutional-based network, Confnext, then Hara, which I think has some really, really beautiful takeaways. And then,

3:55

SPEAKER_00

as always happens with machine learning, we come back, better lesson to the simple thing that scales well, the VIT. So first we're going to start with SWIN. So we have this patchify operation.

4:14

SPEAKER_00

And we take our patches and we split them. Instead of doing global attention across all the patches at once, we just say, OK, we're going to do attention in this window. If we just keep doing attention in this window, though, the tokens will not be able to interact with each other. So these two will never see each other. And so the next layer, in fact, we shift the window a little bit. So we've got these back and forth overlapping windows. And this looks very, very similar to what I described with the convolution. Yes, so we've got this similar looking operation that is happening on these sections of

4:48

SPEAKER_00

the image. And then we end up with like overlap between the filters, the locations that they get applied on. And this is how we proceed. And this actually gets us down to N squared if your window size is independent of your resolution. And it adds a locality inductive bias following the convolutional that. So that that seems logical. That makes sense. Then we go to the someone said, OK, look, there's this transformer operation that has no inherent relationship with vision. Let's go back to the convolutional network. Let's take all the learnings that we've had from the vision transformers and just

5:29

SPEAKER_00

spit them into a convolution network and see what happens. So conv next says, OK, we're going to do a patchify operation. We're going to do a, I think it was a, it was a four by four patch instead of a 16 by 16 patch. And we're going to say our VIT was, as all transformers, a self-attention, feed forward, self-attention, feed forward, et cetera. And that self-attention is mixing your spatial information. OK. So for the convolutional network, what if we just say, OK, we're going to have the convolution mix the spatial information. We're going to do the same pattern. Mixer, feed forward, mixer, feed forward,

6:08

SPEAKER_00

onwards. And we're going to borrow the same hierarchical structure that everyone has been using for these convolutional networks. And also throw in layer norm and a couple other innovations. And that's it. We're going to try that. Turns out that beats VIT and SWIN when you apply it on the standard image net reference. That's great. Finally, we have something that makes a little bit of sense.

6:36

SPEAKER_00

Turns out that's not super fast.

6:42

SPEAKER_00

So someone, Meta, decided, OK, what are the actual important things here? The conv next has a bunch of these beautiful inductive biases. It's following this formula that we got from the transformer. Let's look at what those inductive biases are actually useful for. So we're going to take a really, really good inductively biased transformer model. We're going to strip out the biases one at a time. We're going to get a speed up because we don't have all the specialized equipment anymore for the inductive bias. And we're going to use pre-training to learn the bias instead. So this is, I think, a really, really great example of the balance between pre-training and inherent

7:21

SPEAKER_00

inductive bias, which ends up being how transformers ultimately went out. Here we're using MAE, MAST autoencoder. For those of you who are not familiar, you take your image, you take your patches, you drop a bunch of the patches, and you ask the model to reconstruct what would have been in the patches just based on the context. Very, very similar to BERT, for those of you who come from the language space. You do this at scale, and it turns out the model actually learns back the inductive biases. But you can't actually apply MAE to a convolutional network. How do you drop out a patch when you're doing this convolution that's invariant across patches? So it's a VIT-specific

8:08

SPEAKER_00

pre-training technique that adds inductive bias that would otherwise be missing from the structure.

8:15

SPEAKER_00

So that's great. That's super interesting. That works nicely. It turns out it doesn't... You can take that to an extreme, and you throw in Dynav2, Dynav3, these really, again, VIT-specific pre-training techniques. And you, at the end of it, don't just have these inductive biases towards how to process an image. You actually have really, really, really fit, rich feature maps out of the box. So for example, this is a PCA decomposition of the feature maps produced by a Dynav3 pre-trained VIT. And we see that the paws of the cat have different colors, and that it's tracing the paws correctly for each of the different cats. The satellite imagery is

9:00

SPEAKER_00

decomposed in a way that is semantically meaningful. And in fact, the self-supervised learning objective is catching up with the best that we have from supervised learning. And this is via linear probe. So you have your frozen features. You're just probing into them. You're not training anything specifically on your data, except for that linear projection at the end. And you're getting very, very, very close to the best that we know how to do with fully supervised learning. Okay, but what about the speed issue? This is still n to the fourth power.

9:35

SPEAKER_00

It turns out that people care a lot about attention. So in the LLM world, we start introducing these tools, Flash Attention, the biggest one. And Hera explicitly showed a speedup for the same accuracy versus VIT. But they have a note in their paper where they say, okay, we see the speedup, and we're not going to measure with Flash Attention.

10:05

SPEAKER_00

So then you add back in Flash Attention, and suddenly it doesn't really matter. You're back to this very silly end of the fourth power thing that benefits from this VIT-specific pre-training method. And that's it. We've kind of won. So this is an evolution of the backbones. How it is that the end of the fourth thing ended up beating out everyone that tried to beat it. What does this mean in practice and application? So SAM is a very famous series of models. Again, I don't know where people come from, but SAM was, from my perspective, one of the most important foundation model series in

10:57

SPEAKER_00

vision period. And we actually see the same pattern. So SAM to Mobile SAM to SAM 2 to SAM 3, if you look at the backbones that are underlying it, it's a VIT trained with MAE. Then someone says, okay, well, surely this cannot be the best that we can do. So Mobile SAM actually uses a specialized convolutional transformer hybrid called Tiny VIT and replaces the VIT backbone. And then SAM 2 actually uses HERA with this MAE pre-training. And then SAM 3 just gives up on the architecture ablation and just says, okay, well, we've got this massively pre-trained backbone. Let's just stick it in. That's the best that we can do.

11:44

SPEAKER_00

So that's all good and fine. But where does that actually leave us?

11:49

SPEAKER_00

This is really expensive if we're relying on these huge pre-training strategies in order to recover the performance that is lost due to the fact that our architecture is not biased towards the subject at all. That means we have to spend a huge amount of money every time we want to do a deployment. So no deployment flexibility means that we have these one size fits all models. So SAM 3 is this very, very powerful thing, but it's also 800 million parameters. It takes 300 milliseconds to run on a T4 GPU. So it's not actually usable in a lot of cases, especially because vision historically has been focused on these very low power edge devices,

12:32

SPEAKER_00

these resource-constrained deployment scenarios. So at RoboFlow, what we have done is we've attempted to say, how do we actually take these fixed foundation models and transform them into something that has flexibility? So we introduced a dataset, RF100VL, that measures how well the foundation models transfer to downstream diverse tasks with respect to object detection, which is one of the canonical vision-centric tasks. And we see about a 40x speedup for the same accuracy versus fine-tuning SAM 3. And for merely a 15x speedup, we get a meaningful improvement.

13:29

SPEAKER_00

So this combination of this huge advancement in foundation model pre-training and these strong deployment methods combined with an actual ability to deploy these in hardware-constrained environments ends up being the final nail in the coffin for these classical convolutional-based methods. These, at the time of our publication of RFDetter, these were the best real-time instant segmentation models, and we outperformed them in a meaningful way. It might be. And so this is all those models of that line use actually the same foundation model. We just modify the foundation model using neural architecture search, such that we generate an entire family of

14:26

SPEAKER_00

high-performance models in one go. To do this, we actually introduce a bunch of flexible knobs. Are drop-in compatible with existing foundation model infrastructure. And by mixing and matching them in a way that is dependent on target data and target hardware, we can resolve the issue that these foundation models do not have deployment flexibility. quality.

14:59

SPEAKER_00

So massive VIT-specific pre-training, plus speed ups from LLMs, plus pre-training compatible neural architecture search. And that's it.

15:21

SPEAKER_00

Do we have architectures that support like unified video plus image plus text pre-training already? Is someone working on that?

15:35

SPEAKER_00

So there are a lot of people working on a huge amount of different combinations of things. I actually think SAM 3 is a good example of that. Well, in terms of vision-specific video processing, it does video processing from the perspective of tracking objects through video. So they do massive scale pre-training. They do the perception encoder pre-training for their backbone, and then they do a huge amount of downstream pre-training. Or I guess you would just call it training at that point. What about JEPPA? Yeah, the video JEPPA and the VJEPPA, those are another variety of foundation model.

16:18

SPEAKER_00

I think in terms of single image-centric pre-training, JEPPA doesn't seem to outperform a lot of the other ones. Video JEPPA, I haven't seen anyone use it meaningfully in a video context for downstream transfer yet, but we'll see.

16:40

SPEAKER_00

Any other questions? Cool. Ok. Ok.

16:50

SPEAKER_00

Ok. Ok. Ok. Ok. Ok. Ok. Ok. Ok. Ok. Ok. Ok. Ok. Ok. Ok Ok. Ok Ok.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note