.
Hey, everyone. We are about to start this presentation. It's called the Disk Top Frontier. And it's about where we started and how far we've come with local and open source models. How many, just a quick question, how many of you here follow me on X?
I'm amazing. Love you all.
Love you all. So you know, I sometimes, every now and then, would say a prediction. Here is a new one. Within roughly 18 months, we are going to have the equivalent of GLM 5.2 class intelligence running on a single RTX 5090 with 32 gigabytes of VRAM. That's late 2027. This is conservative. We might actually get there faster. So for a long time, the story has been bigger models, bigger models, bigger models. How can we get to the next $5 trillion?
How can we get to the $20 trillion? And I'm not saying that there won't ever be a gap between frontier intelligence and open source models. There will always be a gap.
But that gap will shrink, and the efficiency of the models will get exponentially better. So the term that I like to think about is impact per parameter. What capability are we talking about? What could the model do? What footprint, hardware footprint, did it have last year in comparison to now? And what hardware does that use? And what hardware did it need to use a year ago? And are we moving down for the same kind of quality on that hardware? Again, as I was saying earlier, I used to run LAMA 2 on RTX 1390. It's now running QN 3.5, 3.6, 27 billion parameter.
That's better than LAMA 3.4.5. That's a 400 billion plus parameters model that you beat with a 27 billion parameter model a year and a half after. So yeah, as I was saying, similar capabilities are moving into smaller hardware footprints. Benchmark scores are one thing. But also, a year ago, this time a year ago, we didn't have any local models that were able to successfully run within cloud code. Right? It wasn't until GLM 4.5 that came out in late July. And GLM 4.5 AIR required at least four RTX 1390s or an RTX Pro 6000. Now, that footprint for hardware is not needed anymore.
All that you need is a single RTX 1390, 1590, and you have something much more capable, much more intelligent. So is this trend just random, or is there more to it? That's a question that everyone should ask. Is it just by random chance that we've gotten this far from models that weren't able to sustain more than 4,000 tokens in terms of context lens? And now we have things that are million tokens locally on your hardware that you own. It's not by chance. It's not just a coincidence that we got here.
There is research being done. There are efficiency gains to be made. There are architecture hacks that compound, and they will continue to compound. And I think I like this line. It's not that small models are beating big models. It's that newer, more efficient models are beating older, less efficient ones. So yeah, capability density is the literature I back this up with. Nature Machine Intelligence calls this pattern density law. And every three and a half months, we are having 50% fewer parameters. Whether that's in dense or activated, that's a different story. But we are getting way more intelligence out of the models that we're running.
So right now, where we're at, it's GLM 5.2. That's our biggest player. And it's 744 billion parameters total, with only 40 billion parameter activated. And that supports up to 1 million context lens. You can run this in MVMV4 on a machine, on a DGX station, or on a server with eight RTX Pro 6000. That's something that you, a DGX station is something that you can set under your desk. And it's running this kind of frontier intelligence. Whether it's on one benchmark, it actually beats GBT 5.5 extra high.
Doesn't that mean that we're getting somewhere with local and open source models, that we can compete with the frontier, that we're not that far off from the best that you can get from the cloud? We also have Nemo Transfree Ultra, which proved that NVFB4 training, more efficient training, can be done on hardware. Right? That's very important. That means that the footprint, even for training these models, for fine-tuning them, for making small and specialized models, as I was talking earlier, could be more efficient, could be done cheaper, and could deliver you value in terms of economics way sooner, or for much less money than you used to. Yeah. So again, Lama 2.
That was a 70-bellion parameter model. If you tried to run that right now, you'd laugh at it. Right? That used to take eight RTX 3090s to load up, and those same eight RTX 3090s could run something like 15 parallel agents right now with QN3.5, 27V.
That's a massive jump in terms of performance gains. So, the vensing law basically means that we have similar or better capabilities with significantly fewer parameters. That's the impact per parameter, as I was saying. I want everybody to leave here thinking about this term and thinking, where are we going to get a year from today? As I was saying earlier, everyone here has a phone, I'm assuming. Raise your hand if you have a phone. If you didn't raise your hand, we know you lie about other things as well. So, come on, guys. You can now run GBT4.0 quality on your iPhone.
That's massive. That thing requires data centers to serve. So, why wouldn't you invest in sovereign AI? Why wouldn't you, as a consumer, as an individual, as a small-sized business, middle-sized business enterprise, why wouldn't you want to be in control of the models that you're on? Why wouldn't you want to make sure that nothing gets taken away from you? That every little thing can be optimized for you later on. That the performance gains can be made specially and specifically for your use cases, and that you can save more money that way in the long run. And ODS for consumers, it's the way that we support individuals.
But enterprises also, and I think that there is something that we, as a community, need to think about deeply. We need enterprises for open source AI to win. We need these people that are using the cloud right now, that are supporting data centers being built for cloud providers, to come on this side, to own their own hardware, to own the stack fully, end-to-end, so that we can keep delivering open source models, so that there is an incentive for open source providers to actually come up with models, so that we can come up with new licenses that allow open source to thrive. So, again, open weight and the frontier, I think I, yep, sorry, that was a missed click.
So, smaller models started bouncing above their weight after Lama 2 with Mr. R7b, one of my favorite models. If you try to put that model right now in cloud code or Robin code, it's not going to work. But it used to take so much in terms of hardware, right, that you would now get from a 9b model that I can run with Telegram, with Oris Hermes, for example, and do a lot of stuff with. So we've come a long way. We had that, we had Mixed Frile 8 by 7b, which everybody knows is an MOE. Then the progression went from that to Lama 3. Lama 3 8b was one of my favorites, still is. It had unique identity, in my opinion.
Then we had the 70 billion, which was the thing that I would run on my 8 RTX 1390s at home. Then there was the 405, the 400 billion plus parameter Lama 3, which, again, required a lot of hardware. And if you put it now against Queen 3.5, the 27 billion parameter would lose against it. That's in the span of, what, two years, two years and some? No, I think less than two years. That's summer 2024 to March 2026.
That's about 21 months. And the next big thing, in my opinion, Gamma 2 27V, and then we had the Queen 2.5. And that was the moment that I was like, okay, we actually are making progress, and the gap was shrinking between open source models and the frontier. Really, Lama 3 helped us a lot. And then Queen 2.5 delivered a massive improvement, and there was a lot of fine-tuning and experiments that could be done on that one. There were amazing papers, and they helped the community immensely, in my opinion. Then the next big thing was DeepSeek R1, in my opinion, and the reasoning becoming something that you can run at home.
That was a massive MOE, almost 700 billion parameters. You had to have a very beefy server to actually get it up and running. And then the improvements that came from just more training on that one, and DeepSeek R1 that was released in May last year, made massive Jamba gains. So it showed that pulse training could deliver more improvements on the same checkpoints. Then GBT open source, GBT OSS 120B. Anyone remembers that one from last summer? Yeah? Nobody here used it? Come on, guys. I need some help here. It was one of the first open source models that were able to successfully do tool calling. And it was a step forward.
It showed us that we can do more with the hardware that we have running at home. That was a footprint shift, right, from that massive 700 billion parameters, DeepSeek R1 that was, yeah, 671 billion parameters, to something that was one-fifth, one-sixth of its size, and GBT OSS was comparable, maybe better, more agentic performance. Then the moment of QN3.5, the 397, the 397 billion parameters. That's a BV MOE. And what's funny is that it's about 15 times the size of that QN3.6, and I'm here, I'm comparing 3.5 to 3.6 of the dense 27 billion parameter model. And that dense model beats it.
And that dense model has 40% higher number of activated parameters, so it's not that far off. That's a massive amount of performance gains in a very small amount of time, with massively different footprint in terms of hardware requirements. And that trend happened in two or three months. So how far could we go from here? How far before we get to a recent model that there were some news about that is finally relaunched again? How far before open source delivers something of that quality that you could run on your own hardware? And you can control and will not be taken away from you and will not refuse a request from you.
So again, these are just some benchmarks where you can see that an iteration on the 27 billion parameter model, a little bit more post-training, proved it across all benchmarks and made it one against a model that is almost 15 size for 15 times its size. Again, remember, this is 27 billion parameters activated versus 17 billion parameter activated. It's still massively the same amount. It's only 40% less in terms of the amount of time it would take to process things, but it's 15 times smaller. That's a lot. So again, how long until the prediction I made earlier becomes possible, when I said that we're going to have the equivalent of GLM 5.2 running on an RTX 5090?
This is the mass 17 months, and this is conservative math. Earlier this year in December, I had a very viral post that I predicted that we're going to have the quality of Obus 4.5 running locally at home on a single RTX Pro 6000. That happened by March. So, a question: hardware purchase today, does it get more valuable as models become more efficient and smaller in size? That's a good question. So why are you funding other people to build data centers so that you can subscribe to them and pay subsidized tokens, and then later on those subsidies are going to go away and you're not going to be able to run those models, and they will have so many limitations?
So might as well ask yourself, why not own the hardware yourself and be in control? So yeah, the forward-looking question is basically what will a DJX station be able to run in three, six, twelve, eighteen months from now? That's something that, there is a reason that I'm not selling any of my RTX 3090s. If you follow me, and I have a lot of hardware, guys, but I'm interested in seeing what I could do with them in a year or two from now more than the amount of money I would get for them today. This is not financial advice, by the way. Let me make that very clear. So yeah, the disk site frontier potential. An individual DJX station could run a lot of today.
It could run GLM 5.2. What would it be able to run tomorrow, six months, 18 months, two years from today? We know that RTX 3090s amber architecture from 2020 sells at higher value than MSRP today and is still being utilized for a lot of use cases. So what will a DJX station, the actively developed blackwell architecture, be able to run in a few months, a couple of years? That's a good question.
So the question you have to ask yourself, if an RTX 3050 90 with 32 gigabytes of VRAM runs the equivalent of a GLM 5.2 in 18 months, and this is the question that everybody should be asking themselves, and I want you all to be looking at the screen taking this very seriously, okay, should you buy a GVU? Thank you And the next big thing, in my opinion, Gamma 2 27V, and then we had the Queen 2.5. And that was the moment that I was like, okay, we actually are making progress, and the gap was shrinking between open source models and the frontier. Really, Lama 3 saved, like, you know, it really helped us a lot. And then Queen 2.5 delivered a massive improvement,
and there was a lot of fine-tuning and experiments that could be done on that one. There was amazing papers, and they helped the community immensely, in my opinion. Then the next big thing was DeepSeek R1, in my opinion, and the reasoning becoming something that you can run at home. That was a massive MOE, almost 700 billion parameters. You know, you had to have, like, a very beefy server to actually get it up and running. And then, you know, the improvements that came from just more training on that one, and DeepSeek R1 that was released in May last year, made massive Jamba gains. So it showed that pulse training could deliver more improvements on the same checkpoints.
Then, GBT open source, like, GBT OSS 120B. Anyone remembers that one from last summer? Yeah? Nobody here used it? Come on, guys. I need some help here.
It was one of the first open source models that were able to successfully do tool calling. And it was a step forward. It showed us that we can do more with the hardware that we have running at home. That was a footprint shift, right, from, like, you know, that massive 700 billion parameters, DeepSeek R1 that was, yeah, 671 billion parameters to something that was one-fifth, one-sixth of its size, and GBT OSS was comparable, maybe better, more agentic performance.
Then the moment of QN3.5, the 397, the 397 billion parameters. That's a BV MOE. And, you know, what's funny is that about, it's about 15 times the size of that QN3.6, and I'm here, I'm comparing 3.5 to 3.6 of the dense 27 billion parameter model. And that dense model beats it. And that dense model has 40% higher number of activated parameters, so it's not that far off. That's massive amount of performance gains in a very small amount of time with massively different footprint in terms of hardware requirements. And that trend happened in, like, what, two or three months? So, you know, how far could we go from here?
How far before we get to, you know, a recent model that there were some news about, you know, that is finally relaunched again? How far before open source delivers something of that quality that you could run on your own hardware? And you can control and will not be taken away from you and will not refuse a request from you.
So, again, these are just some benchmarks where you can see that an iteration on the 27 billion parameter model, a little bit more post training proved it across all benchmarks and made it one against a model that is almost 15 size for 15 times its size again remember this is 27 billion parameters activated versus 17 billion parameter activated it's still massively the same amount like you know it's it's only 40% less in terms of the amount of time it would take to process things but it's 15 times smaller that's a lot so again how long until the prediction I made earlier becomes possible when I said that we're gonna have the equivalent of
GLM 5.2 running on an RTX 5090 this is the mass 17 months and this is a conservative math earlier this year in December I had a very viral post that I predicted that we're gonna have the quality of Obus 4.5 running locally at home on a single RTX pro 6000 that happened by March so a question hardware purchase today does it get more valuable as models become more efficient and smaller in size that's a good question so why are you funding other people to build data centers so that you can subscribe to them and pay subsidized tokens and then later on get those subsidies are gonna go away and you're not gonna be able to run those
models and they will have so many limitations so might as well ask yourself why not own the hardware yourself and be in control so yeah the forward-looking question is basically what will a DJX station be able to run in three six twelve eighteen months from now that's something that there is a reason that I'm not selling any of my RTX 3090s if you follow me and I have a lot of hardware guys but I'm interested in seeing what I could do with them in a year or two from now more than and amount of money I would get for them today this is not a financial advice by the way let me make that very clear so yeah the disk site frontier
potential you know an individual DJX station could run a lot of today it could run GLM 5.2 what would it be able to run tomorrow six months 18 months two years from today we know that you know RTX 3090s amber architecture from 2020 sells at higher value than MSRP today and still being utilized for a lot of use cases so what will a DJX station the actively developed blackwell architecture will be able to run in a few months a couple of years that's a good question so the question you have to ask yourself if an RTX 3050 90 was 32 gigabytes of VRAM runs and the equivalent of a GLM 5.2 and 18 months and this is the question that everybody should be asking themselves and
I want you all to be looking at the screen taking this very seriously okay should you buy a GVU thank you