AI Engineer

The Desktop Frontier — Ahmad Osman, Osmantic

1526 summary words 7 min summary Watch video

Start with the signal

7 min read

Summary

At-a-Glance

  • Verdict: Skim
  • Core thesis: Open-weight models are improving in capability per parameter fast enough that increasingly capable agentic and reasoning workloads will move from cloud infrastructure to locally owned GPU hardware.
  • Why it matters: This is a strategic case for sovereign AI infrastructure—controlling model access, data, cost, customization, and availability—rather than treating cloud frontier-model access as permanently superior.
  • Best use: Use it as a directional input for local-model and GPU-capacity planning, not as a reliable benchmark or hardware-buying forecast.

Executive Summary

Ahmad Osman argues that the important trend is not simply that “small models beat big models,” but that newer architectures, training methods, post-training, quantization, and mixture-of-experts approaches are producing far more capability from the same—or much less—active parameter and hardware footprint. His metric is “impact per parameter”: what a model can do relative to the compute, VRAM, and hardware needed to run it.

The presentation claims that local open models have moved from short-context, limited assistants to models capable of long-context, tool-use, coding, reasoning, and multi-agent workloads. Osman contrasts older setups—such as requiring multiple RTX 3090s for older Llama-class models—with newer 27B-class dense models that he says can outperform much larger prior-generation models and run on a single high-end consumer GPU. He highlights the shift from large DeepSeek-style reasoning models toward smaller models with comparable agentic performance after better training and post-training.

The strategic conclusion is a strong case for owning at least some compute. Osman predicts that within roughly 18 months, a single 32 GB RTX 5090 could run intelligence comparable to today’s very large frontier-class open model. He treats GPUs as potentially appreciating operational assets: as models become more efficient, existing hardware can serve increasingly valuable workloads. This supports a hybrid or local-first posture for workloads with privacy, control, predictable volume, customization, or availability requirements.

The presentation is useful as a thesis and a set of directional examples, but it is not a rigorous procurement analysis. Many claims rely on informal benchmark comparisons, unclear model/version references in the transcript, and optimistic extrapolation. The transcript also repeats a substantial closing segment. Ken should take the underlying capability-density trend seriously while independently validating actual model quality, throughput, VRAM requirements, licensing, and total cost for target workflows.

Key Takeaways

  • Claim: Model capability is becoming dramatically more hardware-efficient. | Evidence: Osman cites a “density law” trend of roughly 50% fewer parameters for similar capability every three and a half months, and compares newer 27B-class models with older 400B-plus models. | Implication: Hardware sizing should be based on the next 12–24 months of model efficiency gains, not just current model requirements. | Caveat: The cited rate is a directional assertion, not a planning-grade forecast.
  • Claim: Local models are increasingly viable for coding, tool use, reasoning, and agentic workflows. | Evidence: The speaker says local models could not reliably run Claude Code-like workflows a year earlier, while later open models enabled tool calling and more capable coding/agent behavior. | Implication: Local inference is now a credible option for bounded internal-agent tasks rather than only for simple chat or offline experimentation. | Caveat: “Works locally” does not establish production reliability, tool-use robustness, latency, or security controls.
  • Claim: The efficiency gains come from architecture and training advances, not random benchmark movement. | Evidence: Osman points to mixture-of-experts designs, reduced activated-parameter footprints, longer context windows, improved training formats, and post-training on existing checkpoints. | Implication: Track active parameters, memory use, tool-use quality, and throughput—not total parameter count alone—when evaluating model routing and deployment.
  • Claim: Owned GPU capacity can become more useful over time as newer models fit into older hardware. | Evidence: Osman notes that hardware once used for a single older model can later host many parallel agents using more efficient models; he also cites continued demand for RTX 3090-class hardware. | Implication: A local GPU fleet can be treated as option value for experimentation, private inference, and future workload migration. | Caveat: Hardware value is exposed to power costs, depreciation, newer accelerators, supply conditions, maintenance, and the possibility that cloud economics remain better for bursty demand.
  • Claim: Sovereign AI is the strategic reason to run local/open models, not merely lower token cost. | Evidence: The speaker emphasizes retaining control over data, model availability, customization, refusal behavior, and dependence on cloud-provider pricing or policy. | Implication: Local deployment is especially relevant for sensitive data, persistent agents, proprietary workflows, regulated environments, and systems that need stable behavior.
  • Claim: A single high-end consumer GPU may soon run near-frontier open-model quality. | Evidence: Osman predicts that an RTX 5090 with 32 GB VRAM could run a GLM-5.2-class model within about 18 months, based on recent footprint reductions. | Implication: Avoid locking architecture assumptions to the idea that frontier-quality inference always requires a cluster. | Caveat: This is the speaker’s speculative prediction; model quality, quantization losses, context length, tokens/sec, and concurrency could make the practical result materially different.

Detailed Brief

Capability Density and the Model-Progression Argument

  • Claims: The speaker frames the past two years as a succession of footprint reductions: early small models proved that capability could exceed their parameter count; mixture-of-experts models reduced active compute; reasoning and tool-use progressively became available in open-weight releases; and post-training drove major gains without requiring entirely new base-model scale.
  • Evidence: Examples cited include Mistral 7B, Mixtral 8x7B, Llama 3 8B/70B/405B, Qwen 2.5, DeepSeek R1, GPT-OSS 120B, and Qwen 3.5/3.6. He specifically contrasts a roughly 671B-parameter reasoning model with a 120B model he says delivered comparable or better agentic performance, and later contrasts a 27B dense model with a model approximately 15 times larger.
  • Caveats: The talk does not provide reproducible benchmark methodology, inference settings, quantization formats, context lengths, tokens-per-second measurements, or task-specific evaluations. Several model names, versions, dates, and parameter comparisons appear ambiguous in the transcript and should be checked against primary releases.
  • Implications: The practical evaluation unit for an agent stack should become “useful work per dollar, watt, and GB of VRAM,” including concurrency and tool success rate, rather than leaderboard score or nominal model size.

Infrastructure Economics and the Local-First Case

  • Claims: Osman argues that cloud token pricing is partly subsidy-supported and may become less attractive as providers seek returns on data-center investment. He positions owned hardware as protection against model withdrawal, pricing changes, policy restrictions, and provider-controlled access.
  • Evidence: He asks what a DGX Station or retained RTX 3090 fleet may run in three, six, 12, and 18 months, and uses the continued utility of 2020-era RTX 3090 hardware as evidence that model progress can extend hardware usefulness.
  • Caveats: The presentation does not compare all-in local costs—electricity, cooling, depreciation, financing, networking, support, redundancy, orchestration, and operator time—with cloud API and managed-inference costs. It also omits the value of cloud elasticity and access to proprietary frontier models.
  • Implications: The strongest architecture is likely hybrid: reserve local capacity for stable, privacy-sensitive, high-volume, or customizable work; retain cloud routing for spikes, difficult tasks, and models that remain meaningfully better than local alternatives.

Notable Concepts & Terms

  • Capability density / impact per parameter: The central metric proposed by the speaker—capability delivered for a given parameter count and hardware footprint. It is more useful than raw parameter count when assessing local deployment viability.
  • Activated parameters: In mixture-of-experts models, only part of the full model is used per token. This helps distinguish storage footprint from actual inference-compute requirements.
  • Mixture of Experts (MoE): An architecture that routes tokens through a subset of specialized model components. It can provide high total capacity while reducing active compute relative to dense models.
  • Post-training: Additional training after pretraining, including instruction tuning, reinforcement learning, distillation, and task optimization. The talk treats post-training as a major driver of rapid capability gains on existing checkpoints.
  • Sovereign AI: Owning or controlling the model, data path, hardware, and serving layer rather than relying entirely on external AI providers. Relevant to privacy, resilience, customization, and policy control.
  • Hardware option value: The idea that a GPU purchase may gain operational usefulness over time as better models fit into its fixed VRAM and compute envelope.

Operator Notes / Why Ken Should Care

  • Establish a quarterly local-model evaluation harness using Ken’s actual agent tasks: tool-call completion, coding-task success, latency, cost per completed workflow, context reliability, and failure/recovery behavior.
  • Maintain a hybrid routing design rather than committing to local-only inference: local models for routine/private workloads; cloud frontier models for escalation and hard cases.
  • Treat 24–48 GB VRAM nodes as an experimental strategic reserve if utilization supports it, but do not buy solely on the presentation’s 18-month RTX 5090 prediction.
  • Require a complete TCO model before hardware expansion: utilization, power, cooling, uptime, orchestration, security, model operations, and the opportunity cost versus API spend.
  • Monitor open-model licenses and tool-use reliability as closely as raw benchmark gains; permissive weights do not automatically mean unrestricted commercial deployment or dependable autonomous operation.

Source/Metadata

  • Title: The Desktop Frontier — Ahmad Osman, Osmantic
  • Transcript words: 3389
  • Timestamp note: No timestamps or chapters were available; the transcript contains a repeated closing section.
Full transcript 2342 words · 15 min read
0:00

.

0:12

Hey, everyone. We are about to start this presentation. It's called the Disk Top Frontier. And it's about where we started and how far we've come with local and open source models. How many, just a quick question, how many of you here follow me on X?

0:34

I'm amazing. Love you all.

0:42

Love you all. So you know, I sometimes, every now and then, would say a prediction. Here is a new one. Within roughly 18 months, we are going to have the equivalent of GLM 5.2 class intelligence running on a single RTX 5090 with 32 gigabytes of VRAM. That's late 2027. This is conservative. We might actually get there faster. So for a long time, the story has been bigger models, bigger models, bigger models. How can we get to the next $5 trillion?

1:07

How can we get to the $20 trillion? And I'm not saying that there won't ever be a gap between frontier intelligence and open source models. There will always be a gap.

1:18

But that gap will shrink, and the efficiency of the models will get exponentially better. So the term that I like to think about is impact per parameter. What capability are we talking about? What could the model do? What footprint, hardware footprint, did it have last year in comparison to now? And what hardware does that use? And what hardware did it need to use a year ago? And are we moving down for the same kind of quality on that hardware? Again, as I was saying earlier, I used to run LAMA 2 on RTX 1390. It's now running QN 3.5, 3.6, 27 billion parameter.

1:51

That's better than LAMA 3.4.5. That's a 400 billion plus parameters model that you beat with a 27 billion parameter model a year and a half after. So yeah, as I was saying, similar capabilities are moving into smaller hardware footprints. Benchmark scores are one thing. But also, a year ago, this time a year ago, we didn't have any local models that were able to successfully run within cloud code. Right? It wasn't until GLM 4.5 that came out in late July. And GLM 4.5 AIR required at least four RTX 1390s or an RTX Pro 6000. Now, that footprint for hardware is not needed anymore.

2:21

All that you need is a single RTX 1390, 1590, and you have something much more capable, much more intelligent. So is this trend just random, or is there more to it? That's a question that everyone should ask. Is it just by random chance that we've gotten this far from models that weren't able to sustain more than 4,000 tokens in terms of context lens? And now we have things that are million tokens locally on your hardware that you own. It's not by chance. It's not just a coincidence that we got here.

2:49

There is research being done. There are efficiency gains to be made. There are architecture hacks that compound, and they will continue to compound. And I think I like this line. It's not that small models are beating big models. It's that newer, more efficient models are beating older, less efficient ones. So yeah, capability density is the literature I back this up with. Nature Machine Intelligence calls this pattern density law. And every three and a half months, we are having 50% fewer parameters. Whether that's in dense or activated, that's a different story. But we are getting way more intelligence out of the models that we're running.

3:27

So right now, where we're at, it's GLM 5.2. That's our biggest player. And it's 744 billion parameters total, with only 40 billion parameter activated. And that supports up to 1 million context lens. You can run this in MVMV4 on a machine, on a DGX station, or on a server with eight RTX Pro 6000. That's something that you, a DGX station is something that you can set under your desk. And it's running this kind of frontier intelligence. Whether it's on one benchmark, it actually beats GBT 5.5 extra high.

4:03

Doesn't that mean that we're getting somewhere with local and open source models, that we can compete with the frontier, that we're not that far off from the best that you can get from the cloud? We also have Nemo Transfree Ultra, which proved that NVFB4 training, more efficient training, can be done on hardware. Right? That's very important. That means that the footprint, even for training these models, for fine-tuning them, for making small and specialized models, as I was talking earlier, could be more efficient, could be done cheaper, and could deliver you value in terms of economics way sooner, or for much less money than you used to. Yeah. So again, Lama 2.

4:21

That was a 70-bellion parameter model. If you tried to run that right now, you'd laugh at it. Right? That used to take eight RTX 3090s to load up, and those same eight RTX 3090s could run something like 15 parallel agents right now with QN3.5, 27V.

4:34

That's a massive jump in terms of performance gains. So, the vensing law basically means that we have similar or better capabilities with significantly fewer parameters. That's the impact per parameter, as I was saying. I want everybody to leave here thinking about this term and thinking, where are we going to get a year from today? As I was saying earlier, everyone here has a phone, I'm assuming. Raise your hand if you have a phone. If you didn't raise your hand, we know you lie about other things as well. So, come on, guys. You can now run GBT4.0 quality on your iPhone.

5:09

That's massive. That thing requires data centers to serve. So, why wouldn't you invest in sovereign AI? Why wouldn't you, as a consumer, as an individual, as a small-sized business, middle-sized business enterprise, why wouldn't you want to be in control of the models that you're on? Why wouldn't you want to make sure that nothing gets taken away from you? That every little thing can be optimized for you later on. That the performance gains can be made specially and specifically for your use cases, and that you can save more money that way in the long run. And ODS for consumers, it's the way that we support individuals.

5:37

But enterprises also, and I think that there is something that we, as a community, need to think about deeply. We need enterprises for open source AI to win. We need these people that are using the cloud right now, that are supporting data centers being built for cloud providers, to come on this side, to own their own hardware, to own the stack fully, end-to-end, so that we can keep delivering open source models, so that there is an incentive for open source providers to actually come up with models, so that we can come up with new licenses that allow open source to thrive. So, again, open weight and the frontier, I think I, yep, sorry, that was a missed click.

5:49

So, smaller models started bouncing above their weight after Lama 2 with Mr. R7b, one of my favorite models. If you try to put that model right now in cloud code or Robin code, it's not going to work. But it used to take so much in terms of hardware, right, that you would now get from a 9b model that I can run with Telegram, with Oris Hermes, for example, and do a lot of stuff with. So we've come a long way. We had that, we had Mixed Frile 8 by 7b, which everybody knows is an MOE. Then the progression went from that to Lama 3. Lama 3 8b was one of my favorites, still is. It had unique identity, in my opinion.

6:11

Then we had the 70 billion, which was the thing that I would run on my 8 RTX 1390s at home. Then there was the 405, the 400 billion plus parameter Lama 3, which, again, required a lot of hardware. And if you put it now against Queen 3.5, the 27 billion parameter would lose against it. That's in the span of, what, two years, two years and some? No, I think less than two years. That's summer 2024 to March 2026.

6:32

That's about 21 months. And the next big thing, in my opinion, Gamma 2 27V, and then we had the Queen 2.5. And that was the moment that I was like, okay, we actually are making progress, and the gap was shrinking between open source models and the frontier. Really, Lama 3 helped us a lot. And then Queen 2.5 delivered a massive improvement, and there was a lot of fine-tuning and experiments that could be done on that one. There were amazing papers, and they helped the community immensely, in my opinion. Then the next big thing was DeepSeek R1, in my opinion, and the reasoning becoming something that you can run at home.

6:51

That was a massive MOE, almost 700 billion parameters. You had to have a very beefy server to actually get it up and running. And then the improvements that came from just more training on that one, and DeepSeek R1 that was released in May last year, made massive Jamba gains. So it showed that pulse training could deliver more improvements on the same checkpoints. Then GBT open source, GBT OSS 120B. Anyone remembers that one from last summer? Yeah? Nobody here used it? Come on, guys. I need some help here. It was one of the first open source models that were able to successfully do tool calling. And it was a step forward.

7:38

It showed us that we can do more with the hardware that we have running at home. That was a footprint shift, right, from that massive 700 billion parameters, DeepSeek R1 that was, yeah, 671 billion parameters, to something that was one-fifth, one-sixth of its size, and GBT OSS was comparable, maybe better, more agentic performance. Then the moment of QN3.5, the 397, the 397 billion parameters. That's a BV MOE. And what's funny is that it's about 15 times the size of that QN3.6, and I'm here, I'm comparing 3.5 to 3.6 of the dense 27 billion parameter model. And that dense model beats it.

7:59

And that dense model has 40% higher number of activated parameters, so it's not that far off. That's a massive amount of performance gains in a very small amount of time, with massively different footprint in terms of hardware requirements. And that trend happened in two or three months. So how far could we go from here? How far before we get to a recent model that there were some news about that is finally relaunched again? How far before open source delivers something of that quality that you could run on your own hardware? And you can control and will not be taken away from you and will not refuse a request from you.

8:36

So again, these are just some benchmarks where you can see that an iteration on the 27 billion parameter model, a little bit more post-training, proved it across all benchmarks and made it one against a model that is almost 15 size for 15 times its size. Again, remember, this is 27 billion parameters activated versus 17 billion parameter activated. It's still massively the same amount. It's only 40% less in terms of the amount of time it would take to process things, but it's 15 times smaller. That's a lot. So again, how long until the prediction I made earlier becomes possible, when I said that we're going to have the equivalent of GLM 5.2 running on an RTX 5090?

9:08

This is the mass 17 months, and this is conservative math. Earlier this year in December, I had a very viral post that I predicted that we're going to have the quality of Obus 4.5 running locally at home on a single RTX Pro 6000. That happened by March. So, a question: hardware purchase today, does it get more valuable as models become more efficient and smaller in size? That's a good question. So why are you funding other people to build data centers so that you can subscribe to them and pay subsidized tokens, and then later on those subsidies are going to go away and you're not going to be able to run those models, and they will have so many limitations?

9:39

So might as well ask yourself, why not own the hardware yourself and be in control? So yeah, the forward-looking question is basically what will a DJX station be able to run in three, six, twelve, eighteen months from now? That's something that, there is a reason that I'm not selling any of my RTX 3090s. If you follow me, and I have a lot of hardware, guys, but I'm interested in seeing what I could do with them in a year or two from now more than the amount of money I would get for them today. This is not financial advice, by the way. Let me make that very clear. So yeah, the disk site frontier potential. An individual DJX station could run a lot of today.

10:09

It could run GLM 5.2. What would it be able to run tomorrow, six months, 18 months, two years from today? We know that RTX 3090s amber architecture from 2020 sells at higher value than MSRP today and is still being utilized for a lot of use cases. So what will a DJX station, the actively developed blackwell architecture, be able to run in a few months, a couple of years? That's a good question.

10:30

So the question you have to ask yourself, if an RTX 3050 90 with 32 gigabytes of VRAM runs the equivalent of a GLM 5.2 in 18 months, and this is the question that everybody should be asking themselves, and I want you all to be looking at the screen taking this very seriously, okay, should you buy a GVU? Thank you And the next big thing, in my opinion, Gamma 2 27V, and then we had the Queen 2.5. And that was the moment that I was like, okay, we actually are making progress, and the gap was shrinking between open source models and the frontier. Really, Lama 3 saved, like, you know, it really helped us a lot. And then Queen 2.5 delivered a massive improvement,

11:01

and there was a lot of fine-tuning and experiments that could be done on that one. There was amazing papers, and they helped the community immensely, in my opinion. Then the next big thing was DeepSeek R1, in my opinion, and the reasoning becoming something that you can run at home. That was a massive MOE, almost 700 billion parameters. You know, you had to have, like, a very beefy server to actually get it up and running. And then, you know, the improvements that came from just more training on that one, and DeepSeek R1 that was released in May last year, made massive Jamba gains. So it showed that pulse training could deliver more improvements on the same checkpoints.

11:52

Then, GBT open source, like, GBT OSS 120B. Anyone remembers that one from last summer? Yeah? Nobody here used it? Come on, guys. I need some help here.

12:08

It was one of the first open source models that were able to successfully do tool calling. And it was a step forward. It showed us that we can do more with the hardware that we have running at home. That was a footprint shift, right, from, like, you know, that massive 700 billion parameters, DeepSeek R1 that was, yeah, 671 billion parameters to something that was one-fifth, one-sixth of its size, and GBT OSS was comparable, maybe better, more agentic performance.

12:47

Then the moment of QN3.5, the 397, the 397 billion parameters. That's a BV MOE. And, you know, what's funny is that about, it's about 15 times the size of that QN3.6, and I'm here, I'm comparing 3.5 to 3.6 of the dense 27 billion parameter model. And that dense model beats it. And that dense model has 40% higher number of activated parameters, so it's not that far off. That's massive amount of performance gains in a very small amount of time with massively different footprint in terms of hardware requirements. And that trend happened in, like, what, two or three months? So, you know, how far could we go from here?

13:40

How far before we get to, you know, a recent model that there were some news about, you know, that is finally relaunched again? How far before open source delivers something of that quality that you could run on your own hardware? And you can control and will not be taken away from you and will not refuse a request from you.

14:07

So, again, these are just some benchmarks where you can see that an iteration on the 27 billion parameter model, a little bit more post training proved it across all benchmarks and made it one against a model that is almost 15 size for 15 times its size again remember this is 27 billion parameters activated versus 17 billion parameter activated it's still massively the same amount like you know it's it's only 40% less in terms of the amount of time it would take to process things but it's 15 times smaller that's a lot so again how long until the prediction I made earlier becomes possible when I said that we're gonna have the equivalent of

15:01

GLM 5.2 running on an RTX 5090 this is the mass 17 months and this is a conservative math earlier this year in December I had a very viral post that I predicted that we're gonna have the quality of Obus 4.5 running locally at home on a single RTX pro 6000 that happened by March so a question hardware purchase today does it get more valuable as models become more efficient and smaller in size that's a good question so why are you funding other people to build data centers so that you can subscribe to them and pay subsidized tokens and then later on get those subsidies are gonna go away and you're not gonna be able to run those

16:00

models and they will have so many limitations so might as well ask yourself why not own the hardware yourself and be in control so yeah the forward-looking question is basically what will a DJX station be able to run in three six twelve eighteen months from now that's something that there is a reason that I'm not selling any of my RTX 3090s if you follow me and I have a lot of hardware guys but I'm interested in seeing what I could do with them in a year or two from now more than and amount of money I would get for them today this is not a financial advice by the way let me make that very clear so yeah the disk site frontier

16:39

potential you know an individual DJX station could run a lot of today it could run GLM 5.2 what would it be able to run tomorrow six months 18 months two years from today we know that you know RTX 3090s amber architecture from 2020 sells at higher value than MSRP today and still being utilized for a lot of use cases so what will a DJX station the actively developed blackwell architecture will be able to run in a few months a couple of years that's a good question so the question you have to ask yourself if an RTX 3050 90 was 32 gigabytes of VRAM runs and the equivalent of a GLM 5.2 and 18 months and this is the question that everybody should be asking themselves and

17:33

I want you all to be looking at the screen taking this very seriously okay should you buy a GVU thank you

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note