AI Engineer

Compression at the Edge — Chris Alexiuk, NVIDIA

2073 summary words 9 min summary Watch video

Start with the signal

9 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Quantization is becoming a core deployment discipline: selectively compressing large open models can preserve most useful capability while making local, private, and lower-cost inference viable across laptops, workstations, and eventually phones.
  • Why it matters: For agent systems and local AI operations, model compression determines whether capable models can be run privately, economically, and with enough throughput to support real harnesses rather than isolated benchmark demos.
  • Best use: Use this as a strategic and technical framing session for local-model selection, quantization validation, model routing, and the failure modes to test before putting compressed models behind agents.

Executive Summary

The panel argues that compression is not simply shrinking models; it is an engineering trade between memory, latency, throughput, and quality that can substantially expand access to capable AI. The most valuable technique discussed is selective quantization: retain high precision in unusually sensitive layers, tensors, or “super weights,” while aggressively reducing precision elsewhere. This can make a model dramatically smaller without proportionately degrading its behavior.

The speakers make a practical case that a larger model quantized to low precision can outperform a smaller model occupying the same memory footprint at higher precision. But this is not universally optimal: native small models can deliver dramatically higher speed, cited as roughly 200 tokens per second versus 5–10 tokens per second for a much larger compressed model on constrained hardware. The right deployment choice therefore depends on whether the task prioritizes planning quality, execution speed, concurrency, privacy, or local ownership.

A recurring warning is that standard accuracy benchmarks are insufficient for deployment confidence. Quantized models can appear acceptable in aggregate evaluations yet fail in real agent harnesses, long-context tasks, or specific architecture-dependent behaviors. The panel recommends behavioral testing in the intended harness and, for quantization assessment, comparing output distributions/logits against the original BF16 model using KL divergence in addition to benchmark scores.

The implementation environment is becoming harder, not easier: new mixture-of-experts, hybrid-attention, sparse-attention, and linear-attention architectures invalidate generic quantization recipes. Some MoE layers can reportedly tolerate one-bit quantization, while quantizing linear-attention layers can look fine initially but collapse into gibberish in long-context production use. The likely next frontier is broader than weight quantization: KV-cache compression, activation sparsity, and compression-aware model releases will matter increasingly for long-horizon agents and edge deployment.

Key Takeaways

  • Claim: Effective compression is selective rather than uniform: model layers, tensors, and even individual weights have sharply different sensitivity to reduced precision. | Evidence: Daniel from Unsloth says early and final layers tend to be important while middle layers may be less sensitive; he cites a “Super Weights” paper in which quantizing one particular number can make a model “20% dumber.” NVIDIA's representative identifies attention projection/QKV-related layers as sensitive and says its tooling uses gradient-based sensitivity analysis plus a linear-programming/knapsack-style solver. | Implication: Do not treat a quantization level such as 4-bit as a complete deployment specification. Select or build quantizations that document layer-level treatment, then validate against the workloads and contexts Ken actually runs. | Caveat: These are architecture- and model-specific heuristics, not universal rules; applying a generic precision policy can preserve benchmark scores while breaking relevant behavior.
  • Claim: A large model quantized to low precision may be a better use of a fixed memory budget than a smaller model at high precision. | Evidence: The panel cites experiments comparing a 35B BF16 model with a roughly four-times-larger 120B model at 4-bit precision; despite similar disk footprint, they claim the larger quantized model is materially more intelligent. GLM 5.2 is offered as an example of a 1.5 TB model that can be reduced to about 250 GB, an 86% size reduction. | Implication: Use compressed large models where planning, difficult reasoning, or quality dominates; use smaller native models where interactive execution, high concurrency, or agent-loop latency dominates. A routed architecture can assign planning to the large model and execution to the smaller one. | Caveat: The panel also stresses that speed can reverse the practical choice: a constrained system may get only 5–10 tokens/sec from a large compressed model versus about 200 tokens/sec from a small model.
  • Claim: Compression is operationally valuable to enterprises, not merely a way to run models on consumer hardware. | Evidence: Panelists cite increased serving concurrency, lower compute costs, in-house deployment for sensitive data, and distilling mid-sized models into very small specialized rerankers that cut costs by millions. Ollama's representative emphasizes enabling organizations to run models not only in central infrastructure but on individual employee machines. | Implication: Compression belongs in AI unit economics and deployment architecture reviews: it can support private/on-prem use cases, specialized low-cost workers, and wider distribution of local AI capabilities.
  • Claim: Benchmark parity alone does not establish that a compressed model is safe or useful in an agentic product. | Evidence: Ollama's representative says benchmark results cannot capture many harness-level failures and describes manually testing whether a model's behavior “feels right” in real use. The panel notes that a model can survive initial checks but become gibberish on long-context benchmarks or production workloads after quantizing linear-attention layers. | Implication: Build an internal compressed-model acceptance suite around Ken's actual tasks: long-horizon tool use, coding, self-repair, context retention, structured output, refusal behavior, and latency/cost under concurrency—not leaderboard scores alone. | Caveat: There is no comprehensive public resource that compares the expanding set of quantized or modified variants against original model behavior across use cases; the panel explicitly acknowledges this gap.
  • Claim: Post-training quantization is relatively accessible for medium and large dense models, but recovering quality in smaller or reasoning-heavy models is materially harder. | Evidence: NVIDIA's representative says a large checkpoint can be quantized to FP4 in a couple of hours on Blackwell nodes, and that post-training quantization often works out of the box for roughly 20B–30B+ dense models with selective heuristics. For smaller models, they say quantization-aware distillation is often needed; reasoning models are harder because they are multi-stage RL and multi-teacher-distilled systems. | Implication: Prefer proven post-training quantizations for mature medium/large base models. Avoid casually retraining or “repairing” compressed reasoning models unless the original-quality training data and a rigorous evaluation loop are available. | Caveat: Quantization-aware distillation with the wrong data can break the model rather than recover performance, because matching the original training distribution and teacher behavior is difficult.
  • Claim: Future compression work will move beyond weights toward the inference state and compute path required by long-context agents. | Evidence: The NVIDIA speaker identifies KV-cache quantization/compaction, dynamic activation sparsity in Rubin, and low-bit math acceleration as next areas. They suggest weight quantization may be nearing a local Pareto frontier around two or three additional bits of reduction, while KV-cache and activation methods remain meaningful expansion areas. | Implication: For long-horizon OpenClaw-style agents, monitor KV-cache memory and attention compute as first-class bottlenecks; model-file size alone will not determine whether a local agent can sustain long tasks. | Caveat: Sparsity has not seen broad adoption because it tends to degrade accuracy more than quantization, despite hardware support.
  • Claim: Compression is a foundational enabler of local ownership, customization, and private agent operation rather than merely an optimization technique. | Evidence: Merve describes running OpenClaw and a Hermes agent on long-horizon tasks using Qwen 3.6 quantizations, including coding and self-harness repair. Multiple panelists frame the destination as capable models running on laptops and phones, allowing local fine-tuning and operation on personal or sensitive data. | Implication: The local-agent opportunity improves as quantized open models improve, but deployment should be tied to concrete privacy, offline-resilience, and cost advantages—not assumed merely because a model fits in memory.

Detailed Brief

Numeric formats and how the panel evaluates quantization fidelity

  • Claims: NVFP4 is presented as more than a simple four-bit storage format: it couples compression with faster low-bit matrix multiplication.; Micro-block scaling is central to making very-low-bit floating-point representations retain useful information.; KL divergence between the original and compressed model's output logits can be a more direct compression-fidelity signal than noisy sampled accuracy benchmarks.
  • Evidence: The NVIDIA speaker describes NVFP4 as a four-bit floating-point format in which groups of 16 values share an FP8 scaling factor, attributing the underlying micro-block scaling idea to Tim Detmers/bitsandbytes.; Their stated objective is to minimize the distance between BF16 and quantized output logits on calibration data while reducing model size.; NVIDIA's model-optimizer team says it typically targets less than 1% overall accuracy degradation across its benchmark suite when publishing quantized checkpoints.
  • Caveats: Low KL divergence is a fidelity metric relative to the original model, not proof that the original model is good for the target task or agent harness.; Benchmark evaluation still consumes substantial engineering effort after the mechanical step of generating a quantized checkpoint.
  • Implications: When selecting artifacts, favor releases that disclose precision formats, calibration/evaluation methodology, and comparisons to the original checkpoint.; Separate two questions in evaluation: whether the quantized model faithfully matches its source and whether either model achieves the intended operational outcome.

Architecture diversity is creating quantization-specific failure modes

  • Claims: The older broadly uniform “Transformer++” environment has given way to heterogeneous architectures with different attention schemes, normalization details, sliding/global-window designs, and MoE patterns.; This diversity increases both model-runtime implementation work and the need for architecture-specific default quantizations.
  • Evidence: Ollama says model labs may provide multiple variants before release, each requiring implementation support and usability testing before a default quantization can be selected.; Daniel contrasts highly quantizable MoE components with linear-attention components that can fail under quantization on long context.; The panel names MLA, sparse attention, index attention, and changing softmax/attention mechanisms as sources of growing complexity.
  • Caveats: The speakers see architectural diversity as worthwhile innovation despite the operational complexity; the conclusion is not to avoid novel architectures, but to avoid assuming prior quantization recipes transfer.
  • Implications: Treat every new model family as a new compatibility and validation target in the local serving stack.; Maintain model-family-specific test fixtures, especially for long-context behavior, rather than defining a single global quantization policy.

Notable Concepts & Terms

  • Post-training quantization (PTQ): Reducing precision after a BF16/base model is released; the panel presents it as comparatively straightforward for medium and large dense models, with evaluation being the harder phase.
  • Quantization-aware distillation (QAD/QAT): Training-based recovery methods used when PTQ harms smaller or more fragile models; quality depends critically on appropriate original or representative training data.
  • NVFP4: NVIDIA's four-bit floating-point format using micro-block scaling; it is positioned as both a storage-compression and low-bit-compute-acceleration mechanism.
  • Micro-block scaling: A shared scaling factor for a small group of quantized values, described here as 16 elements sharing an FP8 scale, which improves usable accuracy at very low precision.
  • Super Weights: A cited finding that isolated model values can be exceptionally important, illustrating why uniform or indiscriminate quantization can create disproportionate quality loss.
  • KL divergence / logit matching: A proposed way to measure distance between BF16 and compressed models using output logits on calibration data, supplementing noisy task-benchmark comparisons.
  • KV-cache compression: Compression of the inference-time context state rather than only model weights; important for reducing memory pressure in long-context and long-horizon agent runs.
  • Model routing: Using different models for different stages—for example, a larger compressed model for planning and a smaller fast model for execution—to optimize quality and latency together.

Operator Notes / Why Ken Should Care

  • Define a compressed-model release gate that includes task-specific harness tests, long-context regression tests, structured-output validity, tool-use success, latency, and memory—not just benchmark deltas.
  • Adopt a two-tier local routing experiment: a high-capability compressed planner plus a small high-throughput executor, and measure whether the added orchestration improves end-to-end agent completion rate.
  • For new hybrid/linear-attention or MoE model families, require architecture-specific quantization validation before making them a default local runtime option.
  • Track KV-cache size, context length, and attention throughput separately from model-weight footprint when sizing long-horizon local-agent infrastructure.
  • Prefer quantization artifacts with reproducible provenance and fidelity reporting; avoid unverified “Frankenstein” conversions where source checkpoint, precision policy, and behavior tests are unclear.
  • Investigate whether sensitive-data workflows justify local compressed deployment even when cloud inference is available, using privacy, concurrency, and unit-cost thresholds as decision criteria.

Source/Metadata

  • Title: Compression at the Edge — Chris Alexiuk, NVIDIA
  • Transcript words: 12398
  • Duration seconds: 2760
  • Timestamp note: No usable timestamps or chapters were present in the supplied transcript; the latter portion also contains repeated transcript segments.
Full transcript 7414 words · 55 min read
0:12

Hello everybody. Welcome to Compression at the Edge, the panel that we'll be conducting for the next bit here. Very nice to meet you all. I'll be your trusty moderator today. My name is Chris Alexiak. I'm a product research engineer at NVIDIA. I work on Nemo Trine. Let's go. Okay. We are joined by Daniel. Yes, hello everyone. I'm from Unsloth. Thanks for coming, everyone. Excellent. And? Hello, hi. I'm Asma. I build NVIDIA model optimizer, and we conduct a lot of models. Let's go. I'm Marve. I work as a machine learning engineer at Hugging Face. Awesome. I'm Parth. I work at Olama.

0:57

So compression: a big topic. We're going to set some context, hopefully, in order to launch into this. So maybe, in each of your own words, if you want to define what you think about compression, let us know how you engage with technology that ultimately is designed to make models that are bigger be a little bit smaller. Right? That's the general idea. So maybe we'll just go in reverse. Parth, if you want to kick us off: what is compression to you?

1:05

Yeah, I think honestly with Olama, and for those of you who are not familiar, Olama is one of the easiest ways to run local models. And for us, what rose us to popularity was being able to run a larger model on a relatively small machine through quantization, which I'm sure we'll talk a lot about today. And to me, compression is so important because it actually makes these giant models viable for most people. I think for me, it's just shrinking something without losing information, but there's absolutely zero free lunch. At the end of the day, you still spend on something, whether it's latency or quality.

1:17

And I agree. I feel like compression is much more than that definition because it democratizes the models for everyone at edge devices, at your computer. I'm sure you are all running some Gemma 4 quant at the moment, or QN 3.6. Those are the hot ones these days. And it just works so well. So yeah, this is my definition: it democratizes things.

1:23

Cool. So the way I think about it is: same cost, more intelligence. So compression accelerates and enables. To give a quick example, originally we started with training in FP32, right? And now we are talking about FP4. So that is 8x more compression and almost the same intelligence, with not much degradation. Yeah: same cost, more intelligence.

1:30

Yeah, how we see quantization is you take a big model like GLM 5.2. It's 1.5 terabytes, which is definitely ginormous. But then the trick is you can actually quantize it and shrink it to 250 GB. So you can make it 86% smaller. But with tricks of quantization, it will not become 86% dumber, right? If you compress it by 86%, it doesn't become terrible, useless. And so what we show with dynamic quantization, if you quantize some layers to higher precision and you leave most of the layers in one-bit or two-bit, and some very important layers in 16-bit, you can still recover 76% of all accuracy. So the trick for quantization is if you quantize the correct layers, you will not make the model literally useless. And compression, I guess, is very important for you to run on your local computers, make democratization of AI.

1:37

Yeah. So we've been talking today about this inflection point, right? Recent moments that led to this resurgence of the importance of local AI and open models, right? Own your own intelligence. You guys were ahead of the curve, though, right? I mean, you guys were thinking about this before it was cool to think about it, right? So I'd love to hear from each of you: when was the moment that you knew that quantization or compression is the path forward? And we're going to talk a lot today about consumer hardware, right? That means your RTX cards, your Macs, things like that, right? But it goes well beyond that. What was the moment you first got the quantization or compression bug that made you say, “Oh, shit, this is going to be big, man”? We'll just go right back down the line. Daniel, take us away.

1:43

Yeah, so I think the biggest moment was Deep Seek R1, definitely. When it got released, it was quite dramatic for the world because we finally had some sort of reasoning model that worked very well, and it was open source. The biggest problem, though, was that it was very big, and running it locally was extremely complex. And so when we started off, we decided, okay, let's do some tricks. Let's quantize some layers to higher precision. And it randomly worked. We were quite surprised that 1.58-bit quant worked well. And we just posted about this, and we were quite surprised. Okay, local models seem to be working well. And obviously over time we had Quinn 3.6, 3.5, Gemma 4. Every single time there's a new, even NVIDIA's open source models, Nemetron, they just keep getting better and better and better. The only problem sometimes is they get bigger and bigger and bigger. So that's the only problem. And so you need to focus more on how to make the model smaller and smaller and smaller. And so recently we have GLM 5.2. It is a very good local model, although it's ginormous, and so making it work very well on the local device is complicated. But I think it's a long history of these models, and we're expecting more. Whatever the next new Metro model is, the next Gemma, the next any model, we're very excited for the future.

1:50

Yeah. It's an ongoing arms race between how big can we make the models versus how small can we make them and they still work, right? Maybe some thoughts from me: what did you think? Like, ah, man, quantization is it?

1:58

So I originally started working on pruning. So, okay, that was maybe three years back. That was when CV models were all the rage, then tragedy happened. So pruning, typically, you lose quality. So you need to fine-tune it a little bit. And then LLMs came out, and quantization was an easy compression thing you could do. You don't lose a lot of accuracy. So what I learned was some kinds of compression are really good. You can architect techniques so that you can chop it down a lot without losing quality. So quantization is really good. Sparsity is really good, but not as much as quantization. So I did not have an exact moment where I got this numeric bug. It slowly grew on me. So, for example, NBF before, right, I was fortunate to work on NBF before experiment math when I joined NVIDIA, and NBF before is very genius. We have been doing a lot of research and experiments to make NBF before better, but it is a really solid numeric format. So yeah, that was basically it. Some compressions are really good, and you can cleverly architect understanding things better and not lose quality. Like mixed-precision quantization, you can compress it a ton without losing much quality. Yeah, that was a slow growth for me.

2:03

Should I? Yes, please. So I think for me, the biggest wow moment of quantization was back in the day. There was something called Q Laura, and the fact that we could actually fine-tune stuff on a toaster, essentially, that is the Colab free-tier T4. For me that was a big wow. Although it's super slow, it's okay. But on top of it, for instance, we have libraries called TRL and bits and bytes and path that enables all of this. And you can even do this with the very advanced not lose quality, dance at mixed precision, quantization, you can compress it at 10, chop it at 10 without losing much quality. Yeah. So that was a slow growth for me. Should I? Yes, please.

2:48

So, I think for me, the biggest wow moment of quantization was back in the day. There was something called QLoRA, and the fact that we could actually fine-tune stuff on a toaster. Essentially, that is the Colab free-tier T4. For me, that was a big wow. Although it's super slow, it's okay. But on top of it, for instance, we have libraries called TRL and bitsandbytes and PEFT that enable all of this. And you can even do this with very advanced techniques like GRPO, for instance. Actually, you can train stuff on very small amounts of VRAM. You need to collocate and stuff, but with very small amounts of VRAM. And then recently, in case you missed it, we have acquired Llama CPP, kind of acquired Llama CPP. And then I kind of pivoted to the Llama CPP world, and I was like, wow, because the second wow moment for me was the fact that I could run OpenClaw and Hermes agent on very long-horizon tasks with Q1 3.6 quants, which wouldn't be possible before. I tried many models, actually, with that, and then Q1 3.5 was like, wow, it can do a lot of coding. It can fix its own harness and stuff. So yeah, those were the two big moments for me, actually.

2:55

Yeah. I think I'll speak to it more from a consumer point of view. This is back in 2023. I was packing around my own stuff, prior to me being at a Llama. But one of the coolest things was I wanted to run my AI on my computer. I was still in school at the time, broke, did not have a lot of money. And so I was like, okay, I need to run this for free somehow. So what's the best option? And so I came across Llama at the time, and I ran the model locally on my computer, and I think it was Llama 3 at the time as well, a very long time ago. Honestly, for me, that's kind of how I got into the whole world of local models. It's like, okay, this is actually workable. I can make it do things. It's able to output something, and back then it wasn't as good as it is now. As Merve said, Q1 3.6 is insane, and Gemma 4 and all these new models can have so much more capability baked into them. But for me, it was honestly the idea of being able to even just run something locally. And it started with Llama 3 with Ollama and being able to actually have it run and complete tasks. Prior to that, I'd built a lot of models before, and I knew how much work goes into building them. And running them has always been hard. So to me, quantization really makes that happen for a lot of people.

3:02

Yeah. I mean, I think everyone's experience is the same, right? Probably most of the people in this room, the first time that you run one of these models that only existed behind an API or a hosted environment on some cloud GPU somewhere, on your computer or on your gaming laptop or whatever it is, I mean, from that moment on, you're like, making these things small is pretty cool. There is something that we have to discuss, right? And anyone can jump in to follow up to this. Well, I'll pose it first to you, Daniel. You said models are 86% smaller, right? But they're not 86% dumber. How? That sounds absurd, right? And beyond how, how do you actually verify or think about verifying that this model has gone through the process of some form of compression, whatever it happens to be? How do I now determine that that model is not garbage, right? That we haven't chopped out 86% of its brain.

3:08

Yeah, that's a great question. So I think generally speaking, if you compress a model down by 86%, you would assume, if you randomly select parts of the model to compress, like delete or something like that, or set them to be, if you round it to the closest number, most likely it will not be 86% dumber. It will be 100% dumber. So if you do that methodology, that will not work. And so the main trick of language models is you should leverage the architecture of the language model itself. So language models generally have 36 layers, 50 layers, many, many layers. Each of the layers has different importance. For example, the first layer is actually very important, and then the last layer is also very important, but then the middle layers are kind of useless. And so the main reason why they're not that useful is because when you train a language model with, let's say, 1 trillion parameters, you have to use many, many tokens, right? So a language model can be trained with 30 trillion tokens, but we're still not there yet in terms of saturating all of the weights. Maybe in the future, once we train to 300 trillion tokens, maybe you can't do compression anymore. Maybe that's another topic. But at the current stage, the trick is 86% of the weights do not need to be there in the model because of the training algorithm, because of back propagation. Some of the weights are very close to zero, and you can literally just set them to zero. So that's one of the tricks. And also, you have to do layer-by-layer analysis. If you quantize layer one, what will happen to accuracy? If you quantize layer two, what will happen to accuracy? And so on, so on, so on. And so you can also think of this as a combinatorial optimization problem. You don't just do, okay, layer one and then do layer two. You also have to do 32 choose two layers or choose three layers. So it becomes very complicated. And so it's a very, then you get some combinatorial explosion problem. It's not just layer by layer. Within the layer, which number of the specific tensor is not quantizable or not? For example, there is something called Super Weights. There is a Super Weights paper which shows that if you quantize one number, just one in the entire model, your model becomes 20% dumber. And so you need to find this specific one number, and then you cannot quantize this. So there are very weird mechanisms in language models during training. And yeah, there's a whole lot of research going into quantizing models correctly. Yeah.

3:15

Anyone else with thoughts to add here? Yeah. So, okay, particularly your question was about how we evaluate. Exactly. Beyond just how do we make the model smaller, how do we know that it's not dumb now?

3:36

We just run all the benchmarks, mostly the benchmarks. So, model optimizer team, we publish a lot of quantized checkpoints on Hugging Face Hub. You can check the NVIDIA model optimizer space. So with that, we target less than 1% accuracy degradation overall on all benchmarks. Yeah. So that is, and they're like, we try to use simple strategies because that way we can push out models faster. On top of that, our learning is that strong, nicely designed number formats like FP4 preserve a lot of accuracy. And then we also see this disproportionate sensitivity not dumb now? We just run all the benchmarks, mostly the benchmarks. So, model optimizer team,

3:47

we published a lot of check quantized checkpoints on Hugging Face Hub. You can check the NVIDIA model optimizer space. So, with that, before we target for less than one percent accuracy degradation overall on all benchmarks. Yeah. So, yeah, yeah. So that is, and, and, they're, we, we, we try to use simple strategies because that way we can push out models faster. On top of that, learning is that the strong simple and strong, nicely designed number formats, like FP4, preserve a lot of accuracy. And then we also see this disproportionate sensitivity to some layers. For example, linear attention projection layers are very sensitive.

4:29

The KV, QKV layers are sensitive, whereas MOE, we, we, we by default keep them in FP4 while we put, say, other players in FP8 or BF16. We use this gradient-based sensitivity analysis. It runs a linear programming solver, but yeah, my main learning has been that FP4, or really nicely designed number formats, help a lot. And we are, yeah, and then a lot of benchmarking, the boring stuff, in which we spent a lot of time here. Maybe just, maybe just for everyone here who is I, I, I think hopefully most of us understand BF16 and FP8. What the hell is NVFP4? Oh, cool. Okay. So, FP4, the number is

5:15

there. It's a floating point 4-bit number, but the genius is, and it is a micro-block scaled number. So, micro-block scaling means every, you choose a group. So, in the case of FP4, you choose 16 elements and you can share one scaling factor, one, one extra FP8 number between these 16 elements. So, this was originally invented by bits and bytes, Tim Detmers. And yeah, so, and we adopted a similar but different design of four bits, but every 16 bits share one eight-bit value to scale them. And yeah, that has been that has been tremendous. It gives significant improvements over other formats. Yeah. Yeah. So, it's not as simple as we just make the number smaller.

6:00

It's a lot more going into it than that. It's interesting to hear from you guys. I think a lot of the time when we talk about compression or quantization or any technique that makes the model more accessible, right, we're talking about it through the lens of so that I can run it on my toaster, like you said, right? But is there any value to compression, or these techniques that exist for a business that presumably has access to, I would hope, more than a toaster? Pardon, Parth, maybe you want to? Yeah, I think for sure, you have a variance in hardware that an employee has versus hardware which scales up and

6:41

they're running their own clusters. And the way that at least we look at it from is you should be able to run whatever model you want locally on that individual's computer as well. And with compression in particular, you run through a lot of different challenges. So accuracy is for sure one of them. But actually making use of it through a harness, seeing what the end result is when you actually try it out, there are so many things that I feel can't be captured by a model optimizer or after quantizing it or certain benchmarks. And it's literally me running through, putting in a quad code or

7:27

something and running the model. It's like, no, it doesn't feel just right. So I think the benchmarks are a great indicator of pointing in the right direction. But having the model behavior kind of stay in line with that, I think is still kind of being worked on. And when it comes to businesses, you want to be able to give them the option to not just deploy their own models on their own infrastructure, but also empower them to be able to run it on individual machines. I think benchmarking oftentimes only works for the verifiable tasks rather than the vibe itself. And now that I think about it, actually, it could have been

8:10

cool to have a quant arena or something, but unfortunately, there was a paper last year by Singetal, I think from Courier, that showed that arenas are very much hacked. So that's also not the way. But anyway, I feel like most of the jobs still do not require a fable level thing. We are kind of in a bubble as software developers, right? So we are like, wow, fable. But at the same time, really most of the jobs do not require that. And on top of it, for the businesses, some of the things that we are saying, I feel like because I'm very open source pilled, it's super obvious, but at the same time, many people don't know about it. You

9:02

can just serve a lot of, you can have much more concurrency. At the same time, you can do in-house deployment and stuff. So increasing concurrency cuts the compute, and then, yeah, most of the time people just don't need that. But, for instance, I've been hearing a lot from Hugging Face users and stuff. And it's not only the quantization, but I know companies that actually distill mid-sized models to very small ones for re-ranking or whatever that doesn't require LLM outputs, cutting millions of costs. So it's not only quantization, but there is so much more that is out there for compression,

9:46

in my opinion. And it just creates a ton of value for businesses, and people aren't aware of it, because people in this room are interested, you read about it, you just assume that people know a lot about it, but actually they don't. And it's kind of shocking to me, but yeah. Another question that comes up all the time, right? We're talking about model compression. You were talking about GLM 5.2, right? And let's shrink it down as much as we can. Back there on the station, we're running a reap quant of, well, compress of GLM. But why would I do that when Nemo Tron 3 Nano exists? Or Gemma Small exists? Or Quentin Tiny exists? Why should I care about compression

10:33

when a lot of these model shops are kind of putting out small enough models that you can run them in their native precision or very close to their native precision? So there is a paper showing that if you want to do compression, it's actually most likely better use of resources if you train a ginormous model, then you quantize it down. And so there is this formula where they show comparing, for example, a very small model, like a 35 billion B float 16, so 16 bit, versus, say, four times bigger, like 120 billion at four bit, and which one's better, right? So essentially they're the same size in terms of disk space, but which one intelligence-wise is better? And from those

11:16

experiments, they show that the bigger model quantized to four bit is actually much better than a 35 billion 16 bit. So in general, most likely what will happen is we get bigger and bigger and bigger and bigger models. And, okay, currently now, GLM, 1.5 terabytes, oh, okay, it's not that big. But what happens if it's 15 terabytes? Then,

11:39

35 billion B flow 16, so 16 bit versus say four times bigger, 120 billion at four bit. And which one's better, right? So essentially they're the same size in terms of disk space, but which one intelligence-wise is better. And from those experiments, they show that the bigger model quantized to four bit is actually much better than a 35 billion 16 bit. So in general, most likely what will happen is we get bigger and bigger and bigger and bigger models. Okay, currently now, GLM, 1.5 terabytes, oh, okay, it's not that big. But what happens if it's 15 terabytes? Then, okay, we must do quantization. We must do compression. This will not fit, and not even enterprises can now service it. And so local models, you must do compression if they're getting bigger and bigger and bigger. And to extract any value out of it, okay, I guess the DJX station has a lot of memory. I guess that's very useful. But once we have 10 trillion parameter models, that will also not fit. And so we need to do compression, quantization for that. But in general, the small ones are very useful. But I think the bigger ones compressed down are slightly more useful. But there's also another trick. You can do model routing. For example, for the small ones, you can use the big ones for planning and then your execution with the small ones. But the small ones are still better because they're much faster. So even if you have a big one and you compress it down, you probably get five tokens to 10 tokens per second if you don't have enough GPU power. And for the small ones, you can get 200 tokens per second. So depending on your use case, you also have to consider speed, throughput, what is the GPU that you have, and stuff like that. Yeah, I guess. Very cool. Very cool, obviously. So Parth here with Ollama. I imagine most of the people here know Ollama. If you're not using it, give it a try. This is one of those topics where, is Ollama useful in a world where we don't have this kind of compression, right? How much is the fact that we can compress intelligence to run on my Mac or my home workstation, right? How much is that ecosystem important to people or businesses or organizations like Ollama? Yeah, I think it's pivotal, obviously, in order to have these models be so viable running on your personal computers and being able to actually make work of them, as whoever was saying, Quin36 just running on computer, running open claw, being able to fix its own harness. That's really possible through some level of compression. Now there are other techniques which some of the model apps will do, especially while releasing smaller size models, things like QAT or, with GPT OSS, it was MXFP4, which is a different format. So I do think as long as there's a demand for being able to run personal language models, even if compression didn't exist, there will be analogous techniques, or maybe model training would look a little bit different. Obviously, everything comes at a trade-off, as Daniel was mentioning. You can have a really large model being condensed down into something smaller so you can run it versus even just having a full precision model but smaller in parameter size. And they both come with different trade-offs.

11:45

And, Marav, how do you see the open source community help drive this, right? I think quantize, I mean, we all remember Tim Detmers, we all remember QLaura. That image, right, where there's a big stack of things that would fall over if it weren't for bits and bytes holding it all up. How have you seen the open source community flourish or grow around not just quantization but compression more generally? I wanted to add something to what Daniel says. By the way, if you don't know about their work, they have immense pipelines to actually do the quants and then verify them and stuff. So I highly recommend you check out OnSloth, which is, in my opinion, the best quants you will find over there, yes. Oh, thank you. I wanted to say, for instance, to add to your point, two years ago, I think two years ago we trained small VLM, where there were no small vision language models. There was lava, and then there was a jump to big models, and then nobody was training smaller ones. We figured out that training, we trained 1B, 500M, and 256M, and 256M was not good, and 500M was actually somewhat good, and it was actually able to run on your iPhone, which was shocking to me. And I noticed that over time it feels like people have been releasing smaller models, which is great, but if you have advanced models and then everybody is just releasing mid-sized models, you can quantize them, and I feel like that's the biggest value that quantization somewhat offers. And, for instance, coming to your question about the community and stuff, this year what happened was that for the Q1 3.6 release, Q1 team asked the people, okay, we are going to release only one checkpoint, which is kind of heartbreaking, and everybody said the mid-sized ones. And this kind of goes to show the value of how quantization is adopted, and everybody is using those checkpoints and then quantizing them instead of asking for smaller models. I feel like you get more and more and more intelligence, and then over time you shrink them, and then suddenly the bigger models make the bigger shifts in terms of how much you can quantize them, and still you do not lose the information, and the quants become smaller than the bigger models of yesterday, which to me is phenomenal.

11:51

As for the community, can you rephrase your question again? I think you answered it, actually. Yeah. It's a great job. Yeah. Thank you.

12:14

I mean, I gotta ask this question since you're up here with us. So NVIDIA NemoTron releases NVFP4 with every release. NVFP4 is, as described, basically a numeric format, right? But the idea is that it's a smaller version of the model, and it retains a lot of accuracy, right? Which is the whole idea of compression and quantization. I'd just like to hear, how hard is it to do that, right? So we kind of heard Daniel talk about it from this dynamic quantization strategy, right? It already sounds very difficult. How difficult as an engineering challenge is it to make the model small without it shitting the butt? I see. Okay. So, if you're doing post-training quantization, which is you take the model, the release model, the BF16 model, and do post-training quantization on it, it is fairly easy to do. So we quantized to FB4 models, like a large GLM or those trillion-size models. We have Blackwell nodes, so in a couple of hours the quantized checkpoint is ready, but then our pain starts there because now we have to evaluate these models, match the model card. Actually, that is where we spend a lot of our time. Now, coming to training-based methods. Okay. So, for large models and medium-sized models, I would say 20 billion

12:20

and do post-training quantization on it, it is fairly easy to do. So we quantized to FB4 models, a large GLM or those trillion-size models. We have Blackwell nodes, so in a couple of hours, the quantized checkpoint is ready, but then our pain starts there because now we have to evaluate these models, match the model card. Actually, that is where we spend a lot of our time. Now, coming to training-based methods.

12:29

Okay. So, for large models and medium-sized models, I would say 20 billion parameters plus or 30 billion parameter plus dense size, this PDQ usually works out of the box, with some selective quantization, the heuristics Dan mentioned, right? We use some heuristics such as sparse, some always can be aggressively quantized. And we also have this auto-quantized, which uses this automatic sensitivity analysis and knapsack solver. So all those, right? So for medium and large models, PDQ works out of the box, very easy, relatively super easy to do. Then if it is a smaller model, say less than 20-bit size model, we have to do some quantization-aware distillation, et cetera, to recover accuracy. That is, yeah. So training-based methods are a little bit more painful. They're becoming more painful, especially with these reasoning models. So you need to have the original data set, and it's not just about the original data set. Internally, we have the original new modern data set, but even then it is a pain because these are multi-trained, multi-stage trained RL models, done with. Now models are being trained with multiple teacher distillation where each teacher is an expert in coding or reasoning or something. So it becomes really difficult to get good training data to train the model with QAD to recover accuracy. If we train, if we do QAD with wrong data, it most commonly breaks the model rather than helping it. So again, it sounds hard. It's still an engineering challenge. No, I want to stay corrected. I want to encourage people to use quantized model. Mostly PDQ will work if you are looking at 30B. It should work out of the box. And there are so many tools right now, model ops, unsloth has tools, Huggy Face and the ecosystem, tons of tools, to make these big models small.

12:35

Something that I want to pick your guys' brains about is architecture for models in the llama era of, let's call it open AI, right, which is every model was the same. The architecture was basically the same, the kinds of things that you saw in the guts of the model were the same. And now we're entering a very cursed era of technology, right, where everyone's doing the architecture just a little bit differently. We're using this hybrid attention, they're using this linear attention variant, we're exploring a lot, and we're trying a bunch of new things. How has that shifted the difficulty of quantization, now that it's not just one problem repeated with different numbers of layers? Is that something that makes it difficult to keep up with? Maybe from Ola's perspective, how hard is it to keep up with this?

12:39

Yeah, there's kind of two facets to it, I'd say. The first is actually just the model implementation itself. There's been times where a model lab would come to us early, and we're kind of working with them to implement the model beforehand, and we do this, except they come, and sometimes there's five different variations of a model, and we need to have them all working. So the implementation's one side of it, and I'm sure other people also have to put a lot of work in it, but the other side is you kind of have to run through the quantization bit and see which one works best, and at Olama, we kind of do a UX thing of giving a default model quantization for most things, for most models, and a big part of that is actually us spending the time, one, quantizing it, but then seeing if it actually works well with different harnesses, and is it actually usable after. And so we find that sometimes when you have a very small parameter-sized model, you don't get great quantization after that, so we sometimes leave it in higher precision as a default, just because we actually want people to have a better experience versus quantizing it down to too little of a precision and not having the model actually work well. So, it's always a challenge, both from the implementation perspective, but then actually you're running into the quantization bit and making sure it still works correctly.

12:44

And Daniel, is it harder to do quantization now that everything is some hybrid or linear variant, and it's also all sparse or variations on sparse MOE? How much harder is it today than it was when everything was just long all the way down?

12:49

Yeah, in the olden days, every model was dense. Yep. The transformer plus plus, so it's called transformer plus plus architecture, it's an old transformer, okay, plus RMS, plus some extra tricks. That's called transformer plus plus. And then now it's like, oh my, it's like transformer plus plus plus, and then plus this thing, plus that thing, minus this thing, minus that thing, a different activation function, linear attention here, sliding window, window attention, how many layers are sliding window, how many layers are global, or let's delete global, do something else, blah, blah, blah, blah, blah. Everyone likes to do their own thing, and they like to compare, okay, this one does better for long context, this one does worse for long context, or something like this. So, there's always reasons why they like to change the architecture. Some folks even change some of the layer norm epsilon, change 1e minus 5 to 1e minus 6, okay, which one's better, and so on. So they do a lot of ablations, they do a lot of testing, this one seems to be better than this one. And yes, it has complicated compression and quantization dramatically. You have your old heuristics, okay, this works well for this model. But then when you go to the MOE world, oh, you can quantize the MOE layers to one bit, and it doesn't break. But then when you go to the linear attention world, you cannot quantize the linear attention layers. So, if you quantize the linear attention layers, we found that if you quantize the linear attention layers, okay, it looks like it's doing good, but then when you do long context benchmarks, when you actually use the model in real production, it becomes gibberish. And so there are some layers you cannot quantize with these new architectures, some layers you can quantize very low to one bit, you can even delete some layers if you like. And so it's like these new architectures complicate the process. But to be honest, very happy with this because we need more different architectures. We don't want everyone to be like,

12:53

and it doesn't break. But then, when you go to the linear attention world, you cannot quantize the linear attention layers. So, if you quantize the linear, we found that if you quantize the linear attention layers, okay, it looks like it's doing good, but then when you do long context benchmarks, when you actually use the model in real production, it becomes gibberish. And so, there are some layers you cannot quantize with these new architectures, some layers you can quantize very low to one bit, you can even delete some layers if you like. And so these new architectures complicate the process. But to be honest, I am

13:26

very happy with this because we need more different architectures. We don't want everyone to be thinking the same way. And open source has been, the open model era has been, there are so many different architectures, and it's very good to have a variety of different opinions and architectures, yeah. That's dope, yeah, hell yeah. We got about five minutes left. Question. Can we ask questions? I want to get one more question with these guys, and then yes. Okay, so compression's an art, not a science right now. Well, it's an artisanal science, let's say. But where is it going? What does compression look like in six months, in 18 months? Maybe we'll just start with Merve,

14:07

and then we'll do a loop around. Can we start from someone? Yeah, sure, we'll start. And would you like to go? Okay, yeah, sure. So, I think the way what we've seen so far is compression being so critical for any model launch that comes around, so it seems that model labs are starting to think more about it as well. The folks at Oncelot do a phenomenal job. Fingers crossed they keep putting some great stuff out. But more than that, I think because the model architecture is changing and this awareness that model labs are having, we're starting to see more things like QAT come out from the labs themselves. And I'm sure

14:40

there are ways to push even that further, but I think it's going to be a mix of the labs becoming a little bit more interested, but also the community doing the great work they already have been and pushing that further. Oh, Tan. And go. Now that I thought about the question. So, what I think is most of, previously we were just running models on servers and stuff, but I think we can finally push the edge because there was this increasing demand for privacy and everything, especially for sensitive data, personal data, and so on. So, I see compression being a super hot topic, and we are developing

15:22

a lot of cool stuff with Lama CPP. So, I would like it if you could stay tuned for that, so that it runs everywhere. So, I think I see the, previously I saw that we were constantly scaling the model, model parameters, and then the data diversity, and then going down. I see it as going even further to the phones and stuff because it wasn't working. We tried a lot, and it wasn't working. I see intelligence going to phones thanks to the quants and everything. So, yeah. Going to phones. Let's go. Yeah. Okay. So, the way I think about compression, the whole space is going to go broader. So, we've focused this discussion mostly on weight compression. So, okay. Going back to

16:10

FP4 once again. So, it also does weight compression and math acceleration because it does the GEM in four bit. So, yeah. So, in terms of weight compression, we might be able to go to maybe two or three bit more. But in terms of quantization alone, we might be close, close to the Pareto optimality. Yeah. Let's see. Then there are more types of compression. So, KV cache compression, right? So, yeah. So, people are still mostly using 8-bit, 4-bit. Large models retain that quality very well, but smaller models, we see some drop. So, yeah. So, looking forward to KV cache, maybe compaction plus quantization, pushing long-horizon reasoning, broader than sparsity. So, okay. So,

17:05

by the way, sparsity has been part of Nvidia hardware, but it has not been broadly adopted. That is because quantization does not degrade accuracy that much, but sparsity causes accuracy degradation a bit more. So, in Rubin, there is this cool feature called dynamic activation sparsity. So, it can improve attention math, et cetera. So, yeah. So, looking forward to that. And then coming back to all these heterogeneous architectures, right? So, with each release, particularly the attention architecture is getting more and more complex. DeepSeek started it. I blame them. They started with MLA. Now, sparse attention,

17:53

index attention, a lot of skips, softmax. Yeah. So, it's getting broader. And, yeah. So, it is part of this process where we make models steeper, and they are having compounding effects, right? Perfect. So, we're going to quantize more things. More is basically the idea. That's pretty dope. I don't mind that. Take us home. Yeah. So, I think the world, we're going to get more and more bigger models, and it would be very cool if we can run them locally on our phones, on our laptops with our GPUs. Imagine a world where the best frontier models will be able to run on your computers. That would be so cool. Now you can control your own

18:45

destiny. You do not need to be controlled by some model labs. And now you can do whatever you like on your computer, right? You can do your own fine tuning. You can customize it. Everything becomes yourself. You own it. And so, I see a world where, in the future, all of the intelligence will be fully democratized via GPUs, via the phone, any single hardware. And you will have very capable AIs on your laptop, on your local devices, running, and it's also going to be efficient, right? You don't want your computer to lose battery and die. But I feel like with all of these techniques, we can have a future where AI is fully democratized.

19:31

And that's where I see it. Yeah. It's not a bad future. Again, we can't do it without all of you in the room. So, big round of applause for the panel here. Thank you so much, guys. We have one question. We'll do one question. All right. So you talked about all the different permutations of craziness that are happening in the model stages. The proliferation of various types needs to do with the press, right? But we often find some weird Frankenstein model out there that you want to try out, but you have no clue how well it actually performs on the original benchmark of the model before all the conversions occurred.

20:18

So I'm just curious, is there any good resource out there? Maybe this is about 100% of the benchmark, again, on the model. Again, what they perform after all the freakish things to see how well they're going. And if there's, where we can find matrices or summaries of what is the best modified model for this or for that, and et cetera, et cetera, right? I'd love a resource like that. Anybody who... Yeah, if you find something like that, let me know, yeah. Yeah, I think that's it. Basically, the question is, what is the resource I can look at to find out

20:58

you find some weird Frankenstein model out there that you want to try out, but you have no clue how well it actually performs on the original benchmark of the model before all the conversations occurred. So I'm just curious, is there any good resource out there? Maybe this is about 100% of the benchmark. Again, on the model. Again, what they perform after all the freakish things for them to see how well they're going. And if there's where we can find matrices or summaries of what is the best modified model for this or for that, et cetera, et cetera, right? I love a resource like that. Anybody who... Yeah, if you find something like that, let me know, yeah.

21:21

Yeah, I think that's it. The question is, what is the resource I can look at to find out what is the of the suite of crazy quants or compresses that exist for a model? How do I find the ones that are good at the things that I care about? Is anyone doing that? I think right now, it is not being done comprehensively. We do have some... So for some models, when we do our dynamic quantizations, we do release benchmarks. And we do not do... So generally, what our view is, accuracy benchmarks can be very complicated, because you have to do sampling, how many trials you need to do, and you have to average.

21:38

So there is another better method in our view, KOR divergence. So KOD is the distance between the unquantized version, which is BFloat16, annual quantized version. And you can calculate some sort of distance between the quantized version and the unquantized version. And your goal is to make the distance zero and the size smaller. Do you do it over output logics? Yes. So you check the output logics, you pass some sort of collaboration data, you shove it in, and then you have some... You check the output logics between the BFloat16 and then the main, the quantized version. And then the goal is make this distance zero. Yeah, yeah. And you make the model smaller.

21:55

That's K-L-D. Yeah, that's K-O-D. And so the LD, you can check out the paper. Accuracy is not all you need if you want to learn more about that one. That's it for us guys. Thank you so much again to our excellent panel. Thank you. when a lot of these model shops are kind of putting out like, uh, small enough models that you can run them in kind of their native precision or, or very close to their native precision? So there is a paper showing that if you want to do compression, um, it's actually most likely better use of resources. If you train a ginormous model, then you quantize it down. Um, and so there is this,

22:30

like formula where they show, um, comparing for example, a very small model, like a 35 billion, um, you know, 35 billion B flow 16, so 16 bit versus say like, um, four times bigger, like 120 billion at four bit. Um, and which one's better, um, right? So essentially they're the same size in terms of disk space, but which one intelligence wise is better. And from those experiments, they show that the bigger model quantized the four bit is actually much better, um, than a 35 billion 16 bit. So in general, most likely what will happen is we get bigger and bigger and bigger and bigger models. Um, and you know, like, okay, currently now, you know, GLM, you know,

23:07

1.5 terabytes, you know, oh, okay, it's not that big. Um, but what happens if it's 15 terabytes? Then, okay, we must do quantization. We must do compression. This will not fit and anyone's not, not even enterprises can now service. And so like local models, you know, you must do compression if they're getting bigger and bigger and bigger. Um, and you know, to extract any value out of it and you know, okay, I guess the DJX station has a lot of memory. I guess that's very useful. Um, but you know, once we have 10 trillion parameter models, that will also not fit. And so like, you know, we need to do

23:38

compression, you know, I guess quantization for that. Um, but in general, the small ones are very useful. Um, but I think the bigger ones compressed down, um, is slightly more useful. Um, but there's also another trick. You can do model routing. For example, for the small ones, you can do it for the, for example, you can use the big ones for planning and then your execution with the small ones. Um, but the small ones are still better because they're much faster. So even if you have a big one and you compress it down, Oh, you probably get like five tokens to 10 tokens per second if you don't have enough GPU

24:08

power. Um, and for the small ones, you can get 200 tokens per second. Um, so I guess like depending on your use case, um, you also have to consider speed throughput, you know, what is the GPU that you have and stuff like that. Um, yeah, I guess. I mean, very cool. Very cool, obviously. Uh, so Parth here with Ollama. I, I, I imagine most of the people here know Ollama. If you're, if you're not using it, give it a try. Uh, you know, this is one of those topics where is Ollama useful in a world where we don't have this kind of compression, right? Like how, how much, uh, is the fact that we can compress

24:48

intelligence to kind of run on like my Mac or my, uh, you know, my, my, my home workstation, right? Like how much is that, uh, ecosystem important to, to, you know, people or businesses or organizations like Ollama? Yeah, I think, uh, it's pivotal obviously for, you know, in order to have like these models be so viable running on your personal computers and being able to actually make work of them as whoever was saying, you know, Quin36 just running on computer, running open claw, being able to fix its own harness. That's really possible through some level of compression. Now there are,

25:25

you know, other techniques which kind of some of the model apps will do, especially while releasing smaller size models, um, things like QAT or, you know, with GPT OSS, it was MXFP4, which is a different format. So I do think as long as there's a demand for being able to run kind of personal language models, even if compression didn't exist, there will be analogous techniques or, you know, maybe model training would look a little bit different. Um, obviously that comes, everything comes at a trade-off, um, as Daniel was mentioning, you know, you can have a really large model being condensed down into

25:59

something smaller so you can run it versus even just having, um, a full precision model but smaller in parameter size. And they both come with different trade-offs. And, it, Marav, like how, how do you see the open source community help drive this, right? Like, uh, I, I think, quantize, I mean, we all remember Tim Detmers, we all remember, uh, QLaura, you know, is like, uh, that, that, that, that image, right, where there's a big stack of things that would fall over if it weren't for bits and bytes holding it all up. Like, like, how have you seen the open source community flourish or grow around, uh, not just quantization but compression more generally?

26:36

I wanted to add something to what Daniel says. By the way, if you don't know about their work, they have, like, immense pipelines to actually do the quants and then verify them and stuff. So, like, I highly recommend you to check out OnSloth, which is, like, in my opinion, the best quants you will find over there, yes. Oh, thank you. Um, I wanted to say, for instance, like, to add to your point, like, two years ago, I think two years ago we trained, like, small VLM, where there were no small vision language models, there was lava, and then there was a jump to big models, and then nobody was training, like, smaller ones. Um, we figured out that, like, um, training,

27:14

like, we trained, like, 1B500M and 256M, and, like, 256M was not good, and, like, 500M was actually somewhat good, and it was actually able to run on your iPhone, which was, like, shocking to me. And, um, I noticed that over time it feels like, uh, people have been releasing, like, smaller models, which is great, but, like, if you have, like, advanced models, and then everybody is just releasing mid-sized models, you can quantize them, and I feel like that's the biggest value that quantization somewhat offers, and, for instance, like, to coming to your question about the community

27:55

and stuff, for instance, this year what happened was that for the Q1 3.6 release, Q1 team asked the people, okay, we are going to release only one checkpoint, which is kind of heartbreaking, and everybody said the mid-sized ones, and this kind of goes to show the value of, like, how quantization is adopted, and everybody is, uh, using those checkpoints and then quantizing them instead of asking for smaller models. I feel like you get more and more and more intelligence, and then over time you shrink them, and then suddenly it's, the, the, the bigger models make the bigger shifts in terms of,

28:31

like, um, how much you can quantize them, and still you do not lose the information, and the quants become smaller than the bigger models of yesterday, which to me is, like, phenomenal. As for the community, can you rephrase your question again? I think you answered it, actually. Yeah. It's a great job. Yeah. Thank you. I mean, you know, I gotta ask this question, uh, since you're, you're on the, you're up here with us. So, uh, NVIDIA NemoTron releases NVFP4 with every, uh, release. NVFP4 is, uh, as described, uh, basically a numeric format, right? Uh, but the idea is that it's, like, a smaller version of the model,

29:14

and it retains a lot of accuracy, right? Uh, which is the whole idea of, of compression and quantization. Uh, I, I just like to hear, like, how hard is it to do that, right? Like, so we, we kind of heard Daniel talk about it from this dynamic, uh, quantization strategy, right? Uh, it's, it already sounds very difficult. Like, how difficult as an engineering challenge is it to, like, make the model small without it shitting the butt? Uh, I see. Okay. So, if you're doing post-training quantization, which is, uh, you take the model that, uh, that, uh, like, the release model, the BF16 model,

29:50

and do, uh, post-training quantization on it, it is fairly, uh, like, you know, easy to do. Uh, so we quantized to FB4 models, like, uh, a, uh, large, uh, GLM or like, you know, those, like, trillion size models, uh, you need, uh, like, uh, we have Blackwell nodes, so we, uh, in, in a couple of hours, uh, the quantized checkpoint is ready, but then our pain starts there because now we have to evaluate these models, match the model card. Uh, actually, that is where we spend a lot of our time. Now, coming to, uh, uh, uh, training-based methods. Okay. So, for large models and medium-sized models, I would say, like, you know, 20 billion

30:33

billion parameters plus or 30 billion parameter plus dense size, uh, this PDQ usually works out of the box, uh, with, uh, some selective quantization, like the heuristics, like, Dan mentioned, right? Like, we use some heuristics such as, uh, sparse, some always, uh, can be aggressively quantized, uh, like, yeah. And we also have this auto-quantized, which uses this automatic sensitivity analysis and, uh, knapsack solver, yeah. Yeah. So all those, right? So for medium and large models, PDQ works out of the box, very easy, relatively super easy to do. Uh, then if it is smaller model,

31:10

say, like, you know, less than 20-bit size model, we have to do some, uh, quantization-aware distillation, et cetera, to recover accuracy. That is, yeah. So training-based methods are little bit more painful. They're becoming more painful, especially with these reasoning models. So you need to have the original data set, and it's not just about the original data set. Internally, we have the original new modern data set, but even then it is a pain because these are multi-trained, uh, multi-stage trained RL models, like, for, done with, now models are being trained with, uh, multiple teacher

31:42

distillation where each teacher is, like, an expert in coding or reasoning or something. So it becomes really difficult to, uh, get, uh, good, uh, training data to, uh, train the model with, with QAD, uh, yeah, to recover accuracy. If we, if we train, if we do QAD with wrong data, it, uh, it, uh, it most commonly breaks the model rather than helping it. Yeah. Yeah. So again, it sounds hard. Uh, I mean, it's still an engineering challenge. No, I want to, no, I want to stay corrected. I want, uh, to encourage people to, like, you know, use quantized model. Uh, mostly PDQ will, uh, work,

32:20

like, you know, if you are, like, looking at, like, 30B, yeah, it should work out of the works. And there are so many tools right now, like, model ops, like, uh, unsloth has tools, uh, Huggy Face and the ecosystem, tons of tools, right, to, to make these big, these big models small. You know, you know, something that I want to pick your guys' brains about is, architecture for models in, like, the llama era of, let's call it open AI, right, which is, like, every model was, like, the same. You know, the architecture was basically the same, the kinds of things that you saw in the guts of the model were the same. And now we're entering,

32:56

like, a very cursed era of technology, right, where, uh, you know, everyone's doing the architecture just a little bit differently. We're using this hybrid attention, uh, they're using this, uh, you know, linear attention, uh, you know, you know, variant, uh, you know, uh, we're, we're exploring a lot, and we're trying a bunch of new things. How has that, like, shifted the difficulty of quantization, now that it's not just, like, one problem repeated with different, uh, you know, you know, numbers of layers? Like, is that something that is, makes it difficult to keep up with? Maybe, maybe, you know,

33:31

from, from Ola's perspective, like, how hard is it to keep up with this? Yeah, um, there's kind of two facets to it, I'd say. Uh, the first is actually just the model implementation itself. Um, there's, you know, been times where a model lab would come to us early, and we're kind of working with them to implement the model, uh, beforehand, and we do this, except, you know, they come, and sometimes there's, like, five different variations of a model, and, you know, we need to have them all working. So, the implementation's one side of it, um, and I'm sure other people also have to put a lot of work in it, but the other, kind of, other side is you kind of have to run

34:07

through the quantization bit and seeing, you know, kind of which one works best, and at Olama, we kind of do a UX thing, of, like, giving a default, uh, model quantization for most things, uh, for most models, um, and a big part of that is actually, you know, us spending the time, one, quantizing it, but then seeing if it actually works well with, like, different harnesses, and, you know, is it actually usable after, and so we find that sometimes when you have a very small parameter-sized model, um, you don't get, like, great quantization after that, um, so we sometimes leave it in higher precision, um, as a default, just because

34:45

we actually want people to have a better experience, uh, versus, you know, quantizing it down, quantizing, yeah, quantizing it down to, too little of a precision, um, and not having the model actually work well. So, it's always a challenge, um, both from, like, the implementation perspective, but then actually, you know, you're running into the quantization bit and making sure it still works correctly. And, and, and, Daniel, like, is it harder to do quantization now that, like, everything is some hybrid or linear variant or, and it's also all, like, sparse or variations on sparse MOE? Like, how much harder is it today than it was when it was, everything was just

35:22

long all the way down? Yeah, like, in the olden days, you know, every model was dense. Yep. The transformer plus plus, so it's called transformer plus plus architecture, you know, it's an old transformer, okay, plus RMS, plus some extra tricks. That's called transformer plus plus. And then now it's like, oh my, it's like transformer plus plus plus, and then plus this thing, plus that thing, minus this thing, minus that thing, a different activation function, linear attention here, sliding window, window attention, you know, how many layers are sliding window, how many layers are global, or let's delete global, do something else, blah, blah, blah, blah, blah. Um, you know,

35:54

everyone likes to do their own thing, and, you know, they like to compare, you know, like, okay, this one does better for long context, you know, this one does worse for long context, or something like this. So, there's always, like, reasons why they like to change the architecture. Um, you know, some folks even change some of the, you know, layer norm epsilon, like, you know, change 1e minus 5 to 1e minus 6, okay, which one's better, and so on. So, they do a lot of ablations, you know, they do a lot of testing, you know, this one seems to be better than this one. Um, and yes, it has complicated compression and

36:22

quantization dramatically. Um, you know, you have your old heuristics, okay, this works well for this model. But then when you go to the MOE world, oh, you can quantize the MOE layers to, like, one bit, and it doesn't break. Um, but then, you know, when you go to the linear attention world, you cannot quantize the linear attention layers. So, if you quantize the linear, we found that if you quantize the linear attention layers, okay, it looks like it's doing good, but then when you do long context benchmarks, you know, when you actually use the model in real production, it becomes gibberish. Um,

36:49

and so, like, there are some layers you cannot quantize with these new architectures, some layers you can, you know, quantize very low to, like, one bit, you know, you can even delete some layers if you like. Um, and so it's like these new architectures complicate the process. Um, but to be honest, very happy with this because we need more different architectures. We don't want everyone to be, like, thinking the same way. And, you know, open source has been, you know, the open model era has been, like, there's so many different architectures, and it's very good to have, like, you know, a variety of

37:17

different opinions and architectures, yeah. I mean, that's dope, yeah, hell yeah. Uh, I, we, we got about five minutes left. Question. Can we ask questions? Uh, I, I, I want to get one more question with these guys, and then yes. Uh, okay, so, compression's an art, not a science right now. Uh, well, it's a, it's a, it's an artisanal science, let's say. Uh, but where is it going? Uh, you know, what, what does, what does compression look like, uh, in six months, in 18 months? Maybe we'll just go, we'll start with Merve, and then we'll, we'll, we'll do a loop around. Can we start from someone? Yeah, sure, we'll start.

37:53

And would you like to go? Okay, yeah, sure. Um, so, I think the way, what we've seen so far is compression being so critical for any model launch that comes around, so it seems that model labs are starting to think more about it as well. Um, the folks at Oncelot do a phenomenal job. Fingers crossed they keep putting some great stuff out. Um, but more than that, I think because of, you know, the model architecture is changing and kind of this awareness that model labs are having, we're starting to see more things like QAT, uh, come out from the labs themselves. And I'm sure, like,

38:29

there's ways to push even that further, but I think it's going to be a mix of the labs kind of becoming a little bit more interested, but also the community kind of doing the great work they already have been and pushing that further. Oh, Tan. And go. Now that I thought about the question. So, um, what I think is most of, so previously we were just running models on servers and stuff, but I think we can just finally push the edge because, like, there was this increasing demand for privacy and everything, especially for, like, sensitive data, personal data and so on. So, like, I see compression being a super hot topic and, um, we are developing

39:10

a lot of cool stuff with Lama CPP. So, I would like, I would like it if you could stay tuned for that, um, so that it runs, like, everywhere. Um, so I think I see the, the, so previously I saw that we were constantly scaling the model, model parameters and then the data diversity and then going down. I see it as, like, going even further to the phones and stuff because it wasn't working. Like, we tried a lot and it wasn't working. I see intelligence going to phones thanks to the quants and everything. So, yeah. Going to phones. Let's go. Yeah. Okay. So, the way I think about compression, uh, the whole space is going to go

39:54

more broader. So, we've, uh, focus this discussion mostly on weight compression. So, uh, okay. Going back to FP4 once again. Uh, so it also do weight compression and, uh, math acceleration because it, it, it do the, uh, gem in four bit. So, uh, yeah. So, for in terms of weight compression, we might be able to go to, like, maybe two or three bit more. But in terms of, uh, quantization alone, we might be, like, close, like, close to the Pareto optimality. Uh, yeah. Let's, let's see. Uh, then there are more, like, you know, more type of compressions. So, KV cache compression, right? So, yeah. So, people are still

40:37

mostly using 8-bit, uh, 4-bit, large models retain that quality very well, but smaller models, we see some drop. So, yeah. So, looking forward to KV cache, uh, maybe compaction plus quantization, uh, like, pushing, uh, the long, uh, horizon reasoning, uh, broader than sparsity. So, uh, so, okay. So, by the way, sparsity is, uh, has been part of Nvidia hardware, but it has not been, uh, like, broadly adopted. Uh, that is because, um, quantization does not degrade accuracy that much, but sparsity causes accuracy degradation a bit more. So, in Rubin, there is this cool feature called dynamic,

41:18

uh, activation sparsity. So, it can, uh, improve attention math, et cetera. So, yeah. So, looking forward to that. And then coming back to, like, you know, all these heterogeneous architectures, right? So, with each release, the, uh, particularly the attention architecture is getting more and more complex. DeepSeek started it. I blame them. They started with MLA. Now, like, you know, sparse attention, index attention, like, you know, a lot of skips, uh, softmax. Yeah. So, it's getting broader. And, yeah. So, it is part of this process where we make models steeper and, and they are, they are having

41:51

compounding effects, right? Perfect. So, it's, so, we're going to quantize more things, uh, more is basically the idea. That's, that's pretty dope. I don't, I don't mind that. Take us home. Yeah. So, I think, like, the world, we're going to get more and more bigger models, um, and, you know, it would be very cool if we can run them locally on our phones, you know, on your laptops with your GPUs, um, and you, like, imagine in a world where, you know, the best frontier models will be able to run on your computers. You know, that would be so cool. Um, you know, now you can control your own

42:23

destiny. You do not need to be, you know, controlled by some model labs. And now you can do whatever you like on your computer, right? You can do your own fine tuning. You can customize it. Um, you can, you know, everything becomes yourself. You own it. Um, and so, like, you know, I see a world where the world in the future, all of the intelligence will be fully democratized, you know, via GPUs, via the phone, you know, any single hardware. And you will, you will have, you know, very capable AIs on your laptop, on your local devices, running, you know, and it's also going to be efficient, right? You

42:57

don't want your computer to, like, you know, lose battery and die. Um, but I feel like, you know, with all of these techniques, um, you know, we can have a future where AIs fully democratized. And that's, yeah, that's where I see it. Yeah. It's not a bad future. Uh, again, we can't do it without any, all of you in the room. So big round of applause for the, the, the panel here.

43:18

Thank you so much, guys. We have one question. We'll do one question. All right. So you talked about all the different permutations of craziness that are happening in the model stages. The proliferation of various types needs to do with the press, right? Um, but we often, you find some weird Frankenstein model out there that you want to try out, but you have no clue how well it actually performs on the original benchmark of the model before all the conversations occurred. So I'm just curious, like, is there any good resource out there? Maybe this is about 100% of the benchmark. Again, on the model. Again, you know, what they perform after all the freakish things

44:02

for them to see how well they're going. And if there's like, where we can find like matrices or summaries of like, what is the best modified model for this or for that, and et cetera, et cetera, right? I love a resource like that. Anybody who... Yeah, if you find something like that, let me know, yeah. Yeah, I think that's it. Basically, the question is, what is the resource I can look at to find out what is, you know, the of the suite of crazy quants or compresses that exist for a model? How do I find the ones that are good at the things that I care about? Is anyone doing that? I think right now, it is not being done comprehensively.

44:40

We do have some... So for some models, when we do our dynamic quantizations, we do release benchmarks. And we do not do... So generally, what our view is accuracy benchmarks can be very complicated, because you have to do sampling, how many trials you need to do, and you have to average. So there is another better method in our view, KOR divergence. So KOD is the distance between the unquantized version, which is BFloat16, annual quantized version. And you can calculate some sort of distance between the quantized version and the unquantized version. And your goal is to make the distance zero and the size smaller. Do you do it over output logics?

45:16

Yes. So you check the output logics, you pass some sort of collaboration data, you shove it in, and then you have some... You check the output logics between the BFloat16 and then the main, you know, the quantized version. And then the goal is make this distance zero. Yeah, yeah. And you make the model smaller. That's K-L-D. Yeah, that's K-O-D. And so the LD, you can check out the paper. Accuracy is not all you need if you want to learn more about that one. That's it for us guys. Thank you so much again to our excellent panel. Thank you.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note