I'm here to talk about what is new in inference engineering. So, hi, I'm Philip, and I'm here because I wrote a book. This is my third year at the AI Engineer World's Fair. This is my favorite conference in the entire world. It's the highlight of the calendar every single year. I really got my start as a speaker and as an engineer here in 2024. I came back in 2025 and did a bunch of stuff. I'm here again. I love it here, and I'm very thankful to the organizers for always having me. I wrote this book called Inference Engineering. We published it three or four months ago, and I've just been overwhelmed by the response. We've done more than 11,000 paper copies.
We're coming up on 30,000 digital copies. And 24 million people around the world, or 24 million Twitter accounts, so we'll see how many people that actually is, have seen something out of inference engineering. And with all of this great reception, there has been one question that people have been asking me. Why in the world would you do this? Why would you write a book about something that's changing so fast? Well, I believe that a lot of the principles of inference engineering at this point have been pretty solidified, and there's a lot that we can learn and repeat over generation and generation of model.
But today, I'm here to talk about what's new in inference engineering. This is the first public addendum of all new information since the book came out. We are going to review the inference engineering principles a little bit, and then we're going to talk about all the stuff that's happened since February 23rd of 2026 in the inference world. We're going to talk about what happened to TurboQuant, talk a little bit about KV compaction. We're going to spend a lot of time on D-Flash and some other new exciting things in specular decoding.
And then, I'm going to do a little bit of prognosticating, a little bit of forecasting of what I think is going to happen in inference coming up here, and what I'm excited about, hopefully being able to talk about next time you guys see me up here. Cool, so let's get started. So, one thing I've been identifying now out of tons and tons of conversations with people about inference is a handful of shared principles. And one of the big ones, I was on this podcast the other day with Sarah. We were talking about inference, and there's two types of inference engineering that have really emerged.
There's local inference where the overwhelming strategy is to get it working on whatever hardware you have by squishing the model with quantization, distillation, pruning, however you can. Splitting it across whatever GPUs you happen to have in your house. And first you get it working, and then you make it less dumb. You take away whatever catastrophic issues all of this compression of the model has created, and you try and get it back to that baseline intelligence running at a batch size of one.
And then there's my world, which is the batch size and data center world, where it's get it working, just day zero, get the build of VLM up, get it working, and then make it less slow. Do stuff like kv-aware routing, speculation, disaggregation, and within these two worlds, I think that we have a lot to learn from each other. I am in this talk going to be focused on advances in data center oriented inference engineering, because that's what I know. But there's a lot of really cool stuff happening in the local world as well. So, in the book, in inference engineering, I generally assume that the weights are a finished product.
And I do think that the handoff from training to inference is an important one to keep in mind, and it's a good way of delimiting the space. However, what I've found more and more recently is that many optimizations for inference come from a dedicated training process. And so, the lines between training and inference are getting blurrier and blurrier. And that's an interesting thing to keep in mind. We're seeing this cycle where you get faster inference, which gives you more data, which you use to train a better model, which gives you faster inference, which gives you more data. And you just keep doing that until you're super rich.
So, with training for inference, we have a bunch of new techniques to talk about across what I like to call the big three. So, we're going to talk about some news in quantization, some news in caching, specifically the KV cache mechanism, and some news in speculation. Because these are the practical day-to-day techniques of how do I make X model faster, usually these are the three techniques that people are reaching for. So, first thing, I publish a book. It's February. I'm feeling awesome about myself. I'm like, wow, everything you need to know about inference in one place. And then, we had some news in the quantization world.
So, just as a quick review, quantization is when we use a smaller, less precise number format in order to save ourselves on bandwidth, save ourselves on compute, make TTFT better, make TPS better. It's usually hardware-specific, gives you cost savings, but potentially degrades model quality a little bit. And, by the way, if you want to hear my whole rant about quantization, I did a talk at AI Engineer Miami last month about how quantization is not necessarily as evil as it sounds, and that there's many things you can do to preserve quality through that process.
So, I was feeling good about my treatment of quantization, and then 20 million people saw TurboQuant, and in fact, it made the memory stock macro dip for a minute just because everyone was thinking, oh, memory's going to be so much more efficient now, we don't need any more flash memory, which was wrong. But, anyway, it was this new quantization approach that was popularized in March of this year that uses polar coordinates for quantization and allows you to quantize the KV cache down to four bits. And it was super hot, and I was thinking, oh, man, there's this whole thing that I left out, and what is this going to look like?
And so, our team did a bunch of research on this. Shout out to Ali from our model performance team at Waterloo. I'm not sure if he's still an intern, actually, but anyway, so he wrote this great piece about the math behind TurboQuant, and basically the benefit you get out of TurboQuant is that you get to represent the KV cache with four bits instead of eight bits. You save half the room and you get effectively double the bandwidth when you're moving KV cache around in your system memory.
But the drawback is pretty big for TurboQuant. It turns out that you need to do additional computation in the forward pass to account for this during decode, and it cuts TPS by more than half, and that's just an unacceptable trade-off for a lot of the production use cases. So, we took a good hard look at TurboQuant, but are not using it for any of these real workloads. We're still on the traditional NVFP4 quantization. That said, it actually is a great technique for the local inference folks.
So, if you are running a model, especially a long context language model on your local computer, on GPUs in your basement, you have a very limited amount of memory. That's the number one bottleneck. And so, anything that can free up memory from KV cache and allow you to put those longer sequences on there is going to be very valuable. And the additional forward pass computation is going to be less of a drawback. So, still, TurboQuant is a fantastic research paper, a really great technique that just ended up not being as applicable in the data center inference world as it might have first appeared.
Instead, we're focused on NVFP4 with a focus on quantizing the weights versus the KV cache. Doing our best to find rough edges in the quantized weights, make sure that we're not flattening out probability distributions. For the KV cache itself, focusing instead on KVAware routing, KV offloading, KV sharing, using Nickel and using NVIDIA Dynamo and other tools in order to move the KV cache around the system and potentially offload to CPU, ordinary memory, etc. Versus trying to use TurboQuant to compress it. And then we're also focused on quantization across modalities.
So, thinking about how can we apply the benefits of NVFP4 not only to language models, but also to image and video models. Ali also wrote a lot of great stuff on Twitter about that, which you should check out. So, that said, the KV cache is still very important. And let's talk about it. Let's talk about KV compaction. Doing our best to find rough edges in the quantized weights, make sure that we're not flattening out probability distributions.
For the KV cache itself, focusing instead on KVAware routing, KV offloading, KV sharing, using Nickel and using NVIDIA Dynamo and other tools in order to move the KV cache around the system and potentially offload to CPU, ordinary memory, etc. Versus trying to use TurboQuant to compress it. And then we're also focused on quantization across modalities. So, thinking about how can we apply the benefits of NVFP4 not only to language models, but also to image and video models. Ali also wrote a lot of great stuff on Twitter about that, which you should check out. So, that said, the KV cache is still very important. And let's talk about it. Let's talk about KV compaction.
Again, quick review, KV cache, if you put in the same prompt with the same prefix, you get to reuse the tokens that you calculated pre-fill last time. That makes your whole system faster and more efficient. Broadly, KV cache is lossless memory. There's only a couple of sources of lossless memory when we think about our inference system. We have the content of the prompt, the context, you have the KV cache. And that's going to scale linearly with the amount of data you pass in. And now, if you're thinking about million token sequence lengths, that actually gets substantial. So, a lot of people are thinking about how do you compress memory? How do you compress context?
Agent harnesses will compress context. RAG, search, all these techniques that we've been talking about for years. Are a compression of a larger context into something that you can give to a model. You can write to files. All of these things scale sublinearly with the amount of data that you have. But what if there was a middle road? What if there was a way where you could get quite a bit of compression in the data that you were remembering with near lossless information retention? So, we have a lot of different ways that we can think of what to keep in the cache. Recent compaction methods have shown that we can replace the cache with a much shorter one.
We've got papers like attention matching and cartridges that have given really promising outcomes here with high compression ratios. But both of these are run at inference time. Again, one of the techniques I want to talk about or one of the themes I want to talk about is training for inference. So, in this case, I want to introduce something called STill by the base 10 research team where the synthesis on top of the cache where we're keeping a learned representation of the information. Rather than the information directly or a deterministic subset of it is amortized via training.
So, Charlie and Mudith from our post training team did a fantastic chalk talk at COSYNE recently. It's up on YouTube. I would encourage you to take a look at it if you're interested in learning about KV compaction. I do not unfortunately have the time or the genius to explain everything up here. But the basic mechanism is that STill is a perceival bottleneck that takes a fixed set of learned query vectors, cross attends it against the full KV cache, and produces a set of compact keys and values in a single forward pass. This creates a fast differentiable compressed memory that the LLM can attend to as if it was real context.
So, if you're interested in KV compaction, definitely check out Charlie and Mudith's work. It's been a fantastic thing to learn about. So, that's two of the techniques. We've talked about quantization. We've talked about caching. The final one is speculation. And there's been a lot of change here. As a review, speculative decoding, we're going to use draft tokens. We're going to verify them during the forward pass. And we're going to use that to generate more than one token per forward pass. It helps a lot with tokens per second. And it is a fully lossless optimization, which is great because we don't have to worry about quality at all.
Now, in the history of speculation, we started with speculative decoding. All of these are in the book. You have spec deck where you use a small model from the same family to generate draft tokens. It turns out small models are not great draft token generators. They're great small models. So, we invented as an industry a bunch of new methods like Medusa where maybe you add draft heads to the model. And then eventually Eagle 3, which was, what if instead of taking a tiny model from the same family, we actually train a billion parameter model on the hidden states of the target model to generate draft tokens. And that actually worked really well.
And so, as of maybe February of this year, Eagle 3 was the best method in speculation. Now we got D-Flash. D-flash is even better. So, it's diffusion for speculation. D-flash creates a sequence of draft tokens instead of a single token. So, the model is a diffusion language model, which means it creates a whole sequence of tokens in the same way that a video or image generation creates a sequence of frames or a sequence of pixels and iterates over it rather than doing an autoregressive token generation. D-flash models might be two or four times slower to run, but they're going to predict eight or 16 tokens at once in that window, while Eagle is only doing one at a time.
So, a single D-flash forward pass is faster than the entire Eagle draft phase and predicts more tokens. These tokens are able to cross-attend to each other and generally create a higher acceptance rate. Because in speculation, acceptance rate is everything. So, in the wild, we're seeing a more than 3x improvement from D-flash. This is measured with a single B200, Qwen38B. And we can see it versus Eagle. It's a substantial improvement in the token acceptance rate and the tokens per second. These D-flash models are trained with an attention mask for bidirectional drafting. So, the target model is going to provide the context.
And within each block, we're going to have a subset of clean tokens that are sampled. And the attention mask is going to enforce causal consistency. And the way we're going to see is still going to allow for bidirectional attention. Where in a traditional autoregressive model, you're only looking at the tokens in a single direction. So, that's why we're able to take advantage of this diffusion-based architecture. And then I thought I was done. And then, a couple days ago, D-Spark came out. Now, D-flash, we do have up and running in production. D-Spark is new research. So, this one, I can basically only say, it exists. It's cool. We're looking at it.
The difference versus D-flash, it still has that diffusion model. But it also pairs it with a sequential model. And the idea is that we're going to improve acceptance rates by having these two models work together. Rather than having just the iterative speculator, just the diffusion speculator, or just the autoregressive speculator. So, D-Spark, very exciting. But we don't have any production results with it yet to share. Hopefully, we'll have those for next time. What we do have production results on is continuous speculator retraining. So, this is, we're back to D-flash here.
And this is the idea that speculative decoding is very dependent on the actual content of the prompts and responses that you're looking at in your system. And so, if you are continuously retraining on those prompts and responses in your live system, you can see a 20% to even 2x improvement in your token acceptance rates. This is actually really hard to do. It takes a lot of storage. And you have to make sure that you have permission to use the data that you're processing in this way. It takes a ton of compute. And you have to move all of this information around. And if you change the underlying model, you also have to change the speculator model.
But when I look forward into the future, I do think that continuous speculation for very large-scale systems is going to be a worthwhile optimization. So, what is next in inference? The following is personal opinion and speculation and public information. And if I knew anything that was actually coming out, I wouldn't be able to talk about it. So, this is just what I think is going to happen. I've been through three hardware cycles through the Ampere release, the Hopper release, the Blackwell release.
And it always takes time for when these chips get shipped to when they get installed in data centers when the entire software stack really is able to take advantage of their capabilities. But some things that I'm excited about are with Rubin, it looks like the NVFP4 performance is going to be fantastic. So, the more we can honestly borrow techniques from local inference Is going to be a worthwhile optimization. So, what is next in inference? The following is personal opinion and speculation and public information. And if I knew anything that was actually coming out, I wouldn't be able to talk about it. So, this is just what I think is going to happen.
I've been through three hardware cycles through the Ampere release, the Hopper release, the Blackwell release. And it always takes time for when these chips get shipped to when they get installed in data centers when the entire software stack really is able to take advantage of their capabilities. But some things that I'm excited about are with Rubin, it looks like the NVFP4 performance is going to be fantastic. So, the more we can honestly borrow techniques from local inference and get a lot of confidence running models in this NVFP4 data format, the more we're going to be able to take advantage of the awesome performance of the upcoming Rubin systems.
I think that disaggregation and system-wide communication is going to be increasingly important. We're seeing really excellent early gains from PD disaggregation. And the ability to move information like KV cache data around the system is going to be increasingly important. And then as I said, the theme of training for inference is going to be something that continues to have a big impact in the industry moving forward. So, thank you all so much for the talk, for coming to the talk. I'm on Twitter, I'm on LinkedIn, and I'm giving out free books. You can download a PDF at the QR code or come down with me to the Base 10 booth to get your free copy of Inference Engineering.
We've got a bunch there, maybe enough for everyone. If not, we will have a career bring some more from the office. So, yeah, I'll be downstairs at the Base 10 booth. Thank you all so much and have a great day. I'll see you now. I'll see you now.
I'll see you now. I'll see you now. Thank you. I'll see you now. Thank you.