AI Engineer

Deep dive on LLM Inference at Scale — Harshul Jain, Audible & Tanmay Sah, Independent AI Researcher

2521 summary words 11 min summary Watch video

Start with the signal

11 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: LLM inference cost and user experience are governed primarily by KV-cache memory, prefill/decode behavior, and batching efficiency; teams should choose model, hardware, and serving engine against explicit quality-latency-throughput trade-offs rather than GPU hourly price alone.
  • Why it matters: The workshop provides a reusable operating model for production agent and LLM systems: measure TTFT, inter-token latency, throughput, KV-cache use, and cache reuse, then select vLLM, SGLang, or hardware-specific tooling based on the workload.
  • Best use: Use it as a practical foundation for capacity planning and inference-engine evaluation, especially before scaling OpenClaw-style agent workloads with repeated system prompts, branching, and long contexts.

Executive Summary

The speakers frame inference as the recurring operating cost of AI systems, distinct from one-time training spend. Their central argument is that token growth, long contexts, and concurrent users make inference expensive because model weights occupy fixed GPU memory while each active request grows a KV cache. The relevant operational problem is not merely fitting a model on a GPU, but allocating the remaining memory between context length and concurrent requests while meeting latency SLOs.

They explain inference through prefill and decode. Prefill creates KV state for the prompt and is compute-bound, so longer prompts increase time to first token (TTFT). Decode produces one token at a time and is memory-bandwidth-bound, so inter-token latency (ITL) is constrained by repeatedly moving weights and cached state through GPU memory. This produces a three-way quality, latency, and throughput trade-off: premium interactive chat should privilege quality and latency, while asynchronous agent work can generally privilege quality and throughput.

On optimization, the workshop separates model-side levers from serving-side levers. Quantization reduces model-weight memory and frees capacity for KV cache, while attention variants such as GQA and MLA reduce KV-cache requirements architecturally. On the serving side, paged attention, continuous batching, KV caching, prefix caching, and KV-cache quantization reduce fragmentation, idle GPU time, or repeated work. The speakers position vLLM as the standard production default because it provides several of these capabilities out of the box.

The most directly relevant conclusion for agent systems is engine selection by workload shape. Their H100 tests found no meaningful vLLM-versus-SGLang difference on a standard ShareGPT-style API workload, but SGLang was reportedly three to four times better in their repeated, branching agent workflow because of its RadixTree-based prefix-cache handling. They repeatedly caveat that benchmark results vary with model, hardware, request mix, and prompt structure, so these findings should trigger a representative internal benchmark rather than a blanket migration.

Key Takeaways

  • Claim: KV-cache growth is the primary capacity constraint after model weights are loaded: context length and concurrent users compete for the same remaining GPU memory. | Evidence: For Mistral 7B, the workshop estimates KV storage at 131 KB per token; a 4K context is roughly 0.5 GB per request, and 80 such users would require about 42 GB of KV memory. The speakers illustrate that a 24 GB GPU can therefore run out of memory well before reaching an intuitively large concurrency target. | Implication: Capacity planning should start with model-weight memory plus KV-cache budget, then solve for context and concurrency under the intended SLO rather than sizing from parameter count or GPU price alone. | Caveat: The exact KV footprint depends on model architecture, layer count, KV-head count, precision, context length, and serving implementation; the Mistral figure is an illustrative calculation, not a universal constant.
  • Claim: Prefill and decode have different bottlenecks, so inference must be measured with separate metrics rather than a single latency number. | Evidence: Prefill processes the full input and constructs KV state, making it compute-bound and the main driver of TTFT. Decode generates tokens sequentially, has low arithmetic intensity, and is limited by high-bandwidth-memory transfer rates; it drives inter-token latency. The workshop identifies memory, TTFT, throughput, and ITL as the four practical performance dimensions. | Implication: For interactive systems, separately set and monitor TTFT and ITL SLOs; prompt length primarily threatens first-response responsiveness, while decode throughput and memory bandwidth shape streaming quality. | Caveat: Decode time is not completely independent of context length because attention still reads cached K/V vectors for prior tokens.
  • Claim: There is no universally optimal GPU configuration: quality/context, latency, and throughput form a practical trade-off triangle. | Evidence: The speakers contrast a premium chat application, where latency and quality should take precedence over high concurrency, with asynchronous agent tasks, where quality and throughput can take precedence. They argue that an H100-class GPU with a high hourly price can still yield lower cost per million tokens if it supports sufficient throughput for the actual workload. | Implication: Choose hardware after fixing the business-critical constraint—such as interactive latency or minimum concurrent agent jobs—and calculate cost per useful token at that operating point. | Caveat: Their illustrative premium-chat configuration cites a 10 ms latency target, but real end-to-end latency includes network, queueing, prefill, and application overhead and should be modeled from actual workload traces.
  • Claim: Quantization and KV-efficient attention architectures are the main model-side methods for fitting larger workloads into limited memory. | Evidence: A Mistral 7B FP16 model is presented as roughly 15 GB; INT8 cuts this to about 7.5 GB and INT4 to roughly 3-4.5 GB, freeing memory for context or concurrency. The speakers describe the attention spectrum from multi-head attention (MHA), through grouped-query attention (GQA), to multi-query attention (MQA) and multi-head latent attention (MLA). GQA reduces the number of KV heads; the corrected workshop estimate is that MLA offers about 14x KV savings versus MHA, not the initially stated 56x. | Implication: Treat model precision and architecture as workload-capacity levers, but gate changes with task-level quality evaluation—especially for reasoning, tool use, and long-horizon agent behavior. | Caveat: Quantization and more aggressive KV compression can affect quality, and the speakers explicitly say apparent quality preservation must be validated on external or use-case-specific benchmarks. Their attention explanations are conceptual rather than an implementation guide.
  • Claim: Serving-engine optimizations often provide much larger practical gains than a naïve Hugging Face inference loop. | Evidence: The speakers report an H100/Mistral 7B baseline of about 51 tokens/sec throughput, 54 TTFT, and 19 inter-token latency units for raw Hugging Face inference. A default vLLM server—using paged attention, continuous batching, and KV caching—reportedly increased throughput by nearly 15x, reduced ITL, and improved KV capacity. Prefix caching further improved throughput and TTFT in their test, while KV-cache quantization mostly reduced KV memory use. | Implication: Avoid using an unoptimized model-serving loop as a production baseline; benchmark a production engine with representative traffic and inspect both latency distribution and cache-memory behavior. | Caveat: The reported units, prompt/output distributions, concurrency levels, and full benchmark methodology are not clearly specified in the transcript, so the 15x result should be treated as directional rather than portable.
  • Claim: vLLM is the default choice for standard API serving, but SGLang deserves targeted evaluation for branching agent workloads with repeated prefixes. | Evidence: On their H100 ShareGPT-style standard API test, the presenters found no statistically meaningful difference between vLLM and SGLang in requests/sec, TTFT, or latency. In a repeated two-stage agent workflow—generate a proposal, then ask for a review and 1-10 rating—they report SGLang as three to four times better, attributing the difference to agentic branching and SGLang's RadixTree prefix-cache strategy. | Implication: For agent control planes, capture prefix-reuse rate, branch structure, and session routing patterns; run a vLLM-versus-SGLang bake-off using real system prompts, tools, retries, and evaluator loops. | Caveat: The speakers explicitly state that the SGLang advantage depends on setup and may not reproduce elsewhere. Static hash-based prefix caching can also miss when prompts differ by even a small edit; cache benefit depends on real prefix reuse and routing behavior.
  • Claim: Speculative decoding is not a reliable universal optimization and should be accepted only where draft-model alignment is high. | Evidence: The technique uses a smaller draft model to propose several tokens that a larger target model accepts or rejects. The speaker says it may work better in constrained domains such as code or predictable syntax, but reports that personal testing did not find conventional speculative decoding useful because of alignment problems. They identify self-speculative decoding, EAGLE, and Medusa as variants, with EAGLE presented as relatively better. | Implication: Do not add speculative decoding by default to agent infrastructure; first instrument acceptance rate, end-to-end TTFT/ITL, and operational complexity on constrained, high-volume tasks. | Caveat: This is largely a practitioner opinion from the speaker rather than a disclosed benchmark result, and performance depends heavily on draft/target acceptance rate and workload characteristics.

Detailed Brief

KV-cache mechanics and serving primitives

  • Claims: KV caching exchanges memory for avoiding repeated K/V computation across generated tokens; without it, repeated attention work creates substantial avoidable compute.; Paged attention treats KV storage as dynamically allocated blocks mapped through logical-to-physical addresses, analogous to operating-system virtual memory.; Continuous batching prevents the GPU from waiting for every request in a fixed batch to finish before accepting more work.
  • Evidence: The illustrative fragmentation case allocates 2 KB of contiguous space to a request that needs only 1 KB, stranding 50% of that reservation.; Prefix caching extends reuse beyond a single request by reusing cached KV state for identical prompt prefixes across requests.; KV-cache quantization reduces the bytes consumed by K/V vectors, primarily increasing possible context or concurrent request count rather than necessarily improving speed.
  • Caveats: Prefix-cache gains require actual common prefixes; slight prompt edits can defeat simple hash-based matching.; Long-term cache engineering introduces eviction, compression, and hybrid-memory design decisions that this beginner/intermediate workshop does not cover in detail.
  • Implications: System prompts, tool schemas, and agent templates should be made stable and shared where possible so serving infrastructure can exploit reusable prefixes.; Cache occupancy, hit rate, eviction behavior, and fragmentation should be first-class production metrics rather than hidden engine internals.

Attention and hardware-level optimization landscape

  • Claims: Attention design changes both model quality behavior and the size of the KV state that must be retained at inference.; FlashAttention improves attention execution by tiling work so smaller chunks can be processed in fast on-chip memory while using online softmax bookkeeping.; TensorRT-LLM is presented as an inference engine, distinct from the TensorRT SDK, focused on NVIDIA-specific low-level optimization.
  • Evidence: The workshop characterizes MHA as preserving quality but carrying larger KV cost, MQA as an aggressive one-KV-head extreme with poorer quality, and GQA as the practical middle ground widely used by contemporary models.; It cites DeepSeek sparse attention as an emerging direction that avoids attending equally to every prior token, and names linear attention and Mamba/state-space models as alternative trajectories beyond conventional attention.; The speakers also name NVIDIA Dynamo for agentic session routing and Stanford's M-star as newer systems to investigate for multi-model serving.
  • Caveats: The workshop does not supply reproducible comparative data for FlashAttention, TensorRT-LLM, Dynamo, M-star, sparse attention, linear attention, or Mamba.; Changing the model architecture is generally a model-selection or training decision, not a drop-in serving configuration.
  • Implications: Separate decisions that can be made at deployment time—engine, batching, cache policy, quantization—from decisions inherited from a selected model's architecture.; NVIDIA-optimized engines may be worth testing when peak utilization on NVIDIA hardware is the binding constraint, but portability and operational complexity should be part of the decision.

Benchmarking discipline and gaps

  • Claims: The speakers recommend returning to first principles whenever a new inference product appears: identify which bottleneck it addresses and whether that bottleneck exists in the target workload.; Distributed inference is materially different from single-GPU serving and is not substantively addressed in this workshop.
  • Evidence: Their demos and capacity calculator use Mistral 7B and an H100/RTX 6000-class GPU context, and the benchmark requires repeatedly restarting vLLM servers and loading models.; They direct viewers to a GitHub repository containing slides, notebooks, a benchmark report, and a capacity calculator, although the transcript does not provide a durable URL.
  • Caveats: Some live demos failed or were constrained by Wi-Fi/GPU detection, and the transcript contains duplicated passages from the recording.; At least one quantitative claim was corrected live: the stated MLA compression was revised from 56x to 14x because a calculation omitted the number of layers.
  • Implications: Use the workshop for its decision framework, not as a source of final performance numbers; reproduce results with controlled versions, warm-up policy, traffic distributions, and task-quality checks.; For multi-GPU or distributed deployments, commission a separate analysis covering parallelism strategy, routing, disaggregation, networking, and failure behavior.

Notable Concepts & Terms

  • KV cache: Stored key/value attention state for previous tokens; it avoids recomputation during decoding but becomes the central memory budget for long context and concurrency.
  • TTFT (time to first token): The user-visible delay before streaming begins, driven primarily by compute-heavy prompt prefill.
  • ITL (inter-token latency): The delay between generated tokens; it reflects sequential decode performance and is often limited by memory bandwidth.
  • PagedAttention: Block-based KV-cache allocation that avoids contiguous-memory fragmentation and underpins vLLM's efficient request management.
  • Continuous batching: Admits new requests as others complete instead of waiting for a fixed batch to drain, increasing GPU utilization and throughput.
  • Prefix caching / RadixTree: Reuse of KV state across requests with shared prompt prefixes; RadixTree-based handling is especially relevant to repeated and branching agent prompts.
  • GQA and MLA: Attention architectures that reduce KV-cache cost: grouped-query attention shares K/V heads among query heads, while multi-head latent attention compresses K/V state through a latent representation.
  • Speculative decoding: A draft model proposes multiple tokens for a target model to verify; it can accelerate predictable outputs but depends on high draft-target alignment.

Operator Notes / Why Ken Should Care

  • Build a production inference scorecard with p50/p95 TTFT, p50/p95 ITL, tokens/sec, requests/sec, GPU utilization, KV occupancy, cache-hit rate, eviction rate, and cost per successful task—not just cost per token.
  • Run a replay benchmark of real agent traces against vLLM and SGLang, preserving shared system prompts, tool schemas, branches, retries, evaluator calls, and session affinity; use the result to decide whether SGLang's cache behavior justifies operational change.
  • Standardize and version high-frequency agent prefixes (system prompt, tool definitions, policy text, common context) to maximize cache reuse; avoid gratuitous per-request variation in those sections.
  • Set explicit workload classes—interactive chat, asynchronous agents, batch processing—and assign different context, concurrency, TTFT, and ITL targets rather than applying one serving configuration globally.
  • Before quantizing a production model or KV cache, require a regression suite covering tool-call validity, long-context retrieval, reasoning quality, safety behavior, and task completion rate.
  • Treat published engine benchmarks as hypotheses. Reproduce them under pinned engine/model versions and document warm-up, request mix, output length, concurrency, GPU type, and queueing conditions.
  • Do not prioritize speculative decoding until a targeted experiment shows a sufficiently high acceptance rate and an end-to-end gain for the actual workload.

Source/Metadata

  • Title: Deep dive on LLM Inference at Scale — Harshul Jain, Audible & Tanmay Sah, Independent AI Researcher
  • Transcript words: 16607
  • Duration seconds: 5292
  • Timestamp note: No usable timestamps or chapters were present in the supplied transcript; the transcript also contains substantial duplicated passages and live-demo interruptions.
Full transcript 10030 words · 73 min read
0:00

Good afternoon, everyone.

0:13

My name is Harshal Jain and he is Tanmisha. And we would like to welcome you all in this two-hour workshop on LLM inference. So the goal of this workshop is to understand this domain from the first principles, dive deeper into it, and understand what's going on throughout the industry. A bit of background about us. So I am a senior software engineer at Audible. I have been building ML/AI data platforms for the past five years. And on the side, I have been writing this open-source handbook on LLM inference. And Tanmisha, he is the senior quantitative modeler at Xions Bank Corporation.

0:52

He recently completed his PhD and he has been actively doing research in agent verifiers and world models. So a quick show of hands here. Who here is brand new to LLM inference?

1:19

Okay, great. And who here has deployed these models in production? They have been tuning it. They have been serving the production traffic. Okay, great. So this workshop is targeted towards the beginner and intermediate level. And all of the slides and exercises are in the repo. I will share that soon. Here is the quick agenda for the workshop. We will start with the problem statement.

1:54

We will try to understand a few of the pain points around LLM inference. Then we understand what causes those pain points and build our foundations from there. Then we will dive into two kinds of optimizations that we do: model optimizations and serving optimizations. And then we start learning about different serving engines that are available to deploy our LLM inference solutions in production. And we will showcase some benchmarks and decision charts on which engine to use. So to understand the pain points, first we need to know what is LLM inference. Probably a lot of us already know this.

2:39

But anything that you ask your AI to do, whether it be generate a video, audio, analyze any text, analyze your medical reports or your tax bills, all of that is LLM inference. And this market is approximately $23 billion today. Semi-analysis recently shared that if you want to model Google search queries with LLMs, you need a profit drain of $36 billion. And query cost has to be less than 0.5 cents to keep your search business profitable. On the other hand, Business Insider mentioned your AI has to be put on diet. And everyone has to start auditing and budgeting their token usage. And all of this is happening. Why? Because your hardware is limited. Compute is expensive.

3:35

Your inference is expensive. And with the growing need for more AI usage, this inference cost is rising. So this stat, it's an old stat from OpenAI, but it's still true. If you look at the training cost of GPT-3, it was around $4.6 million. It was a one-time cost. But if you see the inference cost, that has been a recurring cost because it's an operating cost that scales with every user that comes in, every token that comes in, every session that is being initiated. And there are only two ways to counter this. One way is you reduce your token usage.

4:04

The alternative is you should try to optimize your inference solutions as an inference service provider for your customers and for yourself. And we have been seeing a lot of new solutions coming out every now and then. So the idea would be we will try to build those foundations that will help us understand and evaluate whatever ships next. So to get started, we will do a quick demo of what are the different pain points around inference. This is a repo. You can pull it or you can also open it on GitHub.

4:39

It's called LLM inference at scale. A bit of background here: four months back when I didn't know anything about LLM inference, I started learning it. I saw a lot of resources were scattered. So we started putting it all together in one place so that it could benefit people. So in this repository, if you see a readme file, there is a link to the slides. So this will be this folder where you have a PPTX and there is a benchmark report in there. You can always download it. And then for the demo purposes, we have a couple of Jupyter notebooks. We have collaborated with Molab who are Google Colab alternative. And what they basically provide you is a free RTX 6000 GPU.

5:46

It's a 100 GB RAM GPU. And we have already set up these notebooks so that it becomes easy to experiment with. And all of the assets and everything are preset for you.

6:16

We will start with a simple demo A. So when it comes to inference, you need to do an inference on a certain model. For the workshop purposes, we are using a simple Mistral 7B model. It's a small model of around 15 GB in size. So we are going to load that into the GPU. And we will look at some of the GPU stats as well. We see we are working on the 6000 Blackwell. And you might be thinking I'm not running this because I don't trust the Wi-Fi at conferences. So I would probably just go over the results that we ran previously. We have a GPU which is 102 GB. Now the first thing that comes to my mind is what my memory consumption looks like when I do LLM inference.

7:17

So I load this model and I see I have 15 GB here.

7:20

So I have roughly 87.5 GB.

7:28

And now when I do the inference here, what I notice is the more the number of inputs I pass, the more memory that I need.

7:31

And it's increasing slowly, but it's still increasing. So imagine if you have a context length of around 4000 or 16000 or 32000 tokens. This memory could really grow big and you could actually get out of memory issues. So definitely this is your problem one: memory increasing with the increase in tokens. In form of a simple visualization, it looks like this. The second problem that you would see is the time to your first token is very, very slow. We measure it by a metric called TTFT. It's a short form for time to first token. And when you try to measure the TTFT with the input size, you would see the longer the context, the slower this TTFT becomes.

8:01

So now there are two problems. Your memory increases with the token size. Your TTFT increases with the context size.

8:18

And then the third is the throughput. The throughput is how many tokens can you serve per second and how many users can you serve per second? So if you take a very vanilla implementation on your local system, it would be very sequential. So if you send five requests, all those five requests would be catered sequentially rather than in parallel. So your request basically takes more time to complete if you have multiple users. So these are the three problems: memory, TTFT, and throughput. There is a fourth one. I haven't described it here. Probably we will build that intuition as we move forward. But let's remember these are the three problems. I will go back to the slides.

9:24

Perfect. Within that repository, if you see a workshop folder, you see that readme, and then the readme has all the links to the slides and demos. Does that work? Okay, perfect. So let's start working through the foundations. Let's start understanding what are the reasons behind those pain points. And for that, we have to look at this inference pipeline. We get an input text.

10:06

That text could have any number of words. You convert those into tokens. So for simplicity, you can assume one word equals one token. Then you convert them into embeddings. And then you send it to the transformers. There are 32 layers of transformers, but that's specific to Mistral 7B. We have different models with different numbers of layers. And then you generate a new token. And that token basically goes back to the input. Then you generate another token. And that keeps on going. Now in this entire pipeline, you would see 95% of your compute is taken by these transformer layers. Does that work?

11:07

Okay, perfect. Okay, so let's start working through the foundations.

11:15

Let's start understanding what are the reasons behind those pain points.

11:20

And for that, we have to look at this inference pipeline.

11:26

So we get an input text. That text could have any number of words.

11:32

You convert those into the tokens. So for simplicity, you can assume one word equal to one token. Then you convert them into the embeddings. And then you send it to the transformers. There are 32 layers of transformers, but that's specific to the Mistral 7. We have different models of different number of layers. And then you generate a new token. And that token basically goes back to the input. Then you generate another token. And that keeps on going. Now, in this entire pipeline, you would see 95% of your compute is taken by these transformer layers. So it's worth looking at what goes within this transformer layer.

11:38

Within this transformer layer, you would have more layers. You have a normalization layer. You have an attention layer. You have a feed forward layer and all. And attention layer is the one I think that has been very, very famous. Attention is all you need people. I think that's very well known. So attention is the most compute intensive layer. And we need to understand what goes within that attention layer. So what does attention do? Attention, so if you have an input text, it needs to find the attention scores of every token with respect to all of the previous tokens.

11:38

And to do that, what it needs to do is it needs to project every token into a key, query, and value space. So in simpler terms, just understand this: if you have 10 tokens, then it needs 10 different query, key, and value vectors. If there are 100 tokens, you would need 100 key and value vectors. If there are 1,000 tokens, you would need 1,000 key value vectors. And so your number of key and value vectors increase as you increase the input size. And if you calculate the KV size per token, for a Mistral 7B, it comes out to be 131 KV. This is because you have two vectors, K and V. You have to multiply the size. One vector is 128 dimensions.

11:39

You have to multiply it by 32 transformer layers. And then you have to multiply it by the KV heads. For Mistral 7B, it's KV heads. It's not 32 because it uses a different kind of attention mechanism, which we will talk about for sure. But the KV size per token is 131 KV. Now imagine if you have 4K context, so that size becomes half a GB. If you do 16K context, that size becomes 2.1 GB. Now multiply it by the users. Assume you can serve multiple users together at the same time within that GPU, you could have 42 GB with a 4K context and 80 users. And if your GPU is only 24 GB, you are already running out of memory.

11:43

So you cannot serve that many users with that many contexts. To visualize this, look at GPU memory. So the GPU memory has model weights, which are pretty fixed. These are pre-trained weights. There is overhead that is also fixed. That also changes, but it does not change that much. Overall, you can assume it's fixed. And then there is leftover memory. So this leftover memory is what's being used by your KV memory, like key and value vectors. So assume you have one user. You can only serve that many key and value vectors or that many tokens, which can fit in this entire 80 GB of memory that is left.

12:04

So we can show this with a simple demo too.

12:09

Okay.

12:15

Let me. Okay, great. Okay, great.

12:31

Yeah. So you will see if I can actually run this. Let me see if I can actually run this. What? What? Where the heck is this? Okay, great. Yeah. So you would see the GPU is attached. So here we are just trying to confirm the memory based on the maths and based on the intuition that we have built. So the model memory is, let's say if you have 7 billion parameters, you are doing 16-bit precision. Your total memory comes out to be 14.6 GB. You can basically verify that with the maths. So if you do all that maths, that comes out to be 14.6 GB. Now comes the KV and the KV size. So this KV size is 131 KV per token. So if you do that maths and you try to visualize this. Oh, sure.

13:36

Okay. And then let's just visualize this. Okay, great. Yeah. So this is the memory chart. So if you see as your context increases, your memory keeps increasing. Then another thing to realize is as your users increase, then also your memory increases. So if you want to serve 160 users on a GPU, you can support only lesser context length. So there is always a trade-off between what context length you can serve versus how much cost you can save by putting multiple users or concurrent users into a single GPU. So you have to always take that trade-off. And we will go through that in a couple of more slides. Can you repeat, please?

14:44

I'm sorry. I cannot hear you. Do you create different tools with different product length so you can serve the memory alarming? Yeah. Cool. Okay, good. So let me pull back. So that was light memory. We need to understand why we had slower time to first token when we increased the context length. So for that, we need to understand the two phases of inference, and those phases are the pre-fill and the decode phase. I think you would all see a lot of articles, but we just wanted to explain it. So when you send a lot of input tokens, what you want to do is you want to build those key and value vectors that I mentioned for all the tokens.

15:45

Then you want to compute the attention scores of every token with respect to the previous token. All this operation that you do, it's very metrics heavy. It's very compute heavy operation. And we all know GPUs are very well suited for heavy compute workload.

16:11

So we call pre-fill to be compute bound. And it does take some time to complete. So whatever time that this phase takes to complete, that's your time to the first token. So if you have more input tokens, you have to generate more key value vectors. You have to do a lot more attention math. And because of that, your TTFT becomes slower and slower. Whereas once you generate one token, you need to keep doing this to generate another token sequentially one after another. But in that process, every time you have to build the key and value vectors of all the previous tokens, which is the same as pre-fill, like you were building key value vectors there also and here also.

16:43

But in decode phase, you are only computing the attention math for the new token. And that is why it's very less, it's lesser compute oriented. And it's also called as memory bound.

17:04

We will see it shortly why it's called as memory bound.

17:12

So in a classic timeline, you would see pre-fill and decode phase like this. So time taken by pre-fill, that's your time to first token.

17:18

And then your time taken by every decode step, that's basically your inter-token latency.

17:22

So that's the fourth metric that you need to worry about. What's the time being taken by your decode step? Okay. Okay. Now why does the decode step take time? And why is it being called as a memory bound operation?

17:35

Let's try to understand that. To understand that, we need to look at how the metrics map basically works on the GPU on a high level.

17:48

So GPU has two kinds of memories. You have high bandwidth memory. You have shared memory.

18:07

So the high bandwidth memory is larger size, but lower bandwidth. By lower bandwidth, I mean you can transfer data out of it at a lower rate. Compared to the shared memory, the shared memory is smaller in size, but it has very high bandwidth. That means you can transfer data in and out of it very fast. So when you have to do metrics math, you have to pick the data in chunks from the high bandwidth memory, you have to put it into the shared memory, do that math, write back the result into the high bandwidth memory.

18:33

For the pre-fill phase, when you have to do this, you have to do this metrics math only once, but for the decode phase, you have to do this metrics math again and again, because you are generating each and every token sequentially. And so it doesn't matter how fast your decode is, because now you can transfer your data out of the high bandwidth memory into the shared memory at a certain speed, because you are limited by the high bandwidth memory bandwidth speed. Compared to the shared memory, the shared memory is smaller in size, but it has very high bandwidth. That means you can transfer data in and out of it at a very fast rate.

18:52

So when you have to do matrix math, you have to pick the data in chunks from the high bandwidth memory, put it into the shared memory, do that math, and write back the result into the high bandwidth memory.

19:03

For the pre-fill phase, you have to do this matrix math only once. But for the decode phase, you have to do this matrix math again and again because you are generating each token sequentially. And so it doesn't matter how fast your decode is, because you can transfer your data out of the high bandwidth memory into the shared memory at a certain speed. You are limited by the high bandwidth memory bandwidth speed. And so that governs your token ceiling, like at what rate you can actually generate tokens out of the decode step.

19:09

If we look at this in the roofline plot, there is a left section which is called memory bound. Mathematically, it's governed by the arithmetic intensity. Arithmetic intensity is the number of floating point operations that you perform per byte of data being transferred. For the decode step, since you are transferring a lot of data like the key and value vectors of all the previous tokens and the model weights, but you are doing less computation because you are computing attention map for only one token, its arithmetic intensity is very low.

19:12

But for the pre-fill phase, you are transferring the data once, but then you are doing heavy computation, so its arithmetic intensity is very high. So now you know in terms of mathematics why the arithmetic intensity of pre-fill is very high compared to decode. So this is another small demo. Every time I have to... okay, great. I hope this is already... So again we are loading the model. Now this is the pre-fill cost. What we are basically doing is getting the input text and then trying to generate this. The pre-fill step, the amount of time it takes. We see as we increase the size of the input tokens, this pre-fill is increasing. So that's why your TTFT increases.

19:55

And then your decode time. The decode time on average stays about the same. And so if you ignore the cold start, your decode time is approximately around the average line. It is still impacted by the input size. It's not constant time. And that's because it still needs to pull the key and value vectors from the memory for all the previous tokens. So there is still a small increase in time that you would see with the decode step. And then this is the classic roofline plot.

20:06

Okay.

20:14

Okay, great. So now let's try to understand the throughput dimension. You want to understand how many users you can actually serve. And I think we saw a diagram of the GPU memory where we saw there is some memory that is free for the key and value vectors to grow. So assume you have just a single user. What's the total KV size that you can support? It's defined by your context limit. The max users that you can support is whatever your GPU availability is. Whatever memory is available in the GPU, you divide it by the key and value size per user. And when you do that, it comes out to be the concurrent users.

20:16

Now, assume your GPU is fixed and your model is fixed. So your KV size per token is fixed. There are only two dimensions that are left here, which is context and your concurrent users. If you want to serve more concurrent users, you have to reduce the context length. If you reduce the context length, you could impact your quality.

20:33

So these are the two dimensions that we are trading off. But can you actually serve the max number of concurrent users? In an ideal world, probably not because every business has a latency SLO that we have to meet. So if you remember, in the decode step, I said the time for the decode still increases if you have more inputs. It also increases if you have more users. So ultimately, your inter-token latency also gets impacted if you have a higher batch size and your TTFT also gets impacted. So now there is a third dimension you have to worry about, which is your latency.

20:43

So the three dimensions that you have is quality, latency, and throughput. So it comes out to be a trade-off triangle where you have to choose between two. For a premium chat application, you would want to definitely prioritize quality and latency. You would not want your users to wait for a large latency. You can always sacrifice the number of users you can support on the GPU and probably accept that cost by being more customer obsessed. And if you consider an async agent workload, you would want to prioritize quality and throughput because these are long running tasks and you would want to serve as many concurrent tasks as possible, but with very high quality.

20:57

And often we think, okay, the GPU is a very expensive GPU. That might not be a good fit for us. But it turns out that could actually serve you the lowest cost per million tokens. But you really have to trust your calculations on the max users that you want and you really have to make those estimations correctly. So we do have a capacity calculator. There is a link to the Colab because I was facing certain issues with Molab. I had to migrate out of the whole widget library and I didn't have time, so I just picked Colab. Apologies to Molab.

21:01

Okay, great. So what we have done here is we have shaded some GPUs with their VRAMs and bandwidths, the floating point operations, and the cost per hour. Then we kind of built this simple capacity calculator. This is a KV visualizer where when you increase the number of tokens, you see the KV size increases. And when you increase the number of users, your size increases at a much faster rate. And then in this capacity calculator, let it run. So we have a model which is a 7 billion parameter model that we selected. We set the precision to be FP16. Now we decide the way we basically go by the GPU decision is you have to decide what's your most important dimension first.

21:13

For premium chat, I mentioned latency is definitely the one. And then for async workloads, the minimum batch size that you want to solve from a single GPU is the second dimension. So you want to fix these first. So for a premium chat application, I can go ahead with 10 milliseconds latency. Minimum batch size, I don't care. I'm okay with probably two, or about seven concurrent users on a single GPU. And then my context limit is very important to me because I want to focus on quality as well.

21:20

And so I do see some of the GPUs. The H100 8GB is $8 per hour. But it's around $10 per hour. But if you do all that throughput math that we shared in the mathematics before, you could find your cost per million tokens could be very low. So you need to do such calculations by fixing those dimensions and you need to decide your GPU to reduce your inference costs. This is at least the first step that you can take towards optimizing inference. Okay. Cool.

21:29

So the next slide. Okay, great. And so now the next thing is about model optimization. So we have now built that foundation where we understood some of the pain points, reasons behind those pain points, why those were happening, and how we could address that GPU capacity thing. We need to understand what can we further do about it. So it is about model optimization. And I think I would like to invite Tanmay. He can talk more about these model optimizations provided he has worked on this during his research times. Okay, I can control. Yeah. Okay. Hi everyone. Mic check. Am I audible? At last, yeah.

21:37

Okay, so hi, I'm Tanmay. I work as a senior quant modeler and also as an AI researcher. My work focuses on agent verification and right now building world models. So for this one, model optimization, before we start model optimization, I created a research template so that it will be easy for us to understand all these complex things. So our template is simple. First we will identify the problem. Second step we will solve the problem using two algorithms. These are just...

21:47

We have built that foundation where we understood some of the pain points, the reason behind those pain points, why those were happening, and how we could address that GPU capacity thing. We need to understand what can we do, what can we further do about it. So it is about model optimization, and I think I would like to invite Tanmay. He can talk more about these model optimizations, provided he has worked on this during his research times.

21:50

Okay, I can control. Yeah, okay. Hi everyone, mic check. Am I audible at last? Yeah, okay. So hi, I'm Tanmay. I work as a senior quant modeler and also I'm an AI researcher. My work focuses on agent verification and right now building world models.

21:53

So for this one model optimization, before we start model optimization, I created a research template so that it will be easy for us to understand all these complex things. So our template is simple. First, we will identify the problem. Second step, we will solve the problem using two algorithms. These are just fake algorithms. So first algorithm is called ostrich algorithm. Whenever we see just like ostrich, whenever we see a problem, ostrich put their head into the sand. So same thing we will do. Whenever we face a problem, we will just ignore it. So this is an important algorithm we should follow. Second one is created. It is called world cup algorithm. For example, we don't know who will win this FIFA world cup. So what organizers did, they break the 48 teams into 12 groups, then round 32. So round 32 right now is currently going on, then round 16, then quarterfinals, then semifinals, and finals. So what they are doing is that they are breaking it into smaller problems and the useful results are moving forward. So same analogy or same algorithm we will use to understand this model optimization, all those things. So yeah, let's start.

21:56

So I have one H100 GPU. I have to use this open source model what is called GPT OSS 120 billion parameter model. So right now I think they have trained it on BF float 16 and weight is 240 gigabyte. What should I do? This is the problem we have. So first thing what we have to do is that 240 gigabyte and 80 gigabyte H100. So and I have to fit only in one GPU, not in multiple GPU. So what can we do? I think simple step is that just compress it. But how should we compress it? That's another challenge. So if we compress BF float 16 to FP8, then it will be around 120 gigabyte. But our GPU H100 is still 80 gigabyte. So what I think they did is that they compressed it into further MXFP4 and I think size is around 65 gigabyte. So this is something we can do: compress. But question... so and we will use over this ostrich algorithm. We are assuming that there is no loss in compressing a bigger model into a smaller size.

22:00

Second thing, in this one, okay, yeah. So in this one, in this slide, we have used this Mistral 7B. So 7 billion parameters. So it's a small model. 7 billion parameters. So if you multiply it by 2 bytes, so weight of it's around, is 14, 14.5 gigabyte, which can easily fit into H100 or even A40.

22:03

So, so next, what we can do is that like Mistral 7B, instead of compressing it floating point 16, we can apply different techniques like int 8 or int 4 or nf 4. So basically we have to just use ostrich algorithm and just believe that there is no quality loss. But somehow we also have to mathematically prove that by doing some testing on some external benchmark that whether it is working or not. So and this comes under post-training quantization. One can also do this during fine-tuning. One can also do this quantization. This comes under quant-aware training.

22:05

So let's move to our next problem. So we have this huge matrices. Just imagine 1,000 by 1,000 dimension matrix A and another matrix 1,000 by 1,000. So if you multiply by this two matrices, so number of operations will be 1,000 raised to the power 3. And this is a problem in terms of computing. So we wondered our matrix multiplication should be fast and it should save memory. So what should we do? We have a giant matrix. Okay, let's take this one: 4096 by 4096. What should we do to solve our problem of speeding up the things and saving the memory? 4096 by 4096. So first thing is that we will use just our world cup algorithm. We can decide a random number, just break the block vertically. It does not matter what you are choosing.

22:12

So you have, so let's say we have 4096 columns. We will break it. We will break it into a group of 128 column each. So 128, 128, 128, 128, 128 vertically. So we will get this 32 blocks. If we divide this 4096, then what will happen by doing this thing? So if we just divide this one vertically, then we can use a multiple GPU to speed up the process. So this kind of thing is called multi-head attention.

22:15

So what else can we do? We have a big matrix. As I have mentioned, that ostrich algorithm. So our main problem is sizing. So what we can do is that instead of having all those 32 vertical blocks, we will throw away 31 blocks and we will assume that one block is sufficient enough that all the queries can handle those blocks. Our loss will be almost negligible and we come up with this algorithm and this algorithm is called multi-query attention.

22:18

So as we can see right now, we are at two spectrum. One is multi-head attention where we split it into 32 blocks and use different GPUs or do some parallel processing, and at the same time we are just throwing 31 blocks and we are calling this as multi-query attention. So at both extremes, we should come up with a middle ground. Something we can say that instead of throwing all the 31, maybe we can group some of the blocks together so that and we can assume that similar blocks will attend to similar kind of queries. So this kind of technique comes under grouped query attention, which is very popular right now, even in Mistral or in other models. This grouped query attention works.

22:20

So right now we have understood that we have a big matrix. We can divide it the way we want and doing some mathematical calculation, prove that loss is almost negligible. So what else can we do?

22:22

So after that, after this grouped query attention, see we have a big matrix. One is key and one is value. Let's compress that matrix into a latent vector and then come up with some algorithm to reconstruct from latent vector to our original matrix. So this kind of strategy comes under multi-head latent attention. But again, it has some problems with RoPE because RoPE is position dependent and it is position independent. So one needs to also include some index for keys also so that one can map it. But again, main problem is why we are multiplying all those big matrices? So because that's how this attention mechanism works. Each token will pay attention to every token. So how about let's not pay attention to all the previous tokens. Only pay attention to the important tokens which is important for us. So this kind of field is evolving. So this comes in Deep Seek sparse attention.

22:26

So yeah, and yeah, so okay, next. Yeah, so next one is Flash Attention. So in Flash Attention, so main problem is that so currently, so currently not currently. So right now, almost everyone uses Flash Attention, but way in 2022 or 2023, so that's how it works. That's how it works is that so this Q K query and key matrices, they were in HBM. It loads it. First, it loads into this one our tensor core and it do some calculation and then it will write it back to HBM and then this process goes on multiple times. So in Flash Attention, what they did is that instead of multiplying the whole matrices, so they just divided it into our world cup algorithm, divided the bigger matrices into a small tile and only put those small tiles into SRAM so that it can process multiplication fast and just keep keeping track of this some three variables so that they can calculate this online softmax.

22:28

Yeah, next one. So yeah, so this is just mathematics. So if we have a multi-head attention if it is 524 KV, then it depends upon how much grouping we want. And so if instead of 32 KV head we only want to use 8 KV heads, so we can get a compression of 4x times. And this multi-head latent attention, this formula depends on the model to model how many layers your model have. So in the original Deep Seek paper, I think they have some 128 dimension, 128. I don't remember the exact dimension, but according to that, they have used this one latent vector in which they have used 512 as a dimension and some 64 for RoPE index and then they show that it is 56x more compressed than multi-head attention.

22:30

Okay, yeah. So this is the trade-off diagram. So here I think we have not talked about this linear attention or Mamba. So main problem is just all this matrix multiplication. Right now everyone is using attention. Suppose in future if we don't want to use attention, rather than generating tokens sequentially, just use maybe diffusion models where we can generate everything simultaneously. So all these algorithms will change also. But here I think they have two more. One is linear attention and one is Mamba. So according to this slide, so if we are not compressing anything, so MHA is just we are parallelizing the process. So there is no quality loss, so it's good. And then this group query attention, which is I think almost every model is using just GQA and DSA.

22:33

But according to that, they have used this one latent vector in which they have used 512 as a dimension and some 64 for rope index and then they show that it is 56x more compressed than multi-head attention. Okay, yeah, so this is the trade-off diagram. So here I think we have not talked about this linear attention or mamba. So main problem is just all this matrix multiplication right now. Everyone is using attention. Suppose in future if we don't want to use attention rather than generating tokens sequentially, just use maybe diffusion models where we can generate everything simultaneously. So all these algorithms will change also. But here I think they have two more: one is linear attention and one is mamba. So according to this slide, if we are not compressing anything, so MHA is just we are parallelizing the process, so there is no quality loss, so it's good. And then this group query attention, which I think almost every model is using, just GQA and DSA kind of thing. Yeah, I think same thing we are providing in the attention mechanism scorecard. So I think this one MHA quality is good, throughput is okay. And for grouped query attention, it depends upon your use case also. Though quality is almost similar to multihead attention, but use case also matters a lot. Yeah, multiquery attention is just one extreme. We are, I don't know why, but we are just assuming that we only need one block and all the queries will attend to a smaller block. So quality is not that great for MQA. And this multihead latent attention, if you have tried some deep seek models, I think they are doing great job in quality wise. Besides that, sliding windows. So all these are some techniques which, yeah, all these are some techniques like just slide the windows all those things. And instead of, yeah, instead of multiplying everything, so linear attention is just saying summarize everything first and then look up into it. And then Mamba, this is just a state space model. Yeah, okay, thank you. So for the model like optimizations, we also have like two notebooks here. So there will be, I have to go to this. Okay, so for the quantization like the demo, is this already done? No, let me just run this. Okay, so we are loading the model which is MISTRAL 7B. So this one is like with the FP16 baseline. Wait, did it run? Okay, so it's two milliseconds run. Did this run? Okay, so yeah, this time it's fetching that model with the FP16 precision. The Wi-Fi, it's going to take time. Okay, yeah, because it's downloading the weights from the hugging face. Yeah, so Colab is running online because it needs to make the network call through the hugging face and it's fetching. I don't know, but it's taking time to download, probably. Okay, okay, okay, so here we see the memory size is around 15 GB approximately with the FP16 precision. We are trying to do the 2x compression as Talmeh talked about with the int8. Okay, so we do see your memory size is now 7.5 GB. What that means is now you have more memory for your KV to basically grow. That means you can either serve higher context limit or you can serve higher concurrent users there. If you do the int4, you are doing the 4x compression. So that with the 4x compression it would be more lower. It would be, I think, around 3 to 4 GB, 4.5 GB. And yep, so this is just a basic plot of like these are the theoretical numbers. We are not doing any throughput tests here, but usually you would see memory increases. So you would also have a bit of higher throughput from some of the benchmarks that we studied. We saw the int8 compression it does have a lower throughput. Okay, and then there is a demo on the attention mechanisms. So for the attention, okay, I have to run this. Okay, so it has run. Oh wait, why does it say no GPU detected? I should say the GPU should be detected. Okay, okay, okay. I guess it's not able to detect the GPU for some reason. We do have a GPU here. Okay, never mind. Yeah, so the basic idea here was more like as you try to move towards compressing the computation by using different attention mechanisms like moving from the multi-head to the grouped QA attention and then to the MLA, you would start seeing some optimizations. I think yesterday night we were doing some benchmarking. I wanted to correct this part so it wasn't 56x, it was 14x. Basically the demo had a mistake of a computation where it did not multiply the number of layers. Yeah, so apologies for that. So this MLA is a 14x savings in comparison to your multi-head attention.

22:36

So now that we have understanding of the pain points, the foundations, the one side of the optimizations which is the model optimizations, we want to talk about what can you do on the serving side. So the first thing is we saw when you perform a simple decode step, you are pulling the model weights and then you are recomputing the key and the value vectors for all the previous tokens, even though you already computed those vectors for the tokens. So there is definitely a lot of compute wastage. And if you analyze the time complexity of it, it would come out to be of n squared. And the way to resolve that is a classic trade-off against the memory. You can maintain a memory of those vectors against the tokens and you can reference that memory. So that memory was called KV cache. And the flow looks something like this.

22:40

And then based on this KV cache, there were four optimizations that were really possible. The first one is about page attention. So what's the problem today? When you send multiple requests as the input to the GPU, these requests are in a batch. Every request is allocated a continuous memory storage. Let's say, I'm just taking an example, let's say 2 KB. However, your request needed only, let's say, 1 KB. So there is 50% of that memory fragmentation. And this fragmentation leads to memory wastage. That means there was space in the memory where you could have served more requests, but you could not because you were looking for that contiguous block of memory. So inspiration was being taken from how the OS works. You maintain a logical memory and you basically have a physical memory. So in the logical memory, it would still feel like the KV vector for every token is contiguous, but it will be mapping to a different physical address. So that really helped saving a lot of memory. And it was only possible because they consider memory as a set of blocks and you would be dynamically allocating those blocks as the requests need, as new tokens come in and they need that kind of memory.

22:42

Another lever is when you are sending multiple requests in the batch, GPU is taking those requests for the next batch. But it does not accept the new batch unless all the requests in that batch get completed. So the diagram looks more like page attention, but here it is more about when is GPU available to take the next batch. So there is a time period where GPU is sitting really idle. And you want to resolve that. And for that, the idea was okay, let's do continuous batching. So continuous batching also really helped with throughput because now you can ship more requests pretty quickly, keep making sure GPU is always occupied and it's not sitting idle. So you are saving on that compute.

22:46

The third is prefix caching. So remember the KV cache helped you save the computation for a single request across the tokens. But what if you have the same tokens across multiple requests? How do you basically save against that? So the prefix caching, which was introduced by VLLM, exactly counters that.

22:48

And then the fourth is so we talked about quantizing the model, but you could also quantize the KV weights. So that means now you need lesser space for your key and the value vectors. That means you can serve more key and the value vectors in the memory. That means you can serve more tokens. That means you can serve more context limit. And that means you can serve more model quality. And all of this is already present in VLLM. You don't really need to reinvent that wheel. You can deploy this VLLM in production and you could see that basically growth.

22:54

So next we have a benchmark that we did. So this benchmark was, let me see if I have that. Here. The demos. So doing this benchmark takes around one hour because you have to continuously stop and restart the VLLM servers and you have to load the models and all. So it does take a lot of time in doing the testing, but I can tell you here what we are doing. So we have kept the model the same, the MISTRAL 7B. And then we have a set of input questions that we are sending. Consider them as the prompts. Then we have a couple of helper functions here like checking the server is up or not. This server is the VLLM server. Then there are helper functions to get the VLLM metrics. And I will talk about what those metrics are.

22:57

You can deploy this VLLM in production and you could see that growth. So next we have a benchmark that we did. So this benchmark was, let me see if I have that. Here. The demos. So doing this benchmark takes around one hour because you have to continuously stop and restart the VLLM servers and you have to load the models and all. So it does take a lot of time in doing the testing, but I can really tell you here what we are doing. So we have kept the model the same like the MISL 7B.

23:25

And then we have the set of input questions that we are sending. Consider them as the prompts. Then we have a couple of helper functions here like checking the server is up or not. This server is the VLLM server. Then there are helper functions to get the VLLM metrics. And I will talk about what those metrics are. Then there are a lot of the benchmarks and all. And then you have to measure that KV usage and all. These are the helper functions. So the baseline is very simple. We have a hugging face baseline. This is a raw sending the text to the LLM, getting back the response. We see some results here. We saw hugging face as a throughput of around 51 tokens per second.

24:07

Time to first token was 54. And then the inter-token latency was 19. This was all run on the H100. Right? And then we start a very default VLLM server. So by default, VLLM provides you the page retention, continuous batching, and the KV caching. So three things are present by default. And when you try to compare those benchmarks, you see your throughput is almost 15x. You are able to serve more tokens per second.

24:38

Then your time to the first token, that also rises. And then the inter-token latency goes down. And then your KV versus users and the versus context increases for sure. Now, when you apply the prefix caching to it, so with the prefix caching, you see your throughput increases more. Your TTFT decreases. Your inter-token latency is approximately same. And then your KV cache usage versus the users, it's going down. The versus the context, it's not going down. It's approximately same. I think this is also approximately same. It's not that big of a deal. When you apply the KV quantization on top of it, so it becomes, so you see the throughput is almost similar.

25:22

Your time to first token is similar. Your token latency is similar.

25:32

But then your KV usage actually goes down. This is because you have quantized your key value space. And then there is a concept of speculative decoding that Tanmay will talk about. So when you try to benchmark those, so you also see there is a bit of the less KV usage there. Although the results are approximately same. So yeah, I mean overall these are the metrics across probably I should. Great. So yeah, this is the VLLM benchmarks. It's your production default by the way. We will also share that decision tree when we try to talk about the other engines. So yeah, so we should talk about what are some of the other inference optimizations we can do on top of it.

26:05

So I would like to again invite Tanmay. He is going to talk about some of these optimizations. Oh sorry. I'm sorry. I didn't enable the slides. What was the, okay, great. I'm sorry. Okay. Okay. Which one? The speculative decoding. Yeah. Thank you, Harshal. Yeah.

26:43

So all these are speculative decoding. All these are the different flavors of same kind of soda. So this technique comes under decoding accelerator.

26:57

So first one, so we are only talking about this speculative decoding, but there are other variants like self speculative, eagle, medusa. I only, I think, this one, eagle algorithm. Personally, I don't think speculative decoding works because main problem is alignment. Okay. So let's start with what is speculative decoding. Main problem is that in transformer architecture, all these tokens are generated sequentially one by one by one. So how about just use a smaller model and let a smaller model to generate maybe let's say four or five tokens. And this teacher model, or we can say according to our world cup algorithm, we can say referee.

27:20

So referee will decide how many tokens it accepts. And this loop keeps on going on. And our assumption is that there are certain domain where this kind of things will work, like maybe in decode, maybe in coding, or where there is no creativity. Each code or syntax is almost similar. So maybe it can help it. But based on personal testing, I didn't find this speculative decoding useful at all. But other techniques like self speculative decoding where teacher model also have one head, auxiliary head, and it will do similar kind of things about this base model or small model is doing it. But then this eagle came, eagle one, two, three, I don't know how many versions are.

27:39

But it is saying that instead of creating, instead of generating tokens, let's train a small model and just take features from one of its main models layer so that instead of generating token, it will generate this feature. So eagle is better compared to this other kind of technologies. And then another one is Medusa, which is saying that just generate all these tokens parallelly. Okay, so here in this slide. Yeah. The next slide. Okay. Okay, yeah. Okay, now we will come to this one prefix caching. So I don't know whether people are using this one static prefix caching or not.

28:07

But the thing is that the main problem with prefix caching is that sometimes we type and make a small kind of mistake. And this standard static prefix caching is basically it takes a prompt, do some hashing. And then next time when user asks similar kind of question, it will try to match the hash.

28:21

So if hash is equal, then it will instead of recomputing all those K and V, it will just take it from the storage.

28:27

And then it will take it from the storage.

28:31

But you know that sometimes we make a mistake or maybe we just change a word or letter, something like that. Then we have a very higher cache miss hit rate. So that's why this one, RedixTree. So RedixTree is becoming very popular and also because of agent. So I think almost everyone is doing agent and most of the computation is going during test time, inference kind of thing, where we keep on asking same kind of questions and prompt. For example, you are an expert software engineer, multiplied by 200 times. This kind of loop keeps on going inside this agentic kind of things, where it is necessary to keep or store similar kind of things in a RedixTree.

28:51

So RedixTree is just an advanced version of this prefix tree where we will just collapse and does not have any branch. And for this kind of work where we keep on repeating same thing, this RedixTree helps a lot. And sglang use this kind of algorithm for prefix caching.

29:04

Okay, yeah, then there is another thing. One is tensor RT LLM. This is very confusing. When I first started, I was confused. What is tensor RT LLM? So yeah, so tensor RT is a standard SDK kind of thing. Tensor RT LLM is an inference engine, just like VLLM, sglang. But problem is that it is related to NVIDIA. They optimized each and every layer and every problem.

29:33

And then, as I mentioned in our World Cup algorithm, they just break everything and optimized everything at hardware level also. So, yeah.

29:42

So, okay, next. Yeah, so for this workshop, we also did some benchmarking, which is best. So, our setup was something similar. So, we did two kinds of testing. First one is without agentic testing, where we just, so we use the shared GPT, this one, data set and just ask those questions using VLLM and sglang. Okay. Okay. Okay. Okay. Let me just zoom it up. Okay. Great. Okay, yeah. So, yeah, for this workshop, we used H100 and our first testing was that we just asked, or we take questions from shared GPT and put it into VLLM, sglang, and we found that actually there's no statistical difference between which one is better. So, both have almost similar kind.

30:24

So, both are fulfilling similar kind of request per second, RTTFT and latency. So, but only difference we have seen during agentic branching. So, what we did was that we asked that similar kind of question that you are the best, this one, software engineer in the world. Just solve the problem of traffic congestion in the city kind of thing. So, then we put this into LLM. LLM generates some output. Then we did another round two also. So, once this LLM generates this output, then in round two, we have especially mentioned that provide, review the proposal and give ratings from one to ten. So, these are two terms we did and this loop keeps on repeating it.

30:41

What we found is that for this kind of workflow where everything is standard, all those prompts and context engineering comes into the picture. There's no statistical difference between which one is better. Both have almost similar performance. Both are fulfilling similar requests per second, RTTFT and latency. The only difference we have seen is during agentic branching. What we did was ask a similar question: "You are the best software engineer in the world. Just solve the problem of traffic congestion in the city."

30:49

We put this into the LLM. The LLM generates some output. Then we did another round two. Once the LLM generates this output, in round two, we especially mentioned: provide a review of the proposal and give ratings from one to ten. These are two terms we did and this loop keeps repeating. What we found is that for this kind of workflow where everything is standard, all those prompts and context engineering come into the picture. If we do proper agentic branching, then I think sglang is three to four times better. But again, this depends upon the different setup. If you do it, you may get different results.

30:53

Did we upload it on GitHub? The PDF is also in the drive. It's the same link as the slides. Quick summary here. On a standard API workload throughput, you would see VLLM and sglang would be the same. If you have a standard workload, definitely go with VLLM. It's the production default anyway. But what Tanmay was also saying is when you try to make it agentic workloads, that is where sglang really shines. Keep VLLM as a default. But if you have agentic workloads, probably try to move to sglang if you're not happy with the VLLM part.

30:58

There is also a comparison done at 120 billion for the GPT OSS 120 billion. This is a benchmark prepared by ClivePy. There is a blog link here. They did a similar benchmark and included tensor RT LLM in it. You can always go through these benchmarks and try to understand which suits your use case. As we mentioned, tensor RT tries to optimize the hardware side as well, having peak hardware performance.

31:00

When you want to depict your engines, once you figure out between VLLM, sglang, and tensor RT, there are some new engines popping up. Nvidia Dynamo for sure for agentic session routing. Hugging Face is always there. There is an M-star engine recently proposed by Stanford for multi-model. Definitely you could explore those.

31:02

To give a quick summary, you start with a baseline. Try to find what model could fit your use cases. You could pick DeepSeq or Mistral 7B. You pick your model and want to have smaller memory and fit that bigger model into smaller memory to save cost on GPU. You can do all those quantization techniques. Then you can apply all those serving optimizations by using the right serving engine. That can provide you the throughput you really want.

31:03

Something you can do after going home is reading about source information, different attention mechanisms, different engines, try to read the different benchmarks online. There are in-depth guides on learning about KV eviction strategies. The world is moving towards having a separate KV cache engineering domain. You want to understand what's going on there. KV eviction, cache compressions, hybrid memories. There are many solutions happening there. Always try to stick to those foundations or fundamentals or first principles and try to see which solution solves what problem and whether you actually need that problem solved for your use case.

31:04

There is distributed LLM inference, which is a different topic altogether. You would probably need a two-hour workshop there as well to go over all the internals and hands-on. This is something we are trying to propose for the AI Engine in New York session to dive deeper into advanced sections of LLM inference. This workshop was more for beginner and intermediate level. In this form, we have feedback as well as interest. If you think we need certain improvements in certain sections, definitely give that feedback. If you want to see this workshop in New York Fair, definitely feel free to enroll your interest. The URL works. I forgot to link the QR code together.

31:09

If you can give that feedback, I'll just note it. That will be fine. I think we would like to wrap this workshop. I'm sure many of you would have questions, so we can take those offline and meet to talk about those questions. Thank you, everyone. Thanks for joining. I think it was really meaningful. Thanks a lot. definitely the quality and you want to prioritize the like the latency. You would not want your users to wait infinitely for the like not infinitely, but probably for the larger latency. You can always sacrifice the number of users you can support on the GPU and probably take that costed being more customer obsessed. In form of like, and like if you consider

31:43

like an agent, sorry, the async agent workload, you would want to like prioritize definitely quality and the throughput because these are the long running tasks and you would want to like serve as many as concurrent tasks as possible, but with a very, very higher quality.

32:07

and often like we think like, okay, the GPU is like a very expensive GPU. That might not be a good fit for us, but it turns out that could actually serve you the lowest cost per million tokens, but you really have to trust your kind of calculations on the max users that you want and like you really have to make those estimations correctly.

32:48

So, we do have like let me just okay, great. So, for the capacity calculator, there is like a link to the collab because I was facing certain issues with Molab. I had to migrate out of the whole widget library and I didn't have time so being lazy I just picked collab there. Apologies to Molab.

33:26

So, my vision is connected.

33:34

okay.

33:45

Wi-Fi probably.

33:51

Okay, great.

34:01

So, what we have done over here is we have like shaded some like the GPUs with their VRAMs and bandwidths, the flip-flops and the cost per hours.

34:14

Then, we kind of like build this simple like capacity calculator. This is just a KV visualizer where you kind of like when you increase the number of tokens you see like a KV size it increases and when you increase the number of users your size is like increasing at a much faster rate.

34:37

And then in this capacity calculator let it run. So, we have like a model which is like a 7 billion parameter model that we selected. we set the like precision to be FP16. Now, we decide the way we basically go by the GPU decision is you have to decide what's your like you have to fix one dimension first which you care about the most. For premium chat I mentioned like latency is definitely the one. and then like for the async workloads the minimum batch size that you want to solve from like a single GPU that is the second dimension. So, you want to fix these first. So, I will go about like in a premium chat application so I can go ahead with like 10 milliseconds latency.

35:43

A minimum batch size I don't care like I can so I'm okay with like probably two uh okay so probably with the seven concurrent users on a single GPU and then like my context limit is very important to me because I want to focus on the quality as well. And so like I do see like some of the GPUs so the H108 GB it's like an $8 per hour but like like 300x is it? Yeah so it's like around $10 per hour but if you do all that throughput math that we shared in the mathematics before you could find like your cost per million dollar tokens that could be very very that could be like lesser. So you need to do such calculations by fixing those dimensions and you need to decide your GPU

36:45

to like reduce your kind of inference costs. This is at least the first step that you can take towards optimizing the inference.

36:56

Okay. Cool.

37:00

So the next slide let me okay great and so like now the next thing is about the model optimization so we are now basically have built that foundation where we understood some of the pain points reason behind those pain points why those were happening how we could like address that GPU capacity thing we need to understand what can we do like what can we further do about it so it is about model optimization and I think I would like to invite Tanmay he can talk more about these model optimizations provided he has work on this like during his research times okay I can control yeah okay hi everyone mic check am I audible at last yeah okay so hi I'm Tanmay I work as a senior

38:04

quant modeler and also I'm an AI researcher my work focuses on agent verification and right now building world models so for this one model optimization before we start model optimization so I created a research template so that it will be easy for us to understand all these complex things so our template is simple first we will identify the problem second step we will solve the problem using two algorithms these are just fake algorithms so first algorithm is called ostrich algorithm whenever we see just like ostrich whenever we see a problem ostrich put their head into the sand so same thing we will do whenever we face a problem we will just ignore it so this is an

38:52

important algorithm we should follow second one is created it is called world cup algorithm for example we don't know who will win this FIFA world cup so what organizers did they break the 48 teams into 12 groups then round 32 so round 32 right now is currently going on then round 16 then quarterfinals then semifinals and finals so what they are doing is that they are breaking it into a smaller problems and the useful results are moving forward so same analogy or same algorithm we will use to understand this model optimization all those things so yeah let's start so I have one H100 GPU I have to use this open source model what is called GPT OSS 120 billion parameter model

39:53

so right now I think they have trained it on BF float 16 and weight is 240 gigabyte what should I do this is the problem we have so first thing what we have to do is that 240 gigabyte and 80 gigabyte H100 so and I have to fit only in one GPU not in multiple GPU so what can we do I think simple step is that just compress it but how should we compress it that's another challenge so if we compress BF float 16 to FP8 then it will be around 120 gigabyte but our GPU H100 is still 80 gigabyte so what I think they did is that they compressed it into further MXFP4 and I think size is around 65 gigabyte so this is something we can do compress but question so and we will use

41:00

over this ostrich algorithm we are assuming that there is no loss in compressing a bigger model into a smaller size second thing in this one okay yeah so in this one in this slide we have used this Mistral 7b so 7 billion parameters so it's a small model 7 billion parameters so if you multiply it by 2 bytes so weight of it's around is 14 14.5 gigabyte which can easily fit into H100 or even A40 so so next what we can do is that like Mistral 7b instead of compressing it floating point 16 we can apply different techniques like int 8 or int 4 or nf 4 so basically we have to just use ostrich algorithm and just believe that there is no quality loss kind of things but somehow

42:01

we also have to mathematically prove that by doing some kind of testing on some external benchmark that whether it is working or not so and this comes under post-training quantization kind of thing one can also do this one during fine-tuning one can also do this kind of quantization this comes under a quant-aware training kind of thing so let's move to our next problem so we have this huge matrices just imagine 1,000 by 1,000 dimension matrix A and another matrix 1,000 by 1,000 so if you multiply by this two matrices so number of operations will be 1,000 raised to the power q and this is kind of a problem in terms of computing so we wondered our matrix multiplication

43:08

should be fast and it should save memory so what should we do we have a giant matrix okay let's take this one Mr. 4096 by 4096 what should we do to solve our problem of speeding up the things and saving the memory 4096 by 4096 so first thing is that we will use just our world cup algorithm we can decide a random number just break the block vertically it does not matter what you are choosing it so you have so let's say we have 4096 columns we will break it we will break it into a group of 128 column each so 128 128 128 128 128 vertically so we will get this 32 blocks if we divide this 4096 then what will happen by doing this thing so if we just divide this one vertically

44:17

then we can use a multiple GPU to speed up the process so this kind of thing is called multi-head attention so what else can we do we have a big matrix like as I have mentioned that ostrich algorithm so our main problem is sizing so what we can do is that instead of having all those 32 vertical blocks we will throw away 31 blocks and we will assume that one block is sufficient enough that all the queries can handle those blocks our loss will be almost negligible and we come up with this algorithm and this algorithm is called multi-query attention so as we can see right now we are at two spectrum one is multi-head attention where we split it into 32 blocks and use

45:18

different GPUs or do some parallel processing and at the same time we are just throwing 31 blocks and we are calling this as a multi-query attention so at both extremes we should come up with a middle ground like something we can say that instead of throwing all the 31 maybe we can group some of the blocks together so that and we can assume that similar blocks will attend to similar kind of queries so this kind of technique comes under grouped query attention which is very popular right now even in even in Mistral or in other models this grouped query attention works so right now we have understand that we have a big matrix we can divide it the way we want and doing some

46:18

mathematical calculation prove that loss is almost negligible kind of thing so what else we can do so after that after this grouped query attention see we have a big matrix one is key and one is value let's compress that matrix into a latent vector and then come up with some algorithm to reconstruct from latent vector to our original matrix so this kind of strategy comes under this one multi had latent attention but again it has some problems with rope because rope is position dependent and it is position independent kind of thing so one needs to also include some index for keys also so that one can map it but again main problem is why we are multiplying all those

47:23

big matrices so because that's how this attention mechanism works that each token will pay attention to every token so how about let's don't pay attention to all the previous tokens only pay attention to the important tokens which is important for us so this kind of field is evolving so this comes in deep seek sparse attention so yeah and yeah so okay next yeah so next one is flash attention so in flash attention so main problem is that so currently so currently not currently so right now almost everyone uses flash attention but way in 2022 or 2023 so that's how it works that's how it works is that so this Q K query and key matrices they were in HBM it loads it first

48:37

it loads into this one our tensor core and it do some calculation and then it will write it back to HBM and then this process goes on multiple times so in flash attention what they did is that instead of multiplying the whole matrices so they just divided it into like our world cup algorithm divided the bigger matrices into a small tile and only put those small tiles into SBM so that it can process multiplication fast and just keep keeping track of this some three variables so that they can calculate this online softmax yeah next one so yeah so this is just mathematics so if we have a multi-head attention if it is 524 kV then it depends upon how much how much grouping we

49:39

want and so if instead of 32 kV head we only want to use 8 kV heads so we can get a compression of 4x times and this multi-head latent attention this formula depends on the model to model how many layers your model have so in the original deep seek paper I think they have some 128 dimension 128 I don't remember the exact dimension but according to that they have used this one latent vector in which they have used 512 as a dimension and some 64 for rope index and then they show that it is 56 x more compressed than multi-head attention okay yeah so this is the trade-off diagram so here I think we have not talked about this linear attention or mamba so main problem is just

50:52

all this matrix multiplication right now everyone is using attention suppose in future if we don't want to use attention rather than generating tokens sequentially just use maybe diffusion models where we can generate everything simultaneously so all these algorithms will change also but here I think they have two more one is linear attention and one is mamba so according to this slide so if we are not compressing anything so MHA is just we are parallelizing the process so there is no quality loss so it's a good and then this group query attention which is I think almost every model is using just GQA and DSA kind of thing yeah I think same thing we are providing in the

51:46

attention mechanism scorecard so I think this one MHA quality is good throughput is okay and for grouped query attention it depends upon your use case also though quality is almost similar to multihead attention but use case also matters a lot yeah multiquery attention is just one extreme we are I don't know why but we are just assuming that we only need one block and all the queries will attend to a smaller block so quality is not that great for MQA and this multihead latent attention so if you have tried some deep seat models so I think they are doing great job in quality wise besides that sliding windows so all these are some techniques which yeah all these are some

52:50

techniques like just slide the windows all those things and instead of yeah instead of multiplying everything so linear attention is just saying summarize everything first and then look up into it and then Mamba this is just a state space model yeah okay thank you so for the model like optimizations we also have like two notebooks here so there will be I have to go to this okay so for the quantization like the demo this is is this already done no let me just run this okay so we are loading

54:04

the model which is like MISTRL 7b so so this one is like with the FP16 baseline wait did it run okay so it's two milliseconds run did this run okay so yeah this time it's fetching that model with the FP16 precision the Wi-Fi it's gonna take time okay yeah because it's downloading the weights from the hugging face yeah

55:20

so molab is like running online because it needs to make the network call through the hugging face and like it fetching I don't know like but it's taking time to download probably okay okay okay so here we see like the memory size is like 15 GB around approximately with the FP16 precision we are trying to do the 2x compression as Talmeh talked about with the int 8 okay so we do see like your memory size is now like 0.7 5 GB what that means is now you have a more memory for your KV to basically grow that means you can either serve higher context limit or you can serve the higher concurrent users there

56:35

if you do the like in 4 you are doing the 4x compression so that with the 4x compression it would be more lower it would be I think around 3 to 4 GB 4.5 GB and yep so this is just a basic plot of like so these are the like theoretical numbers we are not doing the like any throughput tests here but usually you would see like a memory increases so you would also have like a bit of higher throughput from some of the benchmarks that we studied we saw like the intake compression it does have like a lower throughput okay and then there is like a demo on the like the attention mechanisms so for the attention okay I have to run this

57:58

okay so it has run oh wait why does it say no GPU detected I should say the GPU should be detected okay okay okay .

58:56

I guess it's not like able to detect the GPU for some reason. We do have a GPU here. Okay, never mind. Yeah, so but the basic idea here was more like as you try to move towards compressing the computation by using different attention mechanisms like moving from the multi-head to the grouped QD attention and then to the MLA. You would start seeing some optimizations. I think yesterday night we were doing some benchmarking. I wanted to correct this part so it wasn't like 56x, it was 14x. Basically the demo had a mistake of like a computation where it did not multiply the number of layers.

1:00:01

Yeah, so apologies for that. So this MLA is like a 14x savings worth in comparison to like your multi-head attention.

1:00:15

So now that we have understanding of the pain points, the foundations, the one side of the optimizations which is the model optimizations, we want to talk about what can you do on the like the serving side. So the first thing is we saw like when you perform like a simple decode step, you are pulling it, you are basically pulling the model weights and then you are recomputing the key and the value vectors for all the previous tokens, even though you already computed those vectors for the tokens. So there is definitely like a lot of compute wastage and if you kind of analyze the time complexity of it, it would come out to be of n squared.

1:01:08

And the way to resolve that is like a classic trade-off against the memory. You can maintain a memory of those vectors against the tokens and you can reference that memory. So that memory was called as like KV cache. And like the flow looks something like this.

1:01:29

And then based on this KV cache, there were like four optimizations that were really possible.

1:01:37

The first one is about the page detention. So what's the problem today? So when you send like multiple requests as the input to the GPU, these requests are in a batch, every request is allocated like a continuous memory storage. Let's say of, I'm just taking an example, like let's say 2 KB. However, like your request needed only let's say 1 KB. So there is like a 50% of that memory fragmentation. And this fragmentation basically leads to the memory wastage. That means there was a space in the memory where you could have served more requests, but you could not because you were looking for that contiguous block of the memory. So an inspiration to was being taken from

1:02:39

like how the OS works. Like you maintain a logical memory and you basically have a physical memory. So in the logical memory, it would still feel like that the KB vector for the like every token is like a contiguous, but it will be mapping to a different physical address.

1:03:06

So that really helped like saving a lot of memory. And it was only possible because they consider like memory as a set of blocks and you would be dynamically allocating those blocks as the request need. As the like new tokens comes in and they need that kind of memory.

1:03:30

Another lever is like when you are sending multiple requests in the batch, GPU is like taking those requests, the next batch. But it does not accept the new batch unless all the requests in that batch gets completed. So the diagram looks more like a paid attention, but here it is more about like when is GPU available to take the next batch. So there is a time period where GPU is like sitting really idle. And you want to like resolve for that. And for that, like the idea was like, okay, let's do that continuous batching.

1:04:17

So the continuous batching also really helped with like throughput because now you can ship more requests pretty quickly. Keep making sure like GPU always is always like occupied and it's not like sitting idle. So you are saving on that compute. The third is the like prefix caching. So you remember like the KV cache held you save the computation for a single request across the tokens. But what if like you have the same tokens across multiple requests? How do you basically save against that? So the prefix caching, which was introduced by the VLLM, exactly counters that. And then the third is like we talked about the fourth, actually.

1:05:10

So we talked about quantizing the model. But you could also, you can also like quantize the KV weights. So that means now you need like a lesser space for your key and the value vectors. That means you can serve more key and the value vectors in the memory. And that means like you can serve more tokens. That means you can serve more context limit. And that means like you can serve more model quality.

1:05:44

And all of this is like already present in the VLLM. You don't really need to reinvent that wheel. You can deploy this VLLM in production and you could see that basically growth. So next we have like a benchmark that we did. So this benchmark was, let me see if I have that. Here. The demos. So doing this benchmark takes like around one hour because you have to continuously stop and like restart the VLLM servers and you have to load the models and all. So it does take a lot of time in doing the testing, but I can like really tell you here what we are doing. So we have kept the model the same like the MISL 7B.

1:06:48

And then we have like the set of input questions that we are sending. Consider them as the prompts. Then we have a couple of helper functions here like checking the server is up or not. This server is the VLLM server. Then there are helper functions to get the VLLM metrics. And I will talk about like what those metrics are. Then there are like a lot of the benchmarks and all. And then you have to measure that KV usage and all. These are the like helper functions. So the baseline is very simple. Like we have a hugging face baseline. This is a raw like sending the text to the LLM, getting back the response. We see some results here.

1:07:38

We saw like hugging face as a throughput of like around 51 tokens per second. Time to first token was like 54. And then the inter-token latency was 19. This was all run on the H100. Right?

1:07:55

And then we start like a very default VLLM server. So by default, VLLM provides you the page retention, continuous batching, and the KV caching. So three things are present by default. And when you try to compare those benchmarks, you see your throughput is like almost 15x. You are able to serve more tokens per second. Then your time to the first token, that also rises. And then the inter-token latency kind of goes down. And then your KV versus users and the versus context increases for sure.

1:08:47

Now, when you apply the prefix caching to it, so with the prefix caching, you see like your throughput increases more. Your TTFT decreases. Your inter-token latency is approximately same. And then your KV cache usage versus the users, it's kind of going down. The versus the context, it's not going down. It's approximately same. I think this is also approximately same. It's like not that big of a deal. When you apply the like KV quantization on top of it, so it becomes like, so you see like the throughput is like almost similar. Your time to first token is similar. Your token latency is similar. But then your KV usage actually goes down.

1:09:42

This is because like you have quantized your key value space.

1:09:49

And then there is a concept of speculative decoding that Tanmay will talk about. So when you try to benchmark those, so you also see like there is a bit of like the less KV usage there. Although like the results are approximately same.

1:10:15

So yeah, I mean overall like these are the like the metrics across probably I should.

1:10:27

Great.

1:10:30

So yeah, this is the like VLLM benchmarks. It's your production default by the way. We will also share that decision tree when we try to talk about like the other engines.

1:10:50

So yeah, so we should talk about like what are some of the other inference optimizations we can do on top of it.

1:11:04

So I would like to again invite Tanmay. He is going to talk about like some of these optimizations.

1:11:18

Oh sorry. I'm so sorry. I didn't enable the slides.

1:11:25

What was the, okay, great. I'm sorry. Okay. Okay. Which one? The speculative decoding. Yeah. Thank you, Harshal. Yeah. So all these are like speculative decoding. All these are the, so what we say, different flavors of same kind of soda. So this technique comes under decoding accelerator. So first one, so we are only talking about this speculative decoding, but there are other variants like self speculative, eagle, medusa. I only like, I think, this one, eagle algorithm. Personally, I don't think speculative decoding works because main problem is alignment. Okay. So let's start with what is speculative decoding.

1:12:17

Main problem is that in transformer architecture, all these tokens are generated sequentially one by one by one. So how about just use a smaller model and let a smaller model to generate maybe let's say four or five tokens. And this teacher model, or we can say according to our world cup algorithm, we can say referee. So referee will decide how many tokens it accepts. And this loop keeps on going on. And our assumption is that there are certain domain where this kind of things will work, like maybe in decode, maybe in coding, or where almost there is no creativity. Each code or syntax is almost similar. So maybe it can help it.

1:13:11

But based on personal testing, I didn't find this speculative decoding useful at all. But other techniques like self speculative decoding where teacher model also have one head, auxiliary head, and it will do similar kind of things about this base model or small model is doing it. But then this eagle came, eagle one, two, three, I don't know how many versions are. But it is just saying that instead of creating, instead of generating tokens, let's train a small model and just take features from one of its main models layer so that instead of generating token, it will generate this feature. So eagle is better compared to this other kind of technologies.

1:14:10

And then another one is Medusa, which is just saying that just generate all these tokens parallelly. Okay, so here in this slide. Yeah. The next slide. Okay.

1:14:29

Okay, yeah. Okay, now we will come to this one prefix caching. So I don't know whether people are using this one static prefix caching or not. But the thing is that the main problem with prefix caching is that sometimes we type and make a small kind of mistake. And this standard static prefix caching is basically it takes a prompt, do some hashing. And then next time when user asks similar kind of question, it will try to match the hash. So if hash is equal, then it will instead of recomputing all those K and B, it will just take it from the storage. And then it will take it from the storage.

1:15:10

But you know that sometimes we make a mistake or maybe we can just change a word or letter, something like that. Then we have a very higher cache miss hit rate. So that's why this one, RedixTree. So RedixTree is becoming very popular and also because of agent. So I think almost everyone is doing agent and most of the computation is going during test time, inference kind of thing, where we keep on asking same kind of questions and prompt. For example, you are an expert software engineer, multiplied by 200 times. This kind of loop keeps on going inside this agentic kind of things, where it is necessary to keep or store similar kind of things in a RedixTree.

1:16:03

So RedixTree is just an advanced version of this prefix tree where we will just collapse and does not have any branch. And for this kind of work where we keep on repeating same thing, this RedixTree helps a lot. And sglang use this kind of algorithm for prefix caching. Okay, yeah, then there is another thing. One is tensor, RT, LLM. This is very confusing. When I first started, I was just confused. What is tensor, RT, LLM? So yeah, so tensor, RT is just a standard SDK kind of thing. Tensor, RT, LLM is just an inference engine, just like VLM, sglang. But problem is that it is related to NVIDIA. They optimized each and every layer and every problem.

1:17:10

And then, as I mentioned in our World Cup algorithm, they just break everything and optimized everything at hardware level also. So, yeah. So, okay, next.

1:17:26

Yeah, so for this workshop, we also did some benchmarking, like which is best. So, our setup was something similar. So, we did two kinds of testing. First one is without agentic testing, where we just, so we use the shared GPT, this one, data set and just ask those questions using VLM and sglang. Okay.

1:18:05

Okay. Okay. Okay. Let me just zoom it up. Okay. Great. Okay, yeah. So, yeah, for this workshop, we used H100 and our first testing was that we just asked, or we take questions from shared GPT and put it into VLM, sglang, and we found that actually there's no statistical difference between which one is better. So, both have almost similar kind. So, both are fulfilling similar kind of request per second, RTTFT and latency. So, but only difference we have seen during agentic branching. So, what we did was that we asked that similar kind of question that you are the best, this one, software engineer in the world.

1:18:59

Just solve the problem of traffic congestion in the city kind of thing. So, then we put this into LLM. LLM generates some output. Then we did another round two also. So, once this LLM generates this output, then in round two, we have especially mentioned that provide, review the proposal and give ratings from one to ten. So, these are two terms we did and this loop keeps on repeating it. What we found is that for this kind of workflow where everything is standard, all those prompts and context engineering comes into the picture. If we do proper this agentic branching, then I think this hglang is three to four times better.

1:19:54

But again, this depends upon the different setup maybe. If you do it, you may get different results. Okay. Yeah, so I think, did we upload it on GitHub? Okay. Yeah, so the PDF is like also in the drive. It's the same link as the slides. So, a quick summary here. So, on a standard API workload throughput, you would see like a VLLM and the sglang would be the same. So, if you don't have, if you have like a standard workload, definitely go with VLLM. It's the production default anyways. But what Tanmay was also saying is when you try to like make it like agentic workloads, that is where like your sglang really shines.

1:20:57

So, yeah, keep like VLLM as a default. But if you have agentic workloads, probably try to move as the sglang. If you're not happy with the VLLM part.

1:21:12

Okay.

1:21:17

Okay.

1:21:25

And then like there is like the, like a comparison that is done at the 120 billion, like for the GPT OSS 120 billion. This is a benchmark that was prepared by ClivePy. So, there is like a blog link here. Oh, nice. Okay. Yeah. So, they did the similar benchmark and they included like a tensor RT LLM in it. Definitely, you can always go through these benchmarks and try to understand which basically suits your use case. As we mentioned like tensor RT, they try to optimize the hardware side as well, having the peak hardware performance.

1:22:16

Wait. This is. Okay. Okay.

1:22:25

Okay. Yeah. And then like in terms of when you want to depict like your engines. Once you figure out like between VLLM, sglangt, and tensor RT. So, there are some new engines that are popping up. Nvidia Dynamo for sure. So, they are also for the agentic session routing. Hugging faces are always there. It's a simple nose over. So, there is like an M-star engine that was recently proposed by Stanford. They are for like multi-model. So, definitely you could explore those. And then you try to basically, just to like give a quick summary, you start with like a baseline. We try to find what model could fit our use cases.

1:23:24

So, you could pick like DeepSeq. You could pick like, don't pick like a Mistral 7B. I mean, it's not good. But, yeah. So, you pick your model and you want to like have a smaller memory and you want to try to fit that bigger model into smaller memory. So, that you could save cost on the GPU cost. So, you can do like all those quantization. Then you can apply all those serving optimizations by using the right serving engine under the hood. So, that can really provide you that throughput that you really want.

1:24:09

And now, something that you can do after going back home, probably because we cannot like actually go over all the material here, is definitely reading about some of the source information, like different attention mechanisms, different like these engines, like try to just read the different benchmarks which are present online as well. And then, there are a lot of like in-depth guides or the next phases of it, which is like learning about some KV eviction strategies. So, the world is moving towards having a separate KV cache engineering domain. So, you want to understand what's going on in there. So, KV eviction, cache compressions, hybrid memories.

1:24:57

So, there are like a lot of solutions that are happening around there. So, always try to stick to those foundations or like the fundamentals or the first principles. And try to see like which solution basically solves what problem. And whether you actually need that problem to be solved for your use case. And then, there is like distributed LLM inference, which is like a different pinpoint altogether. You would probably need like a two-hour workshop there as well. To like go over like all the internals, do all the hands-on.

1:25:40

Yeah, and this is something we are trying to propose for the AI Engine in New York session, which is to like dive deeper into the advanced sections of the LLM inference. So, this workshop was more for the like beginner and the intermediate level. So, in this form, we do have like a feedback as well, plus also the interest. If you think like we need certain improvements in certain sections, definitely give that feedback as well. And if you want to see this workshop in like New York Fair, I mean definitely feel free to enroll your interest. Huh? How is it possible? Boom.

1:26:33

Let me just check. Let me just check. Huh? Yeah. URL works, right? Not the QR code, okay. Probably I forgot to link those two together.

1:26:59

Uh... ZG... G, G, A, F...

1:27:11

Okay, cool. Yeah, so if you can give that feedback, let me just... Oh, okay.

1:27:20

That will be fine. And yeah, I think we would like to wrap this workshop then. Uh... And I'm sure like a lot of you would be having a lot of questions, so we can take all those like offline. Uh... We can meet, uh... And we can, uh... Like talk about those questions. Yeah, sure. Uh... Thank you, everyone. Thanks for joining. Uh... I think it was really meaningful and... All of you like came here. Uh... Thanks a lot. Yeah, thanks. Thank you.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note