AI Engineer

Two Bugs That Hid in Plain Sight: A vLLM Debugging Detective Story — Asaf Gardin & Yuval Belfer

1808 summary words 8 min summary Watch video

Start with the signal

8 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Silent, high-confidence inference corruption in stateful serving systems is best debugged by creating deterministic failures under controlled pressure, comparing token log-probability distributions against a trusted baseline, and tracing request identity through the engine.
  • Why it matters: For production agent or model-serving systems, the most damaging failures may not crash or alert; they can return plausible-looking but wrong output because scheduler, cache, or index bugs contaminate state across requests.
  • Best use: Use this as a concrete debugging playbook for vLLM or other stateful inference stacks, especially when investigating rare quality regressions, unexplained log-prob drift, or multi-tenant cache behavior.

Executive Summary

AI21 recounts two production-adjacent bugs encountered while serving and RL-training Jamba, its hybrid Transformer/Mamba model, through vLLM. Both defects produced silent corruption rather than exceptions: rare gibberish generations and deterministic log-probability spikes despite identical model weights and inputs. The speakers argue that such cases are engineering and systems failures, not model-quality problems, because the system can be confidently wrong with no visible crash.

The first bug appeared roughly once per 1,000 requests, only after sustained concurrent load, and only in vLLM. The team made the failure reproducible by reducing vLLM GPU-memory utilization from 90% to 20%, increasing concurrent traffic, and sampling at temperature zero. They then compared vLLM token log probabilities with a Hugging Face Transformers prefill baseline, instrumented vLLM to carry request IDs into the model forward pass, and discovered that the scheduler could send a never-computed Mamba request into decode before prefill. That decode read stale state left by a prior request.

The second bug surfaced during RL post-training as log-prob differences between rollout and FSDP evaluation before any weight update. Increasing rollouts per prompt from 8 to 128 moved the deterministic failure from every 12 steps to the first step. Contrary to the first case, reducing memory pressure removed the issue because it prevented a state-cache offset from reaching a 32-bit unsigned-integer overflow threshold of roughly 4 billion elements. Changing the cache index type from uint32 to size_t resolved it.

The enduring lesson is methodological: establish an independent numerical baseline, use knobs that change the failure's shape rather than merely its frequency, deliberately stress scheduling and memory boundaries, and add observability where frameworks erase request identity. The talk is especially valuable because it distinguishes a bad kernel from a correctly implemented kernel invoked at the wrong time with stale state.

Key Takeaways

  • Claim: Rare gibberish output in a stateful inference engine should be treated as a systems-integrity incident, not automatically as model-quality degradation. | Evidence: AI21 saw high-confidence gibberish only around one in 1,000 requests, after workload accumulated, and only in vLLM rather than other inference frameworks; there was no crash, warning, or error. | Implication: Ken should ensure quality monitoring captures rare, high-confidence anomalies and numerical drift, rather than relying on crash/error telemetry or aggregate benchmark scores. | Caveat: The diagnosis is specific to Jamba's Mamba state handling in vLLM, but the silent-failure pattern generalizes to other engines with reusable per-request state or caches.
  • Claim: The fastest path to debugging rare nondeterministic-looking inference failures is to make them deterministic under constrained conditions. | Evidence: Reducing vLLM GPU memory utilization from 90% to 20%, sending many concurrent requests, and setting temperature to zero made a previously rare failure recur on the same request, enabling a rapid feedback loop. | Implication: Build stress harnesses that vary memory allocation, concurrency, batch composition, rollout count, and deterministic sampling so failures can be reproduced quickly and repeatedly. | Caveat: Memory pressure is a diagnostic lever, not proof of a memory-capacity root cause; in the second case, reducing allocated memory actually hid the defect.
  • Claim: Comparing token-level log-probability distributions against a simpler trusted implementation can localize silent inference divergence before text output alone reveals its cause. | Evidence: The team generated outputs and log probs in vLLM, fed the full prompt-plus-generation sequence through Hugging Face Transformers' forward prefill pass, applied softmax to its logits, and computed divergence token by token. | Implication: Maintain a correctness oracle for critical models—potentially a slower reference server or offline forward-pass evaluator—and automate log-prob divergence checks in canaries and release validation. | Caveat: A baseline must be configured comparably enough to make numerical differences meaningful; the talk uses Transformers as a deliberately plain implementation relative to vLLM's optimized engine.
  • Claim: The first defect was a scheduler/request-lifecycle bug: a Mamba request could be decoded before it had ever been prefetched, causing stale state from an earlier request to be read. | Evidence: After propagating request IDs into vLLM's forward context and breakpointing on the corrupted request, AI21 found its first forward pass was classified as decode rather than prefill. The eventual fix marked requests with zero previously computed tokens as prefill so they could not be decoded or chunked first. | Implication: For any stateful model backend, make request lifecycle state explicit and assert invariants such as 'no decode before initialization/prefill'; do not assume scheduler semantics designed around attention are safe for recurrent or SSM-style state. | Caveat: Attention models can mask this category of ordering bug because they write token KV data before reading it, whereas Mamba decode first reads persistent state; hybrid or state-space architectures therefore require lifecycle assumptions to be revalidated.
  • Claim: To isolate a failure, prefer knobs that alter when and where it appears, rather than only increasing its raw frequency. | Evidence: For the RL log-prob spikes, increasing rollouts per prompt from 8 to 16, 32, 64, and 128 shifted the deterministic issue from every 12 steps to immediately at step 1. This pointed the team toward a scale- or offset-dependent mechanism. | Implication: During incident investigation, record how each experimental change shifts failure timing, request position, memory address range, or batch location; those changes are stronger causal clues than a simple pass/fail result.
  • Claim: The second defect was a silent uint32 state-cache index overflow, not an ordinary out-of-bounds failure. | Evidence: Mamba kernels used an unsigned 32-bit index pointer; when an offset exceeded about 4 billion numbers, it wrapped around without error. Lowering GPU memory allocated a smaller state buffer, so the index never grew far enough to trigger the overflow. Replacing uint32 with size_t, typically unsigned 64-bit on modern architectures, fixed the issue. | Implication: Audit cache offsets, token counters, slots, and allocator arithmetic for width and overflow behavior under maximum-scale workloads, particularly in custom CUDA kernels and long-lived serving processes. | Caveat: Tools such as NVIDIA Compute Sanitizer may not expose logic-level integer wrapping when the resulting wrapped access remains in a valid addressable region.

Detailed Brief

Why Mamba exposed the stale-state bug while attention could conceal it

  • Claims: The same request-ordering error has architecture-dependent consequences.; A decode-before-prefill condition is particularly dangerous for state-space models because decode consumes stored state before computing new state.
  • Evidence: In attention inference, the system writes the current tokens' key/value data before reading it, so stale cache contents may be overwritten before use.; In Mamba, decode reads the existing state first and then computes over it, allowing stale state from prior requests to influence the current generation.
  • Caveats: The presentation does not quantify whether all Mamba implementations or all scheduler configurations are exposed; it documents the behavior in AI21's vLLM/Jamba path.
  • Implications: Model-serving control planes should be architecture-aware rather than treating all autoregressive models as interchangeable attention workloads.; Cross-request state isolation deserves explicit test coverage for hybrid, recurrent, SSM, and other nonstandard architectures.

Instrumentation and debugging posture

  • Claims: Framework abstractions can obstruct debugging by reducing active requests to anonymous tensors by the time they reach the forward pass.; The speakers recommend modifying the engine rather than waiting for existing observability to expose the root cause.
  • Evidence: AI21 added request IDs to vLLM's forward-context object and propagated them down to the Mamba forward pass, where they could conditionally breakpoint on the known bad request.; They inspected CUDA prefill-kernel math and input/output tensors, ran NVIDIA Compute Sanitizer for out-of-bounds and memory errors, and separated prefill from decode execution before identifying the scheduling issue.
  • Caveats: The initial prefill/decode isolation implicated decode kernels, but that was not the final root cause; the kernel was being called in an invalid lifecycle state.
  • Implications: Observability should preserve a correlation path from external request ID through scheduler decision, cache slot, forward-pass mode, kernel launch, and returned tokens.; Avoid prematurely assigning blame to low-level kernels when a control-plane or scheduling decision can create invalid kernel inputs.

Notable Concepts & Terms

  • Jamba: AI21's hybrid model architecture combining Transformer attention layers with Mamba/state-space-model components; its stateful behavior made attention-oriented scheduling assumptions unsafe.
  • Mamba state: Persistent state used by the Mamba component during inference; stale state can contaminate a request if lifecycle ordering or cache indexing is wrong.
  • Prefill vs. decode: Prefill processes the prompt and initializes inference state, while decode generates subsequent tokens; the first bug occurred when decode ran before required prefill.
  • GRPO: A reinforcement-learning training method used during Jamba post-training, where the second issue appeared as rollout-versus-FSDP log-probability discrepancies.
  • Log-probability forensics: The practice of comparing per-token probability distributions between implementations to detect and localize divergence that may be invisible in aggregate output quality.
  • GPU memory utilization: A vLLM configuration controlling GPU memory allocation across weights, activations, and KV/cache resources; useful as a stress lever but capable of either surfacing or masking bugs.
  • uint32 overflow: An unsigned 32-bit index wraps after approximately 4 billion values; in this case it silently redirected state-cache indexing instead of generating an error.
  • FSDP step: The Fully Sharded Data Parallel training-side computation used as the comparison point for rollout log probabilities before a weight update.

Operator Notes / Why Ken Should Care

  • Add a release-gate test that runs deterministic, concurrent long-lived traffic against the production inference engine and compares token log probs against a reference implementation for a representative model set.
  • Instrument the serving stack with a request-correlated trace containing scheduler classification, prefill/decode mode, cache/state slot, number of previously computed tokens, and kernel-path selection.
  • Create invariant tests for stateful architectures: no decode before state initialization, no reuse of state across request IDs, and safe behavior under chunking, eviction, and scheduler reordering.
  • Run capacity-boundary tests that force cache offsets and counters beyond 32-bit ranges; audit CUDA and host-side indexing types rather than relying only on memory-safety tooling.
  • Treat reduced-memory configurations as bidirectional diagnostic experiments: test whether pressure reveals a bug and whether smaller allocation merely prevents reaching a problematic scale threshold.

Source/Metadata

  • Title: Two Bugs That Hid in Plain Sight: A vLLM Debugging Detective Story — Asaf Gardin & Yuval Belfer
  • Transcript words: 2870
  • Duration seconds: 1086
  • Timestamp note: No timestamps or chapters were present in the supplied transcript.
Full transcript 2893 words · 13 min read
0:00

So, put your model somewhere.

0:16

Put your agent, you're hoping for the best, you're waiting for something to crash. Everything looks good, everything's fine. And you see this. And that's the problem, right, in these type of bugs. There is no crash, there's no warning, no error, and there's high confidence. That's not a quality issue. Because this is something that you don't really know what and why. Welcome to this talk. My name is Yuval. This is Asaf. And together, we're going to take you on a journey of how we ended up fixing those type of bugs. A little bit about us. We work at AI21, which is an AI research lab. We started as a foundation model company, most famously known for Jamba,

1:10

which is a hybrid architecture between Transformers and Mamba, which is an SSM state. And while we were doing those, while we were training those models, while we were shipping those models into production and had users and we had a workload, we got into several interesting bugs. And these are the bugs that I think are the hardest to deal with. Because this is not a quality problem. It's not something you can take your research team and try to optimize or solve or make the model be better at something. This is an engineering problem. This is an issue where there is high confidence, but the output is bad.

1:54

So let's dive deep to the first case, what we call the imposter request, where just to set up the scene, what are we talking about? We are talking about how we, during the training of our Jamba model, more specifically, we did GRPO, which is a type of RL training. And this is a hybrid model. Layers of Mamba and attention. And just to make sure we're all aligned, what is the type of a request? So life of a request. So we start with prompt, tokenization. And then in the forward pass, we're doing both prefill and then decode. After that, we finish the forward pass. Detokenization to go back to text. And the thing about the crime here is that it's bad on so many levels.

2:50

But mainly on these three. This is what we call the one in 1,000 gibberish. It's not something that will happen in the first 500 or 900 requests. But it will happen in the 1,000. Which is rare enough to duplicate it easily, but it's too common to ship it. Also, it only happened in VLLM. Not in other inference frameworks. And it's something which is late on set. It's not something that will happen if you have only few requests. You need some sort of workload. So it's rare. It's late on set. And it's very engine specific. It's a very hard task. And we had to bring one of our best detectives to handle that. So I'll give it to Asaf to explain how. All right. Hey, guys.

3:47

Thank you, Val. Thanks, Yuval. So we're going to start with trying to reproduce something that was very difficult to reproduce. VLLM has a lot of flags, a lot of CLI flags, and a lot of knobs you can turn and tweak. And one of the things that helped us understand how to reproduce it, because when we tried to reproduce it the first time, just sending prompts here, prompts there, a few batches, it didn't really help us manage to get gibberish back from our model. The model responded back just fine. So what we did was we tried to make it happen in a very short amount of time. So we'd get a quick feedback loop when we try to debug it.

4:19

So what we did was we took one of the most default and most common flags that VLLM allows you to play with, which is GPU memory utilization, which basically allows you to choose how much memory, how much GPU memory you want to allocate for your weights, for your activations, and for your KV cache and so on. And we reduced it from 90% to 20%. And once we did that, and then we started running a lot of requests simultaneously, all of a sudden, request number, let's say, 854 suddenly returned gibberish. And when we did that, we sampled all of the batches with temperature zero,

4:58

so we'd be able to deterministically and constantly get the same request to return and respond with gibberish. So like Yuval said, it happened only in VLLM, and we used another way to reproduce it and to understand where the issue really came from, we used Hugging Face's Transformers as a baseline. Since Transformers is a very vanilla and plain implementation of our Mamba kernels, as opposed to VLLM, which has all the kernels and all the engine going through a lot of changes and modifications to support a lot of cool features that VLLM supports. So we used Transformers as our baseline to understand whether or not there is an issue with our inference or not,

5:44

with the model or not. So what we did was we took VLLM and we sent all of our prompts through VLLM and we generated a response, all the responses. Now in our hands we have the response along with the log probs, because in VLLM you're able to get your log probs out and inspect them. Then what we did was we took the full sequence, the prompt and the generation, and we passed it over to Hugging Face's Forward Pass, but all we did was run just the prefill, and then we sample, we took the logits out of the prefill response, we ran it through Softmax, and then we were able to compare the divergence in the distributions of our tokens.

6:23

That's a short pseudocode of how that looked like. You can see here that we take up the prompt, we run it through VLLM's generate, we get the response back along with the log probs, we pass it over to Hugging Face's Forward Pass, we only run it with prefill, we created some function called compute log probs, which runs the Softmax, then you calculate the difference between them, and then you'll be able to tell the divergence between every one of the tokens' log probs. All right, so now that we have the tools in our hand to understand where the issue could maybe come from, we started to look at different parts in VLLM's engine.

7:09

So the first thing we looked at was the CUDA prefill kernel of Mamba. We looked at it, we inspected all of the math that's being done there, and we looked at the tensor in, the tensor's out before we call the prefill and after. Everything looks just fine. Second thing we did was running NVIDIA's compute sanitizer tool to really see if we have any out-of-bound memory, any other memory bugs or issues. Looks okay to me. Then what we did was we tried to isolate between the decode kernels and the prefill kernels. Now we saw that the prefill kernels were working just fine. So we tried to not call the decode kernels because in Mamba you're able to do that.

7:56

So what we did was we moved all of our calls and all of our computations to go through the prefill kernel, and there you have it. The gibberish all of a sudden kind of vanished. So we were, okay, it's got to be the decode kernels. But you know how it is in software. You're getting excited too quickly, and then you figure out it's not what happened. So what we did was we tried to start playing with VLLM. We kind of needed to go and lift the hood up and see what we can do to maybe get a better understanding, and maybe get our hands dirty. Because VLLM didn't really give us more tools to really debug our kernel and our forward pass.

8:37

So once a tensor, once the request gets all the way to your forward pass, and before it goes into your prefill and decode kernels, they don't really have any identity. You can't really tell what prompt is currently being processed. It's all just tensors and numbers and matrices. So what we did was we added the request ID to some class called forward context that we propagated all the way down to Mamba's forward pass, just before the prefill and the decode kernels were called. And there we just managed to have a simple if condition with a request ID, the one that gave us gibberish.

9:13

And put a breakpoint there, and then we were able to infer and to really inspect all the metadata that comes along with it. And the second we did that, we saw that the request was, for the first time, when it went through the forward pass, it's actually doing decode before prefill. The scheduler decided that this request should be doing decode before prefill. And as Yuval said earlier, in the lifecycle of a prompt, a prompt should first be going through prefill and then decode. And what happens was that when in Mamba you run a request with decode first, after a lot of other requests were already computed,

9:50

the state was already overused, and we were using the data and the computations of stale requests, requests that came before it. So now we were actually running decode on previous requests, and that generated gibberish for us. So the kernels weren't doing the wrong thing. They were called at the wrong time for the wrong requests. And why did it matter only for Mamba? The reason is that in attention, when you write the tokens KV, you write the tokens KVs before you actually read it. So even if you have stale data, it's being overwritten. But for Mamba, when you first go through the decode kernels, you first read the state, and then you compute over it.

10:46

So what happens was you just use stale data when you do a decode. And the fix was relatively simple. We just needed to make sure that when a request first, when the scheduler first classifies a request, it's got to make sure that if you see the request whose tokens were never been computed, and they're zero, to mark them as prefill, so when they get to the forward pass, they'll actually just be used for prefill and not decode and not chunked. You can see that it was merged after some time. And that really leads us to, and then we thought everything was fixed, right? We thought everything was fixed, and there you have it. No more issues.

11:38

But that was almost the case, because after a little bit of time, it gets us to case number two, which surfaced another issue that we've faced in our RL, in our inference. So we ran RL, and while we were running our post trainings, and we looked at our evaluations and all of our benchmarks, we thought that we had some log prob spikes between the rollout and the FSDP step. And that was before any weight update. So same weights, same inputs, and the two log probs should be identical. Now they weren't. We saw that every 12 steps constantly, there was a log prob spike, and that was kind of weird. Now what would you do, right?

12:15

What's the first thing to do here? So we wanted to find some lever that changes how things fail and not just how much they fail. We want to tweak some knobs that don't just tell us, hey, this error happens this many times, and so on. We wanted to tweak some knobs that kind of tell us that once we tweak that knob, we understand how it's wired to anything in VLLM, and the engine. And so we'll be able to specifically go and debug that specific part. So what we did was we decided to increase the amount of rollouts per prompt. Since we saw in our default RL engine we have eight rollouts per prompt, and we saw that it happened deterministically every 12 steps,

13:13

we decided, okay, let's try to tweak it up a bit and increase the amount of rollouts per prompt. So we started doubling it from eight to 16 to 32, and 64, and 128. And you can see here that there's a pattern here. The more we increased it, the closer it happened, because what we wanted to achieve here, we wanted to try to reproduce the issue as fast as possible so we'd have a faster debug loop, feedback loop. So when we ran it on 128 rollouts per prompt, it happened immediately on step one, and we didn't have to wait for step 12 and step 24 and so on. Now, you might think, okay, so you guys played with the GPU memory utilization before. You tweaked it, you decreased it.

14:03

It looks like, when you test on pressure, it really surfaced things up. So we thought that as well. And when we reduced the GPU memory from 0.9 to 0.2, it actually caused the issue to go away. So we actually pulled the wrong lever here. And the reason is because we noticed that Mamba kernels used unsigned int 32 index pointer. So once the offset went past 4 billion numbers, it wrapped around instead of throwing an error. So when we shrank the GPU memory, VLLM allocated a small state buffer, and the cache index never got large to hit that slot. So we were just not reaching far enough for the buffer to trigger an overflow. So again, the fix was rather simple.

14:55

All we needed to do was just change one word, one data type variable, from UIN 32 to size T, which basically means for most modern architectures, size T would mean unsigned 64 bit, and that's a very large number. We never reached that number, and that overflow now never happened. So what we can see here is that we had two scenes and one criminal. Both kind of had similar symptoms. Both had silent gibberish and silent log prob spikes, which also kind of sometimes generated gibberish. They were both around the Mamba state cache. They were both surfaced by memory pressure, whether it was for worse or for the best, and both found via log probs forensics.

15:48

Stateful inference systems don't fail loudly. They lie to you confidently. I mean, obviously, sometimes you get crashes, you get out of bounds errors, you get other exceptions and so on.

16:09

But sometimes there are some errors that don't surface up, and you don't get a trace log, you don't get anything. You have to go and dig and understand why things happen. So if there are some takeaways to take from this presentation is build the log probs comparison script. If you need to compare your quality, you need to compare it to understand whether you modeled the issues or not, log probs comparison script with a baseline of some other inference framework that you have or built is always great. Reproducing under pressure, constrain memory, crank the scale up, play with other knobs that the inference framework gives you,

16:36

and really try to understand where the issue comes from. Look for what moves the failure shape, the timing, the space, and the location. And when things don't really have identity, thread identity through, and what also I want you to take from this, don't be afraid to go dig in the code, get your hands dirty. Sometimes, model languages, LLMs are, they might tell you how things work, but without you seeing it in your own eyes, getting your hands dirty, you won't get full understanding of what's going on. Thank you. You guys can add us on LinkedIn. Scan the QR code to read the actual blog that we've published with this finding. Yeah, that's it.

17:32

you won't get full understanding of what's going on. Thank you. You guys can add us on LinkedIn. Scan the QR code to read the actual blog that we've published with this finding. Yeah, that's it.

18:03

I'll see you now, I'll see you now, I'll see you now, I'll see you now, I'll see you now, I'll see you now, I'll see you now, I'll see you now, I'll see you now, I'll see you now, I'll see you now, I'll see you now, I'll see you now, I'll see you now, I'll see you now, I'll see you now, I'll see you now, I'll see you now, I'll see you now, I'll see you now, I'll see you now, I'll see you now, I'll see you now, I'll see you now, I'll see you now, I'll see you now, I'll see you now, I'll see you now, I'll see you now, I'll see you now, I'll see you now, I'll see you now, I'll see you now, I'll see you now, I'll see you now, I'll see you now, I'll see you now Thank you.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note