Are LLM Performance Benchmarks Reliable? — Ashok Chandrasekar & Jason Kramberger, Google
Description
Ask a benchmark harness for 200 queries a second and it may quietly deliver 38, then print results as though it ran 200. Ashok Chandrasekar opens with that experiment, which is why he and Jason Kramberger, both at Google, kept failing to reproduce published numbers. Python's global interpreter lock makes a single process harness CPU bound; the ones they tested capped near 170 and never said so. A thrashing client also inflates the latency it measures, once by 58 seconds, which reads as a bottlenecked server when the server was fine. A shared result claiming 20 percent better throughput turned out to have temperature set to zero, deterministic and faster than any real workload at 0.7. The same public dataset fed to two harnesses produced different input tokens, sampled and truncated differently. The diagnosis is usually your server. Often it is your harness. Kramberger presents the fix they built: Inference Perf, a CNCF project out of the Kubernetes serving working group. A main process schedules requests against a plan, Poisson, constant rate, or fixed concurrency, and fans them across worker processes that report when they actually fired versus when they were meant to. Client side telemetry sits beside server metrics, so you can tell a failing harness from a failing system under test. At 5,000 queries a second it kept up and said so. Configuration is declarative enough to replay multi turn conversations with length distributions, and a published workload catalog defines agentic generation, tree of thought and batch summarization in terms other tools can adopt. He closes with Prism, their UI under the llm-d project, showing combined optimizations against a plain Kubernetes service across eight replicas on TPUs. Speaker info: - https://www.linkedin.com/in/ashokchandrasekar/ - https://ashokc.dev - https://www.linkedin.com/in/jkramberger Timestamps: 0:00 - Two Google engineers who could not reproduce other people's numbers 2:47 - What a production scale benchmark ha
Summary
Generated by gpt-5.6-terraAt-a-Glance
- Verdict: Watch fully
- Core thesis: LLM inference benchmarks are unreliable unless the benchmark client can sustain and verify the intended load, workloads reproduce real production behavior, and datasets/configurations are explicitly controlled.
- Why it matters: Benchmark-harness failures can be mistaken for inference-server bottlenecks, leading teams to make incorrect capacity, optimization, model-serving, and cost decisions.
- Best use: Use this as a practical validation framework for evaluating serving-performance claims and for designing trustworthy load tests for production LLM systems.
Executive Summary
Google engineers Ashok Chandrasekar and Jason Kramberger argue that most LLM performance numbers should not be accepted at face value, particularly when they are produced by simple developer-oriented scripts or generic HTTP load tools. The central issue is that a benchmark is itself a distributed system component: if its client cannot generate the requested traffic, process streaming responses, or accurately measure timing, its results may describe the harness rather than the inference stack.
They distinguish production-scale inference benchmarking from model-server microbenchmarks, accelerator comparisons, and conventional web load testing. Production tests must account for multiple replicas, autoscaling, online and batch traffic, prefill/decode disaggregation, realistic request distributions, SLOs such as P90 time-to-first-token, and the saturation point at which throughput and latency trade off.
The speakers give concrete examples of benchmark distortion. A Python single-process harness requested to send 200 QPS produced only 38 QPS on a small shared-core machine and topped out near 170 QPS even on a more capable machine. In another test, client-side contention inflated apparent latency by as much as 58 seconds against a simulated server that should have had essentially no latency. They also found a claimed 20% throughput improvement that was attributable to setting model temperature to zero, rather than a representative production configuration.
Their proposed remedy, the CNCF-associated inference-perf project, uses declarative workload configurations, multiprocess load generation, scheduled request execution, and client-side observability alongside server metrics. It also publishes a workload catalog for patterns such as multi-turn generation, tree-of-thought agentic workflows, code generation, and batch summarization, aiming to make comparisons reproducible across tools and representative of actual demand.
Key Takeaways
- Claim: A benchmark client that cannot maintain its requested load invalidates conclusions about server throughput and latency. | Evidence: When a separate benchmark harness was configured for 200 QPS, it delivered only 38 QPS on a small shared-core machine; on a more powerful machine, some single-process harnesses still capped around 170 QPS while reporting a completed run. | Implication: Before interpreting an inference benchmark, verify achieved request rate, queueing delay, and client resource saturation; otherwise, do not use the result for capacity planning or vendor/model-server comparisons. | Caveat: The examples demonstrate a failure mode, not that every Python-based or single-process harness is necessarily inaccurate at lower loads.
- Claim: Client-side contention can create false latency regressions that teams incorrectly blame on the inference service. | Evidence: The speakers observed up to 58 seconds of delay caused by a harness thrashing while collecting streaming-token responses; in a 1,000-QPS test using a scalable harness against a simulated server, latency was minimal as expected. | Implication: Latency instrumentation must separate planned send time, actual client send time, response timing, and server-side timing, particularly for streaming generation workloads.
- Claim: Production LLM benchmarking requires workload sweeps and SLO-aware operating-point selection, not a single QPS measurement. | Evidence: The presenters describe sweeping multiple loads to identify saturation and measuring input/output-token throughput, time to first token, time per output token, and SLO conformance such as P90 TTFT. | Implication: Ken should evaluate serving systems through throughput-versus-latency curves at explicit percentile SLOs, then select a safe operating point rather than optimize for peak tokens per second alone.
- Claim: Seemingly minor generation and dataset settings can materially distort benchmark results. | Evidence: A reported 20% throughput improvement was traced to a harness setting temperature to zero instead of a more realistic value around 0.7. Two harnesses using the same shared-GPT dataset produced different input-token counts because they sampled and truncated prompts differently. | Implication: Treat workload definitions as versioned test artifacts: pin sampling, truncation, input/output-length distributions, stopping behavior, temperature, cache assumptions, and conversation-replay logic before comparing results. | Caveat: The talk does not establish one universally correct temperature or dataset; the required settings depend on the intended production workload.
- Claim: A production-grade LLM benchmark must model the full serving environment and real request behavior. | Evidence: The talk cites online and batch traffic, many inference servers, prefill/decode disaggregation, autoscaling, prefix-cache behavior, forced end-of-sequence generation, and multi-turn replay as factors omitted by simple benchmark scripts. | Implication: For agentic or multi-turn systems, test the actual request mix and stateful interaction pattern rather than relying on isolated fixed-length prompts.
- Claim: Inference-perf is designed to make benchmarks reproducible and auditable through declarative configurations and observable multiprocess load generation. | Evidence: Its main process schedules requests by planned execution time, distributes them to multiple worker processes, tracks actual versus planned execution, and reports client metrics together with server metrics. The presenters state it kept up at 5,000 QPS in their comparison. | Implication: Inference-perf is worth evaluating as a benchmark control plane when testing production serving stacks, especially where existing scripts cannot prove load fidelity. | Caveat: The 5,000-QPS result is a project-presented comparison and should be independently reproduced for Ken's own endpoint, protocol, hardware, and workload profile.
- Claim: Standardized workload definitions can make inference-performance results more comparable across tools and infrastructure. | Evidence: The workload catalog includes natural-language scenario descriptions plus detailed generic and inference-perf-specific configurations for multi-turn generation, tree-of-thought agentic generation, batch summarization, and related workloads. Prism displays results such as an eight-replica TPU agentic-code-generation test where combined optimizations outperformed a simple Kubernetes-service baseline. | Implication: Use public workloads as regression and external-comparison baselines, while maintaining a private, trace-derived workload suite for deployment decisions. | Caveat: Standard catalogs improve comparability but cannot substitute for a workload calibrated against an application's own production traces and SLOs.
Detailed Brief
Benchmark ecosystem and the gap the speakers target
- Claims: The speakers divide current tooling into four categories: model-server-native scripts, competitive accelerator-analysis benchmarks, conventional HTTP-scale tools, and production-scale LLM benchmarks.; The first three categories serve legitimate but narrower purposes and do not by themselves establish that an end-to-end inference deployment will meet production demand.
- Evidence: Model-server examples named are vLLM and SGLang; competitive-analysis examples include MLPerf, SemiAnalysis, and Artificial Analysis; HTTP-scale examples include Locust and Grafana k6.; The production reference architecture discussed includes an LLMD inference pool with multiple servers and configurations such as autoscaling and prefill/decode disaggregation.
- Caveats: The presentation is primarily a case for inference-perf and LLMD, so it presents the surrounding tooling through the lens of the gap their project is intended to fill rather than as a neutral feature-by-feature evaluation.
- Implications: Classify a benchmark before relying on it: a model-server script can support local optimization, while an infrastructure or deployment decision requires an end-to-end production-scale test.
Notable Concepts & Terms
- Inference-perf: A CNCF-associated open-source LLM inference benchmarking tool built around declarative workload configuration, multiprocess load generation, and client/server metric visibility.
- Metric fidelity: The requirement that benchmark measurements accurately distinguish behavior of the system under test from limits or delays introduced by the benchmark client.
- Planned versus actual execution time: A load-generator observability mechanism that reveals whether requests were actually sent on the configured schedule rather than merely counted as part of a run.
- TTFT / P90 TTFT: Time to first token and its 90th-percentile value; presented as a key streaming-user-experience SLO rather than a secondary latency statistic.
- Prefill/decode disaggregation: A serving architecture that separates prompt processing from token generation, increasing the need for benchmarks that represent the real distributed stack.
- Workload catalog: Published, reusable workload definitions intended to standardize scenarios and configuration details across benchmarking tools and teams.
- LLMD Prism: A UI in the LLMD project for sharing workload definitions and benchmark results, including comparisons of serving-stack configurations at production scale.
- Stochastic generation settings: Parameters such as temperature that change output behavior and can alter measured throughput, so they must reflect actual application demand.
Operator Notes / Why Ken Should Care
- Require every performance report to include target versus achieved QPS/concurrency, client CPU and process configuration, request scheduling lag, error rate, and separate client/server latency measurements.
- Create a versioned benchmark manifest for each critical serving workload that pins model settings, tokenizer, prompt selection/truncation rules, output/stopping rules, cache assumptions, request distributions, and replay behavior.
- Run a client-only calibration against a low-latency simulated or echo endpoint before using a harness to diagnose inference-server saturation.
- Adopt explicit SLO-based load sweeps for agentic and streaming workloads; record the highest sustainable operating point that meets percentile TTFT and token-generation latency requirements.
- Evaluate inference-perf and its workload catalog as a candidate standardized harness, but validate it against representative internal traces before making it the source of truth.
- Reject vendor or internal throughput comparisons that omit decoding parameters, achieved load, dataset construction details, replica count, or the precise serving architecture.
Source/Metadata
- Title: Are LLM Performance Benchmarks Reliable? — Ashok Chandrasekar & Jason Kramberger, Google
- Transcript words: 2286
- Duration seconds: 967
- Timestamp note: No timestamps or chapters were present in the supplied transcript.
Transcript
Hi, everyone. Welcome to our talk on RLLM performance benchmarks reliable. A little bit about us. I'm Ashok Chandasekar. I'm a Staff Software Engineer at Google. I work on inference performance evaluation and optimization. I lead a couple of open source projects. One is called inference perf, which is a benchmarking tool to do reliable performance benchmarks. I'm also the lead for LLMD SIG benchmarking. LLMD is a distributed inference framework that makes production scale inference possible. Hi, everyone. I'm Jason Kromberger. I'm a software engineer at Google. I'm also a co-maintainer of inference perf and a few of the sub projects that Ashok brought up. And I work on inference performance and benchmarking. Okay, let's get started. Let's look a little bit about how the benchmark ecosystem looks. You have your model server frameworks. These are VLLM, SGLang, and other model servers. And all of these have some benchmark capability within them. These are primarily Python scripts and developer-focused benchmarks to see how you can measure the performance of your model server itself. And then you have your competitive analysis tools. These are MLPerf, Semi-Analysis, Artificial Analysis, and so on. They mainly aim for competitive performance benchmarks to compare chip and accelerator performance. And then you have your typical web benchmarks. These are like Locus, Grafana K6, and so on. These mainly focus on high-scale HTTP benchmarks. And then you have your last segment, which is the production-scale LLM benchmarks. These are to actually benchmark your production inference serving stack. And that is our focus today. We'll be focusing mainly on inference perf and how we solve this production-scale benchmark problem. So if you have run a benchmark before, it typically looks like this. You have some sort of benchmark harness, and then you specify what model you are benchmarking, the number of prompts you want to run, what is the input-output sequence length, and the request rate or the load you want to send. And your output looks something like what is on the right. This is basically your input-token throughput, output-token throughput, some latency metrics, time-to-first token, time-per-output token, and so on. So what are some issues with a simple benchmark like this? So if you want to actually benchmark production-scale workloads, here I have LLMD inference stack as an example. You can have online serving, you can have batch workloads, and if you see the inference pool below, there are a lot of servers that are running. And then you have complex configurations like pre-fill decode disaggregation and workload auto-scaling and other things that are going on under the hood, and usually the scale is much larger. So your normal benchmark harnesses runs into issues when you try to benchmark a setup like this. And if you look at the key characteristics of what we want out of a production-scale benchmark, we need to be able to do high load, which is limited in a lot of tools out there. We need to be able to simulate real-world workloads. What use is it if it is just some synthetic workload that is not accurately representing what your customers are going to run. And then metrics fidelity is very important. Are the metrics accurate and how well they work? This is the set of metrics that LLMD measures by default. I just pulled it from the website there. As you can see, it's not just a single QPS that you are running. You are sweeping a list of various loads, and you try to measure what the baseline is and what optimizations you are making and what the difference there is. You need to find the right point where the server gets saturated, so you know the right optimal point to run your servers on to maximize performance and to save costs. And things like SLOs become more important. What is your time to first token P90 SLO, and are you conforming to that SLO? So when you run a normal benchmark like we saw before, what are some of the pitfalls that you run into? We have been running benchmarks for a couple of years, so we run into all sorts of different results that people share, and a lot of times we are not able to reproduce the results that are shared by other people. So that is what motivated this talk. These four common things that we see as an issue. One is accurate metrics, and two, observability into your benchmark tool itself. Do you know if your benchmark harness is actually failing? Is it not able to maintain the load? And three, reproducibility. There is some inherent randomness in the datasets that you use, so how do you make sure it is reproducible? And four, the dataset quality itself. So this is an experiment we ran. We asked a different benchmark harness to generate 200 QPS, and this was the result. So a couple of things I want to point out. Python has this global interpreter lock, GIL. If you have been working with Python, you know that, which makes everything single-threaded. So even when you have a multi-CPU machine, a lot of times you are limited by the performance of a single CPU, when you are CPU-bound especially. So this shows the single-process benchmark harness and how the QPS you are able to achieve differs based on it. When you run with a really small shared core machine, you can see that even when you request 200 QPS, you are only getting 38 QPS. And then you give it a bigger machine, and then some of these single-process harness, they cap out at 170 QPS. This is a much more powerful machine. But it is a problem because you ask for 200 QPS, and then you don't know whether it actually delivered it. It will just say, I ran it. These are the numbers. So you think, okay, you ran 200 QPS, but in fact, you have not. The other issue that comes out of it is the latency inflation. If your server is saying, okay, this is how much QPS I was able to run, and this was the accurate numbers, that is one thing. But if your benchmark harness is actually inflating latency, because it's thrashing, trying to collect all the streaming token requests. In one of the tests, we noticed the delay was up to 58 seconds. So you might look at this and go, oh, my server is bottled up. It's not able to handle all the requests. But in fact, it's actually your benchmark client that is inflating the latency. We ran a thousand QPS test. When your benchmark harness is actually able to scale out, you can see there is very minimal latency. This is a simulated server, so there shouldn't be any latency at all. And there are other variables that go into it. In one of the benchmarks, someone shared and they said, hey, we are getting 20% better throughput. Then we looked into it and we found out the benchmark harness was setting the model temperature to zero. Which means your model outputs are a lot more deterministic, and it was able to churn out a higher throughput than what you would normally see in a real workload, where your model temperature is somewhere around 0.7. Another thing is we used a shared GPT data set, the same data set across two different benchmark harness, and they produce different input tokens. This is because they sample them differently. They truncate them differently. So as a user, you don't have insight into this. You run it, you trust the numbers it produces, but they are wildly different. And there is much more. Do you actually force it to generate till the end of sequence? Are you looking at prefix cache rates? How do you do multi-turn replay via benchmarks? And how do you actually get high fidelity on the actual workload that would resemble your production workload? So the main thing I wanted to convey here is a lot of times you diagnose it as your server or inference stack problem, but in a lot of cases it could be your benchmark harness. So what is the solution to this? How do we actually do reproducible benchmarks? Jason here will take care of that. Cool. Thanks, Ashok. So yeah, how do you solve these problems? We pulled together inference perf. It's a CNCF project spanned out of Kubernetes working group serving to provide a standardized place for us to work with the community and solve some of these issues together. It enables the ability to have a user defined declarative configuration that allows you to have clear reproducibility across runs. We also added a load generator that solves the GIL problem in Python across multiple processes and reports those client metrics back along with server metrics to ensure that you have the highest metric fidelity and you're able to actually observe when your tool is having an issue versus your system under test. So first going over the load generator, you see that the main process actually queues requests based off of the planned time that they need to execute, which is based off of your configuration. This may be in some poison process or constant rate or maintaining a constant number of concurrent requests. This request queue channel is then spread across multiple processes, which pull and ensure that they execute with minimum overhead, but then also observability about when they execute versus their planned time. And you can see this working at scale. So this is a comparison across other tools, some being the HTTP scale tools like K6. But you see that even at 5,000 QPS, inference perf was able to keep up due to this architecture. And most importantly, actually report that it was able to keep up. Other portion is configuration. So earlier, I showed a brief example of how you might simply run a benchmarking tool. And here on the left, you can see a simple example running a random data set against an endpoint. But the actual configuration that we have in front of inference perf is very detailed with a lot of knobs that allow you to accurately test your configuration off of your workloads. You can see here on the right that this is a configuration for a conversation replay, where you're able to configure not only the input-output length, but their distributions, et cetera. And further beyond just the ability to configure a single run, we've actually worked together to have a published set of some of these workloads and configuration of inference perf that are tied to state-of-the-art inference workloads. For example, here in the workload catalog that we've put out, you're able to access standard multi-turn generation, tree-of-the-art generation, tree-of-thought agentic generation, as well as batch summarization and others. Each one of these has a simple definition in natural language that allows you to understand what the scenario is. But beyond that, there's also pretty detailed configuration metrics, not only in inference perf's configuration, but in generic terms so that this can actually be shared across tools and have a place for standardization of these workloads. So, the culmination of these things leads us to actual results that we can clearly display. And here's a screenshot from Prism, which is a UI we have for sharing not only those workloads that I showed before, but also some benchmarking results. Here's a screenshot from Prism, which is a part of the LLMD project. Here you can see a benchmark result for the agentic code generation workload that we showed before on TPUs. These three lines you see here show you the difference between combined optimizations, as the green line, and a baseline that is just a simple Kubernetes service instead of in front of multiple model server replicas. It's important to note here is that this is at production scale with eight replicas. And you can see that the combined optimizations were measured to be much higher than the baseline, scaling into almost hundreds of thousands of tokens per second. So, the takeaways. The principles for benchmarking validity based off of the pitfalls that Ashok brought up earlier. At production scale, you need client concurrency, and you need observability into your client's behavior and its ability to meet your configuration. The metric fidelity allows you to actually observe your client's behavior as well as your system under test and understand that your scenario was accurately executed and your performance results were valid. The stochastic variables and non-determinism or determinism that you set amongst your run needs to reflect your real-world demands for your workload. And most importantly, your datasets do truly matter. Your workloads need to be as close to what you are intending to test as possible. And we have examples in the workload catalog. So, here we have three links to some of the things that we've presented on here before. Inference Perf is our benchmarking tool, and there's the Git repo for it. LLMD is a project that we work under, and inference perf under it in LLMD benchmark. LLMD is for production scale inference. And then LLMD prism, which was the UI we showed for the benchmarking results. And the workload catalog that defines some of these workloads. So, that will answer questions after, but we appreciate your time. Thank you. Thank you. Thank you. Thank you.