Open Reader

20 days of compute vs 7 hours: rethinking what state-of-the-art means — Bertrand Charpentier, Pruna

completed 19:35 Watch on YouTube

Current Status

completed

Video ID

hqHC6Z_lXyo

RAG / Chat

Enabled
20 days of compute vs 7 hours: rethinking what state-of-the-art means — Bertrand Charpentier, Pruna
Description

Ranking image generation models the way Design Arena does it — 26,000 battles, 62 seconds per generation — takes 20 days of compute, costs $5,000, and consumes roughly 400 marathons worth of energy. Bertrand Charpentier, cofounder and chief scientist at Pruna AI, uses that number to make a point: the same evaluation on a fast compressed model takes 7 hours and $265. Efficiency is a dimension of state of the art, not a footnote. The rest of the talk dismantles the idea that any single model holds the title. Leaderboard rankings disagree with each other — the same model goes from rank 10 on one to rank 5 on another. Most models lose 40% of their head-to-head battles, which means the top-ranked model is the wrong choice for nearly half of real use cases. His answer is the Pareto front: plot quality against latency or cost, find the frontier, and expect three or four models clustered tightly in quality score but varying up to 20x in efficiency. Evaluating this way tends to surface small specialized models rather than large foundation models. Speaker info: - https://www.linkedin.com/in/bertrand-charpentier-76995ab6/ - https://github.com/sharpenb

Summary

Generated by claude-sonnet-4-5

30-second take

Bertrand Charpentier from Pruna argues that "state-of-the-art" AI model selection is broken because people either naively trust inconsistent public leaderboards or run biased internal benchmarks—both of which ignore efficiency and lead to overuse of large foundation models. His core claim: there's no single SOTA model; instead, there's a Pareto frontier of quality vs. efficiency (latency, cost, energy). For image generation, he shows Flux takes 20 days of compute and $5k for 26k eval samples (62 sec/image), while Pruna's optimized model does the same in 7 hours for $265 (<1 sec/image). The talk is a sales pitch for Pruna's compression-as-a-service but makes a legitimate point: benchmarking must account for task-specific quality metrics and compute trade-offs, not just leaderboard rank.

Key takes

  • Public leaderboards are inconsistent and misleading. The same image models rank differently across LM Arena, Design Arena, and Artificial Analysis. A model can be #10 on one leaderboard and #5 on another. Elo score ranges vary wildly (1100–1300 vs. other scales), making cross-leaderboard comparison useless. Implication: Trusting a single leaderboard for model selection is statistically naive.
  • Task-specific rankings diverge from aggregate scores. ChatGPT Image ranks #1 overall on Design Arena but never tops individual task leaderboards (object removal, background editing, text rendering). Different models dominate different subtasks due to training specialization. Implication: If you have a narrow use case, aggregate leaderboards will steer you toward the wrong model.
  • Leaderboard sample sizes are too small for production confidence. Artificial Analysis uses "a few thousand samples" per model; many production APIs see millions of inferences/day. Top models still lose ~40% of head-to-head battles, meaning 40% of use cases might favor a "lower-ranked" model. Implication: Leaderboard rankings are not statistically significant for your specific workload.
  • Manual inspection is doubly biased. Charpentier runs a live poll showing audience preferences split across three generated images (prompt: "little guy and a parrot"). Some people change preferences when shown a second set. Implication: Human eval at small scale reflects individual taste and sample noise, not generalizable quality.
  • Standard automated metrics (e.g., CLIP score) often show trivial differences. Ranking 8 image models by CLIP score on different datasets yields inconsistent rankings and tiny variations (all scores between 0–1 or 0–100 with minimal spread). Task-specific metrics (e.g., text rendering accuracy) show clearer gaps and consistent rankings. Implication: Generic quality metrics can be useless; you need metrics aligned to your use case.
  • Compute cost of SOTA is enormous and rarely discussed. Flux image generation takes 62 sec/image. Generating 26k eval images = 20 days of compute, $5k, 556 kWh energy (equivalent to running 400 marathons). Pruna's optimized model does the same in 7 hours, $265, ~4 marathons of energy (<1 sec/image). Implication: Quality-only comparisons ignore massive efficiency disparities; a slightly lower-quality model can be 20x faster and 20x cheaper.
  • There is no single SOTA—there's a Pareto frontier. Plot quality (y-axis) vs. latency or cost (x-axis). Multiple models sit on the efficient frontier (no model dominates them on both axes). Choosing SOTA means picking a point on that curve based on your quality/efficiency trade-off. Implication: "Best model" is context-dependent; you should optimize for your budget and latency requirements, not just leaderboard rank.

Useful details

  • Compression techniques Pruna uses: Quantization (different schemes per module), pruning (removing unimportant components), and step/denoising reduction (distillation or caching to drop from 50 steps to 20 or even 4 for image/video generation).
  • Example Pareto analysis: When focused on text rendering quality (not general capability), Pruna's optimized Flux models still sit on the Pareto front but are much faster than standard Flux. This shows task-specific tuning beats generic SOTA for real use cases.
  • Pruna's positioning: They offer compression-as-a-service (fastest image/video models, 1–5 sec generation), open-source compression tooling (quantization, pruning, caching algorithms), and educational content (research papers, efficiency courses).
  • Specific energy comparison: Charpentier checked his Strava—one marathon consumes ~1.4 kWh. Flux eval uses 556 kWh = 400 marathons. Pruna model uses ~5.6 kWh = 4 marathons.
  • Leaderboard battle methodology: Design Arena/LM Arena run pairwise comparisons (users pick preferred output). Win rates for top models are ~60%, meaning they lose 40% of the time. This variance implies leaderboard rank is not a strong signal for your specific prompt distribution.

Caveats / counterpoints

  • Pruna conflict of interest. This is a sales pitch for Pruna's compression services. The latency/cost numbers for Pruna models are not independently verified, and the quality trade-offs of their optimizations are not rigorously compared to the original models in this talk.
  • No discussion of quality degradation from compression. Charpentier claims Pruna's models are on the Pareto front, but doesn't show how much quality drops when you go from 50 denoising steps to 4, or what quantization does to edge cases. The talk implies "small quality loss, huge efficiency gain" but provides no nuanced data.
  • Leaderboard critique is valid but doesn't offer a better standard. He correctly identifies that public leaderboards are noisy and task-generic, but his solution is "use Pareto plots and task-specific metrics"—which requires you to already know your task distribution and have eval infrastructure. Not actionable for teams without ML ops resources.
  • Sample size argument is overstated. He says "thousands of samples" is too small compared to "millions of production inferences," but statistical significance for model ranking doesn't require matching production volume—it requires enough samples to differentiate models at your desired confidence level. He doesn't show leaderboard confidence intervals.
  • Live poll is anecdotal. The audience preference experiment is fun but not rigorous (small N, no control for prompt familiarity, no blinding). It illustrates the bias point but doesn't quantify it.
  • No mention of safety, alignment, or API reliability. The talk is purely about quality-efficiency trade-offs. For production, you might care about content moderation, jailbreak resistance, uptime SLAs—none of which are addressed.

Ken relevance

  • Directly applicable to agent system model selection. If Ken is building agents that call vision models (e.g., Artisan's LinkedIn profile analysis, image understanding for workflows), this talk argues for evaluating models on Ken's specific task (not ImageNet or generic benchmarks) and plotting quality vs. latency/cost. Ken should not default to GPT-4V or Claude 3 Opus if a smaller, faster model hits 95% quality at 10x lower latency.
  • Compute cost optimization for AI ops. Ken's AI ops work likely involves inference at scale. The 20x latency difference (62 sec → <1 sec) and 20x cost difference ($5k → $265) for the same eval workload is a forcing function to measure efficiency in any model eval pipeline. Ken should track $/1k inferences and latency p95 alongside accuracy.
  • Investing angle: efficiency moat vs. quality moat. Pruna's pitch is that compression/efficiency is undervalued relative to leaderboard quality. If Ken is evaluating AI infra startups, this suggests winners may not be the ones with the best leaderboard numbers but the ones with the best quality-per-dollar or quality-per-second. Look for companies with Pareto-optimal offerings, not just SOTA claims.
  • Benchmarking bias is a product risk. If Ken is building products that rely on model selection (e.g., choosing a vision model for Artisan), relying on leaderboards or manual inspection will lead to suboptimal choices. He should invest in task-specific eval harnesses with enough samples to matter (thousands, not dozens).
  • Pruna's open-source tools could be useful. If Ken is fine-tuning or deploying models, Pruna's quantization/caching/pruning package might be worth testing. The talk doesn't detail it, but the efficiency techniques (per-module quantization, step reduction) are legit and could apply to Ken's workflows.
  • Low relevance for content/GTM. This is a technical ops talk, not a business strategy or content play. Unless Ken is writing about AI efficiency trends, this won't inform content directly.

Watch verdict

Skim. The core thesis—"SOTA is a Pareto frontier of quality vs. efficiency, not a leaderboard rank"—is valuable and under-discussed in AI discourse. The specific numbers (20 days → 7 hours, $5k → $265) are memorable. But the talk is repetitive (same examples shown 2–3 times in transcript), the Pruna pitch is heavy, and the depth on compression techniques is shallow. Ken can extract the key idea (measure efficiency, not just quality; use task-specific metrics) in 5 minutes without sitting through the full 20+ min presentation. If Ken is actively selecting vision models or building eval pipelines, the Pareto plot framing is worth internalizing. Otherwise, the summary is sufficient.

Transcript

3252 words en Processed in 423.3s

So today we're going to try to ask the question what makes a model state-of-the-art. I guess this is a super important question for everyone because, of course, for our applications in research or when we deploy a product, we want to always have the best performance out of our models. But the problem is that state-of-the-art is a confusing concept and people maybe have different visions on this. So we'll just try to see a bit first what, how people approach this question and how they try to answer it. And usually there are two main methods that people try to use to know if my model is state-of-the-art. The first one is simply they go on the internet and check some publicly available leaderboards to see what is the best model on the public leaderboard. And another method is actually just perform some internal evaluation and again see based on their internal evaluation what model is the best. The problem with these methods is that in most cases if you apply them naively, you will always find a lazy solution which is just to use a large foundation model. So we're going to just try to see with these methods what people tend to do and whether it can be done better. So the first method again is just simply checking a publicly available leaderboard. So for example, if you take the use case of image editing, let's try to find the best image editing model in this case. So usually the first step is just find some leaderboard. In this case, you can use Design Arena, which is a very famous one, and then you just pick the top one, which is ChatGPT Image. And then you feel that you are happy. This is the best model for your use case. In general, it's a good solution. You get a reasonable model at low effort, but the problem is that you don't know exactly a lot of things about how the users will interact with your model and so on, so you can actually make a much better choice. So the first problem is that if you look at many leaderboards, not a single one, you will see that each public leaderboard will have a different ranking. So here maybe it's a bit small, but you can trust me. There are three leaderboards—LM Arena, now called Arena, Design Arena, and Artificial Analysis—and they try to rank image editing models. And if you try to draw the difference between the displayed numbers, you will see that it's not the same ranking. The top model is not the same. Also, relatively, if you compare models between each other, they will be different. For example, there is one model—Union—that goes from rank 10 on one arena to rank 5 on another arena. So it's a bit noisy. What is the information you want to get out of it? There are even some models that appear in the leaderboards and are not in some others. So it's hard to get the main information. And when you check actual details like the Elo scores, which is supposed to be the quality score you use to know what is the best model, we'll see that even these Elo scores are very different. Meaning that for some leaderboards, it will be between 1100 to 1300, but for some it will be a completely different range. So relatively, we don't know how strong the models are between each other. And usually the main solution for this is not to trust a single one, but you need to look at multiple ones. And when you see that there is a lot of difference between different rankings, it means that probably there are some models which are approximately equivalent. It's not because ChatGPT Image is ranked top one on one leaderboard that it means that it's the best overall. Another problem is that in most cases you will have a specific application. What we've seen before is an aggregated score over a lot of different tasks. For example, we can see removing objects, changing backgrounds, editing text. But we can actually build leaderboards for each of these specific use cases. And these are also some leaderboards from Design Arena. And I think there is a problem here. It comes back. And you can see that actually if you draw the difference for each specific use case, you will see that again the rankings are completely different. And ChatGPT Image, if it is ranked top one in the overall ranking, is never at the top in this specific ranking. There are always some new models and some models which are super good at removing objects or models which are super good at doing some other things. There is no model consistently outperforming the others. There are very different models working well for different target use cases. And this is normal because this is just due to the fact that some models have been trained more on some specific tasks than others. And the solution for this is when you check public leaderboards, you should always try to target what your use case will do in the end. If you focus on removing objects, look at this leaderboard and not the others. Another problem is that usually in the leaderboards they are not really statistically significant for your specific use case. So here I try to show two different things. The first thing is on how many samples these leaderboards are built. And if you check on the left, for example, Artificial Analysis, these rankings are built on a few thousand samples for each of them. So it's not much if you compare it to the load of inference you have for many applications. It's probably super low. For some of our models that we have, we have millions of inferences per day. So probably you'll get more information by just evaluating the model on our API rather than just looking at this leaderboard. Another thing is Elo scores. Usually you can also compute what is the win rate of each model. So when you build these rankings, what you do is you actually make models battle against each other and ask people, "Okay, what is the best model between the two?" to a lot of users. And what you can see is actually the win rate. Usually, there are no models which are close to 100 percent win rates. It means that most of the models they lose at least 40 percent of their battles. And if your use case is in this 40 percent of the battles, it means that if you take the best models, you will just take the wrong model. So again, here it's important we need to evaluate on more samples and always have evaluation which is close to the final setup and final use case conditions. Now we can also take the second solution to try to know what is the state-of-the-art for AI models. And the second solution is to do just internal benchmarking. One way to do it, which I see the most actually in the image and video generation, is with certain people who just do manual aspect inspection. And if your use case is in this 40 percent of the battles, it means that you will just take the wrong model if you take the best models. So again, here it's important we need to evaluate on more samples and always evaluation which is close to the final setup, the final use case conditions. Now we can take also the second solution to try to know what is the state of the art for an AI model. The second solution is to do internal benchmark. One way to do it is what I see the most actually in the image and video generation. With certain people, they just do manual aspect inspection. They try a couple of prompts, a couple of models, and they get a feeling intuitively of what is the best model. Another thing that sometimes people do is they just run some automated benchmark out there and then try to see, okay, based on this benchmark, which one has the best performance. So there, you just select the preferred model. It can be, for example, the third model or the one with the highest score. The problem is there are a couple of problems with this, and we can stop this with a little game. So here, I'm just going to show three images. Maybe one question for you is how many people in the room prefer the first image among these three? This is a question. In general, you can ask a lot of questions. The prompt is the question of what image do you prefer and so on. The prompt was, I think, a little guy and a parrot or something like this. Okay, first, who prefers the second image? Okay, couple of people. Who prefers the third image? Okay. So what is great here is we have seen that people have different preferences. It's important to see that if you do manual inspection, you will be super biased to your own preference. So it's very important to not trust only your preference because then you have a big surprise that actually it's not the models that are preferred by everyone. Now we can do it again. Same question: who prefers the first image in this case? I think the prompt was probably a man eating some soup with pasta or something like this. Okay. Who prefers the second image? Okay, great. And who prefers the third? Okay. So that's also super interesting because I've seen some people changing their minds. Always on the left, it was the seed remodel middle Flux One. On the right, with some models we developed, one image. The idea is that you are super biased to all the few samples that you look at. So when you do manual inspections, you are two times biased by you and also by the number of samples, the specific samples you look at. So in general, the idea is you should never only trust the manual inspection. It's good to get the feeling, but it's not enough. You should always ask many people to do it. Human evaluation is usually great, but you have to scale it properly. Another problem is that when you do not human evaluation but more proper automated evaluation with metrics, sometimes you have inconsistent results. So for example, this is a bit small, but you can trust me. We ranked eight models regarding some metrics, like very standard metrics which is called CLIP score. Sometimes when people try to evaluate image models, they check this metric first. And you can see actually that if you check the rankings for the three metrics we looked at, like CLIP score on different datasets, it changed all the time. And these metrics are supposed to be between zero and one or zero and one hundred. Actually, the variations between models are super small, so it means it's hard to know from this metric what is the best model. What you should do is actually first have some clear understanding of what the metric does and also use multiple metrics. So here, for example, this is another type of metric when you know your use case. For example, you know you want to be the best at text rendering. There are a lot of text rendering metrics that would be better to evaluate your models. So here you can see again the ranking is way more consistent. You have always Z image being the first and P image being the second model. Also, the variations are way more significant. So the metrics are supposed to be between zero and one, and there are clear differences between every model. So yes, in general, it's very important to understand your metrics. People usually tend to just use some metrics and see, okay, I did my benchmark, and then I stop here. But it's important to understand what you actually measured with this. And now a last problem, which is actually common to the first and second methods that we've seen before, is that usually quality is driven by compute. So here, this is Flux image. And for the evaluation on design arena, I think, or maybe it's LM Arena, they did 26k battles. So it means they generated 26k images. Each of these images takes one minute to generate. So here I summarize all this information: 62 seconds per image, 26k evaluations. In total, to do these 26k evaluations, it takes 20 days of compute. In terms of cost, it's 5k just to run this evaluation. And in terms of energy, it's approximately 556 kilowatt hours. I know that people might not have the order of magnitude of what this amount of energy represents. So just to give some idea, I checked my Strava and how much energy I was consuming by running a marathon. Actually, it's equivalent to 400 marathons just to generate all these images. So it's a lot. I'm tired after one marathon, so I don't want to do 400 for sure. Now there are some alternatives. You can use different models. So of course, this is a model that we've done that does real-time generation and editing of images in less than one second. For the same amount of evaluation, it takes only seven hours. It also uses way less money, so 265 dollars. And instead of running 500 marathons, I just need to run four marathons. So if three of you want to run a marathon with me, it should be enough to do this. So again, the idea is people tend to just look at quality. But it's important not to look only at quality but also at efficiency. Because sometimes the additional gain you get with quality is not worth the efficiency and the compute cost. So to the question, what model is state-of-the-art? The answer is there are multiple state-of-the-art models. And the tool I prefer for this is usually the Pareto plots, where basically on the x-axis you have one efficiency metric, for example, on the left? 265 dollars and instead of running 500 marathons, I just need to run four marathons. So if three of you want to run a marathon with me, it should be enough to do this. So again, the idea is people tend to just look at quality. But it's important not to look only at quality but also at efficiency. Because sometimes the additional gain you get with quality is not worth the efficiency cost, the compute cost. So to the question what model is state-of-the-art, the answer is there are multiple state-of-the-art models. The tool I prefer for this is usually the Pareto plots. Where on the x-axis you have one efficiency metric. For example on the left, it's the latency for the generation of an image. On the right, it's the price for the generation of an image. On the y-axis you have some quality score. Let's say the CLIP score. And the key, you can draw the Pareto front in red and you can see that there is not one single state-of-the-art model, but there are actually multiple of them. There are three or four. And you can see that even though the quality score is not there, there are no big variations. It's always between 1100 and 1200. There is a big difference in terms of efficiency. So you can be really times, I don't know, 20 times faster just by using a different model. Even better, if you know the specific tasks you want to do, you can draw the Pareto front not with a quality metric which focuses on general capability, but really based on quality metrics which are for the target use case. So this is some Pareto front focusing on text rendering. And here, for example, we optimized a lot the Flux 2 model, the Flux 2 Flux models. We work with BFL for this. And you can see that you can get way faster. You can still be on the Pareto front for the specific use case of text rendering. So is benchmarking dead? The idea is it's not. We can do it properly and get a lot of useful information out of this. And if you use it in a better way, by taking all these rules when using the evaluation, we usually not find a large foundational model but more a lot of small preference models that will be very good for your use case. So I just listed a couple of takeaways which are: create benchmarks on many samples, look at the user use case conditions, use multiple benchmarks for efficiency, which are key things to keep in mind when evaluating models. And how to reach state-of-the-art models in general, this is what we are doing at Puma. We are actually building a lot of what we call performance models with that ourselves. Behind the endpoints, we have the fastest, for example, image models, video models that can run between one second to five seconds. But we also try to give a lot to the open source with a lot of open source contributions, with a package to show you how to compress your models on your own. Also a lot of materials on all the best research papers for efficiency, or even some efficiency courses. So thanks for your attention. I think you are out of time, but if there are any questions, happy to take them. You have a question? Sure, so there are actually multiple compression methods. There are a lot of families of compression methods. Of course you can guess quantization. Things we do a lot and we do a different quantization for every specific module in the model, which is super important. We can also do some pruning where we just remove some components which are not important. And for all these image and video models, something that works quite well is working on the step, the denoiser. When you generate a video or an image, you usually use 20 to 50 steps to generate the content. And you can actually reduce it a lot, either via distillation or caching methods. So instead of doing 50 times the computations using the same backbone, you can do it way less. I don't know, 20 times or even four times, depending on how aggressive you want to be. I want to give you a little bit of caching. Yeah. I want to understand if you guys know something different that I would like to use. Yeah, so we have, in our package, we have a lot of open source algorithms for good caching. But we have also some internal algorithms that we have for the models we serve behind the endpoints. But yes, there are really advanced caching methods and so on. But yes. Sure. Thanks. Okay. One way to do it is what I see the most actually in the image and video generation With certain some people just do manual aspect inspection They try a couple of prompts a couple of models and they they get a feeling intuitively of it. What is the best model? Another thing that sometimes people do is they just like run some benchmark automated benchmark out there and then try to see okay based on this benchmark which is the one that were that has the best performance So there yes, basically then you you just select the preferred model so it can be I don't know for example the third model or the one with the highest score The problem is that so there are a couple of problems with this and we can stop this with a little game so here I'm just going to show like three images and Maybe one question for you is how many people in the room prefer the first team prefer the first image among these three So this is a question like in general you can ask a lot of questions So does it idea of the prompt is the idea what image do you prefer and so on but? the prompt was I think a little guy and a parrot or something like this and Okay, first who prefers the second image? Okay, couple of people who prefers the third image? Okay So what is great here that we have seen that people have different preference so It's important to see that if you do like manual inspection you will be super biased to your own preference So it's very important to not trust only your preference because then you have big surprise that actually it's not the models that are preferred by everyone Now we can do it again same question who prefers the first image in this case? I think the prompt was like probably a man eating some soup with past we with pasta or something like this Okay Who prefers the second image? Okay, great, and who prefer the third? Okay So that's also super interesting because I've seen some people changing their minds so always on the left it was the the seed remodel middle flux one and On the right like with some models we developed one image and The idea is like also you are super biased to all the few samples that you look at So when you do manual inspections you are two times biased by you and by also the number of samples the specific samples you You look at so so in general the idea is like you should never only trust the Manual inspection it's good to get the feeling but it's not enough You should always ask many people to do it and human evolution is usually great, but you have to scale it properly Another problem is that when you do now not human evolution but more like proper like automated evolution with with metrics Sometimes you have like non consistent results So for example, this is a bit small, but you can trust me We ranked like eight models regarding some metrics like a very standard metrics Which is called keep score and sometimes people when they try to evaluate image models they do that they check this metric first And you can see actually that if you check like the rankings for The three metrics we look like clip score on different data sets it changed all the time and And this metrics are supposed to be between zero or zero and one or zero and one hundred and actually the variations between models They are super small so it means it means that it's hard to know from this metrics. What is the best model? What you should do is actually First having some clear understanding of what the metric does and also use multiple multiple of them So here for example, this is another type of metric when you know you you know your use case for example You know to you know you want to be the best at text rendering you there are a lot of text rendering metrics that would be better to evaluate your models So here you can see again like the ranking is way more consistent you have always z image being the first and p image be a being the second model and Also the variations they are way more significant significant So the models are supposed to be zero the metrics are supposed to be between zero and one and there are like clear difference between like every every model So yes in general very important understand your metrics people usually tend to just use some metrics and see okay I did my benchmark and then I stop here, but it's important to understand what you actually measured with this And now a last problem which is actually common to the the the first and second method to that we that we've seen before is that usually quality is driven by compute So here this is 30 pt image and for the evaluation design arena I think or maybe it's a lm arena They did like 27 26 k battles so it means they generated 20 26 k images and each of this image takes one minute to generate so here I summarize all this information 62 seconds a 26 k evaluations and in total to do these 26 k evaluations it takes 20 days of compute In terms of cost it's 5k just 5k just to To to run this evaluation and in terms of energy It's approximately you know 556 kilowatt kilowatt hour So I know that people might not have the order of magnitude of what it represents this amount of energy So just to give like some idea I check my Strava and check how much energy I was consuming by running a marathon and Actually to represent 400 marathon just to generate all these images So it's a lot I'm tired after one marathon, so I don't want to do 400 for sure Now there are some alternative you can use some different models So of course this is a model that we've done that does like time to generation editing of images in less than one second and for the same amount of Evaluation it takes only seven hours It takes also like whether it uses also where way less money so 265 dollars and instead of running 500 500 marathons. I just need to run four marathons So if three of you want to run a marathon with me, it should be enough to do this So again, the idea is like people tend to just look at quality But it's important not to look only at quality but also at efficiency Because sometimes the the additional gain you get with quality is not worth the efficiency the the compute cost So to the question what model is state-of-the-art the answer is there are multiple state-of-the-art model and The tool I prefer for this is usually the Pareto plots Where basically on the x-axis you have one efficiency metric for example on the left? It's The latency for the generation of an image on the right. It's the price for the generation of an image On the y-axis you have some quality score Let's say the yellow score and the key you can draw the Pareto font in red and you can see that there is not one single State-of-the-art model, but there are actually multiple of them and they are like three or four and you can see that Even though the quality score is not there is no big variations. It's always always between 1100 and 1200 there is a big difference in terms of efficiency So you can be like really like times times. I don't know 20 times faster just by using the different model Even better if you know the specific tasks you want to do you can do that the part of our front not with a quality metric which focuses on general capability but really Based on quality metrics which is for the target use case so this is some part of front focusing on text rendering and here for example We optimized a lot like the flux to model the flux to flex models We works with BFL for for this and you can see that you can get way faster You can still be on the part of front for the specific use case of text rendering So is benchmarking dead? The idea is it's not that we can do it properly and get a lot of useful information out of this and if you use it in a better like by taking all these You know rules when using the evaluation We usually usually not find a large long a large foundational model but more like a lot of small Perference models that will be very good for your use case So I just listed a couple of takeaways which are like a great on many samples You look at the user use case conditions use multiple benchmarks or efficiency Which are key things to keep in mind when evaluating models and How to reach like state-of-the-art models in general like this is what we are doing at Puna? We are actually building a lot of what we call performance models with that ourselves behind hand point hand points We have the fastest for example image models video models that can run between one seconds to five seconds and But we also try to give a lot to the open source With a lot of open source contributions with a package to show you how to compress your models on your own Also a lot of materials on all the best research papers for efficiency or even like some efficiency course So thanks for your attention I think you are out of time, but if there are any questions happy to take them You have a question Sure, so actually there are multiple like I mean, you know it as well There are a lot of family of compression methods So of course you can guess like quantization things we do a lot and we do it a different quantization for every specific module in the in the model Which is super important? We can do also some pruning where we just remove some components which are not important and for all these image and video models something that works quite well is working on the step that the denoiser like when you generate a video or an image you usually use like 20 to 50 steps to generate like The the content and you can actually reduce it a lot either via distillation of or caching methods So you instead of doing like 50 times the computations using the same backbone you can do it way less I don't know 20 times or even like four times Depending on how aggressive you want to be I want to give you a much I want to give you a little bit of caching Yeah I want to understand if you guys know something different that I would like to use Yeah So we have a like in our package we have a lot of open source algorithms for good caching But we have also some internal you know algorithms that we have for the models we serve behind the the handpoints But yes, there are really advanced caching methods and so on but yeah Sure Thanks Okay Yo Yo Yo Yo Yo Yo Yo Yo Yo Yo