Open Reader

Evaling Video Slop — Maor Bril, Character.ai

completed 23:13 Jul 25, 2026 Watch on YouTube

Current Status

completed

Video ID

b_PmGocP4rc

RAG / Chat

Enabled
Evaling Video Slop — Maor Bril, Character.ai
Description

A generated clip where the character stands frozen for four seconds can still score well, because the judge rewarded the gloss and the vibe instead of what actually happened. That failure is the whole problem with evaling video: CLIP score misses temporal incoherence, a team watching clips on Friday does not scale, and any AI judge you wire up drifts from human preference unless you measure the drift. Video breaks the text playbook because it has to hold temporal consistency, shot continuity, and a coherent story across frames, not just look good in a single still. The fix that stuck was to stop scoring and start comparing. Absolute scores collapsed to one dimension, but pairwise preference, is B a better story than A, held up, so Maor Bril's team trained a Qwen3-VL judge with Bradley-Terry loss on pairs of real and deliberately broken footage to catch slop before it ships. Drift is cheapest to catch early, especially on longer form video, so the judge runs as a regression gate in CI: every AgentX release at Character.ai clears an eval wall, calibrated against human scores, before users ever see it. Speaker info: - https://x.com/maorbril - https://www.linkedin.com/in/maorbril - https://github.com/character-ai/judgejudy Timestamps: 0:00 - Introduction: evaluating AI generated video 1:19 - Why video generation drifts between frames 3:14 - Story and sound: what a clip has to get right 4:43 - LLM as a judge, and catching drift early 7:01 - Story and sound failure modes 8:28 - Small model vs bigger model as judge 9:20 - Don't score, compare: pairwise preference 10:47 - When the judge scores vibe over substance 11:53 - Pairing real footage to train a quality detector 13:27 - Self verification in the generation loop 15:05 - Q&A

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: High-quality AI video requires evaluation systems that judge story-level and temporal quality inside the generation loop, using relative comparisons, calibrated human feedback, and a fast distilled multimodal judge rather than frame metrics or generic LLM judging alone.
  • Why it matters: This is a concrete eval-system design pattern for any agentic generation workflow: use expensive expert evaluation to build a benchmark and training signal, then distill it into a low-latency verifier that catches defects before downstream assembly makes them costly.
  • Best use: Use this as an architecture reference for building evaluation and self-correction loops for video or other multimodal agent outputs, especially where latency, throughput, explainability, and human taste alignment matter.

Executive Summary

Maor Bril argues that AI video generation has improved faster than AI video evaluation. Existing metrics such as CLIP-style prompt/frame alignment and frame-consistency measures can detect local defects, but they do not reliably determine whether a video tells its intended story, obeys physics over time, maintains character identity across shots, has plausible pacing, or synchronizes sound with events.

Character.ai’s initial answer was a repeatable evaluation harness: combine conventional metrics, an LLM-as-judge, and human annotation that calibrates the judge. This was useful as an offline benchmark but too slow and expensive to sit directly in a high-volume creation workflow. The key operating principle is therefore to move evaluation upstream and evaluate individual short generations or transition frames before they are assembled into a long-form video.

The team distilled its committee of evaluators into a small vision-language model that can score a 15-second video in roughly three seconds and explain why it failed, such as extra limbs, broken physics, or audio-video mismatch. They deliberately chose the smaller model over a stronger but much slower alternative because the incremental quality did not justify the latency cost.

The central modeling lesson is “don’t score, compare.” Absolute one-to-ten quality scores are subjective and poorly calibrated across annotators; pairwise judgments between two videos are more consistent. However, their first model failed confidently because synthetic corruptions taught it to recognize an overall AI-looking “vibe,” not the intended evaluation axis. They repaired this by rebuilding the data around explicitly annotated axes and pairing real footage with AI footage without letting artifacts become the proxy for quality. The resulting system is evolving into an agentic workflow that can validate and repair its own outputs.

Key Takeaways

  • Claim: Frame-level quality metrics are insufficient for video because the important failures are temporal and narrative, not merely visual or prompt-alignment errors. | Evidence: Bril contrasts CLIP-like frame scoring and inter-frame drift checks with video-specific requirements: a character walking downstairs must walk rather than hover, identity must persist across shots, travel and action pacing must make sense, and a door-slam sound must occur at the visual moment of impact. | Implication: Define explicit, task-relevant temporal axes for multimodal outputs rather than treating aggregate visual quality as a proxy for user value. | Caveat: Frame and consistency metrics are still useful components of a broader harness; the critique is that they cannot serve as the sole quality definition.
  • Claim: Evaluation should be embedded as early and as close to generation as possible because defects become more expensive to fix after composition. | Evidence: The speaker gives two examples: catch character drift between starting frames before generating the corresponding shots, and regenerate a bad six-second segment before it is combined into a three-to-five-minute video. | Implication: Design generation systems around small, independently verifiable units with reject/regenerate loops, rather than performing a single final-quality check on a completed artifact.
  • Claim: Use an expensive committee of metrics, frontier models, and human feedback to create a benchmark, then distill that capability into a small, fast multimodal evaluator for production. | Evidence: Character.ai combined conventional metrics, a consistently prompted LLM judge, and human annotations; its distilled small VLM scores a 15-second video in about three seconds and returns reasons for failure, not just a slop/non-slop label. | Implication: Separate offline evaluation fidelity from online evaluation latency: use the former to establish quality and supervision, and the latter to govern real-time generation. | Caveat: A larger model tested better, but was rejected because its quality improvement did not justify the added latency; the right evaluator depends on throughput and serving economics.
  • Claim: Pairwise preference training is more robust than absolute quality scoring for subjective video attributes such as storytelling. | Evidence: Bril notes that annotators will disagree substantially on whether one video deserves a four, five, six, or eight out of ten, but will more often agree when asked whether video A or B tells a better story. The model was trained on A-versus-B comparisons rather than one-to-ten labels. | Implication: For taste-dependent evaluations, collect preference data and optimize ranking behavior instead of asking humans or models for falsely precise absolute scores. | Caveat: Relative comparisons only work when the paired examples and evaluation axis are carefully constructed; otherwise the model can learn superficial correlations.
  • Claim: Synthetic bad-data generation can create a confident but invalid evaluator if the dataset lets it learn visual “vibe” or AI artifacts instead of the criterion being measured. | Evidence: Version 1 gave a 9.2 camera-work score to four seconds of a static image and praised physics in examples involving hovering or flying. Bril attributes this to corruptions that made the model detect coherent-looking video and artificial gloss rather than camera work, physics, or story quality. | Implication: Audit eval models for axis leakage: deliberately construct counterexamples where production source, polish, and the target quality attribute are decoupled. | Caveat: Pairing real video as positive data with AI video as negative data risks creating an AI detector rather than a quality detector.
  • Claim: Human feedback remains necessary, but it should continuously calibrate defined evaluation axes rather than attempt to encode a universal notion of taste. | Evidence: Character.ai periodically has people spend 10–15 minutes annotating videos across randomly assigned axes; those annotations calibrate AI judges and become training data for subsequent model versions. | Implication: Run lightweight recurring annotation programs with narrow rubrics, measure disagreement, and use the data to recalibrate evaluators instead of treating a one-time labeled dataset as durable ground truth. | Caveat: Bril explicitly describes taste as subjective and the process as iterative rather than immediate; even human judges will disagree on what is great.
  • Claim: An agentic evaluator is more adaptable than a fixed pipeline when users bring diverse characters, stories, images, and voices into the generation system. | Evidence: Bril says fixed pipelines work for a narrow use case but drift with heterogeneous user intents; Character.ai shifted from a complex pipeline to an agentic workflow in which agents have tools to validate outputs and fix errors as they proceed. | Implication: Use evaluators as callable tools in an orchestrated generation loop, but require observability and controlled repair policies before assuming an agent can reliably self-correct. | Caveat: The transcript does not specify the agent architecture, repair policy, or measured improvement over the prior pipeline.

Detailed Brief

Audio evaluation and unresolved lip-sync limitations

  • Claims: Audio quality evaluation combines general intelligibility/quality checks with event-level temporal correlation between the visual stream and the audio waveform.; When the prompt specifies an event such as a door slam, the evaluator can identify the relevant visual frame and check for an audio spike at that frame’s timestamp.; Lip-sync remains unsolved, particularly for animated characters whose mouth motion has no natural correspondence to spoken phonemes.
  • Evidence: Bril explains that the system does not need to classify the sound itself as a door slam; it looks for an expected sound spike at the timestamp where the visual event occurs.; Talking-head human characters can potentially be evaluated with lip-focused techniques, unlike stylized talking animation.
  • Caveats: Event timing is not equivalent to semantic audio understanding or reliable lip-sync verification.; Prompt-conditioned checking can only validate events that are specified or otherwise inferable from the video.
  • Implications: Treat audio-video synchronization as a timestamp-alignment problem where possible, but maintain a separate quality track for semantic sound correctness and speech articulation.; Avoid marketing a generalized audiovisual evaluator as a solved lip-sync system.

Scale, deployment economics, and observability

  • Claims: The rationale for a custom small VLM is primarily speed and serving economics, not a claim that it universally outperforms frontier judges.; For hundreds or a thousand videos, an offline expert cohort may be sufficient; for thousands or tens of thousands per day, frontier-model judgment can become cost-prohibitive.; The released repository is a harness that can connect to different agents and LLMs, while Character.ai runs an internal service version with an agentic harness and its chosen metrics.
  • Evidence: Bril frames scaling as a deployment decision: one model instance on one GPU versus many instances, balanced against training, dataset curation, serving cost, and response time.; In response to a question about OTel traces, he accepts the feature request and says OTel telemetry will be added to the harness.
  • Caveats: No total-cost figures, benchmark accuracy, data volume, or evaluator error rates are provided.; The transcript references a likely frontier model name unclearly ('Fable'), so no performance comparison can be relied upon.
  • Implications: Set an explicit routing threshold: reserve expensive judges for benchmark creation, ambiguous cases, and audit samples; use the distilled evaluator for routine gating.; Instrument evaluator inputs, judgments, reasons, model version, retries, and downstream repair outcomes so quality regressions can be traced.

Notable Concepts & Terms

  • Video slop: The speaker’s term for generated video that may look superficially plausible but fails on artifacts, physics, continuity, pacing, narrative, or audio synchronization.
  • LLM-as-a-judge: Using a general foundation model to assess generated output; useful for broad evaluation but slow, prompt-sensitive, and insufficiently repeatable without calibration.
  • Committee of experts: The offline combination of conventional metrics, LLM judging, and human annotation used to create a stronger reference evaluation system.
  • Small VLM: A small vision-language model distilled for fast production video evaluation; it sees visual content and returns failure reasons at usable generation-loop latency.
  • Relative / pairwise evaluation: Training or judging by asking whether A or B is better on a named criterion, rather than assigning subjective absolute scores.
  • Axis leakage: The failure mode where an evaluator learns an unintended proxy—such as AI artifacts or general polish—instead of the specific attribute it is supposed to assess.
  • Agentic workflow: A generation architecture in which agents can invoke validation tools, inspect outputs, and repair or regenerate work rather than executing one fixed pipeline.
  • OTel telemetry: OpenTelemetry-style tracing requested for the evaluation harness, relevant for observing judge calls and integrating the system with external platforms.

Operator Notes / Why Ken Should Care

  • For any multimodal generation product, create a quality taxonomy before choosing models: continuity/identity, physical plausibility, prompt-event fulfillment, pacing, narrative coherence, audio quality, and event synchronization should be separately testable.
  • Build a two-tier eval stack: a high-cost benchmark/audit judge plus a fast production gate. Track disagreement between them and send disagreement cases to human review.
  • Collect pairwise preference labels on isolated axes, and include hard counterexamples that break correlations between AI provenance, visual polish, and actual quality.
  • Put eval calls at shot, scene-transition, and assembly boundaries; make failed outputs regenerate locally rather than forcing full-project reruns.
  • Require reason-coded failures from production evaluators and instrument them with OpenTelemetry-compatible traces, model versioning, prompt/context capture, and repair-result logging.
  • Treat lip-sync as an open risk area; do not use it as a fully automated release gate without domain-specific validation, especially for animated characters.

Source/Metadata

  • Title: Evaling Video Slop — Maor Bril, Character.ai
  • Transcript words: 3619
  • Duration seconds: 1393
  • Timestamp note: No timestamps or chapters were present in the supplied transcript.

Transcript

3365 words en Processed in 174.4s

So, hi. I'm Eor. I've been with Character for a bit over two years, and we'll talk about AI slop. I think that when we look at video generations as a whole, we have two parallel tracks. One is the video generation, which became insanely good from models like Kling and C-Dance and Vio and Sora. We still remember Sora. But the part that got left behind is how we evaluate the quality of the video that was generated. So, on the one hand, we still squint at it and decide whether or not it's good. But on the other hand, we know that the generation has gotten a lot better. And when we look at X or whatever social you're consuming your content on, there are a lot of guides on how to create amazing videos with this model or that. So, the hard part was never how to make video. The hard part was how do we generate good enough video, and how do we judge if the video is good enough? So, now we've gone to a world where the generation of video is basically free, right? Free, especially when you compare it to how much studios would charge. But the problem is the grand majority of video that is generated is not that good, right? We have a lot of hallucinations: a third limb, opening and closing the door at the same time, hovering, physics, et cetera. So, unfortunately, in order to get high quality content, we need a human to judge. And I don't know when was the last time you've seen how someone is creating these long-form generated videos. It's usually a lot of shorter generations and a lot of editing. The problem is because we're using a lot of the tools that we built for the text era, for the image era, for videos, right? We're using things like clip score, which is great to judge a single frame. Things like IPS will help us detect the drift between frames. But the problem is when you combine all these together, all these tools are good at watching the individual frames. They're good at checking this, this, this, this: does this specific frame, does it match the prompt that generated it, right? It will check consistency between frames, and it will check whether or not it matched the prompt that drove it. But what it won't do, it doesn't tell you if you told the story that you meant to tell, right? If you think about what video is, video is a storytelling medium. Video is just another form on how we tell a story, right, for any type of story. So, one of the things we have to look at: does it tell the actual story? Does the physics make sense? For example, if we want a video of a character walking downstairs, does it actually walk or hover? Does the character stay the same character across multiple shots? Does the pacing make sense? For example, people take time going from one place to another. We need to make sure that the pacing makes sense as well. And especially when we add audio, we want to make sure that the audio is synced with the imagery. For example, if someone is slamming a door, we want that sound of the door being slammed to be exactly when the door is actually being slammed. Now, the next iteration we all went to a while ago, we started using LLM as a judge for everything. And we have amazing foundational models that we just throw videos at. The problem with them is that, A, they're slow. B, they're only as good as your prompts. And multiple people will prompt multiple ways. And the same model may respond in a very, very different way. And sometimes the prompt we use is, is it consistent? Does this match the prompt? But then the question we really care about: is it good? And the answer varies. So, oops, sorry about that. So, our first iteration is, let's take all these things and build a repeatable benchmark on how we test video that we can rerun over and over and over again. So, that combines both metrics, as I said earlier, that know how to view individual frames. But also, consistent LLM as a judge, right? Where we also use human annotation to calibrate the LLM as a judge. So, for every report that we generate with that harness, we're able to have humans annotate and basically feed that feedback back into the LLM as a judge prompt to make sure that it's aligned with what I think or what the annotator thought is good. And we use it to score the videos. The problem with this approach, it's very slow, it's very expensive, and especially when we want to bring it for users to be able to generate a lot of video. Because creation is a very hard process. And so, the problem, as I said, the problem is this is a slow process, and we need to bring it as close to the users as possible and also earlier into the process. The reason for that is if we take a look at all the metrics and there are mistakes that we can find earlier than later, then it's a lot cheaper to correct that particular mistake. So, for example, right, on the left, we have two starting frames of different shots, right? But it's easy to correct to view it at this point and see, did the character drift between frame one and frame two, because those frames will be used as starting frames to generate videos. So, if you can catch the drift at this point and correct it, then it's much cheaper to generate the video as a whole because we can correct it at a much cheaper cost. And the same thing applies when we look at longer-form video, right? When we see all these three, four, five minute long videos, they're usually a collection of a lot of shorter videos. And being able to catch a six second generation that drifted and regenerate that before we combine the whole video will end up being a better result as a whole. And now the other problem we're trying to solve is some of these axes, right, only exist across time, right? So, for example, when we look at the, right, we mentioned the story, right? So, does the story that we're trying to tell with that video, does it hold in that video? Does the video tell the exact story? Does the pacing make sense, right? Does the sound, right? Does the sound, right? So, as I said, the underlying goal is to bring that evals closer to the online generation because the sooner we're able to catch those mistakes, the sooner we're able to catch that drift, right, then it's much easier, much cheaper to fix. So, now the problem is that, as I said, this is a very slow process. So, the solution is actually to take all these committee of experts and distill it into one small model that is also very, very fast. But it is able to give us a response that is not whether or not this video is slop or not, but why is it slop, right? Why is that video scored low versus the other? Because, for example, it added an extra limb, because it didn't obey physics, because the audio was out of sync. So, the goal was, A, build it on top of a small VLM. And why is it a VLM? VLM because we needed the model to be able to see the image. But also, we needed it to work fast, right, because we brought it closer to the generation. Where, in fact, with the model we have trained, it takes about three seconds to score a 15 second video. Now, we also tested a bigger model and the results were better, but it was significantly slower. And the decision was to go with the smaller model because the added value from the bigger model didn't justify the slowness. The other very interesting realization we came to is don't score, compare. What does that mean? For example, if I'll ask any person in this room to look at a particular video and rank it from one to 10 on storytelling, right? I'm pretty sure that what will be a six for you will be a five for you, will be a four for you, and an eight for you, right? But if I'll show you two videos and I'll ask you which one of them is telling a better story, the grand majority will probably agree that B is telling a better story than A, right? And if you do it enough times, then it's easy to generalize the model toward detecting what's better versus not. So, we trained on pairs, right? The other very interesting realization we came to is: don't score, compare. What does that mean? For example, if I ask any person in this room to look at a particular video and rank it from one to 10 on storytelling, I'm pretty sure that what will be a six for you will be a five for you, will be a four for you, and an eight for you. But if I show you two videos and ask you which one of them is telling a better story, the grand majority will probably agree that B is telling a better story than A. And if you do it enough times, then it's easy to generalize the model toward detecting what's better versus not. So we trained on pairs, A versus B, as opposed to one through 10. Now we manufactured badness. Luckily, the internet is full of very high-quality videos, and it's very easy to get good videos. And it was very fun to create bad videos, A, by either corrupting good videos or by generating random flop. Now we shipped V1, and it was so wrong. It was wrong, but it was wrong in a very confident way. For example, the frame you see here is from a video that the model scored 9.2 on camera work, and the camera didn't move. For four seconds, it was a still image of the same character, but the model was very happy with the cinematography. So the physics in some other videos, which I'm not showing because of time limitations, it says that the physics look great, but it set it on ghosts hovering and people flying, et cetera. So then the question was, why was it wrong? The reason it was wrong is because of how we generated that data. It scored the vibe as opposed to the axis. So it learned how to detect coherent videos, and it learned how to detect the artificial artifacts, the gloss of the video as opposed to whether or not the video actually told the story. And so the solution was to fix the dataset. And so the way we fixed the dataset, I started pairing real footage versus AI footage. Now, the risk with that, and that's the reason why I avoided doing it at first, is because I didn't want to create an AI detector. Because if you start creating pairs of good as human-generated video and bad as AI video, then there's a very big chance of the model overfitting and becoming an AI detector as opposed to a video quality detector. So there are two things I did in order to avoid that. There are no artificial artifacts for video A versus video B. And I used the exact same method of annotating both videos. So both the axes in those videos were annotated in the same way. And surprise, it turned out pretty awesome. And so now what we're able to do, especially when you're looking at videos, A, we changed from a very complex pipeline to an agentic workflow. The reason behind this is the pipelines work great if you have a very unique use case. But once you put it in front of users, they'll have a very distinct story that they want to tell with their own characters, with their own images, and their own voice. So that's when it starts to drift. But by providing the agents with tools to validate the quality of the output, it's able to adapt to changes better. But it's also able to verify its own work and fix things as they go along. So if you're going to steal from this talk a few things: One, go relative, not absolute. As I explained earlier, the value of comparing video A versus video B will always give you a better result going forward. B, score the real axis that you care about. So if you care about storytelling, if you care about pacing, if you care about physics, score those axes. Don't expect them to miraculously appear. And put eval inside the generation loop, especially if your goal is to have a higher quality of generation. Get the eval as close to the generation loop as possible. Eventually, evaluate it as a story. Videos are stories. Videos are just another way for us to tell stories to others. And thank you very much. All right. Any questions? Okay. Down here. Awesome. All right. I got two down here. Here you go. Hi. How do you eval sound, sound effects, and video matching? I'm sorry. Can you repeat? How do you eval sound, sound effects, and matching with the video? Oh, yeah. That's a fantastic question. So sound is actually a combination of a few things. One, I'm using to make sure that the sound quality is high enough and is understandable. B, the model will learn to identify key frames. And especially because when I feed something into the model, it can be just a video, or it could be the video plus the prompt that generated that video. So, for example, if the prompt says the door slammed, it will look for a door being slammed and will match the sound at that same frame. Did that answer your question? How does the model exercise? So it's both by using Atmos and also to correlate the— So, for example, when it's looking at the frames, it's making sure that, for example, the door being slammed at frame 6. Frame 6 has a specific timestamp, so it's looking for that spike in the sound at that timestamp. It doesn't know that it is that sound, but it's looking for a specific spike of sound at that timestamp. What about lip syncing? Lip syncing is an unsolved problem yet. We're trying, though. Yeah. Yeah, so the question was, what about lip syncing? That wasn't me, but I guess the lip syncing answer would be interesting before I ask my question. Yeah. As I said, it is an unsolved problem still. We're still working through it. Especially for us, some of the characters that we're trying to do are talking head humans, which we can look at with different techniques to try to identify the lips. But some of them are just talking animations that have no real correlation between the movement of the mouth and speech. So, unfortunately, I don't have a solution for that yet. So I'm curious about, for example, if you wanted to further enrich the dataset with human evaluation. Yes. The question of taste and what is good, because I think there is a big question mark about whether that is going to remain the domain of humans. But I've also seen people say that most humans have terrible taste anyway in videos and games and books. Fair. So how would you construct and align human judges? Yeah. So this is actually solved at first at the judge duty part, where every report it will generate, a human can go and annotate it. And we actually do that. But some of them are just talking animations that have no real correlation between the movement of the mouth and speech. So, unfortunately, I don't have a solution for that yet. So I'm curious about, for example, if you wanted to further enrich the data set with human evaluation. Yes. The question of taste and what is good, because I think there is a big question mark about, is that going to remain the domain of humans? But I've also seen people say that, well, most humans have terrible taste anyway in videos and games and books. Fair. So how would you construct and align any human judges? Yeah. So this is actually solved at first at the judge duty part, where every report it will generate, a human can go and annotate it. And we actually do that. We will periodically have sessions where everyone spends 10 to 15 minutes just annotating videos. And that usually happens on multiple axes. I won't ask everyone to annotate the same video on 10 different things. It will be random. And I use the data to calibrate the AI judges. And the results from that are actually being served as a data set for training for the next version of that model. So it's a process that does take a little bit of time, and it does evolve over time. But it's not immediate because also taste is very subjective, and things that are great for me, that I think are fantastic, some people come here and say, are you sure they're great? Because, so yeah, it's a process. And I use the human feedback to calibrate the models all the time. How did you land on the QAN small VLM? Did you try any others? I did. So the intent I had was to, A, find a small enough model. The reason I went with QAN is because we also had a very good experience with post-training QAN on other use cases. So yes, I could have, I did try a few others, but everything was just there and it was good enough. So my question is about scale. Yeah. So obviously Character produces thousands, millions, bajillion videos. Yeah. What scale does this become reasonable for my domain that is not Character? Right. So my domain has hundreds, maybe a thousand videos. Sure. So if you're happy with the cohort of experts and you don't need, right, so I'll rephrase that. The scale is both for speed, right, as well as capacity. Because I can serve this model as one instance on one GPU, or I can serve it as a hundred instances, right? So that determines my scale. The reason I chose to go toward the model is because I wanted to speed up the creation process, right? It would have worked just as well if I didn't have this particular model. I would have used the cohort of experts, right, from metrics that are available both on CPU and GPU, as well as frontier models, right? So it was a balance of, A, how long did it take me to train this model and to cure the dataset and get it to a working set, right? And how much does it cost to serve it versus how much it would have cost me to do this, A, slower. Now, potentially it is better, right? I mean, I assume that if you're going to use the Fable, which came back today, right, it will probably give you a better result. But at what cost, right? If you do it for one or two, that's probably fine. If you do it for thousands or tens of thousands per day, it adds up. So it's a matter of your economics. Cool. Right over here. Yeah. On your right. There you go. Last question. It's very bright. I'm sorry. No worries. No worries. My question is, I looked a bit at the repo. You guys don't export OTL traces of the LMS judges yet. Correct. Is that something, are you open to that? So you can connect to other platforms? Sure. So the repo itself, it's a harness, and you can connect any agents or any LLMs you want. We actually have an internal version of this, which is running it as a service, right, with an agentic harness on top of it that has all the metrics we care about. But I do accept your feature requests, and I'll be adding OTL telemetry to the harness. Awesome. Thank you very much. A warm welcome or round of applause for Mayor. Thank you. Thank you all. Thanks. Thanks. So it's a matter of your economics. Cool. Right over here. Yeah. On your right. There you go. Last question. It's very bright. I'm sorry. No worries. No worries. My question is, I looked a bit at the repo. You guys don't export OTL traces of the LMS judges yet. Correct. Is that something, are you open to that? So you can connect to other platforms? Sure. So the repo itself, it's a harness and you can connect any agents or any LLMs you want. We actually have an internal version of this, which is running it as a service, right, with an agentic harness on top of it that has all the metrics we care about. But I do accept your feature requests and I'll be adding OTL telemetry to the harness. Awesome. Thank you very much. A warm welcome or round of applause for Mayor. Thank you. Thank you all. Thanks. Thanks. . . . . .