Open Reader

Agents at Scale: Inside MiniMax's Model and the Infrastructure Behind It — Olive Song

completed 20:14 Jul 31, 2026 Watch on YouTube

Current Status

completed

Video ID

AVMr9PMINyo

RAG / Chat

Enabled
Agents at Scale: Inside MiniMax's Model and the Infrastructure Behind It — Olive Song
Description

In this conversation, Olive Song, who leads reinforcement learning at MiniMax, opens up the stack behind the company's open weight models and the infrastructure that serves them. Her starting point is a belief in open source: put the weights out, let builders optimize on them, and share the capability widely. From there she walks through what it takes for a model to land well, from agentic coding that also understands and builds games, to computer use problems trained with RL against environments like OS World. Much of the discussion is the unglamorous engineering that makes a launch real. When a model ships, the team wants the inference stack ready on day zero, which means writing and tuning GPU kernels, working through benchmarks like a parallel kernel bench, and threading optimization into everything from KV cache handling to routing. Song also talks through multimodality and the training pitfalls that come with it, such as text and vision collapsing after training unless you train both modalities together, and reflects on longer horizon tasks like replicating a twelve hour run. She closes optimistic that open models are closing the gap faster than a year ago would have suggested. Speaker info: - https://x.com/olive_jy_song Timestamps: 0:00 - Introducing the RL lead at MiniMax 1:17 - Why open source and open weights 3:45 - What builders are doing with the model 4:11 - A model that builds games 5:12 - Computer use and OS World 5:49 - Writing GPU kernels 6:25 - Parallel kernel bench 7:28 - A day zero inference stack 9:34 - Optimization across the stack 10:47 - Multimodality and training collapse 14:10 - Replicating a twelve hour run 17:22 - Where open models go next

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Skim
  • Core thesis: MiniMax and Together argue that open-weight frontier models become practically competitive through a combined loop of environment-centric RL post-training, multimodal capabilities, and continuous inference optimization tailored to long-context agent workloads.
  • Why it matters: The discussion identifies the operating constraints behind scalable agents: long-horizon task evaluation, reward-hacking controls, KV-cache infrastructure, and model-serving optimization matter as much as raw model quality.
  • Best use: Use this as a directional briefing on the MiniMax M3/open-model ecosystem and as a prompt to pressure-test your own agent evaluation, long-context caching, and inference-routing assumptions; it is not an implementation tutorial.

Executive Summary

Olive Song, MiniMax's RL research lead, frames M3 as an open-weight, natively multimodal model whose differentiation is not merely text and code performance but its ability to understand images and video in agent workflows. MiniMax's rationale for releasing weights is strategic: external developers generate usage feedback and contributions, while infrastructure partners can optimize the model for broader availability.

The training argument is that difficult agent capability comes primarily from designing the right environments, task formulations, reward functions, and evaluations—not from applying RL generically. MiniMax cites long-horizon tasks such as reproducing an ICLR paper in a 12-hour run and kernel optimization, where agents need access to GPU-constrained environments, iterative submissions, and anti-hacking validation rather than a single terminal score.

Dan, VP of Kernels at Together AI, describes serving a new model as an ongoing co-design process. Before launch, an inference provider studies architecture choices such as sparse attention, MoE structure, and quantization; it then decides whether existing kernels suffice or new ones are required. After launch, it continuously improves KV-cache handling, attention kernels, quantization, routing, and user-facing quality, sometimes at daily rather than weekly cadence.

The panel's most operational point is that agentic workloads have changed the inference problem. Coding agents may repeatedly use tools and place an entire codebase in context, unlike chat systems with modest system prompts and turn histories. At high concurrency and up to million-token contexts, KV-cache management becomes effectively a distributed-storage problem: determine where cache lives, whether it already exists, how it is retrieved, and how it is moved across infrastructure.

Key Takeaways

  • Claim: MiniMax positions open-weight release as both a distribution strategy and a model-improvement loop, rather than a one-time publication choice. | Evidence: Song says opening M3 lets anyone use it, lets developers contribute feedback and PRs, and lets partners such as Together optimize its inference; Together reports serving MiniMax 2.5 and 2.7 before M3 and holding the largest share of M3 token usage at the time of the panel. | Implication: For an agent platform, open models should be evaluated as an ecosystem asset: model selection, external hosting, kernel optimization, and user feedback can reinforce one another faster than a closed-model-only strategy. | Caveat: The discussion provides no independent adoption, quality, or economics data to substantiate the claimed token-share leadership or the effect of community contributions.
  • Claim: M3's multimodality is intended to improve agent workflows that require visual grounding, especially computer use and website iteration. | Evidence: Song says M3 was trained on text and image data from step zero rather than adding vision later; MiniMax found that text tokens attend to visual tokens in the attention map. Cited use cases include computer-use agents, tool-mediated creation, game development, and agents that inspect a website visually before optimizing it. | Implication: When an agent must act on UI state or verify rendered output, use a natively multimodal model or explicit visual feedback loop rather than treating screenshots as peripheral context. | Caveat: The panel gives qualitative examples but no benchmark scores, reliability rates, or comparison against modular vision-plus-language systems.
  • Claim: Long-horizon agent competence is mainly an environment-design and evaluation problem: models need iterative, constrained tasks with rewards that distinguish genuine progress from exploitation. | Evidence: MiniMax cites a claimed 12-hour run reproducing an ICLR paper and discusses kernel-optimization tasks that require GPUs and hardware constraints. Song identifies data, problem formulation, reward design, environment design, and RL-algorithm adjustments as the core levers; agents can submit multiple iterations, each evaluated, while validation and tests detect reward hacking. | Implication: Do not evaluate operational agents only on final-task pass/fail. Build staged checkpoints, sandbox constraints, repeated submission scoring, and adversarial validation before trusting a model on multi-hour workflows. | Caveat: This is a high-level account; MiniMax does not disclose its specific reward functions, environment setup, RL changes, success rates, or the conditions of the paper-reproduction claim.
  • Claim: Agentic coding workloads change inference priorities because large, persistent context and repeated tool turns create different bottlenecks from ordinary chat. | Evidence: Dan contrasts chat workloads—typically a system prompt plus chat history—with coding agents that may upload an entire codebase and execute hundreds of multi-turn tool calls. He says this changes KV-cache, prompting, routing, kernel, and inference-engine optimization requirements. | Implication: Treat coding-agent traffic as its own workload class in capacity planning and routing. Instrument context size, reuse patterns, tool-turn count, cache hit rate, and time-to-first-token separately from conversational traffic. | Caveat: The panel does not quantify the cost, latency, cache-hit rate, or context-length thresholds at which a different serving design becomes necessary.
  • Claim: At very long context lengths, KV-cache infrastructure should be treated like a distributed data system rather than a local implementation detail. | Evidence: Asked about concurrent requests at 500,000 to 1,000,000 tokens, Dan compares KV-cache management to recreating a distributed file system or large database: the practical questions are where cache is stored, whether it has been seen before, how it is fetched, and how it is transferred. | Implication: For persistent agent sessions, invest in cache identity and reuse semantics, placement/tiering, transfer overhead, and observability; otherwise nominal million-token support may be operationally expensive or slow under concurrency. | Caveat: No concrete architecture, consistency model, tiering approach, or failure-recovery design is provided.
  • Claim: Model-serving performance is not fixed at launch; it is a continuous optimization program that combines model-specific work with reusable lessons from prior architectures. | Evidence: Together begins optimization once it receives early architecture details, examining sparse attention, MoE choices, and quantization. Its launch plan covers quality and UX first, then a backlog involving KV cache, attention kernels, and quantization; Dan says performance improvements can occur from one night to the next. He also says lessons from DeepSeek- and GLM-style sparse attention can transfer to MiniMax's distinct design. | Implication: Avoid locking model/provider decisions based only on launch-day benchmarks. Rebenchmark on a schedule and preserve an abstraction layer that allows provider, kernel, quantization, or routing improvements to flow through without application rewrites. | Caveat: The speakers do not provide before/after throughput, latency, cost, or accuracy-preservation metrics.
  • Claim: The speakers expect open-weight models to close much of the frontier gap as post-training and development cycles accelerate, but argue that today's inference hardware remains underutilized. | Evidence: Dan references a reported SpaceX figure of roughly 10% FLOP utilization and says inference can use current hardware far more effectively. He names M3, GLM, and Kimi as evidence that open models are nearing frontier relevance; Song attributes faster progress partly to models being used internally to improve development speed through 'self-evolution.' | Implication: Maintain a live open-versus-closed model portfolio rather than treating the choice as permanent; the comparative advantage may shift rapidly as serving efficiency and post-training improve. | Caveat: These are forward-looking views from an open-model producer and an infrastructure provider, not an independent comparison of frontier-model capability, safety, or total cost of ownership.

Detailed Brief

Kernel optimization as a benchmark-to-production feedback loop

  • Claims: Together's Parallel Kernel Bench is deliberately designed to include unsolved but production-relevant optimization problems.; The panel rejects the usual concern that benchmark optimization is inherently unproductive when benchmark wins can be directly incorporated into the inference stack.
  • Evidence: Dan says Together surveyed inference-serving approaches and found potential model speedups for which effective kernels do not yet exist.; His stated intent is that if researchers overfit to Parallel Kernel Bench, Together can take the resulting kernels and use them to accelerate inference and development.; Song says kernel-focused RL requires deliberately constructed complex environments in which the model iteratively improves kernel performance.
  • Caveats: A production-useful benchmark still needs robust measurement across hardware, model variants, numerical precision, and real workload distributions; none of those validation details are covered in the panel.
  • Implications: For internal agent tooling that generates low-level code or optimizations, define benchmarks whose success metric maps directly to deployable performance, rather than generic coding scores alone.

Self-evolution and internal-evaluation flywheels

  • Claims: MiniMax describes using its own models to accelerate internal development, then using work-derived tasks as evaluations for subsequent model versions.; This creates a recursive development loop: a better model may speed experimentation, produce more relevant evaluation tasks, and accelerate the next post-training cycle.
  • Evidence: Song links M2.7 to 'self-evolution' and says MiniMax actively uses the model to improve internal development speed.; She says this work yields internal evaluations closely related to MiniMax's own operational tasks.
  • Caveats: Internal task-derived evaluations risk narrowing toward the lab's own workflows and do not establish broad external generalization without held-out, independently designed tests.
  • Implications: Create a governed pipeline that turns real agent failures and high-value operator tasks into versioned evaluations, while keeping an external evaluation set to detect over-specialization.

Notable Concepts & Terms

  • MiniMax M3: The open-weight MiniMax model discussed; positioned as multimodal from pretraining onward, with million-token context and sparse-attention architecture.
  • Native multimodal training: Training text and visual data together from the outset so language and image tokens can interact directly, intended to support visually grounded agents.
  • Long-horizon RL: Reinforcement learning for tasks lasting many iterations or hours, where environment constraints, intermediate submissions, and reward design become central.
  • Reward hacking: An agent exploits an imperfect objective or evaluation without delivering real performance; MiniMax says it uses validation and tests to identify it.
  • KV cache: Cached transformer attention state that enables efficient continuation across tokens and turns, but becomes a distributed-systems problem at large context and high concurrency.
  • Sparse attention: An attention architecture intended to reduce long-context compute; it requires model-specific inference kernels even when lessons transfer across implementations.
  • Parallel Kernel Bench: Together's benchmark of unresolved kernel-optimization problems, explicitly intended to stimulate optimizations that can be moved into production inference.
  • Self-evolution: MiniMax's described loop of using models to improve internal development, then converting closely related work into evaluations and subsequent training signals.

Operator Notes / Why Ken Should Care

  • Separate agent traffic from chat traffic in your telemetry and routing policy; begin with context length, repeated-prefix ratio, tool-turn count, cache-hit rate, and per-session inference cost.
  • Require every long-running agent benchmark to include intermediate checkpoints, isolated execution environments, and an independent validator specifically designed to detect objective gaming.
  • Build or select evaluation tasks from actual internal workflows, but maintain a held-out external suite so model self-improvement does not optimize only for your current harness.
  • Set a recurring open-versus-closed model rebenchmark cadence that measures task success, visual/UI reliability, long-context latency, and total serving cost—not aggregate leaderboard position.
  • For million-token or persistent-context designs, make a deliberate KV-cache architecture decision covering cache keys, ownership, tiered storage, movement across workers, expiration, and failure behavior.

Source/Metadata

  • Title: Agents at Scale: Inside MiniMax's Model and the Infrastructure Behind It — Olive Song
  • Transcript words: 4876
  • Duration seconds: 1214
  • Timestamp note: No usable timestamps or chapters were present in the supplied transcript; the latter portion of the transcript is substantially duplicated.

Transcript

3599 words en Processed in 220.1s

This is a discussion that I'm particularly excited about because the field is moving so fast and we have two people that have this unique vantage point on the field. I want to start off with intros, talk about your role, what you're thinking about, and what you're working on. Maybe Dan, if you can go first. Hey everyone, I'm Dan. I'm the VP of kernels at Together AI. I lead inference, GPU optimization, and trying to figure out how to use GPUs most effectively to serve AI models. One of the things that I wanted to dive in with Dan about is new model drops. What is everything that goes on behind the scenes to serve it so that everybody here, there's a lot of builders here, can use it? Olive, I want to throw it over to you to talk about your role and what you're focusing on. I'm Olive, and I'm the research leader of RL at Minimax, and I'm responsible for the final training of the model and the shipping of the model. So everything before the inference, right? Okay. Awesome. So maybe I wanted to start off, this panel is focusing on open source. I wanted to start off with this is your strongest model yet, Minimax M3. Why open source it? What's the idea behind that as a company as you're releasing these models? We do believe that the open source community as a whole is very strong and powerful. While we open source the model, everyone can use it. So it aligns with our mission that we want to have intelligence with everyone and also different people. So we want to know how different developers can contribute to the model through feedback, through their own PRs, and we can build the models even stronger. And also, for example, then you will be able to optimize on our open weight model and make its inference faster and then serve better for everyone. Yeah. Yeah. We're big believers in open source at Together. Yeah. Yeah. Yeah. I think when we've been following you guys for a while, from way older Minimax models, seeing M3 and seeing how far it's come is really impressive and really great. So I wanted to pick on this a little bit more. Can you explain, so we've got the model creators themselves, Minimax, we've got experts on the inference side of things. How did this partnership come to be? So they launch an open source model and we're now distributing it. I checked this morning. We have the lion's share of token usage for Minimax M3. How does this partnership come to be, and how do we serve a model like this at scale? Yeah. Great question. So at Together, I think one of the things that we're really interested in is how do you make intelligence abundant. So how do you get more tokens for more people to do more useful things and get all these capabilities into more people's hands? So we follow all the open models very closely. I don't remember when exactly we started part—oh, actually, I think I do know this. We had a car event in Las Vegas sometime last year, and someone from Minimax came. He was like, guys, you really got to serve our next model. It's going to be really, really great. So I think from there we started talking. We were serving Minimax 2.5 and I think 2.7 for a while. And then leading up to the launch of M3, we were quite excited about it. We were seeing the usage and what people were doing with it. It was really quite exciting. And so from there, that's really where we partner. We start working on the model, the architecture, optimizing it, figuring out what's the best way to serve inference on it, and all those great pieces. Yeah. I wanted to actually get into more on the model side of things. As the creator of a model, as somebody who's post-training this thing, the model lands and all the builders that are here start using it. From your perspective, what are the unique capabilities that you love to see people use it for? And what are maybe some of the hidden gems that you thought people would love to build, but you haven't seen as much? Could you shed more light on that? So Minimax M3, which was different from the M2 series, was that it was actually multimodal. So it not only understands text and it not only writes code, it also understands videos and images. So we did see a lot of applications on multimodal agents, which is very cool. And I would say there are a couple that we can highlight, right? For example, computer users: the model is able to navigate through a computer and then do some pretty good creations with the tools that they can utilize. And also you can develop games with the model. It's very fun. I think that's one of the hidden gems, that we actually worked on game development. So the model can help you develop really cool games. Yeah. Yeah. When the blog dropped and then the paper dropped, one of the things I noticed was that you guys highlighted SVG bench, you guys highlighted kernel bench, and you also touched on OSWorld. Can you talk more about what it takes to post-train, especially for those particular domains? I would say a very important thing is the data and how we define the problems. And it could be very different for different tasks. For example, let's say the kernel one, right? It would be very important to design the environments of the data so that we can deliberately train reinforcement learning in those very complex environments and let the model optimize the kernels themselves and iteratively improve the performance. And one aspect that I wanted to talk to you about on the kernel development side of things, where are you seeing open models when it comes to kernel development? You recently released a benchmark specifically for this, so I was wondering if you could talk on that a little bit. Yeah. Yeah. It's a great question. So I think we're seeing all sorts of models, the closed frontier models and the open models, get increasingly better at writing kernels. So we use models all the time when we are developing kernels and writing the optimization frameworks. I think the interesting thing that we are starting to look at is this benchmark that we recently released called Parallel Kernel Bench. It actually has a bunch of unsolved problems in it. So we went around, surveyed all the different ways they can serve model inference, and one of the interesting things that we found is that there's a lot of things that we can think of that would actually speed models up that there don't exist good kernels for. So one of the reasons that we put that benchmark out was, one thing that people worry about is bench maxing or overfitting to particular benchmarks. One of our intentions with this benchmark was if you overfit to it, that's great because we'll go take those kernels and use them to accelerate the inference and the development. Yeah. This is a really interesting point. A lot of people have problems with bench maxing. But the way I think about it is if researchers like you put all the really useful benchmarks out and we bench max on all of them and everything is in distribution, then that's a perfect world, right? That's a very useful model that we can then use. Yeah. Okay, cool. So I wanted to touch on the inference side of things now as well. So a new model drops like this. What does it take? Could you take me behind the scenes at the inference stack? And what does it take to go from day-zero launch and then optimizing it week over week, month over month? Yeah. Yeah. Great question. So when we partner with someone like Minimax, we will get some early model details. So for M3, for example, there are things like the Minimax sparse attention and some of those choices that were a little bit different from any model that's what we're doing. And I think if you look at any of the open models now, they're all quite different from each other in different ways. So there's different attention, different MOE choices, differences in quantization, and all these pieces. So as soon as we get those details, we start writing kernels, benchmarking, figuring out, is there existing kernels that work for it? Do we need to modify something? Do we need to write something from scratch? And then day zero, we're trying to think about things like quality. So when this model launches, is it going to have the quality that we all expect? Are we going to be able to provide the right user experience? And then from there, as soon as it launches on day zero, we have a long list of things that we know: hey, we have to do this with the kv cache, we have to do this with the attention kernels, we can look at this part of the quantization, and things like that. So we have that list, and then we start working on it and start optimizing over the course of weeks so that when you use these models, they actually get faster between day zero and day seven and day 14 and et cetera. I was just talking to Ingrid actually yesterday, and I asked her, have we been improving the performance of M3? Do we need to write something from scratch? And then, day zero, we're trying to think about things like quality. So when this model launches, is it going to have the quality that we all expect? Are we going to be able to provide the right user experience? And then from there, as soon as it launches, that day zero, we have a long list of things that we know: hey, we have to do this with the canv cache, we have to do this with the attention kernels, we have to look at this part of the quantization, and things like that. So we have that list, and then we start working on it and start optimizing over the course of weeks so that when you use these models, they actually get faster between day zero and day seven and day 14, and et cetera. I was just talking to Ingrid yesterday, and I asked her, have we been improving the performance of M3? And I meant over the last month, and she said, oh, did you mean from last night? And this is the pace at which these guys work. So it's very real. One aspect that I wanted to touch on with this is we're seeing the workloads shift. We're going from predominantly chat workloads, where you have turns coming in now, to agentic workloads, where you've got this thing sitting inside a harness and you're doing hundreds and hundreds of multi-turn tool calls. Does that change the way you build the inference stack? Yeah, it definitely does. So these agentic churn-based workloads go into everything from informing your KV cache, your prompting, your pieces like this. So it informs what part of the stack you want to optimize because now, I think when we're in the chat world, you have a system prompt of a few thousand, and then you just have the chat logs. So it's a lot of things now with the coding-based agentic workflows. You upload your whole code base to the model. And that's a very different optimization and routing and kernel challenge than just the chat-based workload. So, yeah, we follow these workloads very closely. It's really interesting to see how they evolve and how to adapt the inference stack and the inference engines to really serve them well. Not only do you have agentic workloads, but you've also got multimodal workloads in there. So what I like to do often with these coding agents is get them to optimize a web app and then get it to use it and then do a feedback loop. So one thing that I wanted to come to you all for is Minimax M3 is multimodal, M2.7, all the ones before that were not multimodal. And can you talk a little bit about the optimizations and how you trained it for that aspect? And then also, I want to get into the architecture of it afterwards as well. Right. Definitely. So what's different from before was that it was trained multimodal from scratch. So from step zero, we trained not only text data, we also trained image data. And it was normal for many other labs that the model would collapse after training a little bit. And we managed to solve that problem. And what we found was actually that with this kind of training from scratch, if you look at the attention map, the text tokens would attend to the visual tokens so that they are naturally combined together to naturally understand each other. So, for example, we are developing websites, right? It is better if we train with both modalities. Also, for example, you can look at the website, you can understand how it looks, and then better optimize for it. For example, during reinforcement learning. So, yeah, I think that is pretty cool. So one thing that stuck out with this model for me was the fact that it introduced a lot of new things. The multimodality, the increase of context to 1 million, the fact that you have sparse attention now. Right. So, if you go to the inference side, it is almost a nightmare, isn't it? You get this new model, and there are so many things that you could optimize to speed up inference. Practically, what are the things that you focus on? There are 1,000 things that you could optimize, but where do you get the most bang for your buck? I mean, you focus on 1,001 things. You just go and you keep doing it. You find every edge that you can, and you go and you push on it. So, yeah, I think there is no stone that you leave unturned, and you just keep going at it. If someone tells me you cannot do the 1,000 first thing, I don't know, try harder. Yeah. And then up a little later. Go ahead. So, the other thing that I wanted to ask is there is a whole zoo of open source models. As you are talking about speeding up inference, are there lessons that you can take from one model and apply it to Minimax M3? Or do you have to restart from scratch as you are thinking about the inference engine, the kernels? How does that work? Right. Yeah. So, there are definitely things that you learn from optimizing one model that you take to another. So, I think sparse attentions are something that have become quite popular now. So, the Minimax sparse attention is a little bit different from the Deep Seek and those and the ones that you find in GLM. But there are still similar lessons that you can take from that optimization process and that kernel writing process that you can then bring to the new sparse attentions. And we have been, in some form or another, I have been thinking about this problem for many years. So, going all the way back to my PhD. So, it is great to see some validation that folks can now train it at scale and people are using it. And it is going pretty well. Yeah. Yeah. One of the interesting things, especially about open source, is you have got all these labs that are learning from each other. Mm-hmm. Taking the wins from each other. Right? So, if one lab figures out that Minimax does this really well, then that becomes the golden standard. One of the things on model launch, in the blog post that you guys go into, is that this model was actually able to replicate a 12-hour run where it could reproduce an iClear paper. And so, somebody who is training this model to do this thing, how do you actually go about that? Because that seems like a pretty ludicrous task. Right. So, letting the model do cool stuff like replicating papers, optimizing kernel frameworks, and stuff like that is always exciting for us researchers because it's very related to our job. Yeah. But training it can be very tricky because it's very long horizon. And the task itself would require GPUs. It has hardware constraints. So, it's very interesting to train tasks like that. And I would say the key there is still the environment and the data and how you formulate the problem, how you formulate the rewards, how you formulate the environment, and how you change the reinforcement learning algorithm a little bit so that it's trained more efficiently, so that you can see cool things emerging through the iterations of RL runs. Mm-hmm. Can you talk, maybe if I keep pulling on the thread a little bit, how do you do evaluation over these longer, longer-scale runs? So, if you want the thing to do a 12-hour task, yes, it might or might not do it at the end, but are there intermediate things that you can also look at? Yes, we do. For these tasks, there are iterations, right? So, the model can submit several times, and we would evaluate each of them. Okay. Some of the times, the models would hack, and we do validation and test for it to test if it's really improving on the performance or it's hacking. And also, we design our internal evaluations. So, for example, for the release of M2.7, we touched a bit on self-evolution, right? So, we're actively using the model to improve the speed of development internally, which, out of it, we can build our own evaluations that are closely related to our own work that we can evaluate the models on. Yeah. So, when we start talking about these long horizon tasks that are 12 hours long, we gave an entire workshop on this on Monday, but what I wanted to come to you, Dan, for is KVCache. So, let's say you have concurrent requests that are 500 to 1,000,000 context length long. How do you deal with the KVCache that just keeps on growing, and how does the infrastructure deal with that? Yeah. So, there's a lot of different pieces that you put there. In some sense, it's like recreating a distributed file system. So, we're in some sense building something like that, or a very big database. It's pretty simple in theory. It's like the type of thing that you should have done in your third year of undergrad or something like that. But most of us actually skipped that class, so now we're rediscovering it live in industry. But it's all about where do you store that cache? How do you know? Have you seen this before? So, let's say you have concurrent requests that are 500 to 1,000,000 context length long. How do you deal with the KVCache that just keeps on growing, and how does the infrastructure deal with that? Yeah. So, there's a lot of different pieces that you put there. In some sense, it's recreating a distributed file system. So, we're, in some sense, building something like that, or a very big database. It's pretty simple in theory. It's the type of thing that you should have done in your third year of undergrad or something like that. But most of us actually skipped that class, so now we're rediscovering it live in industry. But it's all about where do you store that cache? How do you know? Have you seen this before? How do you fetch it? How do you send it from one place to another? So, yeah, it's not that complicated, but you do have to make sure that you do a good job. Yeah. One thing I noticed, you gave a lecture at Stanford recently, and one thing that stood out was that if you fast forward two, three years, two, three years is a long time in AI. And if you look back, you said that we'll realize how early we are right now. Yeah. So, from your vantage point three years out, what do you think we'll look back on and be like, why were we doing it this way? Great question. Some things that I hope for. So, I think we underutilize our GPUs a lot right now. SpaceX said they would have 10% flop utilization or something like that. I hope in three years, well, they should already be embarrassed by it, but I hope in three years they're extra embarrassed by it. So, certainly training should be pretty good. I think at inference we can do a lot better with the hardware that we're using, that we have today. So, I hope in a few years we'll have seen the light on some of those pieces. And I think there will be a lot more models. There will be a lot better. I hope finally by then we've put to bed this question about the open models. There's every few months, there's someone like, oh, Anthropic, OpenAI, they're so ahead, yada, yada. But I think we're seeing with models like M3 and GLM and Kimmy and all those models that the open source frontier really can catch up. And it's not even that far behind. So, I think that's quite exciting. Yeah. I wanted to throw that same question over to you, Olive. But you mentioned that for the M2 series and the M3 series, you're using this idea of self-evolution, where the model is building its own harness and then it's training inside of that. And then you get the next checkpoint. If you look three years out, and then you say, what in RL or post-training do you think made the biggest difference? What do you think that is from this vantage point? Great question. But three years ago, I was still in school. I actually didn't start in this industry yet. So, I wouldn't have imagined what's happening right now today. So, it's really exciting. But I can see how models that were developed were already improving the speed of development maybe a year ago, or even further than a year ago. So, I could see how this speed is actually accelerating, how the development is accelerating. And that's how open-weight models can really catch up with frontier labs. And, yeah, that's how we think we are more mission to bring this model to everyone so that everyone can use it. Awesome. Thank you, guys. Thank you, Dan. Thank you, Olive. Thank you, guys, so much. Have a great day. Very cool. Thanks so much. Thank you. Thank you. But training it can be very tricky because it's very long horizon. And, like, the task itself would require GPUs. It has hardware constraints. So, it's very interesting to train tasks like that. And I would say the key there is still the environment and the data and how you formulate the problem, how you formulate the rewards, how you formulate the environment, and how you change the reinforcement learning algorithm a little bit so that it's trained to more efficiently so that you can see cool things emerging through the iterations of RL runs. Mm-hmm. Can you talk, maybe if I keep pulling on the thread a little bit, how do you do evaluation over these longer, longer-scale runs? So, if you want the thing to do a 12-hour task, yes, it might or might not do it at the end, but are there, like, intermediate things that you can also look at? Yes, we do. For these tasks, there are iterations, right? So, the model can submit several times, and we would evaluate each of them. Okay. Some of the times, the models would hack, and we do, like, validation and test for it to test if it's really improving on the performance or it's hacking. And also, we design our internal evaluations. So, for example, for the release of M2.7, we touched a bit on self-evolution, right? So, we're actively using the model to improve the speed of development internally, which, like, out of it, we can build our own evaluations that are closely related to our own work that we can evaluate the models on. Yeah. So, when we start talking about these long horizon tasks that are 12 hours long, we gave an entire workshop on this on Monday, but what I wanted to come to you, Dan, for is KVCache. So, let's say you have concurrent requests that are 500 to 1,000,000 context length long. How do you deal with the KVCache that just keeps on growing, and how does the infrastructure deal with that? Yeah. So, there's a lot of different pieces that you put there. Like, in some sense, it's like recreating a distributed file system. So, we're in some sense building something like that or a very big database. It's pretty simple in theory. It's like the type of thing that you should have done in your third year of undergrad or something like that. But most of us actually skipped that class, so now we're rediscovering it live in industry. But it's all about where do you store that cache? How do you know? Have you seen this before? How do you fetch it? How do you send it from one place to another? So, yeah, it's not that complicated, but you do have to make sure that you do a good job. Yeah. One thing I noticed, you gave a lecture at Stanford recently, and one thing that stood out was that if you fast forward, like, two, three years, two, three years is a long time in AI. And if you look back, you said that we'll realize that how early we are right now. Yeah. So, from your vantage point, three years out, what do you think we'll look back on and be like, why were we doing it this way? Great question. Some things that I hope for. So, I think we underutilize our GPUs a lot right now. You know, SpaceX said they would have, like, 10% flop utilization or something like that. I hope in three years, well, they should already be embarrassed to buy it, but I hope in three years they're extra embarrassed by it. So, certainly training should be pretty good. I think at inference we can do a lot better with the hardware that we're using that we have today. So, I hope in a few years we'll say, we'll have seen the light on some of those pieces. And I think there will be a lot more models. There will be a lot better. I hope finally by then we've put to bed this question about the open models. You know, there's every few months there's someone like, oh, Anthropic, open AI, they're so ahead, yada, yada. But I think we're saying with models like M3 and GLM and Kimmy and all those models that the open source frontier really can't catch up. And it's not even that far behind. So, I think that's quite exciting. Yeah. I wanted to throw that same question over to you, Olive. But you mentioned that for the M2 series and the M3 series, you're using this idea of self-evolution where the model is building its own harness and then it's training inside of that. And then you get the next checkpoint. If you look back three years out and then you say, like, what in RL or post training do you think made the biggest difference? What do you think that is from this vantage point? Great question. But three years ago I was still in school. I actually didn't start this industry yet. So, I wouldn't have imagined what's happening right now today. So, it's really exciting. But I can see how models that were developed were already improving the speed of development maybe a year ago or even further than a year ago. So, I could see how this speed is actually accelerating, how the development is accelerating. And that's how, like, open-weight models can really catch up with frontier labs. And, yeah, that's how we think we are more mission to bring this model to everyone so that everyone can use it. Awesome. Thank you guys. Thank you, Dan. Thank you, Olive. Thank you guys so much. Have a great day. Very cool. Thanks so much. Thank you. Thank you.