AI Engineer

Why AI Agents Need Million-Token Context — Thomas Wolf & Olive Song, MiniMax

1788 summary words 8 min summary Watch video

Start with the signal

8 min read

Summary

At-a-Glance

  • Verdict: Skim
  • Core thesis: MiniMax argues that capable agents require million-token context, native multimodality, and efficient sparse-attention architectures because real agent work accumulates long tool traces and unstructured visual information.
  • Why it matters: The interview offers a useful model-design rationale for agent systems: context length is not merely a retrieval benchmark feature, but a practical constraint on multi-step execution, tool use, video/report understanding, and longer-horizon automation.
  • Best use: Use it to pressure-test context-window requirements and multimodal model selection for agent workflows; skim rather than watch for implementation specifics, which are limited.

Executive Summary

Olive Song presents MiniMax M3 as a relatively compact mixture-of-experts model—about 400 billion total parameters with 20 billion activated—that combines coding ability, vision/video understanding, and a functional one-million-token context window. Her central claim is that these capabilities need to be designed together because future AI applications will be agentic, multimodal, and dependent on long-lived interaction histories.

The strongest technical argument concerns long context. MiniMax had previously demonstrated up to 10 million tokens for largely passive tasks such as reviewing a book, but Song says agentic operation creates a more compelling need: tool results, environmental observations, and multiple interaction rounds rapidly exhaust short context windows. MiniMax Sparse Attention is described as a two-stage mechanism: an index branch identifies relevant regions, then a sparse-attention branch computes over selected blocks.

Song also argues for native multimodal pretraining rather than attaching vision adapters after text pretraining. In MiniMax's account, late-stage multimodal training can degrade text performance, lead to weaker visual convergence, and produce recipes that do not transfer reliably across model sizes or data mixtures. Training text, images, and video jointly from the start required solving training-collapse issues through data work, including interleaved source data, cleaning/masking, and reward modeling.

Operationally, MiniMax describes an unusually open internal research process in which employees propose model improvements, recruit collaborators, run experiments for weeks or months, and ship successful work into training. The closing strategic signal is that Song sees model routing and multi-agent systems—not just bigger single models—as a particularly important near-term application layer.

Key Takeaways

  • Claim: Million-token context is positioned as a practical requirement for complex agents, not simply as a headline capability. | Evidence: Song contrasts earlier MiniMax models that could handle up to 10 million tokens for passive tasks such as reviewing a dumped book with M3's one-million-token context aimed at agents that must retain environmental state, repeated tool responses, and multiple rounds of interaction. | Implication: Ken should treat long-context capacity as especially valuable for long-horizon workflows with accumulating tool traces, but evaluate effective context utilization rather than selecting models on nominal window size alone. | Caveat: The interview does not provide task-level benchmarks showing that one million tokens improves real production-agent success rates relative to retrieval, summarization, or external memory approaches.
  • Claim: MiniMax Sparse Attention is the architectural lever intended to make very long context computationally viable. | Evidence: Song describes an index branch that identifies what matters at a higher level and a sparse-attention branch that performs computation on selected context blocks; Thomas Wolf links this approach to M3's efficiency and low cost. | Implication: For agent infrastructure, sparse or selective attention is a meaningful design category to monitor because it could reduce the cost of keeping raw histories available instead of aggressively compressing them. | Caveat: Neither speaker gives latency, throughput, context-retrieval accuracy, or cost-per-token measurements, so the claimed efficiency cannot be compared from this interview with other long-context implementations.
  • Claim: Native multimodal pretraining is presented as more scalable and stable than adding visual capabilities after text pretraining. | Evidence: Song says that training vision adapters after text pretraining can harm text performance and yield poorer vision convergence, while starting multimodal training midway through pretraining is highly sensitive to architecture, data mix, and learning rate. | Implication: When selecting models for agents that must inspect decks, reports, images, interfaces, or video, prioritize evidence that vision is deeply integrated rather than assuming a nominal multimodal label guarantees robust cross-modal reasoning. | Caveat: The claim reflects MiniMax's internal findings; the discussion supplies no comparative benchmark results against adapter-based or continued-pretraining approaches.
  • Claim: The combination of long context and multimodality expands the range of agent tasks beyond text-centric coding. | Evidence: Song cites reading PowerPoints and poorly structured reports, understanding long videos, and then taking tool-based actions; Wolf frames the extreme case as an agent watching a YouTube tutorial and learning how to use a coding tool. | Implication: Ken can consider multimodal agents for operating procedures embedded in video, slideware, screenshots, and semi-structured documents—domains where text extraction alone loses essential information. | Caveat: These are proposed use cases rather than demonstrated end-to-end agent deployments in the interview.
  • Claim: MiniMax uses its own agentic research harnesses to automate model-development workflows and accelerate iteration. | Evidence: Song says the company has automated much of its research workflow with internal harnesses, citing model-assisted post-training, synthetic/automated data creation, and longer-horizon optimization tasks; she says M3 is already helping build M3.1. | Implication: The relevant pattern is not autonomous research in the abstract, but building a bounded harness around recurring evaluation, data, and optimization loops where agent outputs can be tested before promotion. | Caveat: The interview does not identify which research steps remain human-gated or quantify the quality, speed, or safety gains from these harnesses.
  • Claim: MiniMax sees multi-agent systems and model routing as a key next layer for handling more complex tasks. | Evidence: When asked what is most exciting in the coming months, Song specifically names multi-agents and model routing as mechanisms that unlock more capability and reveal model strengths and limitations. | Implication: Ken should separate the model-layer question—long context and multimodality—from the orchestration-layer question of which specialized model or agent should handle each subtask and when escalation is justified. | Caveat: No routing policy, agent topology, handoff protocol, or evaluation method is provided.

Detailed Brief

M3 positioning and expected scaling path

  • Claims: M3 is positioned as an open model that is unusually competitive while also supporting multimodal inputs.; MiniMax expects to increase model scale substantially because some tasks remain inaccessible to smaller parameter counts.; The team frames one-million-token context as an architectural foundation that can be extended toward substantially larger windows.
  • Evidence: Song describes M3 as roughly 400 billion total parameters and 20 billion activated parameters; Wolf later refers to approximately 428 billion total and 23 billion active.; Song says MiniMax's prior M1 and 01 models could operate on 10 million-token contexts in non-agentic use cases.; She says future ultra-long context will require coordinated advances in both architecture and hardware.
  • Caveats: The parameter figures are stated inconsistently in the conversation, and no independent evaluation or serving profile is discussed.; A stated intention to scale beyond a trillion parameters is a roadmap signal, not a released-product commitment.
  • Implications: M3 should be assessed as a potentially efficient open-model option for multimodal long-context experimentation, rather than as proof that maximum context alone solves persistent memory and planning problems.; The major strategic competition may shift from raw parameter scale toward the economics of serving large, multimodal context windows.

Research and open-source feedback loop

  • Claims: MiniMax allows individuals to identify weaknesses after a release, propose an improvement project, and attract other contributors internally.; The company views external open-source feedback, bug reports, and pull requests as inputs to future model versions.; MiniMax's applications are intended as a distribution and experience layer for its models, not as the company's original raison d'être.
  • Evidence: Song says projects can run for weeks or months, with architecture work requiring extended research, experiments, and repeated pretraining evaluation.; She asks users to report multimodal failures and request features such as adjustable thinking effort.; Song says MiniMax apps have reached more than 300 million people in roughly 200 countries and more than one million companies.
  • Caveats: The adoption numbers are company claims and are not broken down by active use, product, geography, or paid usage.; Community feedback can uncover edge cases but does not substitute for systematic safety, reliability, and benchmark evaluation.
  • Implications: If Ken evaluates or deploys an open model, issue reporting and reproducible failure cases can be a practical channel for influencing upstream capability priorities.; A product surface can generate valuable real-world feedback, but its usage scale should not be interpreted automatically as model quality or enterprise readiness.

Notable Concepts & Terms

  • MiniMax M3: MiniMax's discussed open model, positioned around coding, native multimodality, and one-million-token context.
  • Mixture of Experts (MoE): M3 is described with far more total than active parameters, indicating only a subset of parameters is activated per inference and potentially improving serving economics.
  • MiniMax Sparse Attention: A selective-attention design in which an indexing mechanism identifies relevant context regions before sparse computation is performed on selected blocks.
  • Native multimodality: Joint training on text, images, and video from the beginning of pretraining rather than adding visual adapters later; MiniMax claims this improves scalability and avoids degrading text ability.
  • Interleaved data: Natural multimodal data in which images and videos remain embedded among text rather than being removed or masked out wholesale; cited as part of MiniMax's training approach.
  • Research harness: Internal automation tooling that combines models with repeatable research workflows such as data generation, evaluation, and optimization.
  • Model routing: Selecting among models or specialized agents for different subtasks; Song identifies it as a major enabling layer for complex AI applications.
  • Thinking effort: A requested feature referring to controllable reasoning or compute intensity, mentioned as an example of community feedback MiniMax may incorporate.

Operator Notes / Why Ken Should Care

  • Run a representative long-horizon agent evaluation that compares raw long-context retention against a retrieval-plus-summary memory stack; measure task completion, context misses, latency, and total cost rather than context-window size.
  • Add multimodal test cases to agent acceptance criteria: slide decks, screenshots, scanned or poorly structured reports, and long instructional video followed by a concrete tool action.
  • When evaluating M3 or similar models, demand empirical tests for needle retrieval, cross-modal grounding, tool-trace retention, and effective throughput at the actual context lengths the workflow requires.
  • Design research or operational agents as harnessed loops with explicit evaluators and promotion gates; do not infer safe autonomous iteration merely from claims of internal workflow automation.
  • Treat model routing as a control-plane design problem: define routing triggers, escalation thresholds, shared-state boundaries, and observability before adding multiple agents.

Source/Metadata

  • Title: Why AI Agents Need Million-Token Context — Thomas Wolf & Olive Song, MiniMax
  • Transcript words: 3167
  • Duration seconds: 1248
  • Timestamp note: No usable timestamps or chapter markers were present in the supplied transcript.
Full transcript 3146 words · 14 min read
0:00

Music Joining us on stage is the co-founder and chief science officer at Hugging Face, Thomas Wolfe.

0:21

Music Hello, everyone. Hello, Olive. Nice to have you on stage. Hi. Nice to be here. Thanks for having me. So I think you're in for a treat today because you just saw GLM, which is current number two on the artificial analysis. I take that out because nobody can use it. And now we have number four. So you will have all the top models, at least the top open source model, in a row. And we are very lucky to have Olive, who has a pretty amazing path in life. So she came to the US, Pennsylvania. She was studying during her PhD at NYU in the lab of Yann LeCun, working on JPA. But we decided we won't talk about JPA today, right? Something for another day.

1:19

And then instead of joining Hugging Face, which was in New York also at the time, she decided to go join Minimax. So for those who maybe don't know all the neo labs around the world, and you're forgiven because I think there's like 64 neo labs right now, Minimax is one of the top of what we call the AI dragons in China. So these are the new, there's DeepSeek, which is very well known now, Moonshot, who does Kimi, Z.ai, and GLM that you just saw. And now we have Minimax. They're all extremely good, extremely talented teams fighting for the first spot. So the latest release of Minimax was M3, just earlier in June, which was the top model at that time, top open source model.

2:09

Very impressive. There's a lot of very interesting things about this model, so we'll quickly dive into them. And then talk a little bit about what's specific about Minimax, what's great there. So maybe Olive, to start a little bit, can you give us a little bit of your view of M3, what you like about this model, how is the release? Yeah, we released M3 earlier this month, and it is a smaller model with around 400 billion total parameters and 20 billion activated. But it is very capable in terms of both coding performance, and also it understands vision.

2:52

So that's what open source models don't usually have, that the model can not only deal with coding, but it can also understand videos, images, and it has a super long context of 1 million.

3:13

So we really put these three things together because we know that they will be very important in future AI applications: coding capabilities, agentic capabilities, longer context, and multimodal understanding. Yeah, I think that would be very interesting about the model. Yeah, so there's a lot to unpack in this model, and it's still, I think, the only top five open source model that is actually multimodal, so we need to talk about that.

3:41

But maybe first about the long context, because it was also the first one that really had this real 1 million token long context that's actually functional, and you guys had also the Minimax sparse attention, which is this one technique to make that efficient, that you also published and shared extensively. So can you talk a little bit about this, maybe how the project went from the attention, how to make this long context? Yeah, I would say the story about long context went back to either Minimax M1 and Minimax 01, where the model was actually able to perform tasks of 10 million token context.

3:56

Ten million. Ten million, yes. But then it was not an agentic model, right? It was just, for example, dumping a book, it would be able to give reviews on it, stuff like that. So what we realized was that longer context actually unlocks a lot of capabilities, especially when interacting with users. And now when the agent is interacting with the whole environment and getting all the tool responses, getting multiple rounds, the shorter context wouldn't be enough to perform the complex tasks. So for this version we said, oh, we have to have our longer context back.

4:18

And so what we pursued was with our Minimax sparse attention, which was the architecture that was scalable and had a simple design. So I would say from a higher level, it has an index branch that selects on a higher level what matters more in the context. And then we have a sparse attention branch that performs the calculation on the selected blocks to actually perform the tasks. And so, yeah, like that we really designed an elegant architecture so that we can scale the length and then scale the model size in the future with that.

4:46

That's beautiful. I like how, for those who've been in the field for quite some time, we had a lot of work on attention, right? This n-square, and there was a lot of linear attention. Yeah. And then somehow all of this disappeared at some point when flash attention came around. We discovered we just needed more efficient kernels. And now I like how we come back to thinking, first principle, what is attention? How can we make that more efficient? So 1 million token is crazy, right? GPT-2 was 1,024 and everyone was like, oh, that's really big. We never need more.

5:31

Where do you see this going in the future? Jeff was pitching me the other day a trillion token attention. Do you think we should go to a trillion token attention? That's definitely something we can explore toward, right? Ultra length of the context, definitely. That's something that's very exciting to explore, and something that architecture design along with hardware would require a lot of research on. Yeah. You think there's still a lot of low-hanging fruit? So typically today we saw OpenAI really reducing, I mean, we don't know how confirmed, but reducing their inference bill by half by probably having some more efficient processing around attention.

6:09

You think there is still a lot of low-hanging fruit in how we can process that? So one thing is still very interesting about M3 is how cheap it is in particular because of this sparse attention, or in part because it's a small one, but it's also very efficient. Right. You think we can go even way further? Maybe how did you guys invent Minimax sparse attention? Was it an agent coming up with the idea? Was it a human still coming up with the idea? Tell us a little bit about this.

6:38

Yeah. So we do think there's still a lot of work that can get into architecture and inference optimization so that the model can be more efficient, especially if there are tasks that are very task-sensitive but require very strong capabilities, right? And for those kinds of tasks we really want the model to be efficient. And who came up with this part? Actually, I think an intern from our team worked on that. Me too. Yeah, an intern. That doesn't usually happen in a lot of labs because I think in some labs, interns don't have access to the work and stuff.

7:00

But yeah, we are open to anyone who would like to contribute to our model. So the architecture was actually designed by an intern. That's very good. Still some work for interns here. Good news. That's also a good segue to how Minimax is working internally. So we were discussing before coming on stage, they were saying everyone can propose a project. Can you tell us a little bit about how you are organized, how you do research? I think that is very different from even in school or even in earlier tech companies. It's pretty different.

7:23

It's that what we make sure is that we have a good foundation and good infrastructure so that anyone can play with the model and can think of what they can improve with the model. And then after model releases, when they are free, they can play with the model. They can think of their own evaluations. They can find their own weaknesses and propose a thing that they want to improve on the model. And then other people who are interested in that would propose to join the project. And they will work on it for a couple of weeks or even a couple of months. And when they work it out, the final thing is shipped to our model.

7:52

We use that in our final training and it's shipped out to the audience. Interesting. So you can have people working for a really long time on a project. When you say a couple of months, it can be a really deep exploration of what's possible. Yes. I would say, for example, architecture might require a longer time of investigation, research, experiments, even redoing the evaluations for pre-training. Yes. So it might require longer time. That's very nice. Yeah. And I know you're also very big on evaluation. I agree. We could talk about that. So I think one thing probably related to that is this unique specificity that M3 and your team has around multimodality.

8:29

So not just text, but this model can also understand image and video. And as I understand, but please explain better, when you read the model card on Hugging Face, it says the model was trained from the first step as multimodal, not just like using one after the source, right? Yes. Can you tell us a little bit more about that and why you think it's important, and why starting from the first step on multimodal training and not just training this? So we call it native multimodality. And so it is somehow typical for model labs to train the multimodal, let's say, vision understanding capabilities after the text pre-training is done. They put adapters and then train that part.

9:16

But what we found out was that that would actually harm the text performance. And the vision understanding performance wouldn't converge that well because the model kind of converges toward the text understanding. And it's just not the most optimal and also not the most scalable. If you think about it, we want to scale the data, right? And also some labs train this capability from halfway through the pre-training, for example, continued pre-training. But what we found is that this would be very recipe-sensitive. The recipe would be different for different architectures, different data mixtures, different learning rates.

9:58

It's hard to control, hard to scale. You can't really scale your experiment results and conclusions to a larger model. And so what we thought was, why not just train from the very first step? That comes the most natural. We know that a lot of labs run into problems doing that. The model would collapse after a couple of steps of training, both text and vision understanding. But we managed to solve that problem. We did a lot of work on the IT and we did a lot of work on the data that we actually train. For example, we do interleaved data, what we call interleaved data. It's actually natural data, but we keep the images and videos in instead of masking them out.

10:28

And we do some pretty good cleaning and masking on the data, and we do very good reward modeling so that we train it from the first step and scale up a lot. So yeah, it does not collapse. That's really impressive. Should we expect a much larger model in the future? So this one is still fairly small, right? It's 428 billion parameters, 23 active billion. Well, do you think you will go past the trillion? Definitely. Yeah, definitely in the future. There are many tasks the model wouldn't be able to perform or go at with smaller parameters. We are definitely going more ambitious than this. That's great. Looking forward.

11:20

Another interesting thing I always find fascinating about Minimax is how you also have this whole range of apps and products, right? So I remember already, so Minimax started to open source things on the Hugging Face platform in January last year. So that was 18 months ago. And we were chatting a little bit about the team to understand what you were doing. And I remember you were already having huge usage on some of these apps. Can you tell us a little bit how this started, right? So was it that you had a lot of apps and then you thought, we have all this data, why not train a model? And then they built up a research team. How is the story there?

12:26

Our start is model from the first day. So I believe that multimodality model, a model that can understand all vision and output all modalities, was the first thing that our CEO planned on the first day, even before the company even started. So that was the dream of AGI. I think that was very, very early, even before ChatGPT came out. Wow. Yeah. And then apps were something that came along. Because you have some model capabilities, you want people to experience it well. Not many people can use it with API, right? We can't expect everyone to experience it with API.

13:25

So we need good user interaction interfaces, good apps, good scenarios that people can experience the model with. I think actually those apps covered more than 300 million people around 200 countries globally. And I think over a million companies as well. Yeah. This was mind-blowing when I heard about the size. And we don't often realize the size of this type of usage already. And that kind of brings me to the question around open source business model and all of that, which is always an existing question, which is right now it's nice to open source a model, but you also need to have some revenue stream, right?

13:55

And so I guess M3 is something you decided, for instance, to give for free. And I think it's great for the world. How do you see this? Do you also have some specific models you use for the app? Do you think in the future you'll keep, it's probably hard to say for sure, but do you think you'll keep open sourcing models? How is the culture around open sourcing right now? Personally, and also for the model research team, we always hope to open source the models. That is our plan because we really see how the open source community together can help the model build better.

14:25

For example, we receive a lot of feedback on the model performance from the great community, and we receive PRs on whatever we open source. And those are very valuable when it comes to our later versions. So definitely open sourcing is great. That's great. And actually, do you have some ask for the audience, people who are using M3 or Minimax? Is there something you would love them to send back to you as feedback? For instance, do you read when people try to modify the models or play around with tweaks? What is the best thing you think you can take from the community for future models, for instance?

14:58

I would say whatever issues that people are running into, especially with multimodality, right? This is the first time that we're combining it together. We are definitely going more ambitious on that in the future. It might have some flaws right now, but we are improving on that. So whatever feedback that model is not doing that great, we will definitely improve that in future versions. And also whatever features that people want. Say, for example, thinking effort, right? Some people ask for that. Everyone can ask, and we will try to accomplish that in the future models.

15:41

Yeah. Do you see a lot of users right now already in multimodality in terms of coding agents? I feel like it's a little bit underexplored. It is. But it can actually unlock a lot of capabilities and a lot of agent applications. Say that, for example, you want the model to read the PowerPoints or to read some reports that are not very structured. And you want it to understand a very long video. Say that you've done a long playing video and then you want the model to act using some tools after understanding it. And it unlocks a wide variety of agent use cases.

16:19

So the agent could finally watch my YouTube tutorial and understand how to use my coding tools, how I describe it? Is it something like that? Could the agent finally watch YouTube tutorials and understand things from them? Yeah, yeah, yeah. I think so. Do you use a lot of agent coding tools internally? I mean, coding for sure, but is it also already in terms of research? Is it automated in part? Or how does this evolve? Yes. We have our own research harnesses. We build our own research harnesses that automate our workflows. I would say a lot of our workflows are automated.

17:03

You can see how the latest frontier models all pursue capability-led kernel optimization, right? Let the model post-train other models. Let the model build data, auto-data, stuff like that. You can see how more and more models are capable of doing those, including M3. Actually, we were very good at those cases, longer horizons and kernel organizations. And so we can use that model capability, harness it together, and help with our daily routine and make our iterations even faster. Sure. Is M3 building M4 already?

18:12

Building M3.1. M3.1. Okay. Yes. You had the job. Already. I would love to finish on what you find exciting in the coming months. What do you think? It can be easier in terms of features or things you want to see happening in AI, or more generally in terms of whatever really is top of your mind, I would say, is going to happen. A lot of things are very exciting. But what I recently find the most exciting would be multi-agents. I think a lot of AI applications are using model routing, multi-agents that unlock even more capabilities, even more complex tasks. And also, it tells us what the models are capable and not capable of.

19:10

And you can do a lot of things with that. It's pretty exciting. Thanks a lot, Olive. Pleasure to have you. Thanks for having me. Thanks, everyone. I love you. I love you. I love you. I love you. I love you. I love you. I love you.

20:24

I love you. I love you. Thanks, everyone. I love you. I love you.

20:33

I love you.

20:40

I love you. I love you. I love you. I love you. I love you. I love you.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note