Adaption Labs: Gradient-Free Continual Learning — Sara Hooker, Adaption
Description
Fewer than five thousand people in the world know how to train a frontier model at scale, by Sara Hooker's estimate, and that knowledge travels like an apprenticeship rather than a literature. Modern computer science is 77 years old, two generations, and in that time the route to contributing at the frontier narrowed into one funnel: the right PhD, the right industry lab, the right problem at the right moment. She calls it the unreasonably narrow path, and notes it got compounded in this field by compute, so that a handful of labs build what everyone else uses and whole regions of the world appear nowhere on the map of where breakthroughs happen. Her case that the funnel is about to widen rests on two things. AutoScientist automates the training of models, optimizing the whole loop together from data through alignment and evolving itself per domain, and it beats research staff partly because people carry priors about particular architectures while the search ranges across sizes and across dense and mixture of experts designs. It only started paying off once they controlled data quality alongside the model rather than leaving that to the agent. One nice piece of honesty: the win rates all sit just above 60% because the budget was set to stop there, and lifting that ceiling let them keep climbing. The second reason is her slow death of scaling argument, that pretraining size is no longer the most rewarding axis. That matters for access, because pretraining compute has to be colocated and enormous while the compute that now pays off is distributable. If no lab is going to quadruple model size again on this architecture, recipes and algorithms start to matter more than hoarded GPUs. Speaker info: - https://x.com/sarahookr - https://www.linkedin.com/in/sararosehooker/ - https://www.sarahooker.me/ Timestamps: 0:00 - Seventy seven years of computer science 1:15 - From gentleman scientists to professional labs 1:56 - The unreasonably narrow path 3:13 - GPU poor and GPU r
Summary
Generated by gpt-5.6-terraAt-a-Glance
- Verdict: Watch fully
- Core thesis: Sara Hooker argues that frontier AI is moving away from pretraining-scale dominance toward automated, domain-specific continual adaptation, making high-quality data, post-training, inference-time compute, and optimization recipes more important—and potentially more accessible—than owning the largest GPU cluster.
- Why it matters: This is directly relevant to agent and AI-operations strategy: durable advantage may come less from training a giant general model and more from closed-loop systems that adapt models, data, evaluations, and test-time compute to proprietary workflows.
- Best use: Use it as a strategic framing for evaluating continual-learning infrastructure and model-customization vendors, then pressure-test Adaption's claims around AutoScientist, data co-optimization, safety controls, and the actual economics of supplied compute.
Executive Summary
Hooker frames modern frontier AI as an unusually concentrated system: a narrow professional pipeline and enormous pretraining-compute requirements have limited participation to a small set of labs. Her central proposition is that this concentration is weakening because the highest-return levers for model performance are shifting from ever-larger pretraining runs toward better post-training, task-specific adaptation, agentic/inference-time compute, and algorithmic search.
Adaption's product thesis is AutoScientist, an automated model-training system that co-optimizes data, alignment, model choices, and other training decisions for a specific domain. Hooker says it can outperform research staff because it searches across architectures—including dense and mixture-of-experts models—and changes multiple hyperparameters more aggressively than human researchers typically do. Its reported win rates were initially capped at 60% because the system was configured to stop after reaching that threshold; after removing the stopping rule, performance continued to rise.
The technically consequential point is not merely automated experimentation: Hooker says AutoScientist only produced meaningful returns when data quality and model adaptation were optimized together. She further characterizes this as a long-horizon problem involving choices about what knowledge belongs in model parameters versus non-parametric/external systems, with the training harness and model needing to be co-optimized.
For operators, the opportunity is domain-controlled intelligence rather than one universal model shipped to everyone. Hooker highlights medical, science, legal, code, 242-language coverage, and non-verifiable everyday tasks as areas where last-mile adaptation is especially valuable. However, the talk is primarily a strategic thesis and product presentation, not an independently validated technical evaluation; its claims should be treated as promising directionality rather than settled proof.
Key Takeaways
- Claim: The frontier-AI bottleneck is shifting away from brute-force pretraining scale and toward optimization over a broader set of levers, including post-training, inference-time compute, data quality, and training recipes. | Evidence: Hooker cites her paper, "Slow Death of Scaling," argues that pretraining size is no longer the most lucrative scaling axis, and points to the OpenLLM leaderboard trend in which models under 13B increasingly outperform larger models. | Implication: Ken should prioritize systems that can improve capability through routing, evaluations, data flywheels, post-training, and adaptive inference rather than assuming the largest base model is automatically the strategic moat. | Caveat: She does not claim model size no longer matters: frontier models remain large, distillation remains useful, and a new architecture could reset the relevant scaling ceiling.
- Claim: Automating model-development decisions can accelerate domain innovation because human frontier-training expertise is scarce and manual researchers underexplore the configuration space. | Evidence: Hooker estimates that fewer than 5,000 people worldwide know how to train frontier models at scale. She says AutoScientist tests multiple model architectures, model sizes, dense models, and mixture-of-experts models, while changing more hyperparameters simultaneously than cautious human workflows usually permit. | Implication: The emerging control point may be automated experimentation and training orchestration—not just model weights—particularly for teams with strong domain data but limited ML-research capacity. | Caveat: The claimed performance advantage over research staff is presented by the product builder without benchmark methodology, task definitions, or external replication in the talk.
- Claim: Data must be co-optimized with the model for automated adaptation to generate strong returns; treating data generation or selection as a separate agent is insufficient. | Evidence: Hooker says Adaption did not see adequate performance gains from auto-research approaches that merely let an agent decide whether or how to create data. AutoScientist instead co-optimizes data adaptations alongside model decisions. | Implication: For Ken's agent systems, a model-improvement loop should jointly manage task traces, feedback labels, synthetic data, retrieval/context design, evaluation harnesses, and model updates; optimizing only prompts or fine-tunes will likely leave material performance on the table. | Caveat: The talk does not specify the data-governance, provenance, evaluation, or contamination controls required to safely automate this loop.
- Claim: The target architecture is continuously adaptive intelligence: models should interact with their environment, learn from that interaction, and allocate test-time compute according to task difficulty. | Evidence: Hooker contrasts the prior monolithic-model workflow—one team builds, another serves, another handles front end—with an interactive intelligence model. She says Adaption plans to extend adaptation so that test-time compute is task-adaptive rather than spent uniformly. | Implication: Treat agent deployment as an ongoing learning system with telemetry, feedback capture, regression suites, approval gates, rollback, and per-task compute budgets—not as a one-time model-selection decision. | Caveat: Continual adaptation can introduce regressions, reward hacking, privacy leakage, and evaluation drift unless updates are gated by robust offline and online controls.
- Claim: Domain-specific customization has its highest immediate demand where general models are weakest or where last-mile accuracy has high value. | Evidence: Four weeks after its beta announcement, Hooker says enthusiasm was strongest in medical and science, followed by legal and code. She also identifies non-verifiable tasks as a major near-term opportunity and says Adaption supports 242 languages from day one. | Implication: Focus customization investments on workflows with proprietary feedback loops, meaningful error costs, and domain-specific success criteria rather than generic chat use cases where foundation models are already adequate. | Caveat: Medical, legal, and other high-stakes domains require stronger validation, auditability, liability management, and human review than the talk addresses.
- Claim: Lowering the cost of experimentation changes which questions get pursued, potentially broadening who can contribute to frontier AI. | Evidence: Hooker argues that the cost of asking a research or product question determines what gets asked; reducing that cost increases the number and variety of experiments. Adaption's beta reportedly includes free GPU access to remove the compute hurdle. | Implication: Evaluate vendors on whether they actually compress the full experiment cycle—data preparation, training, evaluation, deployment, and monitoring—not merely whether they subsidize initial GPU usage. | Caveat: Free beta GPU availability is a go-to-market offer rather than proof that production-scale customization will be inexpensive, durable, or broadly available.
Detailed Brief
Strategic argument: the frontier is becoming less centralized
- Claims: Hooker argues that AI research has been filtered through an "unreasonably narrow path": the right PhD program, industry lab, internships, and research credentials were historically prerequisites for frontier contribution.; She sees compute concentration as a second compounding barrier: a small number of recognizable frontier labs have controlled the infrastructure needed to build broadly deployed models.; Her normative position is that builders should own and adapt intelligence locally or privately rather than receive the same static model as billions of other users.
- Evidence: She references Stanford's mapping of statistically significant AI breakthroughs as evidence that whole regions of the world have been excluded from frontier participation.; She contrasts uniform model deployment with the intuitive reality that task difficulty varies substantially, making equal compute allocation inefficient.
- Caveats: The talk treats decentralization primarily as a capability and access opportunity; it offers little operational detail on governance across jurisdictions, data rights, misuse prevention, or model-update accountability.
- Implications: A credible AI platform strategy should give customers controlled adaptation pathways while preserving policy enforcement, observability, provenance, and rollback.; The competitive question is increasingly whether a team can create a superior domain learning loop, not simply whether it can access a frontier API.
Safety and open-access position
- Claims: Hooker rejects a binary framing of openness versus safety: expanding access to model-customization tools creates real risk, but safety concerns can also be used to preserve concentrated participation.; She distinguishes AutoScientist's customization objective from the separate question of whether the resulting models are open source.
- Evidence: In audience Q&A, she explicitly acknowledges that making tools readily available changes their risk profile, while arguing that the access-versus-safety balance must be navigated rather than treated as an absolute.
- Caveats: No concrete safeguards were presented for dangerous-domain screening, authorization, model access controls, audit logging, red teaming, or safeguards against harmful fine-tuning.
- Implications: If considering adaptive-training infrastructure, require explicit answers on tenant isolation, data ownership, authorization boundaries, policy enforcement before training and deployment, audit trails, and rollback mechanisms.
Notable Concepts & Terms
- AutoScientist: Adaption's beta product concept for automating the model-training loop across data, alignment, model configuration, and domain-specific self-improvement.
- Gradient-free continual learning: The title's framing for adaptation without relying solely on conventional gradient-based end-to-end retraining; the transcript emphasizes automated search and iterative optimization more than implementation details.
- Co-optimization of data and model: Hooker's central technical lesson: performance gains appeared only when data quality/adaptation and model choices were optimized jointly.
- Parametric versus non-parametric space: The design trade-off between storing knowledge in model weights and keeping it external, such as in data, retrieval, or other non-parametric mechanisms; Hooker presents balancing these as a core long-horizon learning problem.
- Test-time compute: Compute allocated during inference; Hooker argues it should vary by task difficulty rather than be applied uniformly.
- Post-training: The phase after base pretraining where adaptation, alignment, and task/domain optimization may now offer higher returns than simply increasing model parameter count.
- Distillation: Using larger-model knowledge to train smaller models; Hooker acknowledges it remains valuable even while arguing that pretraining scale has reached diminishing returns under current architectures.
- Unreasonably narrow path: Roseanne Lu's phrase, used by Hooker to describe the credentialed, compute-intensive route historically required to work at the AI frontier.
Operator Notes / Why Ken Should Care
- Define a gated continual-improvement loop for deployed agents: capture task traces and outcomes, generate candidate data/model changes, run domain-specific regression and safety evaluations, then use approval and rollback before production release.
- When evaluating Adaption or comparable vendors, request task-level benchmark methodology, baseline definitions, data requirements, compute costs after beta, reproducibility evidence, and failure/regression rates—not only headline win rates.
- Prioritize pilots in domains where Ken controls proprietary feedback and can define measurable evaluations; avoid starting with high-stakes medical or legal workflows unless audit, human review, and liability controls are already in place.
- Add adaptive inference budgets to agent orchestration: use cheaper/faster paths for routine work and selectively escalate model, tools, retrieval depth, or reasoning budget for ambiguous and high-value cases.
- Require an explicit safety and tenancy design before adopting automated fine-tuning/customization: permissions, data provenance, policy checks, isolated artifacts, audit logs, evaluation records, and reversible deployments.
Source/Metadata
- Title: Adaption Labs: Gradient-Free Continual Learning — Sara Hooker, Adaption
- Transcript words: 4389
- Duration seconds: 1250
- Timestamp note: No timestamps or chapters were provided. The transcript contains duplicated portions of the latter Q&A, which were treated as repetition rather than additional content.
Transcript
It is so lovely to be here. So I wanted to share today some thoughts that I have around who gets to be at the frontier of discovery. So modern computer science as a field has only existed for the last 77 years. It's bizarre when you think about it. So World War II, all the transistor technology that was developed for radio, we finally had our first versions of the computer. But when you think about it, that's only two generations of people working on these tools. However, within that time, even for computer science, who and what and what topics we work on has dramatically changed. And I think it's an interesting setting because actually, if you look back across science as a whole, how we do discovery has been markedly different at different points in time. So when we started, the whole idea of a researcher was what we call a gentleman scientist, typically someone with adequate wealth to dabble in discovery. And these were all individuals, independent researchers. When we see the first associations emerge with the rural society in the 1600s, the idea of being a scientist as a full-time job was very special. And these became the predominant spaces for discovery. Why am I talking about this? Because, to be honest, the professionalization of science led to what I recall, and my good friend Roseanne Lu calls, the unreasonably narrow path. So I'm an AI researcher. This is Jan's actual career. And it's interesting, if you were to be an AI researcher at the forefront, you had to follow this exact narrow path. You had to get into the right PhD program. You had to then go to the right industry lab. You had to do sufficiently interesting work. And then, finally, you got to contribute to the frontier. This was my story too. So a lot of my work has been at efficiency at scale. I did my PhD and worked at DeepMind and a lot of different frontier labs. But this was, in many ways, a very aggressively filtered system. If you did not make it or you were not curious about the right problem at the right time, you didn't have a place to play at a frontier lab. And this is the standard successful scientist. You have a famous advisor, hopefully. You hopefully get one or two important internships. And what's interesting about this is computer science was really about representing the world. And it was all the tools that do that. But the reason why most people go into computer science was the question at the end of it. If you can represent the world, what questions can you answer? So today I'm going to talk about what I think is one of the most profound and interesting topics, which is why this is so important for computer science, that this is changing. So I'll also speak to this. It was double compounded in computer science because of the need for compute. So this resulted in jokes. It's fun that Merve was here. This is her tweet about GPU poor versus GPU rich. It led to barriers of entry on who can contribute frontier AI. So I put here company A, company B, company C. But to be honest, if we polled, I think there would be significant majority votes about who those companies are. But basically, a handful of frontier labs have been able to build the technology we use. It's also determined who gets to participate in breakthroughs and who doesn't. So this is a map of what Stanford calls where statistically significant breakthroughs have come through. And you can see whole sections of the world are completely left out. And so, for me, this is a very important question worth answering. Who gets to shape the frontier? Who gets to answer the questions at the end of the pursuit? We've seen that the shift has dramatically changed from academia to industry. And it's also meant that we ship the same model to everyone. Why that's particularly interesting is that most people intuitively understand that you shouldn't ship the same model to billions of people. And they also understand that it's not a particularly good use of compute, right? You're spending the same amount of compute on everything. And some problems are hard and some are very easy. So where does that leave us? What's my talk for today? I would like to say that we are ripe for a revolution. And we are ripe for a revolution in who gets to participate at the frontier of AI. I'll tell you two reasons why I'm bullish on this. I'll definitely cover one, and then I'll actually see interest and timing because I want to leave plenty of time for questions. I think that's, unless you have a few questions and a bit of banter, these things can be boring. So we'll see. But I'll cover one definitely that I'm actively thinking about. And this is: what if we could allow anyone to build the same frontier intelligence as that in labs? And I've been in a few labs. I've done my tour of duty. And this is the core question I care about now: how do you build intelligence that continuously adapts and that builders everywhere can have more control? So instead of taking years of training to learn how to build the tools, scientists just skip to the questions. So a few weeks ago, we released Autoscientist. And Autoscientist is really about how do you automate the training of models itself. I'll share a few things that are really interesting about this. One, it's co-optimized the entire loop. So it's from data to alignment, and it chooses and self-evolves based upon the domain and the type of data. What's interesting as well is that it actually outperforms research staff, mainly because a lot of our research staff has experience with certain model types. And we're testing it across many different model architectures, different size models, as well as dense and mixture of experts. And that search base is a lot broader. And so exploiting it using how do you self-improve from experience and scale is very effective. I think this is very interesting. This only worked when we co-optimized the data. So there's a lot of auto research projects right now, which basically treat data as the agent. It decides whether to create data or not or what to do. Frankly, we did not get the returns for how much you can squeeze out of performance until you control for data quality. So we actually co-optimized, based on all the adaptation we did with the data, exactly what we would do with the model. And that was super interesting. It speaks to the need to control the entire flow. What's fun about this is really what Autoscientist is doing is it combines all the knowledge it gained from the adaptive data component with also the knowledge of the domain and also the ability to self-improve for a domain and to learn from other components of that domain. What is a cheeky fact, and this is quite fun, you'll notice all these percentages for win rates are 60 plus. And that's because we put the budget stopping it above 60. So once it was above 60, our agentic flow could exit. But we've since removed that barrier, and you can see it just go up over time, which is super fascinating. And then I think what's interesting is it changes a lot of the hyperparameters. Typically, humans are much more wary about changing all at once, and so you get massive exploitation of the search space. I see this as crucial for how do you reduce the amount of compute you use for customization because you train with much more predictability, but also how do you leverage your domain knowledge to really unlock how you build frontier AI. And this was fun. We announced a beta four weeks ago. The excitement is most acute for medical and science. And that's largely, I think, legal and code as well. But these are domains where typically current models fall short and also domains where, in many ways, the degree of last-mile customization is really acute. And this is the core point. I think this is fun. I think I might even have time to cover the other point I want to make about why now is very important for changing who shapes. But really, the main factor that this does is increase your innovation cycle and also increase the likelihood that when you train and spend compute, you'll succeed. And those combined factors are super interesting. One thing that we're doing next is extending that so even your test-time compute should be adaptive based on your task. So this kind of brings me back to where I started and the grumpy statement I said, which is we have this verified super narrow compounding issue of barriers to entry. One is that you need to do this very narrow funnel of who gets to build frontier AI, and the other is that typically compute and cost really dominate. We want to change that. We decided, okay, we're going to cover languages from day one, 242 languages. And also a big interest for us is actually non-verifiable tasks. I think this is super interesting because this is really the bulk of everyday tasks that people do. And it's really where the meat of what is interesting for progress is going to be over the next year. And this leads me into our mandate. We care deeply about how do you accelerate learning in a way that models should be able to learn from their environment. So right now we've moved from an era of the model is monolithic. When I was at different parts of my research career, basically your whole team would be around building a model. One is that you need to do this very narrow funnel of who gets to build Frontier AI, and the other is that typically compute and cost really dominate. We want to change that. We decided, okay, we're going to cover languages from day one: 242 languages. And also a big interest for us is actually non-verifiable tasks. I think this is super interesting because this is really the bulk of everyday tasks that people do. And it's really where the meat of what is interesting for progress is going to be over the next year. And this leads me into our mandate. We care deeply about how do you accelerate learning in a way that models should be able to learn from their environment. So right now we've moved from an era of the model being monolithic. When I was at different parts of my research career, your whole team would be around building a model. You give it to someone else to serve, and you have someone else do the front end. And actually now, the most important intelligence is a model that interacts. And so this idea of how efficiently are you going to interact, how will you continuously learn from the environment, is pretty core. And I think about it a lot. So let's see. I think I do have time, right? How are we doing for time? Oh, I do. I have plenty. This is lovely. So we'll have time for questions, and I'll share a little bit about what I think the next component is. I think core to this: if we just did auto scientists, but it still took enormous compute to do Frontier AI trainings, I think we'd be in a bit of a pickle, right? I'd be saying, oh, great, you can use this agent, but don't worry, just bring your 10,000 GPUs with you. But I think there's another trend which makes this very important timing, and rooms like this probably much more optimistic than they have been a few years ago about who can build Frontier AI. And one of that is the rules of where you get rate of return for compute are totally changing. So I wrote a paper about this called Slow Death of Scaling. But empirically, we do now know that pre-training size in particular is not your most lucrative axis of scale. And what does this mean? If pre-training scale isn't going to dominate performance, it actually really greatly changes who can create the best recipes for innovation. Because pre-training compute typically has to be co-located. It has to be, in many ways, large volume to accommodate for redundancy. Inference compute and other places where you actually apply compute, typically you can have much more distributed. It's also much higher return given the amount of flops. And so it's interesting when we talk about what is the state of pre-training compute. We know it's not giving the same returns, largely because our architecture is saturated. So we see much smaller models outperforming much larger ones. This is the OpenLLM leaderboard. And this is the daily submission of the best small model under 13B versus all the larger models. And you can see over time that ratio totally flips. And also there's the grumpy assessment that most recent models that have severely played with just increasing model size haven't provided the same stepwise change as their predecessors. And a lot of that is because where the most returns for performance are now are on a broader action space. And this is really what I was getting at when we move from an algorithm to expanding optimization space in new places. And what's fun about that is that these are new places where the barriers to entry are much more nimble and where recipe and algorithm and research matter again. And things like how do you automate that discovery. And so this is what I'll state, and I think then we should open up for questions. And I would encourage good grumpy questions or fun positions. Let's make use of the time. I know I was told earlier that almost no talks have time for questions. I find that so disappointing. So we'll need some brave people to start the conversation. But I will say this means all bets are off. And I would say it's a very good time to be working on intelligence because instead of just a handful of people getting to create it, it's much more now about the question you want to answer. At the end of the day, the reason why people did a computer science PhD was to learn the tools to get to the question. And now you can just get to the question, which is super meaningful. Okay, let me open up. Where should we start? We have an abundance. I hear there's no microphone. So if you want to ask a question, you want to make a statement, I will indulge a statement if it's interesting. Yeah, go for it. Just raise your hand and I'll repeat it afterwards. Let me just get to the end of this in case people want to reach me afterwards. Nice. Yes, go ahead. Gentleman in the fourth row, go for it. You mentioned, yeah, I'm here looking for a market price. Can you point to how? So I think how is twofold. One is there's very few people who know how to train frontier models. I would say realistically probably less than 5,000 in the world at scale. I think that type of knowledge, that's a very exploitable search space. And actually as humans, all those configurations, we're not particularly good at. It's kind of like secret knowledge we pass as if we're apprentices. So that's one. Once you automate a lot of that knowledge, you just accelerate innovation cycles, which means that you can explore and do more questions. Typically what people often miss is that the cost of asking something informs what is asked. And if you make it cheaper to ask something, you change the volume of things that are asked, which is super interesting. The other reason, though, I do think it's very much a facet of the changing nature of compute. So agentic compute, post-training compute, matters a significant amount for performance. That does not require the same type of, dare I say, hoarding of GPUs. But I think it's very different compute purchasing dynamics. And again, it means that the person with the best idea has a higher chance of winning, which is fun. Nice. What else? Who wants to go? I see, yeah, we can go up here. And then I saw a hand back there. Okay. Yes, I do see you. The glare is high, but you go first and then we'll come up here. So once you're allowed to figure out a lot of that thinking, one of the challenges is to adapt the whole model. Somebody will take your base model and adapt those to be something. How do you see that? Yeah. So the question, I'll just repeat it, because I think there's probably people in the room who want to know. So the question was, one of the, I guess, counterpoints from some frontier labs about not enabling frontier AI outside is a safety question. So I think it would be, I definitely am not one of those people who says that open source doesn't carry any risk. So when you make a tool more readily available, there's a profile of risk associated with it. Autoscientists is, to be fair, about enabling people to customize their models. You can think of that as a slightly different question from whether those are open source. It's giving people way more control, whether that's local or private or within their company. It's about how do they own their own intelligence? What do I think broadly about the impact of open source on safety? The dynamic has often conflated that real risk of wider access with a slight sense that it restrains who can actually participate. And I think that's a delicate balance. And I think you have to acknowledge risk while also navigating that and acknowledging that it limits who can participate. Yeah, so nuanced answer. So I guess I should be more bombastic on that one. But I guess I have been in this discussion a few times, and I find the binary views on other sides miss a lot. But anyways, okay, go ahead. Are there any specific research ideas or technologies that predict the paradigm of the type of learning? Oh, I think for automating and speeding up learning, one of the core questions is how do you balance what you store in the parametric space and the non-parametric space? And actually, one of the most interesting things, I mentioned that this only worked because we co-optimize data and model. It will only work to do an auto scientist for harnesses if you also co-optimize it with a model. And so it's interesting. It's actually a long-horizon problem. And that's super fascinating to think about, where you're optimizing the choices for each and co-training, which is cool. Nice. I think we have time for maybe two more, and then we can pass on to the next speaker. Nice. Go ahead. So you talked a little bit about this, actually working on the full-training side of the model is cheaper than the first thing, obviously. But still, especially talking about really large models, reinforcement learning, even Oh, I think for automating and speeding up learning, one of the core questions is: how do you balance what you store in the parametric space and the non-parametric space? And actually, one of the most interesting things, I mentioned that this only worked because we co-optimize data and model. It will only work to do an auto scientist for harnesses if you also co-optimize it with a model. And so it's interesting. It's actually a long horizon problem. And that's super fascinating to think about, where you're optimizing the choices for each and co-training, which is cool. Nice. I think we have time for maybe two more, and then we can pass on to the next speaker. Nice. Go ahead. So you talked a little bit about this, actually working on the full-training side of the model, it's cheaper than the first thing, obviously. But still, especially talking really large models, reinforcement learning, even fine-tunes, it's pretty difficult, so you talked a little bit about the smaller models. But I still think most frontier smaller models still rely on the bigger knowledge, like the smaller models, like the smaller models. Yeah, actually that's an excellent point. I think the question amounts to two points. One, are larger models necessary for distillation benefits? And then second, frontier models are still pretty large. So I think for the second one, frontier models are still pretty large. Yes. I don't think I'm arguing that. My argument is slightly different. And my argument is that no frontier AI lab is going to 4X the size of that model again for pre-training. So it's almost like we know we're at an upper ceiling for this architecture. If someone comes out with a new architecture, that's totally different. The architecture determines your ceiling. And I'm saying we are probably at the ceiling of size, which means that that's fun, because it means, okay, it's what you innovate within that. So size does matter. I think that's a very good point to bring up. Meaning I'm not advocating everyone uses 0.8B, but I am saying that we now have a more equal playing field at the top. Second point is interesting, distillation, the impact, certainly. So data quality in general means you use capacity a lot more. So what you will see in pre-training is instead of size, people are just moving post-training further back, which is very fascinating and a bigger lever. So I agree distillation is helpful. It's just that, again, we've hit the ceiling. And so it's almost like no one is going to supersize their model. Or if they do, it's not clear it's beneficial except for a small size of the distribution, which is very much the long tail. And that's interesting, where that tradeoff is worth that much pre-training compute. So very good question. One more, and then I think we are done. Yes, go ahead. Can you add the old installers? Can you add the GPUs to the data? Yeah, it's actually in beta. So you can, I shared here, you can try it in beta. So we actually are offering the GPUs for free. Okay, oh, that's a nice question. I promise I don't know this gentleman. But yes, I think actually we're trying to remove the compute hurdle, and I think it's quite cool to see. So feel free to take a look at the beta. Nice, lovely, thank you so much. Really nice, thank you. Thank you. Nice. I think we have time for maybe two more and then we can pass on to the next speaker. Nice. Go ahead. So you talked a little bit about this, actually working on the full-training side of the model, it's like cheaper than the first thing, obviously. But like, still, especially talking like really large models, like reinforcement learning, even 5-tunes, it's pretty, like, difficult, so like you talked a little bit about the smaller models. But I still think like most, like, frontier smaller models still rely on, like, the bigger knowledge, like, the smaller models, like, the smaller models. Yeah, actually that's an excellent point. I think the question amounts to two points. One, are larger models necessary for distillation benefits? And then second, so frontier models are still pretty large. So I think for the second one, frontier models are still pretty large. Yes. I don't think I'm arguing that you, my argument is slightly different. And my argument is that no frontier AI lab is going to 4X the size of that model again for pre-training. So it's almost like we know we're at an upper ceiling, for this architecture. If someone comes out with a new architecture, that's totally different. You can, the architecture determines your ceiling. And I'm saying we are probably at the ceiling of size, which means that that's fun, because it means, okay, it's what you innovate within that. So size does matter. I think that's a very good point to bring up. Meaning I'm not advocating everyone uses 0.8B, but I am saying that we now have a more equal playing field at the top. Second point is interesting, distillation, like, the impact, certainly. So data quality in general means you use capacity a lot more. So what you will see in pre-training is instead of size, people are just moving post-training further back, which is very fascinating and a bigger lever. So I agree distillation is helpful. It's just that, again, we've hit the ceiling. And so it's almost like no one is going to supersize their model. Or if they do, it's not clear it's beneficial except for a small size of the distribution, which is very much the long tail. And that's kind of interesting, like, where that tradeoff is worth that much pre-training compute. So very good question. One more and then I think we are done. Yes, go ahead. Can you add the old installers? Can you add the GPUs to the data? Yeah, it's actually in beta. So you can, I shared here, you can try it in beta. So we actually are offering the GPUs for free. Okay, oh, that's a nice question. I promise I don't know this gentleman. But yes, I think actually we're trying to remove the compute hurdle and I think it's quite cool to see. So feel free to take a look at the beta. Nice, lovely, thank you so much. Really nice, thank you. Thank you.