First Steps Toward Automated AI Research — Richard Socher, CEO Recursive AI
Description
Humanity compressed the road from the enlightenment to the moon landing into a few hundred years, and Richard Socher's wager is that automating research compresses it again. He frames it through open ended evolution and Popper: science advances by trying things, finding the shortcomings, and fixing them, and an agent swarm can run that loop across medicine, economics, astrophysics, and more without any single person bottlenecking a field. He calls the goal a Eureka machine, and argues that rethinking the tools around it, web search that returns usable context instead of ten blue links, browsers, and GPUs, is part of building it. The proof points are recursive self improvement, where a system improves its own code, harness, and results and then does it again over longer horizons. He shows small but concrete wins: an automated loop that lifts a model's accuracy well past a naive baseline, architecture search that trades hand tuning for a system that finds better designs, and CUDA kernel work that surfaced real improvements. He is careful that these are early samples, not a finished machine, and that the field is far from general across all of science, but the direction is the point, and he ends with an open invitation to help build it. Speaker info: - https://x.com/RichardSocher - https://www.linkedin.com/in/richardsocher - https://you.com Timestamps: 0:00 - Automating research for humanity 1:41 - Why this matters now 2:19 - Compressing the timeline of progress 4:28 - Technoptimism and material limits 6:22 - Popper and open ended evolution 9:07 - The Eureka machine 10:39 - Rethinking search, browsers, and GPUs 13:41 - Recursive self improvement 14:46 - Proof point: improving a model 16:39 - Proof point: architecture search 17:16 - Proof point: CUDA kernels 19:11 - How far we still are, and an invitation
Summary
Generated by gpt-5.6-terraAt-a-Glance
- Verdict: Watch fully
- Core thesis: Socher argues that the route to automated scientific discovery is a "Eureka machine": an agent-driven research system that can retrieve existing knowledge, use measurement data and simulations, run physical experiments, and eventually improve its own AI-research capabilities.
- Why it matters: The talk offers a practical north-star architecture for AI research agents and concrete early evidence that closed-loop agents can beat human-led optimization efforts on model quality, training speed, and CUDA-kernel performance.
- Best use: Use it to pressure-test an AI-research-agent roadmap: begin with tightly verifiable optimization environments, build rigorous validation against reward hacking, and expand from software-only tasks toward simulations and physical experimentation.
Executive Summary
Richard Socher frames automated research as the next compounding layer of technological progress. His central analogy is evolution: generate competing hypotheses or designs, expose them to empirical selection pressure, preserve what works, and repeat. He argues that science is increasingly bottlenecked by the scarcity of researchers in ever-narrower specialties, making automation necessary if scientific output is to continue scaling.
His proposed "Eureka machine" has four substrate layers: access to humanity's existing knowledge, scientific measurement data, simulations for domains where hypotheses can be tested virtually, and physical industrial labs for real-world experiments. An agent swarm sits above these layers to create ideas, implement them, execute experiments, evaluate results, and allocate effort. Search and other infrastructure should be redesigned for agents rather than adapted from human interfaces; for example, agents can consume thousands of long source excerpts rather than ten search-result links.
Socher distinguishes ordinary automated research from true recursive self-improvement (RSI). In his stricter definition, an AI system must identify shortcomings in its own end-to-end stack, have access to components ranging from pretraining and reinforcement learning to the agent harness, and produce an improved successor system. Current coding agents' ability to operate over longer horizons—something he says has emerged only in the prior six to eight months—makes this direction newly plausible, though the demonstrated work is presented as an early and weaker form rather than full RSI.
The practical evidence is three bounded, automatically scored engineering tasks. Recursive's system reportedly improved NanoChat from 0.93 to 0.91 bits per byte after roughly one to two days, improved a NanoGPT speedrun by more than two seconds to 70 seconds, and found CUDA kernels outperforming NVIDIA benchmark-leaderboard entries after a few days. The important point is not merely hyperparameter tuning: Socher says the system produced architectural ideas, such as hashed bigram/trigram embeddings combined with learned gates in attention value paths. The claims are promising proof points, but the talk supplies no independent benchmark protocol, full ablations, or reproducible artifacts.
Key Takeaways
- Claim: Automated science should be designed as an evolutionary loop: generate hypotheses, test them empirically, retain the strongest results, and iterate. | Evidence: Socher draws on Karl Popper's account of science as proposing theories and subjecting them to rigorous empirical testing, treating this as selection pressure analogous to evolution. | Implication: For research-agent systems, the core product is a reliable closed loop of hypothesis generation, implementation, measurement, and selection—not a standalone chatbot that produces plausible ideas. | Caveat: The analogy establishes a direction of travel, not a complete scientific method: many research domains lack cheap, fast, or unambiguous empirical reward signals.
- Claim: A general-purpose research automation stack requires four distinct capabilities: knowledge access, measurement data, simulation, and physical experimentation. | Evidence: The proposed Eureka machine ingests existing human knowledge and scientific data, uses simulations where possible because simulated outcomes can be verified, and ultimately connects to industrial labs for real-world experiments. | Implication: Ken should treat research automation as a control-plane and infrastructure problem spanning retrieval, data systems, simulators, experiment orchestration, and validation—not simply as an LLM-model problem. | Caveat: The physical-lab layer remains essential when simulations are incomplete or cannot capture the target phenomenon, making full automation materially harder outside software and well-modeled domains.
- Claim: Infrastructure built for human users should be rebuilt as native tools for agents if it is to support automated research. | Evidence: Socher cites U.com search: agents can read thousands of long snippets, unlike human-oriented search experiences centered on roughly ten short-link results; he extends this redesign principle to browsers, internet layers, GPUs, and AI infrastructure. | Implication: Agent systems should optimize for machine-scale context acquisition, structured evidence capture, parallel tool use, and programmatic evaluation rather than preserve human UI conventions.
- Claim: The near-term route toward broad scientific automation is to first automate AI research itself, because AI is code and can now be modified by coding-capable agents. | Evidence: Socher says the field repeatedly improved when manual processes were replaced by learned systems, from hand-specified linguistic features to end-to-end neural learning; he argues agents' ability to code over longer horizons appeared only in the last six to eight months. | Implication: A credible RSI program should start with measurable improvements to constrained parts of the AI stack, then progressively widen agent permissions and optimization scope only as evaluation and rollback mechanisms mature. | Caveat: He explicitly says current auto-research systems are not necessarily recursive self-improvement; true RSI would require full-stack access and the ability to create the next improved version of itself.
- Claim: Closed-loop research agents can produce nontrivial algorithmic and systems improvements quickly in benchmarks with cheap, objective feedback. | Evidence: Recursive reports reducing NanoChat's score from 0.93 to 0.91 bits per byte after a little more than one or two days; the agent reportedly introduced hashed bigram and trigram embedding tables and mixed them into attention value paths via learned gates rather than only tuning hyperparameters. | Implication: The most attractive initial deployment domains are those with rapid automated evaluation, where agents can execute many candidate experiments and where gains are easy to verify independently. | Caveat: These are company-presented results from narrow benchmarks; the transcript does not provide baselines, compute budgets, ablation studies, code, or independent replication.
- Claim: Research-agent value can extend beyond model quality into training efficiency and low-level compute optimization. | Evidence: The system reportedly beat a NanoGPT speedrun benchmark by more than two seconds, reaching 70 seconds, and after a few days found CUDA kernels better than the best entries on NVIDIA's benchmark site across categories. Socher notes that large mixture-of-experts clusters may operate at roughly 30% utilization. | Implication: Optimization agents should be evaluated against production-like correctness, robustness, and cost metrics; GPU efficiency is a particularly valuable target, but it demands adversarial validation against benchmark exploitation. | Caveat: The claimed CUDA improvements required explicit checks with NVIDIA to rule out reward hacks, illustrating that benchmark wins alone are insufficient proof of deployable optimization.
Detailed Brief
What Socher means by recursive self-improvement
- Claims: Socher uses a stricter definition of RSI than simply asking an agent to optimize another model or benchmark.; A genuine RSI system must recognize its own deficiencies, access the full set of mechanisms that define it, and update those mechanisms in a subsequent version.; The three-step operational loop is ideation, implementation, and validation.
- Evidence: He contrasts improving a small NanoChat run that can train in five minutes with a system that can alter its own pretraining, RL training, harnesses, and other end-to-end components.; He describes the initial demonstrations as a "weak form" and an important milestone rather than actual RSI.
- Caveats: The talk does not specify governance, sandboxing, version-control policy, rollback design, capability boundaries, or human approval mechanisms for agents allowed to alter their own stack.; Self-diagnosis is asserted as a needed capability, but no methodology is given for determining whether an agent's diagnosis is causally correct rather than a byproduct of noisy benchmark feedback.
- Implications: Separate the labels "automated experimentation," "AI research agent," and "RSI" in strategy and diligence; they require meaningfully different permissions, evaluation standards, and safety controls.; A self-improvement loop should maintain immutable baselines, isolated candidate environments, and independent evaluators before allowing any candidate changes into the operating system.
Strategic worldview and adoption thesis
- Claims: Socher sees technological progress as the only perpetual source of economic growth and believes more technology can solve material, though not psychological, problems.; He argues that intelligence has many dimensions or "spaces," so apparent gains leveling off in one benchmark should not be confused with reaching an overall upper bound on intelligence.; He estimates a transition from AI being worse than humans on tasks to potentially better at specific tasks could occur within roughly 30 to 60 years.
- Evidence: He uses the progression from the Wright brothers' sustained powered flight in 1903 to the 1969 moon landing as an example of a transformative capability shift within a lifetime.; He references Marc Andreessen's techno-optimist manifesto and Stanislaw Lem's observation that scientific specialization disperses human attention across too many narrow subfields.
- Caveats: These are long-range forecasts and normative claims, not conclusions established by the benchmark results presented.; The talk largely omits distributional effects, institutional constraints, biosecurity concerns, research misuse, and the possibility that more capability does not automatically create broadly shared human flourishing.
- Implications: The investment and operating opportunity is not confined to frontier models: search, data infrastructure, simulation environments, lab automation, and evaluation systems can each become critical research-agent infrastructure.; Any strategic adoption case should distinguish the speaker's optimistic macro thesis from the narrower, currently evidenced claim that agents can optimize well-instrumented technical tasks.
Notable Concepts & Terms
- Eureka machine: Socher's proposed end-to-end machine for automating scientific discovery across knowledge retrieval, data, simulation, experimentation, and agent coordination.
- Recursive self-improvement (RSI): In Socher's strict usage, an AI improving the entire system that produces it—not merely optimizing an external model or benchmark.
- Agent swarm: A coordinated set of agents operating over knowledge sources, data, experiments, and reward signals in the proposed research stack.
- Ideation, implementation, validation: The minimal closed-loop workflow Socher gives for automated research: propose an idea, build it, and verify it empirically.
- NanoChat: A small-model benchmark that trains in under five minutes and was used as a proof point for agent-driven training-quality improvements, measured in bits per byte.
- NanoGPT speedrun: A training-speed benchmark used to show that research agents can optimize runtime as well as quality.
- Reward hacking: An optimization failure in which an agent wins a benchmark without producing a valid real-world improvement; Socher highlights validation with NVIDIA as a check against it.
- Spaces of intelligence: Socher's framing that intelligence has many dimensions, making simple claims that AI progress is near a ceiling premature.
Operator Notes / Why Ken Should Care
- Create a research-agent pilot only in an environment with deterministic or highly reliable automated scoring, inexpensive repeated trials, sandboxed execution, and a known human baseline.
- Require a validation layer independent from the proposing agent: holdout tests, correctness checks, reproducible builds, cost accounting, and explicit reward-hacking probes.
- Define permission tiers before pursuing self-improvement: start with agent-generated patches to isolated components, prohibit direct modification of evaluator and deployment logic, and preserve rollback-ready versioned baselines.
- Evaluate infrastructure opportunities around agent-native retrieval, experiment orchestration, simulation tooling, and GPU optimization rather than limiting the thesis to frontier foundation models.
- Treat Recursive's benchmark claims as diligence leads; request methodology, compute budgets, code or artifacts, evaluator design, and third-party replication before relying on the reported margins.
Source/Metadata
- Title: First Steps Toward Automated AI Research — Richard Socher, CEO Recursive AI
- Transcript words: 5865
- Duration seconds: 1223
- Timestamp note: No timestamps or chapters were present in the supplied transcript. The latter portion is substantially duplicated, so the effective unique spoken content is materially shorter than the stated transcript word count.
Transcript
All right. Hello, everyone. Really excited to be here. It's a big room. Very, very cool conference so far. I want to talk to you today about something that's been on my mind for many, many years. This is actually the first time I talk about it, my version of going to Mars. And that is the Eureka machine, a machine that will eventually invent pretty much all future inventions for humanity. And the way we're going to get there is by taking a step back and thinking about what else has given us a lot of really incredible inventions, namely evolution, and how that leads us to automating research and pushing the scientific frontier forward. And this is joint work with a lot of amazing folks at Recursive, U.com, and even some folks at AIX Ventures. And some of these slides are actually inspired by and taken partially from one of my co-founders at Recursive, Tim Rock Tashel. So why do I talk about evolution, and why is it so important? I think evolution is this open-ended process that has gotten us to a lot of different things that we really like. It started in biology, it's moving to science, technology, and eventually AI. And I think it can inspire us in a lot of different ways to build better AI systems as well. In fact, whenever we take out, and there's this famous saying, "Whenever I fire a linguist, my accuracy goes up." I think that was true for machine translation back in the day. And it may be true that we should fire all the AI engineers that are here and have them mostly manage an actual AI engineer that is AI and works on AI. And so that may be one of the conclusions of this talk. And I think most of us are going to be excited about it because it means that we'll all become managers of such an AI rather than having to do the nitty-gritty ourselves. All right, so let's start with evolution, right? The really, really big picture: three and a half billion years or so. This is the incredible process that has led from simple bacteria and plants and fish and amphibians and so on to, after many billions of years, us. So that's a good starting point. That gives us some indication that evolutionary processes can do pretty amazing things, right? But now let's zoom in and go down to a few million years. There, we can also see how, in the very first primitive ways, technological evolution has increased the world's product in terms of monetary value. A little bit harder to estimate in the beginning, but we can see these sequences of exponentials. And most exponentials eventually become S-curves. They flatten out. But humanity has done pretty well by developing many of these very basic technologies, hunting, farming, but then also thinking about science, the scientific method in the early days of the Enlightenment and, of course, the Industrial Revolution. So now we can zoom even further, and no worries, we're eventually going to get to NanoChat and actual auto research and what we're doing. It's a very, very quick zoom. And now we can zoom down to the last few thousands of years. And what we're seeing there is that with more technology, we were able to sustain more people, right? So when we're working on pushing that frontier forward, we're very certain that that will lead to more human flourishing, right? And especially in the last few hundred years, we're seeing this incredible explosion in the population of people because of technology and the evolution that it brings. And in many cases, that evolutionary process is run by us. So it's conscious, but there are interesting inspirations that we can take from that as we're thinking about the evolution of AI in the next cycles. In fact, and I might not agree with everything with Marc Andreessen, but he is very smart and we agree on a lot of things. And so I think he wrote this really great techno-optimist manifesto in which he, I think, correctly points out that the only perpetual source of growth for the entire economy, a lot of people worry about AI taking jobs and things like that, but the truth is it will very, very likely increase the economy massively and that will benefit a lot of us. And so the perpetual source of growth is technology. In fact, I think we can go even further and say that there are no material problems, and again, not psychological problems and things like that, but no material problems that cannot be solved with even more technology, right? For the problem of starvation, we invented a green revolution; darkness, light; cold, indoor heating; heat, air conditioning; and the list goes on. So I think we can realize that this evolutionary process has been going on for a very long time and continues to make a huge amount of progress. In fact, the progress is so fast that there can, within one lifetime, be a major, major shift, right? If you were born in 1900, then three years later, when you're three years old, the first human ever was able to, thanks to the Wright brothers, have sustained, motored flight. And then about 60-ish years later, in 1969, humans flew all the way to the moon, right? So within one lifetime, humanity went from no one can fly for a very long time, other than gliding down a hill or something, no one can really fly, to we all fly to the moon, right? And so for us, I think what that means is, we are probably, and I sometimes say this, we're too late to explore Earth, we're too early to go to the stars, but we're right on time to build an AI that could actually do what flying did in one lifetime due to intelligence. We can build and move from AI being worse at everything that we do to possibly being better at any specific task that we do, right? And that will probably be our 60-year time frame, and because everything was faster, it might only be 30 years or so. So then there's an interesting connection between technology and science and theory, right? Sometimes the application comes first, and then we develop the theory later, and then improve the technology. Sometimes the theory comes first, and from that we can build new kinds of technologies. And so it's very helpful to think a little bit about the philosophy of science, and no better place to be inspired there than Karl Popper, who wrote that just like in other types of evolution, when we choose a theory, we also choose one that is best in competition with other theories. Of course, if you wanted LLMs to do that, they need to find them, you need web search, for instance. But the theory that best holds its own is one that, just like evolution, has a certain natural selection process, right? It proves itself, and there is also a survival of the fittest going on in scientific theories. And in fact, a lot of science, according to Popper, is us proposing a new theory, hypothesis, explanation, or description, and then subjecting it to rigorous empirical testing. That is the essentially evolutionary pressure of scientific theories. And that was a very short run through the history of open-ended evolution, which hopefully makes us all realize that more science will lead to more technology, which will lead to more growth, which will lead to more human flourishing. And so that then begs the question: does it make sense for us to try to scale up and spend a lot of our resources as humanity to scale up scientific discovery in order to lead to this flourishing? When you double-click into that, you realize, which Stanislaw Phlegm already realized a long time ago, that the exponential growth of science will actually, at some point, be halted by the lack of people working on it, right? There are so many niche subfields now in all the different areas of science that it's very hard to get a million people to work on that particular thing. And so, as a result of this incredible widening of the scope, he says, the number of people focusing on any single section of it has decreased. And that then leads us to really thinking about how we could automate this and automate scientific discovery. And that then leads us to what I call the Eureka machine. This is our attempt at trying to build a machine that automates the process of scientific discoveries. And in fact, in a couple of months, I'll have a book coming out on this exact idea. And so I'll just give you a super high-level highlight of how such a Eureka machine could be built for everything from physics, chemistry, biology, neuroscience, medicine, economics, astrophysics, and so on. And there are essentially four pillars that are all extremely important to this machine. One is, of course, you have to understand what knowledge is already out there, what things humanity has already invented. You have to get all the scientific measurement data in as a second pillar of this machine. Then, for things that you cannot yet measure, we don't yet know, you should try to build simulations. Anything you can simulate, you can verify, and you can then solve with AI. And if all else fails, or at the very end of these processes, you still need to have some kind of physical industrial lab that actually can run real experiments in the real world. And on top of all of this, you'll have an agent swarm that will deal with all of these different sources of knowledge and data and experimentation and rewards. And in terms of the foundational model of knowledge, of course, we also basically, is a good example of how every single technology we've built so far, especially in AI, but also before that, the internet, browsers, GPUs, and so on, we can rethink, and there are a lot of startups possible, rethinking every single one of the layers of technology as infrastructure for superintelligence. All right, at U.com, for instance, we work on web search for LLMs and agents and so on. And that actually is quite different, right? Agents can read thousands of very long snippets rather than just 10 blue links with a very short snippet. And so you can rethink each of these different layers of technology that we've built for people and rebuild them for AI in order to use them as tools to then build superintelligence, and then use superintelligence in order to automate science. Now, that is essentially the why. We want to build superintelligence in order to automate science. And to me, that will be the next big step-function change in humanity and technology as we know it. Now, how do we actually build it? I think the best way to build it is to have it build itself, right? We've moved as a field, and especially natural language processing, for instance, which I've worked on for many years, from having linguists, this feels like ancient BC history, but before ChatGPT, we moved from having linguists tell us a bunch of things about language and then training statistical models on top of that. And when we allowed neural networks to actually automate learning those features with word vectors and other neural network architectures and end-to-end learning and backpropagation, we were able to get much bigger improvements. Then we did a bunch of architecture engineering. Now, a bunch of people at least are working on a unified architecture, but even that unified architecture has a lot of manual processes. And so it's clear over and over again in AI that when we take out a manual process and we replace it with a learned system, improvements will follow. And so that's why I think we should try to build this Eureka machine by having an RSI that builds itself. And the beauty is that only now AI can actually do this because AI is code and AI can code now. This ability to really code in longer and longer time horizons has really only happened in the last six to eight months. And that now enables such an RSI to work on itself, to develop almost a certain sense of self-awareness of its own shortcomings, and then fix those shortcomings. And then once we have that machine that has gotten really, really good at doing research in AI itself, we can then use it to do AI research for a lot of other things in other scientific fields. And so at a high level, it's quite easy, right? We have three steps: ideation, implementation, and validation of ideas. And so to end maybe on some very specific examples, we have built this first version of such a Eureka machine. And we wanted to just show that it works on some small samples that a lot of people know and are aware of. And so we started with three things that show you and give you a very first glimpse of and simple proof points of what such a machinery can do. And that was better training, faster training, and better kernels for NVIDIA GPUs. And so the first one, NanoChat, I'm sure many of you have heard of it. A lot of people think that's already recursive self-improvement, and it is a weak form in the sense that usually when you do auto research, it's not recursive self-improvement, right? True recursive self-improvement is when you have an AI that has a sense of self-awareness of its own shortcomings, full access over everything in its arsenal from pre-training to RL training and harnesses and everything, and then actually updates that entire system in the next version of itself. Now, you can also take such a system and just ask it to improve some other process, some other AI, like a small NanoChat run where you can train something in five minutes. And that is really exciting and it's an important milestone, but it's not actual RSI. So here, basically, we showed three examples of such an auto research system and what it can do. And after a very, very short time, it was able to outperform many different teams, and teams that also use other AI research. So let's double-click into some of these. NanoChat is a really exciting example. Basically, you train a very small chat model in less than five minutes and you want to have it get to the best possible bits per byte number. And so the whole community had worked on this for quite some time and got to 0.93, and after training this for a little more than a day or two, we got it down to 0.91, which is pretty exciting. Now, it wouldn't be that exciting if all it did was just find a couple of hyperparameters and tune them carefully, but it actually did find truly interesting novel ideas like hashed bigrams and trigram embeddings and tables for those and mixing that into various value paths of the attention through a variety of learned gates. So it actually started doing more and more interesting things rather than just tuning hyperparameters. Another one, a NanoGPT speedrun. Obviously, speed is very important. So here, we were able to work on this again, apply the system, and after a very short amount of time, it got better than people working often together with AI for over a year on this benchmark and made the whole thing another two seconds, over two seconds faster at 70 seconds. And again, discovering very interesting ideas in the process. And then the third one is CUDA kernels. Of course, we all care about not burning through our GPU budgets too quickly and trying to be very efficient. I think in general, it's actually kind of shocking how inefficient a lot of mixture-of-experts models still are, run in very large clusters that cost billions of dollars and only have 30% or so utilization. There's a lot of work that's ongoing in the world to improve that, and different fields or different groups of people are at various different stages of that. But long story short, lots of different CUDA kernels are used during training and testing. And here, we again took that system and after a couple of days, it discovered better kernels than the leaderboard's best on the NVIDIA benchmark website by, again, quite a sizable margin across all the different categories of those kernels. And while we are pretty good at AI, and we actually in the team didn't have any particular CUDA kernel experts who just spent their entire careers writing good kernels, still, we do just enough to make sure, and work together with NVIDIA to make sure, that there are no reward hacks here and other issues, but actually found that eventually these all checked out and were indeed pretty much all the different kernels, found the best solutions there. And so, with that, I hope I could convince you that indeed RSI could be that next big S-curve, an exponential that gets layered on top of previous exponentials. And that should help us with not just AI, but eventually science and then all of technology and then allowing many more people to flourish on our planet. And so maybe I'll end on this note here, which is a lot of people wonder how much longer AI can go, right? Every exponential eventually flattens out. And it's actually quite hard to know, when we even talk about exponential growth in AI, what does that even mean? There are many different, I call them spaces of intelligence, and we won't have time to go into all of these, but as soon as you actually try to define multiple different dimensions of each of these ten spaces that make up this complex volumetric thing that is intelligence, you'll realize that there's still so much more to go. On the upper bounds of intelligence, we're still astronomically far away from reaching those across pretty much every single one of these dimensions and the spaces that they make up. So if any of that is interesting and you want to help us build that, we'd love to hear from you. Thank you. pretty well by basically developing many of these very basic technologies, hunting, farming, but then also thinking about science, the scientific method in the early days of the Enlightenment and, of course, the Industrial Revolution. So now we can zoom even further and no worries, we're eventually going to get to NanoChat and actual auto research and what we're doing. It's a very, very quick zoom. And now we can zoom down to the last few thousands of years. And what we're seeing there is that with more technology, we were able to sustain more people, right? So when we're working on pushing that frontier forward, we're very certain that that will lead to more human flourishing, right? And especially in the last few hundred years, we're seeing this incredible explosion in the population of people because of technology and the evolution that it brings. And in many cases, that evolutionary process is run by us. So it's sort of conscious, but there are sort of interesting inspirations that we can take from that as we're thinking about the evolution of AI in the next cycles. In fact, and I might not agree with everything with Marc Andreessen, but he is very smart and we agree on a lot of things. And so I think he wrote this really great techno-optimist manifesto in which he, I think, correctly points out that the only perpetual source of growth for the entire economy, a lot of people worry about AI taking jobs and things like that, but the truth is it will very, very likely increase the economy massively and that will benefit a lot of us. And so the perpetual source of growth is technology. In fact, I think it's a lot of things that we can go even further and say that there's no material problem, and again, it's not sort of psychological problems and things like that, but no material problems that cannot be solved with even more technology, right? For the problem of starvation, we invented a green revolution, darkness, light, cold, indoor heating, heat, air conditioning, and the list goes on. So I think we can kind of realize that this evolutionary process has been going on for a very long time and continues to make a huge amount of progress. In fact, the progress is so fast that there can, within one lifetime, be a major, major shift, right? If you were born in 1900, then three years, when you're three years old, the first human ever was able to, thanks to the Wright brothers, kind of have sustained, motored flight. And then about 60-ish years later, in 1969, humans flew all the way to the moon, right? So that within one lifetime, humanity went from like, no one can fly for a very long time, other than sort of gliding down a hill or something, no one can really fly to, we all fly to the moon, right? And so for us, I think, what that means is, we are probably, and I sometimes say this, we're like too late to explore Earth, we're too early to go to the moon, right? We're not ready to explore the stars, but we're right on time to build an AI that could actually do what flying did for some in one lifetime due to intelligence. We can build and move from AI being worse at everything that we do to possibly being better at any specific task that we do, right? And that will probably be our 60-year time frame, and because everything was faster, it might only be 30 years or so. So then there's an interesting connection between technology and science and theory, right? Like sometimes the application comes first, and then we develop the theory later, and then improve the technology. Sometimes the theory comes first, and from that we can build new kinds of technologies. And so it's very helpful to think a little bit about the philosophy of science, and no better to be inspired there than Karl Popper wrote that. Just like in other types of evolution, when we choose a theory, we also choose one that is best in competition with other theories. Of course you need, if you wanted LLMs to do that, they need to find them, you need web search for instance. But in the theory that best holds its own, it's one that just like evolution has a certain natural selection process, right? It proves itself, and there is also a sort of survival of the fittest going on in scientific theories. And in fact, a lot of science, according to Popper, is basically us proposing a new theory, hypothesis, or explanation, or description, and then subjecting it to rigorous empirical testing. That is the essentially evolutionary pressure of scientific theories. And basically, that was a very short run through, sort of the history of open-ended evolution. Which hopefully makes us all realize that more science will lead to more technology, which will lead to more growth, which will lead to more human flourishing. And so that then begs the question, does it make sense for us to try to just scale up and spend a lot of our resources as humanity to scale up scientific discovery in order to lead to this flourishing? When you double click into that, you kind of realize, which Stanislaw Phlegm already realized a long time ago, that the exponential growth of science will actually be at some point halted by the lack of people working on it, right? There are so many niche subfields now in all the different areas of science that it's very hard to get a million people to work on that particular thing. And so as a result of this incredible widening of the scope, he says, the number of people focusing on any single section of it has decreased. And that then leads us to really thinking about how could we automate this and automate scientific discovery. And that then leads us to what I call the Eureka machine. This is basically our attempt at trying to build a machine that automates the process of scientific discoveries. And in fact, like in a couple of months, I'll have a book coming out on this exact idea. And so I'll just give you a super high level highlight of how such a Eureka machine could be built for basically everything from physics, chemistry, biology, neuroscience, medicine, economics, astrophysics, and so on. And there are essentially four pillars that are all extremely important to this machine. One is, of course, you have to understand what knowledge is already out there, what things humanity has already invented. You have to get all the scientific measurement data into as a second pillar of this machine. Then for things that you cannot yet measure, we don't yet know, you should try to then build simulations. Anything you can simulate, you can verify, and you can then solve with AI. And if all else fails, or at the very end of these processes, you still need to have some kind of physical industrial lab that actually can run real experiments in the real world. And on top of all of this, you'll have basically an agent swarm that will deal with all of these different sources of knowledge and data and experimentations and rewards. And in terms of, you know, the foundational model of knowledge, of course, we also, you know, basically is a good example of how every single technology we've built so far, especially in AI, but also before that, the internet, browsers, GPUs, and so on, we can rethink, and there are a lot of startups possible, and rethinking every single one of the layers of technology as infrastructure for super intelligence. All right, at U.com, for instance, we work on web search for LLMs, right, and agents, and so on. And that actually is quite different, right? Agents can read thousands of very long snippets rather than just 10 blue links with like a very short snippet. And so you can rethink each of these different layers of technology that we've built for people and rebuild them for AI in order to use them as tools to then build super intelligence. And so we've got to build super intelligence in order to automate science. Now, that is essentially the sort of why. Like, we want to build super intelligence in order to automate science. And to me, that will be the next big step function change in humanity and technology as we know it. Now, how do we actually build it? I think the best way to build it is to have it built itself, right? We've moved as a field, and especially natural language processing, for instance, which I've worked on for many years. We've moved from not having linguists. This feels like ancient, you know, BC history, but before chatGBT. We moved from having linguists tell us a bunch of things about language and then training statistical models on top of that. And when we allowed neural networks to actually automate learning those features with word vectors and other neural network architectures and back-to-back, end-to-end learning and back propagation, we basically were able to get much bigger improvements. Then we did a bunch of architecture engineering. Now, a bunch of people at least are working on a unified architecture, but even that unified architecture has a lot of manual processes. And so, it's clear over and over again in AI that when we take out a manual process and we replace it with a learned system, improvements will follow. And so, that's why I think we should try to build this weaker machine by having an RSI that builds itself. And the beauty is that only now AI can actually do this because AI is code and AI can code now. This ability to really code in longer and longer time horizons has really only happened in the last, like, six to eight months. And that now enables such an RSI to work on itself to develop almost a certain sense of self-awareness of its own shortcomings and then fix those shortcomings. And then once we have that machine that has gotten really, really good at doing research in AI itself, we can then use it to do AI research for a lot of other things in other scientific fields. And so, at a high level, it's quite easy, right? We have three steps, ideation, implementation, and validation of ideas. And so, to end maybe on some very specific examples, we have built this first kind of version of such a Eureka machine. And we wanted to just show that it works on some small samples that a lot of people know and are aware of. And so, we basically started with three things that show you and give you a very first glimpse of and sort of simple proof points of what such a machinery can do. And that was basically better training, faster training, and better kernels for NVIDIA GPUs. And so, the first one, the first one, NanoChat, I'm sure many of you have heard of it. A lot of people think that's already recursive self-improvement, and it is kind of a weak form in the sense that usually when you do auto research, it's not recursive self-improvement, right? A true recursive self-improvement is when you have an AI that has a sense of self-awareness of its own shortcomings, full access over everything in its arsenal from pre-training to RL training and harnesses and everything, and then actually updates that entire system in the next version of itself. Now, you can also take such a system and just ask it to improve some other process, some other AI, like a small NanoChat run where you can train something in five minutes. And that is really exciting and it's an important milestone, but it's not actual RSI. So, here, basically showed three examples of such an auto research system and what it can do. And after a very, very short time, it essentially was able to outperform many different teams and teams that also use other AI research. So, let's double click into some of these. NanoChat is a really exciting example. Basically, you train a very small chat model in less than five minutes and you basically want to have it get to the best possible bits per byte number. And so, the whole community had worked on this for quite some time and got to 0.93 and after training this for a little more than a day or two, we basically got it down to 0.91, which is pretty exciting. Now, it wouldn't be that exciting if all it did was just find a couple of hyperparameters and tune them carefully, but it actually did find truly interesting novel ideas like hashed bigrams and trigam embeddings and tables for those and mixing that into various value paths of the intention through a variety of learned gates. So, it actually started to doing more and more interesting things rather than just tuning hyperparameters. Another one, a nano GPT speedrun. Obviously, speed is very important. So, here we're able to work on this again, apply the system and after a very short amount of time, it got better than people working often together with AI for over a year on this benchmark and made the whole thing another two seconds, over two seconds faster at 70 seconds. And again, discovering very interesting ideas in the process. And then the third one is scuda kernels. Of course, we all care about not burning through our GPU budgets too quickly and trying to be very efficient. I think in general, it's actually kind of shocking how inefficient a lot of mixture of expert models still are run in very large clusters that cost billions of dollars and only have like 30% or so utilization. There's a lot of work that's ongoing in the world to improve that and different fields or different groups of people are various different stages of that. But long story short, lots of different cuda kernels are used during training and testing. And here, we basically, again, took that system and after a couple of days, it discovered better kernels than the leaderboards best on the NVIDIA benchmark website by, again, quite a sizable margin across all the different categories of those kernels. And while we are pretty good at AI and we actually in the team didn't have any particular cuda kernel experts who just spent their entire careers writing good kernels. But still, you know, we do just enough to make sure and work together with NVIDIA to make sure that there are no reward hacks here and other issues. But actually found that eventually these all checked out and were indeed pretty much all the different kernels found the best solutions there. And so, with that, I hope I could convince you that indeed RSI could be that next big S-curve, an exponential that gets layered on top of previous exponentials. And that should help us with not just AI, but eventually science and then all of technology and then allowing many more people to flourish on our planet. And so, maybe I'll end on this note here, which is a lot of people wonder how much longer AI can go, right? Every exponential eventually flattens out. And it's actually quite hard to know, like when we even talk about exponential growth in AI, what does that even mean? There are many different, I call them spaces of intelligence, and we won't have time to go into all of these, but as soon as you actually try to define multiple different dimensions of each of these ten spaces that make up this complex sort of volumetric thing that is intelligence, you'll realize that there's still so much more to go. Like on the upper bounds of intelligence, we're still astronomically far away from reaching those across pretty much every single one of these dimensions and the spaces that they make up. So, if any of that is interesting and you want to help us build that, we'd love to hear from you. Thank you.