AI Engineer

Morgan Stanley's ALPHALAB: Multi-Agent Research Across Optimization Domains — Brendan Rappazzo

1966 summary words 9 min summary Watch video

Start with the signal

9 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Morgan Stanley's AlphaLab argues that autonomous quant research is becoming viable, but durable enterprise advantage will come less from the agent harness than from proprietary, rigorously designed evaluation environments that encode domain expertise.
  • Why it matters: This is a concrete case study of moving from a multi-agent research workflow to an eval-driven, self-improving agent system, with direct lessons for building reliable AI operations and optimization agents.
  • Best use: Use it as an architecture and operating-model reference for research agents: separate exploration, evaluation, execution, oversight, and meta-optimization; invest heavily in the environment and scoring layer.

Executive Summary

Brendan Rappazzo presents AlphaLab, Morgan Stanley's autonomous research system for quant modeling. Its practical goal is to increase P&L and the number of production algorithms by automating research cycles ranging from greenfield tasks—given a dataset and a natural-language prediction target—to incremental improvement of existing models, evaluations, and data pipelines.

AlphaLab 1.0 uses a three-stage multi-agent harness: a research phase to inspect data and literature, an evaluation-building phase with separate conceptual and programmatic critics, and a Kanban-style experimentation phase in which a strategist proposes work and workers implement, train, submit cluster jobs, analyze outcomes, and feed postmortems back to the strategist. Morgan Stanley reports promising but limited evidence: a top-12% finish in an NVIDIA Kaggle competition with only 10 iterations, better LLM training configuration than a single-agent baseline, and several internal model improvements moving through risk review.

The central lesson from failures is that agent orchestration alone is not a reliable moat or even a sufficient route to reliability. If agents optimize a flawed evaluation, they can produce meaningless progress. AlphaLab 2.0 therefore makes the task environment strict: agents submit containerized models, receive public leaderboard feedback, and are judged by held-out private validation. Morgan Stanley has built roughly 10–20 such environments as the signal for both harness optimization and reinforcement learning.

Rappazzo's strategic conclusion is that general-purpose auto-research will commoditize as frontier and open models improve. The defensible asset is the enterprise environment: proprietary data, verifiable metrics, held-out tests, and qualitative rubrics capturing how expert researchers reason. AlphaLab's intended end state is not a fixed agent design, but a system that can optimize its own orchestration, model mix, and eventually its own research process against those environments.

Key Takeaways

  • Claim: AlphaLab is designed to automate both end-to-end model discovery and iterative optimization of existing quant models. | Evidence: The system can begin with only a dataset/API path and a natural-language target such as predicting an exchange rate one day ahead, or it can take existing data scripts, evaluations, and candidate models and run additional optimization, ensembling, and experimentation cycles. | Implication: Ken should scope autonomous research systems first around domains with structured inputs, measurable outputs, and repeatable evaluation loops; the highest early value may be systematic iteration on already viable workflows rather than unconstrained discovery. | Caveat: The speaker frames this as particularly suitable for relatively well-posed time-series prediction problems with measurable targets, rather than a universal replacement for all research work.
  • Claim: The initial AlphaLab architecture decomposes research into specialized phases rather than relying on one generalist agent loop. | Evidence: Phase one builds context through data inspection, statistical tests, web research, and persistent markdown notes; phase two has an eval-builder reviewed by a conceptual critic and a test-writing programmatic critic; phase three uses a strategist to create Kanban cards and workers to implement, run, and postmortem experiments. | Implication: Use role separation where it creates independently verifiable artifacts—especially evaluation review and experiment postmortems—but treat agent topology as a configurable policy, not settled architecture. | Caveat: Morgan Stanley later questions whether these specific role and workflow choices are intrinsically correct; they regard them as hypotheses that should themselves be optimized against outcomes.
  • Claim: Evaluation quality is the primary constraint on reliable autonomous optimization; a bad eval invalidates the entire system. | Evidence: Rappazzo says AlphaLab encountered meaningful failures and emphasizes that, while LLMs are not malicious, they make serious mistakes; optimizing against a flawed evaluation causes the process to collapse. The 1.0 mitigation was a builder-critic loop checking for conceptual mistakes, forward information leakage, unit-test failures, and integration-test failures. | Implication: Do not let agents optimize business or model metrics without independently designed, adversarially reviewed evaluation gates, leakage controls, and held-out validation. | Caveat: Multi-agent critique improves robustness but is not presented as a proof that an eval is correct; the major 2.0 shift was to make the environment itself more opinionated and measurable.
  • Claim: AlphaLab 2.0 reframes the system as an environment-centered optimization problem: agents submit containerized solutions, while the environment supplies public and private performance signals. | Evidence: The new setup resembles Kaggle: data and a task description go in; the harness's job is to submit containerized models; it receives a public leaderboard score, while the user sees private held-out validation. Morgan Stanley has built approximately 10–20 carefully designed environments. | Implication: For reusable agent platforms, prioritize a standardized submission contract, sandboxed execution, public feedback for iteration, and private evaluation for governance and anti-overfitting. | Caveat: The transcript does not provide performance statistics across these 10–20 environments, so the claimed superiority of the 2.0 approach is architectural direction rather than fully quantified validation.
  • Claim: Once measurable environments exist, the harness itself can be optimized through meta-evaluation rather than manually designed indefinitely. | Evidence: Morgan Stanley is using traces and results to improve the harness, testing choices such as multiple strategists or debate. It is also collecting successful traces and applying GRPO and other on-policy distillation methods to open-source models, while optimizing orchestration across open and closed models. | Implication: Build telemetry and trace capture from day one. They are prerequisites for improving prompts, tool policies, model routing, role structures, and eventually fine-tuning specialized operators. | Caveat: The system remains in a transitional stage: Rappazzo says Morgan Stanley is still manually tuning against the environments, while self-recursive improvement is the longer-term goal.
  • Claim: Morgan Stanley believes enterprise differentiation will reside in proprietary environments and expert rubrics, not in a secret agent framework. | Evidence: Rappazzo expects general auto-research to become a commodity, citing GLM 5.2 as evidence of rapid capability diffusion. Morgan Stanley open-sourced AlphaLab 1.0's code and intends to continue releasing because it views the environment as the location of the real value: proprietary data, metrics, and qualitative grading of research traces. | Implication: Treat domain-specific evals, process rubrics, test datasets, and approval criteria as strategic intellectual property; commodity models and orchestration should be replaceable components. | Caveat: This is a strategic thesis rather than a demonstrated market outcome, and open-sourcing a harness may not be equally sensible for organizations whose orchestration, integrations, or security controls are themselves differentiating.

Detailed Brief

Operational implementation choices in AlphaLab 1.0

  • Claims: The system was built as a custom harness rather than on an off-the-shelf agent framework to retain full control over behavior and future modifications.; Provider portability was an explicit requirement so that frontier closed models and increasingly capable open-source models could be swapped into the same workflow.; Human oversight is designed into the experimental workflow rather than limited to final review.
  • Evidence: Rappazzo says Morgan Stanley wrote the harness itself, with Claude producing the code, and uses functional tool calling to ease adaptation across OpenAI, Anthropic, and open-source providers.; Agents receive full shell access for code and environment setup, web search for public technical context, and a Slurm abstraction through which they request resources such as four H100 GPUs and CPUs without handling lower-level cluster orchestration.; The UI exposes a JIRA/Kanban board: users can cancel experiment cards, add their own cards, inspect artifacts from idea through code, and chat with the strategist to direct it toward particular intuitions or methods.
  • Caveats: Full shell access and cluster job submission materially expand the security and cost-control surface; the transcript does not describe sandboxing, permission boundaries, spend limits, or approval gates.; Web search can accelerate state-of-the-art discovery but may introduce unreliable, irrelevant, or non-reproducible external information unless sources and outputs are governed.
  • Implications: A production research-agent control plane should make compute requests, experiment status, cancellation, provenance, and human steering first-class capabilities.; Model-provider abstraction is worthwhile when the workflow's stable value lies in tools, environment contracts, and evaluation rather than any one model vendor.

Evidence of performance and its limits

  • Claims: AlphaLab has shown early evidence that autonomous iteration can improve technical configurations and existing internal models.; The most credible near-term use case is sustained iteration, because the system benefits from having many experiment cycles.
  • Evidence: In reported academic tests, AlphaLab was applied to CUDA kernels, an academic traffic time-series dataset, and LLM training configuration; it found a better LLM training configuration than a Karpathy-style single-agent loop.; In NVIDIA's Kaggle competition to fine-tune Nemotron into a reasoning model, AlphaLab placed in the top 12% of submissions despite joining late and receiving only 10 iterations.; Internally, it has found meaningful improvements for several already-decent models that are progressing through risk processes toward production.
  • Caveats: The talk does not quantify the internal improvements, compare their economic value to human baselines, or establish that the agent's gains will survive production and risk validation.; The speaker attributes the Kaggle result partly to iteration budget, so it should not be read as a clean benchmark of research capability.
  • Implications: Benchmark an autonomous research platform over long-run throughput, quality-adjusted improvement, and production acceptance rate—not a single leaderboard rank or isolated demo.; Maintain a formal downstream risk/production process even when an agent improves offline metrics.

Notable Concepts & Terms

  • AlphaLab: Morgan Stanley's agentic harness for automating quant research, model evaluation construction, and high-volume experimentation.
  • AlphaLab 2.0: The environment-centered redesign in which strict evaluation tasks, held-out validation, and reinforcement signals replace reliance on a fixed hand-designed multi-agent workflow.
  • Strategist / worker decomposition: A control pattern where one agent generates and adapts an experiment portfolio while execution agents implement, provision compute for, run, and analyze individual experiments.
  • Public versus private leaderboard: A Kaggle-style feedback design that permits iterative optimization against visible scores while using hidden held-out validation to detect overfitting and preserve trustworthy measurement.
  • Environment: More than a benchmark: the combination of data, constraints, verifiable metrics, qualitative rubrics, and feedback that defines what behavior the system learns to optimize.
  • GRPO: A reinforcement-learning approach mentioned as one route for training open models from successful AlphaLab traces and environment rewards.
  • On-policy distillation: Using successful trajectories generated under the current policy/process as learning data to improve model behavior for the task environment.
  • Self-recursive improvement: The intended end state where the agent system uses measured traces and environment outcomes to improve its own orchestration and research process.

Operator Notes / Why Ken Should Care

  • Create a prioritized portfolio of 10–20 representative, sandboxed environments before investing heavily in agent-role sophistication; include hidden tests, anti-leakage checks, cost limits, and production-relevant scorecards.
  • Instrument every run with durable traces linking task context, agent decisions, code changes, compute use, outputs, evaluation results, and human interventions so orchestration can be audited and optimized.
  • Adopt a public-feedback/private-validation pattern for any agent that iterates against a score, and reserve private evaluation for promotion decisions.
  • Make agent-generated experiments subject to explicit resource budgets, sandbox boundaries, source controls, and approval policies before granting shell, network, or production-adjacent tooling.
  • Treat harness choices—single versus multiple strategists, debate, model routing, worker autonomy, and prompts—as experimentally testable configurations rather than design doctrine.

Source/Metadata

  • Title: Morgan Stanley's ALPHALAB: Multi-Agent Research Across Optimization Domains — Brendan Rappazzo
  • Transcript words: 4029
  • Duration seconds: 1206
  • Timestamp note: No timestamps or chapter markers were present in the supplied transcript. The transcript includes a repeated ending segment.
Full transcript 3114 words · 18 min read
0:12

Well, thank you everyone for coming. Today I'll be presenting what we've been building at Morgan Stanley: auto research agent to try and automate quant research. And I wanted to just take the first couple of minutes to explain, give some context on our team and also why we're even pursuing trying to build this auto research agent. And so our group is relatively small. We're about 30 PhD AI researchers. And we operate both half academics. So we're encouraged to publish papers, open source code, share our research. And then the other half is more applied internal work.

0:28

And I think a lot of our problems, or at least some of them, can be fairly well posed. They have a similar shape to Kaggle where we have an input time series data set. And our task is to predict future values, and maybe with some other constraints of wanting the model to be well calibrated. And so I think, being on the sales side, we have maybe less adversarial selection. And so it lends itself naturally to this auto research framework.

0:36

And also I think we have a lot of algorithms in production where we have a feeling if we could just put in more cycles, there might be some improvements to squeeze out, whether it's just better hyper parameter tuning or exploring a lot of different ensembling these different methods together. Also, we work with a lot of different desks, and there's a feeling that something we have built for, say, credit bonds could really transfer well to Moody bonds. And that translation process should be something that can be automated with agents.

0:42

And even going to new trading desks and trying to build them new algorithms, there's a lot of low-hanging fruit where they're not really using a lot of machine learning or AI automation. And so even a reasonably trained model should be able to have a big impact.

0:46

And so in the last year and a half, where these agents became able to do long horizon tasks, we've been interested in trying to do this automation. But it wasn't really until December of 2025, and I think this is a pretty common sentiment now, that it really felt possible for the first time, with Opus 4.5 and with these harnesses like Claude Code and Codex, where it really felt like the models were at a point where they could do these long horizon tasks, and also the idea of putting them in these harnesses that allow them to do so. And so really starting this year, we made it a big effort to build this auto research agent.

0:56

And of course, the top-line goal is just maximizing PNL and maximizing the number of new algorithms we can put in production. But we also had these design concerns. Of course, we wanted to play nicely and integrate with all of our data and the prebuilt scaffolding we have around backtesting and evals. We also wanted it to be able to run the full spectrum. So on one side, maybe it's a new data set and we just want to say, here's the path to the data, a natural language description of what we hope to predict, and the thing should go off and do research and build its own eval and start doing experimentation.

1:02

And then on the other side of the spectrum, maybe it's like, no, we have our data scripts, we have our eval, we actually have a few good models we've already produced, and we just want it to churn and do more cycles and see if it can find an improvement. Also, we wanted to build this to be really model agnostic so we could use any of the frontier providers or increasingly any open source model. And also, we wanted to build it in a way that, as these models get better and better, it's not consuming what we've built, we rise with the tide of the models.

1:12

And lastly, really carefully think about how do we encode our enterprise knowledge as Morgan Stanley and then our human expertise as quant researchers. And so for the rest of the talk, I want to start with what I'm calling Alpha Lab 1.0, which is our first version we released, I think it was early April, and we put out a full 40-page tech report going through all the details and results. We also open sourced all the code on GitHub. So I want to cover that more at a high level because all the details are so public, but I'm happy to talk afterwards in depth about any part.

1:19

But then I really want to cover what's happened since then. So what were the initial results? What were the failure cases since then? Because we have encountered a lot of failures. And then talk about how we're really addressing those by building our own rich set of evals and environments, and how that's allowing us to improve the harness and climb towards this self-recursive improvement and our grand vision now for Alpha Lab 2.0.

1:22

And so to start, Alpha Lab is an agentic harness, and going towards that first side of the spectrum, the goal is you have some data set, let's say it's an exchange rate data set or something, and you can just provide the path to the file, or maybe it lives on an API, and you can just say here's the API access and the API spec. And then just in natural language say what you want to predict. So maybe it's as simple as the simple exchange rate. I will just be curious in predicting one day out what the rate will be.

1:33

And the harness then works in these three phases. So the first phase is research, and I'll cover these all more in depth. The second phase is then actually building its own evaluation or backtesting. And the third phase is the mass experimentation, which is really the heart of the harness. And then as output, you get a suite of trained machine learning models that are trying to do the prediction you care about. And one design choice we made, so the harness is actually all our own code. So we decided not to use any off-the-shelf agent framework. We wrote it all, and really Claude wrote it all.

1:46

And I think in the era of Claude code, I like this approach of building your own from scratch, because you get max freedom and max, you're free to tweak anything you care about. And so all the tool calls are done with these functional tool calling, and that also allows us to nicely really be provider agnostic. So OpenAI, Anthropic, or Open Source providers, it's very easy to adapt the harness to any of those. As far as the actual tools, there are several, but the three main ones are one, full shell access, so it can write any bash command. So this is how it's writing code, setting up its Python environment, editing code, running code.

2:03

Another important one is Web Search, and this allows it to go read archive or technical blogs or anything like that and get up to speed at least in the public domain, what's the state-of-the-art methods. And then the third one is we use Slurm to manage our GPU cluster, but the higher-level idea is just a nice abstraction where the model can say, I'm training a fairly big model, I need four H100s and this many CPUs, and just write the config and submit the job and not have to worry about doing hardware orchestration.

2:13

Yeah, and to go into the actual phases, so again, the first phase is this research phase. And the idea here is almost like a super Cloud MD file or a super init where we just want the system to go off and build enough context such that it can start meaningfully forming hypotheses and testing them. And we built, Alpha Lab is this server-side running thing, but we built this lightweight UI on top. And so how we've done this is build this scaffolding of the to-do list, so it's first prompted to build the to-do list such that if it completed every item, it'd be able to start experimenting. And this allows us to keep re-prompting should it try to exit early.

2:30

And so you can see it's talking about setting up its Python environment, doing different data loading, doing different kinds of statistical testing, and each item of the list is instructed to build, take notes in a markdown file, and this also allows feature agents to smartly query that and manage their context dynamically. And this is really where web search is used most because it will go off and read archive and get good context from the public domain as well. but we built this lightweight UI on top. And so how we've done this is build this scaffolding of the to-do list, so it's first prompted to build the to-do list

3:01

such that if it completed every item, it'd be able to start experimenting. And this allows us to keep re-prompting should it try to exit early. And so you can see it's talking about setting up its Python environment, doing different data loading, doing different kinds of statistical testing, and each item of the list is instructed to take notes in a markdown file, and this also allows feature agents to smartly query that and manage their context dynamically. And this is really where web search is used most because it will go off and read, archive, and get good context from the public domain as well.

3:57

And so this can vary a lot, but it takes roughly around three to four hours. The second phase, and what is most different in 2.0, is the eval building, but just to say our first implementation, of course the evaluation is the most important piece, and LLMs aren't malicious, but they can make very silly mistakes. And if you're optimizing against a bad eval, the whole thing falls apart. So our attempt to be more robust is to have this multi-agent framework. So one is tasked with first building the eval, actually writing all the code, and then that goes off to two critic agents, one that's told to be more high-level, like are there conceptual errors in our evaluation

4:54

or any forward leakage of information, and then one that's more programmatic. So it's writing unit tests and integration tests, and they write up any issues they find. It goes back to the builder to fix, and this loop doesn't end until all of them are happy that the eval is good. And then the third phase, and really the heart of everything, is this mass experimentation. And so we've chosen, both in the code and in our UI, this is formulated as a JIRA board or Kanban board. And so there's this strategist agent that gets to look and query all the context from the previous steps, and it's just supposed to keep coming up with experiments it wants to try,

5:58

and it submits them to this implement column. And then as new cards come in, those get pawned off to worker agents that are tasked with actually writing the code to implement the strategy, writing the Slurm config, what hardware does it need, submitting it to our cluster, waiting for the job to finish, and then looking at the machine learning training curves: did it underfit, overfit, and then also looking at the eval results.

6:34

And then it writes this post-mortem analysis, which goes back to the strategist as each job finishes. So the strategist hopefully can do this self-evolution. So it can see, I suggested three variants of transformers that actually didn't work too well, but XGBoost is working really well, so I want to explore more tree methods or something like that. And also in this step is where we as users can steer, so another reason we picked this JIRA formation is you can cancel cards, you can add your own cards. There's also a chat feature that's cut off here, but you can chat with a strategist and steer it toward more creative

7:43

or give it intuition of different methods it should try. And then we keep this ever-growing leaderboard where you can see, given your eval and given a held-out private validation set, what are the best-performing models? And you can click in and see, from inception of the idea to the code, and you can pull it out and play with it. And then this is just our UI 2.0, which is really just optimized to look cool, but you can see it's suggesting experiments, and this is really sped up, but how each worker pulls it through, building, deploying, and then analyzing the results.

8:55

And in our paper, we did more academic data sets, so we looked at CUDA kernels, we looked at an academic traffic time series data set, and we did the classic Karpathy-style LLM speedrunning. And it's hard to benchmark AlphaLab, but we compared to more of a Karpathy-style, single agent in a loop going, and it did find a better training config for training an LLM. We also put this on a Kaggle competition, which was hosted by NVIDIA to fine-tune their Nemetron model to be a reasoning model, and it got in the top 12% of submissions,

10:04

which, yeah, I think is decent, and also it only had 10 iterations to work with because we joined late, and I think AlphaLab works best the more iterations it can explore, so presumably or hopefully it would have done better had it had more time. And then I can say at a high level, internally, there's been a handful of models, all of the flavor where we had a decent model already, but we just turn it over to AlphaLab to keep churning on it, where it's found meaningful improvements that are now working their way through risk and going into production. And so with my last couple minutes, I want to talk about we ran into some real issues and hard questions.

11:08

I think the first was, okay, so we had some good results, but we've also had cases where it really failed, and so it's always in our head, how real is any of this? If it's failing on these hardest problems, how do we measure how real this is? And the second was, and I'm sure a lot of you might be thinking, okay, a research phase sounds reasonable, having a strategist and a worker sounds reasonable, but isn't it arbitrary? How do you motivate these design choices? And how we feel now is, and I think a big theme of this conference is, should we be making these decisions at all? This itself is a verifiable loop. An LLM should be doing this meta-optimization itself.

12:11

And then lastly, on this 1.0 formulation, it's like we give the data, we give the goal, and the LLM just goes off and does it. So how does that exactly encode our, Morgan Stanley's enterprise knowledge, but just our expertise as quants? It's missing from the picture. And so the answer to all three, we feel, is really in building our own evals and environments. And so to just go one by one through the questions, to the first point, and this is a lesson I learned over and over, monthly working with these models, you have to start with good eval. It's such an obvious thing. I think we were over-eager and wanted to treat it more like a human researcher,

13:02

but you, of course, need a very clear way to measure. And so the biggest change is now we're very opinionated about the eval. It's very much like Kaggle. And so we treat it as data and description in. The harness lives in the middle, and its only job is to submit containerized models. And it gets the feedback of a public leaderboard score, but you as a user get to see a private leaderboard, held-out validation. How is the model performing? And so in this case, you can measure, just for any given task, how well does your model do? But, of course, once you have this strict format, you can think, okay, to me, evals and environments are the same thing.

13:52

You just train in environments. So now what we've done is built on the order of 10 to 20 really careful environments, and that becomes a reinforcement learning signal. And so we're quote, unquote, alpha labbing alpha lab. So once you have that way to measure, you can do human tuning of the harness. So maybe there should be two strategists, and maybe they should debate, or maybe they're, whatever kind of ideas you have, you at least have a way to measure and manually hill-climb. But what we're doing now is really this meta harness optimization, where the LLM is looking at the traces, looking at the results, and improving the harness itself. And also a tangential axis

14:43

is we're now collecting good traces from open source model and really touching weights and doing GRPO or other on-policy distillation methods. And so we see the best-performing thing might be an orchestration of open source and closed source models, and that becomes a reinforcement learning signal. And so we're, quote, unquote, alpha labbing alpha lab. So once you have that way to measure, you can do human tuning of the harness. So maybe there should be two strategists, and maybe they should debate, or maybe they're, whatever kind of ideas you have, you at least have a way to measure and manually hill climb.

15:20

But what we're doing now is really this meta harness optimization, where the LLM is looking at the traces, looking at the results, and improving the harness itself. And also a tangential axis is we're now collecting good traces from open source model and really touching weights and doing GRPO or other on-policy distillation methods. And so we see the best-performing thing might be an orchestration of open source and closed source models, but we're optimizing the whole thing together. And to the last point, how do we as experts encode our expertise? We, again, think it's through environments.

15:30

Building environments, or at least good environments, is really, really hard work. It's get to be... There's, of course, the data that goes in that's proprietary, but designing the verifiable metrics is maybe the easy part, but then we're also doing these qualitative rubrics where we look at the traces and say, what makes a good researcher? What's the thought process? And we can grade each rollout on how well it's following our research process. And so having these good rubrics that become the signal for the model to learn, we feel is really how we're building our own expertise into the system.

15:47

And so just to conclude, the 2.0 version is really having this strict environment and eval setup. And whatever lives in the middle, we almost, in the limit, don't care about. We can initialize it to Alpha Lab 1.0, but it should really be this self-improving system. And just to leave the bigger picture headline result, my feeling is this ability to do general auto research, I think, will become a commodity. I think we've already seen it with GLM 5.2. And so I really think all of your value as an enterprise or a human expert comes from building environments.

16:01

And so temporarily for us, we're still doing manual tuning against the environment, but I think in the limit, the auto research can research itself and just be the self-improving process. And so my site's on there, and then also the project page. Again, the 1.0 version, we released everything, and our plan is to keep releasing because, again, we think the environment encodes all of the value. So thank you. Thank you. and to the floor, and to the floor, and its only job is to submit containerized models. And it gets sort of the feedback of, like, a public leaderboard score, but you as a user get to see a private leaderboard, like, held out validation.

16:29

How is the model performing? And so, you know, in this case, you can measure, like, just for any given task, how well does your model do? But, of course, once you have this strict kind of format, you can think, okay, you know, to me, evals and environments are the same thing. It's just you train in environments. So now what we've done is built on the order of, like, 10 to 20 really careful environments, and that becomes a reinforcement learning signal. And so we're, you know, quote, unquote, alpha labbing alpha lab. So once you have that way to measure, you know, you can do human tuning of the harness. So, like, maybe there should be two strategists,

17:08

and maybe they should debate, or maybe they're, you know, whatever kind of ideas you have, you at least have a way to measure and kind of manually hill climb. But what we're doing now is really this meta harness optimization, where the LLM is looking at the traces, looking at the results, and improving the harness itself. And also sort of a tangential axis is we're now collecting good traces from open source model and really touching weights and doing, like, GRPO or other, you know, on-policy distillation methods. And so we see, like, the best-performing thing might be an orchestration of open source and closed source models, but we're kind of optimizing

17:48

the whole thing together. And to the last point, you know, how do we as experts encode our expertise? We, again, think it's through environments. Like, building environments, or at least good environments, is really, really hard work. You know, it's get to be... There's, of course, like, the data that goes in that's proprietary, but, you know, designing the verifiable metrics is maybe the easy part, but then we're also doing this qualitative rubrics where we kind of look at the traces and say, you know, what makes a good researcher? What's the thought process? And we can grade each rollout on, you know, how well it's following our research process.

18:27

And so having these good rubrics that become the signal for the model to learn, we feel is really how we're building our own expertise into the system. And so just to kind of conclude, the 2.0 version is really having this strict environment and eval setup. And, you know, whatever lives in the middle, we almost, in the limit, kind of don't care about. We can initialize it to Alpha Lab 1.0, but it should really be this self-improving system. And just to leave, kind of the bigger picture headline result, and my feeling is, you know, this ability to do general auto research, I think will kind of become a commodity. Like, I think we've already seen it with GLM 5.2.

19:10

And so I really think all of your value as, like, an enterprise or a human expert comes from building environments. And so, like, you know, temporarily for us, that we're still doing, like, manual tuning against the environment, but I think in the limit, the auto research can research itself and just be the self-improving process. And so my site's on there, and then also the project page. Again, the 1.0 version, we released everything, and our plan is to keep just releasing because, again, we think the environment encodes all of the value. So thank you. Thank you. and to the floor, and to the floor,

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note