Well, thank you everyone for coming. Today I'll be presenting what we've been building at Morgan Stanley: auto research agent to try and automate quant research. And I wanted to just take the first couple of minutes to explain, give some context on our team and also why we're even pursuing trying to build this auto research agent. And so our group is relatively small. We're about 30 PhD AI researchers. And we operate both half academics. So we're encouraged to publish papers, open source code, share our research. And then the other half is more applied internal work.
And I think a lot of our problems, or at least some of them, can be fairly well posed. They have a similar shape to Kaggle where we have an input time series data set. And our task is to predict future values, and maybe with some other constraints of wanting the model to be well calibrated. And so I think, being on the sales side, we have maybe less adversarial selection. And so it lends itself naturally to this auto research framework.
And also I think we have a lot of algorithms in production where we have a feeling if we could just put in more cycles, there might be some improvements to squeeze out, whether it's just better hyper parameter tuning or exploring a lot of different ensembling these different methods together. Also, we work with a lot of different desks, and there's a feeling that something we have built for, say, credit bonds could really transfer well to Moody bonds. And that translation process should be something that can be automated with agents.
And even going to new trading desks and trying to build them new algorithms, there's a lot of low-hanging fruit where they're not really using a lot of machine learning or AI automation. And so even a reasonably trained model should be able to have a big impact.
And so in the last year and a half, where these agents became able to do long horizon tasks, we've been interested in trying to do this automation. But it wasn't really until December of 2025, and I think this is a pretty common sentiment now, that it really felt possible for the first time, with Opus 4.5 and with these harnesses like Claude Code and Codex, where it really felt like the models were at a point where they could do these long horizon tasks, and also the idea of putting them in these harnesses that allow them to do so. And so really starting this year, we made it a big effort to build this auto research agent.
And of course, the top-line goal is just maximizing PNL and maximizing the number of new algorithms we can put in production. But we also had these design concerns. Of course, we wanted to play nicely and integrate with all of our data and the prebuilt scaffolding we have around backtesting and evals. We also wanted it to be able to run the full spectrum. So on one side, maybe it's a new data set and we just want to say, here's the path to the data, a natural language description of what we hope to predict, and the thing should go off and do research and build its own eval and start doing experimentation.
And then on the other side of the spectrum, maybe it's like, no, we have our data scripts, we have our eval, we actually have a few good models we've already produced, and we just want it to churn and do more cycles and see if it can find an improvement. Also, we wanted to build this to be really model agnostic so we could use any of the frontier providers or increasingly any open source model. And also, we wanted to build it in a way that, as these models get better and better, it's not consuming what we've built, we rise with the tide of the models.
And lastly, really carefully think about how do we encode our enterprise knowledge as Morgan Stanley and then our human expertise as quant researchers. And so for the rest of the talk, I want to start with what I'm calling Alpha Lab 1.0, which is our first version we released, I think it was early April, and we put out a full 40-page tech report going through all the details and results. We also open sourced all the code on GitHub. So I want to cover that more at a high level because all the details are so public, but I'm happy to talk afterwards in depth about any part.
But then I really want to cover what's happened since then. So what were the initial results? What were the failure cases since then? Because we have encountered a lot of failures. And then talk about how we're really addressing those by building our own rich set of evals and environments, and how that's allowing us to improve the harness and climb towards this self-recursive improvement and our grand vision now for Alpha Lab 2.0.
And so to start, Alpha Lab is an agentic harness, and going towards that first side of the spectrum, the goal is you have some data set, let's say it's an exchange rate data set or something, and you can just provide the path to the file, or maybe it lives on an API, and you can just say here's the API access and the API spec. And then just in natural language say what you want to predict. So maybe it's as simple as the simple exchange rate. I will just be curious in predicting one day out what the rate will be.
And the harness then works in these three phases. So the first phase is research, and I'll cover these all more in depth. The second phase is then actually building its own evaluation or backtesting. And the third phase is the mass experimentation, which is really the heart of the harness. And then as output, you get a suite of trained machine learning models that are trying to do the prediction you care about. And one design choice we made, so the harness is actually all our own code. So we decided not to use any off-the-shelf agent framework. We wrote it all, and really Claude wrote it all.
And I think in the era of Claude code, I like this approach of building your own from scratch, because you get max freedom and max, you're free to tweak anything you care about. And so all the tool calls are done with these functional tool calling, and that also allows us to nicely really be provider agnostic. So OpenAI, Anthropic, or Open Source providers, it's very easy to adapt the harness to any of those. As far as the actual tools, there are several, but the three main ones are one, full shell access, so it can write any bash command. So this is how it's writing code, setting up its Python environment, editing code, running code.
Another important one is Web Search, and this allows it to go read archive or technical blogs or anything like that and get up to speed at least in the public domain, what's the state-of-the-art methods. And then the third one is we use Slurm to manage our GPU cluster, but the higher-level idea is just a nice abstraction where the model can say, I'm training a fairly big model, I need four H100s and this many CPUs, and just write the config and submit the job and not have to worry about doing hardware orchestration.
Yeah, and to go into the actual phases, so again, the first phase is this research phase. And the idea here is almost like a super Cloud MD file or a super init where we just want the system to go off and build enough context such that it can start meaningfully forming hypotheses and testing them. And we built, Alpha Lab is this server-side running thing, but we built this lightweight UI on top. And so how we've done this is build this scaffolding of the to-do list, so it's first prompted to build the to-do list such that if it completed every item, it'd be able to start experimenting. And this allows us to keep re-prompting should it try to exit early.
And so you can see it's talking about setting up its Python environment, doing different data loading, doing different kinds of statistical testing, and each item of the list is instructed to build, take notes in a markdown file, and this also allows feature agents to smartly query that and manage their context dynamically. And this is really where web search is used most because it will go off and read archive and get good context from the public domain as well. but we built this lightweight UI on top. And so how we've done this is build this scaffolding of the to-do list, so it's first prompted to build the to-do list
such that if it completed every item, it'd be able to start experimenting. And this allows us to keep re-prompting should it try to exit early. And so you can see it's talking about setting up its Python environment, doing different data loading, doing different kinds of statistical testing, and each item of the list is instructed to take notes in a markdown file, and this also allows feature agents to smartly query that and manage their context dynamically. And this is really where web search is used most because it will go off and read, archive, and get good context from the public domain as well.
And so this can vary a lot, but it takes roughly around three to four hours. The second phase, and what is most different in 2.0, is the eval building, but just to say our first implementation, of course the evaluation is the most important piece, and LLMs aren't malicious, but they can make very silly mistakes. And if you're optimizing against a bad eval, the whole thing falls apart. So our attempt to be more robust is to have this multi-agent framework. So one is tasked with first building the eval, actually writing all the code, and then that goes off to two critic agents, one that's told to be more high-level, like are there conceptual errors in our evaluation
or any forward leakage of information, and then one that's more programmatic. So it's writing unit tests and integration tests, and they write up any issues they find. It goes back to the builder to fix, and this loop doesn't end until all of them are happy that the eval is good. And then the third phase, and really the heart of everything, is this mass experimentation. And so we've chosen, both in the code and in our UI, this is formulated as a JIRA board or Kanban board. And so there's this strategist agent that gets to look and query all the context from the previous steps, and it's just supposed to keep coming up with experiments it wants to try,
and it submits them to this implement column. And then as new cards come in, those get pawned off to worker agents that are tasked with actually writing the code to implement the strategy, writing the Slurm config, what hardware does it need, submitting it to our cluster, waiting for the job to finish, and then looking at the machine learning training curves: did it underfit, overfit, and then also looking at the eval results.
And then it writes this post-mortem analysis, which goes back to the strategist as each job finishes. So the strategist hopefully can do this self-evolution. So it can see, I suggested three variants of transformers that actually didn't work too well, but XGBoost is working really well, so I want to explore more tree methods or something like that. And also in this step is where we as users can steer, so another reason we picked this JIRA formation is you can cancel cards, you can add your own cards. There's also a chat feature that's cut off here, but you can chat with a strategist and steer it toward more creative
or give it intuition of different methods it should try. And then we keep this ever-growing leaderboard where you can see, given your eval and given a held-out private validation set, what are the best-performing models? And you can click in and see, from inception of the idea to the code, and you can pull it out and play with it. And then this is just our UI 2.0, which is really just optimized to look cool, but you can see it's suggesting experiments, and this is really sped up, but how each worker pulls it through, building, deploying, and then analyzing the results.
And in our paper, we did more academic data sets, so we looked at CUDA kernels, we looked at an academic traffic time series data set, and we did the classic Karpathy-style LLM speedrunning. And it's hard to benchmark AlphaLab, but we compared to more of a Karpathy-style, single agent in a loop going, and it did find a better training config for training an LLM. We also put this on a Kaggle competition, which was hosted by NVIDIA to fine-tune their Nemetron model to be a reasoning model, and it got in the top 12% of submissions,
which, yeah, I think is decent, and also it only had 10 iterations to work with because we joined late, and I think AlphaLab works best the more iterations it can explore, so presumably or hopefully it would have done better had it had more time. And then I can say at a high level, internally, there's been a handful of models, all of the flavor where we had a decent model already, but we just turn it over to AlphaLab to keep churning on it, where it's found meaningful improvements that are now working their way through risk and going into production. And so with my last couple minutes, I want to talk about we ran into some real issues and hard questions.
I think the first was, okay, so we had some good results, but we've also had cases where it really failed, and so it's always in our head, how real is any of this? If it's failing on these hardest problems, how do we measure how real this is? And the second was, and I'm sure a lot of you might be thinking, okay, a research phase sounds reasonable, having a strategist and a worker sounds reasonable, but isn't it arbitrary? How do you motivate these design choices? And how we feel now is, and I think a big theme of this conference is, should we be making these decisions at all? This itself is a verifiable loop. An LLM should be doing this meta-optimization itself.
And then lastly, on this 1.0 formulation, it's like we give the data, we give the goal, and the LLM just goes off and does it. So how does that exactly encode our, Morgan Stanley's enterprise knowledge, but just our expertise as quants? It's missing from the picture. And so the answer to all three, we feel, is really in building our own evals and environments. And so to just go one by one through the questions, to the first point, and this is a lesson I learned over and over, monthly working with these models, you have to start with good eval. It's such an obvious thing. I think we were over-eager and wanted to treat it more like a human researcher,
but you, of course, need a very clear way to measure. And so the biggest change is now we're very opinionated about the eval. It's very much like Kaggle. And so we treat it as data and description in. The harness lives in the middle, and its only job is to submit containerized models. And it gets the feedback of a public leaderboard score, but you as a user get to see a private leaderboard, held-out validation. How is the model performing? And so in this case, you can measure, just for any given task, how well does your model do? But, of course, once you have this strict format, you can think, okay, to me, evals and environments are the same thing.
You just train in environments. So now what we've done is built on the order of 10 to 20 really careful environments, and that becomes a reinforcement learning signal. And so we're quote, unquote, alpha labbing alpha lab. So once you have that way to measure, you can do human tuning of the harness. So maybe there should be two strategists, and maybe they should debate, or maybe they're, whatever kind of ideas you have, you at least have a way to measure and manually hill-climb. But what we're doing now is really this meta harness optimization, where the LLM is looking at the traces, looking at the results, and improving the harness itself. And also a tangential axis
is we're now collecting good traces from open source model and really touching weights and doing GRPO or other on-policy distillation methods. And so we see the best-performing thing might be an orchestration of open source and closed source models, and that becomes a reinforcement learning signal. And so we're, quote, unquote, alpha labbing alpha lab. So once you have that way to measure, you can do human tuning of the harness. So maybe there should be two strategists, and maybe they should debate, or maybe they're, whatever kind of ideas you have, you at least have a way to measure and manually hill climb.
But what we're doing now is really this meta harness optimization, where the LLM is looking at the traces, looking at the results, and improving the harness itself. And also a tangential axis is we're now collecting good traces from open source model and really touching weights and doing GRPO or other on-policy distillation methods. And so we see the best-performing thing might be an orchestration of open source and closed source models, but we're optimizing the whole thing together. And to the last point, how do we as experts encode our expertise? We, again, think it's through environments.
Building environments, or at least good environments, is really, really hard work. It's get to be... There's, of course, the data that goes in that's proprietary, but designing the verifiable metrics is maybe the easy part, but then we're also doing these qualitative rubrics where we look at the traces and say, what makes a good researcher? What's the thought process? And we can grade each rollout on how well it's following our research process. And so having these good rubrics that become the signal for the model to learn, we feel is really how we're building our own expertise into the system.
And so just to conclude, the 2.0 version is really having this strict environment and eval setup. And whatever lives in the middle, we almost, in the limit, don't care about. We can initialize it to Alpha Lab 1.0, but it should really be this self-improving system. And just to leave the bigger picture headline result, my feeling is this ability to do general auto research, I think, will become a commodity. I think we've already seen it with GLM 5.2. And so I really think all of your value as an enterprise or a human expert comes from building environments.
And so temporarily for us, we're still doing manual tuning against the environment, but I think in the limit, the auto research can research itself and just be the self-improving process. And so my site's on there, and then also the project page. Again, the 1.0 version, we released everything, and our plan is to keep releasing because, again, we think the environment encodes all of the value. So thank you. Thank you. and to the floor, and to the floor, and its only job is to submit containerized models. And it gets sort of the feedback of, like, a public leaderboard score, but you as a user get to see a private leaderboard, like, held out validation.
How is the model performing? And so, you know, in this case, you can measure, like, just for any given task, how well does your model do? But, of course, once you have this strict kind of format, you can think, okay, you know, to me, evals and environments are the same thing. It's just you train in environments. So now what we've done is built on the order of, like, 10 to 20 really careful environments, and that becomes a reinforcement learning signal. And so we're, you know, quote, unquote, alpha labbing alpha lab. So once you have that way to measure, you know, you can do human tuning of the harness. So, like, maybe there should be two strategists,
and maybe they should debate, or maybe they're, you know, whatever kind of ideas you have, you at least have a way to measure and kind of manually hill climb. But what we're doing now is really this meta harness optimization, where the LLM is looking at the traces, looking at the results, and improving the harness itself. And also sort of a tangential axis is we're now collecting good traces from open source model and really touching weights and doing, like, GRPO or other, you know, on-policy distillation methods. And so we see, like, the best-performing thing might be an orchestration of open source and closed source models, but we're kind of optimizing
the whole thing together. And to the last point, you know, how do we as experts encode our expertise? We, again, think it's through environments. Like, building environments, or at least good environments, is really, really hard work. You know, it's get to be... There's, of course, like, the data that goes in that's proprietary, but, you know, designing the verifiable metrics is maybe the easy part, but then we're also doing this qualitative rubrics where we kind of look at the traces and say, you know, what makes a good researcher? What's the thought process? And we can grade each rollout on, you know, how well it's following our research process.
And so having these good rubrics that become the signal for the model to learn, we feel is really how we're building our own expertise into the system. And so just to kind of conclude, the 2.0 version is really having this strict environment and eval setup. And, you know, whatever lives in the middle, we almost, in the limit, kind of don't care about. We can initialize it to Alpha Lab 1.0, but it should really be this self-improving system. And just to leave, kind of the bigger picture headline result, and my feeling is, you know, this ability to do general auto research, I think will kind of become a commodity. Like, I think we've already seen it with GLM 5.2.
And so I really think all of your value as, like, an enterprise or a human expert comes from building environments. And so, like, you know, temporarily for us, that we're still doing, like, manual tuning against the environment, but I think in the limit, the auto research can research itself and just be the self-improving process. And so my site's on there, and then also the project page. Again, the 1.0 version, we released everything, and our plan is to keep just releasing because, again, we think the environment encodes all of the value. So thank you. Thank you. and to the floor, and to the floor,