AI Engineer

Scaling to Long Horizons — Ross Taylor & Chengxi Taylor, General Reasoning

1943 summary words 9 min summary Watch video

Start with the signal

9 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Scaling AI agents from short reasoning tasks to genuinely long-horizon work requires more than stronger models: it requires RL designed for sparse, delayed rewards, persistent memory/context management, realistic environments, and compute-efficient training loops.
  • Why it matters: The talk identifies concrete architectural and training trade-offs that become central when building agents expected to plan, learn, and act over days or weeks rather than complete bounded coding or benchmark tasks.
  • Best use: Use it as a strategic design brief for long-running agent systems and as a checklist of where current agent stacks fail: memory, reward design, value estimation, environment realism, and rollout/training utilization.

Executive Summary

Ross Taylor frames the modern reasoning-model story as evidence that base-model quality alone does not create useful products or capable agents. Galactica was a strong, data-efficient scientific base model, but its public failure versus ChatGPT illustrated the value of post-training: RLHF turned a capable but quirky model into something people could reliably use. His broader claim is that reasoning behavior emerges when the underlying model, context window, and RL compute are sufficiently strong—not merely from applying a clever RL objective to a weak base model.

Taylor describes an unpublished Meta-era recipe that continued pretraining Llama 2 on mathematics and science, then used PPO with verifiable rewards and a strong outcome reward model to initialize a value model. It reportedly achieved state-of-the-art internal math and reasoning results, but did not exhibit the extensive self-reflection and inference-time scaling later associated with o1/R1-style systems. His interpretation is a version of the bitter lesson: better base models, more RL compute, and larger contexts enabled those behaviors to emerge.

Chengxi Taylor argues that long-horizon agent work is a distinct optimization regime, not simply a larger-context engineering problem. As trajectories lengthen, gradient variance rises, terminal rewards become sparse, credit assignment worsens, and variable rollout lengths complicate batching. Their proposed toolkit combines context compaction, persistent external tools such as files/search/archives, and critic/value models that can reduce variance and provide bootstrap learning signals before an episode finishes.

The practical constraint is compute. Pipeline RL improves GPU utilization by training while other rollouts are still being generated, but creates off-policy data; GR's stated experience is that up to eight policy steps of staleness is usually acceptable. For trajectories lasting weeks, that limit breaks down: either GPUs wait for terminal outcomes or training bootstraps from a value model, accepting value-model bias. The speakers also argue that existing coding-heavy benchmarks conceal this gap; in their one-year Premier League trading experiment, frontier models given $100,000 all lost money.

Key Takeaways

  • Claim: Post-training and RL are what turn a strong base model into a dependable product; base-model benchmark strength is insufficient. | Evidence: Taylor contrasts Galactica's public base-model demo with ChatGPT's RLHF-trained product, and cites 2022 InstructGPT results in which a 1B-parameter RLHF model outperformed 175B-parameter models. | Implication: For production agents, invest in behavior shaping, reward/evaluation loops, and reliability layers rather than treating a stronger foundation model as a complete solution. | Caveat: The talk presents this as a historical interpretation rather than a controlled comparison across identical model/data conditions.
  • Claim: High-quality, domain-curated data and repeated training can outperform brute-force token scaling for specialized reasoning domains. | Evidence: Galactica used a 105B-token corpus versus Chinchilla's trillion-token scale and was described as outperforming PaLM, Chinchilla, and GPT-3.5 in scientific tasks; Taylor cites roughly 68% versus GPT-3.5's 49% on LaTeX/science and 36% versus PaLM's 19% on a chain-of-thought comparison, despite Galactica being 30B versus PaLM's 540B parameters. | Implication: A domain-agent program should prioritize curated task data and targeted continued pretraining before assuming generic model scale will close domain gaps. | Caveat: These are historical, speaker-reported benchmark comparisons and do not establish that smaller curated models generally beat current frontier models.
  • Claim: Reasoning RL only produces the most visible inference-time behaviors—extended deliberation, backtracking, and reflection—when prerequisites in model capability, context, and compute are in place. | Evidence: The Meta team reportedly used continued math/science pretraining plus PPO with verifiable rewards and a value model on Llama 2, achieving strong internal math results but not the reflective behavior later seen in DeepSeek R1 and OpenAI o1; Taylor attributes the gap primarily to better base models, more RL compute, and larger context windows. | Implication: Do not expect an RL wrapper around a weak model to create durable autonomous reasoning; qualify the base model and available context/rollout budget first. | Caveat: This is an informed but unverified causal attribution; the talk does not isolate architecture, data, reward, and training differences experimentally.
  • Claim: Long-horizon RL has qualitatively harder optimization problems than short tasks: length increases gradient variance, rewards are sparse and delayed, credit assignment is weak, and trajectories have variable lengths. | Evidence: Chengxi Taylor identifies these four effects as the core barriers when agents must sustain work beyond a single bounded episode. | Implication: Agent evaluation and training should explicitly measure survival and performance across long trajectories, rather than extrapolating from coding-task success or short episodic benchmarks.
  • Claim: Value models/critics are a primary mechanism for making long-horizon training tractable, but they exchange reduced variance and earlier learning signals for complexity and bias. | Evidence: The speakers say critics reduce variance, fit trajectory-level compaction, encourage batch diversity, and enable bootstrapping before an episode ends; the cost is training a separate value model alongside the policy and accepting value-model bias. | Implication: For any long-running agent loop, treat the critic as a first-class system requiring validation, calibration, and monitoring, not as an implementation detail. | Caveat: Bootstrapping can keep GPUs productive, but incorrect value estimates can distort the policy's learning signal—especially when final outcomes are far away.
  • Claim: Persistent external memory is necessary for long work, but retrieval systems can cause shortcutting if agents reuse answers rather than reconstructing reasoning. | Evidence: Proposed tools include context compaction (generate to the context limit, summarize, then continue), files as scratchpads, self-search over prior trajectories, and archives for building on earlier results; the speakers warn that archive access can let an agent simply grab a prior answer without thinking. | Implication: Design memory systems with provenance, intermediate-artifact review, and anti-shortcut evaluations—not just larger context windows or unrestricted historical retrieval. | Caveat: Compaction and archives preserve continuity but can discard critical state, propagate prior mistakes, or reward answer retrieval over genuine problem solving.
  • Claim: Current frontier models remain unreliable on open-ended, economically consequential long-horizon tasks, partly because prevailing benchmarks are overly procedural and coding-centric. | Evidence: General Reasoning let frontier models build ML systems to trade Premier League matches over a one-year horizon; each started with $100,000 and all lost money. The speakers argue that real-world tasks include uncertainty, creative solution spaces, and other strategic actors absent from many standard benchmarks. | Implication: Avoid deploying autonomous agents into open-world financial or operational environments on the basis of coding benchmarks; require realistic simulation, downside controls, and staged human oversight. | Caveat: The transcript does not identify the models, trading rules, costs, controls, benchmark protocol, or whether the result was statistically robust, so it should not be interpreted as a general financial-performance study.

Detailed Brief

Compute scheduling creates a three-way trade-off between policy freshness, utilization, and value bias

  • Claims: Standard on-policy RL waits for complete inference trajectories before training, which wastes accelerator capacity when trajectories are long.; Pipeline RL overlaps rollout generation and model training, improving utilization but producing data from older policies.; For extremely long episodes, neither waiting nor limited-staleness pipelining is fully satisfactory; bootstrapping is the alternative.
  • Evidence: The speakers describe pipeline RL as beginning training while other sequences are still being generated.; They state that, in their experience, off-policy data up to roughly eight steps is normally acceptable.; They note that long-horizon inference may take weeks or longer, inevitably exceeding that staleness threshold if training is continuously pipelined.
  • Caveats: The eight-step tolerance is an empirical claim from the speakers, not a universal stability limit.; Bootstrapping from a critic uses expected returns before terminal completion, increasing utilization at the cost of value-model bias.
  • Implications: Long-running agent training needs an explicit rollout scheduler and a policy-staleness budget, not merely a larger GPU cluster.; Infrastructure decisions should be made jointly with algorithm design because the acceptable amount of stale data determines effective accelerator utilization.

Environment supply is positioned as a bottleneck and an infrastructure layer

  • Claims: The speakers regard diverse training/evaluation environments as essential infrastructure for long-horizon agents, alongside algorithms and compute.; General Reasoning presents OpenReward as its environment access layer.
  • Evidence: OpenReward is described as hosting more than 350 environments through a single API endpoint.; The company says it uses the platform internally for RL and that frontier and newer labs use it.
  • Caveats: The transcript provides no independent validation, environment quality criteria, or evidence that environment count translates into better long-horizon generalization.
  • Implications: The relevant competitive surface may shift from model access alone toward proprietary or well-designed interactive environments with observable outcomes and robust anti-gaming rules.

Notable Concepts & Terms

  • Long-horizon tasks: Tasks whose relevant trajectory spans many decisions, potentially days or weeks, with delayed and uncertain outcomes rather than a short verifiable completion.
  • RLHF: Reinforcement learning from human feedback; used here as the historical example of post-training that made LLMs usable products rather than raw base-model demos.
  • Verifiable rewards: Rewards based on objectively checkable task outcomes, such as mathematical correctness, which the speakers used in their earlier reasoning-RL recipe.
  • PPO: Proximal Policy Optimization, the RL algorithm Taylor says the Meta work used alongside verifiable rewards and a value model.
  • GRPO: Group Relative Policy Optimization; referenced as a simpler alternative against which the speakers contrast a critic/value-model-based approach.
  • Context compaction: Continuing work beyond a context limit by summarizing the prior trajectory and generating onward from that summary; it creates a natural target for RL but risks loss of important state.
  • Critic/value model: A model estimating expected future return, used to reduce gradient variance and bootstrap learning before a long episode receives a terminal reward.
  • Pipeline RL: Overlapping rollout generation with training to improve GPU utilization, at the cost of training on data generated by a stale policy.

Operator Notes / Why Ken Should Care

  • Define a long-horizon evaluation suite for agent initiatives that includes persistent state, delayed outcomes, uncertainty, and adversarial or multi-actor conditions; do not use coding-task pass rates as the release gate.
  • For any agent expected to run beyond one context window, require a memory design review covering compaction loss, retrieval provenance, state versioning, and tests that detect answer-copying from archives.
  • When planning RL investment, budget separately for domain continued-pretraining data, reward/environment construction, policy rollouts, and critic training/validation; a generic base model plus an outcome prompt is not the proposed recipe.
  • Set an explicit maximum policy-staleness threshold and GPU-idle tolerance for asynchronous RL pipelines; compare it against a bootstrapped critic approach rather than optimizing utilization in isolation.
  • Keep autonomous financial, trading, or open-world operational experiments sandboxed with hard loss limits and human approval, even if the agent is strong on procedural benchmarks.
  • Investigate OpenReward's environment catalog only after assessing environment fidelity, reward-gameability, licensing, observability, and whether its API can integrate with the existing agent evaluation/control plane.

Source/Metadata

  • Title: Scaling to Long Horizons — Ross Taylor & Chengxi Taylor, General Reasoning
  • Transcript words: 6008
  • Duration seconds: 1087
  • Timestamp note: No timestamps or chapters were present. The supplied transcript contains a substantial duplicated segment and trailing extraction noise, so the substantive unique content is shorter than the stated word count.
Full transcript 3250 words · 26 min read
0:12

So this talk is called Scaling to Long Horizons. My name is Ross. I'm the CEO of GR. We're a London-based reinforcement learning company. Before GR, I was the reasoning lead at MetaAI, working on Llamas, Galactica, and lots of other models back in the day. I'm joined by Cheng Shi, co-founder and president of GR. Today we're going to talk about algorithms, environments, compute, and all the things you need to do to get agents scaling to longer tasks.

0:18

So we're going to have two parts to this talk today. I'm going to first of all start with a personal perspective about the early days, the golden age of language modeling, between maybe 2020 and 2023. I'll talk about, as I said, all those models and some of our early reinforcement learning efforts for LLMs. And then Cheng Shi is going to talk about what's ahead, what are the next frontiers? And yeah, that's going to be a really interesting talk with a lot of alpha, so I'd encourage you to stick around for that.

0:22

So the journey so far. My journey started here. This was the Papers with Code team. I'm sure many of you used Papers with Code back in the day. We were a London-based startup in 2019. We were acquired by Meta later that year. And then we had a crazy transition within Meta to do research. So we did, as I said, Galactica. Then after ChatTBT came out, we started the post-training for LLAMA 2, LLAMA 3. So all the great work you saw there was folks in this room, and lots of other interesting stuff that never got published as well: Reasoning LLAMA and lots of other things. So yeah, this small team, I like to think the open weight revolution started in this room. And it really hit home this idea to me that small, focused teams, even in the age of scaling, can do amazing things if people are aligned.

0:28

Now for me, things got particularly crazy in 2022. So let me tell you a story. The media perception is that ChatTBT came out of nowhere, shocked the world, and that's how the modern AI wave started. But I have a different personal perspective on this because two weeks before ChatTBT came along, there was another language model called Galactica. So let's talk about Galactica.

0:34

Galactica and ChatTBT were both, in some respects, quite similar. They were both based on pretty good base models. Galactica itself was a base model, and then ChatTBT was based on GPT 3.5. But there was a clear difference in outcomes. Galactica at the time shipped with this base model demo. And as you guys know now, base models come with a lot of quirks. They hallucinate. You prompt them to do silly things, they would do silly things. Whereas ChatTBT wasn't just a base model, but had this crucial reinforcement learning from human feedback pipeline. And this was the key thing that made LLM's products for the first time.

0:40

So I like to think, in a way, this is the first natural experiment showing you that RL provides value, right? And to my misfortune, it was a very personal natural experiment, and Galactica blew up. But that's a good lesson there. A good base model is not enough. So I took that lesson quite early on. As I said, RLHF made LLM's products. They were the thing that made LLM's cross the Rubicon into something that wasn't just a toy, but used by now billions of people.

0:44

But you didn't have to wait until ChatTBT to see this. Even at the time, InstructGPT in 2022 had these pretty stunning results. A 1 billion parameter model with RHF was outperforming 175 billion models. So two orders of magnitude fewer parameters, but getting better results. So that was astonishing. So if you were paying attention closely, maybe we should have been as well, but we were focused on a bloody base model, which is hard work in 2022. But that shows you how important even basic RL is.

0:49

And the Galactica demo itself set up a storm. Ancient history now, but we put out a demo. We let people play around with it. We thought it was cool. And at the time people got scared. So it was, as I said, you prompt it on a research paper on Dyson spheres or a report on the benefits of eating crushed glass, and people were like, oh my God. So that was the state of things in 2022. And yeah, I'll be honest, the Meta association didn't help us either. And yeah, the tragic story, in a way, was a lot of the novel work was maybe overshadowed.

0:55

But the paradox of this whole thing is that, at the time, Galactica was actually a bloody good model. It outperformed Palm, Chinchilla, GPT 3.5, with a lot less compute in scientific domains. It was state of the art. So again, that also shows you how powerful RL is. You can have a SOTA base model, but that is not enough. So here you see on maths, it was beating Chinchilla. On latex equations, science, it was getting around 68% compared to GPT 3.5, 49%, so crushing there. And chain of thought as well. So Palm at the time, which was a Google Brain model, 540 billion, and 30 billion in Galactica was getting 36 versus 19%, so double the performance, all to a magnitude less results. So that again reinforces: good base models are not enough.

0:59

But it introduced some key ideas, which I think are very important. Galactica was the first LLM to really crack data efficiency: 105 billion token corpus compared to trillion tokens in Chinchilla. And it was really contrarian at the time, because at the time, everyone was like, okay, we just need more tokens. And Galactica said no, high-quality curated data sets really matter. And that was a real driver of those results you just saw.

1:03

And it was also the first major LLM to really crack multi-epoch training. It sounds ridiculous now, but at the time, the consensus was you don't do more than an epoch. But this rule of thumb, you may have heard of it, like four epochs of repeated data, that was formalized later. But Galactica was the first real empirical result for that.

1:10

Now, perhaps more importantly, there's this idea of thinking tokens. Some of you might remember this, but it was really quite buried within the paper. So around this time, there were different ideas for reasoning. There was chain of thought, which is one idea, which is where you prompt for the steps. There was scratch pads, where you just put very numerical intermediate steps. But Galactica was really this first idea which said, no, this is an internal working memory process. The internal thinking should be inside these tags. And you should spend the inference compute before you get to an answer, right? So these are all quite pressing ideas.

1:15

But Galactica came and went, blew up, and then we were tasked as a team to spin up the post-training effort for Llama. But I had a personal obsession, which was reasoning. And I had a really simple idea at the time, which is, what if we applied reinforcement learning pressure to these thinking tags? What if we just optimize the thing in between the thinking? And if that sounds familiar, then this is what Deep Secret 01 ended up doing two years later.

1:20

But there's a key difference. At the time, we only had Llama 2 base models, terrible mathematics corpus, terrible results on math. And the context window, we're all very context rich now, 1 million tokens, back in the day it was 4,000, which wasn't too fun.

1:27

But we still had a recipe at the time, and this was unpublished, but it was really good for Meta. Our recipe was this. Number one, continue pre-training on Llama 2 towards mathematics and science data. So that's the first thing. Llama 2 had a poor math corpus, so let's fix that. Number two, PPO with verifiable rewards. But notice this isn't GRPO, right? So we had the time and a strong outcome reward model to initialize the value model. That was a key thing, lots of data on value models at the time.

1:33

difference. So at the time, we only had Llama 2 base models, a terrible mathematics corpus, and terrible results on math. And the context window, we're all very context-rich now, 1 million tokens; back in the day, it was 4,000, which wasn't too fun.

1:37

But we still had a recipe at the time, and this was unpublished, but it was really good for Meta. So our recipe was this. Number one, continue pre-training on Llama 2 towards mathematics and science data. So that's the first thing. Llama 2, ship math corpus, so let's fix that. Number two, PPO with verifiable rewards. But notice this isn't GRPO, right? So we had the time and a strong outcome reward model to initialize the value model. That was a key thing, lots of data on value models at the time. And internally at the time, this recipe led to state-of-the-art results on math and reasoning. So we were, wow, this really shows the power of having the right objective.

1:42

But the really fascinating thing is, we had great results, but we didn't have inference-time scaling. We didn't have this reflective behavior that became the hallmark of R1 and O1: but wait, backtracking, all this stuff. So it begs the question: why? Why didn't we have that moment? And we got an answer around two years later. So there are a couple of things going on here, but essentially, better base models were the thing that really got RL cooking.

1:48

And when DeepSeek came out, I was shocked at the time. I was, holy shit, we just tried the same thing. We didn't have this. What's going on here? And in a weird way, the real lesson was it was just the bitter lesson, the purest form of bitter lesson possible: better base models, more RL compute, bigger context windows, and that's all you need for this emergent behavior. It also says something quite important about the sociology of research, because the fact that OpenAI had this model, GPT-4-level model before anyone else, allowed them to see further, right? So that's a really interesting point. The age of scaling means that if you have certain prerequisites in place, you become smarter, you see further, you see more ideas. So a really interesting point.

1:54

So this was me in 2024. I was a very sad panda, defeated by ChatGPT in O1. But I wasn't deterred. So I wanted to seek the next wave. And I still was convinced that reasoning hadn't been solved. So we started GR to take on truly big tasks. And with that in mind, I'm going to hand over to Chengchi, who's going to talk about what we're thinking about now. Chengchi Taylor-Ross, Ross, Ross. Hi, everyone. I'm Chengchi Taylor, co-founder and president, General Reasoning. I'm going to share what it takes to scale to long horizon. Chengchi Taylor-Ross, Ross, Ross, Ross.

2:21

First, I want to make it clear: long-horizon task is not just an engineering problem. It is a mindset. If we want to solve humanity's biggest problems, such as cure cancer, solve millennium surprise problem, or going to Mars, we have to be patient. It takes time. And if we want AI to move us towards that level of impact, we have to think about long horizon. But here's the first problem. We have a scarce context window. If you take Fermat's Last Theorem as an example, what it takes for the mathematician is over 10 years' time of reading papers, writing thoughts in the scratch pad, or taking a walk to generate creative ideas.

2:30

So we want to make sure that we can use the context window. So we want to make sure that we can use the context window. So we want to make sure that we can use the context window. So one solution is to use compaction. So what it essentially does is generate the token until the end of the context window, summarize, and then on top of that generate more tokens.

2:51

And the beauty of applying RL in this situation is to kill two birds with one stone. You apply RL to the compaction and also the task. But here's the problem. With long horizon, there are three issues. The first is the gradient variance scales with the length. And the second is a sparse reward. And you have this credit assignment problem. And finally, there's also variable length of the trajectory that adds to the problem of optimization. So to solve this issue, we can apply critics, which is the value model. And value model can reduce the variance and also have a couple advantages, such as on the trajectory level, it fits compaction very well, and also encourages batch diversity.

2:56

And also, I'll talk later on bootstrapping: get signal before the end of the episode. But the downside for this is that it's more complicated than GRPO. And you have to train another value model alongside the policy model. And there are some tools to help with the context limitations, such as file system tools, which are essentially a scratch pad for AI to write the reasoning thought, and self-search tools, which allow agents to search over the previous trajectory.

3:02

And then you have archive tools. In cases like auto research, you can build upon your previous result. But we have to be careful. In other scenarios, you don't want AI to cheat by just grabbing the previous answer without thinking. So how good are the current models on those long-horizon tasks? So how good are the current models on those long-horizon tasks? So what we did is that we allowed the agents to build machine learning models to trade in the football matches over a one-year horizon.

3:24

So how good are the current models on the field? And in this case, it's the Premier League if you're interested in football. And we're so fascinated by this because there's real money to be made. And if it was successful, it could make billions. There's a whole industry on sports betting. And unlike things like Kaggle competition, this has a real-world implication. But here's the result. As you can see, we gave all the frontier models 100k to start. All of them lost money. Sad. And that captured the public's imagination. Oh, AI is not as great as they thought. And why are models so bad at long horizon?

3:35

First, I believe now the AI industry is a little too biased towards coding and procedure tasks. What I mean is that the current task is mostly formulated like, do this and fix that. Normally that limits the solution to one or two. There isn't too much space for creativity. And second of all, not enough focus on open-ended tasks. We live in the real world with a lot of complexity, uncertainty. And that's not fully captured by the current benchmark. And also, there isn't enough simulation of the real world. We live in the world, there are other players. Like in today's conference room, there are other real people who have a different thought, different games than you have in your mind. That's the complexity that's not fully captured.

3:42

And another thing I want to talk about is the long-horizon impact on compute. We know that GPU is scarce and precious resources. And in this case of long-horizon reasoning, you have to be careful about how to optimize your use between training and inference. And pipeline RL is a quite popular technique nowadays. So it's a trade-off between off-policy and, in the end, the GPU utilization. So traditionally, you let inference run towards the end, and then you start to train the model. But in the case of long horizon, you have to wait until the inference finishes.

3:51

What the pipeline RL does is let the sequence be generated, and you start to train the model while there are still more sequences being generated. And you see this creates an off-policy. But from our experience, normally, off-policy up to eight steps is okay. So essentially, we made a trade-off between the off-policy and the GPU utilization. But here comes the issue. As the long horizon indicates, sometimes the inference will take weeks or even more. In that case, inevitably, it will go beyond the constraint of eight steps of our policy.

3:58

So your GPU has to just sit there idle and wait for it to finish. And if you don't want to wait, as I mentioned before, applying the value model allows you to bootstrap. What it means is that before the end of the episode, you generate expectation. It's like dopamine in the human brain. And that allows you to train the model. But here's another trade-off. While you utilize the GPU fully, you introduce the value model bias.

4:02

So there's always a bit of trade-off in those solutions. And I want to also mention that in the long horizon, infrastructure is important, especially for the environment. And Open Reward was a product, it's a platform by General Reasoning. If you're interested, you can check it out, openreward.ai. So it's a place that hosts over 350 environments and with a single API endpoint. And we use this for our internal RL, and also some frontier labs and new labs are using this.

4:13

So your GPU has to just sit there idle and wait for it to finish. And if you don't want to wait, as I mentioned before, applying the value model allows you to bootstrap. What it means is that before the end of the episode, you generate expectation. It's like dopamine in the human brain. And that allows you to train the model. But here's another trade-off. While you utilize the GPU fully, you introduce the value model bias. So there's always a bit of trade-off in those solutions.

4:24

And I want to also mention that in the long horizon, infrastructure is important, especially for the environment. And Open Reward was a product, it's a platform by General Reasoning. If you're interested, you can check it out, openreward.ai. So it's a place that hosts over 350 environments with a single API endpoint. And we use this for our internal RL, and also some frontier labs and new labs are using this.

4:30

So to summarize both Ross and my speech, it has been a long journey, as long horizon indicates. We as a team have seen the paradigms in AI reasoning, portraying an agent in the past few years. But looking ahead, what makes us really excited is the long horizon. And it requires us to have new thinking on the algorithm, environments, and compute. There are a lot of challenges and trade-offs. But we find it's really exciting to take on this journey. Because as I mentioned at the very beginning, long horizon is not just an engineering problem. It is a mindset. And if we really are ambitious to solve humanity's biggest problems, this is the journey for everyone.

4:36

And that's also the mission for General Reasoning. So if you are interested, follow us, General Reasoning. We're a London-based AI research company. Thank you.

4:40

Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra

4:46

like maths, it was kind of beating Chinchilla. Kind of latex equations, you know, science, you know, it was getting around 68% compared to GPT 3.5, 49%. So crushing there. And chain of thought as well. So you know, Palm at the time, which was a Google Brain model, 540 billion. You know, as 30 billion in Galactica was getting, you know, 36 versus 19%. So double the performance, all to a magnitude less results. So that again reinforces good base models, not enough. But it introduced some key ideas, which I think are very important. I mean, Galactica was the first LLM to really crack data efficiency. 105 billion token corpus compared to,

5:20

you know, trillion tokens in Chinchilla. And it was really contrarian at the time, because you know, at the time, everyone was like, okay, we just need more tokens. And Galactica said, no, high quality, curated data sets really matter. And that was a real driver of those results you just saw. And it was also like the first major LLM to really crack multi-epoch training. It sounds ridiculous now, but at the time, the consensus was you don't do more than an epoch. But this kind of rule of thumb, you may have heard of it, like four epochs of repeated data, that was formalized later. But Galactica was the first like real empirical result for that.

5:50

Now, perhaps more importantly, there's this idea of thinking tokens. And some of you might remember this, but it was like really quite buried within the paper. So around this time, there were like different ideas for reasoning. There was chain of thought, which is one idea, which is where you prompt for like kind of like the steps. There was scratch pads, where you just like put like very like numerical kind of intermediate steps. But Galactica was really this first idea, which said, no, this is an internal working memory process. This is the internal thinking should be inside these tags. And you should

6:17

spend the inference compute before you get to an answer, right? So these are all like, quite like, pressing ideas. But you know, Galactica came and went, blew up. And then, you know, we were kind of tasked as a team to kind of spin up the post training effort for Lama. But I had like a personal obsession, which was like reasoning. And I had like a really simple idea at the time, which is, what if we applied kind of reinforcement learning pressure to this like kind of thinking tags, this work? Like what if we just optimize the thing in between the thinking? And if that sounds familiar,

6:47

then this is like kind of what Deep Secret 01 ended up doing two years later. But there's a key difference. So at the time, we only had Lama 2 base models, terrible mathematics corpus, terrible results on math. And you know, the context window, you know, we're all like, you know, very context rich now, 1 million tokens, you know, back in the day, it was 4,000, which wasn't too fun. But we still had a recipe at the time, and this was unpublished, but it was really good for the meta. So our recipe was this. Number one, continue pre training on Lama 2 towards mathematics and science

7:18

data. So that's the first thing. Lama 2, ship math corpus, so let's fix that. Number two, PPO with verifiable rewards. But notice this isn't GRPO, right? So we had the time and a strong outcome reward model to initialize the value model. That was a key thing, lots of data on value models at the time. And internally at the time, this kind of recipe led to state-of-the-art results on math and reasoning. So we were kind of like, wow, this is like really shows the power of having the right objective. But the really fascinating thing is like, we had great results, but we didn't have like inference time scaling. We didn't have this

7:51

reflective behavior that became like the hallmark of R1 and O1, you know, butt weights, you know, backtracking, all this kind of stuff. So it begs like the question, like why? Well, why didn't we have that moment? And we got an answer around like two years later. So there's a couple of like things going on here, but essentially better base models were the thing that really got RL cooking. And when DeepSea came out, I was kind of shocked at the time. I was like, holy shit, we just like tried the same thing. We didn't have this. What's going on here? And in a weird kind of way, the real lesson was it was just like the bitter lesson, like the most purest form of bitter

8:22

lesson possible, like better base models, more RL computes, bigger context windows, and that's all you need for this kind of emergent behavior. It also like says something like quite important about the sociology of like research, because the fact that OpenAI had this model, you know, GPT-4 level model before anyone else, it allowed them to see further, right? So that's a really interesting point. Like the age of scaling means that if you have certain prerequisites in place, you become smarter, you see further, you see more ideas. So a really interesting point. So this was me in 2024. I was a

8:53

very sad panda, defeated by chat GPT in 01. But I wasn't deterred. So I wanted to seek the next wave. And I still was like convinced that kind of reasoning hadn't been solved. So we started GR to take on truly like big tasks. And with that in mind, I'm going to hand over to Chengchi, who's going to talk about what we're kind of thinking about now. Chengchi Taylor-Ross, Ross, Ross.

9:20

Hi, everyone. I'm Chengchi Taylor, co-founder and president, Jen Riesling. I'm going to share what it takes to scale to long horizon.

9:31

Chengchi Taylor-Ross, Ross, Ross, Ross. First, I want to make it clear. Long horizon task is not just an engineering problem. It is a mindset. If we want to solve humanity's biggest problems, such as cure cancer, solve millennium surprise problem, or going to Mars, we have to be patient. It takes time. And if we want AI to move us towards that level of impact, we have to think about long horizon.

10:01

But here's the first problem. We have a scarce context window. If you take Furman's last theorem as an example, what it takes for the mathematician with over 10 years' time of reading paper, writing thoughts in the scratch pad, or taking a walk to generate creative ideas.

10:26

So we want to make sure that we can use the context window. So we want to make sure that we can use the context window. So we want to make sure that we can use the context window. So one solution is use compaction. So what it essentially does is generate the token until the end of the context window, summarize, and then on top of that generate more tokens. And the beauty of applying RL in this situation is kind of like kill two birds with one stone. You apply RL to the compaction and also the task. But here's the problem. With long horizon, there are three issues. The first is the gradient variance scales with the length. And the second is

11:04

a sparse reward. And you have this credit assignment problem. And finally, there's also variable length of the trajectory that adds to the problem of optimization. So to solve this issue, we can apply critics, which is the value model. And value model can reduce the variance and also have a couple advantages, such as on the trajectory level, it fits compaction very well, and also encourage the batch diversity. And also I'll talk later on bootstrapping, basically get signal before the end of the episode. But the downside for this is that it's more complicated than GRPO. And basically you have to train another

11:48

value model alongside with the policy model. And there are some tools to help with the context limitations, such as a file system tools, which essentially like a scratch pad for AI to write the reasoning thought, and self-search tools, which allows agent to search over the previous trajectory. And then you have archive tools. In the case like auto research, you can build upon your previous result. But we have to be careful. In other scenarios, you don't want AI to cheat by just to grab the previous answer without thinking.

12:25

So how good are the current model on those long horizon tasks?

12:35

So how good are the current model on those long horizon tasks? So what we did is that we allowed the agents to build machine learning models to trade in the football matches over one year horizon. So how good are the current model on the field? And in this case, it's a Premier League if you're interested in football. And we're so fascinated by this because there's real money to be made. And if it was successful, it could make billions. There's a whole industry on sports betting. And unlike things like Kaggle competition, this has a real world implication. But here's the result. As you can see, we gave all the frontier models 100k to start.

13:24

All of them lost money. Sad. And that captured the public's imagination. Oh, AI is not as great as they thought. And why are models so bad at long horizon? First, I believe now the AI industry is a little bit too biased towards coding and procedure tasks. What I mean is that the current task is most formulated like do this and fix that. Normally that limits the solution like one or two. There isn't just too much space for creativity. And second of all, not enough focus on open-ended tasks. We live in the real world with a lot of complexity, uncertainty. And that's not fully captured by the current benchmark. And also, there isn't enough simulation of the real world.

14:13

We live in the world, there are other players. Like in today's conference room, there are other real people who have a different thought, different games than you have in your mind. That's the complexity that's not fully captured. And another thing I want to talk about is the long horizon impact on compute. We know that GPU is scarce and precious resources. And in this case of a long horizon reasoning, you have to be careful about how to optimize your use between training and inference. And pipeline RL is a quite popular technique nowadays. So basically, it's a trade-off between off-policy

14:53

and the end, the GPU utilization. So traditionally, you let inference run towards the end and then you start to train the model. But in the case of long horizon, you have to wait until the inference finish. What the pipeline RL does is that you let the sequence be generated and you start to train the model, while there's still more sequences being generated. And you see this create an off-policy. But from our experience, normally, off-policy up to eight steps is okay. So essentially, we made a trade-off between the off-policy and the GPU utilization. But here comes the issue. As the long horizon indicates, sometimes the inference will take weeks or even more.

15:39

In that case, inevitably, it will go beyond the constraint of eight steps of our policy. So your GPU have just sit there idle and wait for it to finish. And if you don't want to wait, as I mentioned before, applying the value model allows you to bootstrap. What it means is that before the end of the episode, you generate expectation. It's like a dopamine in human brain. And that allows you to train the model. But here's another trade-off. While you utilize the GPU fully, you introduce the value model bias. So there's always a bit of trade-off in those solutions. And I want to also mention that in the long horizon,

16:21

infrastructure is important, especially for the environment. And Open Reward was a product, it's a platform by general reasoning. If you're interested, you can check it out, openreward.ai. So it's a place where hosts over 350 environments and with a single API endpoint. And we use this for our internal RL and also some frontier labs and new labs are using this.

16:48

So to summarize both Ross and my speech, it has been a long journey as long horizon indicate. We as a team have seen the paradigms in AI reasoning, portraying an agent in the past few years. But looking ahead, what makes us really excited is the long horizon. And it requires us to think, have a new thinking on the algorithm, environments, and compute. There are a lot of challenges and trade-offs. But we find it's really exciting to take on this journey. Because as I mentioned at the very beginning, long horizon is not just an engineering problem. It is a mindset. And if we really are ambitious to solve humanity's biggest problems, this is the journey for everyone.

17:39

And that's also the mission for general reasoning. So if you are interested, follow us, general reasoning. We're a London-based AI research company. Thank you.

17:55

Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra Kra

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note