.
Cool. Hi, everyone. I hope you all have a good time at the conference. I'm Nan from Model. At Model, we spend a lot of time thinking about GPU capacity: where it exists, how do we make it elastic, and what kind of workload can we actually use it for. Today, I want to talk about one place where everything gets really interesting: RL post-training. A lot of our discussion right now is about algorithms and the environments. Sandbox, PPO, GRPO, tool call, maybe low precision training, maybe deterministic kernels. But when you run those experiments at scale, the problem becomes more physical. Where are the GPUs? Are they in the same region? Do they have fast fabric?
Can we get them right now? Maybe the default shape of our compute is too restrictive. Maybe some of the work we usually force into one cluster can actually run scattered, autoscale to capacity. So, can we do our cross-globe? So, this is my talk, mainly about that. So, to make this more concrete, let's start with the RL loop itself. So, in the standard post-training loop, we can see this one trainer, and the trainer updates the policy. Rollout workers, or maybe people call them samplers, use those policies to generate trajectories. The environment returns the reward and observations. Those trajectories go back to the trainer for the next updates.
The important area here is the weight sync. In the default setup, the trainer and the rollout live in the same cluster, and the weight sync is super fast with RDMA. But they also couple the rollout fleet to the trainer cluster. If the rollout needs more capacity, maybe more nodes during runs, you are normally limited by the fixed size of your trainer cluster. So, the next question is, what kind of compute shape did we actually force everything into? On the left side is the cathedral. One region, one fast interconnect. Many GPUs wired together. This is the right shape for the trainer. On the right is the sprawl. This is where a lot of usable compute actually lives.
Different providers, different regions, different prices, and different availability. There's still a lot of capacity out there, but it's not one perfect RDMA island. This is the mismatch. Available compute is distributed, but the default RL loop asks for one tightly coupled cluster. And that cluster is exactly the hard part to get. RL wants all four of these at the same time. So, if the whole RL loop has to live inside one cluster, rollout inherits the hardest part of the trainer. So, that leads to the key question.
Does the whole RL loop actually need this kind of shape? So, let's dive into this. Training is one tightly coupled job. Every step has collectives, all-reduce, and model parallel communication. That part actually wants one fast fabric of RDMA-connected GPUs. Rollout is a fleet of serving jobs. It generates trajectories, calls environments, or maybe tools, and sends data back to the trainer. So, across rollout jobs, there's no global all-reduce. So, the thing I want to move here is not backpropagation. Backpropagation should stay in the cluster. The rollout fleet is the one that can leave. More precisely, the movable unit is the rollout serving island.
A coherent endpoint, or maybe a local group of endpoints, can serve one policy version. Inside the island, a large model may still have local parallelism. They can do PD disaggregation. They can have local serving constraints. So, across islands, the dependence is much lighter right now. Policy version in, and trajectory and metadata out. So, once we define the unit that way, the architecture is much more natural. Once we define the movable unit, the architecture is very straightforward in this case. We just have a trainer standing in the RDMA cluster, and that's where the backprop and the collectives go. The rollout side spreads out across the sprawl.
Each rollout island can be a single engine or maybe a local serving group, depending on the model and the serving topology there. Across islands, there is no global all-reduce. That's the most important thing there. The global interface is very simple. The trainer sends policy weight versions out, and the rollout sends trajectory and metadata back. At this point, the architecture depends on one remaining link: the weight update. So, if we want to send the full parameters, a full checkpoint from disk, or maybe through the network, then everything is minimal and it will break immediately.
So, after this application, the thing that we will be discussing is the size of the full parameters going through the disk. Naively, that means shipping full checkpoints every time rollout needs a new weight version. At this scale, the checkpoint is huge. So, a DeepSeek V3 FP4 checkpoint is like 500 gigabytes. Normally, it takes multiple minutes to hours to just do the weight sync. So, moving that over commodity links might not be the smartest choice, because when you're doing async, maybe even fully async training, you still want the weight update latency to be as low as possible, within seconds. So, the problem here is not whether rollout can leave the cluster.
The problem is that the full checkpoint is the wrong unit of synchronization. So, the next question is, can we keep the exact same served version there, but send a much smaller object? So, this is the bet. What if less than 1% of rollout-visible weights changed from one version to another? By rollout-visible weights, I mean the weights in the served rollout checkpoints, not the FP32 optimizer state, not the Adam moments. The weights are what the rollout engine will actually use to serve, maybe, let's say, the FP8 or maybe NVFP4 format. If that's true, we do not need to ship the entire full parameter set over the network.
We just need to ship the change in the served view, the precision delta, that difference. The important part here is still bitwise reconstruction. The rollout engine gets the same served version you would have gotten from the full checkpoint there. So, if this works, then the link shrinks from hundreds of gigabytes to maybe hundreds of megabytes. And this is something small enough that we can just send it across the network. So, right now we need to justify the less than 1% claim. Why would these rollout-visible weights barely change? Now we will get into this mechanism. So, it's small Adam-like steps meeting finite precision. We need two prerequisites, two ingredients.
Ingredient one is the precision. The optimizer may keep very high-precision master weights, but the next forward pass will be the BF16-visible view. That view has finite resolution. Around a value of magnitude theta, BF16 spacing is roughly theta over 128. That spacing is what people call ULP, the unit in the last place. It's the distance between adjacent representable BF16 values. But the update only needs to cross the nearest rounding boundary to be visible. That boundary is about half of the ULP. So, roughly, it's theta over 256. For a weight around 1, the BF16 ULP is around 0.0078, and the nearest rounding boundary is about 0.0039.
If the optimizer nudges the master weight by something smaller than that, the BF16-visible value will round back. So, you will not see any change from the rollout weight perspective. So, that's the floor. The second primitive there is what we call push. So, for Adam or maybe AdamW here, we ignore the weight decay term. The per-parameter update is the learning rate times the normalized direction. The raw gradient can be dense and can have very different magnitudes across parameters. And Adam divides by running gradient statistics. So, the per-weight push is usually on the order of the learning rate. The paper cited there proves a bound.
The Adam step is at most B times the learning rate. So, you do not need to actually remember the exact bound there. The important note here is Adam makes the push small and very controlled. So, at our post-training learning rates, the push is very, very tiny. So, that's the push. Combining these two primitives, now we have a whole better picture. A served value changes only if the push clears the floor. The push is the Adam step, roughly the learning rate. The floor is the nearest BF16 rounding boundary, roughly theta over 256. Take theta equal to one, the BF16 boundary is about 0.0039. A typical Adam step here is around 3e-6.
So, the update is more than a thousand times smaller than the boundary. So, the BF16-visible value will not change. This is not saying the master weights are frozen forever. It is saying the value that the rollout engine would serve does not change on this step. So, the whole mechanism is push versus floor. Let's visualize this to have a better understanding. The x-axis is the weight magnitude, and the y-axis is the update magnitude. First, we look at the red line. The red line is the BF16-visible boundary, theta over 256. And now we look at the green bound. Asservative value changed only if the push cleared the floor.
The push is the add-on step, roughly the learning rate. The floor is the nearest BF16 rounding boundary, roughly theta over 256. Take theta equal one, the BF16 boundary is about .0039. A typical add-on step here is around three minutes. So, the update is more than a thousand smaller than the boundary. So, the BF visible value will not change. This is not saying the master weight is for them forever. It is saying the value that rollout engine would serve does not change on this part. So, the whole magnet is pushing, is push versus for. Let's visualize this to have better understanding. The x-axis is the weight magnitude, and the y-axis is the update magnitude.
First, we look at the red line. The red line is the BF16 visible boundary. It's like theta over 256. And now we look at the green bound. This is the add-on push. It sits roughly around the learning rate and with a conservative upper bound. So, now we ask where the most point fits. For most of the weights, the red floor is above the green push. Those updates exist in the master weights, but they are not visible in the served BF16 view of this step. Small weights on the left can move, large weights on the red, they will just stay the same. They will be absorbed in. This is the add-on absorption. This is why the serve update become very sparse.
In this case, the object will be shipped as just a diff, not the entire FP32 optimized state. We first look at the rollout view. The weight cast or projected to the D type that the rollout engine will be serving. Then we will be comparing the version T minus one and the version T in this view. The patch is the change of position, plus replacement bits, and also some metadata. There are multiple lossless encoding. People can do selective overwrites. People can also do XOR. The important part is there are bit-level equivalent. It's not a floating point addition, so there's no additive delta drift.
If a rollout engine applies the patch correctly, it reconstructs the same server version bit-wise. Everything so far we explained is about the full parameter, which is hub-hub-rl. In full parameter reinforcement learning, the automizer updates the whole model. But the rollout view patch is sparse, as the thing we just explained. Lower is small for a different reason. The base model is frozen, and the adapter is small enough by construction. So we do not need to have the push versus the full arguments here. So full parameter delta is small by absorption, and the lower updates are small by construction. Let's dive into deeper about the paper itself.
So the paper, they mentioned more stats I will be showing here. The measurement is not gradient sparsity. It's not optimized state sparsity. They cast weights to BF16. Compare consecutive version stillness, and they compare the version bit-wise. And they count what did not change over time. Across model family, the result is around 99% of the time, it's bit identical per step. It also survives stillness. Even when the roll lacks, the changes start to remain very small. The important part is not only the number, it is the patch is lossless. Change index plus the replacement value reconstructs the exact same version.
So a common misconception there is it works because our gradients are sparse. They are not. The paper reports the gradients are dense. About 99% of the parameter gets non-zero gradients. The IP32 master update is also dense. It's just small. The main thing is the rollout weight change is just 1% from the perspective of the rollout engine. So far, we mostly talk about BF16, but the rollout opens serving even lower precision, such as MFFP4, FA, and NVFP4. And we can see many, many model providers doing this in the rollout. This is not the training precision. The training is just in the normal BF16, although people can do Q80 on that.
So for fixed-stale flow format, the visibility for is roughly theta over 2 to the Manteza plus 1. So, as you can see, the FFP4 will be higher, and the FA also will be between BF16 and FFP4, which means in an even lower precision, there will be less weight change. So plain flows are easy to reason about, so each element has its own rounding ceiling. One value crossed the floor, one value they just flipped and changed. Group scales such as int4, they are a bit different. This is the regime where many low precision serving systems are moving towards right now.
For int4, each weight is quantized against a shared group scale, and we can apply the same rationale, and also we can observe a similar thing for NBFV4. It's hierarchical scales, and we can see there are different encoding and displaying mechanism for NBFV4. So this is from one internal run. So here's the model we serve, GRM 4.7 Air in FP8, and we can see in the beginning, there are only 0.15% of weights got changed in the first step, where the learning rate is high. And after we have more training step, when the add-in is going relatively stable, and you can see the entire curve goes stable. We've got only 0.05% weight change during each step.
So we can see this pattern showing more generally. We have different research. We saw different research, so as RL Sync have a similar conclusion, and we saw different model providers such as Cursor, Composer 2, MAI, they are using add-in in their post-training. At this point, assuming we can produce exact rollout weights version cheaply, the next question is, how do we do this in practice? So from sparse data to RL Cross-Globe, how do we do this with elastic rollout engines and also explicit stileness? This is the whole shape. The trainer is saying the RDMA cluster, after it got updated, it published a multiple rollout-based version to a shared bullet board.
Rollout engines live outside of the training cluster, which means they don't need to be RDMA connected with the trainer. They can be in different regions or different providers. This is also a request lane. You can see a request does not just say, give me the completion. You will also say which version you will be sending requests to and which version you will be accepting. And the response will come back with the version and also exact same information as if they are in the same cluster as the trainer. There will be returning tokens, logprop, router replay information, and many more metadata.
The trainer writes the multiple versions to the broad after optimized state, after optimized step. Engine pull a version and materialize it locally in the checkpoints layout so they can just serve directly. The artifact defined version, the engine choose how to load and short it. It does not change, it does not choose the different server version. Since it will be displaying in a HFHF maybe save tensor format, which is accepted widely by muting data. So we have many rollout engines such as SGLAN and VLM. So we can support any compatible backend. Attention backend, MOE backend, different parallelism, compatible serving D-type, and any compatible GPUs there.
We can talk more about the sidecar itself. The sidecar is what makes a normal rollout engine version aware. If the version is already at an acceptable commit version, the sidecar just proxy the request. If the engine is behind but they can catch up, the sidecar just apply the missing translation. If they cannot get there, the sidecar just simply returns not ready. So this will be supporting elastated rollout, and any idle GPU can just be used with this design to support this aggregated rollout. So this is more like a system latency analyst. In cluster way sync is fast because they have RDMA. A full checkpoint across regions through network is pretty slow.
But if we use the delta, if we use exactly what we described previously, we can decrease the number of transfer sites from 500 gigabytes to 500 megabytes. So you will be extremely fast in seconds. So everything above was very general protocol. Stitch is one of very concrete implementation from model that we implemented everything above. So on the trainer side, Stitch published what defines a rollout waste version. And on the contract side, you will be pulling out and report everything on the bullet board. And on the rollout side, you will be pulling the latest waste and then start doing a waste sync across different regions and different providers.
So Stitch itself is a very framework agnostic about trainer and engine and also transport. It's very asyncing first and also agentic first. By doing this, we can have rollout engines auto scale globally. Each one self sync its weights, serve accept the version, and return rollout metadata. That means scattered inference capacity became one elastic rollout fleet. Instead of being limited by the trainer cluster, rollout can be the global pool. So inference capacity can now become our capacity. Last section, we have some ongoing explorations. So we can see a lot of model providers such as Moonshot and also DPC V4. They are adopting Muon in their post training.
Does the spot still hold for Muon? Because a lot of things we discussed previously are only for Adam. Second question is async are at scale. Right now, we can use the compute across the globe.
Then how scalable is the fully asyncing RL? So Stitch itself is very framework agnostic about trainer, engine, and transport.