SPEAKER_01
I'm Ben Kunkel. I'm the Edit Predictions Lead at Zed. We recently announced our model Zeta 2, and this is how we trained it. I'm going to go through a lot. This is obviously a pretty short talk, so I'm going to try and leave enough time for questions at the end, but if you're not familiar with training models, it's going to be a bit of a whirlwind tour. So if you're not familiar with Edit Prediction, it's essentially giving the model a region of code around the cursor, asking them to predict the next edit that you're going to make. We give it various data in, such as your recent edits, your cursor position, the type definitions and variable definitions of things around your cursor, as well as diagnostics errors, etc. It also needs to be very fast because it runs on every keystroke. And so it's ideal for a small specialized model, fine-tuned to do this task and this task only. So that's what we've done. So the pipeline, in essence, is taking these opt-in production data. This works really well because it's snapshots. So all of that data that we have collected, related types and definitions and etc., all of that gets captured. And then we're able to turn that into training data. In order to do that, we use a process called distillation, where we take a frontier model, we give it all of that input, and we say, what prediction would you make? This is a pretty difficult process, as it turns out, even though the frontier models are pretty smart. If you ask them a hundred thousand times, they're going to give you a hundred thousand different answers, right? And so there's a bunch of problems there that we've had to finely tune the prompt that we're giving that frontier model in order to get good things out. One of the things we've done to try and get better predictions to train off of is we run some offline or static evaluations. So we have some heuristics for, is it just undoing what you just typed? Is it ignoring that editable region boundary that we've given it, etc.? And if it does, then we send it to another frontier model with a similar prompt like, hey, it failed in this way. Can you fix it? And so that we call that the repair step. And then once we've repaired the bad predictions, then we can essentially turn what the teacher made into the expected output of the student model, or Zeta 2. Up until that point is reusable across experiments. So this is stuff that we can cache, we can train multiple experiments on top of that by turning what the frontier model predicted into the format that we want the experiment to output. And so that's the next piece is this prompt formatting. This is experiment specific, i.e. are we including diagnostics this time? Are we not? How much of the edit history are we including? Those are the kinds of experiments we're running. And so we'll turn what the teacher gave us into the prompt to distill and train our student model. And then we'll do our final set of offline evaluations. The nice part about this whole process, we've designed it in such a way that it's all JSONL, or a single line has a giant JSON object. These files get huge. But each stage just adds some more fields to it or moves some fields around. So it's a very fluid and dynamic process. We're generally doing 100,000 examples to train a model. That's our peak. For these smaller experiments, we'll cut it down lower to 10 to 50k range. One interesting thing that we're trying right now is to use what we call settled data, which is the idea that eventually the user writes the answer. When you request a prediction as you're typing, eventually you're going to write the code in the way that you wanted it. And so we can wait, given that we're the editor, we can just wait until you stop editing that editable region that we gave the model, snapshot it and save it. And then use that to inform our training. This is actually very noisy. Because by waiting on the edit region to settle, you could change your mind, you could have an agent come in and rewrite it completely. It could be completely different from what it looked like when the prediction was made. So what was maybe a reasonable prediction, it no longer looks reasonable. So we need some way to filter that out. One way that we can do that is by generating 10 of the teacher predictions and seeing if any of them are close, using a Levenshtein distance type of thing, see if any of those are close to the settled state. And if they are, we know a couple of things. We know it's predictable and that it's not noisy and completely different than what the input was. Because we're giving the same input that we gave for the original prediction for this new prediction. That turns out to be quite expensive. For 100,000 examples, you're then doing a million frontier model requests. That is prohibitively expensive. Fortunately, given that we've now trained models using the original teacher predictions, our student models, or Zeta 2, is approaching the teacher in terms of quality of prediction. So instead of running the teacher, we can run our student checkpoint 50 times. That costs us basically nothing. And we can do the same process, and see if any of them are close to the settled region using Levenshtein or something similar. This gives ideal training examples, right? Because there's, by looking at the range of distance to the settled state, there's a region that are super far away. We can be confident that that's just noise. There's a region that's super close. That's it's super obvious what you're going to be doing, right? You typed function add a plus. It's obviously b, right? But then there's this interesting section in the middle where it's almost right. That's the ideal what we want in our training examples. For example, the stuff that's past the training data cutoff of our student model. So new functions, etc., that it's never seen before that you actually wanted. And that's going to show up in these new training examples that we can then train off of. We generally don't train off of the actual settled state just because it's still noisy. But we can train off of what was closest to the settled state. So to run those offline evals, we're running on a held out test set, just making sure we're not training the model on the same stuff we're testing it on. Delta to char F is our Levenshtein. Essentially, it does an n-gram comparison of various sizes of n. And then we're tracking this reversal ratio, reversals being it's undoing exactly what you just typed. And then we can also look at kept rate in production. When we're evaling, we're generally running against three teacher predictions, because a lot of these have no one right answer. And so by generating three distinct answers that were all generated by a frontier model, we can be pretty sure that if it's close to one of those, it's a pretty good prediction. So for our experiments, this is the training and production part of it. Those evals that we have don't necessarily correlate to what users actually want in their editor. And so we have this page set up of our experiments. These are the two that are live right now. You can see over here, we've got this one being sampled at 15%. And that's going to get the rest of production traffic. And so we have a dashboard that I can't show you of the acceptance rate, latency, all of that sort of stuff for these experiments. But this is a page that we created so that we can, once we've deployed it, set it to 15% of traffic, set it to 20, make it our live running model. So this V0211 seed coder, this is what we released as Zeta 2 last week. And so, yeah, we have these dashboards for the acceptance rate. We're trying new diagnostics right now, which is kept rate and diagnostic error counts. Essentially, comparing for kept rate, comparing what was the original text after the prediction and then that settled state and see how many characters between the prediction and the settled state were kept. For diagnostic error counts, it's pretty much exactly what you'd think. We snapshot how many errors there are before the prediction, how many there are after. And then we're trying to use that to judge the quality of the model. So that's it. There was a lot. Happy to answer questions. I think we have five or eight minutes left. So, yeah.
SPEAKER_01
You said it's very noisy to determine like the settled state. Are there any signals that you can share that you use, for example? Sorry, what was the algorithm? So you said that determining the settled state, like when the user is satisfied with that, for example, is very noisy. Are there any particular signals that you can use already that are useful? Like, for example, the Git commit at something?
SPEAKER_01
Sure, yeah. So we don't look at the Git commit. We could. But right now we just do you stop editing that area for 10 seconds. And that serves as a rough enough heuristic that so it's only in the cases where you are consistently editing that location for longer without pausing for 10 seconds that we wouldn't snapshot it. Yeah, gotcha. Yeah, any other questions? All right. I guess you guys get your time back. Thank you for coming. Thank you. this task and this task only. So that's what we've done. So the pipeline, in essence, is taking these opt-in production data. This works really well because it's snapshots. So all of that data that
SPEAKER_01
we have collected, related types and definitions and etc., all of that gets captured. And then we're able to turn that into training data. In order to do that, we use a process called distillation, where we take a frontier model, we give it all of that input, and we say, what prediction would you make? This is a pretty difficult process, as it turns out, even though the frontier models are pretty smart. If you ask them a hundred thousand times, they're gonna give you a hundred thousand one answers, right? And so there's a bunch of problems there that we've had to like finely tune the prompt that
SPEAKER_01
we're giving that frontier model in order to get good things out. One of the things we've done to try and get better predictions to train off of is we run some offline or static evaluations. So we have some heuristics for, you know, is it just undoing what you just typed? Is it ignoring that edible region boundary that we've given it, etc.? And if it does, then we send it to another frontier model with a similar prompt like, hey, it failed in this way. Can you fix it? And so that we call that the repair step. And then once we've repaired the bad predictions, then we can essentially turn what the teacher made into the expected output
SPEAKER_01
of the student model, or Zeta2. Up until that point is reusable across experiments. So this is stuff that we can cache, we can train multiple experiments on top of that by turning what the frontier model predicted into the format that we want the experiment to output. And so that's the next piece is this prompt formatting. This is experiment specific, i.e. are we including diagnostics this time? Are we not? How much of the edit history are we including? Those are the kinds of experiments we're running. And so we'll turn what the teacher gave us into the prompt to distill and train our student model. And then we'll do our final set of offline evaluations.
SPEAKER_01
The nice part about this whole process, we've designed it in such a way that it's all JSONL, or a single line has a giant JSON object. These files get huge. But each stage just adds some more fields to it or moves some fields around. So it's a very like fluid and dynamic process. We're generally doing 100,000 examples to train a model. Like that's our peak. For these smaller experiments, we'll cut it down lower to 10 to 50k range.
SPEAKER_01
One interesting thing that we're trying right now is to use what we call settled data, which is the idea that eventually the user writes the answer. When you request a prediction as you're typing, eventually you're going to write the code in the way that you wanted it. And so we can wait, given that we're the editor, we can just wait until you stop editing that editable region that we gave the model, snapshot it and save it. And then use that to inform our training. This is actually very noisy. Because by waiting on the edit region to settle, you could change
SPEAKER_01
your mind, you could have an agent come in and rewrite it completely. It could be completely different from what it looked like when the prediction was made. So what was maybe a reasonable prediction, it no longer looks reasonable. So we need some way to filter that out. One way that we can do that is by generating 10 of the teacher predictions and seeing if any of them are close, using like a Levenstein distance type of thing, see if any of those are close to the settled state. And if they are, we know a couple of things. We know it's predictable and that it's not noisy and completely different than what the input was.
SPEAKER_01
Because we're giving the same input that we gave for the original prediction for this new prediction. That turns out to be quite expensive. For 100,000 examples, you're then doing a million frontier model requests. That is prohibitively expensive. Fortunately, given that we've now trained models using the original teacher predictions, our student models, or Zeta 2, is approaching the teacher in terms of quality of prediction. So instead of running the teacher, we can run our student checkpoint or something 50 times. That costs us basically nothing. And we can do the same process, and see if any of them are close to the settled region using Levenstein or something similar.
SPEAKER_01
This gives ideal training examples, right? Because there's a, by looking at the range of distance to the settled state, there's a region that are super far away. We can be confident that that's just noise. There's a region that's super close. That's like, it's super obvious what you're going to be doing, right? You typed function add a plus. It's obviously b, right? But then there's this interesting section in the middle where it's almost. That's like the ideal what we want in our training examples. For example, the stuff that's past the training data cutoff of our student model. So new functions, et cetera, that it's never seen before that you actually wanted.
SPEAKER_01
And that's going to show up in these new training examples that we can then train off of. We generally don't train off of the actual settled state just because it's still noisy. But we can train off of, you know, what was closest to the settled state.
SPEAKER_01
So to run those offline evals, we're running on a held out test set, just making sure we're not training the model on the same stuff we're testing it on. Delta to car F is our Levenstein. Essentially, it does a like n-gram comparison of various sizes of n. And then we're tracking this reversal ratio, reversals being it's undoing exactly what you just typed. And then we can also look at kept rate in production. When we're evaling, we're generally running against three teacher predictions, because a lot of these have no one right answer.
SPEAKER_01
And so by generating three distinct answers that were all generated by a frontier model, we can be pretty sure that if it's close to one of those, it's a pretty good prediction. So for our experiments, this is the training and production part of it. Those evals that we have don't necessarily correlate to what users actually want in their editor. And so we have this page set up of our experiments. These are the two that are live right now. You can see over here, we've got this one being sampled at 15%.
SPEAKER_01
And that's going to get the rest of production traffic. And so we have a dashboard that I can't show you of the acceptance rate, latency, all of that kind of stuff for these experiments. But this is a page that we created so that we can, you know, once we've deployed it, set it to 15% of traffic, set it to 20, make it our live running model. So this V0211 seed coder, this is what we released as Zeta 2 last week. And so, yeah, like I said, we have these dashboards for the acceptance rate. We're trying new diagnostics right now, which is kept rate and diagnostic error counts.
SPEAKER_01
Essentially, comparing for kept rate, comparing what was the original text after the prediction and then that settled state and see how many characters between the prediction and the settled state were kept. For diagnostic error counts, it's pretty much exactly what you'd think. We snapshot how many errors there are before the prediction, how many there are after. And then we're then trying to use that to judge the quality of the model. So that's it. There was a lot. Happy to answer questions. I think we have five or eight minutes left. So, yeah.
SPEAKER_01
You said it's very noisy to determine like the settled state. Are there any signals that you can share that you use, like for example, or? Sorry, what was the algorithm? So you said that determining the settled state, like when the user is satisfied with that, for example, is very noisy. Are there any particular signals that you can use already that are useful? Like, for example, the Git commit at something? Sure, yeah. So we don't look at the Git commit. We could. But right now we just do like you stop editing that area for 10 seconds.
SPEAKER_01
And that serves as a rough enough heuristic that... So it's only in the cases where you are consistently editing that location for longer without pausing for 10 seconds that we wouldn't snapshot it. Yeah, gotcha. Yeah, any other questions?
SPEAKER_01
All right. I guess you guys get your time back. Thank you for coming. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you.