Open Reader

How We Built Zeta2: Training an Edit Prediction Model in Production — Ben Kunkle, Zed

completed 10:49 May 30, 2026 Watch on YouTube

Current Status

completed

Video ID

phchDt63qAA

RAG / Chat

Enabled
How We Built Zeta2: Training an Edit Prediction Model in Production — Ben Kunkle, Zed
Description

To validate settled data, Zed ran 10 frontier model predictions per example and measured Levenshtein distance to the final state. For 100,000 training examples that is a million frontier model requests, which is prohibitively expensive. The fix: Zeta 2's student model now approaches teacher quality, so they run it 50 times instead at negligible cost. Ben Conungle, edit predictions lead at Zed, walks through how this pipeline came together. The pipeline pulls opt in production edit traces, distills them through a frontier teacher, and routes bad predictions through a repair step before formatting for the student. The ideal training examples sit in the middle of the Levenshtein distance distribution: too close to the settled state is obvious, too far is noise. A metric called reversal ratio, how often the model undoes exactly what the user just typed, was the key diagnostic for catching bad model behavior before shipping.

Summary

Generated by claude-sonnet-4-5

30-second take

Ben Kunkle explains how Zed trained Zeta 2, a small specialized model for predicting the next edit a developer will make at their cursor position. The core approach is distillation from frontier models using opt-in production snapshots, followed by a "repair step" where failed predictions get fixed by a second LLM call, then training a student model on 100k examples. The innovation is using "settled data" (what users actually typed after 10 seconds) filtered by running the student model 50 times to see if any prediction comes close via Levenshtein distance—this identifies ideal medium-difficulty training examples at near-zero cost. They validate with A/B testing at 15% traffic sampling, tracking acceptance rate and "kept rate" (how much of the prediction survives in the final code). This is a concrete production ML pipeline for a narrow, latency-sensitive task built entirely on real user behavior.

Key takes

  • Distillation with a repair loop: Frontier models generate inconsistent predictions across 100k requests, so Zed uses heuristics (undoing detection, boundary violations) to catch bad outputs and sends them to a second frontier model for correction before using them as training labels.
  • Settled data as ground truth: Waiting 10 seconds after a prediction to snapshot what the user actually wrote creates noisy labels (users change their minds, agents rewrite code), but filtering by generating 50 student model predictions and checking Levenshtein distance to the settled state isolates predictable, non-trivial examples at negligible cost.
  • Student models replace teacher for filtering: Originally generating 10 teacher predictions per example (1M requests for 100k examples) was prohibitively expensive; now they run the student checkpoint 50 times for the same filtering task because Zeta 2's quality approaches the teacher's.
  • The "goldilocks zone" in training data: Examples where predictions are far from settled state are noise; super close ones are trivial (e.g., "function add a + ..." obviously completes to "b"). The ideal training data is predictions "almost right"—like new functions past the model's training cutoff that require reasoning.
  • Production validation over offline evals: Offline metrics (Levenshtein delta, reversal ratio, proximity to 3 teacher predictions) don't fully correlate with user satisfaction, so they A/B test experiments at 15-20% traffic and track acceptance rate, latency, kept rate (chars from prediction retained in settled state), and diagnostic error counts.

Useful details

  • Input to the model: Code region around cursor, recent edits, cursor position, type/variable definitions, diagnostics/errors. Runs on every keystroke, so latency is critical.
  • Training scale: 100k examples at peak; smaller experiments use 10-50k. All data flows through JSONL where each line is a giant JSON object, with stages appending fields.
  • Experiment variables: Whether to include diagnostics, how much edit history to feed, prompt formatting for the student model.
  • Settled state heuristic: User stops editing a region for 10 seconds.
  • Metrics: Delta-to-char-F (n-gram Levenshtein variant), reversal ratio, kept rate (chars between prediction and settled state that survive), diagnostic error counts (errors before vs. after prediction).
  • Deployment: V0211 seed coder is the version released as Zeta 2; they run experiments in production with traffic sampling controls.

Caveats / counterpoints

  • Settled data remains noisy: Even after filtering, they don't train directly on the settled state—they train on the closest teacher/student prediction to it, acknowledging residual noise from user mind changes or agent rewrites.
  • 10-second heuristic is rough: Misses cases where users continuously edit a location for longer than 10 seconds without pausing.
  • Offline evals don't predict user satisfaction: Ben explicitly notes their held-out test metrics don't necessarily correlate with what users actually accept in the editor, hence the need for production A/B tests.
  • No Git commit signal used: They could incorporate commit boundaries as a stronger settled-state signal but currently don't.
  • Transcription includes some repetition: The talk loops partway through (likely a transcript artifact), but the substantive content is all in the first half.

Ken relevance

High relevance for AI ops and agent systems. This is a real production ML pipeline for a latency-critical inference task (every keystroke) that solves classic problems Ken faces: labeling noisy ground truth, managing distillation costs, validating improvements beyond offline metrics, and using the student model itself as a data curation tool. The "settled data" filtering trick—using cheap student inference to replace expensive teacher calls for data quality checks—is directly applicable to agent feedback loops or synthetic data pipelines. The point that offline evals don't correlate with user acceptance reinforces Ken's emphasis on closed-loop measurement. Zed's vertical integration (owning the editor) gives them unique data capture advantages that Ken should consider when thinking about where to build agent tooling vs. integrate with others. The middle-difficulty training example insight (not trivial, not noise) is also relevant for curriculum design in agent training.

Watch verdict

Skim. The transcript already captures all the key technical choices, numbers, and system design. Watching might clarify the UI screenshots of their experiment dashboard, but the substantive ideas are all here. Worth skimming if you're building production ML pipelines for agents or considering distillation workflows; skip if you're not directly working on model training systems.

Transcript

1675 words en Processed in 92.5s

I'm Ben Kunkel. I'm the Edit Predictions Lead at Zed. We recently announced our model Zeta 2, and this is how we trained it. I'm going to go through a lot. This is obviously a pretty short talk, so I'm going to try and leave enough time for questions at the end, but if you're not familiar with training models, it's going to be a bit of a whirlwind tour. So if you're not familiar with Edit Prediction, it's essentially giving the model a region of code around the cursor, asking them to predict the next edit that you're going to make. We give it various data in, such as your recent edits, your cursor position, the type definitions and variable definitions of things around your cursor, as well as diagnostics errors, etc. It also needs to be very fast because it runs on every keystroke. And so it's ideal for a small specialized model, fine-tuned to do this task and this task only. So that's what we've done. So the pipeline, in essence, is taking these opt-in production data. This works really well because it's snapshots. So all of that data that we have collected, related types and definitions and etc., all of that gets captured. And then we're able to turn that into training data. In order to do that, we use a process called distillation, where we take a frontier model, we give it all of that input, and we say, what prediction would you make? This is a pretty difficult process, as it turns out, even though the frontier models are pretty smart. If you ask them a hundred thousand times, they're going to give you a hundred thousand different answers, right? And so there's a bunch of problems there that we've had to finely tune the prompt that we're giving that frontier model in order to get good things out. One of the things we've done to try and get better predictions to train off of is we run some offline or static evaluations. So we have some heuristics for, is it just undoing what you just typed? Is it ignoring that editable region boundary that we've given it, etc.? And if it does, then we send it to another frontier model with a similar prompt like, hey, it failed in this way. Can you fix it? And so that we call that the repair step. And then once we've repaired the bad predictions, then we can essentially turn what the teacher made into the expected output of the student model, or Zeta 2. Up until that point is reusable across experiments. So this is stuff that we can cache, we can train multiple experiments on top of that by turning what the frontier model predicted into the format that we want the experiment to output. And so that's the next piece is this prompt formatting. This is experiment specific, i.e. are we including diagnostics this time? Are we not? How much of the edit history are we including? Those are the kinds of experiments we're running. And so we'll turn what the teacher gave us into the prompt to distill and train our student model. And then we'll do our final set of offline evaluations. The nice part about this whole process, we've designed it in such a way that it's all JSONL, or a single line has a giant JSON object. These files get huge. But each stage just adds some more fields to it or moves some fields around. So it's a very fluid and dynamic process. We're generally doing 100,000 examples to train a model. That's our peak. For these smaller experiments, we'll cut it down lower to 10 to 50k range. One interesting thing that we're trying right now is to use what we call settled data, which is the idea that eventually the user writes the answer. When you request a prediction as you're typing, eventually you're going to write the code in the way that you wanted it. And so we can wait, given that we're the editor, we can just wait until you stop editing that editable region that we gave the model, snapshot it and save it. And then use that to inform our training. This is actually very noisy. Because by waiting on the edit region to settle, you could change your mind, you could have an agent come in and rewrite it completely. It could be completely different from what it looked like when the prediction was made. So what was maybe a reasonable prediction, it no longer looks reasonable. So we need some way to filter that out. One way that we can do that is by generating 10 of the teacher predictions and seeing if any of them are close, using a Levenshtein distance type of thing, see if any of those are close to the settled state. And if they are, we know a couple of things. We know it's predictable and that it's not noisy and completely different than what the input was. Because we're giving the same input that we gave for the original prediction for this new prediction. That turns out to be quite expensive. For 100,000 examples, you're then doing a million frontier model requests. That is prohibitively expensive. Fortunately, given that we've now trained models using the original teacher predictions, our student models, or Zeta 2, is approaching the teacher in terms of quality of prediction. So instead of running the teacher, we can run our student checkpoint 50 times. That costs us basically nothing. And we can do the same process, and see if any of them are close to the settled region using Levenshtein or something similar. This gives ideal training examples, right? Because there's, by looking at the range of distance to the settled state, there's a region that are super far away. We can be confident that that's just noise. There's a region that's super close. That's it's super obvious what you're going to be doing, right? You typed function add a plus. It's obviously b, right? But then there's this interesting section in the middle where it's almost right. That's the ideal what we want in our training examples. For example, the stuff that's past the training data cutoff of our student model. So new functions, etc., that it's never seen before that you actually wanted. And that's going to show up in these new training examples that we can then train off of. We generally don't train off of the actual settled state just because it's still noisy. But we can train off of what was closest to the settled state. So to run those offline evals, we're running on a held out test set, just making sure we're not training the model on the same stuff we're testing it on. Delta to char F is our Levenshtein. Essentially, it does an n-gram comparison of various sizes of n. And then we're tracking this reversal ratio, reversals being it's undoing exactly what you just typed. And then we can also look at kept rate in production. When we're evaling, we're generally running against three teacher predictions, because a lot of these have no one right answer. And so by generating three distinct answers that were all generated by a frontier model, we can be pretty sure that if it's close to one of those, it's a pretty good prediction. So for our experiments, this is the training and production part of it. Those evals that we have don't necessarily correlate to what users actually want in their editor. And so we have this page set up of our experiments. These are the two that are live right now. You can see over here, we've got this one being sampled at 15%. And that's going to get the rest of production traffic. And so we have a dashboard that I can't show you of the acceptance rate, latency, all of that sort of stuff for these experiments. But this is a page that we created so that we can, once we've deployed it, set it to 15% of traffic, set it to 20, make it our live running model. So this V0211 seed coder, this is what we released as Zeta 2 last week. And so, yeah, we have these dashboards for the acceptance rate. We're trying new diagnostics right now, which is kept rate and diagnostic error counts. Essentially, comparing for kept rate, comparing what was the original text after the prediction and then that settled state and see how many characters between the prediction and the settled state were kept. For diagnostic error counts, it's pretty much exactly what you'd think. We snapshot how many errors there are before the prediction, how many there are after. And then we're trying to use that to judge the quality of the model. So that's it. There was a lot. Happy to answer questions. I think we have five or eight minutes left. So, yeah. You said it's very noisy to determine like the settled state. Are there any signals that you can share that you use, for example? Sorry, what was the algorithm? So you said that determining the settled state, like when the user is satisfied with that, for example, is very noisy. Are there any particular signals that you can use already that are useful? Like, for example, the Git commit at something? Sure, yeah. So we don't look at the Git commit. We could. But right now we just do you stop editing that area for 10 seconds. And that serves as a rough enough heuristic that so it's only in the cases where you are consistently editing that location for longer without pausing for 10 seconds that we wouldn't snapshot it. Yeah, gotcha. Yeah, any other questions? All right. I guess you guys get your time back. Thank you for coming. Thank you. this task and this task only. So that's what we've done. So the pipeline, in essence, is taking these opt-in production data. This works really well because it's snapshots. So all of that data that we have collected, related types and definitions and etc., all of that gets captured. And then we're able to turn that into training data. In order to do that, we use a process called distillation, where we take a frontier model, we give it all of that input, and we say, what prediction would you make? This is a pretty difficult process, as it turns out, even though the frontier models are pretty smart. If you ask them a hundred thousand times, they're gonna give you a hundred thousand one answers, right? And so there's a bunch of problems there that we've had to like finely tune the prompt that we're giving that frontier model in order to get good things out. One of the things we've done to try and get better predictions to train off of is we run some offline or static evaluations. So we have some heuristics for, you know, is it just undoing what you just typed? Is it ignoring that edible region boundary that we've given it, etc.? And if it does, then we send it to another frontier model with a similar prompt like, hey, it failed in this way. Can you fix it? And so that we call that the repair step. And then once we've repaired the bad predictions, then we can essentially turn what the teacher made into the expected output of the student model, or Zeta2. Up until that point is reusable across experiments. So this is stuff that we can cache, we can train multiple experiments on top of that by turning what the frontier model predicted into the format that we want the experiment to output. And so that's the next piece is this prompt formatting. This is experiment specific, i.e. are we including diagnostics this time? Are we not? How much of the edit history are we including? Those are the kinds of experiments we're running. And so we'll turn what the teacher gave us into the prompt to distill and train our student model. And then we'll do our final set of offline evaluations. The nice part about this whole process, we've designed it in such a way that it's all JSONL, or a single line has a giant JSON object. These files get huge. But each stage just adds some more fields to it or moves some fields around. So it's a very like fluid and dynamic process. We're generally doing 100,000 examples to train a model. Like that's our peak. For these smaller experiments, we'll cut it down lower to 10 to 50k range. One interesting thing that we're trying right now is to use what we call settled data, which is the idea that eventually the user writes the answer. When you request a prediction as you're typing, eventually you're going to write the code in the way that you wanted it. And so we can wait, given that we're the editor, we can just wait until you stop editing that editable region that we gave the model, snapshot it and save it. And then use that to inform our training. This is actually very noisy. Because by waiting on the edit region to settle, you could change your mind, you could have an agent come in and rewrite it completely. It could be completely different from what it looked like when the prediction was made. So what was maybe a reasonable prediction, it no longer looks reasonable. So we need some way to filter that out. One way that we can do that is by generating 10 of the teacher predictions and seeing if any of them are close, using like a Levenstein distance type of thing, see if any of those are close to the settled state. And if they are, we know a couple of things. We know it's predictable and that it's not noisy and completely different than what the input was. Because we're giving the same input that we gave for the original prediction for this new prediction. That turns out to be quite expensive. For 100,000 examples, you're then doing a million frontier model requests. That is prohibitively expensive. Fortunately, given that we've now trained models using the original teacher predictions, our student models, or Zeta 2, is approaching the teacher in terms of quality of prediction. So instead of running the teacher, we can run our student checkpoint or something 50 times. That costs us basically nothing. And we can do the same process, and see if any of them are close to the settled region using Levenstein or something similar. This gives ideal training examples, right? Because there's a, by looking at the range of distance to the settled state, there's a region that are super far away. We can be confident that that's just noise. There's a region that's super close. That's like, it's super obvious what you're going to be doing, right? You typed function add a plus. It's obviously b, right? But then there's this interesting section in the middle where it's almost. That's like the ideal what we want in our training examples. For example, the stuff that's past the training data cutoff of our student model. So new functions, et cetera, that it's never seen before that you actually wanted. And that's going to show up in these new training examples that we can then train off of. We generally don't train off of the actual settled state just because it's still noisy. But we can train off of, you know, what was closest to the settled state. So to run those offline evals, we're running on a held out test set, just making sure we're not training the model on the same stuff we're testing it on. Delta to car F is our Levenstein. Essentially, it does a like n-gram comparison of various sizes of n. And then we're tracking this reversal ratio, reversals being it's undoing exactly what you just typed. And then we can also look at kept rate in production. When we're evaling, we're generally running against three teacher predictions, because a lot of these have no one right answer. And so by generating three distinct answers that were all generated by a frontier model, we can be pretty sure that if it's close to one of those, it's a pretty good prediction. So for our experiments, this is the training and production part of it. Those evals that we have don't necessarily correlate to what users actually want in their editor. And so we have this page set up of our experiments. These are the two that are live right now. You can see over here, we've got this one being sampled at 15%. And that's going to get the rest of production traffic. And so we have a dashboard that I can't show you of the acceptance rate, latency, all of that kind of stuff for these experiments. But this is a page that we created so that we can, you know, once we've deployed it, set it to 15% of traffic, set it to 20, make it our live running model. So this V0211 seed coder, this is what we released as Zeta 2 last week. And so, yeah, like I said, we have these dashboards for the acceptance rate. We're trying new diagnostics right now, which is kept rate and diagnostic error counts. Essentially, comparing for kept rate, comparing what was the original text after the prediction and then that settled state and see how many characters between the prediction and the settled state were kept. For diagnostic error counts, it's pretty much exactly what you'd think. We snapshot how many errors there are before the prediction, how many there are after. And then we're then trying to use that to judge the quality of the model. So that's it. There was a lot. Happy to answer questions. I think we have five or eight minutes left. So, yeah. You said it's very noisy to determine like the settled state. Are there any signals that you can share that you use, like for example, or? Sorry, what was the algorithm? So you said that determining the settled state, like when the user is satisfied with that, for example, is very noisy. Are there any particular signals that you can use already that are useful? Like, for example, the Git commit at something? Sure, yeah. So we don't look at the Git commit. We could. But right now we just do like you stop editing that area for 10 seconds. And that serves as a rough enough heuristic that... So it's only in the cases where you are consistently editing that location for longer without pausing for 10 seconds that we wouldn't snapshot it. Yeah, gotcha. Yeah, any other questions? All right. I guess you guys get your time back. Thank you for coming. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you.