AI Engineer

What's New in Inference Engineering — Philip Kiely, Baseten

1804 summary words 8 min summary Watch video

Start with the signal

8 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Inference engineering is shifting from standalone serving optimizations toward training-informed systems that jointly optimize quantization, KV-cache management, and speculative decoding for production throughput.
  • Why it matters: The talk gives a current production-oriented view of which inference techniques are delivering real gains, which promising research is not yet viable in datacenters, and where future serving architectures will differentiate.
  • Best use: Use it to update an inference roadmap: prioritize model-specific speculative decoding and KV-aware system design, while treating aggressive KV compression and new research methods as workload-dependent experiments.

Executive Summary

Philip Kiely frames inference engineering as two related but distinct disciplines. Local inference starts with constrained hardware and tries to preserve model intelligence after quantization, pruning, distillation, or multi-GPU splitting; datacenter inference starts with a functioning deployment and then removes latency and throughput bottlenecks through routing, disaggregation, caching, and speculation. His central update is that training and inference are increasingly inseparable: several of the best serving optimizations now require dedicated training.

On quantization, Kiely pushes back on the initial excitement around TurboQuant. Its four-bit KV-cache representation can halve cache storage and effectively double cache-transfer bandwidth, but its added decode computation cut tokens per second by more than half in Baseten's testing, making it unsuitable for many production datacenter workloads. Baseten instead favors NVFP4 weight quantization and operational KV-cache techniques such as routing, sharing, and offloading.

The strongest practical section is speculative decoding. D-Flash, a diffusion-based draft model that generates blocks of draft tokens rather than one autoregressive token at a time, reportedly delivered more than a 3x improvement over Eagle on a single B200 running Qwen3-8B. The mechanism matters because higher draft-token acceptance is the main determinant of speculative-decoding value. Kiely also argues that continuously retraining speculators on live workload distributions can improve acceptance rates by 20% to 2x, though it creates significant data-permission, storage, compute, and model-versioning burdens.

Looking ahead, Kiely expects NVFP4 to become more important with NVIDIA Rubin, prefill/decode disaggregation and system-wide KV movement to grow in importance, and training-for-inference techniques to become a durable competitive advantage. This is a high-signal technical update for anyone operating model-serving infrastructure, though several claims are vendor-presented results rather than independently benchmarked comparisons.

Key Takeaways

  • Claim: Production inference should be approached as a system optimization problem, not merely as serving a fixed set of model weights. | Evidence: Kiely distinguishes local inference, where engineers compress models until they run on available hardware and then recover quality, from datacenter inference, where teams first deploy and then improve speed through KV-aware routing, speculation, and disaggregation. | Implication: Ken should separate optimization strategy by deployment environment rather than assume one quantization, cache, or routing design serves both local and centralized workloads. | Caveat: The talk is explicitly focused on datacenter-oriented inference; its recommendations are not universally optimal for local or memory-constrained deployments.
  • Claim: TurboQuant is a compelling memory-saving technique but not presently a default datacenter optimization because its decode overhead can overwhelm its bandwidth benefit. | Evidence: TurboQuant uses polar-coordinate quantization to reduce KV-cache storage from eight bits to four bits, halving required memory and effectively doubling KV-transfer bandwidth; Baseten found that the added forward-pass work during decode reduced TPS by more than half. | Implication: Do not adopt KV-cache compression solely on memory-efficiency claims; benchmark end-to-end decode TPS, latency, and context capacity on the target workload. | Caveat: Kiely considers it potentially valuable for local long-context inference, where GPU memory capacity is the primary constraint and extra computation is less costly than an inability to fit the context.
  • Claim: For current production systems, NVFP4 weight quantization plus explicit KV-cache movement is a more attractive path than TurboQuant-style KV compression. | Evidence: Baseten is using traditional NVFP4 while checking for quantization-induced probability-distribution degradation, and is prioritizing KV-aware routing, KV sharing, KV offloading, NCCL ('Nickel' in the transcript), NVIDIA Dynamo, and potential CPU-memory offload. | Implication: Treat the KV cache as a distributed systems resource: design placement, reuse, transfer, and offload policies before pursuing lossy cache compression. | Caveat: The appropriate balance depends on hardware, model architecture, context lengths, communication topology, and whether weight bandwidth or KV capacity is the actual bottleneck.
  • Claim: KV compaction is evolving from inference-time selection or compression toward trained, learned memory representations. | Evidence: Kiely cites attention matching and Cartridges as promising inference-time compaction approaches, then describes Baseten Research's STill method: a Perceiver bottleneck cross-attends learned query vectors against the full KV cache and produces compact keys and values in one forward pass. | Implication: For agent systems with very long contexts, learned cache compaction may become a middle path between lossless full-context retention and lossy prompt/RAG compression, but it requires evaluation against task-specific recall and reasoning quality. | Caveat: The talk does not provide production latency, quality-retention, or compression-ratio benchmarks for STill, so it should be viewed as a research direction rather than a validated default.
  • Claim: D-Flash materially advances speculative decoding by drafting token sequences with a diffusion language model rather than generating draft tokens one at a time. | Evidence: D-Flash predicts blocks of roughly eight or 16 tokens, lets draft tokens cross-attend bidirectionally within the block, and is said to outperform Eagle because it produces more mutually consistent drafts with higher acceptance. Baseten reports more than a 3x improvement versus Eagle on one B200 with Qwen3-8B. | Implication: Speculative-decoding evaluation should focus on accepted tokens per expensive target-model pass, not merely the standalone speed of the draft model. | Caveat: The reported result is a specific hardware-and-model benchmark from the presenter; gains will vary with target model, batch/workload distribution, draft-model cost, and verifier acceptance rate.
  • Claim: Continuously retraining speculative draft models on live prompt and response distributions can be a major throughput lever at scale. | Evidence: Kiely reports that continuous retraining for D-Flash can improve draft-token acceptance rates by 20% to 2x because speculation performance depends heavily on the actual prompts and responses seen in production. | Implication: For high-volume model endpoints, a workload-specific draft-model lifecycle may justify investment; for lower-volume or privacy-sensitive systems, the operational and governance cost may exceed the serving gain. | Caveat: This requires permission to use production data, substantial storage and compute, movement of training data, and retraining or replacement of the speculator whenever the underlying target model changes.
  • Claim: The next inference stack will be shaped by lower-precision hardware, prefill/decode disaggregation, and training-designed serving components. | Evidence: Kiely expects strong NVFP4 performance on NVIDIA Rubin, cites early gains from PD disaggregation, and predicts that moving KV-cache data across the system will become increasingly important. | Implication: Infrastructure planning should account for communication fabric, cache locality, and software maturity—not just accelerator FLOPS—when evaluating next-generation GPU platforms. | Caveat: He explicitly labels the hardware and industry outlook as personal forecasting rather than nonpublic product knowledge.

Detailed Brief

Why training is becoming part of inference engineering

  • Claims: The traditional boundary in which training produces finished weights and inference merely serves them is becoming less useful.; Faster inference can create more usable data, which can improve models and specialized serving components, creating a reinforcing loop.
  • Evidence: STill requires training to create a learned compact representation of the KV cache.; Eagle 3 improved upon small-model speculative decoding by training an approximately one-billion-parameter draft model on target-model hidden states.; D-Flash and ongoing speculator retraining extend this training-for-serving pattern.
  • Caveats: The reinforcing data/training loop is easier for large operators with volume, compute, and authorized data access than for smaller teams.; Specialized serving models add versioning dependencies: changing the target model may invalidate or require retraining the draft model.
  • Implications: Serving performance can become a model-training and data-flywheel advantage rather than a commodity configuration exercise.; A mature inference platform should treat draft models, cache compressors, benchmark datasets, and target models as jointly managed artifacts.

Speculation method progression and the D-Spark watch item

  • Claims: The industry moved from conventional speculative decoding with small same-family draft models to draft heads such as Medusa, then hidden-state-trained draft models such as Eagle 3.; D-Spark is an early follow-on research direction that combines a diffusion speculator with a sequential speculator to attempt further acceptance-rate improvement.
  • Evidence: Kiely says small models proved better as small models than as draft-token generators, motivating alternative speculative architectures.; D-Spark had appeared only days before the talk; Baseten had D-Flash in production but no D-Spark production results.
  • Caveats: D-Spark is explicitly not production-validated in this presentation.
  • Implications: Track D-Spark as a research watch item rather than placing it into a near-term serving roadmap.; Acceptance rate remains the common metric linking otherwise different speculative-decoding architectures.

Notable Concepts & Terms

  • Inference engineering: The discipline of making model deployments fast, efficient, and reliable after or alongside training, especially through serving-system and model-aware optimization.
  • KV cache: Stored key/value attention states from prior tokens; it enables prefix reuse but becomes a major memory and data-movement burden at long contexts.
  • TurboQuant: A polar-coordinate approach to four-bit KV-cache quantization that improves cache memory efficiency but, according to Baseten's test, imposes unacceptable decode overhead for many datacenter deployments.
  • NVFP4: NVIDIA's four-bit floating-point format, positioned here as a practical production quantization path, particularly for model weights and future Rubin hardware.
  • STill: Baseten Research's trained KV-compaction approach that uses a Perceiver-style bottleneck to synthesize a compact memory representation from the full cache.
  • Speculative decoding: A lossless decoding optimization in which a cheaper draft mechanism proposes tokens and the target model verifies them, allowing multiple accepted tokens per target-model pass.
  • D-Flash: A diffusion-based speculative draft model that generates a block of tokens with bidirectional within-block attention, intended to raise acceptance rates and throughput.
  • PD disaggregation: Separating prefill and decode serving work so resources can be specialized and KV-cache state moved between them; Kiely expects this to grow in importance.

Operator Notes / Why Ken Should Care

  • Require end-to-end workload benchmarks before approving any KV-cache quantization change; include decode TPS, TTFT, long-context capacity, quality, and interconnect overhead.
  • For high-volume agent or API workloads, test D-Flash-style speculation against the actual prompt and output distribution, with accepted tokens per verifier pass as the primary success metric.
  • Assess whether production-data consent, retention, and tenant isolation permit continuous draft-model retraining before considering workload-specialized speculators.
  • Design model-serving architecture around KV locality and transfer paths—prefix affinity, cache sharing, CPU offload, and prefill/decode separation—rather than treating cache management as an implementation detail.
  • Maintain a research watchlist for trained KV compaction and D-Spark, but avoid treating either as production-ready based on this talk alone.

Source/Metadata

  • Title: What's New in Inference Engineering — Philip Kiely, Baseten
  • Transcript words: 3433
  • Duration seconds: 1149
  • Timestamp note: No timestamps or chapters were present in the supplied transcript.
Full transcript 3165 words · 15 min read
0:13

I'm here to talk about what is new in inference engineering. So, hi, I'm Philip, and I'm here because I wrote a book. This is my third year at the AI Engineer World's Fair. This is my favorite conference in the entire world. It's the highlight of the calendar every single year. I really got my start as a speaker and as an engineer here in 2024. I came back in 2025 and did a bunch of stuff. I'm here again. I love it here, and I'm very thankful to the organizers for always having me. I wrote this book called Inference Engineering. We published it three or four months ago, and I've just been overwhelmed by the response. We've done more than 11,000 paper copies.

0:56

We're coming up on 30,000 digital copies. And 24 million people around the world, or 24 million Twitter accounts, so we'll see how many people that actually is, have seen something out of inference engineering. And with all of this great reception, there has been one question that people have been asking me. Why in the world would you do this? Why would you write a book about something that's changing so fast? Well, I believe that a lot of the principles of inference engineering at this point have been pretty solidified, and there's a lot that we can learn and repeat over generation and generation of model.

1:23

But today, I'm here to talk about what's new in inference engineering. This is the first public addendum of all new information since the book came out. We are going to review the inference engineering principles a little bit, and then we're going to talk about all the stuff that's happened since February 23rd of 2026 in the inference world. We're going to talk about what happened to TurboQuant, talk a little bit about KV compaction. We're going to spend a lot of time on D-Flash and some other new exciting things in specular decoding.

1:49

And then, I'm going to do a little bit of prognosticating, a little bit of forecasting of what I think is going to happen in inference coming up here, and what I'm excited about, hopefully being able to talk about next time you guys see me up here. Cool, so let's get started. So, one thing I've been identifying now out of tons and tons of conversations with people about inference is a handful of shared principles. And one of the big ones, I was on this podcast the other day with Sarah. We were talking about inference, and there's two types of inference engineering that have really emerged.

2:18

There's local inference where the overwhelming strategy is to get it working on whatever hardware you have by squishing the model with quantization, distillation, pruning, however you can. Splitting it across whatever GPUs you happen to have in your house. And first you get it working, and then you make it less dumb. You take away whatever catastrophic issues all of this compression of the model has created, and you try and get it back to that baseline intelligence running at a batch size of one.

2:42

And then there's my world, which is the batch size and data center world, where it's get it working, just day zero, get the build of VLM up, get it working, and then make it less slow. Do stuff like kv-aware routing, speculation, disaggregation, and within these two worlds, I think that we have a lot to learn from each other. I am in this talk going to be focused on advances in data center oriented inference engineering, because that's what I know. But there's a lot of really cool stuff happening in the local world as well. So, in the book, in inference engineering, I generally assume that the weights are a finished product.

3:14

And I do think that the handoff from training to inference is an important one to keep in mind, and it's a good way of delimiting the space. However, what I've found more and more recently is that many optimizations for inference come from a dedicated training process. And so, the lines between training and inference are getting blurrier and blurrier. And that's an interesting thing to keep in mind. We're seeing this cycle where you get faster inference, which gives you more data, which you use to train a better model, which gives you faster inference, which gives you more data. And you just keep doing that until you're super rich.

3:57

So, with training for inference, we have a bunch of new techniques to talk about across what I like to call the big three. So, we're going to talk about some news in quantization, some news in caching, specifically the KV cache mechanism, and some news in speculation. Because these are the practical day-to-day techniques of how do I make X model faster, usually these are the three techniques that people are reaching for. So, first thing, I publish a book. It's February. I'm feeling awesome about myself. I'm like, wow, everything you need to know about inference in one place. And then, we had some news in the quantization world.

4:35

So, just as a quick review, quantization is when we use a smaller, less precise number format in order to save ourselves on bandwidth, save ourselves on compute, make TTFT better, make TPS better. It's usually hardware-specific, gives you cost savings, but potentially degrades model quality a little bit. And, by the way, if you want to hear my whole rant about quantization, I did a talk at AI Engineer Miami last month about how quantization is not necessarily as evil as it sounds, and that there's many things you can do to preserve quality through that process.

5:00

So, I was feeling good about my treatment of quantization, and then 20 million people saw TurboQuant, and in fact, it made the memory stock macro dip for a minute just because everyone was thinking, oh, memory's going to be so much more efficient now, we don't need any more flash memory, which was wrong. But, anyway, it was this new quantization approach that was popularized in March of this year that uses polar coordinates for quantization and allows you to quantize the KV cache down to four bits. And it was super hot, and I was thinking, oh, man, there's this whole thing that I left out, and what is this going to look like?

5:19

And so, our team did a bunch of research on this. Shout out to Ali from our model performance team at Waterloo. I'm not sure if he's still an intern, actually, but anyway, so he wrote this great piece about the math behind TurboQuant, and basically the benefit you get out of TurboQuant is that you get to represent the KV cache with four bits instead of eight bits. You save half the room and you get effectively double the bandwidth when you're moving KV cache around in your system memory.

5:38

But the drawback is pretty big for TurboQuant. It turns out that you need to do additional computation in the forward pass to account for this during decode, and it cuts TPS by more than half, and that's just an unacceptable trade-off for a lot of the production use cases. So, we took a good hard look at TurboQuant, but are not using it for any of these real workloads. We're still on the traditional NVFP4 quantization. That said, it actually is a great technique for the local inference folks.

6:01

So, if you are running a model, especially a long context language model on your local computer, on GPUs in your basement, you have a very limited amount of memory. That's the number one bottleneck. And so, anything that can free up memory from KV cache and allow you to put those longer sequences on there is going to be very valuable. And the additional forward pass computation is going to be less of a drawback. So, still, TurboQuant is a fantastic research paper, a really great technique that just ended up not being as applicable in the data center inference world as it might have first appeared.

6:30

Instead, we're focused on NVFP4 with a focus on quantizing the weights versus the KV cache. Doing our best to find rough edges in the quantized weights, make sure that we're not flattening out probability distributions. For the KV cache itself, focusing instead on KVAware routing, KV offloading, KV sharing, using Nickel and using NVIDIA Dynamo and other tools in order to move the KV cache around the system and potentially offload to CPU, ordinary memory, etc. Versus trying to use TurboQuant to compress it. And then we're also focused on quantization across modalities.

7:11

So, thinking about how can we apply the benefits of NVFP4 not only to language models, but also to image and video models. Ali also wrote a lot of great stuff on Twitter about that, which you should check out. So, that said, the KV cache is still very important. And let's talk about it. Let's talk about KV compaction. Doing our best to find rough edges in the quantized weights, make sure that we're not flattening out probability distributions.

7:57

For the KV cache itself, focusing instead on KVAware routing, KV offloading, KV sharing, using Nickel and using NVIDIA Dynamo and other tools in order to move the KV cache around the system and potentially offload to CPU, ordinary memory, etc. Versus trying to use TurboQuant to compress it. And then we're also focused on quantization across modalities. So, thinking about how can we apply the benefits of NVFP4 not only to language models, but also to image and video models. Ali also wrote a lot of great stuff on Twitter about that, which you should check out. So, that said, the KV cache is still very important. And let's talk about it. Let's talk about KV compaction.

9:06

Again, quick review, KV cache, if you put in the same prompt with the same prefix, you get to reuse the tokens that you calculated pre-fill last time. That makes your whole system faster and more efficient. Broadly, KV cache is lossless memory. There's only a couple of sources of lossless memory when we think about our inference system. We have the content of the prompt, the context, you have the KV cache. And that's going to scale linearly with the amount of data you pass in. And now, if you're thinking about million token sequence lengths, that actually gets substantial. So, a lot of people are thinking about how do you compress memory? How do you compress context?

9:45

Agent harnesses will compress context. RAG, search, all these techniques that we've been talking about for years. Are a compression of a larger context into something that you can give to a model. You can write to files. All of these things scale sublinearly with the amount of data that you have. But what if there was a middle road? What if there was a way where you could get quite a bit of compression in the data that you were remembering with near lossless information retention? So, we have a lot of different ways that we can think of what to keep in the cache. Recent compaction methods have shown that we can replace the cache with a much shorter one.

10:20

We've got papers like attention matching and cartridges that have given really promising outcomes here with high compression ratios. But both of these are run at inference time. Again, one of the techniques I want to talk about or one of the themes I want to talk about is training for inference. So, in this case, I want to introduce something called STill by the base 10 research team where the synthesis on top of the cache where we're keeping a learned representation of the information. Rather than the information directly or a deterministic subset of it is amortized via training.

10:42

So, Charlie and Mudith from our post training team did a fantastic chalk talk at COSYNE recently. It's up on YouTube. I would encourage you to take a look at it if you're interested in learning about KV compaction. I do not unfortunately have the time or the genius to explain everything up here. But the basic mechanism is that STill is a perceival bottleneck that takes a fixed set of learned query vectors, cross attends it against the full KV cache, and produces a set of compact keys and values in a single forward pass. This creates a fast differentiable compressed memory that the LLM can attend to as if it was real context.

11:22

So, if you're interested in KV compaction, definitely check out Charlie and Mudith's work. It's been a fantastic thing to learn about. So, that's two of the techniques. We've talked about quantization. We've talked about caching. The final one is speculation. And there's been a lot of change here. As a review, speculative decoding, we're going to use draft tokens. We're going to verify them during the forward pass. And we're going to use that to generate more than one token per forward pass. It helps a lot with tokens per second. And it is a fully lossless optimization, which is great because we don't have to worry about quality at all.

12:20

Now, in the history of speculation, we started with speculative decoding. All of these are in the book. You have spec deck where you use a small model from the same family to generate draft tokens. It turns out small models are not great draft token generators. They're great small models. So, we invented as an industry a bunch of new methods like Medusa where maybe you add draft heads to the model. And then eventually Eagle 3, which was, what if instead of taking a tiny model from the same family, we actually train a billion parameter model on the hidden states of the target model to generate draft tokens. And that actually worked really well.

12:49

And so, as of maybe February of this year, Eagle 3 was the best method in speculation. Now we got D-Flash. D-flash is even better. So, it's diffusion for speculation. D-flash creates a sequence of draft tokens instead of a single token. So, the model is a diffusion language model, which means it creates a whole sequence of tokens in the same way that a video or image generation creates a sequence of frames or a sequence of pixels and iterates over it rather than doing an autoregressive token generation. D-flash models might be two or four times slower to run, but they're going to predict eight or 16 tokens at once in that window, while Eagle is only doing one at a time.

13:19

So, a single D-flash forward pass is faster than the entire Eagle draft phase and predicts more tokens. These tokens are able to cross-attend to each other and generally create a higher acceptance rate. Because in speculation, acceptance rate is everything. So, in the wild, we're seeing a more than 3x improvement from D-flash. This is measured with a single B200, Qwen38B. And we can see it versus Eagle. It's a substantial improvement in the token acceptance rate and the tokens per second. These D-flash models are trained with an attention mask for bidirectional drafting. So, the target model is going to provide the context.

14:13

And within each block, we're going to have a subset of clean tokens that are sampled. And the attention mask is going to enforce causal consistency. And the way we're going to see is still going to allow for bidirectional attention. Where in a traditional autoregressive model, you're only looking at the tokens in a single direction. So, that's why we're able to take advantage of this diffusion-based architecture. And then I thought I was done. And then, a couple days ago, D-Spark came out. Now, D-flash, we do have up and running in production. D-Spark is new research. So, this one, I can basically only say, it exists. It's cool. We're looking at it.

15:11

The difference versus D-flash, it still has that diffusion model. But it also pairs it with a sequential model. And the idea is that we're going to improve acceptance rates by having these two models work together. Rather than having just the iterative speculator, just the diffusion speculator, or just the autoregressive speculator. So, D-Spark, very exciting. But we don't have any production results with it yet to share. Hopefully, we'll have those for next time. What we do have production results on is continuous speculator retraining. So, this is, we're back to D-flash here.

15:36

And this is the idea that speculative decoding is very dependent on the actual content of the prompts and responses that you're looking at in your system. And so, if you are continuously retraining on those prompts and responses in your live system, you can see a 20% to even 2x improvement in your token acceptance rates. This is actually really hard to do. It takes a lot of storage. And you have to make sure that you have permission to use the data that you're processing in this way. It takes a ton of compute. And you have to move all of this information around. And if you change the underlying model, you also have to change the speculator model.

16:11

But when I look forward into the future, I do think that continuous speculation for very large-scale systems is going to be a worthwhile optimization. So, what is next in inference? The following is personal opinion and speculation and public information. And if I knew anything that was actually coming out, I wouldn't be able to talk about it. So, this is just what I think is going to happen. I've been through three hardware cycles through the Ampere release, the Hopper release, the Blackwell release.

16:36

And it always takes time for when these chips get shipped to when they get installed in data centers when the entire software stack really is able to take advantage of their capabilities. But some things that I'm excited about are with Rubin, it looks like the NVFP4 performance is going to be fantastic. So, the more we can honestly borrow techniques from local inference Is going to be a worthwhile optimization. So, what is next in inference? The following is personal opinion and speculation and public information. And if I knew anything that was actually coming out, I wouldn't be able to talk about it. So, this is just what I think is going to happen.

17:07

I've been through three hardware cycles through the Ampere release, the Hopper release, the Blackwell release. And it always takes time for when these chips get shipped to when they get installed in data centers when the entire software stack really is able to take advantage of their capabilities. But some things that I'm excited about are with Rubin, it looks like the NVFP4 performance is going to be fantastic. So, the more we can honestly borrow techniques from local inference and get a lot of confidence running models in this NVFP4 data format, the more we're going to be able to take advantage of the awesome performance of the upcoming Rubin systems.

17:47

I think that disaggregation and system-wide communication is going to be increasingly important. We're seeing really excellent early gains from PD disaggregation. And the ability to move information like KV cache data around the system is going to be increasingly important. And then as I said, the theme of training for inference is going to be something that continues to have a big impact in the industry moving forward. So, thank you all so much for the talk, for coming to the talk. I'm on Twitter, I'm on LinkedIn, and I'm giving out free books. You can download a PDF at the QR code or come down with me to the Base 10 booth to get your free copy of Inference Engineering.

18:31

We've got a bunch there, maybe enough for everyone. If not, we will have a career bring some more from the office. So, yeah, I'll be downstairs at the Base 10 booth. Thank you all so much and have a great day. I'll see you now. I'll see you now.

19:07

I'll see you now. I'll see you now. Thank you. I'll see you now. Thank you.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note