Text Diffusion — Brendon Dillon, Google DeepMind
Description
GPT-4o answered 40. Gemini 2.5 Flash answered 42 and stuck to it even after working through the reasoning incorrectly. The Gemini Diffusion model, considerably smaller than both, answered 60 on the first forward pass, then 49, then corrected itself to 39 once it finished reasoning. Bidirectional attention means it can see future tokens and go back to fix mistakes. Autoregressive models cannot do that. Brendon Dillon covers why text diffusion is fast (24 denoising steps to generate 256 tokens means roughly 10x fewer memory transfers than autoregressive generation), what the tradeoff is (lower throughput at large batch sizes makes it expensive to serve at scale today), and what gets unlocked when latency drops to 2,000 tokens per second. The demos include a fake Wikipedia generated on the fly, a Reddit clone with AI generated comments and images, an operating system where every click generates the next screen, and a todo app built in 15 seconds by voice.
Summary
Generated by claude-sonnet-4-5At-a-Glance
- Verdict: Watch fully
- Core thesis: Text diffusion models achieve 10x lower latency than autoregressive models by iteratively refining entire token sequences in parallel instead of generating one token at a time, unlocking new real-time applications despite higher serving costs
- Why it matters: Text diffusion represents a fundamental architectural alternative to autoregressive generation with concrete latency advantages (2,000 tokens/sec vs standard GPT speeds) and unique capabilities like bidirectional self-correction and adaptive computation—important for on-device, robotics, and real-time interactive applications
- Best use: Watch the full technical explanation and demos to understand how text diffusion works, why it's memory-bound-optimized, and where it fits strategically (on-device/robotics vs cloud serving); use the detailed brief to extract architecture trade-offs and deployment considerations
Executive Summary
Brendan Dillon from Google DeepMind presents text diffusion as a forward-looking alternative to autoregressive language models. Unlike GPT/Gemini models that generate one token at a time, text diffusion starts with a canvas of random noise tokens and iteratively refines the entire sequence over multiple forward passes (typically 4-31 steps). This approach achieves dramatically lower latency—Gemini Diffusion hit 2,000 tokens/second versus standard autoregressive speeds—because it's optimized for memory-bound hardware. GPUs/TPUs have high compute (flops) but limited memory bandwidth; autoregressive models stream weights for every single token, while diffusion streams weights once per denoising step across hundreds of tokens, achieving roughly 10x fewer memory transfers when generating 256 tokens in 24 passes.
Gemini Diffusion (released as a 100k-user research preview one year prior to this talk) matched Gemini 2.0 Flash quality across benchmarks with superior code performance and much better latency. However, text diffusion has a critical disadvantage: lower throughput for large batches. Autoregressive models serve many users simultaneously by batching queries; diffusion's multiple forward passes hit compute limits earlier, making it more expensive to serve at cloud scale. This throughput concern is why no major model has shipped text diffusion into production cloud LLMs despite the latency win.
Beyond speed, text diffusion enables bidirectional attention within the token canvas—models can see future tokens and self-correct mistakes. Dillon demonstrates a math problem where Gemini Diffusion initially guesses 60, then 49, but after completing the reasoning (step 3), fixes the answer to 39. Both GPT-4o and Gemini 2.5 Flash (much larger models) failed this same prompt at the time. The model also performs adaptive computation: easy queries (100 digits of pi) finish in 4 steps; harder ones (quantum mechanics paragraph) take 31 steps. Harder evals like GPQA Diamond required more denoising steps, while easier MBPP coding tasks finished quickly.
Text diffusion also supports fast in-place editing (like image inpainting): you can mask tokens and prompt the model to fill gaps using surrounding context. Demos included fully generated Wikipedia pages (HTML + text on-the-fly), fake Reddit threads with images (using Imagen), a functioning OS generated per-click, and voice-driven vibe coding that created a dark-mode to-do app in 15 seconds. These demos illustrate how sub-100ms latencies unlock real-time interactive experiences impossible with autoregressive models. Dillon positions text diffusion for on-device and robotics use cases where single-user latency matters more than multi-user throughput, and hints at an upcoming release from the DeepMind team.
Key Takeaways
- Claim: Text diffusion models achieve roughly 10x lower latency than autoregressive models for the same sequence length by doing fewer memory transfers | Evidence: Generating 256 tokens in 24 denoising steps transfers weights 24 times vs 256 times for autoregressive; Gemini Diffusion hit 2,000 tokens/second including pre-fill overhead. GPUs/TPUs are memory-bound (limited bandwidth, high flops), so streaming weights once per step across many tokens is far more efficient than streaming per token. | Caveat: Latency advantage decreases for very short outputs (1-token responses are pre-fill dominated). The 2,000 tokens/sec figure is a year old and included pre-fill; newer models and batch sizes will shift this number. Throughput (multi-user serving cost) is worse than autoregressive. | Implication: For single-user, on-device, or real-time applications where latency is paramount and batch throughput is irrelevant, text diffusion is architecturally superior. Ken should consider text diffusion for agent systems requiring sub-100ms response times or on-device inference where you control the deployment environment. | Timestamp: 03:45
- Claim: Text diffusion models can self-correct answers by seeing future tokens and revising earlier mistakes during iterative refinement | Evidence: Demo: math problem 'sqrt(81)*2/3^2+…'. Step 1: model guesses answer=60. Step 2: changes to 49. Step 3: completes reasoning, gets 39, goes back and fixes answer token to 39. GPT-4o said 40 (later apologized), Gemini 2.5 Flash said 42 and stuck with it ('36+3=42')—both larger models than Gemini Diffusion. | Caveat: Self-correction only works within the denoising budget (number of steps). If the model runs out of steps before catching the error, it won't fix it. Also, this is a cherry-picked example from Google I/O; unclear how often self-correction succeeds across diverse prompts. | Implication: Bidirectional reasoning could reduce hallucinations and math errors in agent workflows. Ken could use text diffusion for tasks requiring internal consistency checks (code generation, structured data) where the model can 'look ahead' and validate outputs. This is a qualitative capability autoregressive models fundamentally lack unless you add separate reasoning steps. | Timestamp: 09:12
- Claim: Text diffusion models perform adaptive computation: they spend more denoising steps on harder problems and fewer on easier ones, as determined by the model itself | Evidence: Easy: 100 digits of pi = 4 steps (100 tokens, trivial memorization). Medium: FizzBuzz code = 18 steps. Hard: quantum mechanics paragraph = 31 steps. Across evals, GPQA Diamond (hardest for that model size) took longest; MBPP (basic Python) took least time. The model decides when to stop, not a fixed step count. | Caveat: Adaptive computation requires training the model to predict its own stopping criterion, which adds training complexity. Dillon doesn't detail how the stopping mechanism is trained or whether it's monotonic (could waste steps on easy tasks if poorly calibrated). | Implication: Adaptive compute could optimize inference cost for heterogeneous workloads—allocate more compute to complex reasoning, less to boilerplate. For Ken's agent systems, this means you could serve a mix of trivial and hard queries without manually tuning per-query budgets. However, you'd need access to the stopping-head training methodology to replicate this. | Timestamp: 12:30
- Claim: Text diffusion models enable fast in-place editing by filling masked tokens using bidirectional context, analogous to image inpainting | Evidence: Demos: 'fix bug in code' → model edits only the buggy indices; 'add documentation' → inserts docstrings in-place; 'add middle paragraph to story' → fills gap consistent with first and third paragraphs. Unlike autoregressive, which generates left-to-right, diffusion can attend to surrounding tokens and insert coherently. | Caveat: Dillon only shows brief demos without failure cases. Unclear how well in-place editing works for large edits, conflicting constraints, or ambiguous prompts. Also, this requires the model to know where to edit, which may need explicit masking or instructions. | Implication: In-place editing could streamline code refactoring, document revision, or structured data updates in Ken's workflows. Instead of regenerating entire responses, you could mask and re-fill specific tokens. This is a UX/product lever: lower latency + surgical edits = real-time collaborative editing experiences. | Timestamp: 14:05
- Claim: Text diffusion's primary deployment blocker is lower throughput for large batches, making it too expensive to serve at cloud scale despite quality parity | Evidence: Gemini Diffusion matched Gemini 2.0 Flash quality (slight code advantage, slight disadvantage elsewhere) but couldn't land in production due to throughput concerns. Autoregressive models batch many users; each does one token at a time, but TPU utilization is high. Diffusion does multiple passes per user, hitting compute ceiling earlier even though per-user latency is lower. 'No one's landing text diffusion into big models primarily because of that disadvantage—too expensive to serve.' | Caveat: This is a serving economics problem, not a fundamental model capability problem. Costs could change with hardware improvements (more bandwidth), better KV-cache techniques, or hybrid architectures. Dillon is speaking from a year-old research preview; current economics may differ. | Implication: Ken should not expect text diffusion in cloud-served GPT/Gemini competitors until serving costs drop or use cases justify the premium. However, for on-device (phones, robots) or single-user applications where you control the hardware, text diffusion is viable and preferable. If Ken builds agent infrastructure that runs locally or in low-concurrency environments, text diffusion is a strategic architectural choice. | Timestamp: 06:20
- Claim: Low-latency text diffusion unlocks new real-time interactive applications that are infeasible with autoregressive models | Evidence: Demos: (1) Wikipedia generated on-the-fly (HTML + text, clicks feel instant). (2) Fake Reddit with text + images (Imagen integration, slightly slower image gen). (3) Entire OS generated per-click (desktop, readme, web pages). (4) Voice vibe-coding: user says 'create to-do app, add 10 random to-dos, dark mode'—15 seconds total, everything working. These experiences require sub-100ms token generation to feel real-time. | Caveat: Demos are curated and optimized. Real-world performance with complex prompts, longer contexts, or edge cases is unknown. The vibe-coding demo is from an external user (Twitter), so it's not fully controlled. Also, image generation (Imagen) was slower, showing multi-modal latency still bottlenecks some demos. | Implication: Ken should explore text diffusion for agent UX requiring real-time feedback loops: live code editing, interactive data exploration, conversational agents with <100ms responses. The OS and Wikipedia demos show that generative models can simulate entire environments if latency is low enough—potential for agent-driven interfaces, synthetic user testing, or on-the-fly content generation in Ken's content/business investing work. | Timestamp: 15:45
Detailed Brief
How Text Diffusion Works and Why It's Faster
- Claims: Text diffusion starts with ground-truth tokens, adds noise (random token replacements) at multiple levels, trains a model to denoise (correct mistakes), then at inference initializes with pure noise and iteratively refines to clean text; Autoregressive models generate one token at a time, conditioning on prior tokens; diffusion generates entire blocks (256+ tokens) in parallel but over multiple forward passes (4-31 steps); GPUs/TPUs are memory-bound: high compute (flops) but limited bandwidth between memory (HBM) and tensor cores; autoregressive streams weights+KV cache for every token, diffusion streams once per denoising step across all tokens; For 256 tokens in 24 steps, diffusion does 10x fewer memory transfers than autoregressive (24 vs 256), yielding roughly 10x speedup if truly memory-bound
- Evidence: Gemini Diffusion hit 2,000 tokens/second including pre-fill; standard autoregressive models (GPT-4o, Gemini Flash) are far slower per-user; Dillon explains bandwidth bottleneck with chip architecture diagram: tensor core does matrix multiplies, HBM stores weights/activations, bandwidth channel is tight; fewer transfers = higher utilization; Longer sequences benefit more: 1-token responses are pre-fill dominated (no latency win); 256+ token responses exploit parallel generation fully
- Caveats: Latency advantage only applies when memory-bound; if batch size is large enough to saturate compute, advantage shrinks; Multiple forward passes mean higher total FLOPs per output token; it's a bandwidth-for-compute trade-off; Dillon's 2,000 tokens/sec figure is a year old and model/hardware-specific; not a universal benchmark
- Implications: Text diffusion is ideal for on-device, robotics, or single-user real-time applications where latency >> throughput; Ken should consider text diffusion for agent systems requiring sub-second response times and not serving thousands of concurrent users; Hybrid architectures (autoregressive for cloud, diffusion for edge) may emerge as the strategic deployment pattern
Bidirectional Reasoning and Self-Correction
- Claims: Autoregressive models use causal attention (past-only); text diffusion uses bidirectional attention (future tokens visible); During iterative refinement, diffusion models can generate an answer, complete reasoning, see the answer was wrong, and revise it; Google I/O demo: sqrt(81)*2/3^2+… = 39. Gemini Diffusion guessed 60 → 49 → completed reasoning → fixed to 39. GPT-4o guessed 40, Gemini 2.5 Flash guessed 42 (never corrected); Both GPT-4o and Gemini Flash were larger models but made errors due to autoregressive constraint (commit to answer before reasoning finishes)
- Evidence: Step-by-step screenshots showing token changes over 3 forward passes: answer=60 (blue, unfixed) → 49 → 39 (final); Reasoning completes '36 + 3 = 39' after which the model revises the initial answer token; Dillon notes modern reasoning models (o1-style) can fix this by punting to separate thinking steps, but text diffusion does it natively
- Caveats: Cherry-picked demo; unclear how often self-correction succeeds across diverse tasks; Self-correction only works within the denoising budget; if steps run out before error is caught, it won't fix; Bidirectional attention doesn't guarantee correctness—it just allows revision; the model must be trained to exploit this capability
- Implications: Fewer hallucinations and math errors in tasks requiring internal consistency (code, structured outputs); Ken could use diffusion for agent workflows where 'thinking ahead' and revising earlier tokens is valuable (multi-step planning, self-debugging code); Qualitative advantage over autoregressive; not just speed but reasoning topology is different
Adaptive Computation and Dynamic Steps
- Claims: Text diffusion models can be trained to decide when they are finished, spending fewer steps on easy tasks and more on hard tasks; Model quality increases monotonically (roughly) with more denoising steps across all evals; you can trade compute for quality at inference; Gemini Diffusion examples: 100 digits of pi = 4 steps (easy memorization); FizzBuzz = 18 steps; quantum mechanics paragraph = 31 steps; Harder evals (GPQA Diamond) took more time; easier evals (MBPP basic Python) took less time, as determined by the model itself
- Evidence: Dillon shows curves: quality vs denoising steps is monotonic upward across six coding evals; Adaptive stopping: model outputs when ready, rather than fixed step count; easy tasks return faster, hard tasks take longer; 100 digits of pi = 100 tokens but only 4 steps; autoregressive would do 100 steps (one per token)
- Caveats: Training the stopping mechanism adds complexity; Dillon doesn't explain how the stopping head is trained or how reliably it generalizes; Monotonicity is 'roughly' true but not strict; occasional quality drops at higher steps suggest noise or overfitting; Adaptive compute requires infrastructure to handle variable-length inference; serving systems must support dynamic budgets
- Implications: Could optimize inference cost for heterogeneous workloads: allocate compute where needed, save where not; For Ken's agent systems, this means you can serve mixed-difficulty queries without manually tuning per-query step counts; Potential to build 'effort-aware' agents that dynamically decide how hard to think based on prompt complexity
In-Place Editing and Use Cases
- Claims: Diffusion models (image or text) can fill in masked regions using surrounding context, enabling surgical edits; Text diffusion can fix bugs in code, add documentation, insert paragraphs in stories—all by masking tokens and re-generating coherently; Unlike autoregressive (left-to-right), diffusion attends bidirectionally so inserted text is consistent with both prefix and suffix
- Evidence: Code demo: 'fix bug' → model edits only the buggy index line; 'add documentation' → inserts docstrings in-place; Story demo: 'add middle paragraph' → model fills gap consistent with first and third paragraphs; Analogous to image inpainting: cut out part of image, model fills using context
- Caveats: Demos are brief and curated; unclear how well it handles large edits, conflicting constraints, or ambiguous masking; Requires explicit masking or instructions to know where to edit; not automatic error detection; No failure cases shown; real-world editing may be less clean
- Implications: Streamlines code refactoring, document revision, structured data updates—surgical edits instead of full regeneration; Could enable real-time collaborative editing UX: user highlights text, model fills/fixes instantly; For Ken's workflows: faster iteration on generated code, content drafts, or agent outputs by masking and re-generating specific spans
Deployment Trade-offs and Strategic Fit
- Claims: Text diffusion has lower latency but worse throughput for large batches, making it too expensive for cloud serving at scale; Autoregressive models batch many users; each generates one token at a time, but TPU utilization is high; diffusion hits compute limits earlier despite per-user speed; Gemini Diffusion matched Gemini 2.0 Flash quality (slight code advantage) but didn't ship to production due to throughput concerns; Text diffusion is ideal for on-device (phones, robots) where single-user latency matters and throughput is irrelevant
- Evidence: Dillon: 'No one's landing text diffusion into big models primarily because of that disadvantage—too expensive to serve, even if much lower latency'; Gemini Diffusion was a 100k-user research preview, not a production release; DeepMind uses text diffusion in on-device applications and robotics internally (within Alphabet ecosystem)
- Caveats: This is a year-old snapshot; serving economics and hardware may have evolved; Throughput concerns could be mitigated with better KV-cache techniques, hybrid architectures, or cheaper bandwidth chips; Quality parity was claimed a year ago; current autoregressive SoTA (GPT-4.5, Gemini 2.0) may have widened the gap
- Implications: Ken should not expect text diffusion in cloud LLMs (ChatGPT, Gemini) until economics change or niche use cases justify premium pricing; For on-device agents, edge inference, or robotics, text diffusion is strategically superior—Ken should prioritize it for local/low-concurrency deployments; Hybrid approach: autoregressive for cloud serving, diffusion for real-time edge; Ken could architect agent systems to route tasks accordingly
Demos and Real-Time Application Unlocks
- Claims: Low latency (sub-100ms token generation) unlocks real-time interactive experiences infeasible with autoregressive models; Internal demos: Wikipedia generated on-the-fly (HTML + text, instant clicks); fake Reddit (text + images via Imagen); entire OS generated per-click; voice vibe-coding (to-do app in 15 seconds); External demo (Twitter user): voice-driven code generation with dark mode conversion in 15 seconds of work
- Evidence: Wikipedia demo: user clicks link, page generates instantly (HTML, text, links all on-the-fly); Reddit demo: generates fake comments and images (Imagen slower, follows after text); user can interact as if real site; OS demo: every click generates next screen/page; desktop, readme, web pages all generated; Vibe-coding demo: user says 'create to-do app, add 10 random to-dos, mark 4 completed, sort by name/state, dark mode'—everything works in 15 seconds
- Caveats: Demos are curated and optimized; real-world performance with complex prompts or edge cases unknown; Image generation (Imagen) was slower, showing multi-modal latency still bottlenecks some experiences; Vibe-coding demo is from external user (not DeepMind-controlled), so replication/reliability unclear
- Implications: Text diffusion enables agent-driven interfaces, synthetic environments, real-time code editing—new product categories; Ken should explore diffusion for interactive agent UX: live data exploration, conversational agents with <100ms responses, on-the-fly content generation; Potential for synthetic user testing, generative environments (OS, websites) for agent training or rapid prototyping in Ken's content/business work
Notable Concepts & Terms
- Text Diffusion: Generative model architecture that starts with noisy tokens, iteratively refines them over multiple forward passes (denoising steps) instead of generating one token at a time autoregressively; analogous to image diffusion but for discrete text tokens
- Memory-Bound Inference: Bottleneck where GPU/TPU performance is limited by bandwidth (data transfer between memory and compute cores) rather than FLOPs; autoregressive models are memory-bound because they stream weights for each token, diffusion reduces transfers by processing many tokens per step
- Bidirectional Attention: Transformer attention mechanism where tokens can attend to both past and future tokens in the sequence (vs causal/autoregressive attention which only sees past); enables text diffusion to self-correct by revising earlier tokens after seeing later context
- Adaptive Computation: Model capability to dynamically decide how many denoising steps to spend on a task based on difficulty; easy prompts finish in fewer steps, hard prompts take more steps, as determined by the model's stopping criterion
- In-Place Editing / Inpainting: Diffusion technique where you mask (remove) specific tokens and regenerate them using surrounding context; analogous to image inpainting; enables surgical edits (fix bugs, add docs) without regenerating entire sequence
- Gemini Diffusion: DeepMind's research preview text diffusion model (100k users, released ~1 year before this talk); variant of Gemini 2.0 Flash architecture using text diffusion instead of autoregressive generation; achieved 2,000 tokens/sec with quality parity
- Denoising Steps: Number of iterative refinement passes the diffusion model performs to transform pure noise tokens into clean output; typically 4-31 steps for Gemini Diffusion depending on task complexity; analogous to diffusion sampling steps in image models
- Throughput vs Latency Trade-off: Core deployment tension: text diffusion has lower per-user latency but worse multi-user throughput (can't batch as efficiently as autoregressive); cheap for single-user (on-device), expensive for cloud serving at scale
- Blockwise Autoregressive: Hybrid generation strategy where diffusion model fixes a window length (e.g., 512 tokens), generates that block with diffusion, then autoregressively generates next block; allows unlimited text generation while maintaining diffusion benefits per block
Operator Notes / Why Ken Should Care
- Text diffusion is architecturally distinct from autoregressive LLMs and offers 10x latency improvements for on-device/single-user scenarios—critical for real-time agent UX, robotics, and edge inference where Ken controls deployment
- Bidirectional self-correction could reduce hallucinations and improve code/structured output quality; Ken should test diffusion models for agent workflows requiring internal consistency checks or multi-step planning
- Adaptive computation (variable steps per task) could optimize inference budgets in heterogeneous agent workloads; useful if Ken serves both trivial and complex queries without manually tuning compute
- In-place editing enables surgical revisions (fix bugs, add docs) without full regeneration—potential for real-time collaborative editing, faster iteration on generated outputs, or agent self-debugging loops
- Low latency unlocks new product categories: real-time code editing (vibe coding), generative environments (OS, Wikipedia on-the-fly), interactive agent UX (<100ms responses); Ken should explore these for content/business use cases
- Throughput blocker means text diffusion won't appear in cloud LLMs (ChatGPT, Gemini) until serving economics change—don't expect API access soon, but watch for on-device releases (phones, robots) where it's viable
- DeepMind plans to 'release something soon' per Dillon; Ken should monitor for Gemini Diffusion successor or API access to test latency/quality for agent systems and investing workflows
- For Ken's GTM/content work: text diffusion could power synthetic user testing (generated environments), rapid prototyping (instant Wikipedia/Reddit/OS generation), or real-time content creation with sub-second feedback loops
- Hybrid strategy: use autoregressive for cloud batch serving, diffusion for on-device or low-concurrency agent tasks; Ken could architect systems to route accordingly based on latency vs throughput needs
Watch Map
- 00:00: Intro: speaker credentials, text diffusion overview; explains noise addition and iterative refinement process
- 01:30: Gemini Diffusion research preview: 100k users, quality parity with Gemini 2.0 Flash, better latency (year-old numbers)
- 03:00: Autoregressive vs diffusion generation: one token at a time vs entire block over multiple passes
- 03:45: Pros: faster inference (2,000 tokens/sec), bidirectional attention, adaptive computation, in-place editing
- 04:30: Cons: lower throughput for large batches, higher serving cost; why it's not in production cloud LLMs
- 05:15: GPU/TPU memory-bound architecture: bandwidth bottleneck, why diffusion does fewer transfers (10x speedup explanation)
- 07:45: Latency deep-dive: streaming weights once per step vs once per token; flops vs bandwidth trade-off
- 09:12: Bidirectional reasoning demo: sqrt(81)*2/3^2 = 39; Gemini Diffusion self-corrects 60→49→39, GPT-4o/Gemini Flash fail
- 11:00: Dynamic computation: quality increases with more steps; model decides when to stop based on task difficulty
- 12:30: Adaptive computation examples: 100 digits of pi (4 steps), FizzBuzz (18 steps), quantum mechanics (31 steps); harder evals take longer
- 14:05: In-place editing demos: fix code bugs, add documentation, insert story paragraphs using bidirectional context
- 15:45: Low-latency application demos: Wikipedia on-the-fly, fake Reddit, generated OS, voice vibe-coding (to-do app in 15 sec)
- 18:00: Q&A: training data (same as autoregressive), distillation, adaptive stopping, window length for long text, on-device use cases, RL compatibility
- 22:30: Q&A continues: setting step limits, multi-window attention, hybrid architectures, discrete vs latent diffusion, throughput outlook
Source/Metadata
- Title: Text Diffusion — Brendon Dillon, Google DeepMind
- Transcript words: 5550
- Duration seconds: 1683
- Timestamp note: Timestamps extracted from transcript speaker tags and estimated from 1683-second duration; specific MM:SS aligned to topic transitions
Transcript
[SPEAKER_02] Some people are still filtering into the room. It's mostly intro stuff for the first couple of sites, so they won't miss anything. Okay, welcome everybody. My name is Brendan. I'm a research scientist at DeepMind. I'm talking today about text diffusion, which is a more forward-looking research area at DeepMind. So you're probably familiar with image and video diffusion, which is state-of-the-art for these modalities right now, where you take ground truth, say image, you add noise to it in training, and then you train a neural network to remove that noise gradually, and then at inference time you just initialize the picture with pure noise, and then you iteratively refine out the noise to recover back to whatever image or video or audio or whatever you're looking for. And the principle is essentially the same for text, for text diffusion, where you start with a clean sequence of tokens, like a clean sentence or something like that, and then you gradually add noise. You corrupt it somehow. There's lots of different ways to do that. You can do it in a continuous or discrete way, but let's just say discrete for now, which would in this case just mean adding random tokens or replacing tokens with other random tokens. And you do that for a bunch of different noise levels, and you train the neural network to try to fill in, to try to correct the mistakes basically in the text. And then at inference time, you initialize the sequence of tokens to just pure noise, like pure random discrete tokens from the vocabulary, and then you iteratively refine through that to fill in the information in the order that the neural network wants to do to recover back to say a clean sentence. And then in practice, so I showed you some GIFs here of what it looks like for images. You get very similar looking outputs for text where it starts off all noisy and then gradually fills in the text and you get clean, relatively clean outputs at the end. Okay, so the team I'm on, we had a research demo release one year ago now called Gemini Diffusion, which was a variant of a Gemini model with text diffusion instead of autoregressive next token generation. And that was a research preview that was open to about 100k people. And we're still keeping posted for new developments in that direction upcoming soon. And we did have some good numbers at the time, but again, it's a year ago, which is prehistoric times in this field. Our main comparator model was Gemini 2.0 Flashlight at the time, because that was the architecture we were branching from the text diffusion model. And we basically had very similar quality across the board there. Mostly a little bit of advantage in code, a little bit of disadvantage in some other areas, but relatively similar performance at much better latencies. But again, this is a year ago, so I wouldn't fixate too much on these numbers. Okay, so what's the difference between autoregressive generation and diffusion? So in the standard vanilla Gemini, Gemma, GPT, whatever, generation of text, you do this: you have some context that comes in and you want to generate some response to that. And the model does it one token at a time. So it generates the first token and then conditions on that generates the next token and so on. Whereas in diffusion, you have a context that will come in, whatever that is, and it'll initialize, like I mentioned, a long sequence of tokens, could be hundreds, could be thousands, could be shorter, depends on the model, to be random noise. And then it iteratively refines that canvas to remove the noise over the course of a few denoising steps. So rather than one token at a time, it does the entire block together, but over a couple of iterations. So it's not just one pass, it does multiple passes, but it gets to attend to the future tokens and so on. So it's a different way of generating text. So that obviously has some pros and cons. So the main pro that people really like, and probably is the biggest advantage that text diffusion models have, is that it's faster inference, it just generates faster tokens per second, because it makes much better use of the hardware, the TPU and the GPU. And I have some slides on that to explain why. But some other advantages are it can do bidirectional attention within this canvas of tokens. So autoregressive models can only attend to the past. They have causal attention within their transformer, whereas a text diffusion model is not restricted. They can attend to the future. And that has some interesting properties, like it can do self-corrected generation based on future tokens. So they could do some reasoning, see that I got the answer incorrect, and then go back and fix the reasoning and do it again. I have a demo of that. Because this process is iterative, it does a number of steps to respond. That means the model can actually do adaptive computation. It turns out that you can train the model to spend more time on harder problems and less time on easier problems. And the diffusion models in general, you can do things in-place editing, where you say, fix the last tokens and give me the prefix that corresponds to those tokens and stuff. But the main disadvantage it has, and the reason why it's not used everywhere right now, is lower throughput for large batches. So autoregressive models, they're slow, but you can have a big batch of queries together. And then if you push that through the neural network on the GPU, each individual user is slow, but you make good use of the TPU by doing that. And so you can serve a lot of queries. And so you keep your costs down. You can serve a lot. Whereas since text diffusion does multiple forward passes on the same data, it hits a compute threshold basically earlier. And even though it's lower latency for any one user, it tends to be lower throughput overall. So higher cost to serve. And right now, if you've played with Claude recently, you'll know that they have some throughput concerns. So people really care about throughput right now. And so no one's landing text diffusion into any of these big models, primarily because of that disadvantage. It's just too expensive to serve, even if it is much lower latency. Okay, so just leaning into why it doesn't have lower latency, in case you're not familiar with the architecture of how GPUs and TPUs run today. So in a GPU, there's a tensor core, which does these big matrix multiplies. It's very efficient, has a lot of flops or hops or whatever. And then the memory that sits on the TPU GPU, this HBM, So people really care about throughput right now. And so no one's landing text diffusion into any of these big models, primarily because of that disadvantage. It's just too expensive to serve, even if it is much lower latency. Okay, so just leaning into why it doesn't have lower latency, in case you're not familiar with the architecture of how GPUs and TPUs run today. So in a GPU, there's a tensor core, which does these big matrix multiplies. It's very efficient, has a lot of flops or hops or whatever. And then the memory that sits on the TPU GPU, this HBM, that's where the weights and the activations and everything are stored. And that has to transfer over from the memory all the weights and the activations and the KV cache into the tensor core in order to do the computation. And so it has to flow through this bandwidth channel. And that bandwidth channel is very tight. It turns out that both GPUs and TPUs have a lot of flops and not that much bandwidth. It's quite hard, it's expensive to put bandwidth onto these chips and it's easy to put flops. So because of that ratio, if you do more flops for each streaming amount of data you put through, the better. So when you're serving an autoregressive model, these chips are memory bound. They're basically bottlenecked by this bandwidth. So when you do autoregressive next token generation, for each token, you're doing one token at a time, let's say batch size one, you have to stream over the entire neural network and all the KV cache and everything to get one token and then you do it again for the next token and so on. Whereas for text diffusion, you're generating, say, 256 tokens, you still stream over everything. But if you can do that less times than the number of tokens, this iterative refinement process, then you'll get a speed up. So if you can do, say, 24 passes to generate 256 tokens, you'll be doing 10 times fewer memory transfers than an autoregressive model. And if you are truly memory bound, then you'll be 10 times faster, something like that. So that's the real reason, that's the hardware reason why text diffusion models are much lower latency than autoregressive models. Okay. So we had this Gemini diffusion demo, maybe some of you got access to it, last year. And that was able to hit something like 2,000 tokens a second, that was really good at the end pretty consistently, depending on the length of the query. Obviously, it depends. The longer sequence it's generating, the less it's pre-fill dominated. [SPEAKER_05] And so you can really lean into these very long sequences of very fast tokens. But if you're only generating one token, for instance, you'll be just dominated by the cost of the pre-fill. And the tokens per second number that was reported on this webpage was incorporated pre-fill and everything like that. So this was 2,000 tokens a second, as genuine raw tokens that you would receive in your web browser. Okay. So that was the whirlwind tour of text diffusion and its main advantage, which is latency. But I want to dig in a little bit into some of the other advantages that text diffusion has, which are not as talked about in the literature. But I think are pretty cool. And this is why I'm excited about it. So at Google I/O last year, they showed this demo for the text diffusion model, which is a really easy prompt. But lots of models actually make mistakes on it. So the prompt is, if you go to the next slide, what is the square root of 81 times 2 thirds squared plus, blah, blah, blah? And I think the answer is 39 to this problem. And so you pass that into the model and you ask the Gemini diffusion to respond to that. And after one forward pass, these are the tokens that have been generated. So one forward pass through the model, it's starting to respond. Now, it's doing this iterative refinement process. So one forward pass is not all it's going to do. But after one forward pass, it has this. So it has answer equals, and then it says 60. It's not correct, but that's what it's guessing for now. And then it starts to do the reasoning. So a solution, calculate the square root of 81, and so on. After two forward passes, it's changed 60 to 49. And it's gotten a little bit further into the reasoning. So it's gotten five steps into the reasoning. Two squared equals four. And some of the blue tokens are still going to change. And then after three forward passes, it's actually gotten all the way through the reasoning. So it gets the answer correct at the end. 36 plus 3 is equal to 39. And it's gone back and fixed the original response to say 39. So it had a mistake twice, 60 and 49. But once it finished the reasoning, it was able to return back and fix the mistake that it made at the start. Now it's going to do a couple more forward passes in order to fix some of these tokens that aren't quite right in the text. But overall, that's basically the structure of the output that it'll return. And this is a property that text diffusion models have. This ability to do bidirectional reasoning. So to not only see the past, but also see the future that it's going to utter, it's going to respond. And also to use that information to do self-correction. So it made a mistake, but it was able to have another forward pass. It was able to go through and fix that mistake. At the time, much bigger models than the one we were serving made a mistake for this problem. So both ChatGPT 4O, which was new at the time, and Gemini 2.5 Flash, which was brand new at the time, both made an error on this exact problem. So you give them the exact same prompt, and then they would say the answer is 39. The GPT 4O said 40, because that's the best guess it can do at that one token. It went through the reasoning, and then it did manage to figure out it was 39. It was able to go through and fix that mistake. At the time, much bigger models than the one we were serving made a mistake for this problem. So both ChatGPT4O, which was new at the time, and Gemini 2.5 Flash, which was brand new at the time, both made an error on this exact problem. So you give them the exact same prompt, and then they would say, remember the answer is 39. The GPT4O said 40, because that's the best guess it can do at that one token. It went through the reasoning, and then it did manage to figure out it was 39. It said, sorry, I made a mistake. It's 39, not 40. Gemini 2.5 Flash also made a mistake. It said 42, and then it actually just stuck to its guns and never changed it. It said 36 plus 3 is 42. So it incorporated the error into its reasoning later. And these are way bigger models than the Gemini diffusion model. So it really is a property, a flaw of autoregressive models that the text diffusion models don't have. And you can fix this with modern reasoning thinking models, but then you're just punting the problem into something else. But anyway, okay, so that's one advantage, which is bidirectional reasoning, self-correction. Another one is what I hinted at before, which is dynamic computation. So you can give the model more time at inference, more forward passes. You give it a bigger budget, and it can just do better. It's not exactly monotonic, but it is roughly monotonic that the quality across every eval basically just continues to go up. Because even if the solution is almost tightly clean and correct, it gets to look at it and see that it made a mistake and then fix it. So you get this nice curve where you always see, as the number of denoising steps, which is the forward passes, increases, overall the quality gets higher. And these are just six coding evals that we monitor internally. [SPEAKER_02] On top of that, a slightly different concept is the model can do adaptive computation, which is that you can allow the model, train it in a way, to determine itself when it is finished. And then for easy responses, it can use a little bit of compute, and for harder responses, it can take longer. So here's just three examples from the Gemini diffusion model, which is what are the first 100 digits of pi? This is actually 100 tokens. It looks like a short response. It's actually 100 tokens. And it only takes four steps to do that. Because the model, it's an easy response, because you've just memorized the 100 digits of pi. You just output it. Whereas an autoregressive model at the same time would have only done four tokens. So that's a very easy one. Slightly more challenging is to write a little bit of code. So that takes 18 forward passes to generate FizzBuzz. And then something more complicated is explaining quantum mechanics in a single paragraph. And that took 31 denoising steps. It just took its time, decided to spend longer on those ones. And so the model naturally gets to decide to determine when it's going to finish and return the response. And typically we see that harder evals take more time. So this is a year ago now. So the evals are old school. But on the harder end is GPQA diamond, which for the model size we were targeting was quite a hard eval. And that took a long time for it to respond to those ones. Whereas on the other end, MBPP, which is mostly basic Python programs, it was very easy to respond to. It took very little time for it to respond to these things. And this is entirely determined by the model itself. Just easier problems, easier prompts it could respond to quickly and harder ones it decided itself to spend more time reasoning. OK, so that's another property, which is the dynamic and adaptive computation. Lastly is the fast in-place editing. So diffusion models in general have this very nice property where you can take an image, cut something out of it, and give it a prompt, and it'll fill it in. And you can use the context that you haven't cut out to fill in the piece you've cut out correctly. And so you can use that for clever image editing. And the reason it can do that is because it's not autoregressive. There are autoregressive image generators, which go left to right, top to bottom, in raster order, generating pixels. But diffusion doesn't work like that. It'll just see the entire image and then start to denoise it. And because of that, because it gets to see every pixel, every pixel gets to see every pixel, it can fill in the missing information and do it in a way that's consistent with whatever prompt you're giving. So we can do something similar. So I have a couple of demos here. Can you see that? Yes. So this is just some code and you say, there's a bug in this code, can you fix it? And it'll just make the edit in the correct place. Like it won't, you can barely see that, but it's a little fix here of the indices. And you can say things like, can you add documentation? It'll go in. It's not just one by one generating all the tokens, it's doing a clever editing procedure to actually fill in the correct edits here. You can do that with more general text. Like you can take a story and then say, add a middle paragraph. And because it can see that the first and third paragraph, it can fill in the paragraph in a way that's consistent with the rest of the story. And this is just in-place editing. Okay. So those are some of the advantages. Don't have a lot of time. The biggest advantage, as I mentioned, is this low latency. And we really lean into that. And I just want to show you a couple of demos of some of the things that people internally have built to show what the advantage of low latency can give you. So it's not just the same thing faster. It can really unlock some new applications. So in your own work, you're all AI engineers. [SPEAKER_05] It'd be interesting to see when the next diffusion model comes out from our team, what the low latency could unlock. And what new applications can be built. So here are just some demos. So this is Wikipedia. Let me just pause it actually. This is Wikipedia where everything is generated on the fly, even the HTML. And I just want to show you a couple of demos of some of the things that people internally have built to show what the advantage of low latency can give you. So it's not just the same thing faster. It can really unlock some new applications. So in your own work, you're all AI engineers. [SPEAKER_05] It'd be interesting to see when the next diffusion model comes out from our team, what the low latency could unlock. And what new applications can be built. So here are just some demos. So this is Wikipedia. Let me just pause it actually. This is Wikipedia where everything is generated on the fly, even the HTML. So this is actually being generated by the model on the fly. So that it's a webpage with the HTML and the text and everything being generated on the fly. So it looks like regular Wikipedia. And when you click on it, the latency is low enough that it can just fill in the page as if it was a real Wikipedia page. So that's Wikipedia generated on the fly by just a very low latency model. We have a similar thing where we did it for Reddit. So now all the responses to your posts will be by bots. They weren't already. And it's generating fake comments. So the Gemini diffusion model was not an image generating model. So this demo links in the startup, state of the art image generator model at the time, which I think was Juno. So it was before nano banana. So it's the two of these models working together to fill in the webpage. So the image generation model is a little slower. But you can see that you can invent any Reddit you want sharks in this case, and it'll generate the page with the text. The images follow a minute later. And then you can interact with this website as if it was a real website with real users and so on. Just entirely, all the comments, all the images, all the HTML, everything is being generated entirely on the fly here. This is being generated by the model. I love this one. This is my favorite one. This is an operating system also being entirely generated on the fly. So every click here is generating the next page of the operating system. So it looks like a real operating system, but it's all being generated by the model on the fly, responding to every click. Every time you enter the readme, it generates the text. It also generates the web pages you can go back to the desktop and so on. I think this, yeah, okay. And then this is a demo from someone on Twitter who used the Gemini Diffusion API. Or sorry, not the API, the web page to do some vibe coding with his voice. So I really like this one. Create a to-do app. Add 10 random to-dos. Allow to-dos to have a completed state. Mark four random to-dos as completed. Allow me to sort to-dos by name and by state. All right, let's see if this works. Sort by name, sort by state. Testing, enter. It was added to the bottom. We'll try deleting a few. Let's add one. Everything's working. Please convert this to dark mode. And this was literally 15 seconds of work. Okay, so that was someone outside of our team, so he could say that. Yeah, vibe coding by voice. But in general, we think that low latency models can really unlock some new experiences for users and new products. [SPEAKER_00] And so we're excited to see what people will do when the next generation comes out. [SPEAKER_00] Okay, and on that note, thank you very much. [SPEAKER_00] Questions? [SPEAKER_00] Yeah? [SPEAKER_00] Yeah. [SPEAKER_00] Yeah, yeah. [SPEAKER_00] So, yeah, we use all the same data. [SPEAKER_00] Yeah. [SPEAKER_00] The algorithms have to change a bit, but we use all the same data. [SPEAKER_00] Yeah. [SPEAKER_00] And also maybe, can you distill these models? [SPEAKER_00] Are there any techniques for them? [SPEAKER_00] You can distill them, yeah. Yeah. I'm not sure if there are any published ones. Are there any published ones? I don't think so. Maybe there's some externally, but. Will we have the new future lecture about the. There's a, yeah, so we're going to release something soon. Yeah. Yeah, yeah. Yeah, so there's a, the bigger models tend to require less steps for the same output. So even if the model is getting bigger and the flops per forward pass are getting bigger, they tend to reduce the forward passes they need. So it's a situation where you have some sort of a diminishing cost of serving even the biggest models. [SPEAKER_05] [SPEAKER_05] So the next question, did you, how do you try to get it? I don't know. We haven't got to that yet. I'm a research scientist. Yeah. How do you define the size of the answer? Because for a few weeks, you expected I have this image and I have the output frame. [SPEAKER_05] How do you do that in coding or text? So there's a few different ways. The easiest way is to just fix some window length and then just iterate on that. [SPEAKER_05] So it's autoregressive, blockwise autoregressive. It's the standard way to do it. But you can have a head that will predict the length of the response and stuff like that. Yeah. Yeah. But you have to fix delta in order to. No, it still can generate unlimited text, but it's just, if you fix a window length, then [SPEAKER_06] it just does that window length autoregressively if it needs to generate many, many windows of text. Yeah? You might have already mentioned this, but when you give that example of the denoising steps being different ranges to different. [SPEAKER_04] Yeah. [SPEAKER_04] Could you ahead of time set a limit on the denoising steps that you apply for a problem so you can almost understand what your latency is going to be ahead of time? Mm-hm. Yeah, yeah. These are all with a limit, but they just finish earlier than the limit. [SPEAKER_06] It just does that window length autoregressively if it needs to generate many, many windows of text. Yeah? You might have already mentioned this, but when you give that example of the denoising steps being different ranges to different. [SPEAKER_04] Yeah. [SPEAKER_04] Could you ahead of time set a limit on the denoising steps that you apply for a problem so you can almost understand what your latency is going to be ahead of time? Mm-hm. Yeah, yeah. These are all with a limit, but they just finish earlier than the limit. Oh, right. Okay. Yeah. [SPEAKER_04] If you have multiple windows like you just said going through, can they then go back and attend to previous windows? [SPEAKER_02] Yeah. [SPEAKER_02] As the window goes, it's. [SPEAKER_02] Yeah. [SPEAKER_02] You could potentially, but for us we just set it in stone and continue. [SPEAKER_03] Yeah. [SPEAKER_03] Lots of versions here. [SPEAKER_03] The version is it can go a lot, but it just wouldn't make sense to make a hybrid. [SPEAKER_03] Yeah. [SPEAKER_02] Yeah. [SPEAKER_02] Yeah. [SPEAKER_02] That's how it works. [SPEAKER_02] Yeah. [SPEAKER_02] Yeah. The start to prepare with the autoregressive style and then go to the division mode or? [SPEAKER_01] Oh, no. [SPEAKER_01] So it's pre-fill is the same. [SPEAKER_01] It's just that you've got some context and you pre-fill. [SPEAKER_01] And then after that, the generation step is typically in blocks of some fixed size, like 512 or 1000 or 32 or whatever you want. And then that's autoregressive. Yeah. [SPEAKER_02] I guess that you're doing the demo thing inside a bi-directional. [SPEAKER_02] So how do you, what is the form supposed to gain back to the open ID? [SPEAKER_05] Is it a latent institution instead of? So you can do that. The easiest way to do it is to just have a logits at the top. Just vanilla prediction head at the top. It's after the demo thing called that? [SPEAKER_02] Well, so this is all discrete diffusion, right? So it's always tokens in, tokens out. If I show you the... For every step. For every step, yeah. So if I go back here. It's always a discrete corruption process. And then you fill in a discrete token back in. [SPEAKER_02] So you're always in a discrete space. [SPEAKER_05] But you can do it in latent spaces, but people have done that. [SPEAKER_05] But most of the text diffusion models and literature is discrete diffusion today. Yeah. Is that only a demo? It might be one day. Yeah. Yeah. [SPEAKER_05] What's your outlook like? [SPEAKER_02] Do you think one of the architectures will win in the future? [SPEAKER_02] Or do they have different use cases? I think for now they have different use cases. So if you think about what a low latency model provides, that's worse throughput. What's that trade off? On-device applications. [SPEAKER_02] So we are in a couple of on-device applications already. [SPEAKER_02] And within the alphabet ecosystem. [SPEAKER_02] So robotics, things like that. We want to run a model on the device itself. So your phone or a robot or whatever. And then you want it to be low latency. And you're not batching with thousands of other queries like Gemini being served on the server side. So you want the lowest latency model. Quality isn't really any difference, they're the same quality basically. [SPEAKER_05] So you may as well pick the low latency one. [SPEAKER_05] Because you don't have the throughput concerns. [SPEAKER_05] What do you mean in the future, can you get quality up to the bar with the current proxy models? [SPEAKER_05] Quality isn't the concern, it's the throughput for serving in a big batch setting. Yeah. Yeah. [SPEAKER_02] So currently you are getting more of the sense of quality of function models in the area. [SPEAKER_02] Yeah. [SPEAKER_02] Yeah. [SPEAKER_02] Yeah. [SPEAKER_02] Yeah. Yeah. Yeah. Yeah. [SPEAKER_05] You can do RL. Yeah. [SPEAKER_05] You just need to change the algorithm. [SPEAKER_05] Yeah. [SPEAKER_02] But you can still do it. [SPEAKER_05] Yeah. [SPEAKER_05] Yeah. Yeah. Yeah. Yeah. [SPEAKER_02] The problem with that is if you train it with noise it expects noise. You'd have to train it with this more model. [SPEAKER_03] As you like the output. [SPEAKER_03] That just adds complexity. So we don't usually do that. Yeah? Is that a paper? [SPEAKER_02] A word you said, and if you want to come in with the idea of our team? [SPEAKER_02] Well, not from our team, but there's a bunch of literature out there, yeah. [SPEAKER_02] Okay. Yeah. Is that a spectrum? Yeah. You get the idea, I think. Yeah. Any other questions? There's a lot of questions. [SPEAKER_02] I've gone through them all. [SPEAKER_02] That's good. [SPEAKER_02] Oh, one more. [SPEAKER_02] Okay. [SPEAKER_06] Just a random one. [SPEAKER_06] What happens if you ask it to generate? [SPEAKER_06] I don't know. [SPEAKER_02] I don't think we've tried to do that. Probably would work. We'd probably do something. Yeah. [SPEAKER_05] Cool. Okay. Thanks, everybody. [SPEAKER_06] I don't know. I don't think we've tried to do that. Probably would work. We'd probably do something. Yeah. [SPEAKER_05] Cool. [SPEAKER_02] Okay. Thanks, everybody. Oh, my God. [SPEAKER_02] Oh, my God. [SPEAKER_02] Oh, my God. [SPEAKER_02] Oh, my God. [SPEAKER_02] Oh, my God. Thank you. Yeah. Yeah. So currently you are getting more like the sense of quality of function models in the area. Yeah. Yeah. Yeah. Yeah. Yeah. Yeah. Yeah. You can do RL. Yeah. You just need to change the algorithm. Yeah. But you can still do it. Yeah. Yeah. Yeah. Yeah. Yeah. Yeah. Yeah. Yeah. Yeah. Yeah. Yeah. Yeah. Yeah. Yeah. The problem with that is if you train it with noise it expects noise. is you'd have to train it with this more model as you like the output. That just adds complexity. So we don't usually do that. Yeah? Is that a paper? Like a sort of word you said, and if you want to come in with the idea of our team? Well, not from our team, but there's a bunch of literature out there, yeah. Okay. Yeah. Is that a spectrum? Yeah. You get the idea, I think. Yeah. Any other questions? There's a lot of questions. I've gone through them all. That's good. Oh, one more. Okay. Just a random one. What happens if you ask it to generate? I don't know. I don't think we've tried to do that. Probably would work. We'd probably do something. Yeah. Cool. Okay. Thanks, everybody. Oh, my God. Oh, my God. Oh, my God. Oh, my God. Oh, my God. Oh, my God. Oh, my God. Thank you.