Open Reader

Research to Reality: Benoit Schillings, Google DeepMind, VP Research (Thinking, Reasoning, Coding)

completed 20:26 Jul 17, 2026 Watch on YouTube

Current Status

completed

Video ID

1P1hJ36rxM0

RAG / Chat

Enabled
Research to Reality: Benoit Schillings, Google DeepMind, VP Research (Thinking, Reasoning, Coding)
Description

A keynote exploring generative AI for code, deep-thinking algorithms, and the future of pre-training and transformer models for Gemini. Speaker: Benoit Schillings leads the Thinking, Reasoning, and Coding teams at Google DeepMind, directing foundational research toward AGI. His work focuses on advancing next-generation model reasoning and integrating software development best practices into AI code generation. Previously, as CTO at X, Benoit guided early-stage teams prototyping Alphabet's moonshot technologies across computing, biochemistry, and clean energy. LinkedIn: https://www.linkedin.com/in/benoit-schillings-2942a5 Timestamps: 0:00 Introduction and speaker background 2:35 The origin story of the Pitchfork project 4:43 Historical eras of software development 7:08 The current state of AI code generation 9:36 The role of self-play in training models 11:13 Changing economics of software engineering 12:41 Implementing guardrails and security 13:48 Inductive architecture and model planning 14:36 Evolution of evaluation benchmarks 15:45 Moving beyond simple chain-of-thought tokens 17:51 Future applications in chemistry and biology Key Takeaways from the talk: Benoit Schillings, VP of Technology at Google DeepMind, discusses the transformative impact of generative AI on software engineering and the future of model reasoning (0:49 - 2:35). The Era of Syntax Generation is Over: (4:43) Coding has shifted from a machine-constrained task to an AI frontier where syntax is effectively solved, moving the bottleneck to architecture and validation. The Power of Self-Play: (9:36) As human-generated training data reaches saturation, DeepMind is utilizing self-play, where models generate and verify their own challenges to reach superhuman performance. Shift in Engineering Economics: (11:13) With writing code becoming nearly free, the focus must transition to active guardrails, security, and managing the explosion of generated code. Inductive Architecture: (13:48) The next

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: AI has already surpassed humans at generating local code syntax, and the next frontier is trustworthy software engineering: planning, architecture, verification, security, and managing codebases whose volume will grow dramatically as generation approaches zero marginal cost.
  • Why it matters: For agent systems and AI-native product development, the durable advantage shifts away from producing code toward specifying intent, designing constraints, evaluating open-ended system quality, and operating continuous security guardrails.
  • Best use: Use this as a strategic framing for how to redesign engineering workflows, coding-agent evaluations, and security controls rather than as an implementation tutorial.

Executive Summary

Benoit Schillings, DeepMind's VP of Research for thinking, reasoning, and coding, argues that coding has entered a new phase: models are already superhuman at routine syntax and function generation, while the hard work has moved to software-engineering tasks that require broad context. Those tasks include understanding large legacy codebases, decomposing ambiguous problems, making architectural tradeoffs, and ensuring systems remain secure and maintainable over years.

He frames this as a reversal of the assumptions underlying much of the software industry. Earlier eras were constrained first by machine performance and then by human cognitive limits, which drove modular programming practices. As AI can hold and process far more context than humans, code generation becomes nearly free; the scarce inputs become high-quality specifications, architectural judgment, and mechanisms that establish whether generated systems actually satisfy their real-world requirements.

Schillings believes code models can continue advancing even as human-authored public code becomes less available for training. He says roughly 80% of new GitHub code is now machine-generated, but argues code is unusually amenable to self-play because models can create programming challenges, compile and test solutions, and score results. This resembles AlphaZero's use of self-play, though he acknowledges the harder target is not merely correct outputs but system-level design.

The operational warning is that organizations will increasingly deploy code that no human has read, much as developers no longer inspect compiler-produced assembly. That makes reactive vulnerability patching insufficient. DeepMind's desired direction is models that produce correct and secure code from the outset, supported by active guardrails and evaluations that test planning and open-ended engineering quality rather than only whether a short solution runs.

Key Takeaways

  • Claim: Routine code writing is no longer the primary bottleneck: frontier models have effectively reached superhuman performance at generating functions and syntax. | Evidence: Schillings says he no longer looks at a Gemini-generated function and believes he could write it better; he characterizes this as "superhuman syntax generation." | Implication: Do not organize engineering advantage around faster manual implementation alone; shift human effort toward intent definition, review criteria, architecture, and operational accountability. | Caveat: This claim applies to local code-generation tasks, not to end-to-end software engineering or long-lived system design.
  • Claim: The central unsolved frontier is multi-step work across complex codebases and architectures, where a correct local patch can still create system-level problems. | Evidence: He contrasts function writing with joining a company and needing to safely change a 35-million-line PHP codebase; architectural decisions must account for hardware optimization, security, and avoiding regrets a decade later. | Implication: Coding agents should be assessed on repository-scale understanding, task decomposition, and architectural consequences—not just on whether they can emit working snippets. | Caveat: He describes frontier models as progressing on this capability, but not yet reliably managing the full complexity.
  • Claim: The marginal cost of producing code is approaching zero, so code volume will explode and specification adequacy will become the limiting constraint. | Evidence: Schillings says the industry's culture, infrastructure, and companies were built on the premise that writing code was expensive; in the new world, the challenge is ensuring a growing mountain of generated or dynamically written code is what was actually needed. | Implication: Treat requirements, invariants, interfaces, tests, and deployment policy as first-class assets; generated output without these constraints increases operational ambiguity rather than reducing it. | Caveat: He calls code generation "nearly" free rather than literally free, and does not quantify the resulting productivity or maintenance gains.
  • Claim: Code-model progress can increasingly come from self-play because programs are comparatively easy to execute and verify. | Evidence: Code has abundant training data, can be compiled and unit-tested, and lets models create challenges, attempt solutions, and judge validity. Schillings compares this to AlphaZero becoming superhuman in Go and chess by playing itself, and says models can conduct hundreds of millions of hours of coding self-play. | Implication: Expect rapid continued gains in coding capabilities even as human-authored training data becomes scarcer; static assumptions based on today's model limitations will age quickly. | Caveat: Verification is much easier for bounded coding tasks than for contextual correctness, architecture, security, or whether a product meets an ambiguous business need.
  • Claim: Public human code is becoming a diminishing training resource because machine-generated code now dominates new GitHub output. | Evidence: Schillings estimates that 80% of new code added to GitHub is machine-generated and says the era of mining new human-written code for knowledge is reaching an end. | Implication: For model strategy, differentiated learning loops, execution environments, test harnesses, and proprietary feedback may matter more than simply accumulating more public code. | Caveat: This is a speaker-provided estimate without methodology in the transcript, so it should be treated as directional rather than a validated industry statistic.
  • Claim: Security must move from post-generation vulnerability detection toward prevention through context-aware generation and active guardrails. | Evidence: He cites "Mythos" detecting an unreasonable number of vulnerabilities and predicts an endless cycle in which models find flaws, teams patch them, and more capable models later discover subtler flaws. His team's stated goal is to teach models to write correct code from the start. | Implication: Do not rely on an agent's apparent coding competence as a security control; build continuous scanning, policy constraints, isolated execution, and explicit security review into the development loop. | Caveat: He explicitly says prevention is very hard because correctness and security are highly context-dependent.
  • Claim: Current code benchmarks reward too narrow a definition of success and should be supplemented with open-ended evaluation of planning and engineering quality. | Evidence: Schillings criticizes what the transcript calls "3bench" for checking whether code runs and returns the right output. His proposed open-ended example asks a model to create a lossless compressor for 10 MB of code, scored by compressed-file size plus source-code size. | Implication: Build evaluation suites that score tradeoffs, decomposition, maintainability, security, and performance under realistic constraints rather than relying solely on pass/fail task completion. | Caveat: His compression example is an illustrative eval rather than a complete measure of production software quality.

Detailed Brief

How the software constraint has shifted

  • Claims: Schillings describes three software eras: machine-constrained programming, human-cognition-constrained modular software development, and an AI frontier in which code writing is no longer the principal constraint.; In the near term, he expects humans to remain most valuable in architecture, in understanding the broader implications of a change, and in inductive reasoning across a wider system context.
  • Evidence: He traces his own progression from Apple II and Commodore 64 assembly language through C++, skepticism of garbage collection, Python, and eventually "vibe coding."; He says traditional human working context is roughly seven to nine rich tokens, while model context is moving toward effectively unlimited scale.
  • Caveats: The claim that model context will become effectively infinite is a forward-looking assertion, not a demonstrated capability described in the talk.; Large context windows alone do not establish reliable cross-domain transfer or sound architectural judgment.
  • Implications: Engineering organizations should preserve and formalize architectural knowledge rather than assuming larger-context agents will automatically convert repository access into sound design decisions.; The relevant human-AI division of labor is likely to evolve from hand-authoring implementation toward framing problems and validating system consequences.

The reasoning gap: inductive architecture and decomposition

  • Claims: Schillings identifies weak knowledge transfer across domains as a major limitation of current models.; To build genuinely complex systems, models must learn planning: choosing the decomposition that provides clarity and correctness before generating implementation details.
  • Evidence: He distinguishes transferring a concept from one domain to another, and combining concepts across contexts, from straightforward code synthesis.; He calls this capability "inductive architecture."
  • Caveats: The talk offers no concrete method or measured benchmark showing that current models can reliably acquire this planning skill.
  • Implications: Agent workflows should make planning artifacts inspectable—such as scoped plans, dependency maps, assumptions, and acceptance criteria—rather than treating a final code diff as the only output.; Complex work should be structured around explicit decomposition and verification stages, not a single unrestricted prompt-to-merge loop.

Notable Concepts & Terms

  • Pitchfork: A Google X project started in 2018 to explore how machine learning could improve how code is written and, especially, compress the slow cycle of small edits and multi-day reviews.
  • Superhuman syntax generation: Schillings's label for the view that models now outperform humans at routine function and syntax production; it is deliberately narrower than full software engineering.
  • Self-play: A training approach in which models generate challenges, solve them, and verify or score solutions, reducing dependence on fresh human-generated code.
  • AlphaZero: The DeepMind precedent Schillings uses to argue that self-play can create superhuman performance without human examples, here applied conceptually to programming.
  • Active guardrails: Continuous controls around generated code, particularly security-oriented checks and constraints, needed when human review of every generated line becomes unrealistic.
  • Inductive architecture: The desired ability to transfer concepts across domains, identify relationships among contexts, and select a sound decomposition for a complex system.
  • 3bench: The benchmark name as rendered in the transcript; Schillings criticizes it as representative of evaluations that only test execution and correct output rather than broader engineering quality.

Operator Notes / Why Ken Should Care

  • Replace code-generation-only scorecards with a gated evaluation stack covering task decomposition, repository-scale change safety, tests, security invariants, performance tradeoffs, and maintainability.
  • Require coding agents to produce a plan, explicit assumptions, acceptance tests, and rollback or containment considerations before they can execute high-impact changes.
  • Treat agent-produced code as untrusted by default in CI/CD: use sandboxed execution, dependency and secret scanning, policy checks, and continuous vulnerability reassessment.
  • Invest in proprietary executable environments and feedback loops; public-code corpus scale may become less differentiating as machine-generated code dominates.
  • Monitor whether agent adoption is increasing code volume faster than the organization can maintain specifications, observability, ownership, and security coverage.

Source/Metadata

  • Title: Research to Reality: Benoit Schillings, Google DeepMind, VP Research (Thinking, Reasoning, Coding)
  • Transcript words: 3013
  • Duration seconds: 1226
  • Timestamp note: No timestamps or chapters were present in the supplied transcript; several passages appear duplicated in the extraction.

Transcript

3070 words en Processed in 171.6s

Please welcome to the stage the Vice President of Research at Google DeepMind, Benoit Schillings. All right, good morning. This is really quite exciting, to be here and have a chance to speak with all of you. My name is Benoit Schillings. I'm actually a bit of a noob when it comes to machine learning. Until a year and a half ago, I was working for Google X, which some of you may know. We've done things like Waymo, which seems to be at every street corner now. We also do things like Glass. So we had a mix of hit and success. But in many ways, this was, for me, an interesting formative experience on how to run a research team in a place like DeepMind. I do have an incredible team. My team's goal in DeepMind is to develop whatever technology will be needed to make Gemini incredible between one month and one year from now. One month because if you start to work on what is needed in one week, that's a very different type of job. And one year because I don't think anybody can really predict anything that far. So that's already pretty ambitious, in my opinion, to think about things that would happen one year in the future. We do many things under that role. A lot of it is related to code, which will be the main subject of my talk today. But we also do a lot of research on what is the evolution of reasoning for models, for instance. Or we do topology research: what are new types of network that might bring better performance. We do fundamental work in the science of reinforcement learning, which is so fundamental to what we're doing today with ML. Let's do a bit of an origin story. We started the project at X named Pitchfork in 2018, which was aimed at looking at how ML could really improve the way code is being written. And this was very interesting because in 2018, when we presented that at Google, honestly nobody would give us the time of day. There was that point: why would you ever need ML to write code? [SPEAKER_00] When we did that project originally, the idea was to look at how we could speed up the evolution of a piece of code. [SPEAKER_01] How could we make many of those small changes which slows down code speed development? The small edit which requires a review that takes three days, and how we could compress that cycle. Some people were talking about Vibe coding, writing code in English. And at the time, honestly, I totally dismissed that. I was, that's why we have programming languages. English is not a programming language. Well, I guess I was pretty wrong on that front. But the resistance we felt at the time reminded me of how my own career was pretty resistive to change. I've been writing code for 45 years. I started by writing video games for Apple II and Commodore 64. So my formation was to write assembly language. And when you spend a long time writing assembly language, you look at compilers with a lot of suspicion. Right? Are those things really working correctly? And then when you switch to C++ and use a compiler, you look at garbage-collected languages as this: Hmm, that's not real programming. You need to manage your memory. Well, today I use Python and Vibe coding. So even old dogs can learn new tricks. But I do understand what happened there. I think that we have a number of eras in what happened with software. And the first one was the one where I started writing code, where the fundamental limit was really the machine. And there was a lot of work to go and extract the last ounce of power out of those machines. And that was the days of assembly language, where you really needed to be incredibly accurate in the way you were writing code. Computing became much cheaper and we switched to the modern cloud era, where getting the best performance is not the most critical aspect. You can actually brute-force many problems. But really what became the limiting factor was the ability for us to design in a modular way. This was the era where software was write it only once. And this was this whole idea of how are you going to build libraries? How are you going to write functions? How are you going to break down that problem into something that is long-term manageable? The limitation there, and that determined a lot of how our software processes are working, were actually the human brain. The traditional human, typical human is able to get the context between seven and nine tokens. We have very rich tokens, but you compare that to modern ML where the context is going to be infinite pretty soon. That fundamental limitation of humans determines a lot of how software was being written. This is over. And we're switching now to that AI frontier where really writing the code is not the challenge anymore. I'll speak some more about it, but the bottlenecks are really how do you ensure that that code is what you really wanted? Because writing the code is easy, but getting what is needed for a specific problem can be much harder to specify. So humans, at least in the near future, will be that role of architecture or thinking of what are really the implication of that piece of code and getting the ML to design. Inductive thinking is another category where I think humans still have a very clear edge, which is to look at a system in a much wider context and to be able to detect patterns and from those patterns take some decision. So where are we today? Superhuman syntax generation. When is the last time I got Gemini to write a function for me and I looked at the function and I was like, I can do that better. It's over. I think that the minutia of code writing, you can fight, you can fight, you can argue, you can find counter example, but that time is gone. Where we have a lot of work to do is multi-step code base. Software engineering is not about writing code. Software engineering is the first time you join a company and you realize that there are 35 million lines of PHP in the code base and that you need to make some changes. That's the day you understand what software engineering is. And that's a place where modules today, our frontier modules are progressing, but this ability to manage that extreme complexity and break it down into manageable pieces is a place where the frontier is still moving. It goes all the way to architecture. You look at, I don't know, the Google architecture. Thank God we have Jeff Dean, who was the key architect there. But that's the level of thinking which has many implications, which can go from how do you do hardware optimization? How do you manage security? How do you build a system so that 10 years later you're not full of regrets? And I think this is really the range of progress we are working on today. So code is over, but there's plenty to do. and break it down into manageable pieces is a place where the frontier is still moving. It goes all the way to architecture. It goes all the way to architecture. You look at, I don't know, the Google architecture. Thank God we have Jeff Dean, who was the key architect there. But that's the level of thinking which has many implications, which can go from how do you do hardware optimization? How do you manage security? How do you build a system so that 10 years later you're not full of regrets? And I think this is really the range of progress we are working on today. So code is over, but there's plenty to do. And break it down into manageable pieces is a place where the frontier is still moving. It goes all the way to architecture. You look at, I don't know, the Google architecture. Thank God we have Jeff Dean, who was the key architect there. But that's the level of thinking which has many implications, which can go from how do you do hardware optimization? How do you manage security? How do you build a system so that 10 years later you're not full of regrets? And I think this is really the range of progress we are working on today. So code is over, but there's plenty to do. And break it down into manageable pieces is a place where the frontier is still moving. It goes all the way to architecture. You look at, I don't know, the Google architecture. Thank God we have Jeff Dean, who was the key architect there. But that's the level of thinking which has many implications, which can go from how do you do hardware optimization? How do you manage security? How do you build a system so that 10 years later you're not full of regrets? And I think this is really the range of progress we are working on today. So code is over, but there's plenty to do. There's plenty of progress to be made. Now code is a very unique problem. And in some way that's the reason we did pitchfork on this. First of all, code is a lot of data. There are other domains where you can find a lot of data to train your model, but code was so incredible. You could go on GitHub and start to scrape GitHub. So this was one of those problems where the amount of training data was a very unique situation. It is also a domain where doing verification is reasonable. You can run a piece of code, you can compile it, you can have unit tests. So the ability to figure out whether the model is generating something correct was something that was pretty reasonable to do. That brought us where we are today. But today what happened is that we ran out of training data. I think that 80% of the new code added to GitHub today is machine-generated. So the notion of humans bringing some knowledge that can be used for mining and to train models is reaching an end. But the good news is that we can do self-play. And self-play is something we always liked a lot at DeepMind. I suppose all know AlphaZero. AlphaZero became a superhuman Go and chess player without any human knowledge, just by playing against itself. We are now at that stage where frontier models for code are able to do the same. Where they can create their own challenge. They can judge the validity of the answer. They can even, to some extent, judge the architecture. So that ability to do those hundreds of millions of hours of self-play, or writing code, is the thing that will bring us to the next layer. It's interesting. Do the experiment. Take a brilliant software engineer. Lock him in a room. Lock him or her in a room for two years. Take a good place and feed pizza. And give the mission: you need to become a better software engineer. What do you do as a person? You give yourself some challenges. Challenges that you can verify. And you keep working and coding on those challenges. We can do the same here. So this is an issue of how much compute, how much self-play time we can have. But that will bring the horizon of how far we go in superhuman coding. So the economics of code are changing dramatically. As I say, we developed a whole software engineering culture and infrastructure and set of companies. Based on the assumption that writing code was the hard part. That this was the expensive part. We are now in a world where writing code is free. Or nearly free. That's why I've got the tilde there. That means that the amount of code that we're going to see produced is going to explode. And there are some hard implications to that. First is the question of design and adequacy. How, in front of that mountain of code, which would be written or written dynamically, how do we keep systems which work and are reliable at the macroscopic level. Great work for human. It is also the issue that we are writing code and we're not reading it very much anymore. I know we still have code review. But I would predict that in one year we'll let Gemini or other model generate the code. And nobody will actually look at it. It's similar to compilers. Who still checks the assembly output of their compiler? And maybe someone there. That's probably the end of it. So the same thing is going to happen to code. And that brings some question of what are the new processes that we need to put in place to keep that manageable. And that's where I've got a bit of a list. Active guardrails. You've all seen the news of Mythos looking at a piece of code and detecting an unreasonable number of vulnerabilities in that code. There is a rush to go and patch those vulnerabilities. But I think that's going to be a never-ending process. We're going to get a certain layer of vulnerabilities discovered by models. We're going to fix those. Models will get smarter. They will go a bit deeper and find even more subtle vulnerabilities. So I think that the first aspect is that we need to think at least as much about code security and the implication of a piece of code as the code writing itself. And the grail, and something my team is working actively on, is instead of detecting the vulnerability and then suggesting some fix, how about teaching model to write correct things from the start. And that is very, very hard to do because it is very context-dependent. The other aspect is that that's what I call inductive architecture. I think that models today are still not very good at transferring knowledge. Of taking knowledge from one domain and applying it to another one. Or taking two concepts and finding the intersection of those contexts to be able to do deductive thinking. So I think that the first aspect is that we need to think at least as much about code security and the implication of a piece of code as the code writing itself. And the grail, and something my team is working actively on, is instead of detecting the vulnerability and then suggesting some fix, how about teaching models to write correct things from the start. And that is very, very hard to do because it is very context-dependent. The other aspect is that that's what I call inductive architecture. I think that models today are still not very good at transferring knowledge. Of taking knowledge from one domain and applying it to another one. Or taking two concepts and finding the intersection of those contexts to be able to do deductive thinking. If we really want to write those very complex software systems using ML, that is a skill that we need to teach. And one aspect of that is to really teach models how to do correct planning in front of a problem. How do you look at a very complex problem and decide what is the right decomposition of that problem that will bring the best clarity or correctness to the problem? We also need to change the way we do evaluation. 3bench is infamous in my book because 3bench verifies if a piece of code runs and produces the right output. That's only a small part of, as I mentioned earlier, code engineering. So, for instance, I think that we need some problems much more in those benchmarks that we use, which are open-ended problems. If we really want to write those very complex software systems using ML, that is a skill that we need to teach. And one aspect of that is to really teach models how to do correct planning in front of a problem. How do you look at a very complex problem and decide what is the right decomposition of that problem that will bring the best clarity or correctness to the problem? We also need to change the way we do evaluation. 3bench is infamous in my book because 3bench verifies if a piece of code runs and produces the right output. That's only a small part of, as I mentioned earlier, code engineering. So, for instance, I think that we need some problems much more in those benchmarks that we use which are open-ended problems. I'll give an example. I love the question of text compression. How many bits per character do you need, and how far can you go? So that's a very simple eval to write. You just take a piece of 10 megabytes of code and you tell the model, write the best compressor you can that is lossless. And the loss function in that case will be the size of the compressed file plus the size of the source code. That's only a small part of, as I mentioned earlier, code engineering. So, for instance, I think that we need some problems much more in those benchmarks that we use, which are open-ended problems. I'll give an example. I love the question of text compression. How many bits per character do you need, and how far can you go? So that's a very simple eval to write. You just take a piece of 10 megabytes of code and you tell the model, write the best compressor you can that is lossless.