SPEAKER_00
Please welcome to the stage the Vice President of Research at Google DeepMind, Benoit Schillings.
SPEAKER_01
All right, good morning. This is really quite exciting, to be here and have a chance to speak with all of you. My name is Benoit Schillings. I'm actually a bit of a noob when it comes to machine learning. Until a year and a half ago, I was working for Google X, which some of you may know. We've done things like Waymo, which seems to be at every street corner now. We also do things like Glass. So we had a mix of hit and success. But in many ways, this was, for me, an interesting formative experience on how to run a research team in a place like DeepMind.
SPEAKER_01
I do have an incredible team. My team's goal in DeepMind is to develop whatever technology will be needed to make Gemini incredible between one month and one year from now. One month because if you start to work on what is needed in one week, that's a very different type of job. And one year because I don't think anybody can really predict anything that far. So that's already pretty ambitious, in my opinion, to think about things that would happen one year in the future. We do many things under that role. A lot of it is related to code, which will be the main subject of my talk today.
SPEAKER_01
But we also do a lot of research on what is the evolution of reasoning for models, for instance. Or we do topology research: what are new types of network that might bring better performance. We do fundamental work in the science of reinforcement learning, which is so fundamental to what we're doing today with ML. Let's do a bit of an origin story.
SPEAKER_01
We started the project at X named Pitchfork in 2018, which was aimed at looking at how ML could really improve the way code is being written. And this was very interesting because in 2018, when we presented that at Google, honestly nobody would give us the time of day. There was that point: why would you ever need ML to write code? [SPEAKER_00] When we did that project originally, the idea was to look at how we could speed up the evolution of a piece of code.
SPEAKER_00
[SPEAKER_01] How could we make many of those small changes which slows down code speed development?
SPEAKER_01
The small edit which requires a review that takes three days, and how we could compress that cycle. Some people were talking about Vibe coding, writing code in English. And at the time, honestly, I totally dismissed that. I was, that's why we have programming languages. English is not a programming language. Well, I guess I was pretty wrong on that front. But the resistance we felt at the time reminded me of how my own career was pretty resistive to change. I've been writing code for 45 years. I started by writing video games for Apple II and Commodore 64. So my formation was to write assembly language.
SPEAKER_01
And when you spend a long time writing assembly language, you look at compilers with a lot of suspicion. Right? Are those things really working correctly? And then when you switch to C++ and use a compiler, you look at garbage-collected languages as this: Hmm, that's not real programming. You need to manage your memory. Well, today I use Python and Vibe coding. So even old dogs can learn new tricks. But I do understand what happened there. I think that we have a number of eras in what happened with software. And the first one was the one where I started writing code, where the fundamental limit was really the machine.
SPEAKER_01
And there was a lot of work to go and extract the last ounce of power out of those machines. And that was the days of assembly language, where you really needed to be incredibly accurate in the way you were writing code. Computing became much cheaper and we switched to the modern cloud era, where getting the best performance is not the most critical aspect. You can actually brute-force many problems. But really what became the limiting factor was the ability for us to design in a modular way. This was the era where software was write it only once. And this was this whole idea of how are you going to build libraries? How are you going to write functions?
SPEAKER_01
How are you going to break down that problem into something that is long-term manageable? The limitation there, and that determined a lot of how our software processes are working, were actually the human brain. The traditional human, typical human is able to get the context between seven and nine tokens. We have very rich tokens, but you compare that to modern ML where the context is going to be infinite pretty soon. That fundamental limitation of humans determines a lot of how software was being written. This is over. And we're switching now to that AI frontier where really writing the code is not the challenge anymore.
SPEAKER_01
I'll speak some more about it, but the bottlenecks are really how do you ensure that that code is what you really wanted? Because writing the code is easy, but getting what is needed for a specific problem can be much harder to specify. So humans, at least in the near future, will be that role of architecture or thinking of what are really the implication of that piece of code and getting the ML to design. Inductive thinking is another category where I think humans still have a very clear edge, which is to look at a system in a much wider context and to be able to detect patterns and from those patterns take some decision. So where are we today?
SPEAKER_01
Superhuman syntax generation. When is the last time I got Gemini to write a function for me and I looked at the function and I was like, I can do that better. It's over. I think that the minutia of code writing, you can fight, you can fight, you can argue, you can find counter example, but that time is gone. Where we have a lot of work to do is multi-step code base. Software engineering is not about writing code. Software engineering is the first time you join a company and you realize that there are 35 million lines of PHP in the code base and that you need to make some changes. That's the day you understand what software engineering is.
SPEAKER_01
And that's a place where modules today, our frontier modules are progressing, but this ability to manage that extreme complexity and break it down into manageable pieces is a place where the frontier is still moving. It goes all the way to architecture.
SPEAKER_01
You look at, I don't know, the Google architecture. Thank God we have Jeff Dean, who was the key architect there. But that's the level of thinking which has many implications, which can go from how do you do hardware optimization? How do you manage security? How do you build a system so that 10 years later you're not full of regrets? And I think this is really the range of progress we are working on today. So code is over, but there's plenty to do. and break it down into manageable pieces is a place where the frontier is still moving. It goes all the way to architecture. It goes all the way to architecture. You look at, I don't know, the Google architecture.
SPEAKER_01
Thank God we have Jeff Dean, who was the key architect there. But that's the level of thinking which has many implications, which can go from how do you do hardware optimization? How do you manage security? How do you build a system so that 10 years later you're not full of regrets? And I think this is really the range of progress we are working on today. So code is over, but there's plenty to do. And break it down into manageable pieces is a place where the frontier is still moving. It goes all the way to architecture. You look at, I don't know, the Google architecture. Thank God we have Jeff Dean, who was the key architect there.
SPEAKER_01
But that's the level of thinking which has many implications, which can go from how do you do hardware optimization? How do you manage security? How do you build a system so that 10 years later you're not full of regrets? And I think this is really the range of progress we are working on today. So code is over, but there's plenty to do. And break it down into manageable pieces is a place where the frontier is still moving. It goes all the way to architecture. You look at, I don't know, the Google architecture. Thank God we have Jeff Dean, who was the key architect there.
SPEAKER_01
But that's the level of thinking which has many implications, which can go from how do you do hardware optimization? How do you manage security? How do you build a system so that 10 years later you're not full of regrets? And I think this is really the range of progress we are working on today. So code is over, but there's plenty to do. There's plenty of progress to be made. Now code is a very unique problem. And in some way that's the reason we did pitchfork on this. First of all, code is a lot of data. There are other domains where you can find a lot of data to train your model, but code was so incredible. You could go on GitHub and start to scrape GitHub.
SPEAKER_01
So this was one of those problems where the amount of training data was a very unique situation. It is also a domain where doing verification is reasonable. You can run a piece of code, you can compile it, you can have unit tests. So the ability to figure out whether the model is generating something correct was something that was pretty reasonable to do.
SPEAKER_01
That brought us where we are today. But today what happened is that we ran out of training data. I think that 80% of the new code added to GitHub today is machine-generated. So the notion of humans bringing some knowledge that can be used for mining and to train models is reaching an end. But the good news is that we can do self-play. And self-play is something we always liked a lot at DeepMind. I suppose all know AlphaZero. AlphaZero became a superhuman Go and chess player without any human knowledge, just by playing against itself. We are now at that stage where frontier models for code are able to do the same. Where they can create their own challenge.
SPEAKER_01
They can judge the validity of the answer. They can even, to some extent, judge the architecture. So that ability to do those hundreds of millions of hours of self-play, or writing code, is the thing that will bring us to the next layer. It's interesting. Do the experiment. Take a brilliant software engineer. Lock him in a room. Lock him or her in a room for two years. Take a good place and feed pizza. And give the mission: you need to become a better software engineer. What do you do as a person? You give yourself some challenges. Challenges that you can verify. And you keep working and coding on those challenges. We can do the same here.
SPEAKER_01
So this is an issue of how much compute, how much self-play time we can have. But that will bring the horizon of how far we go in superhuman coding. So the economics of code are changing dramatically. As I say, we developed a whole software engineering culture and infrastructure and set of companies. Based on the assumption that writing code was the hard part. That this was the expensive part. We are now in a world where writing code is free. Or nearly free. That's why I've got the tilde there. That means that the amount of code that we're going to see produced is going to explode. And there are some hard implications to that. First is the question of design and adequacy.
SPEAKER_01
How, in front of that mountain of code, which would be written or written dynamically, how do we keep systems which work and are reliable at the macroscopic level. Great work for human. It is also the issue that we are writing code and we're not reading it very much anymore. I know we still have code review. But I would predict that in one year we'll let Gemini or other model generate the code. And nobody will actually look at it. It's similar to compilers. Who still checks the assembly output of their compiler? And maybe someone there.
SPEAKER_01
That's probably the end of it. So the same thing is going to happen to code. And that brings some question of what are the new processes that we need to put in place to keep that manageable. And that's where I've got a bit of a list. Active guardrails. You've all seen the news of Mythos looking at a piece of code and detecting an unreasonable number of vulnerabilities in that code. There is a rush to go and patch those vulnerabilities. But I think that's going to be a never-ending process.
SPEAKER_01
We're going to get a certain layer of vulnerabilities discovered by models. We're going to fix those. Models will get smarter. They will go a bit deeper and find even more subtle vulnerabilities. So I think that the first aspect is that we need to think at least as much about code security and the implication of a piece of code as the code writing itself. And the grail, and something my team is working actively on, is instead of detecting the vulnerability and then suggesting some fix, how about teaching model to write correct things from the start. And that is very, very hard to do because it is very context-dependent.
SPEAKER_01
The other aspect is that that's what I call inductive architecture. I think that models today are still not very good at transferring knowledge. Of taking knowledge from one domain and applying it to another one. Or taking two concepts and finding the intersection of those contexts to be able to do deductive thinking. So I think that the first aspect is that we need to think at least as much about code security and the implication of a piece of code as the code writing itself. And the grail, and something my team is working actively on, is instead of detecting the vulnerability and then suggesting some fix, how about teaching models to write correct things from the start.
SPEAKER_01
And that is very, very hard to do because it is very context-dependent. The other aspect is that that's what I call inductive architecture. I think that models today are still not very good at transferring knowledge. Of taking knowledge from one domain and applying it to another one. Or taking two concepts and finding the intersection of those contexts to be able to do deductive thinking. If we really want to write those very complex software systems using ML, that is a skill that we need to teach. And one aspect of that is to really teach models how to do correct planning in front of a problem.
SPEAKER_01
How do you look at a very complex problem and decide what is the right decomposition of that problem that will bring the best clarity or correctness to the problem? We also need to change the way we do evaluation. 3bench is infamous in my book because 3bench verifies if a piece of code runs and produces the right output. That's only a small part of, as I mentioned earlier, code engineering. So, for instance, I think that we need some problems much more in those benchmarks that we use, which are open-ended problems. If we really want to write those very complex software systems using ML, that is a skill that we need to teach.
SPEAKER_01
And one aspect of that is to really teach models how to do correct planning in front of a problem. How do you look at a very complex problem and decide what is the right decomposition of that problem that will bring the best clarity or correctness to the problem? We also need to change the way we do evaluation. 3bench is infamous in my book because 3bench verifies if a piece of code runs and produces the right output. That's only a small part of, as I mentioned earlier, code engineering. So, for instance, I think that we need some problems much more in those benchmarks that we use which are open-ended problems.
SPEAKER_01
I'll give an example. I love the question of text compression. How many bits per character do you need, and how far can you go? So that's a very simple eval to write. You just take a piece of 10 megabytes of code and you tell the model, write the best compressor you can that is lossless. And the loss function in that case will be the size of the compressed file plus the size of the source code. That's only a small part of, as I mentioned earlier, code engineering. So, for instance, I think that we need some problems much more in those benchmarks that we use, which are open-ended problems. I'll give an example. I love the question of text compression.
SPEAKER_01
How many bits per character do you need, and how far can you go? So that's a very simple eval to write. You just take a piece of 10 megabytes of code and you tell the model, write the best compressor you can that is lossless.