AI Engineer

Rewiring the State — Eoin Mulgrew, 10 Downing Street

857 summary words 4 min summary Watch video

Start with the signal

4 min read

Summary

Rewiring the State — Eoin Mulgrew, 10 Downing Street

Main Topics

  • The UK Public Service Crisis: Critical backlogs and inefficiencies across government systems
  • The 10 Downing Street Data Science Team (10DS): Mission, structure, and approach to transformation
  • The Fellowship Program: Recruiting elite technical talent from outside government
  • The Insurgency Model: Operating with autonomy, flexibility, and minimal bureaucratic constraints
  • AI Implementation in Government: Real-world applications and case studies
  • Scaling Challenges: How to expand successful models across the broader civil service

Key Points

The Problem

  • 7.25 million people on NHS waiting lists
  • 350,000 court cases stuck in backlog
  • Only 1 in 5 planning applications decided on time
  • Public sector productivity crisis exacerbated post-pandemic
  • Estimated £40 billion annual productivity gains possible through AI in government
  • Government has 400,000 employees—the largest opportunity for disruption

Barriers to Technical Excellence in Government

  • Low pay compared to private sector (especially tech)
  • Hierarchical structure and excessive bureaucracy
  • Slow decision-making processes
  • Regulations and safeguards necessary but cumbersome
  • Unattractive environment for high-performing technical talent
  • Standard civil service recruitment poorly optimized for technical talent

The Solution: The Insurgency Model

The 10DS operates as a "small insurgent unit" with unique advantages:

  • Political mandate from Number 10 with unusually high backing
  • Market-rate compensation (economically viable to attract talent)
  • High operational autonomy to be opportunistic about challenges
  • Custom recruitment process: 0.7-0.8% success rate, laser-focused on technical skills
  • Recruit exclusively outsiders: Former big tech employees, researchers, Y Combinator founders, serial entrepreneurs
  • Forward-deployed engineers: First embedded engineers in Number 10 history, working directly with policy teams, lawyers, comms staff

Recruitment Philosophy

> "We recruit missionaries, not mercenaries. Pay matters, but it's not alone. A paycheck won't get you out of bed in the morning when stuff gets hard."

The appeal: decisions across a minister's desk are among the most important possible work globally.

Notable Quotes

  • "What gets measured gets improved." — Opening statement setting tone for empirical approach
  • "If any industry is ripe for disruption over the next few years, government is one."
  • "Let us take the shackles off. Let us basically set up a small insurgent unit at the very centre that is not burdened by some of the constraints that I just described to you."
  • "This is almost like a pilot. I would like to see a lot of what we're doing become the norm, become BAU. This at the moment is basically a hack to get around the system."
  • "You've maybe done good stuff in industry. That's brilliant. Come join us and we'll give you the keys to the state and see what you can do." — Describing Will's journey from Harvard dropout/YC founder to working in prisons
  • "Otherwise, everybody would be doing it." — On why external talent + political backing alone isn't sufficient

Takeaways

Proven Early Wins

  • Policy Simulation Tools: Allow policy teams to model decisions (e.g., Universal Credit impacts) before implementation
  • Statutory Analysis: One engineer saved £1.5M statutory book analysis project; made tool reusable and faster than external consultants
  • Delivery Red-Teaming: PMO tool identifying optimism bias and prediction accuracy of departments
  • Public Dashboards: Published transparency dashboards on AI rollout and government delivery progress
  • Planning Application Digitization (Extract): DeepMind collaboration digitalizing handwritten forms; rolling out to all English local authorities
  • AI Tutoring Safeguards: Developing benchmarks to ensure safe classroom AI adoption
  • Justice AI: Former fellow embedding engineers in prisons to reduce drug flow and improve processes
  • New Public Service: Building a service "millions will use" in 2.5 weeks (vs. typical 1+ year discovery phases)

Operational Approach

Two Models:

  • Low-hanging fruit: Direct implementation by 10DS engineers (days to weeks)
  • Complex challenges: Partnership model with prolonged deployment of engineers into departments

Key Risks Acknowledged

  • Confirmation bias in AI tools: Mitigated through red-teaming models and user upskilling
  • Scalability bottleneck: Small team model works but won't alone "turn the oil tanker"
  • Political dependence: Model relies on sustained political will
  • Systemic change needed: Current approach is a "hack"—long-term requires changing how government operates

Scaling Strategy (Next 12-24 Months)

  • Move beyond targeted use cases to horizontal processes applicable across all 400,000 civil servants
  • Target call centers (DWP, HMRC), transcription, manual processes affecting frontline workers
  • Create spin-out teams within departments (e.g., Incubator for AI, Justice AI)
  • Require systemic policy changes, not just individual projects
  • Expand international collaboration (exploring partnerships with US, Singapore; open to Norway)

Recruitment Focus

  • Actively hiring for the fellowship program
  • Seeking talent willing to exchange private sector compensation for high-impact public work
  • Target audience: people "impatient to leave their mark on the world"

Conclusion

The presentation argues that small, elite, autonomously-operating teams with external talent can achieve substantial impact quickly in government, but systemic change requires institutional reform beyond individual projects. The model demonstrates proof-of-concept that government can innovate at unprecedented speed when bureaucratic constraints are removed and high-caliber talent is engaged with meaningful work.

Full transcript 4641 words · 41 min read
0:15

SPEAKER_01

Good to go. Can people hear me okay? Awesome. Cool. Thank you. I'm slightly embarrassed by the grandiose title now that I've had to leave it up for a few seconds. But anyway, who here works in government? Can I ask? Show of hands. Very good. That's what I was hoping for. You might be very well acquainted with some of the stuff that I'm going to grumble about. Who can't think of anything worse? All right. Good. What gets measured gets improved. So I'll do another one at the end. Actually, I won't in case more hands go up. Hi, everyone. I'm Owen Mulgrew. I work in the data science team just down the road in 10 Downing Street. I run our cross-government transformation work, including our fellowship program, which is predominantly what I'm going to talk about today. Yeah, I wanted to come down here and tell you what we're doing in the hope that some of you might decide you want to be a part of it, which would be quite cool. Quick bit about us. So number 10 data science team, 10DS. We were set up during the pandemic, partly in response to the pandemic. Our core business is making sure that the most important decisions in the country are informed by the best possible evidence. However, we are in the process of quite radically scaling up our own AI engineering and development capability, not just with the intention of driving AI adoption within number 10 itself, but also across strategically important parts of the state. And the way that we're going about doing that is quite novel in itself. Before we get into that, just a little bit of context to set the scene. I know some of you are flying in from the West Coast, et cetera. Some of you might not follow the news. Believe it or not, there are some challenges when it comes to public service delivery in the UK. I say that half in jest, but it's pretty serious. You know, at the moment there are seven and a quarter million people that are on NHS waiting lists. There are, I think about 350,000 court cases that are stuck in a backlog. Only one in five planning application decisions in this country are currently decided on time. Sitting behind all of this is a public sector productivity crisis that was bad and has only been exacerbated since the pandemic. There are different figures for the extent of this crisis. I've gone for a Tony Blair Institute figure that says there's a sort of 40 billion prize annual productivity gains from AI and government. But it's clear to anybody that works in the system and most of society that if any industry is ripe for disruption over the next few years, government is one. And I call it an industry rather than an organisation. It's a big complex industry of 400,000 people. And I think we should look at it through that lens. Unfortunately, however, government has traditionally not been great at building and nurturing high performing technical teams. A lot of these issues are not specific to the UK. A lot of our American friends will be familiar with them. But just some of the commonly cited ones. Pay is an obvious one. It makes it quite hard for us to compete for the best talent out there and then retain it. But also some barriers that are both real and perceived. So government is a very hierarchical organisation in many parts. There's a lot of bureaucracy. As a result, it can often move incredibly slowly. Not just for those reasons, but also the fact that there are regulations and safeguards in place that are very sensible because we're ultimately accountable to the public and to Parliament. But all of this can result in a system that is not always that appetising for high performing technical people to join. Especially the sort of people that we want. People that are impatient to leave their mark on the world. So what do we do about that? Changing a lot of these things, it's quite a systemic challenge. It's turning an oil tanker, which is a tired cliche, but it's a good one. We're a little team at the centre. Turning that oil tanker is beyond the remit of any one team, let alone a scrappy little startup like ours. However, I think Cal Bear, our chief AI officer, is going to be closing out the conference this evening. That's very much his job and he's doing great work at the Department of Science and Technology to do just that. So I recommend everybody goes along to it. But there is a lot of political will at the moment to get stuff done and to make sure that this time around we are seizing new technology to actually make a dent in some of those problems I just mentioned. So the question put to us was, well, what can you do about it? And this was sort of the answer. As I said, we're like quite a small team at the centre. In terms of what we can do, we said, okay, well, let us take the shackles off. Let us basically set up a small insurgent unit at the very centre that is not burdened by some of the constraints that I just described to you. What do I mean by an insurgency model? So we're setting up a new team. It operates with a mandate from number 10. We operate with an unusually high level of political backing to go into departments and get stuff done. We're able to pay market rates within reason. We're not paying meta money necessarily. But the thing is, a lot of people will happily take pay cuts if we make it economically viable to come in and work on some of these challenges because they're interesting, right? We operate with an unusually high level of autonomy as well. We're able to be fairly opportunistic about the challenges we take on where we go into a department and see opportunities to have impact. Also, this one's pretty crucial. The civil service standard recruitment process is optimised for a lot of things but not necessarily recruiting exceptional technical talent. We've been allowed to recruit our own way. We've got a fairly grueling selection process that is laser targeted on technical skills. We've got a success rate of about 0.7, 0.8%. And most interestingly of all, and this is what differentiates us, we recruit exclusively outsiders. One of the best ways that I can have impact is by getting some people, the likes of this conference, into government. Because what has happened in the past is they tend not to leave and some of them end up setting up their own teams. I'll get into that later. Just when we're on this though, I don't want to make this sound overly simplistic. A lot of people from the outside, particularly the tech industry, think that this bit alone is the only important thing that you need. That if you have a big enough stick for ministers, you can go in, you can break down data silos, you can do what you want. In practice, it's a lot harder than that. Otherwise, everybody would be doing it. And it's really early days. Like we're only setting out in this journey, but it turns out there's a huge amount of appetite for it. So we have been taking people from the labs. We've been taking people from big tech.

0:22

SPEAKER_01

This sounds overly simplistic. A lot of people from the outside, particularly the tech industry, think that this bit alone is the only important thing that you need. That if you have a big enough stick for ministers, you can go in, you can break down data silos, you can do what you want. In practice, it's a lot harder than that. Otherwise, everybody would be doing it.

0:29

SPEAKER_01

We're only setting out in this journey, but it turns out there's a huge amount of appetite for it. So we have been taking people from the labs. We've been taking people from big tech, from top research institutes. We've been taking YC founders, serial entrepreneurs, people who probably did not think they would be working in the civil service this time last year. But when you think about it, the decisions that go across a minister's desk are some of the most important things you could possibly work on. So if you make it economically viable, and you promise people that you're going to put them in an environment where they can do their best work, it makes it really interesting.

0:34

SPEAKER_01

It's also worth pointing out, we do want to recruit missionaries, not mercenaries. So the pay matters, but it's not alone. Because a paycheck is not going to get you out of bed in the morning when stuff gets hard. And doing the stuff that we do does tend to be difficult. In terms of how we operate then, again, a bit different from normal government teams. There's an abundance of low hanging fruit around the system. As you can imagine, it's a legacy organization. And there's lots of simple AI use cases that you can do in a few days to save money, to improve service delivery, all that good stuff. When it comes to that, we largely do it ourselves. And that's the easiest and most satisfying part. As we speak, we've actually got the first forward deployed engineers in the history of 10 Downing Street, embedding themselves with policy and operational teams, teams of policy advisors, teams of lawyers, teams of comms people, pollsters, and everything in between. They're observing their workflows, their pain points, co-designing solutions with them to help them do their job more efficiently and effectively. And generally taking things from idea to implementation in a couple of weeks and getting new capability into the hands of users quickly. And then some of the other problems we talked about—I mentioned some of the huge backlogs in the system. That is not low hanging fruit. That's really complicated stuff. And normally, when it comes to those, we take more of a partnership model, where we will deploy some of our people into another team or another department, sometimes for prolonged periods. And I'm going to give you examples of both of these in a minute. I'm going to start with some of the low hanging fruit. It's worth pointing out, a lot of the stuff that we do in number 10, I can't really show you. I know that sounds like an easy get out of jail card, but trust me, some of the stuff is a little bit sensitive. We're doing a lot of workflow automation, augmenting existing teams, as you can imagine. And here are a few other examples of stuff we've done just in the past few weeks. So policy simulation, that's turned out to be really interesting. So here we can allow policy teams in the building to test out the impact of different policy decisions before they're made. I think in this one here, we're looking at different decisions around universal credit and how they might impact household finances, amongst other things. But this can be applied to a broad range of stuff. Not replacing human analysis, necessarily. We're not putting ourselves out of a job. But what it has meant is that far more decisions in the building are being informed by high quality modelling and at a far faster rate than otherwise would have been the case. And this is another one just from the past couple of weeks. So the Cabinet Office was about to spend one and a half million pounds on getting an outside firm of lawyers to come in and do analysis of the entire UK statute book. Granted, the statute book is the height of four African elephants of legalese, but still a pretty obvious AI use case. So we were going to spend one and a half million. Instead, one of our engineers embedded with that team of in-house lawyers for a couple of weeks. The benefits of this is not just money saved. Obviously, one and a half million is not nothing, but also speed. So the issue with the analysis that we were going to pay for is that it was going to be done slower than the pace at which new laws and regulations are made, which means you're going to have to do it again after a certain period of time. So now we've got this tool that that team can use and they can do it whenever they want at the drop of a hat. And we can also potentially open source it and share it with other teams in government.

0:37

SPEAKER_01

And then this is another one. So in number 10, we're responsible for the delivery of every major project and manifesto commitment in government. That means a lot of reports come in on how various things are doing. This is a delivery red teaming tool that the team spun up a couple weeks ago that is now being used every day. It's essentially a PMO that we've put in the pockets of delivery teams in number 10, not just so that they can interrogate the delivery reports that are coming across their table, but also give a second judgment on the teams that are reporting them. So it will flag up to decision makers in number 10, does this team, does this department normally have a bit of optimism bias? Do they tend to disproportionately rate their risks as Amber? And are their mitigations usually effective or not?

0:45

SPEAKER_01

Also, aside from AI adoption, having this capability in-house is really good. I think transparency is one thing that this country can do a bit better at. Up until a couple of months ago, the government had never published a public facing dashboard so that you lot can actually see how we're doing when it comes to delivery. But now we've published two in as many months. I think some of you might be familiar with the one on the left. This is the AI Opportunities Action Plan that Matt Clifford drafted about a year ago. This is how the UK is doing when it comes to rolling out compute and generally setting up the UK to be a leader in AI adoption. And now you can go online and see how we're actually doing. Also, another thing that I can't show, but in two and a half weeks time, one of our ministers is going to launch a new public service that millions of people in the country are going to use. I can't go into more detail and steal their thunder. But in their words, it's hard to believe this didn't already exist. I assumed something like it already existed. That's something that we thought of two months ago and is now going to be live and used by the public. It is not an understatement to say that normally in government, a project like that might be in discovery for a year or more. So let's get into the media stuff then.

0:50

SPEAKER_01

how we're actually doing. Yeah. Also, another thing that I can't show, but in two and a half weeks time, one of our ministers is going to launch a new public service that millions of people in the country are going to use. I can't go into more detail and steal their thunder. But in their words, it's hard to believe this didn't already exist. I assumed something like it already existed. That's something that we thought of two months ago and is now going to be live and used by the public. It is not an understatement to say that normally in government, a project like that might be in discovery for a year or more. So let's get into the media stuff then. That's nice, low-hanging fruit within the building. Now I want to talk about some of the work that we're helping other teams with across different parts of the ecosystem. For the purposes of this, I'm just going to focus on three of our partners: the AI Safety Institute, the Incubator for AI, and Justice AI.

0:56

SPEAKER_01

ASIT, I think pretty much everybody in this room will be familiar with it. Massive win for the UK. It's a great thing that we set up. We're a leading government body for evaluating frontier models and it was also the world's first. And we were really proud from day one to support it by putting a couple of our fellows in there to help them set up their cyber security work stream, amongst other things. I'll not dwell too much on this, but one of our early fellows was Dr. Harry Coppock. I don't think Harry is here, but we put him into ASIT from day one. And he led on their inspect tool, amongst other things. So there's a safe, isolated environment for testing what AI agents actually do when you give them autonomy and tools. And the Incubator for AI, which now sits in DSIT. Who here is familiar with the Incubator? A few people. So the Incubator is essentially a spin-out of our program. It's a team that exists in the Department for Science and Technology that does what it says on the tin. It incubates new AI solutions for usage across the public sector. Its original funding team, most of the technical team were our fellows. And what's really cool now is not just seeing the work that they produced while they're there, but also the fact that we're able to collaborate with them when it comes to scaling up some of that work. Here's one recent example. So this tool is Extract. A bunch of our people have worked on Extract. It's a collaboration with DeepMind. It's built on Gemini and it essentially digitizes large swathes of the planning application process, especially those bits that are currently largely handwritten, including handwritten, hand-drawn maps there as well. This was unveiled by the Prime Minister at London Tech Week last year and we're currently in the process of rolling it out to every local authority in England. As I said, only one in five planning applications are currently decided on time. That has a massive impact on economic growth and economic growth is basically the biggest challenge the country faces right now. So anything we can do to make a dent on that is really significant. I think aspirationally as well, this will hopefully get us to a place where more and more planning applications can be decided by AI automatically. And then another interesting one, this is very current, the education gap is a big problem, not just in the UK but elsewhere. Many of you will have read the papers about AI tutors. It's a really exciting moment, the prospect of being able to level the playing field somewhat and put world-class tutors in front of every child regardless of their socioeconomic background. But it's something that has to be done really carefully. So currently we're working on producing safeguards and evaluating various frontier models against benchmarks, not just to make sure that children can interact safely with these in a classroom environment, but also measuring them against various metrics. I think in this one the relevant benchmark is the cognitive load placed on the student.

1:01

SPEAKER_01

And then last but not least, the new kids on the block, Justice AI. Some of you might have been here for the Justice AI talk yesterday. Was anybody here? Okay, good. For most of you this is new. Justice AI are a new team that have been set up in the MOJ, some of whom are over there. Hello. Again, I wouldn't say a spin-off from the fellowship, that gives us way too much credit, but the founder of Justice AI is one of our former fellows, Dan James, who's doing brilliant work in there. And they are deploying forward deployed engineers into prisons and into other parts of the criminal justice system. So taking an approach that we're doing in number 10 with policy people and comms people and lawyers, but instead they're embedding with parole officers and prison wardens. And they're doing loads of really interesting work. I can't go into too much detail. But most of it is around using AI to stop the flow of drugs into prisons, to find efficiencies where currently there's quite manual processes involving lots of people, and generally improving the security and safety within the prison system. And one of those FDEs is over there. It's Will. Will's one of our current fellows. Sorry, Will, I've embarrassed you. I just added in your photo last night because I thought this was a good point to end on. Will is... So to give you an idea, a few months ago, Will was in California getting a tan. That's him outside HMP Wandsworth on a rainy day. Will dropped out of Harvard, started a company, got it into Y Combinator, made a bit of money, but wanted to come and work for us. And that's his second week on the job and he's standing outside a prison with the keys to that actual prison about to go in. And that is exactly what we're trying to do through this program. You've maybe done good stuff in industry. That's brilliant. Come join us and we'll give you the keys to the state and see what you can do.

1:06

SPEAKER_01

So yeah, look, it's really early days. It's an experiment what we're doing. But I think the proof points so far have been that actually small elite teams can actually achieve quite a lot. We're already saving money. We're already shipping new public services at an unprecedented speed. We're already reforming frontline public services and we're already putting new AI capabilities into the hands of other teams at the top of government. So yeah. And surprise, surprise, this was a recruitment pitch. We are hiring. We're hiring. So please do scan the QR code and I'll be here the rest of the day if you want to come up and chat. Thank you.

1:09

SPEAKER_01

I think I've got time for a couple of questions, possibly. Somebody can tell me if not. Oh, sorry. So one on one of the earlier examples you showed, there was this chat to explain policies and, um, try different projections and see how they would behave. Do you have to deal with simulation cofancy, with the fact that, you know, if you have a user that's not necessarily very well versed in AI that just wants to hear what they want to hear, can direct the tool towards... I think I've got time for a couple of questions, possibly. Somebody can tell me if not.

1:33

SPEAKER_01

So one on one of the earlier examples you showed, there was this chat to explain policies and try different projections and see how they would behave. Do you have to deal with the fact that if you have a user that's not necessarily very well versed in AI and just wants to hear what he wants to hear, can direct the tool towards, "Look, I'm an absolutely brilliant mastermind. My policy is going to be fantastic," despite the policy being actually bad, but the AI is basically saying, "Yeah. Should I cut income tax to 0%? You're absolutely right." Exactly.

1:46

SPEAKER_01

Yeah, that's a very real risk. It's not something that we have encountered too much, but it's only because we have red teamed the models for that before we've put it into the hands of users. We also provide quite a bit of upskilling. A lot of the teams that we work with—we're creating tools for them. They're possibly lawyers, they're possibly sociologists, professors, whatever. So we do coach them on some of the risks that this presents. Yeah, it's a good question. Thank you.

1:51

SPEAKER_01

Thank you. Hello. Thank you very much for the speech. I'm Jack from Accenture. I think you beat our pitch for hiring, but we'll try as well. My question on the policy side is, as you progress with this FD type of model actually making an impact, how did you see, or as you said, it's early day experiment, how did you see you start to scale and essentially start dealing with central government or with local governments, the different party lines and so on and so forth—those kind of real governmental and human kind of stances. How do you see that's going to play out?

1:57

SPEAKER_01

Yeah, a hundred percent. So in terms of how this scales, at the end I said, you know, we can do quite a bit, and some of the people that join us then set up their new teams. That's great. Is it enough to turn the oil tanker itself? No, possibly over time, but it would take a long time. And we need to solve these problems quicker than that. So we have been thinking about that. I think realistically, some of this stuff requires strategic intervention. So we need to change the way the rest of government operates. Part of the reason, part of the bargain that we basically made with ministers was, you know, let us take the shackles off. Let us set up a small team at the centre that abides by different rules and use it as a proof point. This is almost like a pilot. I would like to see a lot of what we're doing become the norm, become BAU. This at the moment is basically a hack to get around the system. So we need to change that first of all.

2:02

SPEAKER_01

I think also, if we're talking really about scale, we talked about some fairly targeted use cases there. I think what we want to do over the next 12 to 24 months is do more horizontal work, looking at processes. So I should explain this to people as well. When you think of the civil service, you probably think of policy people working in some of the buildings around here that you can see out the window. That is a very small sliver of the civil service. It's about 400,000 people. Most of them are call centre operators, they're prison wardens, they're nurses, et cetera. There are a lot of processes out there, whether that's transcription, that every police person will tell you is the bane of their existence, or it will be those massive call centres in DWP, HMRC. So yeah, if we want to dial up the ambition, I would like to see us going after more of those horizontal use cases that can be applied en masse across the system. Yeah, that was a bit of a long answer.

2:10

SPEAKER_01

I think this is the last one. Yeah. You can shout if you want, it's up to you.

2:23

SPEAKER_01

Yeah, so I work for an ed tech company that among other things is making AI tutors. So I'd love to talk to you more about that. But one thing we find—the big problem with most kids is that you can make the best AI tutor in the world, but the real problem is motivation. You know, if you sit a kid down, if you sit a 12 year old down in front of a computer, they're going to do everything they can to avoid learning. So my question is, how do you solve that motivation problem? And first of all, I'm just interested to know more about what the vision is—the government's vision for this. Is this going to be going into schools? Are kids going to be sitting down in front of computers and using it?

2:30

SPEAKER_01

Yeah, 100%. So at the moment, I think our plan is largely not to necessarily develop products that compete with yours, but rather set benchmarks and guardrails for how schools can then adopt whatever products they want. In terms of student uptake though, that's not something we've done too much on yet. The test that you just saw was our initial testing, which has been done, I think, with 70 teachers who were then role-playing their pupils. So actually your experience might be quite valuable. So we should chat after.

2:39

SPEAKER_01

I'm from Norway. So it's very great to see your ambitions and I'm sure that many countries across Europe are doing exactly the same thing. Are you doing any form of collaboration with other countries, sharing ideas, et cetera?

2:47

SPEAKER_01

Yeah, a bit. Norway, no. But if you've got contacts, I'm very happy to chat to them. We do a bit. There are a couple of teams that are similar to what we're doing, albeit a little bit differently. There are a couple of different initiatives underway in the US government, which are not a million miles away. Stuff like Tech Force, parts of the US Digital Service. Singapore as well. We talk quite a bit with Singapore. But yeah, we could do more. So if you have any contacts in the Norwegian government, I'd be very open to it. Done. Thank you. industry rather than an organisation. It's a big complex industry of 400,000 people. And I think we

3:15

SPEAKER_01

should look at it through that lens. Unfortunately, however, government has traditionally not been great at building and nurturing high performing technical teams. A lot of these issues are not specific to the UK. A lot of our American friends will be familiar with them. But just some of the commonly cited ones. Pay is a, you know, an obvious one. It makes it quite hard for us to compete for the best talent out there and then retain it. But also some barriers that are both real and perceived. So government is a very hierarchical organisation in many parts. There's a lot of bureaucracy. As a result, it can

3:56

SPEAKER_01

often move incredibly slowly. Not just for those reasons, but also the fact that, you know, there are regulations and safeguards in place that are very sensible because we're ultimately accountable to the public and to Parliament. But all of this can result in a system that is not always that appetising for high performing technical people to join. Especially the sort of people that we want. People that are impatient to leave their mark on the world. So what do we do that? Or what do we do about that? Changing a lot of these things, it's quite like a systemic challenge. It's like turning in an oil

4:33

SPEAKER_01

tanker, which is a bit of a tired cliche, but it's a good one. We're a little team at the centre. You know, turning that oil tanker is beyond the remit of any one team, let alone a scrappy little startup like ours. However, I think Cal Bear, our chief AI officer, is going to be closing out the conference this evening. That's very much his job and he's doing great work at the Department of Science and Technology to do just that. So I recommend everybody goes along to it. But there is a lot of political will at the moment to get stuff done and to make sure that this time around we are seizing

5:07

SPEAKER_01

new technology to actually make a dent in some of those problems I just mentioned. So the question put to us was, well, what can you do about it? And this was sort of the answer. As I said, we're like quite a small team at the centre. In terms of what we can do, we said, okay, well, let us take the shackles off. Let us basically set up a small insurgent unit at the very centre that is not sort of burdened by some of the constraints that I just described to you. What do I mean by an insurgency model? So we're setting up a new team. It operates with a mandate from number 10. We operate with an unusually

5:56

SPEAKER_01

high level of political backing to go into departments and get stuff done. We're able to pay market rates within reason. We're not paying like meta money necessarily. But the thing is, a lot of people will happily take pay cuts if we make it economically viable to come in and work on some of these challenges because they're interesting, right? We operate with an unusually high level of autonomy as well. We're able to be fairly opportunistic about the challenges we take on where we go into a department and see opportunities to have impact. Also, this one's pretty crucial. The civil service

6:31

SPEAKER_01

standard recruitment process is optimised for a lot of things but not necessarily recruiting exceptional technical talent. We've been allowed to recruit our own way. We've got a fairly grueling selection process that is laser targeted on technical skills. We've got a success rate of about 0.7, 0.8%. And most interestingly of all, and this is what differentiates us, we recruit exclusively outsiders. One of the best ways that I can have impact is by getting some people, the likes of this conference, into government. Because what has happened in the past is they tend not to leave and some of them end up

7:09

SPEAKER_01

setting up their own teams. I'll get into that later. Just when we're on this though, I don't want to make this sound like overly simplistic. A lot of people from the outside, particularly the tech industry, think that this bit alone is the only important thing that you need. That if you have a big enough stick for ministers, you can go in, you can, you know, break down data silos, you can do what you want. In practice, it's a lot harder than that. Otherwise, everybody would be doing it.

7:38

SPEAKER_01

And it's really early days. Like we're only setting out in this journey, but it turns out there's a huge amount of appetite for it. So we have been taking people from the labs. We've been taking people from big tech, from top research institutes. We've been taking YC founders, serial entrepreneurs, people who probably did not think they would be working in the civil service this time last year. But when you think about it, the decisions that go across a minister's desk are like some of the most important things you could possibly work on. So if you make it economically viable, and you promise people that you're going

8:13

SPEAKER_01

to put them in an environment where they can do their best work, it makes it really interesting. It's also worth pointing out as well, we do want to recruit missionaries, not mercenaries. So the pay matters, but it's not alone. Because a paycheck is not going to get you out of bed in the morning when stuff gets hard. And doing the stuff that we do does tend to be difficult. In terms of how we operate then, again, a bit different from normal government teams. And some of the there's like an abundance of low hanging fruit around the system. As you can imagine, it's a legacy organization. And there's

8:51

SPEAKER_01

lots of simple AI use cases that you can do in a few days to save money to improve service delivery, all that good stuff. When it comes to that, we largely do it ourselves. And that's the easiest and most satisfying part. As we speak, we've actually got the first forward deployed engineers in the history of 10 Downing Street, embedding themselves with policy and operational teams, teams of policy advisors, teams of lawyers, teams of comms people, pollsters, and everything in between. They're observing their workflows, their pain points, co-designing solutions with them to help them do

9:25

SPEAKER_01

their job more efficiently and effectively. And generally taking things from idea to implementation in a couple of weeks and getting new capability into the hands of users quickly. And then some of the other problems we talked about, I mentioned, you know, some of the huge backlogs in the system. That is not low hanging fruit. That's really complicated stuff. And normally, when it comes to those, we take more of a partnership model, where we will deploy some of our people into another team or another department, sometimes for prolonged periods. And I'm going to give you examples of

9:57

SPEAKER_01

both of these in a minute. I'm going to start with some of the low hanging fruit. It's worth pointing out, actually, a lot of the stuff that we do in number 10, I can't really show you. I know that sounds like an easy get out of jail card, but trust me, some of the stuff is a little bit sensitive. We're doing a lot of workflow automation, augmenting existing teams, as you can imagine. And here are a few other examples of stuff we've done just in the past few weeks. So policy simulation, that's turned out to be really interesting. So here we can allow policy teams in the building to test out the impact of different policy decisions before they're made. I think in this

10:40

SPEAKER_01

one here, we're looking at different decisions around universal credit and how they might impact... Oh, I've paused it. How they might impact household finances, amongst other things. But this can be applied to a broad range of stuff. Not replacing human analysis, necessarily. We're not putting ourselves out of a job. But what it has meant is that far more decisions in the building are being informed by high quality modelling and at a far faster rate than otherwise would have been the case. And this is another one just from the past couple of weeks. So the Cabinet Office was about to spend one and a half million pounds on getting an outside firm of lawyers

11:25

SPEAKER_01

to come in and do analysis of the entire UK statute book. Granted, the statute book is the height of four African elephants of legalese, but still pretty obvious AI use case. So we were going to spend one and a half million. Instead, one of our engineers embedded with that team of in-house lawyers for a couple of weeks. The benefits of this is not just money saved. Obviously, one and a half million is not nothing, but also speed. So the issue with the analysis that we were going to pay for is that it was going to be done slower than the pace at which new laws and regulations are made, which means you're going to have to do it again after a certain

12:04

SPEAKER_01

period of time. So now we've got this tool that that team can use and they can do it whenever they want at the drop of a hat. And we can also potentially open source it and share it with other teams in government.

12:18

SPEAKER_01

And then this is another one. So in number 10, we're responsible for the delivery of every major project and manifesto commitment in government. That means a lot of reports come in on how various things are doing. This is a little sort of delivery red teaming tool that the team spun up a couple weeks ago that is now being used every day. It's essentially a PMO that we've put in the pockets of delivery teams in number 10, not just so that they can interrogate the delivery reports that are coming across their table, but also give a second judgment on the teams that are reporting them.

12:56

SPEAKER_01

So it will flag up to decision makers in number 10, you know, does this team, does this department normally have a bit of optimism bias? Do they tend to disproportionately rate their risks as Amber? And are their mitigations usually effective or not? Also, aside from like AI adoption, having this capability in-house is really good. I think transparency is one thing that this country can do a bit better at. Up until a couple of months ago, the government had never published a public facing dashboard so that you lot can actually see how we're doing when it comes to delivery. But now we've published two in as many months. I think some of

13:40

SPEAKER_01

you might be familiar with the one on the left. This is the AI Opportunities Action Plan that Matt Clifford drafted about a year ago. This is how the UK is doing when it comes to rolling out compute and generally setting up the UK to be a leader in AI adoption. And now you can go online and see how we're actually doing. Yeah. Also, another thing that I can't show, but in two and a half weeks time, one of our ministers is going to launch a new public service that millions of people in the country are going to use. I can't go into more detail and steal their thunder. But in their words,

14:15

SPEAKER_01

it's hard to believe this didn't already exist. I assumed something like it already existed. That's something that we thought of two months ago and is now going to be live and used by the public. It is not an understatement to say that normally in government, that project like that might be in discovery for a year or more. So yeah. Let's get into the media stuff then. So that's nice, low-hanging fruit within the building. Now I want to talk about some of the work that we're helping other teams with across different parts of the ecosystem. For the purposes of this, I'm just going to focus on three

14:52

SPEAKER_01

of our partners. The AI Safety Institute, the Incubator for AI, and Justice AI. AC, I think pretty much everybody in this room will be familiar with it. Massive win for the UK. It's a great thing that we set up. We're a leading government body for evaluating frontier models and it was also the world's first. And we were really proud from day one to support it by putting a couple of our fellows in there to help them set up their cyber security work stream, amongst other things. I'll not dwell too much on this, but one of our early fellows was Dr. Harry Coppock. I don't think Harry is here, but we put

15:28

SPEAKER_01

him into AC from day one. And he led on their inspect tool, amongst other things. So there's a safe, isolated environment for testing what AI agents actually do when you give them autonomy and tools. And the Incubator for AI, which now sits in DESA. Who here is familiar with the Incubator? A few people. So the Incubator is essentially a spin-out of our program. It's a team that exists in the Department for Science and Technology that does what it says on the tin. It incubates new AI solutions for usage across the public sector. Its original finding team, most of the technical team were our fellows. And what's really cool now is not just seeing the work that they

16:14

SPEAKER_01

produced while they're there, but also the fact that we're able to collaborate with them when it comes to scaling up some of that work. Here's one recent example. So this tool is Extract. So a bunch of our people have worked on Extract. It's a collaboration with DeepMind. It's built on Gemini and it essentially digitizes large swathes of the planning application process, especially those bits that are currently largely handwritten, including handwritten, hand-drawn maps there as well. This was unveiled by the Prime Minister at London Tech Week last year and we're currently in the process of rolling it out

16:51

SPEAKER_01

to every local authority in England. As I said, only one in five planning applications are currently decided on time. That has a massive impact on economic growth and economic growth is basically the biggest challenges country faces right now. So anything we can do to make a dent on that is really significant. I think aspirationally as well, this will hopefully get us to a place where more and more planning applications can be decided by AI automatically. And then another interesting one, this is very current, the education gap is a big problem, not just in the UK but elsewhere. Many of you will have read

17:28

SPEAKER_01

the papers about AI tutors. It's a really exciting moment, the prospect of being able to level the playing field somewhat and put world-class tutors in front of every child regardless of their socioeconomic background. But it's something that has to be done really carefully. So currently we're working on producing safeguards and evaluating various frontier models against benchmarks, not just to make sure that children can interact safely with these in a classroom environment, but also measuring them against various metrics. I think in this one the relevant benchmark is the cognitive load placed on the student.

18:11

SPEAKER_01

And then last but not least, the new kids on the block, Justice AI. Some of you might have been here for the Justice AI talk yesterday. Was anybody here? Okay, good. For most of you this is new. Justice AI are a new team that have been set up in the MOJ, some of whom are over there. Hello.

18:32

SPEAKER_01

Again, I wouldn't say a spin-off from the fellowship, that gives us way too much credit, but the founder of Justice AI is one of our former fellows, Dan James, who's doing brilliant work in there. And they are deploying forward deployed engineers into prisons and into other parts of the criminal justice system. So kind of taking an approach that we're doing in number 10 with like policy people and comms people and lawyers, but instead they're embedding with parole officers and prison wardens. And they're doing loads of really interesting work. I can't go into too much detail. But most of it is around using AI to stop the flow of drugs into prisons, to find efficiencies,

19:13

SPEAKER_01

where currently there's quite manual processes involving lots of people, and generally improving the security and safety within the prison system. And one of those FDAs is over there. It's Will. Will's one of our current fellows. Sorry, Will, I've embarrassed you. I just added in your photo last night because I thought this was sort of a good point to end on. Will is... So to give you an idea, a few months ago, Will was in California getting a tan. That's him outside HMP Wandsworth on a rainy day. Will dropped out of Harvard, started a company, got it into Y Combinator, made a bit of money, but wanted to

19:56

SPEAKER_01

come and work for us. And that's his second week on the job and he's standing outside a prison with the keys to that actual prison about to go in. And that is exactly what we're trying to do through this program. You've maybe done good stuff in industry. That's brilliant. Come join us and we'll give you the keys to the state and see what you can do. So yeah, look, it's really early days. It's sort of an experiment what we're doing. But I think the proof points so far have been that actually like small elite teams can actually achieve quite a lot. We're already saving money. We're already shipping new public services at an unprecedented speed.

20:38

SPEAKER_01

We're already reforming frontline public services and we're already putting new AI capabilities into the hands of other teams at the top of government. So yeah. And surprise, surprise, this was a recruitment pitch. We are hiring. We're hiring. So please do scan the QR code and I'll be here the rest of the day if you want to come up and chat. Thank you.

21:09

SPEAKER_01

I think I've got time for a couple of questions, possibly. Somebody can tell me if not. Cause.

21:19

SPEAKER_01

Oh, sorry.

21:25

SPEAKER_01

So one on the, on one of the earlier example you showed, uh, there was this, uh, chat to explain policies and, um, not explain, but try different like projection and see how they would behave. Um, do you have to deal with, uh, uh, sim cofancy, uh, with the fact that, you know, like, if like you have a user that's not necessarily very well versed in AI, like just wants to hear what he wants to hear, can direct the tool towards, uh, look, I'm an absolutely brilliant mastermind. My police is going to be fantastic. Uh, despite the policy being actually bad, but the AI is basically. Yeah. Should I cut income tax to 0%? You're absolutely right. Exactly.

22:06

SPEAKER_01

Um, yeah, that's a very real risk. Um, so it's, it's not something that we have encountered too much, but it's only because we have, um, sort of read team the models for that before we've put it into the hands of users. We also provide quite a bit of upskilling. Um, so like a lot of the teams that we work with, we're creating tools for them. They're possibly lawyers. They're possibly sociologists, professors, whatever. So we do coach them on some of the risks that this presents. Um, but yeah, it's a good question. Thank you.

22:40

SPEAKER_01

Thank you. Hello. Thank you very much for the speech. Uh, I'm Jack from Accenture. I think you beat our pitch for hiring, but, uh, we'll try as well. Um, but my question on the, the policy side is, as you progress with this FD type of model actually making an impact, how did you see, or as you said, it's early day experiment. How did you see you start to scale and essentially start dealing with the central, uh, central government or with local governments, the different party lines and so on and so forth. Those kind of real kind of governmental kind of, you know, human kind of stance.

23:17

SPEAKER_01

How do you see that's going to play out? Yeah, a hundred percent. So in terms of how this scales. So I suppose at the end I said, oh, you know, we can do quite a bit. And some of the people that join us then set up their new teams. That's great. Is it enough to turn the oil tanker itself? No, possibly over time, but it would take a long time. And we need to solve these problems quicker than that. Um, so we have been thinking about that. Um, I think realistically, some of this stuff requires strategic, uh, intervention. Um, so we need to change the way the rest of government operates.

23:51

SPEAKER_01

Part of the reason, part of the bargain that we basically made with ministers was, you know, let us take the shackles off. Let us set up a small team at the centre that abides by different rules and use it as a proof point. So this is, this is almost like a pilot. I would like to see a lot of what we're doing, uh, become the norm, become BAU. This at the moment is basically a hack to get around the system. So we, we need to change that first of all. Um, I think also another, like if we're, if we're talking really about scale, uh, we, we talked about some like fairly targeted use cases there.

24:25

SPEAKER_01

Um, I think what we want to do over the next sort of 12 to 24 months is do more horizontal work, uh, looking at processes. So I should explain this to people as well. Like, um, when you think of the civil service, you probably think of policy people working in some of the buildings around here that you can see out the window. That is a very small sliver of the civil service. It's about 400,000 people. Most of them are call centre operators, they're, uh, prison wardens, they're nurses, et cetera, et cetera. Um, there are a lot of processes out there, whether that's like transcription, uh, that every police, police person will tell you is the bane of their existence,

25:06

SPEAKER_01

or it will be those massive call centres in DWP, HMRC. So yeah, if we want to dial up the ambition, I would like to see us going after more of those like horizontal use cases that can be implied on mass across the system. Um, yeah, sorry, that was a bit of a long answer.

25:30

SPEAKER_01

I think this is the last one. Yeah.

25:38

SPEAKER_01

You can shout if you want, it's up to you.

25:44

SPEAKER_01

Yeah, so I work for an ed tech company that among other things is making AI tutors. So I'd love to talk to you more about that. But one thing we find, I mean, the big problem with most kids is that, you know, you can make the best AI tutor in the world, but the real problem is motivation. You know, if you sit a kid down, if you sit a 12 year old down in front of a computer, they're going to do everything they can to avoid learning. So my question is like how, I mean, first of all, I'm just interested to know more about like what the vision is, the government, the government for this.

26:13

SPEAKER_01

Like, is this going to be going into schools? Are kids going to be sitting down in front of computers and using it? And secondly, like how do you solve that motivation problem? Yeah, 100%. So at the moment, I think our plan is largely not to necessarily develop products that compete with yours, but rather set... Good news. ...benchmarks and guardrails for how schools can then adopt whatever products they want. In terms of student uptake though, that's not something we've done too much on yet. The test that you just saw was our initial testing, which has been done, I think, with 70 teachers who were then role-playing their pupils.

26:52

SPEAKER_01

So actually your experience might be quite valuable. So we should chat after.

27:00

SPEAKER_01

I'm from Norway. So it's very great to see your ambitions and I'm sure that many countries across Europe are doing exactly the same thing. Are you doing any form of collaboration with other countries, sharing ideas, etc.? Yeah, a bit. Norway, no. But if you've got contacts, you're very happy to chat to them. Yeah, we do a bit. There are a couple of teams that are like sort of similar to what we're doing, albeit a little bit differently. There are a couple of different initiatives underway in the US government, which are not a million miles away. Stuff like Tech Force, parts of the US Digital Service. Singapore as well. We talk quite a bit with Singapore.

27:46

SPEAKER_01

But yeah, we could do more. So yeah, if you have any contacts in the Norwegian government, I'd be very open to it.

27:58

SPEAKER_01

Done. Thank you.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note