AI Engineer

Agentic SDLC at Uber — Uday Kiran Medisetty & Adam Huda, Uber

2118 summary words 9 min summary Watch video

Start with the signal

9 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Uber is turning its SDLC into a managed software factory by standardizing model access, tool access, execution environments, reusable skills, organizational context, and agent surfaces—then applying them to agent-assisted feature delivery and maintenance under centralized operational controls.
  • Why it matters: This is a concrete large-scale reference architecture for operating coding agents safely and economically across a monorepo, including gateway design, context engineering, validation placement, autonomous-maintenance governance, and feedback loops.
  • Best use: Use it as an architecture and operating-model benchmark for an enterprise agentic engineering platform, especially for control-plane boundaries and moving validation earlier than CI.

Executive Summary

Uber presents its "Managed Software Factory" as an enterprise platform rather than a single coding agent. Its reported scale is substantial: a few thousand engineers across 12 sites, more than 70% of PRs involving local or cloud agents, twice as many lines of code per engineer year over year, and over 250 automated migrations affecting roughly 9 million lines of code. The speakers explicitly credit prior standardization work—monorepos and Bazel—as foundational to making agents effective.

The platform rests on six integrated building blocks: a model gateway for privacy, policy, cost and attribution; an MCP gateway for controlled agent access to internal APIs and SaaS tools; pre-provisioned remote development environments for long-running, cross-repository work; a governed skills marketplace; a large context graph that unifies organizational and technical knowledge; and Cortana, a cross-surface AI assistant. The unifying design principle is that agents should consume centrally managed services rather than each team independently wiring models, credentials, tools, prompts, and context.

The end-to-end example moves from product ideation in Slack through business research, requirements, Figma variants, implementation by Uber's Minion cloud coding agent, validation, CI, review, and ongoing feature maintenance. The key process change is to shift more validation and review into the agent's inner loop before CI, so autonomous agents produce draft PRs only after static, visual, and integration checks have already been performed.

Uber's most useful caution is that agentic code generation moves the bottleneck rather than eliminating it. CI capacity, experiment capacity, human review bandwidth, and product decision-making become constraints once implementation gets cheap. Its maintenance design also avoids unconstrained autonomy: recurring agent loops are enrolled through a managed surface, scheduled around CI availability, throttled to avoid overwhelming engineers, and improved from the labeled outcomes of accepted or rejected changes.

Key Takeaways

  • Claim: An enterprise model gateway should be the mandatory control plane for every model request, not merely a shared API proxy. | Evidence: Uber routes internal coding and non-coding use cases through one OpenAI- and Anthropic-compatible endpoint. Middleware includes Spire identity/authentication, redaction of 20-plus PII types, and an AI guard using five specialized safety/policy models, all reportedly under 100 ms. The gateway supports attribution by user, caller, team, and catalog project; 800-plus projects make more than 100 million model requests per day through it. | Implication: Treat identity, privacy filtering, policy enforcement, spend management, audit traces, caching, and evaluation data capture as first-class gateway responsibilities before scaling agent usage. | Caveat: The presentation reports internal outcomes but does not provide comparative cost, quality, or latency benchmarks against decentralized model access.
  • Claim: Tool access needs its own gateway and token-efficient abstraction layer; otherwise MCP sprawl becomes an authentication and context-window problem. | Evidence: Uber found that thousands of internal APIs and SaaS systems had incompatible authentication/setup paths and that many MCPs created a significant token tax. Its MCP Gateway projects internal APIs into MCPs via a config change, centralizes hosting and token exchange for tools such as Google, Slack, and Jira, and now exposes 1,000-plus tools. Uber moved from direct MCP use to an Omni MCP discovery layer, then a CLI pattern to keep tool responses out of context, and finally a code-mode skill that generates Python scripts for high-token workflows; it reports over 40% fleet-wide savings from these optimizations. | Implication: Do not equate "more tools" with better agents: build discovery, authorization, response minimization, and task-specific execution patterns into a tool control plane.
  • Claim: Fast, isolated, cross-repository execution environments are a prerequisite for autonomous coding agents to work on real product changes. | Evidence: Uber uses pre-provisioned Kubernetes "balloon pods" with repositories already snapshotted and search indexes prebuilt, enabling an agent environment to start in seconds. It also introduced a mega-DevPort containing all repositories because agents and engineers increasingly need to work across prior language and repository boundaries. | Implication: For multi-service autonomous work, optimize environment startup, repository availability, indexing, isolation, and cross-repo permissions before optimizing agent prompts. | Caveat: This approach relies on Uber's existing remote-development, monorepo, and build-system investments; it is not a lightweight overlay for fragmented repositories and slow builds.
  • Claim: Skills should be managed as a governed, measurable internal product layer rather than as scattered prompts or repository-local scripts. | Evidence: Uber created a lifecycle around core and domain-specific skills: a marketplace with 2,500 skills, linting and automated review for baseline quality, one-command discovery/install, persona-based default installation, and trace/comment/continuous-evaluation feedback to skill owners. It reports more than 20,000 skill executions per day. | Implication: Create an ownership, publishing, quality-gate, distribution, telemetry, and evaluation model for reusable agent capabilities; otherwise duplication and uneven quality will dominate adoption. | Caveat: Uber characterizes skill quality and continuous evaluation as an ongoing investment, indicating that marketplace governance does not itself guarantee reliable skill behavior.
  • Claim: A unified context graph can improve agent reliability more than adding more disconnected MCP tools, because it eliminates repeated context discovery across enterprise systems. | Evidence: Uber observed agents wasting time and tokens locating services, dependencies, on-call owners, conventions, documents, incidents, and data sources across 20-30 systems. Its context graph has 150 node and edge types and 40 million entries spanning mobile, backend, data lake, design documents, Jira, incidents, and bugs. In early evaluations, including a query about cash mobility trips in India, Uber says graph-backed workflows materially improved tokens, turns, and latency. | Implication: Prioritize a permission-aware organizational/technical knowledge layer for planning, incident response, security, data analysis, and code changes instead of asking agents to rediscover enterprise context on every task. | Caveat: The talk does not provide absolute evaluation scores or details on freshness, permissions propagation, or graph-maintenance costs.
  • Claim: The practical agentic SDLC pattern is to validate and self-improve in an inner loop before CI, while preserving CI and human review as outer-loop controls. | Evidence: In Uber's example, Minion works across frontend and backend repositories but stops at a draft PR rather than immediately loading CI. Before CI, it can repair static-analysis issues, launch a simulator, compare screenshots against Figma specifications, and test frontend-backend integration in staging. CI failures can then trigger self-healing, while PRs display a table of completed checks and screenshots to give human reviewers evidence of the autonomous diff's validation history. | Implication: Design agents around evidence-producing validation gates, not raw code generation. Reserve expensive CI and human attention for changes that have already passed cheaper, task-specific checks. | Caveat: The demonstration is an intended system flow; it does not quantify how often visual validation, integration validation, or self-healing CI succeeds in production.
  • Claim: Autonomous maintenance should run as centrally managed, capacity-aware loops that learn from code-review outcomes rather than as unrestricted recurring agents. | Evidence: Uber allows features or services to enroll in maintenance skills such as feature-flag cleanup. The example schedules cleanup on Sunday when CI has more capacity and controls Monday's diff volume so engineers are not overwhelmed. Landed and rejected diffs, plus comments, become labeled data for skill improvement; incident reviews are examined monthly for new maintenance skills that can be applied fleet-wide. | Implication: Put recurring agents behind an explicit enrollment, scheduling, rate-limit, review, and learning system—and shift leadership attention toward prioritization and infrastructure bottlenecks as coding throughput rises. | Caveat: The speakers identify unresolved bottlenecks in CI capacity, experiment capacity, and decision-making; cheap implementation does not establish that a feature is worth building.

Detailed Brief

End-to-end product-flow example: stadium pickup feature

  • Claims: Uber uses a single assistant surface, Cortana, to carry an idea from conversational product discovery into research, requirement formation, initial design exploration, and implementation handoff.; The illustrative feature is a rider pickup-location experience intended to move riders away from crowds after large stadium events.; The intended rollout is North America, and the example explicitly includes two Figma/button-copy variants to support an A/B experiment.
  • Evidence: The workflow begins in Slack, where Cortana is asked to assess whether the opportunity is commercially worthwhile by looking at prior large-scale venue events and candidate stadiums.; Cortana then helps identify reusable application screens and backend components before work transfers to Minion, Uber's cloud coding agent.; Cortana is exposed through Slack, CLI, and web, and employees can create custom personas with team-specific skills and prompts; Uber reports 300 unique personas and more than 20,000 sessions per day in the preceding month.
  • Caveats: The feature example is a demonstration of intended orchestration, not a measured case study showing a completed production launch or business result.; The presentation acknowledges that the ultimate limiting question becomes whether to build a feature, not whether an agent can implement it.
  • Implications: A unified assistant can reduce handoff friction between product, design, and engineering only if it has access to shared business and technical context.; As implementation becomes faster, product research, experiment governance, and prioritization need stronger operating discipline rather than less.

Platform-scale outcomes and enabling foundations

  • Claims: Uber frames agent adoption as already broad across the engineering lifecycle, not an isolated developer-copilot deployment.; Its prior developer-platform standardization is presented as a necessary enabler of the current agentic SDLC.
  • Evidence: Uber reports that more than 70% of PRs are now produced with either local or cloud agents, and that lines of code per engineer doubled year over year.; It reports more than 250 automated migrations and approximately 9 million lines of code migrated automatically.; The stated foundations include six years of investment in monorepos, Bazel, and DevPort remote environments.
  • Caveats: Lines of code and PR involvement are throughput/adoption indicators, not direct evidence of quality, delivery speed, reliability, or customer value.
  • Implications: Agent productivity metrics should be paired with quality, operational-load, review, and business-outcome metrics to avoid optimizing for code volume.; Organizations without standardized build, repository, and environment layers should expect foundational platform work before reproducing Uber's autonomy level.

Notable Concepts & Terms

  • Managed Software Factory: Uber's framing for an integrated, centrally governed agentic SDLC spanning idea generation, implementation, validation, review, and maintenance.
  • Model Gateway: The mandatory model-access control plane for PII redaction, authentication, policy enforcement, attribution, cost controls, audit traces, caching, and model routing.
  • MCP Gateway / Omni MCP: Uber's centralized tool-access layer; Omni MCP provides one installable interface that discovers and invokes the available MCP tools.
  • DevPort / mega-DevPort: Uber's remote development environment, extended into pre-provisioned and cross-repository agent execution environments.
  • Minion: Uber's cloud coding agent, usable interactively or autonomously for multi-repository implementation work.
  • Context graph: A graph of technical, operational, and business knowledge intended to reduce agent context-search time, token consumption, latency, and unpredictability.
  • Inner loop vs. outer loop: The design pattern of moving fast, automated validation and lightweight review before CI, while retaining CI and deeper human/model review as outer-loop gates.
  • Managed maintenance loops: Scheduled, centrally controlled recurring agent tasks whose output volume, timing, and feedback data are managed rather than allowing uncontrolled autonomous jobs.

Operator Notes / Why Ken Should Care

  • Assess whether the current agent stack has explicit control planes for model access, tool authorization, auditability, per-project attribution, and budget enforcement; if not, avoid scaling direct vendor/API integrations team by team.
  • Make context retrieval a platform roadmap item: inventory the systems agents repeatedly search, identify permission-preserving graph or index primitives, and measure token/latency spent on context discovery.
  • Adopt a draft-PR-before-CI workflow for autonomous feature work, with machine-readable validation evidence attached to the PR; prioritize visual and integration checks where applicable.
  • Establish a governed skill registry with named owners, publication checks, execution telemetry, evaluation feedback, and a default-install policy by role or workflow.
  • For recurring remediation or maintenance agents, require enrollment, schedules, CI-capacity controls, output rate limits, and acceptance/rejection feedback capture before enabling unattended operation.
  • Track the post-agent bottlenecks explicitly: CI queue capacity, experimental capacity, reviewer load, and decision throughput. Do not use code-volume growth as the primary success metric.

Source/Metadata

  • Title: Agentic SDLC at Uber — Uday Kiran Medisetty & Adam Huda, Uber
  • Transcript words: 4282
  • Duration seconds: 1105
  • Timestamp note: No timestamps or chapter markers were provided; the transcript includes a repeated closing portion.
Full transcript 3250 words · 19 min read
0:00

.

0:12

Hey, let's get started. Good morning, everyone. I'm Uday. I'm here with my colleague, Adam. We'll talk about our journey towards Managed Software Factory. And in the beginning, in the first part of the talk, I'll talk about the key building blocks that we are investing in. And later, Adam's going to talk about how we take all of these blocks to build an end-to-end cohesive solution for our engineers. To set some context, we have a few thousand engineers across 12 global tech sites. Over the last year, all of the investments we made in Agent IKIA have led to more than 70% of our PRs now either by local or cloud agents.

0:26

And all of this led to twice the number of lines of code per engineer year over year. And this extends way beyond coding, and we see it in every aspect of the engineering lifecycle today. And we are also accelerating toil reduction at a massive pace. We handled more than 250 automated migrations, relatively 9 million lines of code automatically for our engineers. And before even the building blocks, all the investments we made over the last six years on moving to MonoRepos, moving to Bazel, all of that also laid a really solid foundation for us to accelerate this. So first, I'll cover all of these six building blocks.

0:41

And Adam's going to talk about a specific example and show how that feature can be built end-to-end with all of these. And all of these are in various stages of maturity and rollout within Uber. But we want to give everyone a sneak peek of what we are up to. So let's go to the building blocks, the six building blocks, one by one. The first one is model gateway. This is one of our earlier investments.

0:57

The three things that we wanted to ensure were: no PII ever leaves our perimeter to any of the vendors by default, any guardrail that we add here has strictly bounded latency, and every request that goes through this needs to be attributable per user, per project, and per team. So we have a model gateway. We made sure all of our internal use cases, our coding harnesses, our external use cases, they all go through one single OpenAI-Anthropic-compatible endpoint. It goes through a series of middlewares. The first one is identity and authentication using Spire. We have a data anonymizer that redacts 20-plus PII types.

1:12

We have an AI guard that has five specialized models that handle various parts of safety and policy that we want to ensure. And all of that runs under 100 milliseconds. We also are investing in all kinds of caching and token optimization strategies at this layer. And every request that goes through this, we are able to attribute to a specific project in our catalog. And we can attribute per caller, per user, per team, both in real time and also in our data lake.

1:26

This enables us to create all kinds of spend tiers and guardrails in a holistic way across our portfolio. We also use this layer for capturing audit log session traces, which are then plugged into our benchmarking and all kinds of self-improvement loop efforts. And for an engineer at Uber, you take the vanilla client, you set the project ID, and we take care of everything else. Today, we have 800-plus projects internally going through this, cumulatively handling more than 100 million model requests per day. This includes both the frontier models and also open source models, whether that is hosted in our infrastructure or some of our vendors.

1:41

The next is, how do we provide tools to all of these models? Last year, when we started on this journey, we had thousands of internal APIs, but none of them were agent-accessible out of the box. And we had so many other SaaS tools, and each one of them had a different way to authenticate, a different way to set up, which is a lot of hassle for everyone. And once you end up with enough MCPs, they'll all add up and have a massive token tax. Similar to Model Gateway, we have an MCP Gateway that handles a whole bunch of middlewares for engineers, and we have an automated crawler that looks at our internal APIs and projects all of these into MCPs with one single config change.

1:53

And we do the same thing even for our SaaS MCPs. Whether it's Google, Slack, Jira, all of this, they go through the MCP Gateway. We host them. We do the token exchange. So for all the engineers, they go through one single entry point, one common way to install any MCPs. This simplified a lot for all of our engineers and employees. And then a whole bunch of token optimization strategies. We initially had a direct MCP pattern. Earlier this year, we created Omni MCP, which is one single MCP that you install, which can discover and invoke any MCPs within the Gateway.

2:23

And a couple of months ago, we projected all of these MCPs into a CLI pattern so that even the response doesn't eat up your context. And of late, we also have a code mode skill, which is auto-installed, which on the fly creates Python scripts to hyper-optimize some of the top MCP token-consuming use cases. And all of this led to, now we have 1,000-plus MCP tools. And just with these optimization efforts, we've saved more than 40% fleet-wide savings. So once we have the models and the tools, we need a place to run all of this. For many years, we had DevPort, which is our cloud remote environments. We had this because we had large monorepos with millions of lines of code.

2:43

And this is how engineers work at Uber. And now we took what we had with DevPorts and we identified that now we need an environment for agents to run for a longer period of time. They need to be quick. They need to be isolated. We can install any number of them. And they need to be globally available across all of our sites. So we have pre-provisioned Kubernetes balloon pods. When an agent requires a new environment to run, it can take one of those, which is already pre-provisioned. It has all of the repositories already snapshotted. The search index is already built. So the agents can start working within a matter of seconds.

3:19

The next thing we noticed is the roles of engineers are getting blurred. We used to offer a DevPort per language flavor for Go, Java, Android, and so on. Now we need agents to work across repositories and engineers also to work across repositories. So we have a mega-DevPort that has all of the repositories in one common place. And this is what we use for our autonomous coding agents now. And even for our non-engineering employees, we are providing a simple way for them to get started with any of the agent harnesses in a matter of seconds. Now we get to the knowledge part of it. And we jumped on this bandwagon earlier this year.

3:50

We started noticing engineers building tons of skills across many repositories. One of the problems we noticed was there's a lot of duplication, the same skill being built by different engineers in different trepots. And discovery and configuration were a huge hassle. And a lot of skills were of subpar quality. So we built an entire lifecycle around skills. So we have core skills and domain-specific skills. All of that goes into a managed skills marketplace. We have 2,500 skills there right now. And it goes through a whole bunch of lint checks, automated reviews, which ensure a baseline skill quality for any skills that we have.

4:32

And we also simplified the installation and discovery. So there is one single command to discover and install any plugin in our ecosystem. And based on the engineer personas, we even auto-installed some of the default skills. So the agents automatically can pick up the right skill. You don't even have to install them. And of late, we started working on collecting traces and comments and capturing continuous evals so that we can give feedback back to the skill owners for skill improvements. And this is an area of big investment for us right now. And we have 2,500 skills and relatively more than 20,000 skill executions per day across our fleet.

4:55

The next piece of knowledge is context graphs. We started noticing in our execution traces agents spending a lot of time even trying to find basic context, especially in our large monorepos. You need to identify where the service is located, what are the dependencies, who on set, what kind of patterns I need to follow. And all of this context is gathered across scattered systems across Uber. There's 20 to 30 different systems. Each needs its own skills, its own MCPs to gather the context.

5:27

And this burns tokens. This adds a lot of latency.

5:40

And it creates more unpredictable outcomes. So we have one context graph. We took all of the information of how Uber runs into one context graph. This has 150 unique node and edge types. We have 40 million entries there right now. It captures all the way from how our mobile apps are built to our backend to our data lake, all the design docs, Jira, incident bugs. Everything is connected. And this enables agents to quickly find the right context within our ecosystem. We are now plugging all of our skills and use cases into the graph, whether it's our on-call RCA, whether it's our planning or data analysis or security scans.

6:10

And we see across all of this, we are improving the skills by a lot. And I'm just showing a very simple example of asking a simple question of how many mobility trips in India are cache. This needs to understand the concepts of each of these, which tables, what kind of CTEs you need to create for the SQL. With and without graph, we see massive improvement in tokens, turns, and latency. And we see that across any early eval that we did within our infrastructure. And the last thing is, how do we package all of this for everyone in the company to use? So we have our AI assistant called Cortana.

6:31

All of the things that I mentioned so far, whether it's skills, MCPs, and context graphs, they're all plugged into that in every surface possible, whether it's on Slack, CLI, web. So anyone in the company can ask a simple question. It can look up the context graph, invoke any skill, check any code, check any code in any codebase, and give an answer across any of these surfaces. And now we started allowing employees to even personalize that. You can hook up your custom skills, custom prompt, and hook it up into your team Slack channel so that it knows all of the things about that team and works like that teammate.

6:57

And this is a simple example of how you can invoke the same question before in Slack. And all of the employees, one or more people, can even collaborate on the same Slack channel.

7:09

And we have, just in the last one month, 300 unique personas created and more than 20,000 sessions per day. I'll now pass on to Adam, who will talk about how we take all of this and ship a feature end to end. Adam All right, thank you, Udaya. All right, as Udaya said, we've got those building blocks. We're going to use those to power our software factory. So we're going to take a feature here and show it going end to end through this. All right, first up, we need to have an idea, right? A good idea probably for this moment would be something around the World Cup, right?

7:38

Wouldn't it be awesome if you were a rider and you're leaving a busy stadium, if there was a better pickup location to get you away from the crowd? So that's the idea, right? We have our idea, we're jamming on it in Slack here. Let's tag in Cortana, right? That's our AI assistant, to help us with that idea. Cortana, with that context graph, can help us determine whether this is a good business opportunity to go after. So we can go here from Slack and now open Cortana into a web interface. And you'll see an example here of what that business research could look like, right? What other large-scale venue events have happened before?

8:12

What are some stadiums that would make sense here? From there, we start to think about the product requirements, right? This should probably be just a North America rollout since that's where the stadiums are. We can even then bring in Cortana to help us think about the Figma designs. We can create some initial mockups, right?

8:37

Do two variants here. We want to run an experiment, A and B. So the button strings here are different between the two. So we'll test those two variants and see which one performs better. Now we'll start to think a little bit about the design. And Cortana too can help us think about what code changes would need to happen, right? What can we leverage that's in the app already? What screens, and what can we leverage on the back end? So this process before could take a long time. It could take weeks to get everyone aligned. Now we can compress this into a very short amount of time, right, and get to a prototype here very quickly. All right, so now we got to go build this.

9:17

So we hand off from that Cortana agent to what we have at Uber. We have a Minion agent. It's Uber's cloud coding agent solution. All right, so you can use Minion in an interactive mode or you can run it in an autonomous mode as well. So I'm going to show you what this looks like. You and I mentioned the DevPod building block. So this is powered by that DevPod. So it's got a full build environment and it can work across repos. So we're doing back-end changes and the front-end changes here too as well. We're going to see Minion progress here, and it's going to stop at just creating a draft PR. And it's not going to push it to CI yet.

9:59

The reason is that we were seeing that this is great for doing toil workloads, but to build more advanced end-to-end features, we really need to be able to validate the feature first. And we want to prevent a lot of extra load coming onto CI. So if we can validate sooner before we push to CI, that would be a big benefit. So that's what we're going to see here next on validation. Right? In the SDLC, we have an inner loop. We have the outer loop, of course. We can have these be agentified, where we're shifting more checks now to happen in this inner loop. So some of the checks that initially happened, that we've had there previously, are the static analysis checks.

10:31

When those are detected, now we fix those. But we can shift things that happen in the inner loop, things like visual validation. So we can launch a simulator with a skill, grab a screenshot from the simulator, compare it to the Figma specs. We can also bring up the service in our backend staging environment and compare the front-end and the back-end integration together. So now that we've done that part, we move to the outer loop where CI typically happens. Errors can still happen on CI. So self-healing CI is something that we've implemented here where we can fix a lot of the issues that you hit on CI.

11:07

Code review is another thing that happens in the outer loop, but this is another thing that we've shifted. We've moved parts of code review to happen in the inner loop. The outer loop code review can have a powerful model, use reasoning, a skill to do a deeper review. And in the inner loop, we can have a smaller model that runs faster with the medium model. Now, another key thing here is this: if this is an autonomous diff coming from Minion, we want to give a human reviewer some confidence that this diff has gone through a lot of self-improvement already. That did not just test that initial generation that happened, but all these other steps have happened.

11:35

And so on the PR, you will have a table attached that says all these different checks that it went through, including the screenshots. All right, so we've got a lot more code coming through the software factory now. Let's talk about maintenance. Maintenance is even more important. What we have set up now is we can actually enroll our feature or service into maintenance skills. So these are some examples of those skills that we have. Feature flag cleanup, right? We had two variants of that World Cup modal. Now that the B variant is no longer needed, we can have that scheduled on a loop. So the key thing here is that this is actually a managed loop that you go to.

12:06

We don't want thousands of loops being set up across the company without any bounds. You have a managed surface that you go to to set up the loop. So it runs on Sunday when we know we have better CI capacity available. We also don't want to overwhelm engineers Monday morning with a bunch of extra diffs. We want to control how many diffs they're seeing on Monday as well. Another cool key thing here is that when that skill runs and makes those diffs, those diffs will get comments and either get landed or not landed. That's all good labeled data that we can use to improve the skill itself.

12:25

And then at a monthly cadence, we're looking to see what skills we can now learn from in our incident reviews and turn those into new maintenance skills that we can apply to all of our services. All right. So you've seen these parts of the SDLC that we have identified. There are other parts too, like monitoring. You've seen the building blocks that you can use to power those and the architecture underneath them. One of the other things that we're really thinking about now is bottlenecks. We're now putting more strain on our infrastructure. So we're trying to anticipate where our CI capacity needs to be and make the right foundational investments there.

12:42

There's only so many experiments that we can feasibly run as well. So that's another bottleneck. And then lastly, decision-making, right? It's not about can we build. We know we can probably build it now. It's more of a question of should we build it? All right. So that's what we have for you today. And the next talk is going to be from Uber as well. And if you want to learn more about our Agenda code review, Amaya and Will will be presenting that next in this room. Thank you. Thank you. It could take weeks to get everyone aligned. Now we can compress this into a very short amount of time now, right, and get to a prototype here very quickly.

13:17

All right, so now we got to go build this. So we hand off from that Cortana agent to what we have at Uber. We have a Minion agent. It's a Uber's cloud coding agent solution. All right, so you can use Minion in an interactive mode or you can run it in an autonomous mode as well. So I'm going to show you what this looks like. You and I mentioned the DevPod building block. So this is powered by that DevPod. So it's got a full build environment and it can work across repos. So we're doing back end changes and the front end changes here too as well. We're going to see Minion kind of progress here and it's going to stop at just creating a draft PR.

13:50

And it's not going to push it to CI yet. The reason being is that we were seeing that this is great for doing like toil sort of workloads, but to build more advanced like end to end features, we really need to be able to validate the feature first. And we want to prevent a lot of extra load coming on to CI. So if we can validate sooner before we push to CI, that would be a big benefit. So that's what we're going to see here next on validation. Right? In the SCLC, we have an inner loop. We have the outer loop, of course. We can have these be agentified where we're shifting more checks now to happen in this inner loop. So some of the checks that initially happened

14:29

that we've had there previously is like the static analysis sort of checks. When those are detected, now we fix those. But we can shift things that happen in inner loop, things like visual validation. So we can launch a simulator with a skill, grab a screenshot from the simulator, compare it to the Figma specs. We can also bring up the service in our back end staging environment and compare the front end and the back end integration together. So now that we've moved a, now that we've done that part, we move to the outer loop where CI typically happens. Errors can still happen on CI. All right? So self-healing CI is something that we've implemented here

15:07

where we can fix a lot of the issues that you hit on CI. Code review is another thing that happens in the outer loop, but this is another thing that we've shifted. We've moved parts of code review to happen in the inner loop. Right? The outer loop code review can have a powerful model, use reasoning, a skill to do a deeper review. And in the inner loop, we can have a smaller model that runs faster with the medium model. Now, another key thing here, right, is this, if this is an autonomous diff coming from Minion, we want to give a human reviewer some confidence that this diff has gone through a lot of self-improvement already.

15:40

That did not just testing that initial generation that happened, but all these other steps have happened. And so on the PR, you will have a table attached that says all these different checks that it went through, including the screenshots.

15:53

All right, so we've got a lot more code coming through the software factory now. Right? Let's talk about maintenance. Right? Maintenance is even more important. What we have set up now is we can actually enroll our feature or skill into our feature or service into maintenance skills. So these are some examples of those skills that we have. Feature flag cleanup, right? We had two variants of that World Cup modal. Now that the B variant is no longer needed, we can have that scheduled on a loop. So the key thing here is that this is actually a managed loop that you go to. Right? We don't want thousands of loops being set up across the company without any bounds.

16:33

You have a managed surface that you go to to set up the loop. So it runs on Sunday when we know we have better CI capacity available. We also don't want to overwhelm engineers that Monday morning with a bunch of extra diffs. We want to control how many diffs they're seeing on Monday as well. Another cool key thing here is that when that skill runs and makes those diffs, those diffs will get comments and either get landed or not landed. That's all good labeled data that we can use to improve the skill itself. And then at a kind of monthly cadence, we're looking to see what skills can we now learn from in our incident reviews

17:07

and turn those into new maintenance skills that we can apply to all of our services.

17:14

All right. So you've seen these parts of the SDLC that we have identified. There's other parts too, like monitoring. Have you seen the building blocks that you can use to power those and the architecture underneath them? One of the other things that we're really thinking about now is bottlenecks. We're now putting more strain on our infrastructure. So we're trying to anticipate where our CI capacity needs to be and make the right foundational investments there. There's only so many experiments that we can feasibly run as well. So that's another bottleneck. And then lastly, too, right? Decision making, right? It's not about, you know, can we build?

17:47

We know we can probably build it now. It's more of a question of should we build it? All right. So that's what we have for you today. And the next talk is going to be actually from Uber as well. And if you want to learn more about our Agenda code review, Amaya and Will will be presenting that next in this room. Thank you.

18:21

Thank you.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note