From Tokenmaxxing to Trusted Throughput — Mingsheng Hong, Ironclad
Description
An engineer at a large tech company built a voluntary dashboard showing everyone's AI token usage, and colleagues promptly started competing to top it. Mingsheng Hong uses that as the thing not to do. His team runs the same dashboards, but treats them as a smoke detector rather than a leaderboard: a team using surprisingly few tokens is worth a conversation, and nobody should ever be rewarded for burning more. He draws the parallel to lines of code, a number worth tracking and a terrible thing to optimize, given that deleting code is often the better outcome. The pitfall he flags is going straight from measuring cost to cutting it, because that is only one side of a ratio. So Ironclad measures value too, and the metric evolved in public: lines of code, then open pull requests, then merged ones, and now merged pull requests weighted by a complexity score, since a ten line concurrency fix is not a thousand lines of boilerplate. He calls the target trusted throughput, work that clears objective checks, human review, and finally contact with customers. The bottleneck has moved downstream to review and CI, where slow pipelines quietly push engineers toward giant batched pull requests that are harder to review well. His fix is unglamorous: kill flaky tests, cap agent retry loops, and measure the wait from ready to merged. Speaker info: - https://www.linkedin.com/in/mingshenghong/ Timestamps: 0:00 - Token leaderboards, and why they backfire 1:34 - Dashboards as smoke detectors 2:56 - Getting past adoption before managing cost 4:19 - The engineers who lost the craft 5:44 - Trust as the product constraint at Ironclad 7:09 - Measuring cost across several vendors 8:33 - Why cutting cost first is premature 9:58 - Lines of code, and metrics you should not optimize 11:20 - From open pull requests to weighted merges 12:43 - What trusted throughput actually means 14:06 - The bottleneck moves to review and CI 15:30 - AI as the first pass, humans for judgment 16:56 - Flaky tests,
Summary
Generated by claude-sonnet-4-5-20250929At-a-Glance
- Verdict: Watch fully
- Core thesis: Engineering orgs should shift from minimizing token spend to maximizing ROI via 'trusted throughput'—code that passes internal validation (review + CI) and delivers customer value—while proactively addressing new bottlenecks in code review and CI/CD created by AI-abundant code generation.
- Why it matters: Ken is designing agent systems and orchestration control planes where cost optimization, trust validation, and operational loops are first-class concerns; Ironclad's framework (guardrails + learning loops + institutional best practices) and their trust-building approach (objective + subjective + customer validation) directly inform how to operationalize multi-agent workflows at scale without runaway spend or quality collapse.
- Best use: Extract Ironclad's three-bucket trust model (objective metrics/subjective review/customer validation), their pragmatic token optimization framework (guardrails/learning loops/best practices), and their bottleneck diagnosis (review + CI overload from abundant code gen) for agent control-plane design and customer trust strategies.
Executive Summary
Mingsheng Hong, VP of Engineering for AI at Ironclad, argues that the engineering leader's job is not to minimize AI token spend but to maximize ROI by optimizing for 'trusted throughput'—code that earns internal and customer trust through validation. Ironclad, a legal contracting AI company where trust is paramount, applies the same trust-building principles internally: objective checks (tests, security, canary rollouts), subjective human judgment (code/design review for quality and architecture fit), and customer validation (production stability, no rollbacks or usability complaints). Hong cautions against treating usage dashboards as leaderboards (à la Amazon's tokenmaxxing anecdote) and instead positions them as smoke detectors for adoption gaps or usage anomalies, not cost-cutting targets.
The talk introduces two bottlenecks created by AI-abundant code generation: (1) code review overload, and (2) CI/CD infrastructure strain from many smaller PRs. Ironclad's response is to deploy AI tooling as first-level review defense (catching style, missing tests) so human reviewers focus on subjective judgment (architecture, security design, maintainability), and to invest in platform engineering to eliminate flaky tests and reduce wall-clock merge time. Hong warns against the anti-pattern of bundling many changes into giant PRs to avoid CI queue time, since that degrades review quality and increases risk.
To operationalize token optimization, Ironclad uses a three-part framework: (1) Guardrails—budgets, quotas, anomaly detection, and regular human review; (2) Best practices innovation—prompt caching, context pruning, agentic-loop step limits, and institutional prompt libraries for common tasks; (3) Learning loops—leadership reviews metrics with teams, refines guardrails, and feeds insights back into institutional knowledge. Hong emphasizes measuring the value side of ROI: evolving from lines of code → open PR count → merged PR count → complexity-weighted merged PRs (scored by LLM on a T-shirt scale). The goal is to ensure engineers feel AI increases their leverage and preserves high-impact technical work, not just slop review.
Key Takeaways
- Claim: The goal is not to minimize token spend; it is to improve ROI by maximizing 'trusted throughput'—code validated internally (review + CI) and externally (customer production success). | Evidence: Ironclad defines trusted throughput via three buckets: objective metrics (test coverage, security checks, canary rollouts), subjective human judgment (code/design review for quality and architecture), and customer validation (no rollbacks, no usability tickets). They explicitly reject using usage dashboards as leaderboards and instead treat them as smoke detectors for adoption gaps or anomalies. | Implication: For Ken's agent orchestration, the control plane must instrument trust validation at each layer—objective telemetry, human approval checkpoints, and customer outcome tracking—rather than optimizing for cost alone; the framework can inform how to design agent feedback loops that prove value before scaling spend.
- Claim: AI-abundant code generation shifts bottlenecks downstream to code review and CI/CD, risking the anti-pattern of bundling many changes into giant PRs to avoid queue time. | Evidence: Hong observes that if CI takes an hour per run, engineers avoid splitting PRs into 10 smaller ones (which would take 10 hours total), instead submitting large PRs that overload human reviewers and reduce review quality. Ironclad's response: deploy AI review as first-level defense (style, coverage checks) so humans focus on architecture/security, and invest in platform engineering to remove flaky tests and measure wall-clock merge time and retry counts. | Implication: Ken should design agent pipelines with explicit review and validation stages that scale independently of generation throughput, and instrument merge-time and retry metrics as leading indicators of control-plane health; the pattern applies to multi-agent workflows where abundant agent output can overwhelm human or automated validation capacity.
- Claim: Evolve value measurement from lines of code → open PRs → merged PRs → complexity-weighted merged PRs (LLM-scored on T-shirt size) to approximate ROI. | Evidence: Ironclad started with LOC (discarded as gameable), then open PR count (inflection observed but not final value), then merged PR count (ships matter), and finally added a complexity score via LLM prompt to weight each merged PR by business value (e.g., a 10-line concurrency fix vs. 1000-line boilerplate). | Implication: Ken can adapt this progression for agent output: start with agent invocations or raw output tokens, move to successful completions, then add a value score (impact on user goal, downstream tool success, human approval rate) to measure agent ROI; the LLM-as-judge pattern for weighting output complexity is directly applicable. | Caveat: Hong admits there is no traditional definition of complexity and this is an evolving pragmatic approach; they plan to refine the metric further.
- Claim: Token optimization requires a three-part framework: guardrails (budgets, quotas, anomaly detection), best practices innovation (prompt caching, context pruning, agentic-loop step limits, institutional prompt libraries), and learning loops (leadership reviews metrics with teams and refines guardrails). | Evidence: Ironclad sets budgets and anomaly alerts, encourages engineers to structure prompts for caching (fixed system prompt at top, varying user prompt at bottom), prunes context in long sessions, limits agentic retry loops to avoid runaway spend, and maintains an internal playbook of well-crafted prompts for common tasks (bug fixes, new UI features, refactors). Leadership reviews usage dashboards with teams contextually (platform vs. UI teams use AI differently) and feeds learnings back into institutional knowledge. | Implication: Ken should build similar three-layer governance for agent systems: policy/quota layer, operator toolkit (reusable prompts, retry budgets, context summarization), and a learning loop where usage patterns inform agent prompt templates and cost allocation rules; this structure ensures agents scale under control.
- Claim: Before implementing cost controls, orgs must pass the adoption hump; premature optimization can stifle usage and engineer morale. | Evidence: Hong asked the room and saw ~50% had passed the adoption phase. He shared that some engineers resist AI because they lose pride in handcrafting code and feel reduced to reviewing 'AI sloth code,' so leadership must sit down with resisters to identify high-impact technical work that preserves growth and satisfaction in the AI era. | Implication: Ken should sequence agent rollout to ensure adoption first (easy access, clear value demos) before layering on cost controls; for customers, this means early agent pilots should focus on trust and value proof, not cost optimization, to avoid killing adoption.
- Claim: Build versus buy decisions should favor buying non-differentiating tooling (IDE, CI) and building context-specific assets (internal prompt playbooks, custom agents). | Evidence: Ironclad buys IDE and CI infrastructure but builds an internal playbook of AI prompts for domain-specific tasks (bug fixes, UI features, refactors) that get reused and enhanced across teams. They are exploring building a 'builder agent' wrapper around Cloud Code/Codecs but are still evaluating vendor alternatives for ambiguous cases. | Implication: Ken should default to vendor agent platforms (OpenAI Assistants, Anthropic Claude, LangChain) for generic orchestration but build proprietary agent prompts, routing logic, and validation workflows for OpenClaw and client-specific use cases; this maximizes leverage while preserving differentiation.
Detailed Brief
Dashboard Design and Anti-Patterns
- Claims: Ironclad uses AI to build pipelines that extract usage data from multiple vendors (Cloud Code, Codecs) and cross-correlate it by team and individual.; The dashboard's job is to detect adoption gaps (teams/individuals underusing AI) and usage bursts (sudden spikes requiring investigation), not to rank or reward high token usage.; Contextual comparison matters: platform/infrastructure teams use AI differently from UI teams, so raw usage comparisons are misleading.
- Evidence: Hong cited Amazon's tokenmaxxing story and Meta's similar leaderboard culture as cautionary tales, plus a company spending $500M on cloud in a month.; He compared token usage to lines of code: an important metric but not a goal; removing code can be better than adding it, just as efficient token use can be better than maximizing spend.
- Implications: Ken's agent telemetry dashboards should segment by use case (e.g., research vs. customer-facing agents) and flag anomalies, not encourage competition; the smoke-detector metaphor applies to agent cost monitoring.
CI/CD Metrics and Developer Experience
- Claims: Ironclad measures wall-clock time from PR-ready to PR-merged, and the number of retries required to pass tests.; If CI takes an hour and typical submission takes 2-3 hours, that's a red flag for CI infrastructure quality.; Developer experience teams are accountable for reducing flaky tests and improving CI throughput to prevent engineer frustration and workarounds.
- Evidence: Hong described engineers either manually babysitting PRs through flaky test retries or recruiting AI agents to babysit in a loop, both of which are frustrating workarounds, not solutions.
- Implications: Ken should instrument merge latency and retry counts for agent-generated code or multi-agent workflows; high retry rates signal validation or orchestration bottlenecks that need platform investment.
Prompt Engineering Best Practices
- Claims: Prompt caching: structure prompts with fixed system prompts at the top and varying user content at the bottom so vendors can optimize prefix processing.; Context pruning: train engineers to summarize context in long chat sessions to avoid token bloat and improve output quality.; Agentic loop step limits: cap retry loops (e.g., auto-fix-and-retry after test failures) to prevent runaway token spend if logic goes off-track.
- Evidence: Hong mentioned that vendors like Cloud Code now auto-compact context, increasing both efficiency and quality.; He gave the example of engineers writing agentic loops in Cloud Code harnesses that retry until tests pass, with the risk of unbounded spend if not capped.
- Implications: Ken should bake prompt structure (system/user separation), context summarization, and retry budgets into agent orchestration templates and libraries; these are reusable patterns that reduce cost and improve reliability.
Notable Concepts & Terms
- Trusted Throughput: Ironclad's proxy metric for engineering ROI from AI: code that passes objective checks (tests, security, canary), subjective human review (quality, architecture), and customer validation (production stability, no rollbacks/complaints); the goal is to optimize this, not minimize token spend.
- Tokenmaxxing: The anti-pattern (from Amazon/Meta stories) where engineers compete to maximize token usage on a leaderboard, incentivizing spend rather than value; Ironclad explicitly rejects this in favor of treating dashboards as smoke detectors for adoption gaps or anomalies.
- Smoke Detector Dashboard: Ironclad's positioning of usage/cost dashboards: they alert on local underuse (adoption gaps) or spikes (anomalies), not rank or reward high usage; the goal is continuous learning, not cost-cutting or competition.
- Complexity-Weighted Merged PRs: Ironclad's evolving value metric: each merged PR is scored by LLM on a T-shirt complexity scale (simple boilerplate vs. complex concurrency fix) to approximate business value, moving beyond raw PR count or LOC.
- Builder Agent: Ironclad's internal exploration: a Cloud-based code generation wrapper around vendor tools (Cloud Code, Codecs) for context-specific tasks; they are still evaluating build vs. buy for this use case.
- Prompt Caching: Vendor optimization where fixed prompt prefixes (system prompts) are processed once and reused; users should structure prompts with fixed content at the top and varying user input at the bottom to maximize efficiency and reduce cost.
- Context Pruning: The practice of summarizing or compacting context in long chat sessions to avoid token bloat and improve output quality; some tools (e.g., Cloud Code) now auto-compact context.
- Agentic Loop Step Limits: Caps on the number of retry iterations in agent workflows (e.g., auto-fix-and-retest loops) to prevent runaway token spend if logic goes off-track; a key guardrail for autonomous agent systems.
Operator Notes / Why Ken Should Care
- Instrument trust validation in three layers for agent systems: objective telemetry (test pass rates, latency, error codes), human approval checkpoints (review gates for high-stakes outputs), and customer outcome tracking (production success, rollback rates, support tickets); this is Ironclad's trust model applied to agents.
- Design agent pipelines with explicit review and validation stages that scale independently of generation throughput; if agent output is abundant but validation is bottlenecked, the system will collapse or degrade quality (Ironclad's review/CI bottleneck lesson).
- Adopt complexity-weighted output scoring for agent ROI measurement: start with invocations, move to successful completions, then add LLM-as-judge scoring for output business value; this mirrors Ironclad's progression from LOC → merged PRs → complexity-weighted PRs.
- Build three-layer agent governance: (1) guardrails—budgets, quotas, anomaly alerts, (2) operator toolkit—reusable prompts, retry budgets, context summarization, (3) learning loops—usage review with teams, refine prompts and cost rules; this is Ironclad's optimization framework.
- Bake prompt structure (fixed system prompt at top, varying user input at bottom), context summarization, and retry step limits into agent orchestration libraries; these reduce cost and improve reliability without custom engineering per use case.
- Default to vendor platforms (OpenAI, Anthropic, LangChain) for generic orchestration but build proprietary prompt playbooks, routing logic, and validation workflows for OpenClaw and client-specific agents; this is Ironclad's build-vs-buy heuristic.
- Sequence agent rollout to ensure adoption and value proof before layering on cost controls; premature optimization can kill morale and usage (Ironclad's adoption-hump lesson applies to customer pilots).
- Instrument merge latency and retry counts for agent-generated code or multi-agent workflows; high retry rates signal orchestration or validation bottlenecks that need platform investment (Ironclad's CI/CD metrics).
- Treat usage dashboards as smoke detectors for adoption gaps or anomalies, not leaderboards; segment by use case (research vs. customer-facing) and review contextually with teams to avoid misaligned incentives.
- For high-stakes agent domains (legal, compliance, security), adopt Ironclad's trust-building playbook: let users test with known inputs, validate outputs match expectations, then expand scope; this mirrors how Ironclad's lawyers gain confidence in conversational search before using it for redlining or anomaly detection.
Source/Metadata
- Title: From Tokenmaxxing to Trusted Throughput — Mingsheng Hong, Ironclad
- Transcript words: 3709
- Duration seconds: 1384
- Timestamp note: Timestamps unavailable; transcript lacks time markers.
Transcript
All right. Let's get started. Apologies for the delay, but I'm really excited to be here. I'm Ming Shen, VP of Engineering focused on AI at Ironclad. And today I'll be telling you about something that's probably on top of many of your minds: how to control and optimize for your AI token spend. Can I get a quick show of hands that this is a relevant topic? Okay. Awesome. I appreciate that. So we have all heard a few sensational stories from the media. There's an interesting Amazon story where an employee created a voluntary dashboard and everyone starts tracking their own AI token usage. I'm not sure there's explicit encouragement from the leadership, but the fact is, engineers—some of the engineers started competing with each other in maximizing their token usage and getting to the top of the so-called leaderboard. There's a similar story from Meta and then another even more sensational story about some companies spending $500 million on cloud. Oops. Within a month. So while this may not be happening in your companies today, the threats and risks are real. How do we think about the policies? How do we measure the cost? And how do we control and optimize for it? So one initial learning I want to share is it is really important to have a dashboard that tracks every team, every individual's token usage and cost. But that should not be positioned as a leaderboard. We think of the usage dashboard more as a smoke detector. If there are local pockets of teams or individuals that don't use much AI token, that might be a signal worth investigating. But beyond that, certainly we don't want to create indirect incentive to maximize the token usage itself. So how do we think about it then? First, I want to make sure that we position this talk for those of you whose teams have already gone through the hump of getting AI adopted. If you're still in the initial process of provisioning easy access to your engineers or encouraging the teams and individuals to adopt, then you may not be ready to implement some of the ideas for controlling and optimizing for cost. But that's okay. This could still be a good discussion. And frankly, we just got over that hump over the last couple of quarters. So this is a very topical subject that every engineering leader, I believe, is navigating. So I would love to start that dialogue with you all today to explore the best practices. Can I get a quick show of hands for those of you whose teams have gone over the initial adoption phase? Now you are starting to seriously worry about the cost. Okay. I see roughly half of the hands raised. Thank you. So let's talk about how we can control and optimize what we call trusted throughput as a proxy metric as a way to measure your ROI. But before that, just for those of you who are in the process of increasing adoption, one lesson we learned is after the top-down leadership push, sit down with the individual teams and the individuals who may be resistant or struggling with adoption, understand where they came from. For example, there are some legitimate concerns that I heard people say: "Hey, I used to really take pride and joy in hand crafting code. And now a lot of the joy and the pride got taken away and replaced with me reviewing AI sloth code." So that doesn't sound like a very satisfying professional activity. And that's where we need to dig down and understand what are still the high impact and engineering tasks, technical work that we can help our engineers continue to grow themselves in the era of AI. So I wanted to share with you a bit more about what we at Ironclad do. And there's an interesting connection actually with how we think about optimizing for engineering AI token usage. So Ironclad is a legal contracting AI company. We build AI features and native AI products to help lawyers, procurement, and other business users move forward new contracts, move them forward faster with controlled risk. What that means is building trust is the number one priority with our AI product features and products. And for the prior speaker, she did a wonderful job telling you about the importance of trust and how to build it in their domain. In our Ironclad product domain, it often means lawyers, especially, but other personas as well, taking the time to test the water and see if they can trust the AI output. For example, they may feed our conversational search a set of contracts they are familiar with and they run the search and see if the output is towards the expectation. If so, they may expand on searching for things they don't know about or apply other workflows using AI to solve other things like redlining the contract and finding anomalies and so on. And so similarly, using AI and making sure AI is delivering high engineering value also involves a sequence of steps in gaining trust from the internal engineers, the leadership, as well as with our customers. So this is the focus of our talk today. And this probably will not come as a surprise. Here, the goal is not to minimize or not even actually to reduce token spend. So we use the word, it's not about austerity. It's about further improving the ROI of the token spend. So how do we do that? Here, we propose a concept we call trusted throughput. So the trusted throughput comes from having the code reviewed and validated internally and ultimately validated in customer deployments. So how do we think about controlling the cost and measuring and in turn optimizing the ROI? The first step is I'm pretty confident that all of you, your teams who have been adopting AI have been measuring the cost. So if you're using a single tool like Cloud Code or Codecs, then you tend to get very rich analytics from the vendor's dashboard already. If you're us who use a combination of these different coding tools, then we basically use AI to build simple dashboards and pipelines to extract such vendor data. So we can cross correlate them. Then we can break it down, aggregate and then break down by per team, per individual, what is their cost usage across all of these tools. So that's the first step for measuring cost. Now, one pitfall I have seen and we wanted to caution everybody is to then jump from measuring cost to start reducing or minimizing the cost, cutting cost. We think that is premature. Instead, the other important side of the equation for ROI is to measure value. How much value are we getting from burning the tokens? Once we can measure the cost and value side, we understand ROI and then to improve ROI, we want to find and then fix the bottlenecks. In the next couple of slides, I'm going to introduce two new bottlenecks we identified in this whole new software development lifecycle where code generation now becomes abundant, thanks to AI, but the pressure is now getting pushed down to code review and continuous integration, CI/CD, merging the code. So we'll talk about that. And finally, we'll put together these ideas into a pragmatic framework of how we think about optimizing the ROI and thus the leverage in using AI. Okay, so this is a slide on building or using the vendor dashboard to measure the cost. And again, we want to caution that here, the main goal for regularly reviewing the dashboard is to see A, if there's still adoption gap within individual pockets of teams or the individual engineers, and B, if there are any sudden surprises in the usage burst. And if so, understand what's been happening. If they're legitimate. And then also compare teams contextually. So this is important. We don't control just the AI usage per se, because, for example, a platform infrastructure team, the way they use AI and the way they get value may be different from the UI team. So we need to take the context into consideration. All of such review analysis is to help us extract learnings. So there's a self-learning loop that we can then feed back into institutional best practices. What we don't want to use dashboards for is to stack rank people, right? Making it a leaderboard and somehow reward maximization. There's an interesting analogy I want to draw with a traditional engineering productivity metric called lines of code. So I believe all of you will be tracking that metric, but it wouldn't be wise to use that metric as the key goal to measure engineering velocity. Because if we want productive and high quality engineering work, one can argue that removing code is even better. So LOC, line of code, is an important metric, but not something we want to directly optimize for. Same thing for the token usage and spend. So that gets us to the notion of trusted throughput. How do we think about that? How do we define that? First, I want to share the quantified side of things. What are the metrics that we have been involving in defining and tracking? So we talked about line of code is clearly not a good way to measure if AI is generating a lot of value. So the next evolution can be let's count the number of open PRs, pull requests. The intuition being engineers are using AI to generate a lot more code. So let's measure the open PR. So clearly we see a big inflection in the open PR count. So, LOC, line of code is an important metric, but not something we want to directly optimize for. Same thing for the token usage and spend. So, that gets us to the notion of trusted throughput. How do we think about that? How do we define that? First, I want to share the quantified side of the things. What are the metrics that we have been involving in defining and tracking? So, we talked about line of code is clearly not a good way to measure if AI is generating a lot of value. So, the next evolution can be let's count the number of open PRs, pull requests. The intuition being engineers are using AI to generate a lot more code. So, let's measure the open PR. So, clearly we see a big inflection in the open PR count. So, we see a lot of people are using AI to measure the number of the nodes. But eventually, as I assume everyone would agree, over the time, even though people may do one-off RND work to try out things without landing them, but eventually we are all measured by the code we ship. So, therefore, we evolved from tracking the open PR count to tracking the merged PR count. So, that's an improvement. But the next question is, not every merged PR is equal. There can be a PR with only ten lines of code that takes forever that finds and fixes a concurrency bug. Or there can be a thousand line boilerplate code that just takes a lot of time to generate a review. But otherwise, it's not necessarily adding as much business value. So, we then started tagging each merged PR with some sort of complexity score. There's no traditional definition of what that means. We looked at the literature a bit. So, we just took a pragmatic approach of giving AI a well-crafted prompt. And then we feed the PR into basically one or two LLM and say score the complexity based on T-shirt size. So, the idea being if you use AI to generate a more complex PR, we consider that as being more valuable. So, that's how we add a weightage to each merged PR. But that's not the end of the journey. That's just something we're going to evolve, and I would love to discuss with everyone on how we end up creating, defining a set of metrics that approximate the value AI is generating. Now, let's look at the qualitative view. What we think about the way we would define trusted throughput is a high quality output that's entrusted by both internal engineering and leadership and external customers. We think they come from three buckets. The first bucket is all of the objective metrics that we run with checking the test coverage, whether all of the predefined security checks are passing, do we go through the regular canary in practice as we roll our features safely, and so on. In addition, we complement the objective metrics with our subjective human judgment. So, that's where the code review, the design review come in, to look at the code quality, clarity, maintenance. Architecture fit, and so on. And then finally, we want to make sure through all of these internal objective and subjective check, when the rubber meets the road, how customers perceive the changes. Are there production fires that lead to rollbacks? Do customers complain, have tickets that talk about usability, friction, bugs, and so on? So, these are the three buckets that together form what we think is trusted throughput from engineering. Okay, so now let's talk about from a software deployment lifecycle perspective, where we observe the new bottlenecks are. As I mentioned earlier, AI cogeneration is making PR creation abundant. So, now the bottleneck from the whole lifecycle perspective gets shifted onto review, and they are subsequently merging the PR. Does that resonate? Yeah, I see some heads nodding. So, this is where we spend time on figuring out how we can further improve the review process as well as the continuous integration, the CI process. So, we will dive into these two topics in the next couple slides. Here, I just want to say a potential anti-pattern, anti-solution, is that if the CI infrastructure gets overloaded, then a workaround by engineers to stop splitting PR, to stop submitting large PR for review and submission. Because if it takes an hour to run all of your regression tests and submit it, I don't want to break my PR into 10, right? Which might take 10 hours. However, this, in our view, can be pretty risky, because it makes the human review overhead higher and also reduce the quality of the review because the human attention can be spreading. So, that is an anti-pattern I wanted to caution. So, for code review, the key principle we use is to make sure we onboard AI tooling as the first level of defense. They don't replace human reviewers, but we want to offload human reviewers as much as possible. Let the AI review take care of simpler things like coding style issues or if there's missing test coverage. So, make sure the author gets through all of them before then the review gets routed to a human reviewer. And this way, our human engineers can focus on applying their deep judgment on aspects that are somewhat subjective, like if the code is good, if the architecture is sound, if the code passes the security design, and so on. So, that in the end, our engineering team can take the final accountability. Now, let's look at CI. So, I assume all of you deploy some form of CI/CD. And what we're seeing is, thanks to AI, now making it much easier to generate code, splitting code into smaller but more PRs, it puts a lot more pressure on the CI. And this is something that if we don't address at a company level, individual engineers can be struggling. Because that means they have to waste their human time babysitting the PR to get merged. If they run into flaky tests, then they have to manually hit rerun. It's very frustrating. Or they can recruit an AI agent to babysit and do a loop, but that in turn uses AI token as well. So, these are just workarounds, not perfect solution, and also tend to make engineers feel lower morale, a little bit more frustrated. So, what we are doing is we put more developer experience, platform engineering, investing to reducing, removing the flaky test, improving the CI infrastructure. And the key thing here is to also define and measure the right metrics. For example, the wall clock time between when the PR is ready to submit to when it's submitted. If a typical CI run takes an hour, does the typical PR submission take two or three hours? In which case, that's a red flag. And also, the number of times a PR needs to get retried for passing through the test. So, these are the key metrics that we are using to measure our developer experiences and the relevant team who is focused on improving the developer experience. So, with all of the analysis and ideas, here we want to share a pragmatic framework of how we can then measure and optimize token usage. It has three aspects. The first one is to set the right set of guardrails across setting the budget and quota, tracking usage, defining anomalies so that users, leaders can get notified if something feels wrong. This is complimentary to regular human review, which can catch other interesting patterns or learnings and feedback into the institutional knowledge base. Let me just couple that with the third item here, which is the learning loop we talk about. As our leadership work with individuals to define these guardrails, review the metrics, and then refine, that's how we close the learning loop. In addition to that, we want to work with our teams, individual engineers to continue to search for and, if needed, innovate on the best practices of how to use AI. How to use AI to build products and also use it internally. For example, some engineers may be writing an agentic loop as part of the hard nest when they use cloud code. After they generate initial PR, they go and loop around and say, try and pass the set of tests. And then, if some tests don't pass, just auto-fix the test or the code and retry. One thing to watch out for is to put a limit on the number of loop steps to make sure if things go out of control, we don't waste too many tokens on that. Another example is prompt caching. This is becoming increasingly more prevalent by the commercial model vendors, where what they advise is if you send a prompt with the same prefix, they could optimize how they process the prefix of the prompt. What that means then, as a user to those LLM, is that we want to encourage our users to structure their prompt that way. For example, if your prompt consists of a system prompt followed by a user prompt, you want to put a system prompt that's fixed at the top and the varying content at the bottom. Context pruning is also important. We want to drill it into each individual user's new muscle memory so they are aware that as they build out the context through a longer chat session, they would be mindful of summarizing the context and make sure that the token usage is efficient that way. There are increasingly more tools like Cloud Code that will automatically manage and compact the context for you. And so this increases the token usage efficiency but also increases the quality of AI output. What that means then, as a user to those LLM, is that we want to encourage our users to structure their prompt that way. For example, if your prompt consists of a system prompt followed by a user prompt, you want to put a system prompt that's fixed at the top and the varying content at the bottom. Context pruning is also important. We want to drill it into each individual user's new muscle memory so they are aware that as they build out the context through a longer chat session, they would be mindful of summarizing the context and make sure that the token usage is efficient that way. There are increasingly more tools like Cloud Code that will automatically manage and compact the context for you. And so this increases the token usage efficiency but also increases the quality of AI output. There are other ideas we're exploring as well. So I know we're at time, so this is towards the end of the talk. There is sometimes we also face build versus buy decision. The principle is simple. For things that are non-differentiating like IDE, CI infrastructure, we want to buy. But then for things that are specific to our context like how we would generate high quality PR for small bug fixes versus building a new UI feature for refactoring and so on, we have our internal playbook, which is a set of well-crafted AI prompts. So we save that and share across our team so that gets reused and enhanced. So that's something we must build internally. When it comes to case-to-case though, sometimes it's still a bit ambiguous. We're trying to build what we call a builder agent that's a cloud-based code generation that wraps the cloud code codecs and so on. While we know there are also other vendors out there that we're still exploring. So we'd love to exchange thoughts on that. So then to summarize, here are a couple key lessons as we went through the last couple quarters of the journey. I wanted to share so that hopefully you could accelerate your process there. If I were to summarize these three things, it's about learning, planning ahead, and learning from other people's stories and mistakes. So what that means is think about build versus buy early on. As you are encouraging more code gen, think about how that impacts your code review and CI and how you can address these new bottlenecks.