Open Reader

How to build an AI-Native Health Company — Dan Feng, Maven Clinic

completed 17:18 Aug 19, 2026 Watch on YouTube

Current Status

completed

Video ID

WJRdLNhrsLQ

RAG / Chat

Enabled
How to build an AI-Native Health Company — Dan Feng, Maven Clinic
Description

Implementation used to be the expensive step, so teams spent weeks settling requirements before anyone wrote code. Dan Feng's observation is that the cost moved. Building takes minutes now, and arguing is what is expensive. Planning at Maven Clinic changed to match. A one year view survives only as direction, assuming models will handle whatever you need by then, while real commitment runs two to four weeks. Long requirement documents gave way to a page or two meant to be argued with. The awkward casualty is the three to six month plan, which he treats as close to unplannable when nobody knows what models will do by then. The rest is what breaks at that speed. Engineers who once wrote hundreds of lines a day now write thousands, so review had to change rather than scale. Engineers self certify which pull requests need a second reader and stay accountable either way, requests are capped near 500 lines, and large features are stacked into several. The failure he names is the rubber stamp, which buys false confidence rather than none. On reliability he refuses a single bar and sorts failures into tolerable and not. A scheduling action that fails one time in 10,000 is survivable, since the user clicks again. A reimbursement claim is not, because asking for $50 and receiving $200 is an escalation in either direction, so several models read the same receipt and it proceeds only if they agree. Integration tests run many times rather than once, since passing a nondeterministic system on one attempt proves very little. Speaker info: - https://www.linkedin.com/in/dan-feng-2bb5703/ - https://www.mavenclinic.com/ Timestamps: 0:00 - Who here is already AI native 0:51 - Maven Clinic, and starting the journey two years ago 1:29 - Tractors do not replace farmers 2:08 - Adopting internally, then building it into the product 3:25 - Early adopters, the majority, and the reluctant 4:04 - Meeting engineers on whichever tool they moved to 4:43 - Why senior engineers stopped delegating

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Becoming AI-native is an operating-model change—not merely adding features—requiring shared AI infrastructure, AI-fluent independent operators, short delivery cycles, and risk-tiered reliability controls.
  • Why it matters: Maven provides a concrete production-healthcare example of how to move from AI experimentation to organization-wide adoption while preserving controls for high-consequence workflows.
  • Best use: Use it as a practical benchmark for designing an AI operating model, especially engineering workflow, adoption strategy, evaluation, and human-escalation patterns.

Executive Summary

Dan Feng describes Maven Clinic's two-year transition from a conventional digital-health technology company to what it calls an AI-native company. Maven built “Maven Intelligence,” an orchestration layer intended to make AI capabilities available across products, internal teams, and clients. His central argument is that AI adoption is competitively unavoidable, but there is no universal playbook: companies must change tools, talent, incentives, planning, and controls together.

Maven's organizational model focuses on the broad middle of adopters rather than only enthusiasts. It supports the tools employees actually prefer—citing a shift from Cursor to Claude Code—while using shared infrastructure, feedback loops, and a clear company direction to bring along slower adopters. The company expects AI to reduce delegation overhead: senior engineers increasingly solve implementation directly with AI, while new hires must operate more independently, understand product context, and handle ambiguous systems problems.

The operational consequence is a shift away from extensive up-front requirements and medium-range planning. Maven uses one-year ambitions only as directional inspiration, prioritizes two-to-four-week delivery increments, and favors one- or two-page PRDs/TDDs as lightweight communication artifacts. Since AI makes implementation cheap and fast, Feng argues that alignment, judgment, and rapid iteration—not coding throughput—become the primary constraints.

For production AI, especially in healthcare-adjacent financial workflows, Maven does not aim to eliminate hallucinations universally. Instead, it classifies failures by consequence, tolerates recoverable errors in low-risk tasks such as appointment scheduling, and uses redundant model review plus human escalation for reimbursement claims. Its release process combines repeated-run integration testing, pass-rate thresholds, automated conversation evaluation against rubrics, manual review, and heightened review coverage for newly launched features.

Key Takeaways

  • Claim: AI-native transformation requires simultaneous internal automation, customer-facing AI products, and changes to culture and operating processes. | Evidence: Maven created Maven Intelligence as an orchestration layer across products; internally it applies AI to tasks such as summaries, meeting management, and Jira creation, while externally it uses AI to improve user experience and lower operational cost, including always-available chatbot support. | Implication: Ken should treat AI-native status as a control-plane and organizational-design problem, not as a collection of isolated copilots or chatbot features. | Caveat: Feng explicitly says there is no single definition of AI-native and no predefined playbook that can be copied wholesale.
  • Claim: The highest-leverage adoption effort is enabling the mainstream middle, while accommodating tool preferences rather than standardizing prematurely. | Evidence: Feng segments employees into early adopters, a large middle group, and slow adopters. Maven gives early adopters tools and asks them to share learning; for the middle it builds shared infrastructure and easy-to-use tools. It supported both Cursor, widely used the prior year, and Claude Code, adopted by many employees this year. | Implication: A central platform should abstract and govern capabilities while allowing users and teams to adopt the interfaces that maximize their actual usage. | Caveat: Meeting users where they are does not mean tolerating organizational ambiguity; Maven pairs flexibility with a clear statement that the company is headed toward AI adoption.
  • Claim: AI reduces the economic value of delegating well-specified implementation work and raises the value of independent, product-literate technical operators. | Evidence: Maven found that senior engineers who have already determined a solution often use AI to implement it directly because delegating adds overhead. It now seeks hires who are genuinely interested in AI, understand the product and system deeply, and can handle complicated, ambiguous problems; performance reviews also ask what employees have done on the AI side. | Implication: Ken should hire and reward for end-to-end problem ownership, AI learning velocity, and product judgment—not simply capacity to receive and execute delegated tickets. | Caveat: The argument applies most directly once a problem is sufficiently understood for AI-assisted execution; Feng identifies ambiguity and deep systems understanding as areas where AI is weaker.
  • Claim: When AI makes implementation fast, planning should move from exhaustive specifications and quarterly roadmaps toward short, reversible delivery cycles. | Evidence: Maven uses one-year visions only for direction, focuses execution on the next two to four weeks, and prefers one- or two-page PRDs/TDDs over lengthy documents. PMs and designers specify the immediate sprint while engineers ship by the end of it; if the prior decision proves wrong after two weeks, the team changes direction. | Implication: Ken should distinguish durable strategic intent from short-horizon commitments, and design workflows where discovery, building, evaluation, and revision occur continuously. | Caveat: Feng calls three- to six-month planning especially awkward because model capabilities may materially change within that period.
  • Claim: AI coding should be rolled out from low-risk, easily verifiable tasks to general implementation, while shifting engineers toward architecture, review, and evaluation. | Evidence: Maven began with unit tests and documentation, using these tasks to build confidence, rules, skills, and guardrails. It then expanded AI coding tools across implementation; Feng says engineers now focus primarily on reviewing, architecture, and evaluation. | Implication: Ken should sequence agentic coding adoption by verifiability and blast radius, using early use cases to develop organizational guardrails before expanding autonomy. | Caveat: Maven's experience is not that AI code review can replace human judgment: it has tried multiple review tools but does not yet trust them completely.
  • Claim: AI-amplified code output requires redesigning review policy rather than trying to preserve conventional review practices. | Evidence: Feng contrasts hundreds of lines of code per day historically with thousands produced using AI. Maven permits accountable self-merge for simple, high-confidence PRs; when review is needed, it caps PRs at 500 lines, uses stacked PRs for larger features, and explicitly seeks to prevent blind “rubber-stamp” approvals. | Implication: Ken should make review requirements risk- and change-size-sensitive, preserve meaningful reviewer intervention, and avoid process rituals that create false assurance. | Caveat: Self-identified no-review merges transfer accountability to the engineer and therefore require strong ownership norms and monitoring.
  • Claim: Reliable GenAI systems should use consequence-based failure budgets, redundant validation, repeated evaluation, and human fallback rather than a blanket demand for zero hallucinations. | Evidence: Maven considers an appointment-scheduling error roughly once per 1,000 attempts recoverable because a user can retry, but treats reimbursement claims as intolerant of failure. For receipts, multiple models review the same input and Maven proceeds only when they agree; uncertain cases route to human agents. The company runs hundreds of integration tests repeatedly rather than once, targets high pass rates such as 90% over time, auto-evaluates conversations with predefined rubrics, manually spot-checks traffic, and reviews about 20% of conversations for new features. | Implication: Ken should define risk tiers before building an AI workflow, then assign each tier appropriate redundancy, evaluation coverage, approval gates, escalation, and post-launch surveillance. | Caveat: A 90% pass rate is presented as an example rather than a universal deployment threshold; acceptable reliability must be tied to the harm and reversibility of each workflow.

Detailed Brief

Maven's intended end-state for the software delivery lifecycle

  • Claims: Maven already uses AI assistance across nearly every software-development stage, but does not consider the transition complete.; Its desired end-state is an AI system that spans design, implementation, release, production monitoring, early issue detection, and automatic remediation.
  • Evidence: Feng says end-to-end automation through release is the goal.; He specifically identifies monitoring live traffic and automatically fixing detected issues as future capabilities Maven is still working toward.
  • Caveats: Maven is not yet at autonomous production remediation, and current AI code-review tools provide only partial assistance rather than sufficient standalone assurance.
  • Implications: The meaningful maturity model is not “developers use copilots”; it is progressive automation of the full build-to-operate loop with controls appropriate to production risk.; Autonomous remediation should be treated as a later-stage capability contingent on strong observability, evaluation, rollback, and accountability mechanisms.

Evaluation as a continuously calibrated operational system

  • Claims: Evaluation does not end at pre-release testing; Maven evaluates live conversations after launch and uses manual review to tune the evaluator itself.; Human review serves two purposes: finding model or workflow failures and checking whether the automated rubrics are too strict or too loose.
  • Evidence: Maven has a dedicated group responsible for manually reviewing conversations.; For new features, it expands from spot-checking to reviewing approximately 20% of conversations.
  • Caveats: The transcript does not specify the staffing, sampling methodology, evaluation model design, or clinical-governance procedures behind this review process.
  • Implications: Production evals need a feedback loop in which human judgment calibrates automated scoring, rather than treating rubric scores as immutable ground truth.

Notable Concepts & Terms

  • Maven Intelligence: Maven Clinic's orchestration layer for making AI available across its products, internal teams, and clients; it represents the platform layer beneath individual AI applications.
  • AI-native company: In Feng's framing, a company that embeds AI in internal work, product delivery, and organizational processes rather than merely deploying isolated AI features.
  • Farmers and tractors analogy: AI will not simply replace workers; workers and companies able to operate AI effectively will outperform those that do not adopt it.
  • Stacked PRs: A code-review practice in which a large feature is divided into smaller dependent pull requests so review can remain meaningful while implementation continues.
  • Rubber-stamp review: A nominal approval process where reviewers cannot meaningfully change or challenge the submitted code; Feng treats it as dangerous false confidence.
  • Risk-tiered reliability: Setting reliability controls according to the severity and recoverability of an error, rather than applying the same hallucination standard to every AI task.
  • Multi-model agreement: Using different models to independently review the same high-stakes input and proceeding only when their outputs agree, with human escalation for unresolved conflict.
  • Auto-eval: Automated post-launch assessment of AI conversations against predefined good/bad rubrics, supplemented by manual quality review.

Operator Notes / Why Ken Should Care

  • Create a risk register for every agentic workflow that explicitly identifies acceptable failures, intolerable failures, reversibility, escalation owner, and required validation method.
  • Adopt a staged coding-agent rollout: begin with verifiable low-blast-radius work, document the resulting rules and guardrails, then investigate every repeated case where engineers decline AI assistance.
  • Replace universal code-review requirements with a policy that distinguishes self-mergeable low-risk changes from reviewed changes, caps reviewable PR size, and rejects rubber-stamp approvals.
  • Separate annual directional narratives from near-term execution commitments; run delivery in two-to-four-week loops with deliberately lightweight requirement artifacts.
  • Instrument production agents with repeated-run tests, automated rubric-based evaluation, and a human calibration program before increasing autonomy or expanding to high-consequence actions.
  • Review hiring and performance systems for whether they reward AI-leveraged end-to-end ownership, product understanding, and judgment under ambiguity rather than delegated implementation throughput.

Source/Metadata

  • Title: How to build an AI-Native Health Company — Dan Feng, Maven Clinic
  • Transcript words: 4619
  • Duration seconds: 1038
  • Timestamp note: No timestamps or chapters were provided. The latter portion of the supplied transcript substantially repeats earlier material.

Transcript

2660 words en Processed in 100.9s

It's time we can get it started. I'm Dan. I'm from Maven Clinic. Today we'll share the experience of how we transitioned from a traditional technology company to an AI-native company. Before I started, I would like to do a little exercise. Raise your hand if you think you are already an AI-native company. Okay, we saw a few. Raise your hand if you thought about it but haven't started the journey yet. Okay, we saw a few. That means most of us are in between. Hopefully this talk can help you with that one. Maven Clinic is a large digital health platform. We are focused on women and their families. We specialize in maternity, fertility, parenting, and menopause. We started our AI journey just two years back. At this moment, we built something called Maven Intelligence. It's an orchestration layer across all our products to enable AI for everybody in this company and for our clients. AI is here and improving every day. I think adopting it is not optional. Even if you choose not to, your competitors will. This is a quote I heard a couple years back. I would like to share it here again. Tractors won't replace farmers, but the farmers who can operate the tractors will replace the ones who cannot. Hopefully everybody here will become farmers who can operate your tractors. That's the goal. First of all, I don't think there's one single definition of what it means to be AI native. More importantly, there's no predefined playbook you can just follow and, bingo, you become AI native. For us, it really comes down to three parts. One is internal use of AI tools whenever it's possible. It can be as simple as generating your daily summary, managing your meeting, or creating a Jira task. Anything you need to do manually today, you should think about and ask, can I use AI to do it? Whenever you want to ask other people to do something for you, you should ask, can I use AI to do it? A lot of leaders today at Maven, including our CIOs, use AI tools to solve those tasks by themselves now instead of delegating them to other people. Externally, we want to build AI into our product. We achieve two goals there. One is really focused on improving our user experience. Second, maybe help us reduce our operational costs. An AI-based chatbot is a really good example. It's 24/7, always available, can help address issues, and help our customers instantly. It's way better and cheaper compared to human agents. Certainly, and I think more importantly, we need to think about culture, process, and the way we work, and how we can change it so we can maximize what AI offers for us. I will touch on it more in the following slides. When we come to adopting new technologies, there are always three groups of users. One, there are some early adopters. For them, we don't need to do too much. The only thing we need to do is enable the tools for them and encourage them to share what they learn with the company. What we need to really focus on is the one in the middle. That's the majority. We should build a shared infrastructure for them, build easy-to-use tools for them, and make the adoption as seamless as possible. More importantly, we should really listen to them, get feedback, and consistently improve. For example, last year, most folks at Maven were using Cursor. This year, a lot of them switched to Claude Code. For us, we need to support both. We need to meet them where they are and make them feel comfortable using it. And for all places, you always have a few slow adopters. They always have concerns and worries about new technologies. For them, we should meet them where they are and understand what their concern is. But more importantly, we should be crystal clear with them about where the company is heading. AI is really good at execution if we know what we want to do. This will change how we should hire new people and how we should reward them. The way we used to work was to have a senior engineer sense the problem, come up with a solution, and delegate to other engineers for implementation so we could work on it in parallel and be faster. But these days, we found that senior engineers would rather not delegate implementation work to other people. Because they already figured out how to solve the problem, they just use AI to solve it instantly. Delegating to other people means more overheads and is less efficient. That also means when you have new people, you want to make sure they can solve the problem independently. They pretty much have to work as a traditional tech lead. We cannot afford other people to delegate implementation tasks for them. Also, when we hire new people, we should think about what we are looking for. We definitely want to look for somebody genuinely interested in AI. The domain is moving so fast. We want them to keep learning and also help the team stay on track. Secondly, with AI, engineers can do way more than they used to do. The boundaries between PMs and engineers are getting blurry. We found engineers who really understand the product can, in fact, make way more contribution than the traditional engineer who only focuses on software. This is what we are looking for. That deep understanding of the system and the ability to handle complicated, ambiguous problems are also very valuable. This is where AI drops off. When we hire new people, this is also the kind of people we are interested in bringing on board. For people we bring in, we want to reward them in the proper way. Even in our performance review, we start to ask, okay, what have you done on the AI side? We definitely want to reward people who leverage AI to multiply their impact. This is the impact for everybody in the company. Now we have the right tools, we get the right talent in place, and we need to change how we work to maximize the benefit of AI. The way we used to work was to say, okay, we spend weeks, sometimes months, to flesh out the business requirements, finalize the design, and then do the implementation. Because implementation can be really expensive. If we didn't get the other part right in the beginning, it can be very costly to change it later. But in fact, we never get things right in the beginning anyway, for any big projects. With AI, building is super fast. It's probably a couple minutes. You can get it done. Alignment is the really expensive one. We should really think about what's the best way we can work. How can we deliver fast? It's still okay to think about what you want to deliver in one year. You can assume AI models can do anything you want in one year. Based on that, really dream big to think about what you can do in one year. But it should only serve as inspiration and direction. What we really need to focus on is what we want to deliver in the next two to four weeks. We want PMs and designers to say, okay, tell me what I need to do in this sprint. Then the engineers will focus on it and get it released at the end of the sprint, if not sooner. Meanwhile, the PMs have time to flesh out the next batch of requirements. If at the end of the sprint they say, no, what we decided two weeks ago is wrong, it's totally okay. We can switch gears and get it fixed quickly. That also means we prefer people not to write pages and pages of PRDs or TDDs anymore. We prefer them to write just a short one or two pages. That really serves as communication so we can iterate on it. The really awkward part is the midterm goals. Those are three months, six months. It's very hard to plan these days. The reason is, I don't know what AI models will be capable of in three months. There may be multiple releases already. So we prefer not to focus on this one. This may make it very hard for most of the folks who have been in this domain for a long time, because traditionally, we get used to having quarterly planning or planning for six months. But it's our job to get used to the new AI era and learn how to work in it efficiently. I want to talk about coding and software development a little bit more here. AI coding tools are probably the most successful AI application, and they're really good at implementation. You probably heard a lot of people say, okay, I have these AI tools. Now I can even use my phone to implement software and automate every stage. If they feel comfortable doing that, it's totally okay. But you don't have to. What I'm trying to say here is, and maybe this was our journey, how we adopted those AI tools. We started with the lowest-risk tasks, like writing unit tests and documentation. Those things are very easy to verify, and the risk is super known. By doing that, we built confidence. We started to construct our own rules and skills and build our guardrails. Then we pushed to the whole engineering team and said, now you should use these AI coding tools for all the tasks. When they choose not to, that's the time we really want to learn and ask, why don't you do it? At this moment, we pretty much use AI coding tools to do all our implementation. Engineers really focus on reviewing, architecture, and evaluation. With AI coding tools, we are writing so much code these days. Code review becomes really challenging. A good engineer used to probably write hundreds of lines of code every day. These days, they can easily write thousands. If we do code review as we used to do, we won't be able to keep up. We also tried multiple AI code review tools. It helps a little bit, but we don't feel comfortable 100% relying on them yet. We still find the feedback from our engineers very, very valuable. That means we need to really change the way we are doing code review to meet where we are now. A couple things we have done: one is we allow engineers to self-identify whether they still need code review. If they think this PR is simple enough and they feel very confident, they don't need anybody to take a look. We are fine with that. We let them merge, but we still hold them accountable. If they do want code review, we want them to stay with the best practices. For example, each PR shouldn't have more than 500 lines of code because nobody can do a meaningful code review on one that has thousands of lines of code. We also enabled stacked PRs. What it means is that for a big feature, engineers can break it into multiple PRs. While people review the PRs, they can keep working on it. One thing we really want to avoid is a rubber stamp, we call it. It means people submit code for review, but you cannot really do anything to it. You just blindly approve it. This is the worst case we should really avoid because that just gives us false confidence. We think we reviewed it, it's good, and we release it. Meanwhile, we should keep working on our AI code review tools because we think that's the future. At this moment, we use AI tools to pretty much assist in each step of our software development. Our goal is for it to automate the whole life cycle from end to end, from design and implementation until it's fully released. More importantly, we want the AI tools to be able to monitor the live traffic, catch issues early, and automatically fix them. That's what we are still working on, and we are not there yet. The last thing I want to touch on a little bit for this presentation is reliability. What it means is that for traditional software, it does what we implement there. No more, no less. But for GenAI solutions, hallucination is there. We cannot ignore it. Completely eliminating it can be very costly. Sometimes it's not necessary either. So the way we should do it is really have a holistic solution, even from the beginning. For example, we can start with identifying which failures are acceptable and which ones are not acceptable. For our AI system, for example, we have the functionality to help our customers schedule appointments. If we fail one out of one thousand, probably it's okay. I'm not saying it's a good experience, but users can usually just click the button again and we will reschedule for them. Probably it's okay. But if we help the user submit their reimbursement claim, we cannot tolerate failure. Because if people ask for $200 and we issue them $50, or we give them $200 incorrectly, each case will cause an escalation right away. For those cases, we have to put in extra steps. For example, when we receive their receipt, we will use different models to review the same receipt. We only move forward if the results from different models agree with each other. If we really have trouble figuring out which one is right, it's okay to tell the customers, hey, how about processing your stuff? Do you want us to connect you to a human agent? We will move from there. That's our set of solutions. Also, we should have a rigorous process to release our software. For us, we have hundreds of integration tests, which pretty much cover all the use cases we know, and we keep adding to the integration test suite. When we run the integration tests, passing once is not good enough anymore, because the LLM can do different things. So for each test case, we run it many times. We consistently require high pass rates, for example 90% over time. More importantly, after we launch the software, we have our auto-eval system carefully evaluate each conversation. We have predefined a lot of rubrics for what we think is good and what is bad. Then we review the scores. Besides this, we also have a dedicated group whose job is to manually review those conversations. We will spot-check our conversations. That helps us see whether we need to come back and improve our systems, or whether our rubrics are too strict or too loose, and we need to consistently improve them. When we launch new features, that's the time we say spot-checking is probably not enough. We really want to review 20%, and we can do it. This whole process makes sure we feel really confident we never ship something, although we know hallucination is there. That's pretty much all I have for today, and I can stay here to take questions. If you have other things, you can reach out to me. where they are, just make it feel they're comfortable to use it. And for all the places, you always have a few slow adopters. They always have concerns, worries for the new technologies. For them, we just should meet where they are, understand what their concern is. But more important, we should be crystal clear with them, where the company is heading to. So AI is really good at execution if we know what we want to do. So this will change how we should hire new people and how should we reward it. We used to, the way we used to work is we have a senior engineer who will sense the problem, come up with a solution and dedicate to other engineers for implementation. So we can work on it in parallel and be faster. But these days, we found the NASA engineers would like to dedicate implementation work to other people. Because they already figured out how to solve the problem, they just use AI to solve it instantly. Dedicating to other people means more overheads and we manage efficient. That also means like when you have new people, you want to make sure they can solve the problem independently. They pretty much have to work as a traditional technical level. We cannot afford other people to dedicate implementation tasks for them. Also, when we hire new people, we should think about what we are looking for. We definitely want to look for somebody genuinely interested in AI. The domain is moving so fast. We want them to keep learning. Also, help the team to stay on track. Secondly is with AI, engineers can do way more than they used to do. The boundaries between PM and engineers is getting blurry. We found engineers who really understand the product, in fact, they can have way more contribution than the traditional engineer who only focus on software sites. And this is what we are looking for. And those deep understanding of the system, the ability you can handle complicated, ambiguous problem is also very valuable. This is where AI rank off. When we hire new people, this is also the people we are interested in bringing on boards. For people, we bring in, we want to reward them in the proper way. Even in our performance review, we start to ask, okay, what you have done for AI sites. We definitely want to reward people who leverage AI to multiple their impact. Although this is the impact for everybody in the company. So now we have the right tools, we get the right talent in the place. And we need to change how we work to maximize the benefit of AI. The way we used to work is to say, okay, we spend the weeks, sometimes in the months, to flesh out the business requirements, finalize the design, and then do the implementation. Because implementation can be really expensive. If we didn't get the other part right in the beginning, it can be very costly to change it later. But in fact, we never get the things and the rights in the beginning anyway, for any big projects. With AI, building is super fast. It's probably a couple minutes. You can get it done. Argument is really expensive one. So we should really think about what's the best way we can work. How can we deliver fast? It's still okay. You can think about what you want to deliver in one year. You can assume AI models can do anything you want in one year. Based on that one, really dream big to think what you can do in one year. But it should only serve as inspiring and directional. What we really need to focus on is what we want to deliver in the next two to four weeks. What we want to get the PMs and designers to say, okay, tell me what I need to do in this sprint. And the engineer will focus on it and get it released and end of the sprint if not sooner. Meanwhile, the PMs, they have time to flesh out the next bunch of the requirements. If end of the sprint, they say, no, what we decided two weeks ago is wrong, it's totally okay. We can switch the gear, get it fixed quickly. That also means we prefer people not to write pages or pages of PRD or TDD anymore. We prefer them to write just a short one or two pages. That one really serves as communication so we can iterate on it. The really awkward part is the mid-term goals. Those are like three months, six months. It's very hard to plan these days. The reason is, I don't know what AI models will be capable in three months. There may be multiple releases already. So we prefer not to focus on this one. But this one can maybe make it very easy for most of the folks who has been in this domain for a long time. Because traditionally, we get used to have a quarterly planning or we plan it for six months. But it's our job to get used to the new AI error and learn how to work it efficiently. So I want to talk about the coding and software development a little bit more here. AI coding tools is probably the most successful AI application. And it's really good and implementation. So you probably heard a lot of people say, okay, I have this AI tools. Now I can even use my phone to implement software and automate every stage. If they feel comfortable to do that, it's totally okay. But you don't have to. What I'm trying to say here is, and maybe when this is our journey, how we adopt those AI tools, we started with the lowest risk task, like starting with writing unit has documentation. Those things are very easy to verify. And the risk is super known. By doing that one, we build the confidence. And we start to construct our own rules, skills, and build our guardrails. And then we push to the whole engineer team say, now you should use this AI coding tools for all the tasks. When they choose not to do, it's the time we really want to learn, say, why you don't do it. And at this moment, we pretty much use the AI coding tools to do all our implementation. Engineers really focus on reviewing, architecturing, and evaluation. So, and with AI coding tools, we are writing so much code these days. Code review becomes really challenging. So for good engineer, used to, they probably write hundreds of less code every day. These days, they can easily write like thousands. If we can do the code review as we used to do, we won't be able to keep up. We also try the multiple like AI coding review tools. It helps a little bit, but we don't feel comfortable 100% rely on them yet. We still find the feedbacks from our engineers are very, very valuable. That means we need to really change the way we are doing code review to meet where we are now. And a couple things we have done. One is we allow engineers to self-identify whether they still need code review. If they think this PR is simple enough, I feel very confident, I don't need anybody to take a look. We are fine with that one. We need them merge, but we still hold them accountable. And if they do want code review, we want them to stay with the best practices. For example, each PR shouldn't have more than 500 lines of code because nobody can do a meaningful code review with the ones that has like thousands of less code. And we also enabled like stanked PR. What it means is that for big feature and engineers can bring it into multiple PRs while people review the PRs and they can keep working on it. One thing we really want to avoid is a rubber stamp, we call it. It means like people submit code review. You cannot really do anything to it. You just say, blindly prove it. This is the worst case we should really avoid because that just gives us false confidence. We think we reviewed it, it's good, and we release it. Meanwhile, we should keep working on our AI coding review tools because we are thinking that's the future. So at this moment, we use AI tools pretty much assist in each step of our software development. Our goal is it will be automating the whole life cycle from end to end, from designing, implementation, until it's fully released. More importantly, we want the AI tools to be able to monitor the live traffic and be able to catch the issue early and automatically fix it. That's what we are still working on and we are not there yet. The last thing I want to touch a little bit for this presentation is about reliability. So what it means is like for the traditional software, it does what we implement there. No more, no less. But for the GNI solutions, hallucination is there. We cannot ignore it. And completely eliminating them. It can be very costly. Sometimes it's not necessary either. So the way we should do is really have a holistic solution, even from the beginning. For example, we can start with identifying which failures are acceptable, which ones are not acceptable. For our AI system, for example, we have the functionality to help our customers to schedule appointments. If we fail one out of one thousand, probably it's okay. I'm not saying it's a good experience, but the users usually can just click the button again, we will reschedule for them. Probably it's okay. But if we help the user to submit their reimbursement claim, we cannot tolerate the failure. Because if people ask of $200, we issue them $50, we give them $200. Each case will cause an escalation right away. For those cases, we have to put in extra stamps. For example, when we receive their receipt, we will use different models to review the same receipt. We only move forward if the results from different models agree with each other. If we really have trouble to figure it out, which one is right? It's easy. It's okay to tell the customers. Say, hey, how about to process your stuff? Do you want us to get you connect to a human agent? We will move from there. That's our center solutions. And also, we should have a rigorous process to release our software. For us, we have like hundreds of integration tests, which pretty much covered all the use cases we know, and we are keep adding to the integration test suite. And when we run the integration test, not only pass once is not good enough anymore, because the ARM can do different things. So for each test case, we run it many times. We consistently require the high pass rates, like for example, 90% for all the time. And the more important, and after we launch the software, we have our auto evolve system, carefully evaluated each conversation. We have predefined a lot of rubrics. What we think is good, what is bad. And then, we will general results. We will review the score. Besides this one, we also have a dedicated group. Their job is manually review those conversations. We will spot check our conversations. That helps us to say whether we need to come back to improve our systems, or our rubrics is too strict or too loose, and we need to consistently improve it. When we launch new features, then the time we say not on spot check probably not enough. We really want to review, like say, 20%, and we can do it. This whole process makes sure we feel really confident we never reach something, although we know hallucination is there. That's pretty much all I have for today, and I can stay here to take up questions, and if you have other things, you can reach out to me.