It's time we can get it started. I'm Dan. I'm from Maven Clinic. Today we'll share the experience of how we transitioned from a traditional technology company to an AI-native company. Before I started, I would like to do a little exercise. Raise your hand if you think you are already an AI-native company. Okay, we saw a few. Raise your hand if you thought about it but haven't started the journey yet. Okay, we saw a few. That means most of us are in between. Hopefully this talk can help you with that one.
Maven Clinic is a large digital health platform. We are focused on women and their families. We specialize in maternity, fertility, parenting, and menopause. We started our AI journey just two years back. At this moment, we built something called Maven Intelligence. It's an orchestration layer across all our products to enable AI for everybody in this company and for our clients.
AI is here and improving every day. I think adopting it is not optional. Even if you choose not to, your competitors will. This is a quote I heard a couple years back. I would like to share it here again. Tractors won't replace farmers, but the farmers who can operate the tractors will replace the ones who cannot. Hopefully everybody here will become farmers who can operate your tractors. That's the goal. First of all, I don't think there's one single definition of what it means to be AI native. More importantly, there's no predefined playbook you can just follow and, bingo, you become AI native. For us, it really comes down to three parts.
One is internal use of AI tools whenever it's possible. It can be as simple as generating your daily summary, managing your meeting, or creating a Jira task. Anything you need to do manually today, you should think about and ask, can I use AI to do it? Whenever you want to ask other people to do something for you, you should ask, can I use AI to do it? A lot of leaders today at Maven, including our CIOs, use AI tools to solve those tasks by themselves now instead of delegating them to other people.
Externally, we want to build AI into our product. We achieve two goals there. One is really focused on improving our user experience. Second, maybe help us reduce our operational costs. An AI-based chatbot is a really good example. It's 24/7, always available, can help address issues, and help our customers instantly. It's way better and cheaper compared to human agents. Certainly, and I think more importantly, we need to think about culture, process, and the way we work, and how we can change it so we can maximize what AI offers for us. I will touch on it more in the following slides.
When we come to adopting new technologies, there are always three groups of users. One, there are some early adopters. For them, we don't need to do too much. The only thing we need to do is enable the tools for them and encourage them to share what they learn with the company.
What we need to really focus on is the one in the middle. That's the majority. We should build a shared infrastructure for them, build easy-to-use tools for them, and make the adoption as seamless as possible. More importantly, we should really listen to them, get feedback, and consistently improve. For example, last year, most folks at Maven were using Cursor. This year, a lot of them switched to Claude Code. For us, we need to support both. We need to meet them where they are and make them feel comfortable using it.
And for all places, you always have a few slow adopters. They always have concerns and worries about new technologies. For them, we should meet them where they are and understand what their concern is. But more importantly, we should be crystal clear with them about where the company is heading.
AI is really good at execution if we know what we want to do. This will change how we should hire new people and how we should reward them. The way we used to work was to have a senior engineer sense the problem, come up with a solution, and delegate to other engineers for implementation so we could work on it in parallel and be faster. But these days, we found that senior engineers would rather not delegate implementation work to other people. Because they already figured out how to solve the problem, they just use AI to solve it instantly. Delegating to other people means more overheads and is less efficient.
That also means when you have new people, you want to make sure they can solve the problem independently. They pretty much have to work as a traditional tech lead. We cannot afford other people to delegate implementation tasks for them.
Also, when we hire new people, we should think about what we are looking for. We definitely want to look for somebody genuinely interested in AI. The domain is moving so fast. We want them to keep learning and also help the team stay on track. Secondly, with AI, engineers can do way more than they used to do. The boundaries between PMs and engineers are getting blurry. We found engineers who really understand the product can, in fact, make way more contribution than the traditional engineer who only focuses on software. This is what we are looking for. That deep understanding of the system and the ability to handle complicated, ambiguous problems are also very valuable. This is where AI drops off. When we hire new people, this is also the kind of people we are interested in bringing on board.
For people we bring in, we want to reward them in the proper way. Even in our performance review, we start to ask, okay, what have you done on the AI side? We definitely want to reward people who leverage AI to multiply their impact. This is the impact for everybody in the company.
Now we have the right tools, we get the right talent in place, and we need to change how we work to maximize the benefit of AI. The way we used to work was to say, okay, we spend weeks, sometimes months, to flesh out the business requirements, finalize the design, and then do the implementation. Because implementation can be really expensive. If we didn't get the other part right in the beginning, it can be very costly to change it later. But in fact, we never get things right in the beginning anyway, for any big projects.
With AI, building is super fast. It's probably a couple minutes. You can get it done. Alignment is the really expensive one. We should really think about what's the best way we can work. How can we deliver fast? It's still okay to think about what you want to deliver in one year. You can assume AI models can do anything you want in one year. Based on that, really dream big to think about what you can do in one year. But it should only serve as inspiration and direction.
What we really need to focus on is what we want to deliver in the next two to four weeks. We want PMs and designers to say, okay, tell me what I need to do in this sprint. Then the engineers will focus on it and get it released at the end of the sprint, if not sooner. Meanwhile, the PMs have time to flesh out the next batch of requirements. If at the end of the sprint they say, no, what we decided two weeks ago is wrong, it's totally okay. We can switch gears and get it fixed quickly.
That also means we prefer people not to write pages and pages of PRDs or TDDs anymore. We prefer them to write just a short one or two pages. That really serves as communication so we can iterate on it.
The really awkward part is the midterm goals. Those are three months, six months. It's very hard to plan these days. The reason is, I don't know what AI models will be capable of in three months. There may be multiple releases already. So we prefer not to focus on this one. This may make it very hard for most of the folks who have been in this domain for a long time, because traditionally, we get used to having quarterly planning or planning for six months. But it's our job to get used to the new AI era and learn how to work in it efficiently.
I want to talk about coding and software development a little bit more here. AI coding tools are probably the most successful AI application, and they're really good at implementation. You probably heard a lot of people say, okay, I have these AI tools. Now I can even use my phone to implement software and automate every stage. If they feel comfortable doing that, it's totally okay. But you don't have to.
What I'm trying to say here is, and maybe this was our journey, how we adopted those AI tools. We started with the lowest-risk tasks, like writing unit tests and documentation. Those things are very easy to verify, and the risk is super known. By doing that, we built confidence. We started to construct our own rules and skills and build our guardrails. Then we pushed to the whole engineering team and said, now you should use these AI coding tools for all the tasks. When they choose not to, that's the time we really want to learn and ask, why don't you do it?
At this moment, we pretty much use AI coding tools to do all our implementation. Engineers really focus on reviewing, architecture, and evaluation.
With AI coding tools, we are writing so much code these days. Code review becomes really challenging. A good engineer used to probably write hundreds of lines of code every day. These days, they can easily write thousands. If we do code review as we used to do, we won't be able to keep up. We also tried multiple AI code review tools. It helps a little bit, but we don't feel comfortable 100% relying on them yet. We still find the feedback from our engineers very, very valuable. That means we need to really change the way we are doing code review to meet where we are now.
A couple things we have done: one is we allow engineers to self-identify whether they still need code review. If they think this PR is simple enough and they feel very confident, they don't need anybody to take a look. We are fine with that. We let them merge, but we still hold them accountable.
If they do want code review, we want them to stay with the best practices. For example, each PR shouldn't have more than 500 lines of code because nobody can do a meaningful code review on one that has thousands of lines of code. We also enabled stacked PRs. What it means is that for a big feature, engineers can break it into multiple PRs. While people review the PRs, they can keep working on it.
One thing we really want to avoid is a rubber stamp, we call it. It means people submit code for review, but you cannot really do anything to it. You just blindly approve it. This is the worst case we should really avoid because that just gives us false confidence. We think we reviewed it, it's good, and we release it. Meanwhile, we should keep working on our AI code review tools because we think that's the future.
At this moment, we use AI tools to pretty much assist in each step of our software development. Our goal is for it to automate the whole life cycle from end to end, from design and implementation until it's fully released. More importantly, we want the AI tools to be able to monitor the live traffic, catch issues early, and automatically fix them. That's what we are still working on, and we are not there yet.
The last thing I want to touch on a little bit for this presentation is reliability. What it means is that for traditional software, it does what we implement there. No more, no less. But for GenAI solutions, hallucination is there. We cannot ignore it. Completely eliminating it can be very costly. Sometimes it's not necessary either. So the way we should do it is really have a holistic solution, even from the beginning.
For example, we can start with identifying which failures are acceptable and which ones are not acceptable. For our AI system, for example, we have the functionality to help our customers schedule appointments. If we fail one out of one thousand, probably it's okay. I'm not saying it's a good experience, but users can usually just click the button again and we will reschedule for them. Probably it's okay.
But if we help the user submit their reimbursement claim, we cannot tolerate failure. Because if people ask for $200 and we issue them $50, or we give them $200 incorrectly, each case will cause an escalation right away. For those cases, we have to put in extra steps. For example, when we receive their receipt, we will use different models to review the same receipt. We only move forward if the results from different models agree with each other. If we really have trouble figuring out which one is right, it's okay to tell the customers, hey, how about processing your stuff? Do you want us to connect you to a human agent? We will move from there. That's our set of solutions.
Also, we should have a rigorous process to release our software. For us, we have hundreds of integration tests, which pretty much cover all the use cases we know, and we keep adding to the integration test suite. When we run the integration tests, passing once is not good enough anymore, because the LLM can do different things. So for each test case, we run it many times. We consistently require high pass rates, for example 90% over time.
More importantly, after we launch the software, we have our auto-eval system carefully evaluate each conversation. We have predefined a lot of rubrics for what we think is good and what is bad. Then we review the scores. Besides this, we also have a dedicated group whose job is to manually review those conversations. We will spot-check our conversations. That helps us see whether we need to come back and improve our systems, or whether our rubrics are too strict or too loose, and we need to consistently improve them.
When we launch new features, that's the time we say spot-checking is probably not enough. We really want to review 20%, and we can do it. This whole process makes sure we feel really confident we never ship something, although we know hallucination is there. That's pretty much all I have for today, and I can stay here to take questions. If you have other things, you can reach out to me. where they are, just make it feel they're comfortable to use it. And for all the places, you always have a few slow adopters. They always have concerns, worries for the new technologies. For them, we just should
meet where they are, understand what their concern is. But more important, we should be crystal clear with them, where the company is heading to. So AI is really good at execution if we know what we want to do. So this will change how we should hire new people and how should we reward it. We used to, the way we used to work is we have a senior engineer who will sense the problem, come up with a solution and dedicate to other engineers for implementation. So we can work on it in parallel and be faster. But these days, we found the NASA engineers would like to dedicate implementation work to other people. Because they already figured out
how to solve the problem, they just use AI to solve it instantly. Dedicating to other people means more overheads and we manage efficient. That also means like when you have new people, you want to make sure they can solve the problem independently. They pretty much have to work as a traditional technical level. We cannot afford other people to dedicate implementation tasks for them. Also, when we hire new people, we should think about what we are looking for. We definitely want to look for somebody genuinely interested in AI. The domain is moving so fast. We want them to keep learning. Also, help the team
to stay on track. Secondly is with AI, engineers can do way more than they used to do. The boundaries between PM and engineers is getting blurry. We found engineers who really understand the product, in fact, they can have way more contribution than the traditional engineer who only focus on software sites. And this is what we are looking for. And those deep understanding of the system, the ability you can handle complicated, ambiguous problem is also very valuable. This is where AI rank off. When we hire new people, this is also the people we are interested in bringing on boards. For people, we bring in, we want to
reward them in the proper way. Even in our performance review, we start to ask, okay, what you have done for AI sites. We definitely want to reward people who leverage AI to multiple their impact. Although this is the impact for everybody in the company. So now we have the right tools, we get the right talent in the place. And we need to change how we work to maximize the benefit of AI. The way we used to work is to say, okay, we spend the weeks, sometimes in the months, to flesh out the business requirements, finalize the design, and then do the implementation. Because implementation
can be really expensive. If we didn't get the other part right in the beginning, it can be very costly to change it later. But in fact, we never get the things and the rights in the beginning anyway, for any big projects. With AI, building is super fast. It's probably a couple minutes. You can get it done. Argument is really expensive one. So we should really think about what's the best way we can work. How can we deliver fast? It's still okay. You can think about what you want to deliver in one year. You can assume AI models can do anything you want in one year. Based on that one, really dream big to think what you can do in one year.
But it should only serve as inspiring and directional. What we really need to focus on is what we want to deliver in the next two to four weeks. What we want to get the PMs and designers to say, okay, tell me what I need to do in this sprint. And the engineer will focus on it and get it released and end of the sprint if not sooner. Meanwhile, the PMs, they have time to flesh out the next bunch of the requirements. If end of the sprint, they say, no, what we decided two weeks ago is wrong, it's totally okay. We can switch the gear, get it fixed quickly. That also means we prefer people not
to write pages or pages of PRD or TDD anymore. We prefer them to write just a short one or two pages. That one really serves as communication so we can iterate on it. The really awkward part is the mid-term goals. Those are like three months, six months. It's very hard to plan these days. The reason is, I don't know what AI models will be capable in three months. There may be multiple releases already. So we prefer not to focus on this one. But this one can maybe make it very easy for most of the folks who has been in this domain for a long time. Because traditionally, we get used to have a quarterly
planning or we plan it for six months. But it's our job to get used to the new AI error and learn how to work it efficiently. So I want to talk about the coding and software development a little bit more here. AI coding tools is probably the most successful AI application. And it's really good and implementation. So you probably heard a lot of people say, okay, I have this AI tools. Now I can even use my phone to implement software and automate every stage. If they feel comfortable to do that, it's totally okay. But you don't have to. What I'm trying to say here is, and maybe when this is our
journey, how we adopt those AI tools, we started with the lowest risk task, like starting with writing unit has documentation. Those things are very easy to verify. And the risk is super known. By doing that one, we build the confidence. And we start to construct our own rules, skills, and build our guardrails. And then we push to the whole engineer team say, now you should use this AI coding tools for all the tasks. When they choose not to do, it's the time we really want to learn, say, why you don't do it. And at this moment, we pretty much use the AI coding tools to do all our implementation. Engineers
really focus on reviewing, architecturing, and evaluation. So, and with AI coding tools, we are writing so much code these days. Code review becomes really challenging. So for good engineer, used to, they probably write hundreds of less code every day. These days, they can easily write like thousands. If we can do the code review as we used to do, we won't be able to keep up. We also try the multiple like AI coding review tools. It helps a little bit, but we don't feel comfortable 100% rely on them yet. We still find the feedbacks from our engineers are very, very valuable. That means we need to really
change the way we are doing code review to meet where we are now. And a couple things we have done. One is we allow engineers to self-identify whether they still need code review. If they think this PR is simple enough, I feel very confident, I don't need anybody to take a look. We are fine with that one. We need them merge, but we still hold them accountable. And if they do want code review, we want them to stay with the best practices. For example, each PR shouldn't have more than 500 lines of code because nobody can do a meaningful code review with the ones that has like thousands of less code. And we also enabled
like stanked PR. What it means is that for big feature and engineers can bring it into multiple PRs while people review the PRs and they can keep working on it. One thing we really want to avoid is a rubber stamp, we call it. It means like people submit code review. You cannot really do anything to it. You just say, blindly prove it. This is the worst case we should really avoid because that just gives us false confidence. We think we reviewed it, it's good, and we release it. Meanwhile, we should keep working on our AI coding review tools because we are thinking that's the future. So at this moment, we use AI tools pretty much
assist in each step of our software development. Our goal is it will be automating the whole life cycle from end to end, from designing, implementation, until it's fully released. More importantly, we want the AI tools to be able to monitor the live traffic and be able to catch the issue early and automatically fix it. That's what we are still working on and we are not there yet. The last thing I want to touch a little bit for this presentation is about reliability. So what it means is like for the traditional software, it does what we implement there. No more, no less. But for the GNI solutions, hallucination is there. We cannot
ignore it. And completely eliminating them. It can be very costly. Sometimes it's not necessary either. So the way we should do is really have a holistic solution, even from the beginning. For example, we can start with identifying which failures are acceptable, which ones are not acceptable. For our AI system, for example, we have the functionality to help our customers to schedule appointments. If we fail one out of one thousand, probably it's okay. I'm not saying it's a good experience, but the users usually can just click the button again, we will reschedule for them. Probably it's okay. But if we
help the user to submit their reimbursement claim, we cannot tolerate the failure. Because if people ask of $200, we issue them $50, we give them $200. Each case will cause an escalation right away. For those cases, we have to put in extra stamps. For example, when we receive their receipt, we will use different models to review the same receipt. We only move forward if the results from different models agree with each other. If we really have trouble to figure it out, which one is right? It's easy. It's okay to tell the customers. Say, hey, how about to process your stuff? Do you want us to get you connect to a
human agent? We will move from there. That's our center solutions. And also, we should have a rigorous process to release our software. For us, we have like hundreds of integration tests, which pretty much covered all the use cases we know, and we are keep adding to the integration test suite. And when we run the integration test, not only pass once is not good enough anymore, because the ARM can do different things. So for each test case, we run it many times. We consistently require the high pass rates, like for example, 90% for all the time. And the more important, and after we launch the software,
we have our auto evolve system, carefully evaluated each conversation. We have predefined a lot of rubrics. What we think is good, what is bad. And then, we will general results. We will review the score. Besides this one, we also have a dedicated group. Their job is manually review those conversations. We will spot check our conversations. That helps us to say whether we need to come back to improve our systems, or our rubrics is too strict or too loose, and we need to consistently improve it. When we launch new features, then the time we say not on spot check probably not enough. We really want to review,
like say, 20%, and we can do it. This whole process makes sure we feel really confident we never reach something, although we know hallucination is there. That's pretty much all I have for today, and I can stay here to take up questions, and if you have other things, you can reach out to me.