Your Fine-Tuned Model Is Tech Debt: A 50x ROI House of Cards — Dan Bjornn, Lease End
Description
A customer replied good morning to an outreach text and the model called him immediately. Another confirmed a Thursday appointment, said sounds good, and was told a call was happening right now. Both reached production, from a finetuned classifier that had also generated $12 million of revenue at 50 times return inside a year. Dan Bjornn's talk is about what that model was quietly costing underneath those numbers, which he calls the calcification tax. The repair loop is where it accrued. Gather examples of the new failure, synthesize more when there are too few, validate those by hand, sort them into intent buckets, review again, and only then train, which took about an hour and was the shortest step in a process that ran a week. Each round fixed its target and reintroduced something older, so bugs ended up ranked by how much customer pain was tolerable while they waited. The promised portability never arrived either, since training data does not transfer cleanly between model versions, let alone between providers, so they stayed put and could not adopt newer architectures while busy keeping the old one alive. The rebuild swapped the tuned model for skills, prompts, and context on a model agnostic framework. Fixes now ship in under an hour as files uploaded to a bucket, accuracy went up, cost per message went up, and total cost went down. Speaker info: - https://www.linkedin.com/in/dkbjornn Timestamps: 0:00 - Classifying customer intent with retrieval 1:54 - Four reasons to finetune, all of them reasonable 3:32 - The pipeline, and $12 million at 50x return 4:24 - The confused confirmer, and the overeager puppy 6:04 - A week per retrain, with training the shortest step 8:39 - Ranking bugs by tolerable customer pain 9:31 - The calcification tax, in model and architecture 11:18 - The realization from changing skills, not models 12:14 - Rebuilding on skills, tools, and context 13:08 - Fixes in under an hour, deployed as files 14:54 - Cross your reason off the list be
Summary
Generated by gpt-5.6-terraAt-a-Glance
- Verdict: Watch fully
- Core thesis: For production LLM applications that can call frontier models, fine-tuning often creates a hidden maintenance and vendor-lock-in burden that context-driven skills, tools, and resources can outperform operationally.
- Why it matters: LeaseEnd found that a seemingly ideal fine-tuning use case—six-way intent classification at high message volume—produced costly regressions and week-long fixes, while an agentic, context-configured rebuild improved accuracy and cut fixes to under an hour.
- Best use: Use this as a practical architecture and operating-model case study for deciding whether to fine-tune, and for designing fast, model-agnostic remediation loops for customer-facing agents.
Executive Summary
Dan Bjornn describes LeaseEnd's late-2024 SMS application for customers approaching auto-lease maturity. The system answered sales-process questions, scheduled calls, and issued reminders. Its initial workflow/RAG design classified customer messages into six intents using a vector database of previously labeled examples. Because intent accuracy drove every downstream action, LeaseEnd fine-tuned a smaller model to improve classification, reduce inference cost and latency, and gain apparent provider flexibility.
The system produced meaningful business results—$12 million in revenue within a year and a claimed 50x ROI—but also made damaging conversational mistakes. It treated a customer's "Sounds good" response to a next-day appointment confirmation as a request for an immediate call, and interpreted "Hi, good morning" as an invitation to call. The issue was not merely model errors; it was that each correction required collecting or synthesizing examples, manually validating and labeling them, retraining, regression testing, iterating, and deploying. The training job took about an hour, but the end-to-end repair cycle took roughly a week and frequently introduced new regressions.
LeaseEnd ultimately replaced the fine-tuned workflow with an agentic framework built from skills, tools, and loadable resources. Rather than retraining a model for a behavior change, the team adjusted the relevant skill or system prompt, validated it against a curated production-derived set, and deployed Markdown-based configuration files to S3. This reduced problem discovery-to-deployment time to less than an hour, improved accuracy beyond the fine-tuned system, and allowed model-provider switching. Although per-message API cost rose because the new system used better models, total cost fell because maintenance work fell sharply.
Bjornn's conclusion is deliberately strong: fine-tuning should be the exception, primarily when a frontier model cannot be called, such as certain privacy/data-control or offline constraints. His central decision criterion is not token cost or narrow-task fit alone, but whether fine-tuning's full lifecycle cost—data operations, regression risk, retraining latency, model/version coupling, and architectural rigidity—beats a context-and-tools alternative.
Key Takeaways
- Claim: Fine-tuning can look highly justified for a narrow, structured classification task yet still become production tech debt. | Evidence: LeaseEnd classified inbound SMS conversations into six intent categories, handled thousands of real-time messages daily, and expected fine-tuning to improve accuracy, latency, and cost; nevertheless, it accumulated what Bjornn calls a "calcification tax." | Implication: Do not approve fine-tuning solely because a task is bounded, supervised, high-volume, or financially successful; evaluate how quickly the behavior will need to evolve after deployment. | Caveat: The speaker does not claim fine-tuning never works; the initial application generated $12 million in revenue within a year at a claimed 50x ROI.
- Claim: The true cost of a fine-tuned model is dominated by its repair loop, not the model-training runtime or per-call price. | Evidence: The actual fine-tuning run usually took about one hour, but a production fix took about a week because the team had to collect failure examples, synthesize more if needed, manually validate them, label them, create holdouts, retrain, evaluate, and repeat after regressions. | Implication: Account for evaluation-data maintenance, human labeling and review, regression testing, and deployment cadence as first-class costs in the model TCO—not just inference cost and training compute.
- Claim: Fine-tuning made LeaseEnd's customer-facing failure handling a triage exercise based on tolerable harm. | Evidence: The model sometimes immediately called customers after a simple greeting or after they acknowledged a future appointment. Before retraining, the team asked how frequent a bug was, whether it materially hurt customer experience, and whether a temporary band-aid could avoid a retrain; they explicitly ranked bugs by how much customer pain they could tolerate. | Implication: For agents that trigger external actions, separate reversible conversational defects from action-execution failures, and ensure critical-action safeguards do not depend on a slow retraining cycle. | Caveat: Some failures required immediate treatment regardless of frequency, including ignored call-time preferences and malformed scheduling payloads that meant a promised call was never actually scheduled.
- Claim: Fine-tuning did not deliver the expected vendor freedom; it instead coupled LeaseEnd to a particular model and architecture. | Evidence: Bjornn says that model-version changes within a provider altered the appropriate training data, while moving across providers changed data format, data volume needs, and training interfaces. The team retained the same model because retraining and migration were too costly. | Implication: Treat fine-tuned weights and datasets as provider- and version-specific operational assets unless portability has been demonstrated through actual cross-model evaluation and migration work.
- Claim: Contextual skills, tools, and resources can replace many fine-tuning-driven behavior changes with fast, auditable configuration changes. | Evidence: After observing that "Cloud Code" changed performance by changing task context rather than the model, LeaseEnd migrated to skills that loaded relevant tools and resources. A defect could then be addressed by changing the affected system prompt or skill, testing against a curated set, and uploading Markdown files to S3. | Implication: Prioritize a modular control plane for prompts, skills, policies, resources, and evals so behavior can be corrected without weight updates. | Caveat: This approach used better models and raised API cost per message.
- Claim: The context-driven rebuild improved both operational agility and overall economics despite higher inference cost. | Evidence: LeaseEnd reduced fix time from about a week to less than an hour, reported materially better accuracy than with fine-tuning, and reduced total cost because it spent less effort keeping the system operational, even though API cost per message increased. | Implication: Optimize for total system cost and speed of safe iteration; paying more for a stronger frontier model may be rational when it removes recurring ML operations and customer-experience risk. | Caveat: No absolute accuracy metrics, message volumes, or total-cost breakdowns are supplied, so the magnitude and generality of the gain cannot be independently assessed from the talk.
- Claim: Fine-tuning should be a constrained exception rather than the default path to accuracy, latency, lower cost, or model control. | Evidence: Bjornn reports that the rebuild beat fine-tuning on accuracy, the smaller-model latency gain was marginal in practice, and total cost declined despite higher per-message cost. He identifies privacy/data-control and offline requirements as possible cases for fine-tuning, then concludes it should be used only when a frontier model literally cannot be called and the decision still beats the lifecycle tax. | Implication: Require an explicit exception case for fine-tuning, with a comparative prototype against retrieval/context/tooling approaches and a quantified ongoing maintenance plan. | Caveat: The recommendation reflects one company's intent-routing application and architecture transition; it is not evidence that fine-tuning is unsuitable for every offline, private, highly specialized, or extreme-scale workload.
Detailed Brief
The original architecture and why it appeared rational
- Claims: LeaseEnd's first implementation combined deterministic workflows with RAG-based intent retrieval rather than relying on a fully open-ended assistant.; The vector database retrieved previously seen, intent-labeled messages; examples included distinguishing "Call me tomorrow" from "I've got time now."; The team used an LLM-as-judge pipeline to label examples, manually reviewed the labels, maintained holdout sets, and measured fine-tuning results.
- Evidence: The application handled SMS questions, sales-call scheduling, and reminders for customers nearing lease maturity.; The core decision space was six categorical intents, which made supervised fine-tuning appear to be a textbook match.
- Caveats: Retrieval over historical messages did not reliably capture conversational nuance, which is what prompted the shift toward fine-tuning in the first place.
- Implications: A structured intent taxonomy does not remove the need to represent conversation state, preceding outbound messages, and action preconditions.; A formal labeling and holdout workflow is necessary but does not by itself solve post-deployment behavioral drift or regression management.
What the new operating model changes
- Claims: The rebuilt system treats behavioral policy as deployable context rather than embedded training behavior.; LeaseEnd accumulated a curated validation set during production, allowing the team to test targeted updates before release.; The agentic framework was designed to be model agnostic: OpenAI, Anthropic, or other models could be used as long as the relevant context was provided.
- Evidence: The deployment unit was Markdown content uploaded to an S3 bucket.; The team used the rebuilt messaging application as one of the first production tests of an agentic framework that was already under construction.
- Caveats: The transcript does not specify the framework's action-authorization controls, tool schemas, rollback mechanisms, or evaluation thresholds.; Model agnosticism still requires ongoing cross-model evaluation; a common context interface alone does not guarantee equivalent behavior.
- Implications: Configuration should be versioned, testable, and rapidly deployable, but high-impact actions still need deterministic guardrails outside natural-language policy.; An agent architecture can preserve optionality only if its skills, resources, and test suite are decoupled from provider-specific assumptions.
Notable Concepts & Terms
- Calcification tax: Bjornn's term for the increasing rigidity created by repeated fine-tuning: every correction, model change, and architectural upgrade becomes harder because the system depends on accumulated training data and model-specific behavior.
- LLM-as-judge: LeaseEnd used an LLM to classify or label training examples before manual review; it accelerated dataset creation but retained human validation as a quality control step.
- Fine-tuning regression / whack-a-mole: A retraining iteration solved the immediate failure but commonly degraded previously working behavior, creating repeated evaluation and retraining cycles.
- Skills, tools, and resources: The replacement architecture modularized behavior around task-specific instructions, callable capabilities, and relevant context instead of encoding more behavior in model weights.
- Curated validation set: A production-derived test collection used to validate prompt or skill changes before deployment; it is the key safety mechanism enabling rapid configuration changes.
- Model agnosticism: The ability to use different providers or models through a common agent framework; Bjornn argues this comes more from portable context and skills than from transferable fine-tuned datasets.
- Total cost versus per-message cost: The talk distinguishes API inference expense from full operational cost, including data curation, review, retraining, regressions, and engineering time.
Operator Notes / Why Ken Should Care
- Create a fine-tuning exception gate: require teams to document why a frontier-model-plus-context, retrieval, tools, and structured-output design cannot meet the requirement.
- For every agent that triggers a call, booking, payment, message, or other external action, implement deterministic confirmation and payload-validation checks independent of intent classification.
- Maintain a continuously growing, production-derived regression suite segmented by customer harm severity, not just aggregate accuracy.
- Version skills, system prompts, resource files, tool contracts, and evaluation results as deployable artifacts; establish rollback and approval paths for high-impact behavior changes.
- Measure agent TCO as inference spend plus evaluation, labeling, incident handling, regression repair, deployment latency, and model-migration cost.
- Test provider/model portability periodically using the same skill and evaluation suite before an outage or pricing change makes migration urgent.
Source/Metadata
- Title: Your Fine-Tuned Model Is Tech Debt: A 50x ROI House of Cards — Dan Bjornn, Lease End
- Transcript words: 4496
- Duration seconds: 999
- Timestamp note: No usable timestamps or chapters were present in the supplied transcript; the transcript also contains a substantial duplicated segment and applause/filler repetition.
Transcript
Dan Bjorn Reviewer All right. Hello, everybody. Thank you for coming. I'm Dan Bjorn. I'm a senior data scientist at LeaseEnd. At LeaseEnd, we connect people who are coming to the end of their auto lease with financing options so that they can buy out their lease and keep their car. As part of this, we built an LLM-based application in late 2024 to help our customers connect with our sales team. This application allowed them to send messages through text. They could ask questions about the sales process. They could schedule calls. They could get reminders, all of this stuff. Our first solution used a workflow-based approach built on top of a RAG system where we searched a vector database of messages that we had already seen and classified with the customer's intent. For example, a message saying, "Call me tomorrow," would be classified as the customer wants to talk later. A message saying, "I've got time now," would be classified as the customer wants to talk right now. This worked, but not super amazingly. There's a lot of nuance in messages and conversation, and this RAG approach just couldn't quite pick up on that nuance. So we started to look for new options to improve this. Naturally, being a data scientist, my first thought was, "Hey, let's start fine-tuning." This seemed like a fun thing to do, and I was sure that this was the right call. There's a few reasons for that. First of all, we needed better accuracy. Our entire system was built upon us getting the user's intent correct. Did they want to talk now? Did they want to schedule a call? Did they want to opt out? All of this hinged on that decision, and so we needed to make sure that we got that first and foremost. Next, we could use smaller models with fine-tuning, and so this would lower the cost and also lower latency. This was really important for us because we were responding to thousands of messages a day in real time, and so it would help us scale a lot. Then next, as I said, we were classifying the intent of the user, and so this was a very narrow, structured task that we were trying to do. So it lent itself very nicely to supervised fine-tuning. We would bucket that conversation in one of six different categories, and the model would learn the differences between those. So it seemed like a great option there. Lastly, I believed that this would help us have a little bit more control over our destiny with the model providers. The idea was that we had the data, and all we would need to do is pass that into a new model, go through the fine-tuning process, and we could get similar results no matter what we decided to use. So we could be model agnostic. So this was the approach that we took, and I built a pipeline to collect examples, run LLM-as-judge classifications to label our data. I'd manually review that, create holdout sets, go through the fine-tuning process, and check my metrics. This was a data scientist's dream. And the numbers sure helped. Within a year, this application had helped us bring in $12 million of revenue at a 50x ROI. It was pretty awesome. But the whole time, it was quietly accumulating debt underneath that we didn't see. So I want to show a couple examples of how this application could get things wrong. First of all, the confused confirmer is a situation where, when customers set up an appointment with a sales rep, we send them a confirmation message to let them know that it's been scheduled and give them the details of that. So a conversation may look like this. We reach out and say, "Hi, Tracy. Just confirming your lease end call with your advisor is set for Thursday at 2 p.m. We'll call you then." Tracy then sends us a message back saying, "Sounds good." And then our LLM responds with, "Great, I'm calling you right now." It's not what we want. We just confirmed an appointment for the following day, and then all of a sudden we start calling them. This led to frustrated customers and some missed opportunities. The next one I've come to lovingly call the over-eager puppy. The conversation looks like this. First, "Hi, James. This is Alex with Lease End, reaching out about your upcoming lease maturity." James then says, "Hi, good morning." And then, "Good morning. I'm giving you a call." Just like a puppy that gets so excited that somebody's giving it attention, our model decided to give a call right there. Obviously, this is not what James wanted. This actually did happen in production. Very embarrassing there. These are a couple examples of where it went wrong. And don't get me wrong, the app did well. The revenue numbers show that it was working. But it could also mess up pretty spectacularly. The big issue wasn't how to fix it, but how to make the fix manageable. The fine-tuning process was pretty complex. First, we needed to gather examples of the problems that we started to see. Then, we needed to ask ourselves, do we have enough examples to go through fine-tuning? If not, we synthesized those examples. We passed it through an LLM. We created some possible examples there. We'd have to validate those, which was a very manual process, because we wanted to make sure it had the best training data possible. And then, once we had enough, we labeled those with the categorization bins, and we validated those through a manual review. Surprisingly, the fine-tuning process was the shortest part of all of this. Normally, it took about an hour, depending on the size of the data that we had. But we never got it on the first iteration. Normally, what happened was we would fine-tune, and we'd evaluate this. We fixed the problem that we were just trying to solve, but then we caused regressions in other things. And so this turned into a whack-a-mole process, where we would solve something new, but then other old issues kept popping up that we had to whack down. This whole process took about a week to gather the data, label everything, go through the fine-tuning process, iterate, and then deploy. So it was costly. Therefore, we needed to triage all of these issues that we ran into. We asked ourselves three questions before we did any retraining. How frequent is the issue? Is it something that customers are seeing every day? Is it one-off? One big exception to this was if it was hurting the customer experience too much. For example, this would be somebody repeatedly stating what their preference for a call time is, and then the model ignoring that. Another one would be a customer scheduling a call. We tell them that we've scheduled it for them, but we don't return the payload in the proper way, and so the call never gets scheduled, and so we don't follow up with them. So these kinds of things needed to be fixed right away. But before we did that, we asked the last question. Is there anything that we can do in order to prevent a retrain? Can we have some kind of a band-aid fix to get out there so we don't have to go through a whole week-long process for one or two issues? And so we ranked our own bugs based on how much customer pain we could tolerate at the moment. Not a great situation to be in with a production system. This led to what I've come to call the calcification tax. The more we used the model, the more rigid everything became. It manifested in a couple different ways. First, we were locked into our model. You remember when I said that fine-tuning would give us more freedom in what model we used? That was not the case. Within providers, there's nuance between one model version and another, and so that changes the training data that you need to provide it. Across model providers, it's extremely different. The structure of the data you need to pass to it would be different, the amount of training data to get good results, the way to interact with the training interface. All of this caused a lot of complexity, and so it was just too costly for us to switch. So we kept the same model for consistency because we already had a lot to do with each retraining process, and we couldn't afford to upgrade the model. The other way that this locked in was architecture. We built this app in late 2024, when workflows were the gold standard if you wanted good production results. And the AI world moves very fast. We couldn't adapt to that because we were so locked into this, just trying to keep it running. And we couldn't take advantage of the new architectures and improve performance that way. So earlier this year I had an aha moment. We started using Cloud Code for our coding tasks, and I noticed that we never needed to change the model depending on what task we were using. We just changed the skill, the resources that we passed it, the context. You drop in better context, you get better results. And I thought, why can't we do this with our messaging app? This was obviously difficult for me to admit because I was the champion for fine-tuning. Luckily, we were able to piggyback on a project that was already happening, and so we migrated our workflow approach to a series of skills, tools, and resources that the skills could load into, or load up, and get that context. So we pushed this as one of our first production tests of our new agentic framework that was already being built. Now, I want to compare the process before and after our rebuild. Before, we already went through the training cycle, but there was this triage cycle beforehand where we needed to make sure that we had reached a critical mass of problems before we would even attempt to fine-tune again to improve everything. As I said, this took about a week. So it was a long, costly process. After the rebuild, it was a simple process of: you find a problem, you adjust the system prompt or the skill that was affected, we validated performance on a curated set that we had been collecting over the time that this was in production, we iterate a few times, and we deploy that simply by uploading MD files to an S3 bucket. This whole process, from discovering a problem to deploying the fix, we reduced down to less than an hour. So it extremely improved all of this, and we could be far more reactive, giving our customers way better performance, or a better experience there. Now, I'll be honest, it did cost us a little bit more per message. We were using better models, so the API costs were a little higher. But accuracy went way up. I said before that accuracy was the key to getting all of this right, and we did that. Accuracy was far better with this than it ever was with fine-tuning. Next, as I said, we reduced our fix process from days down to minutes. Next, we were able to unfreeze our model and finally get that freedom from a vendor that we never had with fine-tuning. Our agentic framework was built model agnostic, so we could use OpenAI. We can use Anthropic. We can use any other model that we want. The important part is the context that we're providing to that model. And then lastly, while it cost us a little more per message, the total cost went down because we were spending far less time trying to keep it up and running and fine-tuning to keep it working properly. So before you fine-tune, I'd ask you, can you cross your reason off of this list? I thought we would get better accuracy. The rebuild beat the fine-tuned model. I thought we would get lower cost at the volume we were doing. I was looking at the wrong cost. We paid more per message, but the total cost ended up going down with our rebuild. Lower latency: we did see marginal gains on these smaller models, but they were so small that in practice, it really didn't make any difference. And then maybe you've got a narrow or structured task. Our textbook case still became tech debt. And lastly, vendor control. It's not as simple as just plugging the data in. The other two situations where you might have privacy and data control, or you need some offline solution, I would say these are the situations where a fine-tuned model may be useful, but you need to be cautious. There are other solutions out there, but you need to make sure that it's not causing issues in the long run. So finally, fine-tune only when you literally cannot call a frontier model. And even then, your decision still has to beat the tax. Thank you. That yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay But the whole time, it was quietly accumulating debt underneath that we didn't see. So I want to show a couple examples of how this application could get things wrong. First of all, the confused confirmer is a situation where when customers set up an appointment with a sales rep, we send them a confirmation message to let them know that it's been scheduled and give them the details of that. So a conversation may look like this. We reach out and say, hi, Tracy. Just confirming your lease end call with your advisor is set for Thursday at 2 p.m. We'll call you then. Tracy then sends us a message back saying, sounds good. And then our LLM responds with, great, I'm calling you right now. It's not what we want. We just confirmed an appointment for a following day. And then all of a sudden we start calling them. This led to frustrated customers and some missed opportunities. The next one I've come to lovingly call the over-eager puppy. The conversation looks like this. So first, hi, James. This is Alex with Lease End, reaching out about your upcoming lease maturity. James then says, hi, good morning. And, good morning. I'm giving you a call. Just like a puppy that gets so excited that somebody's giving it attention, our model decided to give a call right there. Obviously, this is not what James wanted. This actually did happen in production. Very embarrassing there. But, this is a, these are a couple examples of where it went wrong. And, don't get me wrong, the app did well. The revenue numbers show that it was working. But, it could also mess up pretty spectacularly. The big issue wasn't how to fix it, but how to make it, the fix manageable. The, the fine tuning process was pretty complex. First, we needed to gather examples of the problems that we started to see. Then, we needed to ask ourselves, do we have enough examples for, to go through fine tuning. If not, we synthesized those examples. We passed it through an LLM. We created some, some possible examples there. We'd have to validate those, which was a very manual process, because we wanted to make sure it had the best training data possible. And then, once we had enough, we labeled those with the categorization bins. And, we validated, validated those through a manual review. Surprisingly, the fine tuning process was the shortest part of all of this. Normally, it took about an hour, depending on the size of the data that we had. But, we never got it on the first iteration. Normally, what happened was, we would, we would fine tune, and we'd evaluate this. And, we fixed the problem that we were just trying to solve. But, then, we caused regressions and other things. And, so, this turned into kind of a whack-a-mole process, where we would solve something new. But, then, other old issues kept popping up that we had to whack down. This whole process took about a week to gather the data, label everything, go through the fine tuning process, and iterate, and then deploy. So, it was costly. Therefore, we needed to triage all of these issues that we ran into. We asked ourselves three questions before we did any retraining. How frequent is the issue? Is it something that customers are seeing every day? Is it one-off? One big exception to this was if it was hurting the customer experience too much. So, for example, if this would be somebody repeatedly stating what their preference for a call time is. And, then, the model ignoring that. Another one would be a customer scheduling a call. We tell them that we've scheduled it for them, but we don't return the payload in the proper way. And, so, the call never gets scheduled. And, so, we don't follow up with them. So, these kinds of things needed to be fixed right away. But, before we did that, we asked the last question. Is there anything that we can do in order to prevent a retrain? Can we have some kind of a band-aid fix to get out there so we don't have to go through a whole week-long process for one or two issues? And, so, we ranked our own bugs based on how much customer pain we could tolerate at the moment. So, not a great situation to be in with a production system. This led to what I've come to call the calcification tax. The more we used the model, the more rigid everything became. It's manifested in a couple different ways. First, we were locked into our model. You remember when I said that fine-tuning would give us more freedom in what model we did? That was not the case. Within providers, there's nuance between one model version to another. And, so, that changes the training data that you need to provide it. Across model providers, it's extremely different. The structure of the data you need to pass to it would be different. The amount of the training data to get good results. The way to interact with the training interface. All of this caused a lot of complexity. And, so, it was just too costly for us to switch. And, so, we kept it the same model for consistency because we already had a lot to do with each retraining process. And, we couldn't afford to upgrade the model. So, the other way that this locked in was architecture. And, we built this app in late 2024. When workflows were kind of the gold standard if you wanted good production results. And, the AI world moves very fast. And, we couldn't adapt to that because we were so locked into this. Just trying to keep it running. And, we couldn't take advantage of the new architectures. And, improve performance that way. So, earlier this year I had an aha moment. We started using Cloud Code for our coding tasks. And, I noticed that we never needed to change the model depending on what task we're using. We just changed the skill, the resources that we passed it, the context. You drop in the better context, you get better results. And, I thought, why can't we do this with our messaging app? This was obviously difficult for me to admit. Because, I was the champion for fine tuning. And, luckily we were able to piggyback on a project that was already happening. And, so, we migrated our workflow approach to a series of skills, tools, and resources that the skills could load into. Or, load up. And, get that context. And, so, we pushed this as one of our first production tests of our new agentic framework that was being built already. Now, I want to compare the process before and after our rebuild. Before, we already went through the kind of the training cycle. But, there was this triage cycle beforehand where we needed to make sure that we had reached a critical mass of problems before we would even attempt to fine tune again to improve everything. Like I said, this took about a week. So, it was a long process, costly. After the rebuild, it was a simple process of you find a problem. You adjust the system prompt or the skill that was affected. We validated performance on a curated set that we had been collecting over the time that this was in production. We iterate a few times. We deploy that simply by uploading MD files to an S3 bucket. This whole process from discovering a problem to deploying the fix, we reduced down to less than an hour. So, it extremely improved all of this. And, we could be far more reactive, give our customers way better performance or better experience there. Now, I'll be honest, it did cost us a little bit more per message. We were using better models, so the API costs were a little higher. But, accuracy went way up. I said before that accuracy was the key to getting all of this right. And, we did that. Accuracy was far better with this than it ever was with fine tuning. Next, like I said, we reduced our fix process from days down to minutes. Next, we were able to unfreeze our model and finally get that freedom from a vendor that we never had with fine tuning. Our agentic framework was built model agnostic, so we could use OpenAI. We can use Anthropic. We can use any other model that we want. The important part is the context that we're providing to that model. And then, lastly, while it cost us a little more per message, the total cost went down. Because, we were spending far less time trying to keep it up and running and fine tuning to keep it working properly. So, before you fine tune, I'd ask you, can you cross your reason off of this list? So, I thought we would get better accuracy. The rebuild beat the fine tune model. I thought we would get lower cost of the volume we were doing. I was looking at the wrong cost. We paid more per message, but the total cost ended up going down with our rebuild. Lower latency, we did see marginal gains on these smaller models, but they were so small that in practice, it really didn't make any difference. And then, maybe you've got a narrow or structured task. Our textbook case still became tech debt. And lastly, vendor control. It's not as simple as just plugging the data in. The other two situations where you might have privacy and data control, or you need some offline solution. I would say, these are the situations where a fine tuned model may be useful, but you need to be cautious. There are other solutions out there, but you need to make sure that it's not causing issues in the long run. So, finally, fine tune only when you literally cannot call a front tuner model. And even then, your decision still has to beat the tax. Thank you. That yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay