Why Your Enterprise Tech Stack Isn’t Ready for AI Agents — Christopher Lovejoy & Saul Howard
Description
The proof of concept works. It hits the accuracy targets, it is fast, it is cheap, and the room is happy. Then someone from compliance raises a hand and asks to see the audit trail, and the whole thing stops. Christopher Lovejoy and Saul Howard have watched that meeting happen repeatedly, and their point is that an audit trail is not a developer log. Under the frameworks enterprises actually answer to, it is a complete record of every action an agent took, every place it touched data, and the authorization behind each one, durable enough to stand up as a chain of evidence if the decision were ever examined in court. Their answer is to take the constraints seriously first and rebuild toward the accuracy afterwards, rather than bolting requirements onto a demo. An immutable append only event log makes auditability fall out of the storage model instead of being reconstructed later, at the cost of harder reads. Patient data lives in schema driven object storage alongside that log rather than inside it, so the events hold only references, which lets engineers debug what an agent did without being exposed to the health data itself, and gives a natural place to enforce zero trust and constrain prompt injection. Escalation works because humans and models are both treated as agents, so any action either can take, the other can take too. Evaluation then emerges from those three primitives rather than being attached to the side, including on production data that never leaves the customer's environment. Speaker info: - https://x.com/ChrisLovejoy_ - https://www.chrislovejoy.me - https://x.com/saulhoward - https://linkedin.com/in/saulhoward Timestamps: 0:00 - Why healthcare is hard, and what transfers to other regulated work 1:57 - The enterprise proof of concept 2:49 - What the buildout actually connects to 3:39 - Everyone assumes the hard part is done 4:28 - The questions that arrive the next day 5:19 - An audit trail is not a developer log 7:03 - The immutable event log, an
Summary
Generated by gpt-5.6-terraAt-a-Glance
- Verdict: Watch fully
- Core thesis: Enterprise agent systems fail in production when teams treat auditability, data controls, human escalation, and evaluation as POC add-ons rather than foundational architectural primitives.
- Why it matters: The talk gives a concrete control-plane architecture for regulated, high-consequence agents: an immutable action ledger, separately governed object storage, human-agent interchangeability, and replayable privacy-preserving evals.
- Best use: Use it to pressure-test an agent platform or enterprise deployment design before committing to a promising but non-production-ready POC architecture.
Executive Summary
Christopher Lovejoy of Anthropic and Saul Howard of healthcare-agent company Anterior argue that a successful enterprise agent POC often creates a false sense that the hard problem is solved. A workflow may meet accuracy, latency, and cost targets, yet still be unable to answer the production questions that security, compliance, clinical, and operational stakeholders will ask: what did the agent do, under what authorization, what sensitive data did it see, who can approve consequential actions, and how will performance be evaluated over time?
Their proposed answer is to build around four mutually reinforcing primitives from the outset. First, use an append-only immutable event ledger as the unified system record of every action, authorization, and data reference. Second, keep sensitive domain data in schema-driven, immutable object storage adjacent to orchestration rather than passing raw data freely through agent processes. Third, define humans and LLMs as equivalent types of agents so any step can be escalated to a human without breaking downstream workflows. Fourth, derive evaluation capabilities from those primitives through replay, human-versus-agent comparison, and in-environment testing on protected production data.
The core architectural trade-off is explicit: event sourcing makes writes and audits easy but requires projections, caching, or snapshots to make reads efficient. The speakers consider that trade worthwhile in healthcare and similarly regulated sectors because it enables forensic reconstruction, changing interpretations of historical records, least-privilege data access, and evaluation against real operating conditions.
The practical warning is not to productionize by progressively bolting security, evals, audit logs, and approval mechanisms onto a point solution. That approach produces brittle systems that do not generalize across workflows. Instead, begin with the constraints of scaled production deployment and rebuild upward toward POC-level task accuracy.
Key Takeaways
- Claim: Agent POC success is not evidence that an enterprise system is production-ready; the difficult work begins when the organization requires governance around autonomous actions. | Evidence: The speakers describe a healthcare administrative-workflow POC that hits performance, speed, and cost targets, only for stakeholders to ask for a complete audit trail, PHI handling boundaries, clinician approvals, defenses against untrusted data, ongoing performance monitoring, and integrations such as Epic and Salesforce. | Implication: Treat production requirements as initial architecture inputs and POC acceptance criteria, not as a post-demo hardening backlog. | Caveat: The examples are healthcare-specific, but the speakers explicitly extend the pattern to finance, defense, government, and other process-heavy regulated environments.
- Claim: An append-only immutable transaction or event ledger should be the unified source of truth for agent actions when auditability is a first-class requirement. | Evidence: They distinguish enterprise audit trails under SOC 2, HITRUST, and HIPAA from ordinary developer logs: the record must capture every action, data access, and authorization, sufficient to justify a decision in a legal setting. A timestamped append-only ledger makes it possible to reconstruct system state at any point. | Implication: Design agent execution around durable, queryable action events rather than treating observability logs as the system of record. | Caveat: Event-sourced designs make writes simple but make reads harder because views must be reconstructed; caching and snapshots are needed to make operational reads practical.
- Claim: Separate orchestration events from sensitive data by storing protected records as immutable, schema-driven objects and placing references to them in the action ledger. | Evidence: Healthcare data can be structured or unstructured, large, subject to RBAC, and prohibited from leaving a customer's on-premises VPC. The proposed design records references to protected data blobs in events rather than embedding PHI throughout agent workflows. | Implication: Build the agent control plane so developers can inspect workflow shape and execution behavior without routinely receiving raw customer or regulated data. | Caveat: The architecture requires disciplined schema design and access-control implementation; separation alone does not establish least privilege.
- Claim: Zero-trust, point-of-use access to object storage can reduce data-exfiltration and prompt-injection risk. | Evidence: The speakers propose that agents carry tokens and retrieve data only at the moment it is needed, instead of allowing sensitive data to flow broadly through processes. They position this segregation as a mitigation for the "lethal trifecta," where an agent combines access to sensitive data with untrusted instructions and outbound capability. | Implication: Constrain agent capabilities and data grants per step, and avoid architectures in which one compromised agent context accumulates unrestricted access to unrelated information. | Caveat: This is described as an architectural mitigation rather than a complete prompt-injection solution; the talk does not specify token scopes, egress controls, or policy enforcement details.
- Claim: Human escalation works more reliably when humans and LLMs share a common agent interface, allowing either to perform any workflow action. | Evidence: Escalations are inherently dynamic: an LLM may be uncertain or a treatment may exceed a rule-based approval threshold. By treating both people and models as agents, a human can execute an escalated step and downstream workflow stages do not need to know which performed upstream work. | Implication: Define action contracts and workflow state independently of the executor, then add executor-specific context renderers rather than creating a separate manual exception process. | Caveat: Humans cannot consume the same volume or format of context as an LLM, so the shared context needs separate mappings into prompts for models and usable interfaces for people.
- Claim: The same primitives that support governance can make evaluations a native system property rather than an offline testing layer. | Evidence: An immutable ledger enables exact replay from a historical system state while changing a prompt, model, or code path; human-agent equivalence allows their outputs on the same task to be compared; protected object storage permits evaluations against production data within a customer's environment without exporting that data. | Implication: Prioritize replayable execution records and in-situ evaluation capability so model, prompt, and workflow changes can be tested against realistic data rather than a stale, unrepresentative offline dataset. | Caveat: Replay does not eliminate LLM nondeterminism, and the speakers do not detail statistical evaluation methodology or thresholds for deployment decisions.
Detailed Brief
Why event-sourced records suit evolving healthcare and agent workflows
- Claims: A historical event record is more useful than a single mutable current-state representation when interpretations can change as additional events arrive.; Read models should be treated as computed, potentially ephemeral projections of the canonical event stream.
- Evidence: The speakers note that later healthcare events can change the interpretation of a patient's earlier journey, making it valuable to regenerate views from raw history.; They identify caching and snapshots as patterns for mitigating the cost of reconstructing state from the event log.
- Caveats: The talk does not provide implementation guidance for event schema versioning, retention policies, deletion obligations, or the operational cost of maintaining projections.
- Implications: For long-running agent workflows, preserve the raw sequence of decisions and inputs while allowing reporting, operational dashboards, and case views to evolve independently.
Architecture choice is fundamentally a decision about what must be easy
- Claims: The speakers frame architecture as choosing constraints that make the organization's highest-priority properties simple by design, while consciously accepting other complexity.; Relevant pre-AI patterns already exist in finance, defense, and large-scale enterprise software; AI systems need to combine rather than discard them.
- Evidence: Their central comparison is between a high-performing point solution hardened later and a production-constrained foundation rebuilt toward the same POC accuracy.
- Caveats: The talk is a principles-level architecture presentation, not a vendor comparison, reference implementation, or quantified case study.
- Implications: Avoid evaluating agent platforms solely on model quality or demo metrics; evaluate whether their core data and execution model can support compliance, approvals, and continuous improvement.
Notable Concepts & Terms
- Immutable event ledger / event sourcing: An append-only, timestamped record of system actions that becomes the canonical source for audits, replay, and historical reconstruction.
- Schema-driven object storage: A separate store for sensitive, potentially large structured or unstructured records, allowing the orchestration layer to work mainly with governed references.
- Orchestration-adjacent storage: Keeping protected data accessible to the workflow under explicit controls without freely copying it into agent runtime contexts.
- Human-agent equivalency: A workflow abstraction in which a human and an LLM can execute the same action contract, enabling dynamic approval and escalation.
- Privacy-preserving evals: Evaluation methods that replay recorded execution and test against data in the customer's environment without exporting sensitive records.
- Lethal trifecta: The security condition highlighted here in which an agent may combine sensitive-data access with untrusted inputs and capabilities that enable harmful action or exfiltration.
- Computed projections: Read-oriented views derived from the event log, which can be regenerated as business interpretation or reporting requirements change.
Operator Notes / Why Ken Should Care
- Add a production-readiness gate to agent POCs: require an action-level audit model, data lineage, authorization capture, escalation path, and evaluation plan before declaring a pilot expandable.
- For any OpenClaw or agent-control-plane workflow handling sensitive information, separate execution events from data payloads and enforce scoped, just-in-time retrieval rather than passing raw context through chains.
- Standardize tool/action interfaces so a human reviewer can take over any consequential agent step without creating a workflow fork.
- Require replay tests for prompt, model, and workflow changes, with a plan to run tests in customer-controlled environments when data cannot leave the boundary.
- Investigate event-sourcing operational requirements before adopting the pattern broadly: projection latency, snapshot strategy, storage retention, event-schema evolution, and regulatory deletion obligations.
Source/Metadata
- Title: Why Your Enterprise Tech Stack Isn’t Ready for AI Agents — Christopher Lovejoy & Saul Howard
- Transcript words: 3979
- Duration seconds: 1155
- Timestamp note: No timestamps or chapter markers were present in the supplied transcript; the final evals-and-conclusion section is duplicated.
Transcript
Hello, everybody. My name is Christopher Lovejoy, and I'm a member of technical staff at Anthropic. I work as a full-deployed engineer, so I embed within enterprise organizations and help them get value from using AI agents. I previously worked at Anterior with Saul. Hi, everybody. I'm Saul. I'm VP of engineering at Anterior. We're a New York-based company selling AI, agentic AI, to U.S. health insurance companies. Chris and I have spent a lot of time building in enterprise, and in healthcare enterprises particularly. Healthcare is a very challenging place to develop and deploy AI. Healthcare is so challenging because of the requirements around process and compliance, the regulatory requirements that are so important. Also because of the direct, real impact that your work has on people's lives, which is, of course, also what makes it so rewarding. I think a lot of the learnings you can take from working in enterprise for healthcare, you can take to enterprise in other regulated industries, like finance, defense, government work, anywhere where process is so important and has to be followed. In this talk, we're going to talk about some of the learnings that we've had, and specifically we're going to talk about why enterprise tech stacks aren't ready for AI agents, and some of the primitives that we've built in the past in order to unlock them. To make this concrete, let's start by considering a scenario that might be familiar to many of you, which is the enterprise proof of concept, the enterprise POC. Let's say we have identified a customer that we want to serve, and we've identified a priority use case with them. Obviously, we're on the healthcare track here. Let's consider a large health system and a use case that is some sort of administrative healthcare workflow. You work with them, you scope out a POC, you define the metrics that you are going to care about and benchmark yourselves on. You allocate two engineers, you spend four weeks building it, and the actual buildout might look a little something like this. An enterprise stack is very complicated. It's much more than we're showing here, but generally you can have an application layer, a control plane layer, and the data plane. For your POC, you're going to need some access to the model provider as well. Your POC is going to need access to data across all of these different planes. It may be some in the data lake, some directly from the application layer, for example. So you're going to deploy it something like this. It's going to connect to all these different places. There's going to be some offline data pulling. There's going to be maybe some online. Generally, you'll get access to the data and push toward the results. Things go well. You get great results. The AI performs as you expected. You hit the performance metrics. It's fast. It's relatively cheap. You hold a meeting. You present this to the relevant stakeholders, and everyone seems pretty happy. Your chief of finance in the company is very excited and wants to understand what's going to be the impact on the budget for next year. Your chief medical officer is excited to tell his colleagues how accurate his AI is. The head of sales asks, okay, when can we put Powered by AI on the websites? But the problem is that everyone here is assuming that the hard part is done, that the AI was the challenging part. Actually, as we know, often getting things into production is really where the challenge lies. To get a bit more specific on what that challenge looks like, you hold a meeting the next day. You bring in the relevant stakeholders to discuss productionizing this proof-of-concept application. Somebody raises their hand and says, can I see the audit trail for this? For us, for compliance, it's critical that we can see every step, every action that the agent takes, every piece of data that it accesses. Can you give that to me? And you realize that, with the way things have been implemented in the initial POC, without these true integrations, that's going to be quite challenging. Then somebody else pops up with some other questions. Somebody asks, okay, how is sensitive data being handled here? How is that being passed to the agents? We have a very strict boundary around where our data can go and where it can't go. Is this respecting that? How does that look? Then your chief medical officer says, okay, and who's approving the decisions here? Because we know in certain scenarios we have to escalate to a clinician who will then approve or not agree with what the agent is saying. So how does that happen? What's the mechanism for that? Over the course of the meetings, you can imagine you get more and more questions. Can untrusted data manipulate the model? How do we know that the agent continues to perform well? How do we deal with integrations? How do we connect to Epic, to Salesforce, to the other applications that we care about? For the purposes of this talk, we're going to focus on these four, the highlighted ones. For the other two, feel free to come and chat to me and Saul about these later. We're very happy to talk. But in the interest of time, we'll stay focused. Let's start with this one about the audit trail. This is a question you're guaranteed to get from the security team. They're going to want to see an audit trail. For programmers, an audit trail sounds very much like a typical developer log that you might have in Datadog. Surely it's a similar kind of thing. But for security frameworks that exist in the real enterprise world, like SOC 2, HITRUST, HIPAA, an audit trail is a bit more than that. It has to contain a complete record of absolutely every action that the agent took. It has to contain all of the places where the agent accessed data, all of the authorization by which the agent did something. It's this complete record in a much more fundamental way. One way of thinking about it is in a legal sense. Say our agent's decisions came up in a court of law. Could we show a justifiable chain of evidence for why the particular actions were taken by a decision? And that's something that could easily happen within the healthcare context, for example. When I think about architecting systems like this, I think often about what do I want to make easy? When I'm choosing my constraints, I'm saying, okay, these are the things I want my system to make easy, and let that drive the trade-offs that I'm going to make. A particular pattern that is used in lots of different industries, for example in finance, is a transaction log, an immutable record of events that store all of the transactions that happen throughout the system. And this is append-only, timestamp log. It's complete. This is your source of truth for all of the data of the system, and it's unified. There is only one source of truth across all of the different agents that you might have running in parallel, for example. Architecting this way, making this trade-off, means that auditability becomes trivial. It falls out of your data storage paradigm that you've chosen. It's impossible not to be able to roll back time and see exactly the state of the system at a particular point in time and be able to provide that as an audit trail for what happened at each point in time. Of course, these are trade-offs. So what's a trade-off you're making here? I think we could say that for this kind of event logging, or sometimes called event sourcing pattern, writes become very easy, so you just drop an event. Reads become more difficult because you have to read through all of the events in order to reconstruct a view of what happened, and there are patterns like caching and snapshots that you can bring to make that simpler, but there is more effort there. Although I have seen in the healthcare context that, actually, you're going to want different interpretations of the raw data that your agents recorded after the fact. For example, it might be that more events happened, and that changes the interpretation of the healthcare journey, and you want a different view of the source of truth at that particular time. This pattern makes that easy because all of your views of the data are ephemeral computed projections of the event log. Okay. Next. The compliance officer comes and is asking, how is the sensitive data passed around the system? What's the life cycle of data within our system? Within a healthcare context, as we all know, data means a lot. It's PHI, protected or personal health information. It has legal restrictions around it, not just HIPAA, but other legal restrictions about the use of people's data. You cannot have your agent, just as you cannot have humans, accessing and reading and utilizing healthcare data that they don't absolutely have a necessity to use at that point in time for that particular journey. Again, architecturally, when I think about how am I storing data within a particular system, I would like to think, what is the shape of the data? What kind of characteristics does the data have? For healthcare data, that might be that it's very complicated. It doesn't follow strict hierarchical relationships. It's sometimes unstructured and it's sometimes structured. It could be very large. For example, healthcare data, one piece of healthcare data, can easily be over a megabyte in size, or much more than that. It has strict access controls, as we've been saying. The RBAC comes into play, both for humans and then for agents downstream of that. It may even be, I've seen customers where they're not willing to have their healthcare data leave their own environment, leave their on-prem VPC, for example. So we have tangential access to their data. An architectural paradigm I might go to is object storage. Schema-driven object storage, I think, is a good fit for this. It matches well with the choice of using event logging because you can separate the two. The events we talked about as the record of what the agent is doing at any particular time only contain references to the schema-driven blobs that are the storage of the actual healthcare data itself. It's important, therefore, that the healthcare data is stored immutably, again, so that you can always go back in time and reconstruct what data the agent had access to at that particular point in time. This separation of events for what happened and object storage for the data that was used at that particular point in time has some very useful benefits. For example, with a system like this, it's possible for developers to go back and debug and have observability over what happened, what particular steps the agent took, why it did that, and retrace the agent's steps without having access to the personal health information itself. Because of the schema-driven system, they can see the shape of that data, but they can't and, to be honest, often won't be able to be given access to that healthcare data. So you can separate out observability and orchestration and instrumentation from the healthcare data itself. This then has another benefit, which is zero trust. The object storage becomes a place where you can apply zero trust principles. Your agents can bear tokens and use those tokens to access the data at the point of use and not allow data to flow around the system as it likes. This then leads into a mitigation for prompt injection, for the lethal trifecta. The way I think about the lethal trifecta is, can I solve for the constraint if I have an agent at point A with access to this data? Is it possible within my architecture for the agent to be also accessing data over here? And zero trust principles, tokens are borne by the agents, and object storage segregated from the event stream that has your orchestration logic gives you a place to be able to solve for that constraint. It won't be possible for the agent to access data within the same process that you've given it the previous data. Okay, so then it comes to how do you handle escalation? In many scenarios, you will want to be able to escalate the decision that an agent makes, or an action that an agent makes, to a human. But one of the challenges here is that this is quite dynamic. You don't know in advance when exactly the agent is going to escalate. It could be that you're asking the AI to escalate when it's not sure. It could be that you define some sort of rules in your system. Maybe in a medical context, the treatment's going above a certain threshold means that it needs to be escalated for approval. But this makes it very challenging because of this inability to predict. A second challenge is also that humans and LLMs ultimately process context differently. LLMs will have no problem if you give them massive, massive amounts of text, but for humans, that's not the case. What we've seen is that one pattern that can work very well here is if in your platform you enforce a wider definition of agent, which encompasses both LLMs and humans, then you can make it such that any action that can be taken by an LLM could also be taken by a human. This is helpful because at any point in the chain of actions that your agent is taking, it can escalate to a human, the human could perform that action, and then any step downstream doesn't care about whether it was a human or an LLM that did those actions upstream. On the second point around the context, what this also makes much easier is that you can define methods that take the context, which has some kind of shared definition of context, which is irrespective of whether it's a human or an LLM that's going to be accessing it. You can take those methods to then map into something that's agent-friendly, like a prompt, or into something that's more human-friendly, for example, a UI. Then on this fourth and final question that we're going to talk about, evals. Obviously, we hear a lot about evals. We know that evals can be very helpful, that often they drive decision-making about the types of model you want to use, the type of approach you might want to use within your product. But we also know that evals can be pretty hard, and there are various factors here. We know that LLMs are not deterministic, so it can be quite tricky to pin down the precise change that led to some sort of change in output. We also know that the data that you might put in an offline data set might not necessarily represent production data, and it could be that maybe you sampled from data, but actually that sample isn't truly representative. Then you also have drift of data over time, so maybe your offline data set is now out of date. What we found is that these three primitives that we've described so far in the talk actually give you effective privacy-preserving evals almost as a byproduct, without needing to bolt something onto the side of your architecture. To make that more concrete, the immutable ledger means that you can replay your actions. You can go back to any particular time in this sequence of events, you can see the complete state of the system at that point in time, and if you wanted to, you could then make very specific tweaks. You could tweak a prompt, you could tweak a model, you could tweak the code, and you can see the exact direct impact of that because you have all of that context. Secondly, you have this human-agent equivalency, which means that for any time, you could get both the agent, the LLM agent, and the human to perform it, and your difference is your eval. That gives you the eval scores. Finally, what the object storage enables you to do is to actually run these evals on production data, including inside your customer's environment, without actually ever exposing that data, and you can get your eval results without the sensitive data ever needing to come to where your agent's performing the work. Right. We've gone through four architectural principles that we found useful for building in healthcare and more generally in regulated environments for enterprise: the immutable ledger of actions, the orchestration-adjacent object storage, the human-agent equivalency, and the way that with these three principles, evals can emerge as a first-class property of the system rather than as something you attach onto the side. I think one of the matters here is that I like to think about architecture as taking your constraints very seriously and thinking about what you want to be simple within the system and then choosing the trade-offs for that. Of course, alongside that, some things will become hard, but it's the things that are simple that are most important to you. There are patterns that already exist across enterprises that solve for a lot of these things. Sure, with AI, we need to combine them in new, sometimes radical ways and bring in other pieces. But there are patterns that have worked very well within finance, within defense, within big tech, that can be applied to this kind of system architecture. I'd say the takeaway is that where I've seen it go wrong is taking that initial POC, that point solution that showed so much promise and that showed the high accuracy, for example, and then trying to build up from it, strapping on the enterprise requirements as you come across them. Okay, we need evals, we need security, we need auditability, and bolting these on as additions to the foundations of the POC. You end up with something very brittle, something very hard to externalize and to generalize across different use cases. But where I've seen it go well is if you take the constraints of a production-ready, scaled enterprise system seriously from the beginning and treat those as the architectural principles that you're going to build everything upon, and then build back up toward that POC accuracy using your new primitives. Thank you for your attention. Thank you. applause applause applause applause applause applause applause applause applause Thank you. And on the second point around the context, what this also makes much easier is that you can define methods that take the context, which has some kind of shared definition of context, which is irrespective of whether it's a human or an LLM that's going to be accessing it. And you can take those methods to then map into something that's agent-friendly, like a prompt, or into something that's more human-friendly, for example, a UI. And then on this fourth and final question that we're going to talk about, evals, obviously, you know, we hear a lot about evals. We know that evals can be very helpful, that often they drive decision-making about the types of model you want to use, the type of approach you might want to use within your product. But we also know that evals can be pretty hard, and there's various factors here. We know that LLMs are not deterministic, so it can be quite tricky to pin down the precise change that led to some sort of change in output. We also know that the data that you might put in an offline data set might not necessarily represent production data, and it could be that maybe you sampled from data, but actually that sample isn't truly representative, and then you also have drift of data over time, so maybe your offline data set is now out of date. And what we found is that these three primitives that we've described so far in the talk actually give you effective privacy-preserving evals almost as a byproduct without needing to kind of bolt something onto the side of your architecture. So to make that more concrete, so the invisible ledger, what this means is that you can replay your actions. So you can go back to any particular time, you know, in this kind of sequence of events, you can see the complete state of the system at that point in time, and if you wanted to, you could then make very specific tweaks. So you could tweak a prompt, you could tweak a model, you could tweak the code, and you can see the exact direct impact of that because you have all of that context. Secondly, you have this human agent equivalency, which means that for any time, you could get both the agent, the LLM agent, and the human to perform it, and your difference is your eval. That gives you the eval scores. And then finally, what the object storage enables you to do is to actually run these evals on production data, including inside your customer's environment, without actually ever exposing that data, and you can get your eval results without the sensitive data ever needing to come to where your agent's performing the work. Right. So we've gone through four architectural principles that we found useful for building in healthcare and more generally in regulated environments for enterprise. The immutable ledger of actions, the orchestration adjacent object storage, the human agent equivalency, and the way that with these three principles, evals can emerge as a first-class property of the system rather than as something you attach onto the side. I think one of the matters here is that I like to think about architecture as taking your constraints very seriously and thinking about what you want to be simple within the system and then choosing the trade-offs for that. And, of course, alongside that, some things will become hard, but it's the things that are simple that are most important to you. And that there are patterns that already exist across enterprises that solve for a lot of these things. And, sure, with AI, we need to combine them in new, sometimes radical ways and bring in other pieces. But there are patterns that have worked very well within finance, within defense, within big tech that can be applied to this kind of system architecture. And I'd say the takeaway is that where I've seen it go wrong is taking that initial POC, that point solution that showed so much promise and that showed the high accuracy, for example, and then trying to build up from it, strapping on the enterprise requirements as you come across them. Okay, we need evals, we need security, we need auditability, and bolting these on as additions to the foundations of the POC. You end up with something very brittle, something very hard to externalize and to generalize across different use cases. But where I've seen it go well is if you take the constraints of a production-ready, scaled enterprise system seriously from the beginning and treat those as the architectural principles that you're going to build everything upon and then build back up towards that POC accuracy using your new primitives. Thank you for your attention. Thank you. applause applause applause applause applause applause applause applause applause Thank you.