AI Engineer

How do you diffuse AI into the real world? — Varun Shenoy, Long Lake

2079 summary words 9 min summary Watch video

Start with the signal

9 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: AI will diffuse into real-world services not through a turnkey co-worker product, but by progressively earning autonomy, capturing operational traces, and tightly co-designing software with frontline people and processes.
  • Why it matters: This is a concrete operator-owner perspective on the hard part of enterprise AI: converting model capability into reliable, adopted workflows amid exceptions, legacy systems, and entrenched human habits.
  • Best use: Use it as a strategic framework for designing agent deployment, evaluation, feedback loops, and adoption programs in operationally complex businesses.

Executive Summary

Varun Shenoy argues that the gap between impressive AI demos and adoption in a 200-person services company is normal, not evidence that the technology has failed. He compares AI to electricity: broad economic impact required factories to replace equipment, redesign workflows, and train workers, rather than merely installing a new power source. Long Lake's claimed advantage is that it owns or partners with services businesses, so it bears the operational consequence when AI fails rather than acting as an external software vendor.

His first operating lesson is to treat autonomy as a progression: RAG-style copilot, synchronous agent, asynchronous agent, long-running agent, and ultimately AI co-worker. Companies should not leap to the final stage. They must earn greater autonomy through model reliability and user trust, especially in services work where processes are serial, context-heavy, and full of exceptions. The unresolved design frontier is how to make traditionally serial knowledge work safely parallelizable, analogous to running multiple coding-agent sandboxes.

The second lesson is that valuable enterprise AI data comes from work traces rather than the public internet. By having agents collaborate on live tasks, Long Lake collects tool calls, errors, corrections, and final outcomes; it uses these traces to build automatically scored evaluations, capture implicit feedback through differences between agent output and submitted work, post-train models on domain data, and improve customized agents over time.

Finally, Shenoy treats adoption and continual learning as one reinforcing loop rather than separate research and deployment functions. Usage produces feedback and better agents; better agents produce more usage. The initial adoption problem, however, requires "extreme software-service co-design": embedding tools in existing systems such as ERP, Excel, email, and design software while also working in person with operators to understand their actual processes and lower the behavioral cost of change.

Key Takeaways

  • Claim: The central challenge is AI diffusion: converting improving models into economically valuable work inside real businesses, not producing another compelling demo. | Evidence: Shenoy contrasts flight-booking, ticket-resolution, and code-generation demos with a 200-person property-management firm where he says AI has changed little. He analogizes this to electricity, invented in the 1880s but requiring factory redesign, replacement equipment, and training before broad adoption; he cites a Ford electrified assembly line in 1924. | Implication: Ken should evaluate AI opportunities as organizational redesign and workflow deployment problems, with model quality as only one component of value creation. | Caveat: The electricity comparison is a framing device, not evidence that AI will necessarily follow the same adoption timeline or economic pattern.
  • Claim: Agent autonomy should be deployed as a ladder, from copilot to synchronous and asynchronous agents, then long-running agents and eventually proactive co-workers. | Evidence: The speaker defines a copilot as a fast RAG chatbot; a synchronous agent as a real-time tool-using system that may run for one to five minutes; an asynchronous agent as background work that can be triggered by events or job queues; and a long-running agent as operating for hours, days, weeks, or months. | Implication: Design agent systems with graduated permissions, clear handoffs, and progressively stronger triggers rather than launching an ostensibly autonomous co-worker into uncontrolled production workflows. | Caveat: He explicitly argues that teams must earn the right to add autonomy because models may not yet handle every task reliably and users need iterative exposure to the technology.
  • Claim: The key unsolved product problem in services is making serial, exception-heavy knowledge work parallelizable through asynchronous agents. | Evidence: Coding agents can be wrapped in sandboxes to build, test, and return pull requests, while engineers are accustomed to launching many jobs in parallel. In contrast, services workers generally process work sequentially, such as handling inbox emails one by one. Shenoy asks what the equivalent of code-agent forking is for property management, architecture, and other service operations. | Implication: Look for decomposition opportunities where parallel agents can gather evidence, draft outputs, reconcile records, or surface exceptions, while preserving a controlled human or workflow layer for sequencing and final resolution. | Caveat: He offers the problem framing rather than a proven general solution; he also says form factors will differ substantially by industry.
  • Claim: Operational traces from real work are a strategic data asset because they create domain-specific evaluations, feedback signals, and training material that frontier models lack. | Evidence: Long Lake collects tool calls, failures, paper cuts, user feedback, and outcome data from work such as closing books with missing receipts, coordinating roof repairs, or scoping a construction project. It cites implicit feedback from the diff between AI-generated data and what was ultimately submitted, alongside explicit thumbs-up/down feedback. | Implication: Prioritize agent deployments where final-state data, corrections, and quality signals can be captured automatically, then convert those traces into regression tests and improvement loops. | Caveat: Ground truth is strongest for tasks with observable outcomes, such as whether a roof was repaired or books were closed; many knowledge-work judgments may be less cleanly measurable.
  • Claim: In real work, exceptions are not edge cases; handling accumulated small failures is the job itself. | Evidence: Shenoy contrasts a clean, visible path to success in standard LLM task representations with real operations full of hills, ravines, and "death by a thousand paper cuts." He identifies company-, user-, and client-level variation as core customization requirements in services businesses. | Implication: Measure agent performance on messy end-to-end workflows and exception recovery, not just happy-path task completion or isolated benchmark scores. | Caveat: High customization can improve fit but can also create an expensive, fragmented deployment model if common abstractions and reusable controls are not maintained.
  • Claim: Continual learning and user enablement are one operating loop: adoption generates the feedback needed to improve agents, and agent improvement is required to sustain adoption. | Evidence: Shenoy describes a snowball effect in which greater usage creates continual learning, which produces a better agent and then more usage. He argues that giving a coding tool to an entire enterprise does not create usage automatically, particularly where employees have performed the same process for decades. | Implication: Treat rollout, instrumentation, product improvement, and change management as a single accountable operating system with explicit adoption targets, rather than splitting them between disconnected research and customer-success teams. | Caveat: This loop has a cold-start problem: initial usage must be actively created before feedback and performance gains can compound.
  • Claim: Enterprise AI adoption requires extreme software-service co-design, including both embedding into existing tools and direct, in-person engagement with operators. | Evidence: Examples include building natively into Excel, ERP systems, 3D design software, Outlook, or Gmail; conducting lunch-and-learns; attending company conferences; observing day-to-day work; and providing one-on-one training. Shenoy argues this cannot be done adequately through Zoom or support tickets alone. | Implication: For high-stakes operational deployments, budget for field discovery, workflow embedding, and hands-on enablement rather than assuming self-serve software distribution will drive behavior change. | Caveat: The assertion that co-design is only possible "under the same roof" reflects Long Lake's operator-owner model; external vendors may still achieve deep partnerships, albeit with weaker access and incentives.

Detailed Brief

Long Lake's operator-owner deployment model

  • Claims: Long Lake positions itself as distinct from a conventional AI vendor because it acquires or partners with the operating businesses where it deploys technology.; This ownership model is intended to align product development with business outcomes: if automation fails, the consequences remain with Long Lake rather than being transferred to a customer.; The company frames this as a way to obtain direct access to real workflows, data, users, and operational feedback.
  • Evidence: Shenoy says Long Lake has acquired 35 businesses across areas including HR, property management, architecture, and HR services.; He says its approximately 40-person team spans technology, finance, and operations, with more than half focused on technology, product, data, and field deployment.; He references Long Lake's announced $6.3 billion take-private of American Express Global Business Travel as evidence of operating at substantial scale.
  • Caveats: The transcript presents the scale, financing, acquisitions, and transaction as company claims; it does not independently substantiate their status or operational results.; Owning businesses can improve access and incentives but is capital-intensive and may not be reproducible for a software-first organization.
  • Implications: A non-owner deployment strategy should compensate for weaker access by securing deep workflow access, shared outcome metrics, data rights, and recurring onsite engagement.; The more operational accountability a builder accepts, the more credible its feedback loops and claims of business impact can become.

Representing knowledge work as code

  • Claims: Because frontier models are extensively trained on code and perform strongly at programming, Shenoy proposes exploiting that strength rather than waiting for models to natively master every services domain.; The suggested direction is to represent portions of knowledge work as code-like structures that agents can execute, test, fork, and inspect.
  • Evidence: He asks whether knowledge work can be represented as code so coding-agent capabilities can be applied to services work.; His coding example relies on a sandboxed agent that can build and test before returning a pull request, which provides a concrete execution and review pattern.
  • Caveats: The talk does not specify a representation, compiler, task language, or validation architecture for translating services work into code-like workflows.; Over-formalizing work can miss tacit judgment, customer-specific preferences, and situational exceptions.
  • Implications: Explore workflow representations that make state, tools, constraints, approvals, tests, and outcomes machine-readable without eliminating human judgment where it is genuinely needed.

Notable Concepts & Terms

  • AI diffusion: The process of getting AI models adopted in real organizations and applied to economically relevant tasks; Shenoy presents it as the major challenge beyond model capability.
  • Operator-owner model: Long Lake's approach of owning or closely partnering with services businesses so it directly bears the result of AI deployment and can access live workflows.
  • Autonomy ladder: A staged progression from copilot to synchronous agent, asynchronous agent, long-running agent, and AI co-worker, intended to match autonomy to proven reliability and user trust.
  • Jagged frontier: The idea that model competence is uneven across tasks; agents can be excellent at coding while still weak on particular operational tasks, making blanket autonomy assumptions dangerous.
  • Rich traces: Detailed records of live agent-human work, including tool calls, errors, corrections, and final outcomes, used to build evaluations and improve systems.
  • Implicit feedback: Quality signals inferred from operational outcomes, such as the difference between an AI-generated record and the version ultimately submitted, rather than only user ratings.
  • Hill-climbing benchmarks: Operational evaluations that improve agents iteratively and then become regression tests, preventing future versions from losing previously achieved capabilities.
  • Extreme software-service co-design: Designing AI products alongside the people, processes, systems, and adoption motions of the service operation, modeled after hardware-software co-design.

Operator Notes / Why Ken Should Care

  • Define an autonomy maturity model for each agent workflow, including allowed triggers, run duration, tool permissions, approval gates, escalation conditions, and rollback paths.
  • Instrument production workflows to retain task inputs, tool traces, intermediate artifacts, human edits, final submissions, and outcome labels; make these usable as regression evaluations rather than passive logs.
  • Select one serial operational workflow and test deliberate parallelization: use multiple agents for evidence collection, classification, drafting, reconciliation, or alternative solution generation before a controlled merge step.
  • Make adoption a jointly owned metric across product, deployment, and operations; track active use, completed work, correction rates, time-to-value, and repeat usage rather than licenses provisioned.
  • Require field observation before automating high-value workflows, and favor integrations inside the operator's existing system of record over a standalone chat interface.
  • Avoid presenting exception handling as a later enhancement; explicitly catalog exception classes, required context, human authority boundaries, and recovery procedures during workflow design.

Source/Metadata

  • Title: How do you diffuse AI into the real world? — Varun Shenoy, Long Lake
  • Transcript words: 2959
  • Duration seconds: 1066
  • Timestamp note: No timestamps or chapters were present in the supplied transcript.
Full transcript 2966 words · 13 min read
0:00

Varun Reviewer Hi everyone, I'm Varun. I'm one of the co-founders at Long Lake, and I'm excited to share a little bit about what we've been up to for the last two years. It all comes back to a question all of us have asked time and time again. The models are getting better, but the real question is, how do you actually deploy the AI into the real world? How do you get the models to complete economically relevant tasks?

0:42

Let me start by saying everyone has seen the demo. Think of the agent automatically booking a flight, the agent automatically completing a ticket in some kind of customer service portal. Think of an agent completing a block of code, ready to commit and go. The reality is we've all seen this, and it feels like magic. Two years ago, any of this would have been complete science fiction. The capabilities are real. Now, walk with me into a 200-person property management firm. Real people, real properties, real dollars, real customers all across the US. You would expect AI to show up by now, but the reality is nothing has changed at all.

1:34

Here's the thing. This is totally normal. And maybe, in fact, I'd argue this is what we should expect. This is true for every general-purpose technology. Take electricity, for example. Electricity was invented in the 1880s, and it was first demoed at Edison's Pearl Street Station Dynamo Room over in Manhattan. This was the magic demo of its time. The reality is it took a long time for electricity to be fully adopted. Consider a Ford factory. It's not enough to just have electricity. You have to rip out the existing motors and equipment. You have to bring in the new equipment. You have to go and train everybody to use that very same equipment.

2:20

Here's a picture of a Ford electrified moving assembly in 1924. These things take time. Diffusion of any technology takes a generation. And since everyone here in this room today is talking about AI, I would argue AI diffusion is perhaps the single most important problem for the next 20 years. The models are going to keep getting better. The big question is, how do we actually get these models to be in the real world, complete real tasks, and make people more efficient, happier, and provide better service? So, taking a quick step back, who are we? We are Long Lake.

3:01

Over the last two years, we've raised over $3 billion from Elad Gil, General Catalyst, and Alpha Wave since our founding. Here's the strange part. We don't sell software. We actually go out and acquire and partner with real services businesses in the world. We've acquired 35 businesses across HR and property management, architecture, HR services, and a lot more. To give you a little bit more flavor, we have roughly a 40-person team right now, split between technology, finance, and operations. More than half our team is part of the technology team, focused on building products, data, and deploying the core products into the field.

3:41

We're a collective group of folks, a bunch of ex-founders who've worked in the services before, ex-military folks from Palantir, Ramp, Glean, and, from the finance side, Blackstone, HIG, etc. We are not selling to these companies above from the outside. We're actually deploying into these companies and figuring out how to get the technology to work. And just to show you the scale we're playing at, we announced recently our $6.3 billion take-private of American Express Global Business Travel, the world's largest corporate travel platform. We own these businesses. So when the AI doesn't work, it's not their problem. We're not the vendor. It's our problem.

4:27

Concretely, again, we are not the vendor. We are the operator-owners. And we work very closely with our teams within the businesses to drive real outcomes. Now, I want to step back and get concrete about the how. What are the lessons we've learned over the last two and a half years? And what we've learned from deploying AI into companies we've owned?

4:51

Three quick lessons. One, how we move agents from co-pilots to co-workers. Two, how we leverage real-world data within these businesses. Remember, we're seeing all of the work that's being done in these real services businesses. There's a lot of interesting problems and solutions embedded within that. And then finally, perhaps the most interesting and exciting, is how do you actually get all of this technology to compound over time by learning loops in the enterprise? And we'll get to that at the end over here. So starting off from co-pilots to co-workers, there's a spectrum of how much autonomy you can give an agent.

5:30

On the left here, you see a co-pilot. This is your simple RAG chatbot from two years ago. It's very quick. You can ask a question. Maybe it's integrated with some systems. It can give you information back very, very quickly. The second step is a synchronous agent. Consider something like Cloud Code, Codex, Cloud Cowork. It's real-time. There's this two-way interaction. It's a bit more sophisticated than a co-pilot. You can let it run off for one to five minutes. It'll call tools, maybe use its skills. It's still synchronous. You still need to step in and ask a query. So the next obvious rung of the ladder is the asynchronous agent. You can come in here, still ask a query.

6:08

The agent will go off into the background, do some work, and then come back. And what's really interesting about asynchronous agents is that the user does not have to be the one that triggers them. You can have external triggers as well. Maybe someone completes a certain task and there is an async job queue that allows the async agent to pull from it and proactively offer advice to the end user. Then I'd argue the next step is a long-running agent. How do you get these agents to work for hours, days, weeks, months, etc.? I think this is currently a very core problem that a lot of the labs are focused on, as are we.

6:47

And then finally, at the end, the holy grail. An AI co-worker. This is where most people start off. You want a proactive partner that gets work done just alongside you. This is what everyone wants to sell you. But what we've learned from owning the outcomes in this business is you have to earn the right to do more. It's not enough to jump to the co-worker immediately, for a bunch of reasons. One, for certain tasks, the models might not quite be there yet. And two, you actually have to work with these companies in the field, interact and iterate very, very closely so that they understand that this is the beginning of AI and you can work up the rungs over time.

7:32

I think a really unique lens to look at this problem through is that of the jagged frontier. We all know that agents are incredibly good at writing code. So what does, for example, the synchronous agent for code generation look like? This is super simple. This is just your coding agent. Maybe it's Codex, Cloud Code, just running on your desktop. It has access to a file system. You collaborate with it in real time. You get instant feedback and you iterate. The next step is, if you look at code generation, what is the async agent?

8:03

This is also fairly straightforward and largely solved. You take the exact same coding agent, you wrap it in a sandbox, and you just let it go run. It can build, it can test, and once it's done with its work, it can provide the code in the form of a PR. One thing that's really unique about engineers is folks are incredibly good at already parallelizing their work. It's very commonplace to launch 10 jobs and be comfortable with the fact that job seven might finish before job three. So engineers are incredibly good at using these async agents. Now, when we come to services, the equivalent of a synchronous agent, what we talked about a little bit earlier,

8:39

it's a co-working agent. It's an agent that has deep context about your enterprise. It interacts potentially with MCPs, custom tools, custom integrations, and you can chat with it synchronously just like any of these other products. I think this is the frontier here in the bottom right. What does it mean to build an asynchronous agent for the services? What does it mean to parallelize work in industries where work is traditionally done in a very, very serial manner? This is where we spend a lot of time, and this is what I wake up every morning really excited, thinking about. We've figured out what the async and forking mechanism for code is.

9:17

You just spin up a bunch of sandboxes and do work. What does that look like for the rest of the world? So here's a couple of questions we think about pretty seriously. One, the models are trained on code. They want to write code. They're incredibly good at writing code. How do we leverage these coding agents for actual knowledge work? Rather than wait for the models to catch up on doing services knowledge work, what if we just use that code knowledge and represent knowledge work as code? Two, as I mentioned, engineers are used to parallelizing work. How do you parallelize work that's traditionally serial?

9:49

People clean out their inbox one email by one email, not 10 emails at once. And finally, how do you move up the ladder here, both in terms of product and user enablement? What are the right form factors? And I'd argue this varies dramatically from industry to industry. Just because you have one way of launching an async agent for code doesn't mean that same way is going to work for architecture or property management. The second point I want to cover today is leveraging real-world data. We all know this: frontier models have learned from everything humanity has written down, but the most valuable tasks are not on the internet.

10:25

How do you actually close the books when you're missing receipts? How do you scope a building for construction in a blueprint, potentially collaboratively? How do you coordinate vendors for fixing a broken roof? All of this knowledge lives in people's heads, in 20-year-old software, in the way that one senior person on one of these teams just knows how to do it. How do you make this information explicit and create tasks that you can actually learn from?

10:51

So we've constructed a little bit of a flywheel. We get our agents to collaborate with our employees to do real work. And this allows us to generate rich traces of data and information. Tool calls, the hiccups, the paper cuts, everything that goes wrong with doing real work. This, in turn, allows us to build real-world evals. There is a ground truth here. In the case of the roofing example, the question is, did the roof get repaired? Did the books get closed? And this allows us to hill-climb and build better agents, which leads to more and more impact. And what's really exciting is it ratchets up every week. Our hill-climbing benchmarks become a regression test.

11:33

So our agents get better and better over time. Just to drive a little bit deeper here on the traces, there are three upshots of being able to collect these rich traces. One, we get to generate amazing evals that are built and scored automatically. And we're able to gather both implicit and explicit feedback. Explicit feedback in the sense of thumbs-ups and thumbs-downs. Maybe people provide a note telling us whether this response was good or not. And also implicit feedback. Again, we have the ground truth. Maybe there's some data that the AI generated. And there's a real diff between the data that the AI generated and what was ultimately submitted.

12:10

That's rich information that almost no one else has. Two, we've started post-training models internally on all of the data that these businesses operate on and produce, generally speaking. This is all data that is completely out of distribution for most frontier labs. Think of the tasks I showed at the beginning. A lot of the frontier models today just can't do these tasks yet. And we're trying to post-train our own models internally to be able to do that on the rich source of data that we own. And then finally, the actual agents themselves. The real world is incredibly hairy and messy. And you want customization per company. Every company does things very differently.

12:51

Customization per user. The way each user does their work is very unique. And customization per client. The way you work with every client is different. It's a services business. And you want to uphold those standards. I love this picture. Because it's the whole thing in a single image. The way we usually talk about LLM tasks is the top panel. It's a slope, you've got a bike, but there's clear sight to success. The reality is most work is not like that. And you and I both know that. There are hills and ravines. There's death by a thousand paper cuts. But that's what real work looks like. That's the entire job. The exceptions are the job.

13:42

That's the demo. That's the actual job. Now onto the final thing I want to chat with you guys about today: learning loops within the enterprise. I'd argue there are two hot trends everyone's talking about in 2026. One, it's continual learning. How do you make an agent better over time with feedback? I think there are plenty of sessions this week on how you can use continual learning, whether it's in the prompt or in the weights. And two, enablement. How do you get in these enterprises and actually get them to adopt and use AI?

14:25

Traditionally speaking, these two initiatives are owned by two separate teams. Continual learning is owned by your research team, your platform engineering team. Enablement's owned by growth or deployment or customer experience. Usually pretty siloed, not much interaction between the two. We think these are part of the exact same loop. The agent only improves if people actually use it. And people only use the agent if it's worth adopting. So here's a little graphic of a snowball. More usage drives continual learning, which drives a better agent, which drives more usage again. All this to say, there's still a really big elephant in the room.

15:04

How do you get the initial usage? I think a lot of people will use Cloud Code, or give it to their whole enterprise, and expect folks to just start using it. Everyone assumes the usage just shows up. But as we all know, that's simply not the case. It never does. Getting a hundred-year-old firm to change its processes is hard. You could have the best AI co-worker on the internet or on earth. And if the person who's closed the books for the last 20 years continues to do things the same way, nothing changes, nothing happens. So what can you actually do about it? This seems incredibly hard. What's the upshot? How do you actually get this stuff to work?

15:52

Well, I think a lot about Jensen and how he dominated the market, in his words, with extreme hardware-software co-design, designing the chips and the software together as one system. We look at this through the lens of extreme software-service co-design. How do you co-design our products with the people and the processes at our businesses? And I'd argue this is only possible from being within, under the same roof. We need to meet the people within these companies, both metaphorically, for example, bringing products to their systems so that the energy required for enablement is kept low. And also physically, get on a plane, show up, say hi, learn what people actually do.

16:37

Maybe you build a product that's natively embedded into Excel or into their ERP system, maybe their 3D design software, or maybe even their Microsoft products like Outlook, Gmail, etc. Or you show up in person, you do a lunch and learn with a bunch of folks at one of the companies, you go to their conferences and you create cotton candy and run a stand for them. You go mountain biking and ask them about all the difficulties that they have with their actual day-to-day jobs. Or you show up in person, one-on-one, or sometimes even two-on-one in this case, and just show them how to use the tools.

17:02

And learn from the feedback because this is what the rest of the world really looks like. It's not like the folks in this room or in San Francisco. It's a lot more like this. You cannot co-design software with a services business over Zoom or over a support ticket. You have to be there. You have to be in person. And I'd argue this is the part that actually makes it work. In order to get AI diffusion to work, you have to touch some grass. Thank you so much. And I'll be around for the rest of the day if there's anything I could help with.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note