AI Engineer

What Lies Beneath the API — Benjamin Cowen, Modal

698 summary words 3 min summary Watch video

Start with the signal

3 min read

Summary

30-second take

Modal's Ben Cowen argues that AI startups inevitably cross from generic API use to fine-tuning as they scale and specialize—when frontier models cost more than customers pay, when latency/throughput hits walls, or when evals plateau. The thesis: domain-specific products are custom by definition, and modern serverless platforms plus open-source libraries (VLLM, SGLang, RL frameworks) now let engineers train and serve models in ~300 lines of Python without infrastructure teams or dedicated clusters. He positions Modal as the middle ground between frontier APIs (no customization) and self-managed clusters (huge infrastructure burden), enabling fast iteration with algorithm-level control.

Key takes

  • Economics force the shift: If you're in "caveman mode" (token optimization hacks) and still paying your API provider more than customers pay you, your unit economics don't scale—fine-tuning can deliver 5x–100x cost improvements (Intercom, Pinterest examples).
  • Frontier labs optimize for generality, not your business: GPT-4/Claude want to win on everything; you want to win on your specific task. Domain specialization through fine-tuning is how you separate from competitors using the same API.
  • You probably already have training infrastructure: If you built an agent harness, eval pipeline, or feedback loop, you have the scaffolding for RL or supervised fine-tuning—data collection and evals are the blockers, not code complexity.
  • Serverless unlocks training workflows, not just inference: Hyperparameter sweeps, RL rollouts (50k–100k sandboxes on Modal), and model serving (VLLM/SGLang) all autoscale on-demand, eliminating the "sacred cluster minutes" problem and enabling embarrassingly parallel experimentation.

Useful details

  • Customer examples: Intercom beat frontier API at 1/5 cost; Pinterest claims "orders of magnitude" improvement; Decagon fine-tuned for business logic differentiation.
  • Code complexity claim: Supervised fine-tuning in ~300 lines of Python; RL in similar scope using open-source libraries (no manual gradient taping or linear algebra implementation).
  • Modal's unified API: Same interface for GPU containers/clusters and code execution sandboxes, enabling RL rollouts (practice simulations) and training in one platform.
  • Inference serving: Modal supports VLLM, SGLang, Triton Inference Server, or custom Python inference with autoscaling to match traffic.
  • Open-source examples: Modal publishes training code in their examples repo for users to start immediately.

Caveats / counterpoints

  • No counter-narrative: Cowen doesn't address when not to fine-tune (e.g., small teams, low data volume, rapidly changing product requirements, or when prompt engineering genuinely solves the problem).
  • Vendor pitch embedded: This is a Modal conference talk; claims about "300 lines of code" and ease of use are self-serving and assume Modal's abstraction layer. The actual ML engineering lift (data quality, eval design, hyperparameter tuning) is downplayed.
  • Data quality handwaved: "Garbage in, garbage out" is mentioned as a blocker, but the talk assumes you already have clean, labeled data or RL reward signals—often the hardest part of fine-tuning projects.
  • No discussion of model risk: Fine-tuning introduces deployment complexity, versioning, monitoring drift, and the risk of overfitting to a narrow domain if product requirements shift.

Ken relevance

  • Agent systems timing: If you're building agents with eval loops and feedback, Cowen's "you already have training infrastructure" framing applies. Consider whether your agent harness could be repurposed for RL training as a competitive moat.
  • GTM / positioning angle: The "frontier labs optimize for generality" argument is a strong framing for vertical AI products—if you're advising or investing in AI startups, this is the wedge for domain-specific differentiation.
  • Serverless for AI ops: Modal's value prop (autoscaling training/inference without cluster management) aligns with your interest in removing infrastructure toil. Relevant if you're evaluating platforms for your own AI systems or portfolio companies.
  • Business model red flag: The "paying API more than customers pay you" metric is a sharp signal for when AI startups need to vertically integrate. Useful lens for assessing unit economics in AI businesses.
  • Content opportunity: The "when to fine-tune" decision framework (evals plateauing, cost inversion, latency walls) could be a valuable explainer or checklist for your audience.

Watch verdict

Skip. The transcript delivers all actionable insights—this is a pitch deck disguised as a talk, and the visual code snippets add nothing substantive. You've captured the framework, customer proof points, and Modal's positioning. No novel technical depth or debate to warrant watching.

Full transcript 1648 words · 7 min read
0:15

SPEAKER_00

Yeah, so good afternoon and thanks for coming to the session. I know there's a lot to choose from. My name is Ben Cowan. I'm a forward deployed machine learning engineer at Modal. And I want to talk about an interesting pattern that we've been seeing in AI application development. I'll give you the punchline now. It's about as companies mature and their products mature and specialize, we're seeing more and more turn to fine tuning to get increased performance, better costs and so forth. So this brings up an interesting question of when does your application step over the line into a custom domain? When is fine tuning worth it?

0:59

SPEAKER_00

So if you're not familiar with Modal, we're a general purpose serverless compute platform. We provide basic building blocks like serverless functions and hardened sandboxes for code execution. And so as an FDE on a general purpose platform, I've had the opportunity to work with an extremely wide range of AI applications. So from physics simulations to quantum chemistry and of course voice processing, LLMs and agents. A really interesting one that's been really blowing up is large scale reinforcement learning. And so we started to think about some of these customers as where are they on the model spectrum. So on one side of the spectrum is the frontier API.

1:55

SPEAKER_00

And the frontier APIs, I think everyone here would agree, has unlocked a completely new era of accelerated growth. People can build basically anything exceptionally fast. They're amazing. But you can't customize it at all beyond prompt engineering. And so you might. I love the whole caveman mode thing. If you tell your LLM to speak like a caveman, you can reduce your tokens by a lot. But that's not going to scale if you are a startup 100 X's or 1,000 X's, right? Another interesting thing we see is when startups win large enterprise contracts with very specific latency or throughput requirements.

2:48

SPEAKER_00

There's very little ability to customize for those things, let alone if you have a custom metric that encapsulates your business logic. Okay, so to get this model differentiation, a lot of companies turn to fine tuning. And what that has meant traditionally is this huge jump to the other end of the spectrum, right? Training has a very different scaling and compute characteristic to most production workloads. So if you want to train and serve a production workload, in the past you have to get a big cluster. Now you have to isolate those resources from your production resources.

3:29

SPEAKER_00

You're going to need infrastructure engineers or your AI engineers are going to be working on infrastructure. Maybe even your scientists. So with this extremely customized, powerful option, you also have this massive responsibility for the entire stack. And so you might have a guess who I would recommend for this, but there's a middle ground that's emerging. There's a new type of cloud provider that makes this a lot easier. And we're building this to address this problem that we're seeing, right? So leader after leader in the space are announcing or publishing that they've fine tuned and gotten incredible results.

4:19

SPEAKER_00

Okay, so Intercom is beating their frontier API at one-fifth of the cost. Pinterest says orders of magnitude. I wish I knew the exact amount, but I think one of our customers, Decagon, has summed this up really well, which is that basically the frontier labs probably don't have the exact same goal as you, right? They want their models to win on everything possible. And we want our models to win at our business logic, right? You want to be the best at what you provide your customer. And so, yeah, I'm happy to announce that it's actually a lot easier than you might think to train a model.

5:07

SPEAKER_00

There's some incredible open source libraries out there now that make this extremely accessible. They give you full control over the algorithm, right? So you get to reach across the spectrum to doing it yourself at the algorithm level. Without having to also manage the cluster and so forth. And the most important thing is that this retains the fast iteration cycles of the frontier end of the spectrum. So that's basically our entire mission is to give you algorithm control and fast iteration.

5:46

SPEAKER_00

So this is my hot take that it's just a matter of time until your product steps into being domain specific. So in some sense, if you have a differentiated product, it is custom. So when exactly you cross that line, that's a decision you have to make. But it's something we'd love to talk to you about. So I have here a few signals that might indicate that you're getting close to that time. So if you've moved to caveman mode and you're still paying more for your API than your customers are paying you, that might be a signal that your economics aren't scaling. Right? And that you could probably benefit from a customized inference endpoint.

6:39

SPEAKER_00

Same for latency and throughput. Right? So if you are plateauing on your evals, that's a signal that you might get something out of fine tuning a model. There's a decades old adage in training that if you have garbage data, it's garbage in, garbage out. So if you haven't been collecting data and you don't have mature evals, it's probably not time to train. You need to collect the data. That said, this is one of the main takeaways that I'd love for everyone here to walk out with. Is that if you have built a product, you probably have at least touched all the things you need to train if you haven't already done it. Okay?

7:24

SPEAKER_00

If you've built an agent harness, then you have what you need to have a new model learn through reinforcement learning how to provide your service. Right? So if you're evaluating your products and collecting that data on what's working and what's not, then you have training data to train your model. And with the advent of serverless compute platforms and these open source libraries, a lot of us, when we started training models, we were taping the gradient by hand and implementing the linear algebra. You don't have to do that anymore unless you have a freaky model, which, if you do, I'd love to talk to you about it.

8:13

SPEAKER_00

But you don't need the infrastructure experts and so forth. So this is an exciting time. I knew the video wouldn't play. Anyway, so the next couple of slides are just some snippets of code. I don't expect, yeah, you can't even really read it, but I just want to illustrate what it looks like to set up a training algorithm today. It's not a gigantic monorepo with thousands of lines of code. You can do supervised fine-tuning in 300 lines of Python. Okay, so once you have your data curated, once you have an account on modal or some other serverless platform, you can get started really fast. And this code is on our examples repository.

9:03

SPEAKER_00

The thing in this video, it's just showing how we can scale containers really fast. And so just to bridge these concepts a little bit, people usually associate serverless with inference. But with training, it can be really handy, too, for doing something like hyperparameter tuning. You don't have to, you know, every minute on your cluster isn't sacred anymore. You can fan out to a bunch of containers, get them on demand. As soon as it's not promising, kill it. And it's almost like a meta-evolutionary algorithm at that point. So it's an exciting time to be doing that. And the same goes for reinforcement learning.

9:53

SPEAKER_00

A lot of us who got our graduate degrees in machine learning in the last 10 years didn't do reinforcement learning. Right? This is relatively old. But the stuff we're using today is kind of new. But they have these libraries, too. You can do it in 300 lines of code. And something interesting about modal in particular is we have unified APIs for sandboxes and GPU containers or clusters. So what this means, in a nutshell, when you're training a model with RL, it needs to practice a lot. And so this is massively, embarrassingly parallel kind of evaluation thing called a rollout.

10:38

SPEAKER_00

And so we have one of the most amazing things in the last quarter has been customers scaling up to 50,000, 100,000 sandboxes just to do RL. And you can do it, too. The code is open source. And then I'd be remiss not to mention what comes after the training. You have to serve the model. Right? And this is what the Frontier API is doing under the hood. Well, I don't know if they use VLLM. But my point is that you can do it, too. And it's actually not that much code. VLLM, SGLang, Trident Inference Server, or a custom inference workflow with just Python on modal or other serverless platforms. You can autoscale all of this stuff to match your traffic as it's coming in.

11:33

SPEAKER_00

So, yeah. So just to sum everything up what I'm saying here, I'm not saying go train your model right now. I'm saying it's not something that is like, oh, I'll do that in 10 years. You might want to train your model in one year. Right? You might want to do it in six months. So start thinking about when am I going to know, okay, it's time to train my model. And how can I prepare for that moment by collecting data, developing your evals. And, yeah, I'd love to come by our booth. We're at the end over on that side. I'd love to talk to you more about this. Or you can reach out at my email here. That's it. Thank you. Thank you.

12:35

[SPEAKER_00] Thank you. [SPEAKER_00] Thank you. Thank you.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note