AI Engineer

Notion's Token Town — Sarah Sachs, Notion

1886 summary words 8 min summary Watch video

Start with the signal

8 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: AI product companies avoid becoming "token-poor" by treating model access as a volatile commodity: preserve multi-provider optionality, route work by end-to-end cost/capability/latency, use open weights and deterministic software where possible, and make security and orchestration core product capabilities.
  • Why it matters: This is a practical operating thesis for building agent systems and software factories without surrendering margin, product quality, or security to a single frontier-model vendor.
  • Best use: Use it as an architecture and vendor-strategy briefing for model routing, agent orchestration, governance, and the economics of AI-enabled workflows.

Executive Summary

Sarah Sachs, who leads Notion's AI engineering teams, argues that the primary applied-AI challenge is not simply gaining access to the strongest model. It is building a sustainable product when model suppliers are also potential competitors, model pricing changes unpredictably, and nominally similar upgrades can radically increase token consumption or price. Her prescription is to compete on customer understanding, data flywheels, workflow orchestration, and interface quality—not on reselling a lab's tokens.

Her central operating model is model optionality. Notion maintains multiple model paths, with its Auto Model handling roughly 75% of traffic while preserving access to state-of-the-art models when a task genuinely requires them. The decision unit should not be token price or one-call latency alone, but the end-to-end trajectory: required capability, tool failures, latency, total cost, and user outcome. This also creates negotiating leverage because the company can credibly move traffic rather than accept vendor lock-in in return for discounts.

Sachs separates frontier work from routine work. Expensive frontier models may be justified for difficult tasks such as large-scale data analysis, but routine inbox triage should not automatically receive the most capable model. Open-weight models increasingly cover moderate tasks and can be improved with reinforcement learning; deterministic jobs should bypass LLMs entirely through CPUs, CLIs, workers, SQL, or ordinary code. This reduces cost while improving determinism.

The talk closes by arguing that the next constraint on agentic systems is security and orchestration. Systems become dangerous when they combine private data, untrusted content, and external communication capability—the "lethal trifecta." Notion's software-factory example places agents inside a collaborative, inspectable work surface: one agent scopes work, humans provide missing context, specialist agents provide customer insight, coding agents create PRs, and another model reviews them. The intended result is an auditable multi-agent workflow rather than an autonomous system engineers must continuously babysit.

Key Takeaways

  • Claim: Treat frontier-model vendors as strategically misaligned suppliers, because their first-party products can receive better economics than your resale-based AI product. | Evidence: Sachs describes recurring real-world scenarios in which a model upgrade keeps per-token pricing flat but consumes 3x more output tokens, or raises price by 40% while the prior model is deprecated within four months. She also cites a SemiAnalysis comparison showing major differences between frontier labs' first-party subscription economics and API pricing. | Implication: Do not base unit economics or product positioning on a single vendor's current API pricing; maintain the ability to substitute models and build differentiated value above the token layer. | Caveat: The speaker's "supplier is your competitor" framing is a strategic heuristic, not proof that every model provider or commercial agreement is adversarial.
  • Claim: Model routing should optimize cost per capability per second across a full workflow trajectory, not benchmark rank, raw token price, or latency for an isolated call. | Evidence: Notion's partnership evaluation for Parallel web search considered complete search trajectories rather than a single call; Sachs says this granularity exposes trade-offs that one-call cost or latency obscures. She recommends measuring product-specific needs such as tool-error rates and actual latency requirements. | Implication: Build an eval and routing layer around end-to-end task success, cost, latency, tool reliability, and user impact before making provider or model-selection decisions. | Caveat: This requires proprietary, continuously maintained evaluations tied to real workflows; public benchmarks alone cannot supply the necessary routing judgment.
  • Claim: Use model optionality as both a product capability and commercial leverage, rather than accepting a large vendor discount in exchange for dependency. | Evidence: Notion's Auto Model handles about 75% of its traffic while customers retain access to multiple state-of-the-art models. Sachs advises building a multimodal harness that can switch models, and says Notion exchanges high-value use-case evaluations and early-access feedback with frontier labs instead of relying on extraordinarily large commitments. | Implication: Invest in a provider-neutral model abstraction and switching capability early; the cost of building it may be lower than the strategic and financial cost of being unable to leave a provider. | Caveat: Interoperability has genuine engineering costs, including challenges such as cache invalidation when switching models mid-transcript.
  • Claim: Frontier models should be reserved for genuinely hard work; moderate tasks are increasingly candidates for open-weight models, while deterministic tasks should not use LLMs at all. | Evidence: Sachs contrasts large-scale data analysis, where Notion may recommend Claude Opus, with email triage, where using Opus would be wasteful. She argues open-weight models now handle more than small SFT tasks and may close capability gaps over time; she also cites deterministic alternatives: convert CSV to PDF with code, use a CLI for tool calls, and run SQL directly rather than ask an LLM to do it. | Implication: Create a task taxonomy with at least three paths—frontier, open-weight/cheaper model, and deterministic execution—and progressively migrate tasks downward as capability improves. | Caveat: Open-weight suitability must be tested on the actual workflow; the argument is not that open-weight models are universally the strongest models.
  • Claim: The critical near-term agent problem is security, especially when autonomous systems combine sensitive context, untrusted inputs, and the ability to act externally. | Evidence: Sachs references Simon Willison's "lethal trifecta": access to private data, exposure to untrusted content through sources such as ingestion, MCP, or email, and the ability to communicate externally through actions, payloads, or web search. She stresses that more autonomy means more unsupervised risk. | Implication: Gate any agent that reads sensitive data and consumes external content before it can take outbound actions; isolate execution environments and explicitly define what each agent can see, retain, and do. | Caveat: The transcript names the risk pattern but does not provide a detailed control framework for authorization, sandboxing, auditability, or prompt-injection containment.
  • Claim: A useful software factory is a collaborative orchestration layer, not a collection of autonomous coding agents operating without context or review. | Evidence: Notion's demo starts with Claude scoping a task into an active collaborative document, asks a human team member to resolve a missing question, brings in Decagon for customer-voice data, has Claude create a PR, and asks Codex to review it, finding two issues. Sachs says Notion internally routes substantial feedback to appropriate teams and lets coding agents take the first step; she cites over three minutes saved per task for customers. | Implication: Design agent systems around inspectable state, specialist delegation, human escalation, and independent review—not merely around a single agent's ability to generate code. | Caveat: The stated time-saving figure is presented as a customer ROI claim without methodology, task distribution, or quality/error-rate data.

Detailed Brief

Notion's durable system-of-record thesis

  • Claims: Sachs frames AI adoption as a maturity path from thought partner, to assistant, to teammate, to a critical workflow system in which humans and agents collaborate with each other.; Her explanation for stalled AI transformation is fragmented data and the absence of a durable shared system of record.; She claims 88% of organizations have not progressed beyond AI used as an assistant.
  • Evidence: Early AI usage is described as copy-pasting text from ChatGPT into email; assistant-level AI executes isolated user-requested tasks, while teammate-level AI performs repeatable processes.; Notion positions its collaborative workspace as the place where human-human, human-agent, and agent-agent coordination can persist.
  • Caveats: The 88% figure is asserted without a source or definition of the maturity categories.; A shared workspace alone does not solve data quality, permissions, workflow ownership, or agent reliability.
  • Implications: The control plane for agentic work should preserve task context, decisions, artifacts, feedback, and ownership in a shared inspectable surface rather than scatter them across prompts and point tools.

Governance and persistence as product differentiation

  • Claims: Sachs treats governance as more than policy compliance: model optionality can improve customers' visibility into data usage, maintainability, and control.; She identifies persistent enterprise knowledge—what agents see, what they do, and what remains available afterward—as under-discussed in multi-agent design.; Sandboxes and computer-style execution environments can improve determinism as well as security and token efficiency.
  • Evidence: The presentation contrasts a factory engineers must babysit with a managed orchestration environment where work can be inspected and coordinated through a live collaborative document.
  • Caveats: The presentation does not specify data-retention policy, memory scoping, cross-agent permission inheritance, or audit-log implementation.
  • Implications: Enterprise agent platforms should make state persistence and governance deliberate configuration choices, with clear boundaries between ephemeral execution context, reusable organizational knowledge, and sensitive data.

Notable Concepts & Terms

  • Token-poor: Sachs's label for an AI business whose product economics become unsustainable because it overpays for inference or applies expensive models to work that does not require them.
  • Cost per capability per second: A routing lens that combines task success, capability, end-to-end latency, and total cost rather than optimizing token price alone.
  • Model optionality: The technical and commercial ability to switch models or providers, which protects product quality, limits vendor lock-in, and creates negotiating leverage.
  • Auto Model: Notion's routing approach, said to handle about 75% of traffic, that selects among models rather than binding all work to one provider.
  • Open-weight models: Models available with weights that can lower inference cost, create a credible alternative to frontier APIs, and increasingly address moderate-complexity tasks.
  • Lethal trifecta: A security-risk pattern attributed to Simon Willison: private-data access plus untrusted content plus external communication/action capability.
  • Software factory: An orchestrated workflow in which agents and people jointly scope, enrich, implement, review, route, and close work rather than relying on one autonomous coding agent.
  • AI Switzerland: Notion's positioning as provider-neutral infrastructure that gives customers access to multiple leading models without tying them to a single lab.

Operator Notes / Why Ken Should Care

  • Require every production agent workflow to declare its model-routing policy, fallback model, maximum cost budget, expected latency target, and deterministic non-LLM path where applicable.
  • Create a product-specific evaluation suite that scores complete trajectories, including tool-call failures, retries, user-visible latency, cost, and task success—not just model benchmarks.
  • Audit agent workflows for the lethal trifecta; block or require explicit approval for combinations of sensitive data access, untrusted ingestion, and outbound action.
  • Separate durable organizational memory from per-task context, and define who can inspect, modify, retain, and reuse each category of agent state.
  • Avoid long-term volume commitments or architectural dependencies that prevent provider switching unless the commercial benefit demonstrably exceeds the cost of lost optionality.
  • Use a staged multi-agent delivery pattern for software work: task specification, human clarification where needed, specialist evidence gathering, implementation, independent review, and monitored closure.

Source/Metadata

  • Title: Notion's Token Town — Sarah Sachs, Notion
  • Transcript words: 6360
  • Duration seconds: 1435
  • Timestamp note: No usable timestamps or chapter markers were present. The supplied transcript contains substantial repeated sections and trailing extraction noise.
Full transcript 4119 words · 28 min read
0:02

.

0:21

Okay, hello. Okay, before I get started, you guys, this is a huge keynote room. Can everyone come forward? Because I'm talking to four empty rows and dispersed people. Do me a favor, I'm spending 30 minutes telling you all of our secrets. I can see you still. Thank you, thank you, thank you, thank you. We're just gonna chat. It's a giant room, and there's 500 of us. This room is way larger than that. Thank you. Honestly, I knew you guys had it in you. It's really not so hard. Thank you. I also sit in the back. I also work during talks. I get it. I totally get it. I did all day, but not for me. Okay.

0:24

I'm gonna start, but I'm gonna still point at you if you're in the back, like you. Okay. I'm Sarah. I lead our engineering teams for AI at Notion. Welcome to my talk. It's about Token Town. How do you go from AI-pilled to AI-poor?

0:26

Okay. I know that today's all about software factories. We're gonna talk about that, but we're gonna talk about how to do it sustainably. This is me. This is on my first day at Notion in a very sweaty subway. Like I said, I lead our AI teams at Notion, and I negotiate AI contracts for a living. My team jokes I act like Anna Winter, so this is a nice image of me with AI Anna Winter hair after a press article referred to me that way externally. And that's kind of the idea, right? How do you think about negotiating between different vendors, making sure that you maintain taste for your company?

0:28

I don't do it alone. This is launch day at one of our recent launches. This is just a subset. Any good engineering manager points out that we have a whole company of people building this. I'm just the one that gets to come talk to you about it. So we've been building a lot. This is an example of our AI usage just in 2026. And we've been really proud of how we've been able to grow that usage, and I'm gonna talk to you about how you can build an AI-native product and an AI-native company. But this is just to give me some credit that we're doing it well.

0:31

Okay. So for those of you that don't know, Notion's always been that durable system of record. It's always been the place where you can collaborate with your peers. But today, that point of collaboration is a little bit different. It's not just humans. Notion's always been the place for collaboration. And today that collaboration happens between humans and agents, humans and humans, agents and agents.

0:33

And we like to think about AI transformations going through this journey. And I'm sure some of you are looking at the slide and wondering where you are. AI as a thought partner is when we all started tinkering. We all started just going to the very first version of ChatGPT on Thanksgiving when it came out three years ago, four years ago. And we started saying, how can I send this email to my landlord to say that I shouldn't pay for repainting, right? And then we'd copy-paste it, enter it into our email.

0:35

Eventually, we started getting to a place where we could use AI like an assistant. AI was able to maybe execute individual tasks. That's how Notion AI really took off in the beginning. And it was able to save employee time, but functionally was limited in its capabilities based on what humans asked it to do. AI as teammates is what we were really excited to launch almost a year ago now. But this is true in many products. We can do repetitive work and think about a process and have AI do that process. What I think is really interesting is when AI actually becomes that critical workflow where processes are interfacing with each other and you have entire systems running.

0:38

How many of you guys feel like you have AI as a system down? Aren't you sad you came up now? I'm kidding. Great. None of you. Exactly. We have found that no one has figured out how to do this well. 88% of people can't even get past AI as an assistant. And why is that? We have a thesis at Notion. It's because there's too much siloed data and not a durable system of record for that point of collaboration. And we believe that for your software factory to work, for your company to work, and for your systems to work, you need that durable system of record. And that is Notion's mission.

0:39

So doing that is expensive. You see a lot of companies that try and commit themselves to this vision, and these are just a series of headlines all within a week of how that's painful. So you can put all of your money into a process to try and make a system, and you end up feeling like this, right? You end up using a blow torch to light what is actually a large cigar. But you kind of get the idea. Cost is a structural barrier to entry. It makes it hard for you to serve products. It makes it hard for you to build factories. And it is ultimately, I would posit, one of the largest reasons why things do not happen at scale successfully today. And I would argue, for anyone working at an applied AI company, it's something for them to be really familiar with, to understand the trade-offs that they're making to build durable and exciting and enlightening product for their customers.

0:42

But that's not really how the market is today, right? I'm not going to name names here, but you guys have search engines. You can figure it out. Exhibit A. A reasoning model gets upgraded. Amazing. The per-token pricing is the same. What's not to love? You try it out. It uses three times as many output tokens, right? Exhibit B. A model gets upgraded, but it has an entire new digit, right? Whatever marketing system that model family likes, it's brand new. It's 40% more than its predecessor, which is being deprecated in the next four months.

0:47

These are real scenarios that we face at Notion. All of you are nodding because these are common, pretty much monthly now. But here's the problem. Are you growing 40% in that time period? Are you making 33X more revenue? No. So how do you navigate the system? If you just auto-upgrade your model and everything that you're doing, you're giving someone a bad deal, either your customers or your investors, depending on how you charge and where you get your money. Neither are good.

0:49

Fortune 5 million companies have the capability to navigate this. They can hire large consulting teams, have durable teams on their own, and build expertise on how to navigate these trade-offs. Most people don't. Everyone else has no ability to negotiate with leverage, and they're stuck in these scenarios, right? Part of my job as that Anna Winter joke is to think about advocating for the Fortune 5 million, the non-Fortune 500 companies that don't have the mass to have leverage and negotiate, but need to think about how. I'm going to share some of the lessons that I've learned when I have large amounts of traffic behind me that I think scale to those who don't.

0:50

This is probably less of a secret now than it was when I started giving talks like this, maybe four months ago. What's the problem? Your supplier is your competitor. I know very few people who have convinced me that that's not true. You will always be getting a bad deal on tokens with someone who builds them natively, right? Sometimes the cost of goods served is extremely different. They're serving a first-party product, and then you're buying those tokens at a huge surcharge and then selling them again at another surcharge. That's not really value you can defend. You're getting a really bad deal. And if you tie yourself to one provider, you have no exit. If you build an AI product that you're selling with this structure, you are crossing your fingers and hoping that you are a viable business. I do not encourage that.

0:51

This is really interesting. Dylan in Semi Analysis posted this. I think it says eight hours ago. It wasn't at this point. It was probably a month ago. They purchased a subscription plan, and they just highlighted how different what Frontier Labs charge customers for first-party products is versus what they sell. It's a bad deal. Don't play this game. Or try, and let me know how you win. I don't recommend it.

0:53

Think about everyone else. Think about what that structure means and where you have expertise. I don't think that that's winning on the token economics. I think it's about product. It's about building data flywheels and understanding your customers better than anyone else, understanding when you need capability, when you need low price, when you need latency improvements. I promise you, you don't always need what is usually the slowest but the most capable model out there. And then build compelling UI and orchestration. And I'll show you some examples of that to justify the cost on the bad-deal tokens that you do resell. The job is not to train.

0:55

Some of you might be training the best model, and I'd love to serve it, and come talk to me afterwards. But most of you are not doing that. It's a bad deal. Don't play this game. Or try and let me know how you win. I don't recommend it. Think about everyone else. Think about what that structure means and where you have expertise. I don't think that that's winning on the token economics. I think it's about product. It's about building data flywheels and understanding your customers better than anyone else. Understanding when you need capability, when you need low price, when you need latency improvements.

1:19

I promise you, you don't always need what is usually the slowest but most capable model out there. And then build compelling UI and orchestration. And I'll show you some examples of that to justify the cost on the bad-deal tokens that you do resell. The job is not to train. Some of you might be training the best model, and I'd love to serve it, so come talk to me afterwards. But most of you are not doing that. Stop trying to win that game and think about the best product that uses many models. Help your customers, help your team, bet on the frontier, not on the lab. And we'll talk about what it looks like to do that.

1:42

This cost-per-capability-per-second tradeoff is actually really intense. Citadel came out with this memo a while ago, maybe two weeks ago. I loved it. The idea is that, for the economy at large, simpler models might be the most cost-effective, productivity-augmenting pathway. They talk about this bifurcation on frontier versus everyday usage. I really believe that. And for every product, the definition of frontier versus everyday, the definition of saturated capabilities or model capability overhangs, depends on your expertise on your product. No one can replace that. And not all traffic is equal. It is a huge miss to send all of these to the latest Opus model.

2:08

Some of these, absolutely. Large-scale data analysis, when you do it on Notion, will recommend Opus, right? When you triage an email inbox, if we're charging you to do that on Opus, we're ripping you off and ourselves. Think about where your traffic patterns are. And then think about how frontier lab model providers are structured today. It's functionally an oligopoly, right? And that's fine because they're racing to the top, and I think the top is really hard and really important. This is not to say that products don't have a place for frontier difficult tasks. I want everyone to nod and understand that's not what this talk is about.

2:30

Understand when you need those tasks, and it's not everything. The problem with those tasks is, keep in mind how pricing is incentivized. You can figure out who these players are. Either you are the best model, everything above what AI can't do today is your market. You can basically price it as high as you want. If you're slightly behind that best model, all you need to be is a dollar per million tokens cheaper, and you have the rest of the market. You know that economic theory about gas stations where the best gas stations are the ones that are right next to each other because they cover east and west the most? Yeah.

2:44

It's the same with model pricing, which means that price does not correlate with capability growth. So for this complex task, understand what capabilities you need, but be the expert on what complexity is. And keep in mind that who handles complexity changes. Oftentimes, you'll see applied AI companies be super outspoken in marketing with a specific lab. That's always a red flag for me when they're not model agnostic. Because if you look at this graph, it basically shows that they're behind every month, right? The new model and the new model provider of the best frontier capabilities change.

2:59

And if you hitch your ride with one particular provider in exchange for, for instance, a larger discount, you're doing a disservice to your customers half of the time, right? So really think about if that discount is worth not actually having a frontier product. And remember that that optionality is your leverage. If you don't have the capability to walk at any point, you are stuck. And again, I think that's probably the most expensive decision you'll make, regardless of what discount you get or the engineering work to have model interoperability. One option to navigate this is stay model agnostic.

3:15

Have different models and capabilities in your system so that, at any point, if pricing seems unfair or untenable, you are not out of business. Notion's Auto Model does this really well. We have state-of-the-art models available always. But we also have an auto model there at the top that handles about 75% of our traffic, right? We have the ability to switch between models in our product, and we also offer it to our customers so that they have access to these models without vendor lock-in. That's part of our AI Switzerland approach. You guys love taking photos of slides. This is the slide. Okay. Model-agnostic playbook. This is how you do it. Build for multimodal.

3:35

It is hard to kill the cache and switch models mid-transcript. I understand that. We invest in that technology. It doesn't even have to be per thread. Just think about your harness as model interoperability. Think about the cost per capability per second, not just the tokens. Here's a great example. We posted this review when we announced our partnership with Parallel as our web search provider. If you were to look at just latency of a single call or just cost, Parallel might not be the cheapest. But if you have expertise in entire web search trajectories, you'll see how it differs.

3:59

The granularity of this eval is what lets us make the best decisions for our customers because we understand all of the trade-offs on entire trajectories, not just single calls. Switch fast and often. I think we talked about that. And give them something back. That expertise on use cases is also very valuable to Frontier Labs. We find that our evals and our early access program partnerships actually help us a lot with Frontier Labs and are something that we can exchange instead of extraordinarily large commits. And I don't think the discount is ever worth the loss in optionality. That's a perspective you can choose to keep or not. The second option is moderate tasks.

4:15

Understanding open weights plays there. Open-weight models are really strong enough to handle these tasks. And the possibility to RL on top of them has also expanded the upmarket growth that they can cover. I view open-weight models as basically lowering the barrier to entry on cost for our customers. And they also give you negotiation leverage. So it's a credible alternative that's putting that downward pressure on pricing that, if there's an oligopoly of two or three providers at the top, is unavailable right now otherwise. I think Kimi 26 was probably the first time that we really saw a model that outperformed 5.2, GPT 5.2, GLM 5.2. Now is another 5.2.

4:36

Bombshell in the villa that also probably does best here. But it's no longer the case where open-weight models are good for just SFT on small tasks. Really think about, without RL, if they're capable enough for what you need. And again, don't just think about external benchmarks. Be able to have expertise on your system. What are your tool errors? What's the actual latency that you need, right? Here's an example of a benchmark that we posted. It's a little bit stale on purpose, right? But you get the idea. Phillip at Base Ten's law showed this slide once, and I've stolen it ever since. Thank you. Are you here? Buddy. Okay. We'll chat. Hi.

5:22

Well, he could come up and say it better. But the idea is that you don't have to be at the top, right? I'm not trying to make a case that open weight is the best model out there. The case being made, however, is that the gap gets covered eventually. So if the tasks that you're having today are good enough, then in six months they're probably covered by open weight. So be prepared now. And the last thing is CPUs over GPUs. We've recently launched something at Notion called workers. I don't think that the GPU is necessary for every job. A lot of the jobs that we have are actually serving discrete pieces of code. You don't need an LLM to turn a CSV into a PDF.

5:46

You don't need an LLM to talk to Notion tool calls if we have a CLI. You definitely don't need an LLM to do deterministic SQL queries. This is where people become token-poor very quickly. And I think the last option here, besides open weight, CPUs, and optionality, is actually governance. There's a lot of AI governance. One is visibility. Understanding who's using the data. Understanding its maintainability and control. When you have model optionality, you can offer a lot more to your customers. Here's an example of how that governance works in Notion. So final tips again. Think about architecture. Think about open weight. And build value that transcends tokens.

6:16

So we're going to depart Token Town. I know I said welcome to Token Town. We're going to spend the next 10 minutes really thinking about what to do next. So I think the challenge of the next six months doesn't have to do with capabilities. I think it has to do with security. Let's start there. There's this concept called the lethal trifecta. Simon Wilson, I think, crafted this. This is where people become token-poor very quickly. And I think the last option here, besides open weight, CPUs, and optionality, is actually governance. There's a lot of AI governance. One is visibility. Understanding who's using the data. Understanding its maintainability and control.

6:59

When you have model optionality, you can offer a lot more to your customers. Here's an example of how that governance works in Notion. So, final tips again. Think about architecture. Think about open weight. And build value that transcends tokens. So we're going to depart Token Town. I know I said welcome to Token Town. We're going to spend the next 10 minutes really thinking about what to do next. So I think the challenge of the next six months doesn't have to do with capabilities. I think it has to do with security. Let's start there. There's this concept called the lethal trifecta. Simon Wilson, I think, crafted this.

7:46

If you have access to private data, exposure to untrusted content, whether it be through ingestion, MCP, email, right, and the ability to communicate externally, and that can include payloads and a web search, the second you have that system, you're exposing risk. And in fact, the more autonomous your system is, the more unsupervised this risk is. I think that this is what builds valuable product, not just capability. Same with sandboxes and computers. We talked about this, but it really is something that builds better determinism in your product, and also better token economics for your customers. And multi-agent orchestration.

8:08

Understanding what agents see and do, and what persists. I think persistence of enterprise knowledge is something that's actually really not discussed enough. It's starting to be with some recent launches. There is audio. So, don't have your workflows look like this. And I think this is where most software factories are today, right? It's actually that your entire engineering time just spends time babysitting the factory. I get it. Ours started off like this. Agent orchestration is one of the most difficult tasks of making factories work. So, okay, this is me telling T-Pain to tell people to buy Notion AI.

8:35

And the reason I included this slide is that I'm going to sell Notion for a second. It's my job. Always be closing. Always be selling. Always be hiring. Come find me. But I'm going to talk for a second about how Notion does this. Today, we already have the ability to inspect tasks. And you can imagine any tasks that you look at in a Notion document, you can have Claude actually go ahead and scope out what you need. We've launched this manage agent capability today. So if I go ahead to the top of this task, I can actually ask Claude Agent to scope out the task. Right? Ideally, it's working. And you'll see it'll actually populate an entire spec of what needs to be done.

9:28

In this example, it's not ready. It's going to ask me a question. Keep in mind this isn't a markdown file. This is an active document. Let's say I don't actually know the question, and I go ahead and I ask my team what to do. Imagine that you can tag in your team into these systems. MJs are PM. So in this example, she doesn't know. Usually she does. But multi-agent orchestration is important. Maybe Claude Code isn't the best at customer voice, but Decagon is. Right? You can ask Decagon agents. We're proud partners with them as well. To collect the right data that you need. Okay. In this example, we think we know enough.

10:43

We're going to go ahead and actually iterate through some of this flow. I'm going to skip ahead a little bit. We asked our TL what we needed. He replied, again, it's a collaborative file and not just a markdown. And we can have Claude actually go ahead and spin up the PR. Hopefully this is looking a little familiar now. This is the vision of software factories. That's what we're trying to host. Okay. Claude put up a PR. Maybe that's not enough. Maybe I want to go ahead and ask Codex what it thinks. Great. Found two issues. You can think about this scaling in an actual factory.

11:53

So today in Notion, you're actually able to orchestrate these agents together. And you're not committing to a lab. You're committing to the concept that AI is augmenting and automating what you do. This is real. I asked Rajiv if I could post this. This is how it works today internally at Notion. Almost all of our polish and large feedback like this is actually coordinated through our software factories, both in terms of routing to the right teams and also having coding agents take the first step. Vercel does this as well. Vercel does this as well. From staging to shipping to closing. And we see massive ROI gains from our customers.

12:50

That's over three minutes saved on a given task. Imagine that at scale. So I think we're trying our hardest to think about the factory lens. We cannot do this without optionality. And we cannot do this without conviction that we understand what models are required for which tasks. It's really wild out there. I get it. The market is really young. It's exceptionally opaque. It's moving fast. I'm super grateful for communities like AI Engineer to bring us together and talk openly about these things and how we navigate it. I think we owe it to all of our customers to get it right and to be critical thinkers about how we navigate this together.

13:45

I'm chronically online, fortunately. You can always DM me on Twitter. You can email me. You can find me after this. But thank you for yapping with me and thinking about this problem, and have a good day.

14:06

If you dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare

14:07

You If you were to look at just latency of a single call or just cost, Parallel might not be the cheapest. But if you have expertise in entire web search trajectories, you'll see how it differs. The granularity of this eval is what lets us make the best decisions for our customers because we understand all of the trade-offs on entire trajectories, not just single calls. Switch fast and often. I think we talked about that. And give them something back. That expertise on use cases is also very valuable to Frontier Labs. We find that our evals and our early access program partnerships actually help us a lot with Frontier Labs

14:49

and is something that we can exchange instead of extraordinarily large commits. And I don't think the discount is ever worth the loss in optionality. That's a perspective you can choose to keep or not. The second option is moderate tasks. Understanding open weights placed there. Open weight models are really strong enough to handle these tasks. And the possibility to RL on top of them has also kind of expanded the up-market growth that they can cover. I view open weight models as basically lowering the barrier to entry on cost for our customers. And they also give you negotiation leverage.

15:25

So it's kind of a credible alternative that's putting that downward pressure on pricing that if there's an oligopoly of two or three providers at the top is unavailable right now otherwise. I think Kimi 26 was probably the first time that we really saw a model that outperformed 5.2, GPT 5.2, GLM 5.2, now is another 5.2, Bombshell in the Villa that also probably does best here. But it's no longer the case where open weight models are good for just SFT on small tasks. Really think about without RL if they're capable enough for what you need.

16:00

And again, don't just think about external benchmarks. Be able to have expertise on your system. What are your tool errors? What's the actual latency that you need? Right? Here's an example of a benchmark that we posted. It's a little bit stale on purpose. Right? But you get the idea. Phillip at Base Tenslow showed this slide once and I've stolen it ever since.

16:27

Thank you. Are you here? Buddy. Okay. We'll chat. Hi. Well, he could come up and say it better. But the idea is that you don't have to be at the top. Right? I'm not trying to make a case that open weight is the best model out there. The case being made, however, is that the gap gets covered eventually. So if the tasks that you're having today are good enough, then in six months they're probably covered by open weight. So be prepared now. And the last thing is CPUs over GPUs. We've recently launched something at Notion called workers. I don't think that the GPU is necessary for every job. A lot of the jobs that we have are actually serving discrete pieces of code.

17:12

Like you don't need an LLM to turn a CSV into a PDF. You don't need an LLM to talk to Notion tool calls if we have a CLI. You definitely don't need an LLM to do deterministic SQL queries. This is where people become token poor very quick. And I think the last option here, besides open weight, CPUs, and optionality, is actually governance. There's a lot of AI governance. One is visibility. Understanding who's using the data. Understanding its maintainability and control. When you have model optionality, you can offer a lot more to your customers. Here's an example of how that governance works in Notion. So final tips again. Think about architecture.

17:56

Think about open weight. And build value that transcends tokens. So we're going to depart Token Town. I know I said welcome to Token Town. We're going to spend the next 10 minutes really thinking about what to do next. So I think the challenge of the next six months doesn't have to do with capabilities. I think it has to do with security. Let's start there. There's this concept called the lethal trifecta. Simon Wilson, I think, crafted this. If you have access to private data, exposure to untrusted content, whether it be through ingestion, MCP, email, right? And the ability to communicate externally, and that can include like payloads and a web search.

18:34

The second you have that system, you're exposing risk. And in fact, the more autonomous your system is, the more unsupervised this risk is. I think that this is what builds valuable product, not just capability. Same with sandboxes and computers. We talked about this, but it really is something that builds better determinism in your product, and also better token economics for your customers. And multi-agent orchestration. Understanding what agents see and do, and what persists. I think persistence of enterprise knowledge is something that's actually really not discussed enough. It's starting to be with some recent launches. You know, oh, there is audio.

19:19

So, don't have your workflows look like this. And I think this is where most software factories are today, right? It's like actually your entire engineering time just spends time babysitting the factory. I mean, I get it. Ours started off like this. Agent orchestration is one of the most difficult tasks of making factories work. So, okay, this is me telling T-Pain to tell people to buy Notion AI. And the reason I included this slide is that I'm going to sell Notion for a second. It's my job. Always be closing. Always be selling. Always be hiring. Come find me. But I'm going to talk for a second about how Notion does this.

19:53

Today, we already have the ability to inspect tasks. And you can imagine any tasks that you look at in a Notion document, you can have Claude actually go ahead and scope out what you need. We've launched this manage agent capability today. So if I go ahead to the top of this task, I can actually ask Claude Agent to scope out the task. Right? Ideally, it's working. And you'll see it'll actually populate an entire spec of what needs to be done. In this example, it's not ready. It's going to ask me a question. Keep in mind this isn't a markdown file. This is an active document. Let's say I don't actually know the question and I go ahead and I ask my team what to do.

20:40

Imagine that you can kind of tag in your team into these systems. MJs are PM.

20:50

So in this example, she doesn't know. Usually she does. But multi-agent orchestration is important. Maybe Claude Code isn't the best at customer voice, but Decagon is. Right? You can ask Decagon agents. We're proud partners with them as well. To collect the right data that you need. Okay. In this example, we think we know enough. We're going to go ahead and actually iterate through some of this flow. I'm going to skip ahead a little bit. We asked our TL what we needed. He replied, again, it's a collaborative file and not just a markdown. And we can have Claude actually go ahead and spin up the PR. Hopefully this is looking a little familiar now.

21:29

This is kind of the vision of software factories. That's what we're trying to host. Okay. Claude put up a PR. Maybe that's not enough. Maybe I want to go ahead and ask Codex what it thinks.

21:47

Great. Found two issues. You can think about this scaling in an actual factory. So today in Notion, you're actually able to orchestrate these agents together. And you're not committing to a lab. You're committing to the concept that AI is augmenting and automating what you do.

22:07

This is real. I asked Rajiv if I could post this. This is how it works today internally at Notion. Almost all of our polish and large feedback like this is actually coordinated through our software factories. Both in terms of routing to the right teams and also having coding agents take the first step. Vercel does this as well. Vercel does this as well. From staging to shipping to closing. And we see massive ROI gains from our customers. That's over three minutes saved on a given task. Imagine that at scale. So I think we're trying our hardest to think about the factory lens. We cannot do this without optionality.

22:46

And we cannot do this without conviction that we understand what models are required for which tasks. It's really wild out there, you guys. I get it. The market is really young. It's exceptionally opaque. It's moving fast. I'm super grateful for communities like AI Engineer to bring us together and like talk openly about these things and how we navigate it. I think we owe it to all of our customers to get it right and to be critical thinkers about how we navigate this together. I'm chronically online, fortunately. You can always DM me on Twitter. You can email me. You can find me after this.

23:22

But thank you for yapping with me and thinking about this problem and have a good day.

23:32

If you dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare dare

23:52

You

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note