20VC with Harry Stebbings

OpenRouter CEO: Why Chinese Open Models Are Beating the US | Why Enterprises Fear OpenAI & Anthropic

2238 summary words 10 min summary Watch video

Start with the signal

10 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: OpenRouter CEO Alex Atallah argues that AI will remain structurally multi-model, making independent routing, inference-provider competition, portable safety controls, and cost-aware orchestration enduring infrastructure layers rather than temporary features.
  • Why it matters: The interview offers directly reusable thinking for agent/control-plane design: route at the task level, separate deterministic work from frontier reasoning, treat provider behavior as variable, and avoid strategic dependency on any single model lab.
  • Best use: Use it as a strategic architecture and vendor-risk briefing, then translate its multi-model, subagent, safety, memory, and inference-cost arguments into OpenClaw's routing and governance design.

Executive Summary

Atallah's central case is that neither a single frontier model nor an enterprise's proprietary fine-tune will eliminate demand for a broad model marketplace. New models trained on different data and optimized for different tasks create continuing incentives to test, combine, and substitute models for quality, creativity, cost, and resilience. OpenRouter's role, in his telling, is not merely API aggregation: it is continuous routing across providers whose speed, quality, uptime, and effective token economics materially differ.

The most operationally useful argument is for hierarchical model architecture. A strong frontier model should orchestrate uncertain or open-ended work, while cheap open-weight models handle bounded, deterministic subagent tasks such as classification. Atallah says routing changes continuously as provider performance, pricing, and latency shift, and argues that model selection should be treated as an ongoing systems problem rather than a one-time vendor choice.

He identifies two competing enterprise risks. Startups that are primarily a thin workflow wrapper around intelligence can become targets when a model lab chooses their user segment as strategic, as he speculates Anthropic's Claude Design may do for design teams. Conversely, enterprises can be more uncomfortable with frontier vendors than Chinese open models because closed APIs create uncertainty over prompt retention, access, hosting, and data use. His answer is governed access rather than blanket bans: route broadly but apply centrally managed controls such as prompt-injection detection and PII redaction.

The interview also frames the open-model race as strategically important. Atallah believes Chinese open-weight ecosystems have structural advantages through capital concentration, policy flexibility, strong researchers, and national-champion dynamics, while US open-model labs face harder financing and business-model constraints. He sees distillation and broader access to compute as key tools for building a more competitive US open ecosystem, but notes unresolved concerns around Chinese model cyber posture and censorship.

Key Takeaways

  • Claim: Multi-model use is an enduring equilibrium, including for companies that train proprietary models, because external models continuously offer new capabilities, data exposure, cost improvements, and creative diversity. | Evidence: Atallah argues that a company-specific model must still keep pace with new lab releases and that combining differently trained models can yield ideas unavailable from one model; OpenRouter launched 70 models in July, roughly one every 10 hours. | Implication: Ken should design agent systems for substitutability and experimentation rather than assume one chosen foundation model or proprietary model can be the permanent universal default. | Caveat: Developers do exhibit model retention when an existing app works, its outputs are trusted, and migration risks or eval failures outweigh the benefit of switching.
  • Claim: Inference providers are not interchangeable commodity capacity: the same open-weight model can differ substantially in effective quality, latency, price, and uptime depending on how it is served. | Evidence: OpenRouter continuously benchmarks models across providers and sees materially different results even on static benchmarks; Atallah cites Moonshot's Kimi K3 provider comparison and says routing decisions can change every five minutes. | Implication: Routing should incorporate provider-level evaluation, health checks, fallback policy, and effective-cost measurement—not just a static model-name-to-provider mapping. | Caveat: The market is presently supply constrained, which supports provider differentiation and margins; this could evolve if capacity becomes abundant.
  • Claim: The high-leverage agent architecture is a capable orchestrator delegating well-scoped deterministic work to inexpensive subagents, often running open-weight models. | Evidence: Atallah recommends a frontier orchestrator for unknown, non-deterministic work and low-cost models for tasks with known output shapes, such as text classification; OpenRouter offers a subagent server tool built around this pattern. | Implication: For OpenClaw, explicitly classify tasks by uncertainty and verification requirements, reserve premium inference for planning/judgment, and make cheap parallel subagents the default for structured subtasks. | Caveat: The division of labor depends on having tasks that are genuinely bounded and evaluable; long-horizon and cyber-capable work remains comparatively stronger in frontier models.
  • Claim: Enterprises should manage broad model access with centralized safeguards rather than ban models wholesale, because unrestricted employee experimentation creates security and governance concerns but blanket prohibition forfeits useful capability. | Evidence: OpenRouter provides click-enabled prompt-injection protection and PII redaction across inference, and Atallah compares models to internet access: companies need guardrails rather than an outright ban. | Implication: A model gateway should be a policy enforcement point: log usage, redact sensitive inputs, screen for injections, segment permitted models by data classification, and maintain an explicit removal/escalation process. | Caveat: Atallah explicitly says he cannot know internal practices at organizations such as Moonshot or Alibaba; his safety posture is to follow US best practices and remove models generally considered unsafe.
  • Claim: The main strategic dependency risk is not only technical lock-in but model labs moving up the stack into applications or team workflows that become strategic to them. | Evidence: Atallah says a startup built mainly on integrations and system prompting is vulnerable if a lab targets its vertical; he interprets Claude Design as potentially strategic because it gives Anthropic influence with design teams inside companies, even if direct revenue is limited. | Implication: Ken should avoid businesses or internal platforms whose sole differentiation is a thin interface over one lab's capability; build durable workflow data, controls, proprietary context, distribution, and cross-model portability. | Caveat: He does not claim current evidence that Claude Design is materially cannibalizing Figma, noting Figma's reported financial performance remains strong and that he has not heard a broad repeat-use story.
  • Claim: Falling token prices can expand rather than shrink inference revenue and usage, but organizations need a new operating model for dynamically managing AI spend. | Evidence: Atallah cites OpenAI's GPT 5.6 Luna price falling 10x on OpenRouter over two weeks, followed by 13x usage growth, as a near-Jevons-paradox example. He argues employees' effective costs now vary with their chosen tools and models. | Implication: Build per-task, per-agent, and per-user inference-cost observability tied to outcome quality; treat routing and budget controls as dynamic operational management rather than a fixed software bill. | Caveat: He acknowledges confounders in the Luna example, including concurrent competition from DeepSeek and GLM, so it is an illustrative operating observation rather than a causal market proof.
  • Claim: Chinese open-weight models are advancing faster than US open alternatives, and Atallah expects the gap could widen without better US access to capital, compute, and talent support. | Evidence: He calls America 'very, very behind,' describes GLM 5.2 as a major open-weight step, and attributes China's advantage to concentrated support around national champions such as DeepSeek, fewer constraints, and strong research talent. He proposes distillation from permissively licensed open models and wider access to NVIDIA, TPU, Trainium, and newer-chip compute for US labs. | Implication: Use Chinese open models as serious benchmark and cost options where policy permits, but retain vendor diversification, security review, reproducible evals, and data-boundary controls rather than treating capability leadership as sufficient qualification. | Caveat: He flags unresolved security and governance questions around Chinese models, including cyber posture and potential future censorship constraints; he does not present a verified assessment of their internal operations.

Detailed Brief

Where routing, memory, and harnesses sit in the stack

  • Claims: Atallah rejects the view that routing is easily commoditized: focused routing companies accumulate operational advantages through marketplace coverage, benchmark data, provider relationships, capacity management, and failover behavior.; Memory may become a retention mechanism, but no one layer can own all useful memory. Apps hold the richest application context; model labs may achieve better personalized model performance; infrastructure layers will also try to provide reusable memory.; Harnesses are differentiated from ordinary apps by composability and inspectability: a harness can call or spawn another harness in a sandboxed, Unix-oriented environment that models already understand well.
  • Evidence: OpenRouter's pay-as-you-go take rate is stated as 5.5%, but it offers committed-spend enterprise pricing without fees on that committed spend and removes the fee when customers bring their own inference or keys.; He says prompt-heavy harnesses may be simplifying rather than expanding: Anthropic reportedly found that removing system-prompt material reduced later contradictions, and current harnesses are deleting code as frontier models become more capable.; He contrasts a harness's Bash/Unix-like composition with app orchestration, where an agent may need to launch a browser, authenticate, find an API, read documentation, and navigate many unknowns.
  • Caveats: The optimal location and combination of memory are unresolved; duplicating memory across model, app, and infrastructure layers could create confusion rather than personalization gains.; Focused router superiority is an assertion from a participant with direct commercial interest in the routing layer, not an independently validated competitive analysis.
  • Implications: Keep durable user and workflow memory portable at the application/control-plane layer, with clear provenance and selective sharing into model contexts.; Prefer tool interfaces and agent harnesses that expose deterministic, inspectable primitives over brittle browser-driven composition where possible.

Open ecosystem dynamics and model-lab business incentives

  • Claims: Atallah expects model proliferation to continue as agent companies develop proprietary models to distribute through their products; he names Cognition and Cursor as examples and notes Jeff Dean is starting an agent lab.; He sees open-model distillation as a standard model-building technique rather than inherently improper, while accepting that closed labs can contractually prohibit competitors from using their outputs.; He expects some neo-lab consolidation but rejects the proposition that 70% will die in three years; if acquisitions count as failure, he estimates roughly 50% could consolidate.
  • Evidence: He says models such as Sonnet are partially distilled from larger models such as Opus, framing distillation as a routine method for producing smaller specialized models.; Poolside is singled out as an underrated US lab with small but effective coding models and useful access tooling.; Atallah says NVIDIA seeks to avoid GPU customer concentration and prefers a heterogeneous compute market, which he argues helps preserve an independent inference-provider layer.
  • Caveats: The claim about a lab's use of distillation and his estimates of neo-lab survival are presented without supporting data in the interview.; The alleged Stripe acquisition offer and OpenRouter's valuation are raised by the host but not confirmed by Atallah.
  • Implications: Do not equate a model's open weights with unrestricted downstream use; track licenses, output-use terms, hosting policy, and data residency separately.; Treat agent vendors that begin training models as both potential partners and potential sources of vertical integration risk.

Notable Concepts & Terms

  • Neurodiversity in AI: Atallah's label for a diverse model ecosystem whose different training data, behaviors, and perspectives produce better aggregate creativity and reduce dependence on a single intelligence source.
  • Inference provider layer: Companies such as Fireworks and Together that host and optimize open-weight models; the interview argues this layer affects more than price because serving quality and reliability vary materially.
  • Central router: A model gateway that dynamically selects model/provider combinations based on changing price, speed, quality, capacity, and availability rather than fixed application configuration.
  • Jevons paradox: The proposition that lower unit cost can generate more-than-proportional consumption; Atallah uses Luna's reported 10x price reduction and 13x usage increase as an AI inference example.
  • Harness: A composable agent execution environment, often Unix-oriented, that can invoke other harnesses and exposes inspectable operations more reliably than generic app/browser automation.
  • Subagent architecture: A system in which a stronger orchestrator delegates narrow, low-cost, deterministic tasks to specialized models, then integrates their outputs into broader reasoning.
  • Distillation: Using a stronger model's outputs to train or reinforce a smaller or specialized model; presented as a normal technique, subject to source models' contractual output-use restrictions.
  • Dynamic employee cost: Atallah's proposal to measure an employee's effective cost partly through the variable inference spend generated by their model/tool choices, alongside their productivity and output quality.

Operator Notes / Why Ken Should Care

  • Create a routing policy matrix for OpenClaw that maps task type, uncertainty, data sensitivity, latency target, and evaluation confidence to model/provider tiers.
  • Implement provider-level continuous evals and health signals; measure quality drift, latency, error rate, cost, and fallback outcomes separately for each model-provider pair.
  • Establish a gateway policy pack before expanding model access: PII redaction, prompt-injection screening, model allowlists by data class, request logging, and an incident-driven model disable switch.
  • Keep user/workflow memory application-owned and portable by default; define what context may be sent to a model or provider rather than letting model-specific memory become the system of record.
  • Instrument agent and employee inference spend against task outcomes, but avoid using raw token cost as a performance proxy without quality, leverage, and task-complexity normalization.
  • Run a benchmark track that includes leading Chinese open-weight models alongside US frontier and open alternatives, with explicit security, hosting, licensing, and data-residency gates.

Source/Metadata

  • Title: OpenRouter CEO: Why Chinese Open Models Are Beating the US | Why Enterprises Fear OpenAI & Anthropic
  • Transcript words: 10639
  • Duration seconds: 4105
  • Timestamp note: No usable timestamps or chapters were present in the supplied transcript.
Full transcript 10110 words · 47 min read
0:00

It's going to be the biggest, biggest market in tech ever. A lot of companies are making routers because it's fashionable. The model apps have several incentives to go after you eventually. Today, we have Alex Atala, co-founder and CEO of OpenRouter, the unified interface, the gateway to the world of LLMs. They reportedly have had offers from Stripe for $10 billion. They've raised at a valuation of over a billion and a half. They are the market leader. And this interview could not come at a more prescient time. In July, we launched 70 models, about one model every 10 hours. America is very, very behind. Still, GLM 5.2 was a really big, big step for open weight models.

0:40

There are reports that you are selling to Stripe for $10 billion. Is that going to happen? Ready to go?

0:59

Alex, I am so excited for this, dude. I have wanted to make this one happen for a while. I've heard so many things from Matt at Menlo. I've stalked the shit out of you speaking to Anjini, even your roommate before this show. So thank you for joining me, dude. Thank you. It's great to be here. Now, I want to start with a little bit pre-OpenRouter and start on OpenSea. It was a pretty incredible journey. What did you take with you to OpenRouter, having seen all that you saw with OpenSea? Yeah. So OpenSea started as the first NFT marketplace. And similar to OpenRouter, it was very small for a long time.

1:40

We kept the team very small until the Series A, roughly, or a little bit afterwards. And this was before AI. So right after NFTs started blowing up in 2020, October of 2020, we were like, oh my goodness, we are understaffed. The servers are melting. All kinds of our search index was exploding. We had a couple big outages. It was tough to keep the site up. And it was like, oh my God, we're going to become the Twitter fail whale, but applied to crypto. My biggest goal was to have us not be the Twitter fail whale for crypto. And it took a little bit to create the team, get platform and infrastructure under control, make sure we could predictably scale.

2:27

In other words, do load testing to help the site sustain 10X load, even when we weren't seeing that load. Because with crypto, you just don't know. There were moments where we would get these incredible traffic spikes. And it would be very dependent on the content and the community. And so I built a lot of infrastructure and scaling responsibilities then that I took to OpenRouter and spent a lot of time thinking about, okay, how do we make something that is going to be always up and that people can really count on from an infrastructure point of view, even when there are huge surges in really tumultuous markets?

2:46

Which has been very helpful for AI, of course, because all companies, especially Anthropic, have seen unpredictable growth. And we have as well. And we've had a couple bumps, but overall it's been significantly better. And OpenSea just drilled that into me in a way where I could take it productively to OpenRouter. Can I ask you, when you go back to the founding thesis of the company, what has happened in the ecosystem, in the model landscape, that you did not expect to happen? Okay. Well, one thing that we did not expect was that an ecosystem of companies would emerge to host and serve the open weight models.

3:29

Early on, it wasn't clear that that market wasn't going to be a monopoly where just the three hyperscalers serve all the open weight models and startups don't, they're really far behind. In reality, how often do you hear people running GLM on a hyperscaler? Never. They're using the inference providers like Fireworks and Together. And there's a big list that we see doing the best job of hosting all the open weight models. And in the early days, we had, I think we called it Provider 1 and Provider Fallback. We didn't show which providers were actually doing the hosting because we weren't really a marketplace. We were an exploration tool for finding and discovering new LLMs.

4:00

And we wanted to build a marketplace of model labs, but the inference provider layer, we weren't sure would actually be a marketplace. And it turned out that those companies were doing a way better job than the hyperscalers, were way faster to host the models and figure out these edge cases to hosting them. And uptime was just going to be a constant problem. It wasn't going to magically get solved by the supply side of the market. A lot of people suggest that that inference provider layer is a commoditizable element or layer that will be removed or see margin reduction competed out over time. What would you say to that theory?

4:33

Right now we're in a massively supply-constrained market where, and it's likely going to be supply-constrained for a while, all the inference providers are short, pretty much constantly short. And you're like, okay, so GPUs are really, really beneficial. And why doesn't Google or Amazon or Azure run around and buy up all the GPUs and take all these inference providers out of business? Well, the people making the GPUs don't want that. One of NVIDIA's top priorities is not having customer concentration. They want lots of customers to all have separate allocations of GPUs. They want the heterogeneity of the market. They want competition on the compute layer.

5:20

And this is good for the ecosystem. Users also want this. It's good for NVIDIA, and it's good for end users as well. It allows these inference providers to come up with new innovations on how to serve the models better. Even a single model like Kimi K3, Moonshot just posted a benchmark showing all the inference providers and how well they're serving Kimi K3. And they're pretty different numbers for benchmarks that are really static, that are well known. We post this continuously all the time. We always are benchmarking all of the models on all of the inference providers, all the open weight providers, and finding really different results constantly.

5:57

The results change over time. These models are very emotional. They're very non-deterministic. I had Lynn on the show from Fireworks, and she said that, you know, I said about Gavin Baker, “A token is a token,” is what he said. And she kind of corrected me that a token is not a token, actually, because one provider can make a token go so much further than another token. It's like, how do you get to the store where you can drive around the whole block or you can drive straight to the store? Tokens can be made more efficient and go further. And that's the job of the provider. Yeah, I agree with that.

6:48

I think that, in some ways, we are providing a service to help people discover providers. And ultimately, when one provider is making a token go further, we spend an enormous amount of time on our router, central router tech, so that that provider immediately gets more traffic. As soon as we detect that there’s a quality improvement or a speed-up or a price reduction happening, it immediately starts getting more traffic. And this stuff happens 24/7 every single, every five minutes. There are big changes for the big models. And so it actually does make the experience better.

7:20

You can only invest in one inference provider. Which one do you invest in? I probably have to stay neutral on this. I really like the inference providers that are doing custom hardware and very, very low-level optimizations. I like providers that are also trying to figure out how to make customization easier. So today you fine-tune models and you create this new, fully independent model from the base model.

7:55

Many inference providers are creating these LORAs, or some call them cartridges, that are much more portable potentially between models. And we might see a future where, when you do a fine-tune and you want to change the base model layer, it only costs maybe a few hundred dollars, maybe a few dozen dollars, to change it. It's okay. I understood that Firewise is your favorite. It's okay. I get it. Mine too.

8:36

My question is that when Lynn was on the show, she was like, oh, you don't want to rent your own rent intelligence. You want to own it. And we're going to see companies have specialized models, which is trained on their own data and proprietary to them. In a world of every company having specialized models, that's really tuned to them and their preferences, is that good for an OpenRouter business or not? Oh, definitely. I mean, our goal is— Because you'd stick on one model, which is yours, proprietary, trained on yours, and not be open to the diaspora of models that is available. No, I disagree. It's okay. I understood that Firewise is your favorite.

9:48

It's okay. I get it. Mine too. My question is that when Lynn was on the show, she was like, oh, you don't want to rent your own rent intelligence. You want to own it. And we're going to see companies have specialized models, which is trained on their own data and proprietary to them. In a world of every company having specialized models that's really tuned to them and their preferences, is that good for an open route of business or not? Oh, definitely. Our goal is- Because you'd stick on one model, which is yours, proprietary, trained on yours, and not be open to the diaspora of models that is available. No, I disagree.

10:33

I think our mission from the very beginning has been to increase neurodiversity and AI for the whole ecosystem. And we really believe that a multi-model future is inevitable. When you start, let's say there's one model that, hypothetically, let's say you're right. Let's say there's one model that fulfills all of your desires, either within your company or as a consumer. And more and more people start using that model. And then someone decides, you know what, I'm going to create a neurodivergent model.

10:57

I'm going to create a model that's a little bit different, that talks a little differently, that has ideas that the first model could never have come up with because it's completely different data that's being used to train it. But then it creates inevitable demand to use both models. But creativity is not verifiable. You can't really put an easy number on creative ideas. And when you use two models together, you're more likely to get creative ideas than if you just use one. It's just a fact if that other model was trained in a different way on a different data set or has made a big update. So consolidation on one model just seems like it doesn't make any sense to me.

11:39

Totally get you. So you'll have companies which have a core workflow or their core, which is their own specialized model. And then they'll use a plethora of other models, and they'll use Open Router for those other model selection. Yes. And I think that when companies make, to get back to your question, their own model trained on their own data, the ecosystem around you is all doing the same thing. You have to play out the game theory for these things a little bit.

12:20

If everybody is doing this as well, and all the model labs are creating new models constantly using new data that they've acquired, that they've bought from other companies, that's all potentially data that's valuable to you. What is in your best interest? It's to go and try out those other models and see if you can be more productive with them. If you can merge them together to get better state-of-the-art performance. If you can reduce your costs using these other models. Whether your goal is to reduce your or improve your margins or grow your company, you are incentivized to go use what the ecosystem creates.

12:48

So the model that you made, you're going to have to continuously improve it to keep up, and it's never going to win the whole market. So it's going to be a massive market. That's going to be the biggest market in tech ever, and the biggest market probably in human history. No one's going to win all of it. You're not going to build a model that wins all of it. So you might as well build a model that is known to specialize in something very useful, and that's very important to your company and your business, and be known for that specialty. And I think a lot of enterprises are going to move that direction, make their own models, make their own branded intelligence.

13:34

Your brand is a big part of your moat. And that model will be a way your brand carries around. You mentioned the immense time that you spend on the routing technology that you have. A lot of people are thinking that we're seeing the commoditization of the routing technology. You're seeing ramp release products like this. I mentioned earlier a merge company. We invested in release that product. And several are releasing routing technology similar or claiming to be similar. Are we seeing the commoditization of this layer? I think a lot, yeah, a lot of companies are making routers because it's fashionable.

14:27

I think they're seeing growth happen here, or they're making gateways at least. First, I think there's two issues with that. First, it immediately puts you in the mindset of copying instead of winning something. You're playing to exist rather than playing to win. And maybe you're just trying to serve your existing customer base and you want to see some AI growth happen. I think it immediately puts that gateway many, many months behind the companies that are fully focused on it. I am 100 percent focused on building the best router and gateway and an LLM marketplace.

14:50

And it shows in our product and the benchmarks that we create internally and how we see ourselves compared to the competition. And this is not a side quest for us like it may be for some other companies. The other problem is that it reduces the leverage of all of your users. So I really deeply believe in giving users and developers more leverage.

15:15

Fundamentally, giving them access to more models is about giving them more leverage over all the innovations that happen in AI. You want to be able to access them all. You want to reduce your dependency on any individual one. If you build on top of a router or a gateway that doesn't give you access to the full market or full flexibility or full customizability, it doesn't give you the full leverage of the whole ecosystem. Then you're being cut out. You're cutting out all your employees at your company from things that they need. And so open routers fundamentally about giving people more choice because that gives them more leverage.

16:06

You do that at a price at five point five percent take. That was our pay-go plan. We then added an enterprise plan with a totally different pricing model. And it's been very successful so far. It's based on committed spend and then no fees on that committed spend. Because that was going to be my question. Ultimately, companies will love it small. And then as you scale, you're like, shit, this is really freaking expensive. I'll just build my own routing tech now because it's become such a significant part of my cost base, actually. I kind of figured some of those companies haven't realized we have an enterprise plan.

17:11

And some of it is our fault for not having a better, more detailed pricing model. We're soon going to introduce a business self-serve plan that also makes a lot more sense. And if you have your own inference, if you bring your own inference to open router, if you bring your own keys, that fee goes away. So it's fairly, for inference that we are providing you, when you go into open routers capacity and you're not on our enterprise plan, that's when that fee comes in. Otherwise, we need to be able to predict demand a little bit. So that's why we do these committed spend. What will be the main revenue line of open router in three years' time?

17:39

I think it's going to depend on the economy in so many ways. If the overall AI market keeps growing the way it's been growing over the next four years, with 10 to 15 X every year, or potentially more, it's a lot of growth. I think under that world, I would expect people to continue to underestimate how much inference they're going to need. And thus our revenues are going to be dominated by the same things that dominated today, which is us helping people with unplanned inference capacity, both enterprises and startups. And that's what open router is best at.

18:17

When you need to try models that you weren't expecting you need to try, when you're using more inference than you thought you were going to use on particular models, we make sure that that is not going to be an issue for your company by providing the best failover and best uptime. And this is really, really a good thing to do when the market is continuously underestimating its inference needs and growing at this rate. If this growth rate continues over the next four years, it's going to be a wild amount of growth. The economy has some limits to it. And thus, are our revenues going to be dominated by the same things that dominated today? That dominated it today.

18:51

Which is us helping people with an unplanned inference capacity. Both enterprises and startups. And that's what OpenRouter is best at.

19:10

When you need to try models that you weren't expecting you need to try. When you're using more inference than you thought you were going to use on particular models. We make sure that that is not going to be an issue for your company. By providing the best failover and best uptime. And this is really, really a good thing to do when the market is continuously underestimating its inference needs. And growing at this rate. If this growth rate continues over the next four years. It's going to be a wild amount of growth. The economy. And the economy has some limits to it. I can see major SMB SaaS growing for us.

19:53

And needing to grow for us whenever, if growth does not keep going 10x, 12, 15x per year. We've seen token prices fall 90% in 18 months. Is the reduction of token prices helpful or hurtful to your business? Because obviously you have a take on spend. If they come down and spend is more efficient. Seemingly, it's bad for your business. You have a shrinking pie to take from. Well, a lot of people talk about the Jevons paradox. But when prices go down by 10x, the usage increases by more than 10x. But no one has really done a great job modeling it. We do have a lot of spot stories that confirm it. For example, GPT 5.6 Luna on OpenRouter.

20:57

OpenAI cut prices by 5x and then, in coordination with us, by another 2x. So, in total, the price of Luna has dropped 10x on OpenRouter over the last two weeks. And guess how much usage has grown? 13x. So, it's a close to perfect Jevons paradox story where you drop prices 10x and usage grows by more than 10x. Just a bit more. And also, the usage is pretty stable. It grew. It flattened out at 13x. And then it's been growing at the same rate that it was growing before it hit the 13x multiple. So, that's pretty interesting. And it's a pretty low variable. There are a few other confounding variables in the story.

21:59

And it was also done in the middle of DeepSeq launching and having a really, really good price. And GLM having a really good price. Now Luna is being used more than GLM on OpenRouter. GLM used to be one of the top three, four models by token volume. And now Luna is past it. This is the first time OpenAI has had a model on our platform in the top three to five models by token volume in an extremely long time. So, it was a really big and interesting move. How reflective of the market are your token volumes? Because it's about, I may get this wrong, about maybe one and a half, two percent of token volumes. And so, how reflective are they?

22:39

Because a lot of people, when I say, oh, the top five models, when I look at OpenRouter, are all Chinese. What does that mean? They'll go, oh, well, Harry, no offense to OpenRouter. But it's not reflective of the market. And most people who use Frontier, it doesn't go through that. They use Frontier APIs, and so it's not counted. To what extent are your rankings reflective of true token usage? We try to estimate how they're off by just surveying people sometimes or looking at the surveys other people have done. I think we definitely have a bias toward people who believe our thesis, which is that the future is multi-model. And companies who want multiple models.

23:26

And there are still companies out there. I rarely run into them now. But there are still companies out there that are just like, oh, yeah, we're an OpenAI shop. We only do OpenAI models. And so we're not going to see any of those companies. And I think those companies are primarily focused on the hyperscalers, OpenAI, Anthropic, and Gemini. So we do probably undercount the Frontier models. But I think over time, our thesis is becoming more and more common to see in other companies. And the moment that they're like, oh, yeah, we need to use other models. Then our data becomes more representative. And as we scale up, the data becomes more representative in general.

24:19

So my hope is that it just becomes better and better data over time. Let me ask you, Alex Karp said on CNBC in his rather wonderfully energetic way that companies are terrified of working with Frontier model providers. Do you think they are? I haven't seen what he talked about there when I talk to our customers. But there was definitely a little skittishness, particularly when Claude Design came out around Figma. And that part I did see. And I do think that there are real concerns for a company that is building a thin wrapper around intelligence. Figma is very different. But if a startup is only building a go-to-market wrapper around intelligence.

25:07

Like, hey, we're a company that brings AI to this market and does so by doing the right integrations and customizing the system prompt. You're going to be fine if the model labs don't care about that market, which there will be many markets like that. But the model labs have several incentives to go after you eventually. One is getting multiple teams within companies they do care about to be dependent on them. This is my theory behind why Claude Design was strategic. While it's not a massive amount of revenue for Anthropic, probably not a significant amount of revenue, it does get the design team to really care about Anthropic models.

25:49

And so the companies that they want, they now have another team that really wants to stick to Anthropic. So that team strategy can make you compete with the model labs. And so I think companies like that, that find themselves like, oh, we're building a product for a team that has now become strategic for the model labs, for companies they actually care about. That's where I see probably the most near-term threat. Do you think Claude Design will have a meaningful impact on the Figma business? I speak to many founders today who are bluntly switching from Figma to Claude Design, and it's cannibalizing their Figma usage. Do you see that, and do you think that will happen?

26:53

So I saw a lot of designers try out Claude Design, including our own. But so far I haven't heard the repeat story. I don't know. Honestly, I have not talked to very many designers about this. I certainly haven't heard a lot of chatter about Claude Design. And if you just look at the numbers for Figma, they're quite good. They have very, very incredible earnings. So... This is why you don't want to be public, dude. You see great numbers. Figma down.

28:08

I'm like, poor Dylan. What? Yeah. That was crazy. Do you know what I mean? I was like, really? Come on. We were talking about the different models that we have on offer, and whether companies are willing to work with frontier models. The rate of model development feels immense. Do you think we will see the same rate of model development continue over the next year, two years, three years? Frontier model development or general model? A general model, both frontier and open. Yeah. Yeah. Just because every single day there's two, three, four new models. In July, we launched 70 models. About one model every 10 hours.

28:50

There's some agent labs starting, too, that are all going to probably make models eventually. Jeff Dean is starting an agent lab right now from Google. The companies that are known for making agents have an incentive to create their own model, a very clear incentive to create their own models and distribute it through the agent. And we haven't even seen the start of that. Sorry, we've seen the start of it, but we haven't seen it really pick up. Cognition has a model. Cursor has a model. Does Lovable have a model yet? I don't think so. Not publicly. Yeah.

29:34

Frontier model development or general model? A general model, both frontier and open. Yeah. Yeah. Just because every single day there's two, three, four new models. In July, we launched 70 models. About one model every 10 hours. There's some agent labs starting, too, that are all going to probably make models eventually. Jeff Dean is starting an agent lab right now from Google. The companies that are known for making agents have an incentive to create their own model, a very clear incentive to create their own models and distribute it through the agent. And we haven't even seen the start of that. Sorry, we've seen the start of it, but we haven't seen it really pick up.

30:34

Cognition has a model.

30:44

Cursor has a model. Does Lovable have a model yet? I don't think so. Not publicly. Yeah. So the agent labs are going to, I think, develop models. This pressure from both the GPU makers, like NVIDIA, to create more competition in the space and create more diversity in the space, plus us, plus investors who just want to try new things that all could improve intelligence in some neurodivergent way. I think those are strong incentives. I think they're enough to incentivize more founders to make neolabs. And if American open-weight models pick up steam, then it gives these neolabs a base to train on. That's not Chinese, which will then probably create more American neolabs.

31:25

Do you think we should be concerned by the rate and quality of Chinese open models? We should. We're behind. America is very, very behind still. I think things are picking up. And I think we have Poolside. We have Thinking Machines. We have RC. Do you feel a sense of responsibility for that? And what I mean by that is you are a routing business.

32:40

And you could route a company to a Chinese model that, who knows, people are worried about backdoors, ICCP involvement. You could be the deliverer of that to those models. Do you feel a sense of responsibility for that? So we do feel a responsibility to have safe access for all of these models. Customer trust is our paramount goal. If one of these models is unsafe to use, generally considered unsafe, we pull it from the platform. If there's a way to use it in an unsafe way, there's a way to use all the models in an unsafe way.

33:29

And then we believe in using technology to make it safe and to work with the model labs themselves to figure out how they're doing it on their side so that we can be state-of-the-art or better. We spend an enormous amount of time making sure that our practices match the best things that we're seeing coming out of the labs or better. And because we're a very good way of exploring all the models and finding them for the first time, we're a good focal point for deploying safety measures across your whole company. For example, we have prompt injection protection. You can just turn it on and immediately flag prompts that look like prompt injection that's trying to happen.

34:01

We have PII redaction. We have a couple different things that you can automatically just turn on with a click and get an added safety layer on top of all of your inference. And we build that so that enterprises feel like they can safely deploy new models and that their employees can try them out. I think of the models a little bit like the internet. You can't just ban the internet at your company because there's some bad things on the internet. You can create guardrails, and you should. You need to use AI to build the best possible guardrails that you can. I'm with you, but do you think you actually know what's going on within Moonshot or Alibaba with Quan?

34:53

These are incredibly secretive organizations in the depths of China. Can't pretend I know what's going on inside of them. As a U.S. company, we're going to follow the best practices of what happens in the U.S. to make sure that we're not doing something irresponsible. What do you think U.S. companies are more nervous of, frontier models or Chinese models? I think they're more nervous about frontier models usually, in part because there's much more confusion around the data policy, about what's actually happening to the prompts that I'm sending and where they're being stored and how they're being looked at.

35:27

And you can't run them on your own machine or in a provider of your choice. And so that just immediately creates all of this uncertainty in a lot of enterprises. And it's uncertainty that they can also pattern match. It's very similar to running on their own infra versus running in their VPC and knowing who can see the data. How extraordinary is that, though? They're more nervous of U.S. companies headquartered in Silicon Valley where you can see and touch and feel the headquarters and the leaders. It's just, what a strange world to be in. Yeah, it is very strange, especially with the frontier models having the biggest cyber posture right now.

35:53

I like being the best at cyber effects. What do you make of every company kind of posturing, ha-ha, we hacked someone? First you had OpenAI, then you had Anthropik, and then you had Zark coming out. I don't want to miss the party. We did, too. Yeah. Well, I think they have to talk about it. The right thing to do is to reveal when there's been a cyber incident involving your model. Covering it up doesn't work. It's not going to work in the long term. And it certainly looks like they're all bragging about it.

37:06

But really, if you were in their position and something happened with one of the models and you had to make the choice about whether to publish it or not, I think the right thing to do is to publish it, regardless of how people are going to spin it. So I don't know. I doubt that they're actually thinking of the felony bench or whatever it's called. How significant was the latest Kimi model, which got so much attention? Was it as significant as everyone thought? It's quite good. It's not cyber capable in the same way the frontier models are. And for long-horizon tasks, I think it's still a bit behind the frontier models.

37:40

But GLM 5.2 was a really big, big step for open-weight models. Kimi was kind of like Moonshot getting up to that step. That's a little bit how I see it. And Kimi is also a very good writer. The voice and tone are both pretty good. Whereas some of the frontier models have voice degradation that happens when they get better at coding, especially. And I was like, oh my God. I can't read this output anymore. The output sounds like three of the four arguments you made are right. And one is a turning point. Here's the rub. It's sometimes just impossible to read what they're saying. And this stuff is fixable. But Kimi, I think, has always had pretty interesting writing.

38:25

In 12 months, will the chasm between U.S. open source and Chinese open source be bigger or smaller than it is today? So my fear is that it will be bigger. I wish they were. Bigger. Because when you have DeepSeek, it becomes a national champion in China. And Xi Jinping is going, this is our AI horse. I will concentrate all of my money and efforts behind this. And I will supplement this ecosystem to the end. This is the winner. And then when you see another Moonshot come out, suddenly all regulation gets moved aside. All policy gets pushed aside. All funding becomes available. Everything is allowed. You are free to run.

39:20

And these guys are unabridged in their ability to do whatever they want to get to the end goal. Whereas OpenAI and Anthropic and all the other providers in the U.S., especially open source, fuck, you've got to try raising billions of dollars for a U.S. open source model. Bit tough, actually. Not impossible at all, but tougher. Business model questionable. AI research is super expensive. And you're competing against OpenAI and Anthropic. I think the comparative landscapes they sit in mean that the Chinese open source providers are just inherently advantaged, sadly. And they have very, very good researchers. And I think Americans underestimate that a lot.

40:26

I do think they're going to be concerned about the cyber posture of their models. And then when you see another moonshot come out, suddenly all regulation gets moved aside. All policy gets pushed aside. All funding becomes available. Everything is allowed. You are free to run. And these guys are unabridged in their ability to do whatever they want to get to the end goal. Whereas OpenAI and Anthropic and all the other providers in the U.S., especially open source, fuck, you've got to try raising billions of dollars for a U.S. open source model. A bit tough, actually. Not impossible at all, but tougher. Business model questionable. AI research is super expensive.

41:02

And you're competing against OpenAI and Anthropic. I think the comparative landscapes they sit in mean that the Chinese open source providers are just inherently advantaged, sadly. And they have very, very good researchers. And I think Americans underestimate that a lot. I do think they're going to be concerned about the cyber posture of their models. And they do seem very concerned about censoring the models and censoring the information that the models can provide to people. So while today people complain about American models censoring more due to cyber, I'm not sure that's always going to hold.

41:16

And as the Chinese models grow in importance for China, what are they going to do? Are they going to drop the great firewall? Are they going to give up on putting the firewall around the models? I don't know that much about China, but it does seem strange that they don't seem to care more. That the models are, I've never seen anyone do a profile of what you can do with DeepSeek that you can't do with the internet in China that's available to you within the border. What information you can access. I've never seen anyone do a real deep dive. How far past the firewall does DeepSeek go?

41:45

If the firewall matters to China, if it's going to matter in 10 years, something's going to change. Well, what's interesting is, obviously, the abilities of the Chinese models outside of China are immense. The abilities of the Chinese models inside China are actually relatively limited. The guardrails. The guardrails are incredibly stringent and prohibitive. So it's ironic that they are incredibly superior to us. Shit domestically. Terrible. Interesting. I literally just had my dear friend Jason Lampkin, he runs SaaS, come back and be like, couldn't figure out what time Starbucks opened on DeepSeek.

42:56

Wasn't on offer. Would say, not allowed. Yeah, yeah. Wild. Wow. Very basic rudimentary requests. We're speaking about all of these different models. Yeah, yeah. And the thing I think is, what about loyalty? And you have this incredible seat in the ecosystem where you can see everything.

44:10

Do we see any developer loyalty today with models? Honestly, we do see some. We try to make switching costs close to zero so that when new models come out, people can try them out really easily. But we also measure retention and churn from all the models. We share this data with model labs, too, when they ask for it so that they can know, for my model that just came out, which models drove traffic to it? And for those users, when they leave, which models are they leaving to? And we'll make this more and more available to the world soon.

44:28

And we do notice in the churn data there are developers who continuously stick to models even when there are better models out there, better models for their use cases. I think it's a combination of a couple probable root factors. One is, my app works and I don't want to break it. If the support bot starts saying something weird that I didn't expect, why add more headache? I've already done all this optimization. And I've already put all these guardrails around it. Another is new models are not necessarily going to make your pricing better. In fact, in general, what happens is that the current models' price goes down over time.

45:21

And especially when new advancements in the labs happen, you'll see intelligence jump, but the price curve also jumps and then will start going down over time. So it's not necessarily the most price-effective thing to do to shift over to the newest model, even for open weights. The third reason is it's just fundamental trust in the outputs. If I'm using a model to do my work and I like the way it talks, I probably have some eval, a personal eval. A lot of people have these personal evals that are just these random tests that they give the models. And if the random test doesn't look really good on the new model, they'll just be like, good. I liked Kimi K 2.6 anyway.

46:14

People thought before that memory would be the retentive mechanism. And while OpenAI has all of my previous prompts, it knows that I live in London, I do podcasting, and that will make it a better model for me moving forward. Is memory no longer a retentive mechanism? Memory is really interesting. I've always thought it is a retentive mechanism. And the question is where it lives. Is it going to live with the model? Is it going to live with the inference provider? Is it going to live with the app?

47:00

Is it going to live with the infrastructure provider, the router? My guess is that all of those layers are going to try to own memory in different ways. And there are going to be advantages to sticking your memory in each layer. If you stick it with the app, then the memory has the most app-related context and is model agnostic. If you stick it with the model, the memory might perform the best on personalized benchmarks and perhaps have the best ultimate intelligence. And I think the model labs are going to work on memory. And then the ultimate thing might be, is there a good combination?

47:39

Can I use memory in the model and memory at the infrastructure layer or the app layer at the same time? Is that going to confuse the model? We don't know yet. I do think that it's impossible for one layer to capture all valuable memory because the apps own so much important context that the model labs don't have. And for the model labs to get this to work, they'll have to incentivize the apps to give them that context. Speaking of the apps and the model set, claw code, cursor, bundle, model, and harness. Is the router absorbed into the agent framework before it ever has the chance to be independent when you have the agent and the harness together?

48:19

The harnesses are pretty interesting because in our early days, one of our early bets was that most apps were underestimating the desire for users to choose the model. Most apps in the very early days, in 2023 and 2024, it wasn't even clear which model was being used under the hood. They were like, oh, people are not going to care about that. They just want AI. And one of our strong convictions then was that, no, people are going to want to use particular models. They're going to care about who they're talking to. It's like, I want to know which employees I'm talking to when I'm trying to solve a problem. And models will be kind of like that. And that has played out.

48:45

In Notion, you can choose the model that you talk to. Even though you would think an app like that might want to obscure it completely. A similar thing happened with harnesses where, particularly with developers, they started to build an affinity to different harnesses. And that's because it's a user experience. So I think that is my favorite argument for why harnesses are going to stick around. Not that they're being bundled with the models. Because, in fact, as models get better, they get more resourceful. And the junk that gets thrown in the system prompt just becomes a handicap.

49:32

Anthropic, I think, published a good article about this where they showed that, oh, we got rid of stuff from the system prompt. And suddenly fewer contradictions showed up later on with user prompts and the model performed better. And we're seeing a lot of the harnesses right now are deleting code in order to perform better with the latest frontier models. I don't think it means that harnesses are bad. In fact, I think we'll see more harnesses come up in the future because it's a way of building a user experience on top of models. It's a way for developers who are not model labs to own a user relationship.

50:00

And that is just going to be incredibly valuable for the economy to have that layer. I'm going to get killed for this. And the junk that gets thrown in the system prompt just becomes a handicap. Anthropic, I think, published a good article about this where they showed that, oh, we got rid of stuff from the system prompt. And suddenly fewer contradictions showed up later on with user prompts, and the model performed better. And we're seeing a lot of the harnesses right now deleting code in order to perform better with the latest frontier models. I don't think it means that harnesses are bad.

50:25

In fact, I think we'll see more harnesses come up in the future because it's a way of building a user experience on top of models. It's a way for developers who are not model labs to own a user relationship. And that is just going to be incredibly valuable for the economy to have that layer. I'm going to get killed for this. What's the difference between a harness and an app? It feels like it's word wank. Everyone's talking about harnesses and harnesses.

50:42

I'm like, is that not an app? I'm like, hello? Yeah. The nice thing about the harnesses compared to the apps is that they're more composable. I can have a harness call another harness. I can have a harness spin up another harness in a sandbox in the cloud. I don't know what APIs did for apps. Yes, but it's much more reliable and deterministic and easy for users to grok with a harness because the harnesses are Unix-based. They all have, and the models are so well trained on Unix, on bash commands. Whereas, if I'm telling a harness to go orchestrate an app in the cloud, it's going to be like, oh boy, how do you log into this app? Because do I need your password?

51:40

Do I need to fire up a virtual browser? It's going to be pretty slow. I'll figure it out. Okay, I fired up a browser, and now I need your password, and I'm going to try to find the input where to put it in. And there's probably an API in this app somewhere. I need to look up the docs to figure it out. And okay, now I've got the API, but there's so many unknown unknowns when you're composing around an app. Very, very, very, very few unknown unknowns when you're composing around a harness. So, I think it just gives developers more flexibility, and flexibility that they can inspect. API calls, you're just seeing a whole bunch of code flying around the screen.

52:02

A harness, oh, I can jump into the harness and look at what's going on and talk in English about it. So, it's much more user-friendly. We've seen Meta and Muse really be a focus for Zuck. We've seen Alex Wang front and center much more. Were you impressed by what Meta delivered with Muse? They've been doing a good job, yeah. I mean, it takes a while to set up a whole new model lab from scratch. And I'm sure a lot of organizational debt to deal with. Do you think they will be a serious challenger? I do. I think they're... I think they have the resources.

52:31

I think there's some competitive things they can do around the model that helps people in ways that the model labs are not as interested in doing. Just having a social network and a focus on people. It's something for the brand that maybe Grok and SpaceX AI have too. They do need to find their niche. I'm not quite sure. I think people don't quite know what to do with Muse Spark yet. When to use it or when to go for it or what its core advantages are. They just released a coding harness. They are trying to be a generally capable model right now. I expect that in the future they're going to be like, look, we are way better at this thing.

52:56

And that's going to be a really important moment for them. I see. I was impressed by it, actually. Do you know what I use now? Maybe plug in one of our mutual friends, but Anastasios and Arena. And it's so weird. So, I'll put my prompt in Arena. And then, obviously, it comes back with a load of different model options. Yeah. And I come back with, I used one the other day, Pergamum. Pergamum. Yeah. And it was Kimi and Pergamum. And they offer you four different options. And it takes me to models that I would never have used before. And actually, Muse has come up a couple of times for being pretty impressive.

53:59

But I love that in terms of this discovery mechanism to models that I would never have used. I would never get a Kimi, honestly, dude. I'd just get a fucking chat GPT. It's really interesting. Yeah. It goes to the point of the model layer just becoming a utility layer. What do you mean by that? Well, actually, I have no loyalty to them. I have no affiliation with brand. I go to Arena. And I want to see what you got for me. Show me the results. I don't care if it's Kimi or Muse or Claude or Sonnet or whatever. Do you know what I mean? And actually, I just want to see the options you got.

55:14

And I'll pick the best from there. I'd rather run four in parallel. Do you buy this whole we're going to have one frontier model run four open models? And the frontier model might be 160 IQ points.

55:35

And the open models might be 120 IQ points. But that will be a model infrastructure or structure that we'll work with. I totally think that that is a great architecture that everybody needs to explore. And we've been helping lots of developers do this. We're like, you have subagents. We have a subagent server tool that we tune to be really, really good at using models generally. And then you have an orchestrator model that calls out to the subagents when it wants particular tasks to get done.

56:19

And these subagents are just very, very low cost. And they're focused on deterministic tasks. This is what open weight models are generally really good at compared to frontier models. When you have a deterministic task where you know the shape of the output, you know the type of problem that you're working on. And it's a type of problem that has been solved, like classifying some text, for example. Then you should definitely use a low-cost model from OpenRouter and then have the orchestrator model read the results and then go and continue working on the unknown non-deterministic task that it was set out to do. I want to create an open American ecosystem. Yeah.

57:06

More amazing open American models. And I make you head of this program. What would you do to encourage, incentivize the open US ecosystem to compete more vociferously with the Chinese? I think I would spend time talking to the current American labs a little bit more to figure out what distilling the Chinese models looks like for them and how effective it is. You can probably get pretty far distilling the Chinese models. The nice thing about the open weight models and the Chinese models is that they allow distillation, most of them. And that means that you can take the outputs of these models to do reinforcement learning on top of the model that you're building.

57:32

And this is just a very important and common practice in AI that all labs do. So, I would want to learn a little bit more about how effective it is, but it's one way to catch up with the open weight models because they allow it. The other thing is the American models, and also when you distill, you see the output. So you can inspect them to make sure that they're aligned. So if there's anything about the open weight models that you're worried about not being aligned with the voice or constitution of the model you're creating, you have a much better shot at catching it when you're doing these RL rollouts. The other thing I would try to figure out is the compute question.

57:51

Compute is just a huge advantage that I think we still have relative to China. And these Neo labs need a shot. And there needs to be an easier way to get compute to the right talent in all countries. But especially if we're trying to create a competitive American Neo lab system. Nvidia has been doing a good job of this, but there's Google, there's TPUs, there's Tranium from Amazon. I would work with all of the hardware companies and also the Neo chips to help with compute. I don't think we will have that compute advantage for long. you have a much better shot at catching it when you're doing these RL rollouts.

58:11

The other thing I would try to figure out is the compute question. Compute is just a huge advantage that I think we still have relative to China. And these Neo labs need a shot. And there needs to be an easier way to get compute to the right talent in all countries. But especially if we're trying to create a competitive American Neo lab system, Nvidia has been doing a good job of this, but there's Google, there's TPUs, there's Tranium from Amazon. I would work with all of the hardware companies and also the Neo chips to help with compute. I don't think we will have that compute advantage for long.

58:44

I think you see DeepSeek and ByteDance both aggressively pursuing their own chips. Now, the export controls mean that they have to. And this is the number one problem for Xi Jinping in his race to win the AI war. I agree. If they build a bridge in four weeks, I think they'll manage a chip in six months. Yeah. It's like staying ahead on the chip war is critical for America. Is distillation wrong? I mean, distillation is a technique to build models. But people view it with cynicism and shade. Yeah. Well, and then just distilled models. It's a technique to build models.

59:58

The closed weight model labs distill models too. Sonnet is a partially distilled version of Opus. And this is how you make smaller models out of bigger models. It's an important way to teach your model new things when you find something useful in the ecosystem. We do think that labs have a right to say it's not allowed in their terms of service. A company can cut off access to someone who is trying to build a competitive model. If you're just trying to build a smaller model that's really focused on doing one specific thing that's not competitive, most of the frontier labs don't prohibit that to my knowledge.

1:00:33

But there are going to be markets for companies that allow it and companies that don't. And we make sure that we help both companies uphold their terms of service. I have to ask you one question before we do a quick fire round. I'm going to get killed if I don't ask. There are reports that you are selling to Stripe for $10 billion. Is that going to happen? I can't comment. But whatever happens, we're going to execute on the vision. What we're doing is critical for the ecosystem. And we believe in safe access to AI where one monopoly doesn't take over. And where we have a vibrant ecosystem of models that everyone can explore.

1:01:43

And when new providers and new server tools and new inference-adjacent tech come online, there's a really easy way to discover it and connect it with all of your existing AI. I was thinking in these situations that my response would be, well, I own 22% of the company, $10 billion, $2.2 billion. Now I'm a venture capitalist. But is it hard not to think like that? I don't really think about it. Do you not? I don't spend a lot personally. What I think about when I do with personal capital, I really want to help people work on problems that don't lend themselves very well to venture capital.

1:02:07

They sort of fall in this gray area of problems that people need to solve, but are really tough to fund because they don't come with a business model attached. And I think there are very cool things to do now in the nonprofit space, because you can use AI to review way more data than you ever could before. I'm not quite ready to talk about it publicly yet, but I do want to do something that helps researchers work on those problems and get grants to do it.

1:02:19

One really cool example, I think, of this is David Fialcow, who's one of the founders of General Catalyst, who basically finds incredible stories that won't get funded for movies and funds them to shine a light on them because he thinks they're very important. So the dissident, which obviously told the story of Khashoggi, and then Icarus, which is the story of the Russian doping. And these were films that would not get funded had it not been for his funding because they were politically sensitive, charged. And he's like, I'm going to enable the stories of these forbidden tales. Yeah, it's kind of like that. I love that stuff. He's great. He's fucking awesome.

1:02:49

Anyway, are you ready for a quick fire round? Sure. Okay. So what is the most underrated model on OpenRouter today? Ooh, good one. I mean, first, poolside's models are great. That's probably my fire round answer. Good new American lab building interesting coding models that are small but highly effective. And they're building a lot of useful tools for accessing them.

1:04:02

Good team. 70% of Neo labs will die in the next three years. Agree or disagree? Disagree. 70 seems very high. Of Neo labs, there aren't that many Neo labs. If getting acquired by one of the model labs counts as die, I do think there'll probably be some potential consolidation. If you include the consolidation, I'd say 50. Do you think Dario should be less negative and more positive as a voice in AI?

1:05:12

I think it's important to have somebody who is very paranoid about the future and how things are going to shake up. And I appreciate that. I personally appreciate Anthropic's paranoia. Obviously, there are areas where I want other model labs to not feel like they're just being pushed off the table, but I'm a big believer in neurodiversity. And Anthropic is a part of the neurodiversity map that really matters. And if no one is being extremely paranoid, then no one is offering that voice. And so I appreciate that they're doing it. What's the craziest thing that you see in your seat on top of everyone's usage that you don't think people talk about enough?

1:06:06

A lot of companies are obviously worried about cost management and freaking out about the amount of inference they're spending, and they don't know how to think about it. It's a whole new way of doing business and thinking about your OPEX. The old way of thinking about how much you give your employees, you give them a salary and you kind of forget about it. Someone knows what everyone's making, but it's a static number that gets readjusted on a quarterly basis, maybe after performance reviews. Really, your employees all cost totally dynamic, different amounts now. And I think a lot of companies are putting it on them to do routing.

1:06:22

And I think in the future, there's a good chance that it will get pushed downward to the employee level. Your employees should figure out which tools and models to use that are best for their tasks. And then we should figure out how much you're costing due to the choices that you make as an employee. And your cost as an employee is going to be a dynamic number, and it's going to be dependent on how much that employee is effectively using expensive and cheap models to do their job.

1:06:40

And then I advise companies to still do their normal management work, have their managers assess how effective and productive employees are, but also line it up with how much their employees cost. And then kind of come up with a quadrant of celebration, like these employees are doing a good job and they're pretty price-effective or cost-effective. And then a quadrant of concern: these employees are maybe doing a so-so job, and whoa, they are not cost-effective at all. Their AI use, their AI psychosis, is off the charts. And then you address the quadrant of concern. So I don't think people talk about how you think of employee cost in the age of AI.

1:07:35

And that it should be a dynamic number and not a static thing that only a few people know about and it's gone. Wonderful. But can you imagine going to someone, oh, I'm sorry, you were worth a hundred grand last month, now you're worth 50? I think it would make planning. They are in control of how much they cost. That's the great thing. All employees are in control of how much they cost and can influence that. And then a quadrant of concern: these employees are maybe doing a so-so job, and whoa, they are not cost-effective at all. Their AI use, their AI psychosis, is off the charts. And then you address the quadrant of concern.

1:08:10

So I don't think people talk about how you think of employee cost in the age of AI. And that it really should be a dynamic number and not a static thing that only a few people know about, and it's gone. Wonderful. But can you imagine going to someone, "Oh, I'm sorry, you were worth a hundred grand last month. Now you're worth 50." I think it would make planning. They are in control of how much they cost. That's the great thing. All employees are in control of how much they cost and can influence that. It's now, now you get to think, "Okay, how good am I as an employee and how efficient am I being as well?" Final one.

1:08:19

When you look at the landscape today, there are so many things to be excited about. What are you singly most excited about? One is rare disease research, which I think is one of those things that has been intelligence bottleneck, or really just the inference bottleneck. It involves trying out lots of ideas and seeing if they work. The other is crowdsourcing productive urban life improvements. So for example, imagine if somebody was curious about finding every lead pipe in America or every lead pipe in the UK and had an approach to it, but they really need to make it mature and stress-test it. Now you can use AI to do that.

1:08:25

And we just might solve some weird problems that everyone's just given up on. Because you need a crazy idea to come from somewhere. Brilliant ideas are evenly distributed all over the world. They can come from anywhere. And now you just give them leverage to actually work. So I'm excited about very broad urban or rural quality-of-life improvements that we'll be able to make.

1:08:25

Nice dude.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note