Open Reader

The Open-Source AI Reality | How Token Costs Will Fall 10X & Usage Will Explode 100X | Lin Qiao

completed 1:28:42 Jul 20, 2026 Watch on YouTube

Current Status

completed

Video ID

PCAiqKCfRSk

RAG / Chat

Enabled
The Open-Source AI Reality | How Token Costs Will Fall 10X & Usage Will Explode 100X | Lin Qiao
Description

Lin Qiao is the Co-Founder and CEO of Fireworks AI, the leading specialized intelligence and AI inference platform that last week raised $1.5BN at a whopping $17BN valuation. With just 200 people, the company has hit $1BN in ARR and expects to hit $2BN before the end of the year. Prior to Fireworks, Lin spent several years at Meta including on the founding team of PyTorch. ----------------------------------------------- Timestamps: 0:00 Intro 02:00 - Why Starting a Company at 48 Was an Advantage 03:47 - The AI Layer Everyone Is Overlooking 09:49 - Is AGI Really the End Goal? 11:40 - Open Source vs Frontier Models: Who Wins? 14:30 - Are AI Giants Massively Overvalued? 17:48 - Why Open Models Could Beat Closed AI 19:23 - Should We Trust Chinese AI Models? 22:18 - Do AI Startups Need to Build Their Own Models? 26:49 - Will AI Model Breakthroughs Ever Slow Down? 29:03 - Why One Company Should Never Control Intelligence 32:47 - The Secret Behind Cursor's Explosive Growth 37:33 - Is AI Coding Already Yesterday's Biggest Trend? 41:36 - The AI Infrastructure Race Is Just Getting Started 46:32 - How Cheap Will AI Become? 54:07 - Hypergrowth vs Profit: Why Margins Can Wait 59:30 - Can the West Keep Up With China's Infrastructure Speed? 01:01:04 - Why AI Hardware Depreciates Faster Than Ever 01:04:45 - Why AI Will Create More Jobs, Not Fewer 01:08:35 - The Biggest Mistakes AI Founders Are Making 01:16:55 - Why Great Leaders Stay Close to the Work 01:18:44: Quick-Fire Round ---------------------------------------------------------------------------------------------- Subscribe on Spotify: https://open.spotify.com/show/3j2KMcZTtgTNBKwtZBMHvl?si=85bc9196860e4466 Subscribe on Apple Podcasts: https://podcasts.apple.com/us/podcast/the-twenty-minute-vc-20vc-venture-capital-startup/id958230465 Follow Harry Stebbings on X: https://twitter.com/HarryStebbings Follow Lin Qiao on X: https://twitter.com/lqiao Follow 20VC on Instagram: https://www.instagram.com/20vchq Follow 20VC o

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Lin Qiao argues that AI will not consolidate around a few general models: falling inference costs, open-weight models, and enterprise-private data will produce millions of specialized models owned, tuned, routed, and optimized by the companies using them.
  • Why it matters: The interview offers a concrete control-plane thesis for agent systems: proprietary workflow data, evaluations, model customization, multi-model routing, and inference optimization—not access to a frontier API alone—become the durable operating layer.
  • Best use: Use it to pressure-test OpenClaw and agent-system architecture around model ownership, routing, evaluation-driven tuning, workload-specific deployment, unit economics, and build-versus-buy boundaries.

Executive Summary

Qiao positions Fireworks between foundation-model providers and applications, but rejects the characterization of inference as a commodity. Her argument is that most economically valuable data is private, embedded in enterprise workflows and user interactions, and cannot be absorbed into a public general model. Companies therefore need specialized intelligence: models tuned with their proprietary data, wrapped in bespoke workflow harnesses, and served in a configuration optimized for their particular quality, latency, and cost requirements.

Open-weight models are central to that thesis because they give enterprises control over weights and customization. Qiao contends that both closed and open models have crossed a practical quality threshold for many workflows; Fireworks itself reportedly uses open models for recruiting, finance, debugging, and agents. Once an organization has enough production volume, it can tune an open model against its own evaluations and potentially outperform a general-purpose model on its specific task. The relevant distinction is not simply open versus closed, but whether the organization can control, evaluate, adapt, and economically operate its intelligence.

She expects AI architectures to become hierarchical and multi-model. Expensive, high-capability models should handle complex judgment and decomposition, while smaller specialized models and sub-agents execute narrower tasks. The strategic routing layer is therefore not generic model switching; it is a business-specific mechanism informed by the company’s own task decomposition and evals. Qiao expects automatic routing and automatic tuning eventually to create self-evolving production systems, although she acknowledges that most customers cannot yet build that sophistication without deep technical involvement.

Economically, Qiao forecasts roughly 10x lower task/token costs over the next three years and 100x more usage, while cautioning that tokens are a poor universal cost measure because models differ in verbosity. She argues that cost compression will come from task precision, customized models, inference optimization, and eventual easing of chip and infrastructure constraints. The interview is optimistic and company-promotional, but it contains unusually useful implementation logic on RL infrastructure, distributed training/rollouts, model-routing design, workload maturity, and why optimization choices should depend on scale.

Key Takeaways

  • Claim: The durable AI advantage for enterprises will be specialized intelligence built from private data and proprietary workflows, rather than undifferentiated consumption of one general model API. | Evidence: Qiao argues that public-internet training data is only a small subset of world data, while the majority is private data locked inside enterprises and applications. She cites Jensen Huang's observation that there is no "specialized general company": every company’s differentiated product design, user interactions, and data encode unique operating assumptions. | Implication: For Ken, the defensible layer in an agent product is likely the workflow-specific context, eval suite, policies, orchestration, and feedback loop—not merely a model-provider integration. | Caveat: This is a strategic thesis, not proof that every company should train or host a model; Qiao later says organizations should choose which layers are proprietary versus common infrastructure.
  • Claim: Open-weight models become materially more valuable once they clear a usable capability threshold because they can be tuned and served under the enterprise's control. | Evidence: Fireworks chose to build on open models when they were still immature because its founders' PyTorch background favored openness and user control. Qiao says open and closed models have both improved enough to solve many business problems, and says Fireworks uses open models internally for candidate sourcing, feedback collection, finance processes, debugging, and agents. | Implication: Use open weights where control, tuning, data locality, and unit economics outweigh frontier capability; treat model-level guardrails as insufficient substitutes for application-owned policy controls. | Caveat: Open models do not eliminate governance needs: Qiao says every enterprise should put its own guardrails around any model, open or closed, because providers embed their own judgments, tastes, and policies in training.
  • Claim: Production AI systems will use hierarchical multi-model orchestration, with a high-intelligence layer for difficult judgment and smaller customized models for narrower subproblems. | Evidence: Qiao describes decomposing a task so an expensive open or closed model handles the highest-complexity judgment, while sub-agents solve smaller tasks with smaller open models that can be further specialized. She says companies with deep use-case knowledge and evals may be best positioned to build their own routing mechanisms. | Implication: OpenClaw should make routing, task decomposition, model-policy selection, and outcome evaluation first-class control-plane functions rather than treat a single default model as the system architecture. | Caveat: She believes automated routing that learns from production flows is an emerging opportunity rather than a solved capability, and concedes firms may not need third-party routing vendors once they can build it themselves.
  • Claim: At meaningful scale, AI product-market fit does not ensure a durable business because inference costs can cause a company to "scale into bankruptcy." | Evidence: Qiao contrasts SaaS, where infrastructure costs were comparatively commoditized, with AI products whose usage can create unmanageable COGS. She says incumbents with large existing traffic may be unable to launch AI features broadly if cost forecasting cannot justify the rollout; even a 5% cost reduction matters at millions to billions of users. | Implication: Measure economics per completed task and business outcome, not token price alone; introduce routing, model compression/tuning, caching, and workload-specific serving before usage commitments become structurally unprofitable. | Caveat: Her answer is framed from an inference-platform founder's perspective and favors customization as the solution; lower frontier API prices can also improve economics.
  • Claim: Token costs should fall sharply, but the useful economic metric is cost per successfully completed task rather than nominal price per token. | Evidence: Qiao predicts a 10x overall cost reduction within three years and 100x usage growth. She identifies four levers: fewer tokens needed per task as models become more precise, tuning for a specific problem, inference-platform optimization, and reduced underlying hardware costs after supply constraints ease. She notes a model can appear 2x cheaper per token but consume 2x as many tokens through verbosity. | Implication: Design reporting around quality-adjusted task cost, latency, success rate, and downstream ROI; avoid procurement or routing decisions based on token-list prices alone. | Caveat: She expects supply constraints to remain meaningful for roughly the next year to year and a half, and does not provide an independently verified cost model for the 10x/100x forecast.
  • Claim: The right degree of vertical integration depends on workload maturity: build specialized hardware or data-center infrastructure only after traffic and workload patterns stabilize enough to justify it. | Evidence: Qiao says Fireworks prioritizes agility and wants to run across all available AI chips rather than constrain itself to owned capacity. She argues that Meta's custom-chip investment was justified by long-standing, enormous recommendation workloads, whereas rapidly changing generative-AI models and workloads make a hardware bet fragile. She calls this the question of "workload maturity." | Implication: For Ken, keep the agent platform hardware- and provider-portable now; only hard-code infrastructure or make heavy ownership commitments after demand patterns, model mix, and SLA needs are demonstrably stable. | Caveat: She leaves data-center ownership open as a future option and argues data centers themselves can be specialized, including heterogeneous deployments that combine compute-oriented GPUs with SRAM-heavy accelerators for different inference phases.
  • Claim: AI operations will shift from token maximization to ROI maximization, requiring much stronger attribution and monitoring in production. | Evidence: Qiao identifies ROI monitoring as underinvested relative to the more visible model and application work. She says production AI will force discipline around what was spent, what return was generated, and how value and cost are attributed. | Implication: Instrument agents end to end: track task intent, model and tool choices, token/compute cost, human intervention, success criteria, business outcome, and regression by workflow version.

Detailed Brief

Customization means co-designing the workflow harness, not only fine-tuning a model

  • Claims: In high-stakes vertical products such as legal, the proprietary asset includes the assistant harness that decides which tools to call, in what order, and with what accuracy constraints.; Qiao argues that the orchestration harness and the model powering it should be co-trained or jointly optimized because tool-selection quality is integral to workflow quality.; Coding is presented as an early example of this shift: Qiao says Cursor was among the first application companies to tune models and that many coding companies subsequently followed.
  • Evidence: She describes legal work as error-intolerant and highly varied across case types, making generic application logic insufficient even when the base model is strong.; Her example of a task harness includes integrating tools, deciding which AI tool to invoke, and managing bespoke workflow logic.; She predicts base-model IQ improvements will arrive in periodic step functions, while specialization will accelerate more continuously and rapidly across the branching set of use cases.
  • Caveats: The discussion does not provide comparative performance results showing when model tuning adds more value than retrieval, prompting, tool design, or better evaluations.; Legal examples are illustrative; the speaker explicitly says her direct knowledge of the legal domain is shallow.
  • Implications: Evaluate specialized-agent moats at the level of the complete closed loop: proprietary data, task decomposition, tool policies, evals, feedback collection, and deployment—not the fine-tuned checkpoint alone.; Avoid treating model tuning as an automatic answer; establish whether the bottleneck is domain knowledge, tool reliability, policy compliance, or model capability first.

Cursor RL infrastructure illustrates a cost-conscious distributed-training pattern

  • Claims: Fireworks and Cursor reportedly decoupled reinforcement-learning training from rollout execution, rather than running both as a single tightly interconnected hyperscale cluster.; The design uses geographically distributed, scattered GPUs across five or six data-center regions to run large RL jobs while avoiding dependence on a 10,000- to 100,000-chip InfiniBand-connected cluster.; The key technical trade-off is model-weight freshness: stale weights make rollout rewards less valid, so the system must distribute updated weights fast enough to remain numerically sound.
  • Evidence: Qiao defines the RL trainer as continuously modifying model weights, then deploying each new model version to an RL rollout environment that interacts with synthetic or real coding environments and returns rewards.; She says the joint system was created to preserve startup capital efficiency without slowing research velocity, and links it to Cursor's recent model launches.
  • Caveats: This is an architectural description from a service provider and does not include benchmarks, training cost, convergence data, or an independent account from Cursor.; The pattern is relevant to large-scale post-training and may not be worth the operational complexity for smaller agent products.
  • Implications: For high-volume agent learning, separate online/rollout capacity from parameter-update capacity and explicitly manage freshness, replay, and reward validity.; For lower-volume systems, prioritize a simpler evaluation-and-improvement loop before replicating multi-region RL infrastructure.

Operating lessons: speed comes from context, ownership, and deliberate GTM education

  • Claims: Qiao's leadership lesson from Jensen Huang is that leadership is judgment, and judgment in a fast-moving domain requires direct access to ground truth rather than heavily filtered information.; Fireworks emphasizes hiring for extreme end-to-end ownership over narrowly boxed functional competence; Qiao says these people have the strongest long-term growth curve in the company.; Qiao's stated mistake was underinvesting in marketing because she assumed the product would speak for itself; she reframes marketing as customer education and strategic clarity rather than promotional fluff.
  • Evidence: She says information is inevitably lost as it travels through organizational layers, which makes delayed or abstracted management reporting dangerous in high-velocity environments.; Fireworks grew from about 50 people a year earlier to roughly 200 people, and Qiao says the company became more willing to scale because it learned to hire high-velocity owners and use AI tools aggressively.; She says the GTM team uses AI agents and shares skills to raise productivity while the team scales against a superlinear demand curve.
  • Caveats: The interview offers principles rather than a repeatable hiring assessment method or an operating cadence that would validate these claims.
  • Implications: Create direct telemetry and operating reviews for agent quality, cost, incidents, and user outcomes so strategic decisions retain contact with underlying reality.; Treat agent capabilities, workflow discipline, and model economics as a GTM education problem as well as a product problem.

Notable Concepts & Terms

  • Specialized intelligence: Qiao's category for enterprise-owned models and systems tailored to a particular company's private data, workflow, policies, and user needs.
  • Own intelligence versus rent intelligence: The strategic choice between relying on third-party model APIs and retaining control of customized models and serving; Qiao extends it from companies to national sovereignty.
  • One size fits one: Fireworks' serving philosophy: each customized model deployment should be optimized uniquely for a customer's workload across quality, speed, and cost.
  • Quality-adjusted task cost: The implied alternative to token pricing: compare the total cost and performance required to complete the same task, accounting for model verbosity and accuracy.
  • Zero KLD: A claimed training-to-inference quality property in which the deployed model is bit-equivalent to the trained model, avoiding quality loss at the serving boundary.
  • RL rollout: The execution side of reinforcement learning where a newly updated model interacts with a synthetic or real environment, producing reward signals for further training.
  • Workload maturity: The point at which an application’s traffic and computation pattern are stable enough to justify deeper vertical integration, such as custom chips or specialized data centers.
  • Token maxing to ROI maxing: The predicted operational transition from maximizing AI usage to managing AI as a business system with explicit cost, return, and attribution discipline.

Operator Notes / Why Ken Should Care

  • Define a routing policy for OpenClaw that assigns tasks by risk, latency, cost, privacy, and required reasoning depth; maintain an override path rather than relying on a single automatic router.
  • Build a workflow-level eval harness before committing to model specialization: measure tool selection, task completion, policy compliance, human escalation, and cost per successful outcome.
  • Create an intelligence-ownership map that classifies which assets should remain proprietary (workflow traces, domain evals, policies, fine-tunes, routing logic) versus rented (general reasoning capacity, commodity hosting, standardized tools).
  • Implement per-task economic telemetry now, including token use, model verbosity, retries, tool-call cost, latency, outcome quality, and attributable business value.
  • Maintain model/provider portability and test fallback paths, especially for workflows exposed to API policy changes, regional availability constraints, or sovereign-data requirements.
  • Do not make infrastructure ownership commitments based on projected volume alone; require stable workload patterns and a quantified savings case before considering dedicated capacity or hardware specialization.

Source/Metadata

  • Title: The Open-Source AI Reality | How Token Costs Will Fall 10X & Usage Will Explode 100X | Lin Qiao
  • Transcript words: 14809
  • Duration seconds: 5322
  • Timestamp note: No usable timestamps or chapters were present in the supplied transcript; portions of the interview, especially the closing rapid-fire section, are duplicated.

Transcript

12088 words en Processed in 577.1s

What I don't want to see is that only one company owns intelligence. That doesn't make sense to me. I think last year was the year of coding, and this year is the year of co-work. And in the hot seat today, a founder who I wrote a $10 million check for after just a 15-minute meeting, Lin Kuo, founder of Fireworks. This was one of the easiest investment decisions that I've made in a 10-year investing career. I do think the cost of tokens will go down drastically. 10x cost reduction in the next three years. And this 10x cost reduction will drive 100x usage. We absolutely are not going to move into the application layer. Very clear to us. Whether we will move down into data centers and so on, that could always be on the table. But the question is... Ready to go? Lin, I am so excited for this. I heard so many great things. I just got off the phone with your co-founder, Dima. I spoke to Alfred Lin, Sonia, Matt Miller, and many more. So thank you for joining me. Thanks for having me. Now, I heard that Eric Vishria has a rule: don't invest in big tech directors. But he broke that rule with you, which is very special. I think so, too. So, a funny story. After we decided to handshake, he did call me and said he talked with one of his advisors. And his advisor questioned him, hey, how many big tech executives have you seen being successful in starting a company? Very few. And he told me that. I was surprised. Are we breaking our handshake now? No, but since then, we work very closely with each other. Eric is one of the best. You also started the company when you were 48? Oh, yeah. That's quite late. Can I ask, how do you reflect on being a 48-year-old founder when we glorify starting a company when you're pretty much 15 these days? I didn't think deeply about that. I always wanted to have a tech business myself. I actually wanted to start a business in 2015 because I'm a first-generation immigrant. I came to the U.S. in 2000. I did my PhD in distributed systems, computer science, especially focused on databases. And databases are very complex systems to build, with a lot of different objectives to optimize for. And pretty much after I joined a research lab, I pretty much touched every single aspect of processing data. And then I moved to LinkedIn to further it down, to build systems and products to be used and drive real impact. At that time, I felt I was ready to start a company. I knew all the tech. I knew what product to build. I had a business proposal. I had a list of people I wanted to start a company with. And I spent time thinking about it, and I paused. Because I don't think I had the skill set on people to build a company. It's not just about product. It's not just about tech. It's actually about people. And I decided I wanted to go to a place where I could learn the most about people. And the best company at the time was Facebook. It's a rising star in Silicon Valley. And secretly, I was planning to learn for one year or two and then go back to do my own business. I stayed there for seven years. So with Fireworks, you saw something in inference that the world was not focused on. The world was focused on training. I think it's helpful for people to understand the stack because beneath you, there's obviously chip providers and your NVIDIAs of the world. And then above you, you've got the model providers. And you sit in between. Why is that a valuable part of the stack and not a commodity? That's a really good question. But why bother with specialized intelligence? Why not just use generalized intelligence? And you worry less, right? You just build on top of an API that is provided by Frontier Labs. Wouldn't that be much easier? So the argument is the following. If you think intelligence is a derivative of data, then the majority of the data is actually not used for training a general intelligence model. The training data is coming from the public internet and labeled data. The public internet is very small. Corporate data compared with world data, the majority of world data is actually private data locked inside applications, locked inside enterprises. It will never get shared with anyone else because this is a company's proprietary IP. So then it's interesting. If you look at the space, then it becomes very interesting because the majority of data is not being activated to derive any intelligence. And that's what we believe in, to activate that data. And we believe the future frontier of intelligence is actually private intelligence, or specialized intelligence. So that's where Fireworks, from the beginning, has been focusing on driving the value. I have so many questions to ask you. I totally understand you in terms of the value in private data within some of these largest companies. Is that not the premise of what Anthropic's enterprise business is, though? With Claude Cowork and with a lot of their adjacencies that they're building, would Dario not say that that's exactly what we're going after? That's interesting because I view Anthropic as a company fully believing in AGI. The definition of AGI is that there's this one model that can solve all the problems in the best way. To me, that's the definition of AGI. To me, that means you do not need to specialize. And that one model should be able to solve all the problems. It's so intelligent. It has so much knowledge of every part of businesses, every part of the jobs it can fulfill. Then why do you need to bother specializing? So that itself is a validation that we're living in a world that's not ruled by one principle. We are living in a fully diversified world. I'll give you one example, right? Different regions. We'll have different value systems. We'll have different policies. We'll have different ways of conducting business. We'll have different lifestyles. It's all taste, choices, and judgment combined. I think that's what defines us as human. We are not robots. If our future world is going to be ruled by one standard, a taste dictated by one company, we turn ourselves into an army of robots. And that's very depressing to me. And I think what separates Homo sapiens from other species is creativity, is the deep desire to pursue new things, of discovering new ways of living. That defines us as human beings. And that part cannot be copied. That's my fundamental belief. That's why, in Silicon Valley, there's so much creativity. Across the world, there's so much creativity in building new businesses. What is new business? I had this fun, interesting conversation with Jensen after his GTC keynote. We actually recorded it. It is interesting. I watched it. It was great. Yeah. Recording with Jensen is not really recording. He just started having a conversation with me. I didn't know his crew had already started recording. We just kept talking. It's so easy. We talked about this specialized intelligence. He said one thing to me. Lin, you're right. There's no specialized general company. As in, every company is built on a special belief of doing things. Otherwise, there's no reason they should exist. It feels logical. But then I started to think back about what he said, and it is profound. Because every single company is doing something unique that justifies its existence. And this something unique is deeply baked into their product design. It's deeply baked into their software design and system building. And that's deeply baked into the data and the interaction with their users and their deep understanding of their users' intent, interacting with their product, and engagement, and so on. All of that is the fundamental base of why a company should exist. That is not learnable or shareable by another company sitting outside. Can you help me understand? As a podcaster, I specialize in asking basic questions. So forgive me. But why then do people like Dario, like Sam, like Larry and Sergey talk about AGI in the way that they do, as inevitable? I think what they build is fantastic. Because they are basically building power lines to distribute a really great source of intelligence that everyone else can build on top of. That's how I view their contribution. And if we don't have this fundamental infrastructure, then we will not have all kinds of appliances living in our home. I love my coffee machine. And it's specially branded, right? But without that power, then we don't get to do the things that are fun, that are unique, that are special, that ingrain our taste. So I do think that's very, very important. But the question is, is this power line going to replace everything we do? I don't think so. The question for me as an investor is, are power lines good businesses? You said about PyTorch and open source and the open ecosystem. Open source in the last, I would say, three months, we've all realized is actually accelerating so fast. And the capabilities have increased to such an extent that it's not quite comparable. But it's getting 90% as efficient with, to Chamath's statement, 15 times more cost-effective. Are power lines good businesses in a world of open source? So here's how I view open source. I love my coffee machine. And it's special branded, right? But without that power, then we don't get to do the things that are fun, that's unique, that's special, that ingrains our taste. So I do think that's very, very important. But the question is, is this power line going to replace everything we do? I don't think so. The question for me as an investor is, are power lines good businesses? You said about PyTorch and Open and the Open ecosystem. Open source in the last, I would say, three months, we've all realized is actually accelerating so fast. And the capabilities have increased to such an extent that it's not comparable quite. But it's getting 90% as efficient with 15 times, to Chamath's statement, more cost effective. Are power lines good businesses in a world of open source? So here's how I view open source. So early on, when we founded a company, we had pretty deep debate among the co-founders. What do we do? Do we build our own models or do we build on top of open models? At that time, open model was almost at its infancy. It's a big bet. If we're going to take that direction, it's a huge bet that it's going to do well, right? But with our PyTorch experience, we believe in the open community. We believe in openness. That's a fundamental principle we operate with. Because openness gave control. Openness gave control to the user. Think about open models, right? Once the model is released, you have full control of the weights. You can change it however you want. It's yours. And then you can build on top of it, right? So that is a fundamental different operating principle that we believe in because of our roots in open source before. So we took that bet. And it did pay off in the sense that both open model and closed model, the quality significantly increased, improved over the past two years. To the point, both of these two streams cross the threshold, cross the quality threshold. It can solve so many problems, right? So within Fireworks, obviously with Dogfoot, our own product, we use open model to drive our recruiting process. Candidate sourcing and the feedback collection. We use open model to even drive some internal finance processes. Obviously, for coding, we use open models to help us debug. We have a ton of agents within Fireworks ourselves. And we are cost conscious. So that is important. Both model categories cross the threshold. It solves so many variety of problems. Second is open model cross the threshold. It's so much easier to tune. Okay. So being able to steer a model is intelligence. It's part of the model intelligence. And the model intelligence has passed threshold. It's much easier to steer, especially with small amount of data. A small amount of unique data a particular company has. And then we can heel climb towards your eval. And oftentimes, the end result of heel climbing is to solve your unique problem with your data. You are better than a general purpose model. Okay. When 90% of enterprise workflows can be done, as you said, that the incredible array of functions that you now use open source for with open models. So the usage for frontier models will not be as large as it was if it was needed for everything. So are these companies actually dramatically overvalued and overestimated if the majority can just go through open? I think people start to realize it. I remember from two years ago, I went to different places and talked about an interesting phenomenon. That doesn't exist in the past, in the SaaS era. During SaaS time, product market fit and the durable business almost are equivalent to each other. The hardest thing is finding product market fit. And then once you find it, it just scales as fast as you can, right? Because CPU is a commodity. The infrastructure you're building on top is almost like a commodity. You don't even worry about that as your cogs. And now, product market fit and durable business are two separate concepts. For startups, we have great companies that have power market fit. The customer wants to pay them and they really value their product. But they cannot scale because once they scale, they could scale into bankruptcy. Have you heard about scaling to bankruptcy? So that's a real problem. It's an even bigger problem for incumbents. So the big companies for digital native, because they have the traffic. They have a huge amount of traffic. They're the winner from a decade ago when they were startups. And they have so much traffic. Once they draw out those AI features, they're going to reach all their customer base. And they cannot afford to do it. Because their sales will look at their cost proposal. Cost forecasting, there's no way you can justify this, right? So then it becomes a real problem to all those innovators. Hey, we really want to plug in to this new technology, new disruptive technology, but we cannot afford it. And we need to find an alternative to be able to afford it. And the alternative is to have the control over your open weights model and roll out your own model. It is, or you see what Sam Altman's releasing in the last few days, which is just dramatically lower cost models. I can't remember the amount it is, but I think it's like half as expensive or maybe three times cheaper. Is the next step actually we just see a massive reduction in price from the frontier models? It could be, but I think at the same time, it's just a very different operating principle. Because for open weights model, because it's just there, model acquisition has no cost, right? Obviously, some companies train those models and are willing to open it up. I know within the U.S. there are multiple companies doing that, including NVIDIA training, NemoTron. We're working obviously very closely with them. So once the model is there, whoever is using those model, there's literally no cost. But there's fundamental cost for the frontier labs to invest in those models and recoup the R&D cost back. And second is you just cannot customize those general purpose models. And you use it as is on top of an API. You have no control over. Versus with open model, you have full control. You can tune however you want. You can use it however you want. Especially Fireworks, we are a specialized intelligence platform. We offer all sorts of tools for you to easily customize the model for one specific use case. And after that model is tuned with high quality, then we further optimize for inference deployment. Think about Fireworks. We think about every single model deployment as one size fits one. It's unique for your workload only. It's optimized for your workload only from quality, speed, cost point of view. So we believe that's absolutely needed. Because once you think about a production scale of reaching to millions of users, tens of millions, billions of users, then even 5% of cost reduction means a lot. It's a massive amount. Let alone what we have seen in the past is 5 times to 10 times cost reduction. The one question that I do have to ask is the concern that enterprises have is national security concerns. When you look at OpenRouter, I think the top six models today are Chinese models. And they're incredible quality. The speed of development is incredible. But they are Chinese models. Do we have serious national security concerns when analyzing the power of Chinese open source? I think it's a huge debate happening right now across the industry. Once the model is open, you can put all kinds of guardrails specialized to your business around it. I would say to all models, it doesn't matter if open or closed. You should put your own guardrail around it. The fundamental reason is the following. A model provider will infuse their own judgment, their own taste into the model training process. You cannot guarantee a match is yours. Remember, it goes back to Jensen's comment. There's no specialized general company. Every company is special. Every company will have a special design principle. Every company will have a special taste. Every company will have a special target audience to serve. Because of that specialty, it's guaranteed that the judgment, the taste, the design principle from one company will mismatch, will misalign with your company, which is special, is solve a special problem. So that is a reason you need to tune those models to match yours. And I really believe the future will not be a few small number of AGI models dominant world. I really believe the future will be, it may be scary, but I think that's true. It will be millions of specialized models, one propagation per use case. We saw in the last week actually reports that China were looking at actually restricting access to their open models because they were seeing the development being so fast and so good. What would happen in a world where China actually started restricting access to their open models, given the lack of open models we have in the US? I think it will be a big impact in the short term. But the beauty of open ecosystem is it's not one provider. That's why it's open, right? It usually attracts many, many, many interested parties to participate. I do believe in terms of talent density and resources, I do believe US will be able to build that open system by ourselves, and we should. And I've seen this happening again. Again, in many open systems, there are a thousand flower blossoms. I really believe the future will be, it may be scary, but I think that's true. It will be millions of specialized models, one propagation per use case. We saw in the last week reports that China were looking at restricting access to their open models because they were seeing the development being so fast and so good. What would happen in a world where China started restricting access to their open models, given the lack of open models we have in the US? I think it will be a big impact in the short term. But the beauty of open ecosystem is it's not one provider. That's why it's open, right? It usually attracts many, many, many interested parties to participate. I do believe in terms of talent density and resources, I do believe US will be able to build that open system by ourselves, and we should. And I've seen this happening again. Again, in many open systems, there are a thousand flower blossoms. And that's the beauty of that. When we talk about the specialization of intelligence within enterprises, as you have done just there, if we take a very prime example, which I don't particularly want to take because I'm an investor in Lagora and I think I know which side you're going to fall on here. But you have two companies that compete in the legal space, Harvey and Lagora. And Harvey have committed to building their own model, and then Lagora have not. A year ago, it looked like companies that didn't commit to their own model were right because frontier models were increasing so fast in terms of capability. Now it looks like they're wrong. Should companies like Harvey and Lagora be building their own model? And actually, if you don't, what happens? So here's one observation I had, and many people have, is software development, especially SaaS space, has been significantly disrupted because of the general intelligence of coding. And the application development lifecycle has significantly collapsed in terms of the timeline and the resources needed. In the past, it required tens of very strong product engineers and PMs to convert from idea to implementation to production scale, multiple quarters of years of investment. That's a deep moat. And today, one person, a few weeks, can possibly launch their ideas into a product and scale quickly. This is unprecedented. And that's also created interesting dynamics in redefining where the competition is because it's really hard just to compete on the idea of application, application by itself, because many people have similar ideas. Now implementation is no longer such a big barrier. Is that actually true, though, when you're looking at enterprise deployment, enterprise rollout? If you're working with some of the biggest law firms in the world, the enterprise sales cycle is at least multi-year, with relationship build that's very tough. And then you have deployment that's very customized. It's not like 11 Labs where you pick it up and go. It's different. And also, I think legal space is particularly challenging because lawyers are usually more conservative. Legal is also not tolerant at all on errors, right? Because that's why lawyers get paid, right? You can interview the very rock-solid case. If something holistic is and generates wrong judgment, then you're in trouble. So I do think the legal space is a very interesting space to penetrate, and these companies are both doing a great job. But on the flip side, I do think both companies are owning proprietary knowledge and information, how to build those assistants, to do case studies, to go deep in driving legal research and all this, right? And my understanding of legal is so shallow, but there's so many different versions of flavors of cases. So I do think they are in a unique position to convert that deep understanding, and they all have data. It's not just about how defensive their business is. It's about, hey, oftentimes, when they build those assistants, there's a harness integrating and deciding and orchestrating which AI tool to use, which tools calling to, and this is bespoke. This is customized. And the accuracy of calling those tools and calling to what kind of tools is important. And even that harness needs to be co-trained with the model powering it, right? So there's just ample examples of driving that business to excellence by owning their own intelligence of how to do that in the workflow layer. So maybe it's the timing. Coding, for example, I think in coding space, Cursor probably is one of the pioneers starting to tune their model, and now almost all coding companies tune their own models. Does that pace of model development slow down? Because every single day, it seems like we have a new model with a new capability, and it's like, oh my gosh, Cursor's newest model is amazing. Next, we have someone else, Ms. R's newest model is amazing. Gemini's newest model is amazing. In three years' time, will the pace of model development still be so fast and model superiority be so transient where one day it's one and the next day it's another? So there are a few layers of model advancement. There's base, general IQ advancement. So those will take step functions. So that's why when they release, there's always major release or minor releases, right? The major release of step functions, as you remember beginning of last year, there's a whole, this thinking. The thinking process is new, right? The model just doesn't spit out answer immediately. The model will think by itself and spit out answer is much better that way. So that's one step function. And there are many step functions we have seen too, but I see those as every year or every three quarters there's a major leap. But at the same time, build on top of those, the best base models, and I can see the specialization start to accelerate. Because as I said, it's really like a tree, right? There's so many branches and leaves that can possibly hang out on the trunk. And as the base model quality start to have step function leaps, there's so much more we can do to specialize. So I do see in the world specialization is going to accelerate much faster than the general intelligence part. When we think about the general intelligence part, just before we move further into the stack of multi-model, Sam proffered the 5% gifting of OpenAI and others to the administration. Do you think we've reached a stage where model development is so advanced and so important to society that they will in part be government or administration owned? That's a very interesting question. I think there were precedents of that. If we think about the foundation tier of those general intelligence model as fundamentally a base infrastructure for the big economy to operate around, there has been precedents of PG&E owns electricity and the gas and so on, right? So I actually don't know, but what I don't want to see is there's only one company owns intelligence. I think that doesn't make sense to me because there are different flavors, as I said, there are different flavors of intelligence. There's this general common intelligence that benefits everyone, and then there's a specialized intelligence that actually help us advance in history, to think differently, to create a new paradigm of living or new paradigm of doing business and shaping the industry. I don't want that to die because there's only one company who can do that. I don't think that makes sense. With the many models blooming theory, there's the idea that you will route tasks to different models dependent on what they specialize in. I think so. With that in mind, will you not build your own open router of the world to cater to that? Yes. You can argue they're the best to build it because they deeply understand their use case and they have the evals. So again, my thinking of what is the frontier is not just this one model. The frontier could be your special routing mechanism for your business, and you decompose that based on, hey, in order to fulfill this task, and you usually need a highly intelligent layer, maybe the most expensive open, closed models to judge at the highest complexity. And usually people will also be a sub-agent to solve smaller problems, and those can go to smaller open models, and those can also further be customized to fit into your special design. So I've seen a lot of people already doing that today, and we also think there's a space to build an automatic routing system that can learn by itself, and that compound with automatic tuning system eventually. We think it should all be automated, and then you can see a self-evolving system based on what flows through your product, and your product keeps evolving. Your product is live, right? So you keep deploying and launching new features and to interact with your users, and that just kind of, it will be totally self-evolving automated system. Do you think then that routing layer of the stack is valuable? If it can be automated or it can be built on its own, is that a valuable layer to have? I definitely think so. You do think so? I do think so. If it can be automated or companies can build it themselves, why would you need a requestee or an open router? You probably don't. Yeah. We're not there yet. But I do think this could be an area of innovation. You said Cursor being the frontrunners in terms of how innovative they've been. I completely agree with you. But I heard, and I really stalk you before shows, but I heard that CTO DEMA was embedded at Cursor for months building the RL infrastructure. Is that how it has Your product is live, right? So you keep deploying and launching new features and to interact with your users and that just, it will be totally self-evolving automated system. Do you think then that routing layer of the stack is valuable? If it can be automated or it can be built on its own, is that a valuable layer to have? I definitely think so. You do think so? I do think so. If it can be automated or companies can build it themselves, why would you need a requestee or an open router? You probably don't. Yeah. We're not there yet. But I do think this could be an area of innovation. You said Cursor being the frontrunners in terms of how innovative they've been. I completely agree with you. But I heard, and I really stalk you before shows, but I heard that CTO DEMA was embedded at Cursor for months building the RL infrastructure. Is that how it has to be done? And is that scalable? So what's happening is usually in the early adoption curve of new technology, the early adopters are all hackers. A hacker is not, in a bad way. It doesn't have a negative connotation. They have deep expertise in certain areas and they want to control a lot of things. Versus in the late stage of a new tech adoption curve, it starts to get more accessible to a much bigger cohort user who doesn't have deep expertise and they need less control. So it always goes into deep control first, usually, and little control later. So we definitely are aiming towards the later stage as the ultimate time we want to target, but it's also extremely valuable to understand what is required to get there. So that's why we partner deeply with Cursor. They are the pioneer trying those ideas. They do have researchers from Frontier Labs, and they want to control every single thing, and at the same time we're also pushing to the boundary. We're doing things that never existed before. Webbing system never existed before because we push the boundary that is unique to this particular setting. Okay, what is unique here is: Typically, if you think about training, training happens, training is very capital intensive, and it usually happens in big companies. They have a lot of money. They put that money to buy very expensive training cluster interconnected with each other. Super expensive, and then once you have those expensive large fleet, usually you don't need to think too deeply how to be efficient. You just focus on doing your work. Cursor is like us, they're a startup, right? Both of us are very capital conscious, and we want to be efficient while we don't want to slow down the research innovation. So together we figured out a very smart way to drive their training process: they do massive post training, which is reinforcement learning based, and reinforcement learning we break that into two pieces. One is the trainer that is tweaking the weights of the model, and it basically generates a new model version constantly, and that new model will deploy to what we call the RL rollout. It basically deploys that new version, interacts with a synthetic environment, a synthetic coding environment or real coding environment, and then gets the reward back to judge if that model version is good or bad. So that's a rough process. And we decouple these two. In the past, in large hyperscaler, they ran that all together. If you think about it, you get 10,000, 100,000 chips all interconnected together through Infinity Band. It's extremely expensive. It's really hard to find. But then you go really quickly, and we designed a fully distributed system. We run across five, six data center regions globally and tap into scattered GPUs, and they're able to run massive jobs, RL jobs. But the challenge there is we need to sync model weights across all these different regions. And how hard can that be? It matters because the latency of delay of sending these weights over is going to dictate how fresh the rewards are. And then if it's too stale, then you are too off. So it's a balance. But we innovate a way we can distribute fresh model weights quickly. It's not too off, so numerically it's still sound while we are not limited by a very expensive deployment of GPU fleets. So those are the innovations we work together with Cursor to push the boundary, leading to their recent model launches. We are very proud of them. Can I ask you a question bluntly, which is: incredible customers have amazing progress they've had with you. And it's wonderful to see that partnership. It's a very large customer for you. How do you think about the concern of a Cursor churn in the wake of a SpaceX acquisition? Yeah, everyone's concerned. The whole entire industry in terms of application innovation is by model. In the sense, there are few companies that are very successful. They escape velocity, but few of them. So that's the shape of the whole entire industry. And last year, Cursor is one of the few. I would say all model companies are concentrated on Cursor. We concentrate on the same group of app companies. And since then, it does change. So we do have a very healthy, diversified customer base. Especially, I think last year is the year of coding. I think all major coding companies are on us. And this year is the year of co-work. And co-work is much more diversified by itself than coding. Because they're general purpose co-work. For example, general purpose co-work to help you do all kinds of research. You want to ask, hey, what will be the NVIDIA GPU price two years later? What will be Anthropic's stock price after IPO? So those are deep research, general purpose deep research. Or there are so many different categories of special purpose co-work: legal, we just talked about two great legal companies; finance, customer support, recruiting, sales, marketing, healthcare. So there is a very broad set of co-work space of innovation app company. They are doing really well. And we have them as our customer base. And then more interestingly, we start to see an uptick of consumer-facing companies all start to look into GNI technology. And they are changing how they are thinking about their traditional business of doing recommendation, for example. And that's very interesting to me because we have obviously worked at a huge recommendation system in the world, Meta. And we are very eager to see how that transforms into a new economy for us. I'm sorry for being naive here. Do people work with just one provider in the inference space like you? Or do they work with you and with Together or anyone else in the space? I think people are more in tune to multi-vendor strategy in this space because they don't know what's happening if you're safe to have multiple providers to balance things out. But we don't view ourselves as an inference provider. Again, we view ourselves as delivering specialized intelligence where we help companies tune their model. Give you some numbers. Today we process more than 40 trillion tokens a day. So the majority of those tokens are coming from a customized model, not from off-the-shelf models, but are coming from customized model. What will that token count be end of next year? Anywhere ranging from 20 to 100x could be possible. 20 to 100x? Yeah. We're at a very early stage of S-curve of explosion right now. 20 to 100x? If it's 20 to 100x, the idea that we are in a capex bubble is ridiculous, and we are desperately needing far more capex than we are ever suggesting for compute. Is that right? So that is right. At the same time, I think Jensen has a five-layered cake, five-layered AI cake from top down: application, model, infrastructure, chips, energy. We are bottlenecked by the lower part of the AI cake in terms of supply chain. So... Being energy. Being energy, being chips, I think in the physical world, how fast we can manufacture, because in history, all this industry is not designed for massive scaling. Speaking about 100x scaling, no one was designed for that. I talk with many manufacturers, it's bottlenecked by small parts. Transistor. The smallest tiny parts that hold off the whole manufacturer line of servers that can deploy to data center and be used to generate tokens. Do you have to be, going to Jensen's five-layered AI cake, do you have to then be full stack to win or to reduce dependencies? We've seen OpenAI come out with Jalapeno, terrible name, Anthropic talking to Samsung about building their own chips, DeepSeeker building their own chips, Zuck came out with Meta building their own chips. Do you have to be all of it? It really depends on the company philosophy. To us, agility is everything, and we need to earn the right of building anything. So focus is everything for us. And we want to focus on where we add the biggest amount of value based on our strength. And we would like to leverage other people's strength to build on top of. So in particular, we want to run everywhere, all possible AI chips in the world. We don't want to limit it by how much chips we can bring into Do you have to be going to Jensen's five layered AI cake? Do you have to then be full stack to win or to reduce dependencies? We've seen OpenAI come out with Jalapeno, terrible name, Anthropic talking to Samsung about building their own chips, DeepSeeker building their own chips, Zuck came out with Meta building their own chips. Do you have to be all of it? It really depends on the company philosophy. To us, agility is everything, and we need to earn the right of building anything. So focus is everything for us. And we want to focus on where we add the biggest amount of value based on our strength. And we would like to leverage other people's strength to build on top of. So in particular, dollar, we want to run everywhere, all possible air chips in the world. We don't want to limit it by how much chips we can bring into our data center, whether we're constructed or we rented. But over time, when the business grows very big, I still remember when Meta was young, they didn't build everything. And when they're big, it makes sense to build. You earn the right to build for your own giant traffic, and if it saves five times more cost, then you should go do it, right? But I think at the early stage, that's why I tell you an interesting story in the coding space. I would say Cursor is the first company that decided to work with us early on. I still remember when they worked with us, they were single-digit million-dollar. Very small. That's only two years ago. They grew by a thousand X over two years. Something like that. But they decided to work with us early on because they recognized they only want to focus on product innovation and later on research. They do not want to focus on this platform innovation. They know we are putting all our R&D there, and they want to find the best partner to win big. So I do think that's the right mentality, to specialize. And we want to specialize. We do not want to own the whole entire stack. That's not our goal as a company. I'm sorry to be harping on that. Why does Jensen skip your layer of the cake? Because he's doing Nemotron with models. Why does he not want to cannibalize your business too? Well, Jensen is not building a cloud either. You can say, hey, Jensen probably has all the rights to build an media cloud. So he's not building a cloud infrastructure. I think he mentioned that as well. People asked him that question. And he also mentioned he wants to specialize in what they have the rights to do. Why models? I think it's a pure supply chain question. If U.S. doesn't have a U.S. native open model, it's a problem. It's a supply chain problem. So he is solely there to solve the supply chain problem. But if there's no supply chain problem because a company like us is providing this specialized intelligence platform layer, then he doesn't need to worry about it. So he just wants to make sure the whole entire five layers of AI cake is flowing. There's no blockage. And if there's a blockage, he's interested in solving those problems. Mark Benioff, one of your investors, I think, in the new round, which obviously this will come out after the round is announced, said that he spends about 3.8% of developer salaries at Salesforce on Anthropic and Claw Code. And I think it's a useful analogy because if you assume that that is what's spent on Claw Code and coding tools, that says one side of the market. But if it's 20%, wow, we're underestimating how big these companies can be. When you think forward a year or two, what percent of developer salaries do you think we'll spend? Is it less because these tools will get cheaper, or is it more because they'll get better and better? I do think the cost of token will go down drastically. Because it hasn't so far. It hasn't so far because of supply chain constraint. But we are living in a free economy. So think about whenever there's shortage, price is high. Price is high. High price will invite a lot of people coming in to solve the problem. Competition will bring down the cost. And then eventually it will lead into a very economical solution. But actually that's good for everyone because much more affordable infrastructure will invite more usage. So my prediction is, with the decrease of infrastructure, that's where it comes to my prediction of how far next year will look like. Because the infrastructure cost will go down, and usage will explode because of that. So the moment you don't think about that as a problem for you, if it's a utility, you just use it. How much will token cost come down? Is it like a halving? Is it like it'll be a hundredth of the cost? So there's different ways to think about it, because it's not all tokens are equal. I think we should establish a best practice to evaluate the token economy per task because different models have different way of spitting out tokens. Some are much more verbose than the other. So you can imagine one model is 2x cheaper than the other, but it's 2x more verbose to solve the same task, and then they're the same cost. Right? But overall, I think as the model quality improves, I think being precise is going to be part of the optimization. And so that's one level of optimization, to solve one task, we should need less tokens. Okay? And the second is, for one token, and how to do that is you need to customize the model to solve your problem especially better and more precisely. That goes into model tuning. And second is, for each token spit out from those models and processed by those models, we also specialize in making the unit of economics much better through our platform. And third is underlying infrastructure, like the GPUs, the surrounding memories, and all this. Today is under stark supply chain constraint. It is going to get much better. Situation will get much better. I don't think probably in the next one year or a year half, this situation will change. But in the long term, two to three years, it should change. And that cost will compress. So overall, I can imagine 10x cost reduction in the next three years. And this 10x cost reduction will drive 100x usage. You said there about token efficiency and how you enable your customers to be much more efficient. With that efficiency, you do charge more. When I did the research, when I compared to competitors, I got, like, together's price king. And I don't mean this disparagingly, but they're cheaper. If you want cheap, you go there. And respectfully, if you want better quality product, you go to you. But it is more expensive. Do you think that's a fair assessment and a fair analogy? I think we're probably not comparing Apple to Apple in the sense that, again, it goes back to our business. The majority of our traffic is customized model. And we optimize for quality, number one, always quality. Quality as in model quality towards your applications, your specific business, your use case, and so on. The second is when we deliver those models in inference, it is also quality. And we care about quality so much, we do extreme things. For example, during training time, there's a very hard thing to achieve. It's called zero KLD. It's a little bit technical. The idea here is zero KLD. KLD is a measure of quality. And what it means is between the training system and inference system, when model moves over, we have bit equivalents. So the numerics are fully the same. We do not lose a bit of accuracy. That's really hard to achieve. But the reason we push that, we deliver that, and the reason we push that is because we know our primary business is in model customization and inference of customized model. We want our customers' every single dollar invested in training to maximize it. And then if the training-inference boundary is not bit-wise equivalent, they just drop the quality down. And then it's like you pay your training investment by discounted quality. Why do you do that? So quality first, and quality does bring additional value. And that's why we are not interested in commoditized one-size-fits-all, this off-the-shelf model deployed in the same way for everyone, that kind of business. We'll always customize model deploy in a unique way for your particular workload. Two questions. Do you have to have an FDE model to make the customized model efficient? As a matter of fact, we do have an FDE team. It's called Applied Machine Learning Engineering Team. So their primary job is to accelerate this customized deployment. As a matter of fact, we'll also build an agent to automate a lot of deployments. Given where we are in the stack, a lot of the complexity that we have, we have a margin structure that's a little bit different from traditional SaaS being 80%. down. And then it's you pay your training investment by discounted quality. Why do you do that? So quality first, and quality does bring additional value. And that's why we are not interested in commoditized one-size-fits-all, this off-the-shelf model deployed in the same way for everyone, that kind of business. We'll always customize model deploy in a unique way for your particular workload. Two questions. Do you have to have an FDE model to make the customized model efficient? As a matter of fact, we do have an FDE team. It's called Applied Machine Learning Engineering Team. So their primary job is to accelerate this customized deployment. As a matter of fact, we'll also build an agent to automate a lot of deployments. Given where we are in the stack, a lot of the complexity that we have, we have a margin structure that's a little bit different to traditional SaaS being 80%. I don't know the margins precisely here, but they're traditionally in the 30% to 40% range for where we are. Is that the new normal for where we are? I don't think that's the new normal. I think that is a reflection, at least for us. I don't know other companies. For us, it is a reflection of we are in a hypergrowth phase. During hypergrowth phase, you have the choice, right? You either optimize. To me, margin optimization is a constraint problem, as in, hey, we want to go to 70% margin, we want to go to 80% margin, and then we are going to go backwards and impose those constraints to guarantee those margins. Usually, constraints slow down innovation. For example, during system development, and in a high-velocity system expanding phase, we don't want to overbuild because we're in high experimentation while testing what will stay or not stay. Optimization doesn't make any sense. Once we know this is a system that we want to build 100%, and we are going to scale this a thousand times bigger, then we go optimize the heck out of it. I think you might not think about business the same way. We're in hyper-growth. If our focus is only optimize growth margin, we absolutely can do that, but we are sacrificing the speed of growth as well because we want to go everywhere. We want to go into different geo-regions. We want to go and tackle different use cases. We want to constantly create different product lines, and those are not the time for optimization. That's my opinion. So we will be able to increase margin without moving into different layers of the stack? Not into, we absolutely are not going to move into application layer. Very clear to us. Whether we will move down into, for example, you mentioned build data centers and so on. That could always be on the table, but the question is timing. Isn't the statement, you either die or you live long enough to build your own data centers? As Elon or Zuck now are spending, I think, 10 billion on the latest data center in Canada. Would you like to build data centers? So I have built data centers at Meta. Also, lots of innovation possible there. There's no one-size-fits-all as well. And building a GPU-native data center is also interesting, especially. I think there is a potential direction of building. So it's a trade-off, right? From an operation point of view, it's much better to build a homogeneous deployment. It's all the same chips, all the same SKU, as big as possible, and run multiple workloads, so it's fungible, right? It's very easy to manage it. You build one principle, one process to do maintenance operation. But, again, it goes to optimization. But once it's so big, then any optimization is going to drive a lot of economic return. For example, we're talking about NVIDIA recently acquired a company also called Guac with Q. GPU. It's a large SRAM- based ASIC accelerator. I spoke to Jonathan before this show. Jonathan is excellent. He said what a fan he is of yours. Oh, also a fan of his. So, but it's a great combination between a Flops- Intense GPU and SRAM- Intense ASICs. Because Flop- Intense is really good for first half of LN processing. It's pre-fill. It's called pre-fill. Processing and prompt and so on. And SRAM- Intense is really good for generation. That's just the nature of the model architecture. It's great to combine these two instead of running heterogeneously on the same chip, right? But that requires a very unique system design and deployment into data center. And it is heterogeneous. Actually, before I really mean homogeneous, design is much better for operation, and this is heterogeneous. And then how to operate this heterogeneous design requires unique innovation in data center deployment and so on. So data centers aren't commoditized. You can specialize in data center deployment, and one data center is better than another, and data center deployment can be done well and badly. Data centers are so complicated, right? If you think about the beginning, all the way from construction to power deployment, you have the right power to come in, right fiber channel, the right cooling, especially neural chips requires liquid cooling, to get all this right, and the parts can fall apart, and how to replace them. It is all very deep expertise. It's no joke. It's not tomorrow I can be a data center operator. I cannot. Is that not where you would bet long on China, with the greatest respect? Especially in the US, one of the biggest barriers to data center deployment is policy and local legal infrastructure that prevents it. In China, you don't have any of that, and data center deployment is much faster. I think in general infrastructure, the physical infrastructure construction in China is going really fast. I literally see some kind of crossover bridge. It's being built within a week. The velocity is very high there. And there's a highway close to my home. After one year, it's not done yet. So it's also a crossover. So I do think there's a unique strength, probably because of the population density, and they are specializing in those construction work. But I do think here we also have those specialty people. It's just, I heard even electrician is under severe shortage. We are under global supply chain constraint here. What change would moving into the data center layer cause to margins? Would that take it from 30 to 50? Would it be not that meaningful? What would that change do to margins? How we calculate gross margin is interesting these days. Because how long does hardware depreciate has significantly changed. In the past, it's six years. Solid six years. And hardware release is usually three years. That's fast. And now, within a year, from one vendor alone, we have three SKUs. And the newer model usually runs the best on the newest hardware. Model depreciation is also very fast. Every week we are launching a new model. And the model is peak in its value before the next model comes out. And the new model likes the newest hardware. And imagine this cadence after two years. Which model runs on the two-year- old hardware? It will be a two-year- old model. And are those models still valuable? So I think that's the real dynamics we're facing right now. The hardware will last for six years still. But what you're saying, the speed of model development far outstrips the speed of chip and hardware depreciation. The speed of model development definitely is the fastest. But even the hardware innovation itself is the fastest. So after three years, if every year there's three hardware SKUs, after three years there's nine hardware SKUs in between. Do you still want to go back to nine generations older hardware, running three-year- old model on that? That's questionable. Maybe there's a war, it's still valuable, but with this pace of innovation, it's questionable. Now with a different depreciation cycle, it changes dynamics of build versus buy. And again, it goes back to my original thesis of do you optimize for growth or do you optimize for growth margin? It's all about timing. How do you think about that question for yourself when you're sitting there in an armchair on a Sunday afternoon, thinking we're optimizing for growth now? When is that time to optimize for gross margin? Well, I would say we want to optimize for both. So here's how I think about it. Optimize for growth requires a lot of business planning, assuming there's product market fit. Optimize for growth margin is optimize for differentiation. I think I want to avoid over- optimizing for growth margin, but we should optimize for growth margin continuously, as in, we should optimize for product differentiation continuously. There's no question about it. And I think we want to continue to optimize towards a healthy growth margin, which allow us to grow really fast. And it's a trade-off, and we don't take compromises. The compromise, as in, we over- optimize growth margin to result in a very slow growth, right? One possible way to optimize growth margin, we do not grow at think about that question for yourself when you're sitting there in an armchair on a Sunday afternoon, thinking we're optimizing for growth now. When is that time to optimize for gross margin? Well, I would say we want to optimize for both. So here's how I think about it. Optimize for growth requires a lot of business planning, assuming there's product-market fit. Optimize for growth margin is optimized for differentiation. I think I want to avoid over-optimizing for growth margin, but we should optimize for growth margin continuously, as in we should optimize for product differentiation continuously. There's no question about it. And I think we want to continue to optimize towards a healthy growth margin, which allows us to grow really fast. And it's a trade-off, and we don't take compromises. The compromise is in we over-optimize growth margin to result in very slow growth, right? One possible way to optimize growth margin: we do not grow at all. We just optimize the heck of it. I know we can climb to a high number, but that's an absolute disaster outcome. Okay, interesting. If we just said, hey, sold gross margin, we're going to take it from 30% to 10%, is it a winner-take-all market where we could eat up everyone else's lunch and then optimize? Long-term situation, we do see a particular industry will oscillate and start to settle with a few good ones. Legal, for example. I was at a dinner table, and interesting, it seems like there were a lot of those companies around two years ago, but now it's pretty much two. So I think it's a long game. How do you see the more mature state of your market? Is it like a cloud market where you have Azure, AWS, GCP, or is it an Uber and a Lyft where one takes 90% and the others fight for scraps? Me to specialize in intelligence, to start to seriously think about owning their intelligence is better than renting, because going back to this optimization, when is the good timing, right? So it's the same question we're answering for ourselves when build versus buy, and our customers also think about build versus buy or build versus rent or own versus rent, right? I think AI journey, or AI adoption journey, has gone further along into a lot of companies, has meaningful traffic. A lot of company is deploying AI into production. A lot of company is at the phase of scaling, and that's where optimization kicks in. When optimization kicks in, you need to have control to optimize. If you don't have control, you just don't have the range to optimize. And for you to have the control, you have to build on top of some open model. You have to turn your data into intelligence. That's pretty much the path we have seen. So many companies across industry, they reach the same conclusion. They are moving towards this direction. Speaking, speaking, speaking. Owning your intelligence versus renting it, that does apply to a national layer. And when we've seen Fable be banned in some cases by the administration briefly for 19 days, especially in Europe, we suddenly went, oh my gosh, we cannot be at the hands of OpenAI and Anthropic, where we can just be banned, and our health services sit on the infrastructure of something that large nations or nation blocs own, sovereign models. I definitely see that possibility. I also see, if we think about the general intelligence model as the electricity there, as a power line, every country should own their own power line, right? So I do think that is a very scary moment, is my power line is going to be cut off and all my fundamental day-to-day is not working, because I feel so frustrated whenever there's a power outage in my home alone. I feel so frustrated when I cannot access my Wi-Fi. I feel so anxious. So obviously, operating the country is extremely important, built on top of this fundamental baseline. And for every single company, same thing. It's not just whether a country should have their unique sovereign independence, but every single company should have their independence. You don't want any single person to cut you off. That's an extremely scary moment. Why would you move into the data center space, but you wouldn't move into the chip space? Because I know building a chip is extremely hard. I thought so too. Okay, again, I admit to being a moron, which is why I think the show is a little bit successful. I thought so too, but then how come everyone is seemingly doing it as if it's just another product? As I said, OpenAI, Anthropic, DeepSeek, Meta, we're building our own chips now. I think Meta has been building their chips for more than five years, way more than five years. And MTIA has been a project since 2018, maybe earlier. Because Meta has been investing in AI for a long time, pre-gen AI, and they have a huge AI workload to focus on ranking recommendation. And Meta has been building other hardware as well in the past. So whenever the usage has passed a certain threshold, it makes economic sense for you to build it, build underlying supply, right? And you can specialize towards your workload. And that's another form of specialization, is specialize to bake your logic into hardware. And this hardware is purpose-built for your particular workload. And you better make sure this workload doesn't change, because once the hardware is taped out, it's really hard to go back and change it. It's possible. It's very costly. So once your workload is stabilized, once your business is stabilized, it doesn't change too often, then that's the time to consider building a chip. I still see the whole AI world, especially models customization, is very dynamic, very, very dynamic, workload patterns, very dynamic. So think about how much energy in the application space people are experimenting all kinds of things. You don't know which one is going to take off, and they will just take off quickly, and which one, once they take off, is going to sustain. And a few ones will sustain. Then that's the time, oh, now we know this is the pattern, and now we should probably encode this pattern into hardware and bring this hardware into a data center and so on. It's all cascading, and then it's going to cascade down. To me, it's a fernal question. Where are we in the stage of fernal maturity? Fernal, I mean we're still in the early stage of workload maturity, fernal, to warrant a chip that will be durable. So now you go back to, oh, we have so many accelerators that are successful. Some are really successful. But remember those ASIC company, they started before JNI AI. They started from some ceases to optimize some workload, and they all pivot to AI and trying to fit in AI workload. It's almost like you bet before this AI workload emerges, and now it becomes a serendipity question. Are you lucky enough this just works, right? And some really worked. So fundamental design of putting a lot of XRAM on the chip is great for AI model because they are memory-hungry, and these really accelerate the execution of inference and so on. So those work, and some don't work. What do you see as the greatest bottleneck today? I think it was when I had Jonathan from Grok on the show who said HBM was the greatest bottleneck, and that's why you've seen the 5x increase in price. What do you see as the greatest bottleneck that people don't talk about enough? I still think we don't have a great system for very large model. I really believe the fundamental lower-level infrastructure costs will go down. For solving tasks, we should need less token. That will increase, so collectively the cost will significantly reduce. Therefore, we can run the highest intelligence model much more ubiquitously in the future. But we don't have a system designing for that. For example, we don't have a great system designed for 10 trillion parameter models today, and that will require very smart engineering co-design from the model to the customization serving platform layer all the way to chip layer. The chip is not individual chip, but the system collection chips and system as a total package. I think there's still a lot of innovation we can do. I think recently you announced that you were at 800 million in ARR. Incredible feat and scaled so fast. What is that at the end of this year? We think we can at least double by the end of the year. Wow. What's so interesting for me as a venture investor, I've been investing for 10 years. We used to be in the day where Slack was the golden child by 1 to 10 million in revenue in 18 months, was amazing. And now we have companies like Fireworks where you scale to 800 million in revenue in a matter of years, and you mentioned Cursor scaling to billions in revenue in a matter of years. The speed of company revenue growth is just unparalleled. I think it's because there's a fundamental disruption in this technology that is all empowering, and all empowering in a sense it reach out to every individual one of us to be creative, and it unleashed a lot of creativity that we just don't have access to. And that's why we're seeing this phenomenon of extremely fast growth, because of demand. Final one before we do a quick fire. You hired George Hugh, who was president of Salesforce. He's exceptional. He's one of the most direct, no-BS operators I've ever met. But you met him a couple of years before, a year before, and you were like, oh, we're not ready for you yet. Why did you say that, and why did you decide now was the time, right? So a year ago, I think we were probably just 50 people. 800 million in revenue in a matter of years. And you mentioned Cursor scaling to billions in revenue in a matter of years. The speed of company revenue growth is just unparalleled. I think it's because there's a fundamental disruption in this technology that is all empowering, and all empowering in a sense. It reaches out to every individual, every one of us, to be creative, and it unleashed a lot of creativity that we just don't have access to. And that's why we're seeing this phenomenon of extremely fast growth, because of demand. Final one before we do a quick fire. You hired George Hugh, who was president of Salesforce. He's exceptional. He's one of the most direct, no BS operators I've ever met. But you met him a couple of years before, a year before, and you were like, "Oh, we're not ready for you yet." Why did you say that, and why did you decide now was the time? Right, so a year ago, I think we were probably just 50 people. So today we're at 200 people. We're still not that big. Wow, you're 4 million ahead. Yeah. Wow. So at 50 people, I'm more thinking about scaling the product first than massively scaling the business. And we talked, and I have huge respect for him, but I know he is a legend. He's a legendary operator in Silicon Valley, and I just felt like we're too small for him. And I told him that: we're probably too small for you, but I would like to work with you in some capacity. So he helped me. He helped me actually build out the team, interview a lot of executives. His feedback is always well balanced, very thoughtful, and we started to work together in that capacity until, I think, the end of last year. We were growing really fast, and he knows, and we started talking seriously. And that early relationship paid off. So he's really cool. He's really cool in the sense that he did a lot of things, great accomplishments in the past, but I find a unique character about him is he's extremely experienced, has a high altitude of business vision, but he's also very curious. He doesn't make assumptions: "I know it all. I've seen all the movies. It's the same movie, and let me just direct this movie as I did in the past." So he didn't come with that attitude. He knows AI goes at an insanely fast pace, and he's learning along the way, but also fully embraces AI. Actually, his team, our GTM team, is using all kinds of AI agents. They're sharing skills so they maximize their productivity. And he knows we have a superlinear demand curve, and there's just a certain pace we can build our GTM team. In order for us to catch this curve, we need to build our team, but the team needs to also have increasing productivity to match. So that's a problem he's solving, and I feel very fortunate to work with him. And in general, I feel in the AI space, the unique part is people need to have very special traits, almost contradictory characteristics. For example, very experienced but super curious, and a fast learning curve. Or Dima, we talked a little bit earlier. He is brilliant, high intellectual horsepower, but extremely humble. It's a weird combination. He's amazing, and he's almost cynical in an Eastern European way, but also at the same time very humble. Can I do a quick fire round? Okay, let's do it. Okay. What has changed your mind the most in the last 12 months? I think how fast we grow. I changed my mind because I have been quite worried about too big a team too early. So that's why, when I met George, I told him we're too small for you, because I don't intend to grow very fast in terms of people. I worry very deeply about getting slowed down and losing agility and velocity. But since then, we have been very aggressively using tools. We have developed our own unique way of hiring a certain type of people that we know will be charging forward with high velocity, an extreme sense of ownership, very communicative, and never take no as an answer. So we also learned how to get those people, and now I feel much more comfortable scaling really fast. What's your type of people? And I know that sounds weird, but our type of people is actually really specific. Pretty much only hire immigrants. British people don't work very hard, sorry. Very scientific and rigorous, use data for most things. I actually think creativity often comes from data and is informed by data. And unwaveringly accountable and ownership: nothing is anyone else's fault. It's all my fault, even if it's someone else's fault. That's a 20VC person. What would you say yours is? In a weird way, it's not competence. It's weird. We want people with high competence, but more importantly, the strong indicator of whether they will do well in this way, especially at Fireworks, is whether they are really built for taking extreme ownership. Extreme ownership, as in we are not putting people in any boxes, and we're just stacking the boxes together into a tower. People just automatically claim, "Hey, this is an end-to-end problem. I'm going to see through the whole thing and work with a bunch of people to make it happen, and I'm going to deliver it no matter what." So those kind of people have the highest, longest mileage, and their growth curve is amazing also. What's your biggest lesson from working with Jensen Huang on what makes him so special? He's everywhere. I seriously think he has a clone, hundreds of Jensens somehow plugged in. For example, I send him an email, he will reply in one minute. I just don't understand how. He's constantly in details. But now I operate a company for four years, I understand why he's doing that, is that defines velocity. Because what is leadership? Leadership is just judgment. It's not privilege, it's judgment. You have the context. You need to have the right context to make the right judgment for the team. Especially in a high velocity space, if you do not know what's happening, what works, what doesn't work, what are the gaps, you make the wrong call. In a slow-moving space, you can wait for the cascading information up and down and make those calls. But in a fast iteration space, you just cannot wait, because it's guaranteed there is information lost in transition, layer after layer, people after people. It always happens. And not knowing what exactly is happening, and having the position to make judgment, makes bad leadership. And he has demonstrated through his own example, even before this crazy AI thing, he's operating that way. And before, I was admiring him for his sheer amount of capability, capability of doing that. Now I understand the wisdom behind that, because I also operate that way. I need to know what's happening on the ground to make the judgment for the company. What did you wait on in the Fireworks journey that you wish you hadn't waited on? Marketing. We talked about it. So we are a little bit nerdy in this way, that at the very beginning of our journey, we didn't discuss it, but we feel product will speak for itself. At the end, product stands, and we want to devote all our effort and focus on building product, work with the customer, validate product-market fit, and go from there. And we didn't spend much time on marketing at all. We didn't prioritize educating our customer on what's the right direction to think about the trend and the value. But we do think now, I do think it's important. Marketing is not about fluff. It's more about education. It's more about clarity. And we are working on that. What area of AI is underinvested in today, in your mind? You mentioned cooling or servers. What areas are underinvested in? I think AI has the sexy part. This is such innovative, creative technology, and building something on top of it is the focus. But monitoring the ROI, I think the industry is starting to pay attention to it, but eventually that's what matters. It's not how much you spend. It's what is the return, and what is the cost, and what is the attribution. So I think in the next couple of years, as AI is getting more into production, there will be a lot of focus on getting that clarity and getting that discipline out. So the token maxing is just, I think, a thing in time, but we'll quickly move into ROI maxing, which is all about running a business. What large customer do you not have that you would most like to have? So we haven't spent too much time in the traditional enterprise segment. I think that's just because we were very small. And now, as we build out our company, I do think even without us investing, we have customers like Geico, like Capital One, like Mercury Insurance, and RBI. All these companies, even without us pursuing traditional enterprise, they come to us and they are customers. But I do think that's a very big market. What has to happen before the end of the year that hasn't happened for you to consider it a good year? I'm confident in our capability of driving the business, and to me, this is a year I want to prove we can scale quickly while keeping the same velocity. And that's very important to me. If we reach that point, we should improve point, and next year I have a lot more confidence to continue to scale extremely aggressively. I want to make sure we do it right this year. Final one for you: what does no one see about the next three years that you see very clearly happening or not happening? I really see people very small. And now, as we build out our company, I do think even without us investing, we have customers like Geico, like Capital One, like Mercury Insurance, and RBI. All these companies, even without us pursuing enterprise, traditional enterprise, they come to us, and they are customers. But I do think that's a very big market. What has to happen before the end of the year that hasn't happened for you to consider it a good year? I'm confident in our capability of driving the business, and to me, this is a year I want to prove we can scale quickly by keeping the same velocity, and that's very important to me. If we reach that point, we should improve, point, and next year I have a lot more confidence continue to scale extremely aggressively. I want to make sure we do it right this year. Final one for you: what does no one see about the next three years that you see very clearly happening or not happening? I really see people will own their every single company will own their own intelligence as a must-have. It's not optional. That's a trend I'm seeing because there's an analogy to software. There's a reason why every company builds their own software stack. There's no standardized software you just use off the shelf to solve your problem because every single company is solving a unique problem, and they want to build software because they want to have full control. And obviously, they will pick and choose which part of the stack they want to build themselves, which part of the stack is common knowledge, there's no point of building. But every single company own their own software stack. Obviously, we're talking about in the SaaS time, right? So same, I think every single company should own their intelligence! Linh, it was Matt that introduced us first. I've had the joy of getting to know you, and obviously, George, I can't thank you enough for joining me, for coming in person. It is so wonderful to do it in person, and you've been fantastic. That's an amazing studio you did. You asked a lot of interesting questions. I had a lot of fun talking with you. We did a lot of research before, huh? Yes, you did. Thank to talk to fire round okay let's do it okay what have changed your mind on most in the last 12 months I think how fast we grow I change my mind because I have been quite worried about too big a team too early so that's why when I met George I told me we're too small for you because I don't intend to grow very fast in terms of people I worry about slow down getting slow down and lose agility and velocity very deeply so but since then we have been very aggressively using tools we have developed our own unique way of hiring certain type of people that we know they will be charging forward with high velocity as extreme sense of ownership very communicative and never take no as answer so we also learn how to get those people and now I feel much more comfortable skating really fast what's your type of people and I know that sounds weird but our type of people is actually really specific pretty much only hire immigrants British people don't work very hard sorry very scientific and rigorous use data for most things I actually think creativity often comes from data and is informed by data and unwaveringly accountable and ownership nothing is anyone else's fault it's all my fault even if it's someone else's fault that's a 20 VC person what would you say yours is it's not in weird way it's not competence it's weird we want people with high competence but more importantly the strong indicator whether they will do well in this way especially in fireworks is whether they are really built for taking extreme ownership extreme ownership as in we are not putting people in any boxes and we're just stacking the box together into a tower people just automatically claim hey this is an end-to-end problem I'm going to see through the whole thing and work with a bunch of people to make it happen and I'm going to deliver it no matter what so those kind of people have the highest longest mileage and their growth curve is amazing also what's your biggest lesson from working with Jensen Huang on what makes him so special he's everywhere I serious think he has a clone of like hundreds of Jensen somehow plug for example I sent him an email he will reply in one minute I just don't understand how he's like constantly in details and but now I operate a company for four years I understand why he's doing that is that defines velocity because what is leadership leadership is just judgment it's not privilege it's judgment it's you basically have the context you need to have the right context to make the right judgment for the team if and especially in a high velocity space if you do not know what's happening what works what doesn't work what are the gaps you make the wrong call in a slow moving space you can wait for the cascading information up and down and make those calls but in a fast iteration space you just cannot wait because it's guaranteed there is information lost in transition layer after layer people after people it always happens and not knowing what exactly is happening and having the position of make judgment makes bad leadership and he is demonstrated through his own example even before this crazy AI thing he's operating that way and before I was admiring him in his sheer amount of capability capability of doing that now I understand the wisdom behind that because I also operate that way I need to know what's happening on the ground to make the judgment for the company what did you wait on in the fireworks journey that you wish you hadn't waited on marketing we talk about it so we are a little bit nerdy in this way that at very beginning of our journey we kind of we didn't discuss it but we feel product will speak for itself at the end product stands and we want to devote all our effort and focus on building product work with the customer validate product market fit and go from there and we didn't spend much time marketing at all we didn't prioritize educating our customer what's the right direction to think about the trend and the value but we do think now I do think it's important marketing is not about flows it's more about education it's more about clarity and we are working on that what area of AI is under invested in today in your mind you mentioned like cooling or servers what areas like under invested in I think AI has the sexy part this is such innovative creative technology and build something on top of it is the focus but monitoring the ROI I think the industry start to pay attention to it but eventually that's what matters is not how much spend is what is return and what is the cost and what is the attribution so I think in the next couple of years as AI is getting more into production there will be a lot of focus in getting that clarity and getting that discipline out so the token maxing is just I think a thing in time but we'll quickly move into our I maxing which is all about running a business what large customer do you not have that you would most like to have so we haven't spend too much time in traditional enterprise segment I think that's just because we were very small and now as we build out our company I do think even without us investing we have customers like Geico like Capital One like Mercury Insurance and RBI all these companies even without us pursuing enterprise traditional enterprise they come to us and they are customers but I do think that's a very big market what has to happen before the end of the year that hasn't happened for you to consider it a good year I'm confident in our capability of driving the business and to me this is a year I want to prove we can scale quickly by keeping the same velocity and that's very important to me if we reach that point we should improve point and next year I have a lot more confidence continue to scale extremely aggressively I want to make sure we do it right this year final one for you what does no one see about the next three years that you see very clearly happening or not happening I really see people will own their every single company will own their own intelligence as a must have it's not optional that's a trend I'm seeing because there's an analogy to software is there's a reason why every company builds their own software stack there's no standardized software you just use off the shelf to solve your problem because every single company is solving a unique problem and they want to build software because they want to have full control and obviously they will pick and choose which part of the stack they want to build themselves which part of the stack is common knowledge there's no point of building but every single company own their own software stack obviously we're talking about in the SaaS time right so same I think every single company should own their intelligence ! Linh it was Matt that introduced us first I've had the joy of getting to know you and obviously George I can't thank you enough for joining me for coming in person it is so wonderful to do it in person and you've been fantastic that's an amazing studio you did you asked a lot of interesting questions I had a lot of fun talking with you we did a lot of research before huh yes you did thank to talk to