How Many Will Actually Get Built & Is Energy AI's BIGGEST Bottleneck? | Positron AI Co-founder
Description
Thomas Sohmers is the Co-founder, CTO & Chairman of Positron AI, building chips to make running AI dramatically cheaper and more energy-efficient. The company recently announced an $875 million Series C at a $5 billion valuation, backed by investors including Gavin Baker's Atreides Management, NEA, Valor Equity Partners and Netscape co-founder Jim Clark. ----------------------------------------------- Timestamps: 00:00 Intro 01:03 - What Positron Is Building for AI Inference 01:46 - Why Inference Infrastructure Is Totally Different From Training 04:18 - The Memory Wall: AI’s Next Infrastructure Bottleneck 08:38 - The Hidden Economics Behind AI Tokens 10:32 - Why Anthropic Could Already Be an 80% Gross Margin Business 12:12 - Should Frontier AI Labs Actually Slow Down? 13:59 - Why AI Regulation Could Create a New Class of Technology Monopolies 17:07 - What Happens if the US Paces AI and China Does Not? 22:44 - Is Anti-Data Center Sentiment Becoming a Strategic Risk? 28:24 - How Much of the AI Data Center Boom Actually Gets Built? 30:46 - Is Energy Really the Biggest Bottleneck to AI? 32:41 - Why Sovereign Debt Is a Bigger Risk Than AI Infrastructure Debt 43:16 - Why KV Caching Matters So Much to AI Economics 49:32 - Will Frontier Models Keep Getting Bigger? 51:14 - Will Enterprises Really Own Their Own AI Models? 54:04 - Why Scaling Laws May Still Have a Long Way to Run 55:08 - What Comes After AGI? 57:16 - Why GPT Astra Feels Like a Step Change 01:01:33 - What Happens When Every AI Lab Builds Its Own Chips? 01:04:09 - How Far Can Context Windows Really Expand? 01:09:16 - The Cost of AI Tokens Has Collapsed 60x 01:11:58 - Will AI Stop Being Priced Per Token? 01:13:49 - What Would Actually Burst the AI Bubble? 01:16:31 - How Big Can the AI Data Economy Become? ---------------------------------------------------------------------------------------------- Subscribe on Spotify: https://open.spotify.com/show/3j2KMcZTtgTNBKwtZBMHvl?si=85bc9196860e4466 Subscribe on Ap
Summary
Generated by gpt-5.6-terraAt-a-Glance
- Verdict: Watch fully
- Core thesis: Positron AI co-founder Thomas Somers argues that AI inference is becoming a memory-and-energy infrastructure problem, with KV-cache management, long-context capability, power build-out, and regulatory friction more consequential than raw FLOPS alone.
- Why it matters: For agent systems, the video explains why long-running, high-context, multi-agent workloads can be cheap in compute yet difficult and expensive to operate because persistent context consumes enormous memory and must be intelligently tiered.
- Best use: Use this as a strategic infrastructure briefing: extract the KV-cache and memory-tiering lessons for agent architecture, then treat its market, policy, and geopolitical claims as an opinionated semiconductor founder's investment thesis rather than neutral analysis.
Executive Summary
Somers frames the inference stack as fundamentally different from training. Training is primarily compute-bound and highly parallelizable; autoregressive inference must generate tokens sequentially and repeatedly read model weights, making it memory-bandwidth-bound. He attributes the widening “memory wall” to GPU FLOPS improving roughly 120x from 2014 to 2024 while memory bandwidth improved only about 17x.
His central operational argument is that KV caching is the economic center of modern agentic inference. Reusing prior prompt and generation state avoids recomputing attention, but cache state grows with context length and is unique to each user/session. In SemiAnalysis's Agent X traces of multi-turn Claude Code sessions, Somers cites 96% of tokens as cached, making cache persistence and retrieval more important to operator economics than many teams appreciate.
For agents, he expects the principal constraint to shift from model intelligence to usable context: not merely advertising a large window, but reliably retrieving relevant information from it. He describes a standard hierarchy that keeps active state in accelerator memory, recently inactive state in host RAM, and colder state in NVMe or network storage; the orchestration problem is deciding what lives at each tier without hurting latency or forcing recomputation.
Beyond architecture, Somers is strongly bullish on continued frontier-model scaling, AI demand, vertically integrated custom silicon, and data-center construction. He believes local/on-device models will increase—not cannibalize—frontier-cloud token demand by autonomously filtering personal or enterprise data and escalating harder tasks. His claims about regulation, China, data-center water use, frontier-model capabilities, and AI-company profitability are forceful but frequently speculative, self-interested, or unsupported by independent evidence in the transcript.
Key Takeaways
- Claim: Inference infrastructure should be optimized for memory movement and bandwidth rather than treating raw compute/FLOPS as the primary constraint. | Evidence: Somers says training can parallelize known training tokens and is compute-bound, whereas inference generates tokens autoregressively and must read weights for every output token. He cites approximately 120x GPU FLOPS improvement between 2014 and 2024 versus only 17x memory-bandwidth improvement. | Implication: For OpenClaw-style agent systems, benchmark end-to-end latency, context reuse, memory residency, and bandwidth—not just model tokens/sec or nominal accelerator FLOPS—before selecting routing and serving infrastructure. | Caveat: The 120x and 17x figures are presented by Somers without source methodology in the transcript; workload bottlenecks also vary by model architecture, batch size, serving pattern, and hardware.
- Claim: KV-cache management is a first-order economic and systems-design problem for multi-turn agents because the most valuable agentic workloads are highly cacheable. | Evidence: KV caches store the keys and values derived from prior input and generated tokens, avoiding repeated attention computation. Somers cites SemiAnalysis's Agent X traces of real Claude Code sessions with subagents, where about 96% of tokens across sessions were cached. | Implication: Design agents for stable prefixes, session continuity, subagent context reuse, and cache-aware routing. A provider or self-hosted stack that cannot preserve and retrieve caches efficiently may look inexpensive per token while being materially worse at real agent economics. | Caveat: The 96% rate is from an agentic coding benchmark and should not be generalized to short, stateless chat, one-shot classification, or low-context workflows.
- Claim: Long-context agents face a memory-capacity trade-off: preserving user state can consume more memory than the model weights themselves. | Evidence: Somers estimates that a purported 1.8-trillion-parameter GPT-4 at 4-bit quantization would occupy roughly 900 GB of weights, while a hypothetical 10-trillion-parameter model would require roughly 5 TB. At long contexts, he says an individual session can reach roughly 100 GB, so 50 sessions can exceed model-weight storage. | Implication: Do not equate long-context support with practical long-context serving. Capacity planning for persistent agents must separately budget model weights, active KV cache, inactive session state, retrieval indexes, and concurrency headroom. | Caveat: These model sizes and session-memory figures are illustrative estimates based partly on unverified leaks and expectations, not disclosed production specifications.
- Claim: The standard solution is a tiered memory hierarchy, but it makes cache-placement policy a core orchestration problem. | Evidence: Somers describes active sessions residing in accelerator memory; recently inactive users in host memory, which he estimates at roughly 4–10x accelerator capacity; and older sessions in NVMe, network-attached flash, or colder disk. Each lower tier increases capacity but raises retrieval latency. | Implication: Treat context state as a managed resource: implement explicit TTLs, session affinity, hot/warm/cold promotion rules, cache observability, and privacy boundaries rather than allowing every agent run to accumulate indefinitely in expensive memory. | Caveat: The transcript does not provide concrete eviction policies, cache-hit targets, latency thresholds, or security controls for cross-tier session handling.
- Claim: Compression and attention innovations reduce the cost of long context, but they are trade-offs rather than free efficiency gains. | Evidence: Somers explains quantization from FP16 toward roughly 4.5 bits per value; naive 4-bit reduction can degrade benchmark performance by 20–30%, while more advanced shared-scale methods can keep degradation near 1%. He cites DeepSeek V3's multi-head latent attention as reducing KV-cache size by spending more FLOPS, and says gated delta-net variants can reduce time spent in attention by about 75%. | Implication: Evaluate compression, sparse/linear attention, and cache formats against the actual task suite—especially tool use, coding, long-horizon planning, and recall—not only generic benchmark scores or advertised context length. | Caveat: He asserts that major U.S. labs do not use MLA based on rumor and personal information, and he argues compression carries capability costs; neither point is independently substantiated in the transcript.
- Claim: Local models are likely to become an agentic control layer that increases cloud-frontier usage rather than replacing it. | Evidence: Somers estimates 80–85% of tokens are currently produced by the top four model companies and suggests only about 5% may be on-premises. His thesis is that an on-device or enterprise-local model can continuously inspect email, calendar, messages, and other data, decide what needs escalation, and create many more cloud-model calls than a human manually prompting a chatbot. | Implication: Architect for a two-tier agent model: cheap/private local inference for monitoring, triage, policy checks, and context preparation; selectively route consequential or difficult work to stronger cloud models with explicit autonomy and spend guardrails. | Caveat: This is a demand forecast, not observed evidence; it depends on users trusting autonomous local agents, viable privacy controls, and cloud inference remaining economically attractive.
- Claim: Energy is a durable constraint on intelligence production, but Somers sees construction economics, financing, and permitting—not physical energy scarcity alone—as the nearer-term bottlenecks. | Evidence: He says that if Positron could deliver the compute of a 500 MW Nvidia deployment in 100 MW, operators would likely use the savings to deploy more total intelligence rather than build smaller facilities. He argues data centers increasingly bring dedicated generation, and says the larger constraint is the economic and political ability to finance, permit, and connect new capacity. | Implication: Assume efficiency improvements may expand workload volume rather than reduce total spend or energy use. For AI-ops planning and investment diligence, separate chip efficiency from actual grid interconnection, power contracts, financing conditions, permitting risk, and local community opposition. | Caveat: This view is highly aligned with Positron's commercial position as an inference-hardware supplier and includes broad, contested claims about data-center water use, electricity prices, regulation, and geopolitical influence without supporting data.
Detailed Brief
Context length: headline capacity is not useful context utilization
- Claims: Somers argues that the next meaningful capability gain for coding agents may be enough reliable context to reason across complete and multiple codebases, rather than simply a larger model.; Traditional attention makes expansion beyond current million-token windows difficult because compute and memory costs rise sharply with sequence length.; A stated context-window maximum should be distinguished from a model's ability to retrieve and reason over information placed deep within that window.
- Evidence: He says a one-million-token context can represent only a fraction of some internal large code repositories, while roughly 10 million tokens could hold multiple repositories for cross-codebase agentic work.; He references the RULER and needle-in-a-haystack style evaluations: according to his example, GPT-5.6 retrieved a hidden value about 70% of the time, while GPT-6 Astra exceeded 95%.; He says some early million-context models experienced poor recall above roughly 64,000 tokens.
- Caveats: The model names, performance numbers, and release context cited in the interview are not independently verified within the transcript.; Needle retrieval is a narrow test and does not establish reliable planning, synthesis, source attribution, or software-engineering judgment across a full repository.
- Implications: Use long-context evaluations that measure retrieval, contradiction handling, tool execution, and task completion on Ken's actual repositories and operational documents.; Prefer deliberate context selection and memory systems over indiscriminately loading the maximum available history into every agent run.
Economics, pricing, and market-structure thesis
- Claims: Somers contends that API-model providers have unusually strong economics because cached-token reads cost orders of magnitude less than recomputation while still being monetized.; He believes token price alone is becoming a less meaningful measure because the capability and economic value of a token have improved substantially.; He expects frontier labs eventually to offer outcome- or worker-priced products, while token billing remains useful for commoditized bulk API access.; He sees vertical integration as the eventual risk to data-labeling and data-services companies, since sufficiently capable internal agents could absorb work currently outsourced to vendors.
- Evidence: He claims processing a cached token is on the order of one-thousandth the cost of generating it and references a reported 80-point gross margin for Anthropic's API business.; He contrasts an index falling from about $60 per million tokens five years earlier to below $1 per million, while arguing present-day tokens are conservatively 100x more valuable and perhaps closer to 1,000x in value per unit of intelligence.; He gives a hypothetical of a future superhuman virtual agent sold as unlimited annual access—for example, $1 million per year—rather than token metering.
- Caveats: These margin, cost, and future-pricing statements are directional claims from a supplier to model providers, not audited financial analysis.; Higher model capability does not automatically translate into realized customer value; reliability, integration cost, human review, liability, and workflow adoption determine economic capture.
- Implications: Track cost per successful workflow, cost per accepted artifact, cache-hit rate, and human intervention per completed task alongside cost per token.; Avoid underwriting data-service or middleware moats without testing whether a frontier lab or agent platform can bring the function in-house.
Policy and build-out perspective
- Claims: Somers believes most planned hyperscale capacity will still be built, although individual projects may relocate after local opposition.; He opposes broad “pacing the frontier” approaches, arguing they could centralize technological capability in large companies and governments while adversarial states continue development.; He argues that resolving AI's broader alignment problem requires “human alignment” around economic incentives, energy build-out, data-center siting, and regulation.
- Evidence: He points to U.S. western land, especially Nevada, as suitable for solar, geothermal, nuclear, and data-center development, and mentions Panthalossa's ocean-based data-center/pumped-hydro concept as an alternative to land or space.; He expects political education and relocation to mitigate local data-center resistance rather than cause an existential capacity shortfall.; He identifies sovereign-debt and currency-devaluation risk as a greater macro concern than AI infrastructure companies failing to meet revenue targets.
- Caveats: The discussion is polemical: Somers characterizes anti-data-center sentiment as “almost entirely a Chinese psyop,” asserts several water and grid claims without cited evidence, and offers geopolitical judgments that should not be treated as factual findings.; Data-center development can face legitimate localized constraints involving transmission, water, emissions, land use, reliability, rate design, and community consent even when aggregate generation grows.
- Implications: For infrastructure investments or capacity commitments, perform site-specific diligence rather than relying on aggregate claims about land, water, generation, or grid impact.; Monitor permitting, interconnection queues, local political opposition, debt markets, and utility rate structures as direct delivery risks to AI capacity forecasts.
Notable Concepts & Terms
- Memory wall: The divergence between improving compute throughput and slower memory-bandwidth improvement; Somers uses it to explain why inference increasingly stalls on moving and reading data.
- KV cache: Stored transformer key/value state for prior tokens that prevents repeated attention computation; it is the central serving-cost and session-state issue in the interview.
- Sequence/context length: The ordered token history a model can process; larger windows enable richer agent reasoning but make cache capacity, latency, and recall increasingly difficult.
- Quantization: Reducing numeric precision of weights, activations, or KV cache—such as FP16 toward FP4—to save memory and compute at potential quality loss.
- Multi-head latent attention (MLA): A DeepSeek-associated approach cited as reducing KV-cache size by trading additional compute for lower memory use.
- Gated delta nets: An attention alternative Somers identifies as promising for reducing time spent on attention, illustrating algorithmic routes around hardware memory constraints.
- Tiered memory hierarchy: Keeping active cache in accelerator memory, less-active state in host RAM, and cold state in flash/disk; essential for scalable persistent-agent serving.
- Cost per useful result: A possible future pricing metric for capable AI workers, contrasted with token pricing; useful conceptually but difficult to define and enforce.
Operator Notes / Why Ken Should Care
- Add cache-hit rate, cache-retrieval latency, cache-memory cost, eviction rate, and context reload frequency to agent-platform observability; report these by workflow and model provider.
- Run an internal long-context evaluation on representative codebases, SOPs, and operational threads that measures recall, source fidelity, tool-use success, and completion quality at increasing context sizes.
- Implement a local-to-cloud escalation policy for agents: local model handles ingestion, redaction, classification, and trigger decisions; frontier model receives only high-value tasks plus curated context.
- Set explicit session-memory lifecycle rules: classify state as hot, warm, cold, or delete; define retention and encryption requirements before deploying persistent-agent memory broadly.
- When comparing inference vendors or hardware theses, require workload-level evidence on tokens per joule, memory capacity, cache behavior, concurrency, p95 latency, and quality under quantization—not claims based solely on FLOPS or nominal token price.
- Treat data-center and AI-capacity forecasts as exposed to interconnection, permitting, financing, and local-opposition risk; do not use this interview's confident construction assumptions as a planning baseline.
Source/Metadata
- Title: How Many Will Actually Get Built & Is Energy AI's BIGGEST Bottleneck? | Positron AI Co-founder
- Transcript words: 23716
- Duration seconds: 4794
- Timestamp note: No timestamps or chapters were present in the supplied transcript. The transcript contains substantial duplicated passages.
Transcript
The scariest thing to me on the political spectrum is that it's now become a most unifying issue on left and right about being anti-data centers. And I think that is almost entirely a Chinese psyop. A single In-N-Out uses more water than the largest data centers in the United States. We have a true expert of the space on the show today, Thomas Somers. He's the co-founder and chairman of Positron AI. They just raised an $875 million Series C at a $5 billion valuation. Thomas did not hold back in this episode. Incredible discussion about what we can expect in the next few months from the biggest players in this space. I would say I'm overall opposed to Pace at the Frontier direction it's going in. Ready to go? Thomas, I am so excited for this dude. I said to you just before this that I think there are some big questions that the world doesn't know or things that they know that I think we're going to correct today. So thank you so much for joining me. [SPEAKER_00] Yeah, great to be here. Thanks, Harry. [SPEAKER_01] Not at all, but I'd love to just start. Can we just have a brief description of what Positron is and where does it sit in the stack? [SPEAKER_00] Yeah. So Positron is a fabulous semiconductor startup that's building hardware. So really everything from the chips, the software directly running on top of that, all the way up to the full systems and rack scale deployments to power generative AI inference. So effectively everything in the hardware and the low level direct talking to hardware software stack that powers all of the applications everyone in the world is excited about right now. So everything from the likes of ChatGPT, Claude, etc. [SPEAKER_01] How does the infrastructure stack required for inference, what you're working on, change compared to training? [SPEAKER_00] So training, I would say from the underlying compute level, fundamentally, is a compute bound problem. So it's a workload where the more flops that you have. And if you look at be it from the regulatory and some of the export control frameworks are heavily focused on just flops required. So the how many floating point operations per second can be done. And more or less, the amazing thing that the scaling laws of the past decade have shown is that the more parameters you add to a network, and the more flops you dedicate to that during training, the better that model is going to become. A big difference with inference. And so the deployment of those models is the fact that for the actual math and the steps that you're doing is about half of what you're doing during training in terms of the steps. It shouldn't be thought of as the actual compute involved. But what it turns out to be is that forward pass, that inference portion of it, is heavily memory bound due to the fact that basically for every single token that's generated, every little bit of output that requires going through the weights, the parameters you could, from a biological sense, think of the neurons, you have to read the values of that for every single individual token. And so the way that's, I think, very interesting about this, and I don't know if it says anything about the value or actually the saying that inference is somehow more important because of course you have to train, but fundamentally, when you're training, you already have the corpus, you already have all the training data. And so with all that data, you can massively parallelize the token inputs, all of the sequences of words and sentences, paragraphs, etc., that are going into that. So that's something you can just crush through with a bunch of compute. But when you're inferring, because that's actually generative, you don't know what the token is five words down the line. And so you have to generate each and every one auto-regressively or in order without foresight. And so that becomes a hugely memory bound problem that can't just be massively paralyzed the way training can. So I totally get that in terms of the shift from compute bound to memory bound. Is that what people mean when they say about the memory wall with regards to what you're doing? Partially. I mean, the memory wall as a phrase has been around for a long time before the hype around AI. And really what it's come down to is if you look at the past 50, 60 years of computing, we've been able to have Moore's law giving us more transistors per square millimeter of silicon consistently. And while that has been able to result in greater raw compute, flops, etc. The improvement of the memory technology is not kept up at the same rate. So roughly speaking, between 2014, just very early innings of the new AI era till 2024, you had about a 120x improvement in the flops of GPUs. So a single NVIDIA GPU had about 120 fold improvement in flops. And that's what enabled a whole lot of the improvements over that decade. The improvement of memory bandwidth is only 17x. So I would say the real embodiment of this is we had massive improvements on a per device basis of the flops, and then a whole bunch of elements on the periphery of improving the connectivity, etc. But just the ratio of the compute to memory bandwidth had this divergence. And so you had cases where if a problem was memory bound and you couldn't just scale the compute linearly with that, you were getting more and more memory bound as the decade progressed. Why was there such a misalignment in the progression between the two? One's 100x, one's 17x. Why is that the case? Well, it comes down to a lot of technical implementation details. Like the fact that if you look at the lowest level, the type of memory that is used on the silicon itself is called SRAM, static RAM. And SRAM is made out of six transistors with a bit line and word line, some other control logic around it. But that SRAM cell has not scaled in terms of the sizing of that with Moore's law over the past about 15 years. So they have grown or shrunk, I should say, very much slower than just a group of transistors that you'll use for other purposes. And I would say that there's just been a lot more architectural advancements that could happen on the compute side, while an SRAM more or less has not changed in 30 or 40 years from an architectural primitive perspective. And so on the input side of raw technical capabilities on fabrication, etc., have not been able to improve. But I would also say that there wasn't the right motivations for most of that decade. So with convolutional neural networks, things that powered AlexNet, which would really launch the deep learning revolution in 2012, and then ResNet and all of the advancements during the 2010s was in the realm of machine learning models that were fundamentally compute bound. You could just throw more and more flops at CNNs and get better results. And you didn't really need all that much, be it memory capacity or memory bandwidth. But it was really with the transformer. And even though the attention is all you need paper came out in 2017, I would say it did not really get the attention, pun intended, it deserved until 2020. GPT-1 and GPT-2 came out prior to that in 2018, 2019. But it was really GPT-3 showing that okay, you go from a billion-ish parameter up to 175 billion parameters, and you actually get this massive improvement in capability. And that's really where I would say the transformer revolution started. And most people didn't really catch on to that until the end of 2022, when ChatGPT came out. When you look at token economics and token efficiency today, what does no one know or talk about that you think should be much more front and center? I do find it, compared to a year or two ago, there are now different prices listed for cached versus uncached tokens. But I don't think people realize how any providers that charge the same amount, and even with a lot of people's cached prices, how high margin that is. It's insane. You make all of your money on selling cached input and output tokens. So as a provider— Why is that? Sorry, just so I understand that. Oh, because when we discussed earlier, processing a cached token is essentially free. It's one one-thousandth of the cost, order of magnitude of actually having to generate, to recompute and generate that token. So there is so much you can juice out of selling those cached tokens. Efficiency today, what does no one know or talk about that you think should be much more front and center? I do find that compared to a year or two ago, there are now different prices listed for cached versus uncached tokens. But I don't think people realize how many providers charge the same amount, and even with a lot of people's cached prices, how high margin that is. It's insane. You make all of your money on selling cached input and output tokens. So as a provider— Why is that? Sorry, just so I understand that. Oh, because when we discussed earlier that processing a cached token is essentially free. It's one one thousandth of the cost, order of magnitude of actually having to generate and recompute that token. So there is so much you can juice out of selling those cached tokens. Basically all the providers charge you to cache a token. They charge a higher rate than the normal processing fee for an input token. And then they charge you a lower rate when you read from that. And it's great when you're paying that lower rate, but they're making obscene margin on that cache read. And there's a reason why Anthropic is reported to have 80 points of gross margin right now on their API business. Are you surprised by that 80 points? Not really. I'm impressed by 80 points of margin in any industry. It's difficult to get that margin. And the great thing about capitalism is those margins will compress with competition. So I'm confident and happy for that, even though those people are theoretically my customers and my margin is based on their margin, but I care more about a healthy ecosystem long-term. It's more surprising to me how many people still today think these are horribly unprofitable businesses and the whole market's going to zero. It's absurd to me that the meme of OpenAI, Anthropic, etc., are just burning cash and eventually they'll run out of cash to burn. If they stopped training, they'd be massively profitable overnight. And there are a ton of other levers they have without pacing the frontier, as Dario just had in his essay. I jokingly think that pacing the frontier discussion is a great way to reduce costs ahead of IPO. But I don't think Anthropic or anyone needs to do that. I think they're amazingly profitable businesses with their scaling rates. And it would reduce costs? Just so I understand, because they would spend less on training because they would be slowing down the speed? Right. Yes. Yeah. So I don't think that's actually the intention or anything, but yes, that's the bulk of their costs. How did you analyze the pacing the frontier? You brought it up. How did you analyze it? I have mixed feelings on the safety topic. I am a human, so I would like to live to old age, and more so than that, have humanity continue to the stars and beyond. But I believe much more in the ability for this technology to revolutionize every part of humanity in a positive way. And I do worry that pacing, in a lot of the ways that's being talked about—not necessarily how it will be implemented—has two big risks. One, a major pause, stopping and sort of that playing into the doomerism sentiment that exists. And by pushing for pause, it's actually giving ammunition and a better basis for those that actually just want to stop the technology completely. And I see that as a major risk for humanity. The second piece is I'm very worried about technology and capability being concentrated to relatively few people. And a lot of what's being discussed from a regulatory framework and limitations on technology, etc., I think is the modern road to serfdom. It's the concentration of technological capability. And making it illegal to do matrix multiplications is the thing that will set us back not just to the industrial revolution, but pre-enlightenment capabilities. That is the biggest attack on classical liberal freedom concepts that I can think of. Because while I do think that the vast majority of tokens are going to be produced by the big players, if the technology itself is restricted to just those, then they're going to be the new Lords and Kings and everyone else is back to serfs. With the greatest of respects, is it not just lip service? We'll stick meter in the corner. They can do that compliance. And then we can IPO. Sam can have a reason not to IPO because his numbers aren't as good as Anthropic's. It plays into both our desires. And Elon wants time to catch up as well. And so it plays into everyone. Yeah, I completely agree. And I think that is the biggest internal reason for probably everyone other than Dario. I like Dario, and I would say the vast majority of people in Anthropic are true believers, both in the promise and capabilities of technology and the risks. And if I were in their shoes, I would also take the massive amount of responsibility. Yeah, there's a lot of strategic reasons of saying, okay, by having these auditors, etc., that removes some potential responsibility and culpability from legal perspectives. The risk that I just don't think any of them realize is this is a problem with a lot of very smart people, especially when they've amassed large wealth and power. They think they're going to be able to keep that. And the scariest thing is the centralization of technology. If it gets concentrated with companies and governments, etc., you've got people that think they're the smartest people in the room, not realizing they are not going to be the ones to actually control it when they put these measures in place. Dario basically, on one hand, verbally begging for governments to take over Anthropic. He's in part saying that because he doesn't think it will actually happen. And I would love to see his reaction if and when that actually happens, and he realizes, "Oh shit, I thought if I was begging for regulatory regulations, they would then make me the regulator in some form or fashion." And when that doesn't happen and it just becomes a bureaucracy that halts all progress and the capabilities that currently exist basically get squandered to select bureaucrats. That's the worst outcome I can imagine. I'm very naive. I'm a podcaster and a venture capitalist for a living. So one of the lowest IQs on the spectrum. When you consider the advancements that China are making, especially with their open ecosystem—but they are an incredibly talented ecosystem right now, moving forward. If we pace and they don't, what happens then? I guess when I said that the worst possible outcome, I wasn't counting that as—I wasn't counting the Terminator outcome and I wasn't counting that. So when I think of course everyone can agree Terminator or similar is very bad, but I think it's extremely low probability. And I'm not a believer in that doom scenario. For the vast majority of people, that would result in the same level of serfdom that I worry about with the scenario I described would be significantly worse for some number of people in a Chinese CCP-controlled, super intelligent AI scenario. I think that on one hand their strategic angle right now is to have technology proliferate through open source, etc. I think as soon as they get into pole position, the ladder gets pulled up with them in some way. I don't think they actually want the technology to be easily accessible to everyone. Now, I don't know— So I think of course, everyone can agree Terminator or similar is very bad, but I think is extremely low probability. And I'm not a believer in that doom scenario. For the vast majority of people, it would result in the same level of serfdom that I worry about with the scenario that I described would be significantly worse for some number of people in a Chinese CCP controlled super intelligent AI scenario. I think that on one hand their strategic angle right now is to have technology proliferate through open source, et cetera. I think as soon as they get into pole position, the ladder gets pulled up with them in some way. I don't think they actually want the technology to be easily accessible to everyone. Now, I don't know if they will decide that it's okay if the rest of the world has some access to the technology, but they definitely will not let the billion people that are not CCP party members benefit equally from the technology. So just so I understand, do you agree with it? Because to me, I just didn't get it. Oh, you can't pace the frontier unless the global AI community paces the frontier and I don't see Putin signing up. I agree. And yeah, I think this is a little bit the same naivety that I described by these company leaders and in general people in the Western world thinking, oh, we're so great. We're so advanced so far ahead that we can't get caught up to. I mean, on paper is the US the greatest military force in the world? Yes. But if we had to all of a sudden have a drone incursion, the same level of what's happening from Ukraine and Russia, you know, coming up from Mexico, and if you take Mexico, I'd just say that they developed very naive drone technology, et cetera, at the level of what's happening in Russia, Ukraine, and Iran, how would we respond to that as a country? If we had that coming up across our border? Like, it doesn't matter. Our amazing military might, we've built our military to fight the last war. And I think geopolitically our thinking is, oh, we're the big dog still. And that when it comes to AI technology, there's not the acceptance that export controls and all of the other elements that theoretically would allow us to pace and have people keep pieces behind us just aren't good long-term solutions. Do you think we should have export controls? I am very much a strong believer in free trade and free exchange of ideas. The exception to that is I think China has been a free rider of all of the benefits of a liberal free war, free trade order for the rest of the world. Well, they get to keep everything closed off. So I am very happy and think that any government societies, people that want to embrace free exchange of ideas and trade and everything else, we should have a very vibrant economy and ecosystem. But totalitarian regimes should not be able to participate with that, especially in the case where they get all of the benefits of that and get to export themselves things that make them better able to have that totalitarian system keep up. Can I ask, we mentioned the pacing the frontier and the different people who supported it, you had Zuck and Jensen. Jensen say nothing. Well, Zuck actually come out in opposition to it, saying that we should continue as planned. What should we take from those two seemingly silent and opposing it? Based on what I my overall beliefs right now, as probably evidenced by the conversation so far, I would say I'm overall opposed to the pace of the frontier direction it's going in. I appreciate anyone that is adding to the discussion. That is being realist about the benefits and risks, but you always have to take that with a grain of salt of what are the motives of anyone that's discussing it. And I would say I probably appreciate Zuck or Dario's comments infinitely more than a random politician. And not just random, the quote unquote leading politicians that don't actually understand the technology. And I would, the scariest thing to me on the political spectrum and the way all of this being treated is that it's now become an almost unifying issue on left and right about being anti data centers. And I think that is entirely, almost entirely a Chinese psyop. Can I ask why is it a Chinese psyop being anti data centers? Because it does increase, I'm totally with you on the benefits of them. And you know, Gavin Baker said that, how they're the greatest economic kind of needle mover for large parts of the country. I guess they see increased electricity prices, increased water prices and ugly data centers in their backyard. Why, why is it a Chinese psyop? What am I not seeing? Well, just on the ugly and all that, I totally am supportive of a beautification campaign and really turn them into centerpieces of our society. I think thousands of years from now, future history looks and they should see these massive data centers, the really, really massive impressive ones, should have be like the great pyramids or like the wonders. We need to dress them up to be the world wonders that they are. Technologically the thing from water usage and the amount of power they consume, et cetera, so much of the early information that went out by unsophisticated writers or things that are patently false, like a single data center uses more water than the largest data centers in the United States. And golf courses are orders of magnitude more. If these are closed loop liquid cool systems that you don't even want to use water in a lot of these cases. From the Chinese psyop perspective, they're not, to your point, they're not pacing the frontier. They're adding gigawatts of new electricity generation capacity. Most of it being dirty. They're building massive new data centers, horribly displacing people. Like it irks me so much that we have the freedom in the Western world to criticize companies, governments, everything, and slow down and based on false information. And I love the freedom elements of that, but it is a strategic disadvantage when China can just say, yeah, we're going to bulldoze all these people's homes and do rolling blackouts wherever in order to serve the greater good of new training capacity. I mean, this with the greatest of respects, but I don't understand how anyone thinks the US or Europe can beat China when they have no regulatory or policy restrictions. And I mean, the UK, you can't put up a paper airplane without getting a permit. So we're fucked, but you are getting there and you're becoming a European state in terms of the regulation and policy requirements. Am I wrong? Am I being overly negative? No, I, you're right. I think the greatest advantage the US has in that regard is that there's still a lot of land, a lot of places that do not have all the same levels of restrictions. I don't agree with a lot of things of most administrations of my lifetime. The current administration gets attacked for saying they're destroying our environments and destroying national parks, et cetera. The vast majority, 90 plus percent, I don't know the exact numbers of national federal land is just open empty desert in the West. That is not part of a national park or anything. And the fact that there are so many restrictions to utilizing BLM land for building data centers where it's literally hundreds of miles from a person, from any populated area, my great state of Nevada has still a lot of land, a lot of places that do not have all the same levels of restrictions. I don't agree with a lot of things of most administrations of my lifetime. The current administration gets attacked for saying they're destroying our environments and destroying national parks, et cetera. The vast, vast majority, 90 plus percent—I don't know the exact numbers—of national federal land is just open, empty desert in the West. That is not part of a national park or anything. And the fact that there are so many restrictions to utilizing BLM land for building data centers where it's literally hundreds of miles from any populated area, my great state of Nevada has plentiful geothermal solar, all these green energy technologies, and we could build nuclear and other things in the middle of the desert where it won't impact anyone. And that there's restrictions to that is completely absurd to me. And I will say there has been some political will and push to solve these things, but literally just in the past year, you have Republican governors and other politicians that at least had part of their platform to be pro growth and all of these things backing away because they see from their own political base being anti data center based on completely false premises. And one of the points I want to go back to that you brought up was that people would have higher electricity costs. If we increase generation capacity, this is the most basic supply and demand. If we increase generation capacity and no one's saying we want to be taking energy from what's reserved for people's homes. One of the regulatory problems I see is power companies have to have this buffer of energy availability that is baked into the cost and capabilities for everyone. There is absolutely zero cases where a data center could be potentially pulling power from anything that's already been allocated. So that's just an impossibility. And all these data centers that are getting built right now are coming with generation capacity that covers their own use and beyond that. And we're just not allowing them to hook up to the grid, where they could actually be lowering the prices for everyone. And then you've got people on the power company side that they're lobbying against new generation capacity because that will actually market forces—more capacity will decrease prices, which would be good for consumers. What percentage of data centers that are planned will be completed? From the major providers, I would say that the capacity that they have planned may be in different locations. You've had some local communities that have successfully stopped facilities going in there, but then those data centers just move. I don't think a year ago, the major data center builders and operators were thinking that the political problems were as bad as they were. And so there is a lot more effort being put into education in those communities now, which I think will turn the tide a bit, but it's also just going to mean that those data centers move to locales that aren't going to have those problems as well. And like I said, we've got large tracts of land that can support it. So I'm not too worried that it's going to be an existential threat to capacity build out. And then of course there's space if Elon's successful. Do you believe that space is a viable alternative truly, or is it conference talk and lip service to justify a market cap? I think something can start as one thing and turn into something else. I would never ever bet against Elon. I primarily bet for Elon. If you asked me a year ago, I just would not have thought that there would be a good reason for it in the near term because it's going to be cheaper, easier, et cetera, to build on land. I also think there's great alternative technologies. A company we're partnered with and I'm good friends with the CEO, is a company called Panthalossa that's building ocean based data centers, basically a very interesting pumped hydro solution in the middle of the ocean. So there are alternatives that don't require going into space. I think long-term, part of the reason I'm a long-term big believer in space data centers is I just think we're going to need to have a space economy for humanity to live up to its long-term potential. Love that. Totally agree. I'll never bet against Elon. If we think about the cost of intelligence being tied to the cost of energy, how should we think about energy as a bottleneck moving forwards? To what extent is it? We mentioned policy and regulation being a core bottleneck. Is energy a bottleneck moving forwards or less than people consider? I think there's two pieces to it. One, Positron is trying to deliver more compute, more capabilities per watt per megawatt. And so on our base case, if we can turn what you would have spent 500 megawatts with Nvidia equipment and do that in a hundred megawatts, I don't think that's actually going to mean that you're only going to build a hundred megawatt facility. You're still going to build the maximum amount of compute that you can. You're just getting more tokens, more intelligence per joule. If I go back to the long-term thinking, assuming humanity continues for hundreds, thousands of years, everything turns into an energy problem where we're gaining, and you can go back thousands of years and just look at the progression of mankind. Fundamentally, that is a perfect track of our ability to produce and use energy, from discovery of fire up to nuclear power plants. The simple answer to your question is everything— all progress is gated by energy. And even if there's energy available, it may not be economical. And so it won't be done. So I actually would say the bigger limiter than just saying energy is our ability to build and produce energy. We've got plenty of technologies and capability to do it. I would say we have way more economic limitations. It's how much debt is the world willing to take on to build out everything over the next couple of years, ties into energy, ties into the infrastructure itself, et cetera. So I think economics is a much easier scapegoat to pick. People are already very concerned by the levels of debt being taken out in the debt cycle. Do you think that concerns are justified and then do you share them? I think we've got a major sovereign debt problem that masks a huge amount of second and third order elements in the financial system, just the inflationary consequences of the government that can print infinite amounts of its own currency. And the fact that we are—as we're already seeing in treasuries and the greater bond markets—there is greater and greater perceived risk of the most quote unquote risk-free asset. I think will trickle down to all elements of the financial system. And so when people worry about Oracle's debt and credit rating, I'm like, I believe in Oracle's business model and ability to execute and do everything a whole lot more than United States government. It's just the United States government can issue its own currency and also has guns and nukes to take tax revenue. So my biggest economic concern there is that there will be a more acute specific crisis that arises out of the compounding of national debt leading to devaluation of the currency that has all of the consequences downstream. Rather than, I'm really not worried about any of the companies in the AI debt stream, like not hitting their revenue targets. The past three, four years have shown we're accelerating every aspect of these businesses in terms of revenue, profits, and how they are improving the productivity and value downstream. I'm jumping around, but when I was doing the research, I was reading about KV caching and compression as part of this, and I was honestly getting lost, but I was intrigued and digging deeper. guns and nukes to take tax revenue. So my biggest economic concern there is that there will be a more acute, specific crisis that arises out of the compounding of national debt leading to devaluation of the currency that has all of the consequences down the stream. Rather than, I'm really not worried about any of the companies in the AI debt stream not hitting their revenue targets. I'd like if the past three, four years have shown we're accelerating every aspect of these businesses in terms of revenue profits and how they are improving the productivity and value down downstream. I'm jumping around, but fuck it. When I was doing the research, I was reading about KV caching and compression as part of this, and I was honestly getting lost, but I was intrigued and digging deeper and deeper. And I was wondering, why did I not know this before? And so I don't think many will know this. What should we know about KV caching? Why is it important? Can you explain it to me a little bit? There's always a give and take relationship with innovation. I guess one element I'll have to explain to make all this clear is the concept of a sequence. I kind of already talked about a token. But just to define things, a token effectively as part of the training process, when any of these big model apps are developing a new model, they have a vocabulary that they define. So they take their big giant corpus and they do some statistical work to figure out what is the best encoding method to take all of the text in this and break it into chunks that get reused frequently to have things be more efficient. And so what you end up doing is if you took an English dictionary, you'll find that there are common prefixes and suffixes and groupings of words. And if you just try to think of how would I best compress this, if I just had symbols for these prefixes, suffixes, et cetera, compress this into a thing. And what ends up happening is a token. If you're using ChatGPT and you see text streaming out, if that's going particularly slow or you have a keen eye, you'll see that it's portions of words that come out at times, sometimes a full word, sometimes a small fraction of word. And each of those little flashes that you see is a token. And roughly speaking, it's between half and a token is equivalent to a half to about 75% of a word on average in large English corpuses. So that's token. A sequence or the thing that builds up to being context in a model is the grouping of all those tokens in order. And what happens when you're running an inference, you give it a prompt, you have, what is the capital of France as your input prompt, that is tokenized, that is four or five, six tokens, and that go into the model. And when you do inference, it's going to say the capital of France is Paris and the city of lights, something after that. And so when you have that entire sequence, when transformers originally came out for every token that was generated, you were doing the computation for generating all of those tokens, including the ones that you've already processed and the ones that you've already generated in this turn. And the clever, I would say, kind of obvious based on all of the developments in the past of computer history, but wasn't done initially was that, well, you don't actually have to redo the compute of the things that you've already had as inputs and what you've already generated in this turn. And so the KV cache was born where within the model, there's these two matrices called K and V—keys and values. And those matrices are fully based on the prompt and whatever is generated during a turn. And so by actually storing those two matrices, you can avoid having to do redo computation at the cost of now having to store this thing in memory. And that's the simple example—very small, kilobytes of hundreds of kilobytes of data. But the thing is that these things grow with the sequence length. And the interesting thing is that for the attention mechanism, the compute per token grows quadratically with the sequence length. So you're having to spend more and more compute quadratically. So that's an exponential curve as sequence length grows. But when you store that as just K and V, that's just a linear growth. And so you're really trading off what KV caching does—it means that you don't have to do that compute, which gets very expensive very quickly, at the expense of needing to store these things. And storing that is a complexity in itself because that's a unique KV cache for every single user that you're serving. And it comes questions of how long do you want to keep that, and how do you manage all of that in a large system? So what does it mean then when we hear about compression and uncompression of KV caching and potential entropy within the system? There's two different forms of compression. I'll have a couple more than that, but the two main ones. One is quantization. So the short form of quantization is if you've got each of your values, be it your weights, your KV caches, activations, stored in a particular data type. So before the machine learning revolution, most of the world computation was done in FP32. So you have 32 bits to represent a floating point number and that's broken into mantissa and exponents. Well, it was pretty quickly realized that having 32 bits of precision was overkill for the things that you're wanting to represent. And it cost more from a storage and computational perspective than lower precision. So we went to FP16, Google developed BF16, a little rejiggering of those bits went to FP8. Now we're at FP4 in popular systems. And so we've been reducing the precision quite quickly. But that does lead to some brain loss when these models run just because you are now trying to encode the same information into fewer bits. And so there's been a lot of interesting schemes to say, okay, I'm going to take this group of 16 FP16 values, BF16 values, and I'm going to quantize those. So I'm going to use truncate and rounding down to, let's say, four values. So now you actually saved 75% of your total size of that group of values. You shrunk that down from 16 bits to four bits. But just doing that naively will mean that on a lot of benchmark scores, you'll have them go 20, 30% worse. So you get that 75% savings in space, but you kind of lobotomize the model. But advanced quantization techniques actually say, okay, these 16 values, I'm able to have a shared multiple, a bias or an amount and a multiplier for it. So let's say for those 16 now into four or FP4 values, you store one new FP16 value that gets applied to all of those at compute time. So you get a 75% compression on all those values at the cost of now adding one new FP16. And basically the state of the art here is you're able to get things compressed from FP16, 16 bits per value. get 20, 30% worse. So you get that 75% savings in space, but you know, you kind of lobotomize the model. Now, advanced quantization techniques actually say, okay, these 16 values, I'm able to have a shared bias or an amount and a multiplier for it. So let's say for those 16, now into four or FP four values, you store one new FP 16 value that gets applied to all of those at compute time. So you get a 75% compression on all those values at the cost of now adding one new FP 16. And basically the state of the art here is you're able to get things compressed from FP 16, 16 bits per value down to like four and a half bits per value. And that can be applied to weights, the actual parameters and model that could be applied to the KV caches. But there is no such thing as free lunch. You do still have some lobotomy, but thankfully it's kept within like 1% of unquantized model. [SPEAKER_00] Is KV caching the hardest element of building that inference infrastructure, or is it latency SLOs or load spikes or anything else that we could come up with? Is that the hardest? Like what do we not see that we should see? [SPEAKER_01] You can run a service and do something without having KV caching at all. You're going to economics and performance and everything else wise is going to be much worse. The dark arts and magic with it is the workloads that the industry so far has found the most valuable happen to be very, very highly cachable. So Semi analysis has their Agent X benchmark and suite of test data based on taking a whole lot of Claude code sessions and having dozens to hundreds of turns and those Claude sessions with sub agents and everything else. And what they found is over these massive number of interactions of these real traced code generation agent coding sessions, about 96% of all the tokens that go through these entire sessions are cached. If your workload is going to have this extremely high caching rate where you're going to be reusing the same tokens again and again, that drastically shifts the importance of how you can retrieve those caches because they get to be very, very large. We have gone into trillions of parameters. So if we take the GPT four, which got leaked as 1.8 trillion parameter model, assuming that is in four quantized and rounding down a little bit, that's 900 gigabytes of data size for the model weights. If we take like the high expectations of Claude Fable, that's a 10 trillion parameter model. So around five terabytes of model weights. But the crazy thing is at these long context links for these size models, you have the individual user sessions being in the hundred gigabyte range. So with just 50 users on your service, the user context from just those individual sessions end up being greater than the model weights that you're trying to store. So Claude and OpenAI have a whole lot more than 50 users. And so it becomes a really interesting trade-off of okay, how much of the accelerator memory do you want to dedicate to weights, which you need to process every single token generated. And you want that to be as fast as possible because that sets your SLO, that sets the token latency. But if you don't have their KV caches persistent, you're actually losing a huge amount of efficiency because that was work that you didn't have to actually repeat. And so it saves you as an operator money more than anything. So at some level, having users' KV caches be persistent will give some level of speed improvements that the user perceives, but it's mostly an economics thing for the service provider. If you can return to them and use those tokens again and again, that saves you money as an operator massively. [SPEAKER_00] Totally get that. It saves us money because we don't have to use as much compute, but then it's harder from a memory challenge perspective. How do you think about the right logical next step then? If you appreciate the importance of saving on compute, but the challenge of memory with KV caching, what's the answer then? Do we just have bigger and bigger memory stacks on chip? What does that look like? [SPEAKER_01] The most common deployed solution, and the vast majority of inferences out there are taking place on GPUs, followed up by TPUs and a couple of devices, but the most common paradigm today is you've got your GPU accelerator memory that is primarily responsible for holding the weights and you will keep some number of user sessions on that, the ones that you're actively processing. But the larger group of users has a tiered hierarchy. So you'll have users that were around in the past couple of seconds, but haven't returned and don't have an active request, that's residing in host memory. And that's on the order of anywhere four to 10X more memory on the host than in the accelerators. So you'll be able to store more of those there. And then if someone hasn't been around in a couple of minutes, maybe a couple hours, that's going to be stored in even further away memory. So that could be in NVMe, so flash storage, so a lot slower, but a lot larger capacity on that host. It could be in flash storage on a network attached drive. And eventually, the chat GPT sessions that I had six months ago are somewhere residing on a disc, slow SSD or something, somewhere in a data center. But it would be dumb for them to use expensive memory to store that. So that tiering is the norm, but that introduces a huge amount of complexity of how do you decide when and where you're going to store something for your massive number of users. I would say our solution, kind of how we're trying to go about it, knowing both from our expectation that model sizes are going to drastically increase, the number of users for all these things are going to drastically increase, and the context themselves, two, three years ago the typical context links were on the order of 8,000 to like 64,000 tokens. Then it got up to 128, 256, you know, a million token context links are the norm now in terms of what the model supports. But a million token context links can only hold a portion of some of our internal companies' largest code repositories. It will be a fraction of that. And so if you really want an agent that can take over the capabilities of a whole team of programmers, I think the main limiter today isn't like the model capabilities itself and scaling the model size. It's on how much context can that model have of all the data it needs to make smart decisions. [SPEAKER_00] I just want to break some of the things you set up there. You said that you think model sizes will increase. I thought we were all moving to owning our own intelligence, every enterprise having their own smaller model with proprietary data. Does that go against what you think in terms of model sizes increasing? Can you help me understand? [SPEAKER_01] Yeah, I think you can kind of break it into two tiers, again. So there's going to be the frontier models and capabilities that are being really at the forefront of the development by OpenAI, Anthropic, maybe Google, SpaceX AI, et cetera. And I still think there's a long road to go in terms of getting to pushing the frontier of model capabilities. And those will continue to grow, continue to get better. And there are a lot of workloads where, let's say internal, just speaking for how Positron uses LLMs, I— increase. I thought we were all moving to owning our own intelligence, every enterprise having their own smaller model with proprietary data. Does that go against what you think in terms of model sizes increasing? Can you help me understand? Yeah, I think you can break it into two tiers. So there's going to be the frontier models and capabilities that are being really at the forefront of the development by OpenAI, Anthropic, maybe Google, SpaceX AI, et cetera. And I still think there's a long road to go in terms of pushing the frontier of model capabilities. And those will continue to grow, continue to get better. And there are a lot of workloads where, let's say internal, just speaking for how Positron uses LLMs, I don't today care that much about the cost. If I can get 10 times the output value out of a model today, I very gladly pay 10 times more per token. And I really want the frontier to push that. I think some of it is cost saving. Some of it is just owning, truly owning your proprietary data. There is a push from enterprises to have inference on site and it's a lot more difficult to provide that for largest models. And most companies, if they're adapting from open source or developing their own model, don't have the resources to be pushing the frontier. And so that is what's going smaller models. And do you not buy that reality? I would say right now, somewhere around 80, 85% of all tokens consumed and produced are done by just the top four model companies. And I would say the next five or 10% is done by the three or four after them. So I can totally buy and believe that five percent of all tokens consumed will be done by things on prem, not locked into the big guys, but I mean, from Positron's business perspective and just how I see the world evolving, I'm going to care more about the high volume set of things. That being said, I actually think the small model stuff is actually much more interesting for everything happening on your phone. And the amazing thing about that is I think there's this misconception that if questions can be answered by your phone or any prompts can be done locally, that actually means there's less tokens going to be used with the big guys in the cloud. I think it's the opposite. The reality is that if I have an LLM running on my laptop or phone or in my enterprise secure on-prem cloud, whatever, that is going to be consuming data at such a fast rate of everything coming into it. And it's going to be generating analysis based on that. And it will decide, okay, what is it that I'm going to actually return to the user locally? And what is it that actually requires more intelligence from a better model that isn't self-hosted. So I think for any of the tokens that are being saved by running locally, that's actually going to generate more things. Because in some ways for a simple naive use case of a personal user of LLMs, they're only going to prompt ChatGPT or Claude ever so often. They're limited by their thoughts of when to actually ask an LLM something, but if they have a local LLM that is constantly checking their email, their calendar, messages, et cetera, and deciding to do lookups to cloud hosted models frequently, that's now on a per person basis, a massive increase in the number of tokens being consumed and generated by the cloud models. Even though there's the naive view that there's a shift to this on device LLM. So can I just understand? So I completely hear you in terms of maybe 5% of them will be in this smaller model enterprise owned kind of model landscape. Why are you so bullish then on much larger models and the size is increasing? I'll break it into a portion. So there's the increase of the model sizes, which I think is something that would be shared by a lot of people in the AI space as a gut feel. Based on the fact that we have seen these scaling laws. We call them scaling laws. The fact that going from a hundred million to a billion to 10 billion, a hundred billion, 1 trillion per annum roles. We've seen this amazing increase in capabilities with that. We see that also. Still to this day, going from a trillion to five to 10 trillion at the largest end right now. And there's no way, like we call it a law because we've observed it, but there's no actual mathematical proof that this will continue. So it's on vibes that, okay, this has continued to scale. There's no sign of it slowing down. So is that going to continue to 50 trillion, a hundred trillion and beyond. And I don't see any indication that that's going to stop. So I'll be bullish. What does scaling laws look like at three times what it is now? If AGI has been declared now by Jensen, forgive me, but what is three times this? It's a good question. I mean, there is still, GPT-6 Astra is my first 24 hours with it were basically as magical as when my first experience with ChatGPT with GPT-3.5 in November of 2022. I was at the ChatGPT launch at NeurIPS in 2022. And it was funny because Sam and Ilya were there and it was a party in New Orleans for NeurIPS conference. And basically at the end, they just said, Hey, we launched this little fun experiment called ChatGPT, go check it out. And zero fanfare, it was just a side mention. And I don't think anyone really gave it a thought at the event, but when I went back to the hotel, I loaded it up and I got back at 10, 11 PM or whatever. And I was up for four or five hours straight, just giving random prompts and being amazed that this was the most magical experience that I've ever had with a computer. And I would say I got very close when Sora 2 came out, that I had a similar experience in a short amount of time, but just mind blown by the quality of the videos and especially the weekend that Sora 2 launched when there were no restrictions on what you could generate. Yeah, GPT-6 Astra, I do think is AGI. And to your question of what does that mean going forward? I think my guess is as good as basically anyone's. But why was that? Why was GPT-Astra so good for you? Why was it comparably such a breakthrough? Because I haven't, it's great, but honestly kind of the same as before. In terms of the things that I've found LLMs to fail the most in the past. I'll give a case where it is more linear improvement. So just in terms of general coding capabilities, performance, and analyzing problems, et cetera. It is a step function improvement, but not mind-bogglingly. There are a bunch of things that other models have not been able to fix or went in circles and found inelegant solutions. And with Astra, they're just initially giving it a couple of really hard problems that I've not been able to solve with other LLMs, it was able to do it one shot, having it go through code. In terms of the things that I've found LLMs to fail the most at in the past. I'll give a case where it is more linear improvement. So just in terms of general coding capabilities, performance, and analyzing problems, et cetera. It is a step function improvement, but not mind bogglingly. So there are a bunch of things that other models have not been able to fix or went in circles and found inelegant solutions. And it's still like a human software architect that really understands the problem is able to come up with a better solution with Astra. They're just initially giving it a couple of really hard problems that I've not been able to solve with other LLMs was able to do it one shot, having it go through code base and find both performance improvements, bugs, et cetera, and just solve them without discovering new spaces that I didn't know existed in our portions of our code base. So that's one element step function, but not mind boggling. The second case that was mind boggling just from a whole, wow, is the computer usability is with a lot of set of generic tools. So being able to do blender animations like this, it's become there's a bunch of memes online of it recreating different videos, et cetera, but just the fidelity of that and where that was basically impossible with GPT 5.6, was massive increased capability. And I had it designed, do interior design of my house just based on a couple pictures and just wow, I did not think that at what's fundamentally a text model could do that. And then finally, the biggest thing for Positron was I've been trying with every single new model release to have these models be able to actually take a relatively simple logic design problem, implementing an encryption block in this case, and being able to take that through the full RTL to GDS flow. So from basically the specification of do this encryption function, implement the Verilog, the hardware description language for that, so write that code and then be able to take that code and go through all the way until you have got a chip design that theoretically you could go to tape out. LMs could do different portions of that and could write the scripts and fail a lot of different mid points on the way, but a big problem with the electronic design automation tools, the EDA tools for doing chip design is that they were designed in the nineties, early two thousands. They're really unintuitive. None of the documentation exists out like in the public web. And so that training these models don't have a real good innate view of them, but GPT-6 with both combination of computer use and just an ungodly amazing scripting ability has been able to take this encryption block and implement it with the TSMC three PDKs and take that all the way to GDS and do that in a little over 50 something hours. And meet timing and over a gigahertz, et cetera. And that as a task, if I was giving to someone similarly new to getting the flow mostly working, I would say would take on the order of a week and getting it optimized to the points that Astra is at with that design would maybe be one or two additional weeks depending on the person. So compressing that two to three weeks down to two days and change when the model, it's still mind boggling. It shouldn't be this good at this, just as I would naively think about its training sets, but obviously with opening eyes on chip developments, and in-house, they've, I'm glad that those capabilities are getting added to the models they're releasing to the public and not just being kept inside. We see Jalapeno, terrible name I think personally, but their own chip development, Anthropic are developing their own chips, Deep Seek's supposedly developing their own chips. We see the commoditization of the chip player with everyone building their own chips. How should we think about that? As a consumer of all these things, if I take my Positron hat off, I would say that that's a great thing for the industry having fundamentally that's going to bring costs down and capabilities up and bring it to more people. I think it's such an interesting world where when I got started in the semiconductor space, 13 years ago, silicon was a dirty word in Silicon Valley. And now you have all the biggest companies in the world being somehow connected to the semiconductor industry and most interesting, exciting applications and the companies building them, vertically integrating down to the silicon layer. The interesting thing with all of the ones that you mentioned and the broader set is the companies have the same macro goals. The implementation details are all unique though. And that's, as an engineer and technologist is exciting to me that there are a lot of different ways to skin a cat and people can have their own architectural view and go about implementing it and get different results. And Hot Chips, the biggest conference for this design space, was where opening unveiled Jalapeno last month. And it's still really good. I'm happy that the industry is still pretty open and willing to share, not as much details as people would have shared five or six years ago, but still a good amount in the open discussion of things. And so I think how that applies to Positron is we have our particular architectural views and way that we've decided to do things and that will evolve in the future as well, everyone else's, but there's still plenty of space to make bets and go in different directions. And the great thing about the market is that the market gets to decide what is valuable and those that create value will receive a reward for that. We spoke about context window length earlier and the expansion of it. How much does that expand? Is there infinite expansion capability of context window length? And what does that mean we can do that we can't do today? I'm just fascinated. I think with traditional linear or quadratic attention, there is going to be limits of scale with what the hardware could provide. Now, one of the big things for Positron is we're trying to massively increase the memory capacity per device. So with our upcoming generation, we're going to have eight times more memory capacity than the highest memory skew from Nvidia and Nvidia is actually decreasing the amount of memory per device based on the market memory conditions. I frankly think that context length going from a million tokens that it is today going to 10 or much more than that is really hard with that quadratic expansion of memory cost. The algorithmic advancements that have happened over the past year have been very interesting in terms of being able to further reduce the amount of storage and compute necessary for that context with linear and sparse attention mechanisms. And those have really been innovated by the Chinese model labs. And it's a great example of when you have constraints of export controls on the chips with the highest memory capacity and flops. And so they innovated on not needing that. And Deep Seek beginning of 2025 with Deep Seek V3 had made a lot of waves because they were able to get massive decrease in KV cache size with multi-head latent attention. So you were actually spending more flops to be able to have a smaller KV cache. And that's advanced a lot over the past year and a half. And by the most interesting, or my personal favorite right now is. And this is a great example of when you have constraints of export controls on chips with the highest memory capacity and flops. And so they innovated on not needing that. Deep Seek at the beginning of 2025 with Deep Seek V3 made a lot of waves because they were able to get massive decrease in KV cache size with multi-head latent attention. You were actually spending more flops to be able to have a smaller KV cache. And that's advanced a lot over the past year and a half. And my personal favorite right now is gated delta nets and its derivative versions where you can have a 75% decrease in total time you're spending on the attention portion with this mechanism. So how important then is new hardware if Deep Seek without it, just on architecture innovation alone can cut costs by 80%? I would say that there's no such thing as free lunch. So when they have that MLA compression, it does come at cost of model capabilities in some form. There's a reason why the Chinese labs have really heavily embraced MLA while none of the U.S. labs have. I should say based on rumors, but I also have very good information and belief that none of the major U.S. companies are using MLA and not using some of its brethren. I think that will evolve and change in the future, but the short version of that is there's not just a pure savings on that side. But the reality is to go back to your previous question: everyone does want greater context length. If I had 10 million token context length, I think that would be enough for holding multiple of our largest code bases and really have that cross-pollination happen between them for an agentic coding model. It's not just having the full context. There were some early models that had advertised a million token context length, but as soon as you went above 64,000 tokens—two years ago—its recall ability just went garbage. So just saying that something has this maximum context length is one thing. It's, can it actually use that context length effectively? That's entirely different. And that's actually going back to the Astra thing, which is amazing. There are a couple of different benchmarks measuring long context performance. One of them's called Ruler and there's the find-a-needle-in-a-haystack. So you just flood the context window with a bunch of junk—basically passages from books and all this—and you put somewhere randomly in the middle of all of that a hash, some value that looks out of place. And you prompt the model: "What's the secret value?" A lot of models have done really poorly on this. GPT-5.6, which is only six or seven weeks old, could only do this about 70% of the time. GPT-6 Astra does it over 95% of the time correctly. So there's a lot of room for improvement in these things. Speaking of room for improvement, can I ask you: when I was doing the research for the show, I saw that the Silicon data token price index dropped below $1 per million tokens this month. And five years ago it was $60 per million—$60 to one. What does a million tokens cost in 2028, say two years from now? Do you think? I think the thing I care a lot more about than just that 60 to one is the fact that a $60 token five years ago, no one would pay a cent for today. That was complete garbage relative to today. And the level of quality of a token that you pay a dollar per million tokens for now is so much astronomically more valuable. And that's because of token efficiency and model capabilities. If you say okay, so in 2026, the best model in the world in 2021 was GPT-3. It's kind of crazy at the rate models get released today that GPT-3 was the best in the world basically from August 2020 when it released all the way up to Chat GPT in November 2022. So it was two and a half years between model releases, and really GPT-3.5 was just doing reinforcement learning with human feedback on the same base model. If you remember how bad GPT-3.5 was—the value of that in terms of economic productivity of GPT-3.5 versus GPT-6 today, or whatever comparison points you want—the value per token in terms of what it can improve a person's life, a company's business practices, et cetera, is orders of magnitude. I would say a hundred or thousand fold. So I think there are actually two points to your access of going from $60 to $1. Yes, that's a decrease in cost, but that token today is—to say conservatively—a hundred times more valuable. So I would say that there needs to be some multiplier there as well, where the value per unit of intelligence is probably closer to a thousand fold, not just the 60 fold you're talking about. What does that mean then? If we extrapolate that out to 2028, what does that mean—like does the cost of a token then actually matter? Is that the primary unit that we should measure? Everyone talks about cost per token. Is there actually a different metric that we should measure? It is interesting that with the GPT-6 launch, Greg Brockman had said that they don't think that cost per token—they're not going to be pricing things in tokens much longer and they want to be moving to cost per useful result. I don't think that's where it'll end up because that's really difficult to price and it's a quality question. But I think that cost per token is really great because you can easily calculate the cost, like the cost to generate a token. So determining a margin on that and pricing it in bulk volume to generic customers is really easy. And I think that's going to stick around in large form because that is so easy. We'll see for the largest providers of tokens how they potentially evolve their business models—if you have a GPT-7 or 8 that is superhuman and can fully function as an employee in an amazing capacity and Open AI calculates, through whatever method, running at full tilt, it's only going to cost them however many hundreds of thousands of dollars to produce tokens continuously with that, they may decide that it's actually easier and they'll get more adoption if they just charged a million dollars a year—just using a random number—to have full unlimited usage of that virtual agent worker. So that may be how things evolve. Can I ask you: I'm always very careful of being the naive one—I'm not that young anymore, but being the naive one who's not seen cycles—but Gavin Baker says it well when he says, "I can't speak to a company that doesn't have numbers that are parabolically up and to the right and just everything is better than it's ever been." What would be the first signs of a crack in the chasm? Like, a shift from frontier models? Many hundreds of thousands of dollars to produce tokens continuously with that, they may decide that it's actually easier and they'll be able to get more adoption if they just charged a million dollars a year. Using a random number to have full unlimited usage of that virtual agent worker. So that may be how things evolve. Can I ask you, I'm always very careful of being the young naive one. I'm not that young anymore, but being the naive one who's not seen cycles, but Gavin Baker says it well when he says, I can't speak to a company that don't have numbers that are parabolically up and to the right and just everything is better than it's ever been. What would be the first signs of a crack in the chasm? It's a shift from frontier models to open weight models. Anthropic and Open AI are not continuing in the same level, not quite growth rate because it's impossible at the stage, but level of growth missing numbers next year. And then the bubble getting burst a little bit, whether the two core leaders are having some form of strife. The reason I agree that that's a possibility. The reason I don't think it's likely is I think that the development of open source models and things happening locally will actually drive greater use of greater token volumes for the big guys. I really do think that that's, I thought they were competitive. I thought you, yeah. Yeah, no, I think the smarter and more capable that Siri is on my phone is going to result in it doing a whole bunch of background tasks and things that remove me from having to be the one that instigates it going, having requests and data be processed by even smarter models. I really do think that in a lot of AI applications right now, the bottleneck is actually a human making some sort of decision and different tasks have different levels of autonomy that will result in things getting sent to be processed by a model, by Open AI or Anthropic. But I think the next really big order of magnitude, couple orders of magnitude increase in token volumes is going to come when us humans trust a local LLM because that has access to all of our data all the time to have it decide to do things on its own that it is not smart enough to do. It's right now I trust Astra a lot more than myself on a whole lot of different things, but I still prompt it to do things. And maybe it will go run autonomously for twelve hours or three days. I think the next big leap is when I trust a model running on my laptop or on my phone to prompt the smarter models to do even more wider set of tasks. When we look at the data economy, the powers, the larger models, which you believe in, we see Mercur, we see Surge hitting three billion in revenue. How big do these companies become? Because Anthropic and OpenAI are four to five trillion dollar businesses, say feasible. It's wholly feasible, isn't it? That Mercur and Surge are two hundred billion dollar businesses, which serve both Frontier Labs and some of the world's biggest enterprises building their own models. I think my only skepticism there is on there being vertical integration by the Frontier Labs. I think the reason that hasn't happened is because Anthropic and OpenAI have better uses of their capital and mental power than doing all of that Scale AI Mercur work. But I think the main, that won't always exist. And let's say that GPT seven or eight could be an effective replacement for Sam Altman in terms of being able to manage a large business, why they wouldn't just have agents be taking over those tasks. I like to finish on a tone of optimism. What are you most excited about today that you think the world does not spend enough time on that we should spend time on? I'm very sympathetic to the problems that I think the smarter set of the AI alignment, AI safety community are when it comes to thinking about how do you align incentives? A lot of people just talk about AI alignment being a problem. I think there's a key part of that is human alignment. It's how do we as a society, the human race align ourselves to have a good outcome that I think will be empowered by artificial intelligence. And so we discussed a bunch of the different problems that we're facing geopolitically and socially and how these different things are handled. And I think a lot of smart people are doing good work on the AI alignment problem and thinking through how we solve that, but they may be gated in what can be done there. If we don't get better human alignment on regulatory frameworks and energy production, where we're going to put the data centers, et cetera. And I think framing it as this being a similar sort of technical problem that smart people can work and reason through will hopefully get more people thinking about it that way. And I think a core element that gets discounted by a lot of people in that sphere, there are a lot of economic factors. And I think it's the economic factors that actually will drive real decision-making and actions. And if we don't look at it from a rational, self-interested actors and all these different things, you're just not going to make progress. Thomas, this has been the most varied discussion ever from education on unbelievable infrastructure evolution to Dario and Sam, you are a star. Thank you so much for joining me today. Thank you so much, Harry. Thank you. if we pace and they don't, what happens then? I, I guess when I said that the worst possible outcome, I wasn't counting that as, uh, I wasn't counting the Terminator outcome and I wasn't counting that. So, um, I think of course, everyone can agree Terminator or similar or similar is, is very bad, but I think is extremely low probability. And, and just, I I'm, I'm not a believer in that, that sort of doom scenario for the vast majority of people would result in the same level of serfdom that I worry about with the scenario that I, I described would be significantly worse for some number of people in a, you know, Chinese, uh, CCP controlled, you know, super intelligent AI scenario. Um, I think that, you know, on one hand their strategic angle right now is per have, technology proliferate through open source, et cetera. I think as soon as they get into, you know, pole position, the, the ladder gets pulled up with them in some way. I don't think they actually want the technology to be, you know, easily accessible to everyone. Now, I don't know if they will decide that it's okay if the rest of the world has some access to the technology, but they definitely will not let the billion people, uh, that are not CCP party members, uh, you know, benefit equally from, from the technology. So, so just so I understand, do you agree with it? Because to me, I just didn't get it. Oh, you can't, you can't pace the frontier unless the global AI community paces the frontier and I don't see Putin signing up. I agree. Um, and yeah, I, I think this is a little bit, the same naivety that, that I described by these company leaders and, and in general people in the Western world thinking, oh, we're so great. We're so advanced so far ahead that we can't get caught up to. I mean, on paper is the U S the greatest military force in the world? Yes. Um, if we had to all of a sudden have a drone incursion, the same level of, you know, what's happening from, from, uh, uh, in, in Ukraine, Russia, you know, coming up from Mexico. And if you take Mexico, I'd just say that they developed, you know, very naive drone technology, and et cetera, like on the level of what's happening in Russia, Ukraine, and Iran, how would we respond to that as a country? If we had that coming up across our border? Like, I don't, it doesn't matter. Our amazing military might, I, we we've, we, we built our military to fight the last war. And I think geopolitically our thinking is, oh, we're the big dog still. And that, uh, um, when it comes to AI technology, there's not the acceptance that, um, export controls and all of the other elements that theoretically would allow us to pace and have people keep pieces behind us just aren't good long-term solutions. Do you think we should have that export controls? I am very much a strong believer in free trade and, and free exchange of ideas. I, the, the exception to that is, uh, you know, a little bit, I think China has been a free rider of all of the benefits of, you know, a liberal free war, free trade order for the rest of the world. Well, they get to keep everything closed off. So I am very, very happy and, and think that any government societies, people that want to embrace free exchange of ideas and trade and everything else, we should have a very vibrant, uh, economy and ecosystem. Um, but totalitarian regimes should not be able to participate with that, especially in the case where they get all of the benefits of that and get to export themselves. Uh, you know, things that make them better able to, to have that totalitarian, um, uh, you know, system keep up. Can I ask, we, we mentioned, you know, the pacing, the frontier and, you know, the different people who supported it, you had Zuck and, um, Jensen say nothing. Well, Zuck actually come out in opposition to it, saying that we should continue as planned. What should we take from those two seemingly silence and opposing it? Based on what I, my, my overall beliefs right now, as probably evidenced by the, the conversation so far, I would say I'm overall opposed to, you know, the pace of the frontier direction it's going in. I appreciate anyone that is adding to the discussion. That is, is, is I think being realist about the, uh, benefits and risks and, but you always have to take that with a grain of salt of what are the motives of, of anyone that's, that's discussing in it. And I would say I probably appreciate Zuck or, or Dario's comments infinitely more than a random politician. Um, and not just random, the quote unquote leading politicians that, that don't actually understand the technology. And I would, I, the scariest thing to me on, on the political spectrum and the way all of this being treated is that it's now become a almost unifying issue on left and right about being anti data centers. And I think that is entire, like almost entirely a Chinese psyop. Can I ask why is it a Chinese psyop being anti data centers? Cause like it does increase, I I'm totally with you in the benefits of them. And you know, Gavin Baker said that, how they're the greatest economic kind of needle mover for large parts of the country. I guess they see increased electricity prices, increased water prices and ugly data centers in their backyard. Why, why, why is it a Chinese sub? What am I not seeing? Um, well, just on the ugly and all that, I, I totally am supportive of a, uh, uh, we, we, we need to have beautification campaigns and, and really turn them into a centerpieces of our society. I think if thousands of years in that from now, um, um, you know, future history looks and they should see these massive data centers, like the, the really, really massive, impressive ones, uh, should have, uh, you know, be like the great pyramids or, or, or, like, they, we, we need to dress them up to be the, the world wonders that, that they, they are technologically the thing from water usage and, and, um, like the amount of power they consume, et cetera, uh, like so much of the early information that went out by unsophisticated writers or things that are just patently false, like the, you know, a single in and out uses, you know, more, more water than, than, uh, you know, the largest data centers, uh, in the United States. And like, it's, you know, golf, golf courses are orders of magnitude more like if these are closed loop, you know, liquid cool systems that, that, you know, you don't even want to use water in a lot of these cases from the Chinese psyop perspective is they're not to, to your, your point. They're not pacing the frontier. They're adding, you know, gigawatts of, uh, um, you know, new electricity generation capacity. Most of it being dirty. Um, uh, they're, they're building massive new data centers, horribly displacing people. Like it just, it, it irks me so much that we have the freedom in the Western world to criticize companies, governments, everything, and slow down and based on false information. And I love, I love the freedom elements of that, but it is a strategic disadvantage when China can just say, yeah, we're going to just bulldoze all these people's homes and, and do rolling blackouts, um, wherever in order to, uh, uh, serve, uh, the greater good of, of, uh, new, new training capacity. I mean, this with the greatest of respects, but I don't understand how anyone thinks the US or Europe can beat China when they have no regulatory or policy reg restrictions. And I mean, the UK, you can't, you know, put up a paper airplane without getting a permit. So like we're, we're fucked, but you are getting there and you're getting, you're becoming a European state in terms of the regulation and policy requirements. Am I wrong? Am I being overly negative? I didn't get it. No, I, I, you're right. Um, I think the greatest advantage the US has in that regard is that there's still a lot of land, a lot of places that, um, do not have all the same levels of restrictions. I, I don't agree with a lot of things of most, you know, administrations of my lifetime. The current administration gets attacked for saying like, they're destroying our, our environments and, and, you know, destroying national parks, et cetera, like the vast, vast majority, 90 plus percent. I don't know the exact numbers of national federal land is just open, empty desert in the West. That is not part of a national park or anything. And the fact that there are so many restrictions to utilizing, you know, BLM land for building data centers where it's literally hundreds of miles from, you know, a person from, from any, you know, populated area, my great state of Nevada, um, has plentiful geothermal solar, um, all these green energy technologies, and we could, you know, build, you know, nuclear and other things in the middle of the desert where it won't impact anyone. And that there's restrictions to that is completely absurd to me. And I will say there, there is the, there, there has been some political will and push to solve these things, but literally just in the past year, you have, you know, Republican governors and, and other politicians that at least had part of their platform to be pro growth and all of these things backing away because they see from their own, you know, political base being anti data center based on completely false, uh, premises. And, and one of the points I want to go back to that you brought up was like that people would have higher electricity costs. If we increase, like, this is the most basic supply and demand. If we increase generation capacity and no one's saying we want to be taking energy from what's reserved for people's homes. Like one of the regulatory problems I see is, is power companies have to have this, um, buffer of, of energy availability, uh, that, that is baked into the cost and, and capabilities for everyone. Um, there, there is absolutely zero cases where a data center could be potentially pulling power from anything that's already been allocated. So that that's just an impossibility. And all these data centers that are getting built right now are coming with generation capacity that covers their own use and beyond that. And we're just not allowing them to hook up to the grid, uh, where they could actually be lowering the prices for everyone. And then you've got people on the power company side that they're lobbying to, to against new generation capacity. Cause that will actually, you know, market forces will more, more capacity will, will decrease prices, which would be good for consumers. What percentage of data centers that are planned will be completed. Do you think? From the major providers, I would say that the capacity that they have planned, they may be in different locations. Um, I mean, you, you've had some local communities that have, uh, successfully, you know, stopped, uh, facilities going in there, but then the, those data centers just move. I don't think a year ago, the major data center builders and operators were thinking that the political problems were as bad as they, they were. Um, and so there is a lot more effort being put into education in, in those communities now, which I think will, you know, turn the tide a bit, but I mean, it's also just going to mean that those data centers move to locales that aren't going to have those problems as well. And like I said, we've got, uh, uh, large tracts of land that, uh, can, can support it. So, um, so I, I, I'm not too worried that it's going to be like an existential threat and, and capacity build out. And then of course there's space if, uh, Elon's successful. Do you believe that space is a viable alternative truly, or is it conference talk and lip service to justify a market cap? I think something can start as one thing and turn into something else. Um, I, I would never, ever bet against Elon. I, I, I, I primarily bet I bet for Elon. If you asked me a year ago, I just would not have thought that there would be a good reason for it in the near term because it's going to be cheaper, easier, et cetera, to build on land. I also think there's great alternative technology technologies company we're partnered with. And I'm good friends with, uh, the, the CEO, um, is, is a company called Panthalossa that's building ocean based data centers, basically a very interesting pumped hydro solution in the middle of the ocean. So there are alternatives that don't require going into space. I think long-term part of the, the reason I'm long-term big believer in space data centers is I just think we're going to need to have a space economy for humanity to live up to its long-term potential. Love that. Totally agree. I'll never bet against Elon. If we think about like the cost of intelligence being tied to the cost of energy, how should we think about energy as a bottleneck moving forwards? To what extent is it? We mentioned policy and regulation being a core bottleneck. Is energy a bottleneck moving forwards or less than people consider? I think there's two pieces to it. I mean, one, you know, positron is trying to deliver, you know, more compute, more capabilities per watt per megawatt. Um, and so, uh, you know, sort of on our base case, um, if, if we can turn, uh, what you would have spent 500 megawatts with Nvidia equipment and do that in a hundred megawatt, I don't think that's actually going to, um, mean that you're only going to build a hundred megawatt facility. You're still going to build the maximum amount of compute that you can. You're just getting more tokens, more intelligence per, per a jewel. If, if I go back to the long-term thinking, you know, assuming humanity continues for hundreds, thousands of years, everything turns into an energy problem where we're, we're gay and, and you can go back thousands of years and just look at the progression of mankind. Fundamentally, that is a perfect track of our ability to produce and use energy, you know, discovery of fire up to, uh, you know, nuclear power plants. The simple, um, tongue in cheek answer to your question is everything, uh, all progress is gated by, by energy. And even if there's energy available, it may not be economical. And so it won't be done. So I actually, I would say the bigger limiter than just saying that energy, like our ability to build and produce energy, we've got plenty of technologies and capability to do it. I would say we have got way more economic limitations. It's like how much, how much debt is the world willing to take on to build out everything over the next couple of years, is, ties into energy. It ties into the infrastructure itself, et cetera. So I, I think economics is a much, um, uh, easier sort of scapegoat to pick. I mean, people are already very concerned by the levels of debt being taken out in the debt cycle. Do you think that concerns are justified and then do you share them? I think we've got a major sovereign debt problem that, um, masks a huge amount of, uh, second and third, third order, uh, elements in the financial system, just the, the inflationary consequences of, you know, the government that, that can print, you know, infinite amounts of its own currency. Um, and the fact that we are, as we're already seeing the treasuries and, you know, the greater bond bond markets that there is greater and greater perceived risk of the most quote unquote risk-free asset, um, I think will, uh, trickle down to all elements of the financial system. Um, and so like when people worry about Oracle's, uh, debt and, and credit rating, um, I'm like, I, I believe in Oracle's business model and, and ability to execute and, and do everything a whole lot more than United States government. It's just the United States government can, uh, issue its own currency and, um, also has guns and nukes to, uh, take tax revenue. So my biggest economic concern there is that there will be a, more acute, um, specific, uh, crisis that, that arises out of the compounding of, of national debt leading to, uh, devaluation of the currency that has all of the, uh, consequences down the stream. Um, rather than like, I, I, I, I'm really not worried about any of the companies in the AI debt stream, like not hitting their revenue targets. I'd like if the past three, four years have shown we're accelerating every aspect of these businesses in terms of revenue profits and, uh, uh, how they are improving the productivity and value down downstream. I'm jumping around, but fuck it. When I was doing the research, I was reading about KV caching and compression as part of this, and I was honestly getting lost, but I was intrigued and digging deeper and deeper. And I was like, why did I not know this before? And so I, I don't think many will know this. What should we know about KV caching? Why is it important? Can you explain it to me a little bit? There's, um, always, uh, uh, I would say, uh, you know, give and take, uh, relationship with, with innovation. I guess one element I'll, I'll just have to explain to, to make all this, this clear is the concept of a sequence. I kind of already talked about a, a token. Um, but, uh, uh, you know, just to, to define things, you know, a token, um, effectively as part of the training process, when any of these big, you know, model apps are, are developing a new model, they have a vocabulary that they define. So they take their big giant corpus and they do some statistical work, um, to figure out what is the best encoding method to take all of the text in this and, uh, break it into chunks that get reused frequently to, to, uh, have things be more efficient. And so what you end up doing is if you, if you, if you took a, uh, English dictionary, you'll find that there are common, um, prefixes and suffixes and, and, and, you know, groupings of words. And if you just try to think of how would I best, um, uh, compress this, uh, if, if I just had symbols for these prefixes, suffixes, et cetera, um, compress this into a thing. And basically what, what ends up happening is a token. Um, you know, if you're using chat GPT and you see text streaming out, if that's going particularly slow, or you, you're quite, uh, uh, keen eye, you'll see that it's portions of words that come out at times, sometimes a full word, sometimes a small fraction of word. And each of those little flashes that you see is a token. And, uh, roughly speaking, it's between half and, and, uh, uh, a token is equivalent to a half to like 75% of a word on average and large English, uh, corpuses. So that's token, a sequence or, you know, the thing that builds up to being context in a model is the grouping of all those, those tokens in, in an order. And what happens when you're, you know, running an inference, you, you give it a prompt, you have, uh, you know, what is the capital of France as, as your input prompt, um, that is tokenized, you know, that is, you know, four or five, six tokens, uh, and that, that go into the model. And when you do inference, it's going to say the capital of France is Paris and, you know, the, the city of lights, you know, some, some, you know, uh, uh, thing after that. And so when you have that entire sequence, when transformers originally came out for every token that, that was generated, you were doing the computation for generating all of those tokens, including the ones that you've already processed and the, you know, clever, I would say, kind of obvious based on, you know, all of the developments in the past of, of computer history, but wasn't done initially was that, well, you don't actually have to redo the compute of the things that you've already, you know, had as inputs and, uh, what you've already generated in this turn. And so, um, the KV cache was born where within the model, there's these two matrices called K and V keys and values. And those matrices, um, are, are fully based on the, uh, prompt and whatever is generated during a turn. And so by actually storing those two matrices, you can avoid having to do, redo computation, um, at the cost of now having to store this thing in memory. And that's, you know, the simple example, uh, you know, very, very small, you know, kilobytes, uh, of, of hundreds of kilobytes of data. Um, but the, the thing is that these things grow with, with the sequence length. And so, um, um, but the, the interesting thing is that for the attention mechanism, the, the compute for, um, uh, per token grows quadratically with the, uh, with the sequence length. So you're having to, to spend more and more compute quadratically. So that's, you know, an exponential curve, um, uh, for, as sequence length grows. But, uh, when you store that as, as you just do K and V, that's just a linear growth. And so you're, you're really trading off the, what KV caching does is means that you don't have to do that compute, which gets very expensive, very quickly, quickly at the expense of needing to store these things. And, uh, storing that, uh, is, is a complexity in itself because that's a unique KV cache for every single user that you're serving. And it comes questions of how long do you want to keep that for how long, uh, you know, and how do you manage all of that in a large system? So what does it mean then when we hear about compression and uncompression of KV caching and potential entropy within the system? There's two different forms of compression. I'll, I'll have a couple more than that, but the two main ones. So one is quantization. So, um, one is, so the short form of quantization is if you've got, um, each of your values, be it your weights, your KV caches, um, activations, um, stored in a particular data type. So before the machine learning revolution, you know, most of the world, you know, computation was done in FP32. So you have 32 bits to represent a floating point number and that's broken into, you know, uh, mantissa and exponents. Well, it was pretty quickly realized that having 32 bits of precision was super overkill for the things that you're wanting to represent. And, you know, it's both cost more from a storage and computational perspective than lower precision. So we went to FP 16, you know, Google developed BF 16, you know, a little rejiggering of those bits went to FP8. Now we're at FP4, you know, in, in popular, um, systems. And so we've been reducing the precision quite quickly. Um, but that does lead to, you know, for lack of better words, some brain loss, um, when, when these models, uh, run just because you are now trying to encode the same information into fewer bits. And so there's been a lot of interesting schemes to say, okay, I'm going to take this group of 16, um, FP 16 values, BF 16 values, and I'm going to quantize those. So I'm going to, um, you know, use a, uh, uh, truncate and rounding that, that down to, let's say, uh, into four values. So now you, you actually saved 75% of your, uh, uh, total, uh, size of, of that group of values. You shrunk that down from 16 bits to four bits. Um, but just doing that naively will mean that on a lot of benchmark scores, you'll have them go, you know, get 20, 30% worse. So you get that 75% savings in space, but, um, uh, you know, uh, you kind of lobotomize the model. Um, now, but, you know, advanced quantization techniques actually say, okay, these 16 values, I'm able to have a, you know, shared, um, uh, multiple, a bias or an amount and a multiplier for it. So let's say for those 16, now into four or FP four values, um, you store one new FP 16 value that gets applied to all of those at compute time. So you get a 75% compression on all those values at the cost of now adding to add one new FP 16. And basically the state of the art here is you're able to get things compressed from, you know, FP 16, 16 bits per, per value down to like four and a half bits per value. And that can be applied to weights, the, the, you know, actual parameters and model that could be applied to the KV caches. Um, uh, but, uh, you know, there is, there's no such thing as free lunch. You, you do, uh, uh, still have some lobotomy, but, uh, thankfully it's kept within like 1% of, uh, unquantized model. Adam Chapnick Is, is KV caching the hardest element of building that inference infrastructure, or is it, uh, you name it, uh, latency SLOs or load spikes or anything else that we could come up with? Is that the hardest? Like what do we not see that we should see? Adam Chapnick You, you can run a service and do something without having KV caching at all. You're, you're going to economics and performance and everything else wise is going to be much worse. Um, the, the dark arts and magic with it is the workloads that the industry so far has found the most valuable happen to be very, very highly cashable. So, um, you know, Semi analysis has their, uh, Agent X, uh, benchmark, um, and, and, you know, suite of, of test data based on taking, uh, whole lot of, uh, Claude code sessions and, and having dozens to hundreds of turns and those Claude Claude sessions with sub agents and everything else. And what they found is over these massive number of interactions of these, like real traced, um, uh, code generation, you know, agent, agent, coding sessions, about 96% of all the tokens that go through these entire sessions are cached. If you know, your workload is going to have this extremely high caching rate where you're going to be reusing the same tokens again and again, that drastically shifts the, um, importance of, um, how you can retrieve those caches because it's, it's not just like, and these things get to be very, very large. Like we, we, you know, have gone into trillions of parameters. So if we just take, um, you know, the GPT four, you know, got leaked as, you know, 1.8 trillion parameter model. Now, assuming that that is in four quantized and rounding down a little bit, you know, that's 900 gigabytes of, of data size for, for the model weights. Um, if we take like the high expectations of like Claude Fable, um, you know, that's a 10 trillion parameter model. So around five terabytes of model weights, but the crazy thing is at these long context links for these size models, you have the individual user sessions being in the, uh, let's say in the hundred gigabyte range. So with just 50 users on your service, the, the user context, they're just those individual sessions end up being greater than the model weights that you're trying to store. So that's, that's, uh, uh, you know, Claude and open AI have a whole lot more than 50 users. And so it becomes a really interesting, um, trade-off of, okay, how much of the, uh, you know, accelerator memory do you want to dedicate to, uh, weights, which you need to process every single token generated. And you want that to be as fast as possible because that sets your SLO that sets the, the token latency. But if you don't have their KV caches persistent, you're actually losing a huge amount of efficiency because that was work that you didn't have to actually repeat. And so it saves you as an operator money more than anything. So at, at some level, you know, having, uh, users, KV caches be persistent will give some level of speed improvements to that the user perceives, but it's mostly an economics thing for the service provider, where if you can return to them and use those, those tokens again and again, um, that saves you money as an operator massively. Totally get that. It saves us money because we don't have to use as much compute, but then it's harder from a memory challenge perspective. How do you think about the right logical next step then? If you appreciate the importance of saving on compute, but the challenge of memory with KV caching, what's the answer then that we just have bigger and bigger memory stacks on chip? What does that look like? Most common deployed solution. And you know, the vast majority of, of inferences out there are taking place on GPUs, you know, followed up by TPUs and, you know, a couple of devices, but, um, most common paradigm today is you've got your GPU accelerator memory that is primarily responsible for holding the weights and you will keep some number of user sessions on that the ones that you're actively processing. But you know, the larger group of users has a tiered hierarchy. So you'll have users that were around say in the past, uh, you know, couple of seconds, uh, that, but haven't returned, don't have an active request that's residing in host memory. And let's say that's on the order of, you know, um, anywhere four to 10 X more memory on the host than in the accelerators. Uh, so you'll be able to store more, uh, of those there. And then if someone hasn't been around in a couple of minutes, maybe a couple hours, that's going to be stored and even further away memory. So that could be in, in NVMe. So, so flash storage, so a lot slower, but a lot larger capacity on that host, it could be in flash storage on a network attached, you know, drive. And eventually, like I bet the chat GPT sessions that I had, uh, you know, six months ago, um, you know, somewhere residing on it, on a, you know, disc, you know, slow SSD or something, uh, um, uh, you know, somewhere in a, in a data center. Um, but it would be dumb for them to use expensive memory to, to store that. So, um, that tiering is, is the norm, but that introduces a huge amount of complexity of how do you decide when and where you're going to store something, um, for your massive number of users, I would say our solution, kind of how we're trying to go about it, you know, both from our expectation that model sizes are going to drastically increase the number of users for all these things are going to drastically increase and the context themselves, like two, three years ago, you know, the, the, you know, typical context links were on the order of 8,000 to like 64,000 tokens. Um, then it got up to one 28, 256, you know, a million token context links are the norm now in terms of what the model supports. Um, but a million token context links can only hold, you know, a portion of some of like our internal companies, like largest, uh, code repositories, like it will be a fraction of that. And so if you really want agent that can, can take over, you know, the, the capabilities of a whole team of programmers, I think the main limiter today, isn't like the model capabilities itself and, and scaling the model size it's on how much context can that, that model have of all, all of the, uh, data it needs to make smart decisions. I just want to break some of the things you set up there. You said that you think model sizes will increase. I thought we were all moving to owning our own intelligence, every enterprise having their own smaller model with proprietary data. Does that go against what you think in terms of model sizes increasing? Can you help me understand? Yeah, I think you can kind of break it into two tiers, uh, again. So there's going to be the frontier models and capabilities that are being really at the forefront of the development by opening I anthropic, maybe Google, uh, you know, SpaceX AI, et cetera. And I, I still think there's a long road to go in terms of getting to, you know, pushing the frontier of, of model capabilities. And those, those will continue to grow, continue to get, get better. And there are a lot of workloads where, uh, let's say internal, just speaking for how positron uses LLMs. I, don't today care that much about the cost. I'm on, if, if I can get 10 times the, the output, you know, value out of a model today, I very gladly pay 10 times more, you know, per token. And I really want the frontier to push that. Um, I think, you know, some of it is cost saving. Some of it is just owning, you know, truly owning your proprietary data. Um, there is a push, you know, from, from, from enterprises to, to have inference on site and, and it's, it's a lot more difficult to provide that for largest models. And most companies, you know, if, if they're adapting from open source or developing their own model, don't have the resources to, you know, be pushing the frontier. And so, um, that, that is kind of what's gone smaller models. And do you not buy that reality? I would say like right now, somewhere around 80, 85% of all tokens consumed and produced are done by just the, the, uh, top four, uh, models, uh, model companies. Um, and so, um, and I would say the next, you know, five or 10% is done by the three or four after them. So I can totally buy, believe that five percent of all tokens consumed will be done by things on prem, you know, not, you know, locked into the, to the big guys, but I mean, I'm both from Positron's business perspective and just like how I see the world evolving. I I'm going to care more about the, the high volume set of things that, that being said, I actually think the, the small model stuff is actually much more interesting for everything happening on, on your phone. Um, and, and the amazing thing about that is I think like there's this misconception that, oh, if, um, questions can be answered by your phone or, or, you know, any prompts can be, can be done locally. Um, that actually is meaning that there's less tokens going to be used with the big guys in the cloud. Um, I, I think it's the opposite. The, the reality is that if I have an LLM running on my laptop or phone or in my enterprises, you know, secure on-prem cloud, whatever, um, that is going to be consuming data at such a fast rate of everything coming into it. Uh, and it's going to be generating, uh, you know, analysis based on that. And it will decide, okay, what is it that I'm going to actually return to the user, uh, you know, locally? Um, and what is it actually requires more intelligence from a better model that's isn't self-hosted. And so, uh, I think for, you know, any of the tokens that are being, you know, quote unquote saved by running locally, that's actually going to generate more things. Cause I mean, in, in some ways for, for a simple naive use case of, uh, a personal user of LLMs, they're only going to prompt chat, JBT or Claude ever so often. Like they're, they're sort of limited by, um, their thoughts of, of when to actually ask an LLM something, but if they have a local LLM that is constantly checking their email, their calendar messages, et cetera, and deciding to do these, uh, lookups to, to cloud hosted models frequently, that's now on a per person basis, a massive increase in the number of tokens being consumed and generated by the cloud models. Even though there was, you know, the, the naive view is that, oh, there's the shift to this on device, uh, LLM. So can I just understand? So I completely hear you in terms of maybe 5% of them will be in this smaller model enterprise owned kind of model landscape. Why are you so bullish then on much larger models and the size is increasing? I'll, I'll break it into a portion. So there's the increase of the model sizes, which, um, I think of the, the, you know, thing that would be shared by a lot of people in the AI space is it's kind of a gut feel, um, where, uh, based on the fact that we have seen these scaling laws, like we, we call them scaling laws. The fact that going from, uh, uh, you know, a hundred million to a billion to 10 billion, a hundred billion, 1 trillion per annum rolls. We've seen this amazing increase in capabilities with that. We see that also. Like still to this day, going from a trillion to five to 10 trillion at, that the largest end right now. And there's no way, like we call it a law because we've observed it, but there's no like actual mathematical proof that this will continue. So it's sort of on vibes that, okay, this has continued scale. There's no sign of it slowing down. So is that going to continue to 50 trillion, a hundred trillion and, and beyond. And, um, I don't see any indication that that's going to stop. So, uh, I'll be bullish. What does scaling laws look like at three times what it is now? If AGI has been declared now by Janssen, forgive me, but what is three times this? I had, it's a good question. I mean, there's, it is still, um, I, GPT six Astra is, uh, my, my first, you know, 24 hours with it were basically as magical as when, uh, my, my first experience with chat GPT with GPT 3.5 and, and, uh, November of 22, uh, I w I was at the, uh, uh, chat GPT launch at NeurIPS in, in 2022. And it was so funny because, um, uh, you know, Sam and, and Ilya were there and they, you know, it was a party in, in New Orleans for NeurIPS conference. And, um, uh, basically at the end, they just said, Hey, we launched this little fun, uh, uh, you know, experiments, uh, uh, called chat GPT, uh, go check it out. And like zero fanfare, you know, it was, it was, uh, uh, really just a side mention. And I don't think anyone really gave it a thought at the event, but when I went back to the hotel, I loaded it up and I got back at, you know, 10, 11 PM or whatever. And I was up for four or five hours straight, just giving random prompts and, and just being that this was the most magical experience that I've ever had with a computer. And I would say I got very close when Sora 2 came out, that I had similar experience short amount of time, but just mind blown by the quality of the, the videos and especially the weekend that Sora 2 launched when there being no restrictions on what you could generate. Um, uh, but, uh, yeah, GPT-6 Astra, I do think is, is AGI. And to, to your question of like, what does that mean going forward? I think my guess is as good as basically anyone's, but I, why, why, why was that? Why was GPT-Astra so good for you? Why was it comparably such a breakthrough? Cause I haven't, it's great, but honestly kind of the same as before. Oh, um, in terms of the things that I've found LLMs to fail the most out in the past. So, I'll give a case where it is more linear improvement. So just in terms of general coding capabilities, performance, and analyzing problems, et cetera. Um, it is a step function improvement, but not mind bogglingly. So there are a bunch of things that other models have not been able to fix or, um, kind of went in circles and kind of found inelegant solutions. And it's still like a human software architect that really understands the problem, um, uh, is able to come up with a better solution with Astra. They're just initially giving it a couple of really hard problems that I've not been able to solve with other LLMs was able to do it one shot, um, having it go through code base and find both performance improvements, bugs, bugs, et cetera, and just solve them without like basically discovering new spaces that I didn't know existed in our, our bunch of portions of our code base. So that's one element step function, but not mind boggling. The second case that was mind boggling just from a like whole, like, wow, is the computer usability is with a lot of set of generic tools. So like being able to do blender animations like this, like, you know, it's become there, there's a bunch of memes online of, of it recreating different videos, et cetera, but just the fidelity of that and where that was basically impossible with GPT 5.6 soul, um, was massive increased capability. And like, I had it designed, you know, do interior design of my house just based on a couple pictures and just like, wow, like that I did not think that at what's fundamentally a text model could, could do that. Um, and then finally, like the, the biggest thing for positron was, um, I've been trying with every single new model release to have these models be able to like actually take a relatively simple logic design problem, uh, you know, implementing a encryption block in this case, um, and, uh, being able to take that through the full RTL to GDS flow. So from basically the specification of do this encryption function, um, implements the Verilog. So the hardware description language for that, um, uh, so write that code and then be able to take that code and go through all the way until you have got it chip design that theoretically you could go to tape or tape out. LMS could do different portions of that and could like write the scripts and, you know, fail a lot of different mid points on the way, but a big problem with the, um, electronic design automation tools, the EDA tools for doing chip design is that they were designed in the nineties, early two thousands. They're really unintuitive. None of the documentation exists out like in the public web. Um, and so that, you know, training these models don't have like a real good innate view of them, but GPT-6 with, uh, uh, uh, both combination of computer use and just an ungodly, amazing, uh, uh, scripting ability has been able to take this, this, uh, KKK block and implement it, you know, with the TSMC and three PDKs and take that all the way to GDS and do that in like, uh, a little over like 50 something hours. And so, um, and meet timing and, you know, over a gigahertz, et cetera. And like that, that as a task, if I was giving to someone similarly new to a, uh, uh, uh, a thing like getting the flow mostly working, I would say would take on the order of a week and getting it optimized to the points that Astra is at with, with that design would maybe be one or two, two additional weeks depending on the person. So compressing that two to three weeks down to two days and change when the model, it's still mind boggling. It like, it shouldn't be this good at this, uh, just as I would naively think about, uh, uh, its training sets, but obviously with opening eyes on ship developments, um, and you know, in-house, they've, uh, I'm, I'm glad that those capabilities are getting added to the models they're releasing to the public and not just being kept inside. We see Jalapeno, terrible name, I think personally, but you know, their own ship development, Anthropica are developing their own ships, Deep Seekers supposedly developing their own ships. We see the commoditization of the chip player with everyone building their own ships. How should we think about that? As a consumer of all these things, if I take my Postron hat shirt off, um, I would say that that's a great thing for the industry having, um, uh, fundamentally that's going to bring costs down and, you know, capabilities up and, and bring it to more people. Um, I think it's such an interesting world where when I got started in the semiconductor space, you know, 13 years ago, uh, you know, silicon was a dirty word in Silicon Valley. And now you have all the biggest companies in the world being, you know, somehow connected to the semiconductor industry and most interesting, exciting applications and the companies building them, having it, you know, vertically integrating down to the silicon layer. The interesting thing with all of the, the ones that you mentioned and, and, you know, the broader set is the companies have the same macro goals. Um, the implementation details are all unique though. And that's, uh, just as an engineer and technologist is exciting to me that, that, um, there are a lot of different ways to skin a cat and, uh, people can have their own, um, architectural view and, and go about implementing it and, and get different results. And, um, you know, Hot Chips, the biggest conference for, for this design space and was where, um, opening, I unveiled jalapeno, uh, last month. And, um, uh, it's, it's still really good. Like I'm, I'm happy that, um, the industry is still pretty open and willing to share, uh, not as much details as people would have shared, you know, five, six years ago, but, uh, um, still, still a good amount of, um, in the open, uh, discussion of, of, of things. And so, um, I think how that applies to positron is, you know, we have our particular, um, architectural views and, and, uh, way that, that we've decided to do things and that will, you know, evolve in the future as well, everyone else's, but, you know, there's still plenty of space to make bets and, um, you know, go in, you know, different directions. And the great thing about the market is that the market gets to, to decide what is, is valuable and those that, uh, create, create value will receive a reward for that. We spoke about context window length earlier and the expansion of it. How much does that expand? Is there infinite expansion capability of context window length? And what does that mean we can do that we can't do today? I'm just fascinated. I think with traditional linear or, uh, uh, quadratic attention, there is going to be limits of scale. What it came to what the hardware could provide. Now, one of the big things for positron is we're trying to massively increase the memory capacity per device. So, you know, with our, our upcoming, uh, generation, we're going to have, you know, eight times more memory capacity than the highest memory skew from Nvidia and Nvidia is actually decreasing the amount of memory per device. Uh, you know, based on the market memory conditions, I frankly think that context length going from like a million tokens that it is today going to 10 or even great much more than that is really, really hard with that quadratic expansion of memory cost. The algorithmic advancements that have happened over the past year have been very, very interesting in terms of being able to further, uh, reduce the amount of, of storage and, and compute necessary for that context with linear and sparse attention mechanisms. Um, and those have really, really been innovated by the Chinese model labs. And it's, it's, this is a great example of when you have constraints of, you know, we had export controls on, on, you know, the, the chips with the highest memory capacity, um, and, and flops. And so they innovated on not needing that. And so, um, you know, deep seek beginning of, of 2025 with, uh, deep seek v3 had, um, uh, you know, made a lot of waves because they were able to, to get massive decrease in kv cash size with, um, uh, multi-head latent tension. Uh, so you were actually spending more flops to be able to have a smaller kv cash. And that's advanced a lot over the past year and a half. And by the, the most interesting, or my, my personal favorite right now is, uh, you know, gated delta nets and, and it's derivative, uh, versions where you can have like a 75% decrease in, in, um, total time, uh, you know, you're spending on, on the attention portion, um, you know, with, with this mechanism. So how important then is new hardware if deep seek without it, just on architecture innovation alone can cut costs by 80%. Yeah. I would say that that's like, there's no such thing as free launch. So when, when they have that MLA compression, it does come at cost of, um, model, um, uh, capabilities in, in some form. Um, and like the, you know, there, there's a reason why the Chinese labs have really heavily embraced, um, you know, MLA while none of the U S labs have, I should say based on rumors, but I also have, uh, on, on very good information and belief that, uh, you know, not, none of the major U S companies are, are, uh, uh, uh, they, they definitely are not using MLA and, uh, uh, not using some of its, you know, brethren. Um, I think that, that, that will evolve and change in the future, but, um, uh, it, yeah, that basically the sword version of that is, um, there's not, uh, it's not just a, you know, pure savings on that, on that side. I think that, but the reality is to, and to go back to your previous question of like, everyone does want greater context length. Like if I had 10 million token context length, I think that would be enough for, um, holding multiple of our largest code bases and really have that cross-pollination happen between them for an agentic coding model. It's not just having the full context. Like there were some early models that had that advertised a million token context length, but as soon as you went above like 64,000 tokens that, you know, it's two years ago, um, it's recall ability just went garbage. So just saying that something has this maximum context length is one thing it's, can it actually use that context length effectively is entirely different thing. And that's actually going back to the Astra thing, amazing thing about it. The, there, there's a couple of different benchmarks measuring like the long context performance. Uh, one of them's called ruler and there's like these find a needle in a haystack. So you just flood the, the context window with a bunch of junk, basically like just passages from books and all this. And you just put somewhere randomly in the middle of all of that, um, you know, uh, a, a hash, uh, you know, some, some value that looks out of place. Um, and, uh, you prompt the model say, what, what, what's the, what's the secret value. And a lot of models have done really poorly on this GPT 5.6, which is only like six or seven weeks old. Um, uh, you know, only could do this about 70% of the time. GPT 6 Astra does it like over 95% of the time correctly. So that's, there, there's a lot of room for improvement in these things. Speaking of room for improvement, can I ask you when I was doing the research for the show? I saw that the Silicon data token price index dropped below $1 per million tokens this month. And five years ago, it was $60 per million, $60 to one. What does a million tokens cost in 2028, say two years from now? Do you think? I think the much, the thing I care a lot more about than just that 60 to one is the fact that a $60 token five years ago, no one would pay a cent for today. Like that was a complete garbage token, relatively speaking five years ago. And the level of quality of, or a token that you pay a dollar per million tokens for now, um, is, uh, so much astronomically more valuable. And so, and so that's because of token efficiency and what can be done. No, I'm spaying just in model capabilities. If you, if you say, okay, so, so it's 2026. So in the best model in the world in, in 2021, uh, was GPT-3. Um, it's kind of crazy at the rate models get released today that GPT-3 was the best in the world, basically from, uh, from 2020, I think it was August, 2020, when it released all the way up to, they, they didn't have a new release until chat GPT in November of, of, uh, of 2022. So it was, you know, two, two and a half years, uh, between, between model releases, uh, and really GPT-3.5 was just doing reinforce, reinforcement learning with human feedback on the same base model. If you remember how bad, uh, GPT-3.5 was and like, what was the, the, um, value of that in terms of economic, economic productivity value of GPT-3.5 versus GPT-6 today, or, you know, pick whatever comparison points you want. Um, the value per token in terms of what it can improve a person's life, you know, a company's, you know, business practices, et cetera, et cetera, is orders of magnitude. I would say a hundred or a thousand fold. So I, I think there's actually two points to, to your, your access of going from $60 to $1. Yes. That's a decrease in cost, but that token today is what's to say, I think conservatively a hundred times more valuable. So I would actually be saying that there needs to be some multiplier there as well, where like the, the value per, you know, unit of intelligence is probably closer to a thousand fold, not, not just the 60 fold you're, you're talking about. What does that mean then? If we extrapolate that out to 20, 28, what does that mean that like, does the cost of a token then actually matter? Like, is that the primary unit that we should measure? Cause everyone talks about cost of token. Is there actually a different metric that we should measure? It is interesting that with the GPT-6 launch, Greg Brockman had said that, um, that he doesn't think that cost per token, that, that they're going to be pricing things in tokens much longer and that they, they want to be moving and, uh, to, you know, cost per useful result like, uh, to that. And I don't think that that is, I don't think that's where it'll end up because that's really difficult to price and, um, uh, you know, qualia, et cetera. But I think that the price per token is really great because you can easily calculate the cost, like the, the cost to generate a token. So determining a margin on that and pricing it in, and bulk volume to generic customers is really easy. And I think that's going to stick around in large form because that is so easy. We'll see for the largest providers of tokens, how they potentially evolve their business models in terms of, um, if you have a GPT-7 or 8 that is superhuman and can fully function as an employee in an amazing capacity and open AI calculates through whatever method that, um, running at full tilt, et cetera, it's only going to cost them, you know, however many hundreds of thousands dollars to produce tokens continuously with that, they may decide that it's actually easier and they'll be able to get more adoption if they just charged a million dollars a year. Just, you know, using a random number to have full unlimited usage of that, uh, of that virtual agent worker. So that, that may be how things evolve. Can I ask you, I'm always very careful of being like the young naive one. I'm not that young anymore, but like being the naive one who's not seen cycles, but Gavin Baker says it well when he says, I can't speak to a company that don't have numbers that are parabolically up and to the right and just everything is, is better than it's ever been. What would be the first signs of a crack in the chasm? It's like a shift from frontier models to open weight models, anthropic and open AI are not continuing in the same level, not quite growth rate because it's impossible at the stage, but level of growth missing numbers next year. And then the bubble getting burst a little bit, whether the two core leaders are having some form of strife. The reason I agree that that's a possibility. The reason I don't think it's likely is I think that the development of open source models and things happening locally, et cetera, will actually drive greater use of greater, you know, token volumes for, for the big guys. I, I really do think that that's, I thought they were competitive. I thought you, yeah. Yeah, no, I think that, you know, the smarter and more capable that Siri is on my phone, uh, is going to result on in it doing a whole bunch of background tasks and things that remove me from having to be the one that instigates, um, it going, having requests and data be processed by even smarter models. I, I really do think that in, in a lot of AI applications right now, the bottleneck is actually a human, um, making some sort of decision and different tasks have different levels of autonomy that, that will result in, in things getting sent to be processed by, by a model, you know, by open AI or anthropic. Um, but I think the, the, the next really big, you know, order of magnitude, couple orders of magnitude increase in, um, token volumes is going to come when, when us humans trust a local LLM because that, that has access to all of our data all the time to have it decide to do things on its own that it is not smart enough to do. It's just like, if it right, right now I trust Astra a lot more than myself on a whole lot of different things, but I still prompt it to do things. And maybe it will go run autonomously for 12 hours or three days. I think the next big leap is when I trust a model running on my laptop or on my phone to prompt the smarter models to do even more wider, uh, set of tasks. When we look at like the data economy, the powers, the larger models, which you believe in, we, we see Mercur, we see Surge hitting 3 billion in revenue. How big do these companies become? Because Anthropic and OpenAI are four to $5 trillion businesses, say feasible. It's wholly feasible, isn't it? That Mercur and Surge are $200 billion businesses, which serve both Frontier Labs and some of the world's biggest enterprises building their own models. I think my only skepticism there is, um, on there being vertical integration by, by the Frontier Labs. I think the, the reason that hasn't happened is because Anthropic and OpenAI have better uses of their capital and mental power, et cetera, than, uh, doing all of that, the, you know, scale AI Mercur, et cetera, work. Um, but, uh, the, uh, I, I think the, the main, that won't always exist. And let's say that GPT seven or eight could be an effective replacement for Sam Altman in terms of being able to manage, you know, a large business, why they wouldn't just have agents be taking over those tasks. I like to finish on a tone of optimism. What are you most excited about today that you think the world does not spend enough time on that we should spend time on? I'm very sympathetic to the, um, problems that I think the, the smarter set of the AI alignment, AI safety community, um, are when it, when it comes to thinking about how do you align incentives? Like, uh, a lot of people just talk about AI alignment being a problem. Um, I think there's, uh, a key part of that is human alignment. It's like, how do we, as a society, you know, human race align ourselves to have a good outcome that I think will be empowered by artificial intelligence. And so we discussed all, a bunch of the different problems that we're facing geopolitically and, and socially and, and, and how these different things are handled. And, um, I think a lot of smart people are doing good work on the AI alignment problem and, and thinking through how we solve that, but they may be gated in what can be done there. If we don't get better human alignment on, um, regulatory frameworks and energy production, where we're going to put the data centers, et cetera, et cetera. And I think framing it as this being a similar sort of technical problem that smart people can work and reason through will hopefully get more people thinking about it that way. And I think a core element that gets discounted by, I think a lot of people in, in, um, that sphere, there are a lot of economic factors. And I think it's the economic factors that, that actually will drive real decision-making and, and, and actions. And, um, if, if we don't look at it from, you know, a rational, um, you know, self-interested actors and, and all these different things, um, you're just not going to make progress. Thomas, this has been the most varied discussion ever from education on, you know, uh, unbelievable infrastructure evolution to, uh, Dario and Sam, you are a star. Thank you so much for joining me today. Thank you so much, Harry. Thank you.