Stripe

Reiner Pope of MatX on accelerating AI with transformer-optimized chips

2419 summary words 11 min summary Watch video

Start with the signal

11 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: MatX is betting that transformer-specific co-design—combining HBM and SRAM, a large partitionable systolic array, and mixed low-precision arithmetic—can deliver both superior tokens-per-dollar and low latency, overcoming the trade-off in current AI accelerators.
  • Why it matters: The discussion connects accelerator architecture directly to agent latency, long-context memory, model economics, infrastructure bottlenecks, and emerging model designs, making it unusually relevant to AI systems and investment strategy.
  • Best use: Use it as an architecture and diligence primer for evaluating accelerator vendors, inference economics, state-management choices, hardware-aware model routing, and AI infrastructure opportunities.

Executive Summary

Pope argues that AI compute should be evaluated primarily through usable application economics, not headline FLOPS. Frontier labs operate under fixed multibillion-dollar budgets, so higher throughput lowers dollars per token and permits more training, inference, or model quality. Latency still matters commercially, but existing accelerators often force a trade-off: HBM-based systems such as NVIDIA GPUs and Google TPUs provide strong throughput, while SRAM-heavy systems such as Groq and Cerebras provide lower latency at less competitive throughput.

MatX's proposed answer is a transformer-optimized chip that combines HBM and SRAM, keeps model weights in low-latency SRAM while using HBM for inference state, and pairs that memory system with a very large systolic array that can be partitioned for attention. It also supports mixed low-precision arithmetic, probably centered around 4-bit formats. The broader organizational insight is genuine hardware-model co-design: MatX employs ML researchers who train small models to validate numerical shortcuts and attention choices before those assumptions are frozen into silicon.

Pope challenges the idea that CUDA makes NVIDIA unassailable at frontier labs. Unlike gaming, where thousands of applications require a stable ecosystem, the frontier market has only a handful of major models. A lab spending roughly $10 billion on compute can rationally employ around 50 elite performance engineers and substantially rewrite its software for each hardware generation; Pope says good optimization can sometimes double performance. This creates an opening for specialized hardware, although MatX still faces severe manufacturing, customer-validation, and supply-chain risks.

For agent systems, the most important adjacent argument concerns memory. Long context is constrained by memory bandwidth because token generation repeatedly accesses prior context. Pope expects context windows to grow much more slowly than parameters or inference-time compute, making application-level compaction—such as OpenClaw writing and compressing state into files—a practical architecture rather than merely a temporary hack. He also sees opportunities to separate prefill from decode and potentially train one model representation while serving a transformed, inference-optimized version.

Key Takeaways

  • Claim: Tokens per dollar is the decisive accelerator metric, while latency is a second product-critical metric that current architectures often trade against throughput. | Evidence: Pope frames the buying decision as whether a roughly $30,000 chip produces 10,000 or 100,000 tokens per second. He says reading through HBM takes approximately 20 milliseconds, versus roughly 1 millisecond for SRAM, explaining why HBM-based systems favor throughput and SRAM-heavy systems favor latency. | Implication: Ken should judge inference platforms on end-to-end tokens per dollar at a required latency and quality level, not peak FLOPS or isolated token-speed claims. | Caveat: The comparison is an architectural simplification rather than a published benchmark across matched models, batch sizes, power envelopes, and system costs.
  • Claim: MatX is designed to break the latency-throughput trade-off through a hybrid memory system and transformer-specific compute architecture. | Evidence: The design places model weights in SRAM for low-latency access and inference data in HBM for capacity and concurrency. It adds a very large systolic array that can split into smaller pieces for attention, plus mixed low-precision arithmetic likely centered on 4-bit formats. | Implication: If validated, hybrid HBM/SRAM designs could materially improve interactive agent responsiveness without imposing the unfavorable serving economics associated with latency-only accelerators. | Caveat: These are company architecture claims for a product not yet demonstrated in the transcript with independent silicon benchmarks; Pope only expects limited user exposure around 2027.
  • Claim: Hardware-model co-design is more valuable than treating the model as fixed and merely writing kernels for it. | Evidence: MatX has a dedicated ML research team that trains small models from scratch to test numerics and attention. This allows the hardware team to evaluate nonstandard rounding modes, omit expensive corner cases, and make deliberately 'sloppy' numerical choices only when experiments show model quality remains acceptable. | Implication: For AI infrastructure diligence, Ken should favor teams that can jointly alter models, numerics, kernels, memory layout, and hardware rather than optimizing only one layer of the stack. | Caveat: Results from small-model experiments may not always transfer cleanly to frontier-scale training or unusual workloads.
  • Claim: CUDA's moat is materially weaker for frontier labs than for broad application markets because the economics justify chip-specific software rewrites. | Evidence: Pope contrasts thousands or tens of thousands of games with roughly one major model per frontier lab and perhaps five major labs. A lab with a $10 billion cluster can employ around 50 elite GPU, TPU, or Trainium optimization engineers; he says strong software work can readily double performance depending on the baseline. | Implication: A new accelerator can plausibly enter through a few technically sophisticated anchor customers even without CUDA-level ecosystem breadth, but it may struggle to move down-market. | Caveat: This argument applies most strongly to concentrated frontier buyers and is weaker for smaller enterprises, independent developers, and fragmented workloads that need mature tooling and portability.
  • Claim: Manufacturing scale and supply-chain access are as important as chip architecture for an accelerator startup. | Evidence: MatX raised a $500 million Series B led by Jane Street and Situational Awareness. Pope estimates about $100 million to produce a chip in small volumes, says a full tape-out can cost approximately $30 million, and wants MatX eventually shipping multiple gigawatts annually. Dependencies include TSMC logic dies, HBM from SK Hynix, Samsung, or Micron, packaging, racks, high-speed cables and connectors, cooling, data centers, and grid power. | Implication: Investment diligence should treat committed customers, supplier reservations, packaging capacity, power access, and manufacturing-ramp competence as core proof points rather than operational follow-ons. | Caveat: Customer contracts and financing can improve supplier confidence, but they do not eliminate yield, packaging, HBM allocation, power, interconnect, or high-volume ramp risk.
  • Claim: Persistent agent memory will rely heavily on application-level state compaction because long context is physically expensive. | Evidence: Pope explains that each generated token must access much of the prior context, making memory bandwidth a major constraint. He cites compaction near a roughly 300,000-token limit and describes OpenClaw's file-based summarization as primitive but highly controllable and rapidly iterable. | Implication: Ken should treat structured external memory, selective retrieval, provenance, and repeated compaction as first-class agent-system components rather than wait for context windows to solve state management. | Caveat: Compaction can discard details, preserve incorrect summaries, or hide provenance, so it is not equivalent to perfect native memory.
  • Claim: Transformer architecture still contains removable constraints, especially using the same model for prefill and decode and serving exactly the model that was trained. | Evidence: Pope notes that prefill processes the prompt in parallel while decode generates one step at a time, yet both usually use the same model. He also contrasts compute-intensive training with memory-bandwidth-intensive serving and suggests inference could deliberately perform more computation to exploit otherwise idle resources. | Implication: Model routing and orchestration may evolve below the model-name level, selecting distinct representations or engines for prompt ingestion, decoding, training, and serving. | Caveat: These are research directions rather than demonstrated product architectures, and separation could add routing, consistency, training, and operational complexity.

Detailed Brief

Chip development is a high-cost waterfall process with limited opportunities to recover from mistakes

  • Claims: Chip performance is largely fixed during product definition because array dimensions, operation count, clock target, memory balance, and area budget are chosen before implementation.; Architecture work begins with mental estimates and custom performance simulators; Verilog simulation is primarily used later to verify that the chosen design was implemented correctly.; Pope expects a tick-tock product cadence to be sensible: alternate physical technology upgrades with architectural revisions instead of combining every risk into one release.
  • Evidence: Pope aims to estimate performance within roughly 30% to 40% before writing the design and uses internal references such as 'go slash gates' for the costs of XOR gates, adders, and SRAM cells.; He says first tape-outs become production silicon only about 50% of the time. A full tape-out is about $30 million, while a repair limited to metal layers may cost around $100,000.; After tape-out, masks and wafers take roughly four to five months before first chips return; verification, defect testing, configuration, packaging, and HBM integration follow.
  • Caveats: AI assistance may compress the roughly 9-to-15-month logic-design and verification phase, but physical design, mask creation, fabrication, and deployment remain difficult to accelerate.; A monthly chip cadence would create heterogeneous data centers because deploying a cluster itself can take about a year.
  • Implications: The highest-leverage improvement may be faster and more accurate architecture validation before tape-out, not merely faster Verilog generation.; EDA vendors such as Synopsys and Cadence are strategically positioned to apply specialized AI to physical design, where general coding agents have less direct leverage.

Reliability must be designed as a system property rather than assumed at the chip level

  • Claims: At 100,000-chip scale, failures are continuous, so clusters require spare chips, rerouting, spare racks, and repair procedures.; Space-based data centers could theoretically compensate for the lack of repair by extreme redundancy, but cooling remains a more fundamental unresolved constraint.
  • Evidence: Pope says NVIDIA uses eight spare chips in a rack of 64 and estimates that ordinary serviceable infrastructure can absorb reliability with roughly a 10% tax.; If hardware could never be serviced and average chip life were three to five years, he estimates that deploying twice as many chips could leave approximately half operational after that period.
  • Caveats: Pope explicitly defers on the physics of rejecting heat from a spacecraft, so he does not validate the overall feasibility of space compute.; The redundancy estimates are illustrative and do not include launch cost, radiation effects, communications, or full-system failure correlations.
  • Implications: Cluster topology, fault isolation, repair latency, and graceful degradation should be part of accelerator evaluation alongside nominal performance.; Power abundance alone does not make an infrastructure location viable if heat rejection and maintenance economics remain unsolved.

Google's advantage came from research depth, early specialization, and hardware-aware parallelism

  • Claims: Google benefited from originating transformers, concentrating former Google Brain talent, and building TPUs for neural networks rather than adapting graphics hardware.; TPU v1 succeeded partly because its scope was extremely simple: one large systolic array with adjacent memory.; Mechanical sympathy—the practice of designing software around what hardware naturally does well—is central to achieving a high percentage of peak accelerator performance.
  • Evidence: Pope says TPU v1 was developed in roughly one to one-and-a-half years by a skeleton team of approximately 20 to 30 people and announced in 2016.; He contrasts CPU control overhead with GPU-style wide-vector execution using a motorcycle-versus-truck analogy: the CPU is maneuverable but carries little payload per instruction, while the GPU amortizes control over much more work.; He identifies JAX and the guide 'How to Scale Your Model' as important references for laying out LLMs efficiently across many TPUs, with a GPU version now available.
  • Caveats: Pope says modern AI chips cannot repeat TPU v1's minimalism because current market table stakes and supported workloads are much broader.; His claim that Google largely stopped publishing strong research around 2022 is an interpretation, not substantiated with publication data in the interview.
  • Implications: Google's AI position should be analyzed as an integrated research, compiler, model, chip, and deployment advantage rather than attributed only to Gemini model quality.; Specialized startups can take product-definition bets incumbents avoid, while incumbents retain advantages in workload breadth, manufacturing experience, and installed software.

Notable Concepts & Terms

  • Mechanical sympathy: Designing software and algorithms around the physical strengths and constraints of the machine, especially parallel execution, predictable control flow, memory access, and utilization.
  • HBM: High-bandwidth memory provides capacity and concurrency needed for high-throughput training and inference, but reading a large memory footprint contributes to token latency.
  • SRAM: Fast on-chip memory that can provide approximately millisecond-scale weight access, making it attractive for low-latency inference but expensive in chip area.
  • Systolic array: A regular grid of multiply-accumulate units that is highly efficient in area and power for matrix multiplication; MatX proposes making it large but partitionable for attention.
  • Mixed low-precision arithmetic: Using formats such as 4-bit values where possible while retaining higher precision for sensitive layers, increasing throughput and efficiency without unacceptable model degradation.
  • Prefill versus decode: Prefill processes an input prompt largely in parallel, while decode generates tokens sequentially; their different execution profiles may justify separate models or accelerators.
  • Tape-out: The point at which the completed chip-design files are sent for mask creation and fabrication, initiating a costly process that takes roughly four to five months to return first silicon.
  • Application-level compaction: Summarizing and externalizing accumulated agent context into a smaller persistent state, trading some fidelity for controllability, lower bandwidth use, and longer effective memory.

Operator Notes / Why Ken Should Care

  • Create an accelerator evaluation template that records matched-model throughput, p50 and p99 latency, power, batching assumptions, memory capacity, interconnect cost, utilization, and total dollars per delivered token.
  • Instrument OpenClaw compaction for information loss: retain source pointers, compare pre- and post-compaction task performance, and trigger rehydration when confidence falls.
  • Track MatX's 2027 milestones around first silicon, independent benchmarks, anchor customers, HBM and packaging commitments, software readiness, and production yield.
  • When assessing accelerator investments, require evidence of supply reservations and customer-backed volume commitments rather than relying on architecture presentations or financing alone.
  • Separate proprietary workflow data from any model-improvement partnership unless the contract guarantees isolated training, model ownership, and non-reuse of resulting capabilities.
  • Prototype prompt-ingestion and token-generation telemetry separately so future prefill/decode routing can be adopted without redesigning the entire inference control plane.
  • Monitor Synopsys, Cadence, and specialized EDA startups for tools that reduce physical-design cycle time, since code generation alone does not remove the fabrication bottleneck.

Source/Metadata

  • Title: Reiner Pope of MatX on accelerating AI with transformer-optimized chips
  • Transcript words: 21970
  • Duration seconds: 4397
  • Timestamp note: No timestamps or chapters were provided. The transcript contains duplicated passages, speaker-label inconsistencies, and substantial trailing repetition.
Full transcript 13309 words · 97 min read
0:01

SPEAKER_01

Reynard Pope is the co-founder and CEO of MatX. He's a former math whiz and Haskell programmer who became a TPU architect for Google. And now he's teamed up with Google's former chief chip architect to design a better chip for AI.

0:15

SPEAKER_01

So a year ago, everyone was saying Google is canceled. AI is going to eat their search. No one's going to search for things and therefore the business won't do well. Obviously, that sentiment has really shifted in part helped by Gemini 3 being really good. And then also it's really fast. It's powered by the custom chip hardware Google has. You were inside Google for actually, I think a lot of the foundational period, laying the groundwork for that stuff. What do people not appreciate about what Google did right to lay all the groundwork for their current AI success? [SPEAKER_00] They started with the research, right? [SPEAKER_00] The transformers came from there.

0:52

SPEAKER_00

Pretty much anyone who's maybe over 30 and at a large lab has been at Google Brain at some point. So I think there was and has been a lot of talent there. TPUs are pretty good. I mean, we think there's better you can do, of course, but they at least had the opportunity to design the TPUs for neural nets, rather than graphics applications like NVIDIA. And so the overall architecture, starting with single core, doing what was at the time reasonably large systolic arrays by today's standards, nowhere near as much. But I think those were a lot of really good decisions. [SPEAKER_01] When did the TPU project start? TPU v1 was announced in 2016, I think.

1:15

SPEAKER_00

TPU v1 actually kind of led to the creation of all of those 2016, 2017 startups.

1:20

SPEAKER_01

[SPEAKER_00] So Cerebus, Grok, Graphcore, SambaNova, all of those.

1:26

SPEAKER_00

TPU v1 actually was, I think, a really impressive project. It was done on a very short timeline, maybe about a year or so, maybe a year and a half with a skeleton team of 20, 30 people. Really minimal viable product.

1:37

SPEAKER_01

[SPEAKER_00] More recent TPUs and more recent AI chips in general can't do that because the market has moved and the table stakes are much higher.

1:41

SPEAKER_00

But the first generation product was just one big systolic array, stick a memory next to it, we're done. And it was really simple, a nice, elegant product. [SPEAKER_01] TPU v1 predates the transformer. [SPEAKER_01] Is that just a coincidence that they happened at very similar times or related in some way? Yeah, there was a period of maybe about four years of a lot of ML research or neural net research prior to transformer. So what was popular? LSTMs and Convnets and ResNet and Inception. The big thinking at the time was to adapt it to be used for LSTMs. It's a reasonable fit there. But no, I mean, I think there was just a huge flurry of activity.

2:03

SPEAKER_00

Why did it all happen then and not later is probably just because people stopped publishing. In 2022 was about the time when Google completely stopped publishing its research. And so all the good papers are from before that as a result. [SPEAKER_01] But is there some hand-wavy story you can tell about parallelization where both transformers and TPUs are about really internalizing the importance of parallelization?

2:13

SPEAKER_00

Definitely. I put it somewhat on people, actually. It is just true. Hardware is massively parallel. You've got tens of billions, hundreds of billions of transistors on your chip. And it takes maybe 100 clock cycles to get from one side of the chip to the other. And so you can't do a sequential computation involving transistors on both sides of the chip. So the hardware is just fundamentally parallel. And you have to take advantage of that. TPU v1 and all later TPUs naturally took advantage of that. Just matrix multiply is really nice because it is so parallel.

2:16

SPEAKER_00

[SPEAKER_01] So I think on the hardware side, that's generally understood. I think most ML researchers, especially of the time, were not super deep in what hardware wants and what is mechanical sympathy, sometimes a term that's used for that. So what are the terms? I mean, it kind of makes sense for us. Yeah, it speaks for itself. Think about the poor machine and what does it want?

2:29

SPEAKER_00

I mean, the term actually, I think, originates in maybe high-frequency trading in areas like that, which is areas I haven't worked in, I've just read about the software that people have built from there. And it's like, for them, what does the machine want? It wants a lot of instruction level parallelism. This is CPUs, not GPUs. It wants a lot of don't branch. So unpredictable branches kill your performance. And so think about the things that CPUs do and how to use them best. Can I get to peak performance on a CPU? It's sort of that idea. I think the whole idea of peak performance on a CPU is kind of crazy. Like, no one even says, what is peak performance? What is my percentage of peak on a CPU? Because performance of software running on CPUs is really bad. But on running on GPUs or TPUs or AI chips in general, actually, that is the main focus. It's like, what is my percentage of peak? Can I get 70% or 80%?

2:35

SPEAKER_01

Okay. I feel like many people listening to this know that GPUs perform better for AI workloads than CPUs. And it's kind of a funny history when you think about it, where one day we woke up with all these very mathematically intensive workloads, first crypto mining, and then AI. And so then NVIDIA is extremely well positioned because they've been making GPUs for gamers that you would plug into your Dell PC back in the day and maybe upgrade the graphics card by plugging in a better NVIDIA graphics card than the one the stock computer came with. And they were incredibly well positioned to capture that.

2:46

SPEAKER_01

So I think people know that. What is the intuitive explanation as to why GPUs are better for AI workloads than CPUs? Because people say they're better for these mathematical computations. But that's kind of a tautological answer. Is there some way you can have a mental model for why that is the case? Because software instruction sets also involve doing math?

2:53

SPEAKER_00

Yeah, so intuitions, I'm not sure. Let me try and go to some of the big differences, which is really wide vector instructions is sort of the hallmark of a GPU. the one the Stockdale computer came with. And they were incredibly well positioned to capture that.

3:07

SPEAKER_00

So I think people know that. What is the intuitive explanation as to why GPUs are better for AI workloads than CPUs? Because people say, yeah, they're better for these mathematical computations. But that's a tautological answer. Is there some way you can have a mental model for why that is the case? Because software instruction sets also involve doing math? So, intuitions, I'm not sure. Let me try and just go to some of the big differences, which is really wide vector instructions is the hallmark of a GPU, which I think maybe if you want some intuition, it's like how much is spent on controlling the thing? And control means like if I'm driving a truck, how much is the driver versus the payload? A truck has a huge payload in it. That's more like the GPU, whereas a motorcycle is more like the CPU where you've got the instruction, like actually just processing the instructions, reading what do I have to do next? Okay, how do I do that? That is most of the cost on a CPU, whereas if you just keep the same instructions but make the payload 100 times bigger, then you can shift most of the cost to be in the actual work that you want to do. Okay, so CPUs have been optimized for very complex instruction sets, whereas GPUs optimized for? Yeah, complex instruction sets and fine-grained changing what you want to do. So steering, like in this analogy, a CPU can steer an obstacle course, no problem. Whereas on a GPU, you're just going to go straight line for a really long time. Yes, yes. Okay, so this is getting us into what is MATX? How did you guys get started and which part of this space are you attacking? Yeah, so MATX is making the best chips physically possible for LLMs. What led us into MATX, so Mike is the other founder, Mike and I were both working at Google. And I was working on the inference stack for running LLMs. And I was saying, how can we make the best software on TPUs for running LLMs? And what we really wanted out of hardware was support much, much larger matrices. The matrices have grown from maybe 128 in dimension into the many thousands. And so much larger matrices and much lower precision arithmetic. And we tried to move the TPUs in this direction. TPUs have been moving in this direction, but they're constrained by a lot of other workloads. There was a big ads workload at the time. And so back in 2022, before ChatGPT was released, there was this idea that LLMs were going to be a big thing, but not conviction and really hard to make a big bet on that. I think a startup is more of the right place to make a big bet on a workload. You either fail, it's fine. Another startup will succeed.

3:12

SPEAKER_00

[SPEAKER_01] Yes. Whereas I think a company like Google or NVIDIA, the next chip has to work for sure. And then so... [SPEAKER_01] Ah, you can take more technical risks as it turns out. Yeah. Yeah. Yeah. Well, actually, I would say we're taking product risks rather than technical risks. [SPEAKER_01] But is there actually product risk? Because it seems like LLMs are going to work. I think now we understand it. Two years ago or three years ago, I think it was...

3:54

SPEAKER_00

[SPEAKER_01] Fair. Okay. And when you're going to say the best chips for LLMs, I can think of multiple ways to measure best. It could be best performance per watt. It could be lowest latency, capable of handling the largest models. What is best?

3:58

SPEAKER_00

[SPEAKER_01] In general, there are two metrics which LLM workloads care about, which is throughput, which is really just an economics thing. I buy a chip for $30,000, and can I do 10,000 tokens a second or 100,000 tokens per second of throughput? That determines the dollars per token. So throughput, and then latency, how fast does the thing respond? As I see the market, the economics seems to be most important. Ultimately, the quality of the AI you can train and serve is constrained by: I have only a $10 billion budget, and I want to train and serve the best model I can on that budget. And so if I can have more tokens per dollar, then I can get better quality out. So the product we aim to build is far ahead on latency, on throughput. But the surprising thing is we're competitive with the best on latency as well. And so I think that is a unique thing in offering both in the same place. And is this for AI training and running the models inference? Is this most interesting for inference, or is there a training angle? Is it useful for training, but you're trying to win inference? Is that how you think about it?

4:03

SPEAKER_00

[SPEAKER_01] I think that's a reasonable way to look at it. I think the best inference chip today will be a really good training chip as well. And so our product is both training and inference, but I think the first sales will be inference. That's mostly just a market effect where it's easier to buy. It's not as big of a risk to buy an inference cluster than a training cluster. I think the product is really compelling for training as well. And so I think it should be the best training product.

4:09

SPEAKER_01

[SPEAKER_00] Yeah. And you guys just raised a big new round of financing. [SPEAKER_00] Yeah. That's right. We've raised a series B round. It's led by Jane Street and Situational Awareness. Situational Awareness. That is Leopold Aschenbrenner's fund. He wrote the definitive book on AGI and where it's going. And Jane Street, they're real technical experts. They understand all the details really well. So very happy to have them lead the round. It's a $500 million round. It helps us actually ramp the manufacturing and supply chain for our chip so we can bring our chip to market. That's a lot of money.

4:29

SPEAKER_00

Yeah, it is. I mean, roughly I would say it costs about $100 million to produce a chip in small volumes. But if you want to, you see the orders that are going around like OpenAI, Anthropic, Google are buying multi gigawatt clusters. They cost tens of billions of dollars of chips. And you want to deploy all of that in about a year or so. And so you just need a massive supply chain behind you. And assuming everything works technically, what rate of production could you start to see? We have some estimates of where we'd like to be on this. This is ramping to very large volumes is a huge challenge for anyone. And so obviously for the large players, they've had some

4:37

SPEAKER_00

produce a chip in small volumes. But then if you want to see the orders that are going around like OpenAI, Anthropic, Google are going around buying multi gigawatt clusters. They cost tens of billions of dollars of chips. And you want to deploy all of that in a year or so. And so you just need a massive supply chain behind you. And so assuming everything works technically, what rate of production could you start to see?

4:38

SPEAKER_00

We have some estimates of where we'd like to be on this. This is ramping to very large volumes is a huge challenge for anyone. And so obviously for the large players, they've had some practice in it. Getting to a very large volume for a startup is hard. We would like to be at a place where we're shipping multiple gigawatts a year. Multiple gigawatts per year. Yeah. Speaking of metrics, you know, you talked about tokens per second. We used to measure chips in flops. And I guess there's some kind of custom flop thing for AI chips. But is everyone just using tokens per seconds these days? Is the industry aligning on that as the chip metric?

4:49

SPEAKER_00

Yeah. So I mean, I guess it's an application metric versus the chip itself. Flops of the chip is the key chip metric. There's a little bit of like, if I go and say I've got an exaflop chip to you, then the appropriate suspicion is to say, OK, but can I actually use those flops effectively? [SPEAKER_01] I see. And so then you need to map the application to that. Yeah, yeah, yeah. So this is telling you the usable flops, yeah, for your purposes. [SPEAKER_01] OK. As a consumer of AI, we have known for a long time that lower latency products succeed. Google talked about their internal testing where the differences were down to, was it 50 milliseconds? Something like that.

5:14

SPEAKER_00

Yeah, yeah. In result times where they noticed more Google engagement, the faster the results were. And you'd think that 50 milliseconds is imperceptible to a human. And it almost is, but turns out it's not. And I think Amazon has, certainly they've optimized the latency of the Amazon experience quite a lot. I don't know if they've talked about this stuff publicly, but you know that their internal metrics similarly show that the faster the product page loads, the more people buy it. And yet in AI, Google has carved out a meaningful advantage via Gemini just being really fast for its level of intelligence. And as far as I can tell, ahead of most of the other labs at a latency at a fixed, high level of intelligence.

5:18

SPEAKER_00

Yeah, yeah. Why have you guys or Grok or better chips not being adopted faster to give this product latency? It's just that this will happen and you guys will be powering all the AI products. But I note that Google has an interesting lead there.

5:25

SPEAKER_00

I think there's ultimately for at least for existing chips in the market, there's a really uncomfortable trade-off between latency and throughput. The chips that are best at throughput have historically been the chips that are based on HBM as the memory. So that is Google, Amazon, NVIDIA. In order to have very large throughput, you need a lot of inferences in flight simultaneously. So that needs the large memory, but that hasn't been so good at latency. And then there's the Grok and Cerebris that are much better at latency because they've got the SRAM, weights are in SRAM, very low latency. The problem is, and the challenge when you go to a Grok or Cerebris system is that the throughput you get there, it just is not very good. And so the fundamental dollars per token is just not competitive with Google or NVIDIA or Amazon. It is actually possible to do both in the same chip. It's an obvious thing. You say you take the HBM, you take the SRAM, put them together in the same chip. You put the weights in SRAM, and you put all of the inference data in HBM. That is what we are doing, in fact. And I think that actually hits a really nice sweet spot where you can get low latency and also be very cheap. So I think that's a really attractive point to be. It hasn't happened in the market yet just because of product decisions that have been made by the different chips.

5:28

SPEAKER_00

Got it. But we should expect it. Like, we should expect all the AIs we're using to get significantly faster over the coming three to five years. [SPEAKER_01] Or even faster, I would say. [SPEAKER_01] Yeah. Wow. So, I mean, generally, HBM-based chips tend to be about 10 milliseconds or 20 milliseconds per... I'm sorry. HBM-based chips are things like TPUs. [SPEAKER_01] That's right. That's right. [SPEAKER_01] Yeah. [SPEAKER_01] Yeah. There's just some simple math of how long does it take you to read through all of HBM? It takes about 20 milliseconds. And so that's the amount of time per token it runs. Yes.

6:33

SPEAKER_00

Whereas the amount of time to read through all of SRAM is much faster. And so you can typically get about one millisecond. [SPEAKER_01] Yes. [SPEAKER_01] So that's a lot of magnitude faster. [SPEAKER_01] Yes. [SPEAKER_01] Famously, software used to be old-fashioned deterministic software, the kind that's now out of favor, used to be very easy and quick to scale. And you would have social networks that have some Southwest moment. And they can scale through 10, 100, 1,000 orders of magnitude of a few rows in the database. And it's a very underutilized CPU. That's right. What's interesting about the AI world is there are very real bottlenecks, you know... Yeah.

7:25

SPEAKER_00

You want to spend lots of time talking about power. But it's not just bringing power online. You know, just you mentioned HBM is reminding me of, it seems like there's a view that maybe there's going to be some HBM supply chain crunch. Yes. And so where do you see, are we in for just a crunched world where some limiter is pacing the rate of AI build out over the coming few years where the economics work of the products and everything like that, but ultimately we just can't bring the components online fast enough because we have to build out the factories and things like that. And what are those crunched components?

7:47

SPEAKER_00

Yeah, no, I mean, I think so. And I'll just comment, by the way, this is a great time to be a supplier in this place. [SPEAKER_01] Yeah. [SPEAKER_01] Or just really... [SPEAKER_01] You should have started an HBM company. [SPEAKER_01] I know, right? I think it's also just a fun time to be someone who optimizes software. That's always what I like doing. Always the challenge is, why am I optimizing this if... and everything like that, but ultimately we just can't bring the components online fast enough because we have to build out the factories and things like that. And what are those crunched components?

8:34

SPEAKER_00

Yeah, no, I think so. And I'll just comment, by the way, this is a great time to be a supplier in this place. [SPEAKER_01] Yeah. [SPEAKER_01] Or just really- [SPEAKER_01] You should have started an HBM company. [SPEAKER_01] I know, right? I think it's also just a fun time to be someone who optimizes software. That's always what I like doing. Always the challenge is, why am I optimizing this if no one cares? [SPEAKER_01] Yeah. But finally, there's a place where actually you can, it's actually very meaningful in a very tangible sense. If I can make this 20% more efficient, then it can save that 20% of the build out. Yeah.

9:26

SPEAKER_00

The supply chain, we're going to have crunches on all of the supply chain, really. So if you look at the big components of what any company, but like us, for example, build out, there is dependency on logic dies from typically TSMC, maybe Samsung, or HBM, which are the big three HBM vendors, Hynix, Samsung, and Micron. And then there's also just the whole rack manufacturing, which includes literally just sheet metal and so on that builds the rack, but also cables and connectors because of all the high-speed interconnect. That's what we- [SPEAKER_01] don't sound hard. Are they sneaky hard?

9:33

SPEAKER_00

The big challenge is that you want to bring in a huge amount of power, get a huge amount of heat out and also have phenomenal interconnect, which has very high signal integrity requirements. And so pack a lot of cables in with cables that don't bend too much. They have to have enough copper in them and so on, but you don't lose data rate on the interconnect. Yes. So yeah, if you push it to a limited time. Okay. Wafers, racks, HBM, what else? Data centers, which I think is power primarily, a little bit of build out, but primarily power and grid infrastructure there.

9:40

SPEAKER_01

[SPEAKER_00] Okay. How do you then, as a startup that is looking to acquire all these components, elbow your way in amongst the giants of the Googles and the NVIDIAs and all these people who have long learning relationships and have been buying for much longer? [SPEAKER_00] Yeah. I mean, ultimately what they, what all of these suppliers care about, they do somewhat care about a diversity of their own customers. It's not a great position to be- They don't want monopsony. [SPEAKER_00] That's right. Yeah.

9:52

SPEAKER_01

[SPEAKER_00] Yeah. But then, what is their hesitation or the calculus for one of these large suppliers is, if I reserve some of my capacity for you, a startup, are you going to be around in a year? Is anyone going to even buy your product? Our approach has been to actually find buyers for the product. And then the buyers answer that question ultimately. [SPEAKER_00] Got it. And so if you show up with a bunch of fairly ironclad contracts to a supplier, then that's that. [SPEAKER_00] That's the nature of it. Yeah.

10:12

SPEAKER_01

[SPEAKER_00] I presume also the round you just raised really helps there where showing that you are incredibly well capitalized and not going anywhere also helps from a supplier point of view, a supplier validation point of view. [SPEAKER_00] Yeah, absolutely. Yeah. I mean, it helps just to say that we are around. We, in some cases, are actually, it depends on which part of the supply chain, but some parts of the supply chain, some are fungible. Logic dies are typically pretty fungible. But other parts of manufacturing are, you actually need something specifically set up for you. And so we're also able to cover the capital costs for that.

10:25

SPEAKER_01

Yeah. Yeah. That makes sense. And coming back to the MATX architecture. Okay. You want to build the best your problems. [SPEAKER_00] What is that? [SPEAKER_00] Yeah.

10:47

SPEAKER_01

[SPEAKER_00] Sounds great.

10:53

SPEAKER_01

[SPEAKER_00] Yeah. So, I mean, there's a few aspects to that. I think the first one is just, pick your memory system, right? And so I said, we've seen this HBM family, we've got the SRAM family, put them together is actually the most obvious idea, but you can actually do it. There are a lot of details to make that work well. We've done that work. One of the things that shows up there is you've spent all of this area on your chip on SRAM. How do you fit in the matrix multipliers, which are the other big thing you really need to do. And so somehow create a much more efficient matrix multiply engine. There is a gold standard for that, that is called the systolic array. Make a really large systolic array. You can't beat that in area or power efficiency. Provably so?

10:58

SPEAKER_01

Practically. Practically. Yeah. It is not known a better approach there. The main thing is, where are the inefficiencies typically? The inefficiencies show up when you leave the systolic array. So if you make a really big systolic array, then you just don't leave it as often. So that's the idea. So make a really big systolic array. That is the theme of several of the 2023 era startups, including us. But one of the challenges there is now, there is this part of the neural network as part of the transformer, which is this attention that doesn't map well onto a large systolic array.

11:03

SPEAKER_01

[SPEAKER_00] And so that's attention. The mixture of expert layer maps really well, but the attention does not. And so what we came up with, which is quite different than some of the other startups in this space, is say take a really large systolic array, but have a way to split it up into pieces, without losing efficiency. So that is the core of the design for us. And then the third component. So first was HBM and SRAM. Second is the systolic array. Third component is just an interesting new approach on low precision arithmetic. Low precision arithmetic in general, we've seen number formats get narrower. They get faster as you make them less precise. Sorry, number formats get narrower.

11:08

SPEAKER_01

What does that mean? Yeah. So float 32 was how people used to train neural nets. [SPEAKER_00] up into pieces without losing efficiency. So that is the core of the design for us. And then the third component. So first was HBM and SRAM. Second is the systolic array. Third component is just an interesting new approach on low precision arithmetic. Low precision arithmetic in general, we've seen number formats get narrower and narrower. They get faster and faster as you make them less precise. Sorry, number formats get narrower.

11:15

SPEAKER_01

What does that mean? Yeah. So float 32 was how people used to train neural nets. Yeah. And that's just too much precision. Like it's too much precision. Yeah. It's like saying I've got an image with a billion color bit depth. It's like too many colors. You'd rather have more pixels and fewer colors. And so that trend seems to go all the way, almost all the way down to like one bit even, where you just have very few colors, but a huge number of pixels. And that in net seems to be better, just more efficient way to train models. And so sorry, literally what precision are you dealing with in these designs? So we have a range. The I mean, we actually have an ML team who we hired specifically to research different forms of numerics and how to make them all work together really well. We have a range of precisions. It's not just one precision. We think probably the main thing will be similar to where NVIDIA is at, which is 4-bit precision. But I think a mix of different precisions is useful for when you look at the research, sometimes you want some layers in higher precision or lower precision and so on. Yeah. Yeah. Okay. So 4-bits is 16. Yeah. Yeah. You got 16 choices. That's it. Yeah. That's it. Yeah. Yeah. It was pretty imprecise. Yeah. Yeah. That's really interesting. I didn't know about that dynamic, but it makes sense. Yeah. And half of them are positive, half of them are negative. So yeah, it's even less precise. Yeah. Yeah. How do you design a chip? Like what's that a whiteboard? Like what software are you working in? I just love to know. I understand how you design software and what that looks like. Mm-hmm. Mm-hmm. Mm-hmm. I've actually no sense for what chip design looks like. So the way that you actually type a chip into a computer is similar to software. So you write Verilog. Mm-hmm. Verilog is a programming language. It is a very parallel programming language, which makes it different than like C or Python or something. Yeah. But it is a programming language. So the mechanics of how you express the design are the same as software. And we have continuous integration, get all of those things. But like a program executes, like your Verilog program. We don't really run it, right? Yeah, exactly. How does it run? We synthesize it. Yeah. Okay. So Synopsys and Cadence provide EDA tools. Um, they— So EDA, if you remember, yeah. Electronic design automation. I'm just a humble painter. Okay. I don't even know what it means really. I think it's electronic design automation. Okay.

11:21

SPEAKER_00

It takes the Verilog and says, first turns it into a description of what are the logic gates that are involved and source knots and then the wires between them. And then it runs for days doing some really difficult algorithms and then eventually produces, so gates are the first thing. And then even below that, it literally just produces polygons. It says like P-type semiconductor here, N-type semiconductor here, and polysilicon. Yeah. Okay. So like you write Verilog and then that compiles down into gates and ultimately like the Minecraft, you know, 3D, just this is where your elements should go. But like, then what is the iteration loop? Like when we write code at Stripe, we build a first version of something and you know, then we try it out and then you know, refine it and we add more functionality over time. We're going to write some tests at some point. We'll ship that. We'll find product market fit and then we'll refine it in market. Like, do you just sit down and write the completed chip and it works really well? Yeah. Like every year we tape out a chip and if there's a bug, we just wait till next year. It's not really how we do it. Yeah. Well, so what's the iterative? How do we actually do it? Yeah. It's much more waterfall than software is. So waterfall is almost a bad word in software development. Yeah. But it's just a fact of life in chip design. Yeah. Yeah. So the waterfall goes from architects to logic designers who are writing Verilog. And then there's this design verification and physical design. So there's this really big architecture phase, which happens before even writing any Verilog, which is what do I want the organization of my chip to be? There's in some sense, I mean, what I really like, I came to hardware after doing almost 10 years in software. I really like the blank slate you get in hardware and you've got all of the raw materials you have a much more varied in what you have available. So what is the organization of your chip? Do I have a hundred cores? Do I have one core? Do I have a systolic arrays? Do I have vector units? All of those things. And then we spend a long time coming up with that general principle and then saying, okay, now I've got these applications I want to run. I want to run a transformer of a particular shape. I want to map that onto this architecture that I've got in my head. And so we do a lot of iteration. Well, I've got this architecture in my head. I write it down to communicate to other people that's just like a markdown file. And then still actually a lot in my head, but maybe with Python simulation and so on, I'll see, do my applications map well to it. And so can I run LLM? This is where I was going to go. Okay. So you have a simulator where you write your chip, you can then simulate its performance and you have some battery of tests. Yeah. That you kind of see how this chip design works. Yeah. Is it like an industry standard?

11:26

SPEAKER_00

[SPEAKER_01] and so we do a lot of iteration. Well, I've got this architecture in my head. I write it down to [SPEAKER_01] communicate to other people that that's just like a markdown file. And then still [SPEAKER_01] actually a lot in my head, but maybe with Python simulation and so on, I'll see, do my applications map well to it. And so can I run LLM? This is where I was going to go. Okay. So you have a simulator [SPEAKER_01] where you write your chip, you can then simulate its performance and you have some battery of tests Yeah. That you kind of see how this chip design works. Yeah. Is it like an industry standard,

12:03

SPEAKER_00

you know, is it the X plane of chip testing? Yeah. So, I mean, you, [SPEAKER_01] there's an industry standard thing for the Verilog once you've done the design. Um, they're just [SPEAKER_01] Verilog simulators that you can test against. Okay. Um, that is, but you've already invested a huge amount of work by the time you've got to that point. And so I sure hope you haven't made a [SPEAKER_01] big mistake at that point. Yes. Um, so the thing that everyone does prior to that is we'll write our own performance simulator, um, which I mean, it is very specific to your particular architecture and you can write it quite concisely in just like a normal programming language.

12:35

SPEAKER_00

And so that is where most of the architecture work is done. And then the simulation on Verilog is more, I know what I'm doing. I just want to make sure I didn't have any bugs when I implemented it. Right. But I presume it's a game of inches where different people are trying different things and then you do simulate it to see if it runs 1% better across the battery of tests, or is that not how it works? In this space, not so much. Um, so, I mean, just to

13:03

SPEAKER_01

[SPEAKER_00] characterize what performance of an AI chip is, it is how many really, like, if you're just like, [SPEAKER_00] first thing you care about is flops. Um, how many flops have I got? That's a product of,

13:14

SPEAKER_00

um, how many multiplies? Like I've got a grid of a certain size, like a thousand by a thousand. [SPEAKER_01] So that's, it can do a million multiplies in a clock cycle. And then I have a certain clock [SPEAKER_01] frequency, like a gigahertz. And so I multiply them out. That is the speed of it. Um, I don't even need [SPEAKER_01] to write that and test it to see how fast it is. Um, yeah. So like what I plan in advance is it's [SPEAKER_01] going to be this fast. Um, what I can then optimize on maybe a little bit is clock speed. There's not [SPEAKER_01] a lot I can do there. Um, and then I can optimize a bit on area as well. So there is some room for

13:48

SPEAKER_01

optimization, but actually a lot of it gets set. Like the actually just the speed of the chip gets set very much up front. Got it. And then how many chips do you fab? Is it only the ones going into production or is it just build a few to throw away or how does it work? Yeah. So, um, the ideal, which companies tend to hit about 50% of the time is that your first tape out,

14:15

SPEAKER_00

[SPEAKER_01] tape out costs like $30 million. Um, your first tape out is just production. That's right. Uh, it's the actual manual, like the first chip costs $30 million. The second chip costs $1,000. Yes. Yes. Uh, so tape out is that, is that first chip. Yeah. Okay. [SPEAKER_01] The ideal is that your first tape out is actually is your production thing. So you do a tape out, [SPEAKER_01] you make maybe a thousand chips and test them and then you do production volume. Um, [SPEAKER_01] in the unlucky 50% of the time, you need to redo some more, all of your tape out. So in good

14:45

SPEAKER_00

[SPEAKER_01] cases and in many cases, you can redo just the metal layers, which costs you only like $100,000. Um, [SPEAKER_01] as opposed to the pay the $30 million again. Um, but in bad cases, like if you've made [SPEAKER_01] something serious and you can't fix that at the metal layers, you have to do the whole thing again. [SPEAKER_01] So why can't that be solved in like, is that definitionally an error in simulation where [SPEAKER_01] it turns out these two gates were close together and it just led to some reliability [SPEAKER_01] issues or. Yeah. Um, so yeah. Like what you're describing is like physical, uh,

15:26

SPEAKER_00

[SPEAKER_01] the physical implementation of the chip is wrong. Uh, that's one class. The other class is that the logical specification of the chip is wrong. Um, but shouldn't that be. Shouldn't you have got that before? Yeah. Yeah. Before you spend $30 million on the actual line? Yeah. So so, I mean, yeah, we do a lot of testing. We try not to ship these things. Uh, I hear software companies also ship bugs to production as well. And sometimes things, uh, sometimes things, there's a very good resource. Shouldn't you not be shipping bugs? Um, but I mean, [SPEAKER_01] there is a real trade off in, you can spend more and more time on design verification. Yeah. Um, uh,

16:07

SPEAKER_00

like there's always this question of when do you stop? Yeah. Like and so you stop when your coverage metrics have hit a certain point, but maybe not a hundred percent. Yeah. And then if, um, you know, Apple has to discretize the iPhone release cycle and they've settled on, you know, once per year. And so they'll decide, you know, we've got this better camera, but it's got to wait for the next version. Uh, or you know, we're going to improve the waterproofing, but you know, that's gotta wait for the iPhone eight or whatever. Uh, and so they have taken a continuous process of like all those coming up with ways to make the iPhone better and discretized it into annual iPhone

16:38

SPEAKER_00

releases. What will your discrete cadence be? Yeah. Many chip vendors have this sort of

16:42

SPEAKER_01

[SPEAKER_00] tick tock model, which is, um, you'll do on one generation, maybe you're trying to release [SPEAKER_00] every year. Um, on even numbered years, you'll do a physical technology upgrade. So new transistor technology, new memory technology, a new interconnect. Um, and then on odd numbered years, you might do a

16:46

architecture overall. I think that's a pretty good fit because it um, you have different parts of that you're of your company that are skilled at different areas and it allows you to keep of both of them occupied, um, without having instead every two years doing a massive risk [SPEAKER_00] release. Yeah. Yeah. Okay. And so you think that's probably likely for you. Yeah. That's me. Yeah. [SPEAKER_00] Um, you mentioned interconnect. So there's an art about there that NVIDIA, a huge part of the [SPEAKER_00] defensibility comes not from the chips, which are good, but from the software layer and the ability

16:57

SPEAKER_00

[SPEAKER_01] technology, new memory technology, a new interconnect. And then on odd numbered years, you might do a architecture overall. I think that's a pretty good fit because you have different parts of your company that are skilled at different areas and it allows you to keep both of them occupied without having instead every two years doing a massive risk release. Yeah. Yeah. Okay. And so you think that's probably likely for you. Yeah. That's me. Yeah.

17:08

SPEAKER_01

[SPEAKER_00] You mentioned interconnect. So there's an article about there that NVIDIA, a huge part of the [SPEAKER_00] defensibility comes not from the chips, which are good, but from the software layer and the ability [SPEAKER_00] of the technology for engineers to write these really parallel workloads, and the fact that [SPEAKER_00] they've been refining CUDA for whatever number. Yeah. A decade or something. Exactly. Yeah. [SPEAKER_00] Yeah. A long time. Just how do you think about parallelization and is that narrative true? [SPEAKER_00] Yeah, it's true for sure. It's true in some, in many areas of the market. I think,

17:23

SPEAKER_00

and especially where you look at where NVIDIA entered the market, they're doing PC devices, lots of gaming, and so on. There are thousands of games, maybe tens of thousands of games released, and they all need to be programmed against CUDA. And so, there's such a huge investment in the software that is really important to their compatibility. There are not thousands of LLMs. There's one LLM per Frontier Lab, and there's maybe five Frontier Labs or something like that. And so the economics of that is different.

17:45

SPEAKER_00

The calculation for Frontier Lab roughly goes as I just bought a $10 billion compute cluster. I have hired 50 of the best people who can write optimized GPU or TPU or Tranium software. I pay them less than $10 billion, a lot less. And so, let's put them to [SPEAKER_01] work optimizing the compute. And so they can like good work there can, depending on what your baseline is, but it can very easily double the performance of [SPEAKER_01] the software you write. And so there is a huge amount of custom software written for every generation of chip. When a new chip comes out, software is substantially rewritten to

17:57

SPEAKER_00

optimize for that specific chip. And that's just the right trade-off given the relative costs [SPEAKER_01] of these things. What that means for us is that that ecosystem already exists, [SPEAKER_01] and that way of operating where you say, I'm just going to staff a 50% team to write software for this chip, works really well for if you're trying to sell to Frontier Labs. [SPEAKER_01] Okay. So you're saying CUDA is way more important for the games environment, where it just has a lot of games than this top heavy AI market that we're in. Yeah. Where if people say you need to customize your workload for a Matex chip, it's like, well,

18:25

SPEAKER_00

fine. Yeah. It's the cost of business. Yeah. Yeah. That makes a lot of

18:29

SPEAKER_01

[SPEAKER_00] sense. Where will you fab the chips? TSMC. Okay. Yeah. Why is TSMC so durable? [SPEAKER_00] Yeah. I mean, it's interesting. They don't charge a lot as well. You'd think that if they're [SPEAKER_00] monopoly provider, they should charge a lot of money. They don't. I think that is a big [SPEAKER_00] aspect of why they're so durable. It's cyclical conservatism [SPEAKER_00] crossed with Taiwanese business conservatism means you're at the most conservative part of the [SPEAKER_00] Yeah. Yeah. The matrix. [SPEAKER_00] But I mean, it does, I mean, an American capitalist might say, well, they're just

18:42

SPEAKER_00

screwing up. They could have extracted more money from the market, but you could also say that there's actually this long-term sustaining advantage because they will just stay ahead for a really long time. They don't encourage the creation of competitors. Yeah. Yeah. [SPEAKER_01] But isn't the creation of competitors priced in because of the geopolitical risk. And so [SPEAKER_01] it's not like everyone's fat, dumb, and happy with their TSMC dependence. They're actually [SPEAKER_01] thinking a lot about it. Yeah. I mean, there is real technical advantage there as well. Yeah. It's not just the discouragement.

19:13

SPEAKER_00

But standing chip seems really hard. Building airplanes seems really hard. [SPEAKER_01] There are so many areas where competitive market forces create multiple options. [SPEAKER_01] Yeah. [SPEAKER_01] And yet that has not occurred here.

19:32

SPEAKER_01

[SPEAKER_00] So there are multiple options. You can buy from Intel or Samsung.

19:37

SPEAKER_00

But at leading edge nodes. Yeah. So what do we even care about in leading edge nodes, I guess? The big advantage is on power. The advantage on area is smaller. The leading edge nodes, the density doesn't go up as much as it used to. So when you are really, really sensitive to power, it is a good idea to be on leading edge nodes. So that is AI chips and mobile phone chips. Yeah. But there are a lot of the market where you don't like devices. Sure. Sure. Yeah. Car chips. Yeah, that's fine. But you're saying like, if you exclude the two most interesting parts of the market, you know, yeah, you can,

20:14

SPEAKER_00

just for this super high growth area of the market, it's interesting to me. Like, again, there's a lot of other really complex business problems out there that competition has solved. And yeah, chip design is like, why has someone not left TSMC and gone and built a new fab? Yeah. I mean, I don't know. The cost of a fab is extremely expensive. I mean, I recognize that the cost of a fab is extremely expensive. Um, I don't really understand the technical details of why it's so hard. I mean, there is some [SPEAKER_01] amount of just a $10 billion fab versus a hundred million tape out and chip

20:37

SPEAKER_00

[SPEAKER_01] development. There's a huge difference there, but beyond that, I'm not sure. What's TSMC like to deal with? [SPEAKER_01] So they're very big. So as a startup, we tend to work with not directly with TSMC, but with an ASIC vendor who firstly does a huge amount of the actual backend work for us,

20:51

SPEAKER_01

[SPEAKER_00] the chain of X with them. But then also has their existing relationships with them. Got it. I don't really understand the technical details of why it's so hard. There is some amount of just a $10 billion fab versus a hundred million tape out and chip development. There's a huge difference there, but beyond that, I'm not sure. What's TSMC like to deal with? So they're very big. As a startup, we tend to work with not directly with TSMC, but with an ASIC vendor who does a huge amount of the actual backend work for us, the chain of X with them. But then also has their existing relationships with them.

21:08

SPEAKER_00

Got it. TSMC cares a lot about diversity of their customer pool. And so it gets back to that conservatism. Yeah. So we're, they're great to work with from that perspective. They want to encourage startups. That's right. Yeah. Yeah. That's very cool.

21:10

SPEAKER_00

Why don't the labs design their own chips? Google does. OpenAI is starting. It's really a trade-off of how much advantage you get from vertical integration versus how much advantage do you get by concentration of R and D work. So if you take the five labs and they all buy from one player, then you can put five times as much R and D into that chip. And does that beat the advantage you get from saying, I know exactly what my model is because of the several years delay from designing a chip to being in production. You can't actually say, I know exactly what my model is because models change much faster than that. So even the labs are forced into this position where they have to make predictions and they have to hedge against what they might do two years from now. The calculus is what is the probability distribution of what my model might look like? And then design a chip that gets 90% of that probability distribution or something.

21:18

SPEAKER_00

Yeah. Yeah. Elon is excited about data centers in space. Yeah. The two criticisms I've heard are that cooling is very hard and then just repairing the chips is hard, but I know nothing about chips. You do.

21:23

SPEAKER_00

Yeah. So the repair I think is really interesting. When you look at how NVIDIA deploys their X, we do something pretty similar to what NVIDIA does. In general, you always need to design for the fact that some of your chips are going to be down. Mean time between failure of chips is not that large. And so in a cluster of 100,000 chips, there's going to be chips that are down all the time. One way you can do that is you can make a rack where one rack has some spare chips in it. NVIDIA has eight spare chips in a rack of 64. That's pretty good. The common ontarix works really well for you there. That's the thing where you can actually, because you can pick which ones to avoid, you can with very high probability tolerate a lot of failures. And then the other thing is to say my rack has to work, but I have some spare racks as well. So you can math that out with the tax of reliability here is only 10%. That's pretty good. And you can, but that relies on someone coming in and servicing the device part within a day or something like that. If you say they're going to serve us at never, then I think you actually can get where you want to be, but maybe with a hundred percent tax on reliability rather than 10%. So for example, if you think the average lifetime of a chip is in the range of three to five years, that means if I deploy twice as many chips, then three to five years from now, half of them will still work.

21:24

SPEAKER_00

Yeah. And also the burn-in is particularly failure-prone. And how about the cooling? So most of the challenge, I mean, there's actually really a data center design aspect that then, at the rack level, the challenge of cooling is just getting the heat out as quickly as possible out of the rack into the cooling network. How you get it out of the spaceship, other people would know that better than I do. Okay. Yeah. Yeah. Yeah. Again, that seems to be the main objection, but I don't know.

21:40

SPEAKER_00

Yeah. I mean, I think it's like, if you think the cost of repair is that you need to have deployed twice as many chips, then it's a trade-off of the capital of the chips versus the power saving. Exactly. The repair thing feels like can be solved because also I think part of Elon's claim is that we will just be so power limited that you have no option but to go to space and people can argue about that, but were that to be the case, then yes, it's like, well, you can get power in space and you cannot on earth. And so you might as well go there. Whereas the cooling is a more fundamental, does the product actually work at all?

21:46

SPEAKER_01

[SPEAKER_00] Yeah. Yeah. Yeah.

21:52

SPEAKER_01

[SPEAKER_00] Rayner thinks about AI the unglamorous way—compute systems architecture and what it takes to run models reliably at scale. And if you're building an AI product, the business model similarly has a ton of unglamorous complexity. You're not just selling AI, you're monetizing consumption across API calls, tokens processed, GPU hours. Stripe billing is a scalable system for usage-based billing. It lets you launch token-based pricing, subscriptions, credits, hybrid models, whatever you want. So you can create revenue models based on usage without rebuilding your pricing system every six months. If you're building an AI product, Stripe billing is worth a look.

21:53

SPEAKER_00

What are your AI predictions for 2026? [SPEAKER_01] I mean, what I'm really excited about is being able to, I mean, I'm still excited about the coding. That's what we do as a company. It's what many others do as a company as well. The one aspect of this is expanding into more domains. So for example, in where we spend our time, as a company, we write Rust, we write Verilog, we write Python. I don't know. Haskell. Yeah, no, there's a story there. I used to love Haskell. Rust is my current favorite. Okay. Mutation is good. The models are extremely good at Rust and Python.

21:57

SPEAKER_00

[SPEAKER_01] about the coding. That's what we do as a company. It's what many others do as a company as well. The one aspect of this is expanding into more domains. So for example, where we spend our time, as a company, we write Rust, we write Verilog, we write Python. I don't know. Haskell. Yeah, no, there's a story there. I used to love Haskell. Rust is my current favorite.

22:14

SPEAKER_00

[SPEAKER_01] Okay. Mutation is good. The models are extremely good at Rust and Python. They've done a lot of RL on them. They have not done as much RL on Verilog. They've done almost none on, okay, write me a markdown file that describes a chip architecture. And then how do you even RL on that? You have to say like, what is a good chip architecture? I have to say it somehow, say whether that's a good result or not. Yes. I think one of the things the labs are doing is trying to broaden what they've done RL on, RL on source it from customers and so on, in order to fill out the gaps between the spikes. I presume the labs would love to work with you on improving the models by doing RL on this specific task. However, it's also somewhat—

22:18

SPEAKER_00

It doesn't make sense for us. Yeah. You're a special source of it. So do you want to come up with some AI approaches, but keep them proprietary? Is that— Yeah. So, I mean, we've looked at a few different aspects here. There's the, I mean, what we're able to do by ourselves, our business is not training models. We do it in order to do the research on numerics, but actual production models we don't do. So it's like the biggest mileage I think is on the RL and it's not something we can really do ourselves. We'd love it if we could have a custom model just for us, but that doesn't seem to be— Well, you could, right? The terms we've been offered by labs so far have not been on those terms, but because you have to share the IP back. The way they prefer to do it is that they put it into their mainstream model. Yeah. Because it's good for them. Yeah, yeah, yeah. Which obviously you don't want to do.

22:22

SPEAKER_00

Yeah. I mean, how do you think you, what does you using AI to design a model do you think look like? Because this is actually, I think, an interesting sight glass into recursive self-improvement where we're using the AIs to develop better AIs. And so I'm curious, what do you think that looks like? Is it your own proprietary recursive models? What else? Like, is there day-to-day AI usage that's load-bearing?

22:27

SPEAKER_00

Yeah. I mean, so the stuff that is available today and I think will become even better very quickly is just the stuff that looks most like software. So writing Verilog, running tests, running continuous integration and so on. And that is a big fraction of the development time in the chip. It's probably 9, 12, 15 months or something. There's some stuff that's downstream of that, which is physical design, which is you take that Verilog and you generate the gates and the polygons. We don't have a clear path for it, at least the most obvious thing is not clear for how to compress that. Like, can you tape out a chip in one month? One month would be the goal. In theory, you could compress all of the logic design and design verification down to a short amount of time if you just continue on the same path we're doing now. But if you wanted to take the physical design down, that has to leave code. You're now doing graphical interfaces and saying, well, I want to place stuff and so on.

22:32

SPEAKER_01

Yeah. Actually there has been work on this even prior to LLMs, which is specific model trained for that particular problem. Yeah. And I think the vendors, which is like Synopsis and Cadence, probably will, should move in that direction. Most of the focus has not been do it faster. It's been do it with higher quality. But that is a big bottleneck on can I have a new chip every month? And then there's the practical thing of like, a new chip every month doesn't really make sense because then if I'm deploying, like if it takes me a year to populate a data center, that means I'm going to have different chips in different corners of the data center. Yes. Yes. Sir, when you talk about one month to tape out, so you do all this work to ultimately produce a file. Everything TSMC then does, it's not entirely in software. Like, is there some typesetting that has to happen of moving stuff around? But yeah, what happens when you send your files to TSMC? Yeah. Then what?

22:36

SPEAKER_01

So they create a mask. That is where the ASML tools come in. [SPEAKER_00] And a mask is really just a stencil. You shoot the lasers through the mask and all the x-rays through the mask. And then that produces the different p-type and n-type semiconductors.

22:46

SPEAKER_01

So they produce the mask. That is the expensive part. And then they're building up these 15 or so metal layers. So they place it on the silicon and then there are different layers of metals, which could connect all the transistors together. They do that on a wafer. It happens on a stepping basis. So there's a maximum size of chip you can build, which is constrained by this machinery. The wafer stepper is part of the ASML special sauce, right? Yeah. I guess there's probably some important alignment requirement there.

22:56

SPEAKER_01

[SPEAKER_00] Yeah. I think I remember that being quite, the classic manufacturing throughput problem. And I think they've done a lot of work on optimizing that. [SPEAKER_00] Yeah. Yeah. So they take that. So then you just produce hundreds of copies of your chip. You have to test it because there's defects. You typically, I think the average rates, really depends on process and so on, but small single digit number of defects per— The wafer stepper is part of the ASML special sauce, right?

23:14

SPEAKER_00

[SPEAKER_01] Yeah. I guess there's probably some important alignment requirement there. Yeah. I think I remember that being quite a classic manufacturing throughput problem. And I think they've done a lot of work on optimizing that.

23:25

SPEAKER_00

Yeah. Yeah. So they take that. So then you just produce hundreds of copies of your chip. You have to test it because there's defects. You typically, I think the average rates really depends on process and so on, but small single digit number of defects per chip. So you test the chip and see whether it has any defects in it. Many chips are designed to be able to tolerate a few defects. And so you need to configure it to tolerate the defects. And now you have a die that by itself works. And then you need to package it. So you put it in a package together with memories, typically that's the HBM. And then maybe you escape the wires to connect to other chips.

23:30

SPEAKER_00

Yeah. How long does it take to make a mask? So what we see is time from tape out to first chip to chips back. Again, depends on node, but it's ballpark four or five months. Oh, so tape out is just sending the file or? [SPEAKER_01] Yeah. Well, I mean, I assume tape out is we can send a tape out, send the file, and then there's a whole process of you making the masks for all the layers and then actually just producing the chips. Got it. So producing the masks and producing the chips happens after tape out. That's right. I see. Okay. So is the term tape out from sending a magnetic tape with the instructions or something?

23:49

SPEAKER_00

It could be. I was in software when the film was created. I'm curious what the tape actually means. It feels like when we're thinking about AI predictions, one thing I'm really struck by is how still in 2026, every time you open a chat window, it's contextless. Yeah. It's got no memory. And now to be fair, it's been four years, not even four years. It's been three and a half years. Just calm down. We'll get there. But I also interpret a lot of the current enthusiasm for open Claw and all that stuff as it's this super hacky backdoor into state management where your little Claw will write a markdown file of what it's doing. And then you look at that markdown file the next time and things like that. But it just feels like state management and memory is going to be a huge deal. And that will really change the character of AI products.

23:54

SPEAKER_00

Yeah. It's really interesting. The long context is one of the biggest bottlenecks on model performance on speed of the model. Yes. It just, every single token you generate, it reads through all of the previous tokens, or maybe it reads through a subset of them, but reads through a lot of the previous tokens you've written. And so memory bandwidth for that is really constraining. You can think of model level ways to solve that problem, which is to say maybe I can compress it into a few bytes or something like that. But it's interesting that the most effective way to solve it has been, it's really a combination of everything, but the most effective way to solve it has been, once you hit your 300,000 token limit, have the model go back through it and compact.

23:58

SPEAKER_01

[SPEAKER_00] Yes. Yes. And I mean, it's what Open Claw is doing. It's compacting everything you've done. Yeah. But it's funny that it's so manual. Yeah. I mean, I think manual is the wrong word. I mean, it's so primitive. It's maybe because it's so controllable. You can, if you want to iterate on how you compact, you give a different prompt and you say, compact this way, compact that way. You can iterate on that in seconds or minutes. Yeah. Whereas if you're trying to do some iteration on the model level where you say now I've got a different model architecture, it's going to take months to try and launch something.

24:13

SPEAKER_01

Yes. Yes. Any other AI predictions? I'm generally just interested in what makes models cheaper and faster. So that's at the model architecture level, really tied into this context thing. I think the context size will stay ballpark the same where it is, maybe a few times larger. But the parameter count will go up. Parameter counts will grow much faster than context length, actually, just because of the underlying physics of what's available. So has that been the story, like would that be a reacceleration of parameter count? Because it feels like we've leveled off slightly in the last year or two, and instead we've been focusing on more and better RL.

24:26

SPEAKER_01

Yeah. Okay. Parameter count or token, thinking tokens, I guess. Those are available, but the context length I think is struggling to grow. Yeah. Yeah. Okay. But you think we keep context the same length but we're better at working with large context? Is that what you're saying? Because yeah, I mean, have application level interventions to manage large context, like compacting. Yeah. Yeah. Because I think everyone's had the experience of the chat conversation and the further down in the chat you get it just gets looser. Yeah. It's sloppy. It's really sloppy by the end. And it's making mistakes with our. So you're saying we start to do better with large context. Okay. I buy that.

24:33

SPEAKER_00

When will I be typing into a chat window and it is a Mataxa chip underneath it, powering it? Yeah. Okay. So in 2027, I will be seeing very high performing chats as a result of. In the one percent experiment of the users or something like that. Yeah, exactly. Yeah. I need to find a way to angle myself into the AB test. Yeah. Mataxa is a hundred people. That's right. Yeah. How have you gone about building the team, the culture? Yeah. So what we have on the team is hardware, mostly hardware, but a big software team and also a big ML team. I think the ML team is quite unusual in what we ask them to do. When you look at a typical ML team in a...

24:55

SPEAKER_01

[SPEAKER_00] Yeah. Okay. So in 2027, I will be seeing very high performing chats as a result of, uh, [SPEAKER_00] In the 1% experiment of the users or something like that. Yeah, exactly. Yeah. I need to find a way to [SPEAKER_00] fine angle myself into the AB test. Yeah. MATAX is a hundred people. That's right. Yeah. [SPEAKER_00] How have you gone about building the team, the culture? Yeah. So what we have on the [SPEAKER_00] team is hardware, mostly hardware, but a big software team and also a big ML team. I think [SPEAKER_00] the ML team is quite unusual in what we ask them to do. When you look at a typical ML team in an

25:23

SPEAKER_01

[SPEAKER_00] AI chip company, it will be what I might say, ML engineering or ML performance. They're writing [SPEAKER_00] kernels that actually use your hardware on a given model. [SPEAKER_00] There's a missed opportunity there. If you're saying all we do is we take other [SPEAKER_00] people's models and we write kernels for them. You're optimizing this, but you can't [SPEAKER_00] optimize this at the same time. And so we want to optimize the whole thing at the same time. So [SPEAKER_00] real co-design. Yeah. So our ML team is actual real ML research. What they do every day is [SPEAKER_00] they train small models from scratch, focusing on numerics and attention.

25:58

SPEAKER_01

[SPEAKER_00] And this has really helped us make an interesting product. It's [SPEAKER_00] straight up most strong in our numerics. We often see when people design numerics is they [SPEAKER_00] say, well, back when float 32 was popular, it would be, I'm going to follow the IEEE standard. [SPEAKER_00] Now it is follow the open compute standard. And there's lots of little details where [SPEAKER_00] you say things like, what's the rounding mode I'm going to use, [SPEAKER_00] like round to nearest even or something like that, which is the best known standard for how to round.

26:22

SPEAKER_01

[SPEAKER_00] We want to cut corners anywhere we can. And so maybe don't do the best rounding. Maybe don't

26:26

SPEAKER_00

do the like don't get all the corner cases correctly. That's a very scary proposition if you're just making those choices blind. But if you have the benefit of a research team who can back you up as you do that, it's really powerful and it's really interesting that we can make some sloppy choices in these cases. Yes. I feel like often technical advances come through better iteration loops. A favorite example of this I found recently was that the Wright brothers actually had a failed season before first flight. So I guess first flight was 1904 and they were down in Kitty Hawk in 1903 and not making that

27:11

SPEAKER_00

much progress. And they went back to Ohio and they had a wind tunnel and they were testing their design in the winter. You can imagine not a lot of wind tunnels in 1904. And they did a lot of wind tunnel testing and their successful flight was after that. Is this something you're focused on where, you know, to get better chips, you allow for a better testing and iteration loop and what does that look like? Yeah. I think this mostly happens in the architecture and product definition stage. Maybe even more generally, I think AI chips seem to live or die by product definition and architecture. What is the most extreme form of

27:45

SPEAKER_00

fast iteration? It's doing it in your head. And so can you map a model to hardware in your head? Can you estimate the performance of what it is in your head? You're not going to be 100% perfect, but maybe you can prove some kind of lower bound on performance. And so the simplest possible thing is my model has a trillion parameters. My device can do a billion multipliers per second. So it takes a thousand seconds to run or something like that. Just do that simple division. But then there are much more complicated things. Like, [SPEAKER_01] we tend to look at resource balances and so, how many memory fetches do I need to do per

28:23

SPEAKER_00

[SPEAKER_01] multiply or something like that? So we do. I mean, at least the way I like to do design [SPEAKER_01] and architecture and optimization is to be able to estimate the performance to [SPEAKER_01] within about 30, 40%, before even typing anything in at all. And so we've tried to do that a lot. [SPEAKER_01] A lot of our architecture comes from there. Then the next stage of iteration is,

28:43

SPEAKER_01

that's on the performance side. This also happens on the circuit design side as well. Can you take a circuit and say, what is the gate count on that? So a 16-bit multiplier has approximately 16 squared gates. And you can do that for more complicated things by sorting networks and so on. So we already have a pretty good idea of the costs and speeds

29:03

SPEAKER_00

of things at that point after doing these calculations. Then what we tend to do as

29:08

SPEAKER_01

[SPEAKER_00] the next step of iteration is on the ML side, we run model experiments. You get

29:15

SPEAKER_00

iteration speed just by having small models mostly. And then on the hardware side, we use simulators, performance simulators to do the next level of detail to make sure

29:28

SPEAKER_01

[SPEAKER_00] we're seeing all the things we want to see. Yeah. Yeah. This idea that you should, [SPEAKER_00] the best iteration is in your head, it's reminding me of Jeff Dean's

29:38

SPEAKER_00

numbers every software. Yeah. Like, do you have your equivalent of that? Numbers every MATX?

29:41

SPEAKER_01

[SPEAKER_00] Yeah. We have go slash gates in our company, which says, what is the cost of an XOR gate,

29:49

SPEAKER_00

an AND gate, a full adder, SRAM bit cell and so on. And you want people to be working with that stuff in their head and have an intuitive sense for it because it can't be better iteration. What is the pitch to someone joining MATX? I mean, I think if you are someone who likes optimizing, just optimize something, software, hardware, factorial, whatever. If you're trying to fit something into the smallest budget possible, I think it's a pretty exciting place to be. I think hardware companies in general are really exciting because you have such a broad range of skills of people on the team. You have software

30:09

SPEAKER_00

people, you have hardware people, you've got physical design, you've got people who are looking at the insertion force of a card into a rack. And so there's

30:23

SPEAKER_00

Iteration. What is the pitch to someone joining MATX? I think if you are someone who likes optimizing, just optimize something, software, hardware, factorial, whatever. If you're trying to fit something into the smallest budget possible, I think it's a pretty exciting place to be. I think hardware companies in general are really exciting because you have such a broad range of skills of people on the team. You have software people, you have hardware people, you've got physical design, you've got people who are just looking at the insertion force of a rack into it, of a card into a rack. And so there's so much discussion and learning you can do. I think MATX in particular, we really care about this and I think we extended all the way up into the application and the machine learning as well.

30:28

SPEAKER_00

[SPEAKER_01] And so, really, really, really interesting tactical problems. And I think just generally there's lots of interesting people to talk to. [SPEAKER_01] Yes. And presumably in terms of impact, if you can design a meaningfully higher throughput chip, a 20% higher throughput chip means 20% more AI is happening. You know, if the bottleneck is elsewhere, like power or something like that or cost, you actually just are meaningfully increasing the amount of intelligence in the world, which is presumably exciting to people.

30:38

SPEAKER_01

[SPEAKER_00] Yeah. Yeah. I mean, I think this shows that both as just applying in more applications as well as just how smart is the model. Yes. Yes. Why Rust?

30:42

SPEAKER_01

[SPEAKER_00] So a previous project I worked on at Google, we did a lot of Haskell. I did Haskell going when I was at school. I loved it, very principled, very interesting. I like Haskell, but I also like making stuff fast. And then the question is, what is the first thing you want to do? You want to be able to modify your memory. Haskell, you jump through hoops to do that. Maybe I just want a language that is like programming, like functional programming, that lets me modify my memory. So I think Rust has a lot of the nice things which are like type classes or traits, and a rich type system. One of the things that we have done, interesting ways we use it at MATX are the range of data types that you express on software. Like what are the integer types? Int32, int64, int8, maybe that's all you care about. But it turns out in hardware, you care about every single bit. And so you want to use 17, 18, 19 bit integers. That is quite natural to express. And we build up a whole ecosystem of rich hardware data types in Rust as well.

30:47

SPEAKER_01

Has Rust beaten Go for the position of performant type programming language with modern features or do they actually address different?

30:54

SPEAKER_01

[SPEAKER_00] Yeah. I mean, there's the Rust marketing will say, which is safe without garbage collection, which I think is a real, is the objective thing that you can say is different, but barriers the lead, which is it's also just like it's got nice type system features that Go doesn't have. And then why does garbage collection matter at all? I mean, people often focus on the time it takes to run garbage collector. But the other thing is that every time you allocate an object, you've got the object and then you've got the garbage collector header at the beginning. And so it uses a lot more memory as well. And so if you want to design some data structure that uses the right amount of memory rather than a bit more than.

30:58

SPEAKER_01

Oh, sorry. So I hadn't realized that in Rust you're allocating your memory manually versus in Go you have a garbage collector. Yeah. Yeah, that's right. That's right. Okay. And you prefer that for what you're doing or it just is. [SPEAKER_00] I just really like dealing with the details. Like you give me a puzzle and I'll be like, let me solve every single piece of it. [SPEAKER_00] Yeah. Yeah. [SPEAKER_00] So that tickles that part of my mind with Rust.

31:32

SPEAKER_00

It seems like you're a fan of optimization generally. Is that a fair characterization? Yeah. Yeah.

31:40

SPEAKER_01

[SPEAKER_00] Where else have you, so chip optimization is one domain, where else?

31:45

SPEAKER_01

[SPEAKER_00] Yeah. So I mean, I started, when one of the really exciting things I found about working at Google is the whole Google code base is available and you can look at how does a memory allocator work? How does a mutex work? How does a hash map work? Any of those things. And you can go and look inside the implementations. And Google has excellent implementations of those, some of the best you could write. So one of the things I did on my nights and weekends when I was at Google was just go find those implementations, write a benchmark. How many nanoseconds does it take to allocate eight bytes of memory? And then can I make that faster? Can I, maybe I inline this function. Maybe I look at the assembly and say, looks like there's a few memory moves here, or there's some registers that are being used that I don't need in the fast path. I only need in the slow path. Can I do something there? So that was always my fun and learning activity. Being outside of Google, I mean, I probably could have done this inside of Google as well, but outside of Google, I felt the luxury to be able to talk about these results as well. One of the things I've looked at recently is just hash tables are used so much. One prompt for me was like, what would, if I wanted to design custom CPU instructions for accelerating hash tables, like hash tables are one of the most common things. I'm looking at them up and writing them all the time. What would the optimal CPU be for that? And so then

31:50

SPEAKER_01

[SPEAKER_00] Well, then following down that chain is like, what are the best hash table implementation in the first place? And so I spent some time looking at different SIMD implementations, and there's this really cool technique called cuckoo hashing, where you hash into two different locations and then you use the bucket, which is less full. It's been in the literature for decades, and yet the best hash table implementations don't use it because it's somehow not practical. And so, [SPEAKER_00] I'm sorry, why is it not practical?

32:11

SPEAKER_01

[SPEAKER_00] writing them all the time. What would the optimal CPU be for that? And so then the follow-up is, what is the best hash table implementation in the first place? And so I spent some time looking at different SIMD implementations, and there's this really cool technique called cuckoo hashing, where you hash into two different locations and then you use the bucket which is less full. It's been in the literature for decades, and yet the best hash table implementations don't use it because it's somehow not practical. Why is it not practical?

32:20

SPEAKER_01

[SPEAKER_00] Practical hash tables these days are considered to be ones that use SIMD vector instructions to scan eight buckets at a time. And the way cuckoo hashing is normally described is I look up one bucket here and one bucket there. And so I'm not using the vector instructions. Vector instructions are much faster than scalar instructions. And so there's a missed opportunity. Take the two good ideas and stick them together: do vector instructions on cuckoo hashing. You have to be careful to get the details right, but if you get it right, you can actually just win.

32:26

SPEAKER_01

Cesare, is your claim that one could design a custom CPU that has way better hash table performance, or even on current chips, you could get way better hash table performance? [SPEAKER_00] Both. I'm interested in what you can do in designing custom hardware, but Maddox doesn't make CPUs. We're not going to make CPUs.

32:39

SPEAKER_00

You could. New line of business. We just want to focus on shipping one product well for the time being. Fair. Good answer. I think it's an interesting exercise, but I don't get to feel the endorphins of seeing the number going down. So I first did this on just Intel CPUs. And you can get better performance than some of the best hash table implementations available using cuckoo hashing on Intel CPUs. What are examples of workloads that are really hash table read intensive? JavaScript, I guess. It's a tricky exercise because when you really think about it, you ask yourself, did I really need a hash table there? I probably didn't.

33:04

SPEAKER_01

Yeah. But you just reach for it all the time.

33:09

SPEAKER_00

You could go to the Google JavaScript team and probably help them get better performance in the Chrome JavaScript engine. Potentially. I mean, I'm not going to spend my time on that. Well, if you're listening to this podcast, here's a free idea from Reiner. And then explain the dragon. This is from a book that when I was working on the JAX team. So the JAX team is one of the ML infrastructure teams at Google. I was there as the most recent team before I left. [SPEAKER_01] What does the JAX team do?

33:38

SPEAKER_00

The JAX team develops Google's new, more modern version of TensorFlow or competitive PyTorch. It's how you write models in Python to run on TPUs. A big part of the JAX team, however, is to say, okay, we have JAX, the technical artifact. Can we help enable users to actually use it really well and get high performance? [SPEAKER_01] And so ultimately that became, well, who are the users? It's people writing LLMs. How do you get good performance on LLMs?

33:47

SPEAKER_00

[SPEAKER_01] A really strong team, the JAX team at Google, although as with a lot of brain people are now elsewhere as well. And so we developed a lot of the different techniques for how to lay out models efficiently on many chips. And so ultimately some people at Google, and I contributed after I left Google, wrote this guide called How to Scale Your Model, how to run an LLM as fast as possible. It is the main reference for how to get high performance on TPUs. There is now also a GPU version of this as well. It's a dragon because it's how to train your dragon.

33:54

SPEAKER_00

[SPEAKER_01] I see. Okay. Last question. People might not have thought that there's room for new chip companies. It might have seemed unusual or very hard. And you guys seem like a very good approach with that. Where do you think are other opportunities for companies to be started here in 2026? Where do you think people should be looking for entrepreneurial opportunities or technical challenges that haven't been properly addressed? More labs, I think, is still interesting. Can we do more model architecture is always interesting.

34:06

[SPEAKER_00] You think we have not fully explored model architecture space? [SPEAKER_00] The Frontier Labs have done a pretty good job of exploring it, but I think as the hardware changes, the shape of the model should change for sure. And presumably you're not thinking yet another Frontier Lab pursuing the same architecture. You think there's probably off-the-wall looking architectures that will actually make a lot of sense. There's definitely a little bit off the wall. Do you have a specific architecture in mind? My mentality is always sticking within the Transformer family, but what are the constraints that are currently imposed that you could lift?

34:25

SPEAKER_00

[SPEAKER_01] For example, one of the things is there's this idea when you're doing Transformer inference, you do pre-fill. That is processing what the user said to you. And then there's decode, which is generating the response. And those are totally different in every aspect of how they actually run. One runs a step at a time. The other one runs in parallel. So there is this somewhat artificial constraint today that those are the same model doing both. Maybe lift that constraint. Another example would be there's this idea that the model that you train is the same model that you serve. But training is very different from serving. At training, it's very compute intensive. At serving, it's more memory bandwidth intensive. And so maybe there's a way you can make a model that when you use it at inference time, it increases the amount of compute it does to use some of the available resources.

34:29

SPEAKER_00

Makes sense. Well, Rainer, thank you. [SPEAKER_01] Pleasure. I think this is a great question. [SPEAKER_01] available resources? Yeah. Makes sense. Well, Rainer, thank you. Pleasure. [SPEAKER_01] I think this is a great question. Yeah. A long time. Just how do you think about parallelization and is that narrative true? Yeah, it's true for sure. Um, it's true in some, in, in many, uh, areas of the market. Um, I think, uh, and especially where you look at where NVIDIA entered the market, um, they're doing, uh, like, uh, PC devices, uh, lots of gaming, um, and so on. There are thousands of games, maybe tens of thousands

35:17

SPEAKER_00

of games released, um, and they all need to be programmed against, uh, against CUDA. And so, um, it's, there's such a huge investment in the software that is, that this is really important to their compatibility. There are not thousands of LLMs. There's one LLM per Frontier Lab, and there's maybe five Frontier Labs or something like that. Um, and so the, just the economics of that is different. The calculation, uh, for Frontier Lab roughly goes as I just bought a $10 billion compute cluster. Um, I have hired, uh, 50 of the best, uh, people who can write, um, optimized, uh, GPU or TPU or Tranium

35:54

SPEAKER_00

software. Um, uh, I pay them less than $10 billion, a lot less. Um, and so, um, let's put them to

36:02

SPEAKER_01

it to work, uh, optimizing the, the compute. And so, uh, they can like good work there can, I mean, depends on what your baseline is, but it can very easily double the performance of, of the software you write. And so there is a huge amount of custom software written for every

36:17

SPEAKER_00

generation of chip. Um, when a new chip comes out, uh, software is, uh, like substantially rewritten to optimize for that specific chip. And that's just the right trade-off given the, the relative costs

36:27

SPEAKER_01

of these things. Um, what that means for us is that, uh, that, that ecosystem already exists, uh, and that way of operating where you say, I'm just going to staff a 50% team to, to run on, uh, to write software for this chip, um, works really well for, if you're trying to sell to Frontier Labs. Okay. So you're saying CUDA is way more important for the games environment,

36:49

SPEAKER_00

where it just has a lot of games than this top heavy AI market that we're in, Yeah. Where if people say you need to then customize your workload for a Matex chip, it's like, well, fine. Fine. Yeah. It's the cost of business. Yeah. Yeah. Yeah. Yeah. That makes, that makes a lot of sense. Um, where will you fab the chips? TSMC. Okay. Yeah. Um, why is TSMC so durable? Yeah. I mean, it's interesting. They don't charge a lot as well. You'd think that if they're monopoly provider, they should charge a lot of money. Um, they don't. Um, I think that is a big aspect of why they're so durable. Uh, it's like this cyclical, it's cyclical conservatism

37:37

SPEAKER_00

crossed with Taiwanese business conservatism means you're at the like most conservative part of the, Yeah. Yeah. The matrix. But, but I mean, it does, uh, I mean, like an American capitalist might say, well, they're just, they're just screwing up. They could, they could have extracted more money from the market, but you could also say that, um, there's actually this long-term, uh, sustaining advantage because, um, uh, they will just stay ahead for a really long time. They don't encourage the creation of competitors. Yeah. Yeah.

38:03

SPEAKER_01

But isn't the creation of competitors kind of priced in because the geopolitical risk. And so like, it's not like everyone's fat, dumb, and happy with their TSMC dependence. They're actually thinking a lot about it. Yeah. I mean, so there is real technical advantage there as well.

38:17

SPEAKER_00

Yeah. Yeah. It's not, it's not just like the discouragement. But like, um, standing chip seems really hard. Building airplanes seems really hard.

38:23

SPEAKER_01

There are so many areas where competitive market forces create multiple options. Yeah. And yet that has not occurred here.

38:33

SPEAKER_00

So, I mean, there are multiple options. You can, you can buy from Intel or Samsung. Um, but at leading edge nodes. Yeah. Yeah. So, I mean, what do we even care about in leading edge nodes, I guess? The big advantage is on power. Um, the advantage on area is smaller. Um, the leading edge nodes, the density doesn't go up as much as it used to. Um, so when you are really, really sensitive to power, um, it is a good idea to be on leading edge nodes. So that is AI chips and mobile phone chips. Yeah. Um, but there are, there's a lot of the market where, where you don't like devices. Sure. Sure. Yeah. Car chips. Yeah, that's fine. But, but you're kind of saying like,

39:07

SPEAKER_00

if you exclude the two most interesting parts of the market, you know, yeah, you can, um, just for, for this super high growth area of the market, it's interesting to me. Like, again, there's a lot of other really complex business problems out there that competition has, has solved. And yeah, chip design is like, why has someone not left TSMC and gone and built a new fab? Yeah. I mean, I don't know. It's, it's the cost of a lab of, of a fab is extremely expensive. I mean, I recognize that also the cost of a lab is extremely expensive too. Um, uh,

39:42

SPEAKER_01

I don't really understand the technical details of why it's so hard. Um, uh, I mean, there is some amount of just a $10 billion fab versus a hundred million, um, tape, uh, like tape out and chip development. There's a huge difference there, but beyond that, I'm not sure. What's TSMC like to deal with? So they're very big. So as a startup, we tend to work with, um, uh, not directly with TSMC, but with, uh, um,

40:03

SPEAKER_00

an ASIC vendor who, who, I mean, firstly does a huge amount of the actual backend work for us, the chain of X with them. Um, but then also has their existing relationships with them. Got it. Um, TSMC cares a lot about diversity of their customer pool. And so, uh, It gets back to that conservatism. Yeah. So we're, they're, they're great to work with from, from that perspective. Like they, they, they want to encourage starting. That's right. Yeah. Yeah. That's very cool. Why don't the labs design their own chips? I mean, Google does, but Google does. Um, OpenAI is, is starting. It's really a trade-off of how much advantage you get from vertical integration

40:35

SPEAKER_00

versus how much advantage do you get by concentration of, uh, R and D work. So, um, you take the five labs and if they all buy from one player, then you can put like five times as much R and D into that chip. Um, and does that beat the advantage you get from saying, I know exactly what my model is because of the like several years delay from, from designing a chip to being in production. Um, you can't actually say, I know exactly what my model is because, uh, models change like, uh, much faster than that. So, uh, even the labs are forced into this position where they have to make predictions and they have to hedge against what, what they might do two years from now.

41:12

SPEAKER_00

The calculus is sort of like, what is the probability distribution of what my model might look like? And then sort of, um, design a chip that gets like 90, 90% of that probability distribution or something. Yeah. Yeah. Elon is excited about data centers in space. Yeah. The two criticisms I've heard are that cooling is very hard and then just repairing the chips is hard, but I know nothing about chips. You do. Yeah. So, I mean, the repair I think is really interesting. Um, when you look at how NVIDIA deploys their X, how we, we do something pretty similar to what NVIDIA does. Um, I mean,

41:50

SPEAKER_00

in general, you always need to design for the fact that some of your chips are going to be down. Like, um, uh, mean time between failure of chips is not that large. And so, uh, in a cluster of, uh,

41:59

SPEAKER_01

100,000 chips, there's going to be chips that are down all the time. Um, one way you can do that is you can make a rack with one, where one rack has, um, some spare chips in it. NVIDIA has, uh, eight spare

42:09

SPEAKER_00

chips and a rack of 64. Um, uh, that's pretty good. Uh, the, like the common ontarix works really well for you there. That's the sort of, uh, you can actually like, because you can pick which ones to avoid, you can, um, uh, you can like with very high probability tolerate a lot of failures. And then the other, just for like the other family of things is to say my rack has to work, but I have some spare racks as well. So you kind of, you can math that out with, uh, like, like the tax of reliability here is only like 10%. That's pretty good. Um, uh, and, and you, uh, but that relies on

42:42

SPEAKER_00

someone coming in and servicing the device part within a day or something like that. If you say they're going to serve us at never, then, um, I think you actually can get where you want to be, but maybe with a hundred percent tax on reliability rather than 10%. So, so for example, if you think the average lifetime of a chip is in the range of three to five years, um, so that means if I deploy twice as many chips, then three to five years from now, half of them will still work. Yeah. And also the burn-in is particularly, um, failure-y. And how about the cooling? So most of the challenge uh, I mean, I, I guess there's a, there's actually really a data center design aspect,

43:16

SPEAKER_00

uh, that then, um, at the rack level, the challenge of cooling is just getting the heat out, like as quickly as possible out of the rack into, to the, um, cooling network. Um, how you get it out of the, the, the spaceship, um, other people would know that better than I do. Okay. Yeah. Yeah. Yeah. Um, again, that, that seems to be the main objection, but, um, I don't know. Yeah. I mean, I, I think it's sort of like, if you think the cost of repair is that you need to have deployed twice as many chips, then like, it's a trade-off of the capital of the chips versus the power saving. Exactly. The repair thing, it feels like can be solved because also I think part of

43:52

SPEAKER_00

the best, you know, probably one's claim is that we will just be so power limited that, you know, you have no option but to go, uh, to space and, you know, people can argue about that, but were that to be the case, then yes, it's like, well, you can get power in space and you cannot on earth. And so you, uh, you might as well go there. Whereas like the cooling is a more fundamental, does the product actually work at all? Yeah. Yeah. Yeah. Rayner thinks about AI the unglamorous way, compute systems architecture and what it takes to run models reliably at scale. And if you're building an AI product, the business model similarly has a ton

44:28

SPEAKER_00

of unglamorous complexity. You're not just selling AI, you're monetizing consumption across API calls, tokens processed, GPU hours. Stripe billing is a scalable system for usage-based billing. It lets you launch token-based pricing, subscriptions, credits, hybrid models, whatever you want. So you can create revenue models based on usage without rebuilding your pricing system every six months. If you're building an AI product, Stripe billing is worth a look.

44:55

SPEAKER_00

What are your AI predictions for Twin26?

44:57

SPEAKER_01

I mean, what I'm really excited about is, um, just being able to, I mean, I'm still excited about the coding. Uh, that's, this is what we do as a company. It's what many others do as a company as

45:09

SPEAKER_00

well. Um, the, uh, one aspect of this is expanding into more domains. Um, so for example, in, in where we spend our time, um, uh, we, as a company, we write Rust, we write Verilog, we write Python. Um, I don't know. Haskell. Yeah, no, there's a story there. I, I, I, I used to love Haskell. I, Rust is my

45:30

SPEAKER_01

current favorite. Okay. Mutation is good. Um, uh, um, the models are extremely good at Rust and Python.

45:40

SPEAKER_00

Um, they've done a lot of RL on them. Um, they have not done as much RL, RL on Verilog. Um, they've done almost none on, okay, write, write me a markdown file that describes a chip architecture. Um, and then how do you even RL on that? You have to say like, uh, what is a good chip architecture? I have to say it somehow say whether that's a good result or not. Yes. Um, I think one of the things the labs are doing is trying to broaden what they've done RL on, RL on source it from customers and so on, um, in order to, uh, sort of fill out the, the, the, the knots, the, the, make it less spiky, fill out the

46:16

SPEAKER_00

gaps between the spikes. I presume the labs would love to work with you on, uh, improving the models, uh, by doing RL on this specific, uh, task. However, it's also somewhat- It doesn't make sense for us. Yeah. You're a special source of it. So do you want to come up with some AI approaches, but keep them proprietary? Is that- Yeah. So, I mean, we've looked at a few different aspects here. There's the, I mean, what we're able to do by ourselves, our business is not training models. We do it in order to do the research on numerics, but like

46:50

SPEAKER_01

actual production models we don't do. Um, so it's like the biggest mileage I think is on the RL and, it's not something we can really do ourselves. We'd love it if we could have a custom model just for

47:01

SPEAKER_00

us, but, uh, that doesn't seem to be- Well, you could, right? The terms we've been offered by labs so far have not been on those terms, but, uh, because you have to share the IP back. The way they prefer to do it is that they, that you, um, that they, that they put it into their mainstream model. Yeah. Because it's good for them. Yeah, yeah, yeah. Which obviously you don't want to do. Um, yeah. I mean, how do you think you, what does you using AI to design a model do you think look like? Because this is actually, I think, an interesting sight glass into the, you know, a weak version of

47:29

SPEAKER_00

recursive self-improvement where, you know, we're using the AIs to develop better AIs. And so I'm curious, yeah, what, what do you, what you think that looks like? Is it your own proprietary, um, recursive models? What, what else? Like, is there kind of day-to-day AI usage that's load-bearing? Yeah. I mean, so the, the stuff that is available today and I think, uh, will become even better very quickly is just, um, the stuff that looks most like software. So writing Verilog, running tests,

47:55

SPEAKER_01

running continuous integration and so on. Um, the, uh, and that is a big fraction of the development time in the chip. It's probably nine, 12, 15 months or something. Um, the, there's some stuff that's downstream of that, which is physical design, which is, uh, you take that Verilog and, and you generate the, the gates and the polygons. We don't have a clear path for like, it's not, at least the most obvious thing is not clear for how to, how to compress that. Like, like the goal, can you, can you tape out a chip in one month? One month would be the goal. Um, in theory, you could compress all of

48:29

SPEAKER_01

the logic design and design verification down to a short amount of time if the, uh, just by continuing on the same path we're doing now. But if you wanted to take, uh, the physical design down, that has to leave code. You're now doing like graphical interfaces and saying, well, I want to play stuff and so on. Yeah. Actually there has been work on this even prior to, uh, LLMs, um, uh, which is like specific, um, model trained for that particular problem. Yeah. And, and I think the vendors, uh, which is like synopsis and cadence, um, probably will, should move in that direction. Um, most of the focus has

49:04

SPEAKER_01

not been do it faster. It's been do it with higher quality. Um, uh, but, but that is a big bottleneck on, on, on like, can I have a new chip every month? And then there's just the practical thing of like, a new chip every month doesn't really make sense because then if I'm deploying, uh, like if it takes me a year to populate a data center, that means I'm going to have different chips in different corners of the data center. Yes. Yes. Sir, when you talk about one month to tape out, so you do all this work to ultimately produce a file. Uh, everything TSMC then does, it's not entirely in software. Like, is there some typesetting that has to happen of moving stuff

49:42

SPEAKER_01

around? But yeah, what happens when you, you send your files to TSMC? Yeah. Then what? So they create a mask. So that is, uh, where the ASML tools come in. Um,

49:52

SPEAKER_00

and a mask is, it is really just a stencil. You, you shoot the, the lasers through the mask and all the

49:57

SPEAKER_01

x-rays through the mask. And, and then that, um, that, uh, produces the, uh, different, uh, p-type and n-type semiconductors. Um, so they produce the mask. That is the expensive part. Um, and then, uh, uh, and then they're, they're building up these like 15 or so metal layers. Uh, so they, they place it on the silicon and then there are different layers of metals, um, which could connect all the transistors together. Um, uh, they do that on a wafer. It, it, it happens on, on a, on a stepping basis. Uh, so there's a sort of a maximum size of chip you can build, which is constrained by this machinery. Um, the wafer stepper is part of the ASML special sauce, right?

50:34

SPEAKER_01

Yeah. I guess there's probably some important alignment requirement there.

50:37

SPEAKER_00

Yeah. I think I remember that being quite, uh, like the, you know, it's a classic manufacturing throughput problem. And I think they've done a lot of work on optimizing that. Yeah. Yeah. Um, so, so they take that. So then you just produce hundreds of copies of, of your chip. Um, you have to test it because there's defects. Uh, you typically, I think the average rates, um, really depends on process and, and so on, but small single digit number of defects per, per chip. Um, so you test the chip and see whether it has any defects in it. Um, many chips are designed to be able to tolerate a few defects. Um, and so you need to configure it to tolerate the

51:10

SPEAKER_00

defects. Um, and now you have a die that by itself works. Uh, and then you need to package it. So you put it on, uh, in a package together with, uh, memories, typically that's the HBM. Um, and then, and maybe, uh, some, you escape the wires to, to connect to other chips. Yeah. How long does it take to make a mask? Um, the, so, I mean, what we see is time from, like, tape out to, to first chip, to chips back. Uh, again, depends on node, but it's ballpark four or five months. Oh, so tape out is just, uh, like sending the file or? Yeah. Well, I mean,

51:41

SPEAKER_01

I assume tape out is. We can send a tape out, send the file, and then, and then there's a whole process of you, you make the masks for all the layers and then, and then, uh, actually just producing

51:48

SPEAKER_00

the chips. Got it. So producing the masks and producing the chips happens after tape out. That's right. I see. Okay. So like, is the term tape out from like, you send a magnetic tape with the instructions or something? It could be. I was in software when, when, when the film was created. I'm curious what the tape actually means. It feels like, you know, when we're thinking about AI predictions, one thing I'm really struck by is how, uh, still in 2026, every time you open a chat window, it's contextless. Yeah. It's got no memory. And now to be fair, it's like, guys, it's been four years, like just not even four years. It's been three and a half

52:25

SPEAKER_00

years. Just calm down. We'll get there. But I also interpret a lot of the current enthusiasm for open claw and all that stuff as it's like this super hacky backdoor into state management where, you know, your little claw will write a markdown file of what it's doing. And then, you know, look at that markdown file the next time and things like that. But it just feels like state management and memory is going to be a huge deal. And that will really change the character of AI products. Yeah. It's really interesting. Like the, the, so, I mean, long context is the reason, is one of the biggest bottlenecks on model perform on speed of the model. Yes. Um, it just,

53:09

SPEAKER_00

like every single token you generate, it reads through all of the previous tokens, or maybe it reads through a subset of them, but reads through a lot of the previous tokens you've written. Um, and so memory bandwidth for that is, is, is really constraining. You can think of like model level ways to solve that problem, which is to say, um, maybe I can compress it into a few bytes or something like that. Uh, but it's interesting that the sort of most effective way to solve it has been, I mean, it's really a combination of everything, but the most effective way to solve it has been,

53:36

SPEAKER_00

uh, once you hit your 300,000 token limit, uh, have the model go back through it and compact. Yes. Yes. And, and, uh, I mean, it's kind of what Open Claw is doing. It's like compacting everything you've done. Yeah. Um, but it's funny that it's so manual. Yeah. I mean, I think, uh,

53:54

SPEAKER_01

manual is the wrong word. I mean, it's so primitive. It's maybe because it's so controllable. Um, you can, like, uh, if you want to iterate on how you compact, you, you give a different prompt and you say, compact this way, compact that way. You can iterate that on that in, in seconds or minutes. Yeah. Um, whereas if you're trying to do some iteration on, on the model level where you say, now I've got a different model architecture, it's going to take like months to, to try and launch something. Yes. Yes. Any other AI predictions? I'm generally just interested in what makes

54:20

SPEAKER_01

models cheaper and faster. Um, so that's, uh, just at the model architecture level, really tied into this context thing. I think the context size will stay ballpark the same where it is, maybe a few times larger. Um, but the parameter count will go up. Like parameter counts will grow much, much faster than context length actually, just because of the underlying physics of what's available. So has that been the story, like, would that be a reacceleration of parameter count? Because it feels like we've leveled off slightly in the last year or two, and instead we've been focusing on more and better RL.

54:55

SPEAKER_01

Yeah. Okay. Uh, parameter count or token, thinking tokens, I guess. Those, those are available, but the context length, uh, I think is, is sort of struggling to grow. Yeah. Yeah. Okay. But you think we,

55:08

SPEAKER_00

we say context things are struggling to grow, but you're saying we keep context the same length? Keep context the same length. But we're better at working with large context. Is that what you're saying? Because, uh, yeah, I mean, just have, um, application level interventions to, to manage large context, like compacting. Yeah. Yeah. Cause I think everyone's had the experience, you know, currently of like the chat conversation and the further down in the chat you get. It just gets looser. Yeah. It's sloppy. It's just like really sloppy by the end. And it's like making mistakes with our, so you're saying we start to do better with, with large context. Okay. I buy that.

55:37

SPEAKER_00

When will I be typing into a chat window and it is a MATAX chip underneath it, powering it? Yeah. Okay. So in 2027, I will be seeing very high performing chats as a result of, uh, In the 1% experiment of the users or something like that. Yeah, exactly. Yeah. I need to find a way to fine angle myself into the, uh, the AB test. Yeah. Um, MATAX is a hundred people. That's right. Yeah. How have you gone about building the team, the culture? Yeah. So, I mean, so what we have on the team is hardware, mostly hardware, um, but a big software team and also a big ML team. Uh, I think

56:23

SPEAKER_00

the ML team is quite unusual in, in what we ask them to do. When you look at a typical ML team in a, uh, AI chip company, it will be, um, what I might say, ML engineering or ML performance. Um, they're writing kernels that, uh, that actually, uh, just like use your hardware as well on, on a given model. There's sort of a missed opportunity there. If you're saying, uh, all we do is we take other people's models and we write kernels for them. Uh, you're, you're optimizing this, but you can't optimize this at the same time. And so we want to optimize the whole thing at the same time. So

56:55

SPEAKER_00

like real co-design. Yeah. So our ML team is actual real ML research. Uh, what they do every day is they train, um, small, uh, um, from scratch, focusing on numerics, um, and, and, and attention. Um, and this has really, really helped us, uh, make a, like an interesting product. Um, it's straight up most strong in our numerics. Um, we often what you see when people design numerics is they say, well, back in when, when float 32 was popular, it would be, I'm going to follow the IEEE standard. Now it is like follow the open compute standard. Um, uh, and there's lots of little details where

57:37

SPEAKER_00

you say things like, um, maybe what's the rounding mode I'm going to use, uh, like round to nearest even or something like that, which is like, uh, the, the, the best known standard for how to round. We want to cut corners anywhere we can. Um, and so like, maybe don't do the best rounding. Maybe don't do, uh, the, uh, like don't get all the corner cases correctly. Um, that's a very scary proposition if you're just making those choices blind. But if you have the benefit of a research team who can sort of, uh, back you up as you do that, um, it, it, it, it's really powerful and it's really interesting that we can make, you know, make some, uh, sloppy choices in these cases.

58:11

SPEAKER_00

Yes. I feel like often technical advances come through better iteration loops. Um, a favorite example of this I found recently was that the Wright brothers actually had a failed season before first flight. So I guess first flight was 1904 and they were down in Kitty Hawk in 1903 and not making that much progress. And they went back and they, uh, to Ohio and they had a wind tunnel and they were like testing their design in the winter. You can imagine not a lot of wind tunnels in, uh, in 1904. Um, and they did a lot of wind tunnel testing and their successful flight was, uh, was after that.

58:46

SPEAKER_00

Is this something you're focused on where, you know, to get better chips, you allow for a better testing and iteration loop and what does that look like? Yeah. Uh, I think this mostly happens in the architecture and, uh, product definition stage. Uh, maybe even more generally, I think, uh, AI chips seem to live or die by product definition and architecture. Um, what is the most extreme form of fast iteration? It's doing it in your head. Um, and so can you map a model to a hardware in your head? Can you estimate the performance of what it is, uh, in your head? Um, you're not gonna be 100%

59:19

SPEAKER_00

perfect, but maybe you can prove some kind of lower bound on performance. Um, and so, uh, the, I mean, the simplest possible thing is my model has, uh, a trillion parameters. My, uh, device can do, can, can do a billion multipliers per second. So it takes a thousand seconds to run or something like that. Um, just do that simple division. Um, but then there are much more complicated things. Like,

59:39

SPEAKER_01

we tend to look at, uh, resource balances and so, like, how many memory fetches do I need to do per multiply or something like that? So we do, um, uh, I mean, at least the way I like to do, uh, design and architecture and, and, uh, and optimization is to be able to sort of estimate the performance to within about 30, 40%, um, before even typing anything in at all. Um, and so we've tried to do that a lot. Um, uh, a lot of our architecture comes from there. Um, then sort of the next stage of iteration is, oh, that's kind of on the performance side. This also happens on the, um, on the circuit design side

1:00:15

SPEAKER_01

as well. Can you take a circuit and say, what is the gate count on that? So like a, uh, a 16-bit multiplier has approximately 16 squared many gates. Um, and you can do that for more complicated things by sort, like sorting networks and so on. So we already have a pretty good idea of the costs and speeds

1:00:32

SPEAKER_00

of things, uh, at that point after doing these calculations. Um, then what we tend to do as, as sort of the next step of iteration is, uh, on the ML side, we run model experiments. You get iteration speed just by having small models mostly. Um, and then, um, uh, on the hardware side, we, uh, we use simulators, performance simulators to like do the next level of detail to, to make sure we're, we're seeing all the things we want to see. Yeah. Yeah. This, um, idea that you should, the best iteration is in your head, uh, it's kind of reminding me of Jeff Dean's, you know, numbers every software. Yeah. Like, do you have your equivalent of that? Numbers every MATX?

1:01:06

SPEAKER_00

Yeah. We have go slash gates in our company, which says, what is the cost of an XOR gate, an AND gate, a full adder, um, uh, SRAM bit cell and so on. And you want people to be working with that stuff in their head and have an intuitive sense for it because it can't be better, um, um, iteration. What is the pitch to someone joining MATX? I mean, I think if you are someone who likes optimizing, just optimize something, software, hardware, uh, factorial, uh, whatever. Um, if you're trying to like fit something into the smallest budget possible, I think it's a pretty exciting place to be. Um, uh, I think hardware place, hardware companies in general are really

1:01:48

SPEAKER_00

exciting because you have such a broad range of skills of people on the team. You have software people, you have hardware people, you've got physical design, you've got people who are just like looking at the insertion force of a rack into it, of a card into a rack. Um, and so there's like so much discussion and learning you can do. Um, I think MATX in particular, um, we, we really care about this and I think we extended all the way up into the application and the machine learning as well.

1:02:15

SPEAKER_01

And so, uh, um, uh, really, I mean, really, really, really interesting tactical problems. And, uh, and I think just generally, uh, um, like there's lots of interesting people to talk to. Yes. And presumably in terms of impact, if you can design a meaningfully higher throughput chip, a 20% higher throughput chip means 20% more AI is happening. You know, if the bottleneck is elsewhere,

1:02:38

SPEAKER_00

like, um, you know, power or something like that or cost, um, you actually just are meaningfully increasing the amount of intelligence in the, uh, in the world, which is presumably exciting to people. Yeah. Yeah. I mean, I think, uh, this shows that both as just, uh, kind of apply in more applications as well as just how smart is the model. Yes. Yes. Why Rust? So a previous project I worked on at Google, um, we did a lot of Haskell. Um, I, I did Haskell going, uh, like, uh, when I was at school, uh, I loved it, like very, uh, principled, very interesting. I like Haskell, but I also like making stuff fast. And then the question is, uh, what is the first thing

1:03:16

SPEAKER_00

you want to do? You want to be able to modify your memory. Haskell, you jump through hoops to do that. Uh, maybe I just want a language that is like programming, like functional programming, that lets me modify my memory. Um, so I think, uh, Rust has a lot of the nice things which are like, uh, type classes or traits, um, uh, and, uh, and a rich type system. Um, one of the things that we, we have done, um, like interesting ways we use it at Maddox are the, the range of sort of data types that you express on software. Like what are the integer types? Uh, inch 32, inch 64, inch 8, maybe

1:03:49

SPEAKER_00

that's all you care about. Um, but it turns out in hardware, you, you like, you care about every single bit. And so you'll, you, you want to use like 17, 18, 19 bit integers. Um, the, that, that is quite natural to express. Um, and we, we build up sort of a whole ecosystem of, of, of, um, rich hardware data types in, in Rust as well. Has Rust beaten Go for the position of sort of performant type programming language with modern features or do they actually address different? Yeah. I mean, so there's like, there's the, there's what the Rust marketing will say, which is, uh, um,

1:04:27

SPEAKER_00

uh, safe without a garbage collection, um, which I think is a, is a real, um, I mean, is the objective thing that you can say is different, but sort of barriers the lead, which is, it's also just like, it's got nice type system features that, that, that Go doesn't have. And then like, why is garbage collection, why does it matter at all? Like it's not, I mean, people often focus on the time it takes to run garbage collector. But the, the other thing is that every time you allocate an object, you've got the object and then you've got the garbage collector header at the beginning. And so it uses a lot more

1:04:56

SPEAKER_00

memory as well. And so if you want to design some, I don't know, data structure that like,

1:05:01

SPEAKER_01

uses the right amount of memory rather than like a bit more than. Oh, sorry. So I hadn't realized that in Rust you're allocating your memory manually versus in Go you have a garbage collector. Yeah. Yeah, that's right. That's right. Okay. And you prefer that for what you're doing or it just is.

1:05:12

SPEAKER_00

I just really like dealing with the details. Like you give me a puzzle and I'll be like, let me solve every single piece of it. Yeah. Yeah. So that tickles that part of my mind with Rust. It seems like you're a fan of optimization generally. Is that a fair characterization? Yeah. Yeah. Where else have you, so chip optimization is one domain, where else? Yeah. So, I mean, I started, I mean, when, one of the really exciting things I found about working at Google is like the whole Google code base is available and you can look at how does a memory allocator work? How does a mutex work? How does a hash map work? Any of those things. And you can

1:05:49

SPEAKER_01

go and look inside the implementations. And Google has like excellent implementations of those, like

1:05:53

SPEAKER_00

some of the best you could write. So like one of the things I did on my nights and weekends when I was at Google was just go find those implementations, write a benchmark. How many nanoseconds does it take

1:06:08

SPEAKER_01

to allocate eight bytes of memory? And then can I make that faster? Can I, maybe I inline this function. Maybe I look at the assembly and say, looks like, like there's a few memory moves here,

1:06:20

SPEAKER_00

or there's some registers that are being used that I don't need in the, in the fast path. I only need in the slow path. Can I, can I do something there? So I don't know, that was always my, like just fun and learning activity. Being outside of Google, I feel, I mean, I probably could have done this inside of Google as well, but outside of Google, I felt the luxury to be, to be able to like talk about these results as well. One of the things I've looked at recently is just hash tables are used so much. One prompt for me was like, what would, if I wanted to design like custom CPU instructions for

1:06:57

SPEAKER_00

accelerating hash tables, like hash tables are one of the most common things. I'm looking at them up and writing them all the time. What would the optimal CPU be for that? And so then Well, then the follow, following down that chain is like, what are the, what is the best hash table implementation in the first place? And so I spent some time looking at different SIMD implementations, and there's this really cool technique called cuckoo hashing, where you hash into two different

1:07:23

SPEAKER_01

locations and then you, you use the bucket, which is less full. It's been in the literature for decades, and yet like the best hash table implementations don't use it because it's somehow like not practical. And so,

1:07:39

SPEAKER_00

I'm sorry, why is it not practical? Practical hash tables are these days considered to be ones that use SIMD vector instructions to scan like eight buckets at a time. And the way cuckoo hashing is normally described is I look up one bucket here and one bucket there. And so I'm not using the vector instructions. Vector instructions are much faster than scalar instructions. And so there's kind of a missed opportunity. Again, just like take the two good ideas and stick them together. Do vector instructions on cuckoo hashing. You have to be careful to get the details right, but if you get it right, you can actually just win.

1:08:16

SPEAKER_01

And Cesare, is your claim that one could design a custom CPU that has way better hash table performance, or even on current chips, you could get way better hash table performance.

1:08:27

SPEAKER_00

So both. I mean, I'm interested in what you can do in designing custom hardware, but Maddox doesn't make CPUs. We're not going to make CPUs. You could. New line of business. I mean, we just want to focus on shipping one product well for the time being. Fair. Good answer. So, I mean, I think it's an interesting exercise, but I don't get to feel the endorphins of seeing the number going down. So I first did this on just Intel CPUs. And you can get better performance than like some of the best hash table implementations available using cuckoo hashing on Intel CPUs. And what are examples of workloads that are really

1:09:09

SPEAKER_00

hash table reads intensive? I mean, I know kind of everything, but. I mean, JavaScript, I guess. But yeah, I mean, it's sort of a tricky exercise because like when you really think about it, you're like,

1:09:22

SPEAKER_01

did I really need a hash table there? I probably didn't.

1:09:24

SPEAKER_00

Yeah, yeah, yeah. But you just reach for it all the time. Okay. But you could go to the, you know, Google JavaScript team and probably help them eke out better performance in the Chrome JavaScript engine. Yeah. I mean, potentially, like it's, I mean, I'm not going to spend my time on that. Well, if you're listening to this podcast, here's a free idea from Reiner. And then explain the dragon. Yeah. This is from a book that when I was working on the JAX team. So the JAX team is one of the ML infrastructure teams at Google. I was there as the most recent team before I left. I'm sorry, what does the JAX team do? Oh, yeah. So the JAX team develops,

1:09:54

SPEAKER_00

this is sort of Google's new, more modern version of TensorFlow or competitive PyTorch. It's how you

1:10:00

SPEAKER_01

write models in Python to run on TPUs. A big part of the JAX team, however, is to say, okay, we have JAX, the technical artifact. Can we help enable users to actually use it really well and get high performance? Mm-hmm. And so ultimately that became, well, who are the users? It's people writing LLMs. How do you get good performance on LLMs? Yeah. And so really, really strong team, the JAX team at Google, although as with a lot of brain people are now elsewhere as well. And so we developed a lot of the different techniques for how to lay out models efficiently on many chips. And so ultimately some people at Google, and I contributed after I

1:10:36

SPEAKER_01

left Google, wrote this guide called How to Scale Your Model, how to run an LLM as fast as possible. It is sort of the main reference for how to get high performance on TPUs. There is now also a GPU version of this as well. It's a dragon because it's how to train your dragon. Mm, I see. Okay. Last question. People might not have thought that there's room for new chip companies. It might have seemed unusual or very hard. And you guys,

1:11:06

SPEAKER_00

it seems like a very good approach with that. Where do you think are other opportunities for companies to be started here in 2026? Where do you think people should be looking for entrepreneurial opportunities or just technical challenges that haven't been properly addressed? More labs, I think, is still interesting. Can we do more model architecture is always interesting. You think we have not fully explored model architecture space? Yeah. I mean, the Frontier Labs have done a pretty good job of exploring it, but I think, I mean, as the hardware changes, the shape of the model should change for sure.

1:11:40

SPEAKER_00

Yeah. Okay. And presumably you're not thinking like yet another Frontier Lab pursuing the same architecture. You think there's probably off-the-wall looking architectures that will actually make a lot of sense. Yeah. I think there's a little bit off the wall. Okay. For sure. Yeah. Do you have a specific architecture in mind? My mentality is always sticking within the Transformer family, but what are the constraints that are currently available, like currently imposed that you could

1:12:03

SPEAKER_01

lift? Yeah. So for example, one of the things is there's this idea when you're doing Transformer inference, you do pre-fill. That is sort of processing what the user said to you. And then there's decode, which is generating the response to that. And those are totally different in pretty much every aspect of how they actually run. One runs a step at a time. The other one runs really in

1:12:24

SPEAKER_00

parallel. So there is this somewhat artificial constraint today that those are the same model that's doing both. Maybe lift that constraint. Another example would be there's this idea that the model that you, I mean, this is more fundamental constraint that you have to train the same model as

1:12:40

SPEAKER_01

you serve. But again, training is very different from serving. At training, it's very compute intensive. It's serving. It's more memory bandwidth intensive. And so maybe, is there a way you can make a model that when you use it at inference time, it increases the amount of compute it does to use some of the available resources? Yeah. Makes sense. Well, Rainer, thank you. Pleasure.

1:13:17

SPEAKER_01

I think this is a great question.

1:13:17

SPEAKER_00

I think this is a great question. I think this is a great question. I think I think I think I think I think I think I think I think I think I think I think I think I think I think I think I think I think I think I think I think I think I think

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note