Open Reader

Scaling the Next Paradigm of Heterogeneous Intelligence — Adrian Bertagnoli, Callosum

completed 15:13 May 24, 2026 Watch on YouTube

Current Status

completed

Video ID

WRBNDpUhsJQ

RAG / Chat

Enabled
Scaling the Next Paradigm of Heterogeneous Intelligence — Adrian Bertagnoli, Callosum
Description

A mixture of Qwen 3 VL8B and Kimi K2.5 beat the state of the art on Video Web Arena, outperforming the leading GPT and Gemini models by 18 and 25 percent while costing 3.7 times less and running 3 times faster. The reason it worked is that visual web navigation decomposes into subtasks that do not all need a frontier model: routing zoom and visual parsing to a smaller model alone produced 11x speed and 43x cost improvements on those steps. Adrian Bertagnoli from Callosum makes the case that the GPU cluster era of identical hardware and monolithic models is ending. Heterogeneous intelligence treats model architectures, chip types, and workflows as variables to optimize together. A second result: running recursive long context reasoning tasks on Cerebras instead of a frontier model cuts cost by 7x and latency by 5x while matching accuracy. Callosum is building the automation layer that routes tasks to the right chip and model without bespoke decisions for each subtask. Speaker info: - https://www.linkedin.com/in/adrian-bertagnoli-bb3467178/ Timestamps 0:14 Introduction and definition of heterogeneous intelligence 0:56 Limitations of the current homogeneous intelligence paradigm 1:36 Evolution toward mild heterogeneity (MoE, multi-agent systems, hardware disaggregation) 3:24 The rationale for heterogeneity: complexity and multi-step problem solving 4:26 Mathematical formalization of the production function and skill distribution 5:56 Practical implementation of heterogeneous workflows 6:55 Case study: Recursive language models and context management 9:05 Results on Ulong benchmarks (Cerebras/Sambanova performance) 10:20 Case study: Visual web navigation and Video Web Arena performance 12:02 Offloading subtasks to smaller models for speed and cost efficiency 12:38 The future of compute: Moving to a heterogeneous, multi-agent stack 13:10 Partnership with the UK's Arya institute 13:31 Closing summary and outlook on hardware/software co-evolution 14:01 Q&A: Automation l

Summary

Generated by claude-sonnet-4-5

30-second take

Adrian Bertagnoli from Callosum argues that AI's future lies in heterogeneous intelligence—orchestrating diverse models of different architectures and sizes on optimized hardware, rather than scaling single frontier models on homogeneous GPU clusters. He presents mathematical proof (the "principle of maximum heterogeneity") showing heterogeneous systems outperform homogeneous ones under constraints, then demonstrates it practically: their heterogeneous recursion approach on the Oolong benchmark is 7-12x cheaper and 3-5x faster than GPT-4o while matching performance, and their multi-model approach to visual web navigation beats GPT-4o and Gemini 2.5 by 18-25% while being 3.7x cheaper and 3x faster. They've secured a £3M grant to run the UK's first heterogeneous co-located cluster. This is a concrete pitch for why the next compute paradigm is vertical integration of specialized models/hardware, not just bigger monolithic models.

Key takes

  • Homogeneous scaling is training-era thinking: Neural scaling laws (more data/parameters = better) apply primarily to training; inference workloads decompose into diverse subtasks requiring different types of intelligence, making single-model approaches inefficient
  • Mathematical formalization exists: They proved via the "principle of maximum heterogeneity" (drawing on neuroscience, economics, ecology) that under reasonable constraints, systems with specialized agents communicating outperform generalist agents—the production function analogy shows generalists create short cylinders that don't meet demand peaks
  • Heterogeneous recursion cuts costs dramatically: Extending MIT's recursive language model concept (treating context as environment, not prompt) by mapping subcontexts to different models/chips yields 7x cheaper + 5x faster on Cerebras or 12x cheaper + 3x faster on SambaNova vs GPT-4o on Oolong benchmark, with equivalent performance
  • Visual web navigation shows Pareto shift: Mixing Qwen3-VL-8B and Kimi 2.5 or Qwen3 + GPT-4o beats state-of-the-art GPT-4o/Gemini 2.5 by 18-25% on Video WebArena while being 3.7x cheaper and 3x faster—offloading simple subtasks (zooming) to weak models is 43x cheaper than using ChatGPT
  • Three stages of heterogeneity: Mild (multi-agents on homogeneous clusters), moderate (different chips for different models, SSMs/diffusion mixing), full (co-evolution of hardware/software with vertical integration)
  • Automation layer now exists: Initially hand-tuned task-to-model mapping; now have a system that detects task complexity and auto-predicts optimal model + hardware

Useful details

  • Oolong benchmark specifics: GPT-4o (~$3.75/task, ~2000 seconds); Callosum on Cerebras ($0.54, ~400 seconds); on SambaNova (~$0.31, ~660 seconds)
  • Video WebArena results: 18% better than GPT-4o, 25% better than Gemini 2.5; Qwen3-VL-8B + Kimi 2.5 configuration is 1.3x faster and 18x cheaper than Kimi alone
  • Zooming subtask: 11x faster, 43x cheaper than ChatGPT when offloaded to simpler models
  • Recursive language model background: MIT October paper showing context rot even at low context window occupancy for linear/quadratic information complexity tasks (degrades to 30-60%); solved by treating context as file + Python REPL agent extracting subcontexts
  • Current mild heterogeneity examples: MoE replacing dense models (architecture), multi-agent replacing single LLM calls (workflow), prefill/decode disaggregation (hardware)
  • Grant: £3M from ARIA (UK) to operate first heterogeneous co-located cluster in UK
  • Compute eras: CPU (faster compute), GPU/NVIDIA (parallel compute), heterogeneous (mapped multi-agent workloads)
  • Company: Callosum, actively hiring

Caveats / counterpoints

  • No discussion of orchestration complexity: No acknowledgment of engineering overhead, latency from inter-model handoffs, debugging difficulty, or operational complexity of managing multiple models/hardware types
  • Benchmark selection bias: Only shows Oolong (recursive context) and Video WebArena (visual nav); no demonstration on reasoning-heavy tasks where frontier models excel, or tasks requiring consistent model capability
  • Automation layer is vague: The "detects task complexity and predicts best model" system is mentioned but not explained—no accuracy metrics, failure modes, or details on how it works
  • Mathematical proof is illustrative, not rigorous: The principle of maximum heterogeneity figure is an analogy (production functions, skill distributions) but no formal theorem statement, assumptions, or limitations given
  • Hardware dependencies: Results tied to specific chips (Cerebras, SambaNova); unclear how much is algorithmic vs. just using cheaper/faster inference services
  • No frontier model comparison on hard tasks: Doesn't show whether heterogeneous approach can match GPT-4o/Gemini on tasks requiring deep reasoning where model capability matters most
  • Grant is for infrastructure, not validation: £3M is for running a cluster, not proof the approach works at scale

Ken relevance

High relevance for agent orchestration architecture. This directly challenges the "just use Claude/GPT-4o everywhere" approach Ken might default to for agent systems. The concrete numbers (7-12x cost savings, 3-5x speed gains with equivalent performance) make a strong case for routing strategies in multi-agent systems. For Ken's agent ops:

  • Immediate tactical win: Implement task complexity detection + model routing for long-context or visual tasks to cut costs dramatically without quality loss
  • Architecture principle: Decompose problems into subtasks with different capability requirements rather than assuming frontier models everywhere
  • Investment angle: Callosum is early-stage (hiring, just got grant) but has mathematical framing + benchmarks; could be interesting early bet on post-scaling-law infrastructure
  • GTM lesson: If heterogeneous beats homogeneous with proof, Ken's content/advisory should shift from "scale one model" to "orchestrate specialized models"—counter-narrative to prevailing AI discourse
  • Operational friction: Need to balance savings against complexity of maintaining multiple model integrations, fallback logic, cost tracking across providers

Watch for Callosum's automation layer details—if they solve routing without manual tuning, that's a key unlock.

Watch verdict

Watch fully. This is a rare talk with (1) novel theoretical framing, (2) concrete benchmark wins with numbers, (3) clear architectural implications for Ken's agent work, and (4) a company actively building the infrastructure. The heterogeneous recursion + visual nav results are immediately actionable, and the broader thesis challenges the "bigger model = better" orthodoxy in a falsifiable way. Worth Ken's 15 minutes to absorb the argument and decide if he should test task routing in his own systems.

Transcript

1875 words en Processed in 309.2s

[SPEAKER_01] Thank you for coming to my talk. My name is Adrian Bertagnoli. I'm a founding engineer at Colossum and today I'm going to be talking about scaling the next paradigm of heterogeneous intelligence. So I'm going to start with explaining why we care about heterogeneity in the first place, what particular aspects make it very conducive for scaling AI, how it is actually used in practice today, and how we can utilize it in the future to actually scale the next paradigm of intelligence. So to give you an intuition about what I mean with heterogeneous intelligence, I want to take a step back and explain the current prevailing paradigm of homogeneous intelligence. So homogeneous intelligence in terms of AI mainly refers to scaling single models on a fleet of identical chips. This era was largely brought about by the discovery of neural scaling laws, which showed us that more data and more parameters leads to better models. However, this is primarily rooted in a training domain. And while we move towards an inference domain, this becomes less and less relevant. So it's already changing, and we already see some level of heterogeneity in our current systems. So on the architecture level, we see that mixture of experts are replacing large dense models. On the workflow layer, we see that single LLM calls are being replaced by multi-agent systems. And finally, on the hardware level, single chips are being replaced by pre-fill decode disaggregated systems. So given that we are currently at the state of mild heterogeneity, how can you imagine a greater level of heterogeneity? How will that appear? So initially, what we are currently experiencing is mild heterogeneity. Everything is still running primarily on homogeneous clusters, but we have some variety in the prompts. When we run multi-agent systems, we might use different LLMs for different sub-agents. Again, we have a mixture of experts. When we increase the heterogeneity, we might start to use different chips for different models. So different LLMs might be put on different GPUs. They might be interacting. We might be using different models completely. So we have an increase of state-space models, diffusion models, all interacting with each other, all on optimal hardware that exists currently. And the last stage, where we really see the heterogeneous paradigm unfolding, is when we have a co-evolution of systems, hardware, and software. So there will be a unification where you will have a complete vertical integration of intelligence and hardware. So why heterogeneity? Why is it a good thing in the first place? Real-world problems are complex, multi-step, and open-ended. They decompose into sub-problems which require vastly different types of intelligences. So scaling a singular type of intelligence to solve these is very inefficient and not optimal. So how do we solve them? Solving these actually requires models of different architectures and sizes working together, acting together in long horizons, something we call multi-agent heterogeneous intelligence. Furthermore, new generations of silicon are coming to the market. But currently, there's no interface which allows this new hardware to be unified and constructively help the current compute stack. And so this is what we aim to change. So heterogeneity, the benefit, is not simply a belief that we have. We actually formalized it and proved it mathematically. On the left, you see a figure outlining the principle of maximum heterogeneity. So these are heterogeneous agents where the color indicates a distribution over a skill space. If you have communication between these, indicated here by a ring topology, you can have what we call a production function. And the production function is simply the demand can be well suited for the demand of one problem, but ill suited for another problem. So here we have a production function that's well suited for demand A and ill suited for demand B. If you want to do this in a homogeneous fashion, you would either be able to only scale one peak or in the optimal case to match this demand function, you would have only generalists, with as broad as possible a skill set. But then ultimately, you'd have a very short cylinder that does not meet the production function readily. So we formalized this and we saw that across many domains, including neuroscience, economics, and ecology, these trends hold. And under any reasonable amount of constraints, heterogeneous systems outperform homogeneous ones. So how do we use this in practice? I've been telling you about the benefits of heterogeneity, but I've not told you anything about what it actually means in terms of AI. So we optimize multi-agent systems at three different parts of the workflow. From the hardware where agents run, we choose different hardware depending on the computational demands of the agents. And then how agents interact and what workflow they construct. So we have already demonstrated multiple benefits of this type of orchestration. And I want to go into a couple ones, namely in the workflow, something a primitive we call heterogeneous recursion, and in the agent layer, I want to talk about multimodal video action language models. So heterogeneous recursion. Has anyone here heard of recursive language models? Okay. So for those of you who don't know, recursive language models is a seminal paper that came out of MIT in October. And they basically showed that even if you only occupy a small percentage of the context window, you still can have dramatic context rot, depending on the information complexity you want from the prompt. So if you're doing a needle in a haystack task, that is all of one. The information requirement scales constantly throughout, regardless of how big the prompt is. And then you can imagine adding up the rows. You give rows and columns. Adding up the rows would be O of N, because as the prompt increases, the informational requirement increases linearly. So if you have a constant information requirement, it scales well. You can occupy the full context window and actually get a good answer. However, when you go to linear or quadratic, it degrades at around 60 to 30%. So recursive language models solve this problem by actually treating the context as an environment rather than putting it all into the prompt. So in practice, this looks like you present the context in a file and then a coding agent interacts with it programmatically through Python REPL, basically doing keyword searches, regex, and other tricks to extract subcontexts. And this subcontext is then passed off to an identical recursive agent. So this agent then can answer the question or spawn another recursive agent. And that's why it's called recursive language model. So we simply extended this concept. Instead of using a single model on a single chip, we map based on the subcontext generated towards different chips and different models to emulate the performance while drastically being cheaper and faster. So here are results. You can see this is on the Oolong benchmark. This is basically the benchmark they used in the paper. And GPT 5.2 was the most recent one where we produced this work. It sits around here, where it takes around 2,000 seconds to run through the benchmark. And it costs around $3.75 for one task. Our system, when we go on Cerebras, we are seven times cheaper and five times faster. So you save an incredible amount of time and it's much cheaper. So it's basically having your cake and eating it too. With Samanova, we even get further. We push the price down even further at the cost of some latency. So we're 12 times cheaper and three times faster. So these are architectural decisions that are not simply based on the hardware. You can make huge, impactful price differences and emulate the intelligence you would have from frontier models. So the next problem we wanted to address is basically visual web navigation. So we used a mixture of open and closed video action language models. And we managed to beat the state of the art of video web arena, beating GPT 5.2 and Gemini 2.5 by 18 and 25% respectively. And not only this, the way we did it is instead of treating the problem as a homogeneous one, we acknowledge that the problem is heterogeneous itself. It decomposes into multiple steps of visual reasoning and textual reasoning. And each of these subcomponents requires different models to be completed successfully. So here you see a fundamental shift of the Pareto frontier where singular models like Kimi 2.5 and GPT 5.2 are outperformed by a mixture of heterogeneous models. So when we use QEN3, VL8B Instruct and Kimi 2.5, we're 1.3 times faster than using Kimi alone. We're 18 times cheaper than using GPT 5.2 alone. And if we use QEN3 plus GPT, we're actually three times faster and 3.7 times cheaper. So this is only benefit. There's no downside in constructing this in a heterogeneous manner. So one part of our differentiating factor, how we were able to beat the state of the art, is that we mapped certain subtasks like zooming and created a different visual reasoning for the agent. We offloaded that into less intelligent models because you don't need GPT to zoom for you. So alone on these subtasks, we're able to be 11 times faster and 43 times cheaper than using ChatGPT. And so this is what overall accumulates towards these 3.7 times cheaper and 3 times faster. So, looking ahead, how do we view the future of compute? The first era of scaling compute was dominated by the CPU, where compute got quicker. The second era was making compute massively parallel. This is dominated by NVIDIA. And the third paradigm, compute is going to become heterogeneous, mapping multi-agentic workloads and optimally mapping these workloads onto different chips. We are actually working with ARIA, the UK Institute. We got a three million grant for operating the first heterogeneous co-located cluster in the UK. So we really want to make a difference and spearhead this new era of innovation. So, the era of homogeneous scale delivered extraordinary progress. We should be grateful for it. What comes next is heterogeneous intelligence, where models, workflows and silicon co-evolve and every new source of diversity makes the whole system smarter, faster and cheaper. This is the worst our infrastructure will ever be. Thank you. How do you define which task to run on the faster, cheaper model? For instance, is zooming something that's hard for the technology after zooming? So, initially we started doing bespoke decisions on mapping certain simple sub-tasks to simple models. But since then, we have created an automation layer that detects the task complexity and automatically predicts the best model and the best suited hardware. Any other questions? Great. Thank you so much for your attention. My name is Adrian Bertignoli and if anyone is interested, we are hiring. So, yeah. Great. Thank you very much. Thank you. we move towards an inference domain, this becomes less and less relevant. So it's already changing, and we already see some level of heterogeneity in going into our current systems. So on the architecture level, we see that mixture of experts are replacing large dense models. On the workflow layer, we see that single LLM calls are being replaced by multi-agent systems. And finally, on the hardware level, single chips are being replaced by pre-fill decode disaggregated systems. So given that we are currently at the state of mild heterogeneity, how can you imagine a greater level of heterogeneity? How will that appear? So initially, we'll be what we currently are experiencing is mild heterogeneity. So everything is still running primarily on on homogeneous clusters. But we have some variety in the prompts. When we run multi-agent systems, we might use different LLMs for different sub-agents. Again, we have a mixture of experts. When we increase the heterogeneity, we might start to use different chips for different models. So different LLMs might be put on different GPUs. They might be interacting. We might be using different models completely. So we have an increase of state-space models, diffusion models, all interacting with each other, all on optimal hardware that exists currently. And the last stage, where we really see the heterogeneous paradigm unfolding, is when we have a co-evolution of systems, hardware, and software. So there will be a unification where you will have a complete vertical integration of intelligence and hardware. So why heterogeneity? Why is it a good thing in the first place? So real-world problems are complex, multi-step, and open-ended. They decompose into sub-problems which require vastly different types of intelligences. So scaling a singular type of intelligence to solve these is very inefficient and and not optimal. So how do we solve them? Solving these actually requires models of different architectures and sizes working together, acting together in long horizons, something we like to call multi-agent heterogeneous intelligence. Furthermore, new generations of silicon is coming towards the market. But currently, there's no interface which allows this new hardware to be unified and constructively help the current compute stack. And so this is what we aim to change. So heterogeneity, the benefit, is not simply a belief that we have. We actually formalized it and proved it mathematically. On the right, on the left, you see a figure outlining the principle of maximum heterogeneity. So these are heterogeneous agents where the color indicates a distribution over a skill space. If you take, if you have a communication between these, here indicated by a ring topology, you can have what we like to call a production function. And the production function is simply the demand can be well suited for the demand of one problem, but ill suited for another problem. So here we have a production function that's well suited for demand A and ill suited for demand B. If you want to do this in a homogeneous fashion, you would either be able to only scale one peak or in the optimal case to match this demand function, you would have only generalists, so as broad as possible, the skill set. But then ultimately, you'd have a very short cylinder that does not meet the production function readily. So we formalized this and we saw that across many domains, including neuroscience, economics, and ecology, these trends hold. And under any reasonable amount of constraints, heterogeneous systems outperform homogeneous ones. So how do we use this in practice? Like I've been telling you about the benefits of heterogeneity, but I've not told you anything about what it actually means in terms of AI. So we optimize multi-agent systems at three different parts of the workflow. So all the way from the hardware, where agents run on, we choose different hardware depending on the computational demands on the agents. And then how agents interact and what workflow they construct. So we have already demonstrated multiple benefits of this type of orchestration. And I want to go into a couple ones, namely, in the workflow, something a primitive we like to call heterogeneous recursion. And in the agent layer, I want to talk about multimodal video action language models. So heterogeneous recursion. This is something who's here heard of recursive language models. Okay. So for those of you who don't know recursive language model, it's kind of a seminal paper that came out of MIT in the last October. And they basically showed that even if you only occupy a small percentage of the context window, you still can have dramatic context rot, depending on the information complexity you want from the prompt. So if you're doing a needle in a haystack task, that is all of one. The information requirement scales is constant throughout, regardless of how big the prompt is. And then you can imagine adding up the rows. You give rows and columns. Adding up the rows would be O of N. Because as the prompt increases, the informational requirement increases linearly. So if you have a constant information requirement, it scales well. You can occupy the full context window and actually get a good answer. However, when you go to linear or quadratic, it degrades at around 60 to 30%. So recursive language models solve this problem by actually treating the context as an environment rather than putting it all into the prompt. So in practice, this looks like you present the context in a file. And then a coding agent interacts with it programmatically through Python REPL, basically doing keyword searches, regex, and other tricks to extract subcontexts. And this subcontext is then passed off to an identical recursive agent. So this agent then can answer the question or spawn another recursive agent. And that's why it's called recursive language model. So we simply extended this concept. Instead of using a single model on a single chip, we map based on the subcontext generated towards different chips and different models to emulate the performance while drastically being cheaper and faster. So here are results. You can see this is on the Oolong benchmark. This is basically the benchmark they used in the paper. And GPT 5.2 was the most recent one where we produced this work. It sits around here, where it takes around 2,000 seconds to run through the benchmark. And it costs around $3.75 for one task. Our system, when we go on Cerebras, we are seven times cheaper and five times faster. So you save incredibly much time or a lot cheaper. So it's basically like having your cake and eating it too. With Samanova, we even get further, we push the price down even further at the cost of some latency. So we're 12 times cheaper and three times faster. So these are like making architectural decisions that are not like simply based on the hardware. You can make huge, impactful price differences and emulating the intelligence you would have from frontier models. So the next problem we wanted to address is basically visual web navigation. So we used a mixture of open and closed video action language models. And we managed to beat the state of the art of video web arena, beating GPT 5.2 and Gemini 2.5 by 18 and 25% respectively. And not only this, the way we did it is instead of treating the problem as a homogenous one, we acknowledge that the problem is heterogeneous itself. It decomposes into multiple steps of visual reasoning, of textual reasoning. And each of these subcomponents requires different models to be completed successfully. So here you see a fundamental shift of the Pareto frontier where you see singular models like KimiK 2.5 and GPT 5.2 are outperformed by a mixture of heterogeneous set of models. So when we use QEN3, VL8B Instruct and KimiK 2.5, we're 1.3 times faster than using Kimi alone. We're 18 times cheaper than using GPT 5.2 alone. And if we use QEN3 plus GPT, we're actually three times faster and 3.7 times cheaper. So this is only benefit. There's no downside in constructing this in a heterogeneous manner. So one part of our differentiating factor, how we were able to beat the state of the art, is that we mapped certain subtasks like zooming and creating a different visual reasoning for the agent. We offloaded that into less intelligent models because you don't need GPT to zoom for you. So alone on these subtasks, we're able to be 11 times faster and 43 times cheaper than using ChatGPT. And so this is what overall accumulates towards these 3.7 times cheaper and 3 times faster. So, looking ahead, how do we view the future of compute? The first era of scaling compute was dominated by the CPU, where compute got quicker. The second era was making compute massively parallel. This is dominated by NVIDIA. And the third paradigm, compute is going to become heterogeneous, mapping multi-agentic workloads and optimally mapping these workloads onto different chips. We are actually working with ARIA, the UK Institute. We got a three million grant for the first, for operating the first heterogeneous co-located cluster in the UK. So we really want to make a difference and spearhead this new era of innovation. So, the era of homogeneous scale delivered extraordinary progress. We should be grateful for it. What comes next is heterogeneous intelligence, where models, workflows and silicon co-evolve and every new source of diversity makes the whole system smarter, faster and cheaper. This is the worst our infrastructure will ever be. Thank you. How do you define which task to run on the faster, cheaper model? Like for instance, a Zoom, is that something that's hard for the technology after Zoom? So, initially we started doing bespoke decisions on mapping certain simple sub-tasks to simple models. But since then, we have created an automation layer that detects the task complexity and automatically predicts the best model, the best suited model and hardware. Any other questions? Great. Thank you so much for your attention. My name is Adrian Bertignoli and if anyone is interested, we are hiring. So, yeah. Great. Thank you very much. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you.