The State of Model Routing — NVIDIA, Cognition, OpenRouter
Description
Run terminal bench on Opus and on Haiku and Opus scores about three times better at a tenth of the cost, even though Haiku is far cheaper per token. Alex Atallah's point is that a small model pushed outside its training distribution thrashes, calling tools in loops until it costs more than the expensive model ever would. That inverts the obvious version of model routing, where you send each task to whichever model benchmarks best on it. Walden Yan calls that approach fragile for exactly the reason agents make it worse: a session starts as a question about a codebase, becomes a feature request, then becomes live debugging, and the model you picked at the start is stranded. Cognition's answer keeps a frontier model planning and delegates the implementation, which cut the cost of Fable level intelligence by 40% while going deeper, because a cheaper model can afford to spin off three sub agents to explore a codebase. They also avoid sub agents in favor of one sidekick with a continuous running context, so the KV cache stays warm and cached tokens cost roughly ten times less. Compaction, Yan argues, is worth doing for intelligence rather than cost, since compacting forces a cache miss and model quality falls off a cliff well before the advertised million token window. The most telling story is OpenRouter's: its auto router sat almost unused for two years until openclaw began sending heartbeats every ten minutes, creating one popular app with two completely different intelligence needs. Speaker info: Nader Khalil, moderator (NVIDIA): - https://x.com/naderlikeladder - https://nader.coffee Walden Yan (Cognition): - https://x.com/walden_yan - https://www.linkedin.com/in/waldenyan Alex Atallah (OpenRouter): - https://x.com/alexatallah - https://openrouter.ai Tanay Varshney (NVIDIA): - https://www.linkedin.com/in/tanayvarshney Carter Abdallah (NVIDIA): - https://x.com/Baxate - https://www.linkedin.com/in/carter-abdallah Timestamps: 0:00 - Welcome and the multimodel prem
Summary
Generated by gpt-5.6-terraAt-a-Glance
- Verdict: Watch fully
- Core thesis: Effective model routing is not simple task-to-cheapest-model selection; it is a context-aware orchestration system in which frontier models plan, supervise, and intervene while specialized or smaller models execute bounded work.
- Why it matters: For agentic systems, token price is a poor proxy for completed-task cost: weak models can thrash, overuse tools, lose context, and cost more than stronger models, while coordinated multi-model systems can improve both quality and economics.
- Best use: Use this as a strategic and architectural briefing for designing an agent control plane: routing policies, escalation criteria, context transfer, caching, local/cloud placement, and production-learning loops.
Executive Summary
The panel argues that production AI is inherently multi-model, particularly for agents, local inference, and long-running software workflows. Its central correction to simplistic routing is that models have jagged, complementary capabilities: benchmark leadership does not mean universal superiority, and a cheaper model is only cheaper if it reliably completes its assigned work.
Cognition describes Devin Fusion as a supervisory pattern rather than a pure classifier/router. A frontier model retains planning and oversight while a cheaper implementation model performs delegated work, potentially with more token budget and parallel exploration. Cognition says this reduces the cost of “frontier-level intelligence” by 40%, without claiming that its system exceeds frontier-model capability outright.
The hard engineering problem is not merely choosing a model at request start. Agent tasks evolve through sessions and subtasks, so a model that was suitable for repository exploration may be unfit for implementation, testing, or debugging. Systems need mechanisms for detecting out-of-distribution work, compacting and selectively transferring context, preserving cache value, and escalating before smaller models enter unproductive loops.
The speakers view routing as an immature but increasingly central control-plane layer. It will likely be co-designed with model training, prompts, agent harnesses, caching infrastructure, and deployment topology. OpenRouter’s experience suggests adoption rises when agent workloads contain clearly distinct intelligence tiers—such as cheap heartbeats versus high-value work—and when routing can be invoked through simple developer primitives.
Key Takeaways
- Claim: Route by expected cost per successfully completed task, not token price or broad task labels. | Evidence: OpenRouter reports that on Terminal Bench, Opus performed roughly three times better than Haiku at one-tenth the cost, despite Haiku being substantially cheaper per token; the stated failure mode is that undersized models make excessive tool calls and enter wasteful loops on out-of-domain work. | Implication: Ken should make routing evaluation outcome-based: success rate, retries, tool calls, elapsed time, and total inference cost—not model unit pricing or generic coding benchmark scores. | Caveat: For highly in-distribution, bounded work—such as classifying whether text is a person name or organization name—a small model remains the rational choice.
- Claim: The robust multi-model pattern is frontier supervision plus delegated execution, rather than permanently routing a task to a weaker model. | Evidence: Cognition says Devin Fusion leaves the frontier model responsible for planning and hard decisions while delegating implementation to cheaper/open models; its stated result is a 40% reduction in the cost of frontier-level intelligence. Cheaper delegates can receive far more tokens or launch multiple exploratory agents within the same budget. | Implication: Design agent systems with a persistent high-capability governor that can inspect progress, redirect work, or reclaim execution, rather than relying on a one-time route at the prompt boundary. | Caveat: Cognition explicitly does not claim the fusion system exceeds the absolute capability of the frontier model; the benefit is a better quality-cost trade-off through delegation and coverage.
- Claim: Routing must operate over task phases and sessions because agentic work changes in complexity over time. | Evidence: The panel’s software-engineering example moves from understanding a codebase to implementing features to live testing and debugging edge cases. Cognition calls initial task-type routing “extremely fragile” for these agentic trajectories. | Implication: A router should continuously reassess task state and competence rather than infer a fixed model choice from the user’s initial request.
- Claim: Context transfer and caching are first-class routing constraints, not lower-level implementation details. | Evidence: Cognition warns that indiscriminately sharing a file read with several models can multiply input cost. Its proposed approach keeps most raw context with one ongoing “sidekick” agent, then provides the frontier supervisor compacted state, file references, high-level activity, or selectively retrieved evidence. The sidekick preserves a running context/KV cache rather than behaving as repeatedly reset subagents. | Implication: Treat context ownership, artifact references, cache affinity, and handoff summaries as explicit control-plane primitives; do not broadcast full traces across model calls. | Caveat: Compaction is lossy and can itself force a cache miss; one speaker argues it should primarily be used to preserve intelligence when a cache miss or model handoff is already necessary, not assumed to be a standalone cost optimization.
- Claim: The router’s supervisory model may need to be the larger model, but the correct outer/inner arrangement is workload-specific and empirical. | Evidence: OpenRouter found its best published results for deep research when the smart model was the outer wrapper/orchestrator, while stating that coding may favor a different arrangement. Cognition is experimenting with training a model either as orchestrator or as execution-side “sidekick.” | Implication: Maintain per-workload routing experiments rather than standardizing prematurely on “large orchestrator, small workers” or its inverse. | Caveat: The panel repeatedly characterizes fusion research as early, with limited detailed and historically mixed research results; no universal architecture is established.
- Claim: Production routing should learn from observed correction behavior, not just offline benchmark datasets. | Evidence: Cognition proposes using signals such as users manually upgrading or downgrading models, automatic re-routing after an initial wrong choice, traces of agent failures, and regression tests after prompt changes. It favors asking a capable model to analyze why a routing decision failed over older mechanical gradient-descent-style prompt tuning frameworks. | Implication: Instrument every route, escalation, override, failure recovery, and post-task outcome; use those traces to build a continuous router/prompt improvement loop.
- Claim: Local/cloud routing broadens the problem from cost optimization to privacy, hardware utilization, and workload-specific economics. | Evidence: NVIDIA frames routing as a choice of when sensitive prompts should stay on-device, when they can be anonymized and sent to frontier cloud models, and how to use underutilized local hardware for workloads such as agent heartbeats. The panel also notes that self-hosting permits control over cache duration and optimization for a known workload shape, unlike API pricing that amortizes general demand. | Implication: Include data sensitivity, available local GPU capacity, cache residency, context profile, and cloud marginal cost in a single placement policy rather than treating local inference as a separate stack. | Caveat: Self-hosted inference has its own context-length economics: throughput can deteriorate as context grows, so local deployment does not eliminate the need for compaction and routing.
Detailed Brief
Competence detection and escalation signals
- Claims: A smaller model cannot always reliably recognize that it is out of its depth; a stronger model sometimes must perform the detection.; Scheduled cache refreshes can double as low-incremental-cost supervisory checkpoints, where a frontier model reviews whether the executing model is looping or needs assistance.; Model internal-state probes may provide additional indicators of likely hallucination or confusion beyond token count.
- Evidence: Cognition describes using the inevitable cache-refresh cadence—commonly around five minutes at providers—to ask a frontier model to inspect the smaller model’s trajectory.; NVIDIA describes hallucination probes, including analyses or classifiers over internal/pre-fill-state representations, as proxies for how lost or uncertain a model may be.
- Caveats: The speakers do not present a validated universal escalation metric; they specifically say many routing choices require empirical evaluation.; Long output traces are not necessarily evidence of failure because smaller models may be trained on verbose traces; token volume alone is an incomplete signal.
- Implications: Build escalation as a layered policy: deterministic budget and loop signals, task-state checks, periodic supervisor review, and—where the serving stack permits it—model-state quality probes.; Keep provider-specific cache expiry behavior visible to the control plane because it creates both a cost event and an opportunity to reassess routing.
Market direction: routers as control-plane infrastructure
- Claims: OpenRouter says it had an auto-router for almost two years with little adoption until agent workloads made heterogeneous intelligence needs obvious.; OpenClaw adoption made routing salient because periodic heartbeats created a low-value but persistent token workload that should not use an expensive default model.; The likely end state is not a standalone router divorced from models and harnesses, but coordinated optimization across model training, orchestration, prompts, caches, and infrastructure.
- Evidence: OpenRouter cites its routing offerings: an auto router, Pareto Code for tunable Pareto-optimal coding-model selection, and Fusion for multi-model orchestration.; The panel notes that OpenRouter’s public usage data still shows Opus as the top model by dollars spent for classification tasks, illustrating how much inefficient model selection remains.
- Caveats: A future highly capable model might theoretically be more efficient at nearly every task than a smaller alternative, reducing some routing arbitrage.; Even under that scenario, cache locality and uneven context distribution create reasons for models to collaborate.
- Implications: Avoid building routing as a static procurement rule or a thin model-name switch. It should be an observable, policy-driven orchestration layer with enough authority to arbitrate model behavior.; Prioritize routing first in agent systems with recurring low-value traffic, variable task depth, expensive tool use, or heterogeneous local/cloud capacity.
Notable Concepts & Terms
- Model fusion: The use of multiple models as a coordinated system to obtain a better quality-cost outcome than assigning an entire workload to one model.
- Frontier supervisor / outer model: A high-capability model that plans, monitors delegated work, and decides when to escalate or redirect execution.
- Sidekick: Cognition’s term for a persistent delegated execution agent with a continuing context, designed to retain KV-cache value rather than restarting as a fresh subagent.
- Jagged capabilities: The idea that model ability is uneven across narrow domains and subskills; a model that wins a broad benchmark may lose on a specific library, task type, or distribution.
- In-distribution versus out-of-distribution routing: Routing based on whether a task resembles a model’s learned domain; small models can be economical in-domain but can become costly and unreliable outside it.
- Context compaction: Compressing an agent’s working history before a handoff or long-running continuation; useful for maintaining usable context but inherently lossy and potentially cache-disruptive.
- KV cache / prefix cache: Stored inference state that makes repeated processing of the same prefix far cheaper; cache locality can determine whether a nominally worse model route is economically preferable.
- FlexRunt: NVIDIA’s described approach of distilling a main model into smaller footprints and selecting which model or weight section performs decoding based on task complexity.
Operator Notes / Why Ken Should Care
- Define a routing scorecard at the task level: successful completion, total tokens, cache-hit rate, tool-call count, loop/retry count, latency, and human/model override rate.
- Implement a frontier-supervisor pattern for consequential or open-ended workflows; restrict small models to bounded subtasks with explicit return artifacts and escalation paths.
- Create a context-handoff contract: retain canonical work in files or durable artifacts, pass references and targeted summaries across models, and avoid full-trace fan-out.
- Run controlled experiments for outer-model selection separately across coding, research, classification, and operational heartbeat workloads; do not generalize results between them.
- Log user model upgrades/downgrades and automatic reroutes as labeled training data for router and prompt-policy improvement.
- Add workload placement policy that jointly considers sensitivity/privacy, local GPU utilization, context length, cache expiry, and cloud cost before selecting local versus API inference.
- Treat long context as a quality risk as well as a throughput cost; establish compaction/retrieval thresholds rather than relying on advertised maximum context windows.
Source/Metadata
- Title: The State of Model Routing — NVIDIA, Cognition, OpenRouter
- Transcript words: 10702
- Duration seconds: 2897
- Timestamp note: No usable timestamps or chapters were present in the supplied transcript. The transcript contains substantial duplicated passages near the end.
Transcript
Music We've tried to get a bunch of the industry leaders together to talk about some of the problems that we're facing as we try to run more on local. If you guys were here for the first panel, one of the things that we talked about was model routing. We firmly believe that we're in a multi-model world. I think you heard this from many of the panelists. Anyone who is deploying AI in production and who is doing so locally is seeing that multi-model world. That's why we released these Nemo Tron models at NVIDIA. Everything is released from the data sets to the weights, with recipes so that you can customize them. We do that because we know that people customizing models is going to be huge. And so this panel is really exciting because we're going to talk specifically about model routing. So as you are picking which model to use, how does that tooling itself look? Do you guys want to introduce yourselves? Yeah, sure. I'm Walden. I'm the co-founder of Cognition. We build Devon, AI software engineer. In addition to the product, we spend a lot of time partnering with our customers to figure out how they should deploy these models and these agents. And one of the things they're constantly asking us nowadays is how do I know the ROI of our models, and how do I know which tasks I can actually let our engineers spend the most expensive models on versus letting them use a more cost-efficient model? And so that's why we're also thinking a lot more about multi-model routing nowadays. Totally. I'm Carter. You guys heard from me a little bit earlier, but if you weren't here, I'm a developer tech engineer at NVIDIA. And ultimately, I spend a lot of time thinking about how to get intelligence into as many developers' hands as possible. And something that is continually becoming not an issue, but something that is top of mind for a lot of developers, is as you use more intelligence and the frontier models get more expensive, it becomes somewhat cost-prohibitive to use the best tools, what feels like the best tools, as much as you would like to use them. And so this has become a recent focus: how can we still get the same desired outputs, but actually both as an individual developer, but also imagine startups and small companies, how can you leverage this incredible tool without totally breaking the bank? I'm Taneh. I work on model evaluations, both in terms of its accuracies and efficiency and cost understanding of the model. And then I try and understand those, implement those learnings, and help build a router. So it's, my job is to understand the behavior of the model on an intimate level and then use those learnings to both improve the model and try and design a system of model that can work together with each other. Totally. Yeah, I love a lot of the research that you're doing at NVIDIA as we see the space through. I think what's really interesting is model routing itself is pretty new still. And so what you'll notice is there isn't a very clear solution here. That was something that came up on the first panel, is that there is a lot of space for startups and for companies in the ecosystem to fill in a solution here because we're still figuring out how to best do these patterns. And I think, Walden, I want to ask you. So Cognition just released Fusion, your guys' model router. Yeah. And when you guys released it, your blog said that you're actually getting better performance than Fable, than these frontier models. And I feel like that was a very surprising statement to hear because we're thinking that you're getting as good or close enough, usually when we're running on edge, when we're running local, and these compute-strain smaller footprint models. But you guys are getting better. Can you explain how? Yeah, absolutely. So I also want to be clear about something here. We're not saying that we gap above Fable-level performance. Totally. In the same way that maybe Fable-level performance gaps above other models. I think, actually, there's this really unintuitive dynamic where smarter models actually get better and better at delegating work. And so one of the philosophies we had with building a model router is we don't want to route people to a dumber model, and then suddenly you're stuck with a model that doesn't know how to do your task, and the next thing you know you're switching yourself back to a smarter model anyway and now taking that expensive cost. In general, we think a lot of the existing model routing systems out there are probably the same ones people have been using a year ago. And so we really wanted to put out a new framework that actually lets people still feel like, and still have, a frontier model in their system while getting all these cost benefits. So, yeah, we're reducing the cost of Fable-level intelligence by 40%. The way we do that is we allow Fable to still do the planning and the hard decision-making, but delegate a lot of the work to an implementation model. And the implementation model can be one of these open source models, be it a cheaper mini model. The unintuitive thing is even though it's cheaper, because you're delegating the work to another model, you can let that model go at the task with much more depth and intensity than you might otherwise. You can spin off three sub-agents to go and explore the code base, and maybe that's actually more comprehensive than if you had just let Fable explore the code base itself. So you're actually getting this nice trade-off where it's both more cost-efficient and also more comprehensive overall. Interesting. I see. So you're saying by using a bunch of smaller models, you're essentially, for one example, scouring the code base. You can explore it potentially better than if you were to just have one model, I don't know, figure out with its limited context, with whatever path it's on. Yeah, totally. But also if you think about the budget, if you were to say the frontier model costs this amount per token and the smaller model is this amount per token and it's significantly cheaper, then you can use a lot more tokens from the smaller model, still within the budget, than it would have been from the frontier model. I would also like to encourage everyone to think there are jagged capabilities in most models, right? So coding is not one domain. Within, let's say, data visualization, there'll be scikit-learn, there'll be matplotlib, there'll be something else. It largely comes down to the training corpora that went into each of the models, right? So one model, while you're trying to do X type of work, let's say data visualization, and the other type is Y, that means, let's say, model building. Let's say you're trying to have a data science work stream, you're trying to optimize for some kind of prediction and then visualizing your results. Within that task, different models will have different strengths. So it's not necessary that model A, if it scores higher on a coding benchmark, is just plain better at every task there is. So routing is a task of intimately understanding the behavior and strengths and weaknesses of different models and then applying them thusly. I would encourage everyone to think, hey, models are strong at different things rather than there's one model to rule them all. I see. And by the way, real quick, thank you, Alex, for joining. Sorry, I'm late. No. Oh, is that still? I might need yours. Sorry, I'm late. I'm Alex from Open Router. Thanks for having me, Nani. Yeah, of course. Thank you so much. He came right from the airport. So this is perfect. I think, Taneh, that's super interesting. So the way that you're thinking through model routing, it's not even just delegating to necessarily a smaller model, but maybe this is what you're saying, can you put essentially a swarm of agents to accomplish the same task and suddenly routing the task between them is a problem to solve in and of itself. Yeah. So if you look at, let's take an easy example. Let's take a science or scientific discovery as an example. Right? Usually these are one-shot problems. It's incredibly hard. You have models thing throughout this process, right? So in that, you have tons of subdomains, like tons and tons and tons. So in that aspect, if you think about post-training, like the post-training process of a model, they'd be tuned with different teachers. They'd be tuned on different sub-tasks. So those overlapping strengths will be readily apparent when you're trying to understand failures of each model and different sub-tasks. Once you understand that, you can orchestrate your system to leverage that arbitrage, essentially. And that essentially becomes free. So I think this is on LM RAR bench. There are tons of benchmarks out there. But if you use these techniques, you can get up to 10% higher accuracy even. It depends on the model pool. It depends on the task at hand. But I would encourage to think about the complementary nature of models. I see. Do you see... So in the way that you were describing the way the task is broken up, So in that aspect, if you think about post-training, the post-training process of a model, they'd be tuned with different teachers. They'd be tuned on different sub-tasks. So those overlapping strengths will be readily apparent when you're trying to understand failures of each model and different sub-tasks. Once you understand that, you can orchestrate your system to leverage that arbitrage, essentially. And that essentially becomes free. So I think this is on LM RAR bench. There are tons of benchmarks out there. But if you use these techniques, you can get up to 10% higher accuracy even. It depends on the model pool. It depends on the task at hand. But I would encourage to think about the complementary nature of models. I see. Do you see... So in the way that you were describing the way the task is broken up, do you see that some of the smaller models, because the token cost is cheaper, are they using more tokens? Is it... Are you specifically routing so that they do... So that they are shattier? Oh, yeah. They absolutely do use more tokens. I actually want to riff on something that Tamei was saying, which is, a lot of times when you look at these different benchmarks, you'll see that the small models will perform better than even the frontier models in certain cases. I think a lot of people, they look at this and they immediately jump to, oh, how can we just route the task where the smaller models do better straight to the smaller models? I think that one of the themes we really want to emphasize with our recent blog post and recent DevInfusion was that this naive initial routing based on the task type is extremely fragile, especially the more agentic the task you work on is. So, for example, a real developer, you might ask your agent first, oh, how does this code base work? And then you go deeper and, okay, actually, can you implement some features for me? And then you go deeper and it's, oh, can you now go do a live test of this feature and debug deep cases? The complexity changes and the type of task changes over time, and you don't want to be left with some subpar model for the task that you're now on. I think this is why people like frontier models so much, is they're generally intelligent, and they're capable of shifting between various different domains, even if you can eke out better performance in very specific tasks. And so the challenge is, how do you get a small model to know that it's out of its depth and you need to now go switch to another model or go to a smarter model? And our solution to this is you just always have this main frontier agent that's watching, even if it's not the one doing the work. It should at least be keeping tabs and figuring out, okay, wait, the agent I delegated to now is out of its depth. I need to move it to something else. And overall, just the guarantee of always having frontier intelligence present, I think, reduces the fragility of these systems quite a lot. How does the sharing of context between one of those smaller agents who has basically completed up to some level of a task and decides, actually, I don't think I'm the right person for this. I need to hand it back to the foundational model. Of course, you don't want to have the entire trace of that smaller agent be passed back to the larger model. So how do you get that level of specificity while basically providing the information it needs but not more? Yeah, absolutely. So I think the context here is it's actually very easy to create a system that's more expensive as soon as you're running multiple models together because, oh no, there's one file read. Now every one of these models is now reading this one file read, and you're usually charged three times as much. The trick that we spend a lot of our time on is most of the context by default will only be going to one model. Most of the context, let's say, will be going to the small model. But the thing you need to then tune very well is, okay, maybe you still show what files it's reading. Maybe you show the high-level thinking of what it's doing back to the main model. Maybe you have the small model. You tune its ability to present the context back to the main model. And actually, a lot of these problems have already been well-studied in many domains already, like context compaction is something you already have to solve if you want to do really long-running agents. And so this problem of taking long context, compacting it in a way that is now understandable, it's one that you can also apply to this domain and just give the compacted context back to the main agent. Context compaction is something that I'm familiar with, but I hadn't really thought about. As you're doing model routing and as you're trying to share context across now potentially many models, you're expanding the amount of what could be seen as wasteful tokens or redundant tokens just because you have to process that across the many models. Yeah. Yeah. I think there's... The way I describe it is I think we are early in the model routing domain. I hope that a year from now, even the techniques we use for DevFusion, you can look back on that and are like, oh, these are some really legacy ideas, and now we have much better methods at routing between models. And when people actually start co-designing their models with this in mind, we're going to be in a much better world. Yeah, I echo what you say, right? I think routing will evolve as the task evolves when you start the task, right? So it's more useful to see things in terms of subtasks and sessions than individual problems that you're trying to solve because more than likely, when you're working through a problem, you're asking a lot of questions, you're exploring different things, and it is imperative that people who design routers try and understand these phases of different complexities and then try and apply some logic for essentially side-kicking tasks or leveraging expertise from other models. That's pretty on point. Yeah, I'd love to hear from the router guy. Yeah, I think these are important points, and one of the biggest debates I think we have internally is whether that outer model that's doing the orchestration should be the big model or the small model. Yeah. You get very different results depending on your choice, and it's not even clear what the pricing impact would be because if your outer model that's doing the orchestration is the big model, it can leverage its caching to make more of its decisions, and its caching is going to be a dramatic price savings compared to the small models caching a lot of the time, especially for issues that are on the bright line. Zooming out a little bit, I think what you want from all the models out there when you do model fusion is to benefit from all the data that is being trained on across all the labs and not just the data from one lab or one source. And a model is just a combination of the data and its understanding of the data, both its compute and the quality of its RL. So long-term, I think you want models where they know that, oh, this is in distribution, this is in my data. You can use small models pretty easily and get a cost savings. But if it's out of distribution, small models may actually increase your cost because of how often they'll call tools and how crazy the loops will be. If you run Terminal Bench on Opus and Haiku, Opus will do about three times better at one-tenth the cost of Haiku, even though Haiku is significantly cheaper per token. So it really becomes a huge problem if you use too small of a model, particularly on tasks that are out of domain for the training data. When you're doing something like classifying text, like, hey, is this a person's name or is this an organization's name, that's super in-domain. So you don't want that kind of task to go to a large model. You want it to go to a small model. Everyone has that in their domain. So being able to understand in-domain, out-of-domain is a lot of work that we're doing for open router fusion. And then also figuring out how to orchestrate the outer and inner models for different types of tasks. And it's an early industry. It's an early field of research. Most research on model fusion has not been very detailed, not been very optimistic sometimes. So it really becomes a huge problem if you use too small of a model, particularly on tasks that are out of domain for the training data. When you're doing something like classifying text, hey, is this a person's name or is this an organization's name, that's super in-domain. So you don't want that kind of task to go to a large model. You want it to go to a small model. Everyone has that in their domain. So being able to understand in-domain, out-of-domain is a lot of work that we're doing for OpenRouter Fusion. And then also figuring out how to orchestrate the outer and inner models for different types of tasks. It's an early industry. It's an early field of research. Most research on model fusion has not been very detailed, not been very optimistic sometimes. It's only just recently getting more optimistic. And I think I'm personally very optimistic about it. We're a very ecosystem-driven, collaborative company. And we work with a lot of partners to try to help improve their orchestration pipelines with good primitives, like the sub-agent and the advisor tool, which is similar to what you were talking about. I'm curious. Help me understand. It makes total sense that a small model, if it's in-domain, would be cheaper. But if it's not, then it's going to thrash around as it tries to get an answer. When you're describing whether the main agent should be the local model or the cloud model, is that a decision that's then dependent on whether the task is something that's in-domain or not? Does my question make sense? I don't know. I don't know. It's early to say. I think the results that we published a couple weeks ago, which were focused on deep research, not coding, had the smart model be the wrapper model, be the outer model, and we got the best results from doing that. But for deep research, it works the best. For other tasks, it's unclear. Fusion is not super well-optimized for coding, and it might be that a smaller model ends up being a higher efficiency, or fewer dollars per successfully completed task, but it's early to say. One thing you said earlier is, oh, you get the caching benefit from the mainline agent. You actually can get the caching benefit from the side agent, and this is actually one of the key things we talked about with our Devon Fusion launch, is that you are leaving a lot on the table if you do a main agent and subagents type system. So we don't use subagents. We use what we call a sidekick, which is one subagent that continually has a running context. So the main agent doesn't need to re-provide context from earlier. It's all still in the KV cache, right? It's 10 times cheaper on all those cache tokens. And then if you want to switch the smart model to be the one on the side or the one in charge, that is actually totally fine, and you can do the swapping back and forth. We're also spending a lot of time right now thinking about how you train models to actually work collaboratively with other models. Actually, I think there's a lot of literature out there on how you RL one model to do a task end-to-end. How can you RL a model to also be good at collaboration? And when we think about it, we actually try both of these setups where let's RL the model being the orchestrator and the one deciding what gets delegated to other models, see how well that performs. And we also orchestrate it in a way where the model we're training is actually the executor, the sidekick, and see how well it is at executing other models' instructions. And we expect that to be probably a big lift in this next step of multi-model orchestration, is don't just take models as they are and orchestrate them, but can you actually co-design your models with the orchestration system? Yeah, that makes sense. With NemoTron and with all the foundational models, we're essentially post-training them for the harnesses that they're getting used in. If the harness is going to include a lot of routing, then that makes sense that that makes its way into the post-training. Yeah. Are you guys thinking a lot about model training at NVIDIA for these purposes? Yeah, so we have a technology called... The mics go? So... Testing. Oh, wow. Okay. So we have a technology called FlexRunt. We have a setup where there's a main model. Then we distill it into smaller footprints. And then based on the task at hand, you can switch which model does the decoding. So there's a lot of fancy stuff you can do within a model artifact to essentially only activate a class of model or a section of weights depending on the task at hand or the complexity at hand. In most cases, you can essentially understand the novelty of a question to a model if you have access to the recipe with which it was trained. So this works very well for open models, right? Any model you have access to its data for, right? Because you can literally decide if it's in... see if it's in distribution or not. Again, if you have studies from when it was trained, you can also see how much... essentially... how much was your distillation gap across teachers and the artifact that you trained, right? Because, sure, you have domain data from all the different domains you're tuning, but it's not guaranteed that it absorbs all the data evenly across the model, right? So it becomes very interesting to start thinking about these flexible weights and flexible model sizes, essentially. This also... I wanted to add about the context space, right? So how do you think about ASDs and context compression representations? Compaction, in its very nature, is lossy, right? So just like headroom is there, RTK is there, right? These code bases are usually designed to have representations that we carry forward through life. And you essentially give models the capability to further expand on them. It's more like loss-ish compression, which can retain states of models or status agents. What do you think about that? Yeah, I think this gets to a fundamental philosophy of how agents and context should work. One exercise I like to do is, as a human, how many numbers can you... if I just start spitting out numbers now, right? How many can you remember before you start losing track of them? I think it's actually very few, right? So in some ways you could argue that your context window is actually shorter than these language models, and yet you can actually be very effective at that, right? Your context is very lossy. I think one of the nice things that people are starting to realize with agents is you have a lot of non-lossy systems that you can fall back to. So you have a file system. If in your memory all you remember is that you read some file earlier, you don't need to remember the whole file. You maybe remember the important parts, but you can still have the full version of the file on your system. And that's my goal when I'm thinking about how do we build a good context-engineered harness. The harness should have everything it needs to find what it needs to have, even if it doesn't have everything immediately available. In that case, do you think that the context-sharing problem will become cheaper and less of a problem in future? Yeah, it's definitely possible as well. I've seen cases where the sidekick agent does a bunch of work and tells the main model, oh yeah, here's all the things I found, and instead of dumping the whole thing it just references them by file. And then the main model is actually generally... you find these larger, smarter models, they're actually more token efficient with how they use tools and how they read, and so they actually read the files in a way where they only see the important parts, right? Or they decide that, oh actually I only need to look at a subset of this, or I can run a single command and just know if everything is done properly. It's actually quite amazing, the fact that these multi-model systems actually seem to scale and get better with intelligence, which is not something we should just take for granted, right? It's not obvious that actually more expensive models are creating an overall cheaper system. Yeah, like the scaling laws, if you have a larger model it's going to be more efficient with its tokens, smaller models less efficient with its tokens. Yeah. I guess I had a question for you, Alex. Do you guys place way more importance on the actual it just references them by file. And then the main model is actually generally, you find these larger, smarter models, they're actually more token efficient with how they use tools and how they read, and so they actually read the files in a way where they only see the important parts, right? Or they decide that, oh, actually I only need to look at a subset of this, or I can run a single command and just know if everything is done properly. It's actually quite amazing, the fact that these multi-model systems actually seem to scale and get better with intelligence, which is not something we should just take for granted, right? It's not obvious that actually more expensive models are actually creating an overall cheaper system. Yeah, the scaling laws, if you have a larger model it's going to be more efficient with its tokens, smaller models less efficient with its tokens. Yeah. I guess I had a question for you, Alex. Do you guys lay more importance on the actual infrastructure side of routing? So for instance, KVCacheAware routing, or is that where most of the business is right now? Or are you seeing strong pull of people actually deploying routers in production? So OpenRouter is a marketplace for language models. We exist at, we can't see into the KVCaches of models unless we're running them ourselves, which is pretty rare. We do spend a lot of time optimizing for cache hits, and we pass through cache hits directly to users. But in terms of KVCache optimizations, we can't do any specific work there. What we do for model routing is we try to find the best model or best combination of models for the prompt, and then when we see a cache hit, we will use up the duration of the cache and send the downstream customer the full savings of the cache hit. There's more work that we can do here, where we could say, okay, this looks like something where there's significant benefit to switching the model right now, but you haven't used up the full cache. You still have two minutes left. And we think it's probably worth switching the model and losing the rest of your cache and letting people tweak their tolerance for that behavior. We've been doing a little bit of that, but we haven't exposed it to customers yet. What is next for you guys in terms of your model routing? Because as you mentioned, it is a different direction from the marketplace business that exists today, so I'd love to hear. So we've been doing, we've had an auto router for two years almost, but when we launched it, there was no adoption of it. People really wanted to use specific models, and the auto router just had no real usage. We mostly saw it as a discovery play point, like hey, this is how you discover which model might be good for your prompt. And then around January this year, with OpenClock, it exploded. And the reason it exploded is because there was this fundamental idiosyncrasy in OpenClock where it sends heartbeats every 10 minutes to your model of choice just to see if the client was still active. And that means that if you set Opus to be your default model, it would be using a lot of tokens on this heartbeat process. And so this was the very beginning of a very popular app with two completely different intelligence needs. Completely different. And the models, the open-source models, have improved to a point where it makes sense to segment the market in at least those two areas. And so that's how it got started. And then we saw a lot more segmentation blossom afterwards. And now a whole bunch of agents and apps on OpenRouter use the different routers that we have. And we have a couple of them. We have Pareto code, which gives you the Pareto optimal model for coding tasks given a certain threshold that you can tune. We have Fusion, which orchestrates multiple models and gives you a fused result. And we'll have other experiments in the future. What we want to do is basically create good primitives that developers can use to get really advanced with how they use model orchestration. Kind of like Sidecar, like that. But also give people a really easy thing that they can just set a slug to that works with all harnesses and gets the job done. I feel like it's super interesting how much of a perfect storm there is for model routing right now. Because on one hand, ignore agents, ignore OpenClaw for a second, just to squeeze better performance it seems like we should be smarter about how we tackle problems. That's obvious. Right? If you make a plan, if you make a strategy, that's a better way to go about your day. So I'm not surprised that you're going to see better code get written or more performant code get written, less buggy code get written, if you break the problem down. And so routing specifically for that use case makes a ton of sense. But then hearing this, yeah, the profile of workloads changed with agents, right? They're very, it went from I ask questions, I get a response, then it went to reasoning, where I ask questions, it reasons, and then it comes back, and then it went to, yeah, this heartbeat, right? If my agents are running optimally, there's a token being generated every second, and suddenly that is its own need for model routing, and it feels like hearing the different solutions to tackle each of those is very interesting. Even, well, I was just going to say, I think that the use cases for model routing are, there are many of them. And so one could be getting a better answer, one could be saving money and trying to get the same answer, one that we haven't even talked about yet, which is probably the most relevant, maybe even to this crowd, is when do you want to actually run a model locally versus when do you actually need something like a frontier model to perform that task? And that might be something to the effect of for privacy, protecting information. When I'm running something local, can you detect that my prompt has sensitive information and, if so, do that on device, but then maybe even anonymize some of that information to go do the more advanced workloads on top of that information in the cloud? Another example would be, again, for the cost savings, but it's like, hey, I bought this DJX Spark, and I know I'm not at 100% utilization. How can I make sure that as part of my workloads, whether it's the heartbeat and OpenClaw or what have you, that I'm leveraging that compute to the fullest of its ability? Because I'm only paying for the electrons that are coming in for my power bill, but I'm paying full price for the tokens in the cloud. And I think that's a whole other area of model routing that I know that we're doing some work with at NVIDIA that I think will be really cool as the hybrid of local and cloud starts to really emerge as its own sector. Yeah, I'd be curious to know what you guys' take is on, if you self-host a model, the cost dynamics change. You have a considerably higher cost at a higher context length because your throughput slows down as the context gets deeper. So rather than switching to a cheaper model, even if you have self-hosted models in data center, you can use compaction to bring your throughput back up. Have you guys thought about this paradigm, compaction versus just routing? Because one is you have fewer tokens to work with, one is we have cheaper tokens. Yeah, I think in practice, compacting alone doesn't solve the cost or throughput problems, because a lot of times the differential in model intelligence and cost is just so big. Oh, so by the way, when you compact, you're taking a cache miss, so you're actually now paying 10 times as much for those input tokens. If you didn't compact, the main reason we compact is actually intelligence. All these model providers, they advertise some insane context window, like a million tokens. I would never recommend using these models past 200k tokens, under 100k if you can. The intelligence just falls off a cliff at some point. Sorry Anthropic if you're watching, but I think that compaction is a very useful tool if you are going to have to take a cache miss anyway, one way or another, like when you're routing to another model, and you want to minimize the window there. Do you find that in the Sidecar, yeah, when small models are generating lots of tokens, is that one of the best reasons to switch it to a larger model? Basically, when small models generate lots of tokens, I wonder if that's a crux of the root cause of intelligence problems down the road. You want your big model to generate the big token chunks, the small models to generate smaller token chunks, right? And on that question, you mentioned a small model essentially needing to flag that it needs help from the larger model. What is that mechanism? Because that seems like, what's the indicator, and then what's the mechanism for it to do so? Yeah, totally. So there are a lot of mechanisms. We talk about in our blog post how we just detect that we need to change the model up. To I guess to answer your question first, how does the small model detect? Actually, the thing that we spend a lot of time on is how do we make sure the small model is good at detecting it? Unfortunately, there are a lot of cases where you do need the big model to detect it. One thing that we don't go into in the blog post is you have some kind of cadence on which you're refreshing the cache anyways. a crux of root the root cause of intelligence problems down the road you want your big model to generate the big token chunks the small models to generate smaller token chunks right and on that question, you mentioned a small model essentially needing to flag that it needs help from the larger model. What is that mechanism? Because that seems like what's the indicator, and then what's the mechanism for it to do so? Yeah, totally. So there are a lot of mechanisms we talk about in our blog post about how we just detect that we need to change the model up. I guess to answer your question first, how does the small model detect? Actually, the thing that we spend a lot of time on is how do we make sure the small model is good at detecting it? Unfortunately, there's a lot of cases where you do need the big model to detect it. One thing that we don't go into the blog post is you have some kind of cadence on which you're refreshing the cache anyways, because by default there's some five-minute lifetime on these caches. If you're going to go refresh the cache anyways, you basically can get a free big frontier model call if you ask the right question. So it's at that point where you might say, hey, just take a look at what the small model is doing. Does it feel like it's going into some rabbit hole and need some help? Now, what's the need for the five-minute refresh? It's just a practical you have to pay some kind of cost to keep these KVD caches warm, and so most caches just get evicted on some kind of cadence. How it works is at inference time, you only have so many cache you can keep loaded in the GPU. So once, if a cache is not being used again again, it's offloaded, so it's lost essentially. So that's why the inference provider asks you for money, but if you self-host it, you can get around this problem. You can make it as long as you want based on your business logic. Do you see a world where we'll have much more dynamic cache durations rather than just the five-minute, one-hour, depends on who's deploying the model where? Right, so if you have a GPU which has a lot of memory, the ratio of, let's say, SAMPs to memory is memory more heavily skewed, or if you're working with unified memory and you have systems like Verarubin, you have a lot of tricks to play here, right? The five-minute window is what a lot of providers right now put, but that's more an operational determination rather than a science-based or a core physics determination. So you can technically see over time maybe some APIs are priced differently, but if you self-deploy, again, you can get past a lot of this. The cost economics really change when you move from self-hosted models to API providers because you have a lot more control, and you don't have to guess the shape of your workload. So let's say if your workload is 32k, on average, 32k cache, 1k input, 1k output, and someone else's, let's say, 64k, 1k, 1k. If you use some provider, they are amortizing everyone's use case and then giving you a price, right? And they have optimized, quote unquote, for general use. If you self-host, you can optimize specifically for your use, and you'll likely pay much less. This is the level of hardware-software frontier that we were thinking about when we started Cognition, and we were working on the first agents. I think one reason why no one else worked on agents is they were just extremely expensive. This was before cache tokens was a thing that API providers paid for. If you were sending 100,000 tokens and the same 100,000 tokens, you were paying full price for those tokens. Back in 2024, when we started, one of the key things that let us build Devin and build these first agents was we actually bought direct compute capacity from these providers, and instead of paying on a token basis, we just paid for the underlying compute, knowing that the economics of the compute was that we were actually paying far less for the cache tokens that we'd send over. And nowadays, there's similar things. I would like having a version of the cache that maybe you can just back out to storage in S3 or something and just hold for much longer. Yeah. Now this is not extremely relevant to a DGX setup, but if anyone is looking to do what you guys want to do, try out Dynamo. We have a lot of prefix cache optimizations in there. Yeah, and then going back to your question, Alex, I think you said, oh, when a small model is going off and generating a ton of tokens, is that an interesting time to back off? To be honest, we haven't explored that yet, so that might actually be a very interesting thing to take a look at. It is weird. I think some small models do tend to be less token efficient than others, but they also seem to be trained on their own traces, so maybe in a way it ends up bouncing out. A lot of these things I feel like we have to be very empirical about to actually know. So just to add on that, you have a lot of, these days, there are a lot of hallucination probes, so probes that work on either the internal state, the internal state of the models directly, so you can have some form of either magnitude analysis done, or linear probes, or just the n types of probes that you can see, and you can essentially rate how much you think is tending towards hallucination. So that kind of gives you a proxy for how lost it is, how lost a model is in its thinking. So you can use different kinds of probes to understand the perplexity within a model. That's interesting. So yeah, instead of using the quantity of tokens as indicative of a model being lost, it's hallucinating more. Yeah. So essentially, what is cache, right? It's the pre-fill states, right? So what is a pre-fill state? It's just a vector at the end of the day. So you can tune all kinds of classifiers to understand different aspects of those collections of vectors. So with those kinds of probes, you can guesstimate a lot of states of a model. I see. One question I have is, different models behave differently, and that means that these prompts aren't portable. So as you're doing model routing, how do you handle, essentially, if you're going to a different model architecture, what do you have to do to the prompt, and how much is that a factor into either of your guys' model routing solution? How is the prompt itself a factor into the routing? Yeah. Well, I think with building agents, there are all kinds of paper cuts and edge cases that are domain specific, and the value of an agent company, the value of Devin, is all these doom loops that you've discovered that are across all industries and the best ways to recover from them. And it manifests big time in what the prompts are going to be, both for how the advisor model gets called, the smart friend, how the subtask agents get called. And the best thing is that anyone can, any engineer or any agent can inspect the traces and adjust the prompt, and then see the live accuracy a long time. So I just think that that's part of the startup building process and is also really easy to observe and have multiple people and agents collaborate on them. Yeah. One thing I'd love to do with our Fusion product, and we don't have this yet, and so this is maybe a preview of some things we work on, is you can tune it against a data set, but the real thing you want when you're building a real agent someone uses is to just tune it against what actual people use it for and what actual models they get routed to. And so there's a lot of signals for this. If someone sends a prompt and then Devin is working, and then you see that the user decides themselves to upgrade to a different model, or they decide to downgrade, or the system detects that we originally sent to the wrong one, we now got to replace, that's actually a really useful stream of signals. And we're actually getting to this world of auto research where maybe we can just have a constant stream of prompts, what it should have been, what it was instead, and build a system internally that's just capturing all of this and then reiterating on our routing system until it eventually fits the real production data. That's now that it's public and people are using it, this is now something that we're thinking about. Have you guys looked into prompt tuning, and do you find it useful, like, say, Jepa? Yeah, so there are these prompt tuning frameworks from a few years ago that tried to do some kind of gradient descent type thing. I'm actually personally less bullish on these low-level mechanical prompt tuning harnesses versus just telling a smart model, here is the decision that was made and the context. Figure out why it went wrong. Sometimes you can do something as dumb as asking a model, why did you do this instead of this, and cite the prompts, and then just have your agent, your Devin, just go and update the prompts, rerun the test as a regression, make sure it changes. It's a lot heavier weight of a system, but I trust the intelligence of a system like that a lot more. So we're running out of time, so we're going to wrap up real quick, but I think what's really interesting this is now something that we're thinking about. have you guys looked into prompt tuning, and do you find it useful, say JEPA? yeah, so there are these prompt tuning frameworks from a few years ago that tried to do some kind of gradient descent-type thing. I'm actually personally less bullish on these low-level mechanical prompt-tuning harnesses versus just telling a smart model, here is the decision that was made and the context, figure out why it went wrong. sometimes you can do something as dumb as asking a model, why did you do this instead of this, and cite the prompts, and then just have your agent, your dev, just go and update the prompts, rerun the test as a regression, make sure it changes. it's a lot heavier weight of a system, but I trust the intelligence of a system like that a lot more. so we're running out of time, so we're going to wrap up real quick, but I think what's really interesting is just from talking to you guys, we can see how new this space is, right? how much of this is actually just research. we're starting to see new products come in, and I'm really excited about your guys' solutions as you guys enter this space. the ways and the needs that you need routing for, even on a DGX Spark, when you're doing local inference, you have more compute, and if the memory's filled, or if the memory utilization is high, one thing you need to do is increase the compute utilization. and so one way you can do that is by spawning multiple agents that are working collaboratively. so that collaborative piece is something that not only is optimal for all of these cloud workloads that you guys are doing, but specifically that is how you extract more performance out of this edge hardware. and I think a question here is, maybe to end on, is a router going to be something that we see as a product, or is that going to be seen as part of the plumbing here? are models going to get good at routing to other models because they know they need to be collaborative, or are harnesses going to know that they are working across multiple models? I think we already see this. at Cognition, we're training our models to be able to be good collaborators. I think it's very clear that new frontier models, the Fable models and GPT 5.5, 5.6 models, are themselves naturally collaborative and better at delegation, so I think we're ready there at that point. interesting. yeah, I think that the systems are becoming, not muddied in some sense, but I think that ultimately we're understanding that as we step up the abstraction ladder and build more things to create this smarter blob, which obviously we should hopefully, and we do understand how we're building it and why we're building it, that it's going to become a system that you look at, both the different components of the system, but it's not just going to be just models. there's not going to be a thing as a really great harness that is in absence of a really great model, and vice versa. yeah, makes sense. I think applications, especially built on non-deterministic systems like models, operate in a very low-trust environment. so yes, most of the improvements will likely be distributed across both models and the harnesses, but I think overall there will have to be some form of comptroller trying to have some form of arbitration, because even from the model perspective, you aren't in a perfectly visible world. you don't know the behavior of every model, so it's going to be at the orchestration level where you have these kinds of things. and this has traditionally been shown by other industries. when web launched, you had traffic-based routing, charts different, but all the routing controls have been centralized over time. I think so. I think it's most likely going to be good news in the future, and I think caching is a big reason for that. even if you, I think to take the flip side of this argument, it might be that in the future we have one big model that's, I know I am the most efficient at everything, and I'm way more efficient than Haiku. I'll solve every task better than Haiku can at a lower price. why should I ever delegate to Haiku? something like that actually could be a model that we have in the future. but you're always going to have these, for example, caching. it could be that you tell the model that this other model does have the right context in cache, and the orchestrator model just always has more context, and the models have to be aligned. so I think I don't really see a world where we wouldn't be able to get models to collaborate really well, and I think they're going to get better over time, in part because they just have limited memory. so I think that's one deciding factor, and another is that there will continue to be, if you just look at the rankings on Open Router, if you look at our public data and you look at the top model being used by dollar spent on classification tasks, guess what it is? it's Opus. I think there are big opportunities for using small models for in-distribution, easy tasks, and as time goes on, that's going to be a larger and larger percentage of tasks relative to the most valuable tasks that very smart models spend most of their time on. totally. well, I want to thank you guys so much. can we all give everyone a round of applause? thank you. thank you. thank you. thank you. thank you. thank you. thank you. thank you. thank you. thank you. thank you. thank you. thank you. thank you. thank you. thank you. thank you. thank you. thank you. thank you. thank you. thank you. thank you. thank you. thank you. thank you. thank you. thank you. thank you. thank you. thank you. those collections of vectors so with those kinds of probes you can guesstimate a lot of states of a model I see one question I have is you know different models behave differently and that kind of means that these prompts aren't portable so as you're doing model routing how do you handle essentially if you're going to a different model architecture what do you have to do to the prompt and how much is that a factor into either of your guys' model routing solution like how is the prompt itself a factor into the routing yeah well I think with building agents there are all kinds of paper cuts and edge cases that are domain specific and like the value of an agent company like the value of devin is all these like doom loops that you've discovered that are across all industries and the best ways to recover from them and like it manifests big time in what the prompts are going to be both for like you know how the advisor model gets called you know the smart friend the how the like subtask agents get called and and the best thing is that like anyone can like any engineer or any like agent can inspect the traces and like adjust the prompt and then see the like live accuracy a long time so I mean basically I just think that that's part the prompt is part of the startup building process and is also really easy to observe and like and have like multiple people and agents collaborate on them yeah one thing I'd love to do with our Fusion product and we don't have this yet and so this is maybe a preview of some things we work on is you know you can tune it against a data set but the real thing you want when you're building a real agent someone uses is to just like tune it against what actual people use it for and what actual models they get routed to and so there's a lot of signals for this like if someone sends a prompt and then Devon is working and then you see that the user decides themselves to like upgrade to a different model or they decide to downgrade or the system detects that we originally sent to the wrong one we now got to replace but that's actually a really useful stream of signals and we're actually getting to this world of like auto research where like maybe we can just have like a constant stream of prompts what it should have been what it was instead and build a system internally that's just capturing all of this and then reiterating on our routing system until it eventually kind of like fits the real production data that's kind of like now that it's public and people are using it this is now something that we're thinking about have you guys looked into prompt tuning and do you find it useful like say Jepa yeah so there are like these prompt tuning frameworks from like a few years ago that tried to do some kind of like gradient descent type thing I'm like I'm actually personally less bullish on these kind of like low level mechanical prompt tuning harnesses versus just telling like a smart model like here is the decision that was made and the context figure out why it went wrong sometimes you can do something as dumb as asking a model why did you do this instead of this and cite the prompts and then just have your agent your dev and just go and update the prompts rerun the test as a regression make sure it changes like it's a lot heavier weight of a system but I kind of trust the intelligence of a system like that a lot more so we're running out of time so we're going to wrap up real quick but I think what's really interesting is just from talking to you guys we can kind of see how new this space is right how much of this is actually just research we're starting to see new products come in and I'm really excited about your guys' solutions as you guys enter this space the ways and the needs that you need routing for you know even on a DGX Spark when you're doing local inference you have more compute and if the memory's filled one or if the memory utilization is high one thing you need to do is increase the compute utilization and so one way you can do that is by spawning multiple agents that are working collaboratively so that collaborative piece is something that not only is optimal for all of these cloud workloads that you guys are doing but specifically that is how you extract more performance out of this edge hardware and I think you know a question here is maybe to end on is a router going to be something that we see as a product or is that going to be seen as part of the plumbing here are models going to get good at routing to other models because they know they need to be collaborative or harnesses going to know that they are working across multiple models I think we already see this like you know at Cognition we're training our models to be able to be good collaborators I think it's very clear that new frontier models like the Fable models and GPT 5.5, 5.6 models are like themselves like naturally collaborative and better at delegation so I think we're ready there at that point interesting yeah I think that the systems are kind of becoming not muddied in some sense but I think that ultimately we're understanding that as we step up the abstraction ladder and build more things to create this smarter blob which obviously we should hopefully and we do understand how we're building it and why we're building it that it's going to become a system that you look at kind of both the different components of the system but it's not just going to be just models there's not going to be a thing as like a really great harness that is in absence of a really great model and vice versa yeah makes sense I think applications especially built on non-deterministic systems like models operate in a very low trust environment so yes most of the improvements will likely be distributed across both models and the harnesses but I think overall it's mostly there will have to be some form of comptroller trying to have some form of arbitration because even from the model perspective you aren't in a perfectly visible world you don't know the behavior of every model so it's going to be at the orchestration level where you have these kind of things and this has traditionally been shown by other industries like when web launched you know you had traffic-based routing charts different but all the routing controls have been centralized over time I think so I think it's most likely going to be good news in the future and I think like caching is a big reason for that even if you I think like to take the flip side of this argument you know it might be that in the future we have like one big model that's like I know I am the like most efficient at everything and I'm like way more efficient than Haiku I'll solve every task better than Haiku can at like a lower price why should I ever delegate to Haiku? something like that actually could could be a model that we have in the future but you're always going to have these like you know for example caching it could be that like you tell the model that this other model like does have the right context in cache and you know the orchestrator model just always has more context and the models have to be aligned so I think like it's I don't really see a world where like we wouldn't be able to get models to collaborate really well and I think they're going to get better over time in part because you know they just have limited memory so I think that's kind of one one deciding factor and another is that there will be like there will continue to be like if you just look at like the the rankings on Open Router if you look at our our public data and you look at like the top model being used by dollar spent on classification tasks like guess what it is it's Opus I think there there are there are big opportunities for like using small models for in distribution easy tasks and and like as time goes on that's going to be a larger and larger percentage of tasks relative to like the most valuable tasks that very smart models spend most of their time on totally well I want to thank you guys so much can we all give everyone a round of applause thank you thank you thank you thank you thank you thank you thank you thank you thank you thank you thank you thank you thank you thank you thank you thank you thank you thank you thank you thank you thank you thank you thank you thank you thank you thank you thank you thank you Thank you.