Open Reader

Preferences Over Benchmarks: Model Routing — Archana Kamath & Tyler Gillam, DigitalOcean

completed 15:53 Aug 22, 2026 Watch on YouTube

Current Status

completed

Video ID

FvxY8oPoI8o

RAG / Chat

Enabled
Preferences Over Benchmarks: Model Routing — Archana Kamath & Tyler Gillam, DigitalOcean
Description

Two terminals run the same prompt, build me a spinning wheel app. On the left every request goes to a single premium model. On the right they go through a router that picks a model per task. Both finish at about the same time with comparable output, and by then the router's session has cost 8 cents against 25. The gap widens with every prompt after that. Archana Kamath and Tyler Gillam use it to argue that picking a model by climbing a leaderboard is the wrong instinct, because there is no single best model, only the right one for a given request. What makes a model right is a mix no public leaderboard encodes: the task itself, the system prompt and tools around it, the cost you are willing to spend, the latency the use case needs, and what the end user actually wants. Their router takes those as preferences you declare, in natural language or as decision tree rules, then honors them per request. It runs on a purpose built mixture of experts model that decides in under 200 milliseconds, costs nothing extra, and is open sourced along with the proxy in front of it. Gillam then shows the part that separates it from a vibe check, an evaluation scoring the router at 90% correctness against 95% for the single premium model while using far fewer tokens and returning faster. Routing is the foundation layer, with evaluation, caching and personalization built on top. Speaker info: - https://www.linkedin.com/in/tdgillam Timestamps: 0:00 - Why the one model habit is breaking 2:42 - There is no single best model 4:21 - A router you can customize and evaluate 6:57 - Configuring tasks, model pools and failover 7:48 - Side by side in the playground 9:29 - Proving it with an evaluation 10:18 - Two coding agents, and the session cost gap 13:49 - Under 200ms, open sourced, no code changes 14:43 - Evaluation, caching, personalization

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Model selection should be a per-request, preference-driven control-plane decision—optimized against task, quality, cost, latency, policy, and resilience—not a static choice of the benchmark-leading model.
  • Why it matters: For agentic and software-engineering workflows, task-level routing can materially reduce inference spend and latency while preserving acceptable output quality and adding failover from single-provider/model dependencies.
  • Best use: Use this as a concise architecture and operating-model reference for building an evaluable model-routing layer in front of multi-step agent workflows; treat the vendor performance claims as a prompt to run equivalent evaluations on your own traffic.

Executive Summary

DigitalOcean argues that the default habit of sending all AI work to one frontier model is economically and operationally unsound. It identifies three drivers for routing: escalating inference costs, poor task-to-model fit when premium models handle routine work, and concentration risk when a single model degrades or goes offline. Its central reframing is that there is no universally best model; there is only a model that is right for a particular request under a team’s chosen preferences.

The proposed router accepts a natural-language task description plus explicit preferences—cost, latency, quality, preferred models, and hard rules—and selects from a configured model pool. It supports policies including fixed priority with failover and a “fastest” policy based on recent performance. The company positions its routing model and proxy layer as open source and claims routing takes under 200 ms with no added customer charge.

The demo is most useful as an agent-workflow pattern rather than as proof of universal savings. A coding workflow routes code snippets to Llama Maverick, code generation to GLM 5.2 with GPT 5.2 as failover, performance optimization to GPT 5.2, and test writing to Claude Sonnet. In a small side-by-side OpenCode session, the router used $0.14 versus $0.44 for direct Claude Opus, while the presenters judged the resulting app, tests, and README as broadly comparable.

The strongest operating lesson is that routing must sit inside a continuous loop: define policies, route, evaluate against workload-specific tests, adjust configurations, and eventually add caching and personalization. Public benchmarks remain inputs, but they cannot encode an application’s system prompts, tools, cost ceiling, latency budget, or end-user quality preference.

Key Takeaways

  • Claim: A single-model strategy is increasingly a cost, fit, and availability liability; routing is becoming an AI-era FinOps discipline. | Evidence: The speakers cite Walmart, Uber, and Microsoft as companies actively capping usage to control inference bills, and argue that smaller models can handle classification, labeling, and some code tasks without paying frontier-model rates. | Implication: Treat model orchestration as a first-class cost-and-reliability control plane, with budgets and fallback behavior designed into the product rather than negotiated after spend escalates. | Caveat: The cited company examples are not accompanied by spend figures, implementation details, or evidence that their controls use model routing specifically.
  • Claim: The right model is determined by the individual request and local operating constraints, not by a public leaderboard ranking. | Evidence: The decision inputs named are the task, surrounding system prompts and tools, acceptable cost, latency requirements, and end-user preferences. The talk distinguishes routine classification, code generation/bug fixing, and high-accuracy code review or security work. | Implication: Build routing around an explicit task taxonomy and service objectives rather than declaring one approved model for an entire application or agent. | Caveat: The task categories are heuristics, not guaranteed model assignments; accuracy-critical work still requires application-specific testing.
  • Claim: Routing should be customizable and policy-constrained rather than an opaque auto-routing black box. | Evidence: DigitalOcean’s configuration exposes task definitions, cost/latency/quality preferences, preferred models, hard rules, and decision-tree rules. Its example uses manual ranking to send code generation to GLM 5.2 unless it fails, then fail over to GPT 5.2; bug fixing instead uses a “fastest” policy over the last roughly 30 minutes. | Implication: Separate non-negotiable policy constraints—approved providers, data boundaries, model fallbacks, spend ceilings—from dynamic optimization policies such as recent latency or price. | Caveat: A fastest-recently policy can optimize responsiveness while masking quality variation unless it is bounded by quality thresholds and monitored separately.
  • Claim: Workload-specific evaluations, not benchmark scores or subjective “vibe checks,” are the mechanism for proving and improving routing decisions. | Evidence: The presenters prescribe a route-evaluate-adjust-feedback loop. In one displayed evaluation, their router scored 90% versus 95% correctness for Opus, which they characterized as near the margin of error, while using fewer tokens and completing faster. | Implication: Before deploying a router, maintain a representative evaluation suite by task class and use it to establish minimum quality gates, rather than optimizing only aggregate cost and latency. | Caveat: The transcript does not define the evaluation dataset, sample size, scoring methodology, statistical confidence, or whether the 5-point correctness gap is actually insignificant.
  • Claim: Task-level model switching can yield substantial session-level savings in a multi-step coding agent while retaining similar visible outputs. | Evidence: In the live OpenCode comparison, the routed workflow selected different models for generation, testing, and documentation and ended at $0.14, versus $0.44 for a direct Claude Opus workflow—roughly 3x lower cost. Both completed a spinning-wheel app workflow, though the presenters subjectively preferred the router output in one implementation detail. | Implication: The highest-return routing opportunity is likely not isolated chat prompts but multi-turn agent sessions, where a premium model otherwise receives every planning, implementation, verification, and documentation call. | Caveat: This is a small, vendor-run demo using an informal visual assessment; it does not establish equivalent reliability, maintenance quality, security, or savings across production workloads.
  • Claim: A routing layer is foundational infrastructure, but it becomes more valuable when combined with evaluation, caching, and personalization. | Evidence: The closing architecture identifies evals to validate model-task fit, caching to avoid paying repeatedly for the same answer, and personalization so the router learns what works for a particular team over time. | Implication: Design routing as one component of a broader inference optimization loop; do not expect a model selector alone to solve quality, unit economics, and user-specific behavior. | Caveat: The talk does not explain how personalization is trained, how it avoids feedback loops, or how cache correctness is maintained when prompts, tools, or user context change.

Detailed Brief

DigitalOcean implementation and integration posture

  • Claims: DigitalOcean says a request traverses its open proxy, Plano, and a purpose-built routing model.; The company positions the routing stack as open source and claims customers can adopt it without application-code changes.; Its router is presented as a custom mixture-of-experts model designed specifically for routing rather than for generating end-user answers.
  • Evidence: The claimed routing-decision latency is under 200 milliseconds per request.; The service is described as included at no additional charge to customers.; The presenters claim their routing model outperformed frontier models, including the GPT-5 series, on the routing task itself at a fraction of the latency.
  • Caveats: No independent benchmarks, model-card details, routing accuracy breakdowns, availability terms, or source-code review are supplied in the transcript.; “Zero application code changes” may apply to a compatible proxy/API integration but should be verified against existing observability, provider abstraction, authentication, and failover requirements.
  • Implications: An open proxy and externally inspectable routing component can reduce implementation friction and provider lock-in, but the actual control-plane ownership and portability should be tested before standardizing on the service.; Routing overhead must be assessed against the latency budget of short requests; sub-200 ms may be trivial for long agent steps but meaningful for high-volume interactive paths.

Notable Concepts & Terms

  • Preferences over benchmarks: The talk’s governing principle: model choice should optimize an application’s own quality, cost, latency, policy, and user-preference requirements rather than chase a generalized benchmark leader.
  • Model routing: Selecting a model separately for each request or task instead of fixing one model for an entire application or agent session.
  • Model orchestration as FinOps: The framing that model choice, fallback, and usage controls are now core mechanisms for managing AI unit economics, analogous to cloud-cost optimization.
  • Manual ranking: A deterministic preferred-model policy with ordered fallback; the demo uses GLM 5.2 first for code generation and GPT 5.2 only on failure.
  • Fastest selection policy: A dynamic routing policy that chooses among eligible models using recent latency, described as performance over the prior approximately 30 minutes.
  • Route-evaluate-adjust loop: The proposed operating cycle in which routing policies are tested on private workload evaluations, revised, and fed back into the router.
  • Plano: The open-source proxy/routing component named by DigitalOcean as part of the request path and its no-vendor-lock-in positioning.
  • Task taxonomy: The practical classification layer used to map work such as code snippets, generation, bug fixing, test writing, documentation, review, or security to different model pools and policies.

Operator Notes / Why Ken Should Care

  • Create a routing evaluation harness using real agent traces segmented into planning, generation, debugging, test writing, review/security, and documentation; capture quality, task success, latency, tokens, and fully loaded cost per successful outcome.
  • Define hard routing guardrails before optimizing: approved-model/provider lists, data-residency or confidentiality constraints, maximum per-task spend, minimum quality gates, and deterministic failover paths.
  • Run a controlled multi-step agent comparison between a premium single-model baseline and a task-routed configuration; measure end-to-end task completion and rework, not only per-call latency and token cost.
  • Do not accept the demo’s 3x savings or near-equivalent quality as a production assumption; require sample-size, evaluation-methodology, and failure-mode evidence for each workload.
  • Assess whether a proxy-based router can preserve trace IDs, tool-call observability, authentication boundaries, retry semantics, rate-limit handling, and reproducible routing decisions in the existing stack.
  • Prioritize caching separately from routing for repeated or deterministic requests, while defining cache invalidation around prompt versions, tool state, user context, and model changes.

Source/Metadata

  • Title: Preferences Over Benchmarks: Model Routing — Archana Kamath & Tyler Gillam, DigitalOcean
  • Transcript words: 3326
  • Duration seconds: 953
  • Timestamp note: No usable timestamps or chapters were present in the supplied transcript. The closing portion is duplicated, so the supplied word count includes repeated content.

Transcript

2694 words en Processed in 117.4s

Shri Mataji Reviewer Hello everyone. So, preferences over benchmarks. The talk today is about model routing, and specifically why the way most people think about picking a model, which usually is chasing to the top of a benchmark, is actually the wrong instinct. I'm Archana, VP of Engineering for Inference Engine and AI Infrastructure at DigitalOcean. And I'll be joined by Tyler, who built parts of the router and will actually do a live demo for us today. We both work on the managed agent orchestration and inference engine products at DigitalOcean. So, you may know DigitalOcean as droplets, databases, and our platform. All of that is true. We are also the AI native cloud. There's just five integrated layers, starting from infrastructure all the way up to the managed agents, with the inference engine right in the middle. And that's why we are here talking about inference router. Routing lives in the inference engine. And if you want to know more about our stack and the full story, please come find us at the booth. So, everybody is reaching out for model routing. And let's look at three reasons why, the three reasons that are breaking the one-model habit for most users. The first one I want to talk about is cost. Spend is exploding, and even companies like Walmart, Uber, and Microsoft are actively capping usage to control the inference bills. The second one is fit. One model for every task is likely overkill. You are essentially paying frontier rates for work that a much smaller model will be able to handle really well. And the third one, which for me is the most important one, is the risk. The risk associated with one single model. Models can go down. And if you bet your entire product and production on one model, you have no failover when something degrades. And model orchestration is actually the new FinOps. As you all know, cloud cost optimization took us about 15 years for it to actually become a really good discipline and for companies to get it right. This one actually is arriving in months and not years. And here's the premise that I think everybody gets wrong about this. We all think of, what is the best model for a job? Here's the thing. There is no single best model. The right one depends on the actual request. For example, if you're doing classification and labeling, a small open model may very well work really well for you and will give you really good cost optimizations. And that is where a faster, larger routing model comes into the picture. Think about code generation and bug fixing. You're likely good with a mid open weight model. And again, it will bring you really good cost optimization. So we're using a frontier for something that is likely overkill in this situation. But then you're looking at really accuracy-critical tasks like code review and security, and you're likely going to lean towards a frontier model. So essentially, what makes a model right for a request? It's a mix that no public leaderboard can actually encode for you. Because it's the task itself. What are you actually trying to achieve? What is your model trying to achieve? The system prompts and tools around it. That is the methodology by which you're getting something done using a model. The cost you're willing to spend, this is a very, very important aspect. And latency the use case needs. Not all use cases need the same amount of latency. So depending on what you're trying to do, this can vary widely. And finally, the end-user preference. All of this is driven by what the end user really wants out of your application. So I'm going to welcome Tyler onto the stage so that he can actually show you the inference router live in action and show you how it can really help with all of these key aspects that I'm calling out here. Testing. All right. Thank you, Archana. Okay. So many builders have tried auto-routing before. But the problem was that it feels like a black box. The router makes a choice. And if that choice results in poor performance, you really have no way of improving it. We built ours differently. At the architecture level, which is what you can see on the screen, a request runs through our open proxy, Plano, and our purpose-built routing model. Both open source. There is no vendor lock-in, which is a key DigitalOcean value. You describe what matters for your workflow: cost, latency, quality, preferred models, or hard rules. Then the router uses that context to pick the right model per request. Because the routing model is specialized for this job, it's super fast, under 200 milliseconds. And it costs customers nothing extra. In our evaluations, it actually has beaten frontier models like the GBT 5 series models at that routing task itself, with a fraction of the latency. So the difference is simple. This is routing you can customize, evaluate, and improve without vendor lock-in. So you bring your preferences, and we honor them. You describe a task in natural language and set what matters: cost, latency, and task description. You bring your rules, and we execute them intelligently. Layer decision tree rules on top, start from presets, change anything you want, in a single line of code. And you validate with your own evaluations, not someone else's leaderboard. Route, evaluate, adjust, then feed that back in. That loop is key. Okay, we're going to switch gears here. We're going to do a live demo. Bear with me here. All right, I'm going to show you a couple things. First, I'll show you router configuration in the UI, how to use it, and then how you can use evaluations to measure and improve your router's performance. And then I'll show you a real router that I created inside a coding agent workflow. So I'm here in the Cloud Console, the DigitalOcean Cloud Console. And you can see my routers. We have several presets. You can see software engineering in general, writing, knowledge bases, and document intelligence. In this case, I've actually created my own. So I customized our preset software engineering. If we click into this, we can see that I have several different tasks here. I have bug fixing, code generation, test writing, and a few others. This also shows that you can specify more than one model per task in the bug-fixing case and code-generation case. In the code generation, I have GLM 5.2 and GPT 5.2. And because I really want to always route to GLM 5.2 unless it's down, I use this manual ranking option. So it'll always go to GLM 5.2. If GLM fails, it'll fail over to GPT 5.2. In the bug-fixing one, you can see a little bit of a different one. In this case, I have selection policy fastest. So out of this model pool, if it matches to bug fixing, it'll pick whichever one's been fastest in about the last 30 minutes. Okay, let's do this in action a little bit. Here's our playground, where I'll show a couple of examples side by side. First, I'll start with a simple prompt: write a basic Fibonacci function. And as this runs, we can see on the left, we're routing to Opus. On the right, we're using our software engineering router that I just showed you. And you're going to see that it picks different models on the right. So in this case, it matched to the code snippets task and just used the long before Maverick model that I had configured for that one. And if we scroll down, this is obvious, right? But this model is extremely fast and extremely cheap compared to Opus. Now let's say, optimize my function. And we'll see the same thing happen. In this case, it matched to the code performance optimization task using GPT 5.2. And again, it's obviously significantly faster. If we scroll down here, we can also see that it's significantly cheaper. We'll do one more: write some unit tests. Okay, and in this case, it matched to quad 5 sonnet on the test-writing code verification. And again, we're going to see faster and cheaper. So it's a pattern. It matches my vibe check, right? It still vibes, though. How you actually prove it is working it through evaluations. So I have an evaluation that I ran here, comparing Opus on the left, actually on the right-hand side, to my router on the left-hand side. You can see that the scores, 90% from my router, 95% correctness for Opus, are very, very close. In fact, that's pretty much within margin of error. But what's really interesting is if we scroll down here, we can see that the router used significantly less tokens and was significantly faster than Opus. Okay, let's jump into a real workflow here. This is where the inference router really becomes impactful. Here I have two terminals running open code. On the left, I have a single-model approach using plot Opus. So I have Opus set up, or open code set up, with Opus. On the right, I've configured open code to send requests to our software engineering router that I just showed you configured. Below, I have this custom-built open code where you'll be able to see live observability, essentially. So let's go ahead and get these started. It's just a simple feature request preloaded into here: build me a spinning wheel app. I'll run the same prompt in both. And as this runs, we can focus on the bottom panel. So it'll start to show up here. Hopefully, we can see that on the screen. You'll be able to see token usage in real time, which models are being selected, what tasks those map to, and the cost accumulating live. So on the right, we can already see that we're starting to route to GLM 5.2 because our requests are starting to match the code generation. And on the left, of course, we're just routing to Quad Opus. I think open code sometimes routes to Haiku by itself, so that's what you see there. And we'll notice latency too, how quickly things start to come back. In this case, it wants me to create a temporary directory. So the key difference here is that on the left, we'll see every single request that I write goes to the same premium model. Cost and latency are going to stay high for pretty much every single task. On the right, the router is selecting models based on the task. So we're optimizing both cost and speed. And we can see that our software engineering router already finished. If we look here, it actually matched to two models throughout. So let's go ahead and open this up and see how it looks. Okay, this actually looks really solid to me. And Opus 4.7 finished at a similar time. Let's take a look at that. We can compare them. This is a vibe check, right? But honestly, I would say the software engineering router did better because this is an interesting approach that I'm not even sure works too well. So in this case, the router did a little bit better. So now that step is done, we get similar outputs. But if we look here, the software engineering router has only spent 8 cents on the session, while Opus directly has spent 25 cents. So we have about a 3x in cost and very, very similar quality so far. Let's try another prompt here. What comes next in a software engineering lifecycle? Probably write some unit tests, right? So we'll write this in both. Start up this first. On the right, we have the router again. And we can see that it got matched to the test-writing and code verification, which picked the Claude 5 sonnet model because that's what I configured earlier. And we'll see the same pattern. It's going to be significantly cheaper overall across the entire session than going straight to Opus. So we'll have this finish here. Okay, and that one finished. Let's just queue up one more: write some documentation in a README. And then we'll compare the total session cost. Okay, and as this runs, we'll wait and see what it does. Okay, it created the README. And if we look here, we can see that the total session cost for the router was 14 cents, while the total session cost for Opus was 44 cents. So at this point, we can see the cost is significantly lower. Latency is optimized per step. And the quality remains pretty similar across. So you can see, as you scale this, the cost performance really adds up. Okay, Archana, back to you. Thank you. Thank you. Thank you so much, Tyler. And that was actually a live demo that we ran here. So thanks to Tyler for setting it up and taking us through that. So now that you've seen it work, let's look at some quick facts. Routing decision in under 200 milliseconds per request. It runs on a custom mixture-of-experts model purpose-built for routing. Zero application code changes needed from you to get it to adopt. And it's free and included. So you do not have to roll out your own router. And we open source the whole routing model via Plano. So you can actually check how that looks as well. The last thing I wanted to talk about a bit was routing is the foundation layer. It's not really the destination. And there are three things that we usually build on top of it. The first one is evals to prove that the right model works with your use case and your test well. Caching so that you can stop paying twice or more for the same answer each time. And personalization so that the router learns what works for your team over time. This is a continuous improvement loop maturing over time. That means that the more you route and evaluate, the better the router does for your workload. So to summarize, where does this leave you? There is no single best model. There's only the right model for the request. And benchmarks will only tell you part of the story. Your preferences will tell you the rest. And we built the router to honor your preferences and stay open so that you're never locked into a single stack. And that's how teams actually build. We are DigitalOcean, an AI native cloud. Come find us at the booth and route your next workload with us. Thank you so much for being here. That yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay you Okay, it created the readme. And if we look here, we can see that the total session cost for the router was 14 cents, while the total session cost for Opus was 44 cents. So at this point, we can see the cost is significantly lower. Latency is optimized per step. And the quality remains pretty similar across. So you can see as you scale this, the cost performance really add up. Okay, Archana, back to you. Thank you. Thank you. Thank you so much, Tyler. And that was actually a live demo that we ran here. So thanks to Tyler for setting it up and taking us through that. So now that you've seen it work, let's look at some quick facts. Routing decision and under 200 milliseconds per request. It runs on a custom mixture of experts model purpose built for routing. Zero application code changes needed from you to get it to adopt. And it's free and included. So you do not have to roll out your own router. And we open source the whole routing model via Plano. So you can actually check how that looks as well. The last thing I wanted to talk about was a bit about routing is the foundation layer. It's not really the destination. And there are three things that we usually build on top of it. The first one is evals to prove that the right model works with your use case and your test well. Caching so that you can stop being twice or more for the same answer each time. And personalization so that the router learns what works for your team over time. This is a continuous improvement loop maturing over time. That means that the more you route and evaluate, the better the router does for your workload. So to summarize, where does this leave you? There is no single best model. There's only the right model for the request. And benchmarks will only tell you part of the story. Your preferences will tell you the rest. And we build the router to honor your preferences and stay open so that you're never locked into a single stack. And that's how teams actually built. We are DigitalOcean, an AI native cloud. Come find us at the booth and route your next workload with us. Thank you so much for being here. That yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay yay you