Open Reader

Large clusters for small models — Daniel Svonava, Superlinked

completed 25:06 Sep 19, 2026 Watch on YouTube

Current Status

completed

Video ID

g4SsanB0gMc

RAG / Chat

Enabled
Large clusters for small models — Daniel Svonava, Superlinked
Description

A single mid range GPU can turn half a million tokens per second into embeddings in the low tens of milliseconds, where a managed endpoint costs orders of magnitude more and takes hundreds. Daniel Svonava calls embeddings the no brainer entry point. His real subject is what comes after. A small model fits on one GPU two or three generations old, and for a specific task it is now at or beyond the frontier, which is flattening while small open models climb. You do not prompt one 27 billion parameter model for everything; you slice the workload into tasks and pick the model trained for each, so a contract review agent ends up running nine. The model that has seen the most Vietnamese receipts wins Vietnamese receipt OCR. The models exist; serving a wide fleet is the bottleneck. Three things break. Open source serving tools ship untuned, so adopting one is a research project. Top down routers, built to spread one big model across GPUs, choke on many small requests because their view of worker state is always stale; utilization stalls near 30 percent. And LoRAs and overnight fine tunes turn every deployment into a conversation between AI and infrastructure engineers. Superlinked's answer is open source under Apache 2.0 from control plane to GPU: a gateway annotates a request without fully parsing it and drops it into a shared queue, and workers pull and form their own batches. That inversion doubled cluster throughput. A Rust sidecar abstracts fifty adapters over three runtimes, an autoresearch loop ships every model already tuned, and one output was a LoRA that cost 80 cents and lifted retrieval on German legal text by 18 percent. Speaker info: - https://x.com/svonava Timestamps: 0:00 - Small open source models, do it yourself 2:30 - Small models are catching the frontier 3:42 - One task, one model: a nine model contract agent 6:37 - Serving tools are do it yourself research projects 7:31 - Why top down routing chokes on small requests 8:42 - LoRAs, fine tunes, and th

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Small open-source models are now capable enough for narrowly defined production tasks, but realizing their cost and latency advantages requires a cluster architecture built for large, heterogeneous fleets rather than conventional single-model inference serving.
  • Why it matters: The talk offers a concrete control-plane and scheduling pattern for self-hosting many task-specific models, LoRAs, and fine-tunes without turning model deployment or GPU utilization into an infrastructure bottleneck.
  • Best use: Use it as an architecture reference for an internal small-model serving plane, especially for embeddings, document pipelines, retrieval, labeling, synthetic-data generation, and agent sub-tasks.

Executive Summary

Daniel Svonava argues that the relevant distinction is no longer simply open versus closed models, but generalized frontier models versus small, specialized open models. He defines small models as models that fit on a single, two- to three-generation-old NVIDIA GPU. His claim is that models such as Qwen3 27B are approaching the performance of prior frontier models such as GPT-5.1 for suitable workflows, while specialized models can outperform general models on narrow tasks because they are trained on more task-specific data.

The operational consequence is a shift from one general model endpoint to fleets of models. A contract-review agent may use nine different models; an organization may operate many such agents. Standard inference deployments are optimized for serving one large model across multiple GPUs, where a central router assigns requests to workers. Svonava says this design performs poorly for high-volume, small-model traffic because the routing decision is stale by the time it reaches workers and prevents efficient batch formation.

Superlinked's proposed answer is an Apache-2.0 open-source cluster with a lightweight gateway, centralized shared queue, and worker-driven pulling and batching. Workers select work based on queue state rather than receiving pre-routed requests. The stack uses Rust components, NATS JetStream for the queue, MessagePack rather than base64 JSON for high-throughput multimodal requests, and runtime adapters for PyTorch, Candle, and SGLang. The stated purpose is to separate control-plane concerns from model-runtime diversity and eliminate day-to-day coordination bottlenecks between AI and infrastructure teams.

The most immediately actionable workload is embeddings: Svonava claims a single RTX Pro 6000 can process roughly 500,000 embedding tokens per second with low-tens-of-milliseconds latency, versus managed API latencies in the hundreds of milliseconds. He also identifies controllable batch jobs—synthetic data, annotations, fine-tuning data, and evaluations—as strong candidates for self-hosted task-specific generation. The important qualification is that the performance figures are presenter-reported benchmarks and the approach still demands serious per-model tuning, evaluation, memory management, and operational ownership.

Key Takeaways

  • Claim: Small open models are becoming viable substitutes for earlier frontier models when work is decomposed into narrow, evaluable tasks rather than treated as fully general prompting. | Evidence: Svonava cites Qwen3 27B as being around GPT-5.1-level performance on the referenced intelligence index, while arguing that frontier-model gains are flattening and smaller models are catching up. | Implication: Ken should treat model routing as a task-level portfolio problem: reserve frontier models for ambiguity and hard reasoning, while moving stable sub-workflows to specialized self-hosted models after benchmark validation. | Caveat: This is not a claim that a 27B model can replace a frontier model across arbitrary tasks; the proposed operating model requires task slicing, model selection, evaluation, and sometimes adaptation per task.
  • Claim: Specialized open models can outperform broad frontier models on constrained domains because they are trained or fine-tuned on highly concentrated task data. | Evidence: Examples include OCR for Vietnamese receipts, document question answering, image labeling, SQL generation, and code review; the speaker notes that Hugging Face contains hundreds of thousands of such task-focused models. | Implication: For document and retrieval systems, model discovery and domain-specific evaluation may create more advantage than merely choosing the largest available general LLM. | Caveat: Model availability and permissive licensing are not universal, and task-specific benchmark strength does not establish production reliability on an organization's own data distribution.
  • Claim: Conventional top-down inference routing is structurally inefficient for a large fleet of fast, small models under sustained load. | Evidence: The speaker reports difficulty getting vLLM and SGLang router deployments beyond 20-30% GPU utilization for this traffic pattern, attributing it to stale worker-state visibility and poorly sized batches. | Implication: Before scaling GPU count, Ken should test whether scheduling and batching—not raw capacity—is the binding constraint for high-QPS small-model workloads. | Caveat: The utilization figure is based on Superlinked's experiments, with workload and configuration details not supplied; custom runtime plugins could potentially close some of the gap.
  • Claim: A centralized shared queue with worker-pull scheduling can materially improve cluster saturation because each worker forms batches based on current local conditions. | Evidence: Superlinked's topology uses a gateway to annotate requests and place them into a shared NATS JetStream queue; workers pull jobs rather than receiving pre-assigned work. Svonava claims centralizing the queue can double cluster throughput, not merely improve it by 5%. | Implication: A control plane for heterogeneous agent-model traffic should favor pull-based scheduling, local coordination on multi-GPU nodes, and adaptive batch sizing over globally precomputed routing decisions. | Caveat: Optimal batch size remains difficult to predict. Returning excess work to a network queue adds milliseconds, so Superlinked only enables finer back-and-forth queue negotiation among GPUs co-located on the same machine.
  • Claim: Runtime diversity and rapid model adaptation require an abstraction layer rather than embedding deployment logic into one inference runtime. | Evidence: Superlinked supports roughly 50 model adapters and uses Rust gateway/worker sidecars connected by local sockets to PyTorch, Candle, or SGLang. The speaker says custom SGLang plugins could match some performance, but would lock the implementation to SGLang. | Implication: Ken should avoid making an agent platform dependent on a single serving runtime if frequent model swaps, LoRAs, and custom fine-tunes are strategic; establish a stable internal serving contract above runtime-specific implementations. | Caveat: This abstraction itself is additional platform surface area to maintain, and Candle was still materially behind PyTorch performance in the speaker's testing.
  • Claim: Embeddings and controlled offline generation are the clearest early wins for self-hosting small models. | Evidence: On an RTX Pro 6000, Svonava reports up to roughly 500,000 embedding tokens per second from one GPU and low-tens-of-milliseconds latency; he contrasts this with managed embedding APIs taking hundreds of milliseconds. He also recommends self-hosting synthetic-data and annotation generation because quality can be inspected and controlled. | Implication: Prioritize a self-hosted embedding pilot and batch-generation workloads before attempting latency-sensitive, generalized conversational inference. | Caveat: The claimed economics depend on maintaining sufficient utilization, selecting an adequate model, and accepting the operational burden of the serving stack; no all-in cost comparison is provided.
  • Claim: Small-model operations need dynamic GPU memory management and automated performance research, not static worker pools with one preloaded model each. | Evidence: The speaker recommends combining pinned models with lazy loading and eviction under memory pressure. Superlinked runs automated research loops to add model support, optimize implementations, and ship cluster-level tuned configurations; one cited LoRA reportedly cost $0.80 to train and improved German legal retrieval quality by 18%. | Implication: The durable capability is an eval-and-optimization pipeline that produces deployable configurations, rather than merely standing up GPUs or downloading models. | Caveat: The 18% improvement is a proof of concept on German legal retrieval, not evidence that inexpensive LoRAs reliably create similar gains across domains.

Detailed Brief

Request-path design for high-throughput multimodal inference

  • Claims: The gateway should perform minimal work because it is the first shared component exposed to incoming traffic and can become the system bottleneck.; The API should avoid repeatedly serializing and deserializing requests across internal components.; A single API boundary for both request metadata and multimodal payloads simplifies clients while keeping heavy data out of the core scheduling path.
  • Evidence: Superlinked uses MessagePack, a binary format, rather than OpenAI-style base64-encoded JSON.; For payloads over approximately 1 MB, the gateway separates heavier binary content and defers it to cloud storage while the request remains in flight.; The shared queue is NATS JetStream, which the speaker says can support approximately one million requests per second.
  • Caveats: The discussion does not cover authentication, tenant isolation, data residency, cloud-storage access controls, or failure handling for deferred payload retrieval.
  • Implications: For multimodal agent infrastructure, transport encoding and payload handling can become material throughput design decisions rather than incidental API details.; A clean external API can coexist with a differentiated internal data path that protects schedulers from oversized payloads.

Operational failure mode: organization, not only compute

  • Claims: The main drag on organizational velocity is the human handoff required whenever AI engineers need infrastructure teams to deploy new LoRAs or fine-tunes.; Small models make frequent adaptation economically attractive, which increases the frequency of this handoff unless the serving platform supports self-service deployment.
  • Evidence: The speaker frames recurring requests as: an AI engineer has ten LoRAs, or built a fine-tune overnight, and needs infrastructure to make it production-serving.; He characterizes open-source serving tools as an ongoing research project because they are not pre-tuned for each hardware and model combination.
  • Caveats: Self-service model deployment should still be governed by automated evaluations, provenance, approval controls, rollback procedures, and resource quotas; these controls are not detailed in the talk.
  • Implications: A model-serving platform should make the safe path fast: package model artifacts, attach eval evidence, apply validated serving profiles, and promote or roll back without bespoke infrastructure tickets.

Notable Concepts & Terms

  • Small models: Models that fit on one relatively older, readily available NVIDIA GPU; the talk positions them as economical building blocks for task-specific systems.
  • Task slicing: Decomposing generalized LLM workloads into narrower tasks, then selecting and evaluating the best specialized model for each task.
  • Pull-based worker scheduling: Workers pull requests from a shared queue and form local batches, replacing a router that preassigns requests based on imperfect global state.
  • The knee: The saturation point on a throughput curve where additional requested load no longer raises throughput and instead raises latency; it is the meaningful operating point for benchmark comparisons.
  • LoRA: A lightweight adaptation method that makes small-model specialization cheap, but creates deployment and serving-version-management pressure.
  • Runtime abstraction / Rust sidecar: Superlinked's method of supporting diverse model runtimes—PyTorch, Candle, and SGLang—without hardwiring platform logic into any single runtime.
  • Lazy loading and eviction: Loading models on demand and evicting them under GPU memory pressure, which the speaker presents as more appropriate for small-model fleets than static one-model worker pools.
  • Auto research loop: An automated measurement and optimization system used to build model support, tune performance, and ship validated cluster configurations rather than requiring repeated manual parameter sweeps.

Operator Notes / Why Ken Should Care

  • Run a representative benchmark that compares managed embeddings against a self-hosted open embedding model on quality, p50/p95 latency, utilization, and fully loaded cost—not token price alone.
  • Design the internal inference API as a runtime-independent control-plane contract so model teams can add models, LoRAs, and fine-tunes without platform-specific rewrites.
  • Evaluate pull-based shared-queue scheduling for high-QPS heterogeneous workloads; measure GPU utilization and latency at the saturation knee against current routing behavior.
  • Require an automated promotion package for every adapted model: source/license record, domain eval results, serving configuration, capacity estimate, rollback target, and owner.
  • Treat multimodal payload routing, storage access, and queue isolation as a security review area before adopting the proposed MessagePack/shared-queue architecture.
  • Inspect Superlinked's Apache-2.0 cluster repository as an implementation reference, but independently validate its throughput and operational claims on target models and hardware.

Source/Metadata

  • Title: Large clusters for small models — Daniel Svonava, Superlinked
  • Transcript words: 6429
  • Duration seconds: 1506
  • Timestamp note: No usable timestamps or chapter markers were present in the supplied transcript; substantial portions of the talk were duplicated.

Transcript

3815 words en Processed in 161.3s

All right. I think you guys can hear me. I can certainly hear myself. Whoever came closer gets a T-shirt. I meant it. There's a bag full of T-shirts over here. And also for questions. Maybe there will be some questions at the end. If you ask a question, you'll get a T-shirt as well. And if you can guess what is on the background of this slide, you get a T-shirt as well. Any guesses? What does that visualize? This picture in the background? No? Anybody has seen a transformer model? Yeah, positional encoding. Very good. Very good. You get the T-shirt, sir. All right. So today we'll discuss small open source models and how they're pretty good now and how they create unique challenges when you want to serve a bunch of them in your own cloud. Everything we'll discuss is open source, do it yourself. This is the kind of stuff you can just run a command and own the stack. So there is no proprietary pieces of the puzzle here. Let's get this underway. Wow, this works. Okay. So small models. What do we mean by small models? Depending who you ask, the way I think about it is models that you can run on two, three generations old NVIDIA hardware. The whole model fits into one GPU. And therefore they are easy to serve. Those GPUs are available and they are affordable as well. And then most people think, okay, small models, there will be some kind of trade-off in terms of quality of the results. And hopefully I'll be able to do a good job in this talk to convince you that actually for specific tasks, you can be at Frontier or Beyond Frontier performance and get all the other obvious benefits, right? Orders of magnitudes of cost savings and potentially quite big latency or throughput improvements, of course. So this is one of the charts we like to show. This is the artificial analysis intelligence index over time. And what they typically don't show you is that there is a breakdown of the open source models you should think about, right? There is the GLM 5.2 and so on. Those Frontier open source models with, let's say, 750 billion parameters. But then there are the small open source models trailing the big ones and trailing the Frontier. You can see the Frontier is getting diminishing returns these days. And the small models are catching up, right? So you see this convergence, saturation on top, and growth of the small models. And let's say QN3627B somewhere around the performance of GPT 5.1. So if you have a workflow, if you have a pipeline that can run with GPT 5.1, now you can move it to a small model and get all the benefits we discussed. So small models, not dumb anymore. Now, it is also about how you use the small models, right? So you can't just read that 27 billion parameter QN36 as your totally generalized, I can prompt you to do anything kind of model. Now you need to adopt the approach where you basically figure out slice of tasks from the generalized model workload. And then per task, you figure out which model in the open source fits the task the best. You run some evals, maybe some adaptation we'll discuss. And then that's how you reach the right quality to actually push this into production. So here is some example of a contract review agent that uses nine different models. This is the shape that you will see in your workloads, in your agents, as you move to using small models for your setup. You'll start to see that instead of hammering one API with a bunch of different requests or one model, you would rather use a fleet of models and then your problem is, okay, how do I serve all of these different things in a way that my infra people don't go crazy, right? And this is just one of the agents that you might be running. And there might be 10 of these in your company. So how do we sort of expand that scope of infrastructure, let's say. Now, all of those different tasks that I mentioned, there is an open source model that's sitting there waiting to be used. From OCR to question answering on top of documents to labeling images, generating SQL, reviewing code. There are open source models fine tuned and trained for those tasks. You know, if you use an open source model that's trained to do OCR on receipts in Vietnamese, that project has seen the most receipts in Vietnamese, right? There's somebody who took the time to gather as much data as possible. And on that task, that model will outperform pretty much anything else. And there are hundreds of thousands of models on hugging face that look like that, right? So it's all sitting there and it's all free, mostly quite permissive licenses. So the models exist. There's not the bottleneck. And we have been talking about open source AI since 2024. And it's so far still not really happening. And to the extent it's happening in companies, it basically equals open source AI equals AWS Bedrock. Except when you look at the model catalog in Bedrock, it's very restrained in model types that are available. These models are old, often two, three years behind the state of the art. And when you do any kind of fine tuning in Bedrock, you don't actually own the fine tuned or trained artifacts. So you can't use it as an actual advantage in your business. It stays serving from the Bedrock infra. So on the proprietary side, now if you do small models serving on open source infrastructure, VLLM, SGLang, different solutions, just know that these things are not tuned for any specific model or any specific hardware model combination. You'll have to do the tuning, right? This is the do-it-yourself. All of these tools ship with guides on how to actually do the tuning, the parameter sweep, tailoring to your traffic, and so on. This is an open-ended research project every time you try to adopt one of these tools. So this is not something that you take and it's an engineering project, and a week later you have a high-performance serving infrastructure. It doesn't work like that. And that's the typical problem with open source tools, right? It's a little too much do-it-yourself. And then on top of this not being pre-tuned for small models, the small model workloads and traffic that uses a bunch of different models flips the equation for inference clusters, right? So normally when you try to serve one big model, your problems are how do I share that model across multiple GPUs? How do I have a router sitting on top that understands the state of all these workers, the KVCache state and so on, and then makes a top-down routing decision of, okay, this request goes to this worker or this group of workers and so on, right? It's very top-down setup. But if you have small and fast requests and you have many of them, this top-down routing becomes the bottleneck, right? Because the router has a little bit obsolete version of the worker state, and it's just really hard to saturate the workers if you have that upfront decision on top that has to get it perfectly right in terms of balancing the local queues on each of these workers because there are many small requests, right? And we have experimented with the VLLM and SG-Lang routers for small models and this sort of traffic, and it's very hard to get your GPU utilization beyond 20%, 30% under constant load. And the problem is that those batches are just not correctly sized, because you have that routing bottleneck. And then the third problem is that with small models, you benefit a lot from LORAS and model adaptation in general, and so the traffic that you have to serve contains people coming to you and saying, hey, I have 10 LORAS, how do I use this with our serving stack? Or I have this custom fine tune I made last night, I want to serve this in production. And this conversation between the AI engineer and the infrastructure person in getting those LORAS up there, custom models up there, that's the thing that takes time. And that's the main killer in organizational velocity is talking, right? Ideally, you would want the infrastructure engineers to do their job, and you would want those AI engineers to do their job, and they don't have to talk to operate on the day-to-day mode, so they're not blocking each other. And this model adaptation desire around small models breaks that and creates a lot of back and forth, and that's a problem, right? Hey, I have 10 LORAS, how do I use this with our serving stack? Or I have this custom fine tune I made last night, I want to serve this in production. And this conversation between the AI engineer and the infrastructure person in getting those LORAS up there, custom models up there, that's the thing that takes time. And that's the main killer in organizational velocity is talking, right? Ideally, you would want the infrastructure engineers to do their job, and you would want those AI engineers to do their job, and they don't have to talk to operate on the day-to-day mode, so they're not blocking each other. And this kind of model adaptation desire around small models breaks that and creates a lot of back and forth, and that's a problem, right? So these are some challenges related to: okay, we have a bunch of small models, how do we have a cluster, how do we serve this efficiently? So we have been playing with this problem for a while. I'm Daniel, actually from Superlinked. I kind of skipped the intro. So we are a VC-backed company out of SF, and we have been building AI-powered search and document processing systems and agents for the last couple of years. And our main pain point has always been inference, specifically these problems that I have described. And so we have iterated and iterated and explored different topologies for clusters for running large wide fleets of small models in different environments, because sometimes you need to deploy together with some platform in some environment where who knows what is available there. The small models make it easier because in whatever environment you can get some L4s or some kind of small GPU quota is much easier. So I'll describe a little bit about the topology of the cluster that we have kind of converged to. And by the way, this whole thing is Apache 2.0, completely open source. You guys can just take it and wrap it, and now you are an inference startup. This is open source from the control plane all the way down to the thing that runs on the GPU. So we didn't pull any punches. And the topology is basically there is a gateway, and instead of having a router that pre-decides what goes where, there is a gateway that parses some of the requests and attaches some metadata to the request, inserts that request into a shared queue, and into some side channels. I'll go a little bit into that. And then the workers pull from that centralized queue instead of pushing the data down to the workers. And this way they can saturate themselves better. And then the worker setup, I think I have a slide for that, will describe how we basically absorb the complexity of different model architectures into a coherent set of workers that don't have competing Python requirements and stuff. So that's the overall topology. And this is the life of a request. So maybe I'll call out a couple of things from here. One of the things we don't like about the OpenAI API standard is the base64 encoded JSON. Not good for small models, not good for high throughput. So we use message pack throughout, a binary format. This way we can also push all the multimodal data through the actual API gateway. So there is no binary data over here and then request over there. And then the cluster needs access to your cloud storage to start loading some binary data, images or videos. We kind of encode it all and we push it through the gateway. And then the gateway separates some of these heavier pieces to not clog the internal queue and defers it on cloud storage in-flight while the request is in queue. So it splits up some of these requests that are over a megabyte and then uses cloud storage in the backend. But as a user, you push all your bits and bytes into the API layer and it's a clean interface because of that. Basically the whole stack is Rust. So gateway Rust, the worker is Rust. And then over a socket locally, it attaches to different runtimes. And we have basically PyTorch, Candle and SGLang on the runtime. And then when we do the optimization, I'll go into that on how we make sure that whichever runtime we are using and whichever code is running in that runtime is the most efficient one. We have an auto research loop for that. So the life of a request looks like that. And one tidbit is that you really want to make sure that the gateway that's the first thing that's hit by the request doesn't do too much work. Because then it becomes a bottleneck. So you don't even want to parse the whole request. You want to be able to look at the packets and figure out the general shape of what's coming, do the annotation, and then you have the workers. However many workers you have, hundreds of GPUs that look at the queue state and then pull from there. And the queue uses NATS Jetstream and that thing can do a million requests per second. That's very hard for that to become a bottleneck. So you don't want to serialize, deserialize as you go through all of these different components. That's basically the obvious thing. This is an animation that shows the idea behind the centralized queuing. So instead of the top-down router trying to fill in the local queues just right, which is basically impossible, the whole idea is, can we somehow centralize the queuing and can the workers rather pick up the task of forming their own batches with their own prediction of the cost of the batch and then become much more efficient. Now one tidbit and side note: once you start working on these things, you realize that it's actually really hard to predict how many things to pick up from the shared queue for the batch to be really optimal. And so you would want some mechanism that allows you to put some things back into the queue if you figure out: oh, I pulled a little bit too much. And that's a network hub, right? So that's a problem. And we have special optimization for that for machines that have multiple GPUs locally. So there is additional machine local queuing element that takes advantage of the fact that the local processes that run on the multiple GPUs on one machine can negotiate with the queue a little bit back and forth, which over the network would add milliseconds. And so we don't do it over the network, only when we co-locate the workers on multi-GPU machines. And we are not talking about 5% differences here. You centralize the queue and now you get double the throughput of the cluster. So this is significant. I mentioned three different runtimes. So basically it's either we write, for models that are encoder only, we write the PyTorch code and we optimize it. And we have an auto research loop that optimizes it. Same for Candle. We started to play with Candle not too long ago. We still can't get it to perform anywhere near the PyTorch performance. So it's more of a research project. It's just the dependency. The worker Docker image with PyTorch is like 12 gigabytes and the worker basically binary statically linked binary with Candle is maybe 10% of that. And if you care about waking up from the cold state and loading these images on a bunch of different machines, going from 12 gigs to a gigabyte or something makes a huge difference. So that's the motivation behind Candle is just getting the same performance from PyTorch is really hard. And then SGLang we have there as a go-to baseline. We should perform at least as well as SGLang with optimal tuning of all of those parameters that you have to do tuning on. Here are some numbers. So for example, when we wrap SGLang with the socket and with our Rust sidecar, we can actually improve on the bare SGLang performance just because we do something on the batching side that natively you can't make it do that. If you develop custom plugins into SGLang, probably you can match our performance because you can just push the same logic into the SGLang core server. But now you are developing custom code that only works with SGLang. And the whole lesson here from small models is that the runtimes are super diverse, right? You don't want to necessarily get stuck with any one particular runtime because we have on the order of 50 different adapters now that we parametrize for the different models. # Transcript Here are some numbers. So for example, when we wrap SG Lang with the socket and with our Rast sidecar, we can actually improve on the bare SG Lang performance just because we do something on the batching side that natively SG Lang does not do that. If you develop custom plugins into SG Lang, you can probably match our performance because you can just push the same logic into the SG Lang core server. But now you are developing custom code that only works with SG Lang. And the whole lesson here from small models is that the runtimes are super diverse, right? You don't want to necessarily get stuck with any one particular runtime because we have, I think, on the order of 50 different adapters now that we parametrize for the different models. And so you need to somehow deal with this underlying complexity. And it's probably not by building a bunch of plugins for one specific runtime. It's probably some kind of abstraction, which in our case is this Rast sidecar concept and then the socket. Now I'll talk about a couple of different numbers, but in terms of language around benchmarking, the knee is this concept of when you ramp up traffic on the server, when you request more and more throughput from it, and it gives you more and more throughput, that's when you go linearly up. And then at some point you hit a point where asking for more is not coming. So you flatten out and the latency goes up. So we call that the knee and it's a useful concept in benchmarking because that's the point of saturation, right? That's the maximum performance without hurting latency. So just to give you some ideas of what is possible on relatively small hardware, right? And different types of small models. This is measured on the RTX Pro 6000. We work with NVIDIA L4, A100, RTX Pro 6000, H100, that sort of range. Those GPUs are much more readily available, on demand in any cloud, and most continents have quota. On this kind of stuff, you can basically get, for embedding models, even up to hundreds of millions of parameters, you can get hundreds of thousands of tokens per second and code it into the embedding, right? So imagine you are sitting there now hitting your text embedding three on OpenAI API. Instead, you could be having one GPU and push half a million tokens per second into the thing and get the vectors out, right? So you have half a million tokens that you are pushing into something. You have a single GPU that's not even that big per second and you are getting out vector embeddings for your search system. As opposed to pushing all of that into a managed embeddings endpoint somewhere and paying orders of magnitude more money, right? And you can get latencies of low tens of milliseconds for these calls, right? If you use Cohere, OpenAI APIs and so on, these are hundreds of milliseconds, right? And this is not rocket science. You can have massive cost savings, massive latency improvements and relatively easy operation with a handful of GPUs and some info around them. Right? So this is a really low hanging fruit. If you start anywhere with open source models, small models embeddings are a no brainer, right? But it doesn't end there. So let's say you want to look at named entity recognition. You want to look at multi-vector search, or even generation, right? Of text or structured outputs and so on. You can be getting thousands of tokens per second output from task specific generative models as well per, let's say, half a thousand per second for one GPU there at the bottom. And so let's say you are generating synthetic data. You are generating annotations for your fine tuning, for your evals. Don't do that on a managed endpoint. That's a perfect task because you have it under control. You can survey the quality. That's a perfect task for open source model on your own infra. And then if the infra you have around those GPUs is reasonable, you'll get linear scaling with the number of those GPUs. Now another idea if you are into small model serving is that you don't normally have a worker pool per model, right? You have a set of workers, set of nodes, they have GPUs. You bring those up. You preload the models. The models load for tens of minutes because there are hundreds of billions of parameters. And so you are happy. Okay. They finally loaded. Now I have a worker pool. This mentality doesn't really work with small models. Yeah. How much time is left? Six minutes over. Okay. All right. So pack models on the same GPU is faster. This is a story of how you still want to pin some models, but you want to also do basically lazy loading and eviction as a function of memory pressure. You want to figure out how to combine the two. There is a little bit about auto research. We have auto research loops for adding support for new models and for their performance. We build a lot of internal tooling to do the measurement, to feed into those auto research loops, to basically push the numbers forward. And maybe most importantly, when we ship support for a model, it has all the tuning done, right? So there is no parameter sweep. We bundle basically a config for end to end the whole cluster. This is a setup for the auto research loop. That is a meta loop that builds the harness that then runs the loop. And there is a dashboard on top that helps you understand how it works. We have custom UIs for that. And one of the outputs of that was a lot of that took 80 cents to train and it improved 18%, it improved quality of retrieval on German legalistic as a proof of concept by 18%. And that's it. So small models are good. They're relatively easy to serve. They are actually much cheaper, faster, faster, as smart, and that QR code goes to the GitHub repo of our cluster that I just described. Give us a star, and happy self-hosting. Thank you. This is open source from kind of the control plane all the way down to the thing that runs on the GPU. So we didn't pull any punches. And the topology is basically there is a gateway, and instead of having a router that kind of pre-decides what goes where, there is a gateway that parses some of the requests and attaches some metadata to the request, inserts that request into a shared queue, and into some side channels. I'll go a little bit into that. And then the workers pull from that centralized queue instead of kind of pushing the data down to the workers. And this way they can saturate themselves better. And then the worker setup, I think I have a slide for that, will describe how we basically absorb the complexity of different model architectures into kind of a coherent set of workers that, you know, don't have like competing Python requirements and stuff like that. So that's kind of the overall topology. And this is kind of life of a request. So maybe just I'll call out a couple of things from here. We, you know, one of the things we don't like about the OpenAI kind of API standard is the base64 encoded kind of JSON. Not good for small models, not good for high throughput. So we use a message pack throughout, like a binary format. This way we can also push all the multimodal data through the actual API gateway. So there is no like, hey, you know, binary data over here and then request over here. And then the cluster needs access to your cloud storage to start loading some binary data, images or videos. We kind of encode it all and we push it through the gateway. And then the gateway kind of separates some of these heavier pieces to not clog the internal queue and defers it on cloud storage kind of in-flight while the request is in queue. So it kind of splits up some of these requests that are, let's say, over a megabyte and then uses cloud storage in the backend. But as a user, you push all your bits and bytes into the API layer and it's kind of clean interface because of that. Basically the whole stack is RAST. So gateway RAST, the worker is RAST. And then over a socket locally, it kind of attaches to different runtimes. And we have basically PyTorch, Kendall and SGLang on the, as a runtime. And then when we do the optimization, I'll kind of go into that on how we make sure that whichever runtime we are using and whichever code is running in that runtime is the most efficient one. We have an auto research loop for that basically. But yeah. So life of a request kind of looks like that. And like one tidbit is that you really want to make sure that the gateway that's kind of the first thing that's hit by the request doesn't do too much work. Because then it becomes a bottleneck, right? So you don't even want to parse the whole request. You want to be able to kind of look at the packets and figure out the general shape of what's coming, do the annotation, and then you have the workers. However many workers you have, hundreds of GPUs that look at the queue state and then pull from there. And the queue use NATS Jetstream and that thing can do, you know, million requests per second. Like that's very hard for that to become a bottleneck. So yeah. Like ideally you don't want to serialize, deserialize as you go through all of these different components. That's basically the kind of obvious thing. This is a little animation that shows the idea behind the centralized queuing. So instead of the top-down router trying to, you know, fill in the local queues just right, which is basically impossible, you know, the whole idea is, hey, can we somehow centralize the queuing and can the workers rather pick up the task of forming their own batches with their own prediction of the cost of the batch and then, you know, become much more efficient. Now one tidbit and kind of side note, once you kind of start working on these things, you realize that it's actually really hard to predict how many things to pick up from the shared queue for the batch to be really, like really the optimal size. And so you would want some mechanism that sort of allows you to put some things back into the queue if you figure out, oh, like I pulled a little bit too much. And that's a network hub, right? So that's a problem. And we have special optimization for that for machines that have multiple GPUs locally, right? So there is additional kind of machine local queuing element that takes advantage of the fact that the local processes that run on the multiple GPUs on one machine can kind of negotiate with the queue a little bit back and forth, which over the network, you know, there is like milliseconds extra that that would add. And so we don't do it over the network only when we co-locate the workers on multi GPU machines. And, you know, I mean, we are not talking about like 5% differences here, right? So like you centralize the queue and now you get double the throughput of the cluster. So this is significant. I mentioned three different runtimes. So basically it's either, you know, we write, let's say for models that are encoder only, we write the PyTorch code and we kind of optimize it. And we have an auto research loop that optimizes it. Same for candle. We started to play with candle not too long ago. We still can't get it to perform anywhere near the PyTorch performance. So it's a little bit more of a research project. It's just the dependency. Like, you know, the worker Docker image with PyTorch is like 12 gigabytes and the worker basically binary statically linked binary with candle is maybe like 10% of that, right? And if you care about kind of waking up from the cold state and loading these images on a bunch of different machines, the, you know, going from 12 gigs to a gigabyte or something like this makes a huge difference. So that's kind of the motivation behind candle is just the, the getting the same performances from PyTorch is really hard. And then SG Lang we have there as a kind of a go to baseline. Like we should perform as, at least as well as, as SG Lang with the optimal tuning of all of those parameters that I mentioned that you have to do the tuning. Um, here is some numbers. So for example, when we wrap SG Lang with the socket and with our kind of Rast sidecar, um, actually we can improve on the bare SG Lang performance just because we kind of, uh, do something to do like on the batching side that natively as you can make it to do that. If you do like, if you develop custom plugins into SG Lang and stuff like that, like probably you can match our performance because you know, you can just push the same logic into the SG Lang core server. Uh, but now you are developing custom code that only works with SG Lang. And the whole lesson here from small models is that the runtimes are super diverse, right? You don't want to necessarily get stuck with any one particular runtime because there is, you know, we have, I think on the order of 50 different adapters now that, that we parametrize for the different models. And so you need to somehow deal with this kind of underlying complexity. And it's probably not by building a bunch of plugins for one specific runtime. It's probably some kind of abstraction, uh, which in our case is this, uh, Rast sidecar concept. And then the socket. Um, now I'll talk about a couple of different numbers, but in terms of like language around benchmarking, you know, the knee is this concept of like when you ramp up traffic on the server, uh, when you sort of request more and more throughput from it, and it gives you more and more throughput. That's when you go kind of linearly up. And then some point you hit this point where you kind of ask for more and more is not coming. So you kind of flatten out and the latency goes up. So we call that the, the knee and it's, it's like a useful concept in, in benchmarking, uh, because that's kind of the point of saturation, right? That's, that's kind of the maximum performance without hurting latency. Um, so just to give you some ideas of what is possible on relatively small hardware, right? And different types of small models. So this is measured on the RTX pro 6000. We, we kind of work with NVIDIA L4, you know, A100, RTX pro 6000, H100, that sort of range. Um, again, those GPUs are much more readily available, kind of on demand in any cloud, basically most continents have quota, you know, um, and on this kind of stuff, uh, you can basically get, uh, for embedding models, even up to, let's say, uh, hundreds of millions of parameters. You can get hundreds of thousands of tokens per second and code it into the embedding, right? So imagine you are sitting there now, like hitting your text embedding three on open AI API. Instead, you could be like having one GPU and push half a million tokens per second into the thing and get the vectors out. Right? Like, is this like a connecting, right? You have half a million tokens that you are pushing into something. You have a single GPU that's not even that big per second and you are getting out vector embeddings for your search system. As opposed to like pushing all of that into a managed embeddings endpoint somewhere and paying like orders of magnitude more money. Right? And you can get latencies like, you know, low tens of milliseconds for these calls, right? Like if you use, uh, you know, cohere, open AI APIs and so on, these are hundreds of milliseconds, right? And this is not rocket science. You know, you can have just like massive cost saving, massive latency improvements and relatively easy, easy operation. Um, with, with like handful of GPUs and some, some info around them. Right? So this like really low hanging fruit, if you start anywhere with open source models, small models embeddings are like no brainer. Right? Uh, but it doesn't end there. So let's say, uh, you want to look at, uh, named entity recognition. You want to look at, let's say multi-vector search, uh, even generation, right? Uh, of text or structured outputs and so on. Um, you, you can be getting, you know, thousands of tokens per second output from, uh, you know, task specific generative models as well per, uh, like let's say half a thousand per second. For, for one GPU there at the bottom. Um, and so let's say you are generating synthetic data. You are generating annotations for your fine tuning, for your evals. You know, don't do that on a, on a managed endpoint. That's a perfect task because you have it kind of under control. You can survey the quality. That's a perfect task for, uh, open source model on your own infra. Um, and then you're like, if the infra you have around those GPUs is like reasonable, you'll get linear scaling with, with the number of those GPUs. Um, now another sort of, uh, idea if you are into small model serving, uh, is that you don't, you know, normally, um, you have kind of worker pool per model, right? You have a set of, uh, workers, set of nodes, uh, they have GPUs. You kind of bring those up. You preload the models. The models load for tens of minutes because there are hundreds of billions of parameters. Uh, and so you are happy. Okay. They finally loaded. Now I have a worker pool. This mentality doesn't really work with small models. Yeah. Yeah. Quickly. How, how, what's the time left? Oh, six minutes over. Okay. All right. So pack models on the same GPU is faster. Um, this is a story of how you still want to pin some models, but you want to also do, uh, basically, uh, lazy loading and eviction, uh, as a kind of function of memory pressure. You want to figure out how to combine the two. Um, there is a little bit about kind of auto research. We have auto research loops for, uh, adding support for new models and for their performance. Um, we build a lot of internal tooling to do the measurement, to feed into those auto research loops, to basically push the numbers forward. Um, and maybe perhaps most importantly, when we ship support for a model, it has all the tuning done, right? So there is no, okay, let's do a parameter sweep. We bundle basically a config for end to end the whole cluster. Um, this is a setup for the auto research loop. That is like a meta loop that builds the harness that then runs the loop. And there is a dashboard on top that helps you understand how it works. Um, we have custom UIs for that. And one of the outputs of that was a lot of that took 80 cents to train and it improved 18%, it improved quality of retrieval on German legalistic as a proof of concept by 18%. And that's it. So small models are good. They're relatively easy to serve. Uh, they are actually much cheaper, faster, faster, is as smart, and that QR code goes to the GitHub repo of our cluster that I just described. Give us a star, and happy self-hosting. Thank you. . . . . . .