All right. I think you guys can hear me.
I can certainly hear myself. Whoever came closer gets a T-shirt. I meant it.
There's a bag full of T-shirts over here. And also for questions. Maybe there will be some questions at the end. If you ask a question, you'll get a T-shirt as well. And if you can guess what is on the background of this slide, you get a T-shirt as well. Any guesses? What does that visualize? This picture in the background? No?
Anybody has seen a transformer model? Yeah, positional encoding.
Very good. Very good.
You get the T-shirt, sir. All right. So today we'll discuss small open source models and how they're pretty good now and how they create unique challenges when you want to serve a bunch of them in your own cloud. Everything we'll discuss is open source, do it yourself. This is the kind of stuff you can just run a command and own the stack. So there is no proprietary pieces of the puzzle here. Let's get this underway. Wow, this works. Okay. So small models. What do we mean by small models? Depending who you ask, the way I think about it is models that you can run on two, three generations old NVIDIA hardware. The whole model fits into one GPU.
And therefore they are easy to serve. Those GPUs are available and they are affordable as well. And then most people think, okay, small models, there will be some kind of trade-off in terms of quality of the results. And hopefully I'll be able to do a good job in this talk to convince you that actually for specific tasks, you can be at Frontier or Beyond Frontier performance and get all the other obvious benefits, right? Orders of magnitudes of cost savings and potentially quite big latency or throughput improvements, of course. So this is one of the charts we like to show. This is the artificial analysis intelligence index over time.
And what they typically don't show you is that there is a breakdown of the open source models you should think about, right? There is the GLM 5.2 and so on. Those Frontier open source models with, let's say, 750 billion parameters. But then there are the small open source models trailing the big ones and trailing the Frontier. You can see the Frontier is getting diminishing returns these days.
And the small models are catching up, right? So you see this convergence, saturation on top, and growth of the small models. And let's say QN3627B somewhere around the performance of GPT 5.1. So if you have a workflow, if you have a pipeline that can run with GPT 5.1, now you can move it to a small model and get all the benefits we discussed. So small models, not dumb anymore. Now, it is also about how you use the small models, right? So you can't just read that 27 billion parameter QN36 as your totally generalized, I can prompt you to do anything kind of model.
Now you need to adopt the approach where you basically figure out slice of tasks from the generalized model workload. And then per task, you figure out which model in the open source fits the task the best. You run some evals, maybe some adaptation we'll discuss. And then that's how you reach the right quality to actually push this into production. So here is some example of a contract review agent that uses nine different models. This is the shape that you will see in your workloads, in your agents, as you move to using small models for your setup.
You'll start to see that instead of hammering one API with a bunch of different requests or one model, you would rather use a fleet of models and then your problem is, okay, how do I serve all of these different things in a way that my infra people don't go crazy, right? And this is just one of the agents that you might be running. And there might be 10 of these in your company. So how do we sort of expand that scope of infrastructure, let's say. Now, all of those different tasks that I mentioned, there is an open source model that's sitting there waiting to be used. From OCR to question answering on top of documents to labeling images, generating SQL, reviewing code.
There are open source models fine tuned and trained for those tasks. You know, if you use an open source model that's trained to do OCR on receipts in Vietnamese, that project has seen the most receipts in Vietnamese, right? There's somebody who took the time to gather as much data as possible. And on that task, that model will outperform pretty much anything else. And there are hundreds of thousands of models on hugging face that look like that, right? So it's all sitting there and it's all free, mostly quite permissive licenses. So the models exist. There's not the bottleneck. And we have been talking about open source AI since 2024.
And it's so far still not really happening. And to the extent it's happening in companies, it basically equals open source AI equals AWS Bedrock. Except when you look at the model catalog in Bedrock, it's very restrained in model types that are available. These models are old, often two, three years behind the state of the art. And when you do any kind of fine tuning in Bedrock, you don't actually own the fine tuned or trained artifacts. So you can't use it as an actual advantage in your business. It stays serving from the Bedrock infra.
So on the proprietary side, now if you do small models serving on open source infrastructure, VLLM, SGLang, different solutions, just know that these things are not tuned for any specific model or any specific hardware model combination. You'll have to do the tuning, right? This is the do-it-yourself. All of these tools ship with guides on how to actually do the tuning, the parameter sweep, tailoring to your traffic, and so on. This is an open-ended research project every time you try to adopt one of these tools. So this is not something that you take and it's an engineering project, and a week later you have a high-performance serving infrastructure.
It doesn't work like that. And that's the typical problem with open source tools, right? It's a little too much do-it-yourself. And then on top of this not being pre-tuned for small models, the small model workloads and traffic that uses a bunch of different models flips the equation for inference clusters, right? So normally when you try to serve one big model, your problems are how do I share that model across multiple GPUs? How do I have a router sitting on top that understands the state of all these workers, the KVCache state and so on, and then makes a top-down routing decision of, okay, this request goes to this worker or this group of workers and so on, right?
It's very top-down setup. But if you have small and fast requests and you have many of them, this top-down routing becomes the bottleneck, right? Because the router has a little bit obsolete version of the worker state, and it's just really hard to saturate the workers if you have that upfront decision on top that has to get it perfectly right in terms of balancing the local queues on each of these workers because there are many small requests, right? And we have experimented with the VLLM and SG-Lang routers for small models and this sort of traffic, and it's very hard to get your GPU utilization beyond 20%, 30% under constant load.
And the problem is that those batches are just not correctly sized, because you have that routing bottleneck. And then the third problem is that with small models, you benefit a lot from LORAS and model adaptation in general, and so the traffic that you have to serve contains people coming to you and saying, hey, I have 10 LORAS, how do I use this with our serving stack? Or I have this custom fine tune I made last night, I want to serve this in production. And this conversation between the AI engineer and the infrastructure person in getting those LORAS up there, custom models up there, that's the thing that takes time.
And that's the main killer in organizational velocity is talking, right? Ideally, you would want the infrastructure engineers to do their job, and you would want those AI engineers to do their job, and they don't have to talk to operate on the day-to-day mode, so they're not blocking each other. And this model adaptation desire around small models breaks that and creates a lot of back and forth, and that's a problem, right?
Hey, I have 10 LORAS, how do I use this with our serving stack? Or I have this custom fine tune I made last night, I want to serve this in production. And this conversation between the AI engineer and the infrastructure person in getting those LORAS up there, custom models up there, that's the thing that takes time. And that's the main killer in organizational velocity is talking, right? Ideally, you would want the infrastructure engineers to do their job, and you would want those AI engineers to do their job, and they don't have to talk to operate on the day-to-day mode, so they're not blocking each other. And this kind of model adaptation desire around small models breaks that and creates a lot of back and forth, and that's a problem, right?
So these are some challenges related to: okay, we have a bunch of small models, how do we have a cluster, how do we serve this efficiently?
So we have been playing with this problem for a while. I'm Daniel, actually from Superlinked. I kind of skipped the intro. So we are a VC-backed company out of SF, and we have been building AI-powered search and document processing systems and agents for the last couple of years. And our main pain point has always been inference, specifically these problems that I have described. And so we have iterated and iterated and explored different topologies for clusters for running large wide fleets of small models in different environments, because sometimes you need to deploy together with some platform in some environment where who knows what is available there. The small models make it easier because in whatever environment you can get some L4s or some kind of small GPU quota is much easier.
So I'll describe a little bit about the topology of the cluster that we have kind of converged to. And by the way, this whole thing is Apache 2.0, completely open source. You guys can just take it and wrap it, and now you are an inference startup. This is open source from the control plane all the way down to the thing that runs on the GPU. So we didn't pull any punches.
And the topology is basically there is a gateway, and instead of having a router that pre-decides what goes where, there is a gateway that parses some of the requests and attaches some metadata to the request, inserts that request into a shared queue, and into some side channels. I'll go a little bit into that. And then the workers pull from that centralized queue instead of pushing the data down to the workers. And this way they can saturate themselves better. And then the worker setup, I think I have a slide for that, will describe how we basically absorb the complexity of different model architectures into a coherent set of workers that don't have competing Python requirements and stuff.
So that's the overall topology.
And this is the life of a request. So maybe I'll call out a couple of things from here. One of the things we don't like about the OpenAI API standard is the base64 encoded JSON. Not good for small models, not good for high throughput. So we use message pack throughout, a binary format. This way we can also push all the multimodal data through the actual API gateway. So there is no binary data over here and then request over there. And then the cluster needs access to your cloud storage to start loading some binary data, images or videos. We kind of encode it all and we push it through the gateway. And then the gateway separates some of these heavier pieces to not clog the internal queue and defers it on cloud storage in-flight while the request is in queue. So it splits up some of these requests that are over a megabyte and then uses cloud storage in the backend. But as a user, you push all your bits and bytes into the API layer and it's a clean interface because of that.
Basically the whole stack is Rust. So gateway Rust, the worker is Rust. And then over a socket locally, it attaches to different runtimes. And we have basically PyTorch, Candle and SGLang on the runtime. And then when we do the optimization, I'll go into that on how we make sure that whichever runtime we are using and whichever code is running in that runtime is the most efficient one. We have an auto research loop for that.
So the life of a request looks like that. And one tidbit is that you really want to make sure that the gateway that's the first thing that's hit by the request doesn't do too much work. Because then it becomes a bottleneck. So you don't even want to parse the whole request. You want to be able to look at the packets and figure out the general shape of what's coming, do the annotation, and then you have the workers. However many workers you have, hundreds of GPUs that look at the queue state and then pull from there. And the queue uses NATS Jetstream and that thing can do a million requests per second. That's very hard for that to become a bottleneck.
So you don't want to serialize, deserialize as you go through all of these different components. That's basically the obvious thing. This is an animation that shows the idea behind the centralized queuing. So instead of the top-down router trying to fill in the local queues just right, which is basically impossible, the whole idea is, can we somehow centralize the queuing and can the workers rather pick up the task of forming their own batches with their own prediction of the cost of the batch and then become much more efficient.
Now one tidbit and side note: once you start working on these things, you realize that it's actually really hard to predict how many things to pick up from the shared queue for the batch to be really optimal. And so you would want some mechanism that allows you to put some things back into the queue if you figure out: oh, I pulled a little bit too much. And that's a network hub, right? So that's a problem. And we have special optimization for that for machines that have multiple GPUs locally. So there is additional machine local queuing element that takes advantage of the fact that the local processes that run on the multiple GPUs on one machine can negotiate with the queue a little bit back and forth, which over the network would add milliseconds. And so we don't do it over the network, only when we co-locate the workers on multi-GPU machines.
And we are not talking about 5% differences here. You centralize the queue and now you get double the throughput of the cluster. So this is significant.
I mentioned three different runtimes. So basically it's either we write, for models that are encoder only, we write the PyTorch code and we optimize it. And we have an auto research loop that optimizes it. Same for Candle. We started to play with Candle not too long ago. We still can't get it to perform anywhere near the PyTorch performance. So it's more of a research project. It's just the dependency. The worker Docker image with PyTorch is like 12 gigabytes and the worker basically binary statically linked binary with Candle is maybe 10% of that. And if you care about waking up from the cold state and loading these images on a bunch of different machines, going from 12 gigs to a gigabyte or something makes a huge difference. So that's the motivation behind Candle is just getting the same performance from PyTorch is really hard.
And then SGLang we have there as a go-to baseline. We should perform at least as well as SGLang with optimal tuning of all of those parameters that you have to do tuning on. Here are some numbers. So for example, when we wrap SGLang with the socket and with our Rust sidecar, we can actually improve on the bare SGLang performance just because we do something on the batching side that natively you can't make it do that. If you develop custom plugins into SGLang, probably you can match our performance because you can just push the same logic into the SGLang core server. But now you are developing custom code that only works with SGLang.
And the whole lesson here from small models is that the runtimes are super diverse, right? You don't want to necessarily get stuck with any one particular runtime because we have on the order of 50 different adapters now that we parametrize for the different models. # Transcript
Here are some numbers. So for example, when we wrap SG Lang with the socket and with our Rast sidecar, we can actually improve on the bare SG Lang performance just because we do something on the batching side that natively SG Lang does not do that. If you develop custom plugins into SG Lang, you can probably match our performance because you can just push the same logic into the SG Lang core server. But now you are developing custom code that only works with SG Lang.
And the whole lesson here from small models is that the runtimes are super diverse, right? You don't want to necessarily get stuck with any one particular runtime because we have, I think, on the order of 50 different adapters now that we parametrize for the different models. And so you need to somehow deal with this underlying complexity. And it's probably not by building a bunch of plugins for one specific runtime. It's probably some kind of abstraction, which in our case is this Rast sidecar concept and then the socket.
Now I'll talk about a couple of different numbers, but in terms of language around benchmarking, the knee is this concept of when you ramp up traffic on the server, when you request more and more throughput from it, and it gives you more and more throughput, that's when you go linearly up. And then at some point you hit a point where asking for more is not coming. So you flatten out and the latency goes up. So we call that the knee and it's a useful concept in benchmarking because that's the point of saturation, right? That's the maximum performance without hurting latency.
So just to give you some ideas of what is possible on relatively small hardware, right? And different types of small models. This is measured on the RTX Pro 6000. We work with NVIDIA L4, A100, RTX Pro 6000, H100, that sort of range. Those GPUs are much more readily available, on demand in any cloud, and most continents have quota. On this kind of stuff, you can basically get, for embedding models, even up to hundreds of millions of parameters, you can get hundreds of thousands of tokens per second and code it into the embedding, right?
So imagine you are sitting there now hitting your text embedding three on OpenAI API. Instead, you could be having one GPU and push half a million tokens per second into the thing and get the vectors out, right? So you have half a million tokens that you are pushing into something. You have a single GPU that's not even that big per second and you are getting out vector embeddings for your search system. As opposed to pushing all of that into a managed embeddings endpoint somewhere and paying orders of magnitude more money, right? And you can get latencies of low tens of milliseconds for these calls, right? If you use Cohere, OpenAI APIs and so on, these are hundreds of milliseconds, right?
And this is not rocket science. You can have massive cost savings, massive latency improvements and relatively easy operation with a handful of GPUs and some info around them. Right? So this is a really low hanging fruit. If you start anywhere with open source models, small models embeddings are a no brainer, right?
But it doesn't end there. So let's say you want to look at named entity recognition. You want to look at multi-vector search, or even generation, right? Of text or structured outputs and so on. You can be getting thousands of tokens per second output from task specific generative models as well per, let's say, half a thousand per second for one GPU there at the bottom. And so let's say you are generating synthetic data. You are generating annotations for your fine tuning, for your evals. Don't do that on a managed endpoint. That's a perfect task because you have it under control. You can survey the quality. That's a perfect task for open source model on your own infra.
And then if the infra you have around those GPUs is reasonable, you'll get linear scaling with the number of those GPUs. Now another idea if you are into small model serving is that you don't normally have a worker pool per model, right? You have a set of workers, set of nodes, they have GPUs. You bring those up. You preload the models. The models load for tens of minutes because there are hundreds of billions of parameters. And so you are happy. Okay. They finally loaded. Now I have a worker pool. This mentality doesn't really work with small models. Yeah. How much time is left? Six minutes over. Okay. All right. So pack models on the same GPU is faster.
This is a story of how you still want to pin some models, but you want to also do basically lazy loading and eviction as a function of memory pressure. You want to figure out how to combine the two. There is a little bit about auto research. We have auto research loops for adding support for new models and for their performance. We build a lot of internal tooling to do the measurement, to feed into those auto research loops, to basically push the numbers forward. And maybe most importantly, when we ship support for a model, it has all the tuning done, right? So there is no parameter sweep. We bundle basically a config for end to end the whole cluster.
This is a setup for the auto research loop. That is a meta loop that builds the harness that then runs the loop. And there is a dashboard on top that helps you understand how it works. We have custom UIs for that. And one of the outputs of that was a lot of that took 80 cents to train and it improved 18%, it improved quality of retrieval on German legalistic as a proof of concept by 18%. And that's it. So small models are good. They're relatively easy to serve. They are actually much cheaper, faster, faster, as smart, and that QR code goes to the GitHub repo of our cluster that I just described. Give us a star, and happy self-hosting. Thank you.
This is open source from kind of the control plane all the way down to the thing that runs on the GPU. So we didn't pull any punches. And the topology is basically there is a gateway, and instead of having a router that kind of pre-decides what goes where, there is a gateway that parses some of the requests and attaches some metadata to the request, inserts that request into a shared queue, and into some side channels. I'll go a little bit into that. And then the workers pull from that centralized queue instead of kind of pushing the data down to the workers. And this way they can saturate themselves better.
And then the worker setup, I think I have a slide for that, will describe how we basically absorb the complexity of different model architectures into kind of a coherent set of workers that, you know, don't have like competing Python requirements and stuff like that. So that's kind of the overall topology.
And this is kind of life of a request. So maybe just I'll call out a couple of things from here. We, you know, one of the things we don't like about the OpenAI kind of API standard is the base64 encoded kind of JSON. Not good for small models, not good for high throughput. So we use a message pack throughout, like a binary format. This way we can also push all the multimodal data through the actual API gateway. So there is no like, hey, you know, binary data over here and then request over here. And then the cluster needs access to your cloud storage to start loading some binary data, images or videos. We kind of encode it all and we push it through the gateway.
And then the gateway kind of separates some of these heavier pieces to not clog the internal queue and defers it on cloud storage kind of in-flight while the request is in queue. So it kind of splits up some of these requests that are, let's say, over a megabyte and then uses cloud storage in the backend. But as a user, you push all your bits and bytes into the API layer and it's kind of clean interface because of that.
Basically the whole stack is RAST. So gateway RAST, the worker is RAST. And then over a socket locally, it kind of attaches to different runtimes. And we have basically PyTorch, Kendall and SGLang on the, as a runtime. And then when we do the optimization, I'll kind of go into that on how we make sure that whichever runtime we are using and whichever code is running in that runtime is the most efficient one. We have an auto research loop for that basically. But yeah. So life of a request kind of looks like that. And like one tidbit is that you really want to make sure that the gateway that's kind of the first thing that's hit by the request doesn't do too much work.
Because then it becomes a bottleneck, right? So you don't even want to parse the whole request. You want to be able to kind of look at the packets and figure out the general shape of what's coming, do the annotation, and then you have the workers. However many workers you have, hundreds of GPUs that look at the queue state and then pull from there. And the queue use NATS Jetstream and that thing can do, you know, million requests per second. Like that's very hard for that to become a bottleneck. So yeah. Like ideally you don't want to serialize, deserialize as you go through all of these different components. That's basically the kind of obvious thing.
This is a little animation that shows the idea behind the centralized queuing. So instead of the top-down router trying to, you know, fill in the local queues just right, which is basically impossible, you know, the whole idea is, hey, can we somehow centralize the queuing and can the workers rather pick up the task of forming their own batches with their own prediction of the cost of the batch and then, you know, become much more efficient.
Now one tidbit and kind of side note, once you kind of start working on these things, you realize that it's actually really hard to predict how many things to pick up from the shared queue for the batch to be really, like really the optimal size. And so you would want some mechanism that sort of allows you to put some things back into the queue if you figure out, oh, like I pulled a little bit too much. And that's a network hub, right? So that's a problem. And we have special optimization for that for machines that have multiple GPUs locally, right?
So there is additional kind of machine local queuing element that takes advantage of the fact that the local processes that run on the multiple GPUs on one machine can kind of negotiate with the queue a little bit back and forth, which over the network, you know, there is like milliseconds extra that that would add. And so we don't do it over the network only when we co-locate the workers on multi GPU machines. And, you know, I mean, we are not talking about like 5% differences here, right? So like you centralize the queue and now you get double the throughput of the cluster. So this is significant.
I mentioned three different runtimes. So basically it's either, you know, we write, let's say for models that are encoder only, we write the PyTorch code and we kind of optimize it. And we have an auto research loop that optimizes it. Same for candle. We started to play with candle not too long ago. We still can't get it to perform anywhere near the PyTorch performance. So it's a little bit more of a research project. It's just the dependency. Like, you know, the worker Docker image with PyTorch is like 12 gigabytes and the worker basically binary statically linked binary with candle is maybe like 10% of that, right?
And if you care about kind of waking up from the cold state and loading these images on a bunch of different machines, the, you know, going from 12 gigs to a gigabyte or something like this makes a huge difference. So that's kind of the motivation behind candle is just the, the getting the same performances from PyTorch is really hard. And then SG Lang we have there as a kind of a go to baseline. Like we should perform as, at least as well as, as SG Lang with the optimal tuning of all of those parameters that I mentioned that you have to do the tuning.
Um, here is some numbers. So for example, when we wrap SG Lang with the socket and with our kind of Rast sidecar, um, actually we can improve on the bare SG Lang performance just because we kind of, uh, do something to do like on the batching side that natively as you can make it to do that. If you do like, if you develop custom plugins into SG Lang and stuff like that, like probably you can match our performance because you know, you can just push the same logic into the SG Lang core server. Uh, but now you are developing custom code that only works with SG Lang. And the whole lesson here from small models is that the runtimes are super diverse, right?
You don't want to necessarily get stuck with any one particular runtime because there is, you know, we have, I think on the order of 50 different adapters now that, that we parametrize for the different models. And so you need to somehow deal with this kind of underlying complexity. And it's probably not by building a bunch of plugins for one specific runtime. It's probably some kind of abstraction, uh, which in our case is this, uh, Rast sidecar concept. And then the socket.
Um, now I'll talk about a couple of different numbers, but in terms of like language around benchmarking, you know, the knee is this concept of like when you ramp up traffic on the server, uh, when you sort of request more and more throughput from it, and it gives you more and more throughput. That's when you go kind of linearly up. And then some point you hit this point where you kind of ask for more and more is not coming. So you kind of flatten out and the latency goes up. So we call that the, the knee and it's, it's like a useful concept in, in benchmarking, uh,
because that's kind of the point of saturation, right? That's, that's kind of the maximum performance without hurting latency. Um, so just to give you some ideas of what is possible on relatively small hardware, right? And different types of small models. So this is measured on the RTX pro 6000. We, we kind of work with NVIDIA L4, you know, A100, RTX pro 6000, H100, that sort of range. Um, again, those GPUs are much more readily available, kind of on demand in any cloud, basically most continents have quota, you know, um, and on this kind of stuff, uh, you can basically get, uh, for embedding models, even up to, let's say, uh, hundreds of millions of parameters.
You can get hundreds of thousands of tokens per second and code it into the embedding, right? So imagine you are sitting there now, like hitting your text embedding three on open AI API. Instead, you could be like having one GPU and push half a million tokens per second into the thing and get the vectors out. Right? Like, is this like a connecting, right? You have half a million tokens that you are pushing into something. You have a single GPU that's not even that big per second and you are getting out vector embeddings for your search system.
As opposed to like pushing all of that into a managed embeddings endpoint somewhere and paying like orders of magnitude more money. Right? And you can get latencies like, you know, low tens of milliseconds for these calls, right? Like if you use, uh, you know, cohere, open AI APIs and so on, these are hundreds of milliseconds, right? And this is not rocket science. You know, you can have just like massive cost saving, massive latency improvements and relatively easy, easy operation. Um, with, with like handful of GPUs and some, some info around them. Right?
So this like really low hanging fruit, if you start anywhere with open source models, small models embeddings are like no brainer. Right? Uh, but it doesn't end there. So let's say, uh, you want to look at, uh, named entity recognition. You want to look at, let's say multi-vector search, uh, even generation, right? Uh, of text or structured outputs and so on. Um, you, you can be getting, you know, thousands of tokens per second output from, uh, you know, task specific generative models as well per, uh, like let's say half a thousand per second. For, for one GPU there at the bottom. Um, and so let's say you are generating synthetic data.
You are generating annotations for your fine tuning, for your evals. You know, don't do that on a, on a managed endpoint. That's a perfect task because you have it kind of under control. You can survey the quality. That's a perfect task for, uh, open source model on your own infra. Um, and then you're like, if the infra you have around those GPUs is like reasonable, you'll get linear scaling with, with the number of those GPUs. Um, now another sort of, uh, idea if you are into small model serving, uh, is that you don't, you know, normally, um, you have kind of worker pool per model, right?
You have a set of, uh, workers, set of nodes, uh, they have GPUs. You kind of bring those up. You preload the models. The models load for tens of minutes because there are hundreds of billions of parameters. Uh, and so you are happy. Okay. They finally loaded. Now I have a worker pool. This mentality doesn't really work with small models. Yeah. Yeah. Quickly. How, how, what's the time left? Oh, six minutes over. Okay. All right. So pack models on the same GPU is faster. Um, this is a story of how you still want to pin some models, but you want to also do, uh, basically, uh, lazy loading and eviction, uh, as a kind of function of memory pressure.
You want to figure out how to combine the two. Um, there is a little bit about kind of auto research. We have auto research loops for, uh, adding support for new models and for their performance. Um, we build a lot of internal tooling to do the measurement, to feed into those auto research loops, to basically push the numbers forward. Um, and maybe perhaps most importantly, when we ship support for a model, it has all the tuning done, right? So there is no, okay, let's do a parameter sweep. We bundle basically a config for end to end the whole cluster. Um, this is a setup for the auto research loop. That is like a meta loop that builds the harness that then runs the loop.
And there is a dashboard on top that helps you understand how it works. Um, we have custom UIs for that. And one of the outputs of that was a lot of that took 80 cents to train and it improved 18%, it improved quality of retrieval on German legalistic as a proof of concept by 18%. And that's it. So small models are good. They're relatively easy to serve. Uh, they are actually much cheaper, faster, faster, is as smart, and that QR code goes to the GitHub repo of our cluster that I just described. Give us a star, and happy self-hosting. Thank you. .
. . . . .