The Small Model Infrastructure Nobody Built (So We Did) — Filip Makraduli, Superlinked
Description
Most embedding infrastructure assumes you know exactly which model you want ahead of time. This talk starts where that assumption breaks. Filip Makraduli walks through the real profiling mistakes, infrastructure gaps, and production constraints that led to building an embedding inference engine designed for dynamic model loading, hot-swapping, and memory-aware eviction instead of brittle one-model-per-container deployments. If you're working on small-model inference, embeddings, or GPU infrastructure, this is a practical look at what breaks in the real world and how to design around it. Speaker info: - https://www.linkedin.com/in/filipmakraduli/
Summary
Generated by claude-haiku-4-5-20251001The Small Model Infrastructure Nobody Built (So We Did)
Main Topics
- Small Model Inference: Optimizing inference for small models in AI search and document processing
- Superlinked Inference Engine (SIE): An open-source solution for production-ready small model inference
- Context Management in Agentic Workflows: Using small models to mitigate context rot
- Infrastructure Gaps: Addressing the missing infrastructure between model development and production deployment
- Model Diversity Support: Managing hundreds of heterogeneous open-source models with different architectures
Key Points
Why Small Model Inference Matters
- Context Rot Problem: Quality degrades as context increases in LLM applications. Small models can preprocess data to manage this effectively.
- Agentic Workflows: Small models serve as powerful tools for:
- Data preprocessing before agent execution
- Tool calling and taxonomy classification
- Named entity recognition for knowledge graph generation
- Community Validation: Industry leaders like Andrej Karpathy are building graph-based knowledge bases using similar approaches
What Inference Is NOT
- Not Just More GPUs: Simply adding compute power doesn't solve inference problems
- Not Memory-Focused Alone: Small models (e.g., Stella embeddings, GLiNER NER models) occupy only a few GB, making dedicated GPU provisioning wasteful
- Not Only Server-Side: Production requires routing, auto-scaling, monitoring, and GPU provisioning—infrastructure often missing in open-source tools
The "Yin and Yang" of Model Inference
Yin — Model Support:
- Open-source models are rapidly improving and now outperform managed services on specific tasks
- Hugging Face hosts ~3 million models with varying architectures
- Supporting diverse models requires adapting implementations across:
- Flash attention mechanisms
- Positional embeddings (absolute vs. rotary)
- Normalization techniques
- Query-key-value fusion strategies
- Output formats (vectors vs. scores)
Yang — Infrastructure:
- Three core API primitives:
encode,score,extract - Essential components:
- Router and queuing mechanisms based on load
- Multiple GPU pools (spot instances, specialized hardware)
- Hot-swapping capability with LRU (Least Recently Used) eviction policy
- Auto-scaling with Prometheus metrics and CAD
- Terraform-based GPU provisioning
Technical Innovation
- Variable-Length Flash Attention: Eliminates padding waste in token-based batching
- Model-Agnostic Forward Pass Re-implementation: Adapts attention mechanisms across different architectures
- Flexible Output Handling: Supports embeddings (Colbert), vectors, and scores (cross-encoders, re-rankers)
- GPU Utilization: Hot-swapping multiple small models on a single GPU dramatically improves cost-efficiency
Notable Quotes
> "This part around how models run in production, scheduling GPUs, routing and automation, was a bit of a blind spot for me."
> "On the market currently, there is no open source solution that leads you from creating the model inference to actually productionizing this at a scale of this size."
> "Your inference is worthless if you're not supporting the right models, or you're not offering enough breadth of options for your users."
> "It's not even a sacrifice anymore. It's actually you're being able [to use open source models] for specific tasks [where they beat] managed services."
Takeaways
- Adopt Small Models for Context Management: Use small models as preprocessing and tool-calling layers in agentic workflows to mitigate context rot
- Understand Production Inference Complexity: Inference requires coordinated optimization across model support AND infrastructure—not just algorithmic improvements
- Leverage Open-Source Ecosystem: Small, well-optimized open-source models are now competitive with proprietary alternatives for specific tasks
- Use the Superlinked Inference Engine: For teams building AI search and document processing systems, SIE provides:
- Support for hundreds of heterogeneous models
- Production-ready infrastructure (routing, auto-scaling, monitoring)
- Easy deployment via Helm charts and Terraform
- Significant cost savings through GPU hot-swapping
- Model Heterogeneity Requires Flexibility: Different models need different implementations; a one-size-fits-all approach is insufficient for broad model support
Resources
- Repository: Superlinked Inference Engine (SIE) — open-sourced at soft launch
- Partners: Tested with Chroma, Quadrant, Weaviate, and LensDB
- Infrastructure: Helm charts and Docker images provided for easy deployment
Transcript
Hello everyone, welcome to this talk. I'll be speaking about small model inference and a gap that we've recognized in the market and what we did about it and why we made this approach. And as you can see, this background slide here, this is no accident. So if you can guess what this is, I'll prompt you at the end of the slides, you win a little reward so you can catch me at the break afterwards. So think about this, but also listen to me so don't think too hard. So the story starts with me posting an article a few months ago on Substack that got a bit of traction, got a few people interested, and I explained flash attention, I explained how models worked, how processes can be memory bound, compute bound, and I felt really good because I went deep into this and as a person who's been in AI for a few years, I felt very confident. And that was true. But then some people pointed out that actually I had overlooked one key aspect around what makes these models fast in the real world. And that aspect that I overlooked was inference. So as someone who wants to understand things in first principles and work to understand the problems and the solutions deep, I realized I need to figure this out, I need to know where I've made my mistake. And as an AI researcher and engineer, I need to find out more about inference. So I've done a lot of work with VLLM training models, fine tuning, doing applied ML and AI, did a bit of research in academia as well. But this part around how models run in production, scheduling GPUs, routing and automation, was a bit of a blind spot for me. So I realized this is the time I have to figure this out, learn and make it work. And what better way to do this than to actually build stuff. So I decided to join a team, a team at Superlinked, comprised of very good infrastructure engineers, and actually work with them and build something around inference. And that something is this repo, the Superlinked inference engine that we have open sourced. And this is the soft launch that I'm doing today. So you can have a look at that later. So it's inference for small models around AI search and document processing. And we've tested this out, as you can see, with some of our partners. So we've tested out with Chroma, Quadrant, Weaviate, so a lot of the vector DBs, as well as LensDB. And as you can see, they've tried it out a bit. Sounds fun, sounds interesting. So this is working. And it was the right step for me to figure this out and learn where inference truly is and how that combines with my ML experience. So the three key points I want you to go away with from this talk are these. First, I want to tell you why this matters. So why doing inference for AI search and document processing actually matters when you're building agents or when you're building workflows that involve agents. So this is very important, the why. Then I want to talk to you about the second thing, which is what inference is not about. So there are some misconceptions or ideas of how inference looks like, but it's not all about those things. And the third thing is how we see inference. And I call this the yin and yang of model inference. And it's a way of combining a few things around model support and infrastructure. So why this matters for your agentic workflow? Well, what you have encountered for sure is context rot. And as we probably all know, this research paper from Chroma from some time ago showcases this effect that no matter what you do, there is this effect of context rot. So quality degrades as context increases. So being able to manage this context and do some context management is very important and useful way to solve this. So using small models that can preprocess your data so that then you can actually use your agents and build your workflow is a very powerful technique. You can also use small models for tool calling to do similar things and tackle this problem of context management. And you might say, why would I not use cloud code and grepping? And that's a valid point. However, you can still do that. And having your data being preprocessed is actually making the grepping and file systems that you build even better. And it's not just me saying this. So this is how the community has responded to this problem. So Andre Karpathy is building knowledge bases, graph based. So for example, you can use named entity recognition models to generate ontologies and then build knowledge graphs. This is very effective. And we've done this in production as well. We have a use case where this repo that I showcased, the Superlinked inference engine, we've used it as a tool calling solution where it was around taxonomy classification for an e-commerce store. So going directly with tool calling where the small models are tools that can retrieve and go through your data is also a powerful way to approach this. So now I'll go about inference and what inference doesn't look like. So the traditional perspective is, okay, I will just put in more GPUs, get more compute, all good, I've solved the inference. However, with small models where each model takes up only a small space in memory, you can see, for example, Stella, which is an embedding model, then other re-rankers, the named entity recognition GLiNER model, they only occupy a few gigabytes of memory. And if you provision a GPU for each model, you're wasting a lot of idle space. Your GPU stays idle and it's not used. So in this case of small model inference, it's very important to be able to hot swap models. So what we've done is we've built the ability for you to swap all of these models in one GPU so that per GPU you get much higher utilization. So this lowers your costs, but also enables you to hot swap and switch around between models quickly. If you want to have one tool that's one re-ranker or another tool that's another model, you can switch them around quickly and we have this least recently used eviction policy that we've built in. And also what inference is not about is only the server or only the production situation. So we have solutions for both the server and production, but building something, for example, like using Triton or VLLM or having some API wrapper in order to get that in production, where you actually do routing, auto-scaling, you do monitoring with Prometheus metrics and Grafana, you have to write that code yourself. And on the market currently, there is no open source solution that leads you from tool that's another model, you can switch them around quickly and we have this least recently used eviction policy that we've built in. And also what inference is not about is only the server or only the production situation. So we have solutions for both the server and the production, but building something. So for example, using Vay or VLLM or having even some API wrapper in order to get that in production, where you actually do routing, auto-scaling, you do monitoring with Prometheus metrics and Grafana, you have to write that code yourself. And on the market currently, there is no open source solution that leads you from creating the model inference to actually productionizing this at a scale of this size. So that's why we've wanted to fill this gap as well. So we've included a lot of stuff around routing, auto-scaling, queuing mechanisms, and provisioning GPUs. So I've talked about what inference isn't, why this is important, and let me tell you now about what actually inference is about. And I call this the yin and yang of inference, because I feel this is a holistic approach that has to combine two key things. So first, the yin is model support. So your inference is worthless if you're not supporting the right models, or you're not offering enough breadth of options for your users. And open source models, on Hugging Face currently, there are millions of models. This is from March. Now there might be even more, close to three million models. And open source is moving very quickly, both in size and in accuracy. So you need to support these models because people want to use them, people want to work with open source. And the performance is also getting better and better. So it's not even a sacrifice anymore. It's actually you're being able for specific tasks. So if you look at MTAB or different benchmarks, you can see that for very narrow tasks, open source models are beating managed services. And we see this even with more general purpose models like what Gemma have done. So they've released very low parameter models that have ELO scores that are higher than much bigger models. So open source small models are very relevant. And you need an infra to support them. And we decided to do exactly this. Okay, let's do it. We'll build and support hundreds of models. However, that's not as straightforward as it might sound because all of these models have different run times. So for example, if you use Bert and Quinn as models, they have different implementation of flash attention. They have different positional embeddings. And you need to adjust this in order to make this applicable to your case and make it more general. So there is no universal engine that can handle Bert, Quinn, and modern Bert as well because the architectures are different. So what we've done is we've set up a way of re-implementing this forward pass in order to adapt attention, in order to do padding where needed. So have variable length attention as well. Work with fusion of the query key and value. And this is an underrated component, especially if you look at supporting models like Colbert that has multiple vectors as output. So late interaction models. Also cross encoders and re-rankers that do not output a vector at all but output scores. So it's very important to figure this out. And one example of this is a comparison of a few models that are totally different in five aspects. So normalization is done differently in Bert and Quinn. So that needs to be accounted for. In Colbert, it's also totally different. The query key and values can also be fused somewhere. In others, for example, in Quinn, it cannot be because there is grouped query attention. So that's another problem. The positional embeddings are also different. You can do in Bert absolute lookup with Quinn. There is rotary positional embeddings. So these are problems that people don't think about straightaway but if you want to support a lot of models, there needs to be a consistent way of doing this. So we've set up this with working with agents as well as with humans to be able to support this and reimplement this forward pass in order to make the inference of this model sufficient. And we've also worked to do flash attention, where it's variable length flash attention so that you can do padding and there is no waste. Because what happens is if you want to do token-based batching, you can have tokens that are a lower number of tokens on one request and then another is a higher amount. And then if both are padded at the higher amount, you basically are wasting compute on empty tokens. So you need to adapt this in order to make the models work even better and quicker. And that's a key differentiator in our way of approaching inference. The other part, the yang, is around infrastructure. So this is the complex plot. This is the deep dive. We have this in our website and in our repo. I won't go into each component here but basically, what's up there is the three primitives of our API: encode, score, and extract. And then the infrastructure layer happens. And this is the summary of it. So basically, we have the three primitives of the API and then we have a router and also a queuing mechanism depending on the load. So we're adjusting between these. And also different pools of different GPUs and resources that we use in order to distribute the workload. And we use spot instances but also bigger GPUs. So it's about being able to provision the hardware and also have the metrics to auto scale this. So we're doing CAD auto scaling with Prometheus metrics in order to switch models around, not keep GPUs idle and not waste resources. So the key driver here is that we want you to have the cluster as well as the model support. Not just one or the other. And then you have to merge them together and write all that code. But we want to give you the whole end to end thing so that you can work with these models quickly. And the models are basically just the config that you can switch around and then do Terraform apply. And that's it. And we also have published Helm charts and Docker images as well. And that's how I learned my lesson. We've built size. So this is our soft launch in a way. And we've open sourced both the model inference that I talked about with adapting this forward pass, working with attention, re-implementing different aspects on the model, but also the cluster so that you can actually use this straightaway without having to think about hardware, provisioning GPUs, and all of those problems that arise in the real world. You can scan the QR code to look at the repo. And it's called SIE. So S-I-E. And the company is super linked. And also now we have the background as well. So a quick reminder on that. Does anyone have any ideas from the audience? I'll reveal it in the next slide. But yes, I don't want to do it too quickly. If anyone knows what this is, think of it, machine learning a bit, foundational knowledge. I was talking about attention and all that. So it's a bit around embeddings. straightaway without having to think about hardware, provisioning GPUs, and all of those problems that arise in the real world. You can scan the QR code to look at the repo. And it's called SIE. So S-I-E. And the company is super linked. And also now we have the background as well. So a quick reminder on that. Does anyone have any ideas maybe from the audience? I'll reveal it in the next slide. But yeah, I don't want to maybe do it too quickly. If anyone knows what this is, think of it, machine learning a bit, foundational knowledge. I was talking about attention and all that. So it's around embeddings, but also how transformers work. Okay. No reward, I guess, for you, sadly. Or yeah? Yeah? Sorry? I can't remember exactly. It was like all the embeddings were clustering on one area and using embedding space. Okay. Because of the way we were doing the positional encoding. So I will take that as a correct answer. So it was around positional encodings. And yes, this is, yeah. Congrats. So yes, this is around vector visualization of positional encoding. So this is done when transformers are trained. This is how positions are encoded. And it's sinusoidal. That's why you get this pattern. Because they're embeddings. And yeah. Thank you very much. And this is kind of like the soft launch that I'm doing today. So you can have a look at that later. So basically, it's inference for small models around AI search and document processing. And we've tested this out, as you can see, with some of our partners. So we've tested out with Chroma, Quadrant, VV8, so a lot of the vector DBs, as well as LensDB. And as you can see, they've tried it out a bit. Sounds fun, sounds interesting. So this is working. And it was the right step for me to kind of figure this out and learn where inference kind of truly is and how that combines with my ML experience. So the three key points I want you to go away with from this talk are these. So first, I want to tell you why this matters. So why doing inference for AI search and document processing actually matters when you're building agents or when you're building workflows that involve agents. So this is very important, the why. Then I want to talk to you about the second thing, which is what inference is not about. So there are some misconceptions or ideas of how inference looks like, but it's not all about those things. And the third thing is how we see inference. And I call this as like the yin and yang of model inference. And it's a way of combining a few things around model support and infrastructure. So why this matters for your agentic workflow? Well, what you have encountered for sure is context rot. And as we probably all know, this research paper from Chroma from some time ago showcases this effect that no matter what you do, there is this effect of context rot. So quality degrades as context increases. So being able to manage this context and do some context management is very important and useful way to solve this. So using small models that can preprocess your data so that then you can actually use your agents and build your workflow is a very powerful technique. You can also use small models for tool calling to again do similar things and tackle this problem of context management. And you might say, okay, why would I not use like cloud code and grepping? And that's a valid point. However, you can still do that. And having your data being preprocessed is actually making the grepping and file systems that you build even better. And it's not just me saying this. So this is how the community has responded to this problem. So Andre Carpathie is building knowledge bases, graph based. So for example, you can use named entity recognition models to generate ontologies and then build knowledge graphs. So how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies how the system identifies effective. And we've done this in production as well. We have a use case where this repo that I showcased, the Superlink inference engine, we've used it as a tool calling, basically, a solution where it was around taxonomy classification for an e-commerce store. So going directly with tool calling where the small models are tools that can retrieve and go through your data is also a powerful way to approach this. So now I'll go about inference and what inference doesn't look like. So the traditional maybe perspective is, OK, I will just chuck in more GPUs, get more compute, all good, I've solved the inference. However, with small models where actually each model takes up only a small space in memory. So you can see, for example, Stella, which is an embedding model, then other re-rankers, the name-dentity recognition, gliner model, they only occupy, like, a few gigabytes of memory. And if you provision a GPU for each model, you're wasting a lot of idle space. Your GPU stays idle and it's not used. So in this case of small model inference, it's very important to be able to hot swap models. So what we've done is we've built the ability for you to swap all of these models in one GPU so that per GPU you get much higher utilization. So this lowers your costs, but also enables you to hot swap and switch around between models quickly. If you want to have one tool that's one re-ranker or another tool that's another model, you can switch them around quickly and we have this least recently used eviction policy that we've built in. And also what inference is not about is only the server or only kind of the production situation. So we have solutions for both the server and the production, but building something. So for example, like using Tay or VLLM or having even some API wrapper in order to get that in production, where you actually do routing, auto-scaling, you do kind of monitoring with Prometheus metrics and Grafana, you have to write that code yourself. And on the market currently, there is no open source solution that leads you from creating the model inference to actually productionizing this at a scale of this size. So that's why we've kind of wanted to fill this gap as well. So we've included a lot of stuff around, routing, auto-scaling, queuing mechanisms, and provisioning GPUs. So I've talked about what inference isn't, why this is important, and let me tell you now about what actually inference is about. And I call this the yin and yang of inference, because I feel like this is a holistic approach that has to combine two key things. So first, the yin is model support. So your inference is worthless if you're not supporting the right models, or you're not offering enough breadth of options for your users. And open source models, so on Hugging Face currently, there are millions of models. This is from March. Now there might be even more like close to three million models. And open source is moving very quickly, both in size, but also in accuracy. So you need to support these models, because people want to use them, people want to work with open source. And the performance is also getting better and better. So it's not even a sacrifice anymore. It's actually you're being able for specific tasks. So if you look at MTAB or different benchmarks, you can see that for very narrow tasks, open source models are beating managed services. And we see this even with more general purpose models like what Gemma have done. So they've released very low parameter models that have ELO scores that are higher than much, much bigger models. So open source small models are very relevant. And you need an infra to support them. And we decided to do exactly this. Okay, let's do it. We'll build and support hundreds of models. However, that's not as straightforward as it might sound. Because all of these models have different run times. So for example, with if you use Bert and Quinn as models, they have different implementation of flash attention. They have different positional embeddings. And you need to adjust this in order to make this kind of applicable to your case and make it more general. So there is no universal engine that can handle Bert, Quinn, and modern Bert as well. Because the architectures are different. So what we've done is we've set up a way of re-implementing this forward pass in order to adapt attention, in order to do padding where needed. So have variable length attention as well. Work with kind of fusion of the query key and value. And this is kind of an underrated component, especially if you look at supporting models like Colbert that has multiple vectors as output. So late interaction models. Also cross encoders and re-rankers that do not output a vector at all. But they output kind of scores. So it's very important to figure this out. And one example of this is, for example, this, which is a comparison of a few models that are totally different in this five aspects. So normalization is done differently in Bert and Quinn. So that needs to be accounted for. In Colbert, it's also totally different. The query key and values are also can be fused somewhere. In others, for example, in Quinn, it cannot be because there is grouped query attention. So that's another problem. The positional embeddings are also different. You can do in Bert absolute lookup with Quinn. There is rotary positional embeddings. So these are problems that maybe people don't think about straight away. But if you want to support a lot of models, there needs to be a consistent way of doing this. So we've set up this with kind of working with agents as well as with humans to be able to support this and reimplement this forward pass in order to make the inference of this model sufficient. And we've also worked to do flash attention, where it's variable length flash attention so that you can do padding and that there is no waste. Because what happens is if you want to do token-based batching, you can have tokens that are kind of lower number of tokens. One request and then another is a higher amount. And then if both are padded at the higher amount, you basically are wasting compute on empty tokens. So you need to adapt this in order to make the models work even better and quicker. And that's a key differentiator in our way of approaching inference. The other part, the yang, is around infrastructure. So this is the complex plot. It's kind of this is the deep dive. We have this in our website and in our repo. I won't go into each component here. But basically, what's up there is the three primitives of our API. So encode, score, and extract. And then the infrastructure layer happens. And this is kind of the summary of it. So basically, we have the three primitives of the API. And then we have a router and also a queuing mechanism depending on the load. So we're adjusting between these. And also different pools of different GPUs and resources that we use in order to kind of distribute the workload. And we use spot instances, but also bigger GPUs. So it's about being able to provision the hardware and also have the metrics to auto scale this. So we're doing CAD auto scaling with Prometheus metrics in order to be able to kind of switch models around, not keep GPUs idle and kind of waste resources. So the key driver here is that we want you to have the cluster as well as the model support. So not just one or the other. And then you have to merge them together and write all that code. But we want to give you the whole end to end thing so that you can work with these models quickly. And the models are basically just the config that you can switch around and then do Terraform apply. And that's it. And we also have published Helm charts and Docker images as well. And that's how I learned my lesson. We've built size. So this is kind of our soft launch in a way. And we've open sourced both the model inference that I talked about with kind of adapting this forward pass, working with attention, re-implementing different aspects on the model, but also the cluster so that you can actually use this straightaway without having to think about hardware, provisioning GPUs, and all of those problems that arise in the real world. You can scan the QR code to look at the repo. And it's called SIE. So S-I-E. And the company is super linked. And also now we have the background as well. So a quick reminder on that. Does anyone have any ideas maybe from the audience? I'll reveal it in the next slide. But yeah, I don't want to maybe do it too quickly. If anyone knows what this is, think of it, machine learning a bit, kind of foundational knowledge. I was talking about attention and all that. So it's a bit around embeddings, but also how transformers work. Okay. No reward, I guess, for you, sadly. Or yeah? Yeah? Sorry? I can't remember exactly the hell. It was like all the embeddings were clustering on like one area and using basically embedding space. Okay. Because of the way we were doing the positional encoding. So I will take that as a correct answer. So it was around positional encodings. And yes, this is basically, yeah. Congrats. So yes, this is around vector visualization of positional encoding. So this is done when transformers are trained. This is how positions are encoded. And it's sinusoidal. That's why you get this pattern. Because they're embeddings. And yeah. Thank you very much.