Your Coding Agent Should Do AI System Engineering — Ben Burtenshaw, Hugging Face
Description
An agent written RMSNorm kernel hit 1.88x speedups on H100s. A finetuned Qwen3 0.6B hit 35% on LiveCodeBench. Neither result required a systems engineer. Just coding agents with the right skills loaded. Ben Burtenshaw from Hugging Face walks through three levels: using Claude Code interactively to write and benchmark CUDA kernels distributed as versioned repos on the Hub, a zero-shot task where an agent finetunes a model end to end from a single prompt, and a multi agent research lab running parallel experiments overnight on Hub compute while a reporter agent pushes results to a live Trackio dashboard. The through line is skills: file based context that turns a zero shot failure into a few shot workflow. CUDA programming and ML training pipelines were deep specializations that took years. Skills compress that timeline to hours. Speaker info: - https://x.com/ben_burtenshaw - https://www.linkedin.com/in/ben-burtenshaw/ - https://github.com/burtenshaw
Summary
Generated by claude-haiku-4-5-20251001Your Coding Agent Should Do AI Systems Engineering — Ben Burtenshaw, Hugging Face
Main Topics
- Using Coding Agents for AI Systems Engineering
- Moving beyond basic tasks to tackle harder engineering problems
- Three progressive complexity levels: CUDA kernel writing, model fine-tuning, and multi-agent research labs
- CUDA Kernel Optimization
- Memory bottlenecks in GPU computation
- Custom kernel distribution and management
- Skills-based approach to kernel development
- Model Fine-tuning with Agents
- Zero-shot task execution
- Integration with Hugging Face Hub
- Multi-Agent Automated Research Labs
- Distributed autonomous research teams
- Integration with tracking and compute infrastructure
Key Points
CUDA Kernels & Efficiency
- GPU Bottleneck Reality: Modern GPUs (e.g., H100) can compute at petaflop/second speeds but memory bandwidth is only 3 terabytes/second—memory, not computation, is typically the bottleneck
- Efficiency Breakdown: Efficiency = Compute + Memory + Overhead
- Compute: FLOPs and matrix multiplications
- Memory: Data/tensor movement between memory levels
- Overhead: Python environment, PyTorch dispatch
- Custom Kernels Strategy: Increase arithmetic intensity—maximize GPU operations per memory read/write cycle (e.g., Flash Attention)
- Real-world Result: Generated kernel for Qwen 3 8B on H100 achieved 94X speedup
Infrastructure & Distribution
- Hugging Face Kernels Library: Open-source library for distributing optimized kernels via TOML configuration files specifying hardware compatibility and CUDA versions
- Hub Integration: Kernels can now be published on Hugging Face Hub like models, creating a ecosystem for kernel publishers
- Skills Concept: File-based context that agents can selectively open/close, converting zero-shot tasks into few-shot learning by providing examples
Model Fine-tuning
- Fully Integrated Workflow: Agents can execute prompts like "fine-tune Quen 3 6B on this dataset" with full Hub integration
- Cost Optimization: Tools like Unsloth reduce computational costs while maintaining performance
Multi-Agent Automated Research (AutoLab)
Architecture with Four Agent Types:
- Researcher Agent: Searches papers (via HF Papers with CLI access) and formulates hypotheses
- Planner Agent: Maintains job queue and proposes fresh single-change experiments
- Worker Agents: Implement hypotheses as training script modifications
- Reporter Agent: Monitors jobs and maintains performance dashboard
Infrastructure:
- GitHub-based workflow (main branch maintains current best training scripts)
- HF Jobs integration for distributed execution
- Traceo Dashboard: Open-source metrics tool with completely open Parquet data layer enabling:
- Custom visualizations (Gantt charts, etc.)
- Event tracking and warnings
- Email notifications
- Agent-friendly data access
Efficiency Measurement: Tracks bits-per-byte efficiency improvements across iterative experiments
Notable Quotes
> "We need to go closer to the silicon and tackle harder problems—that's where AI systems engineering comes in."
> "The GPU is often waiting idle for tensors to come back for them to be computed... people like to say we keep the GPUs warm and that's the objective of writing a custom CUDA kernel."
> "Agents work really well with primitives and open primitives... it's more about exposing well [than extracting]."
> "The Hugging Face Hub is ready for these kind of workloads. We have the fundamentals in place like storage, tracking, and compute, which will allow us to scale our engineering to new levels."
> "If you have a verifiable experiment like training a model or writing CUDA kernels, then it's pretty easy to implement a setup and learn some stuff."
Takeaways
For AI Engineers
- Career Development: Shift focus from commodity tasks to systems engineering problems—kernel optimization, architecture improvements, and distributed research
- Leverage Open Primitives: Build with fully open tools (Traceo, Hugging Face Hub, kernels library) that expose rather than abstract complexity
- Low-hanging Fruit: Hardware-model compatibility mismatches create optimization opportunities—use agents to quickly generate compatible optimized kernels
For Infrastructure
- Hub Readiness: Hugging Face Hub has mature storage, tracking, and compute fundamentals ready for agentic workloads
- Open Data Layers Matter: Parquet-based data stores (like Traceo) enable agent flexibility and custom visualizations
- Distributed Workflows: Multi-agent architectures with Git-based coordination scale better than single-agent approaches
Actionable Items
- Explore kernel writing using agent-friendly skills and benchmarking templates
- Try fine-tuning models with Claude/agents via available blog posts with free credits
- Implement own AutoLab setup using OpenCode, Traceo, and HF Jobs for research automation
- Check out publicly available repos with examples for all three approaches
Transcript
Hi everyone as you heard I'm Ben from Hugging Face and the talk that I'm going to present to you today is called your coding agent should do AI systems engineering so there are two main takeaways that I want you to get from this talk one and probably the fun part is that we can use coding agents to tackle the hardest engineering problems in AI so systems engineering and machine learning engineering and maybe the boring part is that in order to do this we're going to need standard repos and we're going to need those on the hub and in many cases we already have them so I think in this case I'm preaching to the choir here but in case you haven't noticed coding agents have been accepted many of us have been using them for a few years but in the last few months they seem to have crossed a acceptance gradient where a broader group of people are using them so with this in mind how do we keep our careers our engineering contemporary and how do we keep challenging ourselves in new areas and my proposal is that we need to go closer to the silicon and tackle harder problems and that's where AI systems engineering comes in I've broken this talk down into three progressively more complex steps and more autonomous steps as well and I've defined those as three bosses from games the first one is a hybrid approach where you interactively use an agent to solve a problem and to write a CUDA kernel the second is a zero shot task where an agent takes a prompt and trains an LLM on Hugging Face the third is a multi-agent auto research setup like an automated AI lab so let's get started on the first boss right this is writing CUDA kernels so for a while writing custom kernels was seen as this unattainable goal for the humble agent they required complex DSLs they required integration with relevant hardware to be benchmarked and to be tested and it was seen as something that couldn't be achieved by agents however that in most cases was wrong if you look at kernel hackathons like those on GPU mode the recent AMD hackathon if you look at papers like kernel bench you'll see that agents are able to write valid and optimized CUDA kernels and that's really cool and something that totally inspires me I'm a part of GPU mode I contribute to that and something that I think everyone should be doing however what do we do with them how do we distribute them and how do we get them into our inference engine so that we actually are using these optimized kernels that we're generating and that's part of the question of this part of the talk let's take a step back now and just say what a kernel is right so when you run an AI model on a GPU the actual work is executed through a kernel this will be defined in a relevant language for that hardware and it will use relevant features to that hardware that may not be available on other hardware we can write custom kernels that will take advantage of that hardware for a specific math operation squeeze everything we can out of it so that the model will infer faster in general this requires a lot of expertise about writing CUDA kernels about the hardware and it's also an installation hell as you deal with a pretty large install matrix from hardware to software to generations and versions of CUDA and these kinds of issues so in short it's hard efficiency in deep learning so efficiency in kernels is split into three main sections one compute two memory and three overhead compute is the flops these are the matrix multiplications and the real math of the process memory is the time spent moving data or tensors around memory typically from slow to fast memory and overhead is basically everything else the Python environment PyTorch dispatch of those kernels these kinds of things in general most people might assume that the compute is the bottleneck here because it's doing most of the math right that's not correct so in most cases memory is usually the bottleneck and that's because a modern GPU let's take an H100 for example can do a petaflop a second of computation but its memory bandwidth is three terabytes so in short the GPU is often waiting idle for these tensors to come back for them to be computed there are custom kernels custom optimized kernels that exist flash attention being the poster child of these and in general what they do is increase arithmetic intensity they basically make the GPU do more operations at once per read and write so we move the tensors across we do as much math as possible on the GPU in one go and then we write it back in short people like to say we keep the GPUs warm and that's the objective of writing a custom CUDA kernel Hugging Face has a library called kernels which is maintained by kernel writers and we're beginning to scale up to agentic workloads so at its core this is a way of distributing kernels it has a TOML file like any kind of project which says which hardware it works on which versions of CUDA and other kinds of software as it requires to work and it's now also a repo on the hub just like models so if you are a kernel writer or you're an aspiring kernel writer with an agent that you want to set up you can now be a kernel publisher just like a model publisher and my point is that this is a kind of super fertile ground for AI engineers looking to scale their career if you check out these repos on the hub you'll see that there's compatibility for different hardware you can configure that so you know this works on my GPU or on my laptop and this is what it looks like here right let's take a look at what this looks like for an agent and how we're helping an agent to do this so first we're going to go to how we do this so skills so I'm sure everyone here is familiar with skills and I'm sure there have been a number of talks that really go deep into skills I like to keep them pretty simple and really they're just kind of file based context with all the benefits of files we can open them and close them we can version them we can source control them and these kinds of things and agents can also do the same they can open them when they need them they can use them when they don't and so in the context of kernels that means that we can give examples of how to write and how to use kernels in skills and they can open those and use them when they need I like to say that it takes a task from being zero shot to being few shot which in ML is a familiar concept right we're just giving the agent examples of how to do things and we can be quite verbose and descriptive about that at Hugging Face we're focusing on integrating skills into their projects so what you'll find is that inside each project there's managed skills by that project which we think is the best way to do this because it means that those projects that the maintainers of those projects are maintaining their skills right that means that they're not necessarily the most YOLO skills because they're well maintained and robust and we have another repo for those kinds of more experimental skills which is called Hugging Face skills go and check that out if you want to try some of these examples you'll see today in kernels this is what the skill looks like it focuses on benchmarking so it has scripts that allow you to benchmark and test the skill sorry to test the kernel and see how performant it is and references with examples of how to do this we benchmarked this skill and we generated a kernel for Qwen 3 8B for H100 and we found that we had a 94X speed up this isn't skills right that means that they're not necessarily the most yolo skills because they're well maintained and robust and we have another repo for those more experimental skills which is called hugging face skills go and check that out if you want to try some of these examples you'll see today in kernels this is what the skill looks like it focuses on benchmarking so it has scripts that allow you to benchmark and test the skill sorry to test the kernel and see how performance it is and references with examples of how to do this we benchmark this skill and we used we generated a kernel for Quen 3 8B for H 100 and we found that we had a 94 speed up this isn't a state-of-the-art speed up on this model by any means it's really just about compatibility and a compatibility matrix so in many cases these models and their kernels won't be optimized for the respective hardware or generation of hardware that you want to use them on so you have some low hanging fruit here where you can just come and pick up some optimizations for that specific hardware maybe because your hardware is cheap on your cloud provider but it's not necessarily the most ideal for that model that you're using so my recommendation would be to come here and pick up some easy speed ups how do we know that these skills are any good and that we should be sharing them and telling people to use them we use an open source library called upskill that we're also maintaining this is a gateway to using cheaper and open models with skills it basically just generates skills generates an eval for the skill and then allows you to compare different models on the same skill so you can see things like this so okay gpt oss is slightly less accurate using the same tokens kimmy is more accurate using less tokens haiku is a bit more accurate using less tokens and these kinds of things so if you've got a skill and you're using it regularly and you're thinking to yourself okay how can I save a few pennies here and get a different model on the go then try out upskill and it allow you to iterate on your skill and improve it right let's move on to boss two I'm going to go through this one pretty quickly this is about fine-tuning models if you're really into this there was a talk yesterday by my colleague Mervé that went into this deeply there's also a blog post here where we got Claude to do this this was from back in November December time now go and check this out basically you can just say fine-tune Quen36B on this data set this is a chain of thoughts data set and you'll improve the models chain of thoughts this is fully integrated to the hub now so you can even run the GPUs on the hub and it uses HF CLI skills so it's all very available I would try this one out you can also try this one out this uses unsloth so it's even cheaper this runs with optimized models and it's maintained by onslaught and by us and it's another blog post and there's also often free credits that you can get around these blog posts so go and check these out okay let's move on to the big one this is auto lab multi-agent research which is a project that basically keeps me up at night Andrej Karpathy a few weeks ago maybe a month ago now released a project called auto research which was based on his other projects nanoGPT and nanochat and it took the nanoGPT architecture and got Claude code to create improve to write improvements to that training script so that it would improve the training process so we can see here the experiments going over and for each experiment there's a change in the training script which increases the efficiency measured in bits per bytes of that run and we can see that the efficiency ends at its best at the end of the process I and everyone thought this was super cool and I had to start implementing it straight away but one of the things that stood out to me was I found it weird that we had one agent working in a single way iterating going and finding improvements and then implementing them and it would make sense to distribute this so that's what I did I distributed the task amongst the research team with four types we have a researcher that basically looks up papers for this we use HF papers but we can also use archive papers HF papers is cool because it has a CLI so you can just pull and search papers from the hub and it acts as a literature scout so it just looks up for papers with ideas and it formulates those as hypotheses we then have a planner which takes those hypotheses and maintains a queue of jobs we then have a set of workers and they pick up those hypotheses and their job is to implement them as training scripts so in many cases just change the architecture or change the parameter or something and then we have a reporter agent that goes and monitors all these jobs and maintains a dashboard that we can use so this is what it looks like if you see here that we have we're working in a GitHub project right so in a git project and we have a main branch that we maintain with our train scripts that we're updating in each branch and then a train original that we keep and then we have a data structure on the main branch that we use to just keep the scores then we implemented this in open code for this example but in the repo which you can also go and check out they it's also implemented in codex and Claude if you want to try those I also implemented it in gas town but that's forward stuff so I did it like a separate project but basically it works really anywhere because it's more just a conceptual implementation right and first you have your planner creating hypotheses you have your researchers looking at papers and then your reporter picking all of this up handing to workers as I said those workers integrate with HF jobs so they start these jobs off on the hub that run with the hardware that they need and then they submit these patches that go back the reporter operates in Tracheo which is an open source dashboard that we use for all metrics Tracheo is useful with agents because it uses a completely open data layer basically parquet so if you don't want the dashboard or your agent doesn't want the dashboard for any reason it can just get into the parquet and just do whatever you want so if you need a Gantt chart or some other visualization it can just go and do that so I would say it's the best agent dashboard tool because it's basically just a data store it's basically just a data structure okay so let's just walk through this now so this is implemented in open code if you don't know open code you have agent configurations so in this one I just set auto lab which is the name of the agent configuration it has skills this is the prompt so it says run one autonomous local research or autonomous research parts in the repo using defined roles I tell it to use planner to propose up two fresh single change experiments use reviewer to reject duplicates or stale ideas I also tell it to use a HF bucket because I want all of the storage to be in the same bucket so that I don't have to upload or download the training scripts every time and then we go and we select one of the sub agents as an ISO interface in open code but it's similar in other tools so I select the planner and then you'll see that the planner receives this prompt and it uses a specific template which I defined in Name of the agent configuration I have. It has skills. This is the prompt. So it says run one autonomous local research or research parts in the repo using defined roles. I tell it to use planner to propose up to two fresh single change experiments, use reviewer to reject duplicates or stale ideas. I also tell it to use a HF bucket because I want all of the storage to be in the same bucket so that I don't have to upload or download the training scripts every time. And then we go and we select one of the sub agents as an ISO interface in Open Code, but it's similar in other tools. So I select the planner, and then you'll see that the planner receives this prompt and it uses a specific template which I defined in my configuration. It's going to have current state. It's going to have a list of the jobs so far, things that have worked which were defined by the reviewer, current hyperparameters that it can change, and it's basically just defining these jobs which will go onto the job list as I mentioned. We then switch over to a reviewer agent which will receive all of these jobs. It has a similar kind of structure based on a template, a reference to where it should be working from, and the latest score that it should be using. It gets an overview of all the failed and successful experiments which it will use to base its decisions of what goes into the next queue on. And it creates this little table which we don't really need to look at. It's really just for the agents to interact with each other and get this information back. To be honest, that's a little bit of a verbose example and we maybe don't need this many tables, and you could probably trim that bit down. But in general, I'd recommend if you think this is cool, go and try that out in the repo. After that, this agent runs in parallel sometimes for hours. And this is the Traceo dashboard that we use, and these are all the runs that are pushed to Traceo. As I said, the main advantage here is that this is fully open source and it's just a data layer. But we get all of these kinds of visualizations. Traceo can also have events and warnings, so we can have all of these events being reported by different agents and we can filter those down. We can also tie those up to notifications so you can get emails from Traceo if you want. If your agents are going rogue or something and you need help. But best of all, Traceo has this free form structure so you can just throw tables in that don't necessarily fit with any other structure. And then on the Hub side, all of these jobs are just run inside Hugging Face so you can explore those jobs. And in most cases, you can tell the agents to use labels and you can sort those labels and review what they're doing, or you can just look at it like this. As I mentioned, you can access that underlying data layer and just create a Gantt chart because this was a convenient way to look at what the agents were doing over time. So you can see this amber agent went off and this was the score that it got. But you could visualize this however you want because you have access to this data layer. The TL;DR of the whole thing is that you can go and just have your own AI lab and you can try it out. And if you have a verifiable experiment like training a model or writing CUDA kernels, then it's pretty easy to implement a setup and learn some stuff. So let's now look at the takeaways. I'd say in simple terms that agents work really well with primitives and open primitives. We want tools that are fully open, things like Traceo, things like kernels that we can expose to agents and they can control in their own way. Even though abstracted APIs are really useful, if we have a layer that we can't necessarily get behind, that is a ceiling. So we don't always need to extract. It's more about exposing well. And the other takeaway is that the Hub is ready. The Hugging Face Hub is ready for these kinds of workloads. We have the fundamentals in place like storage, tracking, and compute, which I think will allow us to scale our engineering to new levels. If you found any of this interesting, I've shared it all on X. I've shared it all on Hugging Face, and there's a blog post about basically each one of the examples that I just shared with you, and they all have repos attached to them so you can go and try that out for yourself. If you find anything that's broken, please tell me. If you think that this was completely wrong, come and find me afterwards and then bully me. That's fine. But most of all, thank you. ourselves in new areas and my proposal is that we need to go kind of closer to the silicon and tackle harder problems and and that's where AI systems engineering comes in I've broken this talk down into three progressively more complex steps and more autonomous steps as well and I've defined those like three bosses from games the first one is a hybrid approach where you interactively use an agent to solve a problem and to write a cuda kernel the second is a zero shot task where an agent takes a prompt and trains an LLM on hugging face the third is a multi-agent auto research setup like a kind of automated AI lab so let's get started on the first boss right this is writing cuda kernels so for a while writing custom kernels was seen as this unattainable goal for the humble agent they required complex DSLs they required integration with relevant hardware to be benchmarked and to be tested and it was seen as something that couldn't be achieved by agents however that in most cases was wrong if you look at kernel hackathons like those on gpu mode the recent AMD hackathon if you look at papers like kernel bench you'll see that agents are able to write valid and optimized cuda kernels and that's really cool and something that totally inspires me I'm a part of gpu mode I contribute to that and and something that I think everyone should be doing however what do we do with them how do we distribute them and how do we get them into our inference engine so that we actually are using these optimized kernels that we're generating and that's part of the question of this part of the talk let's take a step back now and just uh say what a kernel is right so when you run an AI model on a GPU the actual work is executed through a kernel this will be defined in a relevant language for that hardware and it will use relevant features to that hardware that may not be available on other hardware we can write custom kernels that will take advantage of that hardware for a specific math operation kind of squeeze everything we can out of it so that the model will infer faster in general this requires a lot of expertise about writing cuda kernels about the hardware and it's also a bit of an insulation hell as you deal with a pretty large install matrix from hardware to software to generations and versions of say cuda and these kind of issues so in short it's hard efficiency in deep learning so efficiency in kernels is split into three main sections one compute two memory and three overhead compute is the flops this is these are the matrix multiplications and the real math of the process memory is the time spent moving data or tensors around memory typically from slow to fast memory and overhead is basically everything else the python environment pytorch dispatch of those kernels these kinds of things in general most people might assume that the compute is the bottleneck here because it's doing most of the math right that's not correct so in most cases memory is usually the bottleneck and that's because a modern gpu let's take a h100 for example can do a petaflop a second of computation but its memory bandwidth is three terabytes so in short the gpu is often waiting idle for this tensors to come back for it for them to be computed there are custom kernels kind of custom optimized kernels that exist flash attention being the poster child of these and in general what they do is increase arithmetic intensity they basically make the gpu do more sums at once per read and write so we move the tensors across we do as much math as possible in the gpu in one go and then we write it back in short people like to say we keep the gpus warm and that's the objective of writing a custom cuda kernel hugging face has a library called kernels which is maintained by kernel writers and we're beginning to scale up to a kind of agentic workloads so at its core this is a way of distributing kernels it has a toml file like any kind of project which says which hardware it works on which versions of cuda and other kind of software as it requires to work and it's a it's now also a repo on the hub just like models so if you are a kernel writer or you're an aspiring kernel writer with an agent that you want to set up you can now be a kernel publisher just like a model publisher and my point is that this is like a kind of super fervent ground for ai engineers looking to kind of scale their career if you check out these repos on the hub you'll see that there's compatibility for different hardware you can configure that so you know like okay this works on my gpu or on my laptop and this is what it looks like here right let's take a look at what this looks like for an agent and how we're helping an agent to do this so first we're going to go to how we do this so skills so i i'm sure everyone here is familiar with skills and i'm sure there have been a number of talks that really go deep into skills i don't like to i like to keep them pretty simple and really they're just kind of file based context with all the wonders of files we can open them and close them we can version them we can source control them and these kinds of things and agents can also do the same they can open them when they need them they can use them when they don't and so in the context of kernels that means that we can give examples of how to write and how to use kernels in skills and they can open those and use them when they need i like to say that it takes a task from being zero shot to being few shot which in ml is quite a familiar concept right we're just giving the agent examples of how to do things and we can be quite verbose and descriptive about that at hugging face we're focusing on integrating skills into their projects so what you'll find is that inside each project there's managed skills by that project which we think is the best way to do this because it means that those projects that the maintainers of those projects are maintaining their skills right that means that they're not necessarily the most like yolo skills because they're kind of like well maintained and robust and we have another repo for those kind of more experimental skills which is called hugging face skills go and check that out if you want to try some of these examples you'll see today in kernels this is what the skill looks like it focuses on benchmarking so it has scripts that allow you to benchmark and test the skill sorry to test the kernel and see how performance it is and references with examples of how to do this we benchmark this skill and we used we generated a a kernel for quen 3 8b for h 100 and we found that we had a 94 speed up this isn't a state-of-the-art speed up on this model by any means it's really just about compatibility and a compatibility matrix so in many cases these models and their kernels won't be optimized for the respective hardware or generation of hardware that you want to use them on so you have some low hanging fruit here where you can just come and pick up some some optimizations for that specific hardware maybe because your hardware is cheap on your cloud provider but it's not necessarily the most ideal for the for that model that you're using so my recommendation would be to come here and like pick up some easy speed ups how do we know that these skills are any good and and that we should be sharing them and telling people to use them we use an open source library called upskill that we're also maintaining this is a is a gateway to using cheaper and open models with skills it basically just generates skills generates an eval for the skill and then allows you to compare different models on the same skill so you can see things like this so okay gpt oss is slightly less accurate using the same tokens kimmy is more accurate using less tokens haiku is a bit more accurate using less tokens and these kinds of things so if you've got a skill and you're using it regularly and you're thinking to yourself okay how can i save a few pennies here and and get a different model on the go then try out upskill and it allow you to iterate on your skill and improve it right let's move on to boss two i'm going to go through this one pretty quickly this is about fine-tuning models if you're really into this there was a talk yesterday by my colleague Mervé that went into this deeply there's also a blog post here where we got claude to do this this was from back in november december time now go and check this out basically you can just say fine-tune quen36b on this data set this is a chain of thoughts data set and you'll improve the models chain of thoughts this is fully integrated to the hub now so you can even run the gpus on on the hub and it uses hf cli skills so it's all very available i would try this one out you can also try this one out this is uses unsloth so it's even cheaper this runs with like optimized models and it's maintained by onslaught and by us and it's another blog post and there's also often free credits that you can get around these blog posts so i go and check these out okay let's move on to the the big one this is auto lab multi-agent research which is a project that kind of um basically keeps me up at night andre carpathy a few weeks ago maybe a month ago now released a project called auto research which was based on his other projects nano gpt and nano chat and it took the nano gpt architecture and got claude code to create improve to write improvements to that training script so that it would improve the training process so we can see here the experiments going over and for each experiment there's a change in the training script which increases the efficiency measured in bits per bytes of that run and we can see that the efficiency ends at its best at the end of the process i like everyone thought this was super cool and i had to start implementing it straight away but one of the things that stood out to me was i found it kind of weird that we had one agent working in a single way iterating going and finding improvements and then implementing them and it would make sense to kind of distribute this so that's what i did i distributed the task amongst the research team with four types we have a researcher that basically looks up papers for this we use hf papers but we can also use archive papers hf papers is cool because it has a cli so you can just pull and search papers from the hub and it acts as a literate literature scout so it just looks up for papers with ideas and it formulates those as hypotheses we then have a planner which takes those hypotheses and maintains like a queue of jobs we then have a set of workers and they pick up those hypotheses and their job is to implement them as training scripts so in many cases just like change the architecture or change the parameter or something and then we have a report a reporter agent that goes and monitors all these jobs and maintains a dashboard that we can use so this is what it looks like if you see here that we have we're working in a in a github project right so in a git project sorry and we have a main branch that we maintain with our train scripts that we're updating in each branch and then like a train original that we that we keep and then we have a data structure on the main branch that we use to just keep the scores then we implemented this in open code for this example but in the repo which you can also go and check out they it's also implemented in codex and Claude if you want to try those i also implemented it in gas town but that's kind of wild west stuff so i did it like a separate project um but basically it works um really anywhere because it's more just a conceptual implementation right and first you have your planner creating hypotheses you have your researchers looking at paper and then your reporter picking all of this up handing to workers as i said those workers integrate with hf jobs so they start these jobs off on the hub that run with the hardware that they need and then they submit these patches that go back the reporter operates in tracheo which is a an open source dashboard that we use for all metrics tracheo is useful with agents because it uses a completely open data layer basically parquet so if you don't want the dashboard or your agent doesn't want the dashboard for any reason it can just get into the parquet and just do whatever you want so if you need a gantt chart or some other visualization it can just go and do that so i would say it's like the best agent dashboard tool because it's basically just a data store you know it's basically just a data structure okay so let's just walk through this now so this is uh it implemented in open code if you don't know open code you have like agent configurations so in this one i just set uh auto lab which is the name of the agent configuration i have it has skills this is the prompt so it says like run one autonomous local research or also research parts in the repo using defined roles i tell it to use planner to propose up two fresh single change experiments use reviewer to reject duplicates or stale ideas i also tell it to use like a hf bucket because i want all of the storage to be in the same bucket so that i don't have to upload or download the training scripts every time and then we go and we select one of the sub agents as an iso interface in open code but it's similar in other tools so i select the planner and then you'll see that the planner receives this prompt and it uses a specific template which i defined in my configuration of like it's going to have current state it's going to have a a list of the jobs so far things that have worked which were defined by the reviewer current hyper parameters that it can change and it's basically just defining these jobs which will go onto the the job list as i mentioned we then switch over to a reviewer agent which will receive all of these jobs it has a similar kind of structure based on a template a reference to where it should be working from and the latest score that it should be using it gets an overview of all the failed and successful experiments which it will then like use to base its decisions of what goes into the next queue on and it creates this little table which we don't really need to look at it's really just for the agents to interact with each other and get this information back to be honest that's a little bit of a verbose example and we maybe don't need this many tables and you could probably trim that bit down but in general i'd recommend if you think this is cool go and try that out in the repo after that so this agent runs in parallel sometimes for hours and this is the tracheo dashboard that we use and these are all the runs that are pushed to tracheo as i said the main advantage here is that this is fully open source and it's just a data layer but we get all of these kinds of visualizations tracheo can also have like events and warnings so we can have all of these events being reported by different agents and we can filter those down we can also even tie those up to like notifications so you can get emails from tracheo if you want if like your agents are kind of going rogue or something and you need help but best of all tracheo just has this uh like just free form structure so you can just throw tables in that don't necessarily fit with any other structure and then on the hub side all of these jobs are just run inside hugging face so you can explore those jobs and in most cases you can tell the agents to use uh like labels and you can sort those labels and review through what they're doing or you can just look at it like this as i mentioned you can access that underlying data layer and just create a gantt chart because this was a kind of convenient way to look at what the agents were doing over time so you can see like this amber agent went off and this was the score that it got but you could visualize this however you want because you have access to this data layer the kind of tldr of the whole thing is that yeah you can go and just have your kind of own ai lab and you can try it out and if you have a verifiable experiment like training a model or doing uh or writing cuda kernels then it's pretty easy to to implement a setup and and to learn some stuff so let's now look at the the takeaways i'd say so that in simple terms i'd say that agents work really well with primitives and and open primitives and we want tools that are fully open things like tracheo things like kernels that we can expose to agents and they can kind of control in their own way even though abstracted apis are really useful if we have a layer that we can't necessarily get behind that that is a ceiling so we don't always need to extract it's more about exposing well and the other takeaway is that the hub is is ready the hugging face hub is ready for these kind of workloads we have the the fundamentals in place like storage tracking and compute which i think will allow us to scale our engineering to yeah new levels if you found any of this interesting i've shared it all on x i've shared it all on hugging face and there's a blog post about basically each one of the examples that i just shared with you and they all have repos attached to them so you can go and try that out for yourself if you find anything that's broken like please tell me off if you think that this was completely wrong come and find me afterwards and then sort of bully me that's fine but most of all thank you