AI Engineer

Your Agent Can Now Train Models — Merve Noyan, Hugging Face

594 summary words 3 min summary Watch video

Start with the signal

3 min read

Summary

Your Agent Can Now Train Models — Merve Noyan, Hugging Face

Main Topics

  • Open Source in Machine Learning: The value and importance of open-weight and open-source models
  • Hugging Face Hub Ecosystem: Infrastructure supporting 3 million models and various AI resources
  • Agentic AI: How agents can now train, deploy, and manage ML models
  • Local Agent Deployment: Running coding agents locally with open models
  • Hugging Face Skills & MCP: Tools that enable agents to manage ML workflows autonomously
  • Practical Applications: Real-world examples of agents handling complex tasks

Key Points

Open Source Advantages

  • Open models now compete with closed models (GLM 5.1 ranks top on benchmarks)
  • Privacy guarantees through edge deployment without data transmission
  • Model customization through quantization and fine-tuning
  • Performance transparency and no surprise degradation

Hugging Face Hub Features

  • Hosts 3 million+ models, datasets, and spaces (AI applications)
  • Benchmark Datasets: Compare open models on SWE Bench Pro, AIME, and other benchmarks
  • Inference Providers: Route to optimal providers based on cost and speed
  • New Traces Repository Type: Host and explore agent execution traces

Agentic Capabilities

  • Local Coding Agents: Run full agents locally using Pydantic with LLAMA CPP
  • Hermes Agents: Advanced memory management for Slack/WhatsApp integration
  • Vision Language Models (VLMs): Released day-zero with vision capabilities (Gemma 4, Qwen, Kimi)
  • Computer Use Agents: VLMs that understand screenshots and can click/navigate

Skills & MCP Integration

Hugging Face Skills allow agents to:

  • LLM Trainer Skill: Train models on datasets with automatic resource calculation
  • Gradient Skill: Build and host demos
  • Dataset Skill: Explore datasets via API
  • CLI Skill: Manage repositories, run jobs, launch demos

Model Context Protocol (MCP) enables:

  • Access to models, datasets, spaces, and jobs
  • Semantic search across Hugging Face resources
  • Query remote AI apps (image generation, etc.)

Model Serving & Deployment

  • GGUF Format: Quantized models compatible with LLAMA CPP, LMStudio, all-LLAMA
  • Hardware Compatibility: Details on VRAM requirements (e.g., Gemma 4 4-bit fits in L4 GPU with 24GB)
  • "Use This Model" Button: Quick-start commands for local deployment
  • Apps Tab: Filter for locally-servable models

Notable Quotes

> "Your agent can now train models... which to me is absolute sci-fi at this point because it used to not exist and there are so many things going on in the back and the agent actually handles them very well."

> "If you have everything in the open, nothing changes without you knowing. No performance degradation without you knowing. Everything's great."

> "Open source models aren't as good as closed. No, this is not the case... we just caught up and we will catch up even more with the upcoming models."

> "I asked GLM 5.1 to fix it with the Hermes agent and it fixed itself and it was a good day."

Takeaways

  • Agents as ML Engineers: Agents can autonomously handle complex ML workflows—training, inferencing, resource calculation, and infrastructure management
  • Open Models Are Competitive: No longer a gap between open and closed models; GLM 5.1 and others lead benchmarks
  • Accessibility: Local deployment with tools like LLAMA CPP makes agent development frictionless
  • Practical Skills: Use Hugging Face Skills to delegate ML ops to agents—they handle napkin math, cost calculation, and instance setup
  • Vision-First Future: Expect all new models to launch with vision capabilities day-one
  • Start Experimenting: Recommended models: GLM 5.1, Gemma 4; tools: Hermes agents, local LLAMA CPP
  • Resource Discovery: Benchmark datasets and inference provider comparisons simplify model selection
Full transcript 2864 words · 24 min read
0:14

SPEAKER_01

Hello everyone and welcome to this talk on open agent ecosystem and I would like to call it having an AI engineer at your fingertips. I'm Merve and I work in the open source team of Hugging Face. How many of you are using Hugging Face on a daily basis? Oh let's change that, this is not okay.

0:28

SPEAKER_01

But first let's talk a bit about open source and what it is. So when it comes to machine learning, open source is absolutely differential. You have the open weight models that go in with non-commercial licenses, we call them open weight and then we have open source models that have commercially available licenses such as this one from DeepSeek. It's called the MIT license or Apache 2.0 and then there is even more open models that have the code open. If you have agents there the harness is open, everything is open and this matters even more by the fact that yesterday or the other day it was revealed that the cloud performance was going down. So if you have everything in the open nothing changes without you knowing, no performance degradation without you knowing everything's great. But on top of it if you have access to the weights you can shrink them, you can quantize them, you can fine-tune them if you feel like it. And it's absolute guaranteed privacy for your end user because you can deploy it to edge devices, browsers without the data going somewhere else. This matters a lot in my opinion even more these days with the security breaches and everything. And there was this argument maybe a few years ago that open source models aren't as good as closed. No, this is not the case. You see for instance the latest GLM 5.1 is absolutely crashing it and I'm actually using it in my coding setup. This is the artificial analysis intelligence index and the green ones are open models meanwhile the black ones are the closed models. And we just caught up and we will catch up even more with the upcoming models and stuff.

0:37

SPEAKER_01

And let's go back to Hugging Face Hub. So everything is facilitated through Hugging Face Hub, all of the open releases. It's the infra layer for all of your open source workflows and as of now it's hosting even more models. I should have updated the number. It's probably close to three million. A lot of data sets, spaces and everything but that's not all when it comes to the agentic ecosystem and this is what we are going to talk about today.

0:42

SPEAKER_01

So when you go to the models you can filter for agentic models. They are mostly the trending ones and there is two types of models in my opinion. There is the vision LMs and then there is the LLMs and the vision LMs can also act as a computer use agent over the screenshots. They know where to click etc which is pretty cool. And one trend I have recently noticed is the fact that you have labs releasing their LLMs with vision capabilities day zero. Like for instance the Gemma 4 was an Omni model and still it's an agentic model. There is Q1 3.5, there is Kimi-K 2.5, these were VLMs. So I foresee that all of these models will be over time released day zero with vision capabilities.

0:50

SPEAKER_01

And it's super easy to run this actually. You can just use VLLM, MLX or LLAMA CPP, LLAMA server from the get-go with few lines of code. It used to be much more friction-y but these days this is not a big deal. And if you want to compare open models we have recently launched this feature called benchmark data sets. So when you go to the data sets on the left hand side there is on the bottom there is a bunch benchmark button. You just click it and then you can see the popular benchmarks such as SWE Bench Pro or Humanities Last Exam or AI ME and others. And when you go to for instance SWE Bench to see how your agent is good in coding and stuff, you see the open models ranked according to the scores. So currently GLM 5.1 is top of the list. So it's also easy to pick an open model these days because there's three million models out there and it used to be a challenge to pick different models.

0:56

SPEAKER_01

And if you actually want to vibe check it Hugging Face has this service called inference providers which does routing for the best models to best providers like all of the providers are there. There's Grok, Cerebras, Novita and everything. And then it's super easy to compare them as well if you see you have the cheapest or the fastest option. Actually I had to truncate it but also there is the tool use column so you can actually pick one of the open source models for the agentic use case and stuff.

1:03

SPEAKER_01

And going back to agents after all of these Hugging Face Hub content. Hugging Face actually recently has shipped a ton of features for you to use open models with agents and stuff. And first off there is the MCP server where you can plug the hub into your LLM and there is skills which allow you to vibe train models like you just go to your agent and say train Q1 3.5 on this data set for me and then it just trains. Which to me is science fiction at this point because it used to not exist and there are so many things going on in the back and the agent actually handles them very well. And then there is the local agent so you can run full coding agents locally from models with Hugging Face hub because we integrate very well to them.

1:10

SPEAKER_01

Coming to the first one so basically my talk will be consisting of all of these. Coming to the first one there is the local coding agents and your options you have many options but one of my favorites is Pydantic because it's super simple to set up. Basically you can I think you can also use it with inference providers remotely but also if you want to serve a local coding agent you can use LLAMA CPP to serve it and then Pydantic will directly consume that. And something very cool is also LLAMA CPP which is baked into LLAMA CPP as a binary that you can just directly execute and start a model by giving Hugging Face hub ID. So it's super easy as well to get a local agent running. I will share my slides on my Twitter account after so no need to take pictures.

1:17

SPEAKER_01

One of my most favorite things these days is Hermes agents and I will just die on this hill. So this is a bit one step even further from the open claw by means of memory management and everything. And it's actually super easy to get started with that and you can either use it locally or with Hugging Face Inference provider. So for instance I was playing with that. The setup wizard does everything for you, you just give the keys and stuff and then integrate into your Slack or WhatsApp or whatever and you're good to go. And I absolutely recommend using this if you want to use it with an open model. I absolutely recommend GLM 5.1. For instance I actually failed initially to integrate into Slack. I have witnesses in here my colleague Nils this year and I asked GLM 5.1 to fix it with the Hermes agent and it fixed itself and it was a good day. I think GLM 5.1 is a very good model and I absolutely cannot wait to use it with Gemma 4. But also this weekend there is on Twitter there was a

1:26

SPEAKER_01

everything for you. You just give the keys and stuff and then integrate into your Slack or WhatsApp or whatever and you're good to go. And I absolutely recommend using this if you want to use it with an open model. I absolutely recommend GLM 5.1. For instance, I actually failed initially to integrate into Slack. I have witnesses in here, my colleague Nils this year, and I asked GLM 5.1 to fix it with the Hermes agent and it fixed on its own and it was a good day. I think GLM 5.1 is a very good model and I cannot wait to use it with Gemma 4. But also this weekend there is on Twitter there was a rumored minimax model coming up, so I will also probably try with that and share my findings. So I absolutely recommend using Hermes agent with the open models. And one more thing, so Hugging Face Hub now has a new dataset repository type called traces and this is basically all of your codecs cloud code or py traces. They host it. And for instance, if you go to your, if you pushed a trace and then you go over there, you will see in the dataset viewer. If you click on the traces column, it pops up like this. It is very nicely parsed and you can just explore your data and then later if you want, you can even train a model on that, which is pretty cool in my opinion. And if you want to push your agent traces, you can just upload your sessions from these paths and nothing else is needed. And we will also probably have Hermes agent very soon for traces. Going back, if you want to use more options to serve LLM behind the agent locally, so some tips and tricks in finding a good model. You just go to Hugging Face. There is an other tab. Under the other tab, there is the apps. So these apps are like LMStudio, Jean, Lama Cpp, everything that is for local serving is over there. And when you filter for them, you have the models that are supported by these local apps. So whatever you want to serve, we have you covered. And when you go to the model repository, something very cool in my opinion is that on the left and right hand side, there is a gguf section. So basically, gguf, if you don't know it, it's supported. It's basically a file format that is supported in Lama Cpp, the file format that is supported in many things like all Lama, LMStudio, everything. And you have the hardware compatibility. For instance, the Gemma 4 larger model, if you quantize it to 4 bit, it fits inside an L4 GPU with 24 gigabytes of VRAM. So I think this is very cool and this is also served to MLX repositories as well. And when you go again to the model repository, if you have absolutely zero clue on how to serve this model, on top right there is "use this model" and you have the options of the local apps that the model is supported in. And when you click that, you see that with only a few lines of command that you can run, you install, you get the model served and voilà. It's very convenient to run the open models these days. And lastly, supercharging your coding agents using Hugging Face skills. So there is a bunch of skills in order to get you started with training, inferring with the open models, using open models, exploring open datasets, using AI apps, everything. And we have this thing called the Hugging Face CLI skill, which allows coding agents to manage repositories, run jobs, launch demos, and everything. And this is how you can install it. You can just type hf skills on Google and you will find the commands. But we have more skills than that. So basically, this allows you to plug hub into your agents. You get all of the Hugging Face Hub exploration. But the rest of the skills are super cool. There is LLM trainer skill. Basically, this is not only for LLMs but also vision language models. You can just tell the model, okay, train this model on this dataset and it will just kick off the job remotely on our infra or locally wherever you want. And there is Gradient skill, which allows you to build demos. And there is Hugging Face dataset skill, which allows you to explore datasets through our dataset viewer API. And you can install it very easily. Again, we come with more integrations. I just put a clone on Gemini here. So putting this into action, for instance, I asked the model. I asked Claude to say, hey, can you train QVNL VL on Lava Instruct Mix, which is a vision language dataset? And it asked me a few questions. It said, okay, which instance would you like this to go in because you have multiple options? The model actually, in the backend, the agent actually calculates the amount of VRAM required to run fine-tune that model in a given batch size and everything. So it handles everything for you. It just asks you a few questions. Okay, what is your validation split, and then it just launches the job, which to me is absolute sci-fi still to this day as a person who has been training models since the beginning of my career, six years. And at the end, you just find your model on hub. And this is not limited to LLMs and VLMs. I have recently shipped skills for, for instance, training object detectors or segment editing models and everything for vision. It handles, for instance, different bounding box types and everything. You just give the command and let it handle everything.

1:31

SPEAKER_01

And going back to MCP, what do we serve? We have models, datasets, spaces, search for your task, semantic search for spaces. So if you don't know spaces, it's the app store of AI. You have a ton of apps over there for absolutely everything you could see. And also, we have something called jobs, which allows you to kick off one-off jobs that ends. If it fails or if it succeeds, and you pay for the amount of time it was up. And also, you can query these apps from MCP. I'm going to show you shortly, but it plays nicely with all of your favorite platforms. So for instance, in here, I asked the model, generate image of a baklava made of yarn. And then it will call the Hugging Face space of QVAN image, which is an image generation model hosted remotely. And then it will query that and it will bring the output of that. It works very nice.

1:39

SPEAKER_01

But you need to turn on. There is a setting in the MCP called dynamic spaces. If you want more options, if you want absolutely all of the spaces, you need to turn that on, which is a bit experimental. And here are some few ideas that you can use spaces MCP. But you are absolutely not limited to those.

1:47

SPEAKER_01

And tying it all together, my colleague Nils has built something which I found cool, so I wanted to share. So basically, on Hugging Face Hub, there are papers. These papers are basically AI-related papers. We want people to be able to ask questions to these papers or share. But not all of the papers come with markdown, which the model can index and stuff. So we OCR'd 30,000 papers using codecs open OCR models and jobs, all through prompting, which is a bit crazy. So the steps to do that is, firstly, pick an OCR model that is cheap and nice and performs well. Ask the LLM to kick off a processing job and actually write the code for that, and then kick it off on Hugging Face Infra. And then let the skill set up the instance of hosting that model and everything without you going through the pain of the napkin math. And then profits. So to pick an OCR model, you need to go to MTEB, which is a benchmark dataset that I have previously shown you.

1:56

SPEAKER_01

which the model which we can index and stuff so we all see our 30,000 papers using codecs open OCR models and jobs all through prompting which is a bit crazy so the steps to do that is firstly pick an OCR model that is cheap and nice and performance ask the LLM to kick off a processing job and actually write the code for that and then kick it off on Hugging Face Infra and then let the skill set up the instance of hosting that model and everything without you going through the pain of the napkin math and then profits so to pick an OCR model you need to you can go to almost here bench which is a benchmark data set that I have previously shown you

2:41

SPEAKER_01

the first result is Chandra OCR but don't be fooled by this we have just today shipped a skill that you can ask the model okay what is the best model on OCR for fine tuning and it will also make recommendations around fine tuning and stuff so if you need smaller models etc it will handle everything for you with this skill so it's pretty cool check it out once you pick the model okay we in this case we use Chandra we ask model to write the script and it did and then the agents just does the napkin math for the instance and calculates the cost of the running job and everything and then these jobs will be so basically these jobs will be rerun so we have

3:24

SPEAKER_01

recently launched this infra product called buckets which is like an S3 buckets but much cheaper and faster that you can use with mounting and yes basically you can just use that and you can get started in these links I hope you like this talk thank you so much these were VLMs. So I foresee that all of these models will be over time released day zero with vision capabilities. And it's super easy to run this actually. Like you can just use like VLLM, MLX or like LAMA CPP, LAMA server from the get-go with like few lines of code. Like it used to be much more friction-y but these days this is a not a big deal. And if you want to compare open models we have recently launched

4:26

SPEAKER_01

this feature called benchmark data sets. So when you go to the data sets on the left hand side there is like on the bottom there is a bunch benchmark button. You just click it and then you can see the popular benchmarks such as SWE Bench Pro or Humanities Last Exam or AI ME and others. And when you go to for instance SWE Bench to see like how your agent is like good in coding and stuff, you see the open models ranked according to the scores. So like currently GLM 5.1 is top of the list. So it's also easy to pick an open model these days because there's three million models out there and it used to be a challenge to

5:11

SPEAKER_01

pick different models. And if you actually want to vibe check it Hugging Face has this service called inference providers which does routing for the best models to best providers like all of the providers are there. There's Grok, Cerebras, I don't know, Novita and everything. And then it's super easy to compare them as well if you see like you have the cheapest or the fastest option. Actually I had to truncate it but also there is the tool use column so you can actually pick one of the open source models for the agentic use case and stuff. And going back to agents after all of these Hugging Face Hub shill. Hugging Face

5:55

SPEAKER_01

actually recently has shipped a ton of features for you to use open models with agents and stuff. And first off like there is the MCP server where you can plug the hub into your LLM and there is skills which allow you to even vibe train models like you just go to your agent and say train Q1 3.5 on this data set for me and then it just trains. Which to me is like a sci-fi at this point because it used to not exist and like there is so many things going on in the back and the agent actually handles them very well. And then there is the local agent so you can run full coding agents locally from models with

6:43

SPEAKER_01

Hugging Face hub because we integrate very well to them. And coming to the first one so basically my talk will be consisting about all of these. Coming to the first one there is the local coding agents and your options you have like actually many many options but like one of my favorites is Pi because it's like super simple to set up. Basically you can I think you can also use it with inference providers remotely but also if you want to serve like a local coding agent you can use LLMACPP to serve it and then Pi will directly consume that. And something very cool is also LLMACPP which is baked into LLMACPP as a binary

7:25

SPEAKER_01

that you can just directly execute and start a model by giving Hugging Face hub ID. So it's super easy as well to get a local agent running. I will share my slides on my Twitter account after so no need to take pictures. One of my most favorite things these days is Hermes agents and I will just die on this hill. So this is like this is a bit one step even further to from the open claw by means of memory management and everything. And it's actually super easy to get started with that and it is you can either use it locally or with Hugging Face Inference provider. So for instance I was playing with that. Like the setup wizard does

8:14

SPEAKER_01

everything for you you just give the keys and stuff and then integrate into your Slack or WhatsApp or whatever and you're good to go. And I absolutely recommend using this if you want to use it with an open model. I absolutely recommend GLM 5.1. For instance I actually failed initially to integrate into Slack. I have witnesses in here my colleague Nils this year and I asked GLM 5.1 to fix it with the Hermes agent and it's fixed on its own and it was a good day. Like I think GLM 5.1 is a very good model and I cannot I can absolutely wait to use it with Gemma 4. But also this weekend there is like on Twitter there was a

9:03

SPEAKER_01

rumored minimax model coming up so I will also probably try with that and share my findings. So I absolutely recommend using Hermes agent with the open models. And one more thing so basically Hugging Face Hub now has a new dataset repository type called traces and this is basically all of your codecs cloud code or py traces they host it. And for instance if you go to your if you pushed a trace and then you go over there you will see in the dataset viewer if you click on the traces column it pops up like this. It is very nicely parsed and you can just explore your data and then later if you want you can even train a model

9:56

SPEAKER_01

on that which is pretty cool in my opinion. And if you want to push your agent traces you can just upload your sessions from these paths and nothing else is needed. And we will also probably have Hermes agent very soon for traces. Going back if you want to use if you want more options to serve LLM behind the agent locally so some tips and tricks in finding a good model you just go to Hugging Face there is an other tab. Under the other tab there is the apps so these apps are like LMSudio, Jean, Lama Cpp, everything that is for local serving is over there and when you filter for them you have the models that are supported by these

10:46

SPEAKER_01

by these local apps so whatever you want to serve we have you covered. And when you go to the model repository something very cool in my opinion is that on the left and right hand side there is gguf section so basically gguf if you don't know it's supported it's basically comes in Lama Cpp the file format that is supported in many things like all Lama, LMSudio, everything and you have the hardware compatibility for instance the Gemma 4 larger model if you quantize it to 4 bit it fits inside an L4 GPU with the 24 gigabytes of VRAM. So I think this is very cool and this is also served to MLX repositories

11:34

SPEAKER_01

as well and when you go to the again to the model repository if you have absolutely zero clue on how to serve this model on top right there is use this model and you have the options of the local apps that the model is supported in and when you click that you see like only with few lines of command that you can run you install you get the model served and voila it's very very convenient to run the open models these days and lastly supercharging your coding agents using hug and face skills so there is we have like bunch of skills in order to get you started with training I don't know inferring with the open models using open models

12:19

SPEAKER_01

exploring open data sets using AI apps everything and we have this thing called the Hugging Face CLI skill which allows coding agents to manage repositories run jobs launch demos and everything and this is how you can install it you can just type hf skills on google and you will find the commands but we have more skills than that so basically this allows you to plug hub in into your agents like give you all of the hugging face hub exploration but the rest of the skills are super cool there is LLM trainer skill basically this is uh this is not only for LLMs but also vision language models you can just tell the model

13:04

SPEAKER_01

to okay train this model on this data set and it will just kick off the job remotely uh on our infra or like you locally wherever you want and there is gradient skill which allows you to build demos and there is hugging face data set skill which allows you to uh explore data sets through our data set viewer API and you can install it very easily again we come with more integrations I just put a cloud on Gemini here so putting this into action for instance I asked the model to I asked cloud code to say hey can you train QVN2 VL on lava instruct mix which is like a vision language data set and it asked me a few

13:55

SPEAKER_01

questions it said okay which instance would you like this to go in because you have multiple options uh the model model actually like in the backhand the agent actually uh calculates the amount of VRAM required to run fine-tune that model in a given batch size and everything so it handles everything for you it just asks you a few questions okay what is your validation split blah blah and then it just launches the job which to me is absolute sci-fi still to this day as a person who have been training models since I don't know beginning of my career like six six years and you at the end you just find your model on hub and

14:38

SPEAKER_01

this is not limited to LLMs and VLMs I have recently shipped um skills for for instance training object detectors or I don't know segment editing model and everything for vision it handles for instance different bonding box types and everything you just give the command and let it handle everything and going back to MCP what do we serve we have models data set spaces search for your task semantic search for spaces so if you don't know spaces it's like the app store of AI you have a ton of apps over there for absolutely everything you could see and also we have something called jobs which

15:22

SPEAKER_01

allows you to kick off one of jobs that ends like if it fails or if it succeeds and you pay for the amount of time it was up and also you can query these apps from MCP like I'm going to show you shortly but it plays nicely with all of your favorite platforms and so for instance in here I asked the model generate image of a baklava made of yarn and then it will call the hugging face space of QVAN image which is an image generation model hosted remotely and then it will query that and it will bring the output of that it works very nice look but you need to turn on there is a setting in the MCP called dynamic spaces if you want more options of like

16:14

SPEAKER_01

if you want absolutely all of the spaces you need to turn that on which is a bit of bit experimental and here are some few ideas that you can use spaces MCP but you are absolutely not limited to those and tying it all together my colleague Nils has built something which I found cool so I wanted to share so basically on Hugging Face Hub there is papers and these papers basically AI related papers we want people to be able to ask questions to these papers or share but not all of the papers come with markdown which the model which we can index and stuff so we all see our 30 30 000 papers using codecs open OCR models and jobs

17:01

SPEAKER_01

all through prompting which is a bit crazy so the steps to do that is firstly pick an OCR model that is cheap and nice and performance ask the LLM to kick off a processing job and actually write the code for that and then kick it off on Hugging Face Infra and then let the skill set up the instance of hosting that model and everything without you going through the pain of the napkin math and then profits so to pick an OCR model you need to you need you can go to almost here bench which is a benchmark data set that I have previously shown you the first result is Chandra OCR but don't be fooled by this we have just today shipped a skill that you can just ask the model

17:49

SPEAKER_01

okay what is the best model on OCR for fine tuning and it will also make recommendations around fine tuning and stuff so if you need like smaller models etc it will handle everything for you with this skill so it's pretty cool check it out once you pick the model okay we in this case we use Chandra we ask model to write the script and it did and then the agents just does the napkin math for the instance and the calculates the cost of the running job and everything and then these jobs will be so so basically these jobs will be rerun so we have recently launched this infra product called buckets which is like a Ace 3 buckets but much cheaper and

18:37

SPEAKER_01

faster that you can use with mounting and yeah basically you can just use that and you can get started in these links i hope you like this talk thank you so much are are

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note