SPEAKER_01
Good morning everyone. How many of you here have ever tried to run AI on your phone, on your Macbook? How was that experience? Was it good? More or less? Okay, so this talk today is for you. I'm going to show you how you can deploy and manage AI agents or even voice agents, if you will, completely on device using MLX. Today's agenda, of course, we're going to start with why on device. A lot of you use cloud code subscriptions and many other subscriptions. I want to convince you today to offload some of that subscription completely on device and then all you need to pay is your energy bill. Then I'm going to talk about MLX and I'll give you a small demo and I'll show you some of the amazing community projects that I've seen built by the community using the projects that I'll demonstrate today. In 2020, something very magical and also weird happened. It was the best yet also worst year of my life. It was a weird one. On one side my dad became blind and on the other side Apple released one of the most powerful chips for on-device intelligence ever. And I remember I was talking to my dad and I said, hey, I promise you that I'll get you back to reading. He's one of the most voracious readers I know and losing his sight was one of the main things that bummed him out. He could no longer consume information. And I made this really weird promise. I said, hey, I'm going to fix this some way, somehow. And at the same time this happened. And I thought, compute on the cloud doesn't necessarily solve all of these use cases because my dad lives in Africa and there we don't have internet as easy as we have here or the subscription plans there are really bad. So I thought on-device is the future. And then 2023, I was investigating on GitHub some really cool projects and I saw this project here called MLX. It's an array framework for Apple Silicon. You can imagine PyTorch or TensorFlow for Apple Silicon. And I tried out the initial example and I thought, there's a future here. There's a future that was promised for all of us that all of the big companies like Meta and Google could not really deliver because they were trying to optimize for scale for the cloud. And Apple did something different. So then I started contributing to MLX. Three years later, we have over 1.5 million downloads, over 4,000 models ported, and we work with some of the best frontier labs to deliver to you day zero support for all of your open source models. You can imagine Gemma 4, the latest Gemma 4, we had day zero support for that on MLX, meaning you can run all of the best frontier open source models completely on your MacBook, on your iPhone, or your iPad. So let's start with vision, which for him was the biggest sensory deprivation. He cannot see. He cannot navigate the world. And I thought, well, the easiest way to do that is by giving him his vision back or giving him a system that can help him navigate the world. And I think in 2021, I went into this hackathon and I built these goggles with my team that could tell you what's in front of you. But then MLX VLM became the second iteration of that that allows you to do this not only via some weird glasses, but even on your iPhone. You can now just pick up your phone, point it at something, and you'll be able to understand what's in front of you. And then you have Omni models nowadays, beyond just vision models. These are models that can also take in audio. For my dad, in particular, typing is not really a reality, but he can speak. And with his speech, he can control the camera and understand what's in front of him, what he needs to do and navigate the world. So this is MLX VLM. Right now, if you use LM Studio, it's one of the main engines that powers LM Studio, powers Liquid AI models and many other models out there. But when you think on device, you might think, well, that's weird. I cannot run really large models. Well, that's not true anymore. You can now run models of hundreds of billions of parameters even on your initial M1 MacBook. There's a lot of improvements that the community has made that allows you to run even the largest and most abnormal models completely on device. I have some examples that show that you can run models like Gemma 426B on an iPhone using your storage, and you can still get reasonable speeds. Then, after a year or two, I was also experimenting with how can we enable humans more and give them more accessibility, especially the ones that don't have all the senses. And I thought, audio is the next iteration. But it was for that and also for a very selfish reason. I wanted to be able to control my computer without being in front of my computer all the time. What if I could just blur a command and have my computer do it? This is more of the Jarvis vision of the world where you can just speak to your computer and have actions done for you. We started off with text-to-speech and then that became Marvis, one of our custom models that can generate audio in less than 100 milliseconds. Then we have speech-to-text, which allows you to speak to your computer and have it transcribed in real time. So, for example, if you ever use Whisper Flow or Super Whisper, yeah, you can now code that application, just point Cloud Code or Codex into MLX audio, ask it to build it for you and you'll have it in like 10 minutes. And then you also have speech-to-speech. So, beyond the first two capabilities, a big unlock is to have the computer speak back to you. So, speech-to-speech is one of the core capabilities that we recently added and we also support both Python and Swift. We started off with Python because it's much easier, it scales faster, but we also understand that native experiences matter. So, with Swift, you can now build fully native applications enabled by audio intelligence as well as vision intelligence. And on the right side there, you have our modular pipeline that beyond just models that are speech-to-speech natively, you can actually chain a series of different capabilities to create a modular speech pipeline. For instance, you can choose which automatic speech recognition model you want, you can choose which language model you want, and you can also choose which text-to-speech model you want. And this way you can create really custom and modular experiences that fit on every single hardware budget. So, if you have a very simple M1 first generation Apple Silicon or even the latest, you can adjust that to your hardware. But then, you might think, well, speech-to-speech or text-to-speech is not that good. I've seen some videos, audios of speech-to-speech that may be, I don't know, does it sound good? I can promise you that it does and I'll show you a demo in a bit. So, let's start with vision. And with vision, I have a couple of examples. The first one here is real-time image analysis so you can understand what's happening. It's a very simple command and we'll make this even simpler. Right now,
SPEAKER_01
create really custom and modular experiences that fit on every single hardware budget. So, if you have a very simple M1 first generation Apple Silicon or even the latest, you can adjust that to your hardware. But then, you might think, well, speech-to-speech or text-to-speech is not that good. I've seen some videos, audios of speech-to-speech that may be—I don't know, does it sound good? I can promise you that it does and I'll show you a demo in a bit.
SPEAKER_01
So, let's start with vision. And with vision, I have a couple of examples. The first one here is real-time image analysis so you can understand what's happening. It's a very simple command and we'll make this even simpler. Right now, it's in Python, it will come to Swift very soon. But if I run this command, it's going to run the RF Detour model by RoboFlow. And as you can see, this is completely real-time. It's understanding where, and this is all running on my computer. I can actually just, to make sure that this is very clear, I don't know if I turn off the internet, what's going to happen. But here it is continuously running. I can grab a glass and we'll also detect that. Of course, it's thinking it's fine, but I'm sober. So, this is running real-time completely on the device on my Mac. It can also run on your phone. And this is one example. I want to show you another really cool example that you can do with this particular use case. Have you ever tried the, when you're in a meeting, you want to blur the background? You know that Google does this and et cetera. So now, you can actually do this natively and you can build this kind of experiences into your products. I'm not sure you can see that it's blurring the background, but it actually is blurring the background. And it's detecting my mask in real-time and will detect other objects as well.
SPEAKER_01
All right. So, that's example number one. Number two is you can run really large models completely on the device. Here's Gemma 4. It was released, I think, a week ago. A couple, yeah, last week. It was released last week. And with MLXVLM, you just run MLXVLM.chat_UI and you pass in the model you want. And it should load that model and give you a simple interface for you to get started with radio. Wait. Not sure what's happening. This is the problem with demos. Sometimes the gods don't want it to do the demo. Give me a second. Okay. Not sure what's happening. Okay. Well, I need to quickly get something to close off.
SPEAKER_01
So, okay. It stopped. Let's start. Let's try that again. Okay. It's trying to download some model files, but now I'll turn off the internet again just to show you that this is running completely on device. So, now we have a chat window. I don't know if you can see this or I have to zoom in more. And we can choose any image to analyze. Let me see a very simple one. Okay. Okay. Here. Describe this image in detail.
SPEAKER_01
So, here it is. It's saying that it's a profile of a man named Prince Kanuma. And it's picking up all the different details like my bio and et cetera. And all of this is running on device. It's using the GPU in this particular device. So, this particular machine has 96 gigabytes of VRAM. So, I can run actually all of the models I showcase so far in real time, all of them at the same time. So, this is demo number two. Let's now get back. I wanted to show you audio, but it seems like the Swift branch, there's something there I couldn't really figure out this morning. But for the sake of examples, I will show you some of the really cool community use cases. We have one here, one of the creators. I didn't include his particular use case, but I'll show you a video in just a bit. So, let me try and get this in full screen so that you can see better.
SPEAKER_01
One second. Okay. So, we are back. The first example is grounded visual reasoning. You can use Gemma 4. You can use the model I just showcased, the RF Theater or any of our perception models and you can create really cool experiences like here. You can detect all the fires. You can ask it to detect particular items in the video and this is all going to happen completely on device without the use of internet. And now you can have, for example, security systems that run completely on a MacBook on your house and you can analyze even your dash cam. I have a dash cam video and I've had some really cool, or not cool, but crazy experiences. And I usually use this kind of system here that I built to analyze the video footage afterwards.
SPEAKER_01
And then you have example number two. This is more of an honorable mention, which is cartoons generated completely on device using MLX Video, one of the most recent projects I'm running. And here's one of the coolest videos generated by one of our users. This is all generated on device with a simple text prompt.
SPEAKER_01
And he did something very interesting, which is he chained... This is not one video that he generated all at once. What he did is he chained a system that can continuously generate from the video. So he can create a cohesive story, even though it was not one-shotted. And this particular system can run even on a MacBook with 16 gigabytes of VRAM. So yeah, pretty cool. He has a lot more on Twitter. I will put his... Like I'll share, I'll re-share on my Twitter. So if you check my Twitter, I'll re-share all of these videos.
SPEAKER_01
And then I think one of the latest ones that I will show. But before that, let me show you one here from one of our participants on the floor. So if I go to Drive... Let me see... And my Drive... Okay, Neowa Labs... You have... Here. This is actually by this gentleman here in front, Adrian. He builds this application called Locally. And using MLX Audio and Marvis TTS, he now gave that particular application the capability of speaking back to its users. It was a small voice... Can you increase the volume? The chat subo was a bar for professional expatriates. You could drink there for a week and never hear two words in Japanese.
SPEAKER_01
All right. So that is one of the examples where you can actually build really beautiful native experiences. And if you have a good touch of design, a killer application as well. And finally, I think this is one of the most exciting parts of where I think we are going next with on-device AI, which is robotics. So last year, I acquired a robot called Ricci Mini. And the way that I power my Ricci Mini in particular is that I use MLX Audio, MLX Vision to give it all the capabilities or perception capabilities using its camera and audio input. So here it is. It's also doing voice, real-time voice cloning of the original, and never hear two words in Japanese.
SPEAKER_01
All right. So that is one of the examples where you can actually build really beautiful native experiences. And if you have a good touch of design, a killer application as well.
SPEAKER_01
And finally, I think this is one of the most exciting parts of where I think we are going next with on-device AI, which is robotics. So last year, I acquired a robot called Ricci Mini. And the way that I power my Ricci Mini in particular is that I use MLX Audio, MLX Vision to give it all the capabilities or perception capabilities using its camera and audio input. So here it is. It's also doing voice, real-time voice cloning of the original, how do you call it? Iron Man Jarvis voice. So I hope you can hear this. Hey, Jarvis. Hey there. Great to see you. How's it going?
SPEAKER_01
So you can chain and build a lot of really cool applications. This is just a start. And I hope that in the future, you can understand, or from today, you can understand that you can build agents that can hear, see, and sound just like you or one of your loved ones today running on your iPhone, iPad, Mac, or even your robot. Thank you. Questions? Do you already, Apple is promoting their neural engine a lot? Yes. And I've also tried it on MLX. Yeah. But if you look at the usage of the neural engine, it's always stuff like zero.
SPEAKER_01
Yeah. That's a great question. So MLX uses the GPU, not the neural engine. For you to enable the neural engine, you need Core ML. And right now, Core ML, it does not really run like it's not an easy experience for developers. I hope by WWDC, Apple solves the private API issues. And when they do, we have some internal projects that can allow you to run a hybrid inference across both. Yeah. Will it be? It will be. It will be. Yeah.
SPEAKER_01
But we also think that they might be changing the neural engine and putting some components into the GPU with, for example, the M5 series. You can see that it already has some components of it. And we just don't know what direction they are heading. But it's exciting. Let's wait for WWDC. Questions? Yeah. Is there a tool to see your GPUs in the Mac record?
SPEAKER_01
Yes. So the easiest tool for you to get started is MacTop. If you run MacTop, it's going to show you all your usage. This is by Carson. He's a really cool dude. So you can see here what's happening, the GPU, CPU. And you can have this overlay across your device. And if we do run inference, let's say I start a new window and if I can find that command. Okay. So if I do start running inference on this, let me put here for a bit. And okay. You will see that the GPU now is going to start moving up. And if I say, hi, are you? You see, the GPU is already moving up. And then this is one of the easiest ways that I found to track the performance. Any other questions? Yep.
SPEAKER_01
You also mentioned Omni models. Yes. What's the one that you would recommend? What is the state of the art? So you have a couple options. The first one is Gemma 4. The E version. They have this Gemma 4 with the number E or the letter E and then a number E4, E2. Those are Omni models. They take image, audio, and text as input or any of the variation of the three. And then you also have QAN 3 Omni, which is a much larger model, around 30 billion parameters, but you can also run on device. Those are the top ones that I know. What are the key limitations of those models at the moment?
SPEAKER_01
One of the key limitations? I think it depends on your particular use case. So there are certain things the models just cannot do. You're not going to get the performance of Cloud 3 or 4.6 Opus today, but maybe in six months, these open source models will have that performance. So the experience should be adjusted to the expectations of performance. That's the only thing I would say. Outside of that, I don't see any limitations. You can run inference on hundreds of images in parallel. You can run inference on many, many documents. And you now have context of up to a million, thanks to a recent breakthrough that I made with TurboQuant. So now you can actually serve one million context completely on the device, depending on the size of the model and your hardware. But you can do that today.
SPEAKER_01
Male Speaker 1 Yes. Male Speaker 1 Male Speaker 1 Male Speaker 1 Male Speaker 1 Male Speaker 1 Male Speaker 1 Male Speaker 1 Male Speaker 1
SPEAKER_01
So I was one of the first people in the world to implement TurboQuant publicly. So like 30 minutes after the paper was out, I already had implemented it. And I made this tweet at 3 AM. I didn't know that it would go viral. But if you go to my profile and you write TurboQuant, you'll be able to see there's this particular post here. So this was like 25th March and pretty much the same day, but just was midnight. And it got 700,000 views because of that. So TurboQuant does work. Example is the full model takes almost one gigabyte of KV cache or RAM. And by using TurboQuant, you can reduce that by 4X. Male Speaker 1 Male Speaker 1 Male Speaker 1
SPEAKER_01
Yeah. So similar quality. As you can see, when you see exact match, it means that it matches the performance of the responses of the full model. And I also publicize the full results and performance. For example, here, when you get to 300,000 contexts, the performance almost doubles in terms of throughput. So yeah. Yeah. This is one of the many things that I try to do to enable on-device to go even further. Yep. Any other questions? All right. Thank you. Male Speaker 1 analysis so you can understand what's happening. It's a very simple command and we'll make this even simpler. Right now,
SPEAKER_01
it's in Python, it will come to switch very soon. But if I run this command, it's going to run the RF detour model by RoboFlow. And as you can see, this is completely real-time. It's understanding where, and this is all running on my computer. I can actually just, to make sure that this is very clear, I don't know if I turn off the internet, what's going to happen. But here it is continuously running. I can grab a glass and we'll also detect that. Of course, it's thinking it's fine, but I'm sober. So, this is running real-time completely on the device on my Mac. It can also run on your phone. And this is one example. I want to show you another really cool
SPEAKER_01
example that you can do with this particular use case. Have you ever tried, have you ever tried the, when you're in a meeting, you want to blur the background? You know that Google does this and et cetera. So now, you can actually do this natively and you can build this kind of experiences into your products. I'm not sure you can see that it's blurring the background, but it actually is blurring the background. And it's detecting my mask in real-time and will detect other objects as well. All right. So, that's example number one. Number two is you can run really large models completely
SPEAKER_01
on the device. Here's Gemma 4. It was released, I think, a week ago. A couple, yeah, last week. It was released last week. And with MLXVLM, you just run MLXVLM.chat underscore UI and you pass in the model you want. And it should load that model and give you a simple interface for you to get started with radio. Wait. Not sure what's happening. This is the problem with demo. Sometimes the gods don't want it to do the demo. Give me a second. Okay. Not sure what's happening.
SPEAKER_01
Okay. Well, I need to quickly get something to close off.
SPEAKER_01
So, okay. It stopped. Let's start. Let's try that again. Okay. It's trying to download some model files, but now I'll turn off the internet again just to show you that this is not running completely on device. So, now we have a chat window. I don't know if you can see this or I have to zoom in more. And we can choose any image to analyze. Let me see a very simple one. Okay. Okay. Here. Describe this image in detail.
SPEAKER_01
So, here it is. It's saying that it's a profile of a man named Prince Kanuma. And it's picking up all the different details like my bio and et cetera. And all of this is running on device. It's using the GPU in this particular device. So, this particular machine has like 96 gigabytes of VRAM. So, I can run actually all of the models I showcase so far in real time, all of them at the same time. So, this is demo number two. Let's now get back. I wanted to show you audio, but it seems like the Swift branch, there's something there I couldn't really figure out this morning. But for the sake of examples, I will show you some of the
SPEAKER_01
really cool community use cases. We have one here, one of the creators. I didn't include his particular use case, but I'll show you a video in just a bit. So, let me try and get this in full screen so that you can see better. One second. Okay. So, we are back. The first example is grounded visual reasoning. You can use Gemma 4. You can use the model I just showcased, the RF theater or any of our perception models and you can create really cool experiences like here. You can detect all the fires. You can ask it to detect particular items in the video and this is all gonna happen completely on device without the use of internet. And now you can have,
SPEAKER_01
for example, security systems that run completely on a MacBook on your house and you can analyze even your dash cam. I have a dash cam video and I've had some really cool, or not cool, but crazy experiences. And I usually use this kind of system here that I built to kind of analyze the video footage afterwards. And then you have example number two. This is more of like a honorable mention, which is cartoons generated completely on device using MLX video, one of the most recent projects I'm running. And here's one of the, the coolest video generated by one of our users. This is all generated on device with a simple text prompt.
SPEAKER_01
So...
SPEAKER_01
And he did something very interesting, which is he chained... This is not like one video that he generated all at once. What he did is like, he chained a system that can continuously generate from the video. So he can create a cohesive story, even though it was not like one-shotted. And this particular system can run even on a MacBook with 16 gigabytes of VRAM. So yeah, pretty cool. He has a lot more on Twitter. I will put his... Like I'll share, I'll re-share on my Twitter. So if you check my Twitter, I'll re-share all of these videos. And then I think one of the latest ones that I will show. But before that, let me show you one here from one of our
SPEAKER_01
participants in the... on the floor. So if I go to Drive...
SPEAKER_01
Let me see...
SPEAKER_01
And my Drive... Okay, Neowa Labs... You have...
SPEAKER_01
Here. This is actually by this gentleman here in front, Adrian. He builds this application called Locally. And using MLX Audio and Marvis TTS, he now gave that particular application the capability of speaking back to its users. It was a small voice... Can you increase the volume? The chat subo was a bar for professional expatriates. You could drink there for a week and never hear two words in Japanese. All right. So that is one of the examples where you can actually build really beautiful native experiences. And if you have a good touch of design, a killer application as well. And finally, I think this is one of the most exciting parts of where I think we are going next
SPEAKER_01
with on-device AI, which is robotics. So last year, I acquired a robot called Ricci Mini. And the way that I power my Ricci Mini in particular is that I use MLX Audio, MLX Vision to give it all the capabilities or perception capabilities using its camera and audio input. So here it is. It's also doing voice, real-time voice cloning of the original, how do you call it? Iron Man Jarvis voice. So I hope you can hear this. Hey, Jarvis.
SPEAKER_01
Hey there. Great to see you. How's it going?
SPEAKER_01
So you can chain and build a lot of really cool applications. This is just a start. And I hope that in the future, you can understand, or from today, you can understand that you can build agents that can hear, see, and sound just like you or one of your loved ones today running on your iPhone, iPad, Mac, or even your robot. Thank you.
SPEAKER_01
Questions?
SPEAKER_01
Do you already, like, Apple is, like, promoting their neural engine a lot? Yes. And I've also tried it on MLX. Yeah. But if you look at the usage of the neural engine, it's always stuff like zero. Yeah. That's a great question. So MLX uses the GPU, not the neural engine. For you to enable the neural engine, you need Core ML. And right now, Core ML, it does not really run like it's not an easy experience for developers. I hope by WWDC, Apple solves the private API issues. And when they do, we have some internal projects that can allow you to run a hybrid inference across both. Yeah. Will it be? It will be. It will be. Yeah.
SPEAKER_01
But we also think that they might be changing the neural engine and putting some components into the GPU with, for example, the M5 series. You can see that it already has some components of it. And we just don't know what direction they are heading. But it's exciting. Let's wait for WWDC. Questions? Yeah. Is there a tool to see your GPUs in the MacI record? Yes. So the easiest tool for you to get started is MacTop. If you run MacTop, it's going to pretty much show you all your usage. This is by Carson. He's a really cool dude. So you can see here pretty much what's happening, the GPU, CPU. And you can have this overlay across your
SPEAKER_01
device. And if we do run inference, let's say I start a new window and if I can find that command. Okay. So if I do start running inference on this, let me put here for a bit. And okay. You will see that the GPU now is going to start moving up. And if I say, hi, are you? You see, the GPU is already moving up. And then this is one of the easiest ways that I found to track the performance. Any other questions? Yep. Like, you also mentioned Omni models. Yes. Like, what's the one that you would recommend? Like, what is the state of the art? So you have a couple options. The first one is Gemma 4. The E version. They have this
SPEAKER_01
Gemma 4 with the number E or the letter E and then a number E4, E2. Those are Omni models. They they take image, audio, and text as input or any of the variation of the three. And then you also have QAN 3 Omni, which is a much larger model, around 30 billion parameters, but you can also run on device. Those are the top ones that I know.
SPEAKER_01
What are the key limitations of those models at the moment? One of the key limitations? I think it depends on your particular use case. So there are certain things the models just cannot do. You're not going to get the performance of Cloud 3 or 4.6 Opus today, but maybe in six months, these open source models will have that performance. So the experience should be kind of adjusted to the expectations of performance. That's the only thing I would say. Outside of that, I don't see any limitations. You can run inference on hundreds of images in parallel. You can run inference on many, many documents. And you now have context of up to a million,
SPEAKER_01
thanks to a recent breakthrough that I made with TurboQuant. So now you can actually serve one million context completely on the device, depending on the size of the model and your hardware. But you can do that today.
SPEAKER_01
Male Speaker 1 Yes. Male Speaker 1 Male Speaker 1 Male Speaker 1 Male Speaker 1 Male Speaker 1 Male Speaker 1 Male Speaker 1 Male Speaker 1 So I was one of the first people on the world to implement TurboQuant publicly. So like 30 minutes after the paper was out, I already had implemented it. And I made this tweet at like 3 AM. I didn't know that it would go like this viral. But if you go to my profile and you write TurboQuant, you'll be able to see there's this particular post here. So this was like 25th March and pretty much the same day, but just was like midnight. And it got like 700,000 views because of that. So
SPEAKER_01
TurboQuant does work. Example is like the full model takes almost one gigabyte of KV cache or RAM. And by using TurboQuant, you can reduce that by 4X. Male Speaker 1 Male Speaker 1 Male Speaker 1 Yeah. So yeah, similar quality. As you can see, when you see exact match, it means that it matches the performance of the responses of the full model. And I also publicize the full results and performance. For example, here, when you get to like 300,000 contexts, the performance almost doubles in terms of like throughput. So yeah. Yeah. This is one of the many things that I try to do to enable on-device to go even further. Yep. Any other questions? All right. Thank you.
SPEAKER_01
Male Speaker 1