[0:00] Local models just got a massive upgrade. [0:03] Just over the past two months, we've seen some really good local models [0:06] drop that are very capable and relatively easy to run. [0:09] Even on just mid-tier hardware. [0:11] So in this video, I'm going to show you how you can run local models on your own computer [0:15] and then connect them to tools like OpenClaw so you can genuinely save thousands [0:19] of dollars per month and no longer need to rely on any kind of cloud provider. [0:23] Now, I am going to explain this extremely in-depth because there are some trade offs. [0:26] You need to pick the correct local model, and not everyone should be running local models. [0:30] It really depends on your use case, [0:32] what you're looking for, security, and then obviously the type of hardware you have. [0:35] Anyways, let's dive in. [0:37] So the experience that you're going to have running local models is really going to depend [0:40] on the model that you choose and the type of hardware that you're running. [0:44] Now I want to go over both. [0:45] And while it's going to seem a little bit complicated, it's really important to understand this [0:49] because making the wrong decision here can have a massive consequence and force you [0:53] to go back to using these cloud based models, which are going to be really expensive. [0:57] Now, the first thing to understand is that the local models [1:00] are open source, or at least most of the case they are. [1:03] What that means is that these are models that are openly available, [1:06] that are free to use, that anyone can download, modify, view, doesn't matter. [1:10] Okay. [1:11] So you're not going to be able to run, [1:12] you know, opus 4.7 locally on your own computer because that's gated by anthropic. [1:16] And they want you to pay them a massive amount of money to use it instead of tools like OpenClaw. [1:21] Same with GPT 5.4 or whatever the newest model is. [1:24] You can't run that locally. [1:25] You have kind of a limited selection while it is still quite large actually right now. [1:29] And when you are picking those models, there's only certain models [1:32] that are going to be compatible with tool calling or agent kind of orchestration platforms [1:37] like Open Claw, which I promise you, we're going to get into [1:40] now in order to determine which local models you're capable of running, [1:43] you do need to understand the hardware that you're working with. [1:46] These models are very demanding on the hardware [1:49] and specifically on the Ram and the graphics processing unit that you have in your computer. [1:55] So before we go much further, I'm going to ask you to make sure that you understand [1:58] the specs of the computer that you're going to use to run this locally. [2:02] If you're working with a higher end [2:03] MacBook, what you're going to want to do is just go to the little Apple icon here, [2:07] go to about this Mac, and you're going to be looking for this number right here, which is the memory. [2:12] Now assuming that you have a mac, that's no more than maybe 5 or 6 years old. [2:15] This is the unified memory that you have available on your entire machine from your CPU, [2:21] your GPU, and effectively what a local model is going to be able to utilize. [2:25] It's not going to be able to utilize all of the memory. [2:27] In this case, I have 32GB, and probably the highest end model I want to run would take up maybe 20GB. [2:33] Again, we'll look at that in a second. [2:34] But what you want to find is, okay if I'm on Mac, how much Ram do I have in my machine? [2:38] Again, assuming that you're running on a newer MacBook, [2:41] if you have a machine that's six, seven, eight years old, it's going to be very difficult [2:45] to run local models, and you're probably going to have a better experience with cloud [2:49] just because they're going to be extremely slow on that lower end hardware. [2:53] If you're on windows, the situation changes a little bit. [2:56] If you're on windows or even a Linux device, you're going to be looking for your GPU [3:00] or your graphics processing unit, and specifically how much Vram is available there. [3:04] So the Ram in your computer doesn't really matter. [3:06] It's more about just the Ram. Again, if you're on windows. [3:09] So if you're running a 4090, for example, you might have 24GB of Ram. [3:13] If you're running maybe an older GPU, you might have eight gigabytes of Ram. [3:17] Now, typically speaking, these local models can utilize almost 100% of that Vram in your GPU. [3:23] So whatever that number is, is kind of the upper bound of the type [3:27] or performance of local model that you're going to want to be able to run. [3:30] Now we're going to get into all of the details here, but you want to understand what operating system [3:34] on my on for 99% of you it's going to be Mac or it's going to be windows. [3:38] If you're on Mac, you're looking at the total Ram of your device. [3:41] If you're on a newer Macs on M-series Mac, and if you're on windows, [3:44] then you're going to be looking at the total amount of Ram that you have available in your graphics [3:48] processing unit, assuming that that's an Nvidia GPU. [3:51] And again, you're ideally going to want the highest end hardware you possibly can have. [3:55] But even if you're running on a relatively budget device, as long as you have Vram [3:59] available in the GPU or again on a newer series Mac, you still can run these models. [4:04] Okay, so now that we have that out of the way, the next thing is model selection and actually [4:08] setting this up on our machine. [4:09] So what we're going to do in order to run local models is we're going to install a tool [4:13] called Ollama. [4:14] Ollama allows you to pull down different models and then just run them natively on your own machine. [4:19] The only limitation for running these models is that hardware that I talked about before, [4:23] it's completely free. [4:23] You don't need to pay for anything and then you can connect this to tools [4:26] like Open Claw, which we're going to do obviously in just a minute. [4:29] However, before we go any further, I want to make you [4:31] aware of a really cool opportunity, especially if you're someone who likes open claw in. [4:35] It's called ClawComp. [4:36] Now Link Ventures is running it, and the short version is they're going to send you [4:39] a free Mac mini to build with Open Claw, where you can run local models. [4:43] You can keep it no matter where you finish. [4:46] No matter what happens, the hardware is yours for free. [4:49] Now here's how it works. [4:50] It's a multi-month build program, so it's designed to fit around whatever you're already working on. [4:54] You apply with a team of 1 to 3 people. [4:57] You hang out in their discord, you talk in their tality group, and if you stand out, you're in. [5:01] And then you get three months to actually build something real using open clock. [5:05] So not a wrapper, not a demo that falls apart when someone pokes [5:08] in an actual automation with measurable outcomes. [5:11] Now the whole thing ends with something called claw week. [5:13] This is from June 15th to June 18th in Cambridge, Massachusetts. [5:17] At Link Studios. [5:18] They fly you out, [5:19] they cover the housing, and you spend four days pushing your project to the finish line in person. [5:23] Now, founders and investors from the link ecosystem are there. [5:26] Researchers from Harvard and MIT are there, and portfolio [5:29] companies are walking around meeting teams and evaluating the pitches. [5:32] Lock now the prize for $17,500 and the application date closes on May 8th. [5:38] So if you're a student building something or you have an open claw project [5:42] that you've been meaning to take seriously, or if you just want an excuse to go deep [5:45] for three months with a real deadline and real people watching, this is super cool. [5:49] It's an awesome opportunity. Again, it's free to apply. [5:52] Link is in the description. [5:53] Go apply, get in the discord stand out and I hope you guys get the free Mac mini. [5:58] Okay, so that said, let's get back to the video here. [5:59] I want to go through the setup process. [6:01] So the first thing we're going to have to do here, regardless if we're on windows, Mac or Linux, [6:05] and if you're running on a virtual private server, then same thing. [6:08] You're going to go over to your terminal [6:10] and you're just going to run this command that you can get from Ollama website. [6:13] I'm going to leave a link to in the description, and it's just Ollam.com. [6:16] You're going to copy the install command. [6:17] And even if you already have Ollama installed, you should run this again to update Ollama, [6:22] because you will need the newest version in order to use the model I'm going to recommend here. [6:26] Okay, so you open a terminal or command prompt. [6:27] You paste this command, you hit enter and it's going to install Ollam for you. [6:31] Now like I said, Ollama is just a tool runs on your computer that lets you run these local models. [6:36] So we're going to wait for this installation to finish. [6:38] Once it's done, I'll be right back and then we're going to pull a local model to our machine. [6:41] And then once we have the local model running, we're going to be able to actually connect it to OpenClaw. [6:46] And good timing. [6:47] It looks like it's already installed. [6:48] While I was just doing that speech. [6:50] Okay, so we have Ollam installed now just to make sure that it's working. [6:54] We're just going to type the Ollama command. [6:55] If for some reason this Ollama command isn't working for you, just close your terminal [6:59] and reopen it or close your command prompt if you're on windows and reopen it, [7:03] and you should just see something popping up like if the command does some output, you're good. [7:08] And then what you can do is hit escape to exit out of that. [7:10] Okay, get out of this interactive window. [7:13] Now if you're on a virtual private server, which is how I typically recommend running OpenClaw, [7:18] you're going to want to make sure that the virtual private server has the same hardware [7:22] that I talked about before, right? [7:24] So in the case of a Linux machine, you're going to want an Nvidia GPU with a ton of Vram. [7:29] If you have that, then you're going to be able to run the local model on the virtual private server. [7:33] So the steps that I'm showing you here [7:35] work on any operating system, whether it's locally on your own computer, [7:38] whether it's on a VPNs, but it requires that you have adequate [7:41] hardware in order to have a decent experience running these local models. [7:44] Okay, so now that we have Ollama installed, [7:46] what we need to do is select the model that we want to download. [7:49] So what I'm going to recommend is that we use one of the newest best local models, which is GEMA. [7:54] For now. There's a bunch of other models that you can choose from. [7:57] At least when I'm filming this video. [7:58] This is the current, smallest and best model that most of you should be able to run. [8:03] But what we're looking for when we select these models [8:05] is that they have the ability to cull tools for OpenCL. [8:09] Specifically, you need this. [8:10] You don't just want a chat based model, you want one that has tools thinking, right? [8:14] All of these different modes like Gemma for has. [8:17] And if you want to see all the different models available, you can just go to open claw.com/search [8:22] or just go to the models tab and you can look and see all of the different models that are here. [8:26] For example Gwen 3.6. [8:28] Also another great model, just a little bit larger than Gemma four. [8:31] So I'm going to recommend that we go with Gemma for for now. [8:34] So once we've selected that Gemma four is what we want, what we're going to run. [8:38] Is this Ollama run or Ollama pull? [8:41] I'm going to show you the command and a second [8:42] Gemma for this is a command that's going to download the model to our computer. [8:46] But before you do that, you need to select the size of the model that you want to download. [8:51] Now you'll notice that if we go to the models down here, we have a bunch of different options. [8:55] We have Gemma for latest, Gemma for 2 billion, Gemma [8:58] for 4 billion, Gemma for 26 billion Gemma for 31 billion. [9:02] And then the cloud one which we're not going to look at right now. [9:05] Now the billion value here is the number of parameters that this model has. [9:09] The larger the amount of parameters, [9:10] the better performance of the model, but also the larger size of the model. [9:14] So really the thing that you want to look at here is the size okay. [9:18] So we see nine gigabytes, seven gigabytes, nine gigabytes 18GB 20GB. [9:22] Now you want to make sure that this size is smaller than the amount of Ram [9:26] that you have in your computer or that you have available from your graphics card. [9:29] Again, based on what I talked about before, depending on your operating system, in my case, [9:33] I have 32GB of Ram, so I'm fine to go all the way up to this 20 GB model, but I wouldn't [9:39] want to go much higher than that because the bigger the model gets, the slower it's going to be. [9:43] And again, you need to make sure that it fits inside of the Ram. [9:47] Now theoretically, you can run any model you want, even if it was one terabyte, [9:51] as long as you had enough hardware space. [9:52] But it's going to be so incredibly slow [9:54] that you're not going to be able to really even get any use out of it. [9:57] So that's why I'm emphasizing that it needs to be small enough [10:00] that it fits into the Ram and gives you a little bit of a buffer. [10:04] So that's what we're looking for. [10:05] Specifically, you want to pick the biggest model that your computer is capable of running. [10:09] Okay. [10:10] So I'm going to go with let's just go with the 31 B. [10:13] Right. Because that's going to work for my machine. [10:14] Even though it might be a little bit slow. [10:16] And what I'm going to do is I'm going to go to my terminal [10:18] and I'm going to type the command Alama pull. [10:21] And then I'm going to paste this Gemma for 31 B. [10:25] Now, for most of you, you're probably going to want to go with the smaller ones. [10:27] You might go with 4 billion right. [10:29] Or E 4 billion E 2 billion, whatever they've put here for this. Right. [10:32] So you might put E to be right. [10:34] Whatever the name is that you see here, you can just directly copy [10:37] what this is going to do is then start downloading the model on your machine. [10:40] It's going to pull all nine, 15, 20GB, whatever. [10:43] And once that's done we're good to actually right now I already have Gemma four on my machine. [10:48] So I'm not going to run this command. [10:49] But if you don't again you need to download it first. [10:51] It might take a few minutes depending on your internet speed. [10:53] Now once it's downloaded, what you can do to see all of the models [10:57] that you have available is to type the command Ollama list. [11:00] When you do that, it's going to give you a list of all of the models you've downloaded. [11:03] You can see that I just downloaded the latest Gemma for one recently, [11:06] and this is the one that I'll end up using with open Clock. [11:09] But you can see all of these. [11:10] You can also just go Ollama help. [11:12] And if you do that, it's going to give you a list of all of the different commands [11:15] where you can create a new model, you can pull models, you can sign in, you can copy a model. [11:19] There's all kinds of advanced stuff. I'm not going to go into all of the details. [11:22] The point is, awesome is super cool, is a lot of stuff that you can do with it. [11:25] Okay, so now that the model is installed, we can just quickly test it so we can type Ollama run. [11:29] And then we're going to go with whatever the name is. [11:31] So my case I'm just going to go Gemma for but you would put whatever it is that you actually [11:35] installed and it should take a second here and then allow you to communicate with the model. [11:39] So I'm just going to go hello world. [11:40] And then you can see immediately I get the response, hello, how can I help you? [11:45] You know I'm good. [11:45] How are you? Right. Whatever I'm doing. Well. [11:47] And then it gives me the response and you can see [11:49] this is actually quite fast because I'm running one of the lower tier models, [11:52] just the nine gigabyte model on my machine, which is capable of running this. [11:56] If you want to get out of this window, you can hit slash exit. [11:59] This is just kind of a terminal based view where you can chat with the model directly [12:02] if you want to do that. [12:03] Okay, so now that we have the model installed we want to start configuring it inside of OpenClaw. [12:08] So first we need to make sure we have OpenClaw installed on our machine. [12:10] And again if you're running this on a virtual private server, [12:13] that virtual private server would need to have the hardware requirements. [12:17] And you would follow the same steps I just did to install Ollama on the virtual private server. [12:22] So what I'm about to show you, you just do wherever you have OpenClaw installed. [12:26] Okay, I'm going to assume that you have it installed somewhere if you're following along with this video. [12:30] Now what you can do right is we go to open I just ran this command. [12:34] So I have OpenCL installed on my machine. [12:36] And literally all we need to do to get Ollama working with this is we can type OpenCL or [12:42] configure. [12:43] Okay. [12:44] So assuming this is installed on our machine right I don't recommend running it on your local machine. [12:48] But for this tutorial I will show you doing it right here. [12:50] We're going to type OpenCL configure. [12:52] Maybe you guys have a mac mini or something. [12:54] So you're going to do that directly on there. [12:56] And what we're going to do is go through and we're going to select model okay. [12:59] So where it says model we're going to use our arrow keys. [13:01] We're going to press enter. [13:02] And we're going to select down here. [13:04] Mine's a little bit laggy when it first pops up. [13:06] But we're going to go down all the way to where it says Ollama. [13:10] You should see it popping up as an option. [13:12] Now once it pops up we're going to go with just local only. [13:15] We don't want to use the cloud one. [13:17] I'm not going to get into that in this video. [13:18] Ollama has a cloud offering, but I don't recommend using it. [13:21] We're going to go local only. [13:22] We're just going to leave the base URL as it is. [13:24] We don't need to change this and continue. [13:27] And we're going to select the model that we want to use with us. [13:30] Now I apologize for the cut here, but I just wanted to mention that if at this point, [13:33] for some reason Ollama is not appearing or showing, I can't reach the server [13:36] and there's some issue, then what you're going to want to do here is quit out of this. [13:40] You can hit Ctrl C and you want to just run the command alarmist serve. [13:44] Now when you run Ollamaserve, this is just going to start the Ollama service on your computer. [13:48] It's going to run directly in this terminal window. [13:50] So just don't close it. And then you should be good to go. [13:52] Now you probably don't need to do that. [13:54] You also can just go to your spotlight search. [13:56] If you're on something like Mac, and you can just run Ollama, you can just literally run the application [14:01] and you should see you get like a little Ollama icon popping up in the top, [14:04] meaning that it's running in the background. [14:06] Okay, so you can see we have a bunch of options here. [14:08] I'm just going to check this “ollama/gemma4:latest” as well as Java four. [14:12] You can select all these different models. You can see I have a bunch of them right. [14:15] Because these are all the ones that were installed on my machine. [14:18] And I'm just going to go ahead and press on confirm okay. [14:21] So you just select the model that we just installed or whatever you picked. [14:24] For most of you it's going to be that gem of four. [14:26] Okay. So we're going to go ahead and press enter. [14:28] And now both those models or whatever ones we've selected are going to be enabled in OpenCL. [14:32] So I'm going to press on continue. [14:34] And then what we're going to do [14:35] is just restart the gateways that this model will now be available to open class. [14:39] We're gonna go open core and then gateway restart okay. [14:44] Now when we run that command it's going to restart the gateway. [14:47] Both these models should then be available. [14:49] And then what we can do is just start using them directly inside of OpenClaw. [14:52] Now, the same thing goes for windows. [14:54] You can run the Ollama serve command, or you can just go down to your windows [14:57] like application search bar or whatever you want to call it. [14:59] And then you can just run Ollama and it should run automatically in the background for you. [15:03] Okay. So you want to make sure that's running as a background process. [15:06] Now once we do that we can go back to our open cloud control. [15:08] Again I'm going to assume you know how to get here because you already probably set up open [15:12] cloud before. [15:13] And what we can do is just asking the question so I can say something like, hello, how are you doing? [15:17] Can you tell me the meaning of life? [15:18] Right. [15:19] And you'll notice that by default, if you haven't selected any model, it should just be [15:22] using GEMA four. [15:23] And if we go enter here it should just be using this model. [15:27] And you can see we get the response very quickly. [15:29] Now if for some reason you have different models connected, you [15:31] probably want to set the default model to be GEMA for. [15:35] In order to set the default model, you can typically do that from the config. [15:39] I don't know exactly where the configuration is for the default model, but I believe in the open [15:43] cloud dot json file, there is a way to set the default model to just use this local model here. [15:49] And then you're kind of good to go. [15:51] Now, what I typically will do [15:53] when I'm using this is I might have a local model that I use most of the time. [15:57] And then I also configure a cloud model if I want to do something that's a little bit [16:01] more challenging, that requires some higher intelligence, [16:04] because while these models are good, they are going to be stupider than like an opus 4.6 [16:08] or opus 4.7, so you probably want to use them in combination, [16:12] unless you purely just care about privacy and you just want everything running 100% locally. [16:18] So my suggestion is you configure like an open AI model and a traffic model, [16:22] and then the local model, you set the local model as your default, [16:25] and then you switch to the other model when you want to use that. [16:28] Now what you can do is you can actually type slash model here. [16:31] Right. And then you can change the model that you want to use. [16:33] So I just do slash models. [16:34] It will give me a list. You can see that right now we just have two. [16:37] And then I can say slash model. [16:38] And then you know Gemma for right. [16:41] And then switch over to actually use that model. [16:43] In my case I'm already using it. So there is no switch. [16:45] But you also can just directly tell it, hey, I want to switch over to use Gemma for it. [16:50] Can you switch the model and then open? [16:52] Close should be capable of just running the command and automatically switching the model four. [16:55] You can see it called the tool and then it's going to switch it out. [16:58] Now in this case it's saying you can't do it again because we only have that one model. [17:01] By the way, [17:01] if you are wondering what I'm using to dictate here, I'm using a really cool tool called Whisper Flow. [17:06] I can quickly click into it and kind of show you what it looks like here [17:09] if we go to the settings, but effectively it's the best kind of voice dictation. [17:14] It uses AI in the background extremely fast, actually have a long term partnership with them. [17:18] They're free to try out and I just use them anyways, which is why we have the partnership. [17:21] You could see my words per minute here, [17:23] and I'll leave a link to in the description in case you guys want to try it, [17:25] but especially if you're working with a lot of these AI tools, just much faster to speak [17:30] rather than to type, right? [17:31] So if I want to say something, hey, I want you to go do this. [17:33] Here's a bullet pointed list of Abcde. [17:36] Right? [17:36] And then, you know, we go and it just gives me whatever and automatically formats [17:39] it, fix the spelling, punctuation, all of that kind of stuff. [17:43] So anyways, that's pretty much all that I have for you guys in this video. [17:46] Setting it up with OpenClaw is very easy. [17:48] It's a matter of having Ollama installed and then just configuring them all. [17:51] If you want to go further, you can set up inside of like your sold on md file [17:55] as well as your agent start MD file and all of the configuration and open floor, [17:59] and you can tell it when to use the local model versus when to use the cloud model. [18:03] You can install multiple local models. Right. [18:06] So maybe you want Gemma for for something. [18:08] Maybe you want Gwen for something else. [18:09] Maybe there's a specific image or video model you want and you install that. [18:13] You can go crazy with the configuration. [18:15] The important thing is really understanding that hardware constraint. [18:18] And again, the performance and speed that you're going to get [18:21] is going to be dictated by that hardware. [18:23] In combination with that model selection. [18:25] Probably you just want Gemma, but you can mess around with some other ones [18:28] and see what experience you get. [18:30] In my case, Gemma is much, much, much faster than a lot of the other kind of best local models [18:35] that are currently out there. [18:36] So that's it guys. I'm gonna wrap up the video here. [18:38] If you enjoyed, make sure they like subscribe and I will see you in the next one.