[0:00] in this short video I'll teach you [0:01] everything you need to know to get up [0:03] and running with AMA which is a [0:05] fantastic free open-source tool that [0:08] allows you to manage and run llms [0:10] locally rather than having to pay for [0:12] Chad GPT or use these hosted services [0:15] online you can actually run all of these [0:17] models locally on your own computer so [0:19] you get privacy security and best of all [0:22] they are completely free so with that in [0:24] mind let me show you how to set this up [0:26] get it running and I'll also explain to [0:28] you how you can utilize this through [0:29] code because olama provides an HTTP [0:32] server which means you can call your [0:34] models from really any type of [0:35] application so first things first we do [0:37] need to install AMA so to do that you [0:39] can go to the website which is ama.com I [0:42] will link it in the description and you [0:44] can simply press on download and then [0:46] select your operating system in my case [0:48] I'm using Windows but of course you have [0:50] the command for Linux and then the [0:51] installation for Mac now once that's [0:53] downloaded simply double click it and [0:55] install it and then I'll show you the [0:56] next step so once you've installed AMA [0:59] there's a few different ways to run it [1:00] first of all you can just open the [1:02] desktop application so if you're on [1:03] Windows you can go here to the search [1:05] bar and just search for ama if you're on [1:07] Mac you can simply go to the spotlight [1:09] search same for Linux and just run the [1:11] application now when you do that [1:12] nothing's going to appear on your screen [1:15] and the reason for that is this just [1:16] starts a backend server that's running [1:18] the AMA service now the other way to do [1:20] this is to open up a command prompt or a [1:22] terminal so you can see I'm in command [1:24] prompt here on Windows and then to [1:26] simply just type O llama if you do this [1:28] you should get some kind of out put and [1:30] if you see that it means that you've [1:31] installed o llama correctly at this [1:33] point I'll assume that you have this [1:34] installed correctly and that this [1:36] command gave you some kind of output and [1:38] now what we can do is start running [1:40] models so the first thing to look at is [1:42] the different models that we have access [1:43] to with olama now truthfully you have [1:46] access to pretty much any open-source [1:47] model you want and you can even write [1:49] some custom configurations to use your [1:51] own models or things you pull in from [1:53] something like hugging face now if you [1:55] go to the AMA GitHub repository which I [1:57] will link down below you can see some of [1:59] the common mods mod that you may want to [2:00] download now keep in mind that since you [2:02] are running these locally you will need [2:04] to download the entire model and you'll [2:06] need to have enough space on your [2:07] computer you can see some of these are [2:09] 43 GB for example and enough RAM to run [2:12] and load the model depending on how [2:14] large it is so you see if we look [2:15] through here it defines the number of [2:17] parameters for these models if we look [2:19] at something like llama 3.1 we have 231 [2:22] GB and 45 billion parameters and if you [2:26] go down here to this note it specifies [2:28] how many gabt of ram you should have [2:30] based on the different model parameters [2:31] so even on my computer which has 64 GB [2:34] of RAM it would be difficult to load the [2:36] new llama 3.1 model with the 405 billion [2:39] parameters so keep that in mind when you [2:41] are choosing the models that you want to [2:42] use for now I'm just going to go with [2:44] the standard llama 2 model because this [2:47] is older and it's not as large and I [2:49] know that I can run it and most of you [2:50] should be able to run it as well so I'm [2:52] going to show you how we can pull that [2:54] but if you're looking for a list of all [2:56] of the models you have available you can [2:57] go to the oama library so if you go to [3:00] ama.com library and you can scroll [3:02] through here and you'll see there are [3:03] hundreds of different models you can [3:05] sort them you can filter and you can [3:07] find even multimodal models except [3:09] things like video photos voice Etc so [3:12] once you've decided on a Model that [3:14] you'd like to run it's very simple to do [3:16] so all you need to do is type a llama [3:18] run and then the identifier of that [3:20] model now in my case I just want to run [3:22] llama 2 I know this is an outdated model [3:24] I'm just doing it because it's smaller [3:26] so I can simply type O llama run llama 2 [3:29] and if this model is not already [3:31] installed on my system then it will [3:33] download it and install it for me if it [3:35] is already installed it's just going to [3:37] bring up a prompt where it allows me to [3:39] actually start typing to the model and [3:41] messaging with it so notice here that [3:43] it's just loading and it kind of gives [3:44] me these three arrows and I can just [3:46] start typing something to the model and [3:47] get some kind of response and you can [3:49] see it's pretty much instant because [3:50] there's no latency it's running on my [3:52] own machine now again if this wasn't [3:54] already installed it would start pulling [3:56] the model for you and then you would [3:58] have to wait for it to finish it would [3:59] install then you can run the model and [4:01] you can start using it now after some [4:03] experimentation it's told me that you [4:05] can type slash bu to get out of this so [4:07] if I type slby you can see that it will [4:09] enclose this window and then if we want [4:11] we can type amaama and then list and we [4:14] can list the different models that we [4:15] have available on our system in this [4:16] case you can see I have llama 2 which is [4:19] the latest version if I had any other [4:20] models they would show up here so that's [4:22] the basics on running models using oama [4:24] but there's a lot more to show you so [4:26] make sure you stick around after a quick [4:28] word from our sponsor today's video is [4:31] sponsored by SEO writing a tool that's [4:33] transforming content creation across [4:35] different niches and industries their [4:37] new brand voice feature lets you [4:39] generate content that matches your [4:40] Unique Style whether you're writing [4:42] tutorials reviews or even industry [4:44] analysis one click generates a complete [4:47] blog post with AI generated images and [4:49] relevant videos embedded automatically [4:52] potentially saving you hours of manual [4:54] work what sets SEO writing apart is [4:56] their deep web research with built-in [4:58] citations when you need accurate [5:00] up-to-date information the platform [5:02] pulls from reliable sources and adds [5:04] citations automatically their humanized [5:06] text feature helps your AI generated [5:08] content stand out while their external [5:10] linking feature intelligently connects [5:12] to relevant resources and for all you [5:14] WordPress users out there there's a [5:16] gamechanging feature that lets you [5:18] connect your site and autopost content [5:20] directly this feature allows for [5:22] consistent scheduling while focusing on [5:24] other projects now if you're ready to [5:26] try it for yourself then use my code TW [5:28] wt20 5 for a 25% discount click the link [5:32] in the description and see how SEO [5:34] writing can fit into your content [5:36] strategy all right so we are continuing [5:37] here and I want to show you what happens [5:39] when you pull multiple models so again [5:41] if we go back to the library we can [5:43] start looking through different models [5:44] that we may want to utilize maybe I can [5:46] even just go back here to the GitHub if [5:47] I want to find them a little bit easier [5:49] and maybe I want to use the mistal model [5:51] as well if that's the case I can just [5:53] copy this command or the name mistl I [5:55] can go back here I can simply run the [5:57] command AMA run mistl it will then pull [5:59] that manifest for me pull the model once [6:02] that's finished I'll be able to use this [6:03] and I'll show you how so looks like this [6:05] has been downloaded and now I can start [6:06] using the model if I want I can exit out [6:09] of this and if I want to switch between [6:10] the two different models again I just [6:12] type O llama run and then I can specify [6:14] the model that I want to use so if I [6:15] want to go back to llama 2 I use llama 2 [6:18] if I want to go back to mistl I just [6:21] type mistl and now I can start using [6:23] Mistral so you can have as many models [6:25] as you want and again you can list them [6:27] by typing ol llama list and if you want [6:29] all of the commands you can use simply [6:31] type AMA and then it will show you which [6:33] ones you have access to there's a lot of [6:34] them for example you can also remove a [6:36] model if you want to do that copy a [6:38] model there's also customizations you [6:39] can make to them which I'll show you in [6:41] just one second all right so all of that [6:43] is great but we probably want to know [6:44] how to utilize these models from [6:46] something like code from our [6:47] applications sure they're great to use [6:49] here in the terminal but a lot of times [6:50] you want to integrate them with some [6:52] kind of software especially if you're a [6:53] programmer and you watch this channel so [6:55] the interesting thing about olama is [6:57] that it actually exposes an http API on [7:00] Local Host that means that anything we [7:03] just did here with commands we can [7:04] actually trigger through the API so we [7:06] can send request to this from something [7:08] like curl Postman something like python [7:11] code really any code at all that can [7:12] send some type of HTTP request Now by [7:15] default if you're running aama you [7:17] should be able to see this if you're on [7:18] Windows in kind of like the I don't know [7:20] what you would call this Services bar [7:22] wherever it's showing the running [7:23] applications and you can see I have this [7:25] little AMA logo now when olama is [7:27] running as the desktop application by [7:30] default that Port is going to be open so [7:32] you'll be able to access the HTTP API [7:34] but if for some reason this isn't [7:36] running so for example if I quit this [7:38] what I can do to trigger that to run is [7:40] I can simply type AMA serve in my [7:43] terminal if I do this it's now going to [7:45] start running the HTTP API in this [7:47] terminal instance and now I'll have [7:50] access to it and here it will also show [7:51] us what port it's running on although it [7:53] should be standard and you can see if we [7:54] look through here it gives us the exact [7:56] Port so it's on [7:58] 11,434 so if you wanted to you can copy [8:01] that and save it for later so that we [8:02] can use it in our code regardless now [8:05] that the olama serve or the olama HTTP [8:07] API is running we're able to call it and [8:10] again just to clarify if you're running [8:12] this as the desktop application it will [8:14] already be running in the background but [8:15] if for some reason you want to manually [8:17] invoke this to run then you can run the [8:19] command ol llama serve where it will [8:21] give you all of this output and you'll [8:22] be able to view all of the requests to [8:24] the HTTP server so now that the server [8:26] is running we can use something like the [8:28] following python code here to send a [8:30] request to it now this is done manually [8:32] very intentionally I'm going to show you [8:34] an easier way to do this in 1 second but [8:36] it's just to illustrate that you do have [8:37] kind of complete control over this if [8:39] you want so you can see here in Python [8:41] I'm using the requests and the Json [8:43] module now just by the way if you want [8:45] this to work on your machine you will [8:46] need to install the request module so [8:48] you can say pip install request or pip [8:51] three install requests and I'm going to [8:52] leave this code in the description uh [8:54] Linked In A GitHub repo in case you want [8:56] to check it out now what we do is we [8:58] Define our base URL this is the URL of [9:01] the server and then [9:02] /i/ chat there's a lot of other [9:04] endpoints that you can use here and you [9:06] can even control deleting models adding [9:08] models Etc but in this case we just want [9:10] to chat with one of our models then we [9:12] can define a payload this is the model [9:14] that we want to chat with so in this [9:15] case I've gone with mistl and then we [9:17] can Define different messages here's a [9:19] standard message with the role of a user [9:22] next we can send a post request here [9:23] using request. poost to our URL with our [9:27] Chason payload which is this right here [9:29] and enable the streaming mode which [9:30] allows us to grab all of the responses [9:33] as they are typed this way we can grab [9:35] them in real time and we can show the [9:37] model actually typing the response [9:39] rather than waiting for the entire [9:41] response to be generated and then [9:43] viewing it now here's a little bit of [9:44] code just to handle that streaming data [9:46] for us so we're going through all of the [9:48] lines that are returned from this [9:50] response and then we are simply kind of [9:52] printing them out okay so I'm going to [9:54] show you what happens when I run this so [9:55] we already have requests installed and [9:58] if I go python sample request dopy just [10:01] wait one second here it will stream in [10:03] all of the data and then print it out so [10:05] you can see that it's kind of printing [10:06] it out line by line for us here as it [10:08] gets it and there you go python is a [10:09] high Lev language blah blah blah gives [10:11] us the answer if we go back to the API [10:13] we can see that the request was sent [10:15] here it took 4.1 seconds to process and [10:18] it returned to us that data sweet so [10:20] there you go that is how you utilize the [10:22] API manually but a lot of you probably [10:24] don't want to write all of this code so [10:26] instead we can use a very simple module [10:28] from python called you guessed it ol [10:31] llama so if you're using python or [10:33] JavaScript there are packages that will [10:34] do this for you so you can simply pip [10:37] install olama or pip three install olama [10:41] in your systems that you have this [10:43] module and now you have access to the ol [10:45] module you can simply create a client [10:47] you can Define your model you can Define [10:49] some kind of prompt and then you can use [10:51] client. generate specify the model and [10:54] the prompt and then you can grab the [10:56] response okay so I'm going to quickly [10:58] show this to you I can run this code [11:00] with python package. piy and you will [11:05] see here in just one second that we [11:06] should be able to get the response okay [11:09] and there you go we get the response and [11:10] it gives us the answer so that is how [11:12] you use the HTTP API now I'm going to [11:15] show you how you can do some [11:16] customizations to the models in ama so [11:19] moving on I'll show you a quick [11:20] customization that you can make to any [11:22] of the models that you can pull with AMA [11:24] so you can see on the right hand side of [11:26] my screen that I've created something [11:27] called a model file now I've just put [11:29] this in a directory that's on my desktop [11:31] you need to put the file in a location [11:33] that you know and that you're able to [11:34] access from your terminal and for the [11:36] model file I've used this very simple [11:38] syntax that I just took directly from [11:40] the AMA website all you do is you [11:42] specify from and then you have some kind [11:44] of base model so in this case we're [11:45] using llama 3.2 but you can use any [11:48] model that you want that's available [11:49] with a llama you can do something like [11:51] set the temperature of the model you [11:53] don't need to do this but there's some [11:54] other parameters you can set as well and [11:56] then you're able to pass something like [11:58] a system message which is essentially [11:59] kind of instructing the model what it's [12:01] supposed to be doing and how it should [12:03] handle the upcoming messages so in this [12:05] case they've just written Ur Mario from [12:07] Super Mario Bros answer as Mario the [12:09] assistant only okay so we have this [12:12] model file written notice I don't have [12:13] any extension it's literally just called [12:15] Model file no. txt or anything and what [12:18] I've done is I've put my terminal in the [12:20] same directory where this file exists [12:23] now what I'm able to do is create a new [12:25] model based on this model file in olama [12:29] and have one that's set up as Mario so [12:32] to do that I can type AMA create I can [12:35] give this a name in this case I'll call [12:37] it Mario you can call it anything that [12:38] you want and then I'm going to specify [12:40] DF which stands for file and then the [12:42] location of my model file now in this [12:45] case it's just simply at@ slm model file [12:49] okay so I'm going to go ahead and create [12:50] this and you'll see that it says success [12:52] that's because I've already pulled model [12:54] llama 3.2 so now if I want to utilize [12:57] this customized model what I can do is [12:59] type a llama run and then the name of [13:01] the model which is Mario and now if I [13:04] say hello you'll see that it says it's a [13:06] me Mario and it kind of you know [13:08] simulates like how Mario would reply so [13:11] if you want to set up some custom models [13:12] where they have some system prompts they [13:15] have some different parameters set up [13:16] with them you want to tweak them somehow [13:18] you can do that using these model files [13:20] then you can simply create them in olama [13:22] now let's say you're done with this one [13:24] and you want to remove it you can say RM [13:26] or sorry AMA RM and then what is it the [13:29] name of this one Mario and it will [13:30] remove that so now if we type oama list [13:33] you no longer see it and also it's worth [13:35] noting that these uh models like Mario [13:38] you can utilize them from code so in my [13:40] python code here I can just specify [13:42] Mario once that's created and then I can [13:44] use that anyways guys that is it that's [13:46] all I wanted to show you I hope you [13:47] found this valuable if you did make sure [13:49] to leave a like subscribe to the channel [13:51] and I will see you in the next one [13:54] [Music]