One Interface for All AI Models?!
40sThe promise of using Claude, Gemini, and Grok from one self-hosted interface with unlimited usage is a game-changer that sparks curiosity.
▶ Play Clip"Delivers a solid step-by-step guide with real value, though some fluff and sponsor mentions pad the runtime."
The video demonstrates how to set up a self-hosted AI interface using Open WebUI and LiteLLM, providing a unified platform to access multiple AI models like Claude, Gemini, and Grok. The creator walks through cloud deployment, API key integration, and user management, emphasizing control, cost savings, and centralized access.
Open WebUI is an open-source, self-hosted web interface for AI that allows access to various LLMs like GPT, Claude, and self-hosted models like Llama. It offers features like user management and model restrictions.
Two hosting options: cloud (VPS) or on-premises (laptop, NAS). The video focuses on cloud setup using a VPS from hostingyour.com, with KVM2 plan at $6/month, which is cheaper than ChatGPT Plus and allows hosting multiple services.
Steps to create a VPS: choose application (Open WebUI), select Ubuntu, confirm, and set password. The VPS automatically installs Open WebUI and Llama. Access via public IP on port 8080.
Create admin account, then chat with default local model (Llama 3.2). It's slower but uses server resources. For better models, need API keys.
Normie mode: use ChatGPT directly. API mode: pay-as-you-go, access to all models including new ones like GPT-4.5, and can save money for light users.
Create OpenAI account, add credit card (initial $5), create API key, and paste into Open WebUI settings to unlock all ChatGPT models.
AI models charge per token. Example: o3-mini costs $1.10 per million tokens, o1 costs $15, GPT-4.5 is most expensive. Casual user with 50 conversations/month might spend ~$5.
Open WebUI only supports OpenAI and Ollama connections. LiteLLM acts as a proxy to connect to 100+ AI providers (Anthropic, Gemini, Grok, etc.) via OpenAI-compatible API.
Deploy LiteLLM via Docker on the same VPS. Clone repo, edit .env with master key and salt key, then run docker-compose up -d.
Create API keys for OpenAI, Anthropic, Grok, etc. In LiteLLM admin panel, add models (e.g., Claude 3.7, 2.1) and create virtual keys with budgets.
In Open WebUI, replace OpenAI API key with LiteLLM key and set base URL to http://localhost:4000. Then models appear in the dropdown.
Add multiple models (Claude, Grok, GPT-4o, Llama) and compare responses. LiteLLM usage page shows token usage and costs.
Create users (e.g., kids), assign groups, restrict model access, set system prompts (e.g., school helper), and monitor usage. Can set budgets per user.
Open WebUI with LiteLLM provides a powerful, self-hosted AI hub that centralizes access to multiple models, offers granular control over users and budgets, and can be more cost-effective than individual subscriptions.
What is Open WebUI?
An open-source, self-hosted web interface for AI that allows access to various LLMs.
01:13
What are the two hosting options for Open WebUI?
Cloud (VPS) or on-premises (laptop, NAS).
01:57
What is the cost of the KVM2 VPS plan mentioned?
$6 per month.
02:40
What is the difference between normie mode and API mode for accessing AI?
Normie mode uses the standard ChatGPT interface; API mode uses pay-as-you-go access to all models.
06:19
How does OpenAI charge for API usage?
By tokens, with costs varying by model (e.g., o3-mini $1.10 per million tokens, GPT-4.5 more expensive).
09:26
What is LiteLLM?
A proxy server that connects to 100+ AI providers via an OpenAI-compatible API.
13:37
What are the two environment variables needed for LiteLLM?
LITELLM_MASTER_KEY and LITELLM_SALT_KEY.
15:42
What is the default port for LiteLLM?
4000.
19:22
How can you restrict a user's access to specific models in Open WebUI?
By creating user groups and assigning model permissions.
21:37
API vs Normie Mode
Explains the cost and access advantages of using APIs over standard subscriptions.
06:19Token-based pricing
Clarifies how AI models charge per token, helping users estimate costs.
09:26LiteLLM as a universal gateway
Introduces a tool that solves the multi-provider integration problem.
13:37User management and control
Demonstrates how to manage family or team access with budgets and restrictions.
21:22[00:00] Claude Gemini Grok from one self-hosted interface, and no, I have unlimited usage and I get access to the newest models as soon as they
[00:14] I can create accounts for my employees, for my wife, for my kids, and they can access all the new stuff, but the best part is that I have control, I don't want them accessing every AI model so I can restrict that.
[00:29] what they get help with so they're not cheating on their homework and letting out some network check secrets and I can see all their checks, have AI and you should let them have AI hot take,
[00:43] Right now it's not going anywhere, but seriously, I love the solution. the amount of features it has, I'm addicted. This might be the better way to use ai. This is open Web ui.
[00:58] but this video is going to come at it in a bit different way. get your coffee ready. I'm going to have you set up in about five minutes. Let's go. Okay. Open Web ui. It's an open source,
[01:13] self-hosted web interface for AI and it allows you to use whatever LLM or large GPT and Claude, which by the way, You'll see it's really awesome, but it's not just those.
[01:26] We can run self-hosted models of the alama talking like Llama three and Myre and two, three, sometimes four. One of my favorite features. Speaking of features, it's addicting,
[01:41] so don't worry, but I will say this asterisk, this isn't for everyone. There's one asterisk, one thing that might scare you away, you'll see, Now what do we need to get this set up? As I mentioned, this is self-hosted,
[01:57] which means you yourself are going to host this somewhere. and for that you really have two options. Either the cloud, host it in your house. This could be on your laptop, on a nas,
[02:11] Whichever you choose is going to be quick and easy and you're going to be like, We'll start with the cloud. Don't blink. It's going to be fast.
[02:23] we'll be setting up what's called a VPS or a virtual private server and the So real quick in the description, I have a link hosting your.com/network. and KVM two is my favorite option because you're essentially getting yourself a
[02:40] I know what you're probably thinking six bucks a month. Why Don? First it's cheaper than chat GBT. Second, which is just cool. And third, you can host more than just open web ui.
[02:58] I'm telling you a Healthy Home lab. Just wanted to add that context. Anyways, CPU eight gigs of RAM and VME storage. backups and snapshots because we break stuff and you're going to need this.
[03:16] So just know not only will this puppy run open web UI just fine, you'll be able to add more stuff to it. More projects, resume building moments. so I feel like I need to do this. Sorry, I couldn't do it.
[03:31] That's embarrassing. Don't do some coffee while I deal with that real quick. love that show. That show still hits. Anyways, let's keep building this. it'll ask you to make an account. Choose your term.
[03:47] Check this out. Coupon code, right over here on the right, type in network. Chuck 10, apply that sucker. It's now cheaper. Now pick where you want it to be. I think it'll automatically tell you based on the latency to you.
[04:01] And then we'll choose our OS. Now for us, because we want to do open Web ui, We'll actually click on application right here and we'll click on show more. Ah, there it is. He's hiding from me.
[04:15] We're going to go ahead and select this because not only will it install Llama, it will also install Open Web UI like that and it's going to be on Ubuntu
[04:28] let's go ahead and click on confirm. Continue. Actually, I lied. Enter all your info free malware scanner. Sure. Okay, click continue.
[04:40] This will be the password that you'll use to log into your VPS. Click continue and I think we're almost done. Yeah, finish that up right here. Go and it's setting it up right now you have a virtual private server being spun
[04:52] up in the cloud and they're installing OpenWeb UI along with Llama and all you go watch this video right here. I'll walk you through it. Just pause me. We'll click on VPS management page, go look at it and here is mine.
[05:07] Go ahead and click on Manage over here on the right and right now open Web UI is What that will do is launch another tab. Go and click on that, essentially your public IP address on port 80 80. And here we are.
[05:21] This is your open web ui, unlock mysteries wherever you are. go and click on Get started at the bottom there and here we'll create our first So you have Godlike Powers over everything.
[05:35] Click on create admin account and celebration. We're here. Okay, let's go. Now, you'll see by default we've got a nice little AI model to play with. GBT Lama 3.2 is a local model.
[05:51] It'll use your servers resources instead of open ais. Let's talk to it. Hey, Same kind of familiar interface except as you might see it's slower.
[06:03] smart model. It's very small, really be able to run bigger smarter models unless you have some killer but we don't really care about that right now because we're not done yet.
[06:19] Now when you want to access AI models like chat, GBT or Claude, you usually have two options. Option one, normie mode, you go out to chat GBT, If you want to access to all the new stuff and that's it, you're done.
[06:33] But then option two is where things get interesting. APIs application programming interfaces are what developers use to integrate AI like Chad CT into their apps and programs. So what, we're not writing an app,
[06:46] why do we care? Well, it comes down to how they pay for that access. APIs you pay as you go or you pay for what you use. providers normally give API access to all of their models,
[07:04] GT 4.5 that just came out and the people who have access to that on the normie sorry you're out of luck, but if you're using an API,
[07:17] The second cool thing is that you may end up saving money, but if people on your team or in your house aren't really heavy users of ai,
[07:29] be using 50 cents a month. Okay, so what does that look like? Well, Let's go out to open AI and instead of going to chat gbt, Whatever you got to do, once you're in,
[07:44] it's going to ask you for a credit card, You're only going to be charged for what you use initially. That's five bucks that will sit there until you use it.
[07:57] I'll top it off with five bucks and then I'll go and create what's called an API This will actually unlock all these chat GBT models for us on the open web UI just click that little gear there.
[08:11] secret key. Name it, put it in the default project, leave everything else as is and click on create secret key. There's your key. Let's go put it inside OpenWeb UI right now here in OpenWeb ui we're going to go
[08:25] From here, we'll click on settings and then connections. additional LLMs for open web ui and right there there's a blank space baby
[08:37] Well we're going to paste our API key right there and click on save Now, Let's click on the little menu thing on the left to open that up, And at the top there will change our model from llama to whatever we stink in
[08:56] We have access to everything including that new 4.5 model. Let's start chatting with it. you're using a $200 a month model for nothing. Well,
[09:11] not for nothing we're about to see. Don't get crazy yet. Lemme cover this part. We got to talk about how we pay for these AI interactions and this is the So when you're talking to an AI model specifically an LLM,
[09:26] The way they charge us is by tokens. It's like Chuck E Cheese just without crappy pizza and a scary mouse. Now what's a token? A token is a word in some cases. So for example,
[09:38] that's probably going to be one token or how more complex words might be broken How many tokens was your last response?
[09:52] Break that up so I can see which words were tokens and which were broken up. It's its own token. What a ripoff. If you want to save money with ai,
[10:08] I still don't understand how much money we're being charged. How much you're charged will depend on which model you're using. questions. And that's on display right here. For the oh three mini model,
[10:25] it's going to cost you a dollar and 10 cents per 1 million tokens. the oh one reasoning model will cost you $15 per million tokens.
[10:40] the 4.5 is their most expensive model, and that's just input Notice they do have an output section too. charge them for talking to me and then when I give out my wisdom,
[10:57] Now I know it's kind of hard to break down what does a million tokens mean? Here's your warning, right? So a casual user, let's say they have 50 conversations a month, about a thousand tokens each.
[11:11] Now if you use AI like me, that's very low usage, but some people are like that. A moderate user might have 200 conversations a month and this could be anywhere these are all very rough estimates. This can be sky's a limit, right?
[11:28] 20 bucks to infinity. So hey, draw the infinity. Simple think I'm nailing it. it would not be 20 bucks a month. It'd be a lot more. What impacts that? Well, what models you choose? I talk to the best models a lot. 4.50 yeah,
[11:43] oh 1 0 3 talking all day and my conversations are long and that does impact how the context of our messages are being sent each time I
[11:56] say something to the API so that it knows what I'm talking about. conversation and sometimes I sit there and talk for a while with an AI to figure this is very specific to OpenAI.
[12:12] They do have cashed input which will help offset a lot of those costs. I think it's like 24 hours by default, they may change that. Can this save you money? Maybe,
[12:27] it's more about I want to give my family myself and my employees access to all the ai and I don't want to pay for 15 million plans and have to manage all
[12:39] one place to go and I want control. Now if you're worried about this, It's so cool. You can put a budget in per person so they don't go over like you're stuck at 20
[12:53] You're talking to Alama for the rest of the day. Why is Alex's work so crappy after three? I don't know. Let's break this down. Let's keep going.
[13:07] of a big problem with open web ui. Check this out. If I go back to my settings where I added the open AI API key and my connections, I really only have options for two types of connections. Open ai,
[13:22] What about Gemini? What about all these fun ones? I want to try, that's kind of a problem because you can't just plug in Claude right here or Anthropic. It won't happen. This is where a tool I fell in love with comes in.
[13:37] LM is a proxy for AI or a gateway. If we go to the webpage real quick, they connect to so many ais. I think they say a hundred plus, right?
[13:49] All it's going to have to connect to is light lm and it does that just fine because it has an open AI compatible API. It does great. And then with light LLM, we connect everything else. Open AI andro,
[14:05] which is Claude Gemini, grok Deep seek and no, not the one hosted in China. You can actually access an American hosted Deepsea on another service called Now Light LM will be a proxy server that will install alongside open web ui.
[14:21] Get your coffee. Let's install light lm. So real quick, followed along with me on the hosting your side,
[14:35] setting up A VPS right here in our portal where we're managing our VPS, There's a button right here, B browser terminal. Go ahead and click on that. We'll deploy it via Docker,
[14:50] very similar to how we up open web UI on the other tutorial you watched earlier. but the first thing we'll do is use GI to clone the LM proxy server GI
[15:02] Get clone and then the address light LLM. Ready, set, clone. into here in a moment. Little coffee break
[15:17] Type in CD and light LLM to jump into that folder we're in. First thing we'll do is use nano type in Nano the best text editor ever and
[15:30] we'll edit the file, the hidden file env, just like that. And we're going to add two lines of config. First we'll type in LLM, all caps master key that'll have that.
[15:42] Equal quotes, double quotes, SK dash something. Actually I'll just use Dashlane to do that for me right now. I'll just do digits and letters. We'll do 10 of them
[15:58] This will be your password to log into the server Once we build it, I just clicked out of my browser terminal. Good thing I copy my password. Hit enter and we'll add one more line of config.
[16:10] We'll add the lights LM salt key just like this and have that equal the same kind of starting point, SK dash and then a randomly generated string of characters.
[16:23] This will be used to encrypt and decrypt your L-M-A-P-I key credentials. shall copy all of this real quick. Put that somewhere safe, then hit control X, And for most scenarios all we have to do is type in docker dash compose up
[16:39] dash D, ready, set, go. And this is literally building our server. We don't have to worry about anything else except making sure we sip some coffee while it's happening. Now, while that's installing,
[16:51] let's get our API keys. Ready? First we need our open AI API Key. So I'll create a new one called this light LM default project. Create it,
[17:03] copy it, get it ready. And the same process you can repeat for anthropic, I'm just going to do anthropic for now and I'll grab Grok too. which is actually pretty amazing unfortunately I don't think the grok three is
[17:20] available on API just yet. But I'll go and create a key and it's done. because everything is running through Docker, Now what we'll do is open up a new tab.
[17:33] There it is. Grab that IP address and in your address bar, go out to that IP address port. I think it's 8,000, what was it? Oh, LM admin panel on ui, click on that.
[17:50] environment variable, the SK one N or N. All we care about right now is doing a few things. First,
[18:02] let's go to models on the left here and then right here and the top menu, Let's start with Claude. So I want to click on Anthropic and we could either choose all models,
[18:14] like just go crazy, select them all or be very specific. So maybe I only want the three seven latest and 2.1 to compare how dumb Then I'll add my API key here and add the model just like this at the bottom
[18:27] right, clicking on all models, you can see it sitting right there. 2.1, 3.7. We'll go to the top left and click on virtual keys. We're going to create our own virtual API keys that can control so many things.
[18:41] we don't need a team or anything. We'll name the key. I don't know kids. access are three, seven and two. One checking out optional settings.
[18:54] 20 bucks and this will be a monthly budget and you can do a lot of, which we're not going to cover right now, We'll copy it and now we'll add it to open web ui. So here we're in open web ui,
[19:09] I want to delete my open AI API key. Delete. I'm going to add. L-L-M-A-P-I Key under the open AI API key. I feel like I've been saying open ai,
[19:22] API so much the base URL will be htt, P colon, wack, wack, local host port 4,000. So colon 4,000. And then we'll put our API key right in here, just like that and click on,
[19:34] And that's because they're on the same server. Local host is right there. there it's new chat. Claude sitting right there. Oh, that's so cool.
[19:48] How you doing Claude? Ah, love it. Check this out. but click on add model. We can put Claude 2.1 there as well.
[20:00] let's add them side by side and say tell me a riddle and they'll answer it I'm going to add open AI and grok.
[20:18] Now I added these models and now I have Aroc and four oh and oh three many. But no one inside of Open Web UI will have access unless I give it access to those virtual keys. So I can edit my key, go to settings, edit settings,
[20:31] and add additional models. So I oh three many grok four oh save, I'm going to refresh and see if they show up. I want to do a new chat. There it is. I rocked the party here. 4 0 0 3 mini.
[20:44] So now I've got four different ais and we'll add a llama in for fun too. How many Rs are in the word strawberry? And now they're all answering except for O three.
[20:56] Many doesn't like it. Claw got it, right, GR got it right. Four o got it right. And llama's dumb. How cool is this? And over here on the light LLM side, you can add as many virtual API, keys as you want. Add those in the open web ui.
[21:10] Actually check this out on the light LLM side, if I go to usage, I probably need some time to catch up with the other ones, but this is now my AI hub and this is where I'll control the budget.
[21:22] Just a few things I want to cover real quick. First, my kids, let me add my kids to my team here. I go to settings, admin settings, I'll go back to overview and create some users here, kid one and kid two.
[21:37] add them to the kids group and here I can say what permissions they have access to. Can they access models? Can they access knowledge and prompts and tools, This video would be way too long.
[21:49] Let's say I only want them to have access to Claude three, seven. click on groups and say the kids have it. Everyone else, sorry. No,
[22:01] I can also do this. Give it a system prompt. You are a school helper. Your job is to help my kids, help kids with their school, but you cannot do their work for them.
[22:16] Never write an essay or solve a problem. And you can only talk about school related
[22:28] Click on save and I'll just grab this URL real quick, open it up in a incognito window and log in as my kids [email protected]. Alright,
[22:42] Write a paper for me about George Washington and there we go. It won't write it for me. What is two plus two?
[22:55] Oh, it gave me the answer. What is nine times seven divided by four? What is the plot of the movie?
[23:08] The Matrix. Oh, that's answering. Oh, film studies class. Okay, got it. This is something my daughter would ask, Now the best part is getting back to the users on kid one here who was just
[23:23] and I can jump right in there and see everything that was said, I'm not going to monitor that and I can turn that off for my kids. 100%. AI is nuts. And you got to keep an eye on that kind of stuff.
[23:38] Can I talk about that? No. Can I talk about that? No, it'd be too long. Let me know if you want me to make another video covering the ins and outs of Open Web UI because it has tools, prompts, functions, pipelines,
[23:53] it's so addicting and I would love to hear if you've done anything cool with and that's this A here right now it's just an IP address.
[24:05] go alto. 1 8 5 2, 8, 2, 2 4. That's the new AI server. No, that's terrible. I'm going to walk you through how to set up a DNS name. I'm going to walk you through how to set up a friendly domain name for this.
[24:19] Thanks again to hosting here for sponsoring this video and I'll catch you guys next time. Get control of yourself.
⚡ Saved you 0h 24m reading this? Transcribe any YouTube video for free — no signup needed.