TubeSum

Run Claude Code with Ollama Locally — Full Breakdown & Trans

How To Use Claude Code With Ollama (Free Local AI Setup)

0h 08m video Published Jun 28, 2026 Transcribed Jul 28, 2026 Ksk Royal Ksk Royal
Beginner 4 min read For: Developers and tech enthusiasts interested in running local AI models for coding assistance.
AI Trust Score 95/100
✅ Highly Legit

"Title accurately describes the video: it shows how to use Claude Code with Ollama for a free local AI setup."

AI Summary

This video demonstrates how to run Claude Code using Ollama and local AI models on a Mac, Windows, or Linux machine. The presenter shows the setup process, from installing Ollama to downloading and running models like Gemma 2B and Gemma 12B, and then using Claude Code to build a to-do list application. The video highlights the performance differences between smaller and larger models for coding tasks.

[00:00]
Introduction to Local AI Setup

The video shows how to run Claude Code using Ollama and local AI models on an M4 MacBook Air, but the guide works on any modern Mac, Windows, or Linux machine.

[00:32]
Ollama's Anthropic Compatible Endpoint

Ollama recently added native support for the Anthropic compatible API endpoint, allowing Claude Code to talk directly to models running on your own machine.

[00:46]
Installing Ollama

Download the installer from Ollama.com, copy the installation command, paste it into the terminal, and Ollama is installed. A small Ollama icon in the menu bar indicates it's running.

[01:03]
Choosing a Model

Enable 'tools and thinking' filters in Ollama's model section to see compatible models. For 24GB+ RAM, Qwen3.6 27B is recommended. This video uses Gemma 2B (1GB RAM) and Gemma 12B (13GB RAM).

[02:05]
Downloading and Running a Model

Copy the model tag, run the command to download (e.g., Gemma 2B ~7GB), then launch in interactive mode. On M4 MacBook Air, speed is ~30-40 tokens per second.

[03:16]
Increasing Context Window

Change Ollama's context length to 64K in settings. Default is ~8000 tokens, which is too small for large coding tasks. This allows Claude Code to hold conversation history and task instructions.

[03:59]
Connecting Claude Code with Ollama

Go to the provided page to learn how to connect Claude Code with Ollama. Copy the launch command and press Enter. If Claude Code CLI isn't installed, Ollama installs it automatically.

[04:43]
Building a To-Do List App

Using Gemma 2B, the presenter asks Claude Code to create a simple to-do list app with HTML, CSS, and JavaScript. It finishes in about 3 minutes, creating a fully working app.

[05:29]
Adding Dark Mode with 2B Model

Asking to add a dark mode toggle takes ~4 minutes. The dark mode technically works but the UI looks bad, showing the 2B model's limits in reasoning and consistency across multiple files.

[06:32]
Upgrading to Gemma 12B

Download Gemma 12B and launch Claude Code again. Ask it to analyze the existing project, identify issues, and fix them. The 12B model provides better reasoning and output, though it takes longer.

[07:35]
Conclusion

Running Claude Code on your own hardware with Ollama and local AI models is possible. Larger models offer better performance and reasoning.

Running Claude Code locally with Ollama is feasible and free, with smaller models like Gemma 2B handling simple tasks quickly, while larger models like Gemma 12B provide stronger reasoning for complex projects.

Mentioned in this Video

Tutorial Checklist

1 00:46 Go to Ollama.com and download the installer for your OS, then run the installation command in terminal.
2 01:03 Enable 'tools and thinking' filters in Ollama's model section to see compatible models.
3 02:05 Copy the model tag (e.g., Gemma 2B) and run 'ollama run <tag>' to download and start the model.
4 03:16 Open Ollama settings and change context length to 64K.
5 03:59 Copy the launch command from the Claude Code connection page and paste it into terminal to start Claude Code with Ollama.
6 04:43 Use Claude Code to build or modify projects by providing instructions.

Study Flashcards (5)

What is the recommended context window size for Claude Code with Ollama?

easy Click to reveal answer

64K tokens.

03:16

What is the default context window size in Ollama?

easy Click to reveal answer

Around 8000 tokens.

03:28

Which model is recommended for computers with 24GB RAM or more?

medium Click to reveal answer

Qwen3.6 27B.

01:18

What is the generation speed of Gemma 2B on an M4 MacBook Air?

medium Click to reveal answer

Around 30-40 tokens per second.

02:41

Why does the 2B model struggle with adding dark mode to the to-do list app?

hard Click to reveal answer

It lacks reasoning capacity to make consistent changes across multiple files.

06:19

💡 Key Takeaways

🔧

Local AI Setup

Shows that running Claude Code locally is possible on basic hardware like an M4 MacBook Air.

💡

Context Window Importance

Increasing context to 64K is crucial for large coding tasks to prevent the model from losing track.

03:16
📊

2B Model Limitations

Demonstrates that small models struggle with complex multi-file updates.

05:29
💡

12B Model Reasoning

Larger model provides significantly better reasoning and can fix issues the smaller model created.

06:32

[00:00] In this video, I'm going to show you how to run Quad Code using Olima and local AI models. I'm doing this on M4 base MacBook Air, the most basic Apple Silicon Mac.

[00:13] If you have any modern Mac, Windows or Linux machine, you can follow the same guide. The recently Olima added the native support for the Anthropic compatible APR endpoint.

[00:32] That means Quad Code can now talk directly to the model running on your own machine. The first, however, to Olima.com and download the installer for your operating system,

[00:46] I simply copy the installation command, paste it into the terminal, and that's it. Olima is installed. Once it's running, you will see a small Olima icon

[01:03] in the menu bar. That means Olima is running in the background and is ready to serve local models. I start to install a model that works with Quad Code. I go to Olima's model section and enable

[01:18] the tools and thinking filters. This will show you the most popular compatible models. If your computer has 24GB of RAM or more, I recommend Qn3.627b. It's currently one of the best coding

[01:34] models you can run on Kinsium or Hardware. For this video, we're going to use two models from Google's Jamfer family. We will start with Jamfer E2B, a small and fast-edge model that runs on just a GB of

[01:51] RAM. Then later, we will upgrade to Jamfer 12B when we need more reasoning power, which requires around 13GB of RAM. Since I'm using Apple Silicon, I will be downloading the MLX versions because they're

[02:05] optimized for this hardware. I go ahead and copy the model tag you want to use. Then run this command and replace the example tag with your chosen model. We are downloading Jamfer E2B through

[02:20] Olima. It's around 7GB, so it should take just a few minutes. Once it's downloaded, let's launch it in Olima's interactive mode and ask a simple question.

[02:41] Look at the bottom of the terminal. Olima shows your generation speed. On this M4 MacBook Air, we are getting around 30 tokens per second, which feels extremely responsive.

[02:58] If I continue the conversation a little longer, you can see it's almost giving me 40 tokens per second. And when you are done, press Ctrl plus D or simply type by to exit. Before using Cloudcode,

[03:16] there's one setting you must change. If we need to increase the model's context window to 64K, Cloudcode needs more context window to hold your fast conversation history and task instructions

[03:28] at the same time. By default, Olima uses a much smaller context window of around 8000 tokens, with large coding tasks the model quickly loses track of what it's doing. To fix this,

[03:41] open the Olima settings and change the context length to 64K. And don't worry, it won't always use the full 64K. We're simply increasing the maximum available context.

[03:59] Next, head over to this page to learn how to connect Cloudcode with Olima. I simply copy the launch command and press Enter.

[04:12] If Cloudcode CLI isn't already installed, Olima will detect and install it automatically. Once it launches, choose the model you want to use.

[04:29] And that's it. We are now running Cloudcode with Jamal4. Now, let's ask it a few questions.

[04:43] And as you can see, everything works perfectly. Now, let's build something real. I'm going to ask to create a simple to do list application using HTML, CSS and JavaScript. I will also tell it to save the source files

[05:00] explicitly. Otherwise, it may only generate the code instead of writing the files. And as you can see, it's making two calls reading the project directory, planning the structure and creating the files. In just about three minutes, it finishes building

[05:17] the entire application, which is pretty impressive. Now, let's open it in the browser. And there it is, a fully working to do list application. You can add task, delete them and

[05:29] everything works. And just remember, this was built by a 2 billion parameter model running locally and completely free. Now, let's push it a bit further. I'm going to ask it to add a dark mode toggle.

[05:47] This time, it takes around four minutes. And let's see the result. The dark mode technically

[06:05] works, but the UI looks very bad. This is where the 2B model starts reaching its limits, updating the HTML, CSS and JavaScript together requires the model to keep the entire project in memory

[06:19] and make consistent changes across multiple files. A 2 billion parameter model simply doesn't have enough reasoning capacity for more complex tasks. It's great for small jobs, but once the project grows,

[06:32] it starts making mistakes. To fix this, I downloaded Jummo412b, then I launched Gladcode again using the larger model. This time, I asked it to first analyze the existing project, identify

[06:49] what went wrong and then fix it. Now, watch this. It correctly identifies the CSS rules that weren't updating and also checks how the JavaScript class toggles dark mode and applies targeted fixes.

[07:06] It took quite a bit longer this time, but this is the result. The UI now looks much better. However, it accidentally broke the dark mode functionality.

[07:20] So I asked it to fix one last time, and finally, everything works correctly. And that's what a 12B model gives you, not just better output, but significantly stronger reasoning.

[07:35] If you have a faster computer with more RAM, you can even run larger models that perform even better. And that's how you can run Gladcode on your own hardware using Olama and local AI models.

[07:49] Let me know what do you think about this in the comments section down below. If you liked this video, hit that like button and subscribe to see more content. Thank you so much for watching. This is being KSKRIO. I will see you in the next one.

⚡ Saved you 0h 08m reading this? Transcribe any YouTube video for free — no signup needed.