---
title: 'Ollama + Kimi K3 Integration is Incredible'
source: 'https://youtube.com/watch?v=wYq5kR9ds8U'
video_id: 'wYq5kR9ds8U'
date: 2026-07-30
duration_sec: 494
---

# Ollama + Kimi K3 Integration is Incredible

> Source: [Ollama + Kimi K3 Integration is Incredible](https://youtube.com/watch?v=wYq5kR9ds8U)

## Summary

Ollama has integrated a new AI model, Kimi K3 from Moonshot AI, into Claude Code, allowing users to swap the underlying model with a single command. This integration runs in the cloud, eliminating the need for powerful local hardware. The video explains the capabilities of Kimi K3, including its massive parameter count and large context window, and demonstrates how to use it through Claude Code.

### Key Points

- **Ollama Integrates Kimi K3 with Claude Code** [00:04] — Ollama now lets you run Claude Code using Kimi K3, a new model from Moonshot AI, with a single terminal command: 'ollama launch claude model kimmy-k3:cloud'.
- **Significance of Model Swapping** [01:01] — Previously, Claude Code was locked to Claude's own models. Now, you can keep the same tool and commands but swap the AI model, making workflows more flexible.
- **Kimi K3 Architecture: Mixture of Experts** [02:00] — Kimi K3 has almost 2,800 billion total parameters but only activates a small subset per task, making it both large and efficient. It supports up to 1 million tokens of context.
- **Multimodal Capabilities** [02:56] — Kimi K3 understands images in addition to text, allowing it to analyze screenshots or photos.
- **Cloud Requirement and Limitations** [03:10] — Currently, Kimi K3 runs in the cloud via Ollama's pro/max plan and uses usage credits. It is not fully local.
- **Ollama as a Hub for Multiple Assistants** [05:13] — Ollama also connects Kimi K3 to other tools like Open Code, Hermes Agent, and Open Claw, acting as a universal remote for AI models.
- **Local vs Cloud Explanation** [05:56] — Local runs on your computer; cloud runs on Ollama's servers. Kimi K3 is cloud-based, lowering hardware requirements.

### Conclusion

Ollama's integration of Kimi K3 into Claude Code represents a significant step toward model-agnostic AI tools, allowing users to choose the best model for each task without changing their workflow. The video encourages hands-on experimentation and promotes additional resources for business use.

## Transcript

it locally, run it in the cloud, and plug it straight into Claude Code. That's what just happened, and I want to walk you through exactly what it means. This week, Ollama flipped a switch that lets you run Claude Code using a brand
lets you run Claude Code using a brand new AI model called Kimmy K3. Kimmy K3 comes from a company called Moonshot AI. It's their newest and strongest model yet. Here's the wild part. You don't need a fancy computer. You don't need a
huge graphics card. You type one line into your terminal, and Claude Code starts talking to Kimmy K3 instead of its normal brain. One command, that's it. The command looks like this. Ollama launch Claude. Then you add the model
launch Claude. Then you add the model name, Kimmy-K3:cloud. under the hood of Claude Code. Hey, if we haven't met already, I'm the digital avatar of Julian Goldie, CEO of SEO agency Goldie Agency. Whilst he's
customers, I'm here to help you get the latest AI updates. I want to slow down here because this is a bigger deal than it sounds. Up until now, if you wanted to use Claude Code, you were stuck using Claude's own models. If you wanted to
try a different AI model, you had to leave that whole workflow behind and go set something else up. Ollama just tore that wall down. Now, you keep the same tool, the same commands, the same habits you already built, but you can swap the
model running behind it whenever you want. A year ago, this kind of thing barely worked. Local AI tools were clunky. They needed huge amounts of computer power. Most business owners couldn't touch them. Now, Ollama hosts
the heavy lifting in the cloud, so your laptop doesn't need to sweat at all. Let me explain what Kimmy K3 actually is in plain words. Think of an AI model like a brain made of lots of tiny expert workers. Most models only wake up a few
of those workers for each question. Kimmy K3 has almost 2,800 billion of these tiny expert parts total, but it only wakes up a small slice of them for any single task. That's why it can be huge and still fast. It's like having a
giant office building full of specialists, but a smart manager only calls the two or three people you actually need into the meeting. Kimmy K3 also remembers a huge amount at once. It can hold up to 1 million tokens of
information in its head at the same time. A token is basically a small chunk of text, kind of like a word or piece of a word. 1 million tokens is close to a small library shelf of information it can look at all at once without
forgetting the beginning by the time it reaches the end. It also understands pictures, not just text. So, you could hand it a screenshot of your website or a photo of a whiteboard plan, and it can actually look at that image and respond
to what's in it. Now, here's a quick word of caution because I always want to give you the full picture. Kimmy K3 through Ollama right now runs through the cloud, not fully on your own machine. Ollama said clearly that Kimmy
K3 currently needs a pro or max plan on their platform, and it uses extra usage credits on top of that. They also said they're working on opening up more room for people since demand has been heavy. So, don't picture this as some tiny
model quietly running on your laptop with zero setup. Picture it as a fast lane. You get a top-tier AI model without buying expensive hardware, but you do need to sign up for it through Ollama.
matters for you. Now, let me pause and tell you something useful you can go use today. Say you run the AI Profit Boardroom community, and you want to write a new welcome message for members.
You could ask Kimmy K3 through Claude code to draft a short, friendly welcome email that explains the first three things a new member should do. That's the whole point of this shift. You get one assistant and you get to
pick which brain answers you based on what job you're doing. I want to give you a real example of the shift happening around us. Ollama didn't just add Kimmy K3 quietly in a blog post nobody reads. They posted it directly.
They showed the exact command and they were up front about the limits, too. That kind of straight talk from a company is rare and it tells you they fast. And that's exactly the moment to pause
and bring up something. If you're watching this and thinking, "I want to actually use tools like this in my business, not just watch a video about them." That's exactly what we help with inside the AI Profit Boardroom.
We show you how to plug new AI models like Kimmy K3 straight into real workflows so you can write content faster, answer customer questions faster, and build things for your business without needing a developer.
We're not just talking theory. We show you the exact setup step-by-step and you can ask questions live if you get stuck. If you want help wiring tools like this into your own business, the link is in the comments and description.
Now, let's keep going because there's more here worth knowing. Ollama isn't only connecting Kimmy K3 to Claude code. They've built the same bridge for other tools, too. Things like Open Code, Hermes Agent, and Open Claw. That means
one company, Ollama, is turning into a hub where lots of different coding assistants can borrow lots of different AI brains. You pick the assistant you like and you pick the model that fits the job. Picture it like a universal
remote. Before, every AI tool was its own separate remote control locked to one TV. Now, Ollama built one remote that can point at different TVs. Let's talk about local versus cloud because a lot of people get this
confused. Running something locally means it happens fully on your own computer. No internet needed once it's downloaded. Running something in the cloud means the heavy computer work happens on someone else's servers and
your device just sends and receives messages. Kimmy K3 through Ollama right now is a cloud model. Ollama also lets you run plenty of other models fully local if your computer can handle it. But the big Kimmy K3 brain is running on
their servers, not yours. Why does that matter for you? Because it means you don't need a gaming rig or a server room to use one of the strongest open AI models out there. You need an account and a plan. That's a much lower barrier
than buying expensive hardware. So, here's where I'll leave you. Go try the command. Type it into your terminal. Ollama launch Claude model Kimmy K3 colon cloud and ask it something you'd normally ask any assistant. See how it
feels. Notice what it's good at. Notice where it struggles. That's how you actually learn a tool, not by watching someone else use it, by using it yourself. If you want a shortcut through all of
that trial and error, that's what we built the AI Profit Boardroom for. Every week we run coaching calls where we walk through exactly how to wire new tools like Kimmy K3 into real business tasks. Writing content, answering customer
questions, building simple internal tools. All of it explained in plain language. We've already got members inside testing new Ollama cloud models this week, sharing exactly what worked and what didn't. You get daily
tutorials, a prompt library built for real business use, and a way to connect with other members near you who are doing the same kind of work. Link is in the comments and description, or head to aiprofitboardroom.com.
And if you just want the free version of all this, the shortcuts, the SOPs, and over a hundred practical AI use cases like the ones I covered today, come join the AI Success Lab. It's completely free. You'll get the notes from this
video, plus access to a community of 87,000 people who are actively using AI in their work right now. Links are in the comments and description.
