TubeSum โ† Transcribe a video

Ollama + Kimi K3 Integration is Incredible

0h 08m video Published Jul 29, 2026 Transcribed Jul 30, 2026 J Julian Goldie SEO
Intermediate 6 min read For: Business owners, developers, and AI enthusiasts interested in practical AI tool integration without deep technical expertise.
AI Trust Score 45/100
๐Ÿšซ Clickbait / Waste of Time

"Title overhypes a solid but incremental update; video delivers useful info but is padded with community promotions."

AI Summary

Ollama has integrated a new AI model, Kimi K3 from Moonshot AI, into Claude Code, allowing users to swap the underlying model with a single command. This integration runs in the cloud, eliminating the need for powerful local hardware. The video explains the capabilities of Kimi K3, including its massive parameter count and large context window, and demonstrates how to use it through Claude Code.

[00:04]
Ollama Integrates Kimi K3 with Claude Code

Ollama now lets you run Claude Code using Kimi K3, a new model from Moonshot AI, with a single terminal command: 'ollama launch claude model kimmy-k3:cloud'.

[01:01]
Significance of Model Swapping

Previously, Claude Code was locked to Claude's own models. Now, you can keep the same tool and commands but swap the AI model, making workflows more flexible.

[02:00]
Kimi K3 Architecture: Mixture of Experts

Kimi K3 has almost 2,800 billion total parameters but only activates a small subset per task, making it both large and efficient. It supports up to 1 million tokens of context.

[02:56]
Multimodal Capabilities

Kimi K3 understands images in addition to text, allowing it to analyze screenshots or photos.

[03:10]
Cloud Requirement and Limitations

Currently, Kimi K3 runs in the cloud via Ollama's pro/max plan and uses usage credits. It is not fully local.

[05:13]
Ollama as a Hub for Multiple Assistants

Ollama also connects Kimi K3 to other tools like Open Code, Hermes Agent, and Open Claw, acting as a universal remote for AI models.

[05:56]
Local vs Cloud Explanation

Local runs on your computer; cloud runs on Ollama's servers. Kimi K3 is cloud-based, lowering hardware requirements.

Ollama's integration of Kimi K3 into Claude Code represents a significant step toward model-agnostic AI tools, allowing users to choose the best model for each task without changing their workflow. The video encourages hands-on experimentation and promotes additional resources for business use.

Mentioned in this Video

Tutorial Checklist

1 00:04 Ensure you have an Ollama account with a pro or max plan.
2 00:46 Run the command: ollama launch claude model kimmy-k3:cloud
3 03:38 Test the setup by asking Kimi K3 a task, e.g., draft a welcome email.

Study Flashcards (5)

What command allows you to run Claude Code with Kimi K3?

easy Click to reveal answer

ollama launch claude model kimmy-k3:cloud

00:46

What is the total parameter count of Kimi K3?

medium Click to reveal answer

Almost 2,800 billion

02:00

How many tokens of context can Kimi K3 handle?

medium Click to reveal answer

Up to 1 million tokens

02:29

Does Kimi K3 run locally or in the cloud through Ollama?

easy Click to reveal answer

It runs in the cloud on Ollama's servers.

05:56

What are the two promoted communities mentioned in the video?

easy Click to reveal answer

AI Profit Boardroom (paid) and AI Success Lab (free)

04:44

๐Ÿ’ก Key Takeaways

๐Ÿ”ง

Model Swapping with One Command

Demonstrates a significant reduction in friction for using different AI models with existing tools.

00:04
๐Ÿ“Š

Kimi K3's MoE Architecture

Explains how a model can be both huge and efficient by only activating relevant experts.

02:00
๐Ÿ“Š

Million-Token Context Window

Highlights the model's ability to process large amounts of information at once, enabling complex tasks.

02:29
๐Ÿ’ก

Ollama as a Universal Remote

Illustrates the shift towards model-agnostic AI tools, allowing users to choose the best model for each task.

05:13

[00:04] it locally, run it in the cloud, and plug it straight into Claude Code. That's what just happened, and I want to walk you through exactly what it means. This week, Ollama flipped a switch that lets you run Claude Code using a brand

[00:17] lets you run Claude Code using a brand new AI model called Kimmy K3. Kimmy K3 comes from a company called Moonshot AI. It's their newest and strongest model yet. Here's the wild part. You don't need a fancy computer. You don't need a

[00:31] huge graphics card. You type one line into your terminal, and Claude Code starts talking to Kimmy K3 instead of its normal brain. One command, that's it. The command looks like this. Ollama launch Claude. Then you add the model

[00:46] launch Claude. Then you add the model name, Kimmy-K3:cloud. under the hood of Claude Code. Hey, if we haven't met already, I'm the digital avatar of Julian Goldie, CEO of SEO agency Goldie Agency. Whilst he's

[01:01] customers, I'm here to help you get the latest AI updates. I want to slow down here because this is a bigger deal than it sounds. Up until now, if you wanted to use Claude Code, you were stuck using Claude's own models. If you wanted to

[01:16] try a different AI model, you had to leave that whole workflow behind and go set something else up. Ollama just tore that wall down. Now, you keep the same tool, the same commands, the same habits you already built, but you can swap the

[01:30] model running behind it whenever you want. A year ago, this kind of thing barely worked. Local AI tools were clunky. They needed huge amounts of computer power. Most business owners couldn't touch them. Now, Ollama hosts

[01:44] the heavy lifting in the cloud, so your laptop doesn't need to sweat at all. Let me explain what Kimmy K3 actually is in plain words. Think of an AI model like a brain made of lots of tiny expert workers. Most models only wake up a few

[02:00] of those workers for each question. Kimmy K3 has almost 2,800 billion of these tiny expert parts total, but it only wakes up a small slice of them for any single task. That's why it can be huge and still fast. It's like having a

[02:15] giant office building full of specialists, but a smart manager only calls the two or three people you actually need into the meeting. Kimmy K3 also remembers a huge amount at once. It can hold up to 1 million tokens of

[02:29] information in its head at the same time. A token is basically a small chunk of text, kind of like a word or piece of a word. 1 million tokens is close to a small library shelf of information it can look at all at once without

[02:43] forgetting the beginning by the time it reaches the end. It also understands pictures, not just text. So, you could hand it a screenshot of your website or a photo of a whiteboard plan, and it can actually look at that image and respond

[02:56] to what's in it. Now, here's a quick word of caution because I always want to give you the full picture. Kimmy K3 through Ollama right now runs through the cloud, not fully on your own machine. Ollama said clearly that Kimmy

[03:10] K3 currently needs a pro or max plan on their platform, and it uses extra usage credits on top of that. They also said they're working on opening up more room for people since demand has been heavy. So, don't picture this as some tiny

[03:25] model quietly running on your laptop with zero setup. Picture it as a fast lane. You get a top-tier AI model without buying expensive hardware, but you do need to sign up for it through Ollama.

[03:38] matters for you. Now, let me pause and tell you something useful you can go use today. Say you run the AI Profit Boardroom community, and you want to write a new welcome message for members.

[03:51] You could ask Kimmy K3 through Claude code to draft a short, friendly welcome email that explains the first three things a new member should do. That's the whole point of this shift. You get one assistant and you get to

[04:04] pick which brain answers you based on what job you're doing. I want to give you a real example of the shift happening around us. Ollama didn't just add Kimmy K3 quietly in a blog post nobody reads. They posted it directly.

[04:19] They showed the exact command and they were up front about the limits, too. That kind of straight talk from a company is rare and it tells you they fast. And that's exactly the moment to pause

[04:32] and bring up something. If you're watching this and thinking, "I want to actually use tools like this in my business, not just watch a video about them." That's exactly what we help with inside the AI Profit Boardroom.

[04:44] We show you how to plug new AI models like Kimmy K3 straight into real workflows so you can write content faster, answer customer questions faster, and build things for your business without needing a developer.

[04:58] We're not just talking theory. We show you the exact setup step-by-step and you can ask questions live if you get stuck. If you want help wiring tools like this into your own business, the link is in the comments and description.

[05:13] Now, let's keep going because there's more here worth knowing. Ollama isn't only connecting Kimmy K3 to Claude code. They've built the same bridge for other tools, too. Things like Open Code, Hermes Agent, and Open Claw. That means

[05:27] one company, Ollama, is turning into a hub where lots of different coding assistants can borrow lots of different AI brains. You pick the assistant you like and you pick the model that fits the job. Picture it like a universal

[05:41] remote. Before, every AI tool was its own separate remote control locked to one TV. Now, Ollama built one remote that can point at different TVs. Let's talk about local versus cloud because a lot of people get this

[05:56] confused. Running something locally means it happens fully on your own computer. No internet needed once it's downloaded. Running something in the cloud means the heavy computer work happens on someone else's servers and

[06:09] your device just sends and receives messages. Kimmy K3 through Ollama right now is a cloud model. Ollama also lets you run plenty of other models fully local if your computer can handle it. But the big Kimmy K3 brain is running on

[06:23] their servers, not yours. Why does that matter for you? Because it means you don't need a gaming rig or a server room to use one of the strongest open AI models out there. You need an account and a plan. That's a much lower barrier

[06:38] than buying expensive hardware. So, here's where I'll leave you. Go try the command. Type it into your terminal. Ollama launch Claude model Kimmy K3 colon cloud and ask it something you'd normally ask any assistant. See how it

[06:53] feels. Notice what it's good at. Notice where it struggles. That's how you actually learn a tool, not by watching someone else use it, by using it yourself. If you want a shortcut through all of

[07:05] that trial and error, that's what we built the AI Profit Boardroom for. Every week we run coaching calls where we walk through exactly how to wire new tools like Kimmy K3 into real business tasks. Writing content, answering customer

[07:20] questions, building simple internal tools. All of it explained in plain language. We've already got members inside testing new Ollama cloud models this week, sharing exactly what worked and what didn't. You get daily

[07:34] tutorials, a prompt library built for real business use, and a way to connect with other members near you who are doing the same kind of work. Link is in the comments and description, or head to aiprofitboardroom.com.

[07:50] And if you just want the free version of all this, the shortcuts, the SOPs, and over a hundred practical AI use cases like the ones I covered today, come join the AI Success Lab. It's completely free. You'll get the notes from this

[08:03] video, plus access to a community of 87,000 people who are actively using AI in their work right now. Links are in the comments and description.

More from Julian Goldie SEO

View all

โšก Saved you 0h 08m reading this? Transcribe any YouTube video for free โ€” no signup needed.