---
title: 'Kimi K3 Just Broke the Economics of AI'
source: 'https://youtube.com/watch?v=Xj-QdEUxJkE'
video_id: 'Xj-QdEUxJkE'
date: 2026-07-29
duration_sec: 300
---

# Kimi K3 Just Broke the Economics of AI

> Source: [Kimi K3 Just Broke the Economics of AI](https://youtube.com/watch?v=Xj-QdEUxJkE)

## Summary

Kimi K3, a new open-weights AI model with 2.8 trillion parameters, is introduced. It can code a full Mac OS clone and games, leverages novel techniques like KDA and attention residuals for 2.5x scaling efficiency, and is available via cheap API. This marks a shift toward accessible, frontier-level AI.

### Key Points

- **Kimi K3 Capabilities** [00:02] — Can code a mostly working copy of Mac OS, an Animal Crossing-style game, and more. It approaches frontier system performance while being open weights.
- **Open Weights & Free Access** [00:30] — Model weights are free to download and own forever. No one can take it away. Users can try it for free on the web if available, and API pricing is much cheaper than current frontier models.
- **2.8 Trillion Parameters** [00:47] — The model is stupendously large with 2.8 trillion parameters, though most cannot run it locally. Distilled versions are expected soon.
- **Secret Sauce: KDA & Attention Residuals** [01:30] — Two key innovations: Kimi Delta Attention (KDA) which maintains a carefully updated notebook across layers, letting old notes fade gradually; and attention residuals which pass version history between layers, not just the latest document.
- **2.5x Scaling Efficiency Improvement** [03:01] — Combining KDA and attention residuals yields roughly 2.5 times more learning progress from the same training computation, though not directly 2.5x cheaper or faster.
- **Implications for Open Science** [03:32] — Every such paper improves all open models. This is the golden age of open science with free weights, open-source OS, and global human collaboration. AI helps doctors, scientists, and students for free.
- **Sponsor: Lambda GPU Cloud** [04:20] — Lambda provides powerful Nvidia GPUs to reproduce AI research papers, train models, or run inference. Recommended for experimenting with ideas from the paper.

### Conclusion

Kimi K3 represents a major leap in open-weights AI, combining unprecedented scale with novel efficiency techniques that make frontier-level AI more accessible and affordable, heralding a new era of open science.

## Transcript

I'm seeing here. This is a new AI system, Kim K3, that can code up a mostly working copy of a full Mac OS operating system. Also, a cute Animal
Crossing style game, other kinds of games, and so much more. And what are you seeing? What I am seeing? Yes, it is very close to the Frontier systems, some of which are kind of getting banned sometimes, but this one won't. Why?
Well, here comes the best part. This is an open weights model. Yes, we can download and own the weights for free forever. And nobody can take this from us. That is absolutely incredible. Now, wait, wait, wait. It is big. Caro, do
you mean that it's big news? No, I mean it is big. It is absolutely stupendously, humongously big. 2.8 [screaming] trillion parameters. Most of
us can't afford to be running this at home. Not as is, but with a little luck, you can try it for free on the web depending on availability and take it out for a spin. Or if you use the API, it is way way cheaper than current
Frontier models. So even if you don't ever use it, it will be pushing token prices down. Also, don't forget these huge models are often distilled down to smaller, hopefully similarly capable ones on a regular basis. And somehow it
gets even better. They gave us the secret sauce. So, what is the secret sauce? Dear fellow scholars, this is two minute papers with Dr. Koa Eer. One Kimmy Delta attention. Imagine a meeting where every researcher has to reread
where every researcher has to reread everything everyone ever did. Oh, that's not a meeting. That's torture. Basically, instead, this says, "Let's have a carefully updated notebook, read and update only that." And it lets old
and update only that." And it lets old notes gradually fade a bit. This finally lets the institute handle a very long discussion and contribute meaningfully. Now, wait, what happens when this information passes through dozens of
layers? Well, secret sauce number two, attention residuals. Imagine that every document goes through department one. Then department two and three and four only gets the latest version of the document. With attention residuals,
department 4 still gets the latest version of the document, but also a version history as well and see how the document has changed over time. So, KDA
maintains and corrects memory. Attention residuals retrieve useful earlier drafts across layers. And wait until you hear what happens when we combine these two ideas. So, what happens? Well, hold on to your papers, fellow scholars, because
it results in a 2 and 12x improvement in scaling efficiency over Kimmy K2. Wow. Now, this does not mean it is 2 and 1/2x cheaper or 2 and 1/2x faster. No,
it means roughly two and a half times more learning progress out of the same amount of training computation. That is a huge bump over just one version number. So, what does this enable? Well, all of these incredible things here and
something more. Don't forget with every paper like this we make all the other open models work better. We are building this together and you see we get all of this for free forever. Yes, it is not trivial to run yet but I think it will
be trivial to run a distilled version of this hopefully soon. And don't forget this is the golden age of open science. You can have an AI like this running You can have an AI like this running fully free open weights in an operating
system that is fully free open source and all of this developed by humans working together across the planet. And these AI systems help doctors, scientists, students learn and do their work all across the world for free.
Everyone will get access. Isn't that amazing? What a time to be alive. I use Lambda to reproduce AI research papers often in minutes. It's also great to train your own models or fine-tune an existing one. Run inference or text to
image or video. Easy peasy. Running a DeepSeek chatbot or agent. Super fast, super reliable. Lambda gives you powerful Nvidia GPUs to run your own experiments. I test ideas from the papers I cover and moments later,
results. Love it. Seriously, try it out now at lambda.ai/papers.
