Open Weights AI Can Copy Mac OS!
44sShows an open-source AI model replicating a full operating system, sparking debate about open vs. closed AI.
▶ Play Clip"Delivers on the economic disruption promise with real cost reductions and open weights, but hype and sponsor plug slightly dilute the impact."
Kimi K3, a new open-weights AI model with 2.8 trillion parameters, is introduced. It can code a full Mac OS clone and games, leverages novel techniques like KDA and attention residuals for 2.5x scaling efficiency, and is available via cheap API. This marks a shift toward accessible, frontier-level AI.
Can code a mostly working copy of Mac OS, an Animal Crossing-style game, and more. It approaches frontier system performance while being open weights.
Model weights are free to download and own forever. No one can take it away. Users can try it for free on the web if available, and API pricing is much cheaper than current frontier models.
The model is stupendously large with 2.8 trillion parameters, though most cannot run it locally. Distilled versions are expected soon.
Two key innovations: Kimi Delta Attention (KDA) which maintains a carefully updated notebook across layers, letting old notes fade gradually; and attention residuals which pass version history between layers, not just the latest document.
Combining KDA and attention residuals yields roughly 2.5 times more learning progress from the same training computation, though not directly 2.5x cheaper or faster.
Every such paper improves all open models. This is the golden age of open science with free weights, open-source OS, and global human collaboration. AI helps doctors, scientists, and students for free.
Lambda provides powerful Nvidia GPUs to reproduce AI research papers, train models, or run inference. Recommended for experimenting with ideas from the paper.
Kimi K3 represents a major leap in open-weights AI, combining unprecedented scale with novel efficiency techniques that make frontier-level AI more accessible and affordable, heralding a new era of open science.
What is the parameter count of Kimi K3?
2.8 trillion parameters.
00:47
What are the two key technical innovations in Kimi K3?
Kimi Delta Attention (KDA) and attention residuals.
01:30
How does KDA maintain memory?
It uses a notebook that is carefully updated across layers, allowing old notes to gradually fade.
01:46
What advantage do attention residuals provide?
They pass version history (useful earlier drafts) between layers, not just the latest document.
02:31
What is the scaling efficiency improvement of Kimi K3 over Kimi K2?
2.5 times more learning progress from the same training computation.
03:01
Is Kimi K3 available for free?
Yes, the weights are free and open; users can try it for free on the web or use the cheap API.
00:30
Open Weights Revolution
Marks a shift toward accessible frontier AI, free from corporate control.
00:30Novel Architecture Innovations
KDA and attention residuals directly improve scaling efficiency, a key metric.
01:302.5x Efficiency Gain
Quantifies the practical impact of the innovations on training cost.
03:01Golden Age of Open Science
Emphasizes the collaborative, global benefit of open models.
03:32[00:02] I'm seeing here. This is a new AI system, Kim K3, that can code up a mostly working copy of a full Mac OS operating system. Also, a cute Animal
[00:14] Crossing style game, other kinds of games, and so much more. And what are you seeing? What I am seeing? Yes, it is very close to the Frontier systems, some of which are kind of getting banned sometimes, but this one won't. Why?
[00:30] Well, here comes the best part. This is an open weights model. Yes, we can download and own the weights for free forever. And nobody can take this from us. That is absolutely incredible. Now, wait, wait, wait. It is big. Caro, do
[00:47] you mean that it's big news? No, I mean it is big. It is absolutely stupendously, humongously big. 2.8 [screaming] trillion parameters. Most of
[00:59] us can't afford to be running this at home. Not as is, but with a little luck, you can try it for free on the web depending on availability and take it out for a spin. Or if you use the API, it is way way cheaper than current
[01:14] Frontier models. So even if you don't ever use it, it will be pushing token prices down. Also, don't forget these huge models are often distilled down to smaller, hopefully similarly capable ones on a regular basis. And somehow it
[01:30] gets even better. They gave us the secret sauce. So, what is the secret sauce? Dear fellow scholars, this is two minute papers with Dr. Koa Eer. One Kimmy Delta attention. Imagine a meeting where every researcher has to reread
[01:46] where every researcher has to reread everything everyone ever did. Oh, that's not a meeting. That's torture. Basically, instead, this says, "Let's have a carefully updated notebook, read and update only that." And it lets old
[02:01] and update only that." And it lets old notes gradually fade a bit. This finally lets the institute handle a very long discussion and contribute meaningfully. Now, wait, what happens when this information passes through dozens of
[02:15] layers? Well, secret sauce number two, attention residuals. Imagine that every document goes through department one. Then department two and three and four only gets the latest version of the document. With attention residuals,
[02:31] department 4 still gets the latest version of the document, but also a version history as well and see how the document has changed over time. So, KDA
[02:44] maintains and corrects memory. Attention residuals retrieve useful earlier drafts across layers. And wait until you hear what happens when we combine these two ideas. So, what happens? Well, hold on to your papers, fellow scholars, because
[03:01] it results in a 2 and 12x improvement in scaling efficiency over Kimmy K2. Wow. Now, this does not mean it is 2 and 1/2x cheaper or 2 and 1/2x faster. No,
[03:15] it means roughly two and a half times more learning progress out of the same amount of training computation. That is a huge bump over just one version number. So, what does this enable? Well, all of these incredible things here and
[03:32] something more. Don't forget with every paper like this we make all the other open models work better. We are building this together and you see we get all of this for free forever. Yes, it is not trivial to run yet but I think it will
[03:50] be trivial to run a distilled version of this hopefully soon. And don't forget this is the golden age of open science. You can have an AI like this running You can have an AI like this running fully free open weights in an operating
[04:04] system that is fully free open source and all of this developed by humans working together across the planet. And these AI systems help doctors, scientists, students learn and do their work all across the world for free.
[04:20] Everyone will get access. Isn't that amazing? What a time to be alive. I use Lambda to reproduce AI research papers often in minutes. It's also great to train your own models or fine-tune an existing one. Run inference or text to
[04:37] image or video. Easy peasy. Running a DeepSeek chatbot or agent. Super fast, super reliable. Lambda gives you powerful Nvidia GPUs to run your own experiments. I test ideas from the papers I cover and moments later,
[04:53] results. Love it. Seriously, try it out now at lambda.ai/papers.
⚡ Saved you 0h 05m reading this? Transcribe any YouTube video for free — no signup needed.