Open-Source AI Beats GPT-5.6?
39sClaims of a Chinese open-source model outperforming top US models spark geopolitical debate and challenge AI safety narratives.
▶ Play Clip"Delivers on the title but padded with sponsor segment and some hype; overall a solid breakdown."
The video discusses the release of Kimi K3, a massive open-source AI model with 2.8 trillion parameters that rivals top proprietary models like Claude and GPT. It covers the model's architecture, benchmark performance, geopolitical implications, and the growing tension between open-source and regulated AI.
Moonshot AI released Kimi K3, a massive open-source model with 2.8 trillion parameters, instantly becoming the largest open-weight model.
K3's benchmark performance is on par with, and in some cases beats, Claude Fable and GPT-5.6 Soul, causing concern at OpenAI and Anthropic.
K3 is a mixture of experts model with 896 total experts, 16 active per token, making scaling 2.5x more efficient than K2.
K3 was so popular upon release that Moonshot's GPUs ran out, forcing them to turn away paying customers; plans are currently sold out.
K3 ranks #1 on front end code Arena with a 1,679 Elo, ahead of Fable 5 and GPT-5.6 Soul, but trails on other benchmarks like Humanity's Last Exam.
K3 has a 51% incorrect coding rate and tends to produce verbose output, though it excels at UI design and data visualization.
China's Communist Party has become a strong advocate for open AI, while US Silicon Valley pushes for regulation; Dean Ball called open weights 'decelerationist'.
Kimi K3 represents a significant leap in open-source AI, challenging proprietary models and sparking geopolitical debates about control and regulation. The arms race is accelerating, with Alibaba also releasing a competing open model.
What is the total number of parameters in Kimi K3?
2.8 trillion
00:41
How many experts does Kimi K3 have and how many activate per token?
896 total experts, 16 activate per token
00:57
What efficiency improvement does K3 have over K2?
About 2.5 times more efficient
01:09
What Elo rating did Kimi K3 achieve on front end code Arena?
1,679 Elo
02:02
On which benchmark does Kimi K3 trail by about 10 points?
Humanity's Last Exam
02:31
What percentage of Kimi K3's coding outputs are incorrect according to artificial analysis?
51%
02:44
What is the current probability on PolyMarket that the US will ban Chinese AI models?
29%
03:47
Model Specs
Provides concrete numbers on parameters and expert architecture, illustrating the scale and design philosophy.
00:41Coding Benchmark Dominance
Shows K3 outperforming top proprietary models on a practical coding metric, validating its capability.
02:02Geopolitical Shift
Highlights the ironic reversal where communist China advocates open-source while the US pushes for regulation.
03:10Benchmark Caveats
Reminds that K3 still trails on comprehensive exams and has real-world quality issues like high error rate.
02:31[00:02] dropped Kimi K3, a massive open-source monster that instantly parameter mogged every other open model in existence. And not just that, it has OpenAI and Anthropic terrified because its trust-me-bro benchmark performance is on
[00:15] par with, and in some cases beating, Claude Fable and GPT-5.6 Soul. And ago, the government was telling you these models were too dangerous for the Chinese a few weeks to match their performance and then give away all the
[00:28] weights for free. But big AI is not too happy about this, and there are already calls to ban Chinese models completely in the United States. In today's video, breakthrough that is Kimi K3 and find out what it means for the future of
[00:41] 22nd, [music] 2026, and you're watching The Code Report. Like Claude Fable and GPT-5.6 Soul, Kimi K3 is a native multi-model million token context window and a staggering 2.8 trillion parameters, and
[00:57] is optimized for jobs like long-horizon reasoning and, of course, coding. One interesting characteristic of K3 as a mixture of experts model is that it has mixture of experts model is that it has 896 total experts, of which exactly 16
[01:09] activate per token, which means it works just like a big corporation where 16 good programmers do all the work while 880 other managers sit there and do well for Kimi because it makes scaling about 2.5 times more efficient than K2.
[01:23] Despite these gains in efficiency, Kimi was so popular upon release that their GPUs ran out of juice and they had to start turning away paying customers. And plans are currently sold out. That's unfortunate, but in theory, you could
[01:36] are open. The weights are expected to be released on July 27th, but there's no chance in hell you'll be able to run it on your little gaming GPU. To run a monster like this, you'll need a massive array of data center caliber GPUs. But
[01:49] would have unlimited access to a Fable Soul caliber model, and that would be amazing because the benchmark situation with K3 is pretty wild. K3 is ranked number one on front end code Arena at a 1,679
[02:02] Elo, which puts it ahead of Fable 5 and GPT-5.6 Soul. In addition, it lands in intelligence index. And if we look at every other coding benchmark, it's at frontier models. But you should never trust the trust me bro benchmarks
[02:17] because many of the K3 numbers were produced with Moon Shot's own Kimiko different harnesses. That could make Kimiko look slightly better at coding, but to their credit, Moon Shot admits that K3 still trails Fable and GPT-5.6
[02:31] Soul overall, especially on benchmarks like Humanity's Last Exam where it's down by about 10 points. On top of that, artificial analysis measured a 51% not a good thing, especially when it comes to coding. In addition, it also
[02:44] tends to spit out way more tokens than it needs to, which could ultimately end model itself being cheaper. When it comes to things like UI design and data visualization, it's extremely impressive for an open model, but in my opinion,
[02:57] it's still one step behind Fable and GPT Soul. But one of the most interesting things about this release is the geopolitics surrounding it. Recently at Communist Party became the loudest advocate for free and open artificial
[03:10] intelligence. Meanwhile, in the land of the free, Silicon Valley wants to regulate and gate keep it by pushing the fear narrative that it's about to take all of our jobs. In Washington, they're reportedly considering entity listing
[03:22] Chinese AI labs and OpenAI's Dean Ball argued that open weights are inherently decelerationist. >> What? Bro, what are you talking about, man? >> And coincidentally, that's very similar
[03:34] make in the '90s about Linux when he said Linux is communism. And also coincidentally, both of these guys have balls in their names. Frontier labs don't like open models simply because they divert the flow of money from them
[03:47] to someone else. As of today, the odds the US government bans Chinese models is only sitting at 29% on Poly Market, but that could change quickly if they responsible for some kind of cyber attack. But, the best thing about K3 is
[03:59] that it pushes the arms race forward. Alibaba also just released Qwen 3.8, which itself has 2.4 trillion parameters and open weights. And I think this model UI right for Horse Tender. But, before you let AI slop out your UI, you need to
[04:13] check out mobbin.com, the sponsor of today's video. I've been using Mobbin for over 5 years now because it provides highly detailed breakdowns of every screen in thousands of popular web and mobile apps. And they just launched an
[04:25] MCP server, which connects your AI agent to over 600,000 screens and user flows from apps in every category. This gives your agent real-world references, so it can design high-quality UIs for your specific use case, instead of just
[04:39] spitting out generic purple gradient vibes law. It also lets you do deep UI prototype and your agent will use Mobbin's library to provide a ranked list of other apps that do it better. You can also ask it how your competitors
[04:52] paywalls, and it'll show you every screen in their full user flow. And so, if you're tired of your UIs looking like the same as everyone else, I'd highly below. This has been the Code Report. Thanks for watching, and I will see you
[05:06] Thanks for watching, and I will see you in the next one.
⚡ Saved you 0h 05m reading this? Transcribe any YouTube video for free — no signup needed.