TubeSum

DeepSeek V4 Pro: Open Weights & 78% Faster — Full Breakdown & Transcript

DeepSeek Just Made Closed AI Look Ridiculous

0h 05m video Published Aug 19, 2026 Transcribed Aug 19, 2026 Two Minute Papers Two Minute Papers
Intermediate 3 min read For: AI enthusiasts, researchers, and developers interested in open-source models and their implications.
AI Trust Score 75/100
⚠️ Average / Some Fluff

"Delivers exactly what the title promises—a breakdown of DeepSeek's open model that makes closed AI look overpriced."

AI Summary

DeepSeek has released a new model, V4 Pro, which shows significant improvements over its preview version, especially in understanding 3D structures. The weights are free and MIT-licensed, allowing anyone to run the model at their own cost. The video explains the post-training techniques, including distillation and multi-token prediction, that make this possible.

[00:04]
New Model Introduction

The new model is called 0813, and it performs better than the flash version, with improved 3D understanding.

[01:01]
Free Weights and Pricing

Weights are free, but hosting requires hardware; options include Lambda or DeepSeek's hosted service, which has raised prices 2.5-5x.

[01:29]
MIT License and Competition

DeepSeek uses MIT license for open weights, allowing anyone to run the model at their own price, with multiple hosts competing.

[02:09]
Specialist Models

Post-training creates specialist models for math, coding, and agentic work, which are separately trained checkpoints, not mixture of experts.

[02:38]
Distillation Process

Distillation uses 10+ teachers to train a student model, improving its abilities significantly.

[03:19]
Multi-Token Prediction

The model drafts multiple tokens ahead, leading to up to 78% faster generation.

[03:52]
Rapid Adoption

The research paper was published only 6 weeks ago, and now it's being used widely for free.

Mentioned in this Video

Study Flashcards (6)

What is the name of the new model released by DeepSeek?

easy Click to reveal answer

DeepSeek V4 Pro

00:04

What license does DeepSeek use for its weights?

easy Click to reveal answer

MIT license

01:29

How much did DeepSeek raise their prices?

medium Click to reveal answer

2.5 to 5 times

01:17

How do specialist models differ from mixture of experts?

medium Click to reveal answer

They are separately trained model checkpoints, not parts of a single neural network.

02:24

What speed improvement does DeepSeek report for V4 Pro?

medium Click to reveal answer

78% faster generation

03:36

How long ago was the research paper for this model published?

hard Click to reveal answer

6 weeks

03:52

💡 Key Takeaways

📊

MIT Licensed Open Weights

Highlights the open-source nature, allowing anyone to run the model at their own price.

01:29
🔧

Distillation with 10+ Teachers

Explains the key post-training technique that improves the student model significantly.

02:38
📊

78% Faster Generation

Quantifies the real-world speed benefit users can experience immediately.

03:36
💡

From Paper to Practice in 6 Weeks

Demonstrates the rapid adoption of research, making it accessible for free.

03:52

[00:04] the real one. I know it gets confusing. The previous version was called preview, and this is called 0813. The numbers say it performs better than the much smaller flash version. When building a Rubik's Cube, flash did not

[00:19] completely understand the 3D structure of the object. Lots of missing parts, lots of blackness. But, with the pro, look, much better understanding of

[00:31] look, much better understanding of structure. Then, I was also surprised by this. Holy mother of papers. Look at that. It is inching closer and closer to fable quality. Yep, another challenger

[00:46] appeared, and it gets better. They give all this for us for free, which is absolutely incredible. Okay, but what does it mean for us? The weights are available for free for all of us, but, you know, few have the hardware to host

[01:01] it at home. I'd love to, but I don't have that kind of hardware. Other options include Lambda or using it hosted by DeepSeek themselves. But, they just raised their prices dramatically, about 2 and 1/2 to 5x the previous

[01:17] prices. Now, I bet you can already imagine the clickbait headlines saying, I think what they should also say is that DeepSeek has MIT licensed open

[01:29] weights. What does that mean? Well, anyone can run the exact same model at anyone can run the exact same model at their own price. And, look, they do. A bunch of hosts available, and they all compete on price. That is amazing for

[01:44] us, and it is very likely to push the Frontier Labs to give us fellow scholars Frontier Labs to give us fellow scholars something even better, and quickly. And all this improvement comes from the same architecture. But, how is that even

[01:57] possible? The model structure is the same, yet it is massively better than the preview was less than 4 months ago. So, how? Dear fellow scholars, this is

[02:09] Two Minute Papers with Dr. Károly Zsolnai-Fehér. Once again, a lot of the magic happens after pre-training. During post-training, DeepSeek creates several specialist models for mathematics, coding, and agentic work. Now, we have

[02:24] to stop here for a moment. People confuse these with the experts in mixture of experts. That's not quite the same. Those are little pieces within one same. Those are little pieces within one neural network. These are not. These are

[02:38] separately trained model checkpoints. Okay, so what then? Then comes Okay, so what then? Then comes distillation. Yeah. They take more than 10 of these specialist teachers and train one final model to absorb their

[02:53] abilities. So, the student model says, "This is what I would do." Then the teacher says, "Well, this is what I would have done." Then the student adapts its brain [clears throat] to be more like its teacher. Do it with 10

[03:07] teachers and you see that the student indeed improves like crazy. They also added this part to it. Instead of just predicting one token at a time, it

[03:19] drafts several tokens ahead. It does it much better than previous techniques. And hold on to your papers, fellow scholars, because DeepSeek reports up to 78% faster generation for V4 Pro. Real,

[03:36] measurable speed up in real use that you get right now and benefit from it. And get right now and benefit from it. And here is something absolutely insane. This was a research paper, let's see, 6 weeks ago. And now everyone is using

[03:52] it. Let me say it again, a research paper only 6 weeks ago. One of the best papers of the year. And it is coming alive right in our hands, for free.

[04:04] Incredible. Full breakdown video in the description. And don't forget, we own and can run the weights. No one downgrades us to a different model if we downgrades us to a different model if we type the wrong keyword. No games. That

[04:18] is incredible. Even if I can't run it at home, which I would love to do. But, there are options. What a time to be alive. This is open science and open alive. This is open science and open research at its best. And it's important

[04:32] that we talk about it. Why? Because the future belongs to those who understand future belongs to those who understand it. Use DeepSeek and use DeepSpark. Take advantage of them. Oh, and I plan to talk about DeepSeek's incredible no

[04:46] agent harness as well. Novel design, really powerful. If you're interested, consider subscribing and hitting the bell. I use Lambda to reproduce AI research papers often in minutes. It's also great to train your own models or

[05:01] fine-tune an existing one. Run inference or text-to-image or video, easy-peasy. Running a DeepSeek chatbot or agent, super fast, super reliable. Lambda gives

[05:13] you powerful Nvidia GPUs to run your own experiments. I test ideas from the papers I cover and moments later, results. Love it. Seriously, try it out now at lambda.ai/papers.

More from Two Minute Papers

View all

⚡ Saved you 0h 05m reading this? Transcribe any YouTube video for free — no signup needed.