[00:03] see, today AI models have become almost comically large. Some of the biggest open and free AI models, like DeepSeek, reached over 1.6 trillion parameters in [00:15] size. And that's not even the biggest one. So, these cost hundreds of thousands of dollars to run. And then, you find out something crazy. You show them an image and you ask, "What does this image depict?" And you will be [00:28] surprised to hear the answer. It doesn't know. Yes, it cannot see. Now, what if I told you that they promised you a model that is 99% smaller, a speck of dust comparatively, and yet this tiny guy can see? That [00:45] sounds impossible. Yes, maybe in our dreams, wishful thinking. But, it actually exists. It is made by DeepMind and is called Gemma 4. This runs on your [00:57] laptop. It is an absolute gem. Free and open, downloaded more than 300 million times by us fellow scholars. That is insane. I love it. And then, something [01:09] amazing happened. Now, hold on to your papers, fellow scholars, because, yep, they gave us the secret sauce. They finally told us the architecture they used for Gemma 4 to see. And it's [01:23] kind of crazy. So, we finally understand how this can pull off things so easily, like talking as a medieval bard while identifying objects in your video, and so much more. So, how did they add vision and multimodal reasoning into an [01:40] unusually small local model? The secrets are finally out. Dear fellow scholars, this is Two Minute Papers with Dr. Károly Zsolnai-Fehér. So, a conventional AI is really several neural networks connected together. And if you have an [01:55] image, you need a dedicated visual model just for that. Or Okay, what about audio? Giving it ears. Yep, that needs an audio encoder. A specialized part for [02:07] each of these tasks. But this is not really looking or listening. This is just passing a translation around that is created by a different neural network. Their smallest models do that, but when we upgrade to the 12 billion [02:23] model, things get crazy. Scientists at DeepMind say, "Throw that all away. Out. Right now." Instead, mhm, it cuts your picture into small patches. Then it [02:35] projects those pixels directly into the model's internal representation. So, it knows where each patch came from. Okay, but what is the point here? Why do that? [02:47] Well, you don't need a separate neural network, a vision transformer, to interpret the image for you. No. Throw it out. Same for audio. Slice it up into it out. Same for audio. Slice it up into 40-ms chunks, and then comes the magic. [03:02] You just pour all these tokens into your main transformer, and then what happens? Well, this system is forced to learn to be the eyes, ears, and brain at the same [03:14] time. And I think this is one of the most important architectural ideas in most important architectural ideas in Gemma 4. It removes hundreds of millions of specialist parameters, and it blurs the boundary between perception and [03:28] thinking. And the result is an AI system that punches way above its weight. It handles images, it handles audio, and it is bloody smart. And this is something that you can own. Now, two more points. One, the Gemma 4 ecosystem continues to [03:45] get improvements to make it faster and better. I'll give you a link to So in the description. Two, they shared the secret sauce there. So Gemma 4 is not just amazing in and of itself, but it can help Deep Seek and other systems [04:01] learn to see better and more efficiently. That is absolutely amazing. What a time to be alive. And please do not take it for granted that these amazing open models will just keep [04:14] coming in the future. These are gifts to all of us and these gifts may stop coming as capabilities increase. It is not a law of nature that we just get these models for free in the future, too. So, to everyone who is working on [04:29] open models, wherever you are in the world, you are heroes. You help scientists, students, and millions of other people to do their work better. Thank you so much. And we, as a community, have to come together and do [04:44] everything we can to support these open systems. I use Lambda to reproduce AI research papers often in minutes. It's also great to train your own models or [04:56] fine-tune an existing one. Run inference or text-to-image or video, easy-peasy. Running a Deep Seek chatbot or agent, super fast, super reliable. Lambda gives you powerful Nvidia GPUs to run your own experiments. I test ideas from the [05:12] papers I cover and moments later, results. Love it. Seriously, try it out now at lambda.ai/papers.