---
title: 'Google Gemini''s Biggest Update Yet: AI Agents, Omni AI, and Smart Glasses'
source: 'https://youtube.com/watch?v=GiqyqBKmAfA'
video_id: 'GiqyqBKmAfA'
date: 2026-08-12
duration_sec: 500
---

# Google Gemini's Biggest Update Yet: AI Agents, Omni AI, and Smart Glasses

> Source: [Google Gemini's Biggest Update Yet: AI Agents, Omni AI, and Smart Glasses](https://youtube.com/watch?v=GiqyqBKmAfA)

## Summary

Google has released a major update to the Gemini app, introducing a new design, autonomous agents, and advanced multimodal capabilities. The update includes Spark, a feature that can execute multi-step tasks, and the Omni model, which processes text, images, and video together. The video also explores the future of AI with wearables like smart glasses.

### Key Points

- **UI Redesign** [00:28] — The Gemini app has a complete UI overhaul with a new design language called 'neural expressive', featuring a modern look, new icons, and fonts.
- **Spark: Autonomous Agents** [01:38] — Spark is a new feature that brings autonomous agents to the consumer app, allowing it to spin off sub-agents to accomplish larger goals.
- **Scheduled Actions** [03:26] — Spark can handle open-ended tasks and supports scheduled actions called 'heartbeats' for recurring tasks like daily briefings.
- **Omni Model** [04:26] — The Omni model combines text, image, and video inputs, enabling 'anything to anything' generation, such as merging a video with a photo style.
- **Desktop Voice Features** [05:20] — The desktop app uses voice with real-time filtering to remove filler words like 'um' and 'uh', and can access context from files and screen.
- **Future Wearables** [06:37] — The future vision includes wearables like smart glasses with bone-conducting audio, allowing hands-free, context-aware assistance.

## Transcript

Google just dropped a massive wave of AI updates&nbsp; for the Gemini app, completely overhauling the&nbsp;&nbsp; interface and introducing actual autonomous&nbsp; agents that can run tasks in the background.&nbsp;&nbsp; I sat down for an exclusive interview with Jeff&nbsp; Huang, the head of engineering for the Gemini app,&nbsp;&nbsp;
to find out exactly how these new features work,&nbsp; from anything to anything multimodal models,&nbsp;&nbsp; to desktop voice hacks that clean up your speech&nbsp; in real time. Let's go inside the engineering team&nbsp;&nbsp;
to see what's new. I started by asking Jeff about&nbsp; the massive visual change you see the second you&nbsp;&nbsp; launch the Gemini app, which is powered by a new&nbsp; design language Google calls neural expressive.&nbsp;&nbsp;
Here's what he had to say about why they&nbsp; completely rebuilt the UI from the ground up.&nbsp;&nbsp; First, there's a huge redesign of that. As soon&nbsp; as you launch the app, you'll see it. It's more&nbsp;&nbsp; modern. We have a sort of a new design language&nbsp; that came out with it. On iOS, it's liquid glass.&nbsp;&nbsp;
And it's just, I think, a lot cleaner. We were&nbsp; really aiming to sort of modernize and clean up a&nbsp;&nbsp; lot of the UI. So, we redesigned what we call the&nbsp; zero state, like the first screen that you see,&nbsp;&nbsp; the side nav that kind of comes out when&nbsp; you tap the three lines, that's all redone,&nbsp;&nbsp;
brand new set of icons, new fonts across the&nbsp; board. And then as you get a response back,&nbsp;&nbsp; you type in something to Gemini and it gets&nbsp; you, it gives you a response. The response fades&nbsp;&nbsp; all redesigned. So that was just kind of top to&nbsp; bottom. The UI is very different. The reception's&nbsp;&nbsp;
been pretty good. But the absolute biggest&nbsp; announcement from the engineering team is a&nbsp;&nbsp; brand-new feature called Spark. This fundamentally&nbsp; shifts Gemini from a simple chat bot into a true&nbsp;&nbsp;
autonomous assistant that can execute multi-step&nbsp; workflows for you. Here's how Jeff explained it to&nbsp;&nbsp; me. The big one that you know is leading in in&nbsp; Sundar's announcement is called Spark. And so,&nbsp;&nbsp; this is really bringing agents into our consumer&nbsp; AI app. Yeah, it's super exciting. So, I've been&nbsp;&nbsp;
playing with it. If you think about how, you know,&nbsp; eventually you will want something that like kind&nbsp;&nbsp; of can act and be like a true assistant. It's&nbsp; something that should be able to like do work and&nbsp;&nbsp; like kind of spin off multiple tasks to accomplish&nbsp; like a bigger goal versus right now, you know,&nbsp;&nbsp;
most of these chat apps, you kind of give it&nbsp; one thing and it gives you like a response.&nbsp;&nbsp; So, you can kind of give it some like pretty big&nbsp; open-ended hard tasks and it'll know to spin off&nbsp;&nbsp; little sub agents to go accomplish that and then&nbsp; kind of come back. If you were to talk to someone&nbsp;&nbsp;
who has no experience with agents and because I&nbsp; feel like you know you hear of agents and I think&nbsp;&nbsp; some people are a little apprehensive to test it&nbsp; out. What would you say is that first scenario&nbsp;&nbsp; that someone could try where they see value from&nbsp; using an agent? Just tell me what I need to know&nbsp;&nbsp;
to prepare for the day. And actually, we built&nbsp; this as a product that I can talk about too called&nbsp;&nbsp; Daily Brief but Spark can do that as well. And so&nbsp; basically if you kind of think about what you need&nbsp;&nbsp; to do, you know, let's say tomorrow you might wake&nbsp; up and kind of look at your calendar, see what you&nbsp;&nbsp;
have going on, kind of think about what you need&nbsp; to prepare for, like do I need to go, look up&nbsp;&nbsp; Kevin's YouTube channel, same for your email, like&nbsp; what is it that's sort of lingering on your to-do&nbsp;&nbsp; list, do I have bills I have to pay and things&nbsp; like that where it can go and do all that, it can&nbsp;&nbsp;
kind of spin up these like agents or I don't know,&nbsp; you can call them like little minions to like kind&nbsp;&nbsp; of go in like do work for you. And then once it&nbsp; does that, you can ask it to do that repeatedly.&nbsp;&nbsp; So that's the other part of Spark is you can&nbsp; have, well we call them heartbeats internally,&nbsp;&nbsp;
but they can they can repeat so they can be&nbsp; sort of these scheduled actions. When I hear of&nbsp;&nbsp; repeating things or setting up automations, that&nbsp; sounds complicated. Yeah. But I take it it's not&nbsp;&nbsp; morning and that's as far as it goes. Yeah. Okay,&nbsp; as soon as you've done something, or even before&nbsp;&nbsp;
you've done it, it's exactly that it's every day&nbsp; do this, you know, at 8am. So, it almost feels&nbsp;&nbsp; like with Spark, it's almost bringing kind of&nbsp; these agent capabilities to the masses. Exactly.&nbsp;&nbsp; Next, we shifted gears to the brand-new underlying&nbsp; engine driving the app's advanced multimodal&nbsp;&nbsp;
capabilities, the Omni model. I asked Jeff how&nbsp; this completely changes the way Gemini processes&nbsp;&nbsp; different types of media at the same time. So, we&nbsp; call that the Omni model and that one super cool.&nbsp;&nbsp;
It's kind of combining a few of the different&nbsp; capabilities that we had in the app. So, we had&nbsp;&nbsp; Nano Banana, which was image generation, we had&nbsp; Veo, which was video generation. And for those,&nbsp;&nbsp; it was always text to images or text to video.&nbsp; And you could feed it like an image; you could&nbsp;&nbsp;
do image to image. And now we're basically going&nbsp; from Omni, you know, anything to anything. So,&nbsp;&nbsp; it's actually very cool. I've been playing with it&nbsp; a bunch, but you can feed it like a video and like&nbsp;&nbsp; a couple images and then also give it a prompt at&nbsp; the same time. And then it'll output and combine&nbsp;&nbsp;
like all of those things together into a really&nbsp; cool output. So yeah, it's very powerful. It kind&nbsp;&nbsp; of combines a few different tools that we had.&nbsp; So, like as an example, you could take say a&nbsp;&nbsp; video of yourself and then maybe have a photo&nbsp; with a certain style, and then it'll merge the&nbsp;&nbsp;
two of those. Exactly. Okay. From there, Jeff&nbsp; surprised me with one more major announcement,&nbsp;&nbsp; this time specifically built for the desktop&nbsp; experience. It's a feature designed to completely&nbsp;&nbsp;
change how you talk to your computer by using&nbsp; background context and real time speech filtering.&nbsp;&nbsp; It's using the desktop app and using speech to&nbsp; basically do some pretty complex things. So,&nbsp;&nbsp;
one of the benefits of having the desktop app is&nbsp; it has context from your files on your hard drive,&nbsp;&nbsp; what's on your screen, and then voice is really&nbsp; cool because more and more that's a like very&nbsp;&nbsp;
understandable and the models are like it's a&nbsp; very understandable communication method. And&nbsp;&nbsp; the models are getting so good where you can kind&nbsp; of hold down this thing and just speak naturally.&nbsp;&nbsp; And it knows what to filter out in terms of&nbsp; like, you know, I say "um"s and "uh"s a lot,&nbsp;&nbsp;
or you might make a mistake and be like, "Oh, I&nbsp; actually scratch that." And it's good about kind&nbsp;&nbsp; of knowing and understanding how to filter that&nbsp; and output like a clean sort of stream of text.&nbsp;&nbsp; Once you have that coupled with you kind of using&nbsp; the voice now to orchestrate or kind of command&nbsp;&nbsp;
and do things on the on the Mac, which is&nbsp; very cool. To wrap up our conversation,&nbsp;&nbsp; I wanted to look further down the road. I&nbsp; asked Jeff about the ultimate North Star for&nbsp;&nbsp; the Gemini app and how tomorrow's hardware will&nbsp; completely change the way assistants interact&nbsp;&nbsp;
with the physical world around us. As you see&nbsp; the Gemini app continue to evolve and develop,&nbsp;&nbsp; what is your North Star for the experience in&nbsp; the app? Yeah, kind of depends like how far into&nbsp;&nbsp; the future. Eventually, we will have wearables,&nbsp; which can be the kind of the primary interface.&nbsp;&nbsp;
You saw some demos actually with the glasses,&nbsp; which I find actually that's a very natural&nbsp;&nbsp; like interface where it has the audio, you know,&nbsp; sort of bone conducting audio. So, you don't have&nbsp;&nbsp;
to pull out your phone. You don't have to type to&nbsp; it. Like you said, like speech is just so natural.&nbsp;&nbsp; And then also like you can take photos and videos&nbsp; of the context. So, it's actually like super cool.&nbsp;&nbsp;
and I'll turn it on and be like, "Hey, do I&nbsp; need to like fertilize this? Like what plant&nbsp;&nbsp; is this and what kind of fertilizer do I need?"&nbsp; Versus like going and like Googling, you know,&nbsp;&nbsp;
searching and trying to like figure out exactly&nbsp; what fertilizer. So it's kind of nice to have that&nbsp;&nbsp; like kind of in context. That actually I think&nbsp; resolves one of my big pain points because I feel&nbsp;&nbsp;
oftentimes like if I'm gardening or something,&nbsp; I take out my phone, you take the screenshot,&nbsp;&nbsp; and then you add it to the chat. Then you ask your&nbsp; question. It's like all these, you know, steps&nbsp;&nbsp; that it's a lot of work. Yeah. And so just having&nbsp; glasses that are seeing what you're seeing and can&nbsp;&nbsp;
answer your question. And it's hands-free. You can&nbsp; like be actually gardening. So, it's almost like&nbsp;&nbsp; that assistant that's seeing what you're seeing&nbsp; and can answer questions that you have. Yeah,&nbsp;&nbsp; exactly. And then Josh mentioned something in one&nbsp; of his demos where you're sort of like throwing&nbsp;&nbsp;
context to the assistant and it's catching&nbsp; it and sort of like saving it for later. I&nbsp;&nbsp; think that's also really important where the more&nbsp; context it has, the better it gets at, you know,&nbsp;&nbsp; helping you with whatever. These new features are&nbsp; rolling out right now. And if you have access to&nbsp;&nbsp;
Google AI Ultra, keep an eye out for the new&nbsp; Spark beta tab hitting your account to set up&nbsp;&nbsp; your very first autonomous agent. Please consider&nbsp; subscribing and I'll see you in the next video.
