---
title: 'New ChatGPT Voice Update is Insane!'
source: 'https://youtube.com/watch?v=N0Gzq1VqPIw'
video_id: 'N0Gzq1VqPIw'
date: 2026-07-28
duration_sec: 511
---

# New ChatGPT Voice Update is Insane!

> Source: [New ChatGPT Voice Update is Insane!](https://youtube.com/watch?v=N0Gzq1VqPIw)

## Summary

This video tests the new ChatGPT voice update in the desktop app, which allows users to control their computer, run multiple AI agents, and perform tasks via voice commands. The host demonstrates real-world examples, highlights limitations, and provides an honest assessment of the update's capabilities.

### Key Points

- **Introduction to Voice Update** [00:04] — ChatGPT desktop app now has a voice update that can control the computer, run AI agents, and sit on the desktop like a sticky pet.
- **Demo: Posting a Tweet** [01:24] — Using a different account, the host asked ChatGPT to draft and post a tweet about the power of ChatGPT voice. It checked for approval before publishing.
- **Delay Issue** [02:44] — There is a small gap between responses, especially when using a thinking model. Switching to a faster model can reduce delay.
- **Fixing Audio Setup** [03:36] — ChatGPT configured audio routing to record desktop sound and voice simultaneously, saving time.
- **Struggles with Timer** [03:49] — Setting a simple 10-second timer took longer than 10 seconds to create, highlighting a limitation.
- **Coding Task Failure and Recovery** [04:02] — First attempt to improve a habit tracker UI failed due to an issue with the work tree, but it corrected itself on the second attempt.
- **Multi-tasking Capability** [04:28] — ChatGPT can juggle multiple coding tasks simultaneously, such as building a voice system in one project and working on another.
- **Background Operation** [05:11] — Users can walk away while ChatGPT continues tasks in the background, muting the voice if needed.
- **Full Permissions** [05:38] — ChatGPT has access to local files and the entire computer, not just a sandbox browser tab.
- **Settings and Voice Options** [06:04] — Users can choose voice options, microphone, toggle screen context, and set a voice chat hotkey.
- **Voice Speed Limitation** [06:45] — ChatGPT cannot speak 50% faster as requested, only about 10% faster, showing a ceiling on speed adjustment.
- **Final Verdict** [07:10] — Not Jarvis-level yet, but a genuine step up from previous voice control. Works in background, which changes daily utility.

### Conclusion

The ChatGPT voice update is a significant improvement for hands-free computer control, though it has limitations like delay and occasional failures. It shines in background multi-tasking and practical applications, making it a valuable tool for those willing to work through its rough edges.

## Transcript

Would you actually let an AI take control of your entire computer? Because it my mouse, I gave it my keyboard, I gave it my voice, and what happened next surprised me. I'm the digital avatar of Julian Goldie, and I help people learn
how to actually use AI tools in their day-to-day work, not just talk about them. Today, I'm testing the brand new voice update inside the ChatGPT desktop app, and I'm going to show you exactly what it can do, what it can't do yet,
and where it genuinely surprised me halfway through this video. So, here's what's new. There's a chatty voice update inside the desktop app now, and it means you can control your computer and run multiple AI agents coding in the
sits on your desktop like a little sticky pet. You can click on it, mute time, which is handy when you're trying to explain something to someone else and you don't want it talking over you. It's powered by GPT live under the hood. That
means it can speak, it can listen, and it can coordinate work across different tasks at the same time. This is actually a big deal if you've ever tried building setting one up yourself takes a lot of technical effort to get anywhere near
this smooth. Having this built straight into the desktop app skips all of that setup work for you. Is it perfect? No, it breaks sometimes. At one point, it visible on the screen, and it still couldn't tell what it was looking at.
genuinely impressive and moments where it just misses completely. I want to be because that's the only way this is actually useful to you. Let's start with different account and asked it to post a tweet about the power of ChatGPT voice.
It drafted the tweet, checked with me before posting, and once I said publish, it went to X and posted it live. Then I asked it to open the tweet up, and it all without me touching my keyboard once. That's a real working example of
it controlling a browser and publishing content based on a spoken instruction. It didn't just describe what it would do, it actually did it start to finish while I was talking to you. Seeing it actually carry out a task like
that start to finish from a single spoken instruction is exactly why I think this kind of practical testing matters more than just reading a feature hearing a tool can do something and watching it actually happen. That's
exactly the kind of practical real-world testing we go through together inside the iProfit boardroom. Whenever a new tool like this drops, we break it down properly instead of just repeating the headline, and we show members exactly
where it's worth their time and where it isn't yet. If ChatGPT voice control is something you want to actually use properly in your own setup, that's the kind of walk-through and support waiting for you inside the boardroom right now.
The delay is the one thing that bugs me a little. There's a small gap between responds, especially when it's using a thinking model. That makes sense once the scenes, but it does break the illusion a bit. If you're building
something similar yourself, switching to a faster, more instant model instead of a thinking model is probably the move if speed matters more to you than depth. It without being asked twice, which sounds small, but if you've ever manually
dragged five windows around your screen to get your setup right, you know how Here's the part that actually saved me time while making this exact video. My desktop audio wasn't set up correctly, and I needed to record both the sound on
my screen and my own voice at the same time. I asked it to fix it, and it went into my mixer and configured the routing properly so I could hear my desktop and record it both at once. That's something I would have spent ages messing around
with on my own, and it sorted it out directly on screen without me having to let's get into where it actually struggled because I promised you the full picture. I asked it to set a simple 10-second timer while I stepped away,
and it took longer than 10 seconds just to create the timer in the first place. part. They'll show you the highlight reel and skip the bits where it just doesn't work smoothly yet. Then I asked it to open a habit tracker project and
start improving the UI inside it. First attempt, it failed. It told me the work tree wasn't created, so no code changes had actually started. I pushed back and asked again, and that time it caught the issue, restarted the task properly
inside the saved project, and got to work for real. So, it's not flawless, mistake and correct it. While that was running, I also asked it to start building the voice system into a separate Agent OS project I had open
just to see if it could juggle two coding tasks at once. It picked that up, touching what was already built, and kept both tasks moving in the background time. That's honestly the most useful part of this whole update. It's not just
questions, it's more like an orchestrator with a voice. You can have parallel, controlling your browser in a separate tab so it doesn't interrupt notes in the background while you're focused on something else entirely. You
can mute the voice completely and still have it working silently behind the call and don't want to interrupting you out loud. What makes this different from a normal AI chat window is that you can actually walk away from it. You're not
stuck babysitting one conversation waiting for a single response. It can be browser tab for another, and keeping notes updated for a third, all while you're doing something completely unrelated. Then you can come back later,
see exactly what it worked on while you were gone, including anything it never got round to finishing. You also get full permissions with this, which is worth being aware of. It has access to your local files and your whole
computer, not just a sandbox browser tab. That's a big part of why it feels so different from a normal chat window. On the settings side, you can head into the voice section and pick between a handful of different voice options until
to listen to. You can choose which microphone it uses, move the little it completely if you don't want it visible while you work. There are sticky notes attached to it, too, so you can click back at any point and see exactly
what it's currently working on without losing track of your tasks. You can also toggle whether it has screen context, meaning whether it can actually see what voice chat hotkey option, so you can just tap a key and start talking to it
instantly without opening anything first. I went through a few of the voice options myself, and the thing that stood out is how smooth they sound. &gt;&gt; Hey there. I thought &gt;&gt; Hey, I'm ready to hit the ground
&gt;&gt; Hello, it's lovely to meet you. &gt;&gt; Hey, it's great to meet you. How's your robotic. They sound like an actual person talking to you. I ended up down to your own taste, and switching between them takes seconds inside the
settings menu. The speed is worth mentioning again, too, because it matters more than people think. I asked it to speak around 50% faster at one point, and it simply couldn't do that. It did manage to speak about 10% faster
when asked, so there's clearly a ceiling on how far you can push it right now. That's a small but honest limitation, and it's the kind of detail that matters hours at a time. Would I call this Jarvis-level yet? Not
quite, but it is a genuine step up from where voice control inside these tools was before, and the fact that it can keep working in the background while you carry on with your day is the part that actually changes how useful it is
day-to-day. One thing worth testing further is that you can use this voice remote sessions, which means controlling your computer from your phone as long as different way of working when you're away from your desk. If you want the
away from your desk. If you want the full process, SOPs, and 100-plus AI use cases like this one, join the AI Success Lab. Links in the comments and description. You'll get all the video notes from there, plus access to our
community of 85,000 members who are crushing it with AI. If you're planning you're going to run into some of the same rough edges I did. The delay, the second attempt before they work properly. That's exactly what we walk
through together inside the AI Profit Boardroom. We've got coaching calls where you can bring your own setup and get help live, step-by-step tutorials on getting the most out of tools like this, and roadmaps built around exactly what's
new, so you're not figuring it out alone. With over 4,000 members already inside testing, learning, and building with tools like this every single week, it's the fastest way to actually get good at using AI in your real workflow.
good at using AI in your real workflow. Head to AIprofitboardroom.com.
